Evaluation¶
This page shows what has been measured on code the tool was not tuned on, and how. State as of 2026-10-02. Every number below is quoted from a committed record, and the record's path is given with it.
Instruments¶
| Instrument | Measures | Where |
|---|---|---|
| Detection gate | per-language recall/FP on the bundled cheat-sheet corpora (goal: >= 90% recall, < 10% FP) | python -m openultrasast.gate, benchmarks/manifests/ |
| Pair corpus | fire on the vulnerable function, stay silent on its fix | ousast pairs, benchmarks/pairs/ |
| Independent populations | real repositories at vulnerable, fixed and benign pins, scored under a pre-registered protocol | benchmarks/independent/ |
Recall means the share of known bugs found. FP means false positives.
The detection gate and the local pairs are CI gates. They catch regressions. They do not estimate recall on real code, because the rules were written against them.
Pair corpora¶
A pair is a vulnerable function and its fixed version. You can run the catalog
benchmarks/pairs/catalog.toml and the slice directories with ousast pairs --slice <name>.
The slices are local, github, sast, vfc, vibe-py, vfc-js, agent-vfc and owasp.
- Each pair carries a provenance profile and a train/holdout split (
--profile,--split). - Pointer pairs are not vendored.
--pointersfetches them into~/.cache/openultrasast/pairs/. benchmarks/pairs/README.mddocuments the schema and the review tiers.
advisory-fixes (benchmarks/pairs/advisory-fixes/, 2026-10-01) holds 97 pointer pairs,
one per repository. Each is a public CVE/GHSA fix commit. The function the advisory names is
declared on both sides and changed by the fix.
It was assembled to give the decision engine real negatives (fixed, non-vulnerable code). It
targets families the corpus had few of: access control, deserialization, output encoding,
path, untrusted destination. It is a label source of the decision engine
(src/openultrasast/learn/sources.toml), not an ousast pairs slice. The research log and
every rejection are in research-2026-10-01.md.
Independent populations and the M4 gates¶
M4 is the qualification milestone of the pre-push safety net. It requires all of these together:
- at least 95% actionable precision;
- at least 90% supported recall;
- at least 95% supported-check completion;
- the runtime gates;
- zero fixed-side or benign false alerts.
A population is a set of real repositories. Its lifecycle has three steps:
- It is frozen before any of its cases is scanned.
- It is scored once under a pre-registered protocol.
- It is then spent: it may inform changes, but it can no longer qualify anything.
| Population | Status | Result | Record |
|---|---|---|---|
| v1 (11 cases) | spent | engine: recall 1/11, precision 1/12, 4 fixed-side and 1 benign false alert; every M4 gate failed. Eight of eleven cases completed no question at the vulnerable pin. | benchmarks/independent/results-v1.json |
| v2 (17 cases) | spent | engine: recall 0/17, 3 benign alerts, completion 13,206 of 108,374. Tool hunter: recall 1/17, precision 2/14, 1 fixed-side and 16 benign alerts, $8.88. No M4 gate met by either. | results-v2.json, results-v2-hunter.json |
| v2, model-driven pipeline | exploratory (v2 already spent) | stopped after 7 of 17 cases when the account emptied: recall 4/7; precision not adjudicated; no gate measurable | results-v2-model-pipeline.json |
| v3 (PHP, 15 cases; one SSRF slot left empty by the selection rule) | frozen, untouched | none: it is the one-time final check, run once under protocol-v3.md after a prediction is committed |
population-v3-php.toml, freeze-v3-php.json |
v3 is reserved. Development and analysis work never reads its cases.
tests/test_independent_population.py fails when anything in the development tree names one
of its repositories. That covers a recipe, catalog, manifest or measurement.
stateDiagram-v2
[*] --> Frozen: cases selected and frozen before any scan
Frozen --> Scored: scored once under the pre-registered protocol
Scored --> Spent: result recorded
Spent --> [*]: may inform changes, can no longer qualify anything
note right of Frozen
v3 (PHP, 15 cases) waits here as the one-time final check,
run once after a prediction is committed
end note
note right of Spent
v1 and v2 are spent
end note
Folds: how learned decisions are kept out of their own training data¶
The decision engine is evaluated on repositories it was not shown
(src/openultrasast/learn/folds.py). Every split is by repository group. A group is the label
builder's normalised owner/name, merged across URLs, forks and shared advisories. Splits
cover all corpora and are seeded and deterministic.
flowchart TD
groups["All repository groups with labels"] --> c["Compile split: 25% of groups, never evaluated"]
groups --> rest["The other 75%"]
c --> boot["C_boot 60%: demonstrations"]
c --> val["C_val 40%: instruction scoring"]
rest --> outer["Outer grouped 5-fold: calibration cross-fitted, operating points nested"]
rest --> loso["Leave-one-source-out: a source held out whole"]
rest --> lofo["Leave-one-framework-out: frameworks with at least 10 groups"]
assert_disjoint runs before a compile. It checks that the compile split and the evaluation
folds share no group. A framework with fewer than 10 groups gets insufficient data instead of
a fold. As of 2026-10-02, the outer folds have been run for the six evaluable families.
Leave-one-source-out and leave-one-framework-out have not
(decision-engine.md).
The plane increment¶
The first increment of the agentic plane ran on google/ax. It used the model-driven pipeline's validation set: 46 candidates in 15 cases of v2, with 20 declared vulnerable sites. It was compared with the script pipeline on the same cases, candidates, pins and recorded triage.
The pre-registered gate: at least 16 of 20 declared sites agreed, at a cost per candidate under $0.022.
- Increment 1 (2026-09-29,
benchmarks/independent/plane-increment-1.json): 14 of 20 declared sites agreed, against the reference's 16. Cost was $0.0194 per candidate with recorded triage. NOT MET on recall. Correction (2026-09-30, in the record): the first version misread agree'sdeclared_sites_matched. That field is coverage by candidates (19), but it was read as "found in either pass". So the first version concluded detection was unchanged. Recounted from the agreed rows, the plane's passes flagged 17 of 20 declared sites in at least one pass. The reference flagged 19. - Increment 2 (2026-09-30,
benchmarks/independent/plane-increment-2.json): came after the design revisit. It added a tie-break pass c on the 10 disputed candidates, with 2-of-3 agreement. 16 of 20 declared sites agreed (reference 16). Cost was $0.0205 per candidate over 46, and $0.02195 over the 43 after triage (reference: $0.022). MET, with two caveats the record carries: - The tie-break was chosen after seeing the first measurement on the same validation set. So this compares plane with script and qualifies nothing on an independent population.
- The plane's passes still flag fewer declared sites than the reference (17 vs 19).
Precision is not part of this gate and remains the open problem.
The decision engine¶
Two slices have been measured out of repository: injection (2026-10-01) and six families (2026-10-02). The engine is not adopted. Where we stand shows the state against the gates in one place, with the plane increment and the populations above. The per-family tables are on decision-engine.md.
Removal of the earlier agentic extra¶
The earlier agentic extra was removed on 2026-09-30. The removal was checked for byte-identical
deterministic outputs: pair replay, quick benchmark, standard scan, improve --dry-run. Records:
benchmarks/measurements/2026-09-30-harnessx-removal-baseline/ and -equality/.