// PUBLIC BENCHMARKS
Honest lab metrics — with caveats
We publish reproducible harness results: an in-process regression fixture (BENCH-01), multi-lab probes (BENCH-02), and novel-target runs (GEN-01b). Numbers below are from real runs — not marketing targets.
Full methodology & raw JSON on docs →BENCH-01 — in-process regression fixture (3 recall + 2 precision cases)
Starlette micro-lab (Jul 2026). Measures engine regression on 3 injected vulns + 2 safe endpoints — not full-app recall on customer targets.
Loading benchmark data…