// PUBLIC BENCHMARKS

Honest lab metrics — with caveats

We publish reproducible harness results: an in-process regression fixture (BENCH-01), multi-lab probes (BENCH-02), and novel-target runs (GEN-01b). Numbers below are from real runs — not marketing targets.

Full methodology & raw JSON on docs →

BENCH-01 — in-process regression fixture (3 recall + 2 precision cases)

Starlette micro-lab (Jul 2026). Measures engine regression on 3 injected vulns + 2 safe endpoints — not full-app recall on customer targets.

Loading benchmark data…