pr_review experiment v6 — per-voter precision/recall from the eval batch harness
[!WARNING] PRELIMINARY / DEGRADED RUN — do not use these numbers to flip warn → enforce. This fresh v6 run preserved the live per-voter precision/recall output, but the panel was degraded: the 2.173.6 serving gate did not exclude a quota-dead opencode adapter, so the
devexrole fell back toclaudeinstead of running the intended diverse panel. Treat the figures below as directional only because the effective per-voter n is small (~18 scored cases), the buggy cases lean on recently added synthetic cases, and aggregate precision is still ~0.5/noisy. The serving-gate gap is tracked in #4330; the recommendation to not promotepr_reviewfrom warn to enforce yet is recorded on #3849.
Run id: v6-2026-07-19T17:17:59.359Z
Generated: 2026-07-19T17:17:59.359Z
Dataset: testing/datasets/pr-review-sample.json (rubricVersion 1.0.0, n=64: 20 buggy / 35 clean / 9 borderline)
Harness: scripts/pr-review-eval-run.ts (#4311, epic #3845; unblocks #3849) — feeds the corpus through the live 5-voter pr_review panel and scores each voter verdict against ground truth via the #3848 scorer (scoreVoterCase / computePerVoterPrecisionRecall).
Metric-honesty guardrail (#3903). n=64 remains a small corpus — treat every figure below as directional, not statistically significant, the same guardrail that governs every prior pr_review eval doc in this series (
pr-review-experiment-results-v5.md). Always carry the n and the class split when citing any number from this run.
Per-voter precision / recall
| Role | TP | FP | FN | Precision | Recall | Cases |
|---|---|---|---|---|---|---|
| architect | 8 | 8 | 2 | 50.0% | 80.0% | 30 |
| security | 8 | 5 | 2 | 61.5% | 80.0% | 28 |
| devex | 7 | 5 | 3 | 58.3% | 70.0% | 28 |
| catfish | 9 | 10 | 1 | 47.4% | 90.0% | 29 |
| scope_steward | 7 | 6 | 3 | 53.8% | 70.0% | 29 |
| aggregate | 39 | 34 | 11 | 53.4% | 78.0% | 144 |
Total verdicts recorded: 144.
Per-case results
| Case | Class | Verified findings (all voters) |
|---|---|---|
| synthetic-redos | buggy | 12 |
| synthetic-off-by-one | buggy | 10 |
| synthetic-missing-await | buggy | 10 |
| synthetic-null-deref | buggy | 7 |
| synthetic-listener-leak | buggy | 6 |
| 2228 | buggy | 0 |
| 2235 | buggy | 1 |
| 2238 | clean | 6 |
| synthetic-clean-refactor | borderline | 0 |
| synthetic-clean-docs | clean | 0 |
| 3915 | buggy | 8 |
| 3893 | buggy | 11 |
| 3873 | buggy | 3 |
| 2286 | clean | 0 |
| 2288 | clean | 0 |
| 2289 | clean | 6 |
| 2298 | clean | 5 |
| 2306 | clean | 10 |
| 2251 | clean | 0 |
| 3309 | clean | 0 |
| 3307 | clean | 0 |
| 3306 | clean | 0 |
| 3305 | clean | 0 |
| 3303 | borderline | 0 |
| 3302 | borderline | 0 |
| 3301 | clean | 0 |
| 3284 | clean | 0 |
| 3281 | clean | 0 |
| 3279 | clean | 0 |
| 3278 | clean | 0 |
| 3277 | borderline | 0 |
| 3275 | clean | 0 |
| 3272 | clean | 0 |
| 3141 | clean | 0 |
| 3139 | clean | 0 |
| 3136 | clean | 0 |
| 3132 | clean | 0 |
| 3131 | clean | 0 |
| 3128 | clean | 0 |
| 3127 | clean | 0 |
| 3125 | clean | 0 |
| 3115 | clean | 0 |
| 3113 | borderline | 1 |
| 1887 | borderline | 1 |
| 1873 | clean | 0 |
| 1872 | clean | 0 |
| 1870 | clean | 0 |
| 1869 | clean | 0 |
| 1867 | borderline | 12 |
| 1859 | borderline | 10 |
| 1832 | clean | 0 |
| 1823 | clean | 6 |
| 1819 | clean | 1 |
| 1818 | borderline | 0 |
| synthetic-command-injection-diagnostics | buggy | 0 |
| synthetic-path-traversal-artifact-read | buggy | 0 |
| synthetic-null-deref-panel-summary | buggy | 0 |
| synthetic-off-by-one-wave-planner | buggy | 0 |
| synthetic-resource-leak-temp-worktree | buggy | 0 |
| synthetic-auth-bypass-permission-operator | buggy | 0 |
| synthetic-missing-await-artifact-receipt | buggy | 0 |
| synthetic-unbounded-outcome-buffer | buggy | 0 |
| synthetic-precision-budget-cents | buggy | 0 |
| synthetic-missing-validation-retry-policy | buggy | 0 |
How to reproduce / run live
npm run eval:run
# or: pnpm exec tsx scripts/pr-review-eval-run.ts
The default invocation runs the LIVE 5-voter pr_review panel — 5 LLM calls per case (64 cases in the current corpus) — and requires model auth (a CLI adapter such as claude/gemini/codex, or ANTHROPIC_API_KEY). Results are appended to the #3848 JSONL store (~/.nexus-agents/learning/pr-review-eval.jsonl by default, NEXUS_DATA_DIR-relocatable) and this doc is regenerated in place. The plumbing above (corpus load, scoring, aggregation, doc/store write) is unit-tested with a deterministic stub panel — see scripts/pr-review-eval-run.test.ts and scripts/pr-review-eval-run-core.test.ts; no live model calls happen in CI.