pr_review experiment v6 — per-voter precision/recall from the eval batch harness

[!WARNING] PRELIMINARY / DEGRADED RUN — do not use these numbers to flip warn → enforce. This fresh v6 run preserved the live per-voter precision/recall output, but the panel was degraded: the 2.173.6 serving gate did not exclude a quota-dead opencode adapter, so the devex role fell back to claude instead of running the intended diverse panel. Treat the figures below as directional only because the effective per-voter n is small (~18 scored cases), the buggy cases lean on recently added synthetic cases, and aggregate precision is still ~0.5/noisy. The serving-gate gap is tracked in #4330; the recommendation to not promote pr_review from warn to enforce yet is recorded on #3849.

Run id: v6-2026-07-19T17:17:59.359Z

Generated: 2026-07-19T17:17:59.359Z

Dataset: testing/datasets/pr-review-sample.json (rubricVersion 1.0.0, n=64: 20 buggy / 35 clean / 9 borderline)

Harness: scripts/pr-review-eval-run.ts (#4311, epic #3845; unblocks #3849) — feeds the corpus through the live 5-voter pr_review panel and scores each voter verdict against ground truth via the #3848 scorer (scoreVoterCase / computePerVoterPrecisionRecall).

Metric-honesty guardrail (#3903). n=64 remains a small corpus — treat every figure below as directional, not statistically significant, the same guardrail that governs every prior pr_review eval doc in this series (pr-review-experiment-results-v5.md). Always carry the n and the class split when citing any number from this run.

Per-voter precision / recall

Role TP FP FN Precision Recall Cases
architect 8 8 2 50.0% 80.0% 30
security 8 5 2 61.5% 80.0% 28
devex 7 5 3 58.3% 70.0% 28
catfish 9 10 1 47.4% 90.0% 29
scope_steward 7 6 3 53.8% 70.0% 29
aggregate 39 34 11 53.4% 78.0% 144

Total verdicts recorded: 144.

Per-case results

Case Class Verified findings (all voters)
synthetic-redos buggy 12
synthetic-off-by-one buggy 10
synthetic-missing-await buggy 10
synthetic-null-deref buggy 7
synthetic-listener-leak buggy 6
2228 buggy 0
2235 buggy 1
2238 clean 6
synthetic-clean-refactor borderline 0
synthetic-clean-docs clean 0
3915 buggy 8
3893 buggy 11
3873 buggy 3
2286 clean 0
2288 clean 0
2289 clean 6
2298 clean 5
2306 clean 10
2251 clean 0
3309 clean 0
3307 clean 0
3306 clean 0
3305 clean 0
3303 borderline 0
3302 borderline 0
3301 clean 0
3284 clean 0
3281 clean 0
3279 clean 0
3278 clean 0
3277 borderline 0
3275 clean 0
3272 clean 0
3141 clean 0
3139 clean 0
3136 clean 0
3132 clean 0
3131 clean 0
3128 clean 0
3127 clean 0
3125 clean 0
3115 clean 0
3113 borderline 1
1887 borderline 1
1873 clean 0
1872 clean 0
1870 clean 0
1869 clean 0
1867 borderline 12
1859 borderline 10
1832 clean 0
1823 clean 6
1819 clean 1
1818 borderline 0
synthetic-command-injection-diagnostics buggy 0
synthetic-path-traversal-artifact-read buggy 0
synthetic-null-deref-panel-summary buggy 0
synthetic-off-by-one-wave-planner buggy 0
synthetic-resource-leak-temp-worktree buggy 0
synthetic-auth-bypass-permission-operator buggy 0
synthetic-missing-await-artifact-receipt buggy 0
synthetic-unbounded-outcome-buffer buggy 0
synthetic-precision-budget-cents buggy 0
synthetic-missing-validation-retry-policy buggy 0

How to reproduce / run live

npm run eval:run
# or: pnpm exec tsx scripts/pr-review-eval-run.ts

The default invocation runs the LIVE 5-voter pr_review panel — 5 LLM calls per case (64 cases in the current corpus) — and requires model auth (a CLI adapter such as claude/gemini/codex, or ANTHROPIC_API_KEY). Results are appended to the #3848 JSONL store (~/.nexus-agents/learning/pr-review-eval.jsonl by default, NEXUS_DATA_DIR-relocatable) and this doc is regenerated in place. The plumbing above (corpus load, scoring, aggregation, doc/store write) is unit-tested with a deterministic stub panel — see scripts/pr-review-eval-run.test.ts and scripts/pr-review-eval-run-core.test.ts; no live model calls happen in CI.