pr_review eval dataset curation pipeline
Issue: #3847 (part of epic #3845). Depends on #3846 (the
labeling rubric).
Pipeline: scripts/curate-pr-review-dataset.ts
Dataset: testing/datasets/pr-review-sample.json
Goal
Grow the eval set toward n≥50 (target 100) without fabricating cases. (The set started at n=10, reached n=19 after the #3847 outcome-mining pass, and is now n=43 after the 2026-07-19 human-ratified Codex adjudication promotion.) A fabricated PR/bug corrupts the eval — it teaches us nothing about whether the panel catches real bugs. So the pipeline is built to make adding a genuinely-sourced case a one-step, provenance-stamped operation, and the doc is explicit about what cannot be automated.
The pipeline (curate-pr-review-dataset.ts)
One reproducible entry point. Subcommands:
validate— load the dataset, validate every entry against the rubric schema, and check the datasetrubricVersionmatches the rubric doc header. This is what the Vitest test and any CI gate call. Exit non-zero on any violation.stats— print n and the class balance (buggy / clean / borderline) plus the provenance-source breakdown (historical / synthetic / historical-clean). Use this to track progress toward n≥50 and to document class balance per #3847’s acceptance criteria.add <kind>— emit a correctly-stamped skeleton entry (rubricVersion, pre-filledadjudication/provenanceshape) to stdout for a new case, so “adding case N+1” is a documented one-step that cannot forget the stamp.<kind>isbuggy | clean | borderline | synthetic-buggy | synthetic-clean. The author fills the diff/bug fields and the rationale (TODOplaceholders); the skeleton guarantees the shape is rubric-valid.
The schema module (scripts/curate-pr-review-dataset-schema.ts, exporting
parseDataset) is the single source of truth for the dataset shape — used by
validate, add, and the test — so the dataset cannot drift from what the
rubric documents. It is a hand-rolled validator (not Zod) because zod resolves
from the package, not the repo root where the script runs.
Sourcing procedure (how a real case is born)
Cases come from three honest sources, in priority order:
- Historical PRs with a real post-merge bug fix (gold). Find a merged PR
whose diff contained a defect that was fixed by a later PR/commit. The later
fix is the ground-truth label — the bug is real and was confirmed by a human
shipping a fix. Procedure:
gh pr list --state merged --search "fix"/ mine the issue tracker for “fixes #N” where #N’s regression was introduced by an earlier merged PR.- Record
number(the buggy PR), thelocation(file:line in that PR’s diff),severity(rubric Rule 1), andprovenance.fixReference(the fixing PR/commit). Adjudicate per rubric Rule 5. - #2235 is the worked example: shipped a wrong env-var name, fixed in #2255.
- Synthetic diff-readable bugs (generalize v5’s 5). Hand-crafted diffs with a
single planted, statable defect at a known
file:line(ReDoS, off-by-one, missing await, null deref, resource leak …). These are honestly labeledprovenance.source: "synthetic"and never presented as real. They test diff-reading capability, not real-world prevalence. UsecustomDiff. - Verified-clean PRs (honest proportion). Merged PRs that shipped with no post-merge fix and no defensible medium+ objection from the diff (rubric Rule 3). These are the strict-FP denominator — without enough of them an FP rate is not measurable.
Class balance is documented, not assumed. curate-pr-review-dataset.ts stats
prints the buggy/clean/borderline split. The target proportion is roughly
50% buggy / 35% clean / 15% borderline so both recall (needs buggy) and strict
precision (needs clean) are measurable.
Mining #3675 (alibaba/open-code-review)
Issue #3847 calls for mining #3675’s open-code-review evaluation for transferable
cases/method. That eval set is an external corpus: its cases must be
re-adjudicated under this rubric before import (their severity/location
conventions differ), and license/provenance must be recorded
(provenance.source: "external:open-code-review", original-ref retained). This
is a sourcing input, not an automated import — see the blocker below.
Honest assessment: reaching n≥50 autonomously
Current committed n = 43 (10 buggy, 28 clean, 5 borderline) after the 2026-07-19 human-ratified Codex adjudication promotion. The prior states were n=10 after the #3846 re-adjudication and n=19 after the first #3847 outcome-mining pass. The pipeline, schema, and stamping are in place. The remaining ~7 cases to n≥50 cannot be generated autonomously without corrupting the eval, for concrete reasons:
- Real historical/clean cases require live GitHub history mining + human
judgment. Identifying a merged PR + its later regression fix, then confirming
the defect was in-diff, needs
ghaccess to the full repo history and a reviewer to adjudicate severity/reachability (rubric Rule 5). It is exactly the kind of label that, if guessed, poisons the metric. - Synthetic cases could be mass-produced, but they must not dominate. If we padded to 50 with synthetics we would measure diff-reading on toy code, not the real-world bug-catch claim the epic is about. Honest proportion (source 3 above) caps how many synthetics belong in the set.
- The #3675 corpus needs license/provenance review + per-case re-adjudication before any entry is admissible.
The 2026-07-19 promotion adds no fabricated entries: all 24 new cases are
human-ratified, outcome-mined PRs with the real candidate diff preserved in
customDiff. PR #3308 is deliberately excluded from the corpus because its
confirmed BoundedLRUCache issue is tracked separately as bug #4319.
What n≥50 needs (the explicit blocker)
A real PR-sourcing effort, which is human/live-data-bound:
- Additional gold buggy PRs with post-merge fixes, each adjudicated per rubric, to improve the clean-heavy class balance.
- Additional verified-clean or borderline merged PRs only when they have human adjudication and useful subsystem coverage.
- License + provenance review of the #3675 corpus, then per-case re-adjudication of any transferable cases.
- Optionally a bounded number of additional synthetics (≤ the honest-proportion cap) for defect types not represented in the historical set.
Each of (1)–(4) is a curate-pr-review-dataset.ts add invocation followed by
filling real, verifiable fields and re-running validate + stats. None of it
should be auto-generated. This effort is tracked under #3847; this PR scaffolds
it and does not close it.