pr_review eval candidate-mining curation pipeline
Issue: #3847 (part of epic #3845). Depends on #3846 (the
labeling rubric).
Miner: scripts/mine-pr-review-candidates.ts (npm: eval:mine-candidates)
Candidates file: testing/datasets/pr-review-candidates.json
Curated dataset: testing/datasets/pr-review-sample.json
This complements the dataset-validation/stamping pipeline
(pr-review-dataset-curation.md,
scripts/curate-pr-review-dataset.ts). That pipeline validates + stamps cases
the owner has already adjudicated; this one mines candidates to put in front
of the owner. Both share the pure rubric labeler (curate-pr-review-labeling.ts)
so the mining heuristic stays consistent with the adjudication rubric.
The mine -> adjudicate -> promote flow
merged PRs (gh) pr-review-candidates.json
│ eval:mine-candidates │
▼ (weak label = triage hint only) ▼ OWNER reads, adjudicates
┌───────────────┐ ┌────────────────────────┐
│ CANDIDATE │ adjudicated:false │ owner sets the REAL │
│ cases │ ───────────────────────────► │ class buggy/clean/ │
│ + weakLabel │ │ borderline per rubric │
└───────────────┘ └───────────┬────────────┘
│ promote
▼
pr-review-sample.json (toward n>=50)
- Mine. The owner runs
npm run eval:mine-candidateslocally (it shells out toghwith the owner’s auth — CI never runs it). It pulls a window of recently-merged PRs, excludes bot PRs and PRs already in the dataset/candidates file, extracts each PR’s real diff (bounded) + objective signals, computes a weak label (triage hint), and appends NEW candidates totesting/datasets/pr-review-candidates.json. - Adjudicate. The owner reads each candidate and assigns the REAL class
(
buggy/clean/borderline) per the rubric — fillingknownBugs(with severity + location), theadjudication.rationale, and flippingadjudicated: true. TheweakLabelonly orders this queue; it is never the verdict. - Promote. Adjudicated candidates are moved into
testing/datasets/pr-review-sample.json(re-usingcurate-pr-review-dataset.ts addfor a rubric-valid skeleton, then filling the real fields), thencurate-pr-review-dataset.ts validate+statsconfirm the set is still rubric-valid and report progress toward n>=50.
The weak-label heuristic (triage hint, NOT a verdict)
The miner attaches one weakLabel per candidate, derived purely from objective
signals via the same rubric labeler the adjudication uses. It is a triage hint
to order the owner’s queue, never a label that ships:
| weakLabel | Objective signal that produces it |
|---|---|
likely-buggy |
A later fix(/revert PR that touches the same source file AND reads as a genuine defect correction (adds a guard / fail-loud / resolution, or is a revert). Evidence names the fixing PR. |
likely-clean |
Either a later follow-up that only refines (tunes heuristics / no-behavior-change hardening), OR no corrective PR within the long-tenure window (>= 42 days since merge). |
unknown |
A fix( whose nature can’t be established from its title (neither defect nor refinement marker), OR no corrective PR but the PR is too young (< 42 days) to be a clean signal. |
It is deliberately conservative: unknown whenever the signal is ambiguous or
not yet established. The 42-day long-tenure window mirrors the dataset’s existing
clean rationale (a defect would usually have surfaced a corrective PR within ~6
weeks), and deliberately avoids the rejected short (~1-week) no-fix window whose
false-clean risk was too high.
weakLabelEvidence records the exact objective basis (the fixing PR number, the
refinement marker, or the tenure) so the owner can verify the hint quickly.
The no-fabrication guarantee
The pipeline produces candidates + weak labels only. It does NOT fabricate adjudicated eval data — the entire value of the eval is REAL owner-adjudicated cases; a guessed verdict would poison the metric. Concretely, every emitted case:
- is
adjudicated: falsewith a neutral placeholderclass: "borderline", emptyknownBugs/borderlineConcerns, and anadjudication.rationalethat saysUNADJUDICATED; - carries the signal only in
weakLabel/weakLabelEvidence— the miner asserts no buggy/clean class; - has a
customDiffthat is a bounded slice of the REALgh pr diffoutput (truncated with an explicit marker when long), never synthesized.
This invariant is stated in the script header and enforced by the test
(adjudicated:false, neutral class, real diff).
Idempotent + safe
Re-running eval:mine-candidates is safe:
- Dedup against both the curated dataset (
pr-review-sample.json) and the already-emitted candidates file — a PR is never proposed twice. - Bot PRs excluded (changeset-release, dependabot, github-actions, renovate,
any
[bot]). - Never overwrites an adjudicated candidate — a candidate the owner has marked
adjudicated: trueis preserved verbatim; a freshly-mined case that collides on PR number is dropped. - Never invents a diff — if
gh pr difffails for a PR, the excerpt is empty rather than fabricated.
How the owner runs it
# Local only — uses your gh auth. Default: 50 most-recent merged PRs.
npm run eval:mine-candidates
# Or with options:
pnpm exec tsx scripts/mine-pr-review-candidates.ts --limit 80 --diff-cap 6000
Older-window targeting (--search / --min-age-days)
On an active repo the default “most recent N merged PRs” window can be too
young for the weak-label heuristic to say anything: likely-clean requires a
PR to have cleared the CLEAN_TENURE_DAYS (42-day) no-fix window, so a run
whose whole window is younger than that emits nothing but unknown. Two
flags target an older window instead:
-
--min-age-days N— keep only PRs merged at least N days ago. The miner transparently enlarges the rawgh pr list --limitand re-fetches (capped) until it collects up to--limitqualifying PRs,ghruns out of history, or the internal fetch ceiling is hit. -
--search "<gh search query>"— passthrough togh pr list --search, used instead of--state merged(gh rejects combining them); foldis:mergedinto the query yourself, e.g.:pnpm exec tsx scripts/mine-pr-review-candidates.ts \ --search "is:merged merged:<2026-06-06" --limit 50 # or, without a custom search query: pnpm exec tsx scripts/mine-pr-review-candidates.ts --min-age-days 42 --limit 50
406 (“diff too large”) fallback
gh pr diff returns HTTP 406 for PRs touching more than ~300 files. When
that happens the miner falls back to gh api repos/<owner>/<repo>/pulls/<n>/files
(paginated) and assembles a bounded, diff-like excerpt from the real per-file
patch fields (skipping binary/patchless files) — never fabricating content.
If the fallback also yields nothing usable, the PR is recorded as skipped
(printed at the end of the run with a reason) rather than silently dropped.
Then adjudicate the populated testing/datasets/pr-review-candidates.json
in-place (set the real class, fill knownBugs, flip adjudicated: true), and
promote the adjudicated cases into pr-review-sample.json. Run
pnpm exec tsx scripts/curate-pr-review-dataset.ts validate and … stats to confirm
rubric validity and track the n>=50 target.
Architecture (pure core, gh at the edges)
Mirrors the curate-pr-review-harvest / -labeling split so the rubric logic is
unit-tested without a live GitHub round-trip:
scripts/mine-pr-review-candidates-core.ts— pure: bot exclusion, tenure math, the weak-label heuristic (delegating to the rubric labeler), diff bounding.scripts/mine-pr-review-candidates-assemble.ts— pure: candidate-case assembly (schema mirrorspr-review-sample.json+weakLabel/weakLabelEvidence/adjudicated), bot/dedup filtering, adjudication-preserving merge.scripts/mine-pr-review-candidates.ts— the thinghI/O edge + CLI (untested by unit, likebuild-model-registry.ts).scripts/mine-pr-review-candidates.test.ts— fixtures only, no network; collected by the CI Script Tests job.