An eval harness for LLM-as-judge systems. It validates the judge the way you'd validate a panel of human graders: with inter-rater agreement (Fleiss' κ and Krippendorff's α), not accuracy. Before you trust a model to grade AI-generated code, you first prove the grader itself is reliable.
Extracted from a production AI code-governance pipeline. This repo ships the
methodology and tooling (the judge, the κ validation, the rubric, the
benchmark harness) plus a redacted real run: seven frontier models scored
over 23 findings from public-OSS pull requests, with npm run replay:real. The
private corpus stays private, but you can reproduce the real panel's agreement
numbers yourself, and a synthetic sample lets you see the machinery with zero
setup. Self-contained, MIT, no cloud account required.
Start here: the finding · quickstart, no key · the real run · why κ instead of accuracy · gating CI · full write-up
It isn't a precision number: it's that the method catches results you already
believed. An adversarial review by a different model (told to refute, not
confirm) caught a train/serve skew in a calibration scorer: the trainer was
learning from a feature the runtime serves as null, so the lab number could
never reproduce in production. A single model grading its own work misses that; a
second, independent model told to break it doesn't.
The other half is honesty about measurement. A model graded against labels that
correlate with it (cohesion) scores higher than the same model graded against
independent truth (accuracy). The defensible claims are the controlled
comparisons and the methodology, not a precision figure quoted out of context.
See docs/RESULTS.md for the full write-up.
npm install
npm run replayreplay reads a rater panel (a truth.json plus one <model>.jsonl per model),
computes each model's recall and precision against the adjudicated truth, and
reports panel-level agreement with Fleiss' κ and Krippendorff's α (the
recognized statistics for more than two raters; averaging pairwise Cohen's κ is
not). The default panel under data/sample/panel/ is synthetic. Point it at
your own export with --dir=path/to/panel.
On the synthetic panel you'll see per-model recall/precision, then panel agreement:
rater recall (caught real TP) precision (on decided)
model-a 100% (5/5) 71% (5/7)
model-b 40% (2/5) 100% (2/2)
model-c 80% (4/5) 100% (4/4)
Krippendorff's alpha: -0.046 (poor) · 7 findings, 3 raters
Fleiss' kappa: -0.187 (poor)
model-a catches every true positive but over-calls; model-b fires rarely but
is never wrong. They genuinely disagree. A single judge hides that variance; a
panel measured by κ exposes it, which a human adjudication step then resolves.
npm run replay:realThe same tooling over a redacted real panel: seven frontier models (Claude, GPT, Gemini, Grok, Qwen, DeepSeek, GLM) labeling 23 findings from merged PRs in public OSS repos (Cal.com, Discourse), against an independently adjudicated truth set. The model reasoning ships with the labels, so you can read why they split.
findings: 23 · adjudicated TP: 15 · decided (TP or FP): 16
rater recall (caught real TP) precision (on decided)
claude 80% (12/15) 92% (12/13)
gemini 100% (15/15) 100% (15/15)
gpt 67% (10/15) 100% (10/10)
grok 13% (2/15) 100% (2/2)
...
Krippendorff's alpha: 0.046 (poor) · 20 findings, 7 raters
Fleiss' kappa: -0.038 (poor)
NEEDS_INVESTIGATION is an abstention, and is no longer scored as agreement. Two raters who both decline to decide have not agreed about anything. Earlier versions of this table put NI in the label set, which published Fleiss 0.135 / alpha 0.141 — and 90 of the 233 agreeing rater-pairs behind those numbers (38.6%) were both raters saying "I cannot tell." On the question the panel exists to answer, agreement is at or below chance.
Alpha leads because it is built for missing data; Fleiss is reported for continuity and is not designed for uneven coverage. The replay prints both conventions and the NI share, so the difference is visible rather than argued.
One consequence worth seeing: once NI is an abstention, pairs are only
comparable on findings both raters decided, and grok abstains on 20 of 23. It
shares two decided findings with deepseek and they happen to match. That is a
Cohen's kappa of 1.000 that means nothing, so pairs below eight shared decided
findings are excluded from the redundancy means and listed with their counts
instead. grok reports n/a, which is the honest answer.
Real models, real disagreement. Recall separates them cleanly: Gemini catches every true positive, Grok catches two of fifteen.
Precision does not, and an earlier version of this table said otherwise. It scored a rater's TP call against every finding, including the 7 the adjudicator marked NEEDS_INVESTIGATION. NI is the adjudicator abstaining, not a negative, so a rater was charged for deciding something the truth basis declined to. That printed Gemini at 83% and the text above this block used to read "Gemini catches every true positive and over-calls." Gemini never called TP on anything adjudicated FP; restricted to the 16 decided findings it is 15/15.
Precision is now computed over decided findings only, and it still estimates almost nothing here: there is one adjudicated negative in 23 findings, so a rater that called TP on everything decided would score 94%. The replay prints that warning itself. Read the recall column.
Low panel agreement is the finding, not a bug: independent frontier models split on hard findings, and that split is what the quorum and adjudication exist to resolve.
--fail-under turns replay into a check:
npm run replay:real -- --fail-under=0.6It exits 1 unless Fleiss' κ and Krippendorff's α both reach the threshold.
Requiring both matters: on this panel --fail-under=0.14 fails on Fleiss
(0.135) while Krippendorff (0.141) clears it, and that gap is exactly the
borderline case worth stopping for.
| exit | meaning |
|---|---|
| 0 | ran fine, and any requested gate passed |
| 1 | --fail-under was given and the panel fell short |
| 2 | usage or input error |
There is no default threshold, so omitting the flag never fails a build. How much agreement is enough depends on what the panel decides and what a wrong call costs.
Say your corpus of findings is 90% false positives. A lazy judge that always says "false positive" scores 90% accuracy and is worthless. Accuracy lies on imbalanced data. Cohen's κ subtracts out the agreement you'd get by chance, so that rubber-stamp judge scores near zero. It's the correct metric for validating a judge, and most teams reach for accuracy by mistake.
npm test # unit tests, including a case proving accuracy lies where κ doesn'tcp .env.example .env # add one key: OPENROUTER_API_KEY, or ANTHROPIC/OPENAI/GOOGLE
npm run evaleval runs a small set of judge cases through a cross-model panel in parallel.
Each model returns a 0–1 score with a rationale; the panel passes a case by
quorum (default: ≥2 of 3 models score ≥0.8). It writes results.jsonl +
summary.json and exits non-zero if the pass-rate drops below a bar, so you
can wire it into CI and a prompt regression fails the build like a unit test.
One OpenRouter key runs a true cross-family panel (Anthropic + OpenAI + Google). Or set native keys and the judge uses whatever's present, skipping the rest.
PROVIDERS_ENABLED=baseline,qodo npm run benchThis is the harness, not a shipped result. It needs your gh login and API
keys, and writes the comparison to out/ when you run it. Each tool runs over
the same PR corpus and reports findings, cost, and latency in one normalized
shape. baseline is one direct LLM call; qodo shells out to
pr-agent. Hold the base model constant
across both and the only variable left is the orchestration framework, which is
what the benchmark isolates. Add your own tool by implementing the Provider
interface in src/providers/.
data/sample/panel/ SYNTHETIC rater panel (truth.json + <model>.jsonl)
data/panel-real/ REDACTED REAL 7-model panel (npm run replay:real)
│
├── npm run replay ──► src/replay.ts ──► src/kappa.ts
│ recall/precision + Fleiss' κ + Krippendorff's α
│
data/sample/cross-eval-cases.jsonl synthetic judge cases
│
└── npm run eval ───► src/cross-eval.ts ──► src/judge.ts ──► src/providers.ts
quorum-aggregated verdict + CI gate (fetch, any vendor)
labeling/ RUBRIC.md (judge contract + calibration anchors)
cohens-kappa.mts (two-pass Cohen's κ) · rater-reliability.mts
(panel Fleiss' κ / Krippendorff's α + per-rater KEEP/DROP)
RATER_HANDOFF.md · rater-prompts/ (run a blind, multi-vendor panel)
src/bench.ts + src/providers/ multi-tool head-to-head over a PR corpus
scripts/ fetch-corpus.mjs · fetch-multilang-corpus.mjs · fetch-ground-truth.mjs
build-panel-comparison.mts (panel dir → reliability input)
docs/ RESULTS.md (methodology) · corpus-criteria.md
The methodology, not just the math:
RUBRIC.md: the decision contract every rater applies. A verification protocol (verify, don't confirm) plus worked calibration anchors mined from the cases where strong models split. This is what makes κ measure rubric clarity instead of prompt drift.cohens-kappa.mts: reconciles two label passes into a confusion matrix and Cohen's κ (the right statistic for two passes).npm run kappa.rater-reliability.mts: per-rater report headed by the panel-level Fleiss' κ / Krippendorff's α, then abstention, label skew, redundant pairs (high pairwise κ = paying twice for one signal), and a KEEP / DOWN-WEIGHT / DROP call per model.npm run reliability(synthetic) ornpm run reliability:real(the real 7-model panel). All three sharesrc/kappa.ts.RATER_HANDOFF.md+rater-prompts/: how to run a blind, cross-vendor panel through a plain file contract (readfindings.jsonl, writeverdicts-<key>.jsonl, compare).
- Node ≥ 20 (uses built-in
fetch, no model SDKs). - For
eval: one API key (see.env.example). - For
bench: an authenticatedghCLI, plusuvxfor theqodoprovider.
npm test # vitest: κ math + judge parse/aggregate
npm run test:smoke # rate-pacer + pr-agent parser
npm run check-types # tsc, no emitMIT. See LICENSE.