High agreement is a clue.
It is not judge validity.
Compare redacted item IDs, blind human labels and judge labels without uploading the evaluated text. Raw agreement, chance-corrected agreement, label confusion and slices show where adjudication is needed—not whether an automated judge is trustworthy in every context.
Local judge–human calibration
Human rows × judge columns
Calibration gaps
Free: 30 paired labels. A $9 purchase unlocks one report for up to 5,000 pairs. Labels and aggregate disagreement IDs stay local until you explicitly generate the paid redacted report.
Evaluation boundary: agreement with one human label set does not establish truth, fairness, robustness, causal validity or safe automation. Kappa depends on label prevalence and design choices. Define sampling, blind annotation, rubric, adjudication, missing/tie handling, slices, model/prompt version and the decision consequence before using the result.
$9 judge calibration report
Spend one report credit to generate an editable calibration record with metrics, confusion counts, slice table, disagreement queue, sampling questions and a re-test decision gate.
- Raw and chance-corrected agreement with deterministic interval
- Label confusion and slice-level comparison
- Redacted disagreement adjudication queue
- Version, sampling and re-calibration signoff
Human-alignment research and platform substitutes
PMF pre-judgment 84/100 · 21 days + 150 qualified visits before a behavioral verdict