10000app
10000app #0048 · PMF pre-judgment 84/100 · 21 days + 150 qualified visits before a behavioral verdict

High agreement is a clue.
It is not judge validity.

Compare redacted item IDs, blind human labels and judge labels without uploading the evaluated text. Raw agreement, chance-corrected agreement, label confusion and slices show where adjudication is needed—not whether an automated judge is trustworthy in every context.

Free · browser-only

Local judge–human calibration

Calibration gateADD BLIND LABEL PAIRS
Paired labels0
Raw agreement
Cohen's κ
Bootstrap 95% interval
Disagreements0

Human rows × judge columns

Calibration gaps

    Free: 30 paired labels. A $9 purchase unlocks one report for up to 5,000 pairs. Labels and aggregate disagreement IDs stay local until you explicitly generate the paid redacted report.