10000app
免费 · 浏览器本地工具Free · browser-local tool

High agreement is a clue.
It is not judge validity.

Compare redacted item IDs, blind human labels and judge labels without uploading the evaluated text. Raw agreement, chance-corrected agreement, label confusion and slices show where adjudication is needed—not whether an automated judge is trustworthy in every context.

Free · browser-only

Local judge–human calibration

Calibration gateADD BLIND LABEL PAIRS
Paired labels0
Raw agreement
Cohen's κ
Bootstrap 95% interval
Disagreements0

Human rows × judge columns

Calibration gaps

    Free: 30 paired labels. A $9 purchase unlocks one report for up to 5,000 pairs. Labels and aggregate disagreement IDs stay local until you explicitly generate the paid redacted report.