All projects

Judging
the judge

The moment you let a model grade another model's work, you have created a number nobody has checked. A judge that has never been compared to a person is not a measurement — it is an opinion with a number attached.

System
Internal agent operations platform
Scale
Quality scoring for outbound work
Role
Sole architect and engineer
Stack
Python / LLM evaluation / Statistics
01

The challenge

Model-generated quality scores were shaping which work got attention, and nothing established that they agreed with the person whose standard mattered.

02

The approach

Score blind on both sides, measure the agreement with a statistic built for it, and fix the usable threshold before seeing the answer.

03

The result

The quality metric became measurable in its own right, with an explicit rule for when it must be ignored — a loop most systems never close.

A model and a person score the same work without seeing each other's verdicts; the agreement between them is measured statistically and a declared threshold decides whether the model's score is allowed to influence anything. 01 Model score fast, cheap 02 Human score blind to the model 03 Agreement measured, not assumed 04 Threshold decided in advance Scored independently, compared afterwards
A model and a person score the same work without seeing each other's verdicts; the agreement between them is measured statistically and a declared threshold decides whether the model's score is allowed to influence anything. 01 Model score fast, cheap 02 Human score blind to the model 03 Agreement measured, not assumed 04 Threshold decided in advance
Fig. 07 — Two independent verdicts, then a measure of how far apart they are

Hover any step to read what it doesTap any step to read what it does