The moment you let a model grade another model's work, you have created a number nobody has checked. A judge that has never been compared to a person is not a measurement — it is an opinion with a number attached.
System
Internal agent operations platform
Scale
Quality scoring for outbound work
Role
Sole architect and engineer
Stack
Python / LLM evaluation / Statistics
01
The challenge
Model-generated quality scores were shaping which work got attention, and nothing established that they agreed with the person whose standard mattered.
02
The approach
Score blind on both sides, measure the agreement with a statistic built for it, and fix the usable threshold before seeing the answer.
03
The result
The quality metric became measurable in its own right, with an explicit rule for when it must be ignored — a loop most systems never close.
Fig. 07 — Two independent verdicts, then a measure of how far apart they are
Hover any step to read what it doesTap any step to read what it does