Evaluator Integrity
AI systems are increasingly assessed by automated evaluators โ LLM judges, test harnesses, scored pipelines โ that are themselves untested. We audit the measurement layer: openly published criteria, drawn from the published literature, with every claim machine-verifiable against its source.
Publishing September 2026: criteria v0.1, a conformance census of the judge configurations shipped by eight widely used evaluation frameworks, and a controlled demonstration of a shipped judge configuration being gamed โ and fixed.
Idil Gozel ยท London
idil@evaluatorintegrity.com