Evaluator Integrity

AI systems are increasingly assessed by automated evaluators โ€” LLM judges, test harnesses, scored pipelines โ€” that are themselves untested. We audit the measurement layer: openly published criteria, drawn from the published literature, with every claim machine-verifiable against its source.

Publishing September 2026: criteria v0.1, a conformance census of the judge configurations shipped by eight widely used evaluation frameworks, and a controlled demonstration of a shipped judge configuration being gamed โ€” and fixed.


Idil Gozel ยท London
idil@evaluatorintegrity.com