← Learn AI
C_000147 · llms and generative ai · advanced

Evaluator Drift

The evaluation apparatus itself changing over time — judge model updates, rubric edits, annotator turnover — invalidating comparisons.

Step 1 of 4

In words

What it is, why it matters, and what it is like.

Why am I learning this?

Understanding evaluator drift turns evaluation from a single snapshot into a dependable measuring system. Without this understanding, you cannot trust any before-and-after comparison of your AI system's performance. This knowledge lets you build reliable checkpoints that confirm a software update did not break existing features, and it allows you to monitor models in active service to catch slow, invisible drops in quality. It is the difference between a score that looks good on paper and a score you can safely bet a major deployment decision on.

The idea, in plain terms

Suppose you are a manager assessing your team's performance. You have a checklist you use every month to rate their work. Over time, you start marking things differently: you become stricter about punctuality, you care less about the polish of their presentation, and you add a new item about collaboration. Now, when you compare this month's scores to last month's, you cannot tell if the team actually improved or if your personal standards simply shifted. The same thing happens in AI evaluation. The 'judge' that scores your model's outputs—whether it is a human annotator, a written rubric, or another AI model—is not a fixed, unchanging ruler. It is a living tool that changes over time. When the judge changes its mind, the scores change, even if the system being evaluated did not improve or worsen. This invisible shift contaminates every comparison you make, making it impossible to distinguish true progress from mere measurement noise.

An analogy

Think of a weigh scale in a grocery store. You weigh a bag of apples and get 2 kilograms. Next week, you weigh the same sealed bag and get 2.4 kilograms. Did the apples magically get heavier? Or did the scale drift—perhaps its spring loosened or its zero point moved? In a good measurement system, you would check your tool by putting a standard 2-kilogram weight on it. If it reads 2.4, you know the scale is wrong, not the apples. That standard weight is a stable reference point. In AI evaluation, your 'judge' is the scale, and the 'apples' are the outputs of the system you are testing. Evaluator drift is that scale silently changing its calibration. This analogy has one limit: a physical scale usually drifts slowly due to wear, but an AI judge can change suddenly if its underlying code is updated, or subtly if its training data shifts.

Definition

Evaluator drift is the phenomenon where the tool used to measure quality changes over time—such as a scoring model being updated or a human reviewer’s standards shifting—which alters the scores received without any actual change in the system being measured.

Where this sits

This concept is central to LLM Evaluation, which is the broader practice of using software tools and human review to test artificial intelligence systems. Within that field, it directly relates to Benchmarking, the method of testing models against a fixed set of questions to see how they perform on standard tasks; if the benchmark itself changes or becomes 'contaminated' by the model having seen the answers before training, the comparison is no longer valid.

Signal from the Frontier

Get the next essay on mind, machine, and meaning

Essays at the intersection of AI, philosophy, and Indian governance. No promotional content.

We'll send a one-click sign-in link to confirm. No password needed.

Views expressed are personal and do not represent the Government of India or the Government of Uttarakhand.