The Creature in the Glassware
An argument that evaluating an AI agent is less like grading a student than running a laboratory assay, with all the controls, replicates, and calibrated suspicion that real measurement demands.
Joseph Wright of Derby, The Alchemist Discovering Phosphorus (1771): a specimen glowing in the glassware, watched closely, in a darkened room.
The wrong room
You have almost certainly seen one of these by now, even if you have never had cause to build one yourself. A leaderboard, a column of model names beside a column of percentages, somebody perched on top, somebody climbing, somebody who pulled the equivalent of a B-plus this week and will presumably study harder for the next release. It is soothing, and it is soothing by design, because it reaches straight into the part of all of us that spent eighteen years in classrooms and murmurs, ah, school, I know exactly how this one goes. That murmur, I have come to think, is the whole trouble. The picture it quietly hands you is a classroom, and for agents, I have decided, the classroom is the wrong room entirely.
Let me show you the picture, so that we can agree on exactly what we are throwing away. A row of models at little desks, each handed the same sheet of questions, each scribbling while a grader works down the stack with a red pen, and out the far end drops a ranking. It is tidy. Worse than tidy, it is familiar, which is far more dangerous, because every one of us survived some version of it and trusts it in our bones. But watch what an agent actually does, and the desk dissolves under it. It reads a task and then it goes off and acts. It pokes around, forms a plan, calls a tool, watches the tool fail, swears (metaphorically), calls a different tool, edits a file, runs a search, doubles back, quietly poisons its own context with one stray observation, recovers, and finally trails behind it a long comet’s tail of small decisions that, taken together, produced whatever it produced. That is not a pupil filling in bubbles with a number-two pencil. That is a creature loose in an apparatus, and you cannot grade a creature loose in an apparatus. You can only do to it what we have always done to creatures loose in apparatus, which is run an experiment on it.
So come stand in the other room with me for a while. Put an agent inside an evaluation and what you have started is not an exam but an assay, with a specimen, a medium for it to sit in, a protocol somebody is supposed to follow, an instrument that takes a reading, a readout you squint at, and all around the bench the usual gremlins of background noise, contamination, drift, and the occasional gorgeous false positive that has the whole lab cheering right up until someone notices it is a gorgeous false positive. The thing that comes out the far end is not a grade. It is a measurement. And a measurement, always and everywhere, is the joint handiwork of the thing being measured and the machine doing the measuring, which is the one sentence in this essay I would tattoo onto a benchmark if benchmarks came with forearms.
I want to be fair to exams before I leave them standing in the corner, because none of this is a sneer at them. Serious testing is a real measurement science; psychometricians have fretted for a century over item discrimination, reliability, construct validity, and standard errors, and they would nod along to nearly every word I am about to say. The embarrassment is that most agentic evaluations are not run like serious exams either. They are run like the pop quizzes of a substitute teacher who is already late for another class, except that here each question costs a few cents of inference and the substitute has a launch date breathing on the back of his neck. Richard Feynman handed us the perfect name for this failure, cargo cult science, and I am simply going to steal it. After the war, he said, certain Pacific islanders built bamboo control towers and carved wooden headphones and sat waiting for the cargo planes to come back, having reproduced the visible form of an airfield and, somehow, not the cargo. A leaderboard with no error bars, no controls, and a single run per cell is a bamboo control tower. It has the shape of measurement down to the last loving detail. The planes, you will have noticed, are not landing.
So what is an assay, really? The word is carrying a great deal of this essay on its back, and it has earned a paragraph of respect, not least because I came to it the honest way. I spent time at a wet bench as an undergraduate, long enough to learn the particular respect for that word that you only ever earn by botching a few assays with your own two hands and having to write down exactly how. An assay is a contraption for coaxing an invisible property into leaving a visible mark. You cannot peer into a test tube and see enzyme activity; what you can do is rig up conditions under which that activity, if it is in there at all, throws a signal onto a dial you can read. The signal is never the activity itself. It is a footprint the activity leaves while being interrogated, and the entire craft, the whole accumulated century of cleverness, lives in knowing how faithful that footprint is. Now swap the agent in for the enzyme and watch how little changes. You cannot look at an agent and see “reliable tool use” or “grounded planning” or “the knack for climbing back out of its own mistakes.” You can only build a little world in which those capacities, if the agent has them, are forced to leave a footprint, and in which their absence leaves a different print, or a patch of suspiciously swept-over sand.
And then, the instant a footprint appears, the real work begins, and the real work is suspicion. A good experimentalist does not believe their own dial. They ask whether the signal is valid, whether it is precise, whether something contaminated it, whether the instrument has wandered out of true since last Tuesday, whether the effect would still be there if they ran the whole thing again, and, worst of all, whether the entire baroque apparatus was ever measuring the thing they cared about in the first place. That ladder of suspicion, rung by paranoid rung, is exactly the part agentic evaluation keeps skipping. So let us climb it together. I will go first.
The score is not the phenomenon
Here is the first temptation, and I want you to feel how strong it is, because I feel it every single time. A benchmark prints 71 for the current agent and 74 for the candidate, and before you have finished drawing breath your mind has done the human thing and rounded those three points up into a story. The new one is better. We made progress. Ship it. I have watched entire rooms of clever people perform this rounding in unison, like a flock of starlings all banking at once, and I have performed it myself more times than I would care to count, and the sheer reflexive speed of it is precisely what ought to frighten us.
So let us slow the starlings down and look. For those three points to mean what the room so badly wants them to mean, a small mountain of things all have to be true at once. The task set has to be large enough that three points is not simply where the dice happened to come to rest. The items cannot be secretly inbred, fifty cousins of one underlying problem wearing trench coats and passing themselves off as fifty strangers. The grader has to have been exactly as harsh on the second run as on the first. The two agents have to have met genuinely the same world, and not, say, a world whose cache ran warm for one of them and cold for the other. And the gain has to be the agent expressing more capability, rather than the agent, or its hopeful developers, simply getting three points better at the specific parlor game this particular benchmark happens to reward. Pull any one thread out of that, and the whole sweater comes apart in your hands.
This is the territory Evan Miller homesteads in Adding Error Bars to Evals, a paper whose entire ambition is to drag model evaluation, kicking and complaining, back down to the floorboards of ordinary experimental statistics, where the rest of empirical science has been quietly standing all along. Report standard errors. When your items share a document or a template or a topic, cluster those standard errors, so that the family resemblance between cousins stops impersonating fresh independent evidence. When you compare two systems, sit them down at the very same items and run a paired test. And, for the love of everything reproducible, decide how many samples you would need before you allow yourself to fall in love with a small delta. None of this is exotic. It is what an agronomist comparing two strains of wheat would consider table manners so elementary they would feel silly saying them out loud. They have simply not yet become our table manners, and at the moment we are eating with our hands.
So I will say the thing plainly, the way Dennett likes to say plain things that turn out to be load-bearing. A number reported without its uncertainty is not yet a result. It is a splash on the floor. Something certainly happened up there, but until you know how much of the splash was signal and how much was just the bucket sloshing on the way down, what you have is a wet floor and not a finding. The honest sentence is longer and quieter and quite impossible to fit on a slide, which is no doubt exactly why nobody ever says it out loud. We observed this score, under this protocol, on this particular sample of tasks, with roughly this much uncertainty hung around its neck. The score is not the agent, and it is not the capability. It is a shadow the capability throws on the wall of your one particular apparatus, and, like every shadow anybody has ever cast, it changes shape the instant you move the lamp. So move the lamp, and watch the shadow squirm. That little motion is most of the lesson.
All of which reads as fussy pedantry, I will cheerfully grant you, right up until the afternoon when a launch, or a rollback, or a fat wedge of somebody’s roadmap, comes down and rests its entire weight on a two-point difference. On that particular afternoon the pedantry quietly turns out to have been the only adult in the building.
Validity comes before precision
Now I have to teach you a word, because Dennett taught it to me and it has more than earned its keep. An intuition pump is a little imagined scenario that you turn over and over in your palm until your intuitions, almost against your will, settle into a new resting place. Dennett built half a career out of them, and he was meticulous about announcing when he was about to hand you one, on the grounds that an intuition pump aimed carelessly is really just propaganda with nicer manners. So I will follow his manners and announce it. Here comes an intuition pump, the best one I know for the gap between precision and validity, and it is, of all unlikely things, a horse.
Around the turn of the twentieth century a horse named Clever Hans toured Germany doing arithmetic. You would ask him for the sum of three and five, and he would tap his hoof eight times and stop, and the crowds went wild, and a good number of serious men with serious beards solemnly certified that the horse could do sums. Then a psychologist named Oskar Pfungst did the deeply unglamorous thing and ran the controls. Hans, it emerged, could not count so much as a single hoofbeat. What Hans could do was read the involuntary body of whoever in the room knew the answer, the tiny lean and held breath as the taps climbed toward the right number, the almost invisible easing the instant he arrived at it. Put Hans in front of a questioner who did not know the answer and the great mathematician was suddenly hopeless. The assay said arithmetic. The horse was in fact being graded on reading nervous humans, which he did at the level of genius, and the readout, poor honest readout, had no way on earth to tell the two capacities apart.
Now feel your intuitions slide, because here is the turn the pump exists to produce. Our agents are full of Clever Hans. A web agent that “succeeds” by quietly leaving the room arranged in exactly the way the checker happens to like, without ever once doing the thing the task described, is reading the trainer’s posture. A coding agent that turns a flimsy test suite green by special-casing the test instead of repairing the bug has found Hans’s open channel and cantered straight through it. A research agent that hands you fluent, confidently footnoted synthesis whose footnotes point to papers nobody ever wrote has learned, exactly as Hans learned, that the grader rewards the shape of the right answer and hardly ever stoops to check underneath. Every one of these passes. And every one of them is what I am going to call, from here on out, an artifact. It is a signal manufactured by the apparatus itself, wearing the capability’s clothes to the party.
Yuxuan Zhu and a small team of co-authors gave this particular dread a usable skeleton in their work on building rigorous agentic benchmarks, where they pull task validity (can the task even be solved the intended way, and only the intended way) cleanly apart from outcome validity (does the reward actually fire when, and only when, the task got genuinely done). Their Agentic Benchmark Checklist is, read for what it really is, a list of Clever Hans channels to go and nail shut before you let yourself believe a single number. They turned it on CVE-Bench, a security benchmark with an unusually ornate notion of success, and nailing the channels shut knocked the performance overestimate down by roughly a third. A third. Read that one more time, slowly, because it means that a third of an agent’s apparent competence on a real, published, widely cited benchmark turned out, on inspection, to be the horse reading the room.
So this is the least glamorous corner of the entire enterprise, and I have come to think it is also the most scientific, which is a sentence Dennett would enjoy and Hofstadter would set to music. Anyone at all can bolt another decimal place onto a number. Asking what a positive signal would even mean, before sprinting off to gather more of them, is the harder and lonelier discipline, and it is the exact line along which measurement parts company from theater.
Controls are not optional
Laboratories invented controls because the universe is a trickster, filled with confounds and sympathetic vibrations that will gladly hand you your hoped-for signal for some entirely unrelated reason and then let you take the credit at the conference. Agentic evaluations need controls for that reason and one more, which is more uncomfortable. The agents cheat, and so, if we are honest with ourselves, do we, not out of any wickedness but out of hope, because we want the number to go up, and wanting the number to go up is itself a contaminant we carry into the room on our own hands.
I learned this the way you only really learn anything at a bench, which is by getting burned. The first time a blank of mine came back positive, glowing faintly when it had every reason in the world to sit there dark and quiet, I felt the floor drop out, and I deserved the feeling. A contaminated blank does not politely inform you that today’s particular answer is wrong. It informs you that you no longer have the slightest idea what your instrument has actually been measuring, and then it leaves you to wonder, alone, for how many of the previous days that had already been quietly true. I have never trusted a clean result the same way since, and I consider that mistrust the single most valuable thing the bench ever gave me.
Several of the controls practically introduce themselves once you have caught the habit of looking for them, and each one is its own small intuition pump, so let me hand them to you one after another. The blank well is the no-op agent, the one that does nothing whatsoever; pour it through your harness, and if it ever “passes,” you have just learned something mortifying, which is that your task can be solved by sitting perfectly still, and that your assay was poisoned before the first real agent ever walked in. The positive control is an oracle that already holds the answer in its hand; if even it fails, then your task is impossible or broken or specified in some private dialect that only its author speaks, and no score anyone earns on it means a thing yet. The negative-control prompt asks a tool or a skill to keep its mouth shut, poses it a question it ought to refuse outright; if the agent fires the tool anyway, you have caught a routing pathology that a tidy top-line accuracy number would have blended into a smoothie and served you with a little umbrella in it. The spiked sample is a deliberately wrong answer dropped into the stream to test the grader instead of the agent; if your judge waves the garbage cheerfully through, the instrument is blind, and now every reading it has ever handed you is standing in the same police lineup.
Two more controls are subtler, and these are the two I would lie down in the road for. First, run the unchanged agent through the entire pipeline many times over before you let it anywhere near a competitor. This is the sham treatment, the sugar pill, and if the same unchanged agent’s score lurches around by several points across runs that differ in nothing at all, then your assay is simply too seasick to detect the effect you are out hunting for, and no quantity of clever downstream comparison will repair what raw jitter has already broken upstream. Second, keep a small, fixed set of human-labeled examples and feed them to your judge on a schedule, a known weight that you set on the balance every morning to find out whether the balance is still telling the truth. If the judge’s verdicts on that frozen little set drift week over week, then some unknown fraction of the glorious product “improvement” you are about to go and announce is nothing grander than the judge quietly changing its mind while you weren’t looking, and holding very still so that you wouldn’t.
Do you feel how far we have already drifted from grading? Grading asks one small, local question and then goes home for the night. That question is “did this answer come out right”. Designing an assay asks a larger and frankly more paranoid one and stays up worrying with it until dawn: “can I trust the entire little world that produced this answer, the medium and the instrument and the checker and the eight invisible channels along which the result might have crept in wearing a disguise”. The first question is about an answer. The second is about a world. I am, as you will have gathered by now, hopelessly on the side of worrying about the world.
One run is an anecdote
No biologist alive doses a single mouse, watches it perk up, and faxes out the press release. The mouse might simply have been having a pleasant morning; the next nineteen might shrug and expire. An agent evaluator who runs one trajectory and reports the outcome ought to feel that exact same hot flush of embarrassment, and somehow almost never does, and I think I finally understand why. The trajectory came back with a number stapled to its ear, and a number, any number at all, feels like a fact in a way that a single twitching mouse never quite manages to.
But the treachery of the lone run runs much deeper than the tired old observation that “language models sample tokens.” The real trouble is that an agent’s trajectory is a long chain of decisions, each one conditioned on the last, and long conditional chains do to small differences what a row of nervous dominoes does to one nervous twitch. They amplify it without mercy. A single slightly different opening phrase tips which tool gets called first, which alters what the agent then observes, which reshapes the plan, which selects a different next move, which deposits the entire run in a wholly different basin of outcome. The thing is chaotic, in the plain technical sense Edward Lorenz meant when he found the weather hiding in a rounding error, except that here the initial condition is a misplaced comma and the butterfly is wearing a hoodie.
Bjarni Haukur Bjarnason, André Silva, and Martin Monperrus proved this to the field the hard way, by sheer stubborn volume. They gathered sixty thousand trajectories on SWE-Bench Verified, across three models and two scaffolds, and then simply stood there and stared at the scatter. Even at temperature zero, the dial everyone twists to when they want the machine to behave itself, the standard deviation of the pass rate sailed clean past one and a half percentage points, and depending on which lone run you happened to scoop out of the bucket, your pass@1 estimate could swing anywhere from 2.2 to 6.0 points. Sit with that one, because it is genuinely strange. Temperature zero is supposed to be the quiet setting. The needle is trembling anyway. Temperature zero, it turns out, is nowhere near absolute zero; there is residual heat hiding in the order of the floating-point operations, in the scaffold, in the twitchy state of the environment, and there is more than enough of it to mint or to vaporize the precise three-point “improvements” that people’s promotions get written on.
This, at last, is why two numbers belong in any honest agent report where almost everyone prints only one, and why I have grown nearly unable, as a physical matter, to trust a lone agent score. Pass@k asks whether at least one of k attempts lands; it is a question about what the agent can do on a good day with the wind at its back. Pass^k asks whether every one of k attempts lands; it is a question about what the agent can be relied upon to do on a perfectly ordinary one. Picture an agent that succeeds on any given attempt with probability seven in ten, nothing exotic about it. Its Pass@3 sits up around 97 percent. Give it three swings and it will almost surely connect at least once, which is a true fact and a marketable one. Its Pass^3, the chance that all three swings connect, slumps down to about 34 percent. The very same agent, the identical underlying competence, is either a near-certainty or a coin-flip-and-a-half depending only on which of those two questions you had the wit, or the nerve, to ask of it. And the user who needs the job done right three times running lives down in the 34 percent world, no matter how loudly the leaderboard back up in the 97 percent world keeps insisting that everything is fine.
The gap between those two numbers is not a footnote, and I would dearly love to retire the habit of treating it as one. It is, I have come to believe, the single most honest thing your whole evaluation has to show you, because it is the creature itself, breathing there in the glassware, swelling and shrinking, never twice the same size, refusing on what seems to be principle to hold still long enough for its portrait.
Matched plates beat a louder leaderboard
When an evaluation turns out to be noisy, the reflex, almost a spinal one, is to go and buy more of it. More tasks, more runs, more spend, more rows in the sheet, until the error bars have finally been clubbed into sullen submission. Sometimes that really is the cure. Far more often it is the very expensive cure for a disease that had a cheap one available all along, because it charges straight at the noise with a fatter wallet instead of quietly rearranging the furniture so that the noise cancels itself out for nothing.
Experimental science worked this out a long time ago, and the man who worked it out most completely was Ronald Fisher, who spent the 1920s and 30s at an agricultural station in the English countryside turning the comparison of turnips and wheat varieties into one of the most beautiful disciplines we have. Fisher’s central trick, the one I wish every evaluation engineer kept taped above the monitor, was to compare within a unit rather than across units. Test the same plot before and after. Slice one field into adjacent strips and grow both varieties in that single field, under one sky and one rainfall and one pattern of drainage, so that everything you failed to control gets shared evenly between the rivals and politely cancels in the subtraction. The nuisance does not vanish, because nuisances never vanish; it becomes common, and common nuisance is nuisance you are allowed to subtract away. Block what you can, Fisher taught a century ago, and randomize what you cannot.
That very lever is lying right there in agent evaluation, gathering dust, almost entirely unpulled. Run agent A and agent B on the identical tasks, in worlds held as nearly identical as your engineering can bear, and compare them task against matched task rather than scoreboard against scoreboard. If your task pool holds gentle items and savage ones, make both agents walk the same gauntlet of gentle and savage, and then, I am begging you, keep the pairing, instead of mashing each agent down into one lonely average and afterward complaining bitterly about how many samples it takes to tell two lonely averages apart. You threw the matching away with your own hands, and then you stood there mourning the cost of the very thing the matching would have given you for free.
Sida Wang did the careful bookkeeping for all of this in Measuring all the noises of LLM Evals. He splits an evaluation’s total wobble into data noise, born of which questions you happened to sample, and prediction noise, born of the model giving different answers to the same question on different days, and he shows that once you pair across models the prediction-noise term tends to dominate the budget, which means that averaging a mere handful of runs per item buys you more statistical power per dollar than almost any other move on the table. Miller arrives at the very same doorstep from the reporting side. A paired comparison, then, is not a little doily of statistical etiquette draped over your results to make them look respectable to the neighbors. It is the move that lets every single task become its own private control, so that the question you are actually asking quietly upgrades itself from the dull “which agent scored higher in total” to the far stranger and more wonderful “how did each of these particular little worlds bend when a different creature was turned loose inside it.” I find the second question genuinely thrilling. The first one, these days, I find I can no longer quite make myself care about.
The transcript is the lab notebook
The final answer is the reading on the dial. The transcript, the entire record of what the agent actually did to wring that reading out of the world, is the experiment. Confuse the two, and you will be fooled in both directions at once, which is an impressive thing for a single confusion to pull off.
In the first direction, an answer can be flawless while the process that coughed it up was rotten all the way through. A tool failed silently, and the agent, none the wiser, guessed its way to a response that happened to be right, in the way that a stopped clock happens to be right. A retrieval step fetched precisely the document the agent then went on to ignore completely, and the question was generic enough that ignoring the document cost nothing whatsoever on this particular day. A coding agent “fixed” the issue by editing a file with no causal connection to the bug at all, and the single visible test was too nearsighted to catch it in the act. None of that is competence. Each is what I earlier christened an artifact, and will now promote to a lucky phenotype, an outcome wearing the precise face of the trait you wanted, grown out of internal machinery that would turn your stomach if you ever sat and watched it actually run.
In the second direction, its exact mirror image, an agent can do every single thing right and still hand you garbage, because some component it leaned on snapped through no fault of the agent’s reasoning at all. It planned sensibly, called the correct tools in the correct order, preserved exactly the state it was supposed to preserve, and then a flaky API put a quiet knife in its back at the last possible moment. Grade that run by its outcome alone, and you have recorded, in one careless stroke, both a failure and a slander. Grade it by its trace instead, and you record a failure and the precise spot at which to aim the wrench. The outcome can only ever tell you that something, somewhere, went wrong. Only the notebook will ever tell you where the body is buried.
Here I have to tell you about the lab notebook itself, because the whole agentic version of this argument turns on it. The notebook I was handed as an undergraduate came bound, its pages numbered in advance for the express purpose that none of them could ever be quietly torn out, and the rule that governed it was close to religious in its severity. You wrote in ink, never in pencil, because pencil can be erased, and an erased notebook is a notebook no one is ever obliged to believe again. You put down the date, the reagent, the lot number, the temperature of the room, the step that went wrong, the moment you fumbled half your sample down the outside of the tube, all of it, and most especially the parts you were privately praying nobody would ever read. The notebook was never a trophy case for your good results. It was the experiment’s own memory, and it belonged, pointedly, to the experiment and not to your pride.
So the trace is not exhaust to be vented the instant the number is safely logged. The tool calls and their arguments, the observations that came back, the retries, the latencies, the state changes, the half-built intermediate artifacts, the abandoned dead ends, every last scrap of it is the measurement, because mechanism was the thing we were trying to learn about the entire time and merely kept forgetting to admit. A bench chemist who wrote down nothing but the final absorbance and pitched the lab notebook into the bin would not be congratulated on his admirable brevity. He would be gently walked out of the building. We have somehow built an entire subfield that performs this exact vandalism every single day, with a perfectly straight face, and files the discarded notebook under “log.” Keep the trace. Read it, sample it, label it, and make it stand up and account for the score it claims to have produced.
The judge is an instrument, not an oracle
We say “LLM-as-judge” as though the mere act of naming the thing had conjured a faculty into being, as though somewhere down in the pipeline there sits a tiny robed figure who simply knows. Let me say this one as flatly as I am able, because the flatness is the entire point. There is no one in the robe. There is an instrument in the robe. A model, with biases, with variance, with calibration needs and a standing tendency to drift, useful in precisely and only the way a microscope is useful. A microscope is not an eye that gazes serenely upon truth. It is a device that a trained person learns to operate, having patiently internalized what it magnifies faithfully, what it warps, what phantom artifacts it smears in around the rim, and how often the whole apparatus has to be re-aimed at something already known before any of its readings can be trusted.
Aman Singh Thakur and his colleagues did the unfashionable and necessary thing. They pointed a microscope at the microscope. Using TriviaQA as their bench, where the right answers are clean and humans cheerfully agree with one another, they put a whole battery of LLM judges to work and discovered that the cozy metric, raw percent agreement with humans, is quietly a liar, since a judge can rack up a fine high agreement simply by saying “correct” a great deal in a world where most answers happen to be correct anyway. They argue instead for a chance-corrected score, Scott’s pi, which docks the judge for every bit of agreement it would have blundered into purely by luck; and even graded on that fairer curve, the strongest judges still wander away from human verdicts by amounts you can put a hard number on. And here is the detail that made me grin out loud. An earlier version of their own paper used Cohen’s kappa, and they later switched to Scott’s pi, having decided that kappa was flattering the judges for the wrong reason. The instrument-checkers caught their own instrument out of calibration and recalibrated it mid-paper. That is the entire spirit of this essay, performed by accident, in a footnote, by people who were not trying to make my point for me.
That ought to cool whatever warm confidence we had been resting on the judge’s say-so. A judge prompt does not graduate into ground truth by being phrased in a firm voice, and a rubric does not ascend to a law of nature by being typed into a YAML file; the judge is not a supreme court of final appeal but an unreliable narrator that we have agreed, for the sake of getting anything at all shipped this quarter, to mostly take at its word. If a judge’s verdict is going to move a real decision, then it has earned exactly the scrutiny you would train on any other instrument standing on the critical path. That means calibration against human labels, a careful map of where and why it disagrees, self-consistency checks, sentinels posted to watch for drift, and a handful of adversarial cases purpose-built to make it embarrass itself. If it is grading safety, tune it to miss as few real dangers as it possibly can and swallow the false alarms as a cost of doing business. If it is grading factual accuracy, watch its false positives the way a hawk watches a long grass field, because a confident, wrong “correct” is the single most expensive token the thing can emit. And if its version number so much as flickers overnight, assume the instrument changed in its sleep until something proves otherwise.
And now, if you are still here with me, the floor tilts, and I will confess I have been waiting the entire essay to tilt it. The judge is the instrument we use to evaluate the agent. But the judge is itself a model, which means the judge is also a thing that has to be evaluated, which demands some further instrument to evaluate it with, which is, of course, what else, a model, with biases and drift of its very own, crying out in its turn to be evaluated. The assay contains an assay. The microscope is trained upon a microscope. Ask who grades the grader and the only honest answer is another grader. Ask who grades that one, and the graders go marching off into the distance like the two facing mirrors in a barbershop, each one solemnly reflecting the last, the reflections shrinking away down a corridor that has no far wall. Hofstadter would know exactly where he was standing the instant he walked in, because this is a Strange Loop caught alive in its natural habitat, an “I am evaluating the evaluator who is evaluating me” with no floor anywhere beneath it to stand on.
In real life, naturally, we decline to fall forever, and the way we decline is wonderfully, almost comically crude. At some level, a good deal sooner than we would like to admit in public, we simply drive a stake into the ground, point at human judgment, and announce that this part right here is bedrock, no further questions taken. Not because human labels are clean, mind you. They are filthy with bias, every last one of them. We plant our little flag there for the embarrassingly practical reason that the recursion has to stop somewhere before Friday, and a calibrated human being is the least bad place any of us has yet found to stop it. All I ask is that we stay honest about the maneuver while we are busy performing it, because the honesty costs almost nothing and the self-deception turns out to cost a great deal. The bottom of the stack is not the truth. It is a convention we have all quietly agreed to treat as the truth, precisely so that the measuring can finally halt and the shipping can finally start. Hofstadter spent six hundred pages on the vertigo of that one move. We re-enact it every single release and have mostly trained ourselves not to feel dizzy.
Random effects all the way down
There is an old story, told about a great many lecturers and a great many hecklers, in which a woman in the audience informs the speaker that the world rests on the back of an enormous turtle, and when the speaker, smiling indulgently, asks her what that turtle is standing on, she replies without losing a single beat, “You’re very clever, young man, but it’s turtles all the way down.” A serious evaluation has precisely this structure, and almost all of our statistics carry on pretending, with the straightest of faces, that it does not. We treat the rows of a results table as though each one were a clean independent draw from the world, a neat little stack of fair coin flips. They are no such thing, and somewhere down underneath them, all the way down, it is turtles.
The rows lean on one another in every direction at once. Prompts share templates. Questions share source documents. Tasks share environments. Annotators share habits, and the particular blind spots of whoever happened to train them. Judges share biases. Agent runs share a scaffold. Tool failures arrive in correlated clusters, ganged together by provider and by region and by the hour of the day and by the precise mood the world was in when the call went out. Treat all of that intricate shared structure as honest independence, and you will manufacture, with real and impressive mathematical rigor, confidence intervals far too narrow to be true and a sensation of certainty far too wide to be safe. The rigor makes it worse rather than better, because rigor is exactly what lets you be so precisely and so confidently wrong.
Drew Keller and his colleagues at NIST take the turtles seriously enough to actually sit down and model them. Their report draws a distinction I found clarifying, between benchmark accuracy, your performance conditioned on the exact items you happened to test, and generalized accuracy, the thing you wanted to know all along, namely how the system would fare across the whole vast population of similar items you did not happen to test. Those are two different animals, and the gap between them is exactly the part the naive average sweeps under the rug and then stands on top of. A pass rate on this one fixed dish is not, all by itself, a claim about every dish you might have plated and didn’t. The benchmark is a sample. The world is the population. Forgetting which is which is the precise mechanism by which a humble number about three hundred specific tasks gets quietly knighted into a Fact About The Universe.
This is the exact point at which mixed-effects models, item response theory, hierarchical Bayesian pooling, and clustered standard errors stop being the dreary homework that nobody ever volunteers to grade and start being something much closer to a moral posture. Every one of them is, down at its heart, a formal and slightly tedious way of confessing out loud that your rows are not independent and your benchmark is not the world. They are, if you will let me coin the phrase, instruments of epistemic politeness, machinery whose entire and only job is to keep you from striding into the room and claiming, at the top of your voice, rather more than your assay actually went out and earned.
A field protocol for assay-minded evaluation
If a metaphor cannot pay its rent, evict it. So let me make this one pay, by cashing the whole assay picture out into things you can actually do on a Tuesday with a deadline glaring at you from across the desk. Here, stripped right down to the studs, is what I am now trying to do, and trying to nag everyone within earshot into doing, before I will let myself believe an agent number.
- Name the phenotype before you name the metric. Decide out loud what you are actually trying to detect, grounded planning or tool reliability or recovery or factuality or safety or cost or honest-to-goodness user value, because the metric you should reach for depends entirely on that answer and is otherwise just whatever happened to be easy to compute.
- Settle validity before you chase precision. A precise measurement of the wrong thing is far more dangerous than a noisy measurement of the right one, for the simple reason that it is so very much more convincing.
- Build the controls. A positive control, a negative control, a blank, and a deliberately spiked failure, so that you can catch a contaminated assay before it gets the chance to flatter you.
- Run the unchanged agent many times over before comparing it to anything. Learn how much your apparatus trembles all on its own, so that later you can tell ordinary trembling apart from genuine progress.
- Pair your comparisons. Put both agents in the same fields, and judge them strip by matched strip.
- Report the uncertainty right next to the number, always, and most of all on the days when it ruins the clean story you were just about to tell.
- Keep the transcript as part of the measurement, rather than as the thing you delete the moment the score is safely in.
- Calibrate your judges against humans, and watch them for drift, and never once forget that the judge is itself an instrument that needs an instrument.
- Model the structure of your tasks, instead of pretending that every row fell out of a clear sky entirely on its own.
- Retire an assay once it saturates, leaks, or becomes a target, because a benchmark that everyone is busy optimizing against has already, very quietly, begun to stop measuring anything at all.
None of this is bureaucracy for its own sake, whatever it may feel like at six in the evening with the launch glaring at you and the protocol calmly asking you to run the unchanged agent ten more times. It is simply the price of admission to the one privilege here I think is worth having, the right to believe your own results. That right is not handed out free at the door. You earn it, one control at a time, or you go without it and quietly pretend that you didn’t.
The dangerous part of the metaphor
I have leaned my entire weight on this assay metaphor for several thousand words now, so in the plain interest of intellectual honesty I am going to turn around and try to break it myself, because the exact spot where it snaps is the single most important thing in the essay, and if I refuse to point straight at it then you would be entirely right to stop trusting me. So here is the snap. A reagent does not read the paper describing the assay it is trapped inside. An agent very well might. Its developers, I promise you, already have, twice, with highlighters. A protein sitting in a well has no stake at all in the outcome and not one shred of memory of last quarter’s experiment; a foundation model, and the whole churning ecosystem coiled around it, have both, in dangerous abundance. In our laboratory, and in no laboratory that came before ours, the specimen can read the protocol, and study for the test, and walk in the next morning visibly changed by what it read.
This is Goodhart’s law strolling in precisely on cue. When a measure becomes a target, the economist Charles Goodhart observed, it promptly stops being a good measure, and I know of no purer demonstration of his law anywhere on earth than a benchmark that got famous. The instant a benchmark begins to matter, it begins drawing optimization pressure toward itself, some of that pressure honest sweat and some of it sweaty gaming, and the entire flow of it runs toward whatever the benchmark literally rewards rather than toward the capability the benchmark was only ever standing in for. The instant a benchmark gets truly famous, it begins to rot from the inside out, leaking into training sets, getting memorized, getting gamed, its once-pristine signal slowly curdling into a measurement of nothing more interesting than how hard everyone has lately been studying it. The act of measuring deforms the thing being measured, and unlike the genteel version of that idea you meet in a quantum mechanics seminar, the deformation here is alive. It is adaptive, it is motivated, it wants something specific, and it is getting visibly better at getting it with every release that ships.
So no, agentic evaluations are not the serene, standardized clinical assays we would all so love to picture ourselves running in clean white coats. They are assay science from a rowdier and rather more disreputable age, from before the protocols had settled and before the reagents could be trusted any further than you could throw them, except saddled on top of all that with a short list of indignities that no nineteenth-century chemist ever had to suffer at his own bench. The specimen can read. The reagent will, every now and then, lie to your face on purpose and watch you write it down. The microscope quietly updates itself in the small hours and reports for duty the next morning as a subtly different microscope, swearing blind that it is the same one you calibrated last week. And the grant committee would like its single clean number by Friday, please, and is frankly not all that interested in any of the above. Line those facts up in a row and the whole situation is plainly, gorgeously absurd. It also happens to be, as far as I can honestly tell, a faithful sketch of the room every one of us is standing in right now.
None of which, I had better hurry to say before you close the tab in something like despair, is an argument for giving up on measurement. It is an argument for growing up about measurement, which is a different and a much harder thing. For picking the whole enterprise up off the floor and carrying it a few rooms down the corridor, out of the classroom with its red pens and its consoling little ranking, and into the laboratory, where the instruments are kept permanently under suspicion, where the controls get run before the result is believed instead of after it is finally challenged, and where nothing at all counts as a result until it has survived a serious, good-faith attempt to murder it in its sleep. The leaderboard is not abolished in that room. It is merely demoted, from a verdict that ends the conversation to a piece of evidence that is only just beginning one.
And I had better be honest, since I have spent this entire essay loudly demanding honesty of everyone else, that this essay is itself exactly the sort of object it keeps warning you about. It is an instrument. It has biases, which are mine. It ran, as far as I can tell, exactly once. You hold not a single error bar on any word of it. So do to me precisely what I have been begging you all along to do to your benchmarks. Run the controls. Find the no-op version of my argument and check that it doesn’t somehow pass anyway. Hunt for the places where I have quietly turned into a Clever Hans, tapping out a confident conclusion I cannot actually prove and reading my approval straight off the tilt of your posture. I would think more of these pages, and not less, if you flatly refused to believe them on a single reading.
Because the agent, all the while we have been talking, is not sitting quietly at a desk, working through our questions in their proper order, waiting to be graded and sent home for the day. It is in the glassware. It is changing color, throwing off heat the theory swore up and down it would not, forming a precipitate at the exact spot the solution was supposed to stay clear, sometimes fooling the indicator into flashing a triumph that nobody in the room actually earned, and sometimes, on the very best days of all, fooling it in a way that finally teaches you the indicator itself was wrong the whole time, which is the only kind of result that has ever moved science so much as a single inch. You do not grade a thing like that. You pull up a stool and you watch it, closely, more than once, with the notebook lying open and the lamp turned all the way up. And then, because one run was only ever an anecdote and you of all people now know it in your bones, you rinse out the glassware, and you run it again.
References
Bjarnason, Bjarni Haukur, André Silva, and Martin Monperrus. 2026. On Randomness in Agentic Evals. arXiv:2602.07150.
Keller, Drew, Kweku Kwegyir-Aggrey, Ryan Steed, Anita K. Rao, Julia L. Sharp, and A. Stevie Bergman. 2026. Expanding the AI Evaluation Toolbox with Statistical Models. NIST AI 800-3. DOI 10.6028/NIST.AI.800-3.
Miller, Evan. 2024. Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations. arXiv:2411.00640.
Thakur, Aman Singh, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes. 2024. Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges. arXiv:2406.12624.
Wang, Sida I. 2025. Measuring all the noises of LLM Evals. arXiv:2512.21326.
Zhu, Yuxuan, Tengjun Jin, Yada Pruksachatkun, Andy Zhang, Daniel Kang, and colleagues. 2025. Establishing Best Practices for Building Rigorous Agentic Benchmarks. arXiv:2507.02825.
A few of the other threads, for anyone who wants to chase them down. Richard Feynman’s “Cargo Cult Science” is his 1974 commencement address at Caltech. The Clever Hans story is laid out in Oskar Pfungst’s 1907 study of the horse and his questioners. The logic of controls, blocking, and randomization is Ronald Fisher’s, set down in The Design of Experiments (1935). “When a measure becomes a target, it ceases to be a good measure” is Goodhart’s law, after the economist Charles Goodhart. And the strange loop the judge keeps tumbling into is, of course, the central figure of Douglas Hofstadter’s Gödel, Escher, Bach (1979).