We keep scoring AI agents as if they were students filling in bubble sheets, then reading two-point leaderboard jumps as progress. This essay argues the better picture is the assay. An agent dropped into an evaluation behaves like a specimen in glassware, and the number that comes out is a measurement made by a whole apparatus, full of controls, replicates, instrument drift, and the occasional gorgeous false positive. From Clever Hans to Scott’s pi to Goodhart’s law, a tour of why agentic evaluation has to grow up into a measurement science.