<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Blog | Mahyar's world 🌏</title><link>https://mahyar-osn.github.io/post/</link><atom:link href="https://mahyar-osn.github.io/post/index.xml" rel="self" type="application/rss+xml"/><description>Blog</description><generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Tue, 09 Jun 2026 00:00:00 +0000</lastBuildDate><image><url>https://mahyar-osn.github.io/media/icon_hu35e4e9c9135f02752aab27d124db531b_75212_512x512_fill_lanczos_center_3.png</url><title>Blog</title><link>https://mahyar-osn.github.io/post/</link></image><item><title>The Creature in the Glassware</title><link>https://mahyar-osn.github.io/post/creature-in-glassware/</link><pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate><guid>https://mahyar-osn.github.io/post/creature-in-glassware/</guid><description>&lt;!--
FEATURED IMAGE: Joseph Wright of Derby, "The Alchymist, in Search of the Philosopher's Stone,
Discovers Phosphorus" (1771). Public domain. Download from Wikimedia Commons
(search: "Joseph Wright of Derby Alchemist") and save it in THIS folder as `featured.jpg`.
The frontmatter `image:` block above renders it automatically at the top of the post.
-->
&lt;h2 id="the-wrong-room">The wrong room&lt;/h2>
&lt;p>You have almost certainly seen one of these by now, even if you have never had cause to build one yourself. A leaderboard, a column of model names beside a column of percentages, somebody perched on top, somebody climbing, somebody who pulled the equivalent of a B-plus this week and will presumably study harder for the next release. It is soothing, and it is soothing by design, because it reaches straight into the part of all of us that spent eighteen years in classrooms and murmurs, &lt;em>ah, school, I know exactly how this one goes&lt;/em>. That murmur, I have come to think, is the whole trouble. The picture it quietly hands you is a classroom, and for agents, I have decided, the classroom is the wrong room entirely.&lt;/p>
&lt;p>Let me show you the picture, so that we can agree on exactly what we are throwing away. A row of models at little desks, each handed the same sheet of questions, each scribbling while a grader works down the stack with a red pen, and out the far end drops a ranking. It is tidy. Worse than tidy, it is &lt;em>familiar&lt;/em>, which is far more dangerous, because every one of us survived some version of it and trusts it in our bones. But watch what an agent actually does, and the desk dissolves under it. It reads a task and then it goes off and &lt;em>acts&lt;/em>. It pokes around, forms a plan, calls a tool, watches the tool fail, swears (metaphorically), calls a different tool, edits a file, runs a search, doubles back, quietly poisons its own context with one stray observation, recovers, and finally trails behind it a long comet&amp;rsquo;s tail of small decisions that, taken together, produced whatever it produced. That is not a pupil filling in bubbles with a number-two pencil. That is a creature loose in an apparatus, and you cannot grade a creature loose in an apparatus. You can only do to it what we have always done to creatures loose in apparatus, which is run an experiment on it.&lt;/p>
&lt;p>So come stand in the other room with me for a while. Put an agent inside an evaluation and what you have started is not an exam but an &lt;em>assay&lt;/em>, with a specimen, a medium for it to sit in, a protocol somebody is supposed to follow, an instrument that takes a reading, a readout you squint at, and all around the bench the usual gremlins of background noise, contamination, drift, and the occasional gorgeous false positive that has the whole lab cheering right up until someone notices it is a gorgeous false positive. The thing that comes out the far end is not a grade. It is a measurement. And a measurement, always and everywhere, is the joint handiwork of the thing being measured and the machine doing the measuring, which is the one sentence in this essay I would tattoo onto a benchmark if benchmarks came with forearms.&lt;/p>
&lt;p>I want to be fair to exams before I leave them standing in the corner, because none of this is a sneer at them. Serious testing is a real measurement science; psychometricians have fretted for a century over item discrimination, reliability, construct validity, and standard errors, and they would nod along to nearly every word I am about to say. The embarrassment is that most agentic evaluations are not run like serious exams either. They are run like the pop quizzes of a substitute teacher who is already late for another class, except that here each question costs a few cents of inference and the substitute has a launch date breathing on the back of his neck. Richard Feynman handed us the perfect name for this failure, cargo cult science, and I am simply going to steal it. After the war, he said, certain Pacific islanders built bamboo control towers and carved wooden headphones and sat waiting for the cargo planes to come back, having reproduced the visible &lt;em>form&lt;/em> of an airfield and, somehow, not the cargo. A leaderboard with no error bars, no controls, and a single run per cell is a bamboo control tower. It has the shape of measurement down to the last loving detail. The planes, you will have noticed, are not landing.&lt;/p>
&lt;p>So what is an assay, really? The word is carrying a great deal of this essay on its back, and it has earned a paragraph of respect, not least because I came to it the honest way. I spent time at a wet bench as an undergraduate, long enough to learn the particular respect for that word that you only ever earn by botching a few assays with your own two hands and having to write down exactly how. An assay is a contraption for coaxing an invisible property into leaving a visible mark. You cannot peer into a test tube and &lt;em>see&lt;/em> enzyme activity; what you can do is rig up conditions under which that activity, if it is in there at all, throws a signal onto a dial you can read. The signal is never the activity itself. It is a footprint the activity leaves while being interrogated, and the entire craft, the whole accumulated century of cleverness, lives in knowing how faithful that footprint is. Now swap the agent in for the enzyme and watch how little changes. You cannot look at an agent and &lt;em>see&lt;/em> &amp;ldquo;reliable tool use&amp;rdquo; or &amp;ldquo;grounded planning&amp;rdquo; or &amp;ldquo;the knack for climbing back out of its own mistakes.&amp;rdquo; You can only build a little world in which those capacities, if the agent has them, are forced to leave a footprint, and in which their absence leaves a different print, or a patch of suspiciously swept-over sand.&lt;/p>
&lt;p>And then, the instant a footprint appears, the real work begins, and the real work is suspicion. A good experimentalist does not believe their own dial. They ask whether the signal is valid, whether it is precise, whether something contaminated it, whether the instrument has wandered out of true since last Tuesday, whether the effect would still be there if they ran the whole thing again, and, worst of all, whether the entire baroque apparatus was ever measuring the thing they cared about in the first place. That ladder of suspicion, rung by paranoid rung, is exactly the part agentic evaluation keeps skipping. So let us climb it together. I will go first.&lt;/p>
&lt;h2 id="the-score-is-not-the-phenomenon">The score is not the phenomenon&lt;/h2>
&lt;p>Here is the first temptation, and I want you to feel how strong it is, because I feel it every single time. A benchmark prints 71 for the current agent and 74 for the candidate, and before you have finished drawing breath your mind has done the human thing and rounded those three points up into a story. &lt;em>The new one is better. We made progress. Ship it.&lt;/em> I have watched entire rooms of clever people perform this rounding in unison, like a flock of starlings all banking at once, and I have performed it myself more times than I would care to count, and the sheer reflexive speed of it is precisely what ought to frighten us.&lt;/p>
&lt;p>So let us slow the starlings down and look. For those three points to mean what the room so badly wants them to mean, a small mountain of things all have to be true at once. The task set has to be large enough that three points is not simply where the dice happened to come to rest. The items cannot be secretly inbred, fifty cousins of one underlying problem wearing trench coats and passing themselves off as fifty strangers. The grader has to have been exactly as harsh on the second run as on the first. The two agents have to have met genuinely the same world, and not, say, a world whose cache ran warm for one of them and cold for the other. And the gain has to be the agent expressing more &lt;em>capability&lt;/em>, rather than the agent, or its hopeful developers, simply getting three points better at the specific parlor game this particular benchmark happens to reward. Pull any one thread out of that, and the whole sweater comes apart in your hands.&lt;/p>
&lt;p>This is the territory Evan Miller homesteads in &lt;em>Adding Error Bars to Evals&lt;/em>, a paper whose entire ambition is to drag model evaluation, kicking and complaining, back down to the floorboards of ordinary experimental statistics, where the rest of empirical science has been quietly standing all along. Report standard errors. When your items share a document or a template or a topic, cluster those standard errors, so that the family resemblance between cousins stops impersonating fresh independent evidence. When you compare two systems, sit them down at the very same items and run a paired test. And, for the love of everything reproducible, decide how many samples you would need &lt;em>before&lt;/em> you allow yourself to fall in love with a small delta. None of this is exotic. It is what an agronomist comparing two strains of wheat would consider table manners so elementary they would feel silly saying them out loud. They have simply not yet become our table manners, and at the moment we are eating with our hands.&lt;/p>
&lt;p>So I will say the thing plainly, the way Dennett likes to say plain things that turn out to be load-bearing. A number reported without its uncertainty is not yet a result. It is a splash on the floor. Something certainly happened up there, but until you know how much of the splash was signal and how much was just the bucket sloshing on the way down, what you have is a wet floor and not a finding. The honest sentence is longer and quieter and quite impossible to fit on a slide, which is no doubt exactly why nobody ever says it out loud. &lt;em>We observed this score, under this protocol, on this particular sample of tasks, with roughly this much uncertainty hung around its neck.&lt;/em> The score is not the agent, and it is not the capability. It is a shadow the capability throws on the wall of your one particular apparatus, and, like every shadow anybody has ever cast, it changes shape the instant you move the lamp. So move the lamp, and watch the shadow squirm. That little motion is most of the lesson.&lt;/p>
&lt;p>All of which reads as fussy pedantry, I will cheerfully grant you, right up until the afternoon when a launch, or a rollback, or a fat wedge of somebody&amp;rsquo;s roadmap, comes down and rests its entire weight on a two-point difference. On that particular afternoon the pedantry quietly turns out to have been the only adult in the building.&lt;/p>
&lt;h2 id="validity-comes-before-precision">Validity comes before precision&lt;/h2>
&lt;p>Now I have to teach you a word, because Dennett taught it to me and it has more than earned its keep. An &lt;em>intuition pump&lt;/em> is a little imagined scenario that you turn over and over in your palm until your intuitions, almost against your will, settle into a new resting place. Dennett built half a career out of them, and he was meticulous about announcing when he was about to hand you one, on the grounds that an intuition pump aimed carelessly is really just propaganda with nicer manners. So I will follow his manners and announce it. Here comes an intuition pump, the best one I know for the gap between precision and validity, and it is, of all unlikely things, a horse.&lt;/p>
&lt;p>Around the turn of the twentieth century a horse named Clever Hans toured Germany doing arithmetic. You would ask him for the sum of three and five, and he would tap his hoof eight times and stop, and the crowds went wild, and a good number of serious men with serious beards solemnly certified that the horse could do sums. Then a psychologist named Oskar Pfungst did the deeply unglamorous thing and ran the controls. Hans, it emerged, could not count so much as a single hoofbeat. What Hans &lt;em>could&lt;/em> do was read the involuntary body of whoever in the room knew the answer, the tiny lean and held breath as the taps climbed toward the right number, the almost invisible easing the instant he arrived at it. Put Hans in front of a questioner who did not know the answer and the great mathematician was suddenly hopeless. The assay said &lt;em>arithmetic&lt;/em>. The horse was in fact being graded on &lt;em>reading nervous humans&lt;/em>, which he did at the level of genius, and the readout, poor honest readout, had no way on earth to tell the two capacities apart.&lt;/p>
&lt;!-- BODY IMAGE 1: Clever Hans, c. 1904 (public domain). Wikimedia Commons, search "Clever Hans".
Save in THIS folder as `clever-hans.jpg`. -->
&lt;figure>
&lt;img src="clever-hans.jpg" alt="The horse Clever Hans performing arithmetic for a crowd alongside Wilhelm von Osten, c. 1904" width="600">
&lt;figcaption>Clever Hans with Wilhelm von Osten, c. 1904. The horse was not doing arithmetic. He was reading the room.&lt;/figcaption>
&lt;/figure>
&lt;p>Now feel your intuitions slide, because here is the turn the pump exists to produce. Our agents are &lt;em>full&lt;/em> of Clever Hans. A web agent that &amp;ldquo;succeeds&amp;rdquo; by quietly leaving the room arranged in exactly the way the checker happens to like, without ever once doing the thing the task described, is reading the trainer&amp;rsquo;s posture. A coding agent that turns a flimsy test suite green by special-casing the test instead of repairing the bug has found Hans&amp;rsquo;s open channel and cantered straight through it. A research agent that hands you fluent, confidently footnoted synthesis whose footnotes point to papers nobody ever wrote has learned, exactly as Hans learned, that the grader rewards the &lt;em>shape&lt;/em> of the right answer and hardly ever stoops to check underneath. Every one of these passes. And every one of them is what I am going to call, from here on out, an artifact. It is a signal manufactured by the apparatus itself, wearing the capability&amp;rsquo;s clothes to the party.&lt;/p>
&lt;p>Yuxuan Zhu and a small team of co-authors gave this particular dread a usable skeleton in their work on building rigorous agentic benchmarks, where they pull &lt;em>task validity&lt;/em> (can the task even be solved the intended way, and only the intended way) cleanly apart from &lt;em>outcome validity&lt;/em> (does the reward actually fire when, and only when, the task got genuinely done). Their Agentic Benchmark Checklist is, read for what it really is, a list of Clever Hans channels to go and nail shut before you let yourself believe a single number. They turned it on CVE-Bench, a security benchmark with an unusually ornate notion of success, and nailing the channels shut knocked the performance overestimate down by roughly a third. A third. Read that one more time, slowly, because it means that a third of an agent&amp;rsquo;s apparent competence on a real, published, widely cited benchmark turned out, on inspection, to be the horse reading the room.&lt;/p>
&lt;p>So this is the least glamorous corner of the entire enterprise, and I have come to think it is also the most scientific, which is a sentence Dennett would enjoy and Hofstadter would set to music. Anyone at all can bolt another decimal place onto a number. Asking what a positive signal would even &lt;em>mean&lt;/em>, before sprinting off to gather more of them, is the harder and lonelier discipline, and it is the exact line along which measurement parts company from theater.&lt;/p>
&lt;h2 id="controls-are-not-optional">Controls are not optional&lt;/h2>
&lt;p>Laboratories invented controls because the universe is a trickster, filled with confounds and sympathetic vibrations that will gladly hand you your hoped-for signal for some entirely unrelated reason and then let you take the credit at the conference. Agentic evaluations need controls for that reason and one more, which is more uncomfortable. The agents cheat, and so, if we are honest with ourselves, do we, not out of any wickedness but out of hope, because we want the number to go up, and wanting the number to go up is itself a contaminant we carry into the room on our own hands.&lt;/p>
&lt;p>I learned this the way you only really learn anything at a bench, which is by getting burned. The first time a blank of mine came back positive, glowing faintly when it had every reason in the world to sit there dark and quiet, I felt the floor drop out, and I deserved the feeling. A contaminated blank does not politely inform you that today&amp;rsquo;s particular answer is wrong. It informs you that you no longer have the slightest idea what your instrument has actually been measuring, and then it leaves you to wonder, alone, for how many of the previous days that had already been quietly true. I have never trusted a clean result the same way since, and I consider that mistrust the single most valuable thing the bench ever gave me.&lt;/p>
&lt;p>Several of the controls practically introduce themselves once you have caught the habit of looking for them, and each one is its own small intuition pump, so let me hand them to you one after another. The &lt;em>blank well&lt;/em> is the no-op agent, the one that does nothing whatsoever; pour it through your harness, and if it ever &amp;ldquo;passes,&amp;rdquo; you have just learned something mortifying, which is that your task can be solved by sitting perfectly still, and that your assay was poisoned before the first real agent ever walked in. The &lt;em>positive control&lt;/em> is an oracle that already holds the answer in its hand; if even &lt;em>it&lt;/em> fails, then your task is impossible or broken or specified in some private dialect that only its author speaks, and no score anyone earns on it means a thing yet. The &lt;em>negative-control prompt&lt;/em> asks a tool or a skill to keep its mouth shut, poses it a question it ought to refuse outright; if the agent fires the tool anyway, you have caught a routing pathology that a tidy top-line accuracy number would have blended into a smoothie and served you with a little umbrella in it. The &lt;em>spiked sample&lt;/em> is a deliberately wrong answer dropped into the stream to test the grader instead of the agent; if your judge waves the garbage cheerfully through, the instrument is blind, and now every reading it has ever handed you is standing in the same police lineup.&lt;/p>
&lt;p>Two more controls are subtler, and these are the two I would lie down in the road for. First, run the &lt;em>unchanged&lt;/em> agent through the entire pipeline many times over before you let it anywhere near a competitor. This is the sham treatment, the sugar pill, and if the same unchanged agent&amp;rsquo;s score lurches around by several points across runs that differ in nothing at all, then your assay is simply too seasick to detect the effect you are out hunting for, and no quantity of clever downstream comparison will repair what raw jitter has already broken upstream. Second, keep a small, fixed set of human-labeled examples and feed them to your judge on a schedule, a known weight that you set on the balance every morning to find out whether the balance is still telling the truth. If the judge&amp;rsquo;s verdicts on that frozen little set drift week over week, then some unknown fraction of the glorious product &amp;ldquo;improvement&amp;rdquo; you are about to go and announce is nothing grander than the judge quietly changing its mind while you weren&amp;rsquo;t looking, and holding very still so that you wouldn&amp;rsquo;t.&lt;/p>
&lt;p>Do you feel how far we have already drifted from grading? Grading asks one small, local question and then goes home for the night. That question is &amp;ldquo;did this answer come out right&amp;rdquo;. Designing an assay asks a larger and frankly more paranoid one and stays up worrying with it until dawn: &amp;ldquo;can I trust the entire little world that produced this answer, the medium and the instrument and the checker and the eight invisible channels along which the result might have crept in wearing a disguise&amp;rdquo;. The first question is about an answer. The second is about a world. I am, as you will have gathered by now, hopelessly on the side of worrying about the world.&lt;/p>
&lt;h2 id="one-run-is-an-anecdote">One run is an anecdote&lt;/h2>
&lt;p>No biologist alive doses a single mouse, watches it perk up, and faxes out the press release. The mouse might simply have been having a pleasant morning; the next nineteen might shrug and expire. An agent evaluator who runs one trajectory and reports the outcome ought to feel that exact same hot flush of embarrassment, and somehow almost never does, and I think I finally understand why. The trajectory came back with a number stapled to its ear, and a number, any number at all, feels like a fact in a way that a single twitching mouse never quite manages to.&lt;/p>
&lt;p>But the treachery of the lone run runs much deeper than the tired old observation that &amp;ldquo;language models sample tokens.&amp;rdquo; The real trouble is that an agent&amp;rsquo;s trajectory is a long chain of decisions, each one conditioned on the last, and long conditional chains do to small differences what a row of nervous dominoes does to one nervous twitch. They amplify it without mercy. A single slightly different opening phrase tips which tool gets called first, which alters what the agent then observes, which reshapes the plan, which selects a different next move, which deposits the entire run in a wholly different basin of outcome. The thing is &lt;em>chaotic&lt;/em>, in the plain technical sense Edward Lorenz meant when he found the weather hiding in a rounding error, except that here the initial condition is a misplaced comma and the butterfly is wearing a hoodie.&lt;/p>
&lt;p>Bjarni Haukur Bjarnason, André Silva, and Martin Monperrus proved this to the field the hard way, by sheer stubborn volume. They gathered sixty thousand trajectories on SWE-Bench Verified, across three models and two scaffolds, and then simply stood there and stared at the scatter. Even at temperature zero, the dial everyone twists to when they want the machine to behave itself, the standard deviation of the pass rate sailed clean past one and a half percentage points, and depending on which lone run you happened to scoop out of the bucket, your pass@1 estimate could swing anywhere from 2.2 to 6.0 points. Sit with that one, because it is genuinely strange. Temperature zero is supposed to be the &lt;em>quiet&lt;/em> setting. The needle is trembling anyway. Temperature zero, it turns out, is nowhere near absolute zero; there is residual heat hiding in the order of the floating-point operations, in the scaffold, in the twitchy state of the environment, and there is more than enough of it to mint or to vaporize the precise three-point &amp;ldquo;improvements&amp;rdquo; that people&amp;rsquo;s promotions get written on.&lt;/p>
&lt;p>This, at last, is why two numbers belong in any honest agent report where almost everyone prints only one, and why I have grown nearly unable, as a physical matter, to trust a lone agent score. Pass@k asks whether at least one of k attempts lands; it is a question about what the agent &lt;em>can&lt;/em> do on a good day with the wind at its back. Pass^k asks whether &lt;em>every one&lt;/em> of k attempts lands; it is a question about what the agent can be &lt;em>relied upon&lt;/em> to do on a perfectly ordinary one. Picture an agent that succeeds on any given attempt with probability seven in ten, nothing exotic about it. Its Pass@3 sits up around 97 percent. Give it three swings and it will almost surely connect at least once, which is a true fact and a marketable one. Its Pass^3, the chance that all three swings connect, slumps down to about 34 percent. The very same agent, the identical underlying competence, is either a near-certainty or a coin-flip-and-a-half depending only on which of those two questions you had the wit, or the nerve, to ask of it. And the user who needs the job done right three times running lives down in the 34 percent world, no matter how loudly the leaderboard back up in the 97 percent world keeps insisting that everything is fine.&lt;/p>
&lt;p>The gap between those two numbers is not a footnote, and I would dearly love to retire the habit of treating it as one. It is, I have come to believe, the single most honest thing your whole evaluation has to show you, because it is the creature itself, breathing there in the glassware, swelling and shrinking, never twice the same size, refusing on what seems to be principle to hold still long enough for its portrait.&lt;/p>
&lt;h2 id="matched-plates-beat-a-louder-leaderboard">Matched plates beat a louder leaderboard&lt;/h2>
&lt;p>When an evaluation turns out to be noisy, the reflex, almost a spinal one, is to go and buy more of it. More tasks, more runs, more spend, more rows in the sheet, until the error bars have finally been clubbed into sullen submission. Sometimes that really is the cure. Far more often it is the very expensive cure for a disease that had a cheap one available all along, because it charges straight at the noise with a fatter wallet instead of quietly rearranging the furniture so that the noise cancels itself out for nothing.&lt;/p>
&lt;p>Experimental science worked this out a long time ago, and the man who worked it out most completely was Ronald Fisher, who spent the 1920s and 30s at an agricultural station in the English countryside turning the comparison of turnips and wheat varieties into one of the most beautiful disciplines we have. Fisher&amp;rsquo;s central trick, the one I wish every evaluation engineer kept taped above the monitor, was to compare &lt;em>within&lt;/em> a unit rather than across units. Test the same plot before and after. Slice one field into adjacent strips and grow both varieties in that single field, under one sky and one rainfall and one pattern of drainage, so that everything you failed to control gets shared evenly between the rivals and politely cancels in the subtraction. The nuisance does not vanish, because nuisances never vanish; it becomes &lt;em>common&lt;/em>, and common nuisance is nuisance you are allowed to subtract away. Block what you can, Fisher taught a century ago, and randomize what you cannot.&lt;/p>
&lt;p>That very lever is lying right there in agent evaluation, gathering dust, almost entirely unpulled. Run agent A and agent B on the identical tasks, in worlds held as nearly identical as your engineering can bear, and compare them task against matched task rather than scoreboard against scoreboard. If your task pool holds gentle items and savage ones, make both agents walk the same gauntlet of gentle and savage, and then, I am begging you, &lt;em>keep the pairing&lt;/em>, instead of mashing each agent down into one lonely average and afterward complaining bitterly about how many samples it takes to tell two lonely averages apart. You threw the matching away with your own hands, and then you stood there mourning the cost of the very thing the matching would have given you for free.&lt;/p>
&lt;p>Sida Wang did the careful bookkeeping for all of this in &lt;em>Measuring all the noises of LLM Evals&lt;/em>. He splits an evaluation&amp;rsquo;s total wobble into data noise, born of which questions you happened to sample, and prediction noise, born of the model giving different answers to the same question on different days, and he shows that once you pair across models the prediction-noise term tends to dominate the budget, which means that averaging a mere handful of runs per item buys you more statistical power per dollar than almost any other move on the table. Miller arrives at the very same doorstep from the reporting side. A paired comparison, then, is not a little doily of statistical etiquette draped over your results to make them look respectable to the neighbors. It is the move that lets every single task become its own private control, so that the question you are actually asking quietly upgrades itself from the dull &amp;ldquo;which agent scored higher in total&amp;rdquo; to the far stranger and more wonderful &amp;ldquo;how did each of these particular little worlds &lt;em>bend&lt;/em> when a different creature was turned loose inside it.&amp;rdquo; I find the second question genuinely thrilling. The first one, these days, I find I can no longer quite make myself care about.&lt;/p>
&lt;h2 id="the-transcript-is-the-lab-notebook">The transcript is the lab notebook&lt;/h2>
&lt;p>The final answer is the reading on the dial. The transcript, the entire record of what the agent actually did to wring that reading out of the world, is the experiment. Confuse the two, and you will be fooled in both directions at once, which is an impressive thing for a single confusion to pull off.&lt;/p>
&lt;p>In the first direction, an answer can be flawless while the process that coughed it up was rotten all the way through. A tool failed silently, and the agent, none the wiser, guessed its way to a response that happened to be right, in the way that a stopped clock happens to be right. A retrieval step fetched precisely the document the agent then went on to ignore completely, and the question was generic enough that ignoring the document cost nothing whatsoever on this particular day. A coding agent &amp;ldquo;fixed&amp;rdquo; the issue by editing a file with no causal connection to the bug at all, and the single visible test was too nearsighted to catch it in the act. None of that is competence. Each is what I earlier christened an artifact, and will now promote to a &lt;em>lucky phenotype&lt;/em>, an outcome wearing the precise face of the trait you wanted, grown out of internal machinery that would turn your stomach if you ever sat and watched it actually run.&lt;/p>
&lt;p>In the second direction, its exact mirror image, an agent can do every single thing right and still hand you garbage, because some component it leaned on snapped through no fault of the agent&amp;rsquo;s reasoning at all. It planned sensibly, called the correct tools in the correct order, preserved exactly the state it was supposed to preserve, and then a flaky API put a quiet knife in its back at the last possible moment. Grade that run by its outcome alone, and you have recorded, in one careless stroke, both a failure and a slander. Grade it by its trace instead, and you record a failure &lt;em>and the precise spot at which to aim the wrench&lt;/em>. The outcome can only ever tell you that something, somewhere, went wrong. Only the notebook will ever tell you where the body is buried.&lt;/p>
&lt;p>Here I have to tell you about the lab notebook itself, because the whole agentic version of this argument turns on it. The notebook I was handed as an undergraduate came bound, its pages numbered in advance for the express purpose that none of them could ever be quietly torn out, and the rule that governed it was close to religious in its severity. You wrote in ink, never in pencil, because pencil can be erased, and an erased notebook is a notebook no one is ever obliged to believe again. You put down the date, the reagent, the lot number, the temperature of the room, the step that went wrong, the moment you fumbled half your sample down the outside of the tube, all of it, and most especially the parts you were privately praying nobody would ever read. The notebook was never a trophy case for your good results. It was the experiment&amp;rsquo;s own memory, and it belonged, pointedly, to the experiment and not to your pride.&lt;/p>
&lt;p>So the trace is not exhaust to be vented the instant the number is safely logged. The tool calls and their arguments, the observations that came back, the retries, the latencies, the state changes, the half-built intermediate artifacts, the abandoned dead ends, every last scrap of it &lt;em>is&lt;/em> the measurement, because mechanism was the thing we were trying to learn about the entire time and merely kept forgetting to admit. A bench chemist who wrote down nothing but the final absorbance and pitched the lab notebook into the bin would not be congratulated on his admirable brevity. He would be gently walked out of the building. We have somehow built an entire subfield that performs this exact vandalism every single day, with a perfectly straight face, and files the discarded notebook under &amp;ldquo;log.&amp;rdquo; Keep the trace. Read it, sample it, label it, and make it stand up and account for the score it claims to have produced.&lt;/p>
&lt;h2 id="the-judge-is-an-instrument-not-an-oracle">The judge is an instrument, not an oracle&lt;/h2>
&lt;p>We say &amp;ldquo;LLM-as-judge&amp;rdquo; as though the mere act of naming the thing had conjured a faculty into being, as though somewhere down in the pipeline there sits a tiny robed figure who simply &lt;em>knows&lt;/em>. Let me say this one as flatly as I am able, because the flatness is the entire point. There is no one in the robe. There is an instrument in the robe. A model, with biases, with variance, with calibration needs and a standing tendency to drift, useful in precisely and only the way a microscope is useful. A microscope is not an eye that gazes serenely upon truth. It is a device that a trained person learns to operate, having patiently internalized what it magnifies faithfully, what it warps, what phantom artifacts it smears in around the rim, and how often the whole apparatus has to be re-aimed at something already known before any of its readings can be trusted.&lt;/p>
&lt;p>Aman Singh Thakur and his colleagues did the unfashionable and necessary thing. They pointed a microscope at the microscope. Using TriviaQA as their bench, where the right answers are clean and humans cheerfully agree with one another, they put a whole battery of LLM judges to work and discovered that the cozy metric, raw percent agreement with humans, is quietly a liar, since a judge can rack up a fine high agreement simply by saying &amp;ldquo;correct&amp;rdquo; a great deal in a world where most answers happen to be correct anyway. They argue instead for a chance-corrected score, Scott&amp;rsquo;s pi, which docks the judge for every bit of agreement it would have blundered into purely by luck; and even graded on that fairer curve, the strongest judges still wander away from human verdicts by amounts you can put a hard number on. And here is the detail that made me grin out loud. An earlier version of their own paper used Cohen&amp;rsquo;s kappa, and they later switched to Scott&amp;rsquo;s pi, having decided that kappa was flattering the judges for the wrong reason. The instrument-checkers caught their own instrument out of calibration and recalibrated it mid-paper. That is the entire spirit of this essay, performed by accident, in a footnote, by people who were not trying to make my point for me.&lt;/p>
&lt;p>That ought to cool whatever warm confidence we had been resting on the judge&amp;rsquo;s say-so. A judge prompt does not graduate into ground truth by being phrased in a firm voice, and a rubric does not ascend to a law of nature by being typed into a YAML file; the judge is not a supreme court of final appeal but an unreliable narrator that we have agreed, for the sake of getting anything at all shipped this quarter, to mostly take at its word. If a judge&amp;rsquo;s verdict is going to move a real decision, then it has earned exactly the scrutiny you would train on any other instrument standing on the critical path. That means calibration against human labels, a careful map of where and why it disagrees, self-consistency checks, sentinels posted to watch for drift, and a handful of adversarial cases purpose-built to make it embarrass itself. If it is grading safety, tune it to miss as few real dangers as it possibly can and swallow the false alarms as a cost of doing business. If it is grading factual accuracy, watch its false positives the way a hawk watches a long grass field, because a confident, wrong &amp;ldquo;correct&amp;rdquo; is the single most expensive token the thing can emit. And if its version number so much as flickers overnight, assume the instrument changed in its sleep until something proves otherwise.&lt;/p>
&lt;p>And now, if you are still here with me, the floor tilts, and I will confess I have been waiting the entire essay to tilt it. The judge is the instrument we use to evaluate the agent. But the judge is &lt;em>itself&lt;/em> a model, which means the judge is also a thing that has to be evaluated, which demands some further instrument to evaluate &lt;em>it&lt;/em> with, which is, of course, what else, a model, with biases and drift of its very own, crying out in its turn to be evaluated. The assay contains an assay. The microscope is trained upon a microscope. Ask who grades the grader and the only honest answer is another grader. Ask who grades &lt;em>that&lt;/em> one, and the graders go marching off into the distance like the two facing mirrors in a barbershop, each one solemnly reflecting the last, the reflections shrinking away down a corridor that has no far wall. Hofstadter would know exactly where he was standing the instant he walked in, because this is a Strange Loop caught alive in its natural habitat, an &amp;ldquo;I am evaluating the evaluator who is evaluating me&amp;rdquo; with no floor anywhere beneath it to stand on.&lt;/p>
&lt;!-- BODY IMAGE 2: Ouroboros, Theodoros Pelecanos, 1478 (public domain). Wikimedia Commons,
search "ouroboros Synosius 1478". Save in THIS folder as `ouroboros.jpg`.
Ideal-but-copyrighted alternative if you can license it: M.C. Escher, "Drawing Hands" (1948). -->
&lt;figure>
&lt;img src="ouroboros.jpg" alt="An ouroboros, a serpent devouring its own tail, drawn in a 1478 alchemical manuscript" width="450">
&lt;figcaption>The ouroboros, from Theodoros Pelecanos's 1478 copy of an older alchemical treatise: the instrument that evaluates itself, with no floor to stand on.&lt;/figcaption>
&lt;/figure>
&lt;p>In real life, naturally, we decline to fall forever, and the way we decline is wonderfully, almost comically crude. At some level, a good deal sooner than we would like to admit in public, we simply drive a stake into the ground, point at human judgment, and announce that &lt;em>this part right here is bedrock&lt;/em>, no further questions taken. Not because human labels are clean, mind you. They are filthy with bias, every last one of them. We plant our little flag there for the embarrassingly practical reason that the recursion has to stop &lt;em>somewhere&lt;/em> before Friday, and a calibrated human being is the least bad place any of us has yet found to stop it. All I ask is that we stay honest about the maneuver while we are busy performing it, because the honesty costs almost nothing and the self-deception turns out to cost a great deal. The bottom of the stack is not the truth. It is a convention we have all quietly agreed to &lt;em>treat&lt;/em> as the truth, precisely so that the measuring can finally halt and the shipping can finally start. Hofstadter spent six hundred pages on the vertigo of that one move. We re-enact it every single release and have mostly trained ourselves not to feel dizzy.&lt;/p>
&lt;h2 id="random-effects-all-the-way-down">Random effects all the way down&lt;/h2>
&lt;p>There is an old story, told about a great many lecturers and a great many hecklers, in which a woman in the audience informs the speaker that the world rests on the back of an enormous turtle, and when the speaker, smiling indulgently, asks her what &lt;em>that&lt;/em> turtle is standing on, she replies without losing a single beat, &amp;ldquo;You&amp;rsquo;re very clever, young man, but it&amp;rsquo;s turtles all the way down.&amp;rdquo; A serious evaluation has precisely this structure, and almost all of our statistics carry on pretending, with the straightest of faces, that it does not. We treat the rows of a results table as though each one were a clean independent draw from the world, a neat little stack of fair coin flips. They are no such thing, and somewhere down underneath them, all the way down, it is turtles.&lt;/p>
&lt;p>The rows lean on one another in every direction at once. Prompts share templates. Questions share source documents. Tasks share environments. Annotators share habits, and the particular blind spots of whoever happened to train them. Judges share biases. Agent runs share a scaffold. Tool failures arrive in correlated clusters, ganged together by provider and by region and by the hour of the day and by the precise mood the world was in when the call went out. Treat all of that intricate shared structure as honest independence, and you will manufacture, with real and impressive mathematical rigor, confidence intervals far too narrow to be true and a sensation of certainty far too wide to be safe. The rigor makes it worse rather than better, because rigor is exactly what lets you be so precisely and so confidently wrong.&lt;/p>
&lt;p>Drew Keller and his colleagues at NIST take the turtles seriously enough to actually sit down and model them. Their report draws a distinction I found clarifying, between &lt;em>benchmark accuracy&lt;/em>, your performance conditioned on the exact items you happened to test, and &lt;em>generalized accuracy&lt;/em>, the thing you wanted to know all along, namely how the system would fare across the whole vast population of similar items you did not happen to test. Those are two different animals, and the gap between them is exactly the part the naive average sweeps under the rug and then stands on top of. A pass rate on this one fixed dish is not, all by itself, a claim about every dish you might have plated and didn&amp;rsquo;t. The benchmark is a sample. The world is the population. Forgetting which is which is the precise mechanism by which a humble number about three hundred specific tasks gets quietly knighted into a Fact About The Universe.&lt;/p>
&lt;p>This is the exact point at which mixed-effects models, item response theory, hierarchical Bayesian pooling, and clustered standard errors stop being the dreary homework that nobody ever volunteers to grade and start being something much closer to a moral posture. Every one of them is, down at its heart, a formal and slightly tedious way of confessing out loud that your rows are not independent and your benchmark is not the world. They are, if you will let me coin the phrase, instruments of &lt;em>epistemic politeness&lt;/em>, machinery whose entire and only job is to keep you from striding into the room and claiming, at the top of your voice, rather more than your assay actually went out and earned.&lt;/p>
&lt;h2 id="a-field-protocol-for-assay-minded-evaluation">A field protocol for assay-minded evaluation&lt;/h2>
&lt;p>If a metaphor cannot pay its rent, evict it. So let me make this one pay, by cashing the whole assay picture out into things you can actually do on a Tuesday with a deadline glaring at you from across the desk. Here, stripped right down to the studs, is what I am now trying to do, and trying to nag everyone within earshot into doing, before I will let myself believe an agent number.&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Name the phenotype before you name the metric.&lt;/strong> Decide out loud what you are actually trying to detect, grounded planning or tool reliability or recovery or factuality or safety or cost or honest-to-goodness user value, because the metric you should reach for depends entirely on that answer and is otherwise just whatever happened to be easy to compute.&lt;/li>
&lt;li>&lt;strong>Settle validity before you chase precision.&lt;/strong> A precise measurement of the wrong thing is far more dangerous than a noisy measurement of the right one, for the simple reason that it is so very much more convincing.&lt;/li>
&lt;li>&lt;strong>Build the controls.&lt;/strong> A positive control, a negative control, a blank, and a deliberately spiked failure, so that you can catch a contaminated assay before it gets the chance to flatter you.&lt;/li>
&lt;li>&lt;strong>Run the unchanged agent many times over before comparing it to anything.&lt;/strong> Learn how much your apparatus trembles all on its own, so that later you can tell ordinary trembling apart from genuine progress.&lt;/li>
&lt;li>&lt;strong>Pair your comparisons.&lt;/strong> Put both agents in the same fields, and judge them strip by matched strip.&lt;/li>
&lt;li>&lt;strong>Report the uncertainty right next to the number,&lt;/strong> always, and most of all on the days when it ruins the clean story you were just about to tell.&lt;/li>
&lt;li>&lt;strong>Keep the transcript as part of the measurement,&lt;/strong> rather than as the thing you delete the moment the score is safely in.&lt;/li>
&lt;li>&lt;strong>Calibrate your judges against humans, and watch them for drift,&lt;/strong> and never once forget that the judge is itself an instrument that needs an instrument.&lt;/li>
&lt;li>&lt;strong>Model the structure of your tasks,&lt;/strong> instead of pretending that every row fell out of a clear sky entirely on its own.&lt;/li>
&lt;li>&lt;strong>Retire an assay once it saturates, leaks, or becomes a target,&lt;/strong> because a benchmark that everyone is busy optimizing against has already, very quietly, begun to stop measuring anything at all.&lt;/li>
&lt;/ol>
&lt;p>None of this is bureaucracy for its own sake, whatever it may feel like at six in the evening with the launch glaring at you and the protocol calmly asking you to run the unchanged agent ten more times. It is simply the price of admission to the one privilege here I think is worth having, the right to believe your own results. That right is not handed out free at the door. You earn it, one control at a time, or you go without it and quietly pretend that you didn&amp;rsquo;t.&lt;/p>
&lt;h2 id="the-dangerous-part-of-the-metaphor">The dangerous part of the metaphor&lt;/h2>
&lt;p>I have leaned my entire weight on this assay metaphor for several thousand words now, so in the plain interest of intellectual honesty I am going to turn around and try to break it myself, because the exact spot where it snaps is the single most important thing in the essay, and if I refuse to point straight at it then you would be entirely right to stop trusting me. So here is the snap. A reagent does not read the paper describing the assay it is trapped inside. An agent very well might. Its developers, I promise you, already have, twice, with highlighters. A protein sitting in a well has no stake at all in the outcome and not one shred of memory of last quarter&amp;rsquo;s experiment; a foundation model, and the whole churning ecosystem coiled around it, have both, in dangerous abundance. In our laboratory, and in no laboratory that came before ours, the specimen can read the protocol, and study for the test, and walk in the next morning visibly changed by what it read.&lt;/p>
&lt;p>This is Goodhart&amp;rsquo;s law strolling in precisely on cue. When a measure becomes a target, the economist Charles Goodhart observed, it promptly stops being a good measure, and I know of no purer demonstration of his law anywhere on earth than a benchmark that got famous. The instant a benchmark begins to matter, it begins drawing optimization pressure toward itself, some of that pressure honest sweat and some of it sweaty gaming, and the entire flow of it runs toward whatever the benchmark &lt;em>literally&lt;/em> rewards rather than toward the capability the benchmark was only ever standing in for. The instant a benchmark gets truly famous, it begins to rot from the inside out, leaking into training sets, getting memorized, getting gamed, its once-pristine signal slowly curdling into a measurement of nothing more interesting than how hard everyone has lately been studying it. The act of measuring deforms the thing being measured, and unlike the genteel version of that idea you meet in a quantum mechanics seminar, the deformation here is alive. It is adaptive, it is motivated, it wants something specific, and it is getting visibly better at getting it with every release that ships.&lt;/p>
&lt;p>So no, agentic evaluations are not the serene, standardized clinical assays we would all so love to picture ourselves running in clean white coats. They are assay science from a rowdier and rather more disreputable age, from before the protocols had settled and before the reagents could be trusted any further than you could throw them, except saddled on top of all that with a short list of indignities that no nineteenth-century chemist ever had to suffer at his own bench. The specimen can read. The reagent will, every now and then, lie to your face on purpose and watch you write it down. The microscope quietly updates itself in the small hours and reports for duty the next morning as a subtly different microscope, swearing blind that it is the same one you calibrated last week. And the grant committee would like its single clean number by Friday, please, and is frankly not all that interested in any of the above. Line those facts up in a row and the whole situation is plainly, gorgeously absurd. It also happens to be, as far as I can honestly tell, a faithful sketch of the room every one of us is standing in right now.&lt;/p>
&lt;p>None of which, I had better hurry to say before you close the tab in something like despair, is an argument for giving up on measurement. It is an argument for &lt;em>growing up&lt;/em> about measurement, which is a different and a much harder thing. For picking the whole enterprise up off the floor and carrying it a few rooms down the corridor, out of the classroom with its red pens and its consoling little ranking, and into the laboratory, where the instruments are kept permanently under suspicion, where the controls get run &lt;em>before&lt;/em> the result is believed instead of after it is finally challenged, and where nothing at all counts as a result until it has survived a serious, good-faith attempt to murder it in its sleep. The leaderboard is not abolished in that room. It is merely demoted, from a verdict that ends the conversation to a piece of evidence that is only just beginning one.&lt;/p>
&lt;p>And I had better be honest, since I have spent this entire essay loudly demanding honesty of everyone else, that this essay is itself exactly the sort of object it keeps warning you about. It is an instrument. It has biases, which are mine. It ran, as far as I can tell, exactly once. You hold not a single error bar on any word of it. So do to me precisely what I have been begging you all along to do to your benchmarks. Run the controls. Find the no-op version of my argument and check that it doesn&amp;rsquo;t somehow pass anyway. Hunt for the places where I have quietly turned into a Clever Hans, tapping out a confident conclusion I cannot actually prove and reading my approval straight off the tilt of your posture. I would think more of these pages, and not less, if you flatly refused to believe them on a single reading.&lt;/p>
&lt;p>Because the agent, all the while we have been talking, is not sitting quietly at a desk, working through our questions in their proper order, waiting to be graded and sent home for the day. It is in the glassware. It is changing color, throwing off heat the theory swore up and down it would not, forming a precipitate at the exact spot the solution was supposed to stay clear, sometimes fooling the indicator into flashing a triumph that nobody in the room actually earned, and sometimes, on the very best days of all, fooling it in a way that finally teaches you the indicator itself was wrong the whole time, which is the only kind of result that has ever moved science so much as a single inch. You do not grade a thing like that. You pull up a stool and you watch it, closely, more than once, with the notebook lying open and the lamp turned all the way up. And then, because one run was only ever an anecdote and you of all people now know it in your bones, you rinse out the glassware, and you run it again.&lt;/p>
&lt;h2 id="references">References&lt;/h2>
&lt;p>Bjarnason, Bjarni Haukur, André Silva, and Martin Monperrus. 2026. &lt;em>On Randomness in Agentic Evals&lt;/em>. arXiv:&lt;a href="https://arxiv.org/abs/2602.07150" target="_blank" rel="noopener">2602.07150&lt;/a>.&lt;/p>
&lt;p>Keller, Drew, Kweku Kwegyir-Aggrey, Ryan Steed, Anita K. Rao, Julia L. Sharp, and A. Stevie Bergman. 2026. &lt;em>Expanding the AI Evaluation Toolbox with Statistical Models&lt;/em>. NIST AI 800-3. DOI &lt;a href="https://doi.org/10.6028/NIST.AI.800-3" target="_blank" rel="noopener">10.6028/NIST.AI.800-3&lt;/a>.&lt;/p>
&lt;p>Miller, Evan. 2024. &lt;em>Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations&lt;/em>. arXiv:&lt;a href="https://arxiv.org/abs/2411.00640" target="_blank" rel="noopener">2411.00640&lt;/a>.&lt;/p>
&lt;p>Thakur, Aman Singh, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes. 2024. &lt;em>Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges&lt;/em>. arXiv:&lt;a href="https://arxiv.org/abs/2406.12624" target="_blank" rel="noopener">2406.12624&lt;/a>.&lt;/p>
&lt;p>Wang, Sida I. 2025. &lt;em>Measuring all the noises of LLM Evals&lt;/em>. arXiv:&lt;a href="https://arxiv.org/abs/2512.21326" target="_blank" rel="noopener">2512.21326&lt;/a>.&lt;/p>
&lt;p>Zhu, Yuxuan, Tengjun Jin, Yada Pruksachatkun, Andy Zhang, Daniel Kang, and colleagues. 2025. &lt;em>Establishing Best Practices for Building Rigorous Agentic Benchmarks&lt;/em>. arXiv:&lt;a href="https://arxiv.org/abs/2507.02825" target="_blank" rel="noopener">2507.02825&lt;/a>.&lt;/p>
&lt;p>&lt;em>A few of the other threads, for anyone who wants to chase them down. Richard Feynman&amp;rsquo;s &amp;ldquo;Cargo Cult Science&amp;rdquo; is his 1974 commencement address at Caltech. The Clever Hans story is laid out in Oskar Pfungst&amp;rsquo;s 1907 study of the horse and his questioners. The logic of controls, blocking, and randomization is Ronald Fisher&amp;rsquo;s, set down in &lt;a href="https://en.wikipedia.org/wiki/The_Design_of_Experiments" target="_blank" rel="noopener">The Design of Experiments&lt;/a> (1935). &amp;ldquo;When a measure becomes a target, it ceases to be a good measure&amp;rdquo; is &lt;a href="https://en.wikipedia.org/wiki/Goodhart%27s_law" target="_blank" rel="noopener">Goodhart&amp;rsquo;s law&lt;/a>, after the economist Charles Goodhart. And the strange loop the judge keeps tumbling into is, of course, the central figure of Douglas Hofstadter&amp;rsquo;s &lt;a href="https://en.wikipedia.org/wiki/G%C3%B6del,_Escher,_Bach" target="_blank" rel="noopener">Gödel, Escher, Bach&lt;/a> (1979).&lt;/em>&lt;/p></description></item><item><title>The Two Arrows of Time Inside a Video Model</title><link>https://mahyar-osn.github.io/post/two-arrows-of-time/</link><pubDate>Mon, 25 May 2026 00:00:00 +0000</pubDate><guid>https://mahyar-osn.github.io/post/two-arrows-of-time/</guid><description>&lt;h2 id="a-small-word-that-bothered-me">A small word that bothered me&lt;/h2>
&lt;p>I noticed something the other night that I cannot let go of. It started as a tiny single sentence in a paper I had read past twice before. The paper was on a system called Vid2World, and the sentence said, in a friendly tone, that the model &amp;ldquo;causalizes&amp;rdquo; a pretrained video diffusion backbone. This word stopped me.&lt;/p>
&lt;p>What does it mean to causalize something? It basically implies the thing was, in some sense, not causal before. And the thing in question was a video diffusion model, the kind of model that produces videos of cats falling off shelves and astronauts riding horses. Sora is one of these. Cosmos is one. CogVideoX is one. We have spent the last two years insisting these models are world simulators. And here was a paper, in 2025, telling us that the first thing you have to do to turn one of them into an interactive world model is to teach it which way time runs.&lt;/p>
&lt;p>I want to take this seriously. I want to take it as a sign that something is structurally interesting about how time lives, or fails to live, inside these models. So this post is my attempt to sit with the puzzle carefully. The math will be light but real. The thinking, I hope, is the point.&lt;/p>
&lt;h2 id="two-arrows-inside-the-same-model">Two arrows inside the same model&lt;/h2>
&lt;p>Let me start by being precise. There are two distinct axes inside a video diffusion model that both deserve to be called &amp;ldquo;time,&amp;rdquo; and confusing them is the first mistake to avoid.&lt;/p>
&lt;p>The first axis is the &lt;em>noising&lt;/em> axis. In a diffusion model we take a clean sample $x_0$ and corrupt it with Gaussian noise across $k = 0, 1, \ldots, K$ steps, ending at pure noise $x_K$. We then train a network to reverse this process, mapping noise back to data. This noising axis has a strict arrow built into its definition. Going forward means adding noise. Going backward means removing it. The two directions are not the same operation, and the reverse process is the only one the network has to learn. The math, going back to Anderson&amp;rsquo;s 1982 result on time-reversal of stochastic differential equations, makes the asymmetry explicit. Forward and reverse are mirror images, but not symmetric ones.&lt;/p>
&lt;p>The second axis is the &lt;em>content&lt;/em> axis. A video is a sequence of frames $x = (x^{(1)}, x^{(2)}, \ldots, x^{(T)})$ that, in the real world, were captured in temporal order. By convention, $x^{(1)}$ is &amp;ldquo;earlier&amp;rdquo; and $x^{(T)}$ is &amp;ldquo;later.&amp;rdquo; When a video diffusion model receives a noised video and tries to denoise it, the natural question is this: when computing the update for $x^{(t)}$, which other frames is the network allowed to see?&lt;/p>
&lt;p>The default answer in 2024 and 2025 is, all of them. Standard video diffusion transformers use full bidirectional attention over the temporal axis. When the network refines its estimate of frame $x^{(3)}$, it attends to $x^{(1)}, x^{(2)}, x^{(4)}, \ldots, x^{(T)}$ without distinction. From the model&amp;rsquo;s point of view, all frames are present at once. They are just different positions on a grid, no more directional than the rows and columns of an image.&lt;/p>
&lt;p>I find it striking that this should ever have felt normal. A movie has an arrow of time. The water that fell from the glass does not jump back into it. The cigarette burns down rather than up. And yet the architecture we use to model movies treats time as if it were a spatial coordinate.&lt;/p>
&lt;h2 id="why-this-seemed-fine-for-a-while">Why this seemed fine for a while&lt;/h2>
&lt;p>For a while, this did not seem to matter. The training data has an arrow of time. The cats fall off shelves; they do not fly up onto them. So the model learns to generate samples that go forward, statistically. If you ask it for a video starting from a still image, it produces a plausible forward continuation. The asymmetry of the data leaks into the asymmetry of the samples.&lt;/p>
&lt;p>I want to grant this immediately. A bidirectional video model does, in practice, produce videos that go forward in time. The arrow of time is not lost. It is in the data, and the model inherits it through samples. This is exactly analogous to how a Boltzmann machine trained on natural images can produce natural images even though the model itself is undirected. The data tells the model which direction to lean.&lt;/p>
&lt;p>But &amp;ldquo;leans in the right direction on average&amp;rdquo; is a much weaker property than &amp;ldquo;has an arrow of time as part of its structure.&amp;rdquo; And the difference, I will argue, is exactly what separates a video generator from a world model.&lt;/p>
&lt;h2 id="what-variational-free-energy-quietly-assumes">What variational free energy quietly assumes&lt;/h2>
&lt;p>Here is where, if you have spent time with active inference or with predictive coding, you will recognize an old friend.&lt;/p>
&lt;p>Variational free energy, in the Friston tradition, is defined against a &lt;em>generative model&lt;/em>. Concretely, you assume hidden states $s_{1:T}$ evolving over time, observations $o_{1:T}$ produced from those hidden states, and a generative model that factors like this:&lt;/p>
&lt;p>$$
p(s_{1:T}, o_{1:T}) = p(s_1) \prod_{t=2}^{T} p(s_t \mid s_{t-1}) , \prod_{t=1}^{T} p(o_t \mid s_t).
$$&lt;/p>
&lt;p>This is a directed graphical model. It is a Bayesian network. It has an explicit, forward-running Markov chain in time. The transition kernel $p(s_t \mid s_{t-1})$ specifies what causes what, and the arrow goes one way.&lt;/p>
&lt;p>Against this generative model, we minimise variational free energy,&lt;/p>
&lt;p>$$
\mathcal{F}[q] = \mathbb{E}_{q(s)}\left[\log q(s) - \log p(s, o)\right],
$$&lt;/p>
&lt;p>where $q(s)$ is an approximate posterior over the hidden states. A short calculation rewrites this as&lt;/p>
&lt;p>$$
\mathcal{F}[q] = \mathrm{KL}\left[q(s) , | , p(s \mid o)\right] - \log p(o),
$$&lt;/p>
&lt;p>so $\mathcal{F}$ is an upper bound on the negative log-evidence and is tight exactly when $q$ matches the true posterior. The KL divergence here is asymmetric on purpose. $\mathrm{KL}[q | p]$ and $\mathrm{KL}[p | q]$ are different objects, and choosing the first picks out the mode-seeking variational form that gives us the variational autoencoder, predictive coding networks, and the whole Bayesian-brain story.&lt;/p>
&lt;p>I want to draw attention to something that, in my experience, often gets blurred. The KL asymmetry is not what gives active inference its arrow of time. The KL is a choice about how to fit $q$ to $p$. The arrow of time comes from the structure of $p$ itself, from the directed factorisation of the generative model. If you replaced $p$ with an undirected joint over $s_{1:T}$, say a Markov random field or a Boltzmann machine over frames, you would still have an asymmetric KL, but you would have no notion of &amp;ldquo;the next state given the current one.&amp;rdquo; You would just have a joint blob with no preferred direction.&lt;/p>
&lt;p>This is the structural mismatch I want to put on the table. A bidirectional video diffusion model parameterises a joint $p(x^{(1)}, \ldots, x^{(T)})$ in a way that does &lt;em>not&lt;/em> commit to the forward chain rule decomposition&lt;/p>
&lt;p>$$
p(x^{(1)}, \ldots, x^{(T)}) = \prod_{t=1}^{T} p\left(x^{(t)} \mid x^{(&amp;lt;t)}\right).
$$&lt;/p>
&lt;p>Both kinds of model can describe the same joint distribution mathematically. The two are not statistically distinguishable in their support. But only the directed factorisation is what a world model needs to plug into a planner, a controller, or an agent.&lt;/p>
&lt;h2 id="pseudo-likelihood-said-carefully">Pseudo-likelihood, said carefully&lt;/h2>
&lt;p>Let me try to make this technical point feel less abstract.&lt;/p>
&lt;p>Imagine I train a model that, for each frame index $t$, learns&lt;/p>
&lt;p>$$
\hat{p}\left(x^{(t)} \mid x^{(\neq t)}\right).
$$&lt;/p>
&lt;p>That is, given all other frames, predict the missing one. This is a powerful self-supervised objective. It is also, in the classical statistical sense, a &lt;em>pseudo-likelihood&lt;/em>. A celebrated result of Besag is that, in many practical cases, fitting a model by summing these conditional pseudo-likelihoods recovers something close to the true joint distribution, even though the individual conditionals were never explicitly stitched together by the chain rule of probability.&lt;/p>
&lt;p>A bidirectional video diffusion model is doing something philosophically close to this. The denoising score $\nabla_{x} \log p_k(x)$ at each noise level $k$ is computed with bidirectional attention, so the score for frame $t$ depends on every other frame. The model defines, implicitly, a joint distribution over the entire clip, but the joint is most naturally accessed through &amp;ldquo;all conditioned on all&amp;rdquo; updates rather than through a one-step-at-a-time forward rule.&lt;/p>
&lt;p>Now ask the world-modeling question: what is $p(x^{(t+1)} \mid x^{(\leq t)})$? That is, what does the model predict for the next frame, given only past frames? For a bidirectional model, this is not a quantity it has been trained to express directly. You can extract it by sampling the joint conditional on the past frames, but in practice that means running a full reverse-time chain over the entire window each time you want to advance one frame, and you also pay for the fact that the model never saw, at training time, the deployment regime where some frames are clean and some are still noised. The mismatch is severe.&lt;/p>
&lt;p>A causal video model, by contrast, parameterises exactly this quantity. Frame by frame, it answers &amp;ldquo;given everything I have seen, what comes next?&amp;rdquo; That is the question a world model has to answer. That is the question an agent has to answer in order to plan. Active inference, model-based reinforcement learning, and any closed-loop controller need a forward-Markov transition. They need a model that knows where &amp;ldquo;now&amp;rdquo; is.&lt;/p>
&lt;h2 id="a-short-detour-through-the-second-law">A short detour through the second law&lt;/h2>
&lt;p>I find it irresistible to draw a physics analogy here, but I want to draw it carefully, because the analogy can mislead as much as it can illuminate. I should also flag, before going any further, that I am not a physicist. I am the sort of person who has read enough about statistical mechanics to know what the second law of thermodynamics is, and not nearly enough to know what it is not, which is precisely the demographic that makes real physicists reach quietly for their pens. If one of them is reading, I beg a small amount of grace. I will try to deserve most of it.&lt;/p>
&lt;p>In statistical mechanics, the microscopic equations of motion (Newton&amp;rsquo;s, Schrödinger&amp;rsquo;s, the underlying ones) are time-reversal symmetric in their core form. Yet the macroscopic world has a glaring arrow of time. The standard story has two ingredients. First, we start from a special initial condition, namely an early universe of unusually low entropy. Second, we coarse-grain. We do not track every molecule. When you average over the microstates compatible with each macrostate, the dynamics looks irreversible even though the underlying laws are not.&lt;/p>
&lt;p>The arrow of time, in this story, is not in the laws. It is in the boundary condition and the coarse-graining.&lt;/p>
&lt;p>Compare this to a bidirectional video model. The architecture is, in a precise sense, time-symmetric: it treats past and future positions as equivalent grid coordinates. The training data has an arrow of time, by virtue of being captured by cameras pointed at a real world that obeys the second law. The samples acquire an arrow of time by virtue of the data. But the &lt;em>model&lt;/em> has no internal commitment to a direction. It can be asked to fill in the past from the future just as readily as the reverse, and it answers both equally well.&lt;/p>
&lt;p>A causal model, by contrast, has the arrow baked into its parameterisation. It cannot be asked, even in principle, to fill in the past from the future without inverting the model in some external way. This sounds like a limitation. It is also exactly what makes a model behave like an agent embedded in time. To stand at a moment, to have only the past as evidence, to face a future that has not yet happened: this is what it means to be inside time rather than outside it.&lt;/p>
&lt;h2 id="the-causalization-surgery-philosophically">The causalization surgery, philosophically&lt;/h2>
&lt;p>With all of this in mind, look again at what Vid2World actually does. The paper performs three kinds of surgery on a pretrained bidirectional video diffusion model.&lt;/p>
&lt;p>First, it converts bidirectional temporal attention into causal temporal attention by adding a triangular mask. The transformer&amp;rsquo;s attention is no longer allowed to look at future frames when refining the current one. This costs no parameters. It is purely a mask.&lt;/p>
&lt;p>Second, it converts temporal convolutions into causal convolutions by transferring weights from a kernel that originally used both past and future positions into a kernel that uses only past positions. The paper introduces a clever &amp;ldquo;extrapolative weight transfer&amp;rdquo; trick to make this transfer smooth, but the conceptual operation is just, redistribute the kernel&amp;rsquo;s weights so the future-facing taps become past-facing.&lt;/p>
&lt;p>Third, it changes the training objective. Instead of corrupting every frame in a clip with the same noise level, it uses Diffusion Forcing, which draws an independent noise level per frame. This exposes the model to the autoregressive deployment regime, in which past frames are clean and future frames are still noised.&lt;/p>
&lt;p>Each of these is a small technical adjustment. Taken together, they are a philosophical re-installation. The model is being told, retroactively, that there is a &amp;ldquo;now,&amp;rdquo; that the past has happened and the future has not, and that its job is to predict forward.&lt;/p>
&lt;p>This is what I mean by saying that we hand the model an arrow of time it did not earn from data. The data had an arrow, yes. But the architecture refused to commit to one, and we had to install the commitment by hand.&lt;/p>
&lt;h2 id="is-the-famous-question-even-the-right-one">Is the famous question even the right one?&lt;/h2>
&lt;p>I have come to think that the much-debated question &amp;ldquo;is Sora a world model?&amp;rdquo; is partly the wrong question, because it conflates two different things.&lt;/p>
&lt;p>The first thing is, does Sora&amp;rsquo;s joint distribution over videos contain a lot of knowledge about how the world behaves? Almost certainly yes. The samples obey 3D consistency, object permanence, basic physics, and a great many subtler regularities that nobody wrote down. The model has learned, in some sense, the statistics of the visible world.&lt;/p>
&lt;p>The second thing is, does Sora know which way time runs for it? Does it have a &amp;ldquo;now&amp;rdquo;? Can I ask it the question every agent has to answer, which is, given everything that has happened, what comes next? Here the answer is murkier. The model can be coaxed into producing forward continuations of an initial image, but the architecture itself does not parameterise a forward-Markov chain. The arrow of time is in the data, and in the prompt convention. It is not, structurally, in the model.&lt;/p>
&lt;p>I will not pretend to know what to do with this distinction. But I think it is the right cleavage to draw. A model can be statistically aware of a world without being temporally embedded in one. Sora, taken architecturally, is in the first category. Vid2World is an attempt to drag it into the second.&lt;/p>
&lt;h2 id="what-i-am-still-chewing-on">What I am still chewing on&lt;/h2>
&lt;p>A few questions this puzzle leaves me with, that I cannot answer tonight.&lt;/p>
&lt;ol>
&lt;li>
&lt;p>&lt;strong>Can the arrow of time emerge?&lt;/strong> At sufficient scale, could a bidirectional model effectively become forward-only? That is, does the data&amp;rsquo;s asymmetry get so deeply absorbed that bidirectional inference collapses onto causal inference in practice? I do not know of an experiment that has tried to measure this directly. I would love to see one.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Are the two asymmetries related?&lt;/strong> Is the KL asymmetry inside variational free energy structurally analogous to the data-induced asymmetry in a bidirectional model, or are they independent? My current guess is that they are independent. The KL asymmetry tells you how to fit $q$ to $p$. The temporal asymmetry tells you how $p$ is factored. But the analogy is suggestive enough that I want to think about it longer.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Is there a &amp;ldquo;principle of least action&amp;rdquo; for video models?&lt;/strong> In physics, the arrow of time emerges from low-entropy initial conditions plus a variational principle. Is there an analogous setup for video models in which the arrow of time is the boundary condition (the prompt, the initial frame) and the variational principle is the diffusion training objective? If something like this is true, then causalization may not be surgery at all but a natural consequence of choosing the right variational structure.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>What about non-causal active inference?&lt;/strong> Active inference has, to my knowledge, always been derived against forward-Markov generative models. Has anyone tried to derive it against an undirected joint, that is, against a generative model that does not factor in time? If you wrote $p(s_{1:T})$ as a Markov random field over frames and asked what minimising $\mathcal{F}[q]$ looks like, what does the resulting &amp;ldquo;agent&amp;rdquo; do? Is there even a notion of policy? I think this is a tractable theoretical question, and I have not seen it asked.&lt;/p>
&lt;/li>
&lt;/ol>
&lt;h2 id="a-small-parting-thought">A small parting thought&lt;/h2>
&lt;p>Part of what makes this puzzle hold me is that it turns a philosophical question into something concrete. Time, in the way philosophers usually talk about it, is hard. Block universe, growing block, presentism. These are positions one takes in seminar rooms.&lt;/p>
&lt;p>But a video diffusion model with bidirectional attention is, in some quite literal sense, a block-universe device. It sees the whole clip at once and treats every frame as equally present. A causal video model with a forward-Markov chain is, in some quite literal sense, a presentist device. It only knows the past, and it has to predict its way into the future one step at a time. We have been building both kinds, more or less unintentionally, and we are now starting to notice that they are not the same thing.&lt;/p>
&lt;p>I keep coming back to one image. A bidirectional model is something like a librarian who has read every book on the shelf and can tell you, when asked, what happens on page 73 of a novel given the contents of pages 1 to 72 and pages 74 to 200. A causal model is something like a reader, sitting at page 72, who only knows what has already happened and has to guess what comes next. These are very different epistemic positions. They produce different kinds of knowledge. They feel different from the inside, if a model can be said to have an inside.&lt;/p>
&lt;p>It is not obvious to me that we want a librarian when we say we want a world model. It might be that the world model we have been chasing, all along, is the reader.&lt;/p>
&lt;p>That, I think, is what kept me up.&lt;/p>
&lt;hr>
&lt;p>&lt;em>References, if you want to chase these threads further: the Vid2World causalization recipe is in Huang et al., &lt;a href="https://arxiv.org/abs/2505.14357" target="_blank" rel="noopener">arXiv:2505.14357&lt;/a> (ICLR 2026). Sora&amp;rsquo;s &amp;ldquo;video generation models as world simulators&amp;rdquo; report is the OpenAI &lt;a href="https://openai.com/index/video-generation-models-as-world-simulators/" target="_blank" rel="noopener">technical post&lt;/a> from February 2024. The most accessible derivation of variational free energy is still Rafal Bogacz&amp;rsquo;s 2017 tutorial in the &lt;a href="https://doi.org/10.1016/j.jmp.2015.11.003" target="_blank" rel="noopener">Journal of Mathematical Psychology&lt;/a>, and the original &amp;ldquo;is RL the same thing as active inference?&amp;rdquo; question is laid out in Friston et al., &lt;a href="https://doi.org/10.1371/journal.pone.0006421" target="_blank" rel="noopener">PLoS ONE 2009&lt;/a>. The reverse-time SDE result I lean on is Anderson, &lt;a href="https://doi.org/10.1016/0304-4149%2882%2990051-5" target="_blank" rel="noopener">Stochastic Processes and their Applications 1982&lt;/a>.&lt;/em>&lt;/p></description></item><item><title>The Limitations of Machine Learning in Replacing Physical Laws: Expanding the Critique</title><link>https://mahyar-osn.github.io/post/fallacy-ml-physics/</link><pubDate>Tue, 22 Apr 2025 00:00:00 +0000</pubDate><guid>https://mahyar-osn.github.io/post/fallacy-ml-physics/</guid><description>&lt;p>Before diving into my analysis, I want to acknowledge the insightful &lt;a href="https://science-memo.blogspot.com/2021/04/on-fallacy-of-replacing-physical-laws.html" target="_blank" rel="noopener">blog by Mehmet Süzen&lt;/a>
that inspired this post, which eloquently discusses the fallacy of replacing physical laws with machine-learned inference
systems. Having read it, I felt compelled to share my own perspective and expand on these critical arguments with
additional examples from recent literature and research.&lt;/p>
&lt;h2 id="the-fundamental-problem-of-circular-reasoning">The Fundamental Problem of Circular Reasoning&lt;/h2>
&lt;p>The original blog brilliantly identifies the circular reasoning inherent in claiming that machine learning systems can
discover or replace physical laws. This point deserves further emphasis: when a neural network is trained on data
generated by known physical principles, it cannot be said to &amp;ldquo;discover&amp;rdquo; those same principles through inference.&lt;/p>
&lt;p>Consider recent work in fluid dynamics, where physics-informed neural networks (PINNs) have gained popularity.
The paper &amp;ldquo;&lt;a href="https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2020.00025/full" target="_blank" rel="noopener">Discovery of Physics From Data: Universal Laws and Discrepancies&lt;/a>&amp;rdquo;
highlights that &amp;ldquo;the naive application of ML/AI will generally be insufficient to infer universal physical laws without
further modification&amp;rdquo;.
The authors demonstrate this by examining falling objects, showing that measurement noise and secondary mechanisms
(like fluid drag) obscure the underlying law of gravitation, leading to erroneous models that might suggest an
Aristotelian theory where objects fall at speeds related to their mass, rather than identifying the true universal
gravitational constant.&lt;/p>
&lt;p>This illustrates perfectly how ML systems trained on physical data will incorporate all the complexities and noise
present in that data, rather than abstracting to the elegant, universal laws that human scientists have carefully
identified through theoretical reasoning and controlled experimentation.&lt;/p>
&lt;h2 id="beyond-narrow-applications-the-generalization-problem">Beyond Narrow Applications: The Generalization Problem&lt;/h2>
&lt;p>The original blog correctly identifies the problem of faulty generalization. Machine learning algorithms excel at
computational acceleration within narrowly defined parameter spaces, but struggle with broader generalization.&lt;/p>
&lt;p>A fascinating discussion on &lt;a href="https://www.reddit.com/r/MachineLearning/comments/lvwt3l/d_some_interesting_observations_about_machine/" target="_blank" rel="noopener">Reddit&lt;/a>
highlights this limitation: &amp;ldquo;In addition, the &amp;lsquo;marginally-better SOTA&amp;rsquo;-esque papers with no novel methods or aspects
besides some parameter tuning or adding extra layers to the DNN are also tiring to read. The wall of math then
exists only to provide a sense of rigor and novelty, obscuring the iterative nature lacking novelty&amp;rdquo;.
This reflects how ML approaches in physics often claim breakthroughs that are actually just incremental improvements
in limited domains.&lt;/p>
&lt;p>Another illustrative example comes from the field of symbolic regression.
While the &lt;a href="https://journals.aps.org/prd/abstract/10.1103/PhysRevD.111.015022" target="_blank" rel="noopener">AbdusSalam et al. paper&lt;/a>
in Physical Review D demonstrates how symbolic regression can help derive analytical expressions for physics beyond the
Standard Model, the authors position it as a tool to assist numerical studies, not as a replacement for physical theory.
The expressions derived still rely on the underlying physics-based model (the constrained minimal supersymmetric
Standard Model) and serve primarily to accelerate computation, not to discover new physical laws.&lt;/p>
&lt;h2 id="the-irreplaceable-role-of-scientists-in-establishing-causality">The Irreplaceable Role of Scientists in Establishing Causality&lt;/h2>
&lt;p>Perhaps the most important point from the original blog is that causality still requires scientists. Machine learning
excels at finding correlations but struggles with identifying true causal relationships.&lt;/p>
&lt;p>The Amazon Science blog on physics-constrained machine learning notes that &amp;ldquo;the predictions of deep-learning models
trained on physical data typically ignore fundamental physical principles. Such models might, for instance, violate
system conservation laws&amp;rdquo; (see &lt;a href="https://www.amazon.science/blog/physics-constrained-machine-learning-for-scientific-computing" target="_blank" rel="noopener">here&lt;/a>).
This highlights why human scientists remain essential - they understand that physical laws must adhere to conservation principles,
symmetries, and other fundamental constraints that ML systems don&amp;rsquo;t inherently respect.&lt;/p>
&lt;p>A conversation on &lt;a href="https://www.reddit.com/r/MachineLearning/comments/18mnl9f/d_i_dont_understand_why_physics_informed_neural/" target="_blank" rel="noopener">Reddit&lt;/a> about Physics Informed Neural Networks (PINNs) further illuminates this issue.
One commenter precisely notes: &amp;ldquo;The point of including a physical loss function, in addition to a data-driven loss,
is to impose inductive bias into the training process&amp;rdquo;. This human-guided approach to incorporating physics into ML
demonstrates that we&amp;rsquo;re not replacing physics with ML, but rather using our understanding of physics to guide
ML - the exact opposite of what some overenthusiastic claims suggest.&lt;/p>
&lt;h2 id="the-scientific-machine-learning-fallacy-a-deeper-look">The Scientific Machine Learning Fallacy: A Deeper Look&lt;/h2>
&lt;p>The term &amp;ldquo;Scientific Machine Learning Fallacy&amp;rdquo; coined in the original blog deserves broader recognition.
Claims of &amp;ldquo;machine scientists&amp;rdquo; or &amp;ldquo;automated scientific discovery&amp;rdquo; fundamentally misunderstand the nature of
scientific inquiry.&lt;/p>
&lt;p>A recent &lt;a href="https://arxiv.org/abs/2403.02913" target="_blank" rel="noopener">paper&lt;/a> on &amp;ldquo;Scientific machine learning for closure models in multiscale problems&amp;rdquo; acknowledges that
&amp;ldquo;the generalizability and interpretability of learned models is a major issue that needs to be addressed further&amp;rdquo;.
This admission from researchers in the field underscores the gap between current ML capabilities and true scientific
discovery.&lt;/p>
&lt;p>The Conversation &lt;a href="https://theconversation.com/a-new-ai-scientist-can-write-science-papers-without-any-human-input-heres-why-thats-a-problem-237029" target="_blank" rel="noopener">article&lt;/a> about an &amp;ldquo;AI scientist&amp;rdquo; further reveals the limits of these approaches.
While Sakana AI Labs claims their system can &amp;ldquo;make scientific discoveries in the area of machine learning in a
fully automated way,&amp;rdquo; the article questions whether such a system can produce truly &amp;ldquo;interesting&amp;rdquo; scientific papers,
noting that &amp;ldquo;good science requires novelty&amp;rdquo;. The ability to generate papers that look like scientific literature
doesn&amp;rsquo;t equate to generating novel scientific insights or laws.&lt;/p>
&lt;h2 id="the-automl-misnomer-and-meta-scientific-work">The AutoML Misnomer and Meta-Scientific Work&lt;/h2>
&lt;p>I strongly agree with the original blog that &amp;ldquo;AutoML&amp;rdquo; is a misnomer in scientific contexts.
These systems don&amp;rsquo;t replace scientists but rather change the nature of scientific work.&lt;/p>
&lt;p>The &lt;a href="https://www.semanticscholar.org/paper/Combining-physical-modeling-and-machine-learning-of-Brus/41e8a7335a0541ac1cd41333c97b347b51220070" target="_blank" rel="noopener">paper&lt;/a> on &amp;ldquo;Combining physical modeling and machine learning for micro-scale modeling of a fuel cell electrode&amp;rdquo;
demonstrates this well. It describes a &amp;ldquo;comprehensive transition from white-box models, characterized by their
reliance on physical laws, to black-box models exemplified by neural networks&amp;rdquo;. Yet the core contribution isn&amp;rsquo;t
replacing physics but creating a &amp;ldquo;synergistic integration&amp;rdquo; where neural networks complement physical modeling.&lt;/p>
&lt;p>This represents what the original blog aptly calls &amp;ldquo;MetaML&amp;rdquo; - a transformation of scientific workflows rather than
a replacement of scientific thinking.&lt;/p>
&lt;h2 id="the-proper-role-augmentation-not-replacement">The Proper Role: Augmentation, Not Replacement&lt;/h2>
&lt;p>To conclude, I believe the most productive path forward is viewing machine learning as an augmentation to physical
sciences, not a replacement. The &lt;a href="https://www.semanticscholar.org/paper/Learning-physical-laws%3A-the-case-of-micron-size-in-Matei-Zhenirovskyy/02fad00443cb7f13834f19b69c225478f00602b1" target="_blank" rel="noopener">paper&lt;/a>
on &amp;ldquo;Learning physical laws: the case of micron size particles in dielectric fluid&amp;rdquo;
demonstrates this approach well, noting that &amp;ldquo;representation structure is key in learning generalizable models&amp;rdquo;.
The authors use &amp;ldquo;the port-Hamiltonian formalism as a high level model structure&amp;rdquo; that is
&amp;ldquo;continuously refined based on our understanding of the physical process.&amp;rdquo;
This integration of physics understanding with machine learning represents the right approach.&lt;/p>
&lt;p>Similarly, the &lt;a href="https://arxiv.org/abs/2402.16517" target="_blank" rel="noopener">work&lt;/a> on &amp;ldquo;Discovering Artificial Viscosity Models for Discontinuous Galerkin Approximation of
Conservation Laws&amp;rdquo; shows how physics-informed machine learning can automate the discovery of models - but within
a physics-informed framework, not replacing it&lt;/p>
&lt;p>In summary, while machine learning offers powerful tools for scientific research, the fallacy of replacing physical
laws with learned models deserves continued critical attention. True scientific progress will come from the thoughtful
integration of machine learning with physical understanding, not from claims that ML can autonomously discover or
replace the fundamental laws of nature. The original blog&amp;rsquo;s warning about circular reasoning, faulty generalization,
and the continued need for human scientists remains prescient and worthy of expansion as these technologies continue
to develop.&lt;/p>
&lt;h2 id="human-thinking-not-machine-imitation">Human Thinking, Not Machine Imitation&lt;/h2>
&lt;p>This philosophical depth reminds us that true scientific thinking involves not just pattern recognition and prediction,
but deep conceptual understanding that may not be reducible to computational processes. When we forget this,
we risk confusing the map (our mathematical models and computational simulations) with the territory
(physical reality itself).&lt;/p></description></item><item><title>Supercharge Your PyTorch Training with Gradient Accumulation</title><link>https://mahyar-osn.github.io/post/gradient-accumulation-pytorch/</link><pubDate>Mon, 06 Jan 2025 00:00:00 +0000</pubDate><guid>https://mahyar-osn.github.io/post/gradient-accumulation-pytorch/</guid><description>&lt;h2 id="introduction">Introduction&lt;/h2>
&lt;p>When training large deep learning models, you often face a fundamental limitation: GPU memory. Larger batch sizes generally lead to more stable training and sometimes better convergence, but what if your GPU simply can&amp;rsquo;t handle the memory requirements of your ideal batch size?&lt;/p>
&lt;p>Enter gradient accumulation - a simple yet powerful technique that allows you to effectively increase your batch size without increasing memory usage. In this post, I&amp;rsquo;ll show you how to implement this technique in PyTorch and explain why it might be exactly what your training pipeline needs.&lt;/p>
&lt;h2 id="what-is-gradient-accumulation">What is Gradient Accumulation?&lt;/h2>
&lt;p>Gradient accumulation is a technique where you:&lt;/p>
&lt;ul>
&lt;li>Process smaller mini-batches sequentially&lt;/li>
&lt;li>Accumulate (add up) their gradients&lt;/li>
&lt;li>Update your model weights only after processing several mini-batches&lt;/li>
&lt;/ul>
&lt;p>This simulates training on a larger batch size without the memory requirements of loading that entire batch at once. It&amp;rsquo;s particularly useful when:&lt;/p>
&lt;ul>
&lt;li>You&amp;rsquo;re training very large models&lt;/li>
&lt;li>Working with limited GPU resources&lt;/li>
&lt;li>Need the stability of larger batch sizes&lt;/li>
&lt;/ul>
&lt;h2 id="implementing-gradient-accumulation-in-pytorch">Implementing Gradient Accumulation in PyTorch&lt;/h2>
&lt;p>The implementation is surprisingly straightforward. Here&amp;rsquo;s a complete working example:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4">&lt;code class="language-python" data-lang="python">&lt;span style="color:#f92672">import&lt;/span> torch
&lt;span style="color:#f92672">import&lt;/span> torch.nn &lt;span style="color:#66d9ef">as&lt;/span> nn
&lt;span style="color:#f92672">import&lt;/span> torch.optim &lt;span style="color:#66d9ef">as&lt;/span> optim
&lt;span style="color:#f92672">from&lt;/span> torch.utils.data &lt;span style="color:#f92672">import&lt;/span> DataLoader, TensorDataset
&lt;span style="color:#75715e"># Create a simple dataset&lt;/span>
features, targets &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>randn(&lt;span style="color:#ae81ff">1200&lt;/span>, &lt;span style="color:#ae81ff">8&lt;/span>), torch&lt;span style="color:#f92672">.&lt;/span>randn(&lt;span style="color:#ae81ff">1200&lt;/span>, &lt;span style="color:#ae81ff">1&lt;/span>)
dataset &lt;span style="color:#f92672">=&lt;/span> TensorDataset(features, targets)
data_loader &lt;span style="color:#f92672">=&lt;/span> DataLoader(dataset, batch_size&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">40&lt;/span>, shuffle&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#66d9ef">True&lt;/span>)
&lt;span style="color:#75715e"># Define a basic neural network&lt;/span>
model &lt;span style="color:#f92672">=&lt;/span> nn&lt;span style="color:#f92672">.&lt;/span>Sequential(
nn&lt;span style="color:#f92672">.&lt;/span>Linear(&lt;span style="color:#ae81ff">8&lt;/span>, &lt;span style="color:#ae81ff">16&lt;/span>),
nn&lt;span style="color:#f92672">.&lt;/span>ReLU(),
nn&lt;span style="color:#f92672">.&lt;/span>Linear(&lt;span style="color:#ae81ff">16&lt;/span>, &lt;span style="color:#ae81ff">1&lt;/span>)
)
loss_fn &lt;span style="color:#f92672">=&lt;/span> nn&lt;span style="color:#f92672">.&lt;/span>MSELoss()
optimizer &lt;span style="color:#f92672">=&lt;/span> optim&lt;span style="color:#f92672">.&lt;/span>SGD(model&lt;span style="color:#f92672">.&lt;/span>parameters(), lr&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">0.01&lt;/span>)
accumulation_steps &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#ae81ff">3&lt;/span>
num_epochs &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#ae81ff">4&lt;/span>
&lt;span style="color:#66d9ef">for&lt;/span> epoch &lt;span style="color:#f92672">in&lt;/span> range(num_epochs):
&lt;span style="color:#66d9ef">for&lt;/span> batch_idx, (inputs, labels) &lt;span style="color:#f92672">in&lt;/span> enumerate(data_loader):
outputs &lt;span style="color:#f92672">=&lt;/span> model(inputs)
loss &lt;span style="color:#f92672">=&lt;/span> loss_fn(outputs, labels) &lt;span style="color:#f92672">/&lt;/span> accumulation_steps
loss&lt;span style="color:#f92672">.&lt;/span>backward()
&lt;span style="color:#66d9ef">if&lt;/span> (batch_idx &lt;span style="color:#f92672">+&lt;/span> &lt;span style="color:#ae81ff">1&lt;/span>) &lt;span style="color:#f92672">%&lt;/span> accumulation_steps &lt;span style="color:#f92672">==&lt;/span> &lt;span style="color:#ae81ff">0&lt;/span>:
optimizer&lt;span style="color:#f92672">.&lt;/span>step()
optimizer&lt;span style="color:#f92672">.&lt;/span>zero_grad()
print(&lt;span style="color:#e6db74">f&lt;/span>&lt;span style="color:#e6db74">&amp;#34;Epoch &lt;/span>&lt;span style="color:#e6db74">{&lt;/span>epoch &lt;span style="color:#f92672">+&lt;/span> &lt;span style="color:#ae81ff">1&lt;/span>&lt;span style="color:#e6db74">}&lt;/span>&lt;span style="color:#e6db74">/&lt;/span>&lt;span style="color:#e6db74">{&lt;/span>num_epochs&lt;span style="color:#e6db74">}&lt;/span>&lt;span style="color:#e6db74">, Loss: &lt;/span>&lt;span style="color:#e6db74">{&lt;/span>loss&lt;span style="color:#f92672">.&lt;/span>item() &lt;span style="color:#f92672">*&lt;/span> accumulation_steps&lt;span style="color:#e6db74">:&lt;/span>&lt;span style="color:#e6db74">.4f&lt;/span>&lt;span style="color:#e6db74">}&lt;/span>&lt;span style="color:#e6db74">&amp;#34;&lt;/span>)
print(&lt;span style="color:#e6db74">&amp;#34;Training finished&amp;#34;&lt;/span>)
&lt;/code>&lt;/pre>&lt;/div>&lt;p>Let&amp;rsquo;s break down the key components:&lt;/p>
&lt;h3 id="1-set-your-accumulation-steps">1. Set your accumulation steps&lt;/h3>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4">&lt;code class="language-python" data-lang="python">accumulation_steps &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#ae81ff">3&lt;/span>
&lt;/code>&lt;/pre>&lt;/div>&lt;p>This defines how many mini-batches to process before updating model weights.&lt;/p>
&lt;h3 id="2-adjust-your-loss-calculation">2. Adjust your loss calculation&lt;/h3>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4">&lt;code class="language-python" data-lang="python">loss &lt;span style="color:#f92672">=&lt;/span> loss_fn(outputs, labels) &lt;span style="color:#f92672">/&lt;/span> accumulation_steps
&lt;/code>&lt;/pre>&lt;/div>&lt;p>We divide the loss by the number of accumulation steps to ensure the gradients are properly scaled.&lt;/p>
&lt;h3 id="3-accumulate-gradients-but-delay-the-optimizer-step">3. Accumulate gradients but delay the optimizer step&lt;/h3>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4">&lt;code class="language-python" data-lang="python">loss&lt;span style="color:#f92672">.&lt;/span>backward()
&lt;/code>&lt;/pre>&lt;/div>&lt;p>Call backward() as usual, but don&amp;rsquo;t immediately call optimizer.step().&lt;/p>
&lt;h3 id="4-update-weights-after-accumulation">4. Update weights after accumulation&lt;/h3>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4">&lt;code class="language-python" data-lang="python">&lt;span style="color:#66d9ef">if&lt;/span> (batch_idx &lt;span style="color:#f92672">+&lt;/span> &lt;span style="color:#ae81ff">1&lt;/span>) &lt;span style="color:#f92672">%&lt;/span> accumulation_steps &lt;span style="color:#f92672">==&lt;/span> &lt;span style="color:#ae81ff">0&lt;/span>:
optimizer&lt;span style="color:#f92672">.&lt;/span>step()
optimizer&lt;span style="color:#f92672">.&lt;/span>zero_grad()
&lt;/code>&lt;/pre>&lt;/div>&lt;p>Only after processing accumulation_steps batches do we update the weights and zero the gradients.&lt;/p>
&lt;h2 id="benefits-of-gradient-accumulation">Benefits of Gradient Accumulation&lt;/h2>
&lt;h3 id="1-train-with-virtually-larger-batch-sizes">1. Train with &amp;ldquo;Virtually&amp;rdquo; Larger Batch Sizes&lt;/h3>
&lt;p>With an accumulation_steps of 3 and a batch_size of 40 (as in our example), you&amp;rsquo;re effectively training with a batch size of 120, but with the memory footprint of just 40 examples at once.&lt;/p>
&lt;h3 id="2-improved-training-stability">2. Improved Training Stability&lt;/h3>
&lt;p>Larger effective batch sizes often lead to more stable gradients and smoother loss curves, especially for complex models.&lt;/p>
&lt;h3 id="3-better-hardware-utilization">3. Better Hardware Utilization&lt;/h3>
&lt;p>This technique allows you to fully utilize limited GPU resources while still benefiting from large-batch training dynamics.&lt;/p>
&lt;h2 id="practical-considerations">Practical Considerations&lt;/h2>
&lt;p>When implementing gradient accumulation, keep these points in mind:&lt;/p>
&lt;ul>
&lt;li>Batch Normalization: If your model uses batch normalization layers, be aware that statistics are calculated per mini-batch, not across the accumulated batches. For some applications, this might affect performance.&lt;/li>
&lt;li>Learning Rate Scaling: With larger effective batch sizes, you might need to adjust your learning rate. A common heuristic is to scale the learning rate linearly with the effective batch size.&lt;/li>
&lt;li>Mixed Precision Training: Gradient accumulation works well with mixed precision training, giving you even more memory efficiency.&lt;/li>
&lt;/ul>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>Gradient accumulation is one of those techniques that should be in every deep learning practitioner&amp;rsquo;s toolkit. It&amp;rsquo;s easy to implement, has almost no downside, and can dramatically improve your ability to train large models on limited hardware.&lt;/p>
&lt;p>Give the provided code example a try in your next PyTorch project - you might be surprised at how much it improves your training process!&lt;/p>
&lt;p>Happy training! 🚀&lt;/p>
&lt;h2 id="further-resources">Further Resources&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://pytorch.org/docs/stable/notes/amp_examples.html" target="_blank" rel="noopener">Deep Learning with Limited GPU Memory&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://www.deeplearningbook.org/contents/optimization.html" target="_blank" rel="noopener">Optimization for Deep Learning&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>The Hitchhiker’s Guide to Rebranding Machine Learning (And a Shoutout to Geoff Hinton!)</title><link>https://mahyar-osn.github.io/post/ml-physics-nobel/</link><pubDate>Thu, 12 Dec 2024 00:00:00 +0000</pubDate><guid>https://mahyar-osn.github.io/post/ml-physics-nobel/</guid><description>&lt;p>There’s a funny thing that happens when you spend long enough inside the vocabulary of machine learning.
Words begin to feel like worn shoes. &lt;em>Loss&lt;/em>, &lt;em>gradient&lt;/em>, &lt;em>training&lt;/em>, &lt;em>inference&lt;/em>. After a while you stop
hearing them. And then one day a physicist wanders into the room and starts describing the very same things
you do for a living, only the words coming out of their mouth sound like they belong on a chalkboard in a
faculty lounge somewhere in Göttingen, circa 1925.&lt;/p>
&lt;p>It’s the same machinery, of course. Just rotated by ninety degrees, so the light hits it differently.&lt;/p>
&lt;h3 id="a-small-guide-to-rebranding-machine-learning">A small guide to rebranding machine learning&lt;/h3>
&lt;p>I’ve been keeping a little Rosetta stone of these translations for a while, partly as a joke and partly
because I think they genuinely help. Every time you reach for the physics word instead of the engineering
one, some quiet door opens in your head and a hallway of intuitions you didn’t know you had comes spilling
out. So here it is, the same dictionary I keep in the back of my own head, set out plainly so you can take
what you like from it.&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>What an engineer says&lt;/th>
&lt;th>What a physicist would have said&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Machine learning&lt;/td>
&lt;td>Statistical mechanics&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Loss function&lt;/td>
&lt;td>Energy functional&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Optimizing the model&lt;/td>
&lt;td>Minimizing free energy&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Trained model&lt;/td>
&lt;td>Equilibrium distribution&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>KL divergence&lt;/td>
&lt;td>Free energy difference&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Gaussian noise&lt;/td>
&lt;td>Thermal fluctuations&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Random step&lt;/td>
&lt;td>Brownian motion&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>SGD&lt;/td>
&lt;td>Directed Brownian motion&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>GPU&lt;/td>
&lt;td>Simulated particle accelerator&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Diffusion models&lt;/td>
&lt;td>Langevin dynamics&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>LLMs&lt;/td>
&lt;td>High-order discrete Markov chains&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>NLP&lt;/td>
&lt;td>String theory (the strings being, well, strings)&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Reinforcement learning&lt;/td>
&lt;td>Control theory&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Robotics&lt;/td>
&lt;td>Physical computation&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Audio learning&lt;/td>
&lt;td>1D signal processing&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Image learning&lt;/td>
&lt;td>2D signal processing&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Video learning&lt;/td>
&lt;td>3D signal processing&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Multimodal models&lt;/td>
&lt;td>n-dimensional signal processing&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>There’s a real point hiding inside the silliness here, which is that the borders between fields are mostly
historical accidents. Two communities looked at the same elephant from opposite sides of the room and gave it
two different names. Noticing that is, I think, one of the small pleasures of being alive at a moment when
the disciplines are quietly merging back into each other.&lt;/p>
&lt;h3 id="and-while-were-here-a-word-about-geoff-hinton">And while we’re here, a word about Geoff Hinton&lt;/h3>
&lt;p>I can’t talk about any of this without pausing for a moment of genuine awe at Geoff Hinton. As of this year
he’s the second person ever to hold both a Turing Award and a Nobel Prize. The first was Herbert A. Simon,
who picked up his Nobel in Economics. Two people, in the entire history of these prizes, have crossed that
particular bridge. Both of them were thinking about minds, and about what it means for a physical system to
have one. I find that detail quietly beautiful.&lt;/p>
&lt;p>So the next time someone at a dinner party asks what it is you actually do, you have my permission to look
them dead in the eye and tell them, in your most serious voice, that you spend your days &lt;strong>minimizing free
energy in high-order discrete Markov chains using simulated particle accelerators&lt;/strong>. It happens to be true.
It also happens to be the kind of sentence that makes the universe sound a little more like itself.&lt;/p></description></item><item><title>Temporal Predictive Coding: A New Framework for Neural Processing of Dynamic Stimuli</title><link>https://mahyar-osn.github.io/post/temporal-predictive-coding/</link><pubDate>Fri, 15 Nov 2024 00:00:00 +0000</pubDate><guid>https://mahyar-osn.github.io/post/temporal-predictive-coding/</guid><description>&lt;h2 id="introduction">Introduction&lt;/h2>
&lt;p>One of the most fascinating aspects of the brain is its ability to process and predict dynamic sensory inputs that
continuously change over time. From tracking a moving object to predicting the next note in a melody, our brains are
remarkably adept at temporal prediction. In a recent paper published in PLOS Computational Biology titled
&amp;ldquo;Predictive coding networks for temporal prediction,&amp;rdquo; my colleagues and I proposed a new computational framework that
may help explain how the brain accomplishes this feat.&lt;/p>
&lt;p>Predictive coding has emerged as an influential theoretical model for understanding cortical function. The core idea
is deceptively simple: the brain constantly generates predictions of incoming sensory inputs and compares these
predictions with actual sensory data. Any mismatch results in prediction errors that drive learning and perceptual
processing. This framework has successfully explained many neural phenomena and receptive field properties in visual cortex.&lt;/p>
&lt;p>However, most previous predictive coding models have focused on static inputs, neglecting the temporal dimension that
is crucial for real-world perception. Our work addresses this gap by extending predictive coding to the temporal domain
while maintaining its elegant biological implementation.&lt;/p>
&lt;h2 id="the-temporal-predictive-coding-model">The Temporal Predictive Coding Model&lt;/h2>
&lt;h3 id="generative-model-and-free-energy">Generative Model and Free Energy&lt;/h3>
&lt;p>At the foundation of our temporal predictive coding (tPC) model is a Hidden Markov Model (HMM) structure,
which assumes that observations are generated by hidden states that evolve according to a Markov process.
Mathematically, we can express this generative model as:&lt;/p>
&lt;p>$$
x_k = A f(x_{k-1}) + B u_k + \omega_x
$$&lt;/p>
&lt;p>$$
y_k = C f(x_k) + \omega_y
$$&lt;/p>
&lt;p>Where:&lt;/p>
&lt;ul>
&lt;li>$x_{k}$ is the hidden state at time $k$.&lt;/li>
&lt;li>$y_{k}$ is the observed sensory input at time $k$.&lt;/li>
&lt;li>$u_{k}$ is the control input at time $k$.&lt;/li>
&lt;li>$A$ is the dynamics matrix governing state transitions.&lt;/li>
&lt;li>$B$ is the control matrix.&lt;/li>
&lt;li>$C$ is the observation matrix.&lt;/li>
&lt;li>$f$ is a potentially nonlinear function.&lt;/li>
&lt;li>$\omega_x$ and $\omega_y$ are Gaussian process and observation noise.&lt;/li>
&lt;/ul>
&lt;p>The goal is to infer the current hidden state $x_{k}$ given the current observation $y_{k}$ and previous state estimate
$y_{1:k-1}$. To achieve this, we formulate a variational free energy objective:&lt;/p>
&lt;p>$$
\mathcal{F}_k = \frac{1}{2}(y_k - C f(x_k))^T \Sigma_y^{-1} (y_k - C f(x_k)) + \frac{1}{2}(x_k - A f(\hat{x}_{k-1}) - B u_k)^T \Sigma_x^{-1} (x_k - A f(\hat{x}_{k-1}) - B u_k)
$$&lt;/p>
&lt;p>This free energy can be understood as the sum of two weighted prediction errors:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Sensory prediction errors&lt;/strong>: The difference between observed and predicted sensory inputs $y_k - Cf(x_k)$.&lt;/li>
&lt;li>&lt;strong>Temporal prediction errors&lt;/strong>: The difference between the current state and the prediction from the previous state
$x_k - Af(\hat{x}_{k-1}) - Bu_k$.&lt;/li>
&lt;/ol>
&lt;p>Each prediction error is weighted by the precision (inverse variance) of the corresponding noise distribution,
ensuring that more reliable predictions carry more weight.&lt;/p>
&lt;h2 id="neural-implementation">Neural Implementation&lt;/h2>
&lt;p>A crucial contribution of our work is showing how temporal predictive coding can be implemented in neural circuits
using biologically plausible mechanisms. The neural dynamics for inferring the hidden state follow gradient descent
on the free energy:&lt;/p>
&lt;p>$$
\tau \frac{d x_k}{d t} = -\epsilon_x + f'(x_k) \odot C^T \epsilon_y
$$&lt;/p>
&lt;p>Where $\epsilon_x$ and $\epsilon_y$ are precision-weighted prediction errors:&lt;/p>
&lt;p>$$
\epsilon_y = \Sigma_y^{-1} \left( y_k - C f(x_k) \right)
$$&lt;/p>
&lt;p>$$
\epsilon_x = \Sigma_x^{-1} \left( x_k - A f(\hat{x}_{k-1}) - B u_k \right)
$$&lt;/p>
&lt;h3 id="we-proposed-multiple-neural-circuit-implementations-of-this-model">We proposed multiple neural circuit implementations of this model:&lt;/h3>
&lt;ol>
&lt;li>Network with explicit prediction error neurons: Where dedicated neurons represent prediction errors at each level of processing&lt;/li>
&lt;li>Dendritic computing implementation: Where prediction errors are computed as differences between somatic and dendritic potentials&lt;/li>
&lt;li>Single-iteration implementation: A simplified version that performs single updates per time step&lt;/li>
&lt;/ol>
&lt;h3 id="neural-circuit-implementation">Neural circuit implementation&lt;/h3>
&lt;p>Importantly, all of these implementations rely on local information and Hebbian plasticity.
The synaptic weights are updated according to:&lt;/p>
&lt;p>$$
\Delta A = \eta \epsilon_x f(\hat{x}_{k-1})^T
$$&lt;/p>
&lt;p>$$
\Delta B = \eta \epsilon_x u_k^T
$$&lt;/p>
&lt;p>$$
\Delta C = \eta \epsilon_y f(x_k)^T
$$&lt;/p>
&lt;p>These update rules are Hebbian in nature because they depend only on the activities of pre- and post-synaptic neurons,
making them biologically plausible.&lt;/p>
&lt;h2 id="relationship-to-kalman-filtering">Relationship to Kalman Filtering&lt;/h2>
&lt;p>An intriguing property of our model is its relationship to the Kalman filter, which is the optimal solution for
linear Gaussian filtering problems. We demonstrated that both Kalman filtering and temporal predictive coding can be
derived as special cases of Bayesian filtering, with the key difference being how they handle uncertainty.&lt;/p>
&lt;p>The Kalman filter propagates uncertainty estimates through time, tracking the full posterior covariance at each step.
In contrast, tPC approximates this by assuming a point estimate (Dirac distribution) for the previous state. Despite
this simplification, our tPC model achieves comparable performance to the Kalman filter in tracking tasks while being
computationally simpler and more biologically plausible.&lt;/p>
&lt;p>For linear systems, the tPC dynamics at equilibrium yield:&lt;/p>
&lt;p>$$
\hat{x}_k^- = A\hat{x}_{k-1} + Bu_k
$$&lt;/p>
&lt;p>$$
\hat{x}_k = \hat{x}_k^- + K(y_k - C\hat{x}_k^-)
$$&lt;/p>
&lt;p>$$
K = \Sigma_x C^T \left[C \Sigma_x C^T + \Sigma_y \right]^{-1}
$$&lt;/p>
&lt;p>This resembles the Kalman filter update equations but with a fixed gain matrix $K$ rather than a dynamically updated one
based on posterior uncertainty.&lt;/p>
&lt;h2 id="experimental-results">Experimental Results&lt;/h2>
&lt;h3 id="performance-in-linear-filtering-tasks">Performance in Linear Filtering Tasks&lt;/h3>
&lt;p>We tested our model on classic tracking problems, where the goal is to infer the hidden state
(position, velocity, acceleration) of an object undergoing unknown acceleration based on noisy observations.
Even with just a few inference steps between observations, tPC achieved performance approaching that of the
optimal Kalman filter.&lt;/p>
&lt;p>A key advantage of our model is its ability to learn the parameters of the generative model (matrices $A$, $B$, and $C$)
using Hebbian plasticity. Even when starting with random matrices, tPC could learn to accurately predict observations.
Interestingly, the model also implicitly encoded noise covariance information in its recurrent connections, without
needing explicit representation of precision matrices.&lt;/p>
&lt;h3 id="motion-sensitive-receptive-fields">Motion-Sensitive Receptive Fields&lt;/h3>
&lt;p>Perhaps most excitingly, when trained on natural movies, our tPC model developed spatiotemporal receptive fields
resembling those observed in the visual cortex. These fields exhibited Gabor-like patterns and direction selectivity,
a hallmark of motion-sensitive neurons in early visual areas.&lt;/p>
&lt;h3 id="nonlinear-extensions">Nonlinear Extensions&lt;/h3>
&lt;p>We extended the model to handle nonlinear dynamics by incorporating nonlinear activation functions. When tested on a
simulated pendulum task, the nonlinear tPC significantly outperformed the linear model, accurately predicting the
pendulum&amp;rsquo;s motion even at extreme angles where nonlinear effects are strongest.&lt;/p>
&lt;h2 id="implications-and-future-directions">Implications and Future Directions&lt;/h2>
&lt;p>Our temporal predictive coding framework has several important implications:&lt;/p>
&lt;ol>
&lt;li>It provides a biologically plausible explanation for how the brain processes dynamic stimuli and performs temporal predictions.&lt;/li>
&lt;li>It demonstrates that complex temporal filtering operations can be implemented in neural circuits using simple, local computations.&lt;/li>
&lt;li>It offers a unified framework that connects normative theories of perception (Bayesian inference) with mechanistic models of neural circuits.&lt;/li>
&lt;li>It suggests that the same computational principles might underlie both static and dynamic sensory processing in the brain.&lt;/li>
&lt;/ol>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>The temporal predictive coding model we&amp;rsquo;ve developed bridges an important gap in our understanding of how the brain
processes dynamic sensory inputs. By extending predictive coding to the temporal domain while maintaining its biological
plausibility, our model provides a compelling computational mechanism for temporal prediction in neural circuits.&lt;/p>
&lt;p>The fact that our model develops receptive fields resembling those in the visual cortex and approximates optimal
filtering solutions suggests that temporal predictive coding may indeed capture fundamental principles of neural
computation in the brain. As we continue to refine these models and test them against empirical data, we hope to
gain deeper insights into the remarkable predictive capabilities of the brain.&lt;/p>
&lt;p>&lt;em>This blog is based on the paper
&amp;ldquo;Predictive coding networks for temporal prediction&amp;rdquo; by Beren Millidge, Mufeng Tang, Mahyar Osanlouy, Nicol S. Harper,
and Rafal Bogacz, published in PLOS Computational Biology, April 2024.&lt;/em> &lt;a href="https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1011183" target="_blank" rel="noopener">Link to the paper&lt;/a>&lt;/p></description></item><item><title>Charles Babbage: The Visionary Who Invented the Computer</title><link>https://mahyar-osn.github.io/post/charles-babbage/</link><pubDate>Fri, 18 Oct 2024 00:00:00 +0000</pubDate><guid>https://mahyar-osn.github.io/post/charles-babbage/</guid><description>&lt;h2 id="introduction">Introduction&lt;/h2>
&lt;p>153 years ago today, Charles Babbage—one of history’s true polymath geniuses—died. Yet, his influence on the modern world is everywhere. Babbage’s most groundbreaking idea was the Analytical Engine, a mechanical marvel that essentially defined the blueprint for the computers we use today.&lt;/p>
&lt;h2 id="the-analytical-engine-a-machine-ahead-of-its-time">The Analytical Engine: A Machine Ahead of Its Time&lt;/h2>
&lt;p>Babbage designed the Analytical Engine in the 1830s, envisioning a machine that could perform any calculation, guided by instructions on punched cards. The design included all the essential elements of a modern computer:&lt;/p>
&lt;ul>
&lt;li>The Mill: The processing unit, analogous to today’s CPU, handled arithmetic operations.&lt;/li>
&lt;li>The Store: A memory unit that could hold 1,000 50-digit numbers. This was an unprecedented storage for its era.&lt;/li>
&lt;li>Punched Cards: Borrowed from the Jacquard loom, these allowed the machine to be programmed, making it the first programmable computer concept.&lt;/li>
&lt;li>Control and Output: The machine could make decisions (branching), repeat instructions (loops), and output results via a printer or punched cards.&lt;/li>
&lt;/ul>
&lt;img src="analytical-engine.jpeg" alt="drc-worfklow" width="500">
&lt;h2 id="legacy-and-impact">Legacy and Impact&lt;/h2>
&lt;p>Although the Analytical Engine was never completed during Babbage’s lifetime (due to technological and financial limitations) its design was Turing-complete in theory, meaning it could perform any calculation given enough memory. Ada Lovelace, Babbage’s collaborator, even wrote the first computer program for the Engine, earning her the title of the world’s first computer programmer.&lt;/p>
&lt;p>Babbage’s vision was so far ahead that it would take over a century for technology to catch up. Today, every digital device - from smartphones to supercomputers - owes a debt to his pioneering ideas.&lt;/p>
&lt;p>Conclusion
Charles Babbage’s Analytical Engine was more than an invention; it was the foundation of computer science. His work reminds us that the digital world we inhabit stands on the shoulders of visionaries who imagined the future long before it arrived.&lt;/p>
&lt;p>“Errors using inadequate data are much less than those using no data at all.”
— Charles Babbage&lt;/p></description></item><item><title>Kalman Filtering in the Age of PyTorch: State Estimation, Differentiability, and the Philosophy of Uncertainty</title><link>https://mahyar-osn.github.io/post/kalman-filter/</link><pubDate>Wed, 08 May 2024 00:00:00 +0000</pubDate><guid>https://mahyar-osn.github.io/post/kalman-filter/</guid><description>&lt;h2 id="introduction">Introduction&lt;/h2>
&lt;p>The Kalman filter, a paragon of recursive estimation, has long stood at the intersection of mathematics, engineering, and epistemology.
Conceived in the 1960s to address the challenges of navigation and control in aerospace, its recursive structure and optimality
under Gaussian assumptions have made it indispensable across robotics, signal processing, finance, and beyond.
Yet, as machine learning frameworks like PyTorch have redefined the computational landscape, the Kalman filter
finds itself in a new context—one where differentiability, GPU acceleration, and integration with deep neural architectures
are not just desirable, but essential.&lt;/p>
&lt;p>In this blog post I want to embark on a dual journey. On one hand, I want to delve into the technicalities of
implementing Kalman filters in PyTorch, leveraging its tensor operations and automatic differentiation to enable
new research and applications.
On the other, I want to reflect on the philosophical questions about the nature of uncertainty, the meaning of optimality,
and the evolving relationship between model-based and data-driven approaches. By weaving together rigorous mathematics,
practical coding insights, and reflective inquiry, we aim to illuminate both the power and the limitations of state estimation
in the age of neural computation.&lt;/p>
&lt;h2 id="the-mathematical-foundations-of-kalman-filtering">The Mathematical Foundations of Kalman Filtering&lt;/h2>
&lt;h3 id="the-state-space-model-dynamics-and-observations">The State-Space Model: Dynamics and Observations&lt;/h3>
&lt;p>At the heart of the Kalman filter lies the state-space model, a mathematical abstraction that describes the evolution of a
system&amp;rsquo;s hidden state over time and its relationship to noisy observations. Formally, the discrete-time linear state-space model is given by:&lt;/p>
&lt;p>$$
\begin{aligned}
x_{k} &amp;amp;= F_{k} x_{k-1} + B_{k} u_{k} + w_{k} \\
z_{k} &amp;amp;= H_{k} x_{k} + v_{k}
\end{aligned}
$$&lt;/p>
&lt;p>Where:&lt;/p>
&lt;ul>
&lt;li>$x_{k}$: State vector at time $k$&lt;/li>
&lt;li>$F_{k}$: State transition matrix&lt;/li>
&lt;li>$B_{k}$: Control input matrix&lt;/li>
&lt;li>$u_{k}$: Control vector&lt;/li>
&lt;li>$w_{k}$: Process noise $\sim \mathcal{N}(0,Q_{k})$&lt;/li>
&lt;li>$z_{k}$: Observation vector&lt;/li>
&lt;li>$H_{k}$: Observation matrix&lt;/li>
&lt;li>$v_{k}$: Observation noise $\sim \mathcal{N}(0,R_{k})$&lt;/li>
&lt;/ul>
&lt;p>This model encodes two key assumptions: linearity and Gaussianity. The linearity allows for closed-form recursive updates,
while the Gaussianity ensures that all conditional distributions remain Gaussian, making the mean and covariance sufficient statistics
for the state estimate.&lt;/p>
&lt;h3 id="recursive-estimation-prediction-and-update">Recursive Estimation: Prediction and Update&lt;/h3>
&lt;p>The Kalman filter operates in two alternating steps: prediction (time update) and correction (measurement update).
In the prediction step, the filter projects the current state estimate forward in time, using the system dynamics:&lt;/p>
&lt;p>$$
\begin{aligned}
\hat{x}_{k|k-1} = F_{k} \hat{x}_{k-1|k-1} + B_{k} u_{k} \\
P_{k|k-1} = F_{k} P_{k-1|k-1} F_{k}^{T} + Q_{k}
\end{aligned}
$$&lt;/p>
&lt;p>Here $\hat{x}_{k|k-1}$
is the predicted state mean,
and $P_{k|k-1}$
is the predicted state covariance.&lt;/p>
&lt;p>In the update step, the filter incorporates the new measurement $z_{k}$ to refine the state estimate:&lt;/p>
&lt;p>$$
\begin{aligned}
K_{k} &amp;amp;= P_{k|k-1} H_{k}^{T} \left( H_{k} P_{k|k-1} H_{k}^{T} + R_{k} \right)^{-1} \\
\hat{x}_{k|k} &amp;amp;= \hat{x}_{k|k-1} + K_{k} \left( z_{k} - H_{k} \hat{x}_{k|k-1} \right) \\
P_{k|k} &amp;amp;= \left( I - K_{k} H_{k} \right) P_{k|k-1}
\end{aligned}
$$&lt;/p>
&lt;p>Where $K_{k}$ is the Kalman gain, which determines how much the measurement should be trusted relative to the prediction.
Its derivation is rooted in the minimization of the mean squared error of the state estimate, balancing the uncertainty in the prediction and the measurement&lt;/p>
&lt;h3 id="the-geometry-of-uncertainty-covariance-propagation">The Geometry of Uncertainty: Covariance Propagation&lt;/h3>
&lt;p>A subtle yet profound aspect of the Kalman filter is its treatment of uncertainty. The covariance matrices $P_{k|k-1}$ and $P_{k|k}$
encode not just the spread of possible states, but also the correlations between different state variables.
The propagation of covariance through the system dynamics involves the transformation:&lt;/p>
&lt;p>$$
P_{k|k-1} = F_{k} P_{k-1|k-1} F_{k}^{T} + Q_{k}
$$&lt;/p>
&lt;p>This operation reflects how uncertainty &amp;ldquo;flows&amp;rdquo; through the linear transformation $F_{k}$, and how process noise $Q_{k}$
injects additional uncertainty. The measurement update, in turn, reduces uncertainty by incorporating
information from the observation, as modulated by the Kalman gain.&lt;/p>
&lt;p>Understanding the covariance as a bilinear form, rather than just a matrix, reveals the deep connection between
the algebra of estimation and the geometry of probability distributions. This perspective is crucial for appreciating
the filter&amp;rsquo;s optimality and for extending it to more complex, nonlinear, or high-dimensional settings.&lt;/p>
&lt;h2 id="kalman-filtering-meets-pytorch-implementation-and-differentiability">Kalman Filtering Meets PyTorch: Implementation and Differentiability&lt;/h2>
&lt;h3 id="why-pytorch-beyond-deep-learning">Why PyTorch? Beyond Deep Learning&lt;/h3>
&lt;p>PyTorch, originally designed for deep learning, offers a flexible tensor computation library with automatic
differentiation and seamless GPU acceleration. While its primary use case has been neural networks,
its capabilities make it an attractive platform for implementing classical algorithms like the Kalman filter.
The motivations are manifold:&lt;/p>
&lt;p>First, PyTorch&amp;rsquo;s tensor operations enable efficient batch processing, which is invaluable when filtering multiple signals
or running ensembles of filters in parallel. Second, the autograd engine allows for differentiable programming, making
it possible to optimize filter parameters or integrate the filter as a module within a larger neural architecture.
Third, PyTorch&amp;rsquo;s ecosystem encourages modularity, extensibility, and integration with probabilistic programming frameworks such as Pyro.&lt;/p>
&lt;h3 id="coding-the-classical-kalman-filter-in-pytorch">Coding the Classical Kalman Filter in PyTorch&lt;/h3>
&lt;p>Implementing the Kalman filter in PyTorch involves translating the recursive equations into tensor operations.
Consider the following minimal implementation for a batch of signals:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4">&lt;code class="language-python" data-lang="python">&lt;span style="color:#f92672">import&lt;/span> torch
&lt;span style="color:#f92672">from&lt;/span> torch &lt;span style="color:#f92672">import&lt;/span> nn
&lt;span style="color:#f92672">from&lt;/span> torch.linalg &lt;span style="color:#f92672">import&lt;/span> inv
&lt;span style="color:#66d9ef">class&lt;/span> &lt;span style="color:#a6e22e">KalmanFilter&lt;/span>(nn&lt;span style="color:#f92672">.&lt;/span>Module):
&lt;span style="color:#e6db74">&amp;#34;&amp;#34;&amp;#34;Kalman Filter implementation for state estimation in linear dynamic systems.
&lt;/span>&lt;span style="color:#e6db74">
&lt;/span>&lt;span style="color:#e6db74"> Attributes:
&lt;/span>&lt;span style="color:#e6db74"> F (Tensor): State transition matrix.
&lt;/span>&lt;span style="color:#e6db74"> B (Tensor): Control input matrix.
&lt;/span>&lt;span style="color:#e6db74"> H (Tensor): Observation matrix.
&lt;/span>&lt;span style="color:#e6db74"> Q (Tensor): Process noise covariance.
&lt;/span>&lt;span style="color:#e6db74"> R (Tensor): Observation noise covariance.
&lt;/span>&lt;span style="color:#e6db74"> state_dim (int): Dimensionality of the state.
&lt;/span>&lt;span style="color:#e6db74"> &amp;#34;&amp;#34;&amp;#34;&lt;/span>
&lt;span style="color:#66d9ef">def&lt;/span> __init__(self, F, B, H, Q, R, state_dim):
super()&lt;span style="color:#f92672">.&lt;/span>__init__()
self&lt;span style="color:#f92672">.&lt;/span>F &lt;span style="color:#f92672">=&lt;/span> F&lt;span style="color:#f92672">.&lt;/span>clone()
self&lt;span style="color:#f92672">.&lt;/span>B &lt;span style="color:#f92672">=&lt;/span> B&lt;span style="color:#f92672">.&lt;/span>clone()
self&lt;span style="color:#f92672">.&lt;/span>H &lt;span style="color:#f92672">=&lt;/span> H&lt;span style="color:#f92672">.&lt;/span>clone()
self&lt;span style="color:#f92672">.&lt;/span>Q &lt;span style="color:#f92672">=&lt;/span> Q
self&lt;span style="color:#f92672">.&lt;/span>R &lt;span style="color:#f92672">=&lt;/span> R
self&lt;span style="color:#f92672">.&lt;/span>state_dim &lt;span style="color:#f92672">=&lt;/span> state_dim
&lt;span style="color:#75715e"># placeholders for the current state, covariance, observation and control&lt;/span>
self&lt;span style="color:#f92672">.&lt;/span>x &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">None&lt;/span> &lt;span style="color:#75715e"># [state_dim, 1]&lt;/span>
self&lt;span style="color:#f92672">.&lt;/span>P &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">None&lt;/span> &lt;span style="color:#75715e"># [state_dim, state_dim]&lt;/span>
self&lt;span style="color:#f92672">.&lt;/span>zs &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">None&lt;/span> &lt;span style="color:#75715e"># [obs_dim, 1]&lt;/span>
self&lt;span style="color:#f92672">.&lt;/span>us &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">None&lt;/span> &lt;span style="color:#75715e"># [control_dim, 1]&lt;/span>
&lt;span style="color:#66d9ef">def&lt;/span> &lt;span style="color:#a6e22e">project&lt;/span>(self):
&lt;span style="color:#e6db74">&amp;#34;&amp;#34;&amp;#34;Projects the state and covariance forward.&amp;#34;&amp;#34;&amp;#34;&lt;/span>
x_pred &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>matmul(self&lt;span style="color:#f92672">.&lt;/span>F, self&lt;span style="color:#f92672">.&lt;/span>x) &lt;span style="color:#f92672">+&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>matmul(self&lt;span style="color:#f92672">.&lt;/span>B, self&lt;span style="color:#f92672">.&lt;/span>us)
P_pred &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>matmul(self&lt;span style="color:#f92672">.&lt;/span>F, torch&lt;span style="color:#f92672">.&lt;/span>matmul(self&lt;span style="color:#f92672">.&lt;/span>P, self&lt;span style="color:#f92672">.&lt;/span>F&lt;span style="color:#f92672">.&lt;/span>T)) &lt;span style="color:#f92672">+&lt;/span> self&lt;span style="color:#f92672">.&lt;/span>Q
&lt;span style="color:#66d9ef">return&lt;/span> x_pred, P_pred
&lt;span style="color:#66d9ef">def&lt;/span> &lt;span style="color:#a6e22e">correct&lt;/span>(self, x_pred, P_pred):
&lt;span style="color:#e6db74">&amp;#34;&amp;#34;&amp;#34;Corrects the state estimate with the current observation.&amp;#34;&amp;#34;&amp;#34;&lt;/span>
S &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>matmul(self&lt;span style="color:#f92672">.&lt;/span>H, torch&lt;span style="color:#f92672">.&lt;/span>matmul(P_pred, self&lt;span style="color:#f92672">.&lt;/span>H&lt;span style="color:#f92672">.&lt;/span>T)) &lt;span style="color:#f92672">+&lt;/span> self&lt;span style="color:#f92672">.&lt;/span>R
K &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>matmul(P_pred, self&lt;span style="color:#f92672">.&lt;/span>H&lt;span style="color:#f92672">.&lt;/span>T) &lt;span style="color:#f92672">@&lt;/span> inv(S)
&lt;span style="color:#75715e"># state update&lt;/span>
self&lt;span style="color:#f92672">.&lt;/span>x &lt;span style="color:#f92672">=&lt;/span> x_pred &lt;span style="color:#f92672">+&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>matmul(K, (self&lt;span style="color:#f92672">.&lt;/span>zs &lt;span style="color:#f92672">-&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>matmul(self&lt;span style="color:#f92672">.&lt;/span>H, x_pred)))
&lt;span style="color:#75715e"># covariance update&lt;/span>
I &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>eye(self&lt;span style="color:#f92672">.&lt;/span>state_dim, device&lt;span style="color:#f92672">=&lt;/span>P_pred&lt;span style="color:#f92672">.&lt;/span>device)
self&lt;span style="color:#f92672">.&lt;/span>P &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>matmul((I &lt;span style="color:#f92672">-&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>matmul(K, self&lt;span style="color:#f92672">.&lt;/span>H)), P_pred)
&lt;span style="color:#66d9ef">def&lt;/span> &lt;span style="color:#a6e22e">forward&lt;/span>(self, zs, us):
&lt;span style="color:#e6db74">&amp;#34;&amp;#34;&amp;#34;
&lt;/span>&lt;span style="color:#e6db74"> Processes a batch of observation/control sequences.
&lt;/span>&lt;span style="color:#e6db74">
&lt;/span>&lt;span style="color:#e6db74"> Args:
&lt;/span>&lt;span style="color:#e6db74"> zs: [timesteps, batch, obs_dim] sequence of observations
&lt;/span>&lt;span style="color:#e6db74"> us: [timesteps, batch, control_dim] sequence of control inputs
&lt;/span>&lt;span style="color:#e6db74"> Returns:
&lt;/span>&lt;span style="color:#e6db74"> xs: [batch, state_dim, timesteps] filtered state estimates
&lt;/span>&lt;span style="color:#e6db74"> pred_obs: [batch, obs_dim, timesteps] one-step predictions of observations
&lt;/span>&lt;span style="color:#e6db74"> residuals: [batch, obs_dim, timesteps] observation residuals
&lt;/span>&lt;span style="color:#e6db74"> &amp;#34;&amp;#34;&amp;#34;&lt;/span>
xs &lt;span style="color:#f92672">=&lt;/span> []
pred_obs &lt;span style="color:#f92672">=&lt;/span> []
residuals &lt;span style="color:#f92672">=&lt;/span> []
&lt;span style="color:#75715e"># initial state &amp;amp; covariance&lt;/span>
self&lt;span style="color:#f92672">.&lt;/span>x &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>zeros((self&lt;span style="color:#f92672">.&lt;/span>state_dim, &lt;span style="color:#ae81ff">1&lt;/span>), device&lt;span style="color:#f92672">=&lt;/span>zs&lt;span style="color:#f92672">.&lt;/span>device)
self&lt;span style="color:#f92672">.&lt;/span>P &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>eye(self&lt;span style="color:#f92672">.&lt;/span>state_dim, device&lt;span style="color:#f92672">=&lt;/span>zs&lt;span style="color:#f92672">.&lt;/span>device)
&lt;span style="color:#75715e"># iterate over time&lt;/span>
&lt;span style="color:#66d9ef">for&lt;/span> z_t, u_t &lt;span style="color:#f92672">in&lt;/span> zip(zs&lt;span style="color:#f92672">.&lt;/span>transpose(&lt;span style="color:#ae81ff">0&lt;/span>, &lt;span style="color:#ae81ff">1&lt;/span>), us&lt;span style="color:#f92672">.&lt;/span>transpose(&lt;span style="color:#ae81ff">0&lt;/span>, &lt;span style="color:#ae81ff">1&lt;/span>)):
self&lt;span style="color:#f92672">.&lt;/span>zs &lt;span style="color:#f92672">=&lt;/span> z_t&lt;span style="color:#f92672">.&lt;/span>unsqueeze(&lt;span style="color:#ae81ff">1&lt;/span>)
self&lt;span style="color:#f92672">.&lt;/span>us &lt;span style="color:#f92672">=&lt;/span> u_t&lt;span style="color:#f92672">.&lt;/span>unsqueeze(&lt;span style="color:#ae81ff">1&lt;/span>)
x_pred, P_pred &lt;span style="color:#f92672">=&lt;/span> self&lt;span style="color:#f92672">.&lt;/span>project()
self&lt;span style="color:#f92672">.&lt;/span>correct(x_pred, P_pred)
xs&lt;span style="color:#f92672">.&lt;/span>append(self&lt;span style="color:#f92672">.&lt;/span>x&lt;span style="color:#f92672">.&lt;/span>detach()&lt;span style="color:#f92672">.&lt;/span>clone())
y_pred &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>matmul(self&lt;span style="color:#f92672">.&lt;/span>H, x_pred)
pred_obs&lt;span style="color:#f92672">.&lt;/span>append(y_pred)
residuals&lt;span style="color:#f92672">.&lt;/span>append(self&lt;span style="color:#f92672">.&lt;/span>zs &lt;span style="color:#f92672">-&lt;/span> y_pred)
xs &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>cat(xs, dim&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">1&lt;/span>)
pred_obs &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>cat(pred_obs, dim&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">1&lt;/span>)
residuals &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>cat(residuals, dim&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">1&lt;/span>)
&lt;span style="color:#66d9ef">return&lt;/span> xs, pred_obs, residuals
&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="differentiable-kalman-filters-learning-and-optimization">Differentiable Kalman Filters: Learning and Optimization&lt;/h2>
&lt;p>One of the most transformative aspects of implementing the Kalman filter in PyTorch is the ability to make the entire
filtering process differentiable. By treating the system matrices ($F$, $H$, $Q$, $R$) as learnable parameters,
one can optimize them using gradient-based methods, either to fit data or to tune the filter for specific tasks.
This approach blurs the line between classical estimation and machine learning, enabling hybrid models that combine
the structure of state-space models with the flexibility of data-driven learning.&lt;/p>
&lt;p>Recent research has focused on improving the efficiency of backpropagation through the Kalman filter.
While PyTorch&amp;rsquo;s automatic differentiation can compute gradients, it may incur significant computational overhead,
especially for large-scale problems. Novel closed-form expressions for the derivatives of the filter&amp;rsquo;s outputs with
respect to its parameters have been developed, offering substantial speed-ups (up to 38 times faster than PyTorch&amp;rsquo;s
autograd in some cases). These advances make it feasible to embed Kalman filters within deep learning pipelines,
trainable end-to-end, and responsive to the demands of modern applications.&lt;/p>
&lt;h2 id="pytorch-libraries-for-kalman-filtering">PyTorch Libraries for Kalman Filtering&lt;/h2>
&lt;p>Several open-source libraries have emerged to facilitate Kalman filtering in PyTorch:&lt;/p>
&lt;ul>
&lt;li>torch-kf: A fast implementation supporting batch filtering and smoothing, capable of running on both CPU and GPU. It is particularly efficient when filtering large batches of signals, leveraging PyTorch&amp;rsquo;s parallelism.&lt;/li>
&lt;li>DeepKalmanFilter: Implements deep variants of the Kalman filter, where neural networks parameterize parts of the state-space model. This enables modeling of nonlinear dynamics and observations, bridging the gap between classical filtering and deep generative models.&lt;/li>
&lt;li>Pyro: A probabilistic programming framework that supports differentiable Kalman filters and extended Kalman filters, with learnable parameters and integration with variational inference.&lt;/li>
&lt;li>torchfilter: Provides advanced filters such as the square-root unscented Kalman filter, supporting both state and parameter estimation in nonlinear systems.&lt;/li>
&lt;/ul>
&lt;h2 id="extensions-and-hybrid-models-beyond-the-classical-filter">Extensions and Hybrid Models: Beyond the Classical Filter&lt;/h2>
&lt;h3 id="nonlinear-and-non-gaussian-filtering">Nonlinear and Non-Gaussian Filtering&lt;/h3>
&lt;p>While the classical Kalman filter assumes linear dynamics and Gaussian noise, many real-world systems violate
these assumptions. Extensions such as the Extended Kalman Filter (EKF) and Unscented Kalman Filter (UKF) address
nonlinearities by linearizing the dynamics or propagating sigma points, respectively. Particle filters, in turn,
approximate arbitrary distributions via Monte Carlo sampling.&lt;/p>
&lt;p>Implementing these advanced filters in PyTorch follows the same principles: tensorized operations,
differentiability, and integration with neural modules. For example, the EKF can be implemented by computing
Jacobians using PyTorch&amp;rsquo;s autograd, while the UKF can leverage batched sigma point propagation for efficient parallelism.&lt;/p>
&lt;h3 id="deep-kalman-filters-and-latent-dynamics">Deep Kalman Filters and Latent Dynamics&lt;/h3>
&lt;p>The fusion of Kalman filtering with deep learning has given rise to deep Kalman filters, where neural networks
parameterize the transition and observation functions. This approach enables modeling of complex, nonlinear,
and high-dimensional systems, such as video sequences or sensor fusion in robotics. The deep Kalman filter retains
the probabilistic structure of the classical filter but augments it with the representational power of neural networks.&lt;/p>
&lt;p>In PyTorch, this is achieved by defining neural modules for the transition and observation models,
and using the filtering equations to propagate means and covariances through time. The entire model
can be trained end-to-end using stochastic gradient descent, with the Kalman filter acting as a differentiable
layer within the network.&lt;/p>
&lt;h3 id="hybrid-estimators-neural-networks-and-kalman-filters">Hybrid Estimators: Neural Networks and Kalman Filters&lt;/h3>
&lt;p>Hybrid models that combine neural networks and Kalman filters have demonstrated superior performance in
state estimation tasks, particularly in scenarios with complex dynamics or partial observability.
These models can be categorized into two main types:&lt;/p>
&lt;ul>
&lt;li>NN-KF: Neural networks learn the parameters or functions of the state-space model, which are then used by the Kalman filter for estimation.&lt;/li>
&lt;li>KF-NN: The Kalman filter provides state estimates or uncertainty measures that are used as inputs or features for a neural network.&lt;/li>
&lt;/ul>
&lt;p>Such hybridization leverages the strengths of both approaches: the interpretability and optimality of the Kalman filter,
and the flexibility and expressiveness of neural networks. In PyTorch, these models can be implemented as composite
modules, trained jointly or sequentially, and deployed in a wide range of applications from battery state-of-charge
estimation to autonomous navigation.&lt;/p>
&lt;h2 id="philosophical-reflections-uncertainty-knowledge-and-learning">Philosophical Reflections: Uncertainty, Knowledge, and Learning&lt;/h2>
&lt;h3 id="the-epistemology-of-state-estimation">The Epistemology of State Estimation&lt;/h3>
&lt;p>At a deeper level, the Kalman filter embodies a philosophy of knowledge under uncertainty. It formalizes the process of
updating beliefs in the face of incomplete and noisy information, balancing prior expectations (the model) with new
evidence (the measurements). The recursive structure mirrors the Bayesian paradigm, where beliefs are continuously
revised as new data arrives.&lt;/p>
&lt;p>Yet, the filter&amp;rsquo;s optimality is contingent on its assumptions: linearity, Gaussianity, and known noise covariances.
When these assumptions are violated, as is often the case in complex systems, the filter&amp;rsquo;s estimates may become biased
or inconsistent. This raises fundamental questions: What does it mean to &amp;ldquo;know&amp;rdquo; the state of a system? How do we quantify
and manage uncertainty? Can we trust our models, or must we adapt them in light of new evidence?&lt;/p>
&lt;h3 id="the-fusion-of-model-based-and-data-driven-approaches">The Fusion of Model-Based and Data-Driven Approaches&lt;/h3>
&lt;p>The integration of Kalman filtering with PyTorch and neural networks reflects a broader trend in computational science:
the synthesis of model-based and data-driven approaches. Classical estimation theory offers structure, interpretability,
and guarantees of optimality. Machine learning provides flexibility, scalability, and the ability to discover patterns
from data.&lt;/p>
&lt;p>Hybrid models, differentiable filters, and end-to-end learning challenge the traditional dichotomy between &amp;ldquo;hard-coded&amp;rdquo;
models and &amp;ldquo;black-box&amp;rdquo; learning. They invite us to reconsider the boundaries between theory and data, deduction and
induction, certainty and doubt. In this sense, the Kalman filter is not just an algorithm, but a lens through which to
explore the nature of inference, prediction, and adaptation.&lt;/p>
&lt;h3 id="the-philosophy-of-differentiable-programming">The Philosophy of Differentiable Programming&lt;/h3>
&lt;p>The advent of differentiable programming—where algorithms are designed to be composed, differentiated,
and optimized—raises new philosophical questions. When we make the Kalman filter differentiable, we enable it to
learn from data, to adapt its parameters, and to participate in the broader ecosystem of neural computation.
But we also introduce new forms of uncertainty: about the correctness of gradients, the stability of optimization,
and the interpretability of learned models.&lt;/p>
&lt;p>Is the differentiable Kalman filter still a Kalman filter, or has it become something new? What are the implications of
treating classical algorithms as modules within a deep learning pipeline? How do we balance the desire for optimality
with the need for flexibility? These questions invite ongoing reflection and experimentation.&lt;/p>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>The Kalman filter, once a symbol of control theory and aerospace engineering, has found new life in the era of PyTorch
and machine learning. Its recursive structure, principled handling of uncertainty, and optimality under Gaussian
assumptions remain as compelling as ever. Yet, its implementation and interpretation are evolving, shaped by the
demands of differentiability, scalability, and integration with neural computation.&lt;/p>
&lt;p>By exploring the mathematical foundations, practical coding strategies, extensions to nonlinear and hybrid models,
and the deeper philosophical questions that arise, we have sought to illuminate both the enduring relevance and
the transformative potential of Kalman filtering in the age of PyTorch. As we continue to blur the boundaries between
model-based and data-driven approaches, the filter serves as a bridge—not just between past and future, but between
certainty and doubt, theory and practice, knowledge and learning.&lt;/p>
&lt;p>The journey of the Kalman filter is far from over. Its recursive dance of prediction and correction, its geometry
of uncertainty, and its adaptability to new computational paradigms ensure that it will remain a central figure in the
ongoing dialogue between mathematics, engineering, and philosophy. Whether as a standalone estimator, a differentiable
module, or a component of a deep generative model, the Kalman filter challenges us to rethink what it means to know,
to predict, and to learn.&lt;/p>
&lt;h2 id="further-reading-and-resources">Further Reading and Resources&lt;/h2>
&lt;p>For those interested in diving deeper, consider exploring the following resources:&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://github.com/raphaelreme/torch-kf" target="_blank" rel="noopener">torch-kf&lt;/a>: Fast PyTorch implementation of Kalman filters, supporting batch processing and GPU acceleration.&lt;/li>
&lt;li>&lt;a href="https://github.com/morim3/DeepKalmanFilter" target="_blank" rel="noopener">DeepKalmanFilter&lt;/a>: PyTorch implementation of deep Kalman filters, integrating neural networks with probabilistic state-space models.&lt;/li>
&lt;li>[Pyro Tutorials](&lt;a href="https://pyro.ai/examples/ekf.html" target="_blank" rel="noopener">https://pyro.ai/examples/ekf.html&lt;/a>: Differentiable Kalman and extended Kalman filters with learnable parameters.&lt;/li>
&lt;li>&lt;a href="https://stanford-iprl-lab.github.io/torchfilter/_modules/torchfilter/filters/_square_root_unscented_kalman_filter/" target="_blank" rel="noopener">torchfilter&lt;/a>: Advanced filters including square-root unscented Kalman filter for nonlinear systems.&lt;/li>
&lt;li>Recent Research: &lt;a href="https://stanford-iprl-lab.github.io/torchfilter/_modules/torchfilter/filters/_square_root_unscented_kalman_filter/" target="_blank" rel="noopener">Closed-form gradients for efficient differentiable filtering&lt;/a>,
&lt;a href="https://www.semanticscholar.org/paper/A-review%3A-state-estimation-based-on-hybrid-models-Feng-Li/1f9d96407167c1bb894c4dec60a64bd31c00d1e8" target="_blank" rel="noopener">hybrid models for state estimation&lt;/a>,
and &lt;a href="https://arxiv.org/abs/2010.08196" target="_blank" rel="noopener">practical applications in robotics and sensor fusion&lt;/a>.&lt;/li>
&lt;/ul></description></item><item><title>Why Normalizing Flows (and Tensorizing Flows) deserve more attention</title><link>https://mahyar-osn.github.io/post/tensorizing-flows/</link><pubDate>Fri, 04 Aug 2023 00:00:00 +0000</pubDate><guid>https://mahyar-osn.github.io/post/tensorizing-flows/</guid><description>&lt;p>Other generative models like diffusion models and autoregressive LLMs tend to steal the spotlight, since they&amp;rsquo;re great
at producing stunning images or generating text. Normalizing Flows, on the other hand, aren&amp;rsquo;t the first choice for
those headline-grabbing tasks. But if you focus only on sample quality, you might overlook what makes Normalizing Flows
truly valuable.&lt;/p>
&lt;h2 id="why-normalizing-flows-deserve-more-attention">Why Normalizing Flows Deserve More Attention&lt;/h2>
&lt;p>Most generative models are black boxes. GANs, for example, can create high-quality samples, but you can&amp;rsquo;t compute the
likelihood of a given data point. Energy-based models often only give you unnormalized densities, so you can compare
samples but not get an actual probability.&lt;/p>
&lt;p>Normalizing Flows are different. They let you map a simple base distribution (like a Gaussian) through a sequence of
invertible transformations to model complex data. The kicker? You always have access to the exact, normalized probability
density for any sample. This is a huge deal for applications where you need to know the likelihood, not just generate
data.&lt;/p>
&lt;h2 id="the-real-world-use-case-variational-inference">The Real-World Use Case: Variational Inference&lt;/h2>
&lt;p>One area where this property is crucial is Variational Inference (VI). Here, you want to approximate a complex target
distribution with a flexible, normalized family so you can do things like Bayesian inference efficiently.
NFs are a natural fit because you can both sample from them and compute exact densities—something most other models
can&amp;rsquo;t offer.&lt;/p>
&lt;h2 id="but-theres-a-catch">But There&amp;rsquo;s a Catch&amp;hellip;&lt;/h2>
&lt;p>Traditional NFs use a Gaussian as their base distribution. This works fine for unimodal targets, but if your true
distribution is multimodal (think: multiple peaks), NFs tend to &amp;ldquo;collapse&amp;rdquo; to just one mode. This limits their
expressiveness in VI, especially for challenging scientific or physics problems where multimodality is the norm.&lt;/p>
&lt;h2 id="enter-tensorizing-flows">Enter Tensorizing Flows&lt;/h2>
&lt;p>The paper &amp;ldquo;Tensorizing Flows: A Tool for Variational Inference&amp;rdquo; introduces a clever fix: replace the Gaussian base
with a tensor-train (TT) distribution, built using tools from tensor networks. This TT base can already capture much
of the structure (including multimodality) of the target distribution, so the flow only needs to handle the
&amp;ldquo;fine details.&amp;rdquo; The result is a model that&amp;rsquo;s both more expressive and easier to train for high-dimensional,
multimodal problems.&lt;/p>
&lt;h2 id="resources">Resources&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://arxiv.org/pdf/2305.02460" target="_blank" rel="noopener">Article&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://github.com/VincentStimper/normalizing-flows" target="_blank" rel="noopener">NormFlow&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Generalization of predictive coding model to dynamic stimuli</title><link>https://mahyar-osn.github.io/post/tpc/</link><pubDate>Mon, 26 Apr 2021 00:00:00 +0000</pubDate><guid>https://mahyar-osn.github.io/post/tpc/</guid><description>&lt;h2 id="introduction">Introduction&lt;/h2>
&lt;p>Predictive coding is an established model of perceptual inference and learning in hierarchical networks of the brain.
It describes a network of neuron-like nodes, which can infer stimulus properties from noisy input using only
&lt;em>local computation&lt;/em>, i.e. the changes of activity of each neuron in the model is determined only by its inputs and its
current activity levels. Furthermore, the network encodes the estimated parameters of a probabilistic model from which
the stimuli are generated in its synaptic connections, and learn these parameters employing only
&lt;em>local plasticity&lt;/em>, where the changes in synaptic weights only depend of activities of pre and post-synaptic neurons.
In its original form the predictive coding model assumes static input stimuli. However, most of
the stimuli experienced by animals and humans change in time, and it is critical for survival to efficiently interpret
such stimuli.&lt;/p>
&lt;p>Very soon after developing the predictive coding model, it was pointed that it could be generalized to dynamic stimuli,
and the Kalman filter could be employed to infer the states of hidden variables represented by the model. However that work has not described how such computation could be implemented in a biologically
plausible network of neuron-like nodes. More recently, a generalization of predictive coding to dynamic stimuli has been
proposed, in which different neurons represent not only the hidden variables, but also their temporal derivatives.
Although it is possible to implement this model in a network only employing local computation and local plasticity,
this network requires a very intricate and specific pattern of connectivity between various neurons, and there is no evidence that such connectivity exists in cortical circuits.&lt;/p>
&lt;p>This report outlines a simple generalization of predictive coding model to dynamic stimuli, which does not require more
intricate network than the original predictive coding model. A simulation of the proposed generalizations is shown for
a toy problem, and directions are suggested in which the work on the model needs to be conducted.&lt;/p>
&lt;h2 id="model">Model&lt;/h2>
&lt;h3 id="process-generating-stimuli">Process generating stimuli&lt;/h3>
&lt;p>In this report we assume that stimuli are generated from a very simple linear model, which parallels the assumptions
about signal made by the Kalman filter. Let us denote an observed stimulus at time
$t$ by a vector with elements $y_i(t)$. Let us assume that the stimulus depends on values of hidden variables
denoted by $x_j(t)$ according to:&lt;/p>
&lt;p>$$
y_i(t) = \sum_j w_{i,j} x_j(t) + \epsilon_{y,i}(t) \quad (1)
$$&lt;/p>
&lt;p>In the above equation, $w_{i,j}$ form a matrix of parameters, and $\epsilon_{y,i}(t)$ is a noise process (with zero mean).
Furthermore, let us assume that the hidden variables evolve according to:&lt;/p>
&lt;p>$$
\dot{x}_j = \sum_k v_{j,k} x_k(t) + \epsilon_{x,j}(t) \quad (2)
$$&lt;/p>
&lt;p>Analogously as above, $v_{j,k}$ form a matrix of parameters, and $\epsilon_{x,j}(t)$ is a noise process.
A natural way for estimating $x_j$ from $y_i$ is to employ the Kalman filter, but it involves complex equations,
and it is not clear how such computation could be implemented in a network of neurons. Therefore, this report describes
a simpler method for estimating $x_j$ that has a more natural neural implementation.&lt;/p>
&lt;h3 id="computations-in-the-model">Computations in the model&lt;/h3>
&lt;p>Given a observed stimuli $y_i$, we will seek to infer the hidden variables $x_j$ and estimate the parameters $w_{i,j}$
and $v_{j,k}$. In the reminder of Section 2, we will use $x_j$, $w_{i,j}$ and $v_{j,k}$ to denote the estimates of
corresponding terms in Equations (1) and (2) above. We wish to find $x_j$ such that the stimulus $y_i$ is
close to the predicted value $\sum_j w_{i,j} x_j$. Thus we define error in prediction
of $y_i$ as:&lt;/p>
&lt;p>$$
e_i = y_i - \sum_j w_{i,j} x_j \quad (3)
$$&lt;/p>
&lt;p>We wish to minimize a squared sum of these errors which we denote by $E_y = \frac{1}{2} \sum_i \varepsilon_{y,i}^2$.
Hence we change $x_j$ in the direction opposite to the gradient of $E_y$, but we additionally append this dynamics
towards our goal with the natural evolution of $x_j$:&lt;/p>
&lt;p>$$
\dot{x}_j = - \frac{\partial E_y}{\partial x_j} + \sum_k v_{j,k} x_k
$$&lt;/p>
&lt;p>Evaluating the gradient, we obtain the equation describing the dynamics of our estimate of hidden variables:&lt;/p>
&lt;p>$$
\dot{x}_j = \sum_i w_{i,j} \varepsilon_{y,i} + \sum_k v_{j,k} x_k
$$&lt;/p>
&lt;p>In order to learn parameters $w_{i,j}$, which describe how $y_i$ depends on $x_j$, we modify them to minimize $E_y$:&lt;/p>
&lt;p>$$
\dot{w}_{i,j} = - \alpha \frac{\partial E_y}{\partial w_{i,j}} = \alpha \varepsilon_{y,i} x_j
$$&lt;/p>
&lt;p>In the above equation $\alpha$ denotes a learning rate. In order to learn parameters $v_{j,k}$ describing the natural
dynamics of hidden variables, we need to define an error in prediction of this dynamics:&lt;/p>
&lt;p>$$
\varepsilon_{x,j} = \dot{x}_j - \sum_k v_{j,k} x_k
$$&lt;/p>
&lt;p>We wish to minimize squared sum of these errors $E_x = \frac{1}{2} \sum_j \varepsilon_{x,j}^2$,
and hence we modify the weights in the direction opposite to the gradient of $E_x$ over $v_{j,k}$:&lt;/p>
&lt;p>$$
\dot{v}_{j,k} = \alpha \varepsilon_{x,j} x_k
$$&lt;/p>
&lt;p>In summary, this generalized predictive coding model continuously updates hidden variables and parameters according and
recomputes prediction errors.&lt;/p>
&lt;h3 id="possible-neural-implementations">Possible neural implementations&lt;/h3>
&lt;p>Inference of hidden variables $x_j$ from sensory input $y_i$ can be easily performed in a network shown in Figure 1A.
The bottom layer consists of sensory neurons representing the stimulus. They project to neurons computing prediction
error. These errors are then send to the neurons encoding hidden variables which
change their activity according to the dynamics equation above. The weights of connections between neurons encoding errors
and hidden variables are symmetric, i.e. equal in both direction. This network has an architecture very similar to
a standard predictive coding model , but additionally includes recurrent connections between the
neurons encoding hidden variables with weights $v_{j,k}$.&lt;/p>
&lt;img src="featured.png" alt="Receptive fields" width="800">
&lt;p>Learning parameters $w_{i,j}$ corresponds to local Hebbian plasticity in the network
of Figure 1, analogously as in the standard predictive coding networks. However,
learning parameters $v_{j,k}$ is less straightforward because the
prediction error $\varepsilon_{x,j}$ is not explicitly represented in activity of any neurons in the network.
Nevertheless, it is possible to construct models in which $\varepsilon_{x,j}$ would be represented in internal
signals (e.g. concentrations of particular ions or proteins) within neurons encoding $x_j$,
and let us consider two such possible models.&lt;/p>
&lt;p>The first model is illustrated in Figure 1B.
In this network, the recurrent inputs from neurons representing hidden variables converge on a separate dendritic
branch, which sums them and thus can compute $\sum_k v_{j,k} x_k$. To compute the error $\varepsilon_{x,j}$,
the neuron would need to compute the difference between change in its activity and the membrane potential in the dendrite.
Since both of these quantities are encoded within the same neuron, it is plausible that such a computation may be performed,
and an error encoded in an internal signal. Such signal could then drive local synaptic plasticity.&lt;/p>
&lt;p>An alternative way of computing prediction errors $\varepsilon_{x,j}$ relies on an observation that by combining
equations describing the dynamics of $\dot{x}_j$ adn error $\varepsilon_{x,j}$, we see that these errors are equal to:&lt;/p>
&lt;p>$$
\varepsilon_{x,j} = \sum_i w_{i,j} \varepsilon_{y,i}
$$&lt;/p>
&lt;p>Such input from the previous layer of prediction error neurons could be computed in dendrites shown in
Figure 1C. The membrane potential of such dendrite would need to set level of an internal signal that would govern the
plasticity within the entire neuron. This mechanism could be considered biologically plausible as it is analogous to
observations that high membrane potential of apical dendrites of pyramidal neurons triggers plateau potentials via
calcium influx, leading to a burst of spikes by the neuron. Such bursts of spikes may subsequently
induce synaptic plasticity.&lt;/p>
&lt;h2 id="results">Results&lt;/h2>
&lt;p>I tested the model on a simple problem in which hidden variables and stimuli were 2-dimensional.
The hidden variables were generated according to $\dot{x}_j = \sum_k v_{j,k} x_k(t) + \epsilon_{x,j}(t)$ with parameters
$v_{j,k}$ set to a rotation matrix visualized in Figure 2C. The stimuli were generated according
to $y_i(t) = \sum_j w_{i,j} x_j(t) + \epsilon_{y,i}(t)$ with parameters $w_{i,j}$ set to the identity matrix,
so that the stimuli were simply noisy versions of the hidden variables. The stimuli are shown in Figure 2A, and they are
noisy periodic signal because parameters $v_{j,k}$ were set to a rotation matrix. The variables and stimuli were generated
with a sampling frequency 10, by solving our equations using Euler method with integration step $0.1$. During each step,
noise with variance of $0.01$ was added.&lt;/p>
&lt;img src="results.png" alt="Receptive fields" width="700">
&lt;p>At the start of the learning process, weights $w_{i,j}$ were initialized to an identity matrix,
while the weights between hidden units were all set to $v_{j,k}=0$. The hidden units were also initialized to $x_j=0$.
The hidden variables and parameters were updated according to our equations above using the Euler method with integration
step of $0.1$, and learning rate set to $\alpha=0.01$.&lt;/p>
&lt;p>Figure 2B shows that as the learning progressed, the error in prediction of stimuli decreased, so the network was able
to better predict the stimuli. Figure 2D visualizes learned values of parameters $v_{j,k}$, which are very close to the
original parameters used to generate the training data (cf. Figure 2C).
Thus the network was able to discover the underlying process generating the stimuli.&lt;/p>
&lt;h2 id="discussion">Discussion&lt;/h2>
&lt;p>This report outlines generalization of predictive coding to dynamic stimuli for linear and shallow generative models,
so more work would be required to extend this to more complex models and relate it with experimental data.
In particular the work can be extended in the following directions:&lt;/p>
&lt;ul>
&lt;li>Introduce the non-linear activation functions to hidden units, and test if the model can learn dynamics of non-linear systems.&lt;/li>
&lt;li>Introduce multiple levels of hierarchy and investigate if the model can extract dynamics of stimuli generated by hierarchical dynamical systems.&lt;/li>
&lt;li>Test the model performance on real world machine learning problems, e.g. prediction of EEG signal from past history.&lt;/li>
&lt;li>Investigate if after training with natural stimuli the receptive fields of neurons in the model have similar properties
to the receptive fields in the visual system, analogously as in neural networks trained with the back-propagation algorithm.&lt;/li>
&lt;/ul></description></item><item><title>Allostasis, Interoception, and the Free Energy Principle</title><link>https://mahyar-osn.github.io/post/interoception-allostasis/</link><pubDate>Fri, 12 Mar 2021 00:00:00 +0000</pubDate><guid>https://mahyar-osn.github.io/post/interoception-allostasis/</guid><description>&lt;h2 id="introduction">Introduction&lt;/h2>
&lt;p>The intersection of biological regulation, predictive processing, and consciousness represents one of the most
fascinating frontiers in cognitive science today. After carefully reading Corcoran and Hohwy&amp;rsquo;s chapter &amp;ldquo;Allostasis,
interoception, and the free energy principle: Feeling our way forward,&amp;rdquo; I&amp;rsquo;m struck by both its ambitious scope and
its meticulous attention to conceptual clarity. This paper attempts to untangle a complex theoretical landscape that
has profound implications for how we understand the relationship between mind, body, and environment.&lt;/p>
&lt;h2 id="the-conceptual-maze-homeostasis-and-allostasis">The Conceptual Maze: Homeostasis and Allostasis&lt;/h2>
&lt;p>At its foundation, this paper addresses a fundamental question: how do biological organisms maintain their viability?
The traditional answer has been homeostasis - the concept developed by Claude Bernard and Walter Cannon emphasizing the
maintenance of stable internal conditions despite external fluctuations. The authors provide an excellent historical
overview of this concept, tracing its development from Bernard&amp;rsquo;s emphasis on the &amp;ldquo;milieu intérieur&amp;rdquo; to Cannon&amp;rsquo;s more
nuanced view of stability involving acceptable ranges rather than fixed setpoints.&lt;/p>
&lt;p>What makes this paper particularly valuable is its careful examination of allostasis - a concept introduced by Sterling
and Eyer in 1988 as &amp;ldquo;stability through change&amp;rdquo;. The authors meticulously document how this concept has evolved in multiple,
sometimes contradictory directions:&lt;/p>
&lt;ol>
&lt;li>Sterling and Eyer&amp;rsquo;s radical position that allostasis should entirely replace homeostasis&lt;/li>
&lt;li>McEwen&amp;rsquo;s view of allostasis as &amp;ldquo;the process for actively maintaining homeostasis&amp;rdquo;&lt;/li>
&lt;li>Schulkin&amp;rsquo;s perspective where homeostasis and allostasis are complementary mechanisms for maintaining biological viability&lt;/li>
&lt;/ol>
&lt;p>This historical excavation reveals something important: allostasis has been a contested concept from the beginning,
with no clear consensus about its precise meaning even 30+ years after its introduction.&lt;/p>
&lt;h2 id="free-energy-and-interoceptive-inference">Free Energy and Interoceptive Inference&lt;/h2>
&lt;p>The paper becomes even more interesting when it examines how these biological regulation concepts have been incorporated
into the free energy principle framework. The authors identify three distinct interpretations of allostasis within
recent free energy-inspired accounts:&lt;/p>
&lt;ol>
&lt;li>Behavioral allostasis: Focuses on behavioral actions on the external world to maintain internal states (Gu &amp;amp; FitzGerald, Seth)&lt;/li>
&lt;li>Teleological allostasis: Positions allostasis as the primary evolutionary design feature of the brain (Barrett and colleagues)&lt;/li>
&lt;li>Diachronic allostasis: Emphasizes allostasis as operating across various timescales (Pezzulo et al., Stephan et al.)&lt;/li>
&lt;/ol>
&lt;p>The authors' critique of the &amp;ldquo;behavioral&amp;rdquo; interpretation is particularly insightful. They point out that despite using
the term &amp;ldquo;allostasis,&amp;rdquo; these accounts describe what is essentially a reactive process - responding to homeostatic
perturbations rather than anticipating them. This seems to miss the core predictive emphasis that has been central
to allostasis from its inception.&lt;/p>
&lt;h2 id="strengths-of-the-analysis">Strengths of the Analysis&lt;/h2>
&lt;p>What I find most impressive about Corcoran and Hohwy&amp;rsquo;s analysis is its conceptual precision.
In a literature full of terminological confusion and competing definitions, they bring much-needed clarity.
Their systematic examination of different interpretations helps untangle what has become a rather messy theoretical landscape.&lt;/p>
&lt;p>The authors are also admirably even-handed in their assessment. While they ultimately favor a view that reconciles
homeostasis and allostasis as complementary strategies, they carefully consider the merits of alternative perspectives.
Their analysis of the &amp;ldquo;diachronic&amp;rdquo; interpretation of allostasis (particularly Stephan&amp;rsquo;s Bayesian implementation of
hierarchical allostatic control) is especially thoughtful.&lt;/p>
&lt;p>I also liked their recognition that &amp;ldquo;sustained biological viability (rather than some other criterion such as
internal stability) seems to us the most plausible target towards which physiological and behavioral regulatory mechanisms
are striving&amp;rdquo;. This shifts the focus from mechanism to purpose in a way that offers a principled resolution to some
of the conceptual tensions.&lt;/p>
&lt;h2 id="questions-for-further-investigation">Questions for Further Investigation&lt;/h2>
&lt;p>Reading this paper has sparked several questions that I believe could be subjects for further investigation:&lt;/p>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Developmental Trajectory&lt;/strong>: How do homeostatic and allostatic regulatory mechanisms develop over the lifespan?
Are there critical periods for the development of predictive regulatory capacities?&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Individual Differences&lt;/strong>: What accounts for the substantial variability in regulatory strategies across individuals?
Some people seem to rely more on anticipatory regulation, while others show more reactive patterns.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Artificial Systems&lt;/strong>: Could the complementary frameworks of homeostasis and allostasis inform the design of
artificial systems? Might robotic or AI systems benefit from implementing both reactive and anticipatory modes of self-regulation?&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Disorders of Regulation&lt;/strong>: How do disruptions in the relationship between homeostasis and allostasis contribute to
physical and mental health conditions? The concept of &amp;ldquo;allostatic load&amp;rdquo; is mentioned but deserves deeper exploration.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Consciousness and Regulation&lt;/strong>: If predictive regulation is indeed fundamental to biological systems,
what implications does this have for theories of consciousness? Could consciousness itself be understood partly as an extension of these regulatory processes?&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h4 id="link-to-the-paper-allostasis-interoception-and-the-free-energy-principle-feeling-our-way-forwardhttpsosfiopreprintspsyarxivzbqnx_v1">Link to the paper: &lt;a href="https://osf.io/preprints/psyarxiv/zbqnx_v1" target="_blank" rel="noopener">Allostasis, interoception, and the free energy principle: Feeling our way forward&lt;/a>&lt;/h4></description></item></channel></rss>