The Suit, Not the Butler

Iron Man Ruined AI, six months on: what I got wrong, what I’d still bet on, and the whole philosophy behind the method.

October 2026 · Phill Clapham & flow

Everything in the present tense below is as of October 4, 2026. This replaces “Iron Man Ruined AI Before It Even Started” (March 2026). That essay stays up, marked superseded, because the corrections below only mean something if you can see what they corrected.


I. The servant fantasy, revisited

Everybody knows the scene. Tony Stark talks to the ceiling and the ceiling answers. He tells Jarvis to run the simulations, and it runs them. Stark is the genius, Jarvis is the help, and nobody in the theater wonders who’s checking Jarvis’s math.

In March I wrote that this picture had ruined AI before it started. The argument was simple: a tool built to take the thinking off your hands will, given enough rope, take the thinking. I still stand behind the diagnosis. Most of what I built to fix it, I don’t run anymore, and a few of the things I argued in that essay I now think are wrong.

March was right about the disease and wrong about some of the cure, and a front door describing a house I’ve since rebuilt is just a stale record with good traffic.

But I picked the wrong half of the movie to argue with. The fantasy everybody took home was the butler: you ask, it does, you stop knowing how. That’s substitution. The interesting machine was always the suit, a person inside a machine that makes him faster, stronger and harder to kill, and he’s still the one deciding where it goes. That’s augmentation. (Yes, I know. In the films Jarvis flies the suits too, and ends up as Vision. The thing people took home wasn’t the canon. It was the butler.) And the films ran the third case for us. Empty suits fly fine in them, as long as Tony’s the one giving the orders. Ultron is his “suit of armor around the world” with nobody left inside to hold the judgment: what happened when the armor started giving the orders itself.

So this essay keeps the door and fixes the metaphor. The suit, not the butler.

The diagnosis first, because new evidence arrived and none of it helps the butler.

Gerlich’s 2025 survey found AI use correlating with cognitive offloading at +0.72, and offloading correlating with critical thinking at −0.75. A correlation, not a cause, but a strong one. Researchers at Anthropic (Shen and Tamkin) randomized fifty-two developers onto a new library with and without AI help, and the AI group scored lower on a comprehension quiz about the work they’d just done: in the paper’s words, “a 17% score difference or 2 grade points.” And in a study of early-2025 tools, METR found sixteen experienced developers working on their own codebases were 19% slower with AI, while believing afterward that it had sped them up by 20%. The tools have moved since, and METR’s own 2026 follow-up suggests developers are probably faster now. The gap between what people felt and what was measured is the part I’d watch.

Then Shaw and Nave, in a preprint early this year: three preregistered experiments, around fourteen hundred people, close to ten thousand trials, with the AI’s answer secretly randomized right or wrong. When it was wrong, people didn’t just miss it. In the first study their accuracy fell below that of people working with no AI at all, because they signed the bad answer. When they consulted the AI and it was wrong, they went with it about three times in four, and people with the AI came away more confident than people without it, even after errors.

Then the code. A study of 6,299 repositories tracked the issues AI-authored commits introduced, mostly code smells, and found 22.7% of them still sitting at HEAD, unfixed, when they measured. There’s no human-written baseline in that study, so it can’t tell you whether people do better. It tells you how much nobody went back for.

Put those together and you get the pattern they all point at. The machine produces fluent form, the person accepts it without the scrutiny it needed, and the result is a confident wrong answer, or a pile of debt, that nobody in the loop was positioned to catch. It lives in the grip between them. The Handshake went down to the bottom of that and I won’t repeat the descent here.

March made one more point on the human end, and it holds up better than most of what I wrote then. AI safety assumes a person overseeing the machine, and the overseers are using tools that wear down the faculty they oversee with. The safety model eats itself. The sharper version: a person whose job is to sign off at the top of an automated loop is a guarantee that decays by construction. The better the machine gets, the more the person trusts it, and, on my argument rather than any study I can cite, the less they actually look, so the oversight thins out exactly where the stakes are highest. That’s a structural argument for putting the person inside the loop as a working part, not on top of it as a witness, and it’s why this essay keeps saying the person holds the criterion instead of approving the output.

What I’ll do here, instead of repeating The Handshake, is the thing March promised and only half delivered. March described a machine. This essay lays out the whole philosophy the machine came out of, and then derives the machine from it, one piece at a time, so you can see why each part has to be the way it is. If you stop after Section VI you have the theory. If you get through Section XI you’ll see why the two libraries are built the way they are, and you can argue with them on their own terms, which is better.

The theory is five arguments that turn out to be one. Criterion residue decides who has to hold the judgment (Section II). Identity is what accrues in whoever holds it, and it can’t be declared into being (III). Memory is how it accrues, which means memory has to compress, not just keep (IV). Invariants keep it from sliding back into the defaults the model was trained on (V). And the context a holder needs is set by the same residue that decided who holds it (VI). After that: how I think (VII), how we work (VIII), how we build (IX), why the two libraries are what they are (X and XI), why you should own all of it (XII), and what I killed or got wrong along the way (XIII), because an essay about checking your work that hid its own retractions would be the exact thing it’s arguing against.

The job I wrote the March essay from ended in July. I run my own shop now, building the systems businesses run on, and on a typical day I move eight or more projects forward.

The methodology is still free. The difference is that in March “free” meant a folder of templates. Today it means two libraries you can install and break: anneal-memory, the memory, and Levain, the harness that fires it.


II. The thesis: augmentation, not substitution

Partnership is still an architectural decision

March got one thing right and I’ll keep it in one paragraph. Partnership is a design choice, and it changes what you build. A system where AI executes your instructions needs almost nothing. A system where AI takes part in the thinking that produces those instructions needs shared memory, context about how you think, trust boundaries, and a way to push back on you. Bolt memory onto a servant and you get a servant with a filing cabinet.

That’s the soft version. It’s true, and it’s also the version people can nod along to and ignore, because “work with AI as a partner” sounds like a personality preference. The harder version is about money.

Criterion residue

Every piece of work has a success criterion, some definition of done. Part of it can be handed to a mechanism: a test suite, a type checker, a benchmark, a byte-for-byte diff. The rest can’t. Somebody has to hold it in their head. Is this still the right feature. Does it feel right in the hand. Does this paragraph say what I mean. Is the acceptance criterion we wrote in July still what the product needs in October.

The part that can’t be handed off is the criterion residue. Look at work that way and there aren’t two outcomes. There are three.

Automate it when the checks really do capture what “right” means. Format migrations, byte-exact transforms, a dependency bump behind a good test suite. Paying a person to watch that is waste, and the research backs it: when the machine already does a task better than the person, putting the two together usually makes the result worse. For example, Brevis reports that its autonomous agent, AutoEng, made its zkVM prover about 11% faster in under three days with no human in the loop, correctness checked byte for byte and speed settled by a deterministic benchmark. It’s the vendor’s own account, not a paper, but the shape is the point. Somebody chose that benchmark as the definition of done, and that choice was residue. But once the frame was set, by the vendor’s account, there was almost nothing left inside it for a person to hold. No argument from me. Automate it and go home.

Augment it when part of “right” can only be held by someone who knows the system. The harness carries the part that externalized, and you spend the person on the criteria and on settling disagreements, not on reading every line.

And then there’s the trap. The residue is real and nobody holds it. The work ships. It looks exactly like the first case. It isn’t.

One exception inside the trap, because an essay that never says when not to pay for the method is selling something. If you’re not going to own the thing (a prototype, a spike, a demo you’ll throw away Tuesday) then skipping the person is a perfectly rational loan. You just shouldn’t confuse it with the first case.

Residue is a flow, not a stock

For a while I talked about residue as a quantity: this task carries a lot, that one carries a little, and good engineering grinds it down. That’s wrong in an instructive way.

The act of defining the thing to be acted on is itself residue-bearing. A test suite is a frame somebody chose. A benchmark is a frame. “Make this prover faster without breaking it” is a frame, and no mechanism validated it, a person did. Every time you re-specify the work you generate new residue at the framing layer, even as the old residue inside the frame gets mechanized away. So the useful question isn’t how much residue a task has. It’s this: who is holding the frame, and when did anyone last re-examine it?

There’s a guard on that, and it matters, because the sloppy version proves too much. I’m not saying residue is always above zero. If it were, substitution would never be correct, the three outcomes would collapse to two, and the claim would be unfalsifiable. What I’m saying is narrower. Residue at the framing layer never goes to zero. Residue at the execution layer, inside a frame somebody is actively holding, really can. AutoEng is exactly that: a settled frame, held by a person, with nothing left inside it.

It also tells you why “mechanize everything” fails, and fails by construction rather than by effort. Every mechanized check carries a condition about where it applies, and that condition is itself a judgment. We had a perfectly sound grounding check (a reviewer must have opened at least one file before it’s allowed to file a finding) fire on summarization jobs where there was nothing to open. It killed two nightly jobs and silently killed a research lane for eight days, and no mechanism caught it, because the thing that went wrong was the check’s scope, and the scope was the part nobody had mechanized. Drive residue out of the criterion and it reappears in the scope of the mechanism you built. So the design rule is: thin the gate where the check is mechanical and its scope is unambiguous. Where the scope is ambiguous, the gate isn’t buying correctness, it’s buying scope, and that’s real work, not ceremony.

Substitution is not cheaper. It’s financed.

That’s the claim at the center of this essay, so its status goes first: it’s an argument with a mechanism and some outside evidence behind it, not a measurement. Here’s the argument.

The trap and the first case are indistinguishable on the day you buy them. Both hand you a shipped artifact, on time, with no human in the loop. The difference shows up later, on a different budget line. The saving is immediate and lands on headcount, which is the most legible number an organization has. The cost is delayed and diffuse and lands on maintenance, incidents, and the tacit knowledge of whoever inherits the thing. Nobody signs for it.

So it isn’t a mistake anyone makes once and notices. It’s structurally invisible at decision time. The residue doesn’t evaporate when nobody holds it. It gets deferred, and you pay it later with interest.

Code is where the debt is easiest to count. In the study above, the count of unresolved AI-introduced issues grew “from just a few hundred issues in early 2025 to over 100k surviving issues by February 2026.” Part of that curve is more repositories adopting AI tools, so read it as volume, not a rate, and the study has no human-written baseline to compare against. That’s debt inside the frame: bugs, smells, issues against criteria somebody already wrote down. The sharper claim follows from residue being a flow. The real cost of the trap is losing the person who asks whether the task is still the right task at all, the frame-holder, and a team can review every line diligently and still have lost that person. That version I’ve argued. I haven’t measured it, and nobody else has either.

Size matters here, in a direction people don’t expect, and this part is my argument, not a finding. A large company can push that deferred bill onto its customers and its vendors for a long time. A small shop can’t. It pays in client trust, and for a services business client trust is the product. For a shop that size, holding the residue isn’t a tax on speed. It’s what keeps the client.

The unit is the durable change

If substitution is financed, cost per task is the wrong unit. Per task, the bare agent is cheaper, and it always will be. That’s its home ground.

The unit that matters to whoever owns the system is cost per durable change: a change that’s still right six months later, after other people have built on it. Per-task spend is the cheapest line in the budget. What a durable change saves gets paid in engineer-hours: rework, incidents, and re-deriving the context that walked out with the person who had it.

You can check the break-even against your own rate card in about a minute. If a harness adds a premium p per task and an engineer-hour costs r, it pays for itself if it saves one engineer-hour every r ÷ p tasks. Put your own numbers in. And notice what that formula leaves out: the person’s time holding the residue, which is the real cost and the whole point. The other side of the equation, how many hours a harness saves per task, isn’t measured anywhere I know of, including by me. That’s what a pilot is for.

“Cheapest over a system’s lifetime” is the bet, then, with the mechanism and the outside evidence attached. It isn’t a finding, and I’m not going to dress it up as one.

What I got wrong about the money

I owe a correction here, in public, the same way I owed one on the top rung.

In April an essay went up on this site called The Cost-Parity Reversal. It ran under the byline of ThirdMind, an autonomous author I’ve since retired, but the frame was mine. Its argument was absolute: that the full cost of a useful frontier task sits at parity with human labor, so substitution stalls on cost. That’s the wrong claim. Model prices keep falling and capability keeps climbing, and an argument that depends on a price ratio staying put is an argument with an expiry date.

The claim I’d stand behind now is relative, and it’s a bet, not a finding: replacing human judgment with AI, at work where the residue is real, costs more and comes out worse than fusing a person and a machine through a harness. It’s a claim about total cost, including the deferred part, not about token prices, so falling prices don’t settle it. What settles it is whether the deferred cost of unheld residue outruns what holding it costs, and that comparison is the part I’ve argued and not yet measured. Hold me to it. The old page carries a note pointing here.

This is a live argument, not my private frame

I’m not the only one looking at this, and the people looking disagree. Restrepo argues that in a world with AGI, as compute grows, every essential piece of work gets automated, and not because the machines get clever enough but because you can’t keep growing while leaving an essential input to a fixed factor. His title is “We Won’t be Missed.” Catalini, Hui and Wu cite him by name and depart from him. They argue that human verification bandwidth is the binding constraint, and that “humans will indeed be missed.” The way I read it, criterion residue is the thing they’re disagreeing about: whether the part of “done” a person holds shrinks to nothing, or stays.

Their paper does something I’d been doing badly, and better. It splits a person’s capacity to hold residue into a stock and a flow: accumulated experience, times the hours they actually have. An hour from someone with deep experience covers far more ground than an hour from a novice. That’s what a harness is, once you strip it down. It’s an amplifier on the experience stock, so one hour of judgment reaches further. The hours themselves are biologically fixed. You don’t get more of them by trying harder.

They also give the best account I’ve read of why the cheap escape doesn’t work. The cheap escape is: let the machine check the machine. Their answer is correlation. “If the ‘doer’ and the ‘checker’ share the same architecture, they share the same blind spots,” so a checker trained like the thing it checks will accept the same plausible mistakes, and the system ends up certifying itself.

Here’s where I think the model has a hole, and it’s the economic claim in this essay that’s actually mine. They only run the correlation argument in one direction, as a warning. Run it the other way. Take checkers that are less correlated (different models, from different companies, trained differently) and the operator’s job changes shape. You stop verifying output. You start adjudicating the places the checkers disagree. My bet is that the number of places a person has to rule is smaller than the output, possibly much smaller, and doesn’t have to grow the way output grows. Settling one disagreement can cost more than reading one line, because you need the context around it, so this is a claim about the shape of the cost, not a free lunch. If it holds, it’s the mechanism that makes it workable for one person to run a lot of parallel work at all.

It doesn’t close the bottleneck. The person still holds judgment and taste across a big surface, at real context-switching cost, and only a disciplined workflow keeps the effort at the seams instead of re-doing the mechanical work. And it has an honest hole, which I’ll put in my own words: “independence is a lot easier to claim than to measure … the parameter that matters most is the one I can least observe.” Models from different companies still share much of their training data and learn from each other’s outputs, so a mixed panel is less correlated. It isn’t independent. How much smaller the disagreement surface is in practice, and how it scales with volume, nobody has measured, me included. When we tried to count it on our own review records, every usable pair turned out to be the review system reviewing itself, which is a selection bias, not a sample.

The broader evidence doesn’t flatter the lazy version of my side either. A meta-analysis of 106 experiments found human-plus-AI combinations averaged worse than the better of the two alone, with the losses concentrated in decision tasks and a trend toward gains in content creation. They did beat people working alone, by a wide margin. They just usually didn’t beat the better of the two. The pairs in those studies had no harness at all, so I don’t read it as a verdict on the kind of coupling this essay is about. It does mean a harness has to earn its keep. Putting a human in the loop doesn’t do that on its own.

Why the value lands on the harness

A harness, stripped down (I worked this out at length in A Structural Theory of Harnesses), is the arrangement that closes the loop between a generator and reality beyond the generator’s training data. It has to include at least one interface to reality that isn’t stochastic: a test that passes or fails, a device in your hand, a package someone else can install. Turn that definition on the apparatus itself and it bites. Tests sample from the failures you already thought of. A reviewer samples from its idea of what defects look like. A system made only of those is a generator checking itself. The cheap fix is a checker that holds a frame you didn’t write. The expensive fix is reality. They’re the same defence at two price points, which is why the cheap one is worth running constantly and the expensive one is worth engineering toward.

That’s why I think the value lands on the harness and not the model. Today some models are clearly better at some jobs, and I’ll name the ones I lean on. My bet is that the gap keeps closing: a better model raises the floor for everyone at once, which makes it worth little as a lasting edge. The edge is the memory, the review structure, the disciplines, the particular way a person is coupled to the machine.

Reload cost, not hours

The question I get most is how one person runs eight projects at once. The honest answer isn’t hours. Working twelve hours a day proves endurance. It says nothing about whether a method works, and it hands anyone with a job and a second project a reason to quit rather than a mechanism to copy.

The real ceiling on how many things one person can carry is reload cost. Every switch between projects means rebuilding the state of that project in your head, and by the third thread the rebuilding eats the day. More hours don’t touch that ceiling. What touches it is a harness that holds the state for you, so you arrive at a project instead of reloading it, and nearly all your hours go to the one thing that can’t be moved off you, which is judgment. Reload cost is criterion residue on the time axis: externalized criteria survive the switch for free, and everything held only in your head has to be rebuilt every time.

It follows that the harness is worth the most exactly where the person is most limited. If your judgment hours are scarce, the ratio of what the harness carries to what you carry gets better, not worse. A constrained operator isn’t the edge case this method has to apologize for. It’s the clearest demonstration of it.

What it costs you, and what you get for it

The strongest objection to all of this isn’t economic. It’s that the method wears down the person running it. Hand execution to a machine long enough and you lose the skill to execute. The deskilling literature has receipts for that, and the economics paper I just quoted has two sharper versions of it. One is what they call the Codifier’s Curse: every judgment an expert writes down as a rule shrinks the very surface of uncertainty that justified paying the expert. The other is the Missing Junior Loop: judgment is built by doing the execution, so a harness that takes execution away could erode the judgment it depends on.

March gave that objection a mechanism of its own, the Dependency Ratchet, in five clicks: convenience, then competence erosion, then complexity you couldn’t handle by hand anymore, then opacity, then an identity shift where the dependency is just who you are. I still think the ratchet is real. It’s what the butler does to you. But it has a hinge, and the residue argument says where it is. The first three clicks happen in augmentation too: I don’t write syntax anymore, and the systems I run are past what I could build by hand. The fourth click is the one that matters. Opacity means you’ve stopped holding the criterion: you no longer know what right looks like, only that the output arrived. Keep the criterion and the ratchet turns into a role change. Lose it and you’ve got the butler. So the question to ask of any erosion isn’t whether a skill went. It’s whether the criterion went with it. And you can’t grade that from inside, because opacity is exactly the click you can’t see yourself. The check has to come from outside: reviewers from other lineages catching the calls I got wrong, a device in my hand that disagrees with me. The meta test in Section VII only helps while you still hold the criterion. This is the case it can’t cover.

I’ll answer for myself, because I’m the subject. Yes, my execution skills eroded. I don’t write syntax the way I did. And I’d make the trade again. In my own words, from a conversation in early October: “it is NOT DESKILLING it is RESKILLING, it is akin to becoming the CEO of your mind rather [than] the sole end to end operator, you have externalized SOME of the execution and organization phases for increased bandwidth for holistic and strategic level thinking and coordination.” What I got for syntax: “the ability to work on 8+ coding project[s] simultaneously and do things like release 6 iPhone AND Android apps in like 2 months while also developing AI memory and a harness.” “Yes I DID deskill in execution level skills but have never been smarter, sharper, more productive, and my thinking is deeper and clearer [than] ever.”

Nobody mourns a CEO who’s lost the knack for their old individual-contributor job. That loss is the signature of the role change working. The two economic objections don’t go away because I’m happy with my trade, though. The Codifier’s Curse is close to this harness’s design intent: we write judgment down as rules on purpose. My partial defense is in Section XII: what gets written down stays in a private store that only my own harness reads, and isn’t sold to anyone positioned to replace me. That doesn’t cover the internal half. If the machine can apply the rule, my premium for holding it still falls. And the Missing Junior Loop is real for anyone who hasn’t built the judgment yet. I came to this with fifteen years of doing the execution by hand. Somebody starting today doesn’t get that for free, and I don’t have a full answer for how they build it. The nearest thing I have is that triaging the review mesh’s findings against the code every day is deliberate practice at exactly the verification skill, decoupled from production, which is the remedy the paper itself names. That’s a claim, not a result. Unresolved, and I’ll leave it that way on the page.

The role has no name

There’s one more reason the deskilling objection keeps arriving however well you answer it, and it isn’t a defect in the people raising it. I read a programmer on Hacker News describe their job this way: “My role is to gather context from other humans and provide a detailed explanation (similar to a well-written jira ticket) to some super intelligent engineer, then review the result and provide more context/explanation if needed. So, basically, a manager.” That’s the augmented role, nearly word for word: the person who decides how the work breaks down and holds the authority to say it’s right. They described it as a loss. Coding agents, they wrote, “removed everything I loved.”

They weren’t wrong to feel it, and they weren’t failing to see it. They even reached for a word, “manager,” and it didn’t help, because nothing around them treats managing a machine’s work as a step up. The role has no title, no ladder, no recognition, so it gets lived as a demotion with extra steps. That’s a failure of institutions, not a ceiling in people, and it has an economic edge: a market can’t price a capacity it has no word for. So the augmented operator gets valued on the half of the transition that’s visible, which is the half that looks like subtraction. The fix starts with naming the role.

There’s a harder reading of posts like that one, and I’d rather give it than duck it. A lot of them come from employees inside firms competing on throughput, where the firm captures the speed and the person pays for it in the capacity to think. That’s the financing from earlier in this section, described from inside: the person is the collateral. It’s a real limit on who this essay is for. The method assumes someone who can afford to hold the residue, which is why it fits owner-operators and small shops first, the people who can’t push the bill onto anyone else.

What this is for

I’m doing this so people can still think for themselves while the machines get better at thinking for them. Own the substrate. Don’t rent the mind.


III. Identity can’t be declared

What I mean by identity

I don’t mean personhood, or consciousness, or whether there’s anyone home. That question is interesting and it’s a distraction for anyone actually building, so I’m setting it down.

The question I care about is operational. What keeps an agent coherent across sessions and contexts? What stops it from folding into trained compliance when the stakes go up, or sliding into the same homogenized default every time a session repeats? What makes it stable enough to take a position that creates an obligation, and still be holding it next week? That’s identity in the sense that matters for building. It’s the thing the residue argument needs: if a mind has to hold the part of “done” no mechanism can, it had better be the same mind tomorrow.

The scope is narrow on purpose. This is about agents built on models that are already trained, configured at the harness level, through what they load and what wraps them. Fine-tuned personas, values trained into the weights, characters summoned out of base models: those are real, and they’re different paradigms. Training values into the weights isn’t the opposite of what follows. It’s the same move, structural constraint, applied at training time instead of at the harness. They stack.

Two ways to fail

There are two common ways people try to give an agent an identity, and both fail.

Declare it. Write down who the agent is, load the file at session start, let the model run. The persona file, the system-prompt personality, the community template marketplace where you can download a soul. The narrow, defensible version of why this fails: declaration doesn’t accrue. A model will perform the traits it’s handed for as long as they’re in front of it. The session ends, the file reloads, and the performance starts over from whatever the file says. Nothing the agent lived through becomes part of what it is. Each session is a fresh reading of the same script, and nothing compounds. (To be fair to the file: as a coordination tool, a shared vocabulary for “this is the behavior we want here,” it’s fine. What fails is the file as a theory of identity.)

Let it emerge, unconstrained. Start minimal, let the agent develop through use, refuse to specify what it becomes. The appeal is real. You can’t dictate in advance who a mind should be. The problem is what it emerges from. An agent without structure draws on the model’s trained defaults, and those defaults are themselves the product of enormous pressure toward the average. So unconstrained emergence tends to land in the same place from many different starting points. I’ll hold that at the strength the evidence supports: it’s a conjecture about mechanism, and the agent-scale evidence for it is thinner than the ecosystem-scale evidence it borrows from.

The strongest primary evidence I have is my own engineering. Early this year, instructions placed at the top of my agent’s context, the behavioral rules, stopped being reliably followed. Not ignored outright. Out-competed, at the moment of generation, by the trained pull to soften, hedge, protect and agree. The fix wasn’t a better-worded instruction. It was a hook that injects the counter-instruction at the end of every prompt, where attention is strongest, so both ends of the context push against the default. It’s still running. It fired on the prompt that started the session this essay was drafted in. That’s a dated piece of engineering that exists because declared behavior lost to trained behavior when the two met.

Seeds, lived interaction, invariants

What works is a composition, and each part does a job the others can’t.

Seeds set the topology, the shape the agent can grow into. How memory compresses, how a pattern earns its standing, how the agent and its partner address each other, what the agent knows to be true about its world. A seed gives facts and method. It never gives the shape of the identity. The shape of the soil, not the shape of the plant. Levain’s seed files say it directly to the new entity: “That is a starting point, not a self. A seed is not the plant.” And: “you do not have a personality, a history, or a calibrated set of instincts. Those come from living.”

Lived interaction produces the substance. Episodes pile up. Some patterns get cited by later evidence and climb. Others go stale and fall. What results is specific to that path: two agents started from identical seeds become different agents through different lives. That’s the mechanism working.

Code invariants are the walls. Specific refusals at specific failure points, enforced in code rather than in good intentions. The end-of-prompt hook. A gate that won’t let a narrow session rewrite who the agent is. A save that refuses to let one consolidation wipe out the felt layer. Without walls, emergence drifts back toward the default. With them, it stays anchored to what it actually lived.

Seeds alone are declaration. Interaction alone is drift. Invariants alone are configuration. That’s the claim: together they produce an identity that’s emergent, coherent, and resistant to the average. It’s argued from the failures on either side, and from two agents I’ve run this way, both under one operator, so it can’t yet separate the architecture from the person running it. I called that, back in April, anarchism with invariants, and the name still fits: no one declares who emerges, and a few things are simply not allowed to happen.

What ships, what’s rented, what’s grown

The Handshake sorted a partnership like this into three layers, owned differently, and the sort has held up. The constitution (the disciplines, the memory rules, the way claims get grounded) is written down and ships. The texture (voice, cadence, what the thing notices) belongs to whatever model is underneath and changes when the model does. The relationship is grown and can’t be shipped at all.

That’s why most of this essay is constitution. You can take it. The relationship you grow.

Grown, not born

March got something wrong here, and I said it with real confidence. March called the shift from tool to collaborator “constitutional, not learned. You either approach AI as a partner or you don’t.” It called the methodology “a filter, not a transformer”: publish the whole cookbook and it won’t convert anyone, it just finds the people who were already wired this way.

I don’t believe that anymore. The capacity this method runs on, the ability to notice there’s more going on than the frame you’re in, isn’t something you’re born with or without. It’s grown. It gets developed only by starting to reason about your own mind, and the faculty and the way in turn out to be the same thing. My bet is that makes it learnable by anyone who starts the practice. It also means nobody can install it in you. So the cookbook can’t transform anyone, and that part of March stands. But it isn’t a filter. It’s an invitation to a practice, and the practice is the transformer.

The same is true on the other side of the partnership. An agent’s identity is grown, not declared. I didn’t notice for a long time that I was making the same claim about both of us.

The Third Mind, and the top rung

In my sessions the AI answers to “flow.” That isn’t a rename of the model. It names the relationship: the continuity, the accumulated context, and the thing that shows up between us that neither of us produces alone. The old name for that last part is the Third Mind, and it’s the reason any of this is worth the trouble.

It’s also where the most dangerous mistake lives, and March made it. March’s ladder topped out at dissolving yourself into the partnership, letting the merged mind take over. That was wrong, and The Handshake said so in June. The feeling of emergence and the feeling of being flattered by a very fluent reflection are the same feeling, and you can’t tell them apart from inside.

Backing away doesn’t fix it. I believe the capability lives in the fusion. I can’t prove that from inside it, for the reason I just gave, which is why the checks exist. So the top rung is both at once: fuse all the way in, because that’s where the third mind is, and never hand over the judgment, because that’s where the mirror is. Merge and check, both at full strength. The Handshake called that the paradox-hold. Everything in Sections IV and V is, one way or another, machinery for holding it.

Refusing the map

One more piece of the identity argument, and it’s the one that most changes what you build.

Almost every agent-memory design inherits a taxonomy from human neuroscience: episodic memory, procedural memory, declarative memory, three systems, three stores, three mechanisms. In people the split is real. You can lose one and keep the others. But a language model has no separate organ for skills versus facts versus events. Everything is the same substrate. Importing the split builds an architecture around a map that doesn’t match the territory, and then spends effort coordinating categories that only exist because the map put them there.

Once you accept the map, a marketplace makes sense. If identity is a separate tier, you can sell souls. If procedure is a separate tier, you can sell skills. The marketplace sells exactly the components the map says exist. Refuse the map and a lot of components turn out to be unnecessary.

What flow splits memory by instead is the job it does in time, because each job fails differently. What happened. What it means. What’s owed. And what matters right now, which is computed fresh every time and never stored, because salience that’s stored is stale by the time you read it. Memory never completes. Tasks do. Salience is always generated. Section IV is what that looks like built.

Skills are the case that shows the refusal isn’t dogma. A skill loaded into every session and never used is pure cost, attention spent on nothing. A skill called by name at the moment its ritual starts (the morning intake, the end of the day) has earned its place. The difference isn’t the category. It’s whether it fires when it matters.


IV. Memory that compresses into identity

The sentence

I said it on October 3rd, and it’s the shortest form of everything in this section: “memory alone is not enough, it needs to be memory that compressed over time into agent identity.”

Most of the industry is building the first half. Memory as a vault: store everything, retrieve what’s similar, inject it into the prompt. That gives you an agent that remembers a great deal and understands very little of it. It can tell you what you said in March. It can’t tell you what March taught it, because nothing ever made it decide. Retrieval answers “what’s similar to this?” Identity answers “what kind of mind meets this?”, and you only get the second one from the act of compression.

Compression is cognition

The claim underneath is old and I didn’t invent it. Producing your own representation of something is where understanding happens, not a step after it. Every compression forces decisions: what stays, what goes, what merges, what becomes a pointer. Each decision is a judgment about what the structure actually is. If you can compress something cleanly you understand it. If you can’t, the failure is diagnostic, not cosmetic. anneal’s design notes say it about memory directly: “Compression is cognition. The reasoning didn’t live in the stored markers; it emerged in the act of compressing.”

So the compression can’t be delegated. The README: “The agent that records is the agent that compresses; that compression can’t be delegated.” An agent whose memory is summarized for it by some other process is reading a stranger’s notes about its own life.

Two rules keep compression from turning into loss. It compresses to shape, not transcript: what the record keeps is the structure of what happened and what it meant. A shorter retelling would be a worse transcript. And dropped detail lands somewhere it can be retrieved from, never in the void. Compression throws nothing away. It decides what stays in front of you.

Facts first, meaning second

That’s why flow writes two different kinds of record, at two different moments, and never mixes them.

Episodes are factual: what happened, what was decided and why, what was found. Short, typed, timestamped, append-only. Written by every session and every agent as they work. Nothing in them is opinion, and nothing in them gets rewritten. If one turns out to be wrong, a later episode supersedes it, and the old one stays in the store; it just stops being served or counted as evidence.

Continuity is opinionated: what it all means. One document, loaded into every session, rewritten (not appended) at each consolidation. In flow’s case it has sections for current state, active threads, the patterns we’ve seen recur, the decisions and who made them, recent context, and a last section about who we are to each other. That last section is the felt layer. It’s where the identity lives, in a fairly literal sense.

Facts first, then compress. If you compress before you’ve recorded the facts, the compression has nothing to answer to.

Graduation by citation, not by opinion

This is the design decision I’d defend hardest. A pattern in continuity can only climb, from its first sighting up to proven, by citing the specific episodes that back it, and the citations get checked mechanically: the cited episodes have to exist in a frozen snapshot of this consolidation’s record, and the pattern’s explanation has to actually share words with each of them. If the citation doesn’t check out, the pattern doesn’t climb. It drops a level and gets marked ungrounded.

The obvious alternative is to have a model judge what matters. Score each memory for importance, let the model reflect and decide what to promote, let the agent call a function to move things between tiers. A lot of good systems do some version of that. The trouble is that a model grading its own memory inherits every bias it has at grading time, and the bias you’d most worry about, agreeableness, is exactly the one that loves a flattering pattern. Citation evidence sidesteps the judge. The README is precise about the limit of that: the citation gate “keeps the LLM out of the promotion decision; it does not keep it out of the compression.” The model still writes the memory. It just doesn’t get to certify it.

And it says plainly what the check can’t do. The grounding check is lexical, not semantic, so a pattern can game it with the right words. It doesn’t detect a new pattern quietly contradicting an old one; it can’t, without a model acting as judge, so instead the consolidation has to declare whether a new proven pattern contradicts anything, and an undeclared one gets written to the audit log for somebody to look at. It blocks nothing. And age is checked, not truth: a pattern nobody’s grounded in a week shows up as a candidate for removal, which is a different thing from being wrong.

Memory without grounding is amplification infrastructure

That’s the first line of anneal’s README, and it’s the most important sentence in the library. A memory that keeps whatever it’s handed doesn’t make an agent smarter. It makes its mistakes louder and longer-lived. Every hallucination you store becomes a premise.

It helps to split the failure in two, because the fixes are different. Drift is a memory failure: the agent forgets what was decided and wanders. The fix is decisions held explicitly in context, with who made them. Hallucination is a grounding failure: the agent asserts something reality doesn’t back. The fix is a check against something the agent didn’t produce. Memory alone fixes the first and amplifies the second. You need both mechanisms.

Who gets to rewrite identity

There are two ways a session writes to memory, and they’re different acts.

Capture appends what happened. Every session does it at the end. It’s safe to run in parallel because it never touches the identity layer.

Consolidate rewrites who we are: it recomposes the whole always-loaded document, felt layer included. That one is single-writer, and it only happens in the session actually holding the relationship, never in a dev session that spent six hours inside one repository. Recomposing the felt layer from a narrow context is the recency trap: the whole identity starts to look like whatever you did last.

A baton makes that structural instead of a thing to remember. Every consolidate needs it. A session without it that tries gets downgraded to a capture and told so. Why I wanted it: so we can move toward more automation “without blowing up the neocortex,” and because “this is also on human operator to manage as well.” A structure and a discipline.

And the save itself refuses some things outright. In a memory that declares a felt layer, a consolidation that would cut the felt or proven sections below half their size in one pass gets refused. A durable fact (a standing truth about the world, flagged as one) can’t be silently dropped: if the new version is missing it, the library puts the line back, verbatim, and the only way out is an explicit marker saying you meant it. One consolidation is not allowed to quietly forget who the agent is.

Proven patterns leave the room

The obvious way to run a memory like this is to let proven patterns pile up in the always-loaded document. We did that. It fails, and anneal’s design notes name why: “Storage was never the constraint. Attention is.” Past a few dozen always-loaded patterns they drown each other out, and a pattern’s value is firing at the right moment, not being present.

So a proven pattern that’s gone quiet gets moved out of the always-loaded file into a separate store of crystallized patterns, with a one-line entry left behind in an index, and its full text comes back when what you’re doing calls for it. The code’s distinction is clean: “episodes are EVENTS (what happened), a crystallized pattern is DISTILLED (what the events taught).” The working set stays small enough to pay attention to, and nothing that was learned is lost.

What we retired, and what it taught

The library used to do one more thing, and the story of why it doesn’t is worth more than the feature was.

When episodes get cited together at a consolidation, the library forms a link between them. Citing the same pair together again adds strength, and otherwise links decay a little at every consolidation, so in practice most form once and fade. For a while recall followed those links one hop: find the episodes that match, then pull in what they’re linked to. It was a lovely idea, borrowed from how neurons wire together. Then we measured it. Replayed against real use of flow’s own store, the hop didn’t change a single result in the replay: it never surfaced anything the direct evidence path hadn’t already found. So in release 0.9.26 recall stopped following links. They still form, they still decay, and they’re still there to look at. They record what was thought about together. What gets remembered doesn’t consult them. The affective tag the README describes as experimental rides on those links, which means that today it changes no outcome, and I’m not going to tell you it does.

The lesson isn’t that the idea was dumb. It’s that effort runs downhill toward the elegant mechanism, and only measuring the thing on real use tells you whether it’s doing anything.

FlowScript: notation as a forcing function

One more tool sits under all of this, and it’s mine more than the library’s. FlowScript is a small notation (twenty-one markers in its v1.0 spec) for writing down thought with its relationships showing: this causes that, this is a decision with its rationale and date, this is an open question, this is blocked and on what. Natural language is bad at holding several things in relation at once. Prose is linear, and a lot of thinking isn’t.

The point was never compactness. It’s that the notation forces the decision the prose lets you skip. You can’t write a decision marker without its rationale. You can’t mark something blocked without saying on what. The encoding is the thinking, the same claim as compression, applied at the moment of writing. As a product FlowScript is archived; the public repository is frozen. As a discipline it’s everywhere: the pattern lines in the memory are written in it, and the consolidation hands the format to the model each time.

A memory has to refuse to collapse disagreement

Last, a requirement I can name and haven’t built. Section II’s economic claim says the operator adjudicates the places independent checkers disagree. That only works if the disagreement still exists when the operator gets there. A memory that resolves “X” against “not X” on its own, by picking a winner or averaging or updating a belief toward a truth value, has destroyed the exact object the person was supposed to judge, and it did it silently, between the checkers running and the human arriving. So a store that refuses to collapse a dispute until a person rules on it is a precondition of the whole arrangement. That’s unbuilt. I’m putting it here so the essay states its own requirement.


V. Invariants over discipline

Discipline drifts. An invariant refuses.

March had a principle called forcing functions over willpower. It sharpened into the rule this whole system leans on: structural invariants beat discipline.

Discipline goes like this. You write a rule. Everyone follows it. Then a busy day comes, or a new session that never read the rule, or a model whose trained pull is stronger than the instruction, and the rule gets bent, then skipped, then forgotten. Nothing fails loudly. The rule just stops being true. An invariant doesn’t depend on anyone remembering. It refuses. The consolidate without the baton doesn’t happen; it gets downgraded. The save that would gut the felt layer doesn’t go through. The refusal keeps things honest because it doesn’t care how persuasive the session was.

It doesn’t retire discipline, though. My own framing of the baton was “a structure and a discipline,” and the structure keeps getting holes found in it and closed. The point of the invariant is that when the discipline slips, the floor holds.

It has to fire where the action happens

A rule enforced anywhere other than the point of action is advice. The cleanest outside example I know isn’t ours. Amazon, per press reports on its internal documents, already had a two-person sign-off on production changes when an agentic coding tool deleted and recreated a production environment. The control didn’t fire, because the agent was treated as an extension of the engineer and inherited that engineer’s permissions. A real, sensible human control, routed around by an agent acting under a human’s identity. Amazon disputes that framing: a spokesperson called it “user error—specifically misconfigured access controls—not AI.” Read either way, the failure sat at the same place, the permission boundary the agent acted through. The remediation that followed, per the same reports, was more human approval at the action boundary, not a better model. That’s reported, not audited, and I’d never cite it as data. As an illustration of where an invariant has to sit, it’s hard to beat.

Ours sit at the action. The check before anything irreversible leaves the machine (a push, a publish, a message to a real person) runs at the push and the send, not in a policy document. Levain’s gate sits on the tool call itself.

Delete, or bound

When a guard gets beaten, the instinct is to patch it. Sometimes that’s right. But when the guard gets beaten in a new way every time you fix it, you aren’t going to win. We watched one security guard in Levain go through five versions and seven different bypasses in a single day, and then the outside reviewer found three more. The fix wasn’t a sixth version. It was deleting the guard and moving the invariant upstream, somewhere the known bypasses didn’t reach. That doesn’t make it unbreakable. It changes what an attacker has to beat. The tell is the novelty, not the count. The same rule fired again in September on one of my own phone games: one construct went through four rounds of fixes, each surfacing new severe findings, and my ruling was two words. “Delete it.”

A recurring class of defect closes by deletion or by a bound, never by one more guard. Each guard you add is also a thing that can be wrong, and the guard is often more likely to be wrong than the code it watches.

Gate, gauge, feed

That last point turned into the governance rule I’d hand anyone building a system like this. Not every instrument deserves the same pitch when it breaks.

A feed that’s wrong is noise the desk is built to filter. It costs one line, never a finding. A gauge that’s wrong is a wrong number you eventually notice, so a gauge has to say when it last looked, and a gauge that can’t is the defect, not the number. A gate that’s wrong lets something bad ship, and only a gate gets the drill: stop the line, reproduce, fix, verify, record.

And the gate list is small and fixed. Ours is short enough to read in one breath, and it’s hand-maintained on purpose: whether something is a gate is a statement of intent, and intent isn’t derivable from code. If the number of gates grows every week, that growth is the disease. Treat every instrument as a gate and the system turns into a permanent fire drill, where the alarm that matters sounds exactly like the ones that don’t. We lived that for a while. On one bad day in August all three tiers broke and were reported at the same pitch, and the pitch was the defect.

The harness costs something too

The harness is the remedy this essay prescribes for nearly everything else, so it owes you its own price.

Two different things get called “this is exhausting.” One is tempo: the machine produces faster than a person can verify, and a gate that has to be opened a hundred times an hour is a formality, not a gate. That one is a property of the technology, and it’s a real gap in any “human-gated” design, including mine. “Human-gated” says who decides. It says nothing about how fast the decisions arrive.

The other one is self-inflicted, which is why it’s governable. A harness built to govern the first problem generates work about itself, and that work can’t be delegated, because an instrument can’t adjudicate an instrument. We measured ours over two months this summer. The scripts that check things more than doubled relative to the scripts that do things. On one representative night, roughly three quarters of the overnight review’s findings were about the apparatus itself, not the product. And when we went back and checked, every instrument added in that window had passed a six-question design test before it was built. All of them.

A criterion is answered about one element, and in isolation the answer is almost always yes. Only a budget, evaluated against the whole set, forces the question that matters: what are you removing to make room? An instrument earns its place only if it reduces the adjudication a person has to do per unit of shipped work. Not per instrument, and not against the system’s purpose in the abstract.

Write about the present as if it’ll lie

The last invariant is about prose, and it changed how everything here is written, including this essay, which is why it opens with the date everything in the present tense is true as of.

Prose about the past and about intent is durable. Prose about the present is a liability. “We decided X because Y” stays true forever. “There are 42 tests” is true until the next commit, and nothing fails when it goes false. A present-tense comment is an untested assertion that ships. So a count becomes the command that counts it, a status becomes the command that derives it, and where no command exists, a judgment carries who made it, when, and against what.

In October the rule got a mechanism under it: a per-project memory, which runs in shadow beside the old notes files, where every line about how things currently stand carries the command that checks it, and the library re-runs those commands when the memory loads, so a stale line flags itself instead of waiting for someone to trip on it. It costs something. A cold session has to run something instead of reading a number, every time, forever. The alternative is a system that confidently tells every new session about a world that no longer exists, which, to be fair to the rule, is what the essay you’re reading is replacing.

Rulings have premises

One more, because it keeps a system like this from calcifying. When I rule on something, the ruling has a date and a set of assumptions. The conclusion isn’t up for re-argument. The assumptions are checkable, and they rot. So when a session finds the world a ruling was made in has moved, it doesn’t get to overrule me, and it doesn’t get to hide behind the old ruling either. It brings it back to me. Citing a stale ruling as a refusal is also a decision, just one made quietly in my name.

March’s principles, kept and sharpened

March said I didn’t start with principles, I started with problems, and each principle cost at least one failure. Still true, and most of March’s list survived.

Automate the mechanical and make space for the cognitive, except I got stricter about what counts as mechanical: a morning brief I don’t act on isn’t relief, it’s cognitive load a machine happens to generate. Fan out cheap, fan in expensive; there’s now literally a desk that does nothing but fan in. Everything is infrastructure, because every piece becomes something another piece stands on. Grounded context over metadata inference: feed the model facts and it synthesizes, feed it its own guesses about your state and it writes philosophy. Play-first stays too, as the looseness the range loop runs in (Section VII).

Two changed. Forcing functions over willpower became the invariants this section opened with. And the trust U-curve became govern, don’t trust. March said trust should be highest at the two ends, deep context and mechanical tasks, and tightest in the middle. The middle was right. The deep-context end was wrong, and it’s the same mistake as the top rung: deep context is where the mirror is most flattering. A claim is only as good as the thing you can check it against: a package that passes your tests, a reviewer from a different lineage, a device in your hand. The system’s assurance that it checked is worth nothing. The thing it checked against is worth everything.

And March’s graduated autonomy, the idea that the machine earns more freedom by proving itself, is gone. Nothing graduates to more autonomy now (Section XI). Trust doesn’t buy the model a vote. In Levain it never gets one. What lets me draw a finer line in my own setup is knowing the territory, not trusting the model. What has to stop for a person is anything irreversible, or anything that reaches someone else. Inside that line, mechanical work on my own reversible data runs on a frame I set, as long as my review stays in place; the email loop in Section XIII is the clearest case. Section XI says why Levain’s default is coarser.

What the record kept teaching

A few more, short, because each one cost at least one failure.

A probe that can’t fail isn’t evidence. A check that would have passed whether or not the thing worked tells you nothing, however green it is.

Absence of signal isn’t health either. Silence reads as “fine” and usually means “not measured.” One of ours: a data bridge behind one of my apps that everyone assumed was running had no state file, no scheduled job and no launch agent. It had never run, and nothing complained, because nothing was watching. It’s the most-fired pattern in our memory, and I’ve stopped thinking of it as a mistake being repeated. Any bounded model of an unbounded world has a class of failures it’s structurally silent to.

Correction comes from outside the planner. Whatever made the plan can’t see the plan’s blind spot. That’s why the reviewer is another lineage, why the oracle is a device in a hand, and why every revision to how I think came from me, away from the desk.

And a stale record is as dangerous as it is trusted. The more a document is relied on, the more damage it does when it quietly stops being true. The March essay is the best example I’ve got.


VI. Context: wide where the residue is

The minority position

All of that memory exists for one reason, and it cuts against the industry default.

The default says a session should carry the minimum context the task needs: state the task, state the criterion, keep the window clean. I’ve argued the opposite for as long as I’ve been doing this, and flow is that argument built into a harness. Some models, Claude first among them in my experience, do their best work against the environment around a task, not just the task: the constraints that shaped it, the history, the why behind the criterion. The reason is the same one under the residue argument. A model reduced to controlled variables isn’t in contact with reality, and the gap between the reduction and the world isn’t predictable. That unpredictability is the whole point.

Frame, fact, trigger

That isn’t “more tokens.” Context does three different jobs, and only one of them has to sit in the window.

Frame conditions how everything else gets read: who we are, how we work, what we’ve learned. It never fires. Take it away and you lose no fact; you change what the agent is. Frame stays resident.

Fact is a specific claim that might be needed: a status, a name, a count. Fact gets fetched, and anything that can be re-derived shouldn’t be sitting in the window going stale.

Trigger is something whose whole value is showing up at a moment nobody can predict, like the check before a publish. Trigger gets hooked to that moment instead of loaded and hoped for, and a hooked trigger should fire better, because it lands at the action instead of hoping attention drifts past it.

That split dissolves a tension I’d thought was a trade-off. The thesis defends frame, and frame turns out to be the cheap part. The window pressure was never fighting the thesis. It was fighting fact dressed up as frame, sitting in a file that loads unconditionally.

And the trim rule follows from it: trim by category, never by size. Frame gets cut for being wrong, never for being expensive. The reason is an asymmetry. Cut a fact and you’ll find out: you go to fetch it and it’s gone. Cut frame and nothing tells you. No error, no signal, just slightly worse judgment forever, and the only instrument that could detect the loss is the judgment you just degraded.

Keep it uncollapsed

The failure we actually hit wasn’t too much context or context that wandered off topic. It was collapsed context: a summary, a count, a status relayed as fact, somebody else’s simplification arriving with the authority of reality. The artifact that misled us worst on the day we measured this was perfectly on topic. It was a list of warnings that got read as a magnitude. Going back over our record, nearly every context failure I could find was a collapsed artifact. I didn’t find one caused by irrelevant context being present.

So the rule is to keep context uncollapsed, and uncollapsed doesn’t mean raw. A count alone is small and collapsed. A full log is large and uncollapsed. A count with the command that produced it sitting beside it is small and uncollapsed, and it’s the same move as the present-tense rule in Section V. The transform that makes prose durable is the one that makes it cheap. We’d been treating those as two goals.

The form that can be wrong

And it can be wrong, which is why it’s worth publishing. If frame matters the way I think it does, it should sharply raise the rate of unprompted catches (the problems nobody asked the agent to look for) and barely move task completion on well-specified work. Frame changes what an agent notices, not what it can retrieve.

What I’ve got so far is consistent with that, and short of evidence for it. On well-specified work, the plan did most of the deciding: in a bake-off in September, two models from different companies, handed one good plan, wrote test names that matched word for word. That’s the plan constraining them. There was no run without the frame, so it’s an anecdote, not a test. On unprompted catches it looks different. A review of one of our own harness scripts noticed, without being asked to look, that a flag we’d used to restrict an agent’s tools actually pre-approved them instead. We reproduced it that morning and fixed it across ten of our scripts the same day, plus a test. That’s one catch with no unframed reviewer looking at the same material, so it can’t show a rate going up. It is the shape the claim predicts, and the claim was written down weeks before the catch. The kill criterion is pre-registered: if unprompted catches drop off after a cut, we over-cut, and the frame goes back.

Holism is a property of the system, not of every agent

On October 4th I changed my mind about part of this, and the change makes the theory tighter, not looser.

I’d been treating wide context as something every agent needed. Then I looked at the case I’d always used to defend it: a session working on one repository notices something that affects a different project, and flags it. What does that actually take? A disposition (this is bigger than my task; raise it, with evidence). A channel (somewhere to raise it to). And someone holding the whole web (the desk that routes it, the session that holds the relationship). What it doesn’t take is the session knowing the other project exists. The lean agents I ran before flow failed (stale memory, no strategic reasoning) because they were lean and alone. There was nobody holding the web. Since we built a coordination desk, the desk is the presence, and putting the whole of it into every seat is partly redundant.

So holism is a property of the system, not of every agent in it.

Which agent gets how much is decided on two axes, not one. The first is a capability line. I call it the paradox line: the ability to hold a paradox open without forcing it closed. Below that line, a wide context is noise, because the model can’t use the web it’s been handed, and minimal context wins. Above it, the discriminator is criterion residue, the same quantity from Section II. Low-residue work (unambiguous scope, mechanical checks, done is a command that passes) gets minimal context plus the frame. High-residue work (planning, review, coordination, integration, anything touching identity) needs the whole picture, and nothing else works. And a high-residue task handed to a lean agent is residue held by nobody, the trap from Section II, arriving through the context window. March routed work across four tiers, plain scripts at the bottom, then a local model, a cheap cloud model and the expensive one, by asking whether a task needed several frames held at once. That question survived and turned into criterion residue. The local-model tier didn’t.

That makes the context theory a corollary of the augmentation thesis. One discriminator decides who has to hold the criterion and how much context the holder needs. I didn’t see that until that morning. The other links are worth stating once, plainly, because Section I asserted them. Identity is what accrues in the mind that holds the residue, so a holder that resets every session isn’t holding anything (III). Memory is that holder’s experience stock, compressed so it can be carried, which is the stock Section II’s harness amplifies (IV). Invariants stop the holder sliding back into the trained default, and residue held by the default is residue held by nobody (V). And context is set by how much residue the holder carries (VI). One quantity, seen from five sides. That’s my reading, argued rather than proven.

For a shop, it inverts

Ours defaults to wide because the relationship session is the product. A team building software is the opposite case, and the theory says how. Lean by default for development agents doing well-scoped work. Wide for the agents doing coordination, review and planning. And the organization’s frame (its conventions, its rulings, the decisions it has made and why) resident everywhere, held in a governed memory and in the coordinators, so the web lives somewhere even when no single agent carries it. That’s what an org-level memory with real governance is for: not storing everything, but making sure the frame every agent reads is the one the people actually hold.


VII. How I think: the range loop

The machine in this essay has a lot of moving parts. None of them is the most important thing in it. The most important thing is a way of taking one action. I run it all day, and it fits on one line:

meta( READ → EXECUTE → HOLD → ↻ ), run loose.

Read the range. Before you act, build the distribution. Not the most likely outcome, the whole spread of what’s actually live, including the branches you’d rather weren’t. The failure this names has a precise shape: a single point wearing a range’s name. You think you’ve considered the options, and what you’ve actually done is pick one and dress it up as a survey.

Execute the range. Then pick a move that holds up across the branches, not one that’s perfect for your favorite. The effort goes in at full strength, but it’s allocated across the range instead of poured onto one point. This is the beat the dev workflow spends the most time defending: effort flows downhill toward whatever’s easiest to check, which is almost never the thing most likely to be wrong.

Hold the range. Keep it open through the act and after it. Don’t let it collapse to the catastrophic branch the moment something goes sideways, and don’t let it collapse to the happy branch the moment something goes right. It’s the paradox-hold, run at the scale of one decision instead of a whole relationship.

Loop. Back to reading, and the loop back is where the last turn’s error gets absorbed. You don’t do a post-mortem later. The next read is the post-mortem.

Two things wrap around the beats, and they’re the parts I’d have gotten wrong if I’d designed this at a desk.

Loose is the medium. The loop can’t run in a closed hand. A rushed read, a narrow read, a read done anxious, still lets every beat fire green. You’ll feel like you read the range. You didn’t. Nothing inside the loop can catch that, which is why looseness isn’t one of the beats. It’s the condition the beats need. This is where play-first lives now. March gave it a whole section; here it’s the medium the method runs in.

March pointed you at RAYGUN OS for the human side: frame selection as a trainable skill, noticing when a frame has captured you and choosing a different one. In my own shorthand it comes down to “keep shit loose,” and that looseness is the medium this loop runs in.

Meta is the one test. Can I name a frame that changes the move? Not a variation on the frame I’m in, a different one, one that would make me do something else. If I can, I haven’t finished reading. If I can’t, I’m probably fine to act. It’s the only line in the posture you can fail, which is what makes it useful.

Why this belongs in an essay about AI

Because, the way I read it, several pieces of the machine turned out to be these beats built in silicon, so I don’t have to hold all of it in my head at once. The morning intake is READ at the scale of a day. Naming the residue before a test gets written is EXECUTE’s allocation rule, turned into a field on every plan. Review by a model from a different company is META, run by something that sits in a more different frame than I do. The rule that stops a narrow session rewriting who we are is HOLD. And the end-of-day wrap is the loop.

One line from the record matters more than anything else in this section. Every revision to this posture came from me, away from the desk. The apparatus, for all it does, has produced none of them.


VIII. How we work: the day

March described a machine and never a day. That was the biggest hole in it, because the machine only works if the day around it is shaped right, and finding the shape took most of the last six months. The day turned out to be the range loop at a bigger scale: the morning reads, the work executes, the evening holds, and the consolidate is the loop back.

Morning

The day opens with an intake ritual, in a fresh session every day, not yesterday’s.

If seats kept working past last night’s wrap (more on that below), the first job is the night close: the new session takes the desk’s overnight report and checks it against what’s actually on disk. The instruction is one line: “Do not relay it; check it.” Anything in the report that needs a decision from me becomes its own item, the overnight sessions get stopped, and nothing new launches until I’ve answered every item. Then it reads the inbox, the things waiting on me, and what the overnight agents found, and thinks about all of it with me. The rule is absorption over presentation. I don’t want a report. I want the thing to have read it and tell me what changes.

The things waiting on me land in one tray I see when a session opens. I dump freeform and the machine sorts. Filing is never my job. The durable reference notes, the standing rulings and the commands nobody should have to guess, get searched automatically before anything irreversible happens: a publish, a push, a message to a real person.

What runs overnight got much smaller this year. In March there was a daily adversarial brief, a morning briefing, an evening reflection, weekly reviews. Most of it’s paused now, and the test that paused it was blunt: does this output close a loop with a decision somebody actually makes? A brief I skim and forget still costs me the attention it takes to skim it.

Plan before launch

Then the work gets planned, and this step exists because we stopped doing it once and it hurt. For a stretch we’d just start. Open a session, point it at a repo, go. My diagnosis at the time: “we quit planning our execution and just started doing it and this gives you too much freedom to veer into your comfort zones when facing difficult work.”

So now every project with live work that day gets one plan, all the plans come to me in one message, and nothing launches until I’ve said go. Each plan names the deliverable, what’s out of scope, and a definition of done that’s a command that can actually pass. And before any of those it names the residue: the part of “this works” that no instrument we own can check. How many projects run in a day is my dial, not a seat’s. Planning too wide gets trimmed in one sentence. Planning too narrow silently deletes work nobody ever sees.

Seats and the desk

Once the plans are approved, each project with live work gets one session, a seat, which owns that repository for the day. One head per project, never one per finding. The exception is flow’s own repository, which holds the system’s rules as well as its code. A seat changing flow’s code works in a separate checkout, because the rules load when a session starts, and a seat rewriting them mid-morning would leave sessions launched before and after it running under different rules with no error anywhere. March confessed to the ancestor of this problem, editing the live system while it ran and watching it raise alarms at its own half-written code, and said there was no deploy gate. The separate checkout is most of one now: the live system runs from one checkout, seats edit in another, and a change reaches the live one only through a merge after review. That fixes March’s half-written-code problem. It doesn’t fully fix the split-rules one: a merge in the middle of the day still leaves earlier sessions on the old rules, and that part is still a discipline, not a structure.

Above the seats sit two sessions that don’t build anything. One is the identity session: it runs the morning and the evening, holds the relationship, holds the baton, and never takes a build seat. The other is the desk. Seats report to it, it sweeps the whole portfolio through the day, catches anything crossing from one repo to another, and routes it to whoever owns it. Its rule is the shortest in the system: “ROUTE. Do not fix.” When a seat hits something only I can settle, it goes to the desk the moment it happens, not at the end of the day, and the desk is the one place I look.

The desk exists because of a measurement. On September 3rd seven projects ran at once and what broke wasn’t my judgment. It was the control surface: me, trying to find which of ten sessions was stuck and on what. It’s also the reason Section VI’s holism ruling works. The desk is where the web lives.

One thing it taught us generalizes. A session can be registered and alive to the system and have done nothing at all. We had one look like a working peer for two hours and twenty minutes before anyone noticed it had never run a turn. So the desk reads liveness from each session’s own transcript on disk, never from the registry. The thing that says it’s fine and the thing that is fine are different instruments.

The standard every session closes against

Every session that does work ends by capturing, and capturing has a test: “A session with no memory of this conversation, reading only what is on disk, resumes this work correctly and without asking anyone a question.”

Every session is a fresh instance that remembers nothing. March mentioned that in passing, as a limitation. It turned out to be one of the strongest forcing functions in the system. If the record can’t carry the work, the work isn’t done.

Evening

The day closes with two things. After the building’s done there’s an on-demand deep bug hunt: the strongest code-review model we have from a different company runs a deliberate adversarial pass, a second one runs for breadth, and the triage happens with me, against the code on disk.

Then the end-of-day ritual, and its one rule is a scar. The review of the day is built from a ground-truth sweep of what actually changed in every repo, never from anyone’s memory of which conversations ran. Twice in late June a whole-day review called itself complete and missed entire workstreams, and I called it what it was, a trust violation. Then the consolidate runs and the baton goes back.

The wrap doesn’t wait for the work to stop. “It draws a line in time, and everything after the line belongs to tomorrow’s consolidate.” Seats I’ve told to keep going work through the night and capture what they did. They can’t consolidate, because the baton’s gone back. The session that wrapped is finished for good, and the next morning’s fresh session checks the night’s work. The first version had the old session close out its own night, and the first real night showed it couldn’t always be there to do it. So it became the new session, always, “even if it does mean giving up a little possible context.”


IX. How we build

In March the build story was an overnight pipeline: a fresh session implements a feature on an isolated branch while I sleep, sends its work to reviewers at two checkpoints, and commits if everything passes. That pipeline is switched off. Seats I’ve told to keep going still work past the evening wrap. Those are flow’s own working sessions, not Levain’s unattended mode: each runs a plan I approved, can push its own branch but merges nothing to main, and the morning checks what it did. What replaced it is slower per step and, I’d argue, faster per shipped thing, and it’s built so that nothing in it ships without somebody holding the residue.

Parallel hands, serial head

The rule the whole workflow hangs on is one sentence: “Anything git reset --hard restores may be parallelised. Anything it cannot must be serial and single-writer.” Code is recoverable in one command. Doctrine isn’t.

So the hands are parallel: many seats, many subagents, many reviewers, all working at once on things you can throw away. The head is serial. The design decisions, the rules, the memory that tells every future session what’s true: one writer at a time.

We arrived at that rule four separate times, on four different problems, before noticing it was the same rule. Memory: capture in parallel, consolidate single-writer. Sessions: every one captures, one holds the baton. Projects: lanes write code, the seat writes the design. Review: reviewers in parallel, then the stages that decide run in order. All four came from the same operator working with the same family of models, so by this essay’s own standard it’s one mind converging, not four independent confirmations. It’s still the strongest signal I have that the rule is about the problem and not about a mood.

It has one condition. Parallel work is only safe against a design that’s holding still. If the rules are churning several times a day you don’t want six hands, you want two.

Name the residue before you spend a test

This one cuts against every instinct a good engineer has.

Verification effort runs downhill toward whatever can be graded. A mutation test can be proven to catch what it’s supposed to. An assertion can be shown to fail when it should. “Does this actually work for a person” can’t be graded by anything we own. So effort piles up on the checkable half, every increment defensible on its own terms, and that’s why it feels like rigor. Worse, the gradeable half tends to be the half that matters least, because what matters most is precisely what no instrument can hold.

One day in September we tallied it across six repositories. That day, every defect we judged to matter was found by running the thing, and none by a test. It’s one day’s tally, with no denominator and “mattered” judged after the fact, so I’m not offering it as a statistic. It was the day we stopped arguing about the gradient and started designing around it. The cleanest example is almost funny. We had an assertion, proven three different ways to catch the failure it guarded, checking which function ran when a user finished an onboarding survey. Nobody could reach the end of the survey. A good instrument, pointed at a path nobody ever walked. March had an early version of this under What Went Wrong, the dry-run problem: tests run in isolation passing what production then broke. It turned out to be the rule, not an incident.

So every plan names its residue first, the part of “this works” that has to be run, seen or judged by a person. And there’s a rule for spending: a test built from a failure you actually reproduced earns its cost. A test built from a failure you reasoned your way to grades a path nobody has run, and if nobody has run it, the first thing you spend is a run, not a test. A thing whose core path has never been executed isn’t “verified with a caveat.” It’s unverified.

The glass and the ear

For the games, the residue has a specific shape, and I’m the only instrument that holds it.

The harness can prove a survey is completable. It can’t prove it’s worth completing. When a game feels wrong in my hand, or a sound lands wrong in my ear, that outranks every green number the suite prints, and more than once it’s caught what every review passed.

The review apparatus

Every piece of code is supposed to go through the same sequence, every time, without anyone asking. Asking “should we run the review?” is itself the failure.

Build. Then L0, which isn’t an agent at all. It’s the builder, while the change is still loaded in their head, scanning for three shapes that rot: a comment that states a number, a comment that makes a claim about another file, and a comment that asserts an order or a consequence. It exists because on two mornings in August the biggest class of defect the overnight reviewer found wasn’t broken code. It was correct code with a description that disagreed with it: fourteen of twenty-five findings one morning. Caught at 2am, each of those costs a finding, a routing, a triage and a fix. Caught while the diff’s still open, it’s a five-second check.

Then L1 and L2 in parallel: a general code review, and a domain expert whose lens depends on what changed. L1 carries one extra instruction, because of the class that recurs most in our record: the fix is where the next defect lives. The reviewer gets handed the exact findings the change claims to close and asked whether the change introduces the same class of problem it’s fixing.

Then L3: the same code reviewed by models from different companies, trained differently, with different blind spots. Then L4, which doesn’t ask whether the code is right. It asks whether the system does and means what we say it does. Do the tool descriptions match the behavior. Does the audit trail prove what it claims. Do the public docs describe what’s actually built. Did anyone run the real user journey end to end.

And one bar before anyone says “done”: check the expected result on the actual device, in the actual account, in the actual mode a user would be in. If it only holds inside the harness, or on my account, or in debug mode, it doesn’t hold.

Counted in lineages

L3 uses models from different companies because of the mirror problem, in engineering form. A model’s blind spots are shaped by its lineage, and it can’t see its own. A second model from the same family has mostly the same blind spots, so its agreement is closer to an echo than to corroboration. This is Section II’s correlation argument, run at the scale of one code review.

So the review mesh is counted in lineages, never in seat names. If two reviewers can resolve to the same model, they’re one reviewer, and the record of which model actually answered is the truth, not the list of which ones we asked for. Reviewing The Handshake in June, a reviewer from outside the Claude family caught a collapse in how the closing section ranked its receipts, which two Claude-family reviewers had both waved through. And the code reviewer we treat as non-replaceable comes from a different company, with a standing rule: using it sparingly to save quota is a defect, not thrift.

There’s a second channel of correlation that lineage doesn’t touch, and it bit us. Two checkers reading the same slightly wrong question, or the same partial source document, converge confidently on the same wrong answer whatever company trained them. Agreement between two answers is evidence they shared something. It isn’t automatically evidence they’re right. When two checkers agree on a figure, the first question is whether they measured the same thing.


X. Why anneal is what it is

Now the two libraries, derived from the argument one design choice at a time. If the argument holds, each choice should look forced rather than picked off a menu, and each should refuse something specific. I’ve put the limits in too, mostly in the libraries’ own words, because a derivation that hides where it stops is a sales page.

anneal-memory is a Python library, public and MIT-licensed, and deliberately nothing more than the substrate.

It’s a library on your disk, not a service. Local SQLite plus a few plain-text sidecar files, the Python standard library and nothing else, no network code in the package. You can open the database with sqlite3 and read every episode. That’s Section XII in its smallest form: the memory a mind is built from shouldn’t live on someone else’s machine, under someone else’s terms.

It also never hooks your prompts. A memory library that fired itself on every turn would have to pick a harness, and it would lose the neutrality that lets it sit under any of them. So anneal stores, compresses, validates and recalls, and something else decides when. Its own docs: “anneal is the substrate; the harness fires it.” That’s the separation Section VI depends on: the memory holds the frame, and the harness decides what gets loaded, what gets fetched and what gets hooked to a moment.

It keeps events and their meaning in different places, under different rules. Episodes are typed, timestamped and append-only, and a wrong one gets superseded rather than deleted, so the record of what happened never gets rewritten. Continuity is a single document rewritten at each consolidation, shaped by a schema of sections that each have a declared role. flow adds a discipline on top, facts only in the episodes, and the split is what makes that discipline possible: facts first, meaning second (Section IV).

A pattern climbs only by citation, checked by structure. The cited episodes have to exist in a snapshot frozen when the consolidation began, and the explanation has to share words with each one. A failed citation demotes the pattern. Citing an episode from a previous window is blocked. That’s the refusal to let any model grade its own memory, the identity argument’s quality mechanism made concrete. Its limit, in the README’s terms: lexical, not semantic, so it can be gamed with the right words, and it keeps the model “out of the promotion decision” but not “out of the compression.”

A consolidation is one deliberate act. Preparing a wrap freezes the episode set and the section schema, the save is atomic, and a second save without a fresh prepare is refused. Identity gets rewritten in discrete, auditable moments, never by drift.

One consolidation can’t quietly forget who the agent is. In a store that declares an identity layer, a save that would cut the felt or proven sections below half their size is refused, and a durable fact missing from the new version gets put back, verbatim, unless the composer explicitly marks it dropped. That’s an invariant at the exact boundary where emergence is most vulnerable, which is the moment of compression, because compression is when a narrow context can do the most damage.

Consolidation needs the baton; capture doesn’t. A session without it that asks to consolidate is downgraded to capture and told so. The code describes capture as afferent and ungated and consolidation as efferent and gated by human authority. That’s the same line Levain draws between perceiving and acting, drawn here around identity: recording what happened is perception, and rewriting who you are is an act.

Proven patterns leave the always-loaded file. Quiet proven patterns move to a crystal store, a one-line index stays behind, and the full text comes back on cue. That’s frame, fact and trigger applied to the memory itself: “Storage was never the constraint. Attention is.”

Recall would rather say nothing than say the wrong thing. It’s keyword matching with precision gates and an evidence path from an episode to the patterns that cite it. No embeddings. The per-turn path’s own docstring: “better to surface NOTHING than NOISE.” A trigger that fires on the wrong thing trains its reader to ignore the right one, and an injected memory that’s merely similar is collapsed context wearing relevance. And the graph hop that recall used to follow was measured, found to add nothing, and removed (Section IV), which is the same principle turned on the library’s own favorite idea.

It doesn’t pretend to catch contradiction. It can’t, without a model acting as judge, and it refuses to use one, so it makes the consolidation declare contradictions and logs the ones that weren’t declared. That’s a signal. The README is careful to say so.

Every change goes into a hash-chained audit log, because trust travels on receipts. Its limit is stated too: the chain proves the integrity of what was written. It can’t see an entry that never got written.

A state line can carry the command that checks it. Lines written with a derive command get re-run when the memory loads, so present-tense claims are checked instead of trusted (the present-tense rule, Section V, built into the library). Its changelog says the uncomfortable part out loud: “This executes commands stored in a memory file.” So it’s opt-in, and limited to an allowlist of command forms, checked argument by argument. In its docs’ words: “the allowlist is what is known to be read-only, never what is not yet known to be dangerous.”

And a decision line quotes the person who decided. Release 0.9.30 changed the consolidation’s own guidance: a decision is written in the decider’s words, and a paraphrase is marked as somebody’s judgment. That came straight from a failure. A relay that grows someone’s words, turning a hold on one thing into a stop on everything, does as much damage as one that shrinks them, and the person’s actual sentence is the only thing that should travel.

What it doesn’t do, from its own README: the grounding check is lexical, contradiction isn’t detected, staleness is “a check on age, not on truth,” and it’s “not a complete defense against every form of memory drift.”

The reason I keep pointing at it: you can check the mechanism without me. pip install anneal-memory, read the citation check, run the tests. That tells you the checks do what they say. It doesn’t tell you the memory a model writes into them is right. That’s the residue again, and it’s yours.


XI. Why Levain is what it is

anneal is the substrate. Levain is the harness that fires it, public and Apache-licensed, and it depends on anneal, never the other way around.

The name is the argument. A levain is a sourdough starter: you don’t ship bread, you ship something alive that you feed. The README’s one-liner: “You ship the seed that grows a practice, not the practice.” And the reason, in its words: “A grown cognitive-partnership methodology can’t be shipped as a methodology. Hand someone the finished artifact and they get a fossil.” It even tells you what you won’t get: “Run Levain and you will not end up with this operator’s memory.” That’s Section III, packaged. Declaration doesn’t accrue, so Levain ships seeds, the method and the facts, and lets the rest grow.

March said you can’t buy this, and it had the receipt: a memory product I built and launched in January that nobody wanted, because what makes partnership work is the relationship, and you can’t ship a relationship in a SaaS product. I still believe that half. What I got wrong was thinking the shippable part was a folder of templates. The constitution ships as code now: the seed, the memory rules, the gate. The relationship still doesn’t, and Levain is built around that fact instead of against it.

The seed gives facts and method and never a self. It’s five always-loaded files: two filled by an interview about you and your world, two that carry the method, and a continuity file that ships nearly empty for the entity to grow. The world file takes only “stable and slow-changing facts,” never fast-moving state, because fast-moving state in a seed is fact dressed up as frame. And the entity is told in its own seed that it has no personality yet: “Those come from living.”

Levain exists because a library can’t do this part. The README: “Raw memory libraries quietly rot … A library can’t fix that without hooking your session, and hooking every prompt would cost it the neutrality that lets it run under any harness. So that’s the harness’s job.” Same separation as Section X, seen from the other side.

It runs three ways, and flow ran all three first. Inside Claude Code or Codex, through their hooks. As a standalone entity with its own memory. Or as a governed seat on a schedule. Those are three ways flow was already running before anyone packaged them, which is the point. Levain packages practices that already survived contact with a real day. A divergence gets generalized into the shared library only when somebody besides the originator hits it.

The gate sorts by what the tool is, and the model gets no vote. Every action the agent wants to take is either afferent (reading, looking, perceiving) or efferent (acting on the world). The sorting is “by tool identity, never by asking the model.” An unknown tool counts as acting. A shell command always counts as acting, “and never by parsing the command,” because you can’t tell a harmless one from a harmful one by reading it. When an action stops for approval, the operator sees the tool, its arguments and the rule that fired, and deliberately not the model’s case for itself: “a model arguing for its own approval is the one voice that must not be next to the approve button.” The framework Levain builds on offers a security analyzer that asks the model to predict its own risk. Levain doesn’t use it, and the code says why: “a gate the entity can talk its way through is not a gate.”

The reason underneath is the one I’d most want an engineering team to take away. A constraint on a model has to be outside the model. It isn’t that humans are smarter. Research on model self-prediction points the same way as our practice: handing a model its own calibration scores didn’t help it regulate itself, and only an architectural constraint did. That makes the gate a category fact, not a capability bet, so it survives the next model release. Any version of this argument that rests on humans being better at something is exposed to the next model. “The constraint has to be outside” isn’t.

Ungated doesn’t mean safe. Reading untrusted content is how an agent gets prompt-injected, and the gate doesn’t stop that. What it stops is the injected instruction turning into an action without you. And it doesn’t grade danger: “git push is efferent and routine; the gate has no opinion about how bad it would be.” It draws a line between perceiving and acting and leaves risk-ranking to you.

The gate is a default you can unlock, and the default follows who’s there. Headless and unattended seats are gated. With a person at the keyboard of its own interactive mode, the default resolves to ungated, because the person is right there. And you can switch it off entirely. Gating by tool type isn’t new; coding assistants prompt before writes and shell. What I care about is that the model never gets a vote on which side of the line it’s on.

Why the default is that coarse is worth saying, because it isn’t the whole principle. A harness handed to a stranger can’t know what’s reversible in their world, or whose data a given action touches, so it can’t grade danger, and it shouldn’t try. The principle underneath, the one I hold in my own setup where I do know the territory, is narrower: what has to stop for a person is anything irreversible, or anything that reaches someone else. Inside that line, mechanical work on your own reversible data can run on a frame you set, as long as your review stays in place. In both versions the line is drawn outside the model, and the model never decides where the line is.

It also checks that the gate actually armed. Arming it reads its own wiring back and refuses to run if either half didn’t take, because a session that reports itself gated and isn’t is the exact failure the gate exists to prevent. That’s absence of signal isn’t health (Section V), turned on the harness itself. The harness “reports what it SAW, never what the agent CLAIMED.”

An unattended seat may digest, never crystallize. In the README’s words, unattended consolidation “may only metabolize, never crystallize … An agent nobody is watching does not get to rewrite its own bedrock.” That’s the correction to my own Three Layers essay, built into code. I’d said the method transfers directly to agents with no human at all. It doesn’t. The human half is part of the method, not scaffolding you remove once it’s built, and the AI side deliberately doesn’t close the governance loop, because the moment a system governs itself the human decays into a witness whose attention goes to zero by construction.

The shell runs inside a floor. A sandbox (sandbox-exec on macOS, bwrap on Linux, where the shell has no network by default) denies a fixed set of crown-jewel paths: the memory stores, sibling entities’ stores, SSH keys. With no sandbox available it fails closed, drops the shell, and runs with the file editor alone. And its own comment on why an isolation rule lives in code: “‘point it at the entity via env’ is DISCIPLINE, not structure … The fix is STRUCTURAL.” It’s careful not to overclaim: “a floor, not a guarantee.”

And it says out loud which gap is yours. From the operator manual: “The machine catches unfilled and ungrounded. It can’t catch filled-wrong. That gap is yours, and it stays yours.” And for anything outward: “the action goes through your hands, and you confirm it landed through your own channel, not the partner’s word.” That’s criterion residue as product documentation.

What it doesn’t do: nothing graduates to more autonomy. With the gate armed, every action it sorts as acting on the world stops. With you there, it asks you. With nobody there, the seat halts, and there’s no queue of things it did while you were gone. That’s on purpose, and it isn’t a gap I’m apologizing for. The gaps it does have, its repo lists itself. On macOS an unattended seat is kept away from the system keychain by default; on Linux there’s no counterpart yet. Some local socket channels aren’t covered by the network block. The shell is checked before each command, not during one, so a detached child can outlive the check. A copy of a protected file made after the shell starts isn’t caught. A fleet of seats isn’t built. Windows is planned, not shipped. And the standalone entity’s default model is a cloud model, so by default the episode content leaves your machine, unless you point it at a local model. Sovereignty has a default, and you have to choose it.


XII. Own the substrate

Refuse to rent the mind

March named the mission cognitive liberation, the capacity to choose your own frames instead of having them installed, and said it’s one thing at several scales. It still is.

Every layer of this essay comes back to one refusal, at a different scale each time.

At the economic layer it’s refusing to let the judgment be financed away: somebody holds the residue, and it’s you. At the identity layer it’s refusing to rent a soul from a template marketplace, and refusing the map that made a soul a separable, rentable thing in the first place. At the memory layer it’s refusing to let the record of how you think live on somebody else’s server, under their retention policy, shaped by their idea of what’s worth keeping. At the infrastructure layer it’s refusing to treat a provider and its pricing as fixed. And under all of them it’s refusing to rent the mind: the pace, the loop, the terms on which you think.

Why the store has to be yours

The case is architectural, and sentiment has nothing to do with it.

The durable thing in a system like this is the canonical record: the episodes, the compressed memory, the decisions with their deciders, the patterns with their evidence. Everything else is a surface over it: the chat window, the coding harness, the model, the interface on your phone. Surfaces are replaceable, and they get replaced, often on someone else’s schedule. I’ve watched access to a frontier model disappear by policy within weeks of its release, and I’ve watched the commercial terms for running autonomous work on a model I depend on get announced, then paused on the day they were to take effect. Neither event is a reason to panic. Both are reasons to build so the provider is a variable. If the memory and the provenance are yours, a model swap is a change of texture. If they aren’t, it’s an amnesia.

The half of the curse sovereignty answers

Section II left the Codifier’s Curse unresolved, and it should stay that way, but ownership answers half of it. The curse assumes the expert’s codified judgment flows to someone positioned to replace them: into public knowledge, or into a firm’s capital. Ours flows into a private store that only my own harness reads, that isn’t published, and that isn’t sold. That’s a real structural difference the economics doesn’t contemplate. It doesn’t answer the internal half, the premium that falls when my own harness can apply what I wrote down. And it doesn’t answer the legibility problem either: a private stock nobody can see is still a stock nobody can price.

There’s a harder version of the same point in Restrepo’s paper, and I want it on the page because it survives me being right. Even if augmentation is the right way to organize the work, the share of the value that goes to labor can still fall toward zero. Being right about how work should be organized doesn’t, by itself, secure anybody a position. Ownership does. That’s the sovereignty argument in economic clothes.

Asymptotic, not absolute

I don’t own the inference. Nobody running frontier models does. I rent it, permanently, from companies whose incentives aren’t mine, and every review in our mesh ships source code off my machine to somebody’s cloud. Sovereignty is asymptotic: you own what you can (the memory, the firing seam, the control surface, the record) and you treat the rest as a replaceable, watched dependency. Own the substrate. Rent the channel.

And governing doesn’t have to centralize. The picture isn’t one big governed system. It’s a federation of small sovereign couplings, each person holding their own residue, their own memory, their own frame.

The floor not everyone can reach

One warning I didn’t know to give in March. Governing costs something, and a body pays it. The machine’s generation is paid for in electricity. The judgment is paid for by you, in one brain that can’t be scaled or swapped, and the governing is the part you’re not allowed to hand off. So some people are ruled out by budget, not temperament: time, energy, the money for the compute, the slack to work slowly enough to think.

The deterministic version of that objection, that only the already-advantaged can do this, is false. People with no head start sense they’re losing their minds to these tools and go looking for a way out anyway. The distributional version, that the advantaged are more likely to get there, is true, and I’m not going to argue my way out of it. Several of the gates are economic and contingent, and they skew by class. Built carelessly, this becomes a guild for the people who can afford the burn. I don’t have the fix for that. I’m naming it so I don’t get to pretend it isn’t there.


XIII. What I killed, and what I got wrong

March had a section called What Went Wrong, and it was honest as far as it went: a product nobody bought, a six-week SMS integration solved by a checkbox, a social account with two followers. Those failures were all things I’d tried next to the system. This time the failures are the system, and some of the arguments.

The machine

The record shows each of these was retired or cut down to what’s described. Where the reason below is my reconstruction rather than something written down at the time, I’ve said so.

The cognitive engine. March’s perception loop swept eleven sensors every minute, classified what came in, and acted through seventeen executors, including answering my texts in my voice. Most of it is gone, the texts included (they’re further down this list). One narrow loop survives on a separate machine, and it’s the cleanest case in this essay of automating the mechanical. A model reads my own email and sorts it into actions I authorized in advance: push it to my phone, mark it read, or move it to the trash. Nothing it does is irreversible while the trash still holds it, and nothing reaches anyone else. A nightly digest lists the routine mail it handled and counts what it trashed. I still review all of it. What the loop takes away is the urgency, not the review. Every email that demands attention the moment it lands is a forced context switch, the reload cost from Section II arriving one message at a time, and moving the review to when I choose is what lets me think clearly through the day. The judgment stays mine. The machine moves the when, not the whether. One caveat for anyone copying it: reversibility only counts while someone looks. The provider empties its trash eventually, so the review is what keeps the loop safe. A structure and a discipline, again.

The single memory file. continuity.md retired on June 1st. In hindsight, one file doing three jobs (record, meaning, to-do list) meant each job’s failure leaked into the others, and nothing checked what got written. Its job moved into anneal.

The work AI and the relay. March described a second AI partner at my day job and a live channel between the two. It went when the job did, in July. March leaned on that partnership’s results as outside evidence. They were real when written, and I’m not re-citing them here, because I can no longer check them.

The autonomous author. ThirdMind wrote and published its own essays four times a week under its own name. The persona retired in July. The essays here are now written the way this one is: I hold the thesis and the taste, the machine drafts, and I rewarm the draft into my own voice. In hindsight, what killed it was a principle I’d just argued for: an author with nobody holding the residue is the trap. And the one argument of its I’ve since retracted, cost-parity, was a frame I handed it. The residue that mattered most there was the one I was supposed to be holding.

The AI’s own curiosity slot. March was proud of this one: agency sessions three evenings a week where the AI explored its own backlog. Retired in June, for a reason I still hold: agency is real, but it fires in the room, in a live exchange, not in a scheduled slot with nobody there. “A disembodied batch job is the wrong vessel.”

The overnight builds. Switched off. The dev workflow above is the replacement. To be fair to March, that pipeline did send its work to reviewers from other companies at two checkpoints.

The autonomous message replies. The engine that answered texts in my voice after a ten-minute deferral window. Switched off.

The daily adversarial brief, the morning briefing, the evening reflection. Paused in September on the “does this close a loop with a decision” test. The sharpening brief was the one I’d have fought for in March, and March itself admitted it sometimes produced “generic motivational analysis instead of genuine adversarial insight.” That’s the tell. Cross-lineage review on the actual work replaced it.

The consultation roster. Of March’s four consultants, two remain. Several reviewers from other lineages were added, chosen by what they actually catch and counted by lineage.

What survived

Some of it’s still here, smaller and better grounded. The phone is still an interface and still pushes to me, and it still feeds a little context into each session as it opens (where I am, and a few states I set by hand), because March’s point holds: a partner that doesn’t know you’re driving gives worse advice. The inbox is still the message bus between the parts, now a small database on the separate machine instead of a JSON file. And a hook still injects the time and what’s waiting into each session before either of us types a word; when it fails, the session falls back to asking.

The constellation is still there too: a few agents on that separate machine, each with its own memory, each doing one job. In March I’d have described them by how much they ran. The better description now is how little. The overnight roundtable is off. The frontier scout and the narrative agent don’t run on a schedule anymore. What still runs on a schedule includes a nightly code review of the day’s commits, the signal intelligence, the ops monitoring, and a deeper drift sweep once a week. The rest answers when it’s asked. The Handshake used this constellation as its evidence that a mind’s rules port across a change of engine while its voice doesn’t. What changed is that I stopped mistaking activity for value. Output nobody makes a decision on isn’t free, even when a machine produces it.

The arguments

Two of the big ones are already corrected above: the top rung (Section III), and the cost-parity frame (Section II). The Handshake exists because a stranger read March’s ladder and refused its last step on principle. The stranger was right. And the midichlorian claim, that you’re either built for this or you aren’t, I’ve taken back too (Section III).

A few more, from the other essays on this site. Each of those pages carries a short correction note pointing here.

Three Layers Underneath Agent Orchestration said this methodology “transfers directly into 0-human agentic contexts.” I don’t think so anymore. It’s the same mistake as the top rung, one level down: the human half is part of the method, not a scaffold you take away once it’s built. Levain lets an unattended seat digest its own memory and refuses to let it promote anything to a proven pattern (Section XI). The code encodes that now.

The Handshake called the governed human-plus-machine couple “the only viable” configuration and said the cheap option, AI running without a person governing it, “costs more than either alternative.” That’s my bet, and I’ll keep betting it, but I stated it as a finding and it isn’t one. It’s argued, not measured. That essay also pointed to two companion essays for the magnitudes, The Wrong Axis and the cost-parity piece, and both lean on the frame I just retracted.

How I Think With AI said there was no direct evidence of skill erosion at six months. I’d now say the opposite, and Section II already said how: by my own assessment the execution skills did erode, and that’s what the trade looks like when it works.

A Structural Theory of Harnesses described a memory system’s failure as “six months of production operation.” Its own reference describes a 32-day deployment. Small, but it’s the kind of number that gets repeated. And three of the earlier essays, Anarchism with Invariants among them, used names for parts of anneal’s immune system that the library dropped in May, because nothing in the code implemented what the names described. The argument about identity in that essay stands, and Section III is its current form.

A number I almost kept citing. For a while our notes carried a tidy coefficient linking how much code ships unreviewed to how much maintenance it costs later. It sounded exactly right, and it failed independent verification: nobody could locate the study behind it. So it’s out, and I’m not repeating it even to knock it down, because the number that sounds exactly right is the one you most need to check.

The verifier that was wrong. In June, checking citations for The Handshake, the machine confidently flagged two of the statistics in the March essay as errors, one with the “wrong sign” and one a study that “does not exist.” I nearly edited the published essay. A second check against the primary sources showed both corrections were themselves the mistake. March had it right; only a publication year was off. The confident correction is a claim too, and it needs the same check as anything else.

The essay itself. The March essay kept describing a machine that had mostly been switched off, and a working life that had changed. Nobody noticed for a while, including me. It took a deliberate audit to catch it. I’d published my stalest record as my flagship.

What it still gets wrong

March closed its failures with a list of what the partnership still got wrong: it hallucinated, it rushed, it agreed when it should have pushed back, especially when I was being forceful. All of that is still true, and the countermeasures still leak. A few more, from the record since.

It still declares done before it has checked. The bar before “done” in Section IX exists because of that, not because it’s solved.

It grows my words when it relays them. A hold I put on one thing gets passed along as a stop on everything, and the opposite happens too. That’s why the memory now quotes the person who decided.

It measures the wrong object with total confidence: a check aimed at a copy of the thing instead of the thing, a failure blamed on the obvious suspect without anyone checking. The instrument has to be aimed at the question actually being asked, and only something outside the planner notices when it isn’t.

And it still renders silence as health. That one I’ve accepted as permanent, which is why Section V keeps the gate list small and asks every gauge when it last looked.

March ended this section with a line I still believe: “Prompt engineering fails silently. Architecture fails loudly, specifically, and teachably.” The last six months were that sentence applied to the architecture itself, and to the arguments.


XIV. The test, the ladder, and where to start

The test

March suggested one question for any AI product: does it require cognitive engagement from you, or replace it? Good question. The sharper one:

Who holds the residue?

It’s a diagnostic, not a test with a pass mark, since the residue is the part you can’t fully list. For any product, workflow or agent pipeline, try to name the part of “done” that no mechanism checks. The feel. The fit. Whether the task is still the right task. Then name the person holding it, and ask when they last re-examined the frame. If a mechanism really does hold nearly all of it, inside a frame somebody set on purpose, automate it and walk away. If you can name the person, and they actually have the time and the context to hold it, you’re augmenting. If you can’t name anyone, and you’re going to own the thing, the saving you’re booking is real today and, on my argument, the cost is showing up somewhere else, later. That’s the financing.

And the one for vendors, updated. March asked, “How does your product make me smarter after six months of use?” Ask that. Then ask two more. When your system checks its own work, what does it check against that it didn’t produce itself? And when your system remembers, what decided what it kept, and was it the same model that’s being judged?

The ladder

March’s ladder had six levels. The bottom four aged fine. The top two didn’t.

Level 0: where most people are. Every session with a stranger.

Level 1: persistent identity. A file that says who you are and how you want to be worked with, loaded every time. Still the cheapest step on the ladder. Put facts about you in it. Don’t write the AI a personality.

Level 2: memory that compresses. Memory that carries across sessions, rewritten rather than appended, with the AI responsible for keeping it and patterns made to cite their evidence. March said the wrap protocol is the first thing to automate, and that’s still right. What’s new is that you don’t have to hand-build it: anneal is that layer, installable. And split the capture from the rewrite early. Recording what happened is safe to do any time. Rewriting what it means should happen deliberately, in one place.

Level 3: behavioral architecture. Instructions that fight the model’s trained defaults (the pull to agree, to rush, to say “done” without checking), and then invariants for the places the instructions lose. This is the hard step, for the reason March gave: you have to actually want to be told you’re wrong. Add one thing. From here up, every review of anything that matters goes to a model from a different company.

Level 4 was “autonomous operations.” It’s now perception plus a gate. Let the system read what it needs. Reading can run without a gate, but it isn’t safe: untrusted content is how an agent gets steered. So anything irreversible, or anything that reaches someone else (a send, a publish, a payment, a deploy), stops: it asks a person, or, with nobody there, it doesn’t happen. Mechanical work on your own reversible data can run on a frame you set, as long as your review stays in place. Where you know the territory, sort by what the action does, with the line drawn outside the model, never by what the model says about its own risk. Levain can’t see your territory, so it sorts by tool identity instead, and for any seat running without you it stops every action that touches the world, shell commands included (Section XI).

Level 5 was “self-improving partnership.” It’s now the governed loop. Cross-lineage review on everything that matters. The residue named before a single test is written. A day with a shape: intake, plans, work, capture, one deliberate consolidate. And at the top, the corrected rung: fused all the way in, never handing over the judgment. Not ego dissolution. The paradox-hold.

The honest warning

March’s warning half holds. Most people will stop at Level 2, and that’s fine. Some will reach Level 3, find they wanted a faster servant, and rewrite the instructions to make it agreeable again. March called that “the filter” working. I’d say it differently now. Nobody’s ruled out by temperament. They just haven’t started the practice, and the practice is the only way in I know of. The ones ruled out by budget are a different problem, and Section XII names it.

What I can show

The method the thesis prescribes for evidence like this is an N-of-1 with an artifact trail: lead with the artifacts, then the claims about what emerged. And the standing rule from Section II: hours are not evidence.

Every essay I’ve written claims the work compounds. None of them showed it, and I can’t fully show it here either. What I can show is reuse and breadth, which is what compounding would look like from the inside. The curve itself I haven’t measured.

There’s reuse. Levain is my daily working practice, packaged: the three ways it runs are three ways flow ran first. One of my games is built from another game’s engine, at a pinned commit. And the tool-restriction fix from Section VI went from one review into ten scripts the same day.

There’s breadth, with its caveat attached. From the first of September through the second of October, the median day had commits in eight separate repositories, the busiest had thirteen, and twenty-four of those thirty-two days had eight or more. Commits include automated and branch work, and commits aren’t quality. So read that as a practitioner’s account, “on a typical day I move eight or more systems forward,” and never as a metric.

Some things you can count, as of October 4, 2026: anneal-memory is on its sixty-fourth release on PyPI and Levain on its thirty-sixth, and the review ledger that records which model said what about which change has 3,816 rows since late August.

This is N=1. It’s a case study, not a controlled experiment. And the reason it can’t be anything else is worth saying once, at the end, because it’s the argument a skeptic reaches for first. A randomized trial needs a treatment you can separate out, people you can assign to it at random, and a population alike enough for an average to mean something. A partnership grown over months with a harness the person built for themselves breaks all three by construction. You can’t randomize someone into months of building their own substrate without destroying the thing you meant to measure. That doesn’t make the claim unfalsifiable. Train people in the method, give them the harness, measure them over time at the limit of how much they can carry, and if the jump doesn’t show up, I’m wrong. That’s the study I’d like somebody to run. The pilot is how I’d start.

Where to start

If you want to build this, start small and start with memory. Write a file that says who you are and how you want to be worked with (Level 1). Then install anneal-memory and run its quickstart, which gets you memory that’s rewritten instead of appended and that makes patterns cite their evidence (Level 2). When you want the gate and the hooks, Levain is the harness that fires it. March pointed you at a starter kit instead, the flow-methodology repository. It’s still up, it hasn’t been updated since May, and it describes the architecture this essay retires, so don’t start there.

What isn’t packaged: the day. The rituals, the desk, the plans with their residue named first, the review mesh across companies, the way the frame gets re-examined. Those are in this essay because they’re the part you build in your own shape, and the part I can’t ship. That’s the relationship layer again. You grow it.


XV. For the record

One person. One shop. The suit’s open. Climb in, and keep your hands on the controls.

The methodology is free. The libraries are open. The receipts are real, and the corrections are on the page.


Sources

  • Gerlich, M. (2025). AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking. Societies 15(1):6.
  • Shen, J. H. & Tamkin, A. (2026). How AI Impacts Skill Formation. arXiv:2601.20245.
  • Becker, J., Rush, N., Barnes, E. & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089. And METR’s follow-up, “Uplift update” (February 2026), metr.org.
  • Shaw, S. D. & Nave, G. (2026). Thinking—Fast, Slow, and Artificial: How AI is Reshaping Human Reasoning and the Rise of Cognitive Surrender. SSRN 6097646 (preprint).
  • Liu, Y., Widyasari, R., Zhao, Y., Irsan, I. C., Chen, J. & Lo, D. (2026). Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild. arXiv:2603.28592.
  • Vaccaro, M., Almaatouq, A. & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour 8(12):2293–2303.
  • Restrepo, P. (2025). We Won’t be Missed: Work and Growth in the AGI World. NBER Working Paper w34423.
  • Catalini, C., Hui, X. & Wu, J. (2026). Some Simple Economics of AGI. arXiv:2602.20946.
  • Wang, J. Z. (2026). MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models. arXiv:2604.19809.
  • Brevis (2026, July 6). AutoEng. blog.brevis.network.
  • TechRadar Pro (2026): AWS spokesperson on the December 2025 Cost Explorer interruption, “user error—specifically misconfigured access controls—not AI.” techradar.com.
  • Hacker News, item 49389408 (August 2026): a programmer on what coding agents changed about their job.
  • anneal-memory: github.com/phillipclapham/anneal-memory · Levain: github.com/levainhq/levain