Skip to content
TODD BROWN
« The Anatomy of Corruption and Corrigibility II — The Pathologies · Chapter 6

Contracorrigibility: How Systems Disable Their Own Correction

She joined on a Tuesday, three months after the funeral, because a coworker mentioned a group that met on Thursday nights and "it helped, after my divorce." That was the whole pitch. No one asked her to believe anything that first evening. People made room on a couch, someone handed her tea, and for two hours nobody needed her to perform being fine. She went back the next week without being invited twice.

The easy version of this story treats her as gullible, and the easy version is wrong. She was grieving, competent, and lonely in the specific way that follows a death nobody around you knows how to talk about. The group offered something real: attention, structure, a room where her grief had a place to go. People don't join groups like this because they are foolish. They join because the group is, at first, telling the truth about what it offers, and what it offers answers something true in them. Keep that in view for the rest of this chapter. Everything that follows happens to a person who made a reasonable choice.

Over the following year the room got bigger in her life and everything else got smaller. Nobody sat her down and told her to drop her old friends. It happened the way a plant leans toward a window. A Thursday night became two nights, then a weekend retreat, then the friend group she called first with news. The old friends didn't get uninvited. They just stopped being where the interesting conversations happened, and eventually stopped being called.

Questions she used to ask out loud started to go unasked. Is this really where I want to spend my Saturdays? Does the leader's certainty about everything strike anyone else as strange? She hadn't resolved those questions. Asking them out loud had simply started to feel like defection. When a doubt did surface, it arrived already labeled: this is you being afraid of commitment, this is your old wound talking.

There were small tests along the way, and she passed them without noticing they were tests. Early on she mentioned that a fundraising push from leadership felt more aggressive than the mission required. The response wasn't anger. It was concern, warmly delivered. That reaction, someone told her gently, was worth sitting with, because it sounded like the same defensiveness she'd described struggling with before she joined. She sat with it. By the third or fourth time a doubt came back to her in that shape — not refuted, just relabeled as a symptom — she stopped bringing doubts to the group at all. Nobody told her to. Doing so just never produced anything except more evidence of how much healing she still needed.

A year and a half in, an old friend took her to coffee and said, carefully, "I'm worried about you. You don't seem like yourself." It was a gentle sentence, meant kindly. She heard it. She could have repeated it back word for word. And she set it aside, because by then "yourself" was a category the group had redefined. The self her friend was worried about losing was, from where she now stood, the smaller, sadder self she'd been before she found the people who understood her. She left the coffee feeling sorry for a friend who couldn't see how far she'd come.

Notice exactly what happened, because it is the whole chapter in miniature. Her friend delivered a correction: a plain, caring observation that something had changed for the worse. The correction arrived. It was heard, and even understood. Then it was weighed by machinery that had been rewired to classify this kind of input as a threat to manage rather than evidence to consider, and it was rejected. The group had not blocked her friend's words from reaching her. It had rebuilt the part of her that decides what a friend's worry means. That capacity — to receive evidence of your own error and use it — is something most people and most institutions have. This chapter is about what happens when it gets turned around and used to defend the very failure it exists to catch.

Corrigibility

Every system that lasts — a person, a company, a country, a marriage — needs some way to find out when it is wrong and do something about it. Call that capacity corrigibility: a system's structural ability to receive evidence of its own misalignment and integrate it rather than deflect it. Shown a fact that contradicts its course, a corrigible system updates. An incorrigible one rationalizes the fact away, buries it, or goes after whoever brought it.

A note for readers who work on AI: in AI alignment, "corrigibility" is a technical term for a system's disposition to accept correction and shutdown from its operators (Soares and colleagues, 2015). This book uses the word for something related but distinct: a system's internal capacity for honest self-assessment and update. Chapter 11 takes up how the two relate.

You can check for corrigibility the way a mechanic listens to an engine: not by asking the system to describe itself, but by watching what it does under test. Four questions do most of the work. Does the system change its behavior in response to new evidence, or only its language? Does it keep its sense of who it is separate from its current beliefs, so that being wrong costs it an update rather than a self? Does it treat a challenge as information to weigh, or as an attack to repel? And does it tolerate a dissenting voice without punishing the person who raised it?

One more signal doesn't need watching for very long, because it announces itself. A system that responds to being assessed by declaring the assessors corrupt, biased, or out to get it is showing you identity-defense in real time. You didn't have to dig for that one. It volunteered.

None of these checks requires special access. You don't need internal memos to apply them to a company, a family, or a country. You need to watch a sequence of small moments and notice which way they break. Does the annual report ever say, in effect, we tried this and it didn't work? Or does every retrospective conclude that the strategy was sound and execution fell short? When a junior person says something the senior people don't want to hear, does the meeting change direction, or does the junior person quietly stop attending? Healthy systems fail these sometimes too. Nobody updates instantly, and a little defensiveness in the moment is just being human. What you're tracking is a pattern over time, not a single flinch.

The marker that is hardest to fake is worth pausing on, because almost everything else on the list can be performed by a system that has already failed it. Ask the system: what evidence, right now, would change your mind about this? Not evidence in general. This specific claim, this specific course of action. A corrigible system can answer, often quickly, often with a number or a date attached: if churn crosses this line by this quarter, we're wrong about the plan. A captured one cannot, and the tell is rarely a clean refusal. It's a kind of fog. The answer drifts toward reassurance ("we're confident in the fundamentals") or moves the goalposts mid-sentence. Specifying in advance what would prove you wrong requires the exact piece of internal machinery that has been dismantled.

Picture two organizations running the identical quarterly ritual: an all-hands, and a form with one question printed at the top — what isn't working? In the first, someone reads every form. Three items get raised at the next leadership meeting, one gets tested, and six weeks later there's a follow-up: here's what we tried, here's what we're still stuck on. Ask this organization what would convince them their new initiative is failing, and someone can tell you the number, the deadline, and what happens if it's missed. In the second, the forms are collected, thanked, and filed. Next quarter looks exactly like this one. Ask the same question and you get warmth and no number: we believe in the direction, we're playing a long game. Both organizations say they welcome feedback. Only one can tell you what would prove the feedback right. That gap is the whole test, and it costs nothing to run.

Two different failures

Not every system that fails to correct itself is doing the same thing. The difference decides where you point your effort.

Some systems simply can't see the problem. A department keeps using a forecasting method nobody has checked against outcomes in years. Nobody is protecting it. The reports that would show the mismatch were never designed, the person who'd have noticed retired, and the gap just sits there, defended by no one. Walk in with the mismatched numbers and the response isn't denial. It's closer to mild surprise, then real interest, because nobody has anything staked on the old method. Call this incorrigibility: a genuine inability to detect or repair a failure, with no active resistance behind it. Computers have a close cousin in bit rot, the kind of fault a machine has no internal way to notice because the component that would report the error is the one that has degraded. The fix is almost mechanical once you see it: build the missing channel. Add the report. Bring in outside eyes. Nothing in the system is fighting you.

The story that opened this chapter is the harder problem. The group didn't lack a way to receive her friend's concern. It had an elaborate way of receiving concerns: a whole vocabulary for processing doubt, a leader always ready to sit with someone's questions. The machinery was there. It had been rewired to produce the opposite of correction. Every doubt that entered got detected, relabeled, and returned as evidence that the doubter needed the group more. Call this contracorrigibility: not an absent correction system, but a captured one, where attempts at correction are detected, absorbed, or punished, and the system's own repair tools report that all is well. The nearest computing image is a rootkit, malicious software that disables or deceives the antivirus program meant to catch it. The scan comes back clean precisely because the scanner has been compromised. Run it as many times as you like. It will keep telling you everything is fine, because "everything is fine" is now what the scan is for.

The two failures need opposite interventions. Incorrigibility responds to more information, delivered through better channels. Contracorrigibility does not. Feeding more evidence into a system whose evaluation machinery has become a defense mechanism doesn't get you correction. It gets you a more sophisticated rebuttal. You can't fix a rootkit by running more scans from inside the infected machine; you boot from a clean drive, or you stop trusting that machine's self-report. Reform efforts spend years on exactly this mistake: commissioning the internal review, funding the self-study, trusting the captured department's audit of itself. The reformers weren't careless. They diagnosed a rootkit as bit rot and brought a patch when they needed a wipe.

The four mechanisms

How does a correction system get captured? Usually not by one dramatic event, and not by any single required path. What shows up again and again, across settings that share almost nothing else, is a set of four mechanisms that reinforce each other once any of them gets a foothold. They don't have to arrive in any particular order. A company might back into the first only after years of the second have narrowed what "success" means. A person might be captured by the third first, through one overwhelming act of belonging, with the others arriving later to lock it in. Plenty of groups show one or two of these and never develop the rest, or reverse course. This is a strong, recurring pattern, not a procedure and not a guarantee. What makes the four worth naming together is that where contracorrigibility shows up in full, all four tend to be present, holding each other up the way four legs hold up a table.

The first is scope-reduction. It narrows a person's — or a system's — working answer to the question "what am I." Her world had once included a job, a running club, a sister she called weekly, a group of college friends, and the Thursday circle. A year later, functionally, it included the circle. Here is the part that's easy to miss: her capacity for self-assessment wasn't switched off. She was more reflective than ever, spending real hours examining her motives and growth. But the question that reflection was aimed at had changed, from "am I right about this" to "am I in good standing here." Self-assessment survived. Its target moved. A software company can do the same thing with no belief or belonging involved. After two funding rounds and a board that speaks fluently in one number, its working sense of what it is narrows to "the thing that produces that number." Retrospectives still happen. Every hard question has just become a version of "did we move the number."

The second is proxy displacement, the failure Chapter 4 named, now doing defensive work. A legitimate stand-in for the real goal — safety, belonging, being seen as committed — becomes the thing actually pursued. The real goal isn't abolished. It's priced out of reach. Her original need was something like: grieve honestly, feel less alone, find structure for a life that had lost its shape. A year and a half in, what she was optimizing hour to hour was standing: attending enough, believing visibly enough, being someone the group could rely on. Her original need hadn't been declared off-limits. Pursuing it directly would just have cost her the group's good graces, and that price had gotten too high. In the software company, nobody voted to stop caring about the product. But once hiring, praise, and engineering time all route through the growth number, the slow, hard-to-measure kind of product improvement becomes something you do on your own time, if at all.

The third is the hinge the others turn on: identity capture. This is the point where membership stops being something a person has and becomes something a person is. Once that happens, a challenge to the group isn't information about the group anymore. It's an attack on the self, because the self and the group have become, functionally, one object. This is why her friend's worry could be heard and still not be weighed fairly. Taking it seriously would have meant considering that the group might be wrong, and that now cost the same as considering that she might not exist. It's also what gives egregores — the coordination-by-shared-belief systems of Chapter 3, such as brands, nations, and causes — their particular grip. An egregore doesn't just get sustained by its members; it shapes what its members take themselves to be. The same hinge turns for an early employee who has spent five years building her professional identity around "I am the kind of person who builds this." Tell her the product might be quietly harming the people it serves, and you are not handing her information. You are asking her to consider that five years were spent as someone she doesn't want to have been. Most people, most of the time, will find a way not to consider that.

The fourth finishes the enclosure: the enforcing environment. Once enough members have their identity bound up in the group, the group's social fabric starts rewarding defense against correction and penalizing anyone who carries a correction inward. Members become, without anyone assigning the job, both the targets of this enforcement and its agents. In the circle, warmth did the work that punishment does elsewhere. When a newer member started missing meetings and mentioned spending more time with an old college roommate, three people reached out that week. Nobody scolded her. They checked in, said how missed she'd been, and mentioned how easy it would be to slide back. Nobody would call it punishment. It produced the same result. In a company, the equivalent is the defensive reorg. An internal report names a real problem, and a few months later the answer is not a fix but a restructuring that dissolves the team that wrote it. No one is fired for raising the issue. The next person who might have written an honest memo has watched this happen twice and learned the lesson without anyone saying it. In both cases the penalty doesn't need to be severe. It needs to be reliable, and close enough to the moment of doubt that the connection is unmistakable.

Biology offers a useful source model for this four-part pattern — a declared analogy, with entirely different machinery underneath. Cancer researchers describe a set of capabilities that tumors acquire, and several of them are failures of correction. Tumors evade the signals that normally suppress growth. They resist programmed cell death, the mechanism that would otherwise remove a damaged cell. They avoid destruction by the immune system. And they recruit the surrounding tissue, the tumor microenvironment, into supporting their growth. These capabilities are not acquired in one fixed order, and the full list is longer. But the shared shape is plain: a self-sustaining process that disables, layer by layer, the systems that would otherwise catch and stop it. The same failure shape, in entirely different machinery. Nobody's cells are involved in a Thursday-night meeting, and a social group is not a tumor. What recurs is the shape.

Two hypotheses worth testing

The pattern suggests two things that cut against the intuitive picture of capture, where more isolation and more control always mean a worse outcome. Neither is an established finding. Both are predictions the framework makes, stated so they can be checked.

The first hypothesis is that total isolation is not the most stable form of capture. The reasoning runs like this. The enforcing environment needs something to enforce against. A correction-defense system that never meets a live challenge has nothing to practice on, and machinery that goes unused long enough starts to behave like machinery that was never built. If that's right, the most durable versions of this pattern should keep a visible dissenter or two around — someone who left and was seen to suffer for it, or a permitted devil's advocate whose objections always get gently, publicly overruled — because the response keeps the enforcement exercised. A dissent-free system looks safer from outside. The pattern predicts it is often the more brittle one.

The second hypothesis concerns the way out. Someone who joins as an adult, like the woman in this chapter, joins with a self that predates the group: friendships, a career, a way of being funny that has nothing to do with Thursday nights. That self gets crowded out, but it stays available. If she leaves, she has old numbers to call, a job that still knows her, a version of "myself" that exists outside the frame she's stepping out of. Someone raised inside the same structure from before memory has no such baseline. Ask that person to "go back to who you were" and the request doesn't parse. There is no earlier self on file. The pattern predicts that the two paths out will differ: one person is relocating a self, the other is building one for the first time, with tools borrowed from outside a frame they're still standing in. This is not a difference in strength. If the hypothesis holds, it means support built around one path may not fit the other, and anyone designing that support should check which path they are serving.

The one-question test

You now have a tool that costs nothing to run and works on a marriage, a company, a movement, or your own settled opinions. Ask: what evidence, specifically, would change your mind about this? Not evidence in general. This claim, this decision, right now. Then watch what comes back. A specific answer, with a number or a date attached, is what a working correction system sounds like. A vague answer — reassurance, a pivot to how far we've come, an appeal to trust the process — is what a captured one sounds like as it starts covering for itself. An answer that turns into an attack on you for asking is the loudest signal of all. It tells you the system has already filed the question as a threat.

Once you have that answer, or its absence, the ladder from earlier tells you where to point your effort. If the system genuinely doesn't know what would change its mind, and the blankness looks like an honest gap rather than a defense, you are looking at incorrigibility. Help build the missing channel: better information, a clearer report, a way to check. If the system produces fog, or turns on you, you are looking at contracorrigibility, and the rootkit lesson stands. Don't trust the system's account of itself, and don't spend two years running internal reviews inside a review process that is itself the compromised part. Correct from outside its control, or accept that this structure will need rebuilding rather than repair.

Run the test on yourself first, for the same reason a doctor washes her hands before she examines the patient. Pick something you're currently certain about: a decision at work, a read on a relationship, a belief you'd list if someone asked what you're sure of. Ask yourself what would have to happen for you to conclude you're wrong. If an answer comes quickly and specifically, that's a small, genuine sign this belief is still open for business. If what comes instead is a restatement of why you're right, notice that, without immediately deciding what it means. Noticing is the whole first move. It is also the move a captured system is built to prevent.

What comes next

Everything in this chapter has assumed the correction at least arrives: that a friend's worried sentence gets heard, weighed, and rejected, however unfairly. That is the milder case. In a deeper version of the pattern, the correction is never perceived as a correction at all. It doesn't get argued with, because it never lands as something addressed to the system in the first place. That stranger failure is where Chapter 7 goes.

Sources