Skip to content
TODD BROWN
« The Anatomy of Corruption and Corrigibility III — The Instruments · Chapter 9

Markers and Tests

Picture two companies. Both hold a monthly all-hands. Both end it with the same line from the CEO: "We welcome your candid feedback — that's how we get better."

At Company A, three months ago, a mid-level engineer stood up and said the flagship product's onboarding flow was quietly driving away the customers it was meant to convert, and that the dashboard everyone watched every week was measuring the wrong thing. The room went quiet. Then the head of product asked her to walk through the data in more detail the following week. She did. Two engineers were reassigned to fix the flow. The roadmap changed. She was promoted eight months later, and when people asked why, the answer was not "despite the objection." It was "because of it."

At Company B, a mid-level engineer stood up eighteen months ago and raised something similar about a metric everyone there watched too. The room went quiet. He was thanked for his candor. Within two quarters his team had been folded into a different reporting line "for efficiency." The quarter after that, his role was cut in a round of layoffs. Nobody stood up in an all-hands and connected the two events. Everyone connected them privately, in hallway conversations and group chats management wasn't in. The next employee with doubts about the metric kept them to herself.

Same ritual. Same sentence, same tone, same cadence. From outside — from a press release, an investor call, a culture deck, a first-round interview — the two companies are close to indistinguishable. Both say they welcome feedback. Only one can receive it.

That gap, between what a system says about itself and what its structure does with what it's told, is what this book has been circling. This chapter narrows it to something useful on a Tuesday afternoon. If you're standing outside either company, with no access to internal channels or personnel files, what can you actually check? What tells you which one you're looking at?

Performance, not reality

The trouble is that almost every visible sign of health can be manufactured, and a system under enough scrutiny will learn to manufacture it convincingly. A comment box by the break room. An ombudsman's office with its own door. A retrospective where the team "processes" what went wrong. A red team hired at real expense. A system can grow all of these organs, and none of them, alone, proves anything. Company B almost certainly has a comment box. It probably has a whistleblower line printed in the handbook, the kind nobody uses twice after watching what happened to the first person who did.

This isn't cynicism about feedback mechanisms. It's a design constraint: any marker cheap enough to install and display is cheap enough to fake, and any system with something to hide has every reason to display it anyway. A comment box costs a few dollars and a sentence in the handbook. A culture deck costs an afternoon with a designer. Neither requires the reality it claims to represent.

The markers worth trusting are the ones that are hard, expensive, or structurally difficult to fake: markers where producing the appearance of health requires building most of the real machinery along the way, because there's no cheaper route. That's the organizing question for the rest of this chapter. Not "what does health look like from outside," but "which signs of health can only be produced by a system that actually has it."

The disconfirming-conditions test

Chapter 6 introduced the question: what evidence would change your mind? Here it becomes a scoring rule.

Ask both companies the same thing: what would have to happen for you to admit your claim about your own culture is false?

Company A's product lead answers with specifics. "If we go three straight roadmap cycles where engineer-raised concerns never show up in the next planning document, the channel has become decorative. If turnover among our most tenured engineers starts running higher than turnover company-wide, people have stopped speaking up and started leaving. Either one would mean we're wrong about ourselves, and we'd have to say so."

Company B's spokesperson answers somewhere else. "We're always improving. We take every piece of feedback seriously and we're proud of our culture of openness." Push again — what specific evidence, if you saw it, would change your mind? — and the answer drifts toward the person asking. Are you trying to make trouble? Why are you so focused on the negative?

The reason this is so hard to fake is that naming a disconfirming condition requires the machinery that would notice it: something tracking tenured-engineer turnover, something checking whether raised concerns reach a planning document, something willing to report an unflattering number. "We welcome feedback" costs nothing to say. A specific, checkable condition can only be generated by asking that machinery what it would need to see.

That gives the test a portable two-part score:

  • Specific and measurable. Does the answer name a number, a threshold, or a window of time — something a stranger could check later without the system's permission?
  • Costly to the answerer. Does naming it put something on the line for the person giving it, or is it phrased so loosely that no future event could ever satisfy it?

"We're always improving" fails both: no number, no risk. "If turnover among our best engineers crosses fifteen percent, this claim about our culture is false" passes both. An outsider can check it, and if it comes true, the person who said it has to own that.

Watch, finally, for the answer that never touches the question and turns instead to the motives of whoever asked. That redirect isn't a delayed answer. It's a live demonstration that the question couldn't be processed, the identity-defense Chapter 6 described, happening in front of you.

The anatomy of a corrigible system

So far this chapter has been about catching a system failing a test. Turn the question around. What do the healthy ones have in common? Not what they say in a mission statement, but what structural properties keep showing up in systems that go on catching their own mistakes for years rather than months?

Three properties recur, and each comes with a mechanism: a reason it works, not just a description of how it looks.

The first is exercised dissent. Picture a weekly product review where the same junior engineer keeps raising the same uncomfortable question about a metric the leadership team likes. Two months ago her objection got minuted, tested against the numbers, and turned out to be right, and the roadmap moved. Nobody on that team would call the meetings pleasant. They'd call them loud. That loudness isn't a defect the team tolerates alongside its real work. It is the real work. Correction machinery that never gets used degrades the way a muscle does from disuse: slowly and quietly, until it can't do the job when it's finally needed. A grievance process nobody has used in five years and one that was never built are close to the same object. Chapter 6 offered a related hypothesis from the other direction: that captured systems which stop meeting any dissent may become more brittle, not more secure. Either way, dissent isn't noise a healthy system puts up with. It's load-bearing. Take it away and the system's ability to notice its own errors goes with it.

The second property has a name from early cybernetics: requisite variety, from W. Ross Ashby's 1956 book An Introduction to Cybernetics. Stated plainly: a regulator can only counter as many kinds of disturbance as it has kinds of response. A thermostat that can only switch on the heat handles a drafty window and a failing furnace the same way, and that's fine, because both push the room colder. But it has no answer to a heat wave. It has one kind of response, and the room can be pushed in two directions. The same limit shows up in a room full of people. A leadership team that reads the same sources, came up the same career path, and shares the same assumptions about what counts as a risk has, in effect, one kind of response, however sharp its members are. It isn't a failure of intelligence. It's a hole in the detector: a monoculture has no internal channel tuned to the kind of drift nobody in the room expects. This is the mechanism behind a pattern most people have seen at least once: an organization blindsided by something outsiders saw coming for months. The outsiders usually weren't smarter. They were a different sensor, tuned to a different range.

The third property is whether people inside a system can change their minds without it costing them a self. Earlier chapters met this as a trait of individuals. Here it's something a system can be built for. A system builds it by making it not just safe but respected to say "I was wrong" out loud: promoting the engineer who raised the uncomfortable question instead of moving her sideways, treating a reversed decision as evidence the review process works rather than evidence someone failed. Where updating a belief costs only the belief, people update readily. Where it costs standing, income, or belonging, people defend the belief instead. Not out of stubbornness, but because the real stakes on the table are higher than the belief.

None of this is a law, the way a triangle must have three sides. These are strong structural tendencies that recur because of how correction works in practice. They are not guarantees, and not a checklist that certifies a system healthy once all three boxes are ticked. A system can have all three and still fail in a way nobody built a channel to catch. What they give you is a place to look, and, if you're the one building the system, a place to build.

The three also reinforce each other. Variety without exercised dissent is a room full of disagreement nobody voices, which works no better than a monoculture on the day it matters. Exercised dissent without a safe way to change your mind turns into a fight over who's right rather than what's true, because every disagreement becomes a referendum on someone's standing. And a safe way to change your mind means little if there's no variety in the room to generate anything worth changing it about. They're less a checklist than three supports under one structure, each bearing less weight when the others are missing.

The interception pipeline

Not every correction that gets sent actually lands, and where it gets stopped is diagnostic in its own right. Corrective information passes three doors on its way from the outside world into a system's behavior, and it can be stopped at any of them.

The first door is perception. Sometimes a critique never arrives as a critique at all. Chapter 7 covered this in full as anti-memesis: "the whole pattern of how we relate" gets heard as "tell me the specific incidents and I'll fix those."

The second door is evaluation. Here the critique registers — the system understands what's being claimed — and has a ready answer anyway. Established systems build up a repertoire of standard rebuttals the way an immune system builds antibodies: the counterargument, the disclaimer, the "that's a common misunderstanding," rehearsed before the challenge arrives. A well-defended set of beliefs can meet almost any objection with a prepared reply and move on, whether or not the reply addresses what was raised. The critique got in. It just didn't move anything.

The third door is action. The critique gets through, gets understood, is even privately believed, and the person who raised it is penalized anyway, because acting on it openly (a public reversal, a lost promotion cycle, a lawsuit) costs more than quietly ignoring it. This is where exit costs live: the toll a system charges anyone who tries to leave, push back, or change course out loud, set high enough that staying wrong stays cheaper than being right.

Here's the part worth carrying forward. Knowing which door a system has staked its main defense on tells you where to look for the break. Not whether it breaks, and not when. Where.

A system whose authority rests on checkable factual claims — a safety record, compliance numbers, an institution's account of what happened — has staked its defense at the perception door, and it tends to break on information access. Put a clean, hard-to-suppress fact in front of enough people and the defense at that door stops working.

A system whose authority rests on argument and legitimacy — a school of thought, a professional consensus — has staked its defense at the evaluation door. It tends to break on a better argument, or on its own members quietly changing their minds faster than anyone in charge notices. That internal drift often matters as much as any outside refutation.

A system whose authority rests on identity and belonging, where the real cost of dissent is losing a community rather than an argument, has staked its defense at the action door. It tends to break when exit costs collapse: when leaving stops being expensive. That can happen slowly, as alternatives appear, or suddenly, when one visible person leaves, lands well, and shows in public that the advertised cost was smaller than everyone assumed.

So the word "robust" deserves care. No system is robust in general. It is robust to one kind of pressure and brittle to another: robust to public criticism, brittle to a documented internal leak; robust to a legal challenge, brittle to a single credible defector. Naming a defense without naming its matching vulnerability is a different and usually less useful claim.

A concrete version: a regional bank might be genuinely robust to a bad news cycle. Its depositors have stuck with it through negative press before, because the relationship rests on years of checkable, boring reliability. The same bank can be brittle to one leaked internal memo showing it knew about a problem and sat on it, because that fact lands at the perception door its whole defense was staked on. Knowing which door a system relies on tells you which future headline should worry you.

Presence, not absence

There's a trap built into how most people check on a system's health: treating the absence of bad signs as proof of good ones. No scandals this year. No lawsuits. No walkouts. Nobody's complaining, at least not loudly. Surely that means things are fine.

The problem is that "no bad signs" describes two very different situations equally well. One is a healthy system quietly catching small problems before they become visible. The other is a captured system in a lull: defenses intact, simply not under attack, members quiet because the enforcing environment did its work long ago. From outside, both look like silence. Absence alone cannot separate capture from calm, and the two call for opposite responses.

What breaks the tie is a presence marker: something positively happening now that only happens in the healthy state. For a group, there's a sharp version, first raised in Chapter 3: does the group stay coordinated and functional when its enemy disappears?

A community organized around a real shared purpose keeps doing that purpose's work whether or not it's under outside threat. A research group keeps publishing, correcting its errors, and training the next generation whether or not a rival lab is in the news. A team that argues productively keeps doing so in a quiet quarter, not only in a crisis. A group whose main glue is opposition to a shared enemy tends to slacken, drift, or fracture once the enemy fades, or goes looking for a new one, because what held it together was the shared fight. Enemy-independent function is a presence marker: evidence the structure underneath is real, checkable in calm periods. Enemy-dependent function is, as Chapter 3 put it, a signature closer to egregoric glue than to shared, ongoing work.

The practical value is timing. A presence marker can be read before the outcome: before the scandal, the walkout, the collapse everyone later calls obvious. It's a leading indicator you can check on an uneventful Tuesday, without waiting for the postmortem.

Absence markers are seductive because they're free. Checking whether anyone is complaining takes no effort. Checking whether a group still does its real work when there's nothing to fight costs more: watching across a quiet stretch, comparing it honestly to a loud one, and being willing to notice the quiet stretch looked thinner. That effort is the price of telling capture from calm, and it's worth paying before you're the one explaining that nobody saw it coming.

Diagnostic nodes and predictive leaves

One last distinction, smaller than the others, but it quietly ruins a lot of forecasts.

Every marker so far is what's worth calling a diagnostic node: something checkable today, before anyone knows how the story ends. The disconfirming-conditions answer. Whether dissent is actually exercised. Whether a group keeps working with its enemy gone. A diagnostic node reads like a blood test: it tells you which situation you're in now, without waiting to see what happens.

A predictive leaf is different: an outcome at the far end of the story, confirmable only once it has happened. The scandal breaks. The company collapses. The movement splits. Leaves are real, and naming one after the fact can be useful. But a leaf isn't a marker, because something you can only read after the outcome can't help you locate yourself before it. "The culture was obviously toxic" is a fine thing to say after a company implodes. As a diagnostic tool, it's a caption under a photograph that already finished developing.

The mistake to watch for, in your own reasoning as much as anyone's, is smuggling a leaf into a decision as if it were a node. "They're probably fine, nothing's gone wrong yet" treats the absence of an unresolved outcome as a present-tense reading. It tells you nothing about which branch you're on. Every marker in this chapter was chosen to be a node: checkable now, before the tree finishes growing. How far to trust a present-tense reading, and how it should shape a forecast, is a harder question, and it belongs to Chapter 10.

The marker kit

Carried forward, this chapter reduces to four checks you can run on any system — a company, an institution, a team, an online community — using only what's available from outside.

  1. The disconfirming-conditions question. Ask: what evidence would change your mind about this claim? Score the answer on two things only. Is it specific and measurable? Would it cost the person answering something if it came true?
  2. The three anatomy checks. Is dissent actually exercised, not just formally permitted? Is there enough variety in the room to notice a problem nobody expected? Is changing your mind here safe, or better, respected?
  3. The door placement. Where is this system's main defense staked — perception, evaluation, or action — and so where should you look for its break? Never call a system simply "robust." Name what it's robust to, and what it isn't.
  4. The presence rule. Look for what the system does that only a healthy system does, not for the absence of visible trouble. For a group: does it still function when its opposition disappears?

None of these, asked alone, proves anything. Asked together and answered honestly, without softening the answer to protect a system you already like, they're the closest thing available to reading a system's health from outside, before you find out the hard way which company you joined.

These markers describe where a system stands now. They don't tell you where it's headed, how fast, or what it would take to move it. That's a different kind of question, and treating a diagnostic reading as a forecast is exactly the mistake the last section warned against. Chapter 10 takes it up directly: given an honest reading of the present, what can be said about where a system is going, and what can't be said, however badly anyone wants the number.

Sources