Skip to content
TODD BROWN

Todd Paul
Brown Jr.

LOGIC
  • Claude Code
  • Agent orchestration
  • Test & eval gates
  • Revenue operations
  • Web platforms
Illustration of a brain: one half grey circuitry, the other half bright paint

At the frontier of building with AI agents.

CREATIVITY
  • Game development
  • Interactive fiction
  • Audio tools
  • Brand & design
  • Writing
FOR PEOPLE BUILDING WITH AI AGENTS

Field notes from building production software with AI agents: what works, what breaks, and the receipts for both.

Everything here comes from real builds. Build logs with the test counts and the failures left in. Long-form lessons on evals that lied, agents that ignored their rules, and the gates that caught them, written so you can use them on your own work. And a field guide for spotting when a system, human or machine, starts optimizing the wrong thing. New pieces land here first. Who’s writing this

896
automated tests passing on Content Ops
Source

state/repo-state.md, last full run at commit f861f98: 839 TypeScript + 57 Python, 0 failing. Measured. Pulled 2026-09-25.

12 days
Sadyr, from blueprint to production
Source

git log, Sadyr repo: first commit 2026-06-18, live 2026-06-30, 63 commits on 8 working days (claims ledger C-01). Measured.

92 of 95
planned website pages drafted by Content Ops in two days of batch runs
Source

state/manifest.yaml B-03 / B-08, claims ledger C-08. Measured, 2026-09-08 to 2026-09-09.

81,707
lines of gameplay C++ in a game still in development
Source

Game development-stats report, 2026-09-20 (git + line count). Measured. Written by Claude Code under Todd's direction; the game is not shipped.

64
capability packages in LEGION, the system I run Adroit on
Source

state/repo-state.md, filesystem inventory across five capability folders (claims ledger C-11). Measured. Pulled 2026-09-25.

Corrigibility

A field guide to diagnosing systems that drift: what a system really optimizes, and whether it can still be corrected. Software, AI models, companies and the people inside them fail in the same shapes.

III — The Instruments · Chapter 9

Markers and Tests

Signs of health are cheap to fake. The markers worth trusting are the ones only real correction machinery can produce, read now, before the outcome.

Read the chapter »
IV — The Applications · Chapter 11

AI Systems

Specification gaming, sycophancy and eval-gaming: the proxy failure in machinery with no self. Two senses of corrigibility, and instruments for builders.

Read the chapter »

All 15 chapters Or browse the writing »

NOW

This month I'm finishing Sadyr's meeting-booking flow, rebuilding this site, and queuing 32 LinkedIn posts for October and November. More »

Follow on LinkedIn