Reading this with an AI? Start at /for-agents/ (orientation) or /llms.txt (index).

Skip to content
TODD BROWN
« All builds

Interactive Fiction Engine

parked

An on-device interactive fiction engine: a small model on a phone, inside a harness that decides what happens.

Claude Code · TypeScript · node-llama-cpp · SQLite

What it is

An on-device interactive fiction engine: a small model running on a phone, driven by a deterministic TypeScript harness that owns canon state, pacing, branching and knowledge boundaries. The model only writes prose and dialogue inside boundaries the code defines. Content domain: interactive fiction, nothing more.

It's parked. Phases 1–8 of a nine-phase roadmap passed their exit gates on 30 August 2026; Phase 9, "Hardening," was nearly complete except for checks needing a real phone. Last commit: 11 September 2026.

Why it's on this site: parked for two named reasons — prose quality (possibly a model gap, not the harness) and product fit (time and battery life on a phone). Both are measurable, and I haven't finished measuring either.

Why it exists

I wanted to prove the harness mattered more than the model: if a small on-device model can't write beautiful prose alone, the architecture around it should make up the difference. Every passage carries a checklist; the response is rejected before it reaches the screen if it leaks information the player shouldn't have yet, or skips a required beat.

How it's built

58 commits across 6 active days. A 3,020-line requirements document defines nine phases, each with a named exit gate; a 38-entry decision log records what I chose and ruled out. Two eval families run separately: a deterministic suite with no model or GPU, and 19 scored evals against the real model — two of which, knowledge isolation and locked-development, are zero-violation gates re-run on every prompt change.

What broke

Up to six shipped builds ran different sampler settings than every eval battery that had supposedly validated them, because the phone runtime filled in its own defaults wherever the engine didn't set a field. I found it because the app on my phone never matched the eval results. Fixed by setting every sampler field explicitly on every call.

One eval compared the model's own claimed story-beat coverage against a code-computed definition; agreement came back at chance level, so I deleted the self-reported metric. A separate accessibility audit found 73 violations across 28 screens on its first run; all fixed the same day.

Receipts

Figure What Source
1,010 deterministic tests green, highest recorded figure project roadmap, milestone M21
58 / 6 commits / active calendar days git log
38 formal numbered decisions decision log
19 distinct scored evals, each with a numeric threshold evals directory
73 → 0 accessibility violations found → fixed, same day, 28 screens accessibility audit

All measured. More on the receipts page.