Interactive Fiction Engine
What it is
An on-device interactive fiction engine: a small model running on a phone, driven by a deterministic TypeScript harness that owns canon state, pacing, branching and knowledge boundaries. The model only writes prose and dialogue inside boundaries the code defines. Content domain: interactive fiction, nothing more.
It's parked. Phases 1–8 of a nine-phase roadmap passed their exit gates on 30 August 2026; Phase 9, "Hardening," was nearly complete except for checks needing a real phone. Last commit: 11 September 2026.
Why it's on this site: parked for two named reasons — prose quality (possibly a model gap, not the harness) and product fit (time and battery life on a phone). Both are measurable, and I haven't finished measuring either.
Why it exists
I wanted to prove the harness mattered more than the model: if a small on-device model can't write beautiful prose alone, the architecture around it should make up the difference. Every passage carries a checklist; the response is rejected before it reaches the screen if it leaks information the player shouldn't have yet, or skips a required beat.
How it's built
58 commits across 6 active days. A 3,020-line requirements document defines nine phases, each with a named exit gate; a 38-entry decision log records what I chose and ruled out. Two eval families run separately: a deterministic suite with no model or GPU, and 19 scored evals against the real model — two of which, knowledge isolation and locked-development, are zero-violation gates re-run on every prompt change.
What broke
Up to six shipped builds ran different sampler settings than every eval battery that had supposedly validated them, because the phone runtime filled in its own defaults wherever the engine didn't set a field. I found it because the app on my phone never matched the eval results. Fixed by setting every sampler field explicitly on every call.
One eval compared the model's own claimed story-beat coverage against a code-computed definition; agreement came back at chance level, so I deleted the self-reported metric. A separate accessibility audit found 73 violations across 28 screens on its first run; all fixed the same day.
Receipts
| Figure | What | Source |
|---|---|---|
| 1,010 | deterministic tests green, highest recorded figure | project roadmap, milestone M21 |
| 58 / 6 | commits / active calendar days | git log |
| 38 | formal numbered decisions | decision log |
| 19 | distinct scored evals, each with a numeric threshold | evals directory |
| 73 → 0 | accessibility violations found → fixed, same day, 28 screens | accessibility audit |
All measured. More on the receipts page.