Reading this with an AI? Start at /for-agents/ (orientation) or /llms.txt (index).

Skip to content
TODD BROWN
« All builds

Milkman

in development

A local-model benchmarking tool that refuses downloads that won't fit. Built in 31.4 hours.

Claude Code · Python · FastAPI · React · llama.cpp

What it is

Milkman is a local-AI operating layer for a Windows/NVIDIA machine: it installs and manages a local inference runtime, keeps a verified registry of models, derives per-model hardware-fit settings, and runs a benchmark harness that compares a candidate model against an incumbent under a statistical promotion policy.

It's in development. The project's own handoff record lists unresolved risks, including an unrun real GPU candidate sweep from its new web UI.

Why it exists

I wanted to know, with evidence rather than vendor benchmarks, which local model to run on my own 8 GB laptop GPU. Vendor coding benchmarks don't transfer to local, quantized runs: one published score was 73.4; the same model measured independently on this hardware came in at 24.7.

How it's built

15 commits, about 31.4 hours across two calendar days, 22–23 August 2026. Every commit but the first two is agent-co-authored. The repo is built to be handed off cold: a fixed entry-point file, a session-handoff document rewritten (never appended) at the end of every session, and a structured memory of learnings, experiments and decisions, each carrying an epistemic-status ladder — assertion, observation, experimental evidence, validated learning, policy — that may only move up with linked evidence.

The hardware-fit logic runs four rules before touching a candidate model, ending in an outright refusal, before any download starts, if the model's overflow experts would need more RAM than the machine can spare.

What broke

Ten parallel subagents launched for research burned a five-hour session budget in 33 minutes: 8 of 10 were killed mid-write when a shared web-search allowance emptied almost immediately. The fix, codified as a rule: cap concurrent subagents at three, ban sub-forking, tier models to task difficulty.

A subagent-written test's own cleanup call reverted every patch on its sandbox, including the one redirecting it to a fake install root. The rest of the test then ran against the real machine: it deleted a real 4.69 GB model file and overwrote a real registry checksum with a fake one. Recovery worked because the eviction policy writes an inventory record before deleting anything. Fixed with a sandbox fixture a cleanup call can't strip.

Receipts

Figure What Source
15 / ~31.4 h commits / elapsed build time, two calendar days git log
485 / 40,155 files / insertions from the initial commit to tip git diff –stat
290 / 1 tests passed / failed, last recorded run (recorded, not re-run) project handoff record
254 golden-task evaluation runs recorded (recorded, not re-run) project handoff record

More on the receipts page.