Milkman
What it is
Milkman is a local-AI operating layer for a Windows/NVIDIA machine: it installs and manages a local inference runtime, keeps a verified registry of models, derives per-model hardware-fit settings, and runs a benchmark harness that compares a candidate model against an incumbent under a statistical promotion policy.
It's in development. The project's own handoff record lists unresolved risks, including an unrun real GPU candidate sweep from its new web UI.
Why it exists
I wanted to know, with evidence rather than vendor benchmarks, which local model to run on my own 8 GB laptop GPU. Vendor coding benchmarks don't transfer to local, quantized runs: one published score was 73.4; the same model measured independently on this hardware came in at 24.7.
How it's built
15 commits, about 31.4 hours across two calendar days, 22–23 August 2026. Every commit but the first two is agent-co-authored. The repo is built to be handed off cold: a fixed entry-point file, a session-handoff document rewritten (never appended) at the end of every session, and a structured memory of learnings, experiments and decisions, each carrying an epistemic-status ladder — assertion, observation, experimental evidence, validated learning, policy — that may only move up with linked evidence.
The hardware-fit logic runs four rules before touching a candidate model, ending in an outright refusal, before any download starts, if the model's overflow experts would need more RAM than the machine can spare.
What broke
Ten parallel subagents launched for research burned a five-hour session budget in 33 minutes: 8 of 10 were killed mid-write when a shared web-search allowance emptied almost immediately. The fix, codified as a rule: cap concurrent subagents at three, ban sub-forking, tier models to task difficulty.
A subagent-written test's own cleanup call reverted every patch on its sandbox, including the one redirecting it to a fake install root. The rest of the test then ran against the real machine: it deleted a real 4.69 GB model file and overwrote a real registry checksum with a fake one. Recovery worked because the eviction policy writes an inventory record before deleting anything. Fixed with a sandbox fixture a cleanup call can't strip.
Receipts
| Figure | What | Source |
|---|---|---|
| 15 / ~31.4 h | commits / elapsed build time, two calendar days | git log |
| 485 / 40,155 | files / insertions from the initial commit to tip | git diff –stat |
| 290 / 1 | tests passed / failed, last recorded run (recorded, not re-run) | project handoff record |
| 254 | golden-task evaluation runs recorded (recorded, not re-run) | project handoff record |
More on the receipts page.