---
title: "Milkman"
url: "https://toddpaulbrownjr.com/builds/milkman/"
author: "Todd Paul Brown Jr."
description: "Milkman is a local-model benchmarking tool that refuses downloads that won't fit and tunes itself to hardware. Built in 31.4 hours. In development."
kind: "build"
status_word: "in development"
revised: "2026-09-28T06:34:47+00:00"
origin: "written"
canonical: true
updated: "2026-09-28T06:34:47+00:00"
---

# Milkman

## What it is

Milkman is a local-AI operating layer for a Windows/NVIDIA machine: it installs and manages a local inference runtime, keeps a verified registry of models, derives per-model hardware-fit settings, and runs a benchmark harness that compares a candidate model against an incumbent under a statistical promotion policy.

It's **in development**. The project's own handoff record lists unresolved risks, including an unrun real GPU candidate sweep from its new web UI.

## Why it exists

I wanted to know, with evidence rather than vendor benchmarks, which local model to run on my own 8 GB laptop GPU. Vendor coding benchmarks don't transfer to local, quantized runs: one published score was 73.4; the same model measured independently on this hardware came in at 24.7.

## How it's built

15 commits, about 31.4 hours across two calendar days, 22–23 August 2026. Every commit but the first two is agent-co-authored. The repo is built to be handed off cold: a fixed entry-point file, a session-handoff document rewritten (never appended) at the end of every session, and a structured memory of learnings, experiments and decisions, each carrying an epistemic-status ladder — assertion, observation, experimental evidence, validated learning, policy — that may only move up with linked evidence.

The hardware-fit logic runs four rules before touching a candidate model, ending in an outright refusal, before any download starts, if the model's overflow experts would need more RAM than the machine can spare.

## What broke

Ten parallel subagents launched for research burned a five-hour session budget in 33 minutes: 8 of 10 were killed mid-write when a shared web-search allowance emptied almost immediately. The fix, codified as a rule: cap concurrent subagents at three, ban sub-forking, tier models to task difficulty.

A subagent-written test's own cleanup call reverted every patch on its sandbox, including the one redirecting it to a fake install root. The rest of the test then ran against the real machine: it deleted a real 4.69 GB model file and overwrote a real registry checksum with a fake one. Recovery worked because the eviction policy writes an inventory record before deleting anything. Fixed with a sandbox fixture a cleanup call can't strip.

## Receipts

| Figure | What | Source |
|---|---|---|
| 15 / ~31.4 h | commits / elapsed build time, two calendar days | git log |
| 485 / 40,155 | files / insertions from the initial commit to tip | git diff --stat |
| 290 / 1 | tests passed / failed, last recorded run (recorded, not re-run) | project handoff record |
| 254 | golden-task evaluation runs recorded (recorded, not re-run) | project handoff record |

More on the [receipts page](/receipts/).
