Introducing Bakeoff: A Fair Fight Between Local Models
Run one written spec through several local models unattended, on identical baselines, and compare what they actually build — part of the local-llm toolkit.
The Local Model Comparison Problem
You want to know which local model to use for real work. So you check a leaderboard, read some benchmark scores, maybe skim a few generated diffs — and none of it tells you what you actually need to know, which is whether the thing it produces runs.
Worse, informal comparison is almost always unfair. You try model A on Monday with one prompt, model B on Thursday with a slightly different one, from a different starting state, and rescue each of them a few times when they wander. By the end you have an opinion and no evidence.
The two failure modes compound. Scores measure the wrong thing, and hand-rolled comparisons aren’t controlled.
Enter Bakeoff
Bakeoff runs a written specification through several local models unattended, each starting from an identical clean clone, and leaves you the resulting repositories side by side.
./scripts/oneshot.ps1
That’s it. Each model gets the same spec, the same baseline commit, the same prompts, and no human help. What comes back is a repository per model plus a full transcript of how it got there.
The name is deliberate. It isn’t a benchmark — it produces no score. Same recipe, several cooks, look at what comes out of the oven.
Bakeoff is one tool in local-llm, a small toolkit for running and evaluating models on a single consumer GPU.
How It Works
Each model runs four staged prompts with a fresh context per stage:
- decide — read the spec, write
DECISIONS.mdrecording the stack, storage approach, data-handling plan, and deliberate omissions. Then scaffold the project to match. - backend — implement the API described in
DECISIONS.md. - frontend — implement the UI against that API.
- docs — write the README: run instructions, assumptions, architecture.
DECISIONS.md is the handoff between stages, not the model’s memory. That is
what keeps the whole thing inside a 64K context window.
There’s also -Mode single: one message, one response, build the entire
application. Purer, and far less likely to work.
Design Choices That Matter
The prompts say what, never how. No stack, no libraries, no schema. Every one of those is the model’s choice — which is the entire point, because those choices are the most legible signal you get about a model’s judgment.
Genuinely unattended. Every confirmation is auto-accepted. Nobody nudges the model when it goes sideways, because the nudging is exactly what’s being measured out. The cost is real: it will happily create twenty junk files. That’s why every run is an isolated clone.
Sequential, not parallel. On a 12 GB card two resident models don’t fit. Running them concurrently thrashes the loader instead of parallelising anything.
Seeded defects in the task data. A clean dataset lets a model claim it “handles missing values” without ever meeting one. The included generator plants known counts of malformed timestamps, negative durations, duplicate IDs, and inconsistent enum spellings — then keeps itself outside the task repo, because a model that can read the generator can satisfy the requirement by regenerating clean data.
What It Found
The first trial gave qwen3-coder:30b-a3b and gpt-oss:20b the same
underspecified take-home: a full-stack explorer over ~5000 rows of defect-seeded
job history.
Read the artifacts and gpt-oss:20b wins easily — TypeScript, SQLite, a properly
separated backend, incremental commits, a design document with a reason per
choice. qwen3-coder chose plain JavaScript and Create React App, deprecated
since 2023.
Then we installed them.
qwen3-coder:30b | gpt-oss:20b | |
|---|---|---|
| Architecture, on paper | weaker | stronger |
| Repairs to boot | 1 | 11 — still failing |
qwen3-coder’s application needed one line, then worked — and its filters return
counts matching independently computed ground truth exactly. gpt-oss’s needed
five undeclared dependencies, a file moved, two proxy fixes, an import shape
corrected, and repeated parameter-binding repairs, and its primary endpoint still
returned HTTP 500.
Reading the code gave the wrong answer about both of them. That finding is the reason Bakeoff exists in this shape.
Full write-up: I Gave Two Local Models the Same Take-Home. Formal record, with method and threats to validity: Working Paper #3.
Traps It Encodes
Four hazards silently produced empty runs at exit code 0 before they were fixed. They’re baked in now, and they’re documented because they’ll bite anyone writing new prompts:
- Never name a large file in a prompt. Agents auto-detect filenames and offer to add them; auto-accept obliges. The instruction “do not add this file” is what loaded a 792 KB CSV and OOM’d the GPU.
- Never put triple-backtick fences in a context file. Aider picks its edit fence from characters present in context; backticks escalate it to a fence the model ignores, and every edit is silently discarded.
- Ban nested fences in generated markdown — indent code four spaces instead.
A README’s own code blocks otherwise terminate the outer fence early. Demanding
a longer outer fence seems like the fix and isn’t: a model reasoned about the
rule correctly, then closed the block with
######— six hashes. - Never let a model write a lockfile. One spent an entire 15-minute stage
budget emitting
package-lock.jsontoken by token and was killed before writing any code. With lockfiles ignored, the same stage finished in 3.7 minutes. - Pass prompts as files, not arguments.
Start-Process -ArgumentListsplits on whitespace; each word becomes a filename.
The generalizable lesson: judge unattended runs by artifacts, not exit codes.
The Validation That Nearly Didn’t Happen
Before publishing this, we ran the committed pipeline end-to-end — something that had, embarrassingly, never been done. Every fix had been verified in isolation.
It failed twice.
The lockfile trap killed one model’s first stage outright. And the six-backtick
rule — already committed, already documented as the fix — turned out not to work
at all. Both failures returned exit code 0. One reported done (1m) while
producing no files whatsoever.
After fixing both:
| Model | Stages | Wall clock | Files | Deliverables |
|---|---|---|---|---|
gpt-oss:20b | 4/4 clean | 10.4 min | 17 | README 93 lines, NOTES 12 |
qwen3-coder:30b-a3b | 4/4 clean | 15.4 min | 14 | README 76 lines, NOTES 51 |
The validation also surfaced something the original trial had only theorised. Three runs of the same model on the same prompt produced three different build tools — Create React App, hand-rolled webpack, and Vite. Run-to-run variance on stack choice is larger than the between-model difference we built a comparison on. That is now a measured caveat rather than a hypothetical one, and it is the strongest argument for the variance runs on the roadmap.
Get Started
git clone https://github.com/lab-zee/local-llm
cd local-llm
./scripts/configure-ollama.ps1 # agent-sized context; run once
./scripts/serve.ps1
./scripts/oneshot.ps1
Adding a model is one array entry. Adding a task is a git repo with a SPEC.md,
your data, a fence-free excerpt, and an .aiderignore. Full documentation in
docs/bakeoff.md.
Current Status
Bakeoff ships the controlled conditions, not the verdict. It guarantees the artifacts are comparable. It does not tell you which one is better.
That boundary is deliberate. The first trial’s finding was that reading the artifacts gave the wrong answer — so a tool that shipped an evaluator today would almost certainly ship a reading-based one, the exact method the trial discredited. Better to ship the fair fight and leave scoring to you.
Concretely, what it does not yet do:
- It never runs the code. Every
gpt-ossdefect would have surfaced on the firstnpm install. Neither the models nor the harness have a feedback loop. - No acceptance criteria. “Works” is currently a human judgment call.
- One run per model proves little. No variance estimate, and sampling temperature is left at defaults.
What’s Next
Two changes, in order, and they’re the same roadmap as the research:
- An execution stage — boot the result, feed failures back, iterate. This targets the identified mechanism directly.
- Machine-checkable acceptance criteria per task — endpoints return 200, aggregates match ground truth computed independently from the source data.
That second one is when the evaluator ships, and when “works” stops being our opinion. Until then, Bakeoff gets you a fair fight — and in our experience, the fair fight is the part people skip.