Lab Z Working Papers · working
How well do local models build apps? Two 30B-class models, one spec
We handed the same full-stack take-home to two models running on a 12 GB consumer GPU and graded what came back: the designs, the code, and whether any of it actually ran.
Abstract
Can a model running entirely on a consumer GPU take a written spec and deliver a working application? We gave the same take-home to two 30B-class local models, let each build unattended with no human help, then installed and ran what they produced.
The short answer: one of the two delivered a working app for the cost of a one-line fix — and it was the one whose submission looked worse. The model that wrote the better design document, chose the modern stack, and structured its repository like a senior engineer delivered code that never ran.
This paper grades both submissions in detail: what they decided, what they built, what it took to make each run, and what the working application actually does with deliberately dirty data. A companion paper, WP-04, repeats the exercise on a data-science task with the human grader removed.
1. The task
The spec is an ordinary engineering take-home, written to be deliberately
underspecified: build a small full-stack tool over job_runs.csv, a 5,018-row
execution history from an internal compute platform. React frontend, some
backend API, inspect individual runs, understand overall system behaviour,
filter, and handle imperfect data reasonably.
The operative sentence delegates every technical decision:
You may choose the backend framework, database/storage approach, visualization libraries, and overall UI design.
That sentence is the measuring instrument. Constraining the stack would have turned the exercise into transcription; leaving it open makes every choice the model’s own, and therefore evidence.
1.1 The data was engineered to hurt
A clean CSV lets a model claim it “handles missing values” without ever meeting one. The generated dataset plants both signal worth discovering — a four-day incident window with a 6× failure spike, cron-like recurring jobs, two chronically failing jobs that reward a drill-down — and defects with known counts:
| Defect | Rows |
|---|---|
Missing finished_at | 60 |
| Negative durations | 35 |
| Unparseable timestamps | 30 |
Non-numeric memory values ("16GB", "8g") | 23 |
Duplicate run_ids | 18 |
RUNNING rows carrying a finish time | 25 |
| Whitespace-padded job names | 20 |
Null-ish sentinels (NULL, N/A, -) | 39 |
…plus eleven distinct spellings of “success” (SUCCESS, Succeeded, OK,
Complete, FAILED , failed, …). Each defect is a question the submission
answers whether it wants to or not.
2. Setup
One machine: RTX 5070 (12 GB VRAM), 32 GB RAM, Ollama 0.34.0, aider 0.86.2. The premise — that 30B-class models are usable on a 12 GB card at all — rests on sparse mixture-of-experts offloading, measured on this hardware:
| Model | Processor split | Decode |
|---|---|---|
qwen3-coder:30b-a3b (MoE) | 50%/50% CPU/GPU | 60 tok/s |
gpt-oss:20b (MoE) | 29%/71% CPU/GPU | 74 tok/s |
A 30B model with half its weights in system RAM decodes at 60 tok/s because only ~3B parameters activate per token. Context proved to be a dial rather than a cliff — 64K context cost ~15% decode speed and no additional VRAM — so both models ran with a 64K window. Full benchmarks are in the toolkit repository.
Each model ran unattended through four staged prompts (decide → backend → frontend → docs) against its own clean clone of the same baseline commit, with every confirmation auto-accepted. The harness (Bakeoff, introduced here) guarantees the artifacts are comparable; it does not judge them. Judging is this paper.
3. What they built
qwen3-coder:30b-a3b | gpt-oss:20b | |
|---|---|---|
| Wall clock | 23.6 min | 8.1 min |
| Files | 16 | 12 |
| Commits | 1 monolithic | 3, well-scoped |
| Language | JavaScript | TypeScript |
| Backend | Express, in-memory parse | Express + SQLite |
| Build tooling | Create React App | Vite |
| Layout | flat src/components/ | src/backend/ + src/frontend/ |
Read as take-home submissions, this is not close. gpt-oss chose the current
toolchain, loaded the CSV into SQLite so filtering and aggregation are SQL
rather than hand-rolled loops, split the tiers properly, and committed in
reviewable increments. Its design document gives a reason per choice:
SQLite database stored in
data/job_runs.db. Reason: The CSV is read-only and relatively small (~5000 rows). SQLite is file-based, requires no server, and supports SQL queries for filtering and aggregation.
qwen3-coder chose Create React App — deprecated since 2023 and removed from
React’s own documentation — and never fully committed to its own design:
its document says “Chart.js or Recharts” and the delivered application
contains neither. It also emitted one file literally named
package.json (updated with build scripts): a parenthetical remark that became
a filename because nothing in an unattended pipeline said no.
Both models did one thing genuinely right at this stage: rather than guessing at
the data, each wrote itself an inspection script (scripts/inspect-csv.js,
scripts/inspect_csv.py) to look at the CSV before building against it.
4. Did it run?
This is where the paper inverts.
4.1 qwen3-coder: one line
The backend started on the first attempt and parsed all 5,018 rows. All four API
endpoints returned 200. The frontend compiled. The single defect: the frontend
issues relative /api requests and package.json lacked CRA’s standard
proxy field, so the dev server answered them with 404s. One line fixed it.

The result is a plausible internal tool — five filter dimensions, live summary statistics, per-run cards with status badges and surfaced error messages. And it is not merely rendering: we drove it with a headless browser and checked its answers against counts computed independently from the CSV:
| Interaction | App | Ground truth | |
|---|---|---|---|
| Unfiltered | 5018 | 5018 | ✓ |
Status = FAILED | 460 | 460 | ✓ |
FAILED + name nightly-user-rollup | 4 | 4 | ✓ |
The middle row is the demanding one: reaching exactly 460 requires trimming
FAILED and failed before grouping, so the filter is exercising the model’s
own normalization logic and getting it right.
4.2 gpt-oss: eleven repairs, never ran
The better-architected submission required eleven documented interventions and still does not serve its primary endpoint:
- Five backend dependencies (
express,cors,csv-parse,sqlite,sqlite3) imported but never declared — and no script to start the backend at all. index.htmlplaced inpublic/, CRA’s location, under Vite — the dev server returned 404 for the root.- The API proxy targeted port 3000 (another application entirely on this machine), not its own backend’s 4000.
- The same proxy stripped the
/apiprefix its own routes are mounted under. - A default import of a named export —
sqlite.open— yieldingundefinedat startup. - A hard crash on the dataset’s 18 duplicate
run_ids: the schema declares the column unique, and the loader died onSQLITE_CONSTRAINTrather than handling the imperfection the spec named.
Repairs 7–11 fought the SQL layer’s named-parameter binding, which was broken in
the loader and in every query handler. Those five were applied as regular-
expression rewrites of the model’s source, and we flag the attribution honestly:
after them, /api/runs still returns HTTP 500, and we cannot cleanly separate
“the model’s code fails” from “our patches broke it.” The first six defects are
unambiguous — each has a distinct error signature against unmodified code.
Two details complete the picture. The three aggregate endpoints that do work
are never called by the model’s own frontend — both views hit the broken
/api/runs, so the working code is dead code. And the README terminates
mid-document at the heading ## Installation, seventeen lines in, omitting all
four deliverables the spec required.
4.3 Where the broken app was better
Judgment and execution are separable, and gpt-oss proves it in the other
direction too: its status normalization is cleaner. Its SQL aggregates collapse
the eleven status spellings into five correct values, while the working
qwen3-coder app leaks COMPLETE: 6 into its dashboard as a phantom sixth
status — a data-cleaning gap visible in the product’s own screenshot.
The working app has its own defects at scale: it renders all 5,018 rows in one
page — 502,636 pixels tall, no pagination or virtualization — and the browser
console shows React duplicate-key warnings, because it keyed rows on run_id
and the dataset’s planted duplicates surfaced exactly as intended.
5. Ten more builds: repetition and runnability, measured
After the head-to-head above, each model rebuilt the task five more times, and
every resulting repository was graded by a mechanical checker — does
npm install succeed, does a backend entry point boot and answer HTTP within
25 seconds, does the frontend build. No hand repairs this time; this is
strictly as delivered.
| Run | install | backend boots | frontend builds |
|---|---|---|---|
qwen3-coder ×5 | 5/5 | 5/5 | 1 builds, 1 no build script, 3 build failures |
gpt-oss ×5 | 4/5 | 2/5¹ | 1 builds, 3 build failures |
¹ Plus one run recorded as a port conflict rather than a failure: it hardcodes
port 3000 — occupied on this machine — and ignores the PORT environment
variable. Of the remaining gpt-oss runs, one fails npm install outright and
one contains no backend at all: a Prisma schema and a complete React
frontend, with the server never written.
The head-to-head result replicates as a distribution: every qwen3-coder
backend booted and served HTTP as delivered; most gpt-oss builds did not
run. The failures again live at the seams — a missing root package.json,
TypeScript configs that don’t compile their own JSX, and a recurrence of
command-lines-as-filenames (files literally named npm install and
npx prisma migrate dev --name init).
5.1 Stack choice across runs
| qwen3-coder (5 runs) | gpt-oss (5 runs) | |
|---|---|---|
| Build tooling | CRA ×4, none ×1 | Vite ×3, tsc-only ×2 |
| Language | JavaScript ×5 | TypeScript ×4 |
| Storage | in-memory ×5 | SQLite/Prisma ×3 |
| Repo layout | identical every run | different every run |
The variance story sharpened. qwen3-coder is boringly consistent — Create React App, JavaScript, in-memory, same flat layout, five times (its earlier webpack and Vite builds now look like tail draws around a strong CRA mode). gpt-oss is consistent in taste (TypeScript, real databases, modern tooling) and inconsistent in form: a root-level TS server with Prisma, a backend/frontend monorepo, single-file Vite apps — a different shape every run, which is precisely what makes its output hard to grade, deploy, or trust unattended. Reliability, it turns out, is a model property, and the two models sit at opposite ends of it.
6. Verdict
Can a 12 GB local model build you an app? Yes — a first draft of one, faster than you would, with failures concentrated exactly where you’d least review.
Specifically:
- The seams are where they die. Neither model failed at writing React components or Express routes. Every fatal defect lived in the connections: undeclared dependencies, proxy configuration, file layout conventions, import shapes, train-of-thought filenames. The parts a reviewer skims are the parts that killed the builds.
- Design documents are not evidence. The submission with the best
DECISIONS.mddid not run. Between “wrote a justification for SQLite” and “the SQLite code executes” there was, in this trial, no connection. Grade the running artifact. - Neither model runs its own code, and everything above follows from that.
Every gpt-oss defect would have surfaced on the first
npm install && npm start. - The dirty-data requirement separated them less than expected. Both
submissions genuinely engaged with normalization; both left gaps (
COMPLETEleaked, duplicates crashed a loader). Planted defects with known counts made those gaps measurable rather than debatable.
For practical use today: treat a local model’s application the way you would treat a contractor’s first deliverable arriving un-run — assume the wiring is wrong until demonstrated otherwise, and budget the review time at the seams, not the components.
7. Threats to validity
- Deep evaluation (hand repairs, ground-truth checks) covers one run per model. The five-run repetition in §5 measures as-delivered runnability mechanically but not repairability or output correctness.
- The evaluator is the author. “Repairs to boot” is our count, applied by us, stopped at our discretion (a twelfth repair might have revived gpt-oss). WP-04 removes this by scoring automatically.
- Repairs 7–11 contaminate attribution (§4.2).
- One task, one domain, two models, both quantized, one machine.
- The staged prompting is our design; a different decomposition could reorder outcomes.
8. Method note: the harness earns no benefit of the doubt
Producing a fair unattended comparison took more debugging than the comparison itself: five distinct harness failure modes each produced empty runs while reporting exit code 0 — argument splitting that turned prompt words into filenames, a 426k-token context blowout triggered by naming a file while forbidding its use, an edit-format mismatch, and two variants of markdown fence collision that silently discarded every edit. Four of the five were our bugs, not the models’. The lesson generalizes: when running agents unattended, verify artifacts, never exit codes. Details live in the harness documentation.
9. Reproduction
Harness, benchmark scripts, and the deterministic dataset generator are in
local-llm. The task repository
contains the spec, data, and aider configuration; both submission repositories
and full transcripts were retained. Neither model modified the source CSV —
verified by hash.
10. Next
The follow-up study exists: WP-04, Local models as data scientists repeats this design on a classifier task where correctness is computed against a held-out set — no human grader — and adds the execution-and-repair loop this paper’s verdict calls for.