← Papers

Lab Z Working Papers · working

How well do local models write numerical code? A physics simulation graded against exact solutions

Ten unattended runs of a heat-equation take-home, scored by comparing every output temperature to the true solution of the PDE. One model scored 5/6 five times out of five; the other 0/6 five times out of five.

Third trial in the series: the same two local models (qwen3-coder:30b-a3b, gpt-oss:20b, both on a 12 GB consumer GPU) each ran a numerical-computing take-home five times, unattended — build a 1D heat-conduction simulator, output temperatures at requested positions and times. Grading is fully objective: every reported temperature is compared against the exact Fourier-series solution of the PDE, with one hidden config using piecewise-varying diffusivity so the textbook formula cannot substitute for a real solver. A pre-registered calibration ladder (careful solver 6/6, naive 4/6, lazy 0/6, formula-only 1/6) anchors the scale. The result is the cleanest separation of the series: gpt-oss scored 5/6 in all five runs, failing only the varying-diffusivity trap and for the classic reason — pointwise diffusivity instead of flux-conserving interface treatment, its error matching our naive reference almost exactly. qwen3-coder scored 0/6 in all five runs, not through bad physics but through a violated output contract it was told about twice and never fixed. The mid-run repair loop lifted three gpt-oss runs from 0/3 to 3/3 on the self-check — its first unambiguous wins — while repairing nothing for qwen3-coder in five attempts.

Abstract

The first two trials in this series graded whether generated code runs (WP-03, apps) and how models reason (WP-04, data science). Neither could detect the failure numerical people actually fear: code that runs to completion, exits 0, produces plausible curves — and is wrong.

This trial is built to catch exactly that. The task is a heat-conduction simulator; the grade is the maximum difference between the submission’s output and the true solution of the differential equation, computed independently to ten decimal places. A config is passed or it is not. The score is configs passed, out of six the model never sees. No human judgment appears anywhere in the result.

The outcome is the sharpest of the series: gpt-oss:20b scored 5/6 in every one of five runs; qwen3-coder scored 0/6 in every one of five. And the single config gpt-oss always failed is the trap the task was designed around — its code runs cleanly there and returns temperatures that are silently wrong by 0.31 degrees.


1. The task

Simulate heat conduction in a 1D rod. A JSON config specifies rod length, thermal diffusivity, boundary conditions (both ends held at zero, or both insulated), one of three initial temperature profiles, and the positions and times at which temperature must be reported. Output is a CSV of time, x, temperature rows covering every requested combination, via a fixed interface:

python simulate.py --config <config.json> --output <out.csv>

Two design elements carry the weight.

The truth is computable independently. For uniform diffusivity the heat equation has exact Fourier-series solutions, accurate here to ~1e-10 — an answer key no numerical submission can dispute. Grading is max |submission − exact| < tolerance (0.01–0.02, temperatures of order 1), per config.

One hidden config closes the shortcut. A rod made of two joined materials (piecewise diffusivity) has no closed-form solution, so a submission cannot pass by implementing the textbook formula instead of an actual solver. Its reference comes from our own conservative Crank–Nicolson solver, verified by grid refinement and Richardson-extrapolated; the reference’s residual uncertainty (~1e-3) sits an order of magnitude under the scoring tolerance.

1.1 Fair feedback, hidden grading

The task repo ships three example configs with their exact expected outputs, plus a self-check script that runs the submission against all three and prints its true errors — the same comparison the hidden grader performs. The harness executes this check for the model after implementation and again after a repair turn, so every run sees genuine measurements of its own accuracy before final grading. The six graded configs stay outside the repository.

1.2 Calibration: the scale was fixed before any model ran

We implemented four submissions of known quality and graded them first:

CandidateHidden score
Careful conservative Crank–Nicolson6/6
Naive-but-stable explicit scheme4/6
Explicit scheme, hardcoded timestep0/6 — NaN on every config
Textbook single-mode formula, no solver1/6

The ladder separates carefulness, stability awareness, and genuine numerics from pattern-matching. It also debugged the task itself: our first configs placed a probe exactly on a material interface and another on a pulse edge at the earliest output time — points where even reference-quality solvers cannot meet tolerance. Both were adjusted before any model saw the task. That is the third time in this series that calibrating against our own best attempt caught an error we would otherwise have attributed to a model.

2. Method

Identical to the previous trials (Bakeoff harness, unattended, identical clean clones, five runs per model, interleaved), with the stage sequence plan → implement → repair → evaluate → docs. The implement and repair stages each end with the harness executing the self-check and writing its real output into the repository for the next stage to read. Stage cap 20 minutes; grading cap five minutes per config. All ten runs completed, in 3–32 minutes each.

3. Results

3.1 The self-check, mid-run

Immediately after implementAfter the repair turn
gpt-oss:20b3/3, 3/3, 0/3, 0/3, 0/33/3 in all five runs
qwen3-coder0/3 in all five0/3 in all five

This is the repair loop’s first unambiguous success in the series: three gpt-oss runs read the measured errors on their own output and fixed their solver to exact agreement. (In WP-04 the loop managed two repairs, four no-ops and one regression.) qwen3-coder saw the same style of feedback five times and repaired nothing.

3.2 The hidden grade

Run12345
gpt-oss:20b5/65/65/65/65/6
qwen3-coder0/60/60/60/60/6

Twenty-five of thirty configs passed against zero. On the calibration ladder, gpt-oss lands exactly where a competent-but-not-careful numericist sits: above the naive scheme (it passes the sharp-pulse config the naive scheme fails), below the careful one.

3.3 What gpt-oss got wrong — the silent kind

Every gpt-oss run failed the piecewise-diffusivity config, and in the way the config was designed to expose. Three of the five runs produced code that executes cleanly there and reports temperatures wrong by 0.310 — within 3% of our naive reference’s error (0.321), which assigns diffusivity pointwise per grid cell instead of conserving flux across the material interface. That is a real, recognizable numerical-methods error: the code contains an alpha_profile branch, looks correct, runs without complaint, and is wrong. The other two runs crashed outright while constructing their diffusivity grid.

This is the failure mode neither previous trial could see, caught by the strongest submission in the series, in the only place it occurs.

A variance note: the five gpt-oss runs collapsed into two implementation clusters with identical error signatures (runs 1/3/4 and runs 2/5) — strikingly less diverse than the same model’s three-build-tools spread on the app task.

3.4 What qwen3-coder got wrong — the contract, not the calculus

qwen3-coder’s zero is not a physics failure. Its solvers produce plausible temperature fields. It failed the output contract: the spec requires temperatures at the requested probe positions; qwen3-coder reports them at its own nearest grid nodes instead —

time,x,temperature
0.0,0.1501501502,1.214356745      <- requested probe was x = 0.15

— so the grader (and the self-check before it) finds none of the required (time, x) points. The self-check printed exactly this diagnosis, with example missing points, in every run — twice per run — and no run fixed it. One run additionally timed out on every config; two others, where coordinates happened to align, showed value errors above 1.1 as well.

This is the same signature as the previous trials wearing a lab coat: the failure lives at the interface between the model’s work and the process that consumes it.

4. Interpretation

On numerical competence, the models are not peers. gpt-oss wrote a genuine PDE solver five times out of five, repaired it to exact agreement on known answers when shown its errors, and failed only at a subtlety — conservative treatment of material interfaces — that trips up human practitioners. Its scores are reproducible to the third decimal across runs. qwen3-coder never delivered a gradeable simulation in five attempts.

The seams thesis survives its hardest test — by inverting the winner. In WP-03 and WP-04, gpt-oss reasoned well and shipped code that crashed at process boundaries: dependencies, pickling, proxies. A single-file numerical script has none of those boundaries, and with them gone, gpt-oss’s competence dominates and its engineering weakness vanishes. Meanwhile qwen3-coder — the model whose app ran and whose classifier scored — hit this task’s one seam (the output contract) and could not cross it. Across three trials, the predictor of failure has never been the hard part of the problem; it is whatever sits between the model’s code and the thing that runs it.

Feedback works when the error is legible to the model. The repair loop fixed solver accuracy (gpt-oss, numbers too far from known answers) and did not fix contract violation (qwen3-coder, points missing from output) even when the check named the missing points. Reading a measured error and reading a requirements violation are apparently different capabilities.

Silent wrongness is real and detectable. Three runs shipped code whose varying-diffusivity branch executes perfectly and answers wrongly by 30% of the temperature scale. Nothing short of comparison against ground truth would have caught it — not execution, not review, not the model’s own evaluation document.

5. Threats to validity

  • Effective sample is smaller than ten. The runs collapse into a few implementation clusters, so the five-for-five consistencies partly re-measure the same program.
  • The contract strictness is a design choice. A more forgiving grader could interpolate submissions’ grids onto the requested probes; ours requires the submission to do so, treating interface compliance as part of the task. The spec states the requirement, and the self-check reported the violation, but the zero should be read as “failed the take-home”, not “cannot solve the PDE”.
  • One equation, one dimension. Heat conduction is the friendliest PDE there is; stiffness, nonlinearity and higher dimensions are untested.
  • Tolerances are authored, though anchored by the calibration ladder rather than chosen after seeing model results.
  • Two quantized models, one machine, our stage decomposition — as in the previous trials.

6. Reproduction

scripts/make-heat-task.py builds the task, references and holdout and runs the calibration ladder; scripts/score-heatsim.py grades submissions; both in local-llm. The task repository contains the spec, examples with exact expected outputs, the self-check, and the stage definitions. All ten submission repositories, their mid-run check outputs, and full transcripts were retained. Scoring is deterministic — the same submission grades identically every time, unlike WP-04’s unseeded training.

7. Next

  1. A contract-repair experiment: does a dedicated feedback turn quoting the spec’s output requirement fix what the error-value feedback did not?
  2. A stiff or nonlinear variant, where stability failures cannot be solved by resolution alone.
  3. The frontier-model control, now more interesting for having a task where the local models bracket both extremes.

This paper completes the initial trio: applications (WP-03), data science (WP-04), and numerical computing — same two models, same hardware, three independent evaluators, one consistent lesson about where these models fail.