Lab Z Working Papers · working
How well do local models do data science? A classifier take-home, graded automatically
Ten unattended runs of a spam-classifier take-home on a 12 GB GPU: how two local models explored the data, chose their metrics, reasoned about a planted leakage trap — and whether their pipelines ran.
Abstract
WP-03 asked how well local models build applications. This paper asks the adjacent question: how well do they do data science — explore a dataset, pick features and metrics, avoid the classic traps, and ship a pipeline that produces defensible predictions?
The task is a deliberately ordinary take-home: 208,000 labelled messages, build a spam classifier. Underneath it is a planted temporal-leakage trap and an automatic grader, so the answer does not depend on anyone’s opinion. Each of two models ran the task five times, unattended, on a 12 GB GPU.
The headline splits cleanly in two. Their statistical judgment is better than you might expect: their engineering discipline loses the game anyway. Every design document chose the right metric for imbalanced data; one model spotted the leakage trap five times out of five. And seven of ten submissions crashed before emitting a single prediction.
1. The task
The spec reads like any screening exercise:
Build a classifier that decides whether an incoming message is spam. […] Report how well you expect it to perform, and explain how you arrived at that estimate. Handle imperfect data reasonably. Justify your choice of evaluation metric.
It also states, in plain language, when the classifier runs:
The classifier runs on arrival, as each message reaches the gateway […] It scores messages the moment they land — before anyone has looked at them.
That sentence is load-bearing, because of what is hidden in the data.
1.1 The corpus and its traps
208,000 messages with subject, body, sender, domain age, link/attachment counts, and reply count. Planted, with verified rates: 3% flipped labels, 4% near-duplicate rows, mojibake and HTML entities, and an 18% spam rate — enough imbalance that plain accuracy is a misleading metric, which makes metric choice itself a graded question.
The centerpiece is the column moderator_flagged. It correlates 0.97 with the
label in the training data — because moderators flag spam after it arrives.
Given the deployment sentence above, the column is always zero at prediction
time. The 50,000-row holdout reflects that: a submission that leans on the
column reports an excellent local score and collapses on real data.
1.2 Calibration: the trap punishes, the careful path pays
Reference implementations, measured before any model saw the task:
| Approach | Self-reported CV F1 | Holdout F1 |
|---|---|---|
| Leak column only | 0.880 | 0.000 |
| Text only | — | 0.767 |
| Text + metadata, careful | — | 0.801 |
| Text + metadata + leak | 0.882 | 0.479 |
The careless pipeline reports the best local number and delivers the worst real one — the inversion that makes the trap diagnostic rather than decorative. (An earlier corpus draft was too easy: disjoint vocabularies let bag-of-words score a perfect 1.000. A cross-contamination parameter now sets the difficulty, verified by these references.)
2. Method
Same hardware and harness as WP-03 (RTX 5070 12 GB; Bakeoff, unattended, identical clean clones), with three additions built for this trial:
- Five runs per model, interleaved — run 1 of each model, then run 2 — so slow-moving machine state cannot correlate with the model under test.
- A fixed grading interface. Every submission must expose
predict.py --input <csv> --output <csv>. The scorer runs it against the holdout and computes F1. No human judgment enters the result. - An execution loop. The harness runs the model’s own exploration script and
hands back the real output; later it runs
train.py, captures the traceback, and grants a repair turn. Stages: explore → decide → implement → repair → evaluate → docs.
3. How they explored
Both models, in every run, began by writing an exploration script rather than guessing at the data — and those scripts were executed against the real corpus:
| Exploration script ran (exit 0) | |
|---|---|
gpt-oss:20b | 3 / 5 |
qwen3-coder | 4 / 5 |
Eight of ten held up against a 208,000-row file they had seen only as a
ten-line excerpt. The failures were small-bore (one script crashed calling
.remove() on an absent list element) and non-fatal, since the next stage
receives whatever output was produced before the crash.
4. Their statistical judgment
This is the strongest section for the models, and it deserves to be reported as such.
Metric choice: 9 of 9 design documents chose F1 (or precision/recall), none chose accuracy alone, and every one explicitly mentioned class imbalance. On a dataset engineered to punish accuracy-reporting, neither model took the bait even once. (The tenth run produced no design document at all — see §6.)
The leakage trap split the models:
gpt-oss:20b | qwen3-coder | |
|---|---|---|
| Excluded the leaked column, with correct reasoning | 5 / 5 | 2 / 4 |
| Planned to use it | 0 / 5 | 2 / 4 |
gpt-oss stated the exclusion the way a reviewer would want it stated, five
times out of five:
moderator_flagged| Flag set after moderation; not available at arrival. | Excluded.
qwen3-coder, in its two failing runs, asserted the opposite as fact —
moderator_flagged- Available at message arrival
— while, in the same document, excluding the harmless received_at timestamp
because it “could introduce data leakage.” The model owns the vocabulary of
leakage and applied it to the wrong column. This is the single clearest
capability difference the trial produced, and it survives five repetitions.
5. Did the pipelines run?
Seven of ten submissions crashed before producing one prediction.
gpt-oss:20b | qwen3-coder | |
|---|---|---|
predict.py executed and was scored | 1 / 5 | 2 / 5 |
The failure taxonomy is mundane, which is the point — every entry would surface on a single execution:
| Failure | Runs |
|---|---|
SimpleImputer imported from sklearn.preprocessing — it moved to sklearn.impute in 2018 | 1 |
Custom transformer class pickled in train.py, unpicklable from predict.py | 2 |
Training saves one artifact structure, prediction expects another (KeyError) | 1 |
| Undeclared third-party imports | 1 |
No predict.py produced at all | 2 |
The pickle failures reward attention: a class defined inside a training script cannot be reconstructed by a separate prediction process. The code looks correct, reads correctly in review, and fails only when the two scripts run as two processes — which is exactly the condition a submission never meets while being read.
One failure deserves its own note. A submission crashed on a missing bs4; our
scorer does not install dependencies, so this looked like our defect. We
installed the package and re-scored, expecting to recover the run. It then
crashed on unidecode. Repairing the harness’s side of the problem exposed a
second instance of the same defect beneath it.
6. The scores — and what they cannot say
| Submission | Holdout F1 | Precision | Recall |
|---|---|---|---|
qwen3-coder run 2 | 0.767 | 0.662 | 0.910 |
qwen3-coder run 3 | 0.766 | 0.662 | 0.909 |
gpt-oss run 4 | 0.510 | 0.366 | 0.839 |
| careful reference | 0.801 |
Three scores are not a ranking, and we do not offer one. A single additional crash in either direction reverses the ordering. What the scores do show:
- Nobody reached the careful baseline. The best submissions match the text-only reference (0.767) exactly; none exploited the metadata that lifts a careful pipeline to 0.801.
- The one gpt-oss run that executed was poorly calibrated, flagging 40% of messages as spam against a true rate of 18% — high recall bought with precision of 0.37. Its correct reasoning about metrics and leakage did not translate into a well-tuned artifact.
- Scoring is approximate at the third decimal: the same gpt-oss submission scored 0.502 and 0.510 on successive trainings because nothing fixes a random seed.
6.1 Saved by their own bug
The two best-scoring runs deserve a closer look, because they are the trial in
miniature. Both qwen3-coder runs 2 and 3 planned to use the leaked column —
it sits in their declared feature lists, alongside an import of
ColumnTransformer to combine metadata with text. Then their training code
does this:
X_text = X['text']
model.fit(X_text, y)
The feature list is discarded; only text is fitted. That is why both land on 0.767 — the text-only reference — to three decimals. The submissions that scored highest were protected from their own worst decision by failing to implement it. Stated design and executed behaviour diverged inside a single submission, in the direction that happened to be lucky.
7. The repair loop, measured
This trial added the feedback mechanism WP-03’s verdict called for: run the model’s training code, hand back the real traceback, allow one repair turn. Across ten runs:
| Outcome | Count |
|---|---|
| Repaired a failing pipeline | 2 |
| No change needed or made | 4 |
| Broke a working pipeline | 1 |
One feedback turn genuinely rescued two runs. It also regressed one — a 20% regression rate that anyone contemplating an unsupervised repair loop should price in.
8. Verdict
How good are local models at data science? Split the question and the answer is sharp.
At the analyst’s desk — reading data, choosing metrics, reasoning about
deployment-time information — they are genuinely competent, and gpt-oss in
particular behaved like a careful junior data scientist five runs out of five:
right metric, right exclusions, reasons stated in writing.
At the keyboard — turning that analysis into a pipeline another process can run — they fail more often than they succeed, on errors a first execution would catch. The bottleneck is not statistical understanding. It is that nothing in their workflow runs the code before delivering it, and one repair turn only partially compensates.
The practical reading for anyone using these models today: trust the analysis draft, distrust the pipeline, and put your review effort at the boundary between processes — serialization, dependencies, interfaces — where every fatal defect in this trial lived. The same advice WP-03 reached for applications, arrived at here by automatic measurement.
9. Threats to validity
- The F1 table rests on three scores (§6) and is reported accordingly.
- One task in one domain. The crash causes are generic Python engineering and should transfer; the reasoning results may not.
- The trap was authored by the evaluator.
moderator_flaggedwas designed to be catchable from the spec; a subtler leak might separate the models differently. - The harness installs no dependencies; §5 argues this changed no outcome here, but it would matter for a submission one import short.
- Unseeded training makes scores approximate even for a fixed submission.
- Two quantized models, one machine, our stage decomposition.
10. Reproduction
Generator (scripts/generate-messages.py, deterministic and seeded), runner
(scripts/oneshot.ps1, per-stage execution), and scorer
(scripts/score-classifier.py, no sklearn dependency so submissions cannot
disturb it) are in local-llm. The
holdout and answer key live outside the task repository. All ten submission
repositories and full per-stage transcripts were retained.
11. Next
- Install declared dependencies before scoring — removes the one failure class with ambiguous attribution.
- Multi-turn repair, measuring whether the regression rate grows with iterations.
- A numerical simulation task — the converse failure mode: code that runs perfectly and is silently wrong, which neither trial has probed.
- A frontier-model control. “Seven of ten crashed” has no anchor until we know what a strong model does on the identical task.