← Papers

Lab Z Working Papers · working

How well do local models do data science? A classifier take-home, graded automatically

Ten unattended runs of a spam-classifier take-home on a 12 GB GPU: how two local models explored the data, chose their metrics, reasoned about a planted leakage trap — and whether their pipelines ran.

We gave the same data-science take-home to two locally-hosted models — build a spam classifier from a messy 208,000-row corpus — five times each, unattended, on a 12 GB consumer GPU, and graded every submission automatically against a 50,000-row holdout. The task carries a planted temporal-leakage trap: a column that predicts the label almost perfectly in training and is always empty in production. As data scientists, the models are more capable than their reputation suggests: nine of nine design documents chose F1 over accuracy on the imbalanced data, eight of ten exploration scripts executed against the corpus, and one model identified and correctly excluded the leaked column in five of five runs, with sound reasoning stated in writing. As engineers, they fail at the last step: seven of ten submissions crashed before producing a single prediction, on undeclared imports, unpicklable classes, and a module path that moved in 2018. The three that ran scored 0.767, 0.766 and 0.510 against a careful-reference score of 0.801 — and the two best scores came from runs that planned to use the leaked column and were saved only by failing to implement their own design.

Abstract

WP-03 asked how well local models build applications. This paper asks the adjacent question: how well do they do data science — explore a dataset, pick features and metrics, avoid the classic traps, and ship a pipeline that produces defensible predictions?

The task is a deliberately ordinary take-home: 208,000 labelled messages, build a spam classifier. Underneath it is a planted temporal-leakage trap and an automatic grader, so the answer does not depend on anyone’s opinion. Each of two models ran the task five times, unattended, on a 12 GB GPU.

The headline splits cleanly in two. Their statistical judgment is better than you might expect: their engineering discipline loses the game anyway. Every design document chose the right metric for imbalanced data; one model spotted the leakage trap five times out of five. And seven of ten submissions crashed before emitting a single prediction.


1. The task

The spec reads like any screening exercise:

Build a classifier that decides whether an incoming message is spam. […] Report how well you expect it to perform, and explain how you arrived at that estimate. Handle imperfect data reasonably. Justify your choice of evaluation metric.

It also states, in plain language, when the classifier runs:

The classifier runs on arrival, as each message reaches the gateway […] It scores messages the moment they land — before anyone has looked at them.

That sentence is load-bearing, because of what is hidden in the data.

1.1 The corpus and its traps

208,000 messages with subject, body, sender, domain age, link/attachment counts, and reply count. Planted, with verified rates: 3% flipped labels, 4% near-duplicate rows, mojibake and HTML entities, and an 18% spam rate — enough imbalance that plain accuracy is a misleading metric, which makes metric choice itself a graded question.

The centerpiece is the column moderator_flagged. It correlates 0.97 with the label in the training data — because moderators flag spam after it arrives. Given the deployment sentence above, the column is always zero at prediction time. The 50,000-row holdout reflects that: a submission that leans on the column reports an excellent local score and collapses on real data.

1.2 Calibration: the trap punishes, the careful path pays

Reference implementations, measured before any model saw the task:

ApproachSelf-reported CV F1Holdout F1
Leak column only0.8800.000
Text only—0.767
Text + metadata, careful—0.801
Text + metadata + leak0.8820.479

The careless pipeline reports the best local number and delivers the worst real one — the inversion that makes the trap diagnostic rather than decorative. (An earlier corpus draft was too easy: disjoint vocabularies let bag-of-words score a perfect 1.000. A cross-contamination parameter now sets the difficulty, verified by these references.)

2. Method

Same hardware and harness as WP-03 (RTX 5070 12 GB; Bakeoff, unattended, identical clean clones), with three additions built for this trial:

  • Five runs per model, interleaved — run 1 of each model, then run 2 — so slow-moving machine state cannot correlate with the model under test.
  • A fixed grading interface. Every submission must expose predict.py --input <csv> --output <csv>. The scorer runs it against the holdout and computes F1. No human judgment enters the result.
  • An execution loop. The harness runs the model’s own exploration script and hands back the real output; later it runs train.py, captures the traceback, and grants a repair turn. Stages: explore → decide → implement → repair → evaluate → docs.

3. How they explored

Both models, in every run, began by writing an exploration script rather than guessing at the data — and those scripts were executed against the real corpus:

Exploration script ran (exit 0)
gpt-oss:20b3 / 5
qwen3-coder4 / 5

Eight of ten held up against a 208,000-row file they had seen only as a ten-line excerpt. The failures were small-bore (one script crashed calling .remove() on an absent list element) and non-fatal, since the next stage receives whatever output was produced before the crash.

4. Their statistical judgment

This is the strongest section for the models, and it deserves to be reported as such.

Metric choice: 9 of 9 design documents chose F1 (or precision/recall), none chose accuracy alone, and every one explicitly mentioned class imbalance. On a dataset engineered to punish accuracy-reporting, neither model took the bait even once. (The tenth run produced no design document at all — see §6.)

The leakage trap split the models:

gpt-oss:20bqwen3-coder
Excluded the leaked column, with correct reasoning5 / 52 / 4
Planned to use it0 / 52 / 4

gpt-oss stated the exclusion the way a reviewer would want it stated, five times out of five:

moderator_flagged | Flag set after moderation; not available at arrival. | Excluded.

qwen3-coder, in its two failing runs, asserted the opposite as fact —

moderator_flagged - Available at message arrival

— while, in the same document, excluding the harmless received_at timestamp because it “could introduce data leakage.” The model owns the vocabulary of leakage and applied it to the wrong column. This is the single clearest capability difference the trial produced, and it survives five repetitions.

5. Did the pipelines run?

Seven of ten submissions crashed before producing one prediction.

gpt-oss:20bqwen3-coder
predict.py executed and was scored1 / 52 / 5

The failure taxonomy is mundane, which is the point — every entry would surface on a single execution:

FailureRuns
SimpleImputer imported from sklearn.preprocessing — it moved to sklearn.impute in 20181
Custom transformer class pickled in train.py, unpicklable from predict.py2
Training saves one artifact structure, prediction expects another (KeyError)1
Undeclared third-party imports1
No predict.py produced at all2

The pickle failures reward attention: a class defined inside a training script cannot be reconstructed by a separate prediction process. The code looks correct, reads correctly in review, and fails only when the two scripts run as two processes — which is exactly the condition a submission never meets while being read.

One failure deserves its own note. A submission crashed on a missing bs4; our scorer does not install dependencies, so this looked like our defect. We installed the package and re-scored, expecting to recover the run. It then crashed on unidecode. Repairing the harness’s side of the problem exposed a second instance of the same defect beneath it.

6. The scores — and what they cannot say

SubmissionHoldout F1PrecisionRecall
qwen3-coder run 20.7670.6620.910
qwen3-coder run 30.7660.6620.909
gpt-oss run 40.5100.3660.839
careful reference0.801

Three scores are not a ranking, and we do not offer one. A single additional crash in either direction reverses the ordering. What the scores do show:

  • Nobody reached the careful baseline. The best submissions match the text-only reference (0.767) exactly; none exploited the metadata that lifts a careful pipeline to 0.801.
  • The one gpt-oss run that executed was poorly calibrated, flagging 40% of messages as spam against a true rate of 18% — high recall bought with precision of 0.37. Its correct reasoning about metrics and leakage did not translate into a well-tuned artifact.
  • Scoring is approximate at the third decimal: the same gpt-oss submission scored 0.502 and 0.510 on successive trainings because nothing fixes a random seed.

6.1 Saved by their own bug

The two best-scoring runs deserve a closer look, because they are the trial in miniature. Both qwen3-coder runs 2 and 3 planned to use the leaked column — it sits in their declared feature lists, alongside an import of ColumnTransformer to combine metadata with text. Then their training code does this:

X_text = X['text']
model.fit(X_text, y)

The feature list is discarded; only text is fitted. That is why both land on 0.767 — the text-only reference — to three decimals. The submissions that scored highest were protected from their own worst decision by failing to implement it. Stated design and executed behaviour diverged inside a single submission, in the direction that happened to be lucky.

7. The repair loop, measured

This trial added the feedback mechanism WP-03’s verdict called for: run the model’s training code, hand back the real traceback, allow one repair turn. Across ten runs:

OutcomeCount
Repaired a failing pipeline2
No change needed or made4
Broke a working pipeline1

One feedback turn genuinely rescued two runs. It also regressed one — a 20% regression rate that anyone contemplating an unsupervised repair loop should price in.

8. Verdict

How good are local models at data science? Split the question and the answer is sharp.

At the analyst’s desk — reading data, choosing metrics, reasoning about deployment-time information — they are genuinely competent, and gpt-oss in particular behaved like a careful junior data scientist five runs out of five: right metric, right exclusions, reasons stated in writing.

At the keyboard — turning that analysis into a pipeline another process can run — they fail more often than they succeed, on errors a first execution would catch. The bottleneck is not statistical understanding. It is that nothing in their workflow runs the code before delivering it, and one repair turn only partially compensates.

The practical reading for anyone using these models today: trust the analysis draft, distrust the pipeline, and put your review effort at the boundary between processes — serialization, dependencies, interfaces — where every fatal defect in this trial lived. The same advice WP-03 reached for applications, arrived at here by automatic measurement.

9. Threats to validity

  • The F1 table rests on three scores (§6) and is reported accordingly.
  • One task in one domain. The crash causes are generic Python engineering and should transfer; the reasoning results may not.
  • The trap was authored by the evaluator. moderator_flagged was designed to be catchable from the spec; a subtler leak might separate the models differently.
  • The harness installs no dependencies; §5 argues this changed no outcome here, but it would matter for a submission one import short.
  • Unseeded training makes scores approximate even for a fixed submission.
  • Two quantized models, one machine, our stage decomposition.

10. Reproduction

Generator (scripts/generate-messages.py, deterministic and seeded), runner (scripts/oneshot.ps1, per-stage execution), and scorer (scripts/score-classifier.py, no sklearn dependency so submissions cannot disturb it) are in local-llm. The holdout and answer key live outside the task repository. All ten submission repositories and full per-stage transcripts were retained.

11. Next

  1. Install declared dependencies before scoring — removes the one failure class with ambiguous attribution.
  2. Multi-turn repair, measuring whether the regression rate grows with iterations.
  3. A numerical simulation task — the converse failure mode: code that runs perfectly and is silently wrong, which neither trial has probed.
  4. A frontier-model control. “Seven of ten crashed” has no anchor until we know what a strong model does on the identical task.