I ran the same coding prompt at local models 22 times. Five model families, six weight sets. Exactly one result worked, and it doesn’t reproduce. Zero of three fresh seeds repeat it.
But here’s the number that changed how I test things. All 22 runs scored 11 or better out of 13 on static source checks. Six scored a perfect 13/13 while failing behaviorally. Reading the code can’t tell you whether the code runs.
The setup
One prompt, one file, no dependencies. Build a falling-sand simulator as a single HTML page.
It’s a good test. Small enough to one-shot, and it has real invariants. Particles have to fall. A painted stroke has to spread. Mass must not be conserved when acid dissolves sand. A specific color has to appear on the canvas.
Two graders:
- Static. 13 checks against the source text. Does it use
requestAnimationFrame? Is there a double-buffered grid? Does it bind pointer events? Does the mandated hex color appear? - Behavioral. 19 assertions driven through real pointer events in headless Chromium, then a pixel scan of the canvas for that color.
I wrote the static grader first because it was easy. That was the mistake. Running both is what exposed it.
The results
| Measure | Result |
|---|---|
| Local one-shot runs | 22 |
| Behavioral passes | 1 |
| That pass, re-run at 3 fresh seeds | 0 passes |
| Runs scoring ≥11/13 static | 22 (all of them) |
| Runs scoring 13/13 static while failing behaviorally | 6 |
Cloud reference (claude-sonnet-5) | PASS, $0.22, 81 s |
A grader that hands every candidate 11-plus out of 13 is checking formatting. And I’d have quoted it as evidence, because 13/13 reads like proof.
What the failures looked like
None of them were syntax errors. Nothing failed to parse.
They were simulations where sand never leaves the painted stroke. Where the particle count is exactly conserved because the update loop advances zero cells. Where the acid renders and does nothing.
Every one of those files contains a plausible physics function. The function is never reached. Or it’s reached with arguments that make it a no-op. The code is structurally right and dynamically dead. Static checks can’t see that. Neither can you, by reading it.
Three levers, and what they bought
Sampling settings: nothing. Seven profiles. The one that produced the single pass produced three failures at fresh seeds. A lucky sample, not a setting.
Architecture: one clean result. The dense 32B model was the only local one whose failures were legible. I compared it against ~3B-active mixture-of-experts models. It runs 11.8× the active compute per token. Capacity spent on active parameters bought correctness. Capacity spent on sparse total parameters bought tokens per second I didn’t need.
Quantization headroom: negative. The same dense model at Q5 scored 0 of 4. At Q4 it scored 1 of 4. It also cost 20% of decode speed and 3 GB more resident memory. More bits per weight, worse output, slower. Whatever that is, it isn’t a quality dial.
Multi-turn repair is where it stops
I gave a failing run three rounds of targeted repair with the defects named.
Round one produced a genuine 13-line fix with the correct root cause. Round two fixed only the defect I named explicitly and ignored the other. Round three got the function names and the invariant in prose. It returned a byte-identical file and reported success.
That’s worse than a bad draft. A bad draft stops you. A fixed point that reports success is what an unattended loop ships.
One warning if you write a harness like this. My helper functions vline() and hline() only return arrays of points. paint() and stroke() are what dispatch the events. Calling vline() alone is a silent no-op. It produces a clean run and plausible zeros.
I nearly published a whole class of “failures”. They were my own rig not clicking anything. Your harness needs a test that fails. Otherwise you’re grading the grader on vibes too.