py-harness

ask, test, fix, add

GitHub

Experiments

Yauhen Bichel · 6 September 2026 · github.com/YauhenBichel/py-harness

Abstract

A local helper plus a small open model was asked to do four daily Python jobs on one laptop: answer a question, write a test, fix a bug, add one function. Dates: 29–30 August and 5 September 2026.

0.5B means a model with about 500 million weights — Qwen2.5-Coder-0.5B, plus an optional public LoRA. It is the tiny Hub demo, not daily work. 8B means about 8 billion weights — Ollama llama3.1:8b, the default. 7B is qwen2.5-coder:7b.

The 0.5B adapter is a style prior, not an agent (held-out vibe 0 / 4; greedy LoRA 0 / 54). One traceback repair lifts exact-stdout from 7 / 54 to 12 / 54. Daily jobs on the 8B reached 9 / 9 the same evening the 7B coder scored 7 / 9. The helper moved the four Start commands from 0 / 4 to 4 / 4 by doing compiler jobs itself. The everyday-ready bar — beat a plain 8B at the next step and at a real bug the helper cannot write — is not met. This paper does not report a SWE-bench score. A hosted 32B was measured once the benchmark itself was repaired: it scores 9 / 10 on tier 3, where before the repair it scored 1 / 10. A day of repairs to the harness and the benchmark then took the whole suite from 51 / 75 to 64 / 75 without asking more of the model.

12 / 18tiny 0.5B (500 million weights), four drafts then a later loop
0 / 54same 0.5B + LoRA, one greedy try each
8 / 15daily 8B (8 billion weights) picked the right first step
9 / 9daily 8B jobs, evening of 5 Sep

Introduction

The question is whether a laptop helper and a small open model can do everyday Python without a hosted agent. Two sizes are easy to mix up: the public 0.5B (500 million weights, a style prior) and the daily 8B (8 billion weights, what ask / run call).

“Everyday-ready” means: beat a plain 8B at reading the next step, and at fixing a real bug the helper cannot write. The first “real bug” cell was a whole-line return 0.0 on a named sum — the helper can write that, so it is no longer a model job.

This note is the paper. The long run log is the appendix. Cite the software or a score with Cite. The machine: Bench record.

The loop is one Action: block, then tools — a tight read-act loop, not a free shell [1]. SWE-agent showed that the interface around the same model moves the score [2]: previous best retrieval-only 3.8%, agent 12.5%; same model 1.3% → 12.5%. The first-run four jobs here were 0 / 4 then 4 / 4 after the helper. Same shape, smaller tree.

CodeAct lets the model emit Python as the action [3]. This project uses named actions and a write limit. A green suite that never called the bug is not done [4]. run may send one traceback back [5] [6]. Small models can be workers at 7–8B [7]; that is not a claim that a 0.5B style adapter is an agent. SWE-bench is the field’s ruler [8] and the wrong ruler for a one-folder laptop helper. This paper does not report a SWE-bench score.

The longer list of papers: References.

Methods

Machine. One Apple M3 Pro laptop, 18 GB unified memory. Ollama for most cells; MLX for the 0.5B sample-and-run cell [9].

Models. Size names are weight counts, not versions.

Name Weights What it is Role
0.5B 500 million Qwen2.5-Coder-0.5B [10]. Ollama qwen2.5-coder:0.5b or MLX 4-bit. Optional LoRA [11] YauhenBichel/python-vibe-0.5b, step 100 Style prior. Not daily ask / run
7B 7 billion qwen2.5-coder:7b unless a table names another tag Same-night comparison
8B 8 billion Ollama llama3.1:8b [12] Daily default
32B 32 billion Hosted Qwen2.5-Coder-32B-Instruct One GPU comparison, after the fence fix

A “clean” 8B is the same 8 billion weights with no agent system prompt and no loop. The 0.5B file on disk is about 400 MB; the 8B is about 4.9 GB.

Tasks. Four daily jobs on demo/orders unless a table names another fixture. A case counts only if the function runs and does the job — not if a file appeared, and not if the run said done. Writes stay inside the named folder. There is no general shell.

Scoring. One run unless the table says otherwise. A gap of one or two cases is noise. Compiler-bound cells (NameError subtotl, whole-line return 0 on a named sum) finish with no model. They are not model scores. Replay scripts live in scripts/measure/. The full cells are in the appendix.

Results

0.5B — 500 million weights

Held-out vibe (weekday, count-md, jsonl, docstring): 0 / 4. Parsed Action: that day: 0 / 2. Exact stdout, 18 scripts × 3, Ollama qwen2.5-coder:0.5b [13]: base 7 / 54, one traceback repair 12 / 54. Greedy LoRA: 0 / 54. Four drafts then a later loop: 12 / 18 [14]. Sampling found a different set, not a superset.

Daily 8B and 7B

Same evening, 5 September 2026. llama3.1:8b: 9 / 9. qwen2.5-coder:7b: 7 / 9. First-run four on demo/orders: 0 / 4, then 4 / 4 after the helper did the compiler jobs. Live first-Action parse: 8 / 15.

Everyday-ready bar

Beat a plain 8B at the next step and at a ≥1 KB logic fix the helper cannot write (clip). Last recorded night: harness parse 8 / 15, clean 8B 0 / 15; harness fix 0 / 3, clean 8B 3 / 3. Not everyday-ready.

A real tree

4,580 first-party files. Reads worked. Write a test or add a function: 1 / 12.

The instrument, then a hosted 32B

A run that stops to ask needs someone to answer. Local 8B almost never asks; a 7B coder asks in eleven of twenty tier-3 runs. The same week, a hosted 32B wrapped drafts in markdown fences and the fence reached the file. Before the harness stripped it: 1 / 10. After: 9 / 10. Local 8B does not fence, so the same fix leaves it at 10 / 20. The tables are in the appendix [16].

Discussion

The helper is load-bearing for jobs it can finish without a model [2]. First-run four went 0 / 4 to 4 / 4 once the compiler wrote the NameError, the test, and total_lines. The model still has to do the rest. Daily llama3.1:8b is 9 / 9 on small fixtures and 8 / 15 on first-step parse. The everyday-ready bar asks for a ≥1 KB logic fix the helper cannot write (clip). Harness 0 / 3, clean 8B 3 / 3. The loop helps the 8B pick an Action and does not get a patch on clip.

The 0.5B is a style prior. A traceback fixes NameError and SyntaxError, not logic [6]. A 7B coder is close (7 / 9) and not better. Writing on a 4,580-file tree is 1 / 12. Two models of different lineage hit the same wall (51 vs 50 of 75). That pair was measured before the instrument was repaired, so the number to trust is the shape, not the score. Raising the count that works is the target.

The most transferable result is about the instrument, not any model. Both faults had one shape: the harness had specialised to the single model it runs itself. Local weights rarely ask, so nobody noticed there was no one to answer; they do not fence, so nobody noticed the fence reached the file. A hosted 32B looked incapable at 1 / 10 and scores 9 / 10 with four backticks removed [16]. Measure a second model early. It is the cheapest way to find the assumptions the first one hides.

Limitations

One laptop, 18 GB unified memory. One run unless a table says otherwise; gaps of one or two cases are noise. Several cells are compiler binds, not model writes. The instrument was wrong for part of the month: unanswered ask stops, and markdown fences reaching the file. Both are fixed, and every model number from before 6 September is unsafe. Tier 3 has two cases, so its run-to-run spread is wide: the same 8B on the same code scored 14 / 20 and 10 / 20 on different nights.

This paper does not report a SWE-bench score [8]. The public numbers are four jobs on demo/orders and a 4,580-file write rate of 1 / 12. No hosted chat product is named. No claim that the 0.5B LoRA audited a real repository.

Conclusion

Keep llama3.1:8b. Do not train more 0.5B steps. Do not switch the default to a 7B coder or a Hub GGUF that misses the 180s generate cap. The helper should keep finishing compiler jobs. The everyday-ready bar stays: beat a plain 8B at the next step and at a real bug the helper cannot write. It does not, yet.

References

  1. Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. International Conference on Learning Representations. arXiv:2210.03629
  2. Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). SWE-agent: Agent-computer interfaces enable automated software engineering. Advances in Neural Information Processing Systems. arXiv:2405.15793
  3. Wang, X., Chen, Y., Yuan, L., Zhang, Y., Li, Y., Peng, H., & Ji, H. (2024). Executable code actions elicit better LLM agents. International Conference on Machine Learning. arXiv:2402.01030
  4. Liu, J., Xia, C. S., Wang, Y., & Zhang, L. (2023). Is your code generated really correct? Rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems. arXiv:2305.01210
  5. Shinn, N., Cassano, F., Labash, B., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems. arXiv:2303.11366
  6. McAndrews, C. J. (2026). Feedback over form: Why execution feedback matters more than pipeline topology in 1–3B code generation. arXiv:2604.21950
  7. Belcak, P., Heinrich, G., Diao, S., Fu, Y., Dong, X., Muralidharan, S., Lin, Y. C., & Molchanov, P. (2025). Small language models are the future of agentic AI. arXiv:2506.02153
  8. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2024). SWE-bench: Can language models resolve real-world GitHub issues? International Conference on Learning Representations (oral). arXiv:2310.06770
  9. Bichel, Y. (2026). Bench record. In py-harness. yauhenbichel.github.io/py-harness/investigations/bench-record/
  10. Hui, B., Yang, J., Cui, Z., Yang, J., et al. (2024). Qwen2.5-Coder technical report. arXiv:2409.12186
  11. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations. arXiv:2106.09685
  12. Grattafiori, A., et al. (2024). The Llama 3 herd of models. arXiv:2407.21783
  13. Bichel, Y. (2026, September 5). 0.5B exact-stdout eval. In py-harness experiments. yauhenbichel.github.io/py-harness/investigations/held-out-exec-eval/
  14. Bichel, Y. (2026, September 5). 0.5B sample-and-run. In py-harness experiments. yauhenbichel.github.io/py-harness/investigations/sample-and-run/
  15. Bichel, Y. (2026). py-harness [Computer software]. github.com/YauhenBichel/py-harness
  16. Bichel, Y. (2026, September 6). The fence was the whole story. In py-harness experiments. yauhenbichel.github.io/py-harness/investigations/the-fence/

APA and BibTeX for this software: Cite.

Appendix: full tables

Each cell below is one question, what I typed, and what happened. The paper above is the result.

The 0.5B as daily work

The tiny model: 500 million weights, not the daily 8B.

Example. Public adapter YauhenBichel/python-vibe-0.5b on Qwen2.5-Coder-0.5B. Ask it for a weekday-name helper, a markdown file counter, a jsonl line, a docstring. Ask it to emit Action:.

Result

What I asked What I got
Held-out vibe (weekday, count-md, jsonl, docstring) 0 / 4
Parsed Action: that day 0 / 2
Walk a hundred stub files A hundred “no issues”. Not a review
400-step QLoRA Overfit after step 100. Hub file is that checkpoint

That 500-million-weight model is a style prior. It is not daily work. I am not training more 0.5B steps.

Write-up: 0.5B vibe review · Everyday laptop.

Exact stdout on the 0.5B

Same tiny Qwen2.5-Coder (500 million weights). Not llama3.1:8b.

Example. Eighteen held-out scripts. None of the 45 train prompts. Extract the Python block, run it, demand an exact line. Repeat each task three times. Then send the traceback back once.

Result, 5 September 2026, Ollama qwen2.5-coder:0.5b:

Variant Passed
base 7 / 54
one traceback repair 12 / 54

24 of 54 base runs crashed (often sys used, never imported). 23 printed the right number with extra words (Clamped value: 10). Eleven of eighteen tasks never passed. LoRA was not measured (mlx-lm missing). Unit tests for the checkers passed.

Write-up: 0.5B exact-stdout eval. Cite: Cite. Related work: References.

Sample four drafts, then greedy

Example. Same 18 scripts on MLX Qwen2.5-Coder-0.5B-Instruct-4bit. First, up to four independent drafts at temperature 0.7. Then one greedy draft, three repeats, with and without the step-100 LoRA.

Result, 5 September 2026

Variant Four drafts / 18 Greedy unique / 18 Greedy runs / 54
base 6 2 6
one traceback repair 9 3 9
LoRA 2 0 0
LoRA + repair 6 0 0

Sampling found a different set, not a superset. Only one of the +3 from 6 to 9 is a traceback fix; the rest is a new draw. Greedy LoRA printed style notes, not scripts. A later loop (prepend datetime, say when stdout is wrong, one 8B hint) scored 12 / 18. Zero of those twelve were a hint-repair. Stop spending hours on this board.

Write-up: 0.5B sample-and-run.

8B daily jobs

Example. 5 September 2026. Ollama llama3.1:8b. Three jobs that are not the built-in NameErrors, each three times, after the harness started running the suite following a write.

Job What I asked Passed
Write tests write tests for apply_discount in src/app.py 3 / 3 (harness wrote the AAA test)
Add a function add a function clamp(value, lo, hi) … and a unit test 3 / 3 (8B)
Logic bug fix compute_total in src/app.py so it sums the rows 2 / 3 (one hit the step budget)

8 / 9. The miss was a logic-bug run that spent twelve steps without a green suite. Replay one of the wins: Live demo (daily recording).

Everyday-ready is still the older bar: beat a clean 8B on parse and a real ≥1 KB fix the model wrote. This table is the daily loop on small fixtures, not that bar. The first ≥1 KB cell is retired below.

Same-night daily jobs, 7B coder

Example. 5 September 2026, evening. Same script (scripts/measure/eval_daily.py), same twelve steps, same fixtures. llama3.1:8b remasured, then qwen2.5-coder:7b.

Model Write tests Add clamp Logic bug Passed
llama3.1:8b 3 / 3 3 / 3 3 / 3 9 / 9
qwen2.5-coder:7b 3 / 3 1 / 3 (two ask stops) 3 / 3 7 / 9

The two-case gap is inside the noise this page already named. I did not switch the default.

The logic-bug 3 / 3 on both sides is the compiler bind: a whole-line return 0 on a named sum. Same class as the retired ≥1 KB cell. It is not a model writing return sum(rows).

Replay: PYTHONPATH=src python3 scripts/measure/eval_daily.py --model qwen2.5-coder:7b.

The 7B clip bar the same evening:

Check Harness 7B Clean 7B
Live parse 10 / 15 1 / 15
≥1 KB clip fix 0 / 3 (steps; writes [] × 3; turns non-empty) 3 / 3 (one-shot)

everyday_ready stayed false. The 7B coder speaks more first Actions than a clean 7B and still does not write clip. Same wall as the 8B clip remasure.

More 7B–8B on disk, 5 September 2026

Example. Same script (scripts/measure/eval_daily.py), same twelve steps, same fixtures. Tags now on this laptop: DeepSeek-Coder 6.7B, StarCoder2 7B, CodeLlama 7B Python, OpenCoder 8B, SWE-agent-LM 7B. OpenCoder and SWE-agent-LM came from Hub GGUFs (scripts/weights/import_hf_ollama.py), not ollama pull.

Result

Model Write tests Add clamp Logic bug Passed
llama3.1:8b (same night) 3 / 3 3 / 3 3 / 3 9 / 9
qwen2.5-coder:7b (same night) 3 / 3 1 / 3 3 / 3 7 / 9
deepseek-coder:6.7b 3 / 3 (compiler) 1 pass, 1 steps, then 180s timeout not run incomplete
deepseek-coder:6.7b (empty VRAM) 3 / 3 (compiler) 1 pass (steps), then 180s timeout not run incomplete
starcoder2:7b 3 / 3 (compiler) 180s timeout on the first generate not run incomplete
codellama:7b-python 3 / 3 (compiler) 180s timeout on the first generate not run incomplete
opencoder:8b 3 / 3 (compiler) 180s timeout on the first generate not run incomplete
swe-agent-lm:7b 3 / 3 (compiler) 180s timeout on the first generate not run incomplete
swe-agent-lm:7b (empty VRAM) 3 / 3 (compiler) 180s timeout on the first generate not run incomplete

Write-tests 3 / 3 on the extra tags is the harness writing the AAA test. The model is not called. The first job that does call it is clamp, and a cold 7B–8B load plus one generate burned the 180s Ollama cap. DeepSeek got one clamp through, then steps, then the same cap.

Warm remasure, same evening. Each extra tag was loaded first (keep_alive 30 minutes). The pass finished. OpenCoder and StarCoder2: warmup curl got 0 bytes in 300s, then clamp hit 180s. SWE-agent-LM was already in memory and still hit 180s on the first clamp generate. CodeLlama’s warmup returned, then clamp still hit 180s. DeepSeek got one clamp through, then the same cap — same shape as the cold pass. So this is not only a cold start. Write-tests stayed 3 / 3 (compiler). Not a score.

One-word generate, same evening. Prompt: Reply with the single word ok. Cap 180s. No daily job.

Model Reply Wall
llama3.1:8b (already loaded) Ok 3.9 s
qwen2.5-coder:7b (swap from 8B) Ok. 11.5 s
deepseek-coder:6.7b (after 7B coder) Ok 3.8 s
swe-agent-lm:7b (after DeepSeek) OK. 19.6 s
starcoder2:7b timeout 180 s
codellama:7b-python timeout 180 s
opencoder:8b timeout 180 s

The 8B, the 7B coder, DeepSeek, and SWE-agent-LM answer. StarCoder2, CodeLlama, and OpenCoder do not finish even this prompt inside the cap that daily run uses. DeepSeek and SWE-agent-LM still timed out on daily clamp — a one-word reply is not a score.

Bare clamp prompt, same night. The daily task text only. No harness system prompt. Cap 180s. No file write.

Model Tokens Wall
llama3.1:8b 384 22.3 s
qwen2.5-coder:7b 565 35.5 s
deepseek-coder:6.7b 281 24.0 s
swe-agent-lm:7b 450 39.6 s

Those four finish a short clamp ask. Daily clamp still timed out.

First helper chat, same night. The real first daily clamp request: system 897 characters, user 5,935 characters, about 1,700 tokens (num_ctx 8,192). Same builder as eval_daily.py. Cap 180s.

Model Prompt tokens Reply tokens Wall
llama3.1:8b 1,698 40 16.6 s
qwen2.5-coder:7b 1,706 53 12.4 s
deepseek-coder:6.7b 2,059 334 39.1 s
swe-agent-lm:7b 1,706 66 33.4 s

The 8B and the 7B coder opened with Action: patch. DeepSeek opened with Action: skill then a patch. SWE-agent-LM opened with prose. The helper first turn is not too big to finish once a generate has already succeeded on the machine.

Cold first helper chat, same night. Unload (keep_alive 0), then the same first clamp chat. Cap 180s.

Model Load Wall
llama3.1:8b 8.0 s (7B still listed) 17.1 s
deepseek-coder:6.7b 25.2 s (empty) 53.6 s
swe-agent-lm:7b 28.2 s (empty) 37.9 s

A clean cold first turn is well under 180s. Daily clamp still timed out when a load returned no bytes.

Clean daily remasure, same night. keep_alive 0 did not evict qwen2.5-coder:7b. Then eval_daily.py --model deepseek-coder:6.7b: write-tests 3 / 3 (compiler), first clamp generate 180s timeout. After the timeout, qwen2.5-coder:7b was still the loaded tag. DeepSeek never sat in memory. That swap is the 180s miss — the same first chat is 54 s from empty VRAM.

Empty VRAM daily, same night. ollama stop qwen2.5-coder:7b left {"models":[]}. Then the same script:

Job Result
Write tests 3 / 3 (compiler)
Add clamp 1 pass (steps after 12), then 180s timeout on the second
Logic bug not run

The same empty-VRAM start on swe-agent-lm:7b (ollama stop DeepSeek first): write-tests 3 / 3 (compiler), first clamp generate 180s timeout. After the timeout the tag was loaded. An isolated first chat on that tag was 38 s; the daily first generate did not return in 180s. A follow-up ok generate while /api/ps still listed the tag hit 60s. Listed is not the same as answering. After /api/ps was empty again, the same ok prompt finished in 6.8 s (load 6.5 s). The wedge ended when the listed load expired.

After expiry, same night. Empty VRAM. The real first helper clamp chat finished in 14.5 s (load 4.9 s, 1,708 prompt tokens, prose). Then eval_daily.py on that same loaded tag: write-tests 3 / 3 (compiler), first clamp generate 180s timeout. The isolated chat answers; the daily first generate did not return in 180s.

Agent body, same night. OllamaGenerate.body() sends model, stream, messages, and options.num_ctx. It does not send keep_alive. Two identical first-turn POSTs of that body (about 6,800 characters, 8,192 context) both hit 180s. The same body with keep_alive 30 minutes added finished in 115.7 s (load 105.6 s, 1,709 prompt tokens, prose). After that reply, /api/ps listed llama3.1:8b, not SWE. Still not a nine-cell table.

Two keep_alive chats, same night. Empty VRAM. The same Agent body with keep_alive 30 minutes: first POST 180s timeout (/api/ps still empty). Immediate second POST 44.8 s (load 18.0 s, 1,708 prompt tokens, prose). After that reply /api/ps was empty again; a later check listed llama3.1:8b. One 115.7 s finish does not make keep_alive a first-turn fix. Still not a nine-cell table.

Concurrent 8B bench, same night. lsof on port 11434 showed scripts/measure/bench.py --tier 3 --model llama3.1:8b --repeat 10 holding /api/chat. That is why /api/ps listed the 8B after SWE chats. The same-night SWE first-turn 180s walls were taken while that generate was in flight. Do not treat keep_alive as the cause. Do not remasure SWE until that bench is idle.

Same rerun, next tag. The 8B --repeat 10 ended. The same rerun immediately started bench.py --tier 3 --model qwen2.5-coder:7b --repeat 10. /api/ps listed the 7B coder. The laptop is still not idle. SWE was not remasured.

Idle local remasure, same night. Unloaded the leftover 7B coder. /api/ps empty. No local client on port 11434. Two identical Agent bodies (no keep_alive, about 6,800 characters): first POST 180s timeout, then /api/ps listed swe-agent-lm:7b. Immediate second POST 2.15 s (load 0.01 s, 1,709 prompt tokens, prose). The first-turn 180s is not only the 8B bench. Still not a nine-cell table.

Tier-6 bench, same night. A new local bench.py --tier 6 --model llama3.1:8b --repeat 5 held /api/chat. When that arm ended, --tier 6 --model qwen2.5-coder:7b --repeat 5 was already starting and /api/ps listed the 7B coder. A first SWE chat past the 180s client cap was not run.

Past 180s, same night. After the 7B tier-6 arm, /api/ps was empty and no client was on port 11434. One Agent body (no keep_alive): 300s timeout. /api/ps was still empty. A later check listed llama3.1:8b. Raising the client cap past 180s did not get a first-turn reply. Still not a nine-cell table.

Later, same laptop. A new local bench.py --tier 3 --model llama3.1:8b --repeat 10 was holding /api/chat. /api/ps listed the 8B. A one-word SWE generate from empty VRAM was not run.

Idle one-word, later. That 8B bench ended. Unloaded the 8B. /api/ps empty. No client on port 11434. Prompt: Reply with the single word ok. Cap 180s. 7.42 s (load 5.09 s, 36 prompt tokens, prose). After that /api/ps listed swe-agent-lm:7b. The tag loads. The 300s miss is the helper-sized first chat, not a dead weight file. Still not a nine-cell table.

Helper chat off the 8B, later. The one-word SWE load had expired. /api/ps listed llama3.1:8b. One Agent body (no keep_alive, about 6,800 characters): 36.69 s (load 5.73 s, 1,709 prompt tokens, prose). After that /api/ps listed SWE. The helper-sized first chat can finish when swapping off the 8B. The 300s miss was empty VRAM, not prompt size alone.

Already listed, later. /api/ps still listed SWE. No client on port 11434. The same Agent body: 74.83 s (load 0.02 s, 1,710 prompt tokens, 1,700 eval tokens, prose). The wall is the long reply, not the load. Still not a nine-cell table.

Later. A local Python client was holding /api/chat. /api/ps listed llama3.1:8b. An empty-VRAM helper remasure was not run.

Still later. A second local Python client was holding /api/chat. /api/ps still listed the 8B. Empty helper still not remasured.

Empty helper remasure, later. That client ended. Unloaded the 8B. /api/ps empty. No client on port 11434. One Agent body (no keep_alive): 43.45 s (load 5.62 s, 1,709 prompt tokens, 797 eval tokens, prose). After that /api/ps listed SWE. The earlier 300s empty miss is not stable.

One daily clamp, later. Same eval_daily.py add-function job, twelve steps, swe-agent-lm:7b. /api/ps listed the 8B at start. Stopped steps in 464 s. No def clamp. Suite on the untouched fixture was green. That is a fail, not a score. Still not a nine-cell table.

Empty daily clamp, later. Another local Python client then held /api/chat. After two idle checks, unloaded the 8B. /api/ps empty. No client on port 11434. Same add-function job, twelve steps. First generate timed out at 180 s. After that /api/ps listed SWE. No def clamp. The 464 s clamp started with the 8B listed and finished twelve steps; empty VRAM stopped on the first generate. Still not a nine-cell table.

Listed daily clamp, later. Another local Python client then held /api/chat. After two idle checks, unloaded the 8B. /api/ps empty. One-word SWE generate: 4.93 s (load 4.61 s, 36 prompt tokens, OK.). /api/ps listed SWE. No client on port 11434. Same add-function job, twelve steps. First generate timed out at 180 s. After that /api/ps was empty. No def clamp. Listing the tag is not enough for the helper-sized first turn. Still not a nine-cell table.

8192-listed daily clamp, later. A local bench.py --tier 4 --tier 5 --tier 6 on the 8B then held /api/chat. After two idle checks, unloaded the 8B. /api/ps empty. One Agent body (OllamaGenerate.body(), 6,832 characters, num_ctx 8192, no keep_alive): 29.76 s (load 5.35 s, 1,713 prompt tokens, 388 eval tokens, prose). /api/ps listed SWE. No client on port 11434. Same add-function job, twelve steps. First generate timed out at 180 s. After that /api/ps still listed SWE. No def clamp. A successful 8192 helper load is not enough for the daily first turn. Stop SWE. Still not a nine-cell table. Do not switch. Default stays llama3.1:8b.

Replay one finished table: PYTHONPATH=src python3 scripts/measure/eval_daily.py --model llama3.1:8b.

Write-up: Which model · Hub models.

8B greenfield CLI

Example. 5 September 2026. Ollama llama3.1:8b. Empty folder. Typed: design and develop a small cli app for reviewing github PRs.

Before the app checklist the 8B treated it as a ship job:

Check Result
First Action locate open-pr
Files written none
Suite never ran
Stop ask

After scaffold + checklist (init, urllib and an env token, list, show, mocked tests), three repeats at the default twenty steps. Comment, pagination, and Path.home() config are overflow — a later typed run, not --steps.

Repeat First Action Files Checklist Suite Stop
1 patch weekday test pkg/pr_review.py (list + show via get_prs), tests mocked_tests (wanted list_pulls) never ran steps
2 write-script, then edit pkg/pr_review.py list_pulls + show_pull + mock test list / show ready red (GITHUB_TOKEN, then os) steps
3 edit tests first stub pkg/pr_review.py (37 B) http, list, show, tests missing steps

1 / 3 list/show checklist. 0 / 3 suite green. 0 / 3 done.

Then three repeats at twelve steps, same budget as the daily jobs, after the harness started scaffolding pkg/ and refusing locate / ask:

Repeat list + show + mocks Stopped What it wrote
1 yes steps pkg/pr_review.py, pkg.py, tests
2 yes steps pkg/pr_review.py, tests
3 no (show, mocked_tests) steps pkg/pr_review.py, pkg/pull_viewer.py

2 / 3. The miss spent the budget on a second module.

Later the same day, after #206 (refuse locate until list and show exist), twelve steps again:

Repeat list + show + mocks Stopped What it wrote
1 yes steps pkg/pr_review.py, tests
2 yes steps pkg/pr_review.py, tests
3 yes steps pkg/pr_review.py, tests

3 / 3 on the checklist. 0 / 3 said done. Every run hit the step cap with the files already on disk. Replay: python scripts/measure/eval_cli_app.py (twelve steps; pass --steps 20 for the first cell).

Finish was the gap: the files were on disk and the model kept writing. Once list and show exist, the harness now writes the mocked urlopen test (token via patch.dict) and runs the suite — the same idea as the add-feature cover test. Overflow (comment / pagination / config) is a later typed run, not more --steps.

Same prompt, twelve steps, after that mock-test write (#214). 5 September 2026. Ollama llama3.1:8b.

Repeat Checklist Suite Stopped Wrote
1 no (mocked_tests) red steps pkg/pr_review.py × 4
2 no (show, mocked_tests) red steps pkg/pr_review.py
3 yes green done pkg/pr_review.py, tests

1 / 3 checklist. 1 / 3 suite green. 1 / 3 done. Two of three stayed red after one repair, so I stopped adding product copy.

Same prompt, twelve steps, after the mock test bound the list/GET name the 8B wrote (#220). Same evening. Ollama llama3.1:8b.

Repeat Checklist Suite Stopped Wrote
1 yes green done pkg/pr_review.py × 2, tests
2 yes green done pkg/pr_review.py, tests
3 yes green done pkg/pr_review.py × 4, tests

3 / 3 checklist. 3 / 3 suite green. 3 / 3 done. Replay: PYTHONPATH=src python scripts/measure/eval_cli_app.py.

Later the same day, overflow from a runnable list+show tree. Typed: add the comment subcommand and a mocked test. After #216.

Check Result
First try grep comment (add-feature hint). 20 steps. No comment
Timed cell (12 steps × 3) 0 / 3 closed the comment gap. Every repeat hit the cap
After hint tighten first Action edit; def comment_on on disk; done refused because nothing called it

def comment now counts. Overflow done is allowed once that piece exists — argparse wiring is not demanded by the unused-function guard.

Same prompt, twelve steps, after that unused-function skip (#222). Same evening. Ollama llama3.1:8b. Seeded list+show+mocks tree.

Repeat Comment gap Stopped Wrote
1 closed done pkg/pr_review.py × 2
2 closed done pkg/pr_review.py
3 closed done pkg/pr_review.py

3 / 3 closed comment. Pagination and config stayed leftover — later typed runs, not --steps. Replay: PYTHONPATH=src python scripts/measure/eval_cli_overflow.py.

py-harness run "add the comment subcommand and a mocked test"

Same evening, pagination from a list+show+comment tree. Typed: add pagination to the GitHub PR CLI. After #228 (indented page= counts; a module-level pulls?page= NameErrors on import). Twelve steps × 3. Seeded list+show+comment tree.

Repeat Pagination gap Stopped Wrote
1 open steps none
2 open steps none
3 open steps none

0 / 3 closed pagination. Config stayed leftover. An earlier twenty-step try on a leftover comment tree wrote pkg/pagination.py and a module-level ?page=, then drifted. The timed cell wrote nothing.

Same prompt, after the harness put page= on the list URL (#233). No model. Seeded list+show+comment tree.

Repeat Pagination gap Stopped Wrote
1 closed done pkg/pr_review.py
2 closed done pkg/pr_review.py
3 closed done pkg/pr_review.py

3 / 3 closed pagination. Config stayed leftover — a later typed run, not --steps. Replay: PYTHONPATH=src python scripts/measure/eval_cli_overflow_page.py.

py-harness run "add pagination to the GitHub PR CLI"

Same evening, config from a list+show+comment+page= tree. Typed: add a config file via Path.home. Twelve steps × 3.

Repeat Config gap Stopped Wrote
1 open steps none
2 open steps none
3 open steps none

0 / 3 closed config. The tree already looked finished, so the 8B wrote nothing — the pagination 0/3 shape.

Same prompt, after the harness wrote pkg/config.py with Path.home() (#241). No model. Seeded list+show+comment+page= tree.

Repeat Config gap Stopped Wrote
1 closed done pkg/config.py
2 closed done pkg/config.py
3 closed done pkg/config.py

3 / 3 closed config. Comment, pagination, and config are all later typed runs that the harness can finish without the 8B. Replay: PYTHONPATH=src python scripts/measure/eval_cli_overflow_config.py.

py-harness run "add a config file via Path.home"

Everyday-ready is still the older bar.

Everyday-ready bar

Example. Same evening, 5 September 2026. Ollama llama3.1:8b. Fifteen action_prompts.jsonl rows for first Action. Then fix compute_total in pkg/util_stats.py so it sums the rows on a 2.8 KB file that returns 0.0 — not tota, not subtotl. Three repeats, twelve steps. Clean 8B is the same model with no AGENT_SYSTEM and no agent loop (one-shot draft).

Check Harness 8B Clean 8B
Live parse 11 / 15 0 / 15
≥1 KB logic fix 0 / 3 (steps; two writes were tests only) 3 / 3 (one-shot)

After #229 (refuse rewriting a covering test). Same evening, same script, same twelve steps.

Check Harness 8B Clean 8B
Live parse 10 / 15 0 / 15
≥1 KB logic fix 0 / 3 (steps; writes [] × 3) 3 / 3 (one-shot)

#229 stopped the test rewrite. It did not get a patch on compute_total.

After #238 (refuse explore once the named impl is open). Same evening, same script, same twelve steps.

Check Harness 8B Clean 8B
Live parse 11 / 15 0 / 15
≥1 KB logic fix 0 / 3 (steps × 2, done × 1; writes [] × 3) 3 / 3 (one-shot)

#238 did not get a patch on compute_total. Harness still beats clean on parse. Clean still beats harness on the real fix.

Same evening, after overflow closed (#243). Same script, same twelve steps.

Check Harness 8B Clean 8B
Live parse 12 / 15 0 / 15
≥1 KB logic fix 0 / 3 (steps × 2, done × 1; writes [] × 3) 3 / 3 (one-shot)

Parse moved. The fix did not: still no write to compute_total.

After #246 (bind a zero return to a sum) and #248 (print turns). Same evening, same script, same twelve steps.

Check Harness 8B Clean 8B
Live parse 9 / 15 0 / 15
≥1 KB logic fix 3 / 3 (done; pkg/util_stats.py; turns []) 3 / 3 (one-shot)

The model never ran. The harness wrote the sum and stopped. Parse still beats clean. The fix ties clean, so the script’s harness_fix > clean_fix is false. Not everyday-ready.

That whole-line return 0 / return 0.0 on a named sum is the same class as subtotl and page=: the compiler writes it. The ≥1 KB cell that used that shape is retired as a model job. Do not remasure eval/fixtures/everyday_fix. Replay of the last recorded night: PYTHONPATH=src python scripts/measure/eval_everyday_bar.py.

The live ≥1 KB cell is clip in eval/fixtures/everyday_live: it filters outliers instead of clamping them. The compiler leaves that shape alone. Score it only when #248 turns are non-empty.

After #254 (never-autofix clip cell). Same evening, same script, same twelve steps.

Check Harness 8B Clean 8B
Live parse 8 / 15 0 / 15
≥1 KB logic fix 0 / 3 (steps × 2, done × 1; writes [] × 3; turns non-empty) 3 / 3 (one-shot)

The model ran. It did not write clip. Parse still beats clean. Clean still one-shots the file. Not everyday-ready. Replay: PYTHONPATH=src python scripts/measure/eval_everyday_bar.py.

Same evening, same script, qwen2.5-coder:7b: harness parse 10 / 15 vs clean 1 / 15; harness fix 0 / 3 vs clean 3 / 3. Not everyday-ready. Detail under same-night daily jobs.

Four jobs, as typed

Example. Sample tree demo/orders. Two NameErrors sit in the code:

# src/orders.py
subtotal = compute_total(prices)
return subtotl + (subtotl * TAX_RATE)

# src/orders_controller.py
class OrdersController:
    def status(self) -> str:
        return stauts

Result, first typing (evening)

I typed What I got
ask "what does compute_total return?" "int"
run "write tests for apply_discount" A second test below if __name__. It never ran
run "find the NameError and fix it" Three files edited
run "add a function total_lines and a test" Opened a file. Suite red. Then it asked

0 / 4 I would ship without reading the diff.

Result, after the harness did the compiler jobs first

I typed What I got Check
same ask A sentence that quotes int and says it sums line prices nothing written
same write-tests already has a test. No model suite green
same NameError subtotlsubtotal in orders.py. No model total_with_tax([10]) is 12.0
same add def total_lines(prices) and an AAA test. No model total_lines([10, 20]) == 2
run "find the NameError in src/orders_controller.py" Asks. Does not write return status. Answer okreturn "ok". No model status as an answer is refused

Four of those five finish with no model. That is why they are the same every time. Live first-Action parse the same night (eval_everyday.py --live, llama3.1:8b): 8 / 15. Offline fixtures were clean. Those fifteen cases changed verdict on ten of them across three unchanged reruns. A single parse pass is not a score.

Write-up: First-run four · Scenarios.

Which small open model

Example. scripts/measure/bench.py. A case counts only if the function runs and does the job — not if a file appeared.

Result

Model Write a test / add / fix Platform paths
llama3.1:8b 6–9 / 9 (six runs) 1 / 4
qwen2.5-coder:7b 7 / 9 (one run) 2 / 4
30B-class coder timed out 0 / 4
1B and 1.5B on disk no Action: (prose or # patch)

I did not switch the default. The 7B coder trades two of the daily jobs for one extra platform task.

One run is not a score

The nine cases above were run six times against unchanged code:

9/9   6/9   8/9   7/9   8/9   7/9

Five of the nine pass every time — clamp, double, cover-discount, cover-shout, fix-nameerror — and three of those five finish without calling the model at all. The other four come and go. Over the whole fifteen-case bench, ten of fifteen changed verdict between identical runs, and the totals ranged from 7 to 12.

So the single figures on this page are worth reading as a rough size, not a rank. The comparison between models rests on one run each, which is enough to see that the 30B never finished and not enough to separate 8b from coder:7b. Anything smaller than about a four-case gap is inside the noise.

Same eleven demo tasks against a hosted IDE agent, same wording: the laptop column does not match. No browser, no free shell, no any-language tree.

Write-up: Which model · Same jobs.

Hub GGUFs that Ollama does not ship

Example. 5 September 2026. Two small code models on Hugging Face that this laptop can hold and that ollama pull cannot see. scripts/weights/import_hf_ollama.py downloads the Q4_K_M GGUF (~4.7 GB) and runs ollama create.

Local tag Source What it is
opencoder:8b infly/OpenCoder-8B-Instruct Code-instruct 8B
swe-agent-lm:7b SWE-bench/SWE-agent-LM-7B Qwen2.5-Coder-7B plus 5k traces from their agent
python3 scripts/weights/import_hf_ollama.py --name opencoder
python3 scripts/weights/import_hf_ollama.py --name swe-agent-lm
py-harness --model opencoder:8b run "add a function clamp and a unit test"

Result. Both tags are on disk. A one-word generate hit 180s on OpenCoder and finished in 19.6 s on SWE-agent-LM. The first helper clamp chat (~1,700 tokens) finished in 38 s on a clean cold load of SWE-agent-LM. Daily clamp timed out when a load returned no bytes (write-tests 3 / 3 is the compiler bind, no model). That is not a score. Default stays llama3.1:8b. Other 7B–8B weights that fit this laptop, and the ones that do not, are listed on Hub models.

Write-up: Hub models.

Train more, or not

Example. 35 short train pairs. 30 handwritten Action traces. train.py --everyday is a 7B-class LoRA config. It has not been run.

Result

Idea What it would teach Do it?
More 0.5B steps Tone. Already overfit No
8B LoRA on 30 traces The first Action: line No
7B LoRA after ~2k oracle-clean --record turns The protocol and a finish, if it beats the 8B Later

“Patch the leftover name, write a test that calls it, refuse done” is a harness job. That is what moved the four Start commands from 0 / 4 to 4 / 4.

Write-up: Fine-tune or harness.

On a real repository

Example. Everything above uses demo/orders, a fixture with two built-in bugs. This is the same tool pointed at a working repository of 4,580 first-party files that nobody wrote for this benchmark. Nothing was written inside it: reads ran against it directly, writes against a fresh copy of one module.

Result

Job Score
brief, layout, ask --scope Correct. 6–7 s each
Import cycles reported by layout 4 reported, 0 real — then 4 reported, 4 real after the fix
Write a test, add a function 1 / 12 verified, four tasks, three runs each
Undefined-name guard across 3,658 files 3% flagged, every one correct code — now 0

Reading a real repository works. Writing to one does not, and the same tasks pass on the fixture, which is worth knowing about the fixture.

Detail: Bench record.

When a run says done and means nothing

The worst outcome is not a failure. It is a run that finishes, reports success, and leaves the file exactly as it was — because the only way to find that out is to go and look.

Counting why each run stopped, across 45 benchmark runs, put a number on it: two of the nine failures reported done.

Result

One task, ten runs each side Reported success having changed nothing
Before 5 of 10
After two fixes 0 of 10

Neither fix was a missing guard. One guard existed and its escape hatch was a sentence the refusal itself handed the model, which the model handed back. The other cause was not in the model at all: the harness took a word out of the task, found it as a substring in a test file, and finished. The word is in 17 of this project’s test files and called in 5.

Write-up: When a run says done and means nothing.

Asking a bigger model, rarely

If the harness could put a question to a larger model the user has registered, when should it? The call is easy; knowing when to make it is not.

Result

Why a run stopped, 45 runs Share Was it really stuck?
Asked a question 7% 3 of 3
Ran out of steps 18% 4 of 8
Said done, was wrong 4% no stop reason catches it

A run that stops to ask has earned it: asking is capped at two, and refused outright once files have changed. Running out of steps means much less. Seven of the nine failures were platform and operations work, the tier that moved 37% to 70% on harness fixes alone — gaps in the tool, which sending them away would hide.

Write-up: Asking a bigger model, rarely.

A chain of easy tasks

If the model is not very good, is it better to give it several small instructions than one composite one?

Result

Same work, same fixture, 8 runs each Worked Average
One instruction 5 of 8 20s
Split in two, sent blind 4 of 8 46s
Split, each step checked and retried 4 of 8 42s

Splitting bought nothing and cost twice the clock. A run is already up to twenty turns, each a single action, so splitting from outside adds a second copy of the decomposition rather than more of it — and each run builds its own memory, so every step started from nothing.

Write-up: Small steps, measured.

The instrument was broken

A day spent asking whether a bigger model breaks the wall found two faults in the benchmark instead. Both were invisible while only local models were measured; both would have made a fine-tune evaluation wrong.

Result

Tier 3, ten passes Nobody answering Answered
llama3.1:8b 13 of 20 14 of 20
qwen2.5-coder:7b 7 of 20 13 of 20

A run that stops to ask needs somebody to answer, and nobody was there, so the question ended the run as a failure. qwen2.5-coder:7b asks in eleven runs of twenty where llama3.1:8b asks in one, so the benchmark was measuring willingness to act without asking.

The second fault only appears against a hosted model: it wraps drafts in markdown fences, the fence reaches the file unchanged, and the result is a SyntaxError. Nine of ten runs then produced nothing that would load. A 14B, meanwhile, times out on this machine three times out of three.

Every model number published before this is unsafe.

Write-up: The instrument was broken.

The fence was the whole story

Example. The same hosted 32B, the same two tier-3 cases, ten runs each side. The only difference is whether the harness takes the markdown fence off a draft before writing it to a file.

Result

Qwen2.5-Coder-32B-Instruct, tier 3 worked
before the fence was stripped 1 of 10
after 9 of 10

Per case after the fix: slugify 5 of 5, wordcount 4 of 5. The one failure is an ordinary word_count not found in any module, not a file the model had broken. Median run 15.3 s. Both columns were measured after the answerer was fixed, so the jump belongs to the fence alone.

The model was never the problem. Four backticks were.

The control says the same thing from the other side. Local weights do not fence their code — zero of twenty recorded turns contain one — so the fix cannot move them, and it does not:

llama3.1:8b, tier 3, twenty runs worked
without the fence fix 10 of 20
with it 10 of 20

Write-up: The fence was the whole story.

A day of repairs

Example. Six September was spent on the tool rather than the model: the benchmark’s own faults, the parser, and the paths the harness takes without asking a model at all. Same fifteen cases, five passes, llama3.1:8b, before and after.

Result

  worked
5 September 51 of 75
6 September 64 of 75
tier    
1, 2, 4 15/15, 10/10, 10/10 clean
3 8/10 wordcount 3/5
5 7/10 fix-offbyone 2/5
6 14/20 env-flag 2/5

Almost none of it was the model. A benchmark that scored a question as a failure; a parser that fed markdown to Python; a mechanical path that wrote def word_count(prices): return len(prices) and its own passing test; a suite that ran nothing and counted as passing; a refusal whose advice named a field the draft did not have. On the reported “write tests” reproduction, suites that actually ran a test went from 2 of 8 to 8 of 8.

Two of the nine fixes moved no score and are kept anyway. Guarding against a draft that deletes the function it was sent to fix took that outcome from 4 of 12 runs to 0 of 12, while the pass rate stayed inside the noise — it converts deleted the function into did not fix the bug.

Write-up: A day of repairs.

What the totals were hiding

Example. Tier 3, twenty runs an arm, llama3.1:8b. A change to how the harness reads the subject of a task — preferring a name written with brackets, word_count(text) — measured before and after.

Result

tier 3, twenty runs worked slugify wordcount zero-step runs
before 9 of 20 6/10 3/10 0
subject fixed 10 of 20 10/10 0/10 10
subject fixed and guarded 19 of 20 0

One case looked like noise. Underneath were two large effects cancelling out, and the second arm scored exactly 1 on all ten passes — a flat line on a benchmark that changes verdict two thirds of the time. All ten wordcount runs used no model steps: a mechanical path was writing def word_count(prices) -> int: return len(prices), writing its test to match, and reporting “Tests passed”.

The bug was already in the harness. The change only removed the accident hiding it, because the subject of that task used to read as module. Guarding it — a task that spells its own signature has answered the question the guess was for — took the tier to 19 of 20, ten cases against a floor of two.

Write-up: What the totals were hiding.

The wall two local models share

Example. Tier six is platform and operations work: environment flags, virtualenv paths, KEY=VALUE files, retries. It is the tier two local models were said to stop at together. Same harness, same four cases, five passes each, measured after the benchmark was repaired.

Result

tier six, twenty runs each worked
Qwen2.5-Coder-32B-Instruct (hosted) 18 of 20
qwen2.5-coder:7b (local) 11 of 20
llama3.1:8b (local, the default) 8 of 20

Applying this project’s own rule — under about four cases is noise — the two local models are not distinguishable from each other, and the larger model’s lead over both is. So the wall is real, shared, and the part of it that harness work did not reach is not reachable by another rule: the tool holding all three models is the same tool.

That inverts which stuck moments are worth a remote call. Tier six was the tier to keep local, because sending it away would hide tool gaps. It is the tier where the local ceiling is lowest.

Write-up: The wall two local models share.

Two models, one wall

Before training anything, the cheap question: is the base model the constraint? The benchmark takes a model name, so it costs one command.

Result

Seventy-five runs each Worked Wrote nothing Wrote the wrong thing
llama3.1:8b 51 of 75 8 16
qwen2.5-coder:7b 50 of 75 18 7

One case apart on the score, and almost opposite failures. Two models of different lineage meeting the same wall says something about the size rather than about either model.

It also moves the bar for a fine-tune. Wrong-code failures can be more than halved without a single extra run working — they just become refusals to act. Raising the count that works is the target; improving the manner of failing is not.

The per-tier splits in that run suggested sending some task types to one model and some to the other. Checked at ten passes, the bugfix tier came out level at 18 of 20 each — the apparent gap was one case in a five-run sample — while tier 3 widened to 13 against 7. So there is no task type worth routing to qwen2.5-coder, and five passes turns out to be too few to compare two models per tier at all.

Write-up: Two models, one wall.

Where the failures are

Seven harness changes measured, six moved nothing. So rather than measure an eighth, seventy-five runs were classified by what they left behind, and eight hundred and thirty-six model turns by what the model was sent.

Result

Of the 24 failures in 75 runs Share
wrote something, but not the thing asked for 42%
wrote nothing at all 33%
wrote something, it did not do the job 25%
claimed success having written nothing 0%

Two thirds of what fails is plausible, wrong code, and nothing deterministic separates that from plausible, right code — only running it does, and the suite already runs. The harness has taken the failures it can take.

The last row is the week’s one measured gain: that shape was two of nine failures a week ago and is nought of twenty-four now. Not a higher pass rate — no lies about it.

A quarter of every run is the harness saying no: 23% of turns are a refusal or a nudge, most often “run the tests before finishing” (58), “read the file before patching it” (36) and “that is the wrong file” (32).

Write-up: Where the failures are.

What the harness cannot fix

Most gaps here close when the harness stops guessing and starts checking. Four did not, and they are more informative than the ones that did.

Result

Measurement Outcome
Refusing a bot’s major version bump 0 of 5 — five merged safely, nothing caught. Since fixed: 2 of 6 now allowed, the rest name the workflow nobody ran
Telling the model what the project already has Pointer correct, ignored 3 of 3
Platform work on stock llama3.1:8b 6 of 8 over two passes, no new weights
This project’s own fine-tune 0 of 4 held-out, worse than its base model
Training data collected in a week of real work 0 rows — recording was behind a flag
Centring a long file’s excerpt on the task’s subject Defect real and fixed; 0 of 5 either side
Showing the model how long its functions are Rule existed as a merge gate only; 21 of 30 either side

Five of the six are cases where the harness knew something and it made no difference. The one that worked, worked by running something: a dependency’s major bump was cleared by installing the version and calling every function the project uses against it.

What closes a gap is an oracle. What does not is telling the model more.

Write-up: What the harness cannot fix.

A larger open model

Example. The 30B already timed out on this laptop. --engine openai sends only the generate call to a GPU. The write limit stays here.

Result

Run Score
30B on this laptop Timeout. 0 / 4 platform cases
14B on this laptop Could not be measured. 9 GB of weights on 18 GB put the machine into 12–13 GB of swap; no run finished
32B on a GPU, tier 3 9 / 10 after the fence was stripped; 1 / 10 before. Not the four daily jobs

The 14B result is about the machine, not the model. Weights are only part of the budget: the key-value cache grows with context and the operating system wants its share, so the practical ceiling here is about 11–12 GB, not 18. If you are choosing hardware, reckon on roughly twice the size of the model you mean to run.

Write-up: Cloud weights · Bench record.