Cloud weights
Question. The laptop 8B is the everyday brain. A 30B timed out here. How do we experiment with a stronger model, or a later fine-tune, without giving up the write limit or training the 0.5B again?
Answer. Keep the harness on the laptop. Move only the generate
call to a GPU you rent. Same Action: schema, same oracles, same
demo/orders checks. Do not train on the 30 seed traces, locally or in
the cloud. Publish new adapters only on YauhenBichel.
Related: Experiments · fine-tune or harness · which model · hub models · local loop vs hosted agents.
What stays local
python-vibe still reads and writes only inside --project. serve.py
still binds 127.0.0.1. Every file the run writes is written here, by
this machine.
What does travel is the prompt, and the prompt contains code. The harness
opens the file the task names and puts it in the first turn, and a later
read or grep sends what it found. Asked to fix
src/billing.py, the first request carried the whole of that file —
constants, comments and all. The cloud box is not given the working tree,
but it is given whatever the run reads out of it. Point it at a host you
would show that code to.
A bigger remote model does not become a hosted IDE agent. There is still no browser Action and no free shell. See local loop vs hosted agents.
Two ways to reach a GPU
Ollama on a box you rent. The client is already written.
OLLAMA_HOST points at that box. --engine stays ollama.
export OLLAMA_HOST=https://gpu.example:11434
python-vibe --model qwen2.5-coder:32b run "find the NameError and fix it"
OpenAI-compatible HTTP. Hugging Face Inference, vLLM, or any other
/v1/chat/completions host. --engine openai. The token is an
environment variable. It is never written into a trace.
export HF_TOKEN=… # or PYTHON_VIBE_API_KEY
export PYTHON_VIBE_BASE_URL=https://router.huggingface.co/v1
python-vibe --engine openai \
--model Qwen/Qwen2.5-Coder-32B-Instruct \
run "find the NameError and fix it"
If HF_TOKEN is set and PYTHON_VIBE_BASE_URL is not, the harness uses
the Hugging Face router. PYTHON_VIBE_TIMEOUT defaults to 180 seconds.
The 0.5B sidecar and the laptop 8B stay the defaults. A remote model is
an experiment flag, not a new everyday brain, until it beats untuned
llama3.1:8b on the checks below.
The experiment
Same wording, same fresh demo/orders copy, same independent file
checks as live scenarios. The only
changed cell is the generate backend.
| Cell | Engine | Weight | Where |
|---|---|---|---|
| A | ollama |
llama3.1:8b |
this laptop (already measured) |
| B | openai |
Qwen/Qwen2.5-Coder-14B-Instruct or 32B |
Hugging Face Inference or a rented GPU |
| C | openai |
a 70B-class instruct model | rented GPU only |
Score is the same as the laptop tables: would a daily user ship the
diff, and does the planted NameError stay gone. First-Action parse
from scripts/measure/eval_everyday.py --live is a second number. It flips
between runs; do not treat one pass as a win.
Everyday-ready still means: beat cell A on parse and on a real
≥1 KB fix. A faster 32B that still says done with subtotl in the
file is not a win. The oracles stay on.
This page has no live B/C numbers yet. Fill them in after a run. Until then the laptop 8B column on which model is the only score.
Fine-tune later, on that same GPU
The in-tree later LoRA is still
mlx-community/Qwen2.5-Coder-7B-Instruct-4bit via
configs/python-vibe-8b.yaml. On a CUDA box the same recipe is a
Hugging Face Qwen/Qwen2.5-Coder-7B-Instruct LoRA. The data rule does
not change with the cloud bill:
- Record only turns the oracles already accept
(
--record, gitignoreddata/agent-loop/extra.jsonl). - Wait until that file is on the order of ~2k clean turns, not 30.
- Train on the rented GPU. Publish the adapter as a YauhenBichel repo. Official weights stay on that account.
- Serve the adapter with vLLM or Ollama on the same box. The laptop
uses
--engine openai(orOLLAMA_HOST) against that URL. - Evaluate with
scripts/measure/eval_everyday.py --liveandscripts/run/demo.py. If the LoRA loses to cell A, delete the adapter.
train.py --everyday on the thirty seed rows is still a no, including
when someone else is paying for the GPU. That is the C4 cell in
fine-tune or harness.
Do not
- Train more 0.5B steps because a GPU is cheap this week.
- Change
serve.pyto bind0.0.0.0. The sidecar stays loopback. - Put tokens, hostnames, or adapter folders in git.
- Call a remote 32B everyday-ready from one parse pass.
- Drop
PythonVibeGuardbecause the model is larger. The write limit is the product.
Order of spend: measure cell B on the four Start commands and the controller leftover bind. Only then rent a bigger box. Only then train.