Architecture · 5 of 5
Experiment infrastructure: a GPU you can only rent by the notebook cell
Training runs on a Colab A100, which is cheap, fast, and hostile to long unattended jobs: it has no shell you can keep, it kills idle VMs, and a cell interrupt takes every child process with it. The infrastructure below turns that into something a laptop can drive for a day — and this page records the three lessons it taught, one of which changed a result.
How to read this. Time runs downward. Each vertical line is a machine; each arrow is a file moving or a command being issued. The two loops are daemons that run for the life of the VM: one polls for commands, one pushes logs and results back. Drive is the only durable store — everything on the VM disappears when the session ends.
- The Mac never runs a model; it serves files and reads results
- Every stage on the VM is a detached process
- Checkpoints are the only thing that survives a kill
The pieces
| Component | Where | What it does |
|---|---|---|
| File bridge | Mac, 127.0.0.1:8766 | GET serve/ hands the VM its payload — packs, scripts, the pool index; PUT in/ receives logs, results.json, encoders and ONNX exports |
| Tunnel | cloudflared | exposes the bridge to the VM without opening a port; the VM only ever sees one URL |
| Bootstrap cell | Colab | fetches the payload, installs dependencies (ONNX is not preinstalled on a fresh VM), starts the daemons — and then keeps executing, because an idle notebook is a dead VM |
| Stage runner | VM, stage_runner.py | runs one stage (pretrain, fine-tune, export, gate) as a detached process; several can share the GPU |
| Command daemon | VM, every 30 s | polls queue.txt on the bridge for run:<stage> or sh:<base64> lines — the Mac drives the VM without touching the notebook |
| Status daemon | VM, every 60 s | pushes <stage>.log, results.json and status.txt back, so the Mac can tail a run it cannot see |
| Drive | Google Drive | checkpoints every 5 epochs; encoders and ship bundles; the only home of the Phase B champion base |
Lessons
1 · The idle timeout, and why the bootstrap cell never returns
Colab reclaims a VM whose notebook is not executing. Daemons alone do not count; a cell has to be running. So the bootstrap cell ends in a loop that stays alive for the session, and every stage checkpoints to Drive every five epochs with auto-resume. This was tested by the platform rather than by us: Phase C proper was killed at epoch 11 and resumed from its checkpoint without intervention.
2 · Loading the pool
The full pool is 123,340 clips. The Phase C pilot pretrained on a random 30k subset; an earlier version sliced the head of the pool instead, which biased the subset toward whichever corpus happened to be first — a bug now fixed. Exact memory figures for a full-pool load on the A100 VM are pending; Phase C proper is running on the full pool.
3 · The name collision
Bugs fixed along the way
| Symptom | Cause | Fix |
|---|---|---|
| LP-FT arm never improved | the encoder's learning rate was stuck at 0 — a cosine scheduler recursion | scheduler rewritten; the arm still screened negative (−5.3 pp), but for a real reason |
| Subset pilots favoured one corpus | the pool subset was a head slice, not a sample | random subset |
| Interrupting a cell killed the training | Colab cell interrupts propagate to child processes | every stage runs detached from the notebook |
| Export failed on a fresh VM | ONNX is not in the base image | installed in bootstrap |
| Gate reported for the wrong run | scorer looked up results by a stale stage name | names unified between runner and scorer |