Skip to content

Architecture · 5 of 5

Experiment infrastructure: a GPU you can only rent by the notebook cell

Training runs on a Colab A100, which is cheap, fast, and hostile to long unattended jobs: it has no shell you can keep, it kills idle VMs, and a cell interrupt takes every child process with it. The infrastructure below turns that into something a laptop can drive for a day — and this page records the three lessons it taught, one of which changed a result.

Figure 5One experiment, end to end — the Mac bridge, the tunnel, the VM daemons, and Drive

How to read this. Time runs downward. Each vertical line is a machine; each arrow is a file moving or a command being issued. The two loops are daemons that run for the life of the VM: one polls for commands, one pushes logs and results back. Drive is the only durable store — everything on the VM disappears when the session ends.

  • The Mac never runs a model; it serves files and reads results
  • Every stage on the VM is a detached process
  • Checkpoints are the only thing that survives a kill
Figure 5One experiment, end to end — the Mac bridge, the tunnel, the VM daemons, and Drive

The pieces

ComponentWhereWhat it does
File bridgeMac, 127.0.0.1:8766GET serve/ hands the VM its payload — packs, scripts, the pool index; PUT in/ receives logs, results.json, encoders and ONNX exports
Tunnelcloudflaredexposes the bridge to the VM without opening a port; the VM only ever sees one URL
Bootstrap cellColabfetches the payload, installs dependencies (ONNX is not preinstalled on a fresh VM), starts the daemons — and then keeps executing, because an idle notebook is a dead VM
Stage runnerVM, stage_runner.pyruns one stage (pretrain, fine-tune, export, gate) as a detached process; several can share the GPU
Command daemonVM, every 30 spolls queue.txt on the bridge for run:<stage> or sh:<base64> lines — the Mac drives the VM without touching the notebook
Status daemonVM, every 60 spushes <stage>.log, results.json and status.txt back, so the Mac can tail a run it cannot see
DriveGoogle Drivecheckpoints every 5 epochs; encoders and ship bundles; the only home of the Phase B champion base

Lessons

1 · The idle timeout, and why the bootstrap cell never returns

Colab reclaims a VM whose notebook is not executing. Daemons alone do not count; a cell has to be running. So the bootstrap cell ends in a loop that stays alive for the session, and every stage checkpoints to Drive every five epochs with auto-resume. This was tested by the platform rather than by us: Phase C proper was killed at epoch 11 and resumed from its checkpoint without intervention.

2 · Loading the pool

The full pool is 123,340 clips. The Phase C pilot pretrained on a random 30k subset; an earlier version sliced the head of the pool instead, which biased the subset toward whichever corpus happened to be first — a bug now fixed. Exact memory figures for a full-pool load on the A100 VM are pending; Phase C proper is running on the full pool.

3 · The name collision

Bugs fixed along the way

SymptomCauseFix
LP-FT arm never improvedthe encoder's learning rate was stuck at 0 — a cosine scheduler recursionscheduler rewritten; the arm still screened negative (−5.3 pp), but for a real reason
Subset pilots favoured one corpusthe pool subset was a head slice, not a samplerandom subset
Interrupting a cell killed the trainingColab cell interrupts propagate to child processesevery stage runs detached from the notebook
Export failed on a fresh VMONNX is not in the base imageinstalled in bootstrap
Gate reported for the wrong runscorer looked up results by a stale stage namenames unified between runner and scorer