Architecture · 2 of 5
The training program: from corpora to a 4.57 MB bundle
Every model in the lineage is produced by the same pipeline: pretrain an encoder on unlabelled ISL landmarks, fine-tune it on INCLUDE with the extra supervision we can manufacture, ensemble four seeds, distil them into one student, quantise, and ship it as a Lab bundle. This page walks the pipeline top to bottom; the numbers each stage produced live on the results page.
How to read this. Top to bottom is the order things run. Grey boxes are data; green boxes are models or training stages; the outlined boxes at the bottom are artifacts that ship. The two corpora on the right are never trained on: every fine-tuned seed and the distilled student are scored on them through the accuracy test, drawn dotted because scoring changes nothing upstream.
- Corpora & synthetic data
- Pretraining, fine-tuning, distillation
- Held-out evaluation
- Shipped artifacts
1 · Data
| Corpus | Role | Size |
|---|---|---|
| INCLUDE-260 | the labelled core; classes of the 260-sign set | 4,284 clips (4,276 after extraction) · 260 classes |
| CISLR | dictionary corpus; gloss-matched infusion into fine-tuning | 4,616 units in the pool · 379 test-role clips never trained on |
| ISLRTC dictionary | dictionary corpus; infusion | 4,928 units |
| RKMVU-220 | the cross-signer test: one signer the model has never seen | 220 clips (104 odd-class clips form the strict held-out split); RKMVU pool 1,898 clips / 2,002 dictionary queries |
| iSign | unlabelled pretraining | 99,457 units |
| YouTube-SL-25 (ISL subset) | unlabelled pretraining | 8,165 units |
Unlabelled pool (pool_v2) | JEPA pretraining | 119,064 units → 123,340 with INCLUDE |
| Synthetic signers | fine-tuning augmentation | 2 "full" variants per INCLUDE take → 8,568 clips, plus view/shape variants |
Gloss-matched infusion adds roughly 2,000 dictionary clips (CISLR + ISLRTC) to the labelled fine-tuning set: a dictionary clip whose gloss is one of the 260 INCLUDE classes becomes a second, differently-performed example of that class.
Synthetic signers
synth_signers_geo.py re-poses INCLUDE's own MediaPipe landmarks in 3D: a yaw/pitch rotation
using MediaPipe's z-coordinate, and a bone-length retarget to a different body. Because it is a
landmark-space operation it costs nothing to render, and it was validated for parity 0.0005
against the real extraction pipeline — a synthetic clip is numerically indistinguishable from
what the browser would have produced for that pose.
2 · Self-supervised pretraining (ISL-JEPA)
The encoder is pretrained on the full pool with a latent-prediction objective — no labels, no coordinate reconstruction. The objective, the collapse it suffered in Phase B, and the all-layer teacher that fixes it in Phase C have their own page. What the training program needs from it is a transferable encoder; the finding that only layers ≤ 4 of the Phase B encoder transfer is why every fine-tune below is Cut@4.
3 · The fine-tune recipe
The encoder is cut at layer 4 and an attentive head is trained on top for the 260 classes.
The recipe that produced the current best is written all+synth:full@2 + rr:
| Ingredient | What it is | Measured effect |
|---|---|---|
all — gloss infusion | the ~2k gloss-matched dictionary clips join INCLUDE as labelled examples | with synthetic signers: 5-seed mean cross-val 0.3018 → 0.3927 |
synth:full@2 | both "full" synthetic variants per INCLUDE take | (measured jointly with infusion) |
rr — recover-and-resample | drop up to 25% off either end of a clip and stretch it back to 64 frames, with p = 0.5 | +1.7 pp gate, +2.5 pp RKMVU over the same recipe without it; 3 of 4 seeds up; half the seed spread |
4 · Four seeds → one teacher
The recipe is trained four times with different seeds. Their softmaxes are averaged online to
form an ensemble teacher. With rr, the four seeds gate at .4057 / .4174 / .4124 / .4107; the
ensemble gates at 0.4574, and 0.4641 with mirror test-time augmentation.
5 · Distillation
A single student of the same shape is trained on every INCLUDE clip against the ensemble
teacher — cross-entropy on the labels plus a temperature-scaled KL term to the teacher's
softmax. The point of distillation is deployment: the student is one model, not four, at
ensemble-level accuracy. The rr student (jepa-rr4) gates at 0.4591 — slightly above
the ensemble it was distilled from.
6 · Export and the Lab bundle
The student is exported to ONNX (opset 17) with velocity, bone, normalisation and softmax inside the graph, then dynamically quantised to int8 and parity-checked against the fp32 export. The int8 model is 4.57 MB and, on the cross-signer set, scores 0.4864 to fp32's 0.4818 — quantisation costs nothing.
The bundle (include-classifier + model card) drops into the studio's Lab as a selectable
model. It is promoted to the default only after the gate
says it is better, not merely equal, to what is deployed.