Skip to content

Architecture · 2 of 5

The training program: from corpora to a 4.57 MB bundle

Every model in the lineage is produced by the same pipeline: pretrain an encoder on unlabelled ISL landmarks, fine-tune it on INCLUDE with the extra supervision we can manufacture, ensemble four seeds, distil them into one student, quantise, and ship it as a Lab bundle. This page walks the pipeline top to bottom; the numbers each stage produced live on the results page.

Figure 2The training program — data, self-supervision, the fine-tune recipe, ensemble, distillation, export

How to read this. Top to bottom is the order things run. Grey boxes are data; green boxes are models or training stages; the outlined boxes at the bottom are artifacts that ship. The two corpora on the right are never trained on: every fine-tuned seed and the distilled student are scored on them through the accuracy test, drawn dotted because scoring changes nothing upstream.

  • Corpora & synthetic data
  • Pretraining, fine-tuning, distillation
  • Held-out evaluation
  • Shipped artifacts
Figure 2The training program — data, self-supervision, the fine-tune recipe, ensemble, distillation, export

1 · Data

CorpusRoleSize
INCLUDE-260the labelled core; classes of the 260-sign set4,284 clips (4,276 after extraction) · 260 classes
CISLRdictionary corpus; gloss-matched infusion into fine-tuning4,616 units in the pool · 379 test-role clips never trained on
ISLRTC dictionarydictionary corpus; infusion4,928 units
RKMVU-220the cross-signer test: one signer the model has never seen220 clips (104 odd-class clips form the strict held-out split); RKMVU pool 1,898 clips / 2,002 dictionary queries
iSignunlabelled pretraining99,457 units
YouTube-SL-25 (ISL subset)unlabelled pretraining8,165 units
Unlabelled pool (pool_v2)JEPA pretraining119,064 units → 123,340 with INCLUDE
Synthetic signersfine-tuning augmentation2 "full" variants per INCLUDE take → 8,568 clips, plus view/shape variants

Gloss-matched infusion adds roughly 2,000 dictionary clips (CISLR + ISLRTC) to the labelled fine-tuning set: a dictionary clip whose gloss is one of the 260 INCLUDE classes becomes a second, differently-performed example of that class.

Synthetic signers

synth_signers_geo.py re-poses INCLUDE's own MediaPipe landmarks in 3D: a yaw/pitch rotation using MediaPipe's z-coordinate, and a bone-length retarget to a different body. Because it is a landmark-space operation it costs nothing to render, and it was validated for parity 0.0005 against the real extraction pipeline — a synthetic clip is numerically indistinguishable from what the browser would have produced for that pose.

2 · Self-supervised pretraining (ISL-JEPA)

The encoder is pretrained on the full pool with a latent-prediction objective — no labels, no coordinate reconstruction. The objective, the collapse it suffered in Phase B, and the all-layer teacher that fixes it in Phase C have their own page. What the training program needs from it is a transferable encoder; the finding that only layers ≤ 4 of the Phase B encoder transfer is why every fine-tune below is Cut@4.

3 · The fine-tune recipe

The encoder is cut at layer 4 and an attentive head is trained on top for the 260 classes. The recipe that produced the current best is written all+synth:full@2 + rr:

IngredientWhat it isMeasured effect
all — gloss infusionthe ~2k gloss-matched dictionary clips join INCLUDE as labelled exampleswith synthetic signers: 5-seed mean cross-val 0.3018 → 0.3927
synth:full@2both "full" synthetic variants per INCLUDE take(measured jointly with infusion)
rr — recover-and-resampledrop up to 25% off either end of a clip and stretch it back to 64 frames, with p = 0.5+1.7 pp gate, +2.5 pp RKMVU over the same recipe without it; 3 of 4 seeds up; half the seed spread

4 · Four seeds → one teacher

The recipe is trained four times with different seeds. Their softmaxes are averaged online to form an ensemble teacher. With rr, the four seeds gate at .4057 / .4174 / .4124 / .4107; the ensemble gates at 0.4574, and 0.4641 with mirror test-time augmentation.

5 · Distillation

A single student of the same shape is trained on every INCLUDE clip against the ensemble teacher — cross-entropy on the labels plus a temperature-scaled KL term to the teacher's softmax. The point of distillation is deployment: the student is one model, not four, at ensemble-level accuracy. The rr student (jepa-rr4) gates at 0.4591 — slightly above the ensemble it was distilled from.

6 · Export and the Lab bundle

The student is exported to ONNX (opset 17) with velocity, bone, normalisation and softmax inside the graph, then dynamically quantised to int8 and parity-checked against the fp32 export. The int8 model is 4.57 MB and, on the cross-signer set, scores 0.4864 to fp32's 0.4818 — quantisation costs nothing.

The bundle (include-classifier + model card) drops into the studio's Lab as a selectable model. It is promoted to the default only after the gate says it is better, not merely equal, to what is deployed.