Review 2 · Deliverable 1 of 4
Implementation of algorithms and techniques
Eleven techniques, each implemented, in use, and tied to a measurement. They fall into four groups: what runs in the browser, how the model is trained, how it is packed for the browser, and how the data and the experiments are kept honest. The two formulas that decide what a user sees — the distillation loss and the calibration — are written out in full.
Where each technique sits
| # | Technique | Stage | Where it is measured |
|---|---|---|---|
| 1 | Landmark extraction and the canonical window | runtime + training | every number (one contract for both) |
| 2 | Mirror test-time augmentation | runtime | in-app accuracy |
| 3 | Full-vocabulary retrieval | runtime | CISLR open-vocabulary protocol |
| 4 | Shape match (no learned weights) | runtime | shape-match results |
| 5 | ISL-JEPA self-supervised pretraining | training | model lineage |
| 6 | Fine-tune recipe: infusion, synthetic signers, recover-and-resample | training | seed tables |
| 7 | Ensemble → knowledge distillation | training | ensemble vs student |
| 8 | int8 ONNX export | packing | torch ↔ ONNX parity, then the accuracy test |
| 9 | Temperature calibration, cross-checked | packing | calibration cross-fit |
| 10 | Dataset quality system | data | +7.0 pp accuracy test, 4/4 seeds |
| 11 | Compute harness | experiments | ≥ 4 seeds per claim |
How to read this. Everything inside the browser region happens on the user's device. The camera frame is turned into 107 landmark points and dropped; from then on only numbers move. The 64-frame window feeds three answers — the 260-sign set, the full vocabulary and shape match — and the surfaces (Learn, Practice, Dictionary, Studio) read from them. The static host only serves the bundle.
A · What runs in the browser
1 · Landmark extraction and the canonical window
MediaPipe (the Holistic-class landmarkers) finds the body pose (33 points), both hands (21 each) and a subset of the face in every camera frame. The project keeps 107 canonical 2-D points per frame, normalises each frame by the signer's shoulders — so distance from the camera and position in the frame stop mattering — resamples a clip or a live stream to a 64-frame window, and writes zeros for any frame in which a hand was not found, so the model sees "absent" rather than a guess. The same code produces training packs and the live window: a model is never trained on one geometry and run on another.
2 · Mirror test-time augmentation
A signer whose left hand is dominant produces the mirror image of the training data. Rather than learn both, the app runs every window twice — as captured and flipped left-to-right — and keeps the answer of whichever pass is more confident:
pTTA(x) = p(x) if max p(x) ≥ max p(x̃), otherwise p(x̃)
Mirror TTA is part of what the app shows, so the in-app numbers on the results page include it.
3 · Full-vocabulary retrieval
The full vocabulary (8,188 words, of which the bank reaches 8,165) is answered by retrieval, not classification: a supervised encoder trained with an ArcFace margin maps each window to a 256-dimensional embedding, and the answer is the nearest rows of a prototype bank of 8,497 rows by cosine similarity, shown only if it clears a confidence bar. A word signed more than one way (833 words) keeps one row per variant, which rebuilt the bank and added +10.7 pp.
4 · Shape match — no learned weights
Shape match compares the hand and arm movement of a window against stored reference clips with dynamic time warping over hand-crafted kinematic features (the fingerprint). It has no trained parameters, which makes it a control: it shows what the data gives away without any learning. On the accuracy test it went from 7.0% to 16.7%, and to 25.0% on the current all-260 bench. How it behaves as a helper to the AI model is on the results page.
B · How the model is trained
How to read this. Four regions, read top row then bottom row: data, pretraining, fine-tuning, shipping. The corpora are extracted into one canonical pack; ISL-JEPA pretrains an encoder without labels; the encoder is cut at layer 4 and fine-tuned twelve times; the twelve are distilled into one student, exported int8, calibrated, and promoted only through the accuracy test.
5 · ISL-JEPA self-supervised pretraining
The encoder is a transformer over articulator tokens — each token is one body part at one moment — with eight layers. It is pretrained on a 123,340-clip unlabelled pool with a joint-embedding predictive objective: parts of a clip are hidden, and the encoder, through a small predictor, must reproduce the latent representation that a slowly moving copy of itself (the EMA teacher) computes for the hidden parts from the full clip. Because the target is a representation rather than raw coordinates, tracker jitter is not something the model is asked to reproduce.
ℒJEPA = (1 / |M|) · Σi ∈ M ‖ ŝi − sg(s̄i) ‖² · ξ ← τ ξ + (1 − τ) θ
Phase B used the teacher's last layer; the upper layers collapsed, and only layers up to four transfer, so the encoder is cut at layer 4. Phase C's all-layer teacher removed the collapse, but at full budget it was a wash against the Phase B recipe and its students lost on the unseen signer — recorded under negative results. The objective and the collapse are explained on the JEPA page.
6 · The fine-tune recipe all+synth:full@2+rr
The pretrained encoder is fine-tuned on the INCLUDE training split with three additions:
- Gloss infusion. Clips from CISLR and the ISLRTC dictionary whose word matches one of the 260 classes are added for all 260 classes — other people signing the same word, filmed differently.
- Synthetic signers. Each real clip is lifted to 3-D with MediaPipe's depth estimate, rotated (yaw and pitch) and bone-retargeted to other limb proportions, then projected back — two variants per clip.
- Recover-and-resample. With probability 0.5, up to 25% of the clip is trimmed from either end and the remainder stretched back to 64 frames, so the model stops depending on where in the window a sign begins.
How to read this. One real clip's landmarks go in on the left. They are lifted into 3-D, turned and re-proportioned, and projected back to the 2-D points the model reads — two new signers per clip. The boxes below are measured: stacked with the real infusion clips (jepa-infuse4) the five-seed cross-validation mean rose 9.1 pp; changing body proportions by ±10% changed nothing, and child proportions cost 2.3 pp.
The synthetic signers are an example of a technique kept for a measured reason, not an assumed one: alone they gave about nothing, stacked with infusion the recipe's five-seed cross-validation mean went from 0.3018 to 0.3927 (+9.1 pp), and scaling body proportions by ±10% changed the result by 0.0 pp — the model was already body-invariant.
7 · Ensemble → knowledge distillation
Fine-tuning is noisy from seed to seed, and an average of several fine-tunes is more accurate than any one of them — but twelve models do not fit in a browser. So the twelve (four on Phase B encoders, eight on Phase C encoders) become a teacher: at every training step their temperature-softened probabilities are averaged, and one student — a Phase B encoder cut at layer 4, the size of the deployed model — is trained against both the label and the teacher:
ℒ = (1 − α) · CE0.1(y, σ(zs)) + α · T² · KL( pt(T) ‖ σ(zs / T) ), pt(T) = (1/N) Σk σ(zk / T)
How to read this. Twelve fine-tunes on the left are averaged into one soft teacher. The student on the right learns from the label and from the teacher at once, then is exported int8. The annotations compare the ensemble with its student: accuracy test 0.5576 → 0.5509, unseen signer 0.5045 → 0.5000 — almost all of the ensemble kept in one 4.57 MB model.
The student keeps almost all of the ensemble: accuracy test 0.5576 → 0.5509, unseen signer 0.5045 → 0.5000, in a single 4.57 MB model.
C · How the model is packed for the browser
8 · int8 ONNX export
The student is exported to ONNX (opset 17) with the softmax inside the graph, then quantised to int8 with dynamic quantisation. Before a bundle is accepted, the exported graph must agree with the PyTorch model to within 1.3 × 10⁻⁶; the int8 graph is then scored on the accuracy test like any other model, because quantisation can move a number either way (jepa-x12: unseen signer 0.5000 fp32, 0.4909 int8).
9 · Calibration, cross-checked
A probability shown to a learner in Practice mode must mean what it says. One temperature T is fitted to the log-probabilities of the mirror-pass posterior:
p̂(c | x) = exp(ℓc / T) / Σc′ exp(ℓc′ / T)
The cross-check is the point: a temperature that only works on the set it was fitted on is fitted noise. With the calibrated confidence, a Practice-mode floor of 0.5 credits 69% of correct signs at 83% precision for jepa-x12 (jepa-rr4: 61% at 82%).
D · How the data and the experiments stay honest
10 · The dataset quality system
Every clip of the four corpora can carry marks — at the clip, word or corpus-rule level, from
a catalogue of known issues, raised by an agent or a person, with a needsReview flag. A command
(dataset-marks.ts apply) turns the marks into a contract that the training packs are built
from, so a decision about a clip is made once and every later model inherits it. Hard cases go
through rounds of vision-model review and then a blind review of the real videos with hidden
controls — clips whose answer is already known, mixed in to check the reviewer (Claude answered
10/10 of the committed controls correctly). Whatever is still open goes to a Review tab for a
human decision. The loop and its numbers are on the report page.
11 · Compute harness
Fine-tunes and distillations run on Modal A10G GPUs through one harness script
(gauntlet.py): about 5 minutes per fine-tune seed and about 17 minutes per distillation, which
is what makes ≥ 4 seeds per claim affordable. Earlier phases ran on Colab Pro and Beam; the
infrastructure page records that history.