Review 2 · Deliverable 2 of 4
Intermediate results obtained
One browser model of the same size and speed went from 0.4591 to 0.5509 on the accuracy test. This page shows how that was measured, what produced it, how much of it is the unseen signer (less than the headline suggests), and the experiments that did not work — which shape the plan as much as the ones that did.
The current best: jepa-x12
0.5509
Accuracy test
599 clips · jepa-rr4 was 0.4591
0.5000
Unseen signer (RKMVU)
int8 0.4909 · jepa-rr4 0.4818
0.5805
CISLR test
379 clips never trained on
0.5546
In the app
int8 + mirror pass · was 0.4538
0.6909
Unseen signer, top-5
right sign among the first five
4.57 MB
int8 model
same size as jepa-rr4
12 → 1
Distilled
twelve fine-tunes, one student
0.788
Temperature
calibration, cross-checked
jepa-x12 is the student of a twelve-model ensemble (see distillation). It is the model frozen for Review 2 and the app's default 260-sign model from 27 September (it replaced jepa-rr4).
The accuracy test
How to read this. Two test sets feed one number. RKMVU is 220 clips of one signer the model has never seen — the stranger number. CISLR test is 379 clips never trained on, but its presenters also appear in CISLR training clips. Together they are 599 clips, so one clip is 0.17 pp. The side checks are what a result must also pass before it is believed; the rule at the end is the one that promotes a model.
The accuracy test (called "the gate" in the earlier reports) is the clip-weighted top-1 accuracy over 599 clips: RKMVU-220, one signer who appears nowhere in training, and CISLR test-role 379, clips the model is never trained on. One clip is 0.17 pp, so differences under about half a point are single clips. Because CISLR's test clips share presenters with its training clips, the RKMVU number is reported beside every accuracy-test number — it is the only one that measures a true stranger.
Model lineage
All numbers fp32 unless marked; top-1.
| Model | Date | Accuracy test | RKMVU (unseen) | CISLR test | Notes |
|---|---|---|---|---|---|
| Supervised v3.5.2 (full-vocabulary encoder; the old 260 model) | Aug | .1469 (browser) | .3318 | .0396 | 9.90 MB · 34.6 ms/window |
| Review 0 in-corpus reference | Aug | — | — | — | INCLUDE-263 top-1 0.958 in-corpus; CISLR 8,165-word open vocabulary top-1 0.2429 / top-20 0.4105 |
| Review 1 finding | 30 Aug | — | 0.33 cross-corpus vs 0.82 in-corpus | — | left-dominant signer 0.12 → 0.33 with the mirror pass |
| jepa-ship4 (Phase B JEPA, cut at layer 4) | 1 Sep | — | .3273 · int8 .3318 | — | 4.57 MB |
| jepa-infuse4 (+ gloss infusion + synthetic signers) | 2 Sep | .4124 | .4318 | — | 5-seed mean cross-val .3018 → .3927 (+9.1 pp) |
| jepa-distil4 (4-seed ensemble → student) | 3 Sep | .4424 | .4545 | .4354 | |
| jepa-rr4 (+ recover-and-resample) — deployed since 4 Sep | 3 Sep | .4591 | .4818 · int8 .4864 | .4459 | 12.5 ms/window · 4.57 MB |
| jepa_c4 (Phase C all-layer teacher) | 4 Sep | .4524 | — | — | a wash against jepa-rr4 |
| jepa-v4 (reviewed v4 pack, 4 teachers → student) | 26 Sep | .5359 | .4818 · int8 .5000 | .5673 | in the app .5412 |
| jepa-x12 (12 teachers → Phase B student) | 26 Sep | .5509 | .5000 · int8 .4909 | .5805 | in the app .5546 · T = 0.788 · top-5 RKMVU .6909 |
| 12-model ensemble (Mac only) | 26 Sep | .5576 · + mirror .5609 | .5045 | .5884 | not shippable to a browser as it is |
| Frozen for Review 2 = jepa-x12 | 27 Sep | .5509 | .5000 · int8 .4909 | .5805 | kept by the rule stated before scoring: the 16-teacher student could not finish once Modal ended |
- Accuracy test (599 clips)
- RKMVU — unseen signer (220)
- CISLR test (379)
- Review 1 target 0.50
How to read this. Time runs left to right, with the review dates as milestones. Each marker is a model with the number that decided it. Between 4 and 26 September the work was the dataset review; the two 26 September models are trained on its reviewed contract.
Where the gain came from
The accuracy test rose by +9.2 pp (0.4591 → 0.5509). Two causes can be separated because each was measured on its own:
-
The dataset contract: +7.0 pp over four seeds. The same recipe was trained four times on the old pack and four times on the reviewed v4 pack, in the same run:
Pack Seed 1 Seed 2 Seed 3 Seed 4 Mean Control (champion recipe, old pack) .4057 .4174 .4124 .4107 .4116 Reviewed v4 pack .4825 .4942 .4691 .4808 .4817 All four seeds went up. The review was about sense conflation — whether the clips filed under one word really show one sign, and whether two words are really the same sign — see the dataset review loop.
-
Ensembling and distillation: the rest. Four-seed ensembles of the same two packs:
Ensemble Accuracy test RKMVU CISLR Control, 4 seeds .4574 .4727 .4485 v4 pack, Phase B, 4 seeds .5376 .4909 .5646 Twelve fine-tunes distilled into one student (jepa-x12) took it to .5509.
What the app shows
The harness scores fp32 models; the app runs the int8 model with the mirror pass. Scored the way the app runs them, on the same 599 clips:
| Model | In the app (int8 + mirror pass) |
|---|---|
| jepa-rr4 (current default) | .4538 |
| jepa-v4 | .5412 |
| jepa-x12 | .5546 |
In Practice mode a sign is credited when its calibrated confidence clears 0.5. At that floor jepa-x12 credits 69% of correct signs at 83% precision; jepa-rr4 credited 61% at 82%. More correct attempts are recognised, and a credited attempt is slightly more likely to be right.
The honest reading
- The Review 1 target (≥ 0.50) is met on the accuracy test (0.5509) and borderline on the unseen signer (0.5000 fp32, 0.4909 int8).
- The ensemble is still ahead of its student (0.5576 vs 0.5509), so a better student is possible without new data — but on the unseen signer the gap is 0.5045 vs 0.5000, about one clip.
- Every claim on this page rests on at least four seeds.
Negative results kept
Each of these was a reasonable idea, measured properly, that did not help. They are kept because each one removes a direction from the plan.
| Idea | What was measured | Result |
|---|---|---|
| RTMW whole-body landmarks instead of MediaPipe | same trainer, 5 seeds, all 599 clips | MediaPipe wins the accuracy test (.257 vs .238). RTMW's +13 pp on INCLUDE's own signer split (.497 → .625) does not transfer: CISLR −3 to −8 pp, and RTMW invents legs on waist-framed video |
| Image pre-processing for missing hands (gamma, CLAHE, sharpen, denoise, upscale) | hand-detection recovery | null or negative. Hands go missing because they move (wrist speed 5.3 → 12.5 px/frame, in 97% of clips) or leave the frame (68% of misses) |
| Shape match as an RL reward / pseudo-label checker | precision at equal coverage | worse than the model's own confidence: 55.3% vs 68.0% |
| Body proportions ±10% as augmentation | accuracy test | 0.0 pp — the model is already body-invariant; child proportions −2.3 pp |
| Phase C objective (all-layer teacher) at full budget | accuracy test, unseen signer | a wash (jepa_c4 .4524 vs .4591); Phase C students lose on RKMVU |
| The boundary head for trimming idle frames | frame labels | labels 94.9% of frames "inside a sign", so it cannot trim idle lead-in or tail; the v3.5 export also dropped its output |
Shape match, measured as a helper
How to read this. Top: the AI model proposes its top words and shape match chooses among them — better than shape match over all 260, never better than the AI model alone. Bottom: as a checker of pseudo-labels, shape match is less precise than the model's own confidence at the same coverage (55.3% vs 68.0%), so it is not used as a training reward.
Shape match has no learned weights, so it was tested in the two roles where it could help the AI model rather than compete with it:
| Role | Shape match | Compared with |
|---|---|---|
| Choosing among all 260 signs | 25.0% | AI model alone: 45.7% |
| Choosing within the AI model's top 3 | 34.7% | AI model alone: 45.7% |
| Checking pseudo-labels (precision at equal coverage) | 55.3% | the AI model's own confidence: 68.0% |
A shortlist helps shape match by almost ten points, but it never overtakes the model it is reading from, and as a label checker it is worse than the model's confidence. It stays in the product as a transparent second opinion and as a control, not as a training signal.
What the data review found
- ISLRTC dictionary: 64% of files are longer than 8 s — lessons, not single words; 17% are silently cut at 36 s; 767 "words" exist only as file names.
- INCLUDE: 76% of frames are idle (before or after the sign).
- Vision review: 75 escalations; 35 settled by blind review of the real videos (21 by INCLUDE's own ground truth, 6 confirmed different signs, 8 rejected as the same sign); 40 open in the Review tab for a human decision.