Skip to content

Review 2 · Deliverable 2 of 4

Intermediate results obtained

One browser model of the same size and speed went from 0.4591 to 0.5509 on the accuracy test. This page shows how that was measured, what produced it, how much of it is the unseen signer (less than the headline suggests), and the experiments that did not work — which shape the plan as much as the ones that did.

The current best: jepa-x12

0.5509

Accuracy test

599 clips · jepa-rr4 was 0.4591

0.5000

Unseen signer (RKMVU)

int8 0.4909 · jepa-rr4 0.4818

0.5805

CISLR test

379 clips never trained on

0.5546

In the app

int8 + mirror pass · was 0.4538

0.6909

Unseen signer, top-5

right sign among the first five

4.57 MB

int8 model

same size as jepa-rr4

12 → 1

Distilled

twelve fine-tunes, one student

0.788

Temperature

calibration, cross-checked

jepa-x12 is the student of a twelve-model ensemble (see distillation). It is the model frozen for Review 2 and the app's default 260-sign model from 27 September (it replaced jepa-rr4).

The accuracy test

Figure 1The accuracy test — two sets, one number

How to read this. Two test sets feed one number. RKMVU is 220 clips of one signer the model has never seen — the stranger number. CISLR test is 379 clips never trained on, but its presenters also appear in CISLR training clips. Together they are 599 clips, so one clip is 0.17 pp. The side checks are what a result must also pass before it is believed; the rule at the end is the one that promotes a model.

Figure 1The accuracy test — two sets, one number

The accuracy test (called "the gate" in the earlier reports) is the clip-weighted top-1 accuracy over 599 clips: RKMVU-220, one signer who appears nowhere in training, and CISLR test-role 379, clips the model is never trained on. One clip is 0.17 pp, so differences under about half a point are single clips. Because CISLR's test clips share presenters with its training clips, the RKMVU number is reported beside every accuracy-test number — it is the only one that measures a true stranger.

Model lineage

All numbers fp32 unless marked; top-1.

ModelDateAccuracy testRKMVU (unseen)CISLR testNotes
Supervised v3.5.2 (full-vocabulary encoder; the old 260 model)Aug.1469 (browser).3318.03969.90 MB · 34.6 ms/window
Review 0 in-corpus referenceAug———INCLUDE-263 top-1 0.958 in-corpus; CISLR 8,165-word open vocabulary top-1 0.2429 / top-20 0.4105
Review 1 finding30 Aug—0.33 cross-corpus vs 0.82 in-corpus—left-dominant signer 0.12 → 0.33 with the mirror pass
jepa-ship4 (Phase B JEPA, cut at layer 4)1 Sep—.3273 · int8 .3318—4.57 MB
jepa-infuse4 (+ gloss infusion + synthetic signers)2 Sep.4124.4318—5-seed mean cross-val .3018 → .3927 (+9.1 pp)
jepa-distil4 (4-seed ensemble → student)3 Sep.4424.4545.4354
jepa-rr4 (+ recover-and-resample) — deployed since 4 Sep3 Sep.4591.4818 · int8 .4864.445912.5 ms/window · 4.57 MB
jepa_c4 (Phase C all-layer teacher)4 Sep.4524——a wash against jepa-rr4
jepa-v4 (reviewed v4 pack, 4 teachers → student)26 Sep.5359.4818 · int8 .5000.5673in the app .5412
jepa-x12 (12 teachers → Phase B student)26 Sep.5509.5000 · int8 .4909.5805in the app .5546 · T = 0.788 · top-5 RKMVU .6909
12-model ensemble (Mac only)26 Sep.5576 · + mirror .5609.5045.5884not shippable to a browser as it is
Frozen for Review 2 = jepa-x1227 Sep.5509.5000 · int8 .4909.5805kept by the rule stated before scoring: the 16-teacher student could not finish once Modal ended
The accuracy test, the unseen signer and CISLR test across the lineage
  • Accuracy test (599 clips)
  • RKMVU — unseen signer (220)
  • CISLR test (379)
  • Review 1 target 0.50
v3.5.2 · Augsupervised · test run in browser
.1469
.3318
.0396
jepa-ship4 · 1 Sep
not measured
.3273
not measured
jepa-infuse4 · 2 Sep
.4124
.4318
not measured
jepa-distil4 · 3 Sep
.4424
.4545
.4354
jepa-rr4 · 3 Sepdeployed since 4 Sep
.4591
.4818
.4459
jepa_c4 · 4 Sep
.4524
not measured
not measured
jepa-v4 · 26 Sep
.5359
.4818
.5673
jepa-x12 · 26 Sepcurrent best
.5509
.5000
.5805
0.00.10.20.30.40.50.6
Figure 2Model lineage, Review 0 → Review 2

How to read this. Time runs left to right, with the review dates as milestones. Each marker is a model with the number that decided it. Between 4 and 26 September the work was the dataset review; the two 26 September models are trained on its reviewed contract.

Figure 2Model lineage, Review 0 → Review 2

Where the gain came from

The accuracy test rose by +9.2 pp (0.4591 → 0.5509). Two causes can be separated because each was measured on its own:

  1. The dataset contract: +7.0 pp over four seeds. The same recipe was trained four times on the old pack and four times on the reviewed v4 pack, in the same run:

    PackSeed 1Seed 2Seed 3Seed 4Mean
    Control (champion recipe, old pack).4057.4174.4124.4107.4116
    Reviewed v4 pack.4825.4942.4691.4808.4817

    All four seeds went up. The review was about sense conflation — whether the clips filed under one word really show one sign, and whether two words are really the same sign — see the dataset review loop.

  2. Ensembling and distillation: the rest. Four-seed ensembles of the same two packs:

    EnsembleAccuracy testRKMVUCISLR
    Control, 4 seeds.4574.4727.4485
    v4 pack, Phase B, 4 seeds.5376.4909.5646

    Twelve fine-tunes distilled into one student (jepa-x12) took it to .5509.

What the app shows

The harness scores fp32 models; the app runs the int8 model with the mirror pass. Scored the way the app runs them, on the same 599 clips:

ModelIn the app (int8 + mirror pass)
jepa-rr4 (current default).4538
jepa-v4.5412
jepa-x12.5546

In Practice mode a sign is credited when its calibrated confidence clears 0.5. At that floor jepa-x12 credits 69% of correct signs at 83% precision; jepa-rr4 credited 61% at 82%. More correct attempts are recognised, and a credited attempt is slightly more likely to be right.

The honest reading

  • The Review 1 target (≥ 0.50) is met on the accuracy test (0.5509) and borderline on the unseen signer (0.5000 fp32, 0.4909 int8).
  • The ensemble is still ahead of its student (0.5576 vs 0.5509), so a better student is possible without new data — but on the unseen signer the gap is 0.5045 vs 0.5000, about one clip.
  • Every claim on this page rests on at least four seeds.

Negative results kept

Each of these was a reasonable idea, measured properly, that did not help. They are kept because each one removes a direction from the plan.

IdeaWhat was measuredResult
RTMW whole-body landmarks instead of MediaPipesame trainer, 5 seeds, all 599 clipsMediaPipe wins the accuracy test (.257 vs .238). RTMW's +13 pp on INCLUDE's own signer split (.497 → .625) does not transfer: CISLR −3 to −8 pp, and RTMW invents legs on waist-framed video
Image pre-processing for missing hands (gamma, CLAHE, sharpen, denoise, upscale)hand-detection recoverynull or negative. Hands go missing because they move (wrist speed 5.3 → 12.5 px/frame, in 97% of clips) or leave the frame (68% of misses)
Shape match as an RL reward / pseudo-label checkerprecision at equal coverageworse than the model's own confidence: 55.3% vs 68.0%
Body proportions ±10% as augmentationaccuracy test0.0 pp — the model is already body-invariant; child proportions −2.3 pp
Phase C objective (all-layer teacher) at full budgetaccuracy test, unseen signera wash (jepa_c4 .4524 vs .4591); Phase C students lose on RKMVU
The boundary head for trimming idle framesframe labelslabels 94.9% of frames "inside a sign", so it cannot trim idle lead-in or tail; the v3.5 export also dropped its output

Shape match, measured as a helper

Figure 3Shape match as a helper, measured

How to read this. Top: the AI model proposes its top words and shape match chooses among them — better than shape match over all 260, never better than the AI model alone. Bottom: as a checker of pseudo-labels, shape match is less precise than the model's own confidence at the same coverage (55.3% vs 68.0%), so it is not used as a training reward.

Figure 3Shape match as a helper, measured

Shape match has no learned weights, so it was tested in the two roles where it could help the AI model rather than compete with it:

RoleShape matchCompared with
Choosing among all 260 signs25.0%AI model alone: 45.7%
Choosing within the AI model's top 334.7%AI model alone: 45.7%
Checking pseudo-labels (precision at equal coverage)55.3%the AI model's own confidence: 68.0%

A shortlist helps shape match by almost ten points, but it never overtakes the model it is reading from, and as a label checker it is worse than the model's confidence. It stays in the product as a transparent second opinion and as a control, not as a training signal.

What the data review found

  • ISLRTC dictionary: 64% of files are longer than 8 s — lessons, not single words; 17% are silently cut at 36 s; 767 "words" exist only as file names.
  • INCLUDE: 76% of frames are idle (before or after the sign).
  • Vision review: 75 escalations; 35 settled by blind review of the real videos (21 by INCLUDE's own ground truth, 6 confirmed different signs, 8 rejected as the same sign); 40 open in the Review tab for a human decision.