Current state · as of 2 October 2026
Results and model lineage
Every model this program has produced, scored the same way, in the order it was produced. The unseen-signer number moved from 0.33 to 0.48 in four days of early September, to 0.50 by the 26th and to 0.51 on 2 October, while the accuracy test rose to 0.55; this page records each step and what caused it. The Review 2 pages carry the full reading up to 26 September.
The deployed model: jepa-asl-b3
0.5526
Accuracy test
599 held-out clips · clip-weighted top-1
0.5136
RKMVU top-1
unseen signer · int8 0.5091
0.5752
CISLR test top-1
379 clips never trained on
0.5136
In the app, unseen signer
int8 + mirror pass · jepa-x12 0.5000
jepa-asl-b3 (2 October) is one student with the phonological head, distilled from twelve fine-tunes whose encoders first learned American Sign Language: Google's Isolated Sign Language Recognition corpus — 94,290 landmark clips of 250 ASL signs by 21 Deaf signers. It replaced jepa-x12 as the deployed 260-sign model by one to three clips on identical inputs, and its confidence is better calibrated (expected calibration error 0.039 against 0.056). What the ASL data did and did not do is below.
What changed on 2 October — ASL transfer
Twenty-one Deaf ASL signers are more signer variety than every Indian corpus this program has combined, so the question was whether learning ASL first makes the encoder more signer-invariant for ISL. All arms ran on the v4 recipe, four seeds each, against a same-harness control:
| Arm | RKMVU (4-seed mean) | Gate | 4-model ensemble RKMVU / gate |
|---|---|---|---|
| Control | 0.4296 | 0.4733 | 0.5000 / 0.5309 |
| ASL auxiliary head only | 0.4375 | 0.4934 | 0.4773 / 0.5326 |
| ASL first (6 epochs), then ISL | 0.4523 | 0.5109 | 0.5000 / 0.5659 |
| ASL first + ASL head kept on | 0.4648 | 0.5050 | 0.5091 / 0.5509 |
| ASL first (16 epochs) + phonological head | 0.4568 | 0.5209 | 0.4818 / 0.5509 |
Two follow-ups on 3 October, both negative and kept here:
- A small ensemble in the browser. Two or three of the int8 models run side by side and combined with the app's own mirror rule, on sets named before scoring: none beat jepa-asl-b3 alone on the unseen signer (best 0.5091 against 0.5136); the best sets add about 0.8 pp on the accuracy test for three times the download and inference. Not shipped.
- Self-supervised pre-training on the ASL clips. The Phase B encoder trained further on the unlabelled pool with the 94,290 ASL clips added, against the same training on the pool alone, each then fine-tuned four times: unseen signer 0.4273 and 0.4250 against 0.4296 for Phase B itself. The ASL clips help only when their labels are used.
Before it: jepa-x12
0.5509
Accuracy test
599 held-out clips · clip-weighted top-1
0.5000
RKMVU top-1
unseen signer · int8 0.4909
0.5805
CISLR test top-1
379 clips never trained on
0.5546
In the app
int8 + mirror pass · jepa-rr4 0.4538
jepa-x12 (26 September) is twelve fine-tunes distilled into one Phase B student of the same size as jepa-rr4, trained on the reviewed v4 dataset pack. It was the deployed default from 27 September (the model frozen for Review 2) to 2 October. Most of its gain over jepa-rr4 is on CISLR (+13.5 pp); the unseen signer moved +1.8 pp. The measurements, the seed tables and that reading are on the Review 2 results page.
Deployed until 27 September: jepa-rr4
0.4591
Two-corpus gate
599 held-out clips · clip-weighted top-1
0.4818
RKMVU top-1
unseen signer · int8 0.4864
0.5481
Held-out classes
104 clips never in any pool
0.4459
CISLR test top-1
379 clips never trained on
4.57 MB
int8 model
was 9.90 MB supervised
+15.0 pp
vs deployed supervised
RKMVU 0.3318 → 0.4818
0.4641
4-seed ensemble + TTA
the teacher it was distilled from
4
seeds
rr recipe · gates .4057 / .4174 / .4124 / .4107
jepa-rr4 is a distilled student of a four-seed ensemble, fine-tuned with the
all+synth:full@2 + rr recipe on the Phase B encoder cut at layer 4, exported to int8. How it
is made is the training program; how it is scored is the
gate.
Model lineage
All numbers are cross-signer RKMVU top-1 unless the column says otherwise. "Gate" is the two-corpus number that decides.
| Model | When | RKMVU | Gate | Notes |
|---|---|---|---|---|
| Supervised v3.5.2 — the deployed retrieval-lane encoder | Aug | 0.3318 | — | 9.90 MB; the baseline every later model is measured against |
Phase B JEPA final100@4 (jepa-ship4) | 1 Sep | 0.3273 fp32 / 0.3318 int8 | — | 4.57 MB; last-layer JEPA collapsed (target-std 0.98 → 0.52); only layers ≤ 4 transfer |
+ gloss infusion + synthetic signers all+synth:full@2 (jepa-infuse4, seed 1) | 2 Sep | 0.4318 (int8 0.4364) | 0.4124 | 5-seed mean cross-val 0.3927, from 0.3018 |
| 4-seed ensemble → distilled student (jepa-distil4) | 3 Sep 01:00 | 0.4545 | 0.4424 | held-out 0.5385 · CISLR 0.4354 |
+ recover-and-resample rr → distilled (jepa-rr4) | 3 Sep 14:25 | 0.4818 (int8 0.4864) | 0.4591 | held-out 0.5481 · CISLR 0.4459 · 4-seed rr ensemble 0.4574, +TTA 0.4641 |
Phase C proper — all-layer teacher, 100 ep, full pool → jepa_c4 | 4 Sep | — | 0.4524 | a wash against jepa-rr4; pilot had shown +4.8 pp at equal budget |
| Reviewed v4 pack, 4 teachers → student (jepa-v4) | 26 Sep | 0.4818 (int8 0.5000) | 0.5359 | CISLR 0.5673 · in the app 0.5412 |
| 12 teachers → Phase B student (jepa-x12) | 26 Sep | 0.5000 (int8 0.4909) | 0.5509 | CISLR 0.5805 · in the app 0.5546 · top-5 RKMVU 0.6909 |
| 12-model ensemble (Mac only, not browser-shippable) | 26 Sep | 0.5045 | 0.5576 (+ mirror 0.5609) | CISLR 0.5884 |
| ASL-pretrained teachers (Google ISLR) → phonological-head student (jepa-asl-b3) | 2 Oct | 0.5136 (int8 0.5091) | 0.5526 | CISLR 0.5752 · in the app 0.5136 unseen signer · seeds 1–3 RKMVU .4955 / .4955 / .5136 |
What changed on 3 September
A night and a day, in the order it happened.
The recipe screen on the wrong encoder
Ten candidate recipe changes were screened, two seeds each, against an in-harness control — and nothing was positive on the gate:
| Arm | Δ gate | Arm | Δ gate |
|---|---|---|---|
| swad (weight averaging) | +0.0 | interp (interpolated resample) | −2.6 (RKMVU +3.2, CISLR −5.9) |
| lw8 (8-layer) | −0.3 | vocab | −3.0 |
| rr | −0.3 | lpft (linear-probe-then-fine-tune) | −5.3 |
| angles | −0.5 | rot | −1.3 |
| nojit (no jitter) | −1.2 | cut6 | −1.9 |
The re-screen on the real encoder
On the Phase B encoder, the in-harness control reproduces the champion exactly (per-seed gates
.4140 / .3923 / .3773 / .3940). Against it, rr is the one winner: four seeds at
.4057 / .4174 / .4124 / .4107, +1.7 pp gate, +2.5 pp RKMVU, three of four seeds up, and half
the seed spread. swad −0.3, angles +0.1 (mixed), lw8 −1.4. rr drops up to 25% off either end
of a clip and stretches it back to 64 frames, with probability 0.5.
Ensemble, distil, ship
The four rr seeds were ensembled (gate 0.4574; 0.4641 with mirror TTA) and distilled into jepa-rr4 at 14:25 — the current best, exported and dropped into the Lab.
The Phase C pilot
In parallel, four pretraining objectives were compared at 12 epochs on a random 30k pool subset: the all-layer teacher gives +4.8 pp gate at equal budget; motion targets hurt. Phase C proper was launched on the full pool. The table and the reasoning are on the JEPA page.
Also fixed on 3 September
- Demo-table leak. Stale windows from a previous capture epoch could be voted into the next
row ("husband 77%" on the "dead" row). Fixed with capture epochs in
LandmarkRing, acapture-startmessage to every worker, and a per-row window vote — runtime page. - Five training-side bugs: the LP-FT encoder learning rate stuck at 0 (cosine scheduler recursion); a head-slice bias in pool subsets; Colab cell interrupts killing child processes; ONNX missing on fresh VMs; a gate-scorer name mismatch.
Phase C proper — a wash at full budget
| Epoch | 0 | 10 | 20 | 30 | 40 | 50 | 60 | 70 |
|---|---|---|---|---|---|---|---|---|
| Pretext loss | 0.654 | 0.235 | 0.203 | 0.191 | 0.179 | 0.157 | 0.158 | 0.145 |
Target standard deviation pinned at 0.999 (no collapse); checkpoints to Drive every five epochs;
auto-resume proven at the epoch-11 idle-timeout kill. Fine-tuned, jepa_c4 scored 0.4524 on
the gate against jepa-rr4's 0.4591 — a wash — and Phase C students lose on the unseen signer. The
pilot's +4.8 pp did not survive the full budget. Phase C encoders still serve as eight of the
twelve teachers behind jepa-x12.
What changed between 4 and 26 September
The work moved from the model to the labels. A review of the training data for sense conflation — clips filed under one word that are really different signs, and the reverse — produced a reviewed dataset contract (the v4 pack). The same recipe trained on it scored +7.0 pp on the gate over four seeds, all four up; ensembling twelve fine-tunes and distilling them gave jepa-x12. The loop, the seed tables and the negative results are on the Review 2 pages.
What this means for the review targets
Review 1 pre-registered a Phase B target of ≥ 0.50. Phase B alone missed it (0.3273). The fine-tuning program on the Phase B encoder reached 0.4818 on the unseen signer by 4 September, and jepa-x12 stood at 0.5509 on the gate — met — and 0.5000 on the unseen signer (0.4909 int8) — borderline. Its successor jepa-asl-b3 reads 0.5526 and 0.5136 (0.5091 int8), which clears the unseen-signer target by a few clips. The targets are scored in full on the Review 2 outcomes page.