Skip to content

Architecture · 3 of 5

Evaluation: the gate, and what it refuses to be fooled by

One number decides whether a model ships: clip-weighted top-1 over 599 clips the model has never seen, from a signer it has never seen and a corpus it has. Everything else — in-corpus accuracy, frozen probes, single-seed wins — has been shown to mislead on this problem, and this page records how.

Figure 3The accuracy test — two held-out corpora, one weighted number, three side reports

How to read this. The two grey boxes on the left are the only inputs; neither is ever trained on. They combine into one clip-weighted top-1 — the gate. The three outlined boxes are reported beside it but never decide. The INCLUDE split enters only as a dotted guardrail: it can veto a model that forgot the 260-sign set, never promote one.

  • Held-out corpora
  • The deciding number
  • Side reports
Figure 3The accuracy test — two held-out corpora, one weighted number, three side reports

The gate

Two-corpus gate = clip-weighted top-1 over RKMVU-220 (220 clips, one unseen signer) and CISLR test-role (379 clips, unseen clips from a corpus the encoder has seen during pretraining) — 599 clips, so one clip is worth 0.17 pp. Every model is reported with:

ColumnWhat it measuresCurrent best (jepa-rr4)
Gateclip-weighted top-1 over the 5990.4591
RKMVU top-1 / top-5the cross-signer number on its own0.4818 (int8 0.4864) / pending
Held-out odd classesthe 104 RKMVU clips whose classes never entered any pool, even unlabelled — the anti-transductive check0.5481
Mirrored inputthe same clips flipped left↔right: hand-dominance invariancereported per run
Mirror-TTA averageaverage of p(x) and p(mirror x)4-seed rr ensemble: 0.4574 → 0.4641
CISLR-test top-1the seen-corpus, unseen-clip number0.4459

Two corpora rather than one because they fail differently. RKMVU asks does this generalise to a new person; CISLR-test asks does this generalise to new clips of people it has seen. A recipe can trade one for the other — the interp arm in the September screen gained +3.2 pp on RKMVU and lost 5.9 pp on CISLR — and the weighted gate refuses that trade.

Why in-corpus is only a guardrail

The INCLUDE held-out split sits at roughly 0.80–0.83 for every model in the lineage. It is reported, and a model that dropped far below it would be rejected for forgetting the closed set, but it has no power to promote a model — because it already failed to predict the number that matters. At Review 1 the supervised model measured 0.82 in-corpus and 0.33 on the unseen signer. Within-corpus validation, even signer-disjoint, cannot see a distribution-wide style shift; only a held-out corpus can.

Mirror test-time augmentation

A left-hand-dominant signer produces the mirror image of the reference performance. Because the canonical normalisation is mirror-symmetric, the mirror map is exact in landmark space, so the studio scores every window in both orientations and the more confident side answers; in evaluation the two probabilities are averaged. The ensemble's +0.7 pp from TTA (0.4574 → 0.4641) is the residual asymmetry that training on both orientations has not yet removed.

Why probes mislead on this encoder

Pretraining objective (Phase C pilot, 12 ep, 30k pool subset)Probe RKMVUFT gate seed 1 / seed 2Mean gateRKMVUHeld-out
Phase B (last-layer target)0.1864.3472 / .3740.3606.3341.4135
+ motion targets0.1091.3656 / .3222.3439.3295.3894
+ all-layer teacher0.1773.4007 / .4157.4082.3955.4808
+ motion + all-layer0.1682.3740 / .3973.3856.3841.4519

The same lesson appeared earlier in Phase B itself: the full-depth probe read 0.0182 on the unseen signer — chance — while a layer-4 probe read 0.2409 and a layer-4 fine-tune reached 0.3273. The representation was never "collapsed"; the probe was reading the wrong layer.

Seeds, noise, and what counts as a win

With 599 clips, a 1 pp difference is six clips. Every recipe decision is made on at least two seeds, and the current best on four. A change ships when the mean moves and most seeds move with it — rr qualified with +1.7 pp mean, 3 of 4 seeds up, and half the seed spread of the control. A single-seed +3 pp is logged as a lead, not a result.