Architecture · 3 of 5
Evaluation: the gate, and what it refuses to be fooled by
One number decides whether a model ships: clip-weighted top-1 over 599 clips the model has never seen, from a signer it has never seen and a corpus it has. Everything else — in-corpus accuracy, frozen probes, single-seed wins — has been shown to mislead on this problem, and this page records how.
How to read this. The two grey boxes on the left are the only inputs; neither is ever trained on. They combine into one clip-weighted top-1 — the gate. The three outlined boxes are reported beside it but never decide. The INCLUDE split enters only as a dotted guardrail: it can veto a model that forgot the 260-sign set, never promote one.
- Held-out corpora
- The deciding number
- Side reports
The gate
Two-corpus gate = clip-weighted top-1 over RKMVU-220 (220 clips, one unseen signer) and CISLR test-role (379 clips, unseen clips from a corpus the encoder has seen during pretraining) — 599 clips, so one clip is worth 0.17 pp. Every model is reported with:
| Column | What it measures | Current best (jepa-rr4) |
|---|---|---|
| Gate | clip-weighted top-1 over the 599 | 0.4591 |
| RKMVU top-1 / top-5 | the cross-signer number on its own | 0.4818 (int8 0.4864) / pending |
| Held-out odd classes | the 104 RKMVU clips whose classes never entered any pool, even unlabelled — the anti-transductive check | 0.5481 |
| Mirrored input | the same clips flipped left↔right: hand-dominance invariance | reported per run |
| Mirror-TTA average | average of p(x) and p(mirror x) | 4-seed rr ensemble: 0.4574 → 0.4641 |
| CISLR-test top-1 | the seen-corpus, unseen-clip number | 0.4459 |
Two corpora rather than one because they fail differently. RKMVU asks does this generalise to
a new person; CISLR-test asks does this generalise to new clips of people it has seen. A
recipe can trade one for the other — the interp arm in the September screen gained +3.2 pp
on RKMVU and lost 5.9 pp on CISLR — and the weighted gate refuses that trade.
Why in-corpus is only a guardrail
The INCLUDE held-out split sits at roughly 0.80–0.83 for every model in the lineage. It is reported, and a model that dropped far below it would be rejected for forgetting the closed set, but it has no power to promote a model — because it already failed to predict the number that matters. At Review 1 the supervised model measured 0.82 in-corpus and 0.33 on the unseen signer. Within-corpus validation, even signer-disjoint, cannot see a distribution-wide style shift; only a held-out corpus can.
Mirror test-time augmentation
A left-hand-dominant signer produces the mirror image of the reference performance. Because the canonical normalisation is mirror-symmetric, the mirror map is exact in landmark space, so the studio scores every window in both orientations and the more confident side answers; in evaluation the two probabilities are averaged. The ensemble's +0.7 pp from TTA (0.4574 → 0.4641) is the residual asymmetry that training on both orientations has not yet removed.
Why probes mislead on this encoder
| Pretraining objective (Phase C pilot, 12 ep, 30k pool subset) | Probe RKMVU | FT gate seed 1 / seed 2 | Mean gate | RKMVU | Held-out |
|---|---|---|---|---|---|
| Phase B (last-layer target) | 0.1864 | .3472 / .3740 | .3606 | .3341 | .4135 |
| + motion targets | 0.1091 | .3656 / .3222 | .3439 | .3295 | .3894 |
| + all-layer teacher | 0.1773 | .4007 / .4157 | .4082 | .3955 | .4808 |
| + motion + all-layer | 0.1682 | .3740 / .3973 | .3856 | .3841 | .4519 |
The same lesson appeared earlier in Phase B itself: the full-depth probe read 0.0182 on the unseen signer — chance — while a layer-4 probe read 0.2409 and a layer-4 fine-tune reached 0.3273. The representation was never "collapsed"; the probe was reading the wrong layer.
Seeds, noise, and what counts as a win
With 599 clips, a 1 pp difference is six clips. Every recipe decision is made on at least two
seeds, and the current best on four. A change ships when the mean moves and most seeds move
with it — rr qualified with +1.7 pp mean, 3 of 4 seeds up, and half the seed spread of the
control. A single-seed +3 pp is logged as a lead, not a result.