Review 1 · Deliverable 2 of 4
Novelty justification
Review 0 justified novelty as a proposed composition. Review 1 justified it with measurements: a deployed-pipeline evaluation contract nobody else runs, a failure decomposition it produced, inference-side fixes it already shipped, and a pretraining design whose first two pilot runs behaved exactly as pre-registered. The days since have added the strongest evidence: the pretraining path now beats the supervised model it was measured against.
N1 · The measurement: a byte-exact, deployed-pipeline cross-corpus harness
Every published ISL number this project knows of is computed in a training framework. This project instead evaluates through the exact numeric contract the deployed browser executes — the same landmark canonicalisation, the same aspect correction, the same upper-median shoulder clamp, the same int8 graph — replayed offline over a held-out corpus (220 dictionary recordings by a signer, studio, and framing the model has never seen).
The result is the finding this review was built on (at review 1):
| Condition (220 real-signer clips) | Top-1 |
|---|---|
| In-corpus, unseen signers (group-disjoint INCLUDE) | 0.82 |
| Cross-corpus, deployed contract | 0.33 |
| Cross-corpus, translated +0.15 / scaled ×1.25 | 0.33 — bit-identical outputs |
| Cross-corpus, mirrored (left-hand-dominant signer) | 0.12 |
The decomposition is the novelty: position, height and framing are provably already solved (the anchor-and-scale normalisation cancels them exactly — the perturbed evaluations return the same bits), hand dominance is catastrophic and unhandled, and the residual gap has flat confusions — no pair repeats — meaning a distribution-wide style shift, not a few broken classes. Within-corpus validation, even signer-disjoint, is structurally unable to see any of this.
Since review 1 the harness became the program's decision rule — a second held-out corpus was added so that a recipe cannot buy cross-signer accuracy by overfitting one signer:
How to read this. Two held-out corpora on the left combine into one clip-weighted top-1 — the gate. The three outlined boxes are reported beside it and never decide. In-corpus accuracy enters only as a dotted guardrail.
- Held-out corpora
- The deciding number
- Side reports
N2 · Shipped inference novelties (measured, in production)
Two consequences run in the deployed studio rather than living in a plan:
- Always-on mirror test-time augmentation. Every window is scored in both orientations in the dictionary lane and the more confident side answers. Measured at review 1: a left-hand-dominant signer recovers from 0.12 to the model's full 0.33, at ≈ 0.5 pp cost to the majority orientation. The mirror map is exact in canonical space because the normalisation is mirror-symmetric. On the current 4-seed ensemble, mirror TTA still adds +0.7 pp (0.4574 → 0.4641) — the residual asymmetry training on both orientations has not yet removed.
- Vocabulary-conditional posteriors. The user's vocabulary restriction renormalises the closed-set posterior to the honest conditional ("given it is one of my words") instead of masking scores — a filter disjoint from the dictionary declines rather than fabricates.
N3 · ISL-JEPA: latent self-supervision for sign landmarks, with invariance as the pretext
The literature survey shows the ingredients exist separately; the composition does not. The pretraining is, to our knowledge, the first JEPA formulated on sign-language landmarks — and its genuinely new element is the third mask:
- M1 · articulator-stream masking — hide a whole hand (or the body) over a span; predict its latents from the remaining streams (the SHuBERT adaptation, in latent space).
- M2 · motion-aware temporal spans — hide 8–16-frame spans biased toward transitions and holds; predict their latents (the V-JEPA/S-JEPA recipe).
- M3 · nuisance-view prediction — the context is a deliberately corrupted view (mirrored, non-uniformly tempo-warped, morphology-rescaled, tracker-noised) that must predict the clean clip's latents. The measured real-world nuisances stop being augmentation regularisers and become the pretext task itself: same sign ⇒ same latents, by construction.
Pilot evidence at review 1, pre-registered. Both pilots ran against gates written down before launch:
| Model (single seed, 150 epochs) | In-corpus | Cross-corpus | Mirrored |
|---|---|---|---|
| Supervised baseline | 0.82 | 0.33 | 0.12 |
| A1 · JEPA, INCLUDE-only, fine-tuned | 0.75 | 0.23 | 0.21 |
| A2 · JEPA + 2.5× unlabeled data, fine-tuned | 0.76 | 0.28 | 0.28 |
A1's gate passed on the mirror criterion — the dominance cliff is eliminated by the objective, not patched at test time. A2's gate ("a clearly positive slope from added unlabeled diversity") passed on every metric, including the held-out class half.
Evidence since review 1. Phase B (100 epochs, the full 123k-clip pool) missed its pre-registered 0.50 target at 0.3273 — and the miss was informative: the last-layer target had collapsed (target-std 0.98 → 0.52) while layers ≤ 4 still transferred, so the shipped encoder is cut at layer 4 (int8 0.3318, 4.57 MB — a tie with the 9.90 MB supervised model). Fine-tuning that encoder with gloss-matched dictionary clips, 3D-re-posed synthetic signers and a recover-and-resample augmentation, then ensembling four seeds and distilling, reaches 0.4818 on the unseen signer — the JEPA path is now 15 pp above the supervised model it was measured against at review 1. The Phase C objective (an all-layer teacher) adds +4.8 pp in pilot and is in training. Every step is on the results page.
N4 · The corpus machine
Serving N3 required corpus engineering that is itself non-trivial: a 119,064-unit unlabelled ISL pool in the deployed canonical space — the two Review 0 corpora joined by the re-extracted ISLRTC dictionary (4,928 units), 99,457 iSign units (streamed clip-by-clip out of a gated archive via ranged reads, so the archive is never downloaded), 8,165 hand-presence-filtered windows from YouTube-SL-25's ISL subset, and 1,898 RKMVU pool clips. Every clip passes through the same extraction contract as the browser, so pretraining, evaluation, and deployment share one numeric space — the property that makes N1's harness trustworthy in both directions.
Since review 1, the same discipline produced two more data instruments: gloss-matched infusion (~2,000 dictionary clips joining INCLUDE fine-tuning) and a 3D-aware synthetic-signer generator validated at parity 0.0005 with the real pipeline — together worth +9.1 pp of 5-seed cross-validation (0.3018 → 0.3927).