Skip to content

Review 1 · Deliverable 2 of 4

Novelty justification

Review 0 justified novelty as a proposed composition. Review 1 justified it with measurements: a deployed-pipeline evaluation contract nobody else runs, a failure decomposition it produced, inference-side fixes it already shipped, and a pretraining design whose first two pilot runs behaved exactly as pre-registered. The days since have added the strongest evidence: the pretraining path now beats the supervised model it was measured against.

N1 · The measurement: a byte-exact, deployed-pipeline cross-corpus harness

Every published ISL number this project knows of is computed in a training framework. This project instead evaluates through the exact numeric contract the deployed browser executes — the same landmark canonicalisation, the same aspect correction, the same upper-median shoulder clamp, the same int8 graph — replayed offline over a held-out corpus (220 dictionary recordings by a signer, studio, and framing the model has never seen).

The result is the finding this review was built on (at review 1):

Condition (220 real-signer clips)Top-1
In-corpus, unseen signers (group-disjoint INCLUDE)0.82
Cross-corpus, deployed contract0.33
Cross-corpus, translated +0.15 / scaled ×1.250.33 — bit-identical outputs
Cross-corpus, mirrored (left-hand-dominant signer)0.12

The decomposition is the novelty: position, height and framing are provably already solved (the anchor-and-scale normalisation cancels them exactly — the perturbed evaluations return the same bits), hand dominance is catastrophic and unhandled, and the residual gap has flat confusions — no pair repeats — meaning a distribution-wide style shift, not a few broken classes. Within-corpus validation, even signer-disjoint, is structurally unable to see any of this.

Since review 1 the harness became the program's decision rule — a second held-out corpus was added so that a recipe cannot buy cross-signer accuracy by overfitting one signer:

Figure 1The two-corpus gate that now decides every model

How to read this. Two held-out corpora on the left combine into one clip-weighted top-1 — the gate. The three outlined boxes are reported beside it and never decide. In-corpus accuracy enters only as a dotted guardrail.

  • Held-out corpora
  • The deciding number
  • Side reports
Figure 1The two-corpus gate that now decides every model

N2 · Shipped inference novelties (measured, in production)

Two consequences run in the deployed studio rather than living in a plan:

  • Always-on mirror test-time augmentation. Every window is scored in both orientations in the dictionary lane and the more confident side answers. Measured at review 1: a left-hand-dominant signer recovers from 0.12 to the model's full 0.33, at ≈ 0.5 pp cost to the majority orientation. The mirror map is exact in canonical space because the normalisation is mirror-symmetric. On the current 4-seed ensemble, mirror TTA still adds +0.7 pp (0.4574 → 0.4641) — the residual asymmetry training on both orientations has not yet removed.
  • Vocabulary-conditional posteriors. The user's vocabulary restriction renormalises the closed-set posterior to the honest conditional ("given it is one of my words") instead of masking scores — a filter disjoint from the dictionary declines rather than fabricates.

N3 · ISL-JEPA: latent self-supervision for sign landmarks, with invariance as the pretext

The literature survey shows the ingredients exist separately; the composition does not. The pretraining is, to our knowledge, the first JEPA formulated on sign-language landmarks — and its genuinely new element is the third mask:

  1. M1 · articulator-stream masking — hide a whole hand (or the body) over a span; predict its latents from the remaining streams (the SHuBERT adaptation, in latent space).
  2. M2 · motion-aware temporal spans — hide 8–16-frame spans biased toward transitions and holds; predict their latents (the V-JEPA/S-JEPA recipe).
  3. M3 · nuisance-view prediction — the context is a deliberately corrupted view (mirrored, non-uniformly tempo-warped, morphology-rescaled, tracker-noised) that must predict the clean clip's latents. The measured real-world nuisances stop being augmentation regularisers and become the pretext task itself: same sign ⇒ same latents, by construction.

Pilot evidence at review 1, pre-registered. Both pilots ran against gates written down before launch:

Model (single seed, 150 epochs)In-corpusCross-corpusMirrored
Supervised baseline0.820.330.12
A1 · JEPA, INCLUDE-only, fine-tuned0.750.230.21
A2 · JEPA + 2.5× unlabeled data, fine-tuned0.760.280.28

A1's gate passed on the mirror criterion — the dominance cliff is eliminated by the objective, not patched at test time. A2's gate ("a clearly positive slope from added unlabeled diversity") passed on every metric, including the held-out class half.

Evidence since review 1. Phase B (100 epochs, the full 123k-clip pool) missed its pre-registered 0.50 target at 0.3273 — and the miss was informative: the last-layer target had collapsed (target-std 0.98 → 0.52) while layers ≤ 4 still transferred, so the shipped encoder is cut at layer 4 (int8 0.3318, 4.57 MB — a tie with the 9.90 MB supervised model). Fine-tuning that encoder with gloss-matched dictionary clips, 3D-re-posed synthetic signers and a recover-and-resample augmentation, then ensembling four seeds and distilling, reaches 0.4818 on the unseen signer — the JEPA path is now 15 pp above the supervised model it was measured against at review 1. The Phase C objective (an all-layer teacher) adds +4.8 pp in pilot and is in training. Every step is on the results page.

N4 · The corpus machine

Serving N3 required corpus engineering that is itself non-trivial: a 119,064-unit unlabelled ISL pool in the deployed canonical space — the two Review 0 corpora joined by the re-extracted ISLRTC dictionary (4,928 units), 99,457 iSign units (streamed clip-by-clip out of a gated archive via ranged reads, so the archive is never downloaded), 8,165 hand-presence-filtered windows from YouTube-SL-25's ISL subset, and 1,898 RKMVU pool clips. Every clip passes through the same extraction contract as the browser, so pretraining, evaluation, and deployment share one numeric space — the property that makes N1's harness trustworthy in both directions.

Since review 1, the same discipline produced two more data instruments: gloss-matched infusion (~2,000 dictionary clips joining INCLUDE fine-tuning) and a 3D-aware synthetic-signer generator validated at parity 0.0005 with the real pipeline — together worth +9.1 pp of 5-seed cross-validation (0.3018 → 0.3927).