Skip to content

Review 1 · Deliverable 1 of 4

Literature survey

Review 0 surveyed six threads: datasets, appearance- and pose-based recognition, the one-shot regime, template matching, and on-device inference. This survey extends it with the literature the project's own measurements sent us to: self-supervised latent prediction, skeleton and sign-specific self-supervision, and the field's turn toward signer-independent evaluation.

1 · Latent-predictive self-supervision: the JEPA line

Joint-Embedding Predictive Architectures learn by predicting the latent representation of hidden content from visible context, never reconstructing pixels or coordinates. I-JEPA established the claim for images: latent targets beat pixel reconstruction for semantic features [1]. V-JEPA carried it to video with temporal span masking [2], and V-JEPA 2 scaled it to a 1.2B-parameter model over ~1M hours of internet video, reaching state-of-the-art motion understanding (Something-Something v2, Diving48) and zero-shot robot planning — evidence that latent video prediction learns transferable dynamics rather than dataset texture [3]. V-JEPA 2.1 refined the recipe with a dense predictive loss (every token carries loss, not only masked ones), deep self-supervision (the loss applied at intermediate encoder depths), and released distilled checkpoints down to an 80M ViT-B — small enough to serve as a teacher for compact student models [4].

Why this matters here: the predictor can only succeed on what is predictable across the corpus. Signer-specific style — the very thing our measurement shows the supervised model memorises — is unpredictable noise to a JEPA objective, and is structurally discarded.

2 · Self-supervision on skeletons

Sign recognition in this project is landmark-based, so the skeleton SSL literature is the directly applicable form. S-JEPA is the blueprint: given a partial skeleton sequence, predict the latents of masked joints, with motion-aware masking, an EMA target encoder, and target centering against collapse; a vanilla transformer pretrained this way set state of the art on NTU-60/120 and PKU-MMD, with its largest margins in the low-label regime — exactly our regime [5]. Skeleton2vec and the masked-skeleton line replicate the core finding: latent targets beat coordinate reconstruction for skeletal action features [6], and the IJCV benchmark of skeleton SSL confirms it systematically across methods and datasets [7].

3 · Sign-language-specific self-supervision

Between generic skeletons and full video sits sign-specific SSL. SignBERT+ pioneered hand-model-aware masked pretraining for sign [8]. SHuBERT is the strongest recent result: masked multi-stream cluster prediction — separate hand, face, and body streams located by MediaPipe — over ~1,000 hours of ASL video, transferring to state of the art across translation, isolated recognition, and fingerspelling [9]. SignMAE shows masking should follow the articulators (segmentation-driven masking of moving body parts) rather than fall on uniform patches [10]. SSL-SLR adds a caution this project takes seriously: contrastive objectives with negatives are ill-suited to sign, because different signs share sub-movements and negatives become false; its negative-free design also transfers across sign languages [11]. Uni-Sign demonstrates the value of large-scale sign pretraining in a unified encoder across tasks [12]. On the data side, YouTube-SL-25 released video IDs for 3,207 hours across 25+ sign languages — including 3,023 Indian Sign Language videos, the largest open ISL source this project has found [13]; its ISL subset contributes 8,165 units to this project's pretraining pool.

4 · The signer-independent turn

The evaluation methodology literature has converged on this project's central concern. The Multimodal Sign Language Recognition workshop at ICCV 2025 and CVPR 2026 runs an explicit signer-independent track: train on some signers, evaluate on entirely unseen signers [14]. In continuous SLR, signer-removal and consistency-constraint methods treat each signer as a domain and penalise signer identity in the features [15]. This validates the protocol choice this project made empirically: within-corpus accuracy — even group-disjoint — does not predict cross-corpus, cross-signer behaviour, and only a held-out corpus (different signers, studio, framing) measures what a deployed webcam will see. The project's two-corpus gate is that protocol made mechanical.

5 · What the threads leave unsolved

The literature gives usIt does not give usThis project's answer (4 Sep)
Latent prediction as the anti-texture objective (JEPA line)Any JEPA on sign-language landmarks, or any ISL applicationISL-JEPA: Phase B run, Phase C in training
Skeleton SSL recipes with collapse rails (S-JEPA)Skeleton SSL confronted with cross-corpus sign evaluationPhase B's collapse measured through the gate; all-layer teacher adopted
Multi-stream sign SSL at ASL scale (SHuBERT)The same for ISL, or on-device deployability of the result4.57 MB int8 encoder in the browser
A signer-independent evaluation ethos (MSLR)A byte-exact deployed-pipeline evaluation contractThe two-corpus gate
3,000+ h of open ISL video IDs (YouTube-SL-25)That corpus turned into training-ready landmark data8,165 units in the 119,064-unit pool

References

  1. [1]M. Assran, Q. Duval, I. Misra, et al. “Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture (I-JEPA)” — CVPR 2023.
  2. [2]A. Bardes, Q. Garrido, J. Ponce, et al. “Revisiting Feature Prediction for Learning Visual Representations from Video (V-JEPA)” — TMLR 2024.
  3. [3]M. Assran, A. Bardes, D. Fan, et al. “V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning” — arXiv:2506.09985, 2025.
  4. [4]L. Mur-Labadia, M. Muckley, A. Bar, et al. “V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning” — arXiv:2603.14482, 2026.Dense predictive loss + deep self-supervision; its intermediate-depth argument was reproduced here by Phase B's depth probe.
  5. [5]M. Abdelfattah, A. Alahi. “S-JEPA: A Joint Embedding Predictive Architecture for Skeletal Action Recognition” — ECCV 2024.The architectural blueprint for this project's pretraining: latent joint targets, motion-aware masking, EMA target encoder, centering.
  6. [6]R. Xu, et al. “Skeleton2vec: A Self-supervised Learning Framework with Contextualized Target Representations for Skeleton Sequences” — arXiv:2401.00921, 2024.The multi-layer, normalised teacher target that Phase C adopts.
  7. [7]J. Zhang, et al. “Self-Supervised Skeleton-Based Action Representation Learning: A Benchmark and Beyond” — IJCV 2026 (arXiv:2406.02978).
  8. [8]H. Hu, W. Zhao, W. Zhou, H. Li. “SignBERT+: Hand-Model-Aware Self-Supervised Pre-Training for Sign Language Understanding” — IEEE TPAMI 2023.
  9. [9]S. Gueuwou, X. Du, G. Shakhnarovich, K. Livescu, A. H. Liu. “SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction” — ACL 2025.Multi-stream (hands/face/body) masked prediction over ~1,000 h ASL — the sign-specific adaptation this project's stream mask follows.
  10. [10]K. Xie, Z. Cai, K. Stefanov. “SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition” — ICPR 2026 (arXiv:2605.02094).
  11. [11]A. Basso Madjoukeng, J. Fink, P. Poitier, E. B. Kenmogne, B. Frenay. “SSL-SLR: Self-Supervised Representation Learning for Sign Language Recognition” — arXiv:2509.05188, 2025 (v2 2026).Evidence that contrastive negatives are ill-posed for sign; motivates this project's negative-free objective.
  12. [12]Z. Li, et al. “Uni-Sign: Toward Unified Sign Language Understanding at Scale” — ICLR 2025 (arXiv:2501.15187).
  13. [13]G. Tanzer, B. Zhang. “YouTube-SL-25: A Large-Scale, Open-Domain Multilingual Sign Language Parallel Corpus” — ICLR 2025 (arXiv:2407.11144).Source of the open-domain ISL videos entering this project's unlabelled pool.
  14. [14]MSLR Workshop Organisers. “Multimodal Sign Language Recognition Workshop — signer-independent track” — ICCV 2025 / CVPR 2026.The community's adoption of the evaluation protocol this project reached independently.
  15. [15]R. Zuo, B. Mak. “Improving Continuous Sign Language Recognition with Consistency Constraints and Signer Removal” — ACM TOMM 2024.