Review 0 · Deliverable 3 of 5
Literature review on relevant topics
Six threads of prior work meet in this case study: sign-language datasets, appearance-based and pose-based recognition, the low-resource one-shot regime, template matching, and on-device inference. Each section closes with what the thread leaves unsolved. This review stands as submitted; the Review 1 survey extends it with the self-supervision literature the project's own measurements sent us to.
1 · Word-level SLR datasets, and where ISL stands
Word-level sign language recognition matured on American Sign Language corpora: WLASL collected 2,000 glosses across 21,000+ videos [1] and MS-ASL 1,000 classes from over 200 signers in unconstrained conditions [2]. Indian Sign Language entered this scale of study much later. INCLUDE [3] is the standard ISL word-level benchmark: 4,287 studio recordings covering 263 word classes across 15 semantic categories — enough examples per class for supervised classification, but a vocabulary two orders of magnitude smaller than the language; its best reported baseline reaches 85.6% accuracy (94.5% on the 50-class subset). CISLR [4] inverted the trade-off: 7,050 videos spanning roughly 4,700 words, most with a single reference performance, explicitly framing large-vocabulary ISL as a one-shot retrieval problem; its baseline is a prototype-based one-shot learner that borrows features from resource-rich ASL. iSign [5] recently added 118K+ video–sentence pairs and a task suite for sentence-level ISL processing. Beyond academic corpora, the ISLRTC dictionary programme publishes thousands of official per-word demonstration videos — an underused supervision source this project treats as first-class reference material.
Takeaway: for ISL, the realistic regime is thousands of words × one or two references each, not hundreds of classes × dozens of examples. Methods that only work in the second regime do not serve the language. (In this project's data today: 260 INCLUDE classes with 4,284 clips against 4,806 dictionary words; iSign alone contributes 99,457 unlabelled units to pretraining.)
2 · Appearance-based recognition
The strongest closed-set results historically come from video CNNs — I3D pretrained on Kinetics [6] transferred to sign corpora set the WLASL and INCLUDE baselines [1],[3]. These models ingest raw RGB, which brings three costs this project cannot pay: they are heavy for commodity in-browser inference, they bind the model to studio appearance (clothing, lighting, background), and they require shipping the user's camera feed into the model — the exact privacy surface this work is built to remove.
3 · Pose-based recognition
Skeleton-first methods replace pixels with estimated keypoints. ST-GCN introduced graph convolutions over joint sequences [7]; SAM-SLR carried multi-modal skeleton ensembles to the top of large SLR challenges [8]; SPOTER showed a compact transformer over pose alone is competitive on word-level benchmarks [9]. Most relevant here, OpenHands [10] demonstrated pose-based pretraining across sign languages including ISL, and argued explicitly that pose is the accessibility-friendly substrate: light, appearance-invariant, and privacy-preserving. MediaPipe [11],[12] made high-quality holistic pose and hand landmarks available on-device at camera frame rates, which is what makes the browser a viable deployment target at all.
Takeaway: pose is the right representation for this problem — but pose-based work still mostly targets closed-set classification, leaving the open-vocabulary regime to retrieval formulations.
4 · The low-resource, one-shot regime
With one reference per word, classification collapses into matching: learn an embedding where performances of the same sign cluster, then recognise by nearest prototype [13]. Margin-based losses from face recognition — ArcFace [14] and its curriculum variant CurricularFace [15] — are the standard tools for making such embeddings discriminative. CISLR's own baseline applies exactly this prototype-retrieval framing to ISL [4]. Reported open-vocabulary numbers remain low in absolute terms (top-1 in the 0.1–0.3 range over thousands of classes), which the literature attributes to signer variation against single references — a data problem more than an architecture problem, a diagnosis this project's own experiments independently reproduce.
5 · Template matching and fingerprinting
Before deep learning, isolated-gesture recognition was alignment: dynamic-time-warping distance between a query trajectory and stored exemplars [16]. Shazam's audio fingerprinting [17] is the industrial proof that retrieval against a reference library can be more robust than open-ended learning when the reference and query really are performances of the same underlying object. This project revives that lineage as its training-free pipeline: dictionary demonstrations become landmark fingerprints, and live signing is matched by an orientation-aware, dwell-weighted subsequence alignment. Its role is diagnostic as much as functional — where matching succeeds and the model fails, the model is the problem; where both fail, the reference data is.
6 · On-device inference
ONNX Runtime Web executes quantised models in WebAssembly with SIMD and threads [18], and int8 quantisation keeps models small enough to ship — the current dictionary model is 4.57 MB int8, with zero measured accuracy loss against fp32. Prior SLR literature almost never treats deployment as part of the research object; here, latency, memory, and privacy are first-class evaluation axes alongside accuracy.
7 · Scoping note: recognition, not translation
Continuous sign language translation — gloss sequences to grammatical sentences — is its own field with its own benchmarks [19]. This case study deliberately stops at word-level recognition plus rule-based gloss assembly (with an optional, opt-in LLM polish), and cites SLT only to mark the boundary.
8 · Citation and originality discipline
This review — like every document in this set — is written to be the seed of the research papers this case study intends to publish in the following semester, and is held to publication standards from the start:
- All prose is original. Nothing is copied or closely paraphrased from any source; prior work is described in this project's own words and judged against this project's question.
- Every factual claim about prior work carries a numbered citation, and the dataset statistics quoted above (video counts, class counts, baseline accuracies, venues) were verified against the publishing venue's records (ACL Anthology, ACM DL, project pages) in August 2026 — not reproduced from memory or from secondary summaries.
- Quotation is avoided entirely; where a source's specific numbers are used, they are attributed inline.
- These review documents are non-archival: reusing their text and structure in a future paper submission creates no dual-publication conflict, and any such reuse will still follow the target venue's originality and self-overlap policies.
The gap this work occupies
| Prior work gives us | Prior work does not give us | Status (4 Sep) |
|---|---|---|
| Closed-set ISL classification at 263 classes (INCLUDE) | Recognition over the actual dictionary of the language | Retrieval lane over 4,806 words / 8,188 entries, deployed |
| One-shot retrieval framing for ISL (CISLR) | That framing deployed — running on end-user devices at camera rate | Deployed at studio.sanket.kadal.cc |
| Pose-based architectures and pretraining (OpenHands, SPOTER) | Pose pipelines engineered for the browser with privacy as a construction property | Four workers, landmarks only — runtime |
| Metric-learning tools (ArcFace family, prototypes) | An account of why accuracy saturates for a real signer | Measured: 0.82 → 0.33 cross-signer; now 0.4818 via self-supervision + fine-tuning |
| Classical exemplar alignment (DTW, Shazam) | An alignment designed for signs: orientation-aware, dwell-weighted, chirality-invariant, with reference hygiene | Shipped as the fingerprint witness |
References
- [1]D. Li, C. Rodriguez-Opazo, X. Yu, H. Li. “Word-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods Comparison (WLASL)” — WACV 2020.
- [2]H. R. Vaezi Joze, O. Koller. “MS-ASL: A Large-Scale Data Set and Benchmark for Understanding American Sign Language” — BMVC 2019.
- [3]A. Sridhar, R. G. Ganesan, P. Kumar, M. Khapra. “INCLUDE: A Large Scale Dataset for Indian Sign Language Recognition” — ACM Multimedia 2020.263 ISL word classes; 260 of them are the classes of this project's dictionary lane.
- [4]A. Joshi, A. Bhat, P. S. Basu, et al. “CISLR: Corpus for Indian Sign Language Recognition in a Low-Resource Setting” — EMNLP 2022.One-shot, large-vocabulary ISL; its 379 test-role clips are half of this project's two-corpus gate.
- [5]A. Joshi, et al. “iSign: A Benchmark for Indian Sign Language Processing” — ACL Findings 2024.
- [6]J. Carreira, A. Zisserman. “Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset (I3D)” — CVPR 2017.
- [7]S. Yan, Y. Xiong, D. Lin. “Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition (ST-GCN)” — AAAI 2018.
- [8]S. Jiang, et al. “Skeleton Aware Multi-modal Sign Language Recognition (SAM-SLR)” — CVPR Workshops 2021.
- [9]M. Boháček, M. Hrúz. “Sign Pose-based Transformer for Word-level Sign Language Recognition (SPOTER)” — WACV Workshops 2022.
- [10]P. Selvaraj, G. NC, P. Kumar, M. Khapra. “OpenHands: Making Sign Language Recognition Accessible with Pose-based Pretrained Models across Languages” — ACL 2022.
- [11]C. Lugaresi, et al. “MediaPipe: A Framework for Building Perception Pipelines” — arXiv:1906.08172, 2019.
- [12]F. Zhang, et al. “MediaPipe Hands: On-device Real-time Hand Tracking” — CVPR Workshop on Computer Vision for AR/VR, 2020.
- [13]J. Snell, K. Swersky, R. Zemel. “Prototypical Networks for Few-shot Learning” — NeurIPS 2017.
- [14]J. Deng, J. Guo, N. Xue, S. Zafeiriou. “ArcFace: Additive Angular Margin Loss for Deep Face Recognition” — CVPR 2019.
- [15]Y. Huang, et al. “CurricularFace: Adaptive Curriculum Learning Loss for Deep Face Recognition” — CVPR 2020.
- [16]H. Sakoe, S. Chiba. “Dynamic Programming Algorithm Optimization for Spoken Word Recognition” — IEEE Trans. ASSP, 1978.The alignment lineage the direct-match pipeline modernises.
- [17]A. Wang. “An Industrial-Strength Audio Search Algorithm (Shazam)” — ISMIR 2003.
- [18]Microsoft. “ONNX Runtime Web — WebAssembly execution provider (SIMD, threads)” — onnxruntime.ai, accessed Aug 2026.
- [19]N. C. Camgöz, S. Hadfield, O. Koller, H. Ney, R. Bowden. “Neural Sign Language Translation” — CVPR 2018.Cited to mark the recognition/translation scope boundary.