Review 0 · Deliverable 2 of 5
One-page abstract
The whole case study on one page: the problem, the two-lane on-device system, what has been measured, and what the rest of the semester adds. Refreshed on 4 September; the numbers as they stood at review 0 are kept in the note below.
Abstract
India has an estimated 63 million deaf and hard-of-hearing citizens and only a few hundred certified Indian Sign Language interpreters, so most everyday interactions — a bank counter, a clinic, a classroom — happen without one. Automatic word-level ISL recognition could close part of that gap, but the systems the literature offers are constrained in two ways that matter in India: they recognise a small closed set of classes fixed at training time, and they run server-side, which imposes connectivity, cost, and a camera feed leaving the room. A third constraint appeared only once this project measured it: models that score well on their own corpus fall apart on a signer they have never seen.
This case study builds and evaluates Sanket, a recognition system that runs entirely on the user's own device in a web browser. Video frames are reduced immediately to a 107-point landmark skeleton (MediaPipe Holistic: 33 pose, 21 + 21 hand, a face subset); the frames are discarded, and every later stage operates on 64-frame landmark windows in Web Workers. Two lanes share that window. The dictionary lane classifies 260 INCLUDE signs with a 4.57 MB int8 model whose encoder was pretrained by latent prediction (JEPA) on a 119,064-unit unlabelled ISL pool and fine-tuned with gloss-matched dictionary clips, 3D-re-posed synthetic signers, and a recover-and-resample augmentation. The retrieval lane embeds the same window into 256 dimensions and retrieves by cosine over a 40,439-row bank covering 4,806 dictionary words (8,188 vocabulary entries), with a training-free fingerprint matcher beside it as a diagnostic witness.
Every model is judged by one pre-registered number: clip-weighted top-1 over 599 clips it has never seen — 220 from one unseen signer (RKMVU) and 379 CISLR test-role clips. On that gate the program moved from a supervised baseline of 0.3318 on the unseen signer (August) to 0.4818 (gate 0.4591) by 3 September, at less than half the model size, while in-corpus accuracy stayed near 0.80–0.83. Along the way the project measured that a frozen-probe ranking of encoders disagrees with the fine-tuned ranking, and that Phase B's last-layer JEPA target collapsed (target-std 0.98 → 0.52) while its middle layers still transferred — findings that redirected pretraining to an all-layer teacher (Phase C, in training, +4.8 pp in pilot). The remaining semester adds Phase C, re-embedding of the retrieval bank, personal enrolment, and a signer study, targeting the final report by end of November 2026, aligned with SDG 10, Reduced Inequalities.
Keywords: Indian Sign Language · open-vocabulary recognition · on-device inference · self-supervised learning · signer-independent evaluation · accessibility