Review 0 · Deliverable 1 of 5
Title identification
A title is a scope commitment. This page states the proposed title, unpacks what each phrase commits the work to — and, a month on, records how each commitment has been kept — and lists the alternates considered and why they lost.
Proposed title
Sanket: On-Device, Open-Vocabulary Indian Sign Language Recognition in the Browser
with the formal long-form variant for the report cover:
Privacy-Preserving, Open-Vocabulary Indian Sign Language Recognition on Commodity Devices using Pose-Based Metric Learning and Training-Free Exemplar Matching
Sanket (सङ्केत) is Sanskrit-derived and shared across Hindi, Marathi and Nepali: a sign, a signal, a gesture that carries meaning — a name from the languages the product serves rather than an English coinage.
The long-form variant doubles as the working title for the semester-2 research papers: a title whose every phrase is a defensible claim survives the trip from case study to paper without renaming the work midway. One phrase has grown since review 0 — the encoder is now self-supervised then fine-tuned rather than purely metric-learned — and the paper title will say so; the case-study title stands.
What each phrase commits us to
| Phrase | Commitment | Kept how (4 Sep) |
|---|---|---|
| On-device / in the browser | Every stage — pose extraction, recognition, sentence assembly — executes in the user's own tab (WebAssembly) or desktop app. No video, and no landmark stream, ever leaves the device. An architectural property, not a policy promise. | Deployed at studio.sanket.kadal.cc; the runtime page shows where each hop enforces it. |
| Open-vocabulary | Recognition is not fixed to a closed training set. The system retrieves over a dictionary-supervised vocabulary and grows it by adding references, not retraining. | At review 0: 8,165 words. Now: 4,806 words / 8,188 vocabulary entries over a 40,439-row bank, beside a 260-class dictionary lane. |
| Indian Sign Language | Word-level ISL specifically — not ASL work re-labelled. Datasets, references and evaluation are ISL throughout (INCLUDE, CISLR, ISLRTC dictionary material). | Plus RKMVU-220 as the cross-signer test, and a 119,064-unit unlabelled ISL pool for pretraining. |
| Pose-based metric learning | The learned pipeline embeds landmark windows and recognises by distance to prototypes — the metric-learning formulation that makes one-shot vocabulary growth possible. | The retrieval lane still is (ArcFace scale 30, cosine, tier A/B). The encoder beneath it is now JEPA-pretrained on unlabelled ISL and fine-tuned — the JEPA page. |
| Training-free exemplar matching | A parallel pipeline — direct match — matches live trajectories against dictionary references with an orientation-aware, dwell-weighted alignment; no model in the loop. It separates what is hard because of learning from what is hard because of data. | Runs beside the retrieval lane as the fingerprint witness in the recognition worker. |
Scope boundaries (explicitly out of Review 0 scope)
- Continuous signing / full sentence translation — the system recognises word-level signs and assembles gloss streams; full grammatical ISL-to-English translation is an application layer, with an optional, opt-in LLM polish step.
- Fingerspelling as a dedicated modality.
- Regional dialect coverage beyond what the source dictionaries encode.
Alternates considered
| Candidate title | Why it lost |
|---|---|
| "Indian Sign Language Recognition with Pose-Based Spatio-Temporal Networks" | The original shortlist phrasing. Accurate about one pipeline, silent about the two properties that make the work distinctive — open vocabulary and on-device execution. |
| "Real-Time ISL-to-Text Translation in the Browser" | Overclaims: translation implies sentence-level grammar transfer, which is scoped as an application layer, not the research core. |
| "A Shazam for Indian Sign Language" | The honest inspiration for the exemplar pipeline, and the clearest one-line explanation of it — kept as an explanatory device in the abstract, but too informal for a title. |
| "Few-Shot Sign Language Recognition for Low-Resource Vocabularies" | Generic: loses the language, the deployment target, and the product. Reads like a survey. |
Keywords
Indian Sign Language · word-level sign recognition · open-vocabulary retrieval · metric learning · self-supervised learning · JEPA · pose estimation · MediaPipe · on-device inference · WebAssembly · exemplar matching · signer-independent evaluation · accessibility · SDG 10