Skip to content

Review 0 · Deliverable 1 of 5

Title identification

A title is a scope commitment. This page states the proposed title, unpacks what each phrase commits the work to — and, a month on, records how each commitment has been kept — and lists the alternates considered and why they lost.

Proposed title

Sanket: On-Device, Open-Vocabulary Indian Sign Language Recognition in the Browser

with the formal long-form variant for the report cover:

Privacy-Preserving, Open-Vocabulary Indian Sign Language Recognition on Commodity Devices using Pose-Based Metric Learning and Training-Free Exemplar Matching

Sanket (सङ्केत) is Sanskrit-derived and shared across Hindi, Marathi and Nepali: a sign, a signal, a gesture that carries meaning — a name from the languages the product serves rather than an English coinage.

The long-form variant doubles as the working title for the semester-2 research papers: a title whose every phrase is a defensible claim survives the trip from case study to paper without renaming the work midway. One phrase has grown since review 0 — the encoder is now self-supervised then fine-tuned rather than purely metric-learned — and the paper title will say so; the case-study title stands.

What each phrase commits us to

PhraseCommitmentKept how (4 Sep)
On-device / in the browserEvery stage — pose extraction, recognition, sentence assembly — executes in the user's own tab (WebAssembly) or desktop app. No video, and no landmark stream, ever leaves the device. An architectural property, not a policy promise.Deployed at studio.sanket.kadal.cc; the runtime page shows where each hop enforces it.
Open-vocabularyRecognition is not fixed to a closed training set. The system retrieves over a dictionary-supervised vocabulary and grows it by adding references, not retraining.At review 0: 8,165 words. Now: 4,806 words / 8,188 vocabulary entries over a 40,439-row bank, beside a 260-class dictionary lane.
Indian Sign LanguageWord-level ISL specifically — not ASL work re-labelled. Datasets, references and evaluation are ISL throughout (INCLUDE, CISLR, ISLRTC dictionary material).Plus RKMVU-220 as the cross-signer test, and a 119,064-unit unlabelled ISL pool for pretraining.
Pose-based metric learningThe learned pipeline embeds landmark windows and recognises by distance to prototypes — the metric-learning formulation that makes one-shot vocabulary growth possible.The retrieval lane still is (ArcFace scale 30, cosine, tier A/B). The encoder beneath it is now JEPA-pretrained on unlabelled ISL and fine-tuned — the JEPA page.
Training-free exemplar matchingA parallel pipeline — direct match — matches live trajectories against dictionary references with an orientation-aware, dwell-weighted alignment; no model in the loop. It separates what is hard because of learning from what is hard because of data.Runs beside the retrieval lane as the fingerprint witness in the recognition worker.

Scope boundaries (explicitly out of Review 0 scope)

  • Continuous signing / full sentence translation — the system recognises word-level signs and assembles gloss streams; full grammatical ISL-to-English translation is an application layer, with an optional, opt-in LLM polish step.
  • Fingerspelling as a dedicated modality.
  • Regional dialect coverage beyond what the source dictionaries encode.

Alternates considered

Candidate titleWhy it lost
"Indian Sign Language Recognition with Pose-Based Spatio-Temporal Networks"The original shortlist phrasing. Accurate about one pipeline, silent about the two properties that make the work distinctive — open vocabulary and on-device execution.
"Real-Time ISL-to-Text Translation in the Browser"Overclaims: translation implies sentence-level grammar transfer, which is scoped as an application layer, not the research core.
"A Shazam for Indian Sign Language"The honest inspiration for the exemplar pipeline, and the clearest one-line explanation of it — kept as an explanatory device in the abstract, but too informal for a title.
"Few-Shot Sign Language Recognition for Low-Resource Vocabularies"Generic: loses the language, the deployment target, and the product. Reads like a survey.

Keywords

Indian Sign Language · word-level sign recognition · open-vocabulary retrieval · metric learning · self-supervised learning · JEPA · pose estimation · MediaPipe · on-device inference · WebAssembly · exemplar matching · signer-independent evaluation · accessibility · SDG 10