Reference
Glossary
Every term the case-study documents lean on — sign-language and ISL domain language, datasets, institutions, machine learning, and this project's own vocabulary — each explained in a sentence or two. Dotted-underlined terms across the docs link here; the same entries ship in the PDF handout.
Sign language & ISL
- Indian Sign Language (ISL)
- The sign language used by the Deaf community across India — a complete natural language with its own grammar and word order, not a signed encoding of Hindi or English.
- Gloss
- The written label for a single sign, conventionally in capitals (YESTERDAY, SCHOOL). A gloss names the sign; it is not a translation of the sentence the sign appears in.
- Word-level (isolated) sign language recognition
- Recognising one sign at a time from a video clip or a live window — the task this case study addresses. Contrast with continuous recognition.
- Continuous sign language recognition
- Recognising an unsegmented stream of signing, where sign boundaries must be discovered. Harder than word-level recognition and out of this case study's scope.
- Sign language translation (SLT)
- Producing grammatical spoken-language sentences from signing — a sequence-to-sequence task beyond recognition. This project stops at gloss streams plus rule-based assembly.
- Fingerspelling
- Spelling out words letter by letter using a manual alphabet, typically for names and loanwords. Scoped out of this case study as a dedicated modality.
- Signing space
- The region in front of the signer's torso and face where signs are articulated. This project uses it operationally: hands below signing space are treated as idle.
- Dominant hand
- The hand a signer leads with; left- and right-dominant signers produce mirrored performances of the same sign, which recognition must treat as identical.
- Chirality
- Handedness of a performance — whether it is produced (or captured) left- or right-handed, including mirror flips introduced by selfie cameras. The matcher is chirality-invariant because handedness is not part of a sign's identity.
- Static hold (held sign)
- A sign whose content is a sustained handshape and orientation rather than movement — most number signs, for example. Holds are first-class citizens in this project's matcher.
- Non-manual markers
- Facial expression, mouth shape, head and torso movement that carry grammatical meaning alongside the hands. Tracked as face reference points here, but not yet modelled as grammar.
Institutions & programmes
- ISLRTC
- The Indian Sign Language Research and Training Centre, a Government of India institute under the Ministry of Social Justice and Empowerment. Its official ISL dictionary videos are one of this project's reference sources.
- SRMIST
- SRM Institute of Science and Technology — the university this M.Tech (Artificial Intelligence) case study is conducted at.
- AI4Bharat
- An IIT Madras research initiative building open AI resources for Indian languages; the group behind the INCLUDE dataset and the OpenHands sign-language toolkit.
- SDG 10 — Reduced Inequalities
- The UN Sustainable Development Goal this case study aligns with: technology access should not depend on hearing, connectivity, or the availability of an interpreter.
- Review 0 / 1 / 2 / final review
- The SRM case-study review cadence: proposal (Review 0), mid-term progress (Reviews 1–2) with the support faculty, and the final review conducted by the main faculty.
Datasets & benchmarks
- INCLUDE
- The standard ISL word-level benchmark (AI4Bharat, ACM Multimedia 2020): 4,287 studio recordings over 263 word classes in 15 categories. This project's closed-set head is trained and evaluated on it.
- CISLR
- Corpus for Indian Sign Language Recognition (EMNLP 2022): 7,050 videos across roughly 4,700 words, mostly one performance per word — the dataset that frames large-vocabulary ISL as one-shot retrieval, and the source of this project's open-vocabulary protocol.
- WLASL
- Word-Level American Sign Language dataset (WACV 2020): 2,000 glosses across 21,000+ videos — the scale reference for word-level recognition in a resource-rich sign language.
- MS-ASL
- A Microsoft ASL benchmark (BMVC 2019) of 1,000 classes from over 200 signers in unconstrained recording conditions.
- iSign
- An ISL processing benchmark (ACL Findings 2024) of 118K+ video–sentence pairs with a suite of sentence-level tasks. Cited here to mark the boundary between recognition and translation.
- ISLRTC dictionary
- The official public ISL dictionary programme: thousands of per-word demonstration videos by certified signers. This project processes them into landmark reference clips — the video itself is never redistributed.
- RKMVU validation set (the unseen signer)
- 220 clips of one signer who appears nowhere in any training data — the project's stranger test. Its top-1 is reported beside every accuracy-test number because it is the only part of the test no training clip resembles.
Machine learning
- Pose estimation / landmarks
- Detecting body and hand keypoints (landmarks) in each video frame, reducing appearance to a skeleton. This project's entire pipeline operates on an 85-point landmark skeleton, never on pixels.
- Embedding
- A learned fixed-size vector representation (here, 256 dimensions per landmark window) in which distance means dissimilarity — the substrate that makes recognition-by-retrieval possible.
- Metric learning
- Training an embedding so that performances of the same sign land close together and different signs land far apart — recognition then reduces to nearest-neighbour lookup instead of a fixed classifier.
- ArcFace / CurricularFace (margin losses)
- Training losses from face recognition that enforce an angular margin between classes, producing highly discriminative embeddings; CurricularFace adds a curriculum that emphasises hard examples over time.
- Prototype
- A single representative embedding standing for a class — here, the embedding of a word's dictionary reference. Recognising by distance to prototypes is what lets one example per word suffice.
- One-shot learning
- The regime where a class must be recognised from a single labelled example. Large-vocabulary ISL is effectively one-shot: most words have exactly one dictionary performance.
- The 260-sign set
- The fixed list of signs a model is trained to name outright. Answers are probabilities and they are the most reliable the system produces — but the list cannot grow without retraining. Called closed-set recognition in the literature.
- The full vocabulary
- All 8,188 words the system can reach, by comparing a sign against stored examples rather than naming it from a fixed list. Adding a word needs an example, not a retraining run. Scores here are similarity, not probability. Called open-vocabulary recognition in the literature.
- Retrieval
- Answering "which word was signed?" by ranking stored references by similarity to the query, instead of running a fixed classifier — the formulation that scales to dictionary-sized vocabularies.
- Calibration
- Making a model's confidence scores mean what they say — a calibrated 0.9 is right about nine times in ten — so that emission thresholds behave predictably.
- Quantisation (int8)
- Storing weights and reference vectors as 8-bit integers instead of 32-bit floats, shrinking artifacts roughly fourfold so an 8,165-word bank ships to a browser.
- Fine-tuning
- Continuing training of an existing model on new data or a new objective. Three fine-tuning tracks in this project produced ≤0.3pp — the measurement behind the data-not-model diagnosis.
- Ablation
- Removing or swapping one component at a time to measure what each part actually contributes — the experimental discipline the semester's retrain phase commits to.
- Top-k accuracy / enrolled top-1
- A prediction counts if the right word appears in the model's k best guesses. "Enrolled" means every one of the 8,165 vocabulary words is a candidate — nothing is excluded to flatter the number.
- Frozen protocol / held-out test
- An evaluation set and procedure fixed in advance and never trained or tuned on. This project treats its 2,285-query CISLR test as sacred for exactly this reason.
- Domain shift
- The gap between the data a system was built on (studio dictionary recordings) and the data it meets in use (webcams, home lighting, untrained signers) — a named risk in the proposal.
- JEPA / ISL-JEPA (self-supervised pretraining)
- Joint-embedding predictive architecture: hide parts of a clip and train the encoder to predict the hidden parts' latent representation rather than their raw coordinates. ISL-JEPA applies this to landmark tokens over a large unlabelled pool before any labels are used.
- EMA teacher (target encoder)
- A copy of the encoder whose weights trail the trained ones as an exponential moving average; it supplies the prediction targets in JEPA pretraining, which keeps them stable while the student learns.
- Knowledge distillation (ensemble → student)
- Training one small model to match the averaged, temperature-softened answers of several larger or more numerous ones, alongside the true labels — how twelve fine-tunes become one 4.57 MB browser model.
- Synthetic signers
- New training clips made from real ones by lifting the landmarks to 3-D, turning the body and changing limb proportions, then projecting back — two variants per clip. They help only when stacked on real extra clips.
- Recover-and-resample (rr)
- A training augmentation: with probability 0.5, trim up to a quarter of the clip from either end and stretch the rest back to 64 frames, so the model stops depending on where in the window a sign starts.
System & project terms
- Accuracy test ("the gate")
- The one number that decides whether a model ships: top-1 over 599 held-out clips — 220 from an unseen signer (RKMVU) and 379 CISLR clips never trained on. One clip is 0.17 pp. Written "the gate" in earlier reports.
- Mirror pass (mirror test-time augmentation)
- Scoring every window twice — as captured and flipped left to right — and keeping the more confident answer, so a left-hand-dominant signer is read as well as a right-handed one.
- Excalidraw (hand-drawn diagrams)
- An open-source whiteboard whose drawings look hand-sketched. The Review 2 diagrams are Excalidraw scenes written once as code and shown in the deck, on these pages and in the report.
- On-device inference
- Running the entire recognition pipeline on the user's own hardware, inside the browser tab — no server round-trip, no video upload, and it works offline. In this project a construction property, not a policy.
- WebAssembly (WASM)
- A portable binary format browsers execute at near-native speed; with SIMD and threads it is what makes running a neural network inside a tab practical.
- ONNX Runtime Web
- Microsoft's inference engine for the ONNX model format compiled to WebAssembly — the runtime this project's quantised encoder executes on, at ~14 ms per window.
- MediaPipe
- Google's on-device perception framework; its holistic pose and hand models supply the per-frame landmarks this project consumes at camera rate.
- Canonical space
- The normalised coordinate frame all landmarks are mapped into — origin at the mid-shoulder point, scale set by shoulder width — so camera distance, framing, and signer height stop mattering.
- Recognition window
- A 64-frame slice of the landmark stream — the unit the encoder scores. Overlapping windows are voted over so a sign is recognised once, not once per window.
- Prototype bank
- The matrix of 256-dimensional embeddings that full-vocabulary retrieval scans on every window: 8,497 rows covering 8,165 words, with one row per variant when a word is signed more than one way.
- The two lists
- The AI model answers from two places: the 260-sign set, which it names outright with a probability, and the full vocabulary, which it matches by similarity. A word is only said when one of them clears its confidence bar. Written tier A / tier B in the code and the earlier reports.
- Shape match (no AI)
- Recognition with no neural network at all: the movement of your hands and arms is compared frame by frame against stored example clips, the way Shazam matches a song. It runs beside the AI model as a second opinion, and as the instrument that separates what the model learned from what the data already gave away.
- Fingerprint (trajectory)
- A compact orientation-aware feature track distilled from a reference clip or live window — wrist paths, forearm and finger directions, handshape ratios — the unit the direct-match aligner compares.
- Dynamic time warping (DTW)
- A classic alignment algorithm that matches two sequences that unfold at different speeds. This project's aligner descends from it, with sign-specific changes: condensed poses, dwell weighting, coverage penalties.
- Dwell weighting
- Weighting each pose by how long it is held, so a sustained handshape carries its full duration of evidence while one-frame transitions barely count — how holds and movements coexist in one matcher.
- Resting detection
- A cheap check run before anything else: are the hands raised into signing position, and is anything actually moving? If not the window is skipped, which is what stops idle video from producing words. Shown in the app as "resting".
- Audio fingerprinting (Shazam)
- Identifying a recording by matching a compact signature against a reference library rather than by open-ended learning — the industrial precedent the direct-match pipeline is modelled on.
- Gloss stream / sentence assembly
- The sequence of recognised glosses, and the rule layer that reorders them into a readable sentence. An optional LLM polish exists strictly as user opt-in.
- Temporal voting
- Aggregating predictions across overlapping windows before emitting a word, trading a few hundred milliseconds of latency for large stability gains.
- Teach pack / .isl format
- The project's 8,187 per-word landmark reference clips in a compact binary format (.isl). It drives the 3D-mannequin dictionary, the practice mode, and the direct-match fingerprints alike.
- 3D mannequin (Learn / Practice)
- A rigged 3D avatar that performs any dictionary word by replaying its landmark clip — the dictionary you can watch from any angle, and the reference the practice mode scores you against.
- Natural-signing harness
- A committed regression suite that matches synthetic independently-constructed signers — not replayed references — against the pack, and fails the build if naturally-performed signs stop ranking first.
Research & publishing
- Non-archival
- A document or venue whose publications do not count as prior publication — these review documents are non-archival, so future papers can build on their text without dual-publication conflict.
- Self-overlap / dual publication
- Reusing one's own published text or results across submissions. Venues restrict it; this project's rule is to check the target venue's policy at every submission.
- Reproducibility
- The property that every reported number can be regenerated by running the committed evaluation harnesses — the standard this document set holds itself to, and papers inherit.
- Informed consent
- Written agreement from study participants who understand what is collected and why, with the right to withdraw. The signer study proceeds only under it — and collects no video at all.
- Dataset licence audit
- Re-reading every source dataset's release terms and documenting what use they permit — scheduled in this project before any paper submission or public artifact.
- Venues referenced
- CVPR, WACV, BMVC (computer vision); ACL, EMNLP, LREC-COLING (language); NeurIPS, AAAI (machine learning); ACM Multimedia, ISMIR (media); ASSETS, W4A, CHI (accessibility and HCI); ICVGIP (Indian vision & graphics). The publication path targets the accessibility and language groups.