Review 0 · Deliverable 4 of 5
Novelty proposal
Five claims, each stated the same way: what the field currently does, what this work adds, the evidence at review 0, and the evidence now. None of them depends on inventing a new architecture — the novelty is in formulation, composition, and measurement. Each claim is written as a falsifiable contribution statement, because these are the contributions the semester-2 research papers will defend.
Claim 1 — Open-vocabulary ISL recognition shipped on-device
Prior practice. ISL recognition systems in the literature classify into closed sets (typically INCLUDE's classes) and run as offline experiments or server-side services. The one-shot large-vocabulary framing exists (CISLR) but only as a benchmark protocol — no deployed system.
This work. Retrieval over the dictionary vocabulary — a bank of 40,439 rows covering 4,806 words (8,188 vocabulary entries) — executing entirely in the user's browser tab via WebAssembly, beside a 260-class dictionary lane. Vocabulary grows by adding references, not by retraining.
Evidence at review 0. Live system at ~30 FPS pose / ~14 ms per window on a consumer laptop; 0.958 top-1 on the INCLUDE closed set; 0.243 top-1 / 0.411 top-20 over 8,165 words on a held-out 2,285-query protocol.
Evidence now. Deployed at studio.sanket.kadal.cc; the dictionary lane is a 4.57 MB int8 model at 0.4818 top-1 on a signer it has never seen (from 0.3318 in August).
Claim 2 — A training-free counterfactual pipeline as a scientific instrument
Prior practice. Learned-only studies cannot say why accuracy saturates: model capacity, training procedure, and reference data quality are confounded.
This work. A parallel exemplar-matching pipeline (no model in the loop) over the same references and the same live input. Because it has no trainable parameters, every failure it shares with the learned pipeline is attributable to data, and every failure it avoids is attributable to learning.
Evidence at review 0. Three fine-tuning tracks each produced ≤ 0.3 pp open-vocabulary improvement while the matcher traced the shared failures to reference-side defects.
Evidence now. The instrument's role widened: the decisive measurement of the semester was a deployed-contract replay of an unseen signer's clips through the exact bytes the browser computes, which exposed the 0.82 → 0.33 cross-signer cliff that in-corpus validation could not see. The fingerprint matcher runs in the recognition worker as a witness beside the learned lane.
Claim 3 — A sign-specific exemplar alignment
Prior practice. Classical DTW gesture matching treats all frames and all coordinates equally, breaks on static holds, and is chirality- and position-sensitive.
This work. An alignment designed from the phenomenology of signs, each property forced by a measured failure: holds are first-class (trajectories condense into pose nodes before alignment); coverage is the vote (skipping a reference's trajectory costs its duration); orientation is part of the sign (world-frame index-finger and forearm direction features); chirality is not (left/right-handed and mirrored performances match equally); idle hands are a shared state; position is a tiebreaker, not the signal.
Evidence at review 0. A committed regression harness — the natural-signing harness — with synthetic independent signers: natural one/zero/gun rank top-1 with 4–28× distance margins; an inverted one is rejected by every word.
Evidence now. The chirality principle became a training principle: the learned encoder is pretrained with a mirrored nuisance view and scored with mirror test-time augmentation (4-seed ensemble 0.4574 → 0.4641 with TTA), so hand dominance is handled in both pipelines.
Claim 4 — Reference hygiene as a first-class method
Prior practice. Dictionary supervision is taken as given; nobody audits what the reference clips actually contain.
This work. Systematic reference-quality auditing: signing-space cropping, extraction-failure detection, per-word geometry checks.
Evidence at review 0. Signing-space cropping removed approach/retreat footage from 4,132 of 8,187 dictionary clips.
Evidence now. The same discipline applies to the labelled side: gloss-matched infusion adds ~2,000 audited dictionary clips to INCLUDE fine-tuning, and 3D-re-posed synthetic signers are admitted only after a parity check (0.0005) against the real extraction pipeline. Together they lifted 5-seed cross-validation from 0.3018 to 0.3927.
Claim 5 — Privacy and access as construction properties
Prior practice. "Privacy-preserving" in deployed recognition usually means a policy promise attached to a server.
This work. There is no server in the recognition path to make promises about: frames are reduced to landmarks in the pose worker and discarded; models, bank, and fingerprint pack are static assets cached locally; the system works offline. The same property makes the full suite (recognition studio, 3D-mannequin dictionary, practice mode) deployable to any modern browser, including on low-cost hardware.
Evidence now. The runtime page tabulates what exists at each hop and what can leave — nothing but landmarks after the first worker, and nothing at all over the network.
Publication path
These claims are scoped so that, if the semester's results hold, they decompose into publishable papers rather than one unpublishable everything-paper:
| Candidate paper | Built from | Venue class |
|---|---|---|
| Systems & accessibility paper — the deployed on-device suite, its privacy-by-construction argument, and the signer study | Claims 1 & 5 + the signer-study findings | Accessibility and HCI venues (ACM ASSETS, W4A, CHI late-breaking) |
| Measurement & method paper — the deployed-contract cross-signer audit, the two-corpus gate, and ISL-JEPA with its fine-tuning program | Claims 2–4 + the Review 1 measurement + the results lineage | NLP/vision venues and workshops (LREC-COLING resource track, ACL/EMNLP findings, WACV workshops, ICVGIP) |
Three commitments follow from treating this document as the seed of those papers: every number a paper inherits from here must re-run from the committed harnesses at submission time; the dataset licence audit (see the proposed system's ethics section) completes before any submission; and because these review documents are non-archival, their text can be built upon freely — subject to each venue's originality policy, never assumed.