Review 0 · Deliverable 5 of 5
Proposed system with timeline
The pipeline as it runs today, what the semester has added so far and what remains, the evaluation protocol that keeps the numbers honest, and the plan to the final report at end of November 2026. The system diagram is the current one; the review-0 version of each component is noted in the table beneath it.
System architecture
Everything below the camera runs inside the user's browser tab. There is no server in the recognition path.
How to read this. Left to right is time. The pose worker is the only box that sees an image; everything after it receives 107 landmarks per frame. The two recognition workers answer two different questions — which of 260 signs, and which of 4,806 words — from the same 64-frame window.
- Capture & buffers
- Models & encoders
- Diagnostic / witness
- Decisions & UI
Components and artifacts
| Component | At review 0 (9 Aug) | Now (4 Sep) |
|---|---|---|
| Landmarks | 85-point skeleton, 48-frame windows | 107 canonical points (33 pose, 21 + 21 hands, face subset), 64-frame LandmarkRing with capture epochs |
| Dictionary lane | calibrated 263-class head | 260-class JEPA encoder Cut@4 + attentive head, ONNX int8, 4.57 MB, preprocessing inside the graph |
| Retrieval lane | 8,165 × 256-d int8 prototype bank (~2 MB) | 256-d embedding vs a 40,439-row bank — 4,806 words / 8,188 entries; cosine; tier A 0.6 · floor 0.65 · shortlist 0.55 · gap 0.05; ArcFace scale 30 |
| Fingerprint matcher | 8,055 orientation-aware trajectory fingerprints (~7 MB) | runs beside the retrieval lane as a witness |
| Teach pack | 8,187 per-word landmark reference clips (.isl) | drives the 3D-mannequin dictionary & practice mode |
| Evaluation | frozen CISLR retrieval protocol (2,285 queries) + natural-signing regression suite | the two-corpus gate — RKMVU-220 + CISLR test-role 379 — plus the earlier harnesses |
The suite around the recogniser
The same landmark substrate powers three user-facing surfaces: the Studio (live recognition), Learn (a 3D mannequin performs any dictionary word from its .isl clip), and Practice (sign a prompted word, get scored against the reference). This is deliberate — recognition research artefacts double as accessibility features.
Evaluation protocol
Ethics, licensing and reproducibility
Because this document seeds research papers, the obligations a venue would audit are owned from the start, not retrofitted:
- Dataset terms. INCLUDE and CISLR are used strictly for research, under the terms their authors released them with; ISLRTC dictionary videos are official public reference material processed into landmark skeletons — no source video is redistributed by this project. iSign and YouTube-SL-25 enter only as landmarks in the unlabelled pool. A formal licence re-audit of every source is scheduled before any paper submission or public dataset artifact.
- Human-participant ethics. The planned signer study collects no video (recognition is on-device by construction): sessions record only recognition outcomes and structured feedback, under written informed consent with the right to withdraw, and proceed only after whatever institutional approval SRMIST norms require.
- Privacy. No camera frame or landmark stream leaves the participant's device at any point in the study or in normal use — the same property the product argues, applied to its own evaluation.
- Reproducibility. Every number in these documents runs from committed harnesses; packs and models are versioned artifacts; papers inherit numbers only by re-running the harnesses at submission time, never by quoting this document.
- Originality. All prose across this document set is written for this case study; prior work is cited, never borrowed. The set is non-archival, so building the semester-2 papers on it raises no dual-publication conflict — target-venue self-overlap policies still apply and will be checked per submission.
What the semester adds — and where each item stands
| # | Proposed at review 0 | Status (4 Sep) |
|---|---|---|
| 1 | Reference expansion — complete the ISLRTC dictionary ingestion | 4,928 ISLRTC units in the pool; ingestion continues as quota allows |
| 2 | Reference hygiene at scale — audit pipeline across all sources | Applied; extended to gloss-matched infusion (~2k audited dictionary clips into fine-tuning) |
| 3 | Personal enrolment — record-your-own references in the Studio | pending |
| 4 | Encoder retrain with quantisation and ablations | Became the JEPA program: 0.3318 → 0.4818 on the unseen signer; Phase C encoder in training |
| 5 | Decode improvements — temporal voting, calibration | Window vote per capture epoch shipped (3 Sep); mirror TTA shipped |
| 6 | Signer study — structured sessions with ISL users | pending — scheduled after Phase C lands |
Timeline — 9 August to 30 November 2026
- Aug 9
Review 0 — proposal freeze
Review 0 · done- Document set approved with the support faculty; scope and evaluation protocol locked.
- Aug 10 – Aug 30
Cross-signer measurement & the corpus machine
Review 1 · 30 Aug · done- Deployed-contract harness built; 0.82 → 0.33 cross-signer cliff measured and decomposed
- Mirror TTA and vocabulary-conditional posteriors shipped
- 119,064-unit unlabelled ISL pool assembled; ISL-JEPA designed and piloted (A1/A2)
- Sep 1 – Sep 4
Phase B, the fine-tuning program, Phase C pilot
- Phase B JEPA: gate missed at 0.3273, depth finding → Cut@4, int8 tie at 4.57 MB
- Infusion + synthetic signers + rr → 4-seed ensemble → distilled jepa-rr4: RKMVU 0.4818, gate 0.4591
- Phase C pilot: all-layer teacher +4.8 pp; Phase C proper launched on the full pool
- Sep 5 – Sep 27
Phase C, bank re-embedding, enrolment
- Score jepa_c4 on the gate; rerun the recipe screen on the Phase C encoder if it leads
- Re-embed the retrieval bank in the new space; re-score the frozen CISLR protocol
- Record-your-own-reference flow in the Studio
- Sep 28 – Oct 18
Signer study & accessibility
- Consent protocol and institutional approvals settled, then structured sessions with ISL signers
- Learn/Practice loop polish from study findings
- Oct 12 – Oct 31
Evaluation freeze
Review 2 (≈ mid-Oct) · support faculty- Final measured results across all axes (accuracy, latency, footprint)
- Feature freeze; report skeleton drafted from these docs
- Nov 1 – Nov 21
Final report & viva preparation
- Full case-study report written (these documents are its living draft)
- Demo hardening; reproducibility pass on all quoted numbers
- Nov 22 – Nov 30
Submission buffer
Final review · main faculty- Slack for review feedback and corrections; the final review/viva is conducted by the main faculty.
Risks and mitigations
| Risk | Impact | Mitigation |
|---|---|---|
| Phase C does not beat jepa-rr4 | The pending row stays pending | jepa-rr4 is already the shipping candidate; Phase C is upside, not a dependency |
| Colab idle timeouts and VM loss | Lost training time | Checkpoints to Drive every 5 epochs with auto-resume — proven at the epoch-11 kill; the bridge and daemons drive runs unattended |
| Encoder provenance | A result rests on the wrong base | Every reported number names its encoder's origin after the name collision |
| Webcam domain shift vs studio references | Live accuracy trails benchmark | The gate replays the deployed contract; mirror TTA and the nuisance-view mask target the measured shifts |
| Solo-developer timeline | Any phase can slip | Phases overlap by design; each produces a shippable increment, so a slip degrades scope, not the deliverable |