Case Study Final Review · 11 October 2026 · 2:00–6:00 PM · SRM KTR campus, in person
Final review — the complete case study
The case study concludes with the overall design, the justification for each choice, the results and their discussion, and the report. Since Review 2 the recogniser has learned from signers of another language: a supervised American Sign Language stage, measured over four seeds per recipe and distilled into one browser model, took the unseen signer past 0.50 in the graph the browser runs. The complete report is below.
- Candidate
- Mithunish Prabhakaran
- Registration No.
- RA2512049015065
- Programme
- M.Tech Artificial Intelligence, Department of Computational Intelligence, SRMIST (2025–2027)
- Course
- 21CSC601T – Case Studies · School of Computing
- Guide
- Dr. S. Saravanan, Professor, Department of Computational Intelligence
- Review
- 11 Oct 2026 · 2:00–6:00 PM · SRM KTR campus, in person
- SDG alignment
- SDG 10 — Reduced Inequalities (secondary: SDG 4 — Quality Education)
- Deployed model
- jepa-asl-b3 · 5.23 MB int8 · since 2 Oct 2026
- Presentation
- deck.kadal.cc/sanket-final
- Contact
- pmithunish@gmail.com
Deliverables
| # | Deliverable | Where it is answered |
|---|---|---|
| 1 | Overall design | Report Chapter 3 — the approach, data, requirements and architecture; Architecture |
| 2 | Justification for the proposed system | Report Section 3.5 — each choice against its alternative, with the measurement that decided it |
| 3 | Results and discussion | Report Chapter 5 — lineage, transfer from ASL, the deployed model, sign by sign, negative results; Results |
| 4 | Report submission | The PDF above: 6 chapters, 26 figures, 26 tables, 7 algorithms, 49 references; appendices with code, a manuscript, screenshots, reproducibility and MoM |
How to read this. Read left to right: landmarks only; a self-supervised encoder learned without labels; transfer from ASL and targeted data; many models distilled into one; and a stranger test before anything ships. The bar underneath is the constraint that holds throughout — the same canonical landmarks in the browser, in training and in evaluation, and no video leaves the device.
From Review 1 to the final review
| Thread | Review 1 (30 Aug) | Review 2 (27 Sep) | Final (11 Oct) |
|---|---|---|---|
| Unseen signer (RKMVU-220) | 0.3318 | 0.5000 (jepa-x12) | 0.5136 with the app's mirror rule · int8 0.5091 |
| Accuracy test (599 clips) | 0.1469 | 0.5509 | 0.5526 int8 · 0.5559 in the browser |
| Calibration error (ECE) | — | 0.056 | 0.039 (T = 0.804, cross-fitted) |
| Model | supervised, 9.9 MB | 12 teachers → one student, 4.57 MB | 12 ASL-pretrained teachers → one student, 5.23 MB |
| Continuous signing | sliding windows | sliding windows | sign by sign: word accuracy 0.230 → 0.498 |
How to read this. Left: Google's ISLR corpus — 94,477 ASL clips from 21 Deaf signers — streamed by byte range on a CPU task and canonicalised with the same rules as the Indian corpora. Right: the transfer recipes — a supervised ASL stage before the ISL fine-tune, an ASL auxiliary loss, the phonological head — and their distillation into the deployed model. Bottom: what was measured, four seeds per recipe.
How to read this. Each browser model on the review timeline with its accuracy test (gate) and unseen-signer (RKMVU) numbers, in order but not to scale. The pink marker is the deployed jepa-asl-b3.
How to read this. The webcam stream is split where both hands rest below the hip line; each span is resampled to 64 frames and classified once. Offline, on joined unseen-signer clips, word accuracy rose from 0.230 with sliding windows to 0.498, against a ceiling of 0.509 with the true boundaries.
Future scope
How to read this. Four next steps, from a signer study to continuous signing. Each is kept only if it beats the deployed model on the unseen signer.