Review 2 · Deliverable 4 of 4
Expected outcomes
What the final review should expect, stated so it can be checked: the targets set at Review 1 scored against today's numbers, the four next steps in the order they will be tried, and the rule every one of them must pass — the unseen signer must not get worse.
1 · By the final review (11 October 2026, tentative)
- A frozen model in production, with its accuracy-test and unseen-signer numbers measured on the deployed bytes and published beside each other.
- The report complete, from the 50% draft to all chapters, with the results and discussion written around the final model.
- A live demonstration on a webcam, in the browser, with no video leaving the device.
2 · The Review 1 targets, scored
Review 1 wrote its targets down before the work. Here they are against what was measured.
| Target set at Review 1 | Measured now | Verdict |
|---|---|---|
| Phase B target: ≥ 0.50 | accuracy test 0.5509 · unseen signer 0.5000 fp32, 0.4909 int8 (jepa-x12) | Met on the accuracy test; borderline on the unseen signer — exactly at the line in fp32, below it in int8 |
| Stretch: ≥ 0.65 on the unseen signer by the final review | 0.5000 · top-5 0.6909 | Not yet met — the right sign is among the first five 69% of the time |
| Model no larger than the deployed one | 4.57 MB int8 (was 9.9 MB supervised) | Met |
| A new model is a bundle drop, not a rewrite | jepa-v4 and jepa-x12 shipped as Lab bundles in the existing app contract | Met — jepa-x12 promoted to the default on 27 Sep |
3 · Next levers
How to read this. Four steps in the order they will be tried, each aimed at the unseen signer. Under every arrow is the same condition: a step is kept only if RKMVU does not get worse.
In the order they will be tried:
- Settle the 40 open escalations and rebuild the pack. The v4 lift (+7.0 pp over four seeds) came from exactly this kind of cleanup; the remaining cases are the ones a person must decide in the Review tab.
- A canonical 3-D body frame. Rotate every clip into the signer's own body frame using MediaPipe's world landmarks, so the camera angle stops being something the model has to learn to ignore.
- A motion-style prior from the 123k-clip pool. Learn how different signers move from the unlabelled pool, and re-perform labelled clips in those styles — synthetic signers that differ in timing and dynamics, not only in geometry.
- Per-user adaptation from Practice mode. In Practice the target word is known, so every attempt is a free, correct label for that user's signing — the one source of data about the person actually in front of the camera.
Every step is judged by the same rule as every model so far: promote only if the unseen signer does not regress, on at least four seeds.
4 · Research outcomes: the semester-2 paper seeds
| Seed | Working claim | Venues in view | Fed by |
|---|---|---|---|
| Systems / accessibility paper | Open-vocabulary ISL recognition that runs entirely on the user's device, in a browser, with a measured cost model | ASSETS · W4A · CHI Late-Breaking Work | the system, the in-app numbers |
| Method / analysis paper | A cross-corpus audit of ISL recognition, ISL-JEPA pretraining, and the reviewed dataset contract as an evaluation instrument | LREC-COLING · ACL / EMNLP Findings · ICVGIP | the accuracy test, the seed tables, the negative results |
Both are written to the posture held since Review 0: original prose, venue-checked citations, ethics and dataset licences audited before submission, and every number re-run from committed harnesses at submission time.
5 · Societal outcome — SDG 10
The outcome this project answers to is SDG 10, Reduced Inequalities: a free, private tool for learning and recognising Indian Sign Language, for India's Deaf community and for the hearing people who want to sign with them. Privacy is by construction — no video leaves the device — and because nothing is computed on a server, the tool can stay free for anyone with a browser and a webcam. (A secondary alignment is SDG 4, Quality Education, through the Learn and Practice surfaces.)