Skip to content

Review 0 · Deliverable 5 of 5

Proposed system with timeline

The pipeline as it runs today, what the semester has added so far and what remains, the evaluation protocol that keeps the numbers honest, and the plan to the final report at end of November 2026. The system diagram is the current one; the review-0 version of each component is noted in the table beneath it.

System architecture

Everything below the camera runs inside the user's browser tab. There is no server in the recognition path.

Figure 1Recognition pipeline — all stages on-device (as deployed, 4 September)

How to read this. Left to right is time. The pose worker is the only box that sees an image; everything after it receives 107 landmarks per frame. The two recognition workers answer two different questions — which of 260 signs, and which of 4,806 words — from the same 64-frame window.

  • Capture & buffers
  • Models & encoders
  • Diagnostic / witness
  • Decisions & UI
Figure 1Recognition pipeline — all stages on-device (as deployed, 4 September)

Components and artifacts

ComponentAt review 0 (9 Aug)Now (4 Sep)
Landmarks85-point skeleton, 48-frame windows107 canonical points (33 pose, 21 + 21 hands, face subset), 64-frame LandmarkRing with capture epochs
Dictionary lanecalibrated 263-class head260-class JEPA encoder Cut@4 + attentive head, ONNX int8, 4.57 MB, preprocessing inside the graph
Retrieval lane8,165 × 256-d int8 prototype bank (~2 MB)256-d embedding vs a 40,439-row bank — 4,806 words / 8,188 entries; cosine; tier A 0.6 · floor 0.65 · shortlist 0.55 · gap 0.05; ArcFace scale 30
Fingerprint matcher8,055 orientation-aware trajectory fingerprints (~7 MB)runs beside the retrieval lane as a witness
Teach pack8,187 per-word landmark reference clips (.isl)drives the 3D-mannequin dictionary & practice mode
Evaluationfrozen CISLR retrieval protocol (2,285 queries) + natural-signing regression suitethe two-corpus gate — RKMVU-220 + CISLR test-role 379 — plus the earlier harnesses

The suite around the recogniser

The same landmark substrate powers three user-facing surfaces: the Studio (live recognition), Learn (a 3D mannequin performs any dictionary word from its .isl clip), and Practice (sign a prompted word, get scored against the reference). This is deliberate — recognition research artefacts double as accessibility features.

Evaluation protocol

Ethics, licensing and reproducibility

Because this document seeds research papers, the obligations a venue would audit are owned from the start, not retrofitted:

  • Dataset terms. INCLUDE and CISLR are used strictly for research, under the terms their authors released them with; ISLRTC dictionary videos are official public reference material processed into landmark skeletons — no source video is redistributed by this project. iSign and YouTube-SL-25 enter only as landmarks in the unlabelled pool. A formal licence re-audit of every source is scheduled before any paper submission or public dataset artifact.
  • Human-participant ethics. The planned signer study collects no video (recognition is on-device by construction): sessions record only recognition outcomes and structured feedback, under written informed consent with the right to withdraw, and proceed only after whatever institutional approval SRMIST norms require.
  • Privacy. No camera frame or landmark stream leaves the participant's device at any point in the study or in normal use — the same property the product argues, applied to its own evaluation.
  • Reproducibility. Every number in these documents runs from committed harnesses; packs and models are versioned artifacts; papers inherit numbers only by re-running the harnesses at submission time, never by quoting this document.
  • Originality. All prose across this document set is written for this case study; prior work is cited, never borrowed. The set is non-archival, so building the semester-2 papers on it raises no dual-publication conflict — target-venue self-overlap policies still apply and will be checked per submission.

What the semester adds — and where each item stands

#Proposed at review 0Status (4 Sep)
1Reference expansion — complete the ISLRTC dictionary ingestion4,928 ISLRTC units in the pool; ingestion continues as quota allows
2Reference hygiene at scale — audit pipeline across all sourcesApplied; extended to gloss-matched infusion (~2k audited dictionary clips into fine-tuning)
3Personal enrolment — record-your-own references in the Studiopending
4Encoder retrain with quantisation and ablationsBecame the JEPA program: 0.3318 → 0.4818 on the unseen signer; Phase C encoder in training
5Decode improvements — temporal voting, calibrationWindow vote per capture epoch shipped (3 Sep); mirror TTA shipped
6Signer study — structured sessions with ISL userspending — scheduled after Phase C lands

Timeline — 9 August to 30 November 2026

AugustSeptemberOctoberNovember
  1. Aug 9

    Review 0 — proposal freeze

    Review 0 · done
    • Document set approved with the support faculty; scope and evaluation protocol locked.
  2. Aug 10 – Aug 30

    Cross-signer measurement & the corpus machine

    Review 1 · 30 Aug · done
    • Deployed-contract harness built; 0.82 → 0.33 cross-signer cliff measured and decomposed
    • Mirror TTA and vocabulary-conditional posteriors shipped
    • 119,064-unit unlabelled ISL pool assembled; ISL-JEPA designed and piloted (A1/A2)
  3. Sep 1 – Sep 4

    Phase B, the fine-tuning program, Phase C pilot

    • Phase B JEPA: gate missed at 0.3273, depth finding → Cut@4, int8 tie at 4.57 MB
    • Infusion + synthetic signers + rr → 4-seed ensemble → distilled jepa-rr4: RKMVU 0.4818, gate 0.4591
    • Phase C pilot: all-layer teacher +4.8 pp; Phase C proper launched on the full pool
  4. Sep 5 – Sep 27

    Phase C, bank re-embedding, enrolment

    • Score jepa_c4 on the gate; rerun the recipe screen on the Phase C encoder if it leads
    • Re-embed the retrieval bank in the new space; re-score the frozen CISLR protocol
    • Record-your-own-reference flow in the Studio
  5. Sep 28 – Oct 18

    Signer study & accessibility

    • Consent protocol and institutional approvals settled, then structured sessions with ISL signers
    • Learn/Practice loop polish from study findings
  6. Oct 12 – Oct 31

    Evaluation freeze

    Review 2 (≈ mid-Oct) · support faculty
    • Final measured results across all axes (accuracy, latency, footprint)
    • Feature freeze; report skeleton drafted from these docs
  7. Nov 1 – Nov 21

    Final report & viva preparation

    • Full case-study report written (these documents are its living draft)
    • Demo hardening; reproducibility pass on all quoted numbers
  8. Nov 22 – Nov 30

    Submission buffer

    Final review · main faculty
    • Slack for review feedback and corrections; the final review/viva is conducted by the main faculty.

Risks and mitigations

RiskImpactMitigation
Phase C does not beat jepa-rr4The pending row stays pendingjepa-rr4 is already the shipping candidate; Phase C is upside, not a dependency
Colab idle timeouts and VM lossLost training timeCheckpoints to Drive every 5 epochs with auto-resume — proven at the epoch-11 kill; the bridge and daemons drive runs unattended
Encoder provenanceA result rests on the wrong baseEvery reported number names its encoder's origin after the name collision
Webcam domain shift vs studio referencesLive accuracy trails benchmarkThe gate replays the deployed contract; mirror TTA and the nuisance-view mask target the measured shifts
Solo-developer timelineAny phase can slipPhases overlap by design; each produces a shippable increment, so a slip degrades scope, not the deliverable