Review 1 · Deliverable 3 of 4
Architectural design for the proposed system
Three architectures, one numeric contract: the inference system as deployed, the training system that produces its next model, and the evaluation system that decides whether that model ships. The shared canonical landmark space is what binds them. Diagrams updated 4 September to the system as built; each section notes what the review-1 version said.
A · The deployed inference architecture
Everything below the camera executes in the user's browser tab. No video, frame, or landmark leaves the device; privacy is a property of the construction, not a policy.
How to read this. Left to right is time. The pose worker is the only box that sees an image. The classifier worker is the dictionary lane (260 signs); the recognition worker is the retrieval lane (4,806 words) with the training-free fingerprint matcher beside it as a witness.
- Capture & buffers
- Models & encoders
- Diagnostic / witness
- Decisions & UI
| Worker | At review 1 (30 Aug) | Now (4 Sep) |
|---|---|---|
| Pose | MediaPipe hand + pose landmarkers, 85-point canonical frame, ~24–30 FPS; frame discarded immediately | MediaPipe Holistic, 107 canonical points (33 pose, 21 + 21 hands, face subset); frame discarded in the worker |
| Recognition (open-vocabulary router) | 6.3M-parameter int8 encoder over 64-frame windows vs an 8,188-word bank; tiered cascade; DTW matcher beside it | 256-d embedding vs a 40,439-row bank (4,806 words / 8,188 entries); cosine with tier A 0.6 · floor 0.65 · shortlist 0.55 · gap 0.05; fingerprint witness |
| Classifier (dictionary lane) | 260-sign closed-set model, 9.9 MB int8, ~2 Hz, mirror TTA | JEPA encoder Cut@4 + attentive head, 4.57 MB int8; velocity, bone, normalisation and softmax inside the graph; input pos 64×107×2 |
| Expert (hand-shape) | lazily-downloaded second opinion on RGB hand crops | unchanged; structurally absent until first escalation |
| Main thread | windowed voting, rule-based assembly | LandmarkRing with capture epochs and a per-row window vote (the 3 Sep fix) |
| Property | Mechanism |
|---|---|
| Privacy | on-device inference; the only opt-in egress paths carry gloss text or consented keypoints, never frames — where each hop enforces it |
| Honesty | silence is a first-class outcome; posteriors and cosines are rendered in their own units, never conflated |
| Portability | every threshold and landmark index comes from the model bundle at runtime — a retrained model is a Lab bundle drop, not an app release |
| Latency (measured at review 1) | router p50 ≈ 10–16 ms/window; dictionary + mirror TTA ≈ 75–92 ms at 2 Hz. The int8 JEPA dictionary lane is under half the size of the model those numbers were measured on; a re-measurement is pending |
B · The training architecture (ISL-JEPA and the fine-tuning program)
At review 1 this section proposed: a corpus machine, JEPA pretraining with masks M1–M3, an optional video-teacher distillation, heads, and export. What was built follows that plan with two changes forced by evidence — the encoder is cut at layer 4, and the "teacher" that matters turned out to be a four-seed ensemble of fine-tuned heads, distilled into one student.
How to read this. Top to bottom is the order things run. Grey boxes are data; green boxes are models or training stages; outlined boxes are shipped artifacts. The two corpora on the right are never trained on: every fine-tuned seed and the distilled student are scored on them through the gate, drawn dotted.
- Corpora & synthetic data
- Pretraining, fine-tuning, distillation
- Held-out evaluation
- Shipped artifacts
| Stage | At review 1 (proposed) | As built |
|---|---|---|
| Corpus machine | 118,168 canonical units (~100 h) through the browser-parity contract | 119,064-unit pool, 123,340 with INCLUDE; plus gloss-matched infusion (~2k clips) and 8,568 synthetic-signer clips (parity 0.0005) |
| Pretraining | articulator-token transformer (~6M), EMA target, masks M1–M3, stratified sampling | as proposed; Phase B's last-layer target collapsed (target-std 0.98 → 0.52) → Phase C uses an all-layer teacher, see the JEPA page |
| Teacher distillation | V-JEPA 2.1 ViT-B video teacher (option) | not used; the teacher is a 4-seed ensemble of fine-tuned heads, distilled into one student |
| Heads | attentive probe → light fine-tune, selected by cross-corpus accuracy | Cut@4 + attentive head, recipe all+synth:full@2 + rr; probes logged but never decide |
| Export & verify | int8 ONNX through the existing export path; scripted live harness | ONNX opset 17 → int8 with preprocessing inside the graph; parity-checked; 4.57 MB; Lab bundle with model card |
C · The evaluation architecture (what decides)
At review 1 the decision metric was cross-corpus top-1 through the deployed contract on the 220 RKMVU clips, with a held-out class half, an in-corpus guardrail, and pre-registered gates. It has since become the two-corpus gate: clip-weighted top-1 over RKMVU-220 and 379 never-trained CISLR test-role clips (599 clips), reported with RKMVU top-1/top-5, the 104-clip held-out odd classes, mirrored input, mirror-TTA average and CISLR-test top-1. Its discipline:
- Held-out half. Odd-indexed dictionary classes never enter the unlabeled pool, even as raw footage — the slope measured on them cannot be transductive credit.
- Guardrail, not promoter. In-corpus accuracy (~0.80–0.83) can veto a model that forgot the closed set; it cannot promote one, because it already failed to predict the real-signer number.
- Fine-tuned gates only. Frozen probes ranked the Phase C objectives in the wrong order; no decision rests on a probe.
- Seeds. Two seeds minimum per decision, four for the shipping candidate; one clip is 0.17 pp.
- Pre-registered targets. Phase B's 0.50 was written before the run and reported as missed (0.3273); the program best is now 0.4818 against it.