Architecture · 1 of 5
The runtime: everything happens in the tab
A camera frame exists in exactly one place — the pose worker that turns it into 107 landmarks and throws it away. Everything after that point is landmark arithmetic in Web Workers, and nothing leaves the device. This page follows one frame from the camera to the dictionary card.
How to read this. Left to right is time. Solid arrows carry data every frame or every window; dotted arrows are optional or diagnostic. The pose worker is the only box that ever sees an image — everything to its right receives 107 landmarks per frame and nothing else.
- Capture & buffers
- Models & encoders
- Diagnostic / witness
- Decisions & UI
The path of a frame
- Capture — the studio reads the camera (or a footage file) at roughly 30 fps on the main thread and hands each frame to the pose worker.
- Landmarks — the pose worker runs MediaPipe Holistic and keeps 107 canonical points: 33 pose, 21 + 21 hand, and a face subset. The frame is discarded the moment the landmarks exist. Landmarks are re-anchored to the mid-shoulder origin and scaled by shoulder width (canonical space), which is why camera distance, framing and signer height cancel out exactly.
- The ring —
LandmarkRingon the main thread keeps a rolling 64-frame window. Every capture (a demo row, a practice attempt, a live take) starts a new capture epoch, and a window is only ever voted inside the epoch it was recorded in. - Two lists, two workers — each new window is posted to the classifier worker (the 260-sign set) and to the recognition worker (the full vocabulary plus the training-free fingerprint witness).
- Vote and decode — per-window probabilities are voted per epoch for the dictionary card; ranked retrieval candidates flow through the decoder into the gloss stream.
The two answer lists
| 260-sign set | Full vocabulary | |
|---|---|---|
| Question it answers | Which of the 260 INCLUDE signs is this? (260-sign set) | Which of the 8,165 reachable words is this? (full vocabulary, 8,188 vocabulary entries) |
| Worker | classifier worker | recognition worker |
| Model | ONNX int8 · JEPA encoder Cut@4 + attentive head · 4.57 MB | 256-d embedding encoder · bank of 8,497 rows |
| Input | pos only — 64 × 107 × 2. Velocity, bone vectors, normalisation and the softmax are all inside the graph, so the JavaScript side does no preprocessing | the same 64-frame window |
| Decision | softmax over 260 → window vote per capture epoch | cosine similarity vs the bank, then the confidence bars: 0.6 to answer from the 260-sign set · 0.65 similarity floor · 0.55 to offer a choice · 0.05 minimum gap; ArcFace scale 30 |
| Output | probs (260) → dictionary card, demo table | ranked words / shortlist / silence → gloss stream |
Worker messages
The engine is four workers and a main thread; the contract between them is a small set of typed messages. This is the whole surface through which data moves.
| Worker | Inbound (main → worker) | Outbound (worker → main) |
|---|---|---|
| Pose | init · frame · groups · resize · capture-start · crops-enabled · crops-request · stop | ready · stats · landmarks (one canonical frame) · crops (dev only) · error |
| Classifier (260-sign set) | init · capture-start · window (pos 64×107×2) · classify-batch | classifier-ready · classifier-window (probs 260) · classifier-batch · classifier-benchmark · classifier-error |
| Recognition (full vocabulary) | init · capture-start · landmark frame · decoder · reset · benchmark · fingerprint · vocab · fusion · stop | ready · window (ranked candidates) · gloss (decoded event) · segment · benchmark · error |
| Expert (hand-shape second opinion, lazy) | init · score · stop | loaded · scores |
capture-start is broadcast to all three recognition-path workers at once: it advances the
capture epoch so that a window recorded before the message can never be voted after it.
Where privacy is enforced
Privacy here is not a policy attached to a server; it is the shape of the data at each hop.
| Stage | What exists | What can leave the device |
|---|---|---|
| Camera → pose worker | RGB frames, in memory, one at a time | nothing — the frame never reaches a network API |
| Pose worker → ring | 107 landmarks per frame (the frame is already gone) | nothing |
| Ring → workers | 64 × 107 landmark windows | nothing |
| Models, bank, fingerprint pack | static assets cached by the browser | fetched once, like any file; no upload path exists |
| Optional egress | gloss text, or consented keypoints in a study | only on an explicit user action, never frames |
The capture-epoch bug (fixed 3 September)
Deployment
The studio is served at studio.sanket.kadal.cc; the landing site at sanket.kadal.cc. The dictionary model in the shipped bundle is the 4.57 MB int8 export; which encoder that is, and how it was chosen, is the subject of the training and evaluation pages.