Skip to content

Architecture · 1 of 5

The runtime: everything happens in the tab

A camera frame exists in exactly one place — the pose worker that turns it into 107 landmarks and throws it away. Everything after that point is landmark arithmetic in Web Workers, and nothing leaves the device. This page follows one frame from the camera to the dictionary card.

Figure 1The in-browser pipeline: one place pixels exist, two answer lists, no network

How to read this. Left to right is time. Solid arrows carry data every frame or every window; dotted arrows are optional or diagnostic. The pose worker is the only box that ever sees an image — everything to its right receives 107 landmarks per frame and nothing else.

  • Capture & buffers
  • Models & encoders
  • Diagnostic / witness
  • Decisions & UI
Figure 1The in-browser pipeline: one place pixels exist, two answer lists, no network

The path of a frame

  1. Capture — the studio reads the camera (or a footage file) at roughly 30 fps on the main thread and hands each frame to the pose worker.
  2. Landmarks — the pose worker runs MediaPipe Holistic and keeps 107 canonical points: 33 pose, 21 + 21 hand, and a face subset. The frame is discarded the moment the landmarks exist. Landmarks are re-anchored to the mid-shoulder origin and scaled by shoulder width (canonical space), which is why camera distance, framing and signer height cancel out exactly.
  3. The ring — LandmarkRing on the main thread keeps a rolling 64-frame window. Every capture (a demo row, a practice attempt, a live take) starts a new capture epoch, and a window is only ever voted inside the epoch it was recorded in.
  4. Two lists, two workers — each new window is posted to the classifier worker (the 260-sign set) and to the recognition worker (the full vocabulary plus the training-free fingerprint witness).
  5. Vote and decode — per-window probabilities are voted per epoch for the dictionary card; ranked retrieval candidates flow through the decoder into the gloss stream.

The two answer lists

260-sign setFull vocabulary
Question it answersWhich of the 260 INCLUDE signs is this? (260-sign set)Which of the 8,165 reachable words is this? (full vocabulary, 8,188 vocabulary entries)
Workerclassifier workerrecognition worker
ModelONNX int8 · JEPA encoder Cut@4 + attentive head · 4.57 MB256-d embedding encoder · bank of 8,497 rows
Inputpos only — 64 × 107 × 2. Velocity, bone vectors, normalisation and the softmax are all inside the graph, so the JavaScript side does no preprocessingthe same 64-frame window
Decisionsoftmax over 260 → window vote per capture epochcosine similarity vs the bank, then the confidence bars: 0.6 to answer from the 260-sign set · 0.65 similarity floor · 0.55 to offer a choice · 0.05 minimum gap; ArcFace scale 30
Outputprobs (260) → dictionary card, demo tableranked words / shortlist / silence → gloss stream

Worker messages

The engine is four workers and a main thread; the contract between them is a small set of typed messages. This is the whole surface through which data moves.

WorkerInbound (main → worker)Outbound (worker → main)
Poseinit · frame · groups · resize · capture-start · crops-enabled · crops-request · stopready · stats · landmarks (one canonical frame) · crops (dev only) · error
Classifier (260-sign set)init · capture-start · window (pos 64×107×2) · classify-batchclassifier-ready · classifier-window (probs 260) · classifier-batch · classifier-benchmark · classifier-error
Recognition (full vocabulary)init · capture-start · landmark frame · decoder · reset · benchmark · fingerprint · vocab · fusion · stopready · window (ranked candidates) · gloss (decoded event) · segment · benchmark · error
Expert (hand-shape second opinion, lazy)init · score · stoploaded · scores

capture-start is broadcast to all three recognition-path workers at once: it advances the capture epoch so that a window recorded before the message can never be voted after it.

Where privacy is enforced

Privacy here is not a policy attached to a server; it is the shape of the data at each hop.

StageWhat existsWhat can leave the device
Camera → pose workerRGB frames, in memory, one at a timenothing — the frame never reaches a network API
Pose worker → ring107 landmarks per frame (the frame is already gone)nothing
Ring → workers64 × 107 landmark windowsnothing
Models, bank, fingerprint packstatic assets cached by the browserfetched once, like any file; no upload path exists
Optional egressgloss text, or consented keypoints in a studyonly on an explicit user action, never frames

The capture-epoch bug (fixed 3 September)

Deployment

The studio is served at studio.sanket.kadal.cc; the landing site at sanket.kadal.cc. The dictionary model in the shipped bundle is the 4.57 MB int8 export; which encoder that is, and how it was chosen, is the subject of the training and evaluation pages.