Skip to content

On-device by construction, not by policy

Indian Sign Language recognition that never uploads your video

Sanket watches your camera and reads Indian Sign Language into text. Every step — pose tracking, recognition, sentence assembly — happens inside your own browser or desktop app. There is no upload endpoint to disable, because there is no upload endpoint.

Video bytes uploaded
0
Accounts required
None
Works offline
Yes
Price
Free

How it works

Three stages, each running where you are. Follow a single frame through them and you can see for yourself where the video ends up: nowhere.

  1. 01 · Pose

    Your camera becomes a skeleton

    A worker in your own tab reads each frame and extracts 107 landmark points — both hands, the arms and shoulders, and a face reference — into a canonical signing-space layout that is the same on a phone camera and on a laptop.

  2. 02 · Model

    A model on your device reads the motion

    Windows of 64 skeleton frames go to a 4.6 MB self-supervised encoder running on WebAssembly. It answers from two lists at once: a 260-sign set it names outright with a probability, and the full vocabulary of 8,165 words it matches against stored examples — 12 ms per window, on the device.

  3. 03 · Sentence

    Glosses become a sentence

    Recognition emits glosses — dictionary words in sign order, like YESTERDAY I SCHOOL GO. A rule layer reorders them into English, handles negation and questions, and punctuates. That layer is local too. Only if you explicitly opt in does a language model polish the wording, and only the gloss text is sent.

camera → landmarks (worker) → 48-frame window → classifier (worker) → gloss stream → sentence

The video never leaves

Sign language is spoken with your face and your body. A recognition service that uploads video is asking you to hand a stranger a recording of yourself saying something private — at a clinic, to your bank, to your family. Sanket was built the other way round: the model comes to your device, and the recording stays there.

This is an architectural property, not a promise in a policy document. There is no code path that transmits a frame or a landmark. The one thing that can ever be sent — if, and only if, you switch on the optional polish tier — is the recognised gloss text, and the app shows you the exact request body before you enable it.

Read exactly what is and is not sent
  • No upload endpoint

    Frames and landmarks have nowhere to go. The workers that handle them have no network access to a backend at all.

  • Works with the network off

    Once the model is cached, pull the plug and it keeps recognising. That is the honest test of an on-device claim.

  • No account for the core product

    Sign-in exists only to sync your practice progress if you want that. Recognition never asks who you are.

UN Sustainable Development Goal 10 — Reduced Inequalities

Why this exists

India has roughly 63 million deaf and hard-of-hearing people and only a few hundred certified Indian Sign Language interpreters. Access to an interpreter should not decide whether someone can be understood at a bank counter, a clinic, or a classroom.

Software cannot train interpreters. What it can do is make the ordinary exchanges — asking a question, filling a form, being understood by someone who does not sign — possible without one. Running on a mid-range phone, free, and without a data connection is not a technical flourish here; it is the difference between a tool that reaches people and a demo that does not.

Model quality and speed

Model jepa-rr4 (v0.4.0-jepa) 260-sign set · v3.5.2 full vocabulary. A 260-sign set the model names outright, beside a full vocabulary of 8,165 words matched against stored examples; all numbers below are signer-independent.

Run it on your own device

Figures marked “target” have not been measured yet. They are the budgets the system is being built against, shown so the claims are checkable rather than vague. Every number on this page comes from a single file, src/data/benchmarks.json, with the method recorded beside it.

Recognition quality on signers it has never seen

Isolated-sign top-1 on two corpora the model never trained on — one unseen signer (RKMVU-220) and 379 dictionary clips (CISLR test role).

Two-corpus gate
measured
45.9%

Clip-weighted top-1 over RKMVU-220 + CISLR test-role 379 (599 clips), int8 graph in the browser; the shipped model before this one scored 44.2%, the supervised baseline 14.7%.

Unseen signer, top-1
measured
48.2%

RKMVU-220: 220 clips from one signer absent from training; top-5 is 65.9%. The August baseline was 33.2%.

Vocabulary
measured
260 + 8,165signs

260 INCLUDE signs answered with a calibrated probability; 8,165 of the 8,188 known words reachable by matching against stored examples.

On-device speed

Measured in the browser by the studio's own benches, on the visitor's hardware.

Inference per window
measured
12.5 ms

Mean over 599 real 64-frame windows, int8 ONNX in onnxruntime-web WASM, ONE thread (headless Chromium, Apple M-series); the previous 9.9 MB dictionary model took 34.6 ms on the same windows.

Pose tracking
target
15–30 fps

MediaPipe hand + pose landmarkers, GPU delegate with WASM fallback, mid-range Android — read it live on the studio's /benchmark page.

Working set
target
≤ 250 MB

Peak JS heap + WASM memory during a 60 s session, read from the studio's /benchmark page.

What you download

Once. Then it is cached and works offline.

Dictionary model
measured
4.6 MB

jepa_rr4_probs.int8.onnx, int8 dynamic quantization, velocity/bone/normalisation and softmax inside the graph; the full vocabulary adds its 9.9 MB encoder and stored word examples.

Video bytes sent to a server
measured
0

Landmarks are extracted and classified in Web Workers on the device; no frame, landmark or embedding leaves the browser. The only optional egress is the recognised gloss text, and only if you switch the cloud sentence tier on.

Last reviewed 2026-09-04.

Questions

Does Sanket really not upload my video?

It really does not. The camera frames are read by a Web Worker in your own tab, turned into skeleton landmarks there, and classified by a model that was downloaded to your device. There is no upload endpoint for video or landmarks — not one that is disabled, one that does not exist. You can verify it yourself: open your browser's network panel and sign; nothing leaves.

Then what is the optional cloud tier for?

Recognition produces glosses — bare dictionary words in sign order, like YESTERDAY I SCHOOL GO. Turning those into fluent English reads better with a language model. If you switch that tier on, and only then, the recognised gloss *text* is sent to be rewritten. Never a frame, never a landmark, never audio. The settings screen shows you the exact JSON body before you enable it, built by the same function that sends it.

Which signs does it know?

Two lists at once. The 260-sign set is what the model names outright, with a probability attached — on a signer it has never seen it gets 48% of isolated signs right first time and 66% within its top five. The full vocabulary of 8,165 words is reached by matching the sign against stored examples, which gives a similarity score instead. Both are versioned artifacts the app downloads, so a better model ships as a file, not a rebuild; the studio's Lab page lets you switch between versions and see each one's measured numbers.

What hardware do I need?

A device with a camera and a reasonably current browser. It is built to a mid-range Android phone budget: a 4.6 MB dictionary model plus a 9.9 MB retrieval encoder downloaded once, under 250 MB of working memory, and 15–30 frames per second of pose tracking. Recognition itself takes about 12 ms per window on a laptop CPU with a single thread.

Does it work offline?

Yes, once the model and runtime have been cached. That is a direct consequence of the design — if inference happens on your device, a network connection is only needed to fetch the model the first time.

Is this a replacement for a human interpreter?

No, and it should not be presented as one. It recognises isolated signs from a fixed vocabulary and assembles them into simple sentences. A certified interpreter handles regional variation, fingerspelling, classifier constructions, register and context — none of which a 50-sign classifier attempts. This is a tool for the many everyday moments where no interpreter is available at all.

Is it free?

Yes. It is a research and accessibility project, not a subscription. There is no account to create for the core recognition features, and no telemetry unless you switch it on.

Try it in your browser

Nothing to install, and nothing to sign up for. The model downloads once.

Open the live studio