Review 1 · Deliverable 4 of 4
Expected outcomes
Unusually for an outcomes section, the baselines here were measured before the targets were set — and the first target has now been tested. The table keeps the review-1 column as presented, adds the measured column for 4 September, and restates each target with the gate that decides it, the date that bounds it, and the paper it seeds.
1 · Primary outcome: a real-world-credible recogniser
All numbers are top-1 on the deployed-contract harness unless marked otherwise.
| Metric | At review 1 (30 Aug) | Now (4 Sep) | Phase B target | Stretch (final) |
|---|---|---|---|---|
| Cross-signer accuracy (RKMVU-220) | 0.33 supervised · 0.28 JEPA-A2 | 0.4818 jepa-rr4 (int8 0.4864) | ≥ 0.50 | ≥ 0.65 |
| Two-corpus gate (with CISLR test 379) | — | 0.4591 | tracks | tracks |
| Same, mirrored signer | 0.12 supervised · 0.28 JEPA-A2 | parity by construction (M3); ensemble + mirror TTA 0.4641 | parity | parity |
| In-corpus guardrail | 0.82 | ~0.80–0.83 | ≥ 0.80 | ≥ 0.80 |
| Held-out class half (anti-transductive) | 0.30 JEPA-A2 | 0.5481 | tracks cross-signer | tracks |
| Model size (dictionary lane) | 9.9 MB int8 | 4.57 MB int8 | ≤ deployed | ≤ deployed |
| Latency in-browser | 10–16 ms router · 75–92 ms dictionary + TTA | budget unchanged; re-measurement of the int8 JEPA lane pending | unchanged | unchanged |
The decision rule was pre-registered, and it was applied: Phase B alone missed the 0.50 target (0.3273) and the miss was reported against the gate. The equal-compute fallback did not ship, because the depth finding and the fine-tuning program lifted the same encoder to 0.4818 with the guardrail held — 1.8 pp from the target, at less than half the deployed size. Phase C (in training; +4.8 pp gate in pilot) is the candidate to cross it.
2 · System outcomes (by final review, end of November)
- A bundle drop, not a rewrite — done for the dictionary lane: jepa-rr4 is exported int8 into the existing app contract as a Lab bundle with a model card. Pending: promotion to the default after the gate says it is better, not merely equal.
- The open-vocabulary bank re-embedded in the invariant space — pending; the 4,806-word / 8,188-entry retrieval tier re-scored on the frozen CISLR protocol, measured whatever direction it moves.
- The corpus as an artifact — the 119,064-unit browser-parity unlabelled ISL pool, its builders, the synthetic-signer generator and the gate harness released with the report (licences permitting per source), so the measurement is reproducible by others.
- An honest model card — gate, RKMVU, held-out, mirrored, CISLR-test and in-corpus numbers stated side by side in the shipped bundle, continuing this project's practice of publishing the unflattering number next to the flattering one.
3 · Research outcomes: the semester-2 paper seeds
| Seed | Working claim | Fed by | Status |
|---|---|---|---|
| P1 · Systems paper | On-device, open-vocabulary ISL in the browser: architecture + measured cost model | Review 0 §system, runtime architecture | architecture stable; cost model pending re-measurement |
| P2 · Measurement paper | The deployed-contract cross-corpus audit: within-corpus evaluation systematically overstates ISL readiness | Deliverable 2; the gate; the probes-mislead finding | evidence complete |
| P3 · Method paper | ISL-JEPA: nuisance-view latent prediction for signer-invariant landmark encoders — and the depth/collapse finding | Deliverables 1–4; JEPA page; the results lineage | Phase B done; Phase C pending |
4 · Risks, stated with their mitigations
| Risk | At review 1 | Now |
|---|---|---|
| Phase B misses the 0.50 gate | Equal-compute augmentation baseline built; the gate decides mechanically | It did miss (0.3273). The gate reported it; depth-aware fine-tuning and the recipe program recovered to 0.4818 |
| Single-seed pilot noise (±2–3 pp) | Phase B runs 2 seeds | Decisions on ≥ 2 seeds; the shipping candidate is a 4-seed distillation; one clip = 0.17 pp is stated with every number |
| Tracker failure bounds any landmark model | Tracker-noise channel in M3 | unchanged; JEPA cannot repair garbage landmarks |
| Encoder provenance | — | Found and fixed: a name collision put the pilot encoder where the champion base should be (0.3623 vs 0.4124); every number now names its base |
| Colab VM loss | — | 5-epoch Drive checkpoints with auto-resume, proven at the epoch-11 kill |
| Corpus licence constraints | Per-source terms recorded; gated sources never redistributed as video | unchanged |
| Compute budget | Pilots ≈ $2; Phase B ≈ $15–40 | Phase B ran; Phase C proper is running within the same reserved credits |
5 · Societal outcome (SDG 10)
The outcome this project is accountable to: an ISL recogniser that works for the signer in front of the webcam — either hand dominant, any body type, any camera framing — with their video never leaving their device. The dominance result remains the concrete case: a left-handed signer went from a 12%-accuracy user of this system to a full-accuracy one, first by a shipped inference fix, now by an objective that never learns the bias at all. The remaining gap is quantified — 0.4818 against a 0.50 target on the unseen signer — its closure is funded and running, and the honest number will ship in the product's own model card.