Review 2 · Deliverable 3 of 4
Report completion (50%)
The report is being written as the final case-study submission and as the seed of the semester-2 papers, so it is drafted from the same measured record as these pages. Half of it exists as draft text today. This page lists every chapter with its status, links the draft, and walks through the one part of the work that is new since Review 1 and has no earlier document: the dataset review loop.
Chapter plan and status
| Part | Chapter | Where to read it on this site | Status (27 Sep) |
|---|---|---|---|
| — | Front matter, abstract, Review 2 status page | Review 2 hub | done |
| 1 | Introduction — problem, objectives, SDG 10, scope | Review 0 title and abstract | done |
| 2 | Literature survey | Review 1 survey | done |
| 3 | System design and architecture — runtime, training, evaluation, data loop | Architecture | done |
| 4 | Implementation of algorithms and techniques | Implementation | done |
| 5 | Intermediate results — incl. negative results, frozen model jepa-x12 | Intermediate results | done for Review 2; multi-seed runs of the next levers to add |
| 6 | Expected outcomes and roadmap | Expected outcomes | in progress — updated after the panel's feedback |
| 7 | Conclusion | — | outline only; written for the final review |
| A–C | Appendices — screenshots, reproducibility, glossary | Glossary | done |
How the report is kept honest
- One source for numbers. Every figure in the report comes from the same measured record as these pages; nothing is re-computed or rounded differently for print.
- Original prose and verified citations. Prior work is described in the project's own words; references are carried over from the venue-checked Review 0 and Review 1 lists.
- The stranger number beside every headline. Wherever the accuracy test appears, the RKMVU unseen-signer number is next to it.
- The same diagrams. The hand-drawn figures on these pages are the report's figures, exported from one set of scenes.
The dataset review loop
How to read this. Start at the Dataset tab, where every clip of the four corpora can be inspected. Marks record what is wrong; a vision model reviews the hard cases and escalates what it cannot settle; a blind review of the real videos, with hidden controls, settles some; the rest go to a person in the Review tab, whose decisions return as marks. The contract turns marks into the training pack. The loop arrow is the point: a decision is made once and every later model inherits it.
Between Review 1 and Review 2 the largest single gain came from checking the labels rather than changing the model. The loop has five parts.
- Look at every clip. The dev-only Dataset tab shows every clip of INCLUDE, CISLR, the ISLRTC dictionary and RKMVU with its video, landmarks and skeleton, worst first.
- Mark what is wrong. A mark attaches to a clip, a word, or a rule for a whole corpus,
names an issue from a fixed catalogue, records who raised it (an agent or a person) and can be
flagged
needsReview. - Escalate the hard cases. Vision-model review rounds, working from stick-figure packets, raised 75 escalations — cases it could not settle from the landmarks alone.
- Settle what can be settled blind. The escalations were reviewed on the real video, with hidden controls mixed in — cases whose answer was already known — to check the reviewer; the controls came back 10 of 10 correct. 35 escalations were settled: 21 by INCLUDE's own ground truth, 6 confirmed as different signs, 8 rejected as the same sign. 40 remain open in the Review tab for a human decision.
- Compile the contract.
dataset-marks applyturns the marks into the contract the training packs are built from. The v4 pack was compiled this way (828 of 874 rows).
The effect, measured with the same recipe on four seeds in the same run: accuracy test 0.4116 → 0.4817 (+7.0 pp), all four seeds up. Settling the remaining 40 escalations and rebuilding the pack is the first item on the roadmap.