Skip to content

Review 2 · Deliverable 3 of 4

Report completion (50%)

The report is being written as the final case-study submission and as the seed of the semester-2 papers, so it is drafted from the same measured record as these pages. Half of it exists as draft text today. This page lists every chapter with its status, links the draft, and walks through the one part of the work that is new since Review 1 and has no earlier document: the dataset review loop.

Case study report — Review 2 draft (50%)PDF — the same content as these pages, typeset for printing and approval records.

Chapter plan and status

PartChapterWhere to read it on this siteStatus (27 Sep)
—Front matter, abstract, Review 2 status pageReview 2 hubdone
1Introduction — problem, objectives, SDG 10, scopeReview 0 title and abstractdone
2Literature surveyReview 1 surveydone
3System design and architecture — runtime, training, evaluation, data loopArchitecturedone
4Implementation of algorithms and techniquesImplementationdone
5Intermediate results — incl. negative results, frozen model jepa-x12Intermediate resultsdone for Review 2; multi-seed runs of the next levers to add
6Expected outcomes and roadmapExpected outcomesin progress — updated after the panel's feedback
7Conclusion—outline only; written for the final review
A–CAppendices — screenshots, reproducibility, glossaryGlossarydone

How the report is kept honest

  • One source for numbers. Every figure in the report comes from the same measured record as these pages; nothing is re-computed or rounded differently for print.
  • Original prose and verified citations. Prior work is described in the project's own words; references are carried over from the venue-checked Review 0 and Review 1 lists.
  • The stranger number beside every headline. Wherever the accuracy test appears, the RKMVU unseen-signer number is next to it.
  • The same diagrams. The hand-drawn figures on these pages are the report's figures, exported from one set of scenes.

The dataset review loop

Figure 1The dataset review loop

How to read this. Start at the Dataset tab, where every clip of the four corpora can be inspected. Marks record what is wrong; a vision model reviews the hard cases and escalates what it cannot settle; a blind review of the real videos, with hidden controls, settles some; the rest go to a person in the Review tab, whose decisions return as marks. The contract turns marks into the training pack. The loop arrow is the point: a decision is made once and every later model inherits it.

Figure 1The dataset review loop

Between Review 1 and Review 2 the largest single gain came from checking the labels rather than changing the model. The loop has five parts.

  1. Look at every clip. The dev-only Dataset tab shows every clip of INCLUDE, CISLR, the ISLRTC dictionary and RKMVU with its video, landmarks and skeleton, worst first.
  2. Mark what is wrong. A mark attaches to a clip, a word, or a rule for a whole corpus, names an issue from a fixed catalogue, records who raised it (an agent or a person) and can be flagged needsReview.
  3. Escalate the hard cases. Vision-model review rounds, working from stick-figure packets, raised 75 escalations — cases it could not settle from the landmarks alone.
  4. Settle what can be settled blind. The escalations were reviewed on the real video, with hidden controls mixed in — cases whose answer was already known — to check the reviewer; the controls came back 10 of 10 correct. 35 escalations were settled: 21 by INCLUDE's own ground truth, 6 confirmed as different signs, 8 rejected as the same sign. 40 remain open in the Review tab for a human decision.
  5. Compile the contract. dataset-marks apply turns the marks into the contract the training packs are built from. The v4 pack was compiled this way (828 of 874 rows).

The effect, measured with the same recipe on four seeds in the same run: accuracy test 0.4116 → 0.4817 (+7.0 pp), all four seeds up. Settling the remaining 40 escalations and rebuilding the pack is the first item on the roadmap.