Skip to content

Architecture · 4 of 5

The JEPA objective: what the encoder learns before it sees a label

The encoder is pretrained to predict the latent representation of hidden parts of a clip from the visible parts. Phase B did this against the target encoder's last layer and the target collapsed while the loss kept falling. Phase C predicts an average of every layer instead, and the collapse is gone. This page explains the objective, the collapse, and the one-line change.

Figure 4The objective, and the one line that changed between Phase B and Phase C

How to read this. A clip enters twice: masked, through the context encoder that is being trained, and clean, through an EMA copy of it that produces targets. The predictor tries to reproduce the targets at the masked positions. The two lower arrows are the only difference between phases: which layer(s) of the target encoder the targets come from. Red is the arm that collapsed; green is the arm now in training.

  • Input & masks
  • Encoders & predictor
  • Phase B target (collapsed)
  • Phase C target (stable)
Figure 4The objective, and the one line that changed between Phase B and Phase C

The objective

  • Input. A clip is 64 frames × 107 landmarks × 6 channels (position, velocity and bone vectors, two coordinates each). Tokens are articulators over time.
  • Masks. Three masking families are applied to the context view: m1 hides an articulator stream (a hand, or the body) so it must be inferred from the rest; m2 hides temporal spans; m3 feeds a nuisance view — mirrored, tempo-warped, morphology-rescaled, tracker-noised — and asks it to predict the clean clip's latents, so those nuisances become things the encoder is trained to ignore.
  • Two encoders. The context encoder is trained by gradient; the target encoder is an exponential-moving-average copy of it and produces the targets. Nothing is reconstructed in coordinate space — the loss is L2 between predicted and target latents at the masked tokens.
  • Transfer. After pretraining, the encoder is cut at a depth and fine-tuned with an attentive head (training program).

Phase B: the target collapsed from the top down

Phase B took targets from the target encoder's last layer. Over 100 epochs the pretext loss fell steadily — and the per-dimension standard deviation of the targets fell with it, from 0.98 to 0.52. A target that varies less is easier to predict; the objective was rewarding the target encoder for becoming uninformative. The symptom downstream was a frozen probe at chance on the unseen signer.

Depth probing showed the damage was confined to the top of the network: only layers ≤ 4 transfer. Cutting at layer 4 and fine-tuning (final100@4, shipped as jepa-ship4) reached 0.3273 fp32 / 0.3318 int8 on the unseen signer — a tie with the deployed supervised model at under half its size. That result is why every later model is Cut@4, and why Phase C was designed around the target rather than the encoder.

Phase C: an all-layer teacher

Phase C's target is the mean of all 8 teacher layers, each instance-normalised before averaging and the average layer-normalised — the data2vec-2.0 / Skeleton2vec construction. Averaging over depth makes the target carry both the low-level geometry the early layers keep and the abstraction the late layers add, and a single layer can no longer collapse the target on its own.

The pilot ablation

Four objectives, each pretrained 12 epochs on a random 30k subset of the pool plus INCLUDE (about 34k clips), one seed, then fine-tuned with two seeds each through the standard recipe:

ObjectiveProbe RKMVUFT gate s1 / s2Mean gateRKMVUHeld-out
Phase B (last layer)0.1864.3472 / .3740.3606.3341.4135
+ motion targets0.1091.3656 / .3222.3439.3295.3894
+ all-layer teacher0.1773.4007 / .4157.4082.3955.4808
+ motion + all-layer0.1682.3740 / .3973.3856.3841.4519

Three readings:

  1. The all-layer teacher carries the gain: +4.8 pp gate at equal budget, on both seeds.
  2. Motion targets hurt, alone and in combination — predicting velocity latents adds noise rather than signal for this encoder.
  3. 12 epochs on 34k clips fine-tunes to the level of the 150-epoch pilot encoder. The objective, not the schedule, was the bottleneck.

Phase C proper — in training

The full run: all-layer teacher, 100 epochs, the full 123k-clip pool, on a Colab A100 with checkpoints to Drive every five epochs.

Epoch010203040506070
Pretext loss0.6540.2350.2030.1910.1790.1570.1580.145

The target standard deviation is pinned at 0.999 throughout — no collapse. The run survived an idle-timeout kill at epoch 11 and resumed from its Drive checkpoint, which is the infrastructure working as designed. Its fine-tuned gate (jepa_c4) is pending — expected early on 4 September — and will appear on the results page with the rest of the lineage.