AI Research NotesLeVJEPA · SIGReg · Causal Video Encoder · 27 Aug 2026
LeVJEPA · arXiv:2608.27395v1 · 27 Aug 2026

비디오를 더 복잡하게 학습하는 대신
95%를 버리고도 더 잘 배우는
하나의 인코더

Efficient & Scalable Video Pretraining without the Heuristics — collapse-free joint embedding, sparse observation, and causal representations by construction.

GLOBAL VIEWLOCAL VIEWS95% RANDOMTOKEN DROP95% RANDOMTOKEN DROPSHARED BLOCK-CAUSAL ENCODERSHARED BLOCK-CAUSAL ENCODER[CLS] MSEINVARIANCESIGRegGAUSSIANDISTRIBUTIONNo target encoder · no predictor · no stop-gradient · no masked reconstruction
Central Thesis

LeVJEPA는 비디오 self-supervised learning에서 collapse를 막기 위해 쌓아온 구조적 비대칭을 제거하고, 단일 encoder + projector와 단일 composite objective만으로 학습한다. 그 결과 pretraining 비용은 사실상 encoder가 실제로 보는 token 수에 의해 결정되고, 95% uniform random token dropping이 비용을 크게 줄이는 동시에 downstream representation을 오히려 개선한다.

논문은 세 가지를 동시에 주장한다. 첫째, ViT-S/B/L에서 동일 epoch·동일 data 기준 V-JEPA 2와 비슷하거나 더 나은 정확도를 5.6–20.8× 낮은 total pretraining compute로 얻는다. 둘째, block-causal attention이 bidirectional attention과 사실상 같은 frozen-probe 정확도를 보여 causal frame representation을 별도 temporal model 없이 encoder 자체에 넣을 수 있다. 셋째, 동일 FLOP budget에서 image-pretrained DINOv2와 appearance 성능 격차를 몇 point 수준으로 줄이면서 motion-centric 성능은 거의 두 배로 만든다.

페이지 1의 Figure 1은 이 단순화를 가장 잘 보여준다. global/local view 모두 같은 encoder를 통과하고, 오직 [CLS] embedding의 MSE invariance와 SIGReg만 학습신호로 사용된다. target encoder, predictor, stop-gradient, masked-token reconstruction은 없다.

이 논문의 가장 흥미로운 메시지는 “video를 더 싸게 학습한다”가 아니다. collapse prevention의 heuristics를 제거하자 masking, temporal aggregation, attention topology가 다시 자유로운 설계변수가 되었고, 가장 단순하고 sparse한 설정이 오히려 가장 강력한 축에 들어섰다는 점이다.
Part I · Why Simpler?

비디오가 비싼 이유와 self-supervised recipe가 비싼 이유를 분리한다

§1 · Video as Physical Supervision

motion, causality, object permanence는 정지영상이 제공하지 않는 구조다

저자들은 video를 physical world representation learning의 자연스러운 substrate로 본다. annotation 없이 풍부하게 존재하고 motion, causality, object permanence 같은 temporal structure를 담기 때문이다. 그러나 한 clip은 image보다 order-of-magnitude 더 많은 token을 포함하고, 기존 joint-embedding method는 여기에 target encoder, predictor, stop-gradient, structured masking까지 더해 계산량을 키웠다.

§2 · Two Previous Routes

collapse prevention 또는 pixel reconstruction

Joint Embedding

BYOL/DINO/V-JEPA 계열은 online/EMA target branch, stop-gradient, predictor 같은 비대칭으로 collapse를 피한다.

Masked Reconstruction

VideoMAE 계열은 pixel reconstruction task 자체가 trivial collapse를 허용하지 않으므로 decoder와 structured tube masking을 사용한다.

LeVJEPA는 둘 다 택하지 않는다. embedding distribution 자체를 명시적으로 제약해 collapse를 배제한다. 이 때문에 이전 recipe에 종속돼 있던 여러 heuristic이 더 이상 필수가 아니게 된다.

§3 · Related Work Map
Research lineRepresentative methodsLeVJEPA와의 차이
Image joint embeddingBYOL, DINO, DINOv2, I-JEPAglobal–local view는 계승하지만 teacher–student/stop-grad/masked prediction machinery는 제거
Embedding constraintsVICReg, Barlow Twins, LeJEPALeJEPA의 isotropic-Gaussian characterization과 SIGReg를 video에 이전
Pixel reconstructionVideoMAE, VideoMAEv2imputation이 없으므로 tube mask가 필요하지 않으며 random sparse observation이 더 유리
Video feature predictionV-JEPA, V-JEPA 2, V-JEPA 2-AC별도 target encoder/predictor와 post-hoc temporal model 없이 pretraining encoder 자체에 causality를 삽입
Generative / world modelsGenie, Dream-to-Control, DINO-WMaction-conditioned dynamics를 직접 학습하는 것이 아니라 causal state encoder의 foundation을 제공
Part II · Collapse-Free Objective

global–local invariance와 SIGReg, 딱 두 항의 목적함수

§4 · Views & Readout

16-frame clip에서 한 global view와 여러 local view

16-frame clip 하나에서 full-resolution global view \(x_0\)와 aggressive spatial crop 및 photometric augmentation으로 만든 \(V\)개의 local view \(x_1,\dots,x_V\)를 만든다. 모든 view는 동일한 temporal window를 공유하며 spatial/photometric aspect만 다르다.

모든 view는 동일 encoder \(E_\theta\)를 공유한다. learnable [CLS] token이 clip-level readout을 제공하고 작은 projector \(h_\phi\)가 이를 \(K\)-dimensional embedding으로 바꾼다. projector는 pretraining 후 버리고 downstream은 encoder representation만 사용한다.

§5 · Objective

단 하나의 trade-off hyperparameter, \(\lambda=0.02\)

\[\mathcal{L}=\mathcal{L}_{\mathrm{inv}}+\lambda\mathcal{L}_{\mathrm{SIGReg}},\qquad \lambda=0.02\]

invariance term은 local view embedding이 같은 clip의 global embedding과 가까워지도록 MSE를 사용한다.

\[\mathcal{L}_{\mathrm{inv}}=\frac{1}{V+1}\sum_{v=0}^{V}\lVert z_0-z_v\rVert_2^2\]

중요한 점은 양쪽 branch 모두 gradient가 흐른다는 것이다. global target에도 stop-gradient가 없고 target network도 없다. 따라서 이 항만 사용하면 constant embedding이라는 trivial collapse가 가능하다.

§6 · SIGReg

고차원 분포를 random 1D projection의 normality test로 제약

SIGReg는 embedding distribution을 isotropic Gaussian \(\mathcal{N}(0,I_K)\)에 맞춘다. LeJEPA의 분석에 따르면 이 분포는 mild assumptions 아래 worst-case downstream probing risk를 최소화하며, variance가 어느 방향에서 0이 되는 collapsed solution은 이 분포에서 멀어진다.

Cramér–Wold theorem을 이용해 고차원 제약을 random univariate projection의 표준정규성 검사로 바꾸고, Epps–Pulley statistic으로 각 projection의 deviation을 penalize한다. empirical characteristic function과 gradient가 bounded되어 whitening/centering 없이 outlier에 robust하고 distributed setting에서도 communication overhead가 작다.

Part III · Architecture & Evaluation

per-frame tokenization, 95% random dropping, block-causal attention

§7 · Default Architecture
ComponentDefaultMeaning
BackboneViT-B/16video-adapted Vision Transformer
Clip16 framesglobal/local views share same temporal window
Temporal patch extent\(\tau=1\)한 token이 한 frame의 16×16 spatial patch에 대응
Global / local token count3,136 / 576 before dropping224² global, 96² local view
Drop ratio\(\rho=0.95\)patch token의 95%를 uniform random discard
Attentionblock-causalframe 내부는 bidirectional, frame 간은 causal
Positionfactorized 3D rotarytemporal/vertical/horizontal relative position
Supervised token[CLS] onlypatch token에는 직접 loss가 없음
Trainable partsencoder + projectorpredictor/target encoder 없음
§8 · Why Causal?

frame representation을 과거관측만으로 계산

block-causal attention에서는 현재 frame token이 같은 frame과 과거 frame만 본다. [CLS]는 모든 token을 읽지만 patch token이 [CLS]를 보지는 않는다. 결과적으로 frame representation은 current/past observation의 함수가 된다.

이 구조는 streaming perception과 autoregressive world model에 중요하다. 새 frame이 들어올 때 과거 frame을 다시 encoding하지 않고 representation을 frame-by-frame 확장할 수 있기 때문이다. temporal ordering을 post-hoc temporal model이 아니라 pretrained encoder 자체의 속성으로 만든다.

§9 · Data & Evaluation

controlled comparison을 위해 동일 source data, frozen probing

기본 pretraining은 Kinetics-400/600/700 training set의 union인 K710에서 validation overlap을 제거한 뒤 class-balanced 20% subset을 사용한다. 확대실험에서는 K710 + Something-Something-v2 + Walking Tours + Perception Encoder Video Dataset을 결합한다.

encoder는 frozen 상태로 평가한다. ImageNet-1K와 Something-Something-v2는 learnable query의 cross-attention attentive probe를 쓰고, Kinetics-400은 mean-pooled frozen token 위 linear classifier를 사용한다. 평가축은 appearance-centric static recognition, motion-centric understanding, action recognition을 함께 포함한다.

Part IV · What Actually Matters?

가장 과감한 token dropping이 가장 싸고, ImageNet에서는 가장 정확하다

§10 · Token Dropping

0% 처리보다 95% 버릴 때 33.9 → 47.6

33.9
39.6
44.8
47.4
47.6

페이지 6의 Figure 4(a)는 token dropping이 단순 approximation이 아니라 augmentation처럼 작동함을 보여준다. sparse random observation만 보고도 같은 clip의 global/local representation이 일치해야 하므로 encoder가 부분관측에 강한 clip-level representation을 배운다는 해석이다.

다만 motion-centric SSv2에서는 short schedule에서 drop ratio가 0.3을 넘으면 성능이 떨어진다. longer training이 이를 상당 부분 회복하지만, high sparsity에서 temporal correspondence를 더 잘 보존하는 dropping scheme은 open problem으로 남는다.

§11 · Random Beats Tube

structured mask는 reconstruction task의 산물이지 video 자체의 법칙이 아니다

uniform random dropping은 ImageNet에서 50.7, tube dropping은 39.6이다. SSv2에서도 \(\tau=2\) 설정에서 28.8 대 26.4로 random이 앞선다. masked prediction에서는 tube mask가 adjacent-frame interpolation shortcut을 막지만, LeVJEPA는 imputation을 하지 않는다. 여기서는 retained token이 encoder가 보는 전부이므로 tube pattern은 scene의 대부분을 시간 전체에 걸쳐 영구적으로 가려버린다.

§12 · Local View Budget

V=10까지는 작은 비용으로 추가 이득

Local views VIN1K attentive top-1
447.6
649.1
849.2
1050.2
1249.8

local view 하나를 추가하는 비용은 약 29 processed tokens로 global view 비용의 1/5보다 작다. 논문은 비교실험의 기본값은 V=4로 유지하지만, V=10까지는 modest extra compute로 추가 accuracy를 얻을 수 있음을 보여준다.

§13 · Temporal Patching & Attention
Design choiceIN1KSSv2Interpretation
\(\tau=2,\rho=0.90\)47.428.82-frame temporal aggregation
\(\tau=1,\rho=0.95\)50.730.4per-frame tokenization; same retained-token budget
AttentionIN1K top-1
Bidirectional50.7
Block-causal51.2

temporal patch aggregation은 필요하지 않았고, causal attention도 measurable accuracy penalty가 없었다. 즉 temporal causality를 얻기 위해 representation quality를 희생해야 한다는 가정이 이 실험에서는 성립하지 않았다.

Part V · Compute-Matched Evidence

같은 epoch에서는 훨씬 싸고, 같은 FLOPs에서는 더 오래 학습해 더 강해진다

§14 · Epoch-Matched

240 epochs, effective batch 3,072, identical 20% K710

ViT-S/B/L을 같은 data와 같은 epoch 수로 다시 학습한 비교에서 LeVJEPA는 V-JEPA 2와 비슷하거나 더 높은 정확도를 5.6×–20.8× 적은 total pretraining compute로 달성한다. ViT-B에서는 4.8 ExaFLOPs 대 36.4 ExaFLOPs로 정확도 차이가 1 point 미만이며, ViT-L에서는 LeVJEPA가 1.9 point 높으면서 compute는 5.6× 낮다.

페이지 2의 Figure 2가 이 관계를 시각화한다. x축을 total ExaFLOPs의 역방향 log scale로 놓으면 LeVJEPA가 V-JEPA 2와 VideoMAEv2 대비 오른쪽, 즉 훨씬 싼 compute 영역에서 경쟁력 있는 정확도를 형성한다.

§15 · FLOP-Matched

같은 total compute를 주면 1,085 epochs까지 학습

MethodIN1KSSv2K400
VideoMAEv253.443.637.4
V-JEPA 251.642.540.7
LeVJEPA61.040.444.6

equal-FLOP budget에서는 LeVJEPA가 sample당 싸기 때문에 더 긴 schedule을 사용할 수 있다. V=10 local views로 1,085 epochs를 학습했고 ImageNet에서 가장 강한 video baseline보다 7.6 point 높고 K400도 최고다. motion-centric SSv2에서는 최고 baseline보다 3.2 point 낮아, motion은 여전히 가장 큰 headroom임을 보여준다.

§16 · Video vs Image Pretraining

appearance는 3.1 point 차이, motion은 거의 두 배

MethodIN1KSSv2
DINOv2 · video frame-based image pretraining53.816.9
LeVJEPA · video pretraining50.730.4

동일 source video에서 뽑은 frame으로 DINOv2를 동일 total FLOPs만큼 학습했다. image method가 ImageNet에서 3.1 point 앞서지만 LeVJEPA는 SSv2에서 거의 두 배다. 저자들은 video pretraining의 appearance 비용이 몇 point로 줄어든다면, temporal structure까지 포함하는 video가 general-purpose representation의 더 효율적인 substrate가 될 수 있다고 해석한다.

§17 · Consumer Hardware

RTX 5080 16GB 한 장, 12시간

12 hSingle GPU
620kWalking Frames
~5MProcessed Clips
25.2%IN1K Top-1

ViT-Tiny를 Walking Tours의 8개 unlabeled egocentric video로 12시간 학습했다. frozen ImageNet top-1은 initialization 8.9%에서 25.2%로 상승했다. 같은 RTX 5080 16GB에서 LeVJEPA는 batch 128을 8GB 미만으로 학습하는 반면, 같은 encoder size의 V-JEPA configuration은 batch 28에서 card memory를 포화시킨다.

§18 · Scaling Data

objective를 바꾸지 않고 corpus만 키워도 계속 개선

ViT-L/16을 K710 + SSv2 + Walking Tours + PE Video Dataset의 combined corpus에서 100 epochs 학습하면 frozen attentive probe가 69.5% ImageNet, 55.0% SSv2에 도달한다. 20% K710 ViT-L 대비 ImageNet은 9.5 point 상승한다. objective와 \(\lambda=0.02\)는 그대로다.

Part VI · Emergent Structure & Implications

[CLS]만 학습했는데 patch token이 객체와 배경을 스스로 분리한다

§19 · Dense Semantics Without Dense Loss

페이지 3 Figure 3과 페이지 8 Figure 5

LeVJEPA의 loss는 clip-level [CLS]에만 걸린다. patch token에는 direct supervision이 없다. 그런데 patch-token PCA를 RGB로 시각화하면 animal과 furniture/background가 분리되고, query patch와의 cosine similarity도 object boundary 안에 sharp하게 모인다.

논문은 이 구조가 patch-level auxiliary loss가 없는 V-JEPA 2보다 훨씬 선명하고, explicit dense objective를 도입한 V-JEPA 2.1과 comparable한 시각적 조직성을 보인다고 보고한다. clip-level invariance와 distributional regularization만으로 spatially precise semantic token이 emergent할 수 있다는 관찰이다.

§20 · World-Model Relevance

representation learning과 temporal structure를 pretraining 단계에서 결합

기존 V-JEPA 2-AC나 DINO-WM 계열은 frozen visual encoder 위에 별도의 temporal/action model을 적합해 planning/control을 수행한다. LeVJEPA는 action-conditioned dynamics 자체를 해결하지는 않지만 block-causal representation을 pretraining encoder에서 직접 만든다.

따라서 streaming perception, autoregressive world modeling, online observation encoding에서 과거 frame 재인코딩을 줄일 수 있는 foundation이 된다. 효율성은 누가 train할 수 있는지를, causality는 무엇에 사용할 수 있는지를, simplicity는 새로운 regime으로 얼마나 안정적으로 옮길 수 있는지를 결정한다는 것이 저자들의 결론이다.

§21 · Open Problems

아직 입증되지 않은 세 축

Motion-Preserving Sparsity

95% dropping에서 temporal correspondence를 유지하는 scheme이 필요하다.

Internet-Scale Regime

very large batch/model과 internet-scale video에서 SIGReg의 거동은 아직 통제실험이 없다.

Dense Prediction

emergent patch semantics가 segmentation·tracking에 충분한지는 평가되지 않았다.

또한 current controlled experiment는 최대 ViT-L과 제한된 corpus 범위다. favorable data scaling은 saturation의 증거가 아니라 headroom의 신호로 해석해야 한다.

§22 · Final Takeaway

video가 general-purpose pretraining substrate가 될 조건

LeVJEPA가 직접 보여준 범위는 “single encoder, single loss, one fixed hyperparameter”로 competitive self-supervised video representation을 훨씬 적은 compute에 학습할 수 있다는 것이다. video가 image pretraining을 완전히 대체했다고 증명한 것은 아니지만, video를 선택하지 못하게 했던 compute penalty가 생각보다 구조적 필연이 아니었음을 보여준다.

video는 가장 풍부한 supervision source였지만 가장 비싼 선택지였다. LeVJEPA는 그 비용의 상당 부분이 video 자체가 아니라 우리가 collapse를 피하기 위해 덧붙인 training recipe에 있었을 수 있음을 보여준다.
Part VII · Appendix & Reproducibility Details

12페이지의 마지막까지 보면 단순성의 구현 조건이 더 선명해진다

§23 · SIGReg Implementation
DetailValue
Quadrature17 knots, \(t\in[0,3]\)
Random directions\(M=1,024\) per step
Distributed evaluationfull global batch with one all-reduce
Communication payloadsingle \(2\times M\times17\) tensor; independent of batch size and embedding dimension
Whitening / centeringnot required
§24 · Projector & Checkpoint
ComponentSpecification
3D rotaryattention-head dimensions partitioned into temporal / vertical / horizontal groups
ProjectorLinear \(d\to2048\) → BatchNorm → GELU → Linear \(2048\to K\), \(K=256\)
Polyak averagedecay 0.9999, updated every 32 optimizer steps
Role of Polyak modelevaluation checkpoint only; no forward pass and no role in training objective
§25 · Probe Details

representation information content를 보는 frozen evaluation

attentive probe는 learnable query 하나가 frozen encoder의 모든 output token에 cross-attend하고, residual + 2-layer MLP + GELU + layer norm 뒤 linear classifier를 학습한다. K400은 frozen token mean pooling 뒤 linear classifier를 사용한다. 저자들은 K400의 mean-pooling linear probe가 attentive probe보다 약한 adaptation이므로 K400 score를 보수적 추정으로 본다.

§26 · Authors & Provenance

DKFZ, Mila, Brown, NYU, AMI Labs의 공동연구

Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner가 참여했다. 소속은 German Cancer Research Center, German Cancer Consortium, Goethe University Frankfurt, Mila, Université de Montréal, Brown University, Courant Institute at NYU, Advanced Machine Intelligence (AMI Labs)이다. Randall Balestriero와 Florian Buettner는 equal advising이다.

연구는 hessian.AI Service Center 및 Innovation Lab, EU ERC TAIPO의 지원을 받았다고 acknowledgments에 명시한다.

Full bibliography map · 33 references
  1. Arnab et al. — ViViT: A Video Vision Transformer (ICCV 2021).
  2. Assran et al. — Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture (I-JEPA, CVPR 2023).
  3. Assran et al. — V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (2025).
  4. Ba, Kiros & Hinton — Layer Normalization (2016).
  5. Balestriero & LeCun — LeJEPA: Provable and Scalable Self-Supervised Learning without the Heuristics (2025).
  6. Bardes, Ponce & LeCun — VICReg (2021).
  7. Bardes et al. — V-JEPA: Latent Video Prediction for Visual Representation Learning (2024).
  8. Bolya et al. — Perception Encoder: The Best Visual Embeddings Are Not at the Output of the Network (NeurIPS 2026).
  9. Bruce et al. — Genie: Generative Interactive Environments (ICML 2024).
  10. Busbridge et al. — Distillation Scaling Laws (2025).
  11. Caron et al. — Emerging Properties in Self-Supervised Vision Transformers (DINO, ICCV 2021).
  12. Chen et al. — SimCLR (ICML 2020).
  13. Dosovitskiy et al. — An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale (ICLR 2021).
  14. Epps & Pulley — A Test for Normality Based on the Empirical Characteristic Function (Biometrika 1983).
  15. Goyal et al. — Something-Something Video Database (ICCV 2017).
  16. Grill et al. — BYOL (NeurIPS 2020).
  17. Hafner et al. — Dream to Control: Learning Behaviors by Latent Imagination (ICLR 2020).
  18. Kay et al. — The Kinetics Human Action Video Dataset (2017).
  19. Kuhn et al. — LeVLJEPA: End-to-End Vision-Language Pretraining without Negatives (2026).
  20. Maes et al. — LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels (2026).
  21. Mur-Labadia et al. — V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning (2026).
  22. Oquab et al. — DINOv2: Learning Robust Visual Features without Supervision (2023).
  23. Oquab et al. — DINOv2, TMLR version (2024).
  24. Rao & Ballard — Predictive Coding in the Visual Cortex (Nature Neuroscience 1999).
  25. Russakovsky et al. — ImageNet Large Scale Visual Recognition Challenge (IJCV 2015).
  26. Spelke et al. — Spatiotemporal Continuity, Smoothness of Motion and Object Identity in Infancy (1995).
  27. Su et al. — RoFormer: Enhanced Transformer with Rotary Position Embedding (2024).
  28. Tong et al. — VideoMAE (NeurIPS 2022).
  29. Venkataramanan et al. — Is ImageNet Worth 1 Video? Learning Strong Image Encoders from 1 Long Unlabelled Video (ICLR 2024).
  30. Wang et al. — VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking (CVPR 2023).
  31. Zbontar et al. — Barlow Twins (ICML 2021).
  32. Zhou et al. — DINO-WM: World Models on Pre-Trained Visual Features Enable Zero-Shot Planning (ICML 2025).
  33. Zhou et al. — Image BERT Pre-Training with Online Tokenizer (ICLR 2022).
Primary Source & Project

논문과 공개 코드

01
LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
Kuhn et al. · arXiv:2608.27395v1 · 27 Aug 2026
02
LeVJEPA Project Page
Code and models

본 게시물은 첨부된 12페이지 논문의 본문, Figure 1–5, Table 1–4, Discussion, Appendix A–C와 bibliography를 기준으로 재구성했다. 원문이 직접 검증하지 않은 dense prediction, internet-scale scaling, downstream planning 성능은 open problem 또는 해석으로 명시했다.