TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Abstract
Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched 4 times 6 = 24 architecture-objective study at roughly 170M ~ 190M encoder scale on sim1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something-Something V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54% ~ 121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2. HMDB51, IARD, and EPIC-Kitchens bound the claim.
Community
We introduce TT-VidT, a video pretraining framework that treats spatial and temporal separately instead of stacking per-frame embeddings or running one global 3D model: it outputs a per-frame "motion token," and a decoder rebuilds each frame from a key frame plus that token, reaching strong results with only small-scale training. When we flip or time-reverse SSv2 mirror-class videos, TT-VidT keeps its original answer on ≤1% of clips while baselines stay on 9–43%, so it really reads the motion. We think this is a promising direction with lots of room to grow, and we'd love to see you build on it!
project page:
https://kohakublueleaf.github.io/TTVidT/
Good
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- LeVJEPA: Efficient&Scalable Video Pretraining without the Heuristics (2026)
- ACE-Ego-Hand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery (2026)
- V-RAE: Rethinking Video Latent Spaces for Generation (2026)
- Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs (2026)
- Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation (2026)
- Motion-Saliency Complementary Masked Modeling for Point Cloud Video Understanding (2026)
- Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.33419 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 2
KBlueLeaf/TTVidT-decoders
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper