On September 30, 2026, a paper was published on arXiv. Six researchers from the Indian Institute of Technology and other institutions proposed the VideoMSN framework, in which image classifiers can directly process video without modification, surpassing prior methods designed specifically for video on three standard benchmarks and reducing pretraining cost by more than 100x. The paper has been accepted at BMVC 2026.
VideoMSN-DINO (ViT-B specification) achieved 83.3% top-1 accuracy on the Kinetics-400 action recognition benchmark, with only 10 pretraining epochs. VideoMAE needs 1,600 epochs to reach 81.5%, while SMILE needs 600 epochs to reach 83.1%. On UCF101, VideoMSN reaches 95.4% with 10 epochs, whereas VideoMAE's 91.3% corresponds to 3,200 epochs. On HMDB51, VideoMSN reaches 70.4% with 10 epochs, while VideoMAE's 62.6% consumes 4,800 epochs. VideoMSN's FLOPs are about 1/8.9 of VideoMAE's, and its wall-clock time is 12.3x faster.
Squashing Video into Images
VideoMSN arranges multiple frames of a video in a grid and stitches them into one large image, called a “super image.” This operation converts 3D spatiotemporal data into a 2D image input, allowing off-the-shelf image ViTs to process it directly.
From the super images generated from the same video, it constructs two views: one masks image patches at spatial positions (spatial masking), and the other masks several frames along the temporal dimension (temporal masking), with no information leakage between them. A shared ViT encoder processes both views and aligns embeddings through a masked siamese loss. The pipeline requires no decoder and performs no pixel-level reconstruction.
The base encoders chosen are Meta's DINO-v3 and DeiT-v3. DINO-v3 was released in summer 2025 and trained using a 7-billion-parameter teacher model on 1.7 billion images, making it one of the strongest image self-supervised vision encoders. VideoMSN performs video fine-tuning directly on these encoders, transferring prior knowledge from image understanding to temporal scenarios.
Counterintuitive Signals
For the past decade, the dominant view in video understanding has held that image models cannot truly understand motion and that 3D convolutions or temporal attention must be introduced. Methods such as VideoMAE and TimeSformer designed architectures along this line. VideoMSN's experimental results contradict this view.
When multiple frames are tiled into the same super image, spatial relative positions carry temporal information, and the positional relationships between frames implicitly encode motion direction and speed changes. ViT's global attention mechanism captures patch-to-patch associations in a 2D plane, allowing it to learn cross-frame motion patterns. Temporal masking forces the model to reconstruct embeddings when some frames are missing, enabling the encoder to learn inter-frame continuity.
VideoMSN does not modify the image model; instead, it recasts the video problem as one the image model already excels at.
Sources of Efficiency
VideoMSN's efficiency comes from three aspects: encoders such as DINO-v3 have already learned robust visual semantics from massive image data, so video fine-tuning only needs to supplement temporal information; no decoder is needed, eliminating VideoMAE's computation for reconstructing pixels in masked regions; and the super image packs multiple frames into a single forward computation, yielding higher batch efficiency than frame-by-frame 3D approaches.
On Kinetics-400, VideoMAE by default needs 800 epochs to reach a competitive level, and its best configuration uses 1,600 epochs; under the same ViT-B specification, VideoMSN has already exceeded VideoMAE's 1,600-epoch accuracy at 10 epochs. On UCF101 and HMDB51, the gap is even larger. In a low-sample UCF101 experiment with only 1,000 training samples, VideoMSN-DINO reaches 89.0%, 2.6 percentage points higher than SMILE.
The Compute Barrier Shifts
The compute requirements of video self-supervised learning are a major obstacle to scaling. VideoMAE requires tens of GPU hours to complete 800 epochs of pretraining on Kinetics-400, and thousands-of-epoch scenarios cost thousands of dollars, keeping research in this field largely limited to large laboratories.
VideoMSN's efficiency gains make it theoretically feasible to complete video self-supervised pretraining on a single consumer-grade GPU within one day. If the results are validated in broader scenarios, industrial demand for customized video models could be met with a lower compute budget.
Current Limitations
The paper focuses on three action recognition benchmarks: Kinetics-400, UCF101, and HMDB51. Video clips in these datasets are usually shorter than 10 seconds, with clear action categories. Super images effectively encode temporal structure in short-video, concentrated-action scenarios, but there is a lack of data on their performance on complex scene videos lasting several minutes.
DINO-v3's pretraining data consists mainly of internet images, so its performance when transferred to non-natural-image video domains such as infrared, ultrasound, and satellite requires additional verification.
Conclusion
VideoMSN demonstrates that, with an appropriate input representation, image ViTs can effectively encode video spatiotemporal information with very little specialized training. The results of 10 epochs versus 1,600 epochs and 83.3% versus 81.5% are already difficult to explain by error on standard benchmarks.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接