Research collection · maintained with the survey

Video SSL & VideoSSL Papers by Year & Venue

A researcher-maintained catalog of video self-supervised learning and self-supervised video representation learning papers, organized by verified publication year and latest confirmed venue.

Collection snapshot

Repository statistics

Public counts are generated from verified publication years and venues.

282representation-learning papers
2016–2026years covered
64normalized venues
Bar chart of video self-supervised learning papers by publication year
Papers by year
Bar chart of Video SSL papers by publication venue
Papers by venue
Research coverage

Video SSL terminology, datasets and venues

The catalog keeps per-paper cards focused on year and venue while this overview explains the broader research scope.

Video self-supervised learning

Video SSL, VideoSSL and SSL video research refer to self-supervised learning for video, including self-supervised video representation learning, video representation pretraining and masked video modeling.

Video understanding datasets

The collection covers research using UCF101, HMDB51, Kinetics-400, Kinetics-600, Kinetics-700, Something-Something V1, Something-Something V2, Diving48, EPIC-KITCHENS, AVA, FineGYM, Charades, Ego4D and related datasets.

Conferences and journals

Verified venues include CVPR, ICCV, ECCV, NeurIPS, ICLR, AAAI, WACV, ACM Multimedia, specialist workshops and peer-reviewed journals, with arXiv retained only when no later publication is confirmed.

Interactive research index

Browse the paper collection

Filter the catalog instead of scrolling through a single long bibliography. The GitHub README keeps the traditional year-by-year list.

282 papers
2026CVPR

Progressive Mask Distillation for Self-supervised Video Representation

Kewei Wu, Chong Liang, Zhao Xie, Dan Guo

2026CVPR

TrackMAE: Video Representation Learning via Track Mask and Predict

Renaud Vandeghen, Fida Mohammad Thoker, Marc Van Droogenbroeck, Bernard Ghanem

2026arXiv / Preprint

V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

Lorenzo Mur-Labadia, Matthew Muckley, Amir Bar, Mido Assran, Koustuv Sinha, Mike Rabbat, Yann LeCun, Nicolas Ballas, Adrien Bardes

2026CVPR

From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning

Yang Liu, Qianqian Xu, Peisong Wen, Siran Dai, Xilin Zhao, Qingming Huang

2026arXiv / Preprint

The TIME Machine: On The Power of Motion for Efficient Perception

Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara

2026arXiv / Preprint

TrAction: Action Recognition with Sparse Trajectories

Jan F. Meier, Felix B. Mueller, Alexander Ecker, Timo Lüddecke

2026arXiv / Preprint

OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

Feilong Tang, Xiang An, Yunyao Yan, Yin Xie, Bin Qin, Kaicheng Yang, Yifei Shen, Yuanhan Zhang, Chunyuan Li, Shikun Feng, Changrui Chen, Huajie Tan, Ming Hu, Manyuan Zhang, Bo Li, Ziyong Feng, Ziwei Liu, Zongyuan Ge, Jiankang Deng

2026arXiv / Preprint

Factorized Latent Dynamics for Video JEPA: An Empirical Study of Auxiliary Objectives

Santosh Premi

2026arXiv / Preprint

Self-Supervised Learning of Structured Dynamics from Videos

Lukas Knobel, Andrew Zisserman, Yuki M. Asano

2026arXiv / Preprint

Depth-Wise Representation Development Under Blockwise Self-Supervised Learning for Video Vision Transformers

Jonas Römer, Timo Dickscheid

2026Engineering Applications of Artificial Intelligence

Beyond reconstruction: Enhancing masked autoencoders with contrastive learning for video representation learning

Yawei Feng, Lijun Guo, Guitao Yu, Rong Zhang, Jiangbo Qian, Chong Wang, Shangce Gao

2026ECCV

Structured-Noise Masked Modeling for Video, Audio and Beyond

Aritra Bhowmik, Fida Mohammad Thoker, Carlos Hinojosa, Bernard Ghanem, Cees G. M. Snoek

2026ICLR

Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers

Xianhang Li, Chen Huang, Chun-Liang Li, Eran Malach, Josh Susskind, Vimal Thilak, Etai Littwin

2026CVPR

Recurrent Video Masked Autoencoders

Daniel Zoran, Nikhil Parthasarathy, Yi Yang, Drew A. Hudson, Joao Carreira, Andrew Zisserman

2026ICLR

Dual Perspectives on Non-Contrastive Self-Supervised Learning

Jean Ponce, Martial Hebert, Basile Terver ;

2026International Journal of Computer Vision

Self-Supervised Video Representation Learning in a Heuristic Decoupled Perspective

Zeen Song, Jingyao Wang, Jianqi Zhang, Changwen Zheng, Wenwen Qiang

2026IEEE TCSVT

BIMM: Brain Inspired Masked Modeling for Video Representation Learning

Zhifan Wan, Jie Zhang, Changzhen Li, Shiguang Shan

2025CVPR Workshops

Efficient VideoMAE via Temporal Progressive Training

Xianhang Li, Peng Wang, Xinyu Li, Heng Wang, Hongru Zhu, Cihang Xie

2025ICCV

An Empirical Study of Autoregressive Pre-training from Videos

Jathushan Rajasegaran, Ilija Radosavovic, Rahul Ravishankar, Yossi Gandelsman, Christoph Feichtenhofer, Jitendra Malik

2025ICCV Workshops

Reinforcement Learning Meets Masked Video Modeling: Trajectory-Guided Adaptive Token Selection

Ayush K. Rai, Kyle Min, Tarun Krishna, Feiyan Hu, Alan F. Smeaton, Noel E. O'Connor

2025ACROSET

Entropy-Guided Masked Autoencoding for Self-Supervised Human Action Recognition Using Video Swin Transformer

Kollu Praveen Kumar; Guduri Baby Harshitha; Koti Vijay; Angothu Sravika

2025ICCV Workshops

Privacy Preservation Using Superimposed 3D-Models for Self-Supervised Training in Action Recognition

Asfandyar Azhar, Nidhish Shah, Shaurjya Mandal, Yongjie Jessica Zhang;

2025Machine Learning

Learning Complementary Knowledge via Trusted Multi-view Space Decomposition for Self-Supervised Contrastive Learning

Jiangmeng Li, Yunze Zhao, Yifan Jin, Changwen Zheng & Wenwen Qiang;

2025NeurIPS

OSKAR: Omnimodal Self-supervised Knowledge Abstraction and Representation

Mohamed O Abdelfattah, Kaouther Messaoud, Alexandre Alahi;

2025arXiv / Preprint

MME: Video Representation Learning as World Model for Understanding and Planning

Xinyu Sun, Changhao Li, Chen Jian, Chuang Gan, Peihao Chen, and Mingkui Tan;

2025ICCV Workshops

Hashtag2Action: Data Engineering and Self-Supervised Pre-Training for Action Recognition in Short-Form Videos

Yang Qian, Ali Kargarandehkordi, Yinan Sun, Parnian Azizian, Onur Cezmi Mutlu, Saimourya Surabhi, Zain Jabbar, Dennis Wall, Peter Washington, Huaijin Chen;

2025Cluster Computing

Kdhiera: boosting self-supervised masked video modeling via hierarchical knowledge distillation

Yunlong Wang, Hong Liang, Mingwen Shao & Qian Zhang;

2025Proceedings of SPIE

Self-supervised video representation learning based on foreground and temporal information.

Zhongliang Zhou, Jiayong Fang ;

2025IJCV

Feature Hallucination for Self-supervised Action Recognition

Lei Wang, Piotr Koniusz; ;

2025CVPR Workshops

ViDROP: Video Dense Representation through Spatio-Temporal Sparsity

Sepehr Sameni, Simon Jenni, Paolo Favaro; ;

2025CVPR

SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding

Yangliu Hu, Zikai Song, Na Feng, Yawei Luo, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang; ;

2025arXiv / Preprint

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba, Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, Xiaodong Ma, Sarath Chandar, Franziska Meier, Yann LeCun, Michael Rabbat, Nicolas Ballas ;

2025CVPR

When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning

Yang Liu, Qianqian Xu, Peisong Wen, Siran Dai, Qingming Huang;

2025NeurIPS

Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals

Stefan Stojanov, David Wendt, Seungwoo Kim, Rahul Venkatesh, Kevin Feigelis, Jiajun Wu, Daniel LK Yamins

2025ICMR

Label Ranker: Self-aware Preference for Classification Label Position in Visual Masked Self-supervised Pre-trained Model

Peihao Xiang, Ou Bai

2025CVPR

AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing

Niu Lian, Jun Li, Jinpeng Wang, Ruisheng Luo, Yaowei Wang, Shu-Tao Xia, Bin Chen

2025AAAI

Efficient Self-Supervised Video Hashing with Selective State Spaces

Jinpeng Wang, Niu Lian, Jun Li, Yuting Wang, Yan Feng, Bin Chen, Yongbing Zhang, Shu-Tao Xia1

2025Image and Vision Computing

Exemplar-free class incremental action recognition based on self-supervised learning

Chunyu Hou, Yonghong Hou, Jinyin Jiang, Gunel Abdullayeva

2025CVPR

Learning from Streaming Video with Orthogonal Gradients

Tengda Han⋄, Dilara Gokay, Joseph Heyward, Chuhan Zhang, Daniel Zoran, Viorica Patraucean, Joao Carreira, Dima Damen, Andrew Zisserman

2025arXiv / Preprint

Intuitive physics understanding emerges from self-supervised pretraining on natural videos

Quentin Garrido, Nicolas Ballas, Mahmoud Assran, Adrien Bardes, Laurent Najman, Michael Rabbat, Emmanuel Dupoux, Yann LeCun

2025Pattern Analysis and Applications

ST-HViT: spatial-temporal hierarchical vision transformer for action recognition

Limin Xia, Weiye Fu

2025Pattern Recognition Letters

Advancing video self-supervised learning via image foundation models

Jingwei Wu, Zhewei Huang, Chang Liu

2025CVPR

SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning

Fida Mohammad Thoker, Letian Jiang, Chen Zhao†, Bernard Ghanem

2025CVPR Workshops

A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning

Akash Kumar, Ashlesha Kumar, Vibhav Vineet, Yogesh S Rawat

2025Proceedings of SPIE

Progressive self-supervised spatio-temporal feature learning based on video sequence saliency

Jinlong Kang, Tao Xu, Boting Qu, Xiang Wang, Xiaoli Lian, Jing Guo, Yuan Gao

2025arXiv / Preprint

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders

Shihab Aaqil Ahamed∗, Malitha Gunawardhana∗, Liel David, Michael Sidorov, Daniel Harari, Muhammad Haris Khan

2025EURASIP Journal on Image and Video Processing

Motion-driven Adaptive Frame Selection Strategy for Video Action Recognition

Hao Ding, Chen Guo, Jing Sun, Xiaoping Jiang, Hongling Shi, Jianjin Li

2025Signal, Image and Video Processing

Mitigating background bias in self-supervised video representation learning

Arif Akar, Ufuk Umut Senturk & Nazli Ikizler-Cinbis

2025TMLR

ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning

Sucheng Ren, Hongru Zhu, Chen Wei, Yijiang Li, Alan Yuille, Cihang Xie

2025ICCV Workshops

Learning Video Representations without Natural Videos

Xueyang Yu, Xinlei Chen, Yossi Gandelsman;

2025IEEE TIFS

Collaboratively Self-supervised Video Representation Learning for Action Recognition

Jie Zhang, Zhifan Wan, Lanqing Hu, Stephen Lin, Shuzhe Wu, Shiguang Shan

2024CVPR

Asymmetric Masked Distillation for Pre-Training Small Foundation Models

Zhiyu Zhao, Bingkun Huang, Sen Xing, Gangshan Wu, Yu Qiao, Limin Wang

2024ECCV

Data Collection-free Masked Video Modeling

Yuchi Ishikawa, Masayoshi Kondo, Yoshimitsu Aoki

2024ECCV

Text-Guided Video Masked Autoencoder

David Fan, Jue Wang, Shuai Liao, Zhikang Zhang, Vimal Bhat, Xinyu Li

2024BMVC

FILS: Self-Supervised Video Feature Prediction In Semantic Language Space

Mona Ahmadian, Frank Guerin, Andrew Gilbert

2024NeurIPS

Extending Video Masked Autoencoders to 128 Frames

Nitesh Bharadwaj Gundavarapu, Luke Friedman, Raghav Goyal, Chaitra Hegde, Eirikur Agustsson, Sagar M. Waghmare, Mikhail Sirotenko, Ming-Hsuan Yang, Tobias Weyand, Boqing Gong, Leonid Sigal

2024arXiv / Preprint

Scaling 4D Representations

João Carreira et al.

2024CVPR

VideoMAC: Video Masked Autoencoders Meet ConvNets

Gensheng Pei, Tao Chen, Xiruo Jiang, Huafeng Liu, Zeren Sun, Yazhou Yao

2024arXiv / Preprint

Self-supervised Video Object Segmentation with Distillation Learning of Deformable Attention

Quang-Trung Truong,Duc Thanh Nguyen, Binh-Son Hua, Sai-Kit Yeung

2024ECCV

Towards Latent Masked Image Modeling for Self-supervised Visual Representation Learning

Yibing Wei, Abhinav Gupta & Pedro Morgado

2024ECCV

SIGMA: Sinkhorn-Guided Masked Video Modeling

Mohammadreza Salehi, Michael Dorkenwald, Fida Mohammad Thoker, Efstratios Gavves, Cees G. M. Snoek & Yuki M. Asano

2024CVPR Workshops

ST2ST: Self-Supervised Test-time Adaptation for Video Action Recognition

Masud An-Nur Islam Fahim, Mohammed Innat, Jani Boutellier;

2024CogSci

Self-supervised learning of video representations from a child's perspective

A. Emin Orhan, Wentao Wang, Alex N. Wang, Mengye Ren, Brenden M. Lake;

2024ECCV

ViC-MAE: Self-supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders

Jefferson Hernandez, Ruben Villegas, Vicente Ordonez;

2024CVPR

Learning to Predict Activity Progress by Self-Supervised Video Alignment

Gerard Donahue, Ehsan Elhamifar;

2024Pattern Recognition

Repeat and learn: Self-supervised visual representations learning by Repeated Scene Localization

Yuanhang Zhang, Shuang Yang, Shiguang Shan, Xilin Chen;

2024CVPR

ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations

Yuanhang Zhang, Shuang Yang, Shiguang Shan, Xilin Chen;

2024WACV

Self-supervised Learning of Semantic Correspondence Using Web Videos

Donghyeon Kwon, Minsu Cho, Suha Kwak;

2024IPEC

Video Compression and Action Recognition in Self-supervised Learning

Zongbo Hao; Conghui Hao; Kecheng He

2024WACV

CycleCL: Self-supervised Learning for Periodic Videos

Matteo Destro, Michael Gygl

2024ICME Workshops

Self-Supervised Learning via Multi-Transformation Classification for Action Recognition

Duc-Quang Vu; Ngan Le; Jia-Ching Wang

2024Pattern Recognition

Motion-guided spatiotemporal multitask feature discrimination for self-supervised video representation learning

Shuai Bi, Zhengping Hu, Hehao Zhang, Jirui Di, Zhe Sun

2024CVPR

What When and Where? Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions

Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann, Samuel Thomas, Shih-Fu Chang, Rogerio Feris, James Glass, Hilde Kuehne

2024Applied Intelligence

Clustering-based multi-featured self-supervised learning for human activities and video retrieval

Muhammad Hafeez Javed, Zeng Yu, Taha M. Rajeh, Fahad Rafique & Tianrui Li

2024ICASSP Workshops

Positive and negative sampling strategies for self-supervised learning on audio-video data

Shanshan Wang, Soumya Tripathy, Toni Heittola, Annamaria Mesaros

2024AAAI

No More Shortcuts: Realizing the Potential of Temporal Self-Supervision

Ishan Rajendrakumar Dave, Simon Jenni, Mubarak Shah.

2024Pattern Recognition Letters

GLOCAL: A self-supervised learning framework for global and local motion estimation

Yihao Zheng , Kunming Luo , Shuaicheng Liu , Zun Li , Ye Xiang , Lifang Wu , Bing Zeng , Chang Wen Chen

2024IEEE TCSVT

Self-supervised Video Representation Learning via Capturing Semantic Changes Indicated by Saccades

Qiuxia Lai, Ailing Zeng, Ye Wang, Lihong Cao, Yu Li, Qiang Xu, IEEE

2024IEEE TMM

MAR: Masked Autoencoders for Efficient Action Recognition

Zhiwu Qing, Shiwei Zhang, Ziyuan Huang, Xiang Wang, Yuehuan Wang, Yiliang Lv, Changxin Gao, Nong Sang

2024CVPR

VicTR: Video-conditioned Text Representations for Activity Recognition

Kumara Kahatapitiya, Anurag Arnab, Arsha Nagrani, Michael S. Ryoo

2024IEEE Transactions on Cybernetics

Self-Supervised Video Representation Learning by Video Incoherence Detection

Haozhi Cao, Yuecong Xu, Kezhi Mao, Lihua Xie, Jianxiong Yin, Simon See, Qianwen Xu, and Jianfei Yang

2024ICLR

Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding

Xiong, Y., Zhao, L., Gong, B., Yang, M. H., Schroff, F., Liu, T., ... & Yuan, L.

2024BMVC

MotionMAE: Self-supervised Video Representation Learning with Motion-Aware Masked Autoencoders

Haosen Yang, Deng Huang, Bin Wen, Jiannan Wu, Hongxun Yao, Yi Jiang, Xiatian Zhu, Zehuan Yuan

2024ICML

EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens

Sunil Hwang, Jaehong Yoon, Youngwan Lee, Sung Ju Hwan

2024AAAI

XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning

Pritam Sarkar, Ali Etemad

2024Visual Intelligence

Controllable Augmentations for Video Representation Learning

Rui Qian, Weiyao Lin, John See, Dian Li

2023NeurIPS

Self-supervised object-centric learning for videos

Görkay Aydemir, Weidi Xie, Fatma Guney

2023NeurIPS

Language-based Action Concept Spaces Improve Video Self-Supervised Learning

Kanchana Ranasinghe, Michael S Ryoo

2023NeurIPS

Uncovering the Hidden Dynamics of Video Self-supervised Learning under Distribution Shifts

Pritam Sarkar, Ahmad Beirami, Ali Etemad

2023NeurIPS

Self-supervised video pretraining yields robust and more human-aligned visual representation

Nikhil Parthasarathy, S. M. Ali Eslami, João Carreira, Olivier J. Hénaff.

2023CVPR

AdaMAE: Adaptive Masking for Efficient Spatiotemporal Learning with Masked Autoencoders

Wele Gedara Chaminda Bandara, Naman Patel, Ali Gholami, Mehdi Nikkhah, Motilal Agrawal, Vishal M. Patel

2023ICCV

Spatio-Temporal Crop Aggregation for Video Representation Learning

Sepehr Sameni, Simon Jenni, Paolo Favaro

2023ICCV

Motion-Guided Masking for Spatiotemporal Representation Learning

David Fan, Jue Wang, Shuai Liao, Yi Zhu, Vimal Bhat, Hector Santos-Villalobos, Rohith MV, Xinyu Li

2023ACM Multimedia

Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation Learning

Minghao Zhu, Xiao Lin, Ronghao Dang, Chengju Liu, Qijun Chen

2023ICCV

Unmasked Teacher: Towards Training-Efficient Video Foundation Models

Kunchang Li, Yali Wang, Yizhuo Li, Yi Wang, Yinan He, Limin Wang, Yu Qiao

2023arXiv / Preprint

Concatenated Masked Autoencoders as Spatial-Temporal Learner

Zhouqiang Jiang, Bowen Wang, Tong Xiang, Zhaofeng Niu, Hong Tang, Guangshun Li, Liangzhi Li

2023ICTAI

AV-MaskEnhancer: Enhancing Video Representations through Audio-Visual Masked Autoencoder

Xingjian Diao, Ming Cheng, Shitong Cheng

2023CVPR

OmniMAE: Single Model Masked Pretraining on Images and Videos

Rohit Girdhar, Alaaeldin El-Nouby, Mannat Singh,Kalyan Vasudev Alwala, Armand Joulin , Ishan Misra

2023CVPR

TimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action Recognition

Ishan Rajendrakumar Dave, Mamshad Nayeem Rizve, Chen Chen, Mubarak Shah

2023Image and Vision Computing

Attentive spatial-temporal contrastive learning for self-supervised video representation

Xingming Yang, Sixuan Xiong, Kewei Wu, Dongfeng Shan, Zhao Xie

2023ICCV

MGMAE: Motion Guided Masking for Video Masked Autoencoding

Bingkun Huang, Zhiyu Zhao, Guozhen Zhang, Yu Qiao, Limin Wang

2023MVA

Cross-modal Manifold Cutmix for Self-supervised Video Representation Learning

Srijan Das; Michael Ryoo

2023ACM Multimedia

CHAIN: Exploring Global-Local Spatio-Temporal Information for Improved Self-Supervised Video Hashing

Rukai Wei, Yu Liu, Jingkuan Song, Heng Cui, Yanzhao Xie, Ke Zhou

2023ACM Multimedia

Data-Efficient Masked Video Modeling for Self-supervised Action Recognition

Qiankun Li, Xiaolong Huang, Zhifan Wan, Lanqing Hu, Shuzhe Wu, Jie Zhang, Shiguang Shan, Zengfu Wang(

2023IEEE Internet of Things Journal

Temporal Transformer Networks with Self-Supervision for Action Recognition

Yongkang Zhang, Jun Li, Guoming Wu, Han Zhang, Zhiping Shi, Member, IEEE, Zhaoxun Liu, Zizhang Wu

2023arXiv / Preprint

CMAE-V: Contrastive Masked Autoencoders for Video Action Recognition

Cheng-Ze Lu, Xiaojie Jin, Zhicheng Huang, Qibin Hou, Ming-Ming Cheng, Jiashi Feng

2023Computer Vision and Image Understanding

Learning Representational Invariances for Data-Efficient Action Recognition

Yuliang Zou, Jinwoo Choi, Qitong Wang, Jia-Bin Huang

2023Neurocomputing

SOR-TC: Self-attentive octave ResNet with temporal consistency for compressed video action recognition

Junsan Zhang, Xiaomin Wang, Yao Wan, Leiquan Wang, Jian Wang, Philip S. Yu

2023CVPR

Masked Motion Encoding for Self-Supervised Video Representation Learning

Xinyu Sun, Peihao Chen, Liangwei Chen, Thomas H. Li, Mingkui Tan, Chuang Gan

2023Signal, Image and Video Processing

Spatiotemporal consistency enhancement self-supervised representation learning for action recognition

Shuai Bi, Zhengping Hu, Mengyao Zhao, Shufang Li & Zhe Sun

2023IEEE TIP

Self-Supervised Video-Based Action Recognition With Disturbances

Wei Lin, Xinghao Ding, Yue Huang, Huanqiang Zeng

2023CVPR

Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning

Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Lu Yuan, Yu-Gang Jiang

2023Engineering Applications of Artificial Intelligence

Enhancing motion visual cues for self-supervised video representation learning

Mu Nie, Zhibin Quan, Weiping Ding, and Wankou Yang

2023Advanced Engineering Informatics

Continuous frame motion sensitive self-supervised collaborative network for video representation learning

Shuai Bi, Zhengping Hu, Mengyao Zhao, Hehao Zhang, Jirui Di, and Zhe Sun

2023Signal, Image and Video Processing

Self-supervised pretext task collaborative multi-view contrastive learning for video action recognition

Shuai Bi, Zhengping Hu, Mengyao Zhao, Hehao Zhang, Jirui Di, and Zhe Sun

2023IEEE TPAMI

Self-Supervised Learning from Untrimmed Videos via Hierarchical Consistency

Zhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yi Xu, Xiang Wang, Changxin Gao, Rong Jin, and Nong Sang

2023AAAI

Audio-Visual Contrastive Learning with Temporal Self-Supervision

Simon Jenni, Alexander Black, and John Collomosse

2023CVPR

Video Test-Time Adaptation for Action Recognition

Wei Lin, Muhammad Jehanzeb Mirza, Mateusz Kozinski, Horst Possegger, Hilde Kuehne, and Horst Bischof

2023AAAI

Self-Supervised Video Representation Learning via Latent Time Navigation

Di Yang, Yaohui Wang, Quan Kong, Antitza Dantcheva, Lorenzo Garattoni, Gianpiero Francesca, and Francois Bremond

2023ICASSP

Temporal Contrastive Learning with Curriculum

Shuvendu Roy and Ali Etemad

2023ICLR Workshops

Nearest-Neighbor Inter-Intra Contrastive Learning from Unlabeled Videos

David Fan, Deyu Yang, Xinyu Li, Vimal Bhat, and Rohith MV

2023ICCV

Tubelet-Contrastive Self-Supervision for Video-Efficient Generalization

Fida Mohammad Thoker, Hazel Doughty, and Cees Snoek

2023ICASSP

Multi-scale Compositional Constraints for Representation Learning on Videos

Georgios Paraskevopoulos, Chandrashekhar Lavania, Lovish Chum, and Shiva Sundaram

2023WACV

Flavr: Flow-agnostic Video Representations for Fast Frame Interpolation

Tarun Kalluri, Deepak Pathak, Manmohan Chandraker, and Du Tran

2023arXiv / Preprint

HomE: Homography-Equivariant Video Representation Learning

Anirudh Sriram, Adrien Gaidon, Jiajun Wu, Juan Carlos Niebles, Li Fei-Fei, and Ehsan Adeli

2023WACV

ViewCLR: Learning Self-supervised Video Representation for Unseen Viewpoints

Srijan Das and Michael S Ryoo

2023CVPR

Videomae v2: Scaling Video Masked Autoencoders with Dual Masking

Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao

2023AAAI

Self-Supervised Audio-Visual Representation Learning with Relaxed Cross-Modal Synchronicity

Pritam Sarkar, Ali Etemad

2023WACV

Previts: contrastive pretraining with video tracking supervision

Chen, B., Selvaraju, R. R., Chang, S. F., Niebles, J. C., & Naik, N.

2023CVPR

Modeling Video As Stochastic Processes for Fine-Grained Video Representation Learning

Zhang, H., Liu, D., Zheng, Q., & Su, B.

2023ICCV

Learning Fine-Grained Features for Pixel-wise Video Correspondences

Li, R., Zhou, S., & Liu, D.

2023CVPR Workshops

Cali-NCE: Boosting Cross-Modal Video Representation Learning With Calibrated Alignment

Zhao, N., Jiao, J., Xie, W., & Lin, D.

2023IEEE TNNLS

Self-supervised motion perception for spatiotemporal representation learning

Chang Liu, Yuan Yao, Dezhao Luo, Yu Zhou, Qixiang Ye

2023Machine Vision and Applications

Similarity Contrastive Estimation for Image and Video Soft Contrastive Self-Supervised Learning

Julien Denize, Jaonary Rabarisoa, Astrid Orcesi, Romain H´erault

2023ICIP

Self-Supervised Contrastive Learning for Audio-Visual Action Recognition

Yang Liu, Ying Tan, Haoyuan Lan

2023IEEE TMM

Self-Supervised Scene-Debiasing for Video Representation Learning via Background Patching

Maregu Assefa, Wei Jiang, Kumie Gedamu, Getinet Yilma, Bulbula Kumeda, Melese Ayalew

2023IEEE TMM

LgNet: A local-global network for action recognition and beyond

Jiaqi Zhou, Zehua Fu, Qiuyu Huang, Qingjie Liu, Yunhong Wang

2023IEEE TCSVT

Unsupervised Video-Based Action Recognition With Imagining Motion and Perceiving Appearance

Wei Lin , Xiaoyu Liu , Yihong Zhuang , Xinghao Ding , Xiaotong Tu , Yue Huang , Huanqiang Zeng

2023AAAI

Spatiotemporal Augmentation on Selective Frequencies for Video Representation Learning

Jinhyung Kim, Taeoh Kim, Minho Shim, Dongyoon Han, Dongyoon Wee, Junmo Kim

2023IEEE TCSVT

Consistent Intra-video Contrastive Learning with Asynchronous Long-term Memory Bank

Zelin Chen, Kun-Yu Lin, Wei-Shi Zheng

2022CVPR

BEVT: BERT Pretraining of Video Transformers

Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, Lu Yuan

2022NeurIPS

Masked Autoencoders As Spatiotemporal Learners

Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming He

2022CVPR

SPAct: Self-supervised Privacy Preservation for Action Recognition

Ishan Rajendrakumar Dave, Chen Chen, Mubarak Shah

2022AAAI

Suppressing Static Visual Cues via Normalizing Flows for Self-Supervised Video Representation Learning

Manlin Zhang, Jinpeng Wang, Andy J. Ma

2022CVPR

Self-supervised Video Transformer

Kanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan, Michael S. Ryoo

2022ACM TOMM

Exploring Relations in Untrimmed Videos for Self-Supervised Learning

Dezhao Luo, Bo Fang, Yu Zhou, Yucan Zhou, Dayan Wu, Weiping Wang

2022ACM Multimedia

MaMiCo: Macro-to-Micro Semantic Correspondence for Self-supervised Video Representation Learning

Bo Fang, Wenhao Wu, Chang Liu, Yu Zhou, Dongliang He, Weiping Wang

2022IEEE TIP

TCGL: Temporal Contrastive Graph for Self-Supervised Video Representation Learning

Yang Liu , Keze Wang , Lingbo Liu , Haoyuan Lan, and Liang Lin

2022CVPR

Cross-Architecture Self-supervised Video Representation Learning

Sheng Guo, Zihua Xiong, Yujie Zhong, Limin Wang, Xiaobo Guo, Bing Han, Weilin Huang

2022AAAI

Contrastive spatio-temporal pretext learning for self-supervised video representation

Yujia Zhang, Lai-Man Po, Xuyuan Xu, Mengyang Liu, Yexin Wang, Weifeng Ou, Yuzhi Zhao, Wing-Yin Yu

2022CVPR

Transrank: Self-supervised video representation learning via ranking-based transformation recognition

Haodong Duan, Nanxuan Zhao, Kai Chen, Dahua Lin

2022CVPR

Learning from untrimmed videos: Self-supervised video representation learning with hierarchical consistency

Zhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yi Xu, Xiang Wang, Mingqian Tang, Changxin Gao, Rong Jin,Nong Sang

2022CVPR

Motion-aware contrastive video representation learning via foreground-background merging

Shuangrui Ding, Maomao Li, Tianyu Yang, Rui Qian, Haohang Xu, Qingyi Chen, Jue Wang, Hongkai Xiong

2022ICME

Self-Supervised Video Representation Learning with Motion-Contrastive Perception

Jinyu Liu, Ying Cheng, Yuejie Zhang, Rui-Wei Zhao, Rui Feng

2022IEEE TCSVT

Self-supervised video representation learning using improved instance-wise contrastive learning and deep clustering

Yisheng Zhu, Hui Shuai, Guangcan Liu, Senior Member, Qingshan Liu

2022Computer Vision and Image Understanding

TCLR: Temporal contrastive learning for video representation

Ishan Dave, Rohit Gupta, Mamshad Nayeem Rizve, Mubarak Shah

2022AAAI

Self-supervised spatiotemporal representation learning by exploiting video continuity

Hanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen, Peng Dai, Juwei Lu, Yang Wang

2022CVPR

Probabilistic representations for video contrastive learning

Jungin Park, Jiyoung Lee, Ig-Jae Kim, Kwanghoon Sohn

2022CVPR

Contextualized spatio-temporal contrastive learning with self-supervision

Liangzhe Yuan, Rui Qian, Yin Cui, Boqing Gong,Florian Schroff,Ming-Hsuan Yang, Hartwig Adam, Ting Liu

2022NeurIPS

VideoMAE: Masked Autoencoders Are Data-Efficient Learners for Self-Supervised Video Pre-Training

Zhan Tong, Yibing Song, Jue Wang, Limin Wang

2022WACV

Self-supervised video representation learning with cross-stream prototypical contrasting

Martine Toering, Ioannis Gatopoulos, Maarten Stol, Vincent Tao Hu

2022CVPR

SLIC: Self-supervised learning with iterative clustering for human action videos

Salar Hosseini Khorasgani, Yuxuan Chen, Florian Shkurti

2022ECCV

GOCA: guided online cluster assignment for self-supervised video representation Learning

Huseyin Coskun, Alireza Zareian, Joshua L. Moore, Federico Tombari, Chen Wang

2022ACCV

TCVM: Temporal Contrasting Video Montage Framework for Self-supervised Video Representation Learning

Fengrui Tian, Jiawei Fan, Xie Yu, Shaoyi Du, Meina Song, Yu Zhao

2022ECCV

Static and Dynamic Concepts for Self-supervised Video Representation Learning

Rui Qian, Shuangrui Ding, Xian Liu, Dahua Lin

2022ECCV

SOS! Self-supervised Learning over Sets of Handled Objects in Egocentric Action Recognition

Victor Escorcia, Ricardo Guerrero, Xiatian Zhu, Brais Martinez

2022CVPR Workshops

Self-Supervised Video Representation Learning with Cascade Positive Retrieval

Cheng-En Wu, Farley Lai, Yu Hen Hu, Asim Kadav

2022IEEE JSTSP

Self-Supervised Learning of Audio Representations From Audio-Visual Data Using Spatial Alignment

Shanshan Wang, Archontis Politis, Annamaria Mesaros

2022WACV

Hierarchically decoupled spatial-temporal contrast for self-supervised video representation learning

Zehua Zhang, David Crandall

2022ICME

Spatio-temporal self-supervision enhanced transformer networks for action recognition

Yongkang Zhang, Han Zhang, Guoming Wu, Jun Li

2022ICPR

Inter-Intra Cross-Modality Self-Supervised Video Representation Learning by Contrastive Clustering

Jiutong Wei. Guan Luo, Bing Li, Weiming Hu

2022CVPR Workshops

SCVRL: Shuffled Contrastive Video Representation Learning

Michael Dorkenwald, Fanyi Xiao, Biagio Brattoli, Joseph Tighe, Davide Modolo

2022arXiv / Preprint

InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang,Jilan Xu, Yi Liu, Zun Wang, Sen Xing, Guo Chen, Junting Pan, Jiashuo Yu,Yali Wang, Limin Wang, Yu Qiao

2022ICANN

Video Motion Perception for Self-supervised Representation Learning

Wei Li, Dezhao Luo, Bo Fang, Xiaoni Li, Yu Zhou, Weiping Wang

2022IEEE TCSVT

An improved inter-intra contrastive learning framework on self-supervised video representation

Li Tao, Xueting Wang, Toshihiko Yamasaki

2022CVPR Workshops

Auxiliary Learning for Self-Supervised Video Representation via Similarity-based Knowledge Distillation

Amirhossein Dadashzadeh, Alan Whone, Majid Mirmehdi

2022ECCV

Motion Sensitive Contrastive Learning for Self-supervised Video Representation

Jingcheng Ni, Nan Zhou, Jie Qin, Qian Wu, Junqi Liu, Boxun Li, Di Huang

2022NCC

Unsupervised Learning of Spatio-Temporal Representation with Multi-Task Learning for Video Retrieval

Vidit Kumar

2022ECCV

Federated Self-supervised Learning for Video Understanding

Yasar Abbas Ur Rehman, Yan Gao, Jiajun Shen, Pedro Porto Buarque de Gusmão , Nicholas Lane

2022Neurocomputing

Contrastive predictive coding with transformer for video representation learning

Yue Liu, Junqi Ma, Yufei Xie, Xuefeng Yang, Xingzhen Tao, Lin Peng, Wei Gao

2022Applied Intelligence

Video representation learning by identifying spatio-temporal transformation

Sheng Geng, Shimin Zhao , Hu Liu

2022BMVC

On temporal granularity in self-supervised video representation learning

Rui Qian, Yeqing Li, Liangzhe Yuan, Boqing Gong, Ting Liu, Matthew Brown, Serge Belongie, Ming-Hsuan Yang, Hartwig Adam, and Yin Cui

2022ICML Workshops

LAVA: Language Audio Vision Alignment for Data-Efficient Video Pre-Training

Sumanth Gurram , Andy Fang , David Chan , John Canny

2022arXiv / Preprint

It Takes Two: Masked Appearance-Motion Modeling for Self-supervised Video Transformer Pre-training

Yuxin Song, Min Yang, Wenhao Wu, Dongliang He, Fu Li, Jingdong Wang

2022BMVC

MAC: Mask-Augmentation for Motion-Aware Video Representation Learning

Arif Akar, Ufuk Umut Senturk, and Nazli Ikizler-Cinbis.

2022AVSS

Temporal-Invariant Video Representation Learning with Dynamic Temporal Resolutions.

Seong-Yun Jeong, Ho-Joong Kim, Myeong-Seok Oh, Gun-Hee Lee, Seong-Whan Lee

2022ACM Multimedia

Dual Contrastive Learning for Spatio-temporal Representation

Shuangrui Ding,, Rui Qian, and Hongkai Xiongo

2022ECCV Workshops

MoQuad: Motion-focused Quadruple Construction for Video Contrastive Learning

Yuan Liu, Jiacheng Chen, Hao Wu

2022arXiv / Preprint

On Negative Sampling for Audio-Visual Contrastive Learning from Movies

Mahdi M. Kalayeh, Shervin Ardeshir, Lingyi Liu, Nagendra Kamath, Ashok Chandrashekar

2022CVPR

Frame-wise Action Representations for Long Videos via Sequence Contrastive Learning

Minghao Chen, Fangyun Wei, Chong Li, Deng Cai

2022CVPR

Masked Feature Prediction for Self-Supervised Visual Pre-Training

Wei, C., Fan, H., Xie, S., Wu, C. Y., Yuille, A., & Feichtenhofer, C.

2022arXiv / Preprint

Pixel-level Correspondence for Self-Supervised Learning from Video

Yash Sharma, Yanchao Zhu, Chris Russell, Thomas Brox

2022CVPR

Temporal Alignment Networks for Long-Term Video

Han, T., Xie, W., & Zisserman, A.

2022arXiv / Preprint

SimVTP: Simple Video Text Pre-Training with Masked Autoencoders

Ma, Y., Yang, T., Shan, Y., & Li, X.

2022ICLR

Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Shi, B., Hsu, W. N., Lakhotia, K., & Mohamed, A.

2022IEEE TPAMI

Self-supervised video representation learning by uncovering spatio-temporal statistics

Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Wei Liu, Yun-hui Liu

2021BMVC

Inter-intra Variant Dual Representations for Self-supervised Video Recognition

Lin Zhang, Qi She, Zhengyang Shen, Changhu Wang

2021NeurIPS

VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning

Hao Tan, Jie Lei, Thomas Wolf, Mohit Bansal

2021arXiv / Preprint

Watching too much television is good: Self-supervised audio-visual representation learning from movies and tv shows

Mahdi M. Kalayeh, Nagendra Kamath, Lingyi Liu

2021ICPR

Temporally coherent embeddings for self-supervised video representation learning

Joshua Knights, Ben Harwood, Daniel Ward, Anthony Vanderkop, Olivia Mackenzie-Ross, Peyman Moghadam

2021CVPR

Audio-visual instance discrimination with cross-modal agreement

Pedro Morgado, Nuno Vasconcelos, Ishan Misra

2021CVPR

Removing the background by adding the background: Towards background robust self-supervised video representation learning

Jinpeng Wang, Yuting Gao, Ke Li, Yiqi Lin, Andy J. Ma, Hao Cheng, Pai Peng, Feiyue Huang, Rongrong Ji, Xing Sun

2021AAAI

Enhancing unsupervised video representation learning by decoupling the scene and the motion

Jinpeng Wang, Yuting Gao, Ke Li, Jianguo Hu, Xinyang Jiang, Xiaowei Guo, Rongrong Ji, Xing Sun

2021AAAI

SeCo: Exploring Sequence Supervision for Unsupervised Representation Learning

Ting Yao, Yiheng Zhang, Zhaofan Qiu, Yingwei Pan, Tao Mei

2021ICCV

Enhancing self-supervised video representation learning via multi-level feature optimization

Rui Qian, Yuxi Li, Huabin Liu, John See, Shuangrui Ding, Xian Liu, Dian Li, Weiyao Lin

2021AAAI

RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning

Peihao Chen, Deng Huang, Dongliang He, Xiang Long, Runhao Zeng, Shilei Wen, Mingkui Tan, Chuang Gan

2021CVPR

VideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples

Tian Pan, Yibing Song, Tianyu Yang, Wenhao Jiang, Wei Liu

2021ICCV

On compositions of transformations in contrastive self-supervised learning

Mandela Patrick, Yuki M. Asano, Polina Kuznetsova, Ruth Fong, João F. Henriques, Geoffrey Zweig, Andrea Vedaldi

2021CVPR

Unsupervised visual representation learning by tracking patches in video

Guangting Wang, Yizhou Zhou, Chong Luo, Wenxuan Xie, Wenjun Zeng, Zhiwei Xiong

2021CVPR

A large-scale study on unsupervised spatiotemporal representation learning

Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, Kaiming He

2021CVPR

CoCon: Cooperative-Contrastive Learning

Nishant Rai, Ehsan Adeli ,Kuan-Hui Lee, Adrien Gaidon, Juan Carlos Niebles

2021NeurIPS

VATT: Transformers for multimodal self-supervised learning from raw video, audio and text

Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, Boqing Gong

2021ICCV

ASCNet: Self-supervised video representation learning with appearance-speed consistency

Deng Huang, Wenhao Wu, Weiwen Hu, Xu Liu, Dongliang He, Zhihua Wu, Xiangmiao Wu, Mingkui Tan, Errui Ding

2021IEEE Access

Self-supervised visual learning by variable playback speeds prediction of a video

Hyeon Cho, Taehoon Kim, Hyungjin Chang, Wonjun Hwang

2021ICCV

Self-supervised video representation learning with meta-contrastive network

Yuanze Lin, Xun Guo, Yan Lu

2021ICCV

Long short view feature decomposition via contrastive video representation learning

Nadine Behrmann, Mohsen Fayyaz, Juergen Gall, Mehdi Noroozi

2021ICCV

Time-equivariant contrastive video representation learning

Simon Jenni, Hailin Jin

2021CVPR

Self-supervised video representation learning by context and motion decoupling

Lianghua Huang, Yu Liu, Bin Wang, Pan Pan, Yinghui Xu, Rong Jin

2021WACV

Unsupervised video representation learning by bidirectional feature prediction

Nadine Behrmann, Juergen Gall, Mehdi Noroozi

2021ICLR

Self-supervised learning of compressed video representations

Youngjae Yu, Sangho Lee, Gunhee Kim, Yale Song

2021CVPR

Spatiotemporal contrastive video representation learning

Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, Yin Cui

2021arXiv / Preprint

MoDist: Motion Distillation for Self-Supervised Video Representation Learning

Fanyi Xiao, Joseph Tighe, Davide Modolo

2021ICCV

Broaden your views for self-supervised video learning

Adria Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang, Ross Hemsley, Florian Strub, Corentin Tallec, Mateusz Malinowski, Viorica Patraucean, Florent Altche, Michal Valko, Jean-Bastien Grill, Aaron van den Oord, Andrew Zisserman

2021ICCV

Vi2CLR: Video and image for visual contrastive learning of representation

Ali Diba, Vivek Sharma, Reza Safdari, Dariush Lotfi, M. Saquib Sarfraz,Rainer Stiefelhagen, Luc Van Gool,

2021ICCV

Contrast and order representations for video self-supervised learning

Kai Hu, Jie Shao, Yuan Liu, Bhiksha Raj, Marios Savvides, Zhiqiang Shen

2021ICCV

Motion-augmented self-training for video recognition at smaller scale

Kirill Gavrilyuk, Mihir Jain, Ilia Karmanov, Cees G. M. Snoek

2021ICCV Workshops

Video contrastive learning with global context

Haofei Kuang, Yi Zhu, Zhi Zhang, Xinyu Li, Joseph Tighe,Soren Schwertfeger, Cyrill Stachniss, Mu Li

2021ICCV

Motion-focused contrastive learning of video representations

Rui Li, Yiheng Zhang, Zhaofan Qiu, Ting Yao, Dong Liu, and Tao Mei

2021BMVC

Back to the Future: Cycle Encoding Prediction for Self-supervised Video Representation Learning

Xinyu Yang, Majid Mirmehdi,Tilo Burghardt

2021ICCV

Composable augmentation encoding for video representation learning

Sun, C., Nagrani, A., Tian, Y., & Schmid, C.

2021ICCV

Learning temporal dynamics from cycles in narrated video

Epstein, D., Wu, J., Schmid, C., & Sun, C.

2021ICCV

CrossCLR: Cross-Modal Contrastive Learning for Multi-Modal Video Representations

Zolfaghari, M., Zhu, Y., Gehler, P., & Brox, T.

2021ICLR

Watching the World Go By: Representation Learning from Unlabeled Videos

Daniel Gordon, Kiana Ehsani, Dieter Fox, Ali Farhadi

2021ICLR

Parameter Efficient Multimodal Transformers for Video Representation Learning

Lee, S., Yu, Y., Kim, G., Breuel, T., Kautz, J., & Song, Y.

2021ICLR

Active Contrastive Learning of Audio-Visual Video Representations

Ma, S., Zeng, Z., McDuff, D., & Song, Y.

2020Sensors

Self-Supervised Learning to Detect Key Frames in Videos

Xiang Yan,Syed Zulqarnain Gilani,Mingtao Feng ,Liang Zhang,Hanlin Qin and Ajmal Mian

2020ECCV

Self-supervised motion representation via scattering local motion cues

Yuan Tian, Zhaohui Che, Wenbo Bao, Guangtao Zhai, Zhiyong Gao1

2020ACM Multimedia

Self-supervised video representation learning using inter-intra contrastive framework

Li Tao, Xueting Wang, Toshihiko Yamasaki

2020arXiv / Preprint

Video representation learning with visual tempo consistency

Ceyuan Yang, Yinghao Xu, Bo Dai, Bolei Zhou

2020arXiv / Preprint

Self-supervised temporal discriminative learning for video representation learning

Jinpeng Wang, Yiqi Lin, Andy J. Ma,Pong C. Yuen

2020NeurIPS

Self-supervised learning by cross-modal audio-video clustering

Humam Alwassel, Dhruv Mahajan, Bruno Korbar ,Lorenzo Torresani, Bernard Ghanem, Du Tran

2020ECCV

Self-supervised video representation learning by pace prediction

Jiangliu Wang, Jianbo Jiao, Yun-Hui Liu

2020CVPR

Unsupervised learning from video with deep neural embeddings

Chengxu Zhuang, Tianwei She, Alex Andonian, Max Sobol Mark, Daniel Yamins

2020ECCV Workshops

Unsupervised learning of video representations via dense trajectory clustering

Pavel Tokmakov, Martial Hebert, Cordelia Schmid

2020ECCV

Video representation learning by recognizing temporal transformations

Simon Jenni, Givi Meishvili, Paolo Favaro

2020CVPR

Video playback rate perception for self-supervised spatio-temporal representation learning

Yuan Yao, Chang Liu, Dezhao Luo, Yu Zhou, Qixiang Ye

2020NeurIPS

Self-supervised co-training for video representation learning

Tengda Han, Weidi Xie, Andrew Zisserman

2020AAAI

Video cloze procedure for self-supervised spatio-temporal learning

Dezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang, Can Ma, Qixiang Ye, Weiping Wang

2020CVPR

End-to-end learning of visual representations from uncurated instructional videos

Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira,Ivan Laptev, Josef Sivic, Andrew Zisserman

2020CVPR

SpeedNet: Learning the Speediness in Videos

Sagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri, William T. Freeman, Michael Rubinstein, Michal Irani, Tali Dekel

2020ECCV

Contrastive multiview coding

Yonglong Tian, Dilip Krishnan, Phillip Isola

2020Signal Processing: Image Communication

Self-supervised video representation learning by maximizing mutual information

Fei Xue, Hongbing Ji, Wenbo Zhang, Yi Cao

2020ECCV

Memory-augmented dense predictive coding for video representation learning

Tengda Han, Weidi Xie, Andrew Zisserman

2020CVPR

Evolving losses for unsupervised video representation learning

AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo

2020arXiv / Preprint

AudioVisual SlowFast Networks for Video Recognition

Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, Christoph Feichtenhofer

2020NeurIPS

Cycle-Contrast for Self-Supervised Video Representation Learning

Quan Kong, Wenpeng Wei, Ziwei Deng, Tomoaki Yoshinaga, Tomokazu Murakami

2020arXiv / Preprint

Can temporal information help with contrastive self-supervised learning?

Yutong Bai, Haoqi Fan, Ishan Misra, Ganesh Venkatesh, Yongyi Lu

2020NeurIPS

Self-supervised multimodal versatile networks

Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira Sander Dieleman, Andrew Zisserman

2020arXiv / Preprint

Pretext-Contrastive Learning: Toward Good Practices in Self-Supervised Video Representation Learning

Li Tao, Xueting Wang, Toshihiko Yamasaki

2020arXiv / Preprint

UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Luo, H., Ji, L., Shi, B., Huang, H., Duan, N., Li, T., ... & Zhou, M.

2020ECCV

Self-supervised learning of audio-visual objects from video

Afouras, T., Owens, A., Chung, J. S., & Zisserman, A.

2020CVPR

Speech2Action: Cross-Modal Supervision for Action Recognition

Nagrani, A., Sun, C., Ross, D., Sukthankar, R., Schmid, C., & Zisserman, A.

2020ACM Multimedia

Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation Learning

Cheng, Y., Wang, R., Pan, Z., Feng, R., & Zhang, Y.

2019CVPR

Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics

Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Yunhui Liu, Wei Liu

2019ICCV Workshops

Video representation learning by dense predictive coding

Tengda Han, Weidi Xie, Andrew Zisserman

2019CVPR

Self-supervised spatiotemporal learning via video clip order prediction

Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, Yueting Zhuang

2019WACV

Video Jigsaw: Unsupervised Learning of Spatiotemporal Context for Video Action Recognition

Unaiza Ahsan, Rishi Madhok, Irfan Essa

2019AAAI

Self-supervised video representation learning with space-time cubic puzzles

Dahun Kim, Donghyeon Cho, In So Kweon

2019arXiv / Preprint

Learning Video Representations Using Contrastive Bidirectional Transformer

Chen Sun, Fabien Baradel, Kevin Murphy, Cordelia Schmid

2019ICCV

DynamoNet: Dynamic Action and Motion Network

Ali Diba, Vivek Sharma, Luc Van Gool, Rainer Stiefelhagen

2019CVPR

Temporal Cycle-Consistency Learning

Dwibedi, D., Aytar, Y., Tompson, J., Sermanet, P., & Zisserman, A.

2019ICCV

VideoBERT: A Joint Model for Video and Language Representation Learning

Sun, C., Myers, A., Vondrick, C., Murphy, K., & Schmid, C.

2018CVPR

Geometry Guided Convolutional Neural Networks for Self-Supervised Video Representation Learning

Chuang Gan, Boqing Gong, Kun Liu, Hao Su, Leonidas J. Guibas

2018arXiv / Preprint

Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction

Longlong Jing, Xiaodong Yang, Jinggen Liu, Yingli Tian

2018NeurIPS

Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization

Bruno Korbar, Du Tran, Lorenzo Torresani

2018ECCV

Audio-Visual Scene Analysis with Self-Supervised Multisensory Features

Andrew Owens, Alexei A. Efros

2018CVPR

Compressed Video Action Recognition

Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu, R. Manmatha, Alexander J. Smola, Philipp Krahenb

2018ECCV

Improving Spatiotemporal Self-Supervision by Deep Reinforcement Learning

Uta Buchler, Biagio Brattoli, Bjorn Ommer

2018CVPR

Learning and Using the Arrow of Time

Donglai Wei, Joseph Lim, Andrew Zisserman, William T. Freeman

2017ICCV

Unsupervised Representation Learning by Sorting Sequences

Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, Ming-Hsuan Yang

2017CVPR

Self-Supervised Video Representation Learning With Odd-One-Out Networks

Basura Fernando, Hakan Bilen, Efstratios Gavves, Stephen Gould

2016ECCV

Shuffle and Learn: Unsupervised Learning Using Temporal Order Verification

Ishan Misra, C. Lawrence Zitnick, Martial Hebert

Surveys and evaluation

Research resources

Survey and benchmarking work that helps place individual methods in context.

2024 · Preprint

Unifying Video Self-Supervised Learning across Families of Tasks: A Survey

Ishan Dave*, Malitha Gunawardhana*, Limalka Sadith, Honglu Zhou, Liel David, Daniel Harari, Mubarak Shah, Muhammad Hairs Khan

2022 · ACM Computing Surveys

Self-Supervised Learning for Videos: A Survey

Madeline C. Schiappa, Yogesh S. Rawat, And Mubarak Shah

2024 · arXiv preprint

SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning

Fida Mohammad Thoker, Letian Jiang, Chen Zhao, Piyush Bagad, Hazel Doughty, Bernard Ghanem, Cees G. M. Snoek

2024 · arXiv preprint

How Effective are Self-Supervised Models for Contact Identification in Videos

Malitha Gunawardhana, Limalka Sadith, Liel David, Daniel Harari, Muhammad Haris Khan

2023 · arXiv preprint arXiv:2306.06010

Benchmarking self-supervised video representation learning

Akash Kumar, Ashlesha Kumar, Vibhav Vineet, Yogesh Singh Rawat

2023 · arXiv preprint arXiv:2303.13505

A Large-scale Study of Spatiotemporal Representation Learning with a New Benchmark on Action Recognition

Deng, A., Yang, T., & Chen, C.

2022, October · In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022

How Severe Is Benchmark-Sensitivity in Video Self-supervised Learning?

Fida Mohammad Thoker, Hazel Doughty, Piyush Bagad, Cees Snoek

Field guide

Video self-supervised learning FAQ

What is video self-supervised learning?

Video self-supervised learning learns useful spatial and temporal representations from videos without requiring a manually annotated label for every training example. Common objectives include masked reconstruction, contrastive learning, temporal prediction, motion modeling and cross-modal learning.

What do Video SSL, VideoSSL and SSL video mean?

Video SSL and VideoSSL are common abbreviations for video self-supervised learning. SSL video is another search phrasing for the same research area, which is also described as self-supervised video representation learning or video representation pretraining.

Which publication metadata is shown?

Each catalog entry displays its confirmed publication year and latest verified venue, alongside the title, authors and research links.

How are later conference or journal publications handled?

When a preprint is later published at a conference or journal, the public entry is updated to the latest confirmed venue while the supporting audit record remains in the repository.