2026CVPR
Progressive Mask Distillation for Self-supervised Video Representation
Kewei Wu, Chong Liang, Zhao Xie, Dan Guo
2026CVPR
TrackMAE: Video Representation Learning via Track Mask and Predict
Renaud Vandeghen, Fida Mohammad Thoker, Marc Van Droogenbroeck, Bernard Ghanem
2026arXiv / Preprint
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
Lorenzo Mur-Labadia, Matthew Muckley, Amir Bar, Mido Assran, Koustuv Sinha, Mike Rabbat, Yann LeCun, Nicolas Ballas, Adrien Bardes
2026CVPR
From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning
Yang Liu, Qianqian Xu, Peisong Wen, Siran Dai, Xilin Zhao, Qingming Huang
2026arXiv / Preprint
The TIME Machine: On The Power of Motion for Efficient Perception
Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara
2026arXiv / Preprint
TrAction: Action Recognition with Sparse Trajectories
Jan F. Meier, Felix B. Mueller, Alexander Ecker, Timo Lüddecke
2026arXiv / Preprint
OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence
Feilong Tang, Xiang An, Yunyao Yan, Yin Xie, Bin Qin, Kaicheng Yang, Yifei Shen, Yuanhan Zhang, Chunyuan Li, Shikun Feng, Changrui Chen, Huajie Tan, Ming Hu, Manyuan Zhang, Bo Li, Ziyong Feng, Ziwei Liu, Zongyuan Ge, Jiankang Deng
2026arXiv / Preprint
Factorized Latent Dynamics for Video JEPA: An Empirical Study of Auxiliary Objectives
Santosh Premi
2026arXiv / Preprint
Self-Supervised Learning of Structured Dynamics from Videos
Lukas Knobel, Andrew Zisserman, Yuki M. Asano
2026arXiv / Preprint
Depth-Wise Representation Development Under Blockwise Self-Supervised Learning for Video Vision Transformers
Jonas Römer, Timo Dickscheid
2026Engineering Applications of Artificial Intelligence
Beyond reconstruction: Enhancing masked autoencoders with contrastive learning for video representation learning
Yawei Feng, Lijun Guo, Guitao Yu, Rong Zhang, Jiangbo Qian, Chong Wang, Shangce Gao
2026ECCV
Structured-Noise Masked Modeling for Video, Audio and Beyond
Aritra Bhowmik, Fida Mohammad Thoker, Carlos Hinojosa, Bernard Ghanem, Cees G. M. Snoek
2026ICLR
Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
Xianhang Li, Chen Huang, Chun-Liang Li, Eran Malach, Josh Susskind, Vimal Thilak, Etai Littwin
2026CVPR
Recurrent Video Masked Autoencoders
Daniel Zoran, Nikhil Parthasarathy, Yi Yang, Drew A. Hudson, Joao Carreira, Andrew Zisserman
2026ICLR
Dual Perspectives on Non-Contrastive Self-Supervised Learning
Jean Ponce, Martial Hebert, Basile Terver ;
2026International Journal of Computer Vision
Self-Supervised Video Representation Learning in a Heuristic Decoupled Perspective
Zeen Song, Jingyao Wang, Jianqi Zhang, Changwen Zheng, Wenwen Qiang
2026IEEE TCSVT
BIMM: Brain Inspired Masked Modeling for Video Representation Learning
Zhifan Wan, Jie Zhang, Changzhen Li, Shiguang Shan
2025CVPR Workshops
Efficient VideoMAE via Temporal Progressive Training
Xianhang Li, Peng Wang, Xinyu Li, Heng Wang, Hongru Zhu, Cihang Xie
2025ICCV
An Empirical Study of Autoregressive Pre-training from Videos
Jathushan Rajasegaran, Ilija Radosavovic, Rahul Ravishankar, Yossi Gandelsman, Christoph Feichtenhofer, Jitendra Malik
2025ICCV Workshops
Reinforcement Learning Meets Masked Video Modeling: Trajectory-Guided Adaptive Token Selection
Ayush K. Rai, Kyle Min, Tarun Krishna, Feiyan Hu, Alan F. Smeaton, Noel E. O'Connor
2025ACROSET
Entropy-Guided Masked Autoencoding for Self-Supervised Human Action Recognition Using Video Swin Transformer
Kollu Praveen Kumar; Guduri Baby Harshitha; Koti Vijay; Angothu Sravika
2025ICCV Workshops
Privacy Preservation Using Superimposed 3D-Models for Self-Supervised Training in Action Recognition
Asfandyar Azhar, Nidhish Shah, Shaurjya Mandal, Yongjie Jessica Zhang;
2025Machine Learning
Learning Complementary Knowledge via Trusted Multi-view Space Decomposition for Self-Supervised Contrastive Learning
Jiangmeng Li, Yunze Zhao, Yifan Jin, Changwen Zheng & Wenwen Qiang;
2025NeurIPS
OSKAR: Omnimodal Self-supervised Knowledge Abstraction and Representation
Mohamed O Abdelfattah, Kaouther Messaoud, Alexandre Alahi;
2025arXiv / Preprint
MME: Video Representation Learning as World Model for Understanding and Planning
Xinyu Sun, Changhao Li, Chen Jian, Chuang Gan, Peihao Chen, and Mingkui Tan;
2025ICCV Workshops
Hashtag2Action: Data Engineering and Self-Supervised Pre-Training for Action Recognition in Short-Form Videos
Yang Qian, Ali Kargarandehkordi, Yinan Sun, Parnian Azizian, Onur Cezmi Mutlu, Saimourya Surabhi, Zain Jabbar, Dennis Wall, Peter Washington, Huaijin Chen;
2025Cluster Computing
Kdhiera: boosting self-supervised masked video modeling via hierarchical knowledge distillation
Yunlong Wang, Hong Liang, Mingwen Shao & Qian Zhang;
2025Proceedings of SPIE
Self-supervised video representation learning based on foreground and temporal information.
Zhongliang Zhou, Jiayong Fang ;
2025IJCV
Feature Hallucination for Self-supervised Action Recognition
Lei Wang, Piotr Koniusz; ;
2025CVPR Workshops
ViDROP: Video Dense Representation through Spatio-Temporal Sparsity
Sepehr Sameni, Simon Jenni, Paolo Favaro; ;
2025CVPR
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
Yangliu Hu, Zikai Song, Na Feng, Yawei Luo, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang; ;
2025arXiv / Preprint
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba, Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, Xiaodong Ma, Sarath Chandar, Franziska Meier, Yann LeCun, Michael Rabbat, Nicolas Ballas ;
2025CVPR
When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning
Yang Liu, Qianqian Xu, Peisong Wen, Siran Dai, Qingming Huang;
2025NeurIPS
Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals
Stefan Stojanov, David Wendt, Seungwoo Kim, Rahul Venkatesh, Kevin Feigelis, Jiajun Wu, Daniel LK Yamins
2025ICMR
Label Ranker: Self-aware Preference for Classification Label Position in Visual Masked Self-supervised Pre-trained Model
Peihao Xiang, Ou Bai
2025CVPR
AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing
Niu Lian, Jun Li, Jinpeng Wang, Ruisheng Luo, Yaowei Wang, Shu-Tao Xia, Bin Chen
2025AAAI
Efficient Self-Supervised Video Hashing with Selective State Spaces
Jinpeng Wang, Niu Lian, Jun Li, Yuting Wang, Yan Feng, Bin Chen, Yongbing Zhang, Shu-Tao Xia1
2025Image and Vision Computing
Exemplar-free class incremental action recognition based on self-supervised learning
Chunyu Hou, Yonghong Hou, Jinyin Jiang, Gunel Abdullayeva
2025CVPR
Learning from Streaming Video with Orthogonal Gradients
Tengda Han⋄, Dilara Gokay, Joseph Heyward, Chuhan Zhang, Daniel Zoran, Viorica Patraucean, Joao Carreira, Dima Damen, Andrew Zisserman
2025arXiv / Preprint
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
Quentin Garrido, Nicolas Ballas, Mahmoud Assran, Adrien Bardes, Laurent Najman, Michael Rabbat, Emmanuel Dupoux, Yann LeCun
2025Pattern Analysis and Applications
ST-HViT: spatial-temporal hierarchical vision transformer for action recognition
Limin Xia, Weiye Fu
2025Pattern Recognition Letters
Advancing video self-supervised learning via image foundation models
Jingwei Wu, Zhewei Huang, Chang Liu
2025CVPR
SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning
Fida Mohammad Thoker, Letian Jiang, Chen Zhao†, Bernard Ghanem
2025CVPR Workshops
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
Akash Kumar, Ashlesha Kumar, Vibhav Vineet, Yogesh S Rawat
2025Proceedings of SPIE
Progressive self-supervised spatio-temporal feature learning based on video sequence saliency
Jinlong Kang, Tao Xu, Boting Qu, Xiang Wang, Xiaoli Lian, Jing Guo, Yuan Gao
2025arXiv / Preprint
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders
Shihab Aaqil Ahamed∗, Malitha Gunawardhana∗, Liel David, Michael Sidorov, Daniel Harari, Muhammad Haris Khan
2025EURASIP Journal on Image and Video Processing
Motion-driven Adaptive Frame Selection Strategy for Video Action Recognition
Hao Ding, Chen Guo, Jing Sun, Xiaoping Jiang, Hongling Shi, Jianjin Li
2025Signal, Image and Video Processing
Mitigating background bias in self-supervised video representation learning
Arif Akar, Ufuk Umut Senturk & Nazli Ikizler-Cinbis
2025TMLR
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
Sucheng Ren, Hongru Zhu, Chen Wei, Yijiang Li, Alan Yuille, Cihang Xie
2025ICCV Workshops
Learning Video Representations without Natural Videos
Xueyang Yu, Xinlei Chen, Yossi Gandelsman;
2025IEEE TIFS
Collaboratively Self-supervised Video Representation Learning for Action Recognition
Jie Zhang, Zhifan Wan, Lanqing Hu, Stephen Lin, Shuzhe Wu, Shiguang Shan
2024CVPR
Asymmetric Masked Distillation for Pre-Training Small Foundation Models
Zhiyu Zhao, Bingkun Huang, Sen Xing, Gangshan Wu, Yu Qiao, Limin Wang
2024ECCV
Data Collection-free Masked Video Modeling
Yuchi Ishikawa, Masayoshi Kondo, Yoshimitsu Aoki
2024ECCV
Text-Guided Video Masked Autoencoder
David Fan, Jue Wang, Shuai Liao, Zhikang Zhang, Vimal Bhat, Xinyu Li
2024BMVC
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
Mona Ahmadian, Frank Guerin, Andrew Gilbert
2024NeurIPS
Extending Video Masked Autoencoders to 128 Frames
Nitesh Bharadwaj Gundavarapu, Luke Friedman, Raghav Goyal, Chaitra Hegde, Eirikur Agustsson, Sagar M. Waghmare, Mikhail Sirotenko, Ming-Hsuan Yang, Tobias Weyand, Boqing Gong, Leonid Sigal
2024arXiv / Preprint
Scaling 4D Representations
João Carreira et al.
2024CVPR
VideoMAC: Video Masked Autoencoders Meet ConvNets
Gensheng Pei, Tao Chen, Xiruo Jiang, Huafeng Liu, Zeren Sun, Yazhou Yao
2024arXiv / Preprint
Self-supervised Video Object Segmentation with Distillation Learning of Deformable Attention
Quang-Trung Truong,Duc Thanh Nguyen, Binh-Son Hua, Sai-Kit Yeung
2024ECCV
Towards Latent Masked Image Modeling for Self-supervised Visual Representation Learning
Yibing Wei, Abhinav Gupta & Pedro Morgado
2024ECCV
SIGMA: Sinkhorn-Guided Masked Video Modeling
Mohammadreza Salehi, Michael Dorkenwald, Fida Mohammad Thoker, Efstratios Gavves, Cees G. M. Snoek & Yuki M. Asano
2024CVPR Workshops
ST2ST: Self-Supervised Test-time Adaptation for Video Action Recognition
Masud An-Nur Islam Fahim, Mohammed Innat, Jani Boutellier;
2024CogSci
Self-supervised learning of video representations from a child's perspective
A. Emin Orhan, Wentao Wang, Alex N. Wang, Mengye Ren, Brenden M. Lake;
2024ECCV
ViC-MAE: Self-supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders
Jefferson Hernandez, Ruben Villegas, Vicente Ordonez;
2024CVPR
Learning to Predict Activity Progress by Self-Supervised Video Alignment
Gerard Donahue, Ehsan Elhamifar;
2024Pattern Recognition
Repeat and learn: Self-supervised visual representations learning by Repeated Scene Localization
Yuanhang Zhang, Shuang Yang, Shiguang Shan, Xilin Chen;
2024CVPR
ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations
Yuanhang Zhang, Shuang Yang, Shiguang Shan, Xilin Chen;
2024WACV
Self-supervised Learning of Semantic Correspondence Using Web Videos
Donghyeon Kwon, Minsu Cho, Suha Kwak;
2024IPEC
Video Compression and Action Recognition in Self-supervised Learning
Zongbo Hao; Conghui Hao; Kecheng He
2024WACV
CycleCL: Self-supervised Learning for Periodic Videos
Matteo Destro, Michael Gygl
2024ICME Workshops
Self-Supervised Learning via Multi-Transformation Classification for Action Recognition
Duc-Quang Vu; Ngan Le; Jia-Ching Wang
2024Pattern Recognition
Motion-guided spatiotemporal multitask feature discrimination for self-supervised video representation learning
Shuai Bi, Zhengping Hu, Hehao Zhang, Jirui Di, Zhe Sun
2024CVPR
What When and Where? Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann, Samuel Thomas, Shih-Fu Chang, Rogerio Feris, James Glass, Hilde Kuehne
2024Applied Intelligence
Clustering-based multi-featured self-supervised learning for human activities and video retrieval
Muhammad Hafeez Javed, Zeng Yu, Taha M. Rajeh, Fahad Rafique & Tianrui Li
2024ICASSP Workshops
Positive and negative sampling strategies for self-supervised learning on audio-video data
Shanshan Wang, Soumya Tripathy, Toni Heittola, Annamaria Mesaros
2024AAAI
No More Shortcuts: Realizing the Potential of Temporal Self-Supervision
Ishan Rajendrakumar Dave, Simon Jenni, Mubarak Shah.
2024Pattern Recognition Letters
GLOCAL: A self-supervised learning framework for global and local motion estimation
Yihao Zheng , Kunming Luo , Shuaicheng Liu , Zun Li , Ye Xiang , Lifang Wu , Bing Zeng , Chang Wen Chen
2024IEEE TCSVT
Self-supervised Video Representation Learning via Capturing Semantic Changes Indicated by Saccades
Qiuxia Lai, Ailing Zeng, Ye Wang, Lihong Cao, Yu Li, Qiang Xu, IEEE
2024IEEE TMM
MAR: Masked Autoencoders for Efficient Action Recognition
Zhiwu Qing, Shiwei Zhang, Ziyuan Huang, Xiang Wang, Yuehuan Wang, Yiliang Lv, Changxin Gao, Nong Sang
2024CVPR
VicTR: Video-conditioned Text Representations for Activity Recognition
Kumara Kahatapitiya, Anurag Arnab, Arsha Nagrani, Michael S. Ryoo
2024IEEE Transactions on Cybernetics
Self-Supervised Video Representation Learning by Video Incoherence Detection
Haozhi Cao, Yuecong Xu, Kezhi Mao, Lihua Xie, Jianxiong Yin, Simon See, Qianwen Xu, and Jianfei Yang
2024ICLR
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
Xiong, Y., Zhao, L., Gong, B., Yang, M. H., Schroff, F., Liu, T., ... & Yuan, L.
2024BMVC
MotionMAE: Self-supervised Video Representation Learning with Motion-Aware Masked Autoencoders
Haosen Yang, Deng Huang, Bin Wen, Jiannan Wu, Hongxun Yao, Yi Jiang, Xiatian Zhu, Zehuan Yuan
2024ICML
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
Sunil Hwang, Jaehong Yoon, Youngwan Lee, Sung Ju Hwan
2024AAAI
XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning
Pritam Sarkar, Ali Etemad
2024Visual Intelligence
Controllable Augmentations for Video Representation Learning
Rui Qian, Weiyao Lin, John See, Dian Li
2023NeurIPS
Self-supervised object-centric learning for videos
Görkay Aydemir, Weidi Xie, Fatma Guney
2023NeurIPS
Language-based Action Concept Spaces Improve Video Self-Supervised Learning
Kanchana Ranasinghe, Michael S Ryoo
2023NeurIPS
Uncovering the Hidden Dynamics of Video Self-supervised Learning under Distribution Shifts
Pritam Sarkar, Ahmad Beirami, Ali Etemad
2023NeurIPS
Self-supervised video pretraining yields robust and more human-aligned visual representation
Nikhil Parthasarathy, S. M. Ali Eslami, João Carreira, Olivier J. Hénaff.
2023CVPR
AdaMAE: Adaptive Masking for Efficient Spatiotemporal Learning with Masked Autoencoders
Wele Gedara Chaminda Bandara, Naman Patel, Ali Gholami, Mehdi Nikkhah, Motilal Agrawal, Vishal M. Patel
2023ICCV
Spatio-Temporal Crop Aggregation for Video Representation Learning
Sepehr Sameni, Simon Jenni, Paolo Favaro
2023ICCV
Motion-Guided Masking for Spatiotemporal Representation Learning
David Fan, Jue Wang, Shuai Liao, Yi Zhu, Vimal Bhat, Hector Santos-Villalobos, Rohith MV, Xinyu Li
2023ACM Multimedia
Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation Learning
Minghao Zhu, Xiao Lin, Ronghao Dang, Chengju Liu, Qijun Chen
2023ICCV
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
Kunchang Li, Yali Wang, Yizhuo Li, Yi Wang, Yinan He, Limin Wang, Yu Qiao
2023arXiv / Preprint
Concatenated Masked Autoencoders as Spatial-Temporal Learner
Zhouqiang Jiang, Bowen Wang, Tong Xiang, Zhaofeng Niu, Hong Tang, Guangshun Li, Liangzhi Li
2023ICTAI
AV-MaskEnhancer: Enhancing Video Representations through Audio-Visual Masked Autoencoder
Xingjian Diao, Ming Cheng, Shitong Cheng
2023CVPR
OmniMAE: Single Model Masked Pretraining on Images and Videos
Rohit Girdhar, Alaaeldin El-Nouby, Mannat Singh,Kalyan Vasudev Alwala, Armand Joulin , Ishan Misra
2023CVPR
TimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action Recognition
Ishan Rajendrakumar Dave, Mamshad Nayeem Rizve, Chen Chen, Mubarak Shah
2023Image and Vision Computing
Attentive spatial-temporal contrastive learning for self-supervised video representation
Xingming Yang, Sixuan Xiong, Kewei Wu, Dongfeng Shan, Zhao Xie
2023ICCV
MGMAE: Motion Guided Masking for Video Masked Autoencoding
Bingkun Huang, Zhiyu Zhao, Guozhen Zhang, Yu Qiao, Limin Wang
2023MVA
Cross-modal Manifold Cutmix for Self-supervised Video Representation Learning
Srijan Das; Michael Ryoo
2023ACM Multimedia
CHAIN: Exploring Global-Local Spatio-Temporal Information for Improved Self-Supervised Video Hashing
Rukai Wei, Yu Liu, Jingkuan Song, Heng Cui, Yanzhao Xie, Ke Zhou
2023ACM Multimedia
Data-Efficient Masked Video Modeling for Self-supervised Action Recognition
Qiankun Li, Xiaolong Huang, Zhifan Wan, Lanqing Hu, Shuzhe Wu, Jie Zhang, Shiguang Shan, Zengfu Wang(
2023IEEE Internet of Things Journal
Temporal Transformer Networks with Self-Supervision for Action Recognition
Yongkang Zhang, Jun Li, Guoming Wu, Han Zhang, Zhiping Shi, Member, IEEE, Zhaoxun Liu, Zizhang Wu
2023arXiv / Preprint
CMAE-V: Contrastive Masked Autoencoders for Video Action Recognition
Cheng-Ze Lu, Xiaojie Jin, Zhicheng Huang, Qibin Hou, Ming-Ming Cheng, Jiashi Feng
2023Computer Vision and Image Understanding
Learning Representational Invariances for Data-Efficient Action Recognition
Yuliang Zou, Jinwoo Choi, Qitong Wang, Jia-Bin Huang
2023Neurocomputing
SOR-TC: Self-attentive octave ResNet with temporal consistency for compressed video action recognition
Junsan Zhang, Xiaomin Wang, Yao Wan, Leiquan Wang, Jian Wang, Philip S. Yu
2023CVPR
Masked Motion Encoding for Self-Supervised Video Representation Learning
Xinyu Sun, Peihao Chen, Liangwei Chen, Thomas H. Li, Mingkui Tan, Chuang Gan
2023Signal, Image and Video Processing
Spatiotemporal consistency enhancement self-supervised representation learning for action recognition
Shuai Bi, Zhengping Hu, Mengyao Zhao, Shufang Li & Zhe Sun
2023IEEE TIP
Self-Supervised Video-Based Action Recognition With Disturbances
Wei Lin, Xinghao Ding, Yue Huang, Huanqiang Zeng
2023CVPR
Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning
Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Lu Yuan, Yu-Gang Jiang
2023Engineering Applications of Artificial Intelligence
Enhancing motion visual cues for self-supervised video representation learning
Mu Nie, Zhibin Quan, Weiping Ding, and Wankou Yang
2023Advanced Engineering Informatics
Continuous frame motion sensitive self-supervised collaborative network for video representation learning
Shuai Bi, Zhengping Hu, Mengyao Zhao, Hehao Zhang, Jirui Di, and Zhe Sun
2023Signal, Image and Video Processing
Self-supervised pretext task collaborative multi-view contrastive learning for video action recognition
Shuai Bi, Zhengping Hu, Mengyao Zhao, Hehao Zhang, Jirui Di, and Zhe Sun
2023IEEE TPAMI
Self-Supervised Learning from Untrimmed Videos via Hierarchical Consistency
Zhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yi Xu, Xiang Wang, Changxin Gao, Rong Jin, and Nong Sang
2023AAAI
Audio-Visual Contrastive Learning with Temporal Self-Supervision
Simon Jenni, Alexander Black, and John Collomosse
2023CVPR
Video Test-Time Adaptation for Action Recognition
Wei Lin, Muhammad Jehanzeb Mirza, Mateusz Kozinski, Horst Possegger, Hilde Kuehne, and Horst Bischof
2023AAAI
Self-Supervised Video Representation Learning via Latent Time Navigation
Di Yang, Yaohui Wang, Quan Kong, Antitza Dantcheva, Lorenzo Garattoni, Gianpiero Francesca, and Francois Bremond
2023ICASSP
Temporal Contrastive Learning with Curriculum
Shuvendu Roy and Ali Etemad
2023ICLR Workshops
Nearest-Neighbor Inter-Intra Contrastive Learning from Unlabeled Videos
David Fan, Deyu Yang, Xinyu Li, Vimal Bhat, and Rohith MV
2023ICCV
Tubelet-Contrastive Self-Supervision for Video-Efficient Generalization
Fida Mohammad Thoker, Hazel Doughty, and Cees Snoek
2023ICASSP
Multi-scale Compositional Constraints for Representation Learning on Videos
Georgios Paraskevopoulos, Chandrashekhar Lavania, Lovish Chum, and Shiva Sundaram
2023WACV
Flavr: Flow-agnostic Video Representations for Fast Frame Interpolation
Tarun Kalluri, Deepak Pathak, Manmohan Chandraker, and Du Tran
2023arXiv / Preprint
HomE: Homography-Equivariant Video Representation Learning
Anirudh Sriram, Adrien Gaidon, Jiajun Wu, Juan Carlos Niebles, Li Fei-Fei, and Ehsan Adeli
2023WACV
ViewCLR: Learning Self-supervised Video Representation for Unseen Viewpoints
Srijan Das and Michael S Ryoo
2023CVPR
Videomae v2: Scaling Video Masked Autoencoders with Dual Masking
Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao
2023AAAI
Self-Supervised Audio-Visual Representation Learning with Relaxed Cross-Modal Synchronicity
Pritam Sarkar, Ali Etemad
2023WACV
Previts: contrastive pretraining with video tracking supervision
Chen, B., Selvaraju, R. R., Chang, S. F., Niebles, J. C., & Naik, N.
2023CVPR
Modeling Video As Stochastic Processes for Fine-Grained Video Representation Learning
Zhang, H., Liu, D., Zheng, Q., & Su, B.
2023ICCV
Learning Fine-Grained Features for Pixel-wise Video Correspondences
Li, R., Zhou, S., & Liu, D.
2023CVPR Workshops
Cali-NCE: Boosting Cross-Modal Video Representation Learning With Calibrated Alignment
Zhao, N., Jiao, J., Xie, W., & Lin, D.
2023IEEE TNNLS
Self-supervised motion perception for spatiotemporal representation learning
Chang Liu, Yuan Yao, Dezhao Luo, Yu Zhou, Qixiang Ye
2023Machine Vision and Applications
Similarity Contrastive Estimation for Image and Video Soft Contrastive Self-Supervised Learning
Julien Denize, Jaonary Rabarisoa, Astrid Orcesi, Romain H´erault
2023ICIP
Self-Supervised Contrastive Learning for Audio-Visual Action Recognition
Yang Liu, Ying Tan, Haoyuan Lan
2023IEEE TMM
Self-Supervised Scene-Debiasing for Video Representation Learning via Background Patching
Maregu Assefa, Wei Jiang, Kumie Gedamu, Getinet Yilma, Bulbula Kumeda, Melese Ayalew
2023IEEE TMM
LgNet: A local-global network for action recognition and beyond
Jiaqi Zhou, Zehua Fu, Qiuyu Huang, Qingjie Liu, Yunhong Wang
2023IEEE TCSVT
Unsupervised Video-Based Action Recognition With Imagining Motion and Perceiving Appearance
Wei Lin , Xiaoyu Liu , Yihong Zhuang , Xinghao Ding , Xiaotong Tu , Yue Huang , Huanqiang Zeng
2023AAAI
Spatiotemporal Augmentation on Selective Frequencies for Video Representation Learning
Jinhyung Kim, Taeoh Kim, Minho Shim, Dongyoon Han, Dongyoon Wee, Junmo Kim
2023IEEE TCSVT
Consistent Intra-video Contrastive Learning with Asynchronous Long-term Memory Bank
Zelin Chen, Kun-Yu Lin, Wei-Shi Zheng
2022CVPR
BEVT: BERT Pretraining of Video Transformers
Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, Lu Yuan
2022NeurIPS
Masked Autoencoders As Spatiotemporal Learners
Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming He
2022CVPR
SPAct: Self-supervised Privacy Preservation for Action Recognition
Ishan Rajendrakumar Dave, Chen Chen, Mubarak Shah
2022AAAI
Suppressing Static Visual Cues via Normalizing Flows for Self-Supervised Video Representation Learning
Manlin Zhang, Jinpeng Wang, Andy J. Ma
2022CVPR
Self-supervised Video Transformer
Kanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan, Michael S. Ryoo
2022ACM TOMM
Exploring Relations in Untrimmed Videos for Self-Supervised Learning
Dezhao Luo, Bo Fang, Yu Zhou, Yucan Zhou, Dayan Wu, Weiping Wang
2022ACM Multimedia
MaMiCo: Macro-to-Micro Semantic Correspondence for Self-supervised Video Representation Learning
Bo Fang, Wenhao Wu, Chang Liu, Yu Zhou, Dongliang He, Weiping Wang
2022IEEE TIP
TCGL: Temporal Contrastive Graph for Self-Supervised Video Representation Learning
Yang Liu , Keze Wang , Lingbo Liu , Haoyuan Lan, and Liang Lin
2022CVPR
Cross-Architecture Self-supervised Video Representation Learning
Sheng Guo, Zihua Xiong, Yujie Zhong, Limin Wang, Xiaobo Guo, Bing Han, Weilin Huang
2022AAAI
Contrastive spatio-temporal pretext learning for self-supervised video representation
Yujia Zhang, Lai-Man Po, Xuyuan Xu, Mengyang Liu, Yexin Wang, Weifeng Ou, Yuzhi Zhao, Wing-Yin Yu
2022CVPR
Transrank: Self-supervised video representation learning via ranking-based transformation recognition
Haodong Duan, Nanxuan Zhao, Kai Chen, Dahua Lin
2022CVPR
Learning from untrimmed videos: Self-supervised video representation learning with hierarchical consistency
Zhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yi Xu, Xiang Wang, Mingqian Tang, Changxin Gao, Rong Jin,Nong Sang
2022CVPR
Motion-aware contrastive video representation learning via foreground-background merging
Shuangrui Ding, Maomao Li, Tianyu Yang, Rui Qian, Haohang Xu, Qingyi Chen, Jue Wang, Hongkai Xiong
2022ICME
Self-Supervised Video Representation Learning with Motion-Contrastive Perception
Jinyu Liu, Ying Cheng, Yuejie Zhang, Rui-Wei Zhao, Rui Feng
2022IEEE TCSVT
Self-supervised video representation learning using improved instance-wise contrastive learning and deep clustering
Yisheng Zhu, Hui Shuai, Guangcan Liu, Senior Member, Qingshan Liu
2022Computer Vision and Image Understanding
TCLR: Temporal contrastive learning for video representation
Ishan Dave, Rohit Gupta, Mamshad Nayeem Rizve, Mubarak Shah
2022AAAI
Self-supervised spatiotemporal representation learning by exploiting video continuity
Hanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen, Peng Dai, Juwei Lu, Yang Wang
2022CVPR
Probabilistic representations for video contrastive learning
Jungin Park, Jiyoung Lee, Ig-Jae Kim, Kwanghoon Sohn
2022CVPR
Contextualized spatio-temporal contrastive learning with self-supervision
Liangzhe Yuan, Rui Qian, Yin Cui, Boqing Gong,Florian Schroff,Ming-Hsuan Yang, Hartwig Adam, Ting Liu
2022NeurIPS
VideoMAE: Masked Autoencoders Are Data-Efficient Learners for Self-Supervised Video Pre-Training
Zhan Tong, Yibing Song, Jue Wang, Limin Wang
2022WACV
Self-supervised video representation learning with cross-stream prototypical contrasting
Martine Toering, Ioannis Gatopoulos, Maarten Stol, Vincent Tao Hu
2022CVPR
SLIC: Self-supervised learning with iterative clustering for human action videos
Salar Hosseini Khorasgani, Yuxuan Chen, Florian Shkurti
2022ECCV
GOCA: guided online cluster assignment for self-supervised video representation Learning
Huseyin Coskun, Alireza Zareian, Joshua L. Moore, Federico Tombari, Chen Wang
2022ACCV
TCVM: Temporal Contrasting Video Montage Framework for Self-supervised Video Representation Learning
Fengrui Tian, Jiawei Fan, Xie Yu, Shaoyi Du, Meina Song, Yu Zhao
2022ECCV
Static and Dynamic Concepts for Self-supervised Video Representation Learning
Rui Qian, Shuangrui Ding, Xian Liu, Dahua Lin
2022ECCV
SOS! Self-supervised Learning over Sets of Handled Objects in Egocentric Action Recognition
Victor Escorcia, Ricardo Guerrero, Xiatian Zhu, Brais Martinez
2022CVPR Workshops
Self-Supervised Video Representation Learning with Cascade Positive Retrieval
Cheng-En Wu, Farley Lai, Yu Hen Hu, Asim Kadav
2022IEEE JSTSP
Self-Supervised Learning of Audio Representations From Audio-Visual Data Using Spatial Alignment
Shanshan Wang, Archontis Politis, Annamaria Mesaros
2022WACV
Hierarchically decoupled spatial-temporal contrast for self-supervised video representation learning
Zehua Zhang, David Crandall
2022ICME
Spatio-temporal self-supervision enhanced transformer networks for action recognition
Yongkang Zhang, Han Zhang, Guoming Wu, Jun Li
2022ICPR
Inter-Intra Cross-Modality Self-Supervised Video Representation Learning by Contrastive Clustering
Jiutong Wei. Guan Luo, Bing Li, Weiming Hu
2022CVPR Workshops
SCVRL: Shuffled Contrastive Video Representation Learning
Michael Dorkenwald, Fanyi Xiao, Biagio Brattoli, Joseph Tighe, Davide Modolo
2022arXiv / Preprint
InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang,Jilan Xu, Yi Liu, Zun Wang, Sen Xing, Guo Chen, Junting Pan, Jiashuo Yu,Yali Wang, Limin Wang, Yu Qiao
2022ICANN
Video Motion Perception for Self-supervised Representation Learning
Wei Li, Dezhao Luo, Bo Fang, Xiaoni Li, Yu Zhou, Weiping Wang
2022IEEE TCSVT
An improved inter-intra contrastive learning framework on self-supervised video representation
Li Tao, Xueting Wang, Toshihiko Yamasaki
2022CVPR Workshops
Auxiliary Learning for Self-Supervised Video Representation via Similarity-based Knowledge Distillation
Amirhossein Dadashzadeh, Alan Whone, Majid Mirmehdi
2022ECCV
Motion Sensitive Contrastive Learning for Self-supervised Video Representation
Jingcheng Ni, Nan Zhou, Jie Qin, Qian Wu, Junqi Liu, Boxun Li, Di Huang
2022NCC
Unsupervised Learning of Spatio-Temporal Representation with Multi-Task Learning for Video Retrieval
Vidit Kumar
2022ECCV
Federated Self-supervised Learning for Video Understanding
Yasar Abbas Ur Rehman, Yan Gao, Jiajun Shen, Pedro Porto Buarque de Gusmão , Nicholas Lane
2022Neurocomputing
Contrastive predictive coding with transformer for video representation learning
Yue Liu, Junqi Ma, Yufei Xie, Xuefeng Yang, Xingzhen Tao, Lin Peng, Wei Gao
2022Applied Intelligence
Video representation learning by identifying spatio-temporal transformation
Sheng Geng, Shimin Zhao , Hu Liu
2022BMVC
On temporal granularity in self-supervised video representation learning
Rui Qian, Yeqing Li, Liangzhe Yuan, Boqing Gong, Ting Liu, Matthew Brown, Serge Belongie, Ming-Hsuan Yang, Hartwig Adam, and Yin Cui
2022ICML Workshops
LAVA: Language Audio Vision Alignment for Data-Efficient Video Pre-Training
Sumanth Gurram , Andy Fang , David Chan , John Canny
2022arXiv / Preprint
It Takes Two: Masked Appearance-Motion Modeling for Self-supervised Video Transformer Pre-training
Yuxin Song, Min Yang, Wenhao Wu, Dongliang He, Fu Li, Jingdong Wang
2022BMVC
MAC: Mask-Augmentation for Motion-Aware Video Representation Learning
Arif Akar, Ufuk Umut Senturk, and Nazli Ikizler-Cinbis.
2022AVSS
Temporal-Invariant Video Representation Learning with Dynamic Temporal Resolutions.
Seong-Yun Jeong, Ho-Joong Kim, Myeong-Seok Oh, Gun-Hee Lee, Seong-Whan Lee
2022ACM Multimedia
Dual Contrastive Learning for Spatio-temporal Representation
Shuangrui Ding,, Rui Qian, and Hongkai Xiongo
2022ECCV Workshops
MoQuad: Motion-focused Quadruple Construction for Video Contrastive Learning
Yuan Liu, Jiacheng Chen, Hao Wu
2022arXiv / Preprint
On Negative Sampling for Audio-Visual Contrastive Learning from Movies
Mahdi M. Kalayeh, Shervin Ardeshir, Lingyi Liu, Nagendra Kamath, Ashok Chandrashekar
2022CVPR
Frame-wise Action Representations for Long Videos via Sequence Contrastive Learning
Minghao Chen, Fangyun Wei, Chong Li, Deng Cai
2022CVPR
Masked Feature Prediction for Self-Supervised Visual Pre-Training
Wei, C., Fan, H., Xie, S., Wu, C. Y., Yuille, A., & Feichtenhofer, C.
2022arXiv / Preprint
Pixel-level Correspondence for Self-Supervised Learning from Video
Yash Sharma, Yanchao Zhu, Chris Russell, Thomas Brox
2022CVPR
Temporal Alignment Networks for Long-Term Video
Han, T., Xie, W., & Zisserman, A.
2022arXiv / Preprint
SimVTP: Simple Video Text Pre-Training with Masked Autoencoders
Ma, Y., Yang, T., Shan, Y., & Li, X.
2022ICLR
Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Shi, B., Hsu, W. N., Lakhotia, K., & Mohamed, A.
2022IEEE TPAMI
Self-supervised video representation learning by uncovering spatio-temporal statistics
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Wei Liu, Yun-hui Liu
2021BMVC
Inter-intra Variant Dual Representations for Self-supervised Video Recognition
Lin Zhang, Qi She, Zhengyang Shen, Changhu Wang
2021NeurIPS
VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning
Hao Tan, Jie Lei, Thomas Wolf, Mohit Bansal
2021arXiv / Preprint
Watching too much television is good: Self-supervised audio-visual representation learning from movies and tv shows
Mahdi M. Kalayeh, Nagendra Kamath, Lingyi Liu
2021ICPR
Temporally coherent embeddings for self-supervised video representation learning
Joshua Knights, Ben Harwood, Daniel Ward, Anthony Vanderkop, Olivia Mackenzie-Ross, Peyman Moghadam
2021CVPR
Audio-visual instance discrimination with cross-modal agreement
Pedro Morgado, Nuno Vasconcelos, Ishan Misra
2021CVPR
Removing the background by adding the background: Towards background robust self-supervised video representation learning
Jinpeng Wang, Yuting Gao, Ke Li, Yiqi Lin, Andy J. Ma, Hao Cheng, Pai Peng, Feiyue Huang, Rongrong Ji, Xing Sun
2021AAAI
Enhancing unsupervised video representation learning by decoupling the scene and the motion
Jinpeng Wang, Yuting Gao, Ke Li, Jianguo Hu, Xinyang Jiang, Xiaowei Guo, Rongrong Ji, Xing Sun
2021AAAI
SeCo: Exploring Sequence Supervision for Unsupervised Representation Learning
Ting Yao, Yiheng Zhang, Zhaofan Qiu, Yingwei Pan, Tao Mei
2021ICCV
Enhancing self-supervised video representation learning via multi-level feature optimization
Rui Qian, Yuxi Li, Huabin Liu, John See, Shuangrui Ding, Xian Liu, Dian Li, Weiyao Lin
2021AAAI
RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning
Peihao Chen, Deng Huang, Dongliang He, Xiang Long, Runhao Zeng, Shilei Wen, Mingkui Tan, Chuang Gan
2021CVPR
VideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples
Tian Pan, Yibing Song, Tianyu Yang, Wenhao Jiang, Wei Liu
2021ICCV
On compositions of transformations in contrastive self-supervised learning
Mandela Patrick, Yuki M. Asano, Polina Kuznetsova, Ruth Fong, João F. Henriques, Geoffrey Zweig, Andrea Vedaldi
2021CVPR
Unsupervised visual representation learning by tracking patches in video
Guangting Wang, Yizhou Zhou, Chong Luo, Wenxuan Xie, Wenjun Zeng, Zhiwei Xiong
2021CVPR
A large-scale study on unsupervised spatiotemporal representation learning
Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, Kaiming He
2021CVPR
CoCon: Cooperative-Contrastive Learning
Nishant Rai, Ehsan Adeli ,Kuan-Hui Lee, Adrien Gaidon, Juan Carlos Niebles
2021NeurIPS
VATT: Transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, Boqing Gong
2021ICCV
ASCNet: Self-supervised video representation learning with appearance-speed consistency
Deng Huang, Wenhao Wu, Weiwen Hu, Xu Liu, Dongliang He, Zhihua Wu, Xiangmiao Wu, Mingkui Tan, Errui Ding
2021IEEE Access
Self-supervised visual learning by variable playback speeds prediction of a video
Hyeon Cho, Taehoon Kim, Hyungjin Chang, Wonjun Hwang
2021ICCV
Self-supervised video representation learning with meta-contrastive network
Yuanze Lin, Xun Guo, Yan Lu
2021ICCV
Long short view feature decomposition via contrastive video representation learning
Nadine Behrmann, Mohsen Fayyaz, Juergen Gall, Mehdi Noroozi
2021ICCV
Time-equivariant contrastive video representation learning
Simon Jenni, Hailin Jin
2021CVPR
Self-supervised video representation learning by context and motion decoupling
Lianghua Huang, Yu Liu, Bin Wang, Pan Pan, Yinghui Xu, Rong Jin
2021WACV
Unsupervised video representation learning by bidirectional feature prediction
Nadine Behrmann, Juergen Gall, Mehdi Noroozi
2021ICLR
Self-supervised learning of compressed video representations
Youngjae Yu, Sangho Lee, Gunhee Kim, Yale Song
2021CVPR
Spatiotemporal contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, Yin Cui
2021arXiv / Preprint
MoDist: Motion Distillation for Self-Supervised Video Representation Learning
Fanyi Xiao, Joseph Tighe, Davide Modolo
2021ICCV
Broaden your views for self-supervised video learning
Adria Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang, Ross Hemsley, Florian Strub, Corentin Tallec, Mateusz Malinowski, Viorica Patraucean, Florent Altche, Michal Valko, Jean-Bastien Grill, Aaron van den Oord, Andrew Zisserman
2021ICCV
Vi2CLR: Video and image for visual contrastive learning of representation
Ali Diba, Vivek Sharma, Reza Safdari, Dariush Lotfi, M. Saquib Sarfraz,Rainer Stiefelhagen, Luc Van Gool,
2021ICCV
Contrast and order representations for video self-supervised learning
Kai Hu, Jie Shao, Yuan Liu, Bhiksha Raj, Marios Savvides, Zhiqiang Shen
2021ICCV
Motion-augmented self-training for video recognition at smaller scale
Kirill Gavrilyuk, Mihir Jain, Ilia Karmanov, Cees G. M. Snoek
2021ICCV Workshops
Video contrastive learning with global context
Haofei Kuang, Yi Zhu, Zhi Zhang, Xinyu Li, Joseph Tighe,Soren Schwertfeger, Cyrill Stachniss, Mu Li
2021ICCV
Motion-focused contrastive learning of video representations
Rui Li, Yiheng Zhang, Zhaofan Qiu, Ting Yao, Dong Liu, and Tao Mei
2021BMVC
Back to the Future: Cycle Encoding Prediction for Self-supervised Video Representation Learning
Xinyu Yang, Majid Mirmehdi,Tilo Burghardt
2021ICCV
Composable augmentation encoding for video representation learning
Sun, C., Nagrani, A., Tian, Y., & Schmid, C.
2021ICCV
Learning temporal dynamics from cycles in narrated video
Epstein, D., Wu, J., Schmid, C., & Sun, C.
2021ICCV
CrossCLR: Cross-Modal Contrastive Learning for Multi-Modal Video Representations
Zolfaghari, M., Zhu, Y., Gehler, P., & Brox, T.
2021ICLR
Watching the World Go By: Representation Learning from Unlabeled Videos
Daniel Gordon, Kiana Ehsani, Dieter Fox, Ali Farhadi
2021ICLR
Parameter Efficient Multimodal Transformers for Video Representation Learning
Lee, S., Yu, Y., Kim, G., Breuel, T., Kautz, J., & Song, Y.
2021ICLR
Active Contrastive Learning of Audio-Visual Video Representations
Ma, S., Zeng, Z., McDuff, D., & Song, Y.
2020Sensors
Self-Supervised Learning to Detect Key Frames in Videos
Xiang Yan,Syed Zulqarnain Gilani,Mingtao Feng ,Liang Zhang,Hanlin Qin and Ajmal Mian
2020ECCV
Self-supervised motion representation via scattering local motion cues
Yuan Tian, Zhaohui Che, Wenbo Bao, Guangtao Zhai, Zhiyong Gao1
2020ACM Multimedia
Self-supervised video representation learning using inter-intra contrastive framework
Li Tao, Xueting Wang, Toshihiko Yamasaki
2020arXiv / Preprint
Video representation learning with visual tempo consistency
Ceyuan Yang, Yinghao Xu, Bo Dai, Bolei Zhou
2020arXiv / Preprint
Self-supervised temporal discriminative learning for video representation learning
Jinpeng Wang, Yiqi Lin, Andy J. Ma,Pong C. Yuen
2020NeurIPS
Self-supervised learning by cross-modal audio-video clustering
Humam Alwassel, Dhruv Mahajan, Bruno Korbar ,Lorenzo Torresani, Bernard Ghanem, Du Tran
2020ECCV
Self-supervised video representation learning by pace prediction
Jiangliu Wang, Jianbo Jiao, Yun-Hui Liu
2020CVPR
Unsupervised learning from video with deep neural embeddings
Chengxu Zhuang, Tianwei She, Alex Andonian, Max Sobol Mark, Daniel Yamins
2020ECCV Workshops
Unsupervised learning of video representations via dense trajectory clustering
Pavel Tokmakov, Martial Hebert, Cordelia Schmid
2020ECCV
Video representation learning by recognizing temporal transformations
Simon Jenni, Givi Meishvili, Paolo Favaro
2020CVPR
Video playback rate perception for self-supervised spatio-temporal representation learning
Yuan Yao, Chang Liu, Dezhao Luo, Yu Zhou, Qixiang Ye
2020NeurIPS
Self-supervised co-training for video representation learning
Tengda Han, Weidi Xie, Andrew Zisserman
2020AAAI
Video cloze procedure for self-supervised spatio-temporal learning
Dezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang, Can Ma, Qixiang Ye, Weiping Wang
2020CVPR
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira,Ivan Laptev, Josef Sivic, Andrew Zisserman
2020CVPR
SpeedNet: Learning the Speediness in Videos
Sagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri, William T. Freeman, Michael Rubinstein, Michal Irani, Tali Dekel
2020ECCV
Contrastive multiview coding
Yonglong Tian, Dilip Krishnan, Phillip Isola
2020Signal Processing: Image Communication
Self-supervised video representation learning by maximizing mutual information
Fei Xue, Hongbing Ji, Wenbo Zhang, Yi Cao
2020ECCV
Memory-augmented dense predictive coding for video representation learning
Tengda Han, Weidi Xie, Andrew Zisserman
2020CVPR
Evolving losses for unsupervised video representation learning
AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo
2020arXiv / Preprint
AudioVisual SlowFast Networks for Video Recognition
Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, Christoph Feichtenhofer
2020NeurIPS
Cycle-Contrast for Self-Supervised Video Representation Learning
Quan Kong, Wenpeng Wei, Ziwei Deng, Tomoaki Yoshinaga, Tomokazu Murakami
2020arXiv / Preprint
Can temporal information help with contrastive self-supervised learning?
Yutong Bai, Haoqi Fan, Ishan Misra, Ganesh Venkatesh, Yongyi Lu
2020NeurIPS
Self-supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira Sander Dieleman, Andrew Zisserman
2020arXiv / Preprint
Pretext-Contrastive Learning: Toward Good Practices in Self-Supervised Video Representation Learning
Li Tao, Xueting Wang, Toshihiko Yamasaki
2020arXiv / Preprint
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation
Luo, H., Ji, L., Shi, B., Huang, H., Duan, N., Li, T., ... & Zhou, M.
2020ECCV
Self-supervised learning of audio-visual objects from video
Afouras, T., Owens, A., Chung, J. S., & Zisserman, A.
2020CVPR
Speech2Action: Cross-Modal Supervision for Action Recognition
Nagrani, A., Sun, C., Ross, D., Sukthankar, R., Schmid, C., & Zisserman, A.
2020ACM Multimedia
Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation Learning
Cheng, Y., Wang, R., Pan, Z., Feng, R., & Zhang, Y.
2019CVPR
Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Yunhui Liu, Wei Liu
2019ICCV Workshops
Video representation learning by dense predictive coding
Tengda Han, Weidi Xie, Andrew Zisserman
2019CVPR
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, Yueting Zhuang
2019WACV
Video Jigsaw: Unsupervised Learning of Spatiotemporal Context for Video Action Recognition
Unaiza Ahsan, Rishi Madhok, Irfan Essa
2019AAAI
Self-supervised video representation learning with space-time cubic puzzles
Dahun Kim, Donghyeon Cho, In So Kweon
2019arXiv / Preprint
Learning Video Representations Using Contrastive Bidirectional Transformer
Chen Sun, Fabien Baradel, Kevin Murphy, Cordelia Schmid
2019ICCV
DynamoNet: Dynamic Action and Motion Network
Ali Diba, Vivek Sharma, Luc Van Gool, Rainer Stiefelhagen
2019CVPR
Temporal Cycle-Consistency Learning
Dwibedi, D., Aytar, Y., Tompson, J., Sermanet, P., & Zisserman, A.
2019ICCV
VideoBERT: A Joint Model for Video and Language Representation Learning
Sun, C., Myers, A., Vondrick, C., Murphy, K., & Schmid, C.
2018CVPR
Geometry Guided Convolutional Neural Networks for Self-Supervised Video Representation Learning
Chuang Gan, Boqing Gong, Kun Liu, Hao Su, Leonidas J. Guibas
2018arXiv / Preprint
Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction
Longlong Jing, Xiaodong Yang, Jinggen Liu, Yingli Tian
2018NeurIPS
Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization
Bruno Korbar, Du Tran, Lorenzo Torresani
2018ECCV
Audio-Visual Scene Analysis with Self-Supervised Multisensory Features
Andrew Owens, Alexei A. Efros
2018CVPR
Compressed Video Action Recognition
Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu, R. Manmatha, Alexander J. Smola, Philipp Krahenb
2018ECCV
Improving Spatiotemporal Self-Supervision by Deep Reinforcement Learning
Uta Buchler, Biagio Brattoli, Bjorn Ommer
2018CVPR
Learning and Using the Arrow of Time
Donglai Wei, Joseph Lim, Andrew Zisserman, William T. Freeman
2017ICCV
Unsupervised Representation Learning by Sorting Sequences
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, Ming-Hsuan Yang
2017CVPR
Self-Supervised Video Representation Learning With Odd-One-Out Networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, Stephen Gould
2016ECCV
Shuffle and Learn: Unsupervised Learning Using Temporal Order Verification
Ishan Misra, C. Lawrence Zitnick, Martial Hebert