IJMEMES logo

International Journal of Mathematical, Engineering and Management Sciences

ISSN: 2455-7749 . Open Access


Learning Unsupervised Visual Representations using 3D Convolutional Autoencoder with Temporal Contrastive Modeling for Video Retrieval

Learning Unsupervised Visual Representations using 3D Convolutional Autoencoder with Temporal Contrastive Modeling for Video Retrieval

Vidit Kumar
Department of Computer Science and Engineering, Graphic Era Deemed to be University Dehradun, India.

Vikas Tripathi
Department of Computer Science and Engineering, Graphic Era Deemed to be University, Dehradun, India.

Bhaskar Pant
Department of Computer Science and Engineering, Graphic Era Deemed to be University, Dehradun, India.

DOI https://doi.org/10.33889/IJMEMS.2022.7.2.018

Received on September 04, 2021
  ;
Accepted on January 25, 2022

Abstract

The rapid growth of tag-free user-generated videos (on the Internet), surgical recorded videos, and surveillance videos has necessitated the need for effective content-based video retrieval systems. Earlier methods for video representations are based on hand-crafted, which hardly performed well on the video retrieval tasks. Subsequently, deep learning methods have successfully demonstrated their effectiveness in both image and video-related tasks, but at the cost of creating massively labeled datasets. Thus, the economic solution is to use freely available unlabeled web videos for representation learning. In this regard, most of the recently developed methods are based on solving a single pretext task using 2D or 3D convolutional network. However, this paper designs and studies a 3D convolutional autoencoder (3D-CAE) for video representation learning (since it does not require labels). Further, this paper proposes a new unsupervised video feature learning method based on joint learning of past and future prediction using 3D-CAE with temporal contrastive learning. The experiments are conducted on UCF-101 and HMDB-51 datasets, where the proposed approach achieves better retrieval performance than state-of-the-art. In the ablation study, the action recognition task is performed by fine-tuning the unsupervised pre-trained model where it outperforms other methods, which further confirms the superiority of our method in learning underlying features. Such an unsupervised representation learning approach could also benefit the medical domain, where it is expensive to create large label datasets.

Keywords- Contrastive learning, Convolutional autoencoder, Content-based search, Deep learning, Video retrieval, Future prediction, Unsupervised learning

Citation

Kumar, V., Tripathi, V., & Pant, B. (2022). Learning Unsupervised Visual Representations using 3D Convolutional Autoencoder with Temporal Contrastive Modeling for Video Retrieval. International Journal of Mathematical, Engineering and Management Sciences, 7(2), 272-287. https://doi.org/10.33889/IJMEMS.2022.7.2.018.