Long Video Understanding

Streaming and long-form video: efficient memory, hierarchical representations, and knowing what to answer and when.

References

2026

  1. StreamReady: Learning What to Answer and When in Long Streaming Videos
    Shehreen Azad, Vibhav Vineet, and Yogesh Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  2. Forget, Anticipate and Adapt: Test Time Training for Long Videos
    Rajat Modi, S. Noel, Xin Liang, and Yogesh S. Rawat
    In European Conference on Computer Vision (ECCV), 2026

2025

  1. HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
    Shehreen Azad, Vibhav Vineet, and Yogesh Singh Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
  2. video-action.jpg
    A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
    Akash Kumar, Ashlesha Kumar, Vibhav Vineet, and Yogesh S Rawat
    In Proceedings of the Computer Vision and Pattern Recognition Conference, 2025

2024

  1. survey.jpg
    Foundation models for video understanding: A survey
    Neelu Madan, Andreas Møgelmose, Rajat Modi, Yogesh S Rawat, and Thomas B Moeslund
    arXiv preprint arXiv:2405.03770, 2024

2023

  1. Self-supervised learning for videos: A survey
    Madeline C Schiappa, Yogesh S Rawat, and Mubarak Shah
    ACM Computing Surveys, 2023

2022

  1. vlm.jpg
    Svgraph: Learning semantic graphs from instructional videos
    Madeline C Schiappa and Yogesh S Rawat
    2022 IEEE Eighth International Conference on Multimedia Big Data (BigMM), 2022