Publications

publications by categories in reversed chronological order. generated by jekyll-scholar.

2026

  1. Forget, Anticipate and Adapt: Test Time Training for Long Videos
    Rajat Modi, S. Noel, Xin Liang, and Yogesh S. Rawat
    In European Conference on Computer Vision (ECCV), 2026
  2. Robust Onion: Peeling Open Vocab Object Detectors Under Noise
    Priyank Pathak, M. Karuppasamy, A. Baranwal, Shruti Vyas, and Yogesh S. Rawat
    In European Conference on Computer Vision (ECCV), 2026
  3. CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
    Shresth Grover, Priyank Pathak, Akash Kumar, and Yogesh S. Rawat
    In European Conference on Computer Vision (ECCV), 2026
  4. vlm.jpg
    Learning to Deny: Action Denial in Multimodal Large Language Models
    Raiyaan Abdullah, Shehreen Azad, and Yogesh Singh Rawat
    In European Conference on Computer Vision (ECCV), 2026
  5. StreamReady: Learning What to Answer and When in Long Streaming Videos
    Shehreen Azad, Vibhav Vineet, and Yogesh Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  6. Sky2Ground: A Benchmark for Site Modeling under Varying Altitude
    Zengyan Wang, Sirshapan Mitra, Rajat Modi, Hui Xian Grace Lim, and Yogesh S. Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  7. ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction
    Sirshapan Mitra and Yogesh S. Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2026
  8. RobustGait: Robustness Analysis for Appearance Based Gait Recognition
    Reeshoon Sayera, Akash Kumar, Sirshapan Mitra, Prudvi Kamtam, and Yogesh S Rawat
    In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026
  9. health.jpg
    ACuRE: Accurate Continuity-Regularized SpO2 Estimation Using Liquid Time-Constant Networks
    Shahzad Ahmad, Divya Mishra, Sania Bano, Sukalpa Chanda, and Yogesh Singh Rawat
    In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2026
  10. vlm.jpg
    VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
    Chaoyang Wang, Wenrui Bao, Sicheng Gao, Bingxin Xu, Yu Tian, Yogesh S Rawat, Yunhao Ge, and Yuzhang Shang
    arXiv preprint arXiv:2603.14523, 2026
  11. vlm.jpg
    VISTA: A Benchmark for Spatio-Temporal Interaction Understanding in Videos
    Alejandro Aparcedo, Akash Kumar, Aaryan Garg, Dalton Pham, Wen-Kai Chen, Anirudh Bharadwaj, Aman Chadha, and Yogesh Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2026
  12. molecule.jpg
    MolSight: Molecular Property Prediction with Images
    Aaditya Baranwal, Akshaj Gupta, Yogesh S Rawat, and Shruti Vyas
    arXiv preprint arXiv:2605.10157, 2026

2025

  1. MolVision: Molecular Property Prediction with Vision Language Models
    Deepan Adak, Yogesh Singh Rawat, and Shruti Vyas
    In Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2025
  2. Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement
    Priyank Pathak and Yogesh S. Rawat
    In IEEE/CVF International Conference on Computer Vision (ICCV), 2025
  3. DisenQ: Disentangling Q-Former for Activity-Biometrics
    Shehreen Azad and Yogesh Rawat
    In IEEE/CVF International Conference on Computer Vision (ICCV), 2025
  4. Punching Bag vs. Punching Person: Motion Transferability in Videos
    Raiyaan Abdullah, Jared Claypoole, Michael Cogswell, Ajay Divakaran, and Yogesh Rawat
    In IEEE/CVF International Conference on Computer Vision (ICCV), 2025
  5. STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
    Aaryan Garg, Akash Kumar, and Yogesh Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
  6. HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
    Shehreen Azad, Vibhav Vineet, and Yogesh Singh Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
  7. DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-ID
    Xin Liang and Yogesh Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
  8. LR0.FM: Low-Resolution Zero-Shot Classification Benchmark for Foundation Models
    Priyank Pathak, Shyam Marjit, Shruti Vyas, and Yogesh S. Rawat
    In International Conference on Learning Representations (ICLR), 2025
  9. Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
    Akash Kumar, Zsolt Kira, and Yogesh Singh Rawat
    In International Conference on Learning Representations (ICLR), 2025
  10. T2L: Efficient Zero-Shot Action Recognition with Temporal Token Learning
    Shahzad Ahmad, Sukalpa Chanda, and Yogesh S. Rawat
    Transactions on Machine Learning Research (TMLR), 2025
  11. Stable Mean Teacher for Semi-supervised Video Action Detection
    Akash Kumar, Sirshapan Mitra, and Yogesh Singh Rawat
    In AAAI Conference on Artificial Intelligence, 2025
  12. solar.jpg
    Advancing automatic photovoltaic defect detection using semi-supervised semantic segmentation of electroluminescence images
    Abhishek Jha, Yogesh Rawat, and Shruti Vyas
    Engineering Applications of Artificial Intelligence, 2025
  13. Understanding Depth and Height Perception in Large Visual-Language Models
    Shehreen Azad, Yash Jain, Rishit Garg, Yogesh S Rawat, and Vibhav Vineet
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2025
  14. Scaling Open-Vocabulary Action Detection
    Zhen Hao Sia and Yogesh Singh Rawat
    In IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2025
  15. video-action.jpg
    A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
    Akash Kumar, Ashlesha Kumar, Vibhav Vineet, and Yogesh S Rawat
    In Proceedings of the Computer Vision and Pattern Recognition Conference, 2025
  16. health.jpg
    PULSE: Physiological Understanding with Liquid Signal Extraction
    Shahzad Ahmad, Sania Bano, Sachin Verma, Yogesh Singh Rawat, Sukalpa Chanda, Santosh Kumar Vipparthi, and Subrahmanyam Murala
    In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2025
  17. video-action.jpg
    Coarse attribute prediction with task agnostic distillation for real world clothes changing reid
    Priyank Pathak and Yogesh S Rawat
    arXiv preprint arXiv:2505.12580, 2025
  18. iSafetyBench: A video-language benchmark for safety in industrial environment
    Raiyaan Abdullah, Yogesh Singh Rawat, and Shruti Vyas
    In IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2025
  19. Re: Verse-Can Your VLM Read a Manga?
    Aaditya Baranwal, Madhav Kataria, Naitik Agrawal, Yogesh S Rawat, and Shruti Vyas
    In IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2025
  20. video-action.jpg
    OmViD: Omni-supervised active learning for video action detection
    Aayush Rana, Akash Kumar, Vibhav Vineet, and Yogesh S Rawat
    In IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2025
  21. GaitCrafter: Diffusion Model for Biometric Preserving Gait Synthesis
    Sirshapan Mitra and Yogesh Singh Rawat
    In IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2025

2024

  1. Asynchronous Perception Machine for Efficient Test-Time-Training
    Rajat Modi and Yogesh Rawat
    In Advances in Neural Information Processing Systems (NeurIPS), 2024
  2. Navigating Hallucinations for Reasoning of Unintentional Activities
    Shresth Grover, Vibhav Vineet, and Yogesh S. Rawat
    In Findings of the Association for Computational Linguistics: EMNLP, 2024
  3. Activity-Biometrics: Person Identification from Daily Activities
    Shehreen Azad and Yogesh Singh Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
  4. survey.jpg
    Foundation models for video understanding: A survey
    Neelu Madan, Andreas Møgelmose, Rajat Modi, Yogesh S Rawat, and Thomas B Moeslund
    arXiv preprint arXiv:2405.03770, 2024
  5. vlm.jpg
    Probing conceptual understanding of large visual-language models
    Madeline Schiappa, Raiyaan Abdullah, Shehreen Azad, Jared Claypoole, Michael Cogswell, Ajay Divakaran, and Yogesh Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2024
  6. robustness.jpg
    Robustness analysis on foundational segmentation models
    Madeline Chantry Schiappa, Shehreen Azad, Sachidanand Vs, Yunhao Ge, Ondrej Miksik, Yogesh S Rawat, and Vibhav Vineet
    In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2024
  7. Semi-supervised active learning for video action detection
    Ayush Singh, Aayush J Rana, Akash Kumar, Shruti Vyas, and Yogesh Singh Rawat
    In Proceedings of the AAAI Conference on Artificial Intelligence, 2024
  8. AirSketch: Generative Motion to Sketch
    Hui Xian Grace Lim, Xuanming Cui, Ser-Nam Lim, and Yogesh S Rawat
    In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

2023

  1. Self-supervised learning for videos: A survey
    Madeline C Schiappa, Yogesh S Rawat, and Mubarak Shah
    ACM Computing Surveys, 2023
  2. A large-scale robustness analysis of video action recognition models
    Madeline Chantry Schiappa, Naman Biyani, Prudvi Kamtam, Shruti Vyas, Hamid Palangi, Vibhav Vineet, and Yogesh S Rawat
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023
  3. Hybrid active learning via deep clustering for video action detection
    Aayush J Rana and Yogesh S Rawat
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023
  4. robustness.jpg
    Efficiently robustify pre-trained models
    Nishant Jain, Harkirat Behl, Yogesh Singh Rawat, and Vibhav Vineet
    In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023
  5. robustness.jpg
    PRAT: PRofiling Adversarial aTtacks
    Rahul Ambati, Naveed Akhtar, Ajmal Mian, and Yogesh S Rawat
    In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023
  6. On occlusions in video action detection: Benchmark datasets and training recipes
    Rajat Modi, Vibhav Vineet, and Yogesh Rawat
    Advances in Neural Information Processing Systems, 2023
  7. Revealing the unseen: Benchmarking video action recognition under occlusion
    Shresth Grover, Vibhav Vineet, and Yogesh Rawat
    Advances in Neural Information Processing Systems, 2023
  8. EZ-CLIP: Efficient Zero-shot Video Action Recognition
    Shahzad Ahmad, Sukalpa Chanda, and Yogesh S Rawat
    arXiv preprint arXiv:2312.08010, 2023

2022

  1. surveillance.jpg
    GabriellaV2: Towards better generalization in surveillance videos for action detection
    Ishan Dave, Zacchaeus Scheffer, Akash Kumar, Sarah Shiraz, Yogesh Singh Rawat, and Mubarak Shah
    In IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW), 2022
  2. End-to-End Semi-Supervised Learning for Video Action Detection
    Akash Kumar and Yogesh Singh Rawat
    In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022
  3. Don’t Pour Cereal into Coffee: Differentiable Temporal Logic for Temporal Action Segmentation
    Ziwei Xu, Yogesh S Rawat, Yongkang Wong, Mohan Kankanhalli, and Mubarak Shah
    In Advances in Neural Information Processing Systems, 2022
  4. Robustness analysis of video-language models against visual and language perturbations
    Madeline Schiappa, Shruti Vyas, Hamid Palangi, Yogesh Rawat, and Vibhav Vineet
    Advances in Neural Information Processing Systems, 2022
  5. vlm.jpg
    Svgraph: Learning semantic graphs from instructional videos
    Madeline C Schiappa and Yogesh S Rawat
    2022 IEEE Eighth International Conference on Multimedia Big Data (BigMM), 2022
  6. Are all Frames Equal? Active Sparse Labeling for Video Action Detection
    Aayush Rana and Yogesh S Rawat
    In Advances in Neural Information Processing Systems, 2022
  7. Video Action Detection: Analysing Limitations and Challenges
    Rajat Modi, Aayush Jung Rana, Akash Kumar, Praveen Tirupattur, Shruti Vyas, Yogesh Singh Rawat, and Mubarak Shah
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2022

2021

  1. In defense of pseudo-labeling: An uncertainty-aware pseudo-label selection framework for semi-supervised learning
    Mamshad Nayeem Rizve, Kevin Duarte, Yogesh S Rawat, and Mubarak Shah
    In The International Conference on Learning Representations (ICLR), 2021
  2. Modeling Multi-Label Action Dependencies for Temporal Action Localization
    Praveen Tirupattur, Kevin Duarte, Yogesh Rawat, and Mubarak Shah
    In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021
  3. video-action.jpg
    We Don’t Need Thousand Proposals: Single Shot Actor-Action Detection in Videos
    Aayush J Rana and Yogesh S Rawat
    In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2021
  4. video-action.jpg
    Unsupervised Discriminative Embedding for Sub-Action Learning in Complex Activities
    Sirnam Swetha, Hilde Kuehne, Yogesh S Rawat, and Mubarak Shah
    In 2021 IEEE International Conference on Image Processing, 2021
  5. video-action.jpg
    Plm: Partial label masking for imbalanced multi-label classification
    Kevin Duarte, Yogesh Rawat, and Mubarak Shah
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021
  6. video-action.jpg
    Novel View Video Prediction Using a Dual Representation
    Sarah Shiraz, Krishna Regmi, Shruti Vyas, Yogesh S Rawat, and Mubarak Shah
    In IEEE International Conference on Image Processing, 2021
  7. video-action.jpg
    Tinyaction challenge: Recognizing real-world low-resolution activities in videos
    Praveen Tirupattur, Aayush J Rana, Tushar Sangam, Shruti Vyas, Yogesh S Rawat, and Mubarak Shah
    arXiv preprint arXiv:2107.11494, 2021
  8. NoisyActions2M: A Multimedia Dataset for Video Understanding from Noisy Labels
    Mohit Sharma, Raj Patra, Harshal Desai, Shruti Vyas, Yogesh Rawat, and Rajiv Ratn Shah
    In ACM International Conference on Multimedia in Asia, 2021
  9. Pose-guided Generative Adversarial Net for Novel View Action Synthesis
    Xianhang Li, Junhao Zhang, Kunchang Li, Shruti Vyas, and Yogesh S Rawat
    In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021
  10. video-action.jpg
    "Knights": First Place Submission for VIPriors21 Action Recognition Challenge at ICCV 2021
    Ishan Dave, Naman Biyani, Brandon Clark, Rohit Gupta, Yogesh Rawat, and Mubarak Shah
    arXiv preprint arXiv:2110.07758, 2021
  11. LARNet: Latent Action Representation for Human Action Synthesis
    Naman Biyani, Aayush J Rana, Shruti Vyas, and Yogesh S Rawat
    In The British Machine Vision Conference (BMVC), 2021
  12. Reformulating zero-shot action recognition for multi-label actions
    Alec Kerrigan, Kevin Duarte, Yogesh Rawat, and Mubarak Shah
    Advances in Neural Information Processing Systems, 2021

2020

  1. video-action.jpg
    Multi-view Action Recognition using Cross-view Video Prediction
    Shruti Vyas, Yogesh S Rawat, and Mubarak Shah
    In European Conference on Computer Vision (ECCV), 2020
  2. video-action.jpg
    TinyVIRAT: Low-resolution Video Action Recognition
    Ugur Demir, Yogesh S Rawat, and Mubarak Shah
    In 25th International Conference on Pattern Recognition (ICPR), 2020
  3. robustness.jpg
    Adversarial Learning for Personalized Tag Recommendation
    Erik Quintanilla, Yogesh Rawat, Andrey Sakryukin, Mubarak Shah, and Mohan Kankanhalli
    IEEE Transactions on Multimedia, 2020
  4. capsule-net.jpg
    Visual-textual Capsule Routing for Text-based Video Segmentation
    Bruce McIntosh, Kevin Duarte, Yogesh S Rawat, and Mubarak Shah
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020
  5. surveillance.jpg
    Gabriella: An Online System for Real-Time Activity Detection in Untrimmed Security Videos
    Mamshad Nayeem Rizve, Ugur Demir, Praveen Tirupattur, Aayush Jung Rana, Kevin Duarte, Ishan Dave, Yogesh Singh Rawat, and Mubarak Shah
    In 25th International Conference on Pattern Recognition (ICPR), 2020
  6. video-action.jpg
    A Recurrent Transformer Network for Novel View Action Synthesis
    Kara Marie Schatz, Erik Quintanilla, Shruti Vyas, and Yogesh S Rawat
    In Proceedings of the European Conference on Computer Vision (ECCV), 2020
  7. video-action.jpg
    View-invariant action recognition
    Yogesh S Rawat and Shruti Vyas
    In Computer Vision: A Reference Guide, 2020

2019

  1. capsule-net.jpg
    CapsuleVOS: Semi-Supervised Video Object Segmentation Using Capsule Routing
    Kevin Duarte, Yogesh S Rawat, and Mubarak Shah
    In Proceedings of the IEEE International Conference on Computer Vision, 2019
  2. photography.jpg
    Photography and exploration of tourist locations based on optimal foraging theory
    Yogesh Singh Rawat, Mubarak Shah, and Mohan S Kankanhalli
    IEEE Transactions on Circuits and Systems for Video Technology, 2019

2018

  1. capsule-net.jpg
    VideoCapsuleNet: A Simplified Network for Action Detection
    Kevin Duarte, Yogesh S Rawat, and Mubarak Shah
    In Advances in Neural Information Processing Systems, 2018
  2. brain-bci.jpg
    ThoughtViz: Visualizing Human Thoughts Using Generative Adversarial Network
    Praveen Tirupattur, Yogesh Singh Rawat, Concetto Spampinato, and Mubarak Shah
    In Proceedings of the 2018 ACM on Multimedia Conference, 2018
  3. video-action.jpg
    Time-aware and view-aware video rendering for unsupervised representation learning
    Shruti Vyas, Yogesh S Rawat, and Mubarak Shah
    arXiv preprint arXiv:1811.10699, 2018

2017

  1. photography.jpg
    A spring-electric graph model for socialized group photography
    Yogesh Singh Rawat, Mingli Song, and Mohan S Kankanhalli
    IEEE Transactions on Multimedia, 2017

2016

  1. photography.jpg
    ConTagNet: Exploiting user context for image tag recommendation
    Yogesh Singh Rawat and Mohan S Kankanhalli
    In Proceedings of the 24th ACM international conference on Multimedia, 2016
  2. photography.jpg
    ClickSmart: A Context-Aware Viewpoint Recommendation System for Mobile Photography
    Yogesh Rawat and Mohan Kankanhalli
    2016
  3. photography.jpg
    Real Time Assistance in Photography Using Social Media
    YOGESH SINGH RAWAT
    National University of Singapore, 2016

2015

  1. photography.jpg
    Context-aware photography learning for smart mobile devices
    Yogesh Singh Rawat and Mohan S Kankanhalli
    ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 2015
  2. photography.jpg
    Real-time assistance in multimedia capture using social media
    Yogesh Singh Rawat
    In Proceedings of the 23rd ACM international conference on Multimedia, 2015

2014

  1. photography.jpg
    Context-based photography learning using crowdsourced images and social media
    Yogesh Singh Rawat and Mohan S Kankanhalli
    In Proceedings of the 22nd ACM international conference on Multimedia, 2014
  2. photography.jpg
    Mode of teaching based segmentation and annotation of video lectures
    Yogesh Singh Rawat, Chidansh Bhatt, and Mohan S Kankanhalli
    In 2014 12th International Workshop on Content-Based Multimedia Indexing (CBMI), 2014

2008

  1. handwriting.jpg
    An analytic scheme for online handwritten Bangla cursive word recognition
    U Bhattacharya, A Nigam, YS Rawat, and SK Parui
    Proc. of the 11th ICFHR, 2008