Vision-Language Models

Probing, grounding, and extending vision-language models for spatial, temporal, and safety-aware reasoning.

References

2026

  1. CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
    Shresth Grover, Priyank Pathak, Akash Kumar, and Yogesh S. Rawat
    In European Conference on Computer Vision (ECCV), 2026
  2. vlm.jpg
    VISTA: A Benchmark for Spatio-Temporal Interaction Understanding in Videos
    Alejandro Aparcedo, Akash Kumar, Aaryan Garg, Dalton Pham, Wen-Kai Chen, Anirudh Bharadwaj, Aman Chadha, and Yogesh Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2026
  3. vlm.jpg
    VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
    Chaoyang Wang, Wenrui Bao, Sicheng Gao, Bingxin Xu, Yu Tian, Yogesh S Rawat, Yunhao Ge, and Yuzhang Shang
    arXiv preprint arXiv:2603.14523, 2026

2025

  1. LR0.FM: Low-Resolution Zero-Shot Classification Benchmark for Foundation Models
    Priyank Pathak, Shyam Marjit, Shruti Vyas, and Yogesh S. Rawat
    In International Conference on Learning Representations (ICLR), 2025
  2. Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
    Akash Kumar, Zsolt Kira, and Yogesh Singh Rawat
    In International Conference on Learning Representations (ICLR), 2025
  3. STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
    Aaryan Garg, Akash Kumar, and Yogesh Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
  4. Understanding Depth and Height Perception in Large Visual-Language Models
    Shehreen Azad, Yash Jain, Rishit Garg, Yogesh S Rawat, and Vibhav Vineet
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2025
  5. iSafetyBench: A video-language benchmark for safety in industrial environment
    Raiyaan Abdullah, Yogesh Singh Rawat, and Shruti Vyas
    In IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2025
  6. Re: Verse-Can Your VLM Read a Manga?
    Aaditya Baranwal, Madhav Kataria, Naitik Agrawal, Yogesh S Rawat, and Shruti Vyas
    In IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2025

2024

  1. vlm.jpg
    Probing conceptual understanding of large visual-language models
    Madeline Schiappa, Raiyaan Abdullah, Shehreen Azad, Jared Claypoole, Michael Cogswell, Ajay Divakaran, and Yogesh Rawat
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2024
  2. Navigating Hallucinations for Reasoning of Unintentional Activities
    Shresth Grover, Vibhav Vineet, and Yogesh S. Rawat
    In Findings of the Association for Computational Linguistics: EMNLP, 2024