Multimodal Learning

Combining vision with other modalities, from generative sketch synthesis to molecular property prediction.

References

2026

  1. molecule.jpg
    MolSight: Molecular Property Prediction with Images
    Aaditya Baranwal, Akshaj Gupta, Yogesh S Rawat, and Shruti Vyas
    arXiv preprint arXiv:2605.10157, 2026

2025

  1. MolVision: Molecular Property Prediction with Vision Language Models
    Deepan Adak, Yogesh Singh Rawat, and Shruti Vyas
    In Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2025

2024

  1. AirSketch: Generative Motion to Sketch
    Hui Xian Grace Lim, Xuanming Cui, Ser-Nam Lim, and Yogesh S Rawat
    In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

2021

  1. NoisyActions2M: A Multimedia Dataset for Video Understanding from Noisy Labels
    Mohit Sharma, Raj Patra, Harshal Desai, Shruti Vyas, Yogesh Rawat, and Rajiv Ratn Shah
    In ACM International Conference on Multimedia in Asia, 2021

2020

  1. capsule-net.jpg
    Visual-textual Capsule Routing for Text-based Video Segmentation
    Bruce McIntosh, Kevin Duarte, Yogesh S Rawat, and Mubarak Shah
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020
  2. robustness.jpg
    Adversarial Learning for Personalized Tag Recommendation
    Erik Quintanilla, Yogesh Rawat, Andrey Sakryukin, Mubarak Shah, and Mohan Kankanhalli
    IEEE Transactions on Multimedia, 2020

2018

  1. brain-bci.jpg
    ThoughtViz: Visualizing Human Thoughts Using Generative Adversarial Network
    Praveen Tirupattur, Yogesh Singh Rawat, Concetto Spampinato, and Mubarak Shah
    In Proceedings of the 2018 ACM on Multimedia Conference, 2018