Long Video Understanding Streaming and long-form video: efficient memory, hierarchical representations, and knowing what to answer and when. Human Actions Recognizing, localizing, and detecting human actions in video, from fully-supervised to weakly- and semi-supervised settings. Robustness Making vision and vision-language models robust to noise, occlusion, adversarial attacks, and distribution shift. Vision-Language Models Probing, grounding, and extending vision-language models for spatial, temporal, and safety-aware reasoning. Multimodal Learning Combining vision with other modalities, from generative sketch synthesis to molecular property prediction. Biometrics Identifying people from video, including clothes-changing re-identification, gait, and activity-based biometrics. 3D Vision Reconstructing and understanding 3D scenes from imagery captured at varying altitudes and viewpoints. Applications Taking vision and multimodal learning into practice, from healthcare and industry to early work in social photography. Generative AI Video and image synthesis with diffusion models, GANs, and generative video prediction. Label-Efficient Learning Getting strong performance from limited annotations, through semi-supervised, weakly-supervised, and active learning. Efficient Learning Reducing the compute and data cost of vision models, from test-time training to low-resolution robustness. Video Understanding Foundational video representation learning, cross-view and novel-view synthesis, and video segmentation.