[CVPR2021] SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning over Traffic Events
-
Updated
Aug 19, 2024 - JavaScript
[CVPR2021] SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning over Traffic Events
[EMNLP 2023] TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding
Codes and Models for COSA: Concatenated Sample Pretrained Vision-Language Foundation Model
Unifying the Video and Question Attentions for Open-Ended Video Question Answering
Video Question Answering via Hierarchical Spatio-Temporal Attention Networks
The teaches you to integrate text, images, and videos into applications using Gemini's state-of-the-art multimodal models. Learn advanced prompting techniques, cross-modal reasoning, and how to extend Gemini's capabilities with real-time data and API integration.
Add a description, image, and links to the video-qa topic page so that developers can more easily learn about it.
To associate your repository with the video-qa topic, visit your repo's landing page and select "manage topics."