Introducing agentic video understanding with Gemini

Our new agentic feature for video analysis cuts token consumption by up to 88%, reduces costs by up to 66%, and boosts quality by up to 7%. Google just launched agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. This feature allows the model to dynamically scan video segments, which improves accuracy…
Opening of the original on Google DeepMind Blog
Summary
Google Gemini now offers agentic video understanding, cutting token use up to 88% and costs up to 66% while improving accuracy by 7%. This feature allows models to dynamically scan video segments, improving performance on long-form content. It is available via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, with plans to integrate into the Gemini app and YouTube's 'Ask YouTube' feature.
Why it matters
This update makes video analysis significantly more efficient and accurate for developers using Gemini. By dynamically scanning video instead of processing fixed frames, it reduces costs and token consumption, especially for long videos. This impacts developers building applications that rely on video understanding. While competitors are also advancing multimodal capabilities, Gemini's agentic approach for video appears to offer a distinct efficiency advantage. Future developments will likely focus on broader integration into Google products and further performance enhancements.
Related: OpenAI: Fable 5.1 just smoked ASTRA... · Anthropic: Claude Fable AI Is Much Stranger Than The Headlines Suggest · Microsoft: Orchard: An open framework for scalable agentic AI · Meta: Meta Muse Code & Muse Spark Course – Build AI Agents, APIs, and Full-Stack Apps · xAI: Ajeya Cotra – "This might be the clearest warning shot we ever get"
Rated middle: a real update, not a headline event.
- Google says its new Gemini 3.8 Flash model ‘works harder’ but might cost moreThe Verge AI · Web · Sep 2, 2026
- llm-gemini 0.34Simon Willison · Web · Sep 2, 2026
- LWiAI Podcast #255 - Gemini 3.7, Jalapeño, Qwen 3.8, DronesLast Week in AI · Web · Aug 31, 2026
- Google’s latest AI weather model gives you no excuse to forget your umbrellaTechCrunch AI · Web · Sep 3, 2026
- Claude Fable 5.1 made me a really nice animated pelicanSimon Willison · Web · Sep 1, 2026
- The Most Overhyped and Underhyped New AI ModelsMatt Wolfe · Web · Sep 2, 2026
- Agentic video understanding in GeminiGoogle for Developers on YouTube
- Pairing Google Antigravity with Gemini 3.7 Flash solves notable multi-agent math and engineering problems.The Keyword: Developers
- LWiAI Podcast #255 - Gemini 3.7, Jalapeño, Qwen 3.8, DronesLast Week in AI
- Koray Kavukcuoglu on frontier models, coding agents, and building AGIGoogle for Developers on YouTube
- Intelligence EXPLOSION: Harness Engineering with Pi Agent, Deepseek, and GeminiIndyDevDan
- When millions of AI agents meetGoogle DeepMind on YouTube
Questions people ask
- How does agentic video understanding improve efficiency?
- It allows Gemini models to dynamically scan and inspect specific video segments across frames, audio, and transcripts, rather than ingesting the entire video at a fixed frame rate. This targeted approach reduces token usage and costs.
- What are the benefits for developers?
- Developers can achieve higher accuracy and significantly lower costs for video analysis, especially with long-form content. It also reduces development overhead by automating the process of identifying relevant video moments.
- Where is agentic video understanding available?
- It is available via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. It will also roll out to the Gemini app and power YouTube's 'Ask YouTube' feature.
More from Google DeepMind Blog 99 more
Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to Google DeepMind Blog.









