Google's new Gemini agentic video understanding cuts token consumption by up to 88%
Google introduced agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite models. The new capability dynamically searches video segments, reducing token consumption by up to 88% and costs by up to 66% while improving accuracy by up to 7%. The feature is available immediately via the Gemini API at no additional cost and supports uploads and YouTube videos.
I don't generally consider Gemini the best model available and it lags significantly behind competitors from OpenAI and Anthropic. However, it seems to be finding use cases where it makes sense. I'll definitely try it.
How does agentic video understanding differ from static processing?
Agentic video understanding allows the model to dynamically decide which video segments to analyze, at what speed, and through which modality (frames, audio, or transcript), whereas static processing ingests video at a fixed frame rate. This results in significantly lower token consumption without quality loss.
What real-world use cases does agentic video analysis enable?
The feature enables sub-second moment retrieval for precise editing, efficient searching across multi-hour videos, anomaly detection through adaptive frame-rate sampling, and accurate counting of repeated actions or objects.
What development benefits does this provide?
Developers no longer need to manually implement logic for dynamic video traversal—Gemini now handles this automatically through agentic loops, significantly reducing development overhead.
- Google introduces Gemini 3.7 Flash — blog.google 83 % match
- LFM2.5-2.6B: small and capable local AI model — liquid.ai 73 % match
- Gemini Spark now integrates with Chrome — blog.google 72 % match
- Gemini7
- Gemini 3.7 Flash
- Gemini 3.6 Flash3
- Gemini 3.5 Flash-Lite
- Google37
- Google AI Studio3
- Gemini Enterprise Agent Platform
- Google DeepMind3