AI News Feed
Market watch
Large Language Models

Gemini introduces agentic video understanding with lower token use and better accuracy

Gemini now offers agentic video understanding, with up to 88% lower token usage and 7% better accuracy, per Android Authority.

The feature is available in Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Google said the approach cuts token usage by up to 88 percent and offers up to 7 percent better accuracy than earlier video processing methods.

Previously, Gemini could only process videos in what Google called a static way: splitting the footage into individual frames, which caused slower performance and higher costs. With agentic video understanding, Gemini can decide what to watch and at what speed, and choose between frames, audio, and transcripts to analyze videos more efficiently. It can pinpoint split-second changes, answer complex questions about multi-hour videos, inspect videos for visual artifacts, and count and track physical movements and objects.

The new capability is currently available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Google also said the feature will roll out to the Gemini app soon, and will power YouTube's Ask YouTube feature, easing video analysis for creators.