Agentic video search
The model can dynamically search, scan, and inspect target video segments instead of processing every clip at a fixed frame rate.
Agentic video understanding with Gemini enables selective video analysis for better accuracy and lower token usage via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.
Agentic video understanding is a Gemini video analysis mode that lets the model actively decide what parts of a video to inspect, rather than relying on fixed frame sampling. Google says this approach improves accuracy while lowering token usage and cost for long-form and detail-heavy video tasks.
The launch covers Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Developers can use the feature for video uploads and YouTube videos through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, with processing set to "agentic". The same post says the feature uses standard Gemini API token pricing with no additional feature fee.
The model can dynamically search, scan, and inspect target video segments instead of processing every clip at a fixed frame rate.
It uses Gemini’s native video tools across visual frames, audio, and transcripts to improve analysis quality.
Developers can request video analysis through the Gemini API in Google AI Studio or the Gemini Enterprise Agent Platform.
The feature supports video uploads and YouTube videos, which broadens the set of inputs available for analysis.
The workflow is designed to reduce token usage and cost compared with static frame sampling.
The page says developers enable the capability by setting processing to "agentic," reducing the need to manually implement their own video search loop.
Find split-second changes or tight cut boundaries that are hard to catch with one-frame-per-second sampling, such as edits in recorded content or short state changes in a process video.
Search across hours of footage for a specific answer or event without spending tokens on full static coverage of every frame.
Inspect fast motion, subtle artifacts, or irregular behavior by resampling selected windows at higher frame rates when the model needs more detail.
Track repeated actions or objects over time, where counting accuracy matters more than simply summarizing the clip.
Use the model as a more efficient first-pass analyzer for uploaded videos or YouTube content, then refine the result with a targeted follow-up query if needed.
It is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. The launch post also says it will roll out to the Gemini app across Flash and Flash-Lite models soon.
The launch post says you enable it by setting video processing to "agentic" in the API configuration.
The launch announcement names Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite as the models that support it at launch.
The page says it uses standard Gemini API token pricing and that there is no additional feature fee.
The source describes it as an active video analysis workflow that can search, scan, and inspect target segments across frames, audio, and transcripts. It is intended for tasks such as moment retrieval, anomaly detection, and counting.
AakarDev AI helps teams manage AI provider access, project setup, logs, and analytics in one dashboard. BYOK support included.
DeepMotion is a web-based AI motion capture and 3D animation platform with Animate 3D for video-to-animation and SayMotion for text-to-animation.
ByteAsk is a terminal-first AI coding agent for C and C++ that edits repos and verifies changes with compilers, debuggers, sanitizers, and tests.
CreateOS Sandbox is an isolated compute environment for running code and agent workloads in Firecracker micro-VMs with private networking and SDK, CLI, or MCP control.
hob is an independent workspace for coding agents, with local control over sessions, terminals, history, routing, and follow-up work.
Ably Chat is a chat API platform for custom realtime chat apps, with rooms, typing indicators, presence, reactions, message updates and usage-based pricing.