Agentic video understanding with Gemini icon

Agentic video understanding with Gemini

Agentic video understanding with Gemini enables selective video analysis for better accuracy and lower token usage via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.

Agentic video understanding with Gemini

Overview

Agentic video understanding is a Gemini video analysis mode that lets the model actively decide what parts of a video to inspect, rather than relying on fixed frame sampling. Google says this approach improves accuracy while lowering token usage and cost for long-form and detail-heavy video tasks.

The launch covers Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Developers can use the feature for video uploads and YouTube videos through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, with processing set to "agentic". The same post says the feature uses standard Gemini API token pricing with no additional feature fee.

Core capabilities

Agentic video search

The model can dynamically search, scan, and inspect target video segments instead of processing every clip at a fixed frame rate.

Multi-modal video inspection

It uses Gemini’s native video tools across visual frames, audio, and transcripts to improve analysis quality.

API and platform access

Developers can request video analysis through the Gemini API in Google AI Studio or the Gemini Enterprise Agent Platform.

Video uploads and YouTube support

The feature supports video uploads and YouTube videos, which broadens the set of inputs available for analysis.

Lower token and cost usage

The workflow is designed to reduce token usage and cost compared with static frame sampling.

Simple configuration switch

The page says developers enable the capability by setting processing to "agentic," reducing the need to manually implement their own video search loop.

Practical use cases

  • Moment retrieval and precise editing

    Find split-second changes or tight cut boundaries that are hard to catch with one-frame-per-second sampling, such as edits in recorded content or short state changes in a process video.

  • Long-form video question answering

    Search across hours of footage for a specific answer or event without spending tokens on full static coverage of every frame.

  • Anomaly detection in video streams

    Inspect fast motion, subtle artifacts, or irregular behavior by resampling selected windows at higher frame rates when the model needs more detail.

  • Counting actions and objects

    Track repeated actions or objects over time, where counting accuracy matters more than simply summarizing the clip.

  • Video workflows in Gemini apps and API projects

    Use the model as a more efficient first-pass analyzer for uploaded videos or YouTube content, then refine the result with a targeted follow-up query if needed.

Pros and Cons

Pros

  • Can reduce token consumption by up to 88% and costs by up to 66% compared with static processing, according to Google’s benchmarks.
  • Improves accuracy by up to 7% on standard video analysis benchmarks.
  • Handles long-form video more efficiently by searching only the segments needed instead of ingesting every frame at a fixed rate.
  • Supports practical tasks such as sub-second moment retrieval, anomaly detection, and counting.
  • Available through familiar Gemini workflows in Google AI Studio and the Gemini Enterprise Agent Platform without an additional feature fee.

Cons

  • The launch post does not provide a detailed limitations list or supported-format matrix beyond video uploads and YouTube videos.
  • The strongest gains are described for long-form video and benchmark-style tasks, so shorter or simpler clips may see less visible benefit.
  • Availability is described for specific Gemini models at launch, so support may vary by model and product surface.

FAQ

Where can developers use agentic video understanding?

It is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. The launch post also says it will roll out to the Gemini app across Flash and Flash-Lite models soon.

How do you turn it on?

The launch post says you enable it by setting video processing to "agentic" in the API configuration.

Which Gemini models support it?

The launch announcement names Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite as the models that support it at launch.

Does it have separate pricing?

The page says it uses standard Gemini API token pricing and that there is no additional feature fee.

What kinds of tasks is it designed for?

The source describes it as an active video analysis workflow that can search, scan, and inspect target segments across frames, audio, and transcripts. It is intended for tasks such as moment retrieval, anomaly detection, and counting.

Quick Facts

Category
Developer Tool
Product domain
blog.google
Primary task
Agentic video analysis
Availability
Google AI Studio and Gemini Enterprise Agent Platform
Launch models
Gemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite
Pricing shape
Standard Gemini API token pricing; no additional feature fee