NVIDIA's Metropolis Blueprint for Video Search and Summarization (VSS) version 3.3 introduces two features aimed at cutting both development and runtime costs for visual AI agents: the Build Vision Agent skill (vss-build-vision-ai) and Adaptive Efficient Video Sampling (Adaptive EVS). The update helps convert vision-language model capabilities into maintainable, cost-effective systems that combine ingestion, stream processing, event detection, retrieval, summarization, and reporting.
What VSS does and why it matters
VSS connects vision-language models (VLMs) such as NVIDIA Cosmos, large language models such as NVIDIA Nemotron, retrieval-augmented generation (RAG), and Model Context Protocol (MCP) tools to transform live and recorded video into natural-language search, visual Q&A, verified alerts and automated reporting.
A production visual AI agent typically spans multiple workflows (ingest, detection, alerts, search, summarization, reporting), and integrating these workflows creates recurring costs in development, operations, and change management.
Key updates in 3.3
Two additions in VSS 3.3 reduce costs on both sides of deployment:
- Build Vision Agent skill (vss-build-vision-ai): composes multiple VSS workflows (alerting, search, summarization, etc.) into a single application from a natural-language description and can extend a running deployment without rebuilding the whole stack.
- Adaptive EVS: reduces redundant VLM processing by pruning unchanged visual tokens per patch and batching VLM work around moments of activity. EVS was previously shipped with fixed pruning rates in vLLM and Cosmos NIM microservices; the adaptive variant in VSS 3.3 decides which tokens to keep per patch and per frame inside the real-time VLM microservice.
Reducing development cost with the Build Vision Agent skill
Previous VSS skills handled individual tasks (deployment, camera setup, summarization, search, alerts, analytics). vss-build-vision-ai organizes skills into deployment skills, operation skills, tools, and benchmarks, and composes them based on the developer's intent.
Instead of generating a deployment from scratch, the skill starts from the closest of four validated developer profiles (the Foundation) and applies the smallest delta required by the request. The four profiles are:
- base: VLM dense captioning and Q&A on clips
- alerts: real-time VLM alerting or RT-CV detection with behavior analytics and VLM alert verification
- lvs: long-video summarization
- search: object and video embeddings with agentic search
The skill minimizes duplicated infrastructure: it adds or removes only exact service keys, keeps only services that a requested capability reaches, and converges shared roles to a single instance. For example, if two capabilities require a detector, they share one detector; Kafka and Elasticsearch instances are shared with separate indices when needed. When choices are ambiguous, the skill asks a structured question.
What the skill automates:
- Maps application goals to required VSS workflows and microservices,
- Combines workflows like alerting, search, and summarization into one deployment plan,
- Reuses shared infrastructure (VIOS, Kafka, Redis, Elasticsearch, HAProxy ingress, MCP services),
- Produces a self-contained build ( _builds/<name>/override.env with the Foundation, effective Compose profiles and only changed settings; compose.yml and resolved.yml ) without modifying the repository's deploy/docker/ tree,
- Shows an architecture diagram for review, then runs validation, deployment, and readiness checks,
- Offers to deploy an agent harness (default: NemoClaw). Declining yields a headless stack driven by the VSS CLI,
- Extends a running deployment through a smaller delta that reuses existing services.
This shortens discovery, makes composition repeatable, and avoids duplicate ingestion, storage, messaging, and analytics infrastructure across workflows.
Example: building a bottling-line visual AI agent
In a demo, an agent for an orange juice bottling line watched filler and capper cameras, alerted on bottle overflows and juice spills, verified each alert with a VLM, made alert clips searchable, and generated shift reports.
Sample prompt used:
"Build a VSS vision agent for an orange juice bottling line. Use two RTSP cameras on the filler and capper. Detect bottle overflows and juice spills, verify each alert with the VLM, make alert clips searchable, and generate a shift report for the line supervisor."
The skill assembled:
- VIOS-backed ingestion for RTSP cameras and recorded clips,
- Real-time detection, tracking, captioning, or VLM-based alerts depending on the selected Foundation,
- Behavior analytics or rules for overflow, spill, and line-stoppage events,
- VLM verification that confirms alerts and explains reasoning,
- Natural-language search across validated clips and indexed video,
- Summarization and reporting for operator handoff and incident review,
- Shared messaging, storage, APIs, and observability.
On a two-GPU NVIDIA RTX PRO 6000 Blackwell host, the skill combined search and alerts by reusing services and adding only an alert bridge and a real-time VLM. FP8 Cosmos 3 Nano shared the detector’s GPU to avoid duplication. A recorded alerts build reached a live, previewable deployment in under 30 minutes.
The pattern: ingest video once, share evidence across workflows, and provide operators with natural-language search, alerts, summaries, and reports.
How Adaptive EVS reduces runtime cost
A running agent’s cameras produce continuous video, but most of each frame is unchanged (for example, the filler, guards, floor). VLMs still read those frames for context, so much compute repeats on regions identical to the previous frame. That VLM processing is a major runtime cost driver; Adaptive EVS targets that waste.
Adaptive EVS changes the pipeline by:
- Dynamic pruning: comparing each patch to the prior frame via cosine similarity and dropping unchanged patches before they reach the language model.
- Event-aware batching: using token retention to indicate activity—clips above ~70% retention are batched as events, clips below ~30% are dropped or flushed, and the remainder run normally.
Performance impact (reported tests):
On an NVIDIA RTX PRO 6000 Blackwell running Cosmos 3 Super FP8, Adaptive EVS:
- Cut alert contextualization latency by 17%, from 1,021 ms to 844 ms, while similarly reducing token usage,
- Increased concurrent real-time VLM streams by 46%, from 13 to 19,
- Summarized a 60-minute video in about half the time using 80% fewer VLM input tokens.
Results vary with scene motion, chunk length, and similarity threshold; benchmark representative footage before selecting production defaults.
Adaptive EVS is most useful when VLMs read many frames and produce short responses (dense captioning, long-video summarization, alert verification). It offers less benefit for long outputs from few frames. It runs inside the RT-VLM container (not against remote endpoints) and is optional. Enable it in override.env and tune via:
VIA_EVS_SESSION=true VLM_VIDEO_PRUNING_RATE=0.5 # 0.0 to 1.0; higher prunes more VLLM_EVS_SIMILARITY_THRESHOLD=0.2
Practical impact for teams
- Lower integration effort via natural-language composition of multi-workflow applications,
- Less duplicated infrastructure across alerting, search, summarization, reporting, and Q&A,
- Higher GPU efficiency by pruning unchanged visual regions,
- Lower summarization and incident-review latency by skipping uneventful video,
- Easier extension through incremental deltas that reuse running deployments.
Together these reduce upfront development work and the GPU work needed to process video with VLMs across environments and applications.
Getting started with VSS 3.3
- Clone the VSS Blueprint repository and check out the branch containing the 3.3 skills.
- Install the VSS Agent Skills into your coding agent’s standard skills directory.
- Describe the desired agent (video sources, workflows, deployment constraints), or say "build a vision agent" for guided composition.
- Review the architecture diagram and _builds/<name>/override.env—pay attention to GPU placement, model endpoints, ports, storage, and security boundaries.
- For RT-VLM workloads, enable Adaptive EVS, tune pruning rate, and benchmark accuracy, throughput, and latency on representative video.
- Deploy on a trusted, isolated network with authentication, TLS, rate limiting, and external controls.
Example setup commands are provided in the repository. NVIDIA also offered a live build demonstration on Oct. 1 at 9:00 a.m. PT where a visual AI agent was constructed from a single prompt.
Conclusion
VSS 3.3 provides tools to reduce both the developer effort and GPU cost of production visual AI agents. The Build Vision Agent skill accelerates and standardizes composition and change, while Adaptive EVS reduces unnecessary VLM processing by focusing compute on moments that matter. Tests on representative hardware show shorter deployment times, lower latency, fewer tokens used, and higher concurrent stream capacity.



