VentureBeat Research's June survey polled 573 technical leaders at companies with 100 or more employees across five parallel surveys focused on the agentic stack. The central finding: enterprises are knowingly deploying AI agents before the control functions needed to manage them are fully in place.
What was measured and why it matters
The study examined five control layers: agent identity (which agent may do what under which credentials), evaluation of agent outputs (whether the work is good), cost telemetry (what each agent costs to run), the context layer (business data and definitions agents use), and the orchestration control plane (software coordinating multi-step agent work).
Companies are now retrofitting controls and budgeting for them: roughly six in 10 enterprises plan to switch or add vendors in each of the five layers within 12 months, and about a third — depending on the layer — plan to move within a quarter.
Key findings and figures
Idle expensive hardware and limited cost tracking
- 86% of enterprises that operate their own GPUs report utilization of 50% or less, meaning the most expensive hardware is often running at half capacity or less.
- Only 44% rigorously track the actual cost and return of their AI compute; the remainder estimate costs. Despite that, 45% said the compute option they are most likely to evaluate in the next 12 months is an AI-specialized cloud (for example, CoreWeave, Lambda, Crusoe, Nebius), while under 2% report using such neoclouds today.
- Roughly one in three (32%) said they are most likely to evaluate non-Nvidia accelerators (AWS Trainium, Google TPUs, AMD) in the next 12 months, and 28% named next-generation Nvidia GPUs. The survey recommends measuring utilization and per-workload GPU cost before committing budget to new compute options.
Most deployed "agents" are single-prompt chatbots
- 71% of respondents say a quarter or fewer of their deployed agents can complete multi-step work on their own; most deployed agents are single-prompt chatbots. Only 10% said true multi-step agents are the majority of what they run.
- This gap matters because a single-prompt chatbot requiring human review needs fewer identity, evaluation, and cost controls than a true autonomous multi-step agent.
Automated evaluations and removal of human review
- 34% already allow an agent to push a code or system change to production based solely on automated evaluation results with no human review; another 33% are engineering their pipelines to allow this within the next 12 months. Only 5% fully trust those automated evaluations.
- The distrust is supported by experience: half of enterprises shipped an agent that passed internal evaluations and then caused a customer-facing failure in the past year; a quarter saw this happen more than once. The top cited weakness in current evaluations was “poor alignment with real-world outcomes” (29%).
- Most checking stops after deployment: only 23% run real-time quality checks on agent answers in production; 51% monitor only system health (uptime, request traces, gateway logs), which reveals whether the agent is running but not whether its answers are correct.
- Recommendation: before removing human review, test evaluations against production outcomes and instrument answer quality, not just uptime.
Credential sharing increases security incidents
- 69% of companies allow credential sharing somewhere in their agent fleet during runtime (multiple agents operating under one API key or service account).
- Organizations with credential sharing experienced a security incident or near-miss at a rate of 63.5% (47 of 74), versus 40.9% (9 of 22) for organizations where every agent has its own scoped identity.
- Recommendation: assign each agent a scoped identity, starting with agents that touch production systems.
Missing or inconsistent business context causes wrong answers
- 57% traced at least one confident but incorrect agent answer in the past six months to missing or inconsistent business context (wrong metrics, stale definitions, absent documents), and many saw this occur multiple times.
- Status: 25% already run a governed semantic layer in production (a single governed set of business definitions that every AI reads from), 34% are building one, and 41% have not started.
- Recommendation: govern the definitions your agents rely on — metrics and entities first — before scaling agents that depend on them.
Portability rose as a priority this quarter
- Portability and vendor lock-in became a more prominent concern: in the June survey, 51% now expect their primary control plane for enterprise agents to be hybrid (provider-native plus external orchestration) by the end of 2026, up from 34% in the spring wave. The share relying purely on provider-managed agent services fell from 12% to 7%.
- The June survey followed market events that may have influenced attitudes: a June 12 U.S. Commerce Department export order temporarily took Anthropic’s Claude Fable 5 offline for enterprises; multiple open-weight model releases (for example GLM-5.2, Tencent Hy3) and OpenAI previews of GPT-5.6 occurred in June–July. The report notes these events as contextual timing that may help explain rising interest in portability and open-weight models.
No established incumbents — a wide buying window
Across all five layers, 57%–64% of enterprises plan to switch or add vendors within 12 months (64% in infrastructure and evaluations, 59% in agent security, 57% in retrieval and context), and 26%–38% plan to move within a quarter depending on the layer. No layer shows a clear incumbent: the most common evaluation tooling is either the model provider’s built-in evals or no dedicated tooling (17% each); 82% name provider-native or hyperscaler controls as their primary agent security layer; and provider-native retrieval leads the context layer.
Today, most enterprises default to the built-in tools that come with their primary AI platforms (Anthropic, OpenAI, Google, Microsoft, AWS) for convenience. Over the next four quarters, spending decisions and contracts in these five layers will determine whether enterprises stay with those built-in tools or shift toward specialist vendors.
Next steps and follow-up
The Q3 survey wave will measure whether enterprises implemented their planned changes: whether agents received scoped identities, whether evaluations were tested against production outcomes, whether GPU utilization rose, and whether semantic layers shipped. VentureBeat will publish the full Q2 reports across all five VB Pulse trackers at VB Transform, held July 14–15 at Hotel Nia in Menlo Park. VentureBeat produces both this research and the VB Transform event.



