VentureBeat Research's five parallel surveys conducted in June 2026 conclude that companies knowingly rolled out AI agents faster than they implemented the controls required to manage them. The research examined five control layers organizations need to trust agentic systems: identity, evaluation, cost telemetry, the context layer, and orchestration.
What the five controls mean
- Identity: which agent can do what and under which credentials.\
- Evaluation: whether an agent's output is acceptable.\
- Cost telemetry: tracking how much each agent costs to run.\
- Context layer: the business data and definitions agents use to answer.\
- Orchestration: the control plane that coordinates multi-step agent workflows.
Each of the five reports in the series focuses on one of these controls and together they reveal common gaps.
Many "agents" are actually simple chatbots
Seventy-one percent of enterprises reported that a quarter or fewer of their deployed "agents" can complete multi-step work autonomously; only 10% said true multi-step agents make up the majority of their deployments. The respondents are informed: 81% recommend or decide on AI purchases at their companies. A single-prompt chatbot that has a human read every answer requires few of the controls; a genuine multi-step agent requires all of them — and most organizations cannot reliably say which type they have deployed.
Autonomy is outpacing trust in evaluations
Two-thirds of enterprises either already allow an agent to push code or system changes to production based solely on automated evaluation results with no human review, or are actively engineering toward that capability within 12 months. Only 5% fully trust the evaluations that would enable such automated pushes. Half of enterprises shipped an agent that passed internal evaluations and then caused a customer-facing failure in the past year. The reports recommend testing evaluations against production outcomes rather than only internal benchmarks before removing human review from any workflow.
Shared credentials correlate with more incidents
Sixty-nine percent of companies allow at least some agents to share credentials (multiple agents operating under one API key or service account). Organizations that permit credential sharing experienced a security incident or near-miss at a rate of 63.5% (47 of 74), compared with 40.9% (9 of 22) at companies where every agent has its own scoped identity. The recommended mitigation is scoped identity for every agent, prioritized for those that touch production systems.
Underutilized GPUs and incomplete cost tracking
More than eight in ten enterprises that run their own GPUs reported utilization of 50% or less, and only 44% rigorously track the actual cost and return of their AI compute. The immediate opportunity is not simply buying more GPUs but improving utilization and per-workload cost of existing hardware.
Agents draw confidently from ungoverned data
Fifty-seven percent of enterprises traced a confident, incorrect agent answer in the past six months to their own missing or inconsistent business context — incorrect metrics, stale definitions, or absent documents — and most saw it happen more than once. Governing the definitions agents use (metrics and entities first) needs to precede scaling the agents that depend on them.
No entrenched incumbents; orchestration shows highest switching intent
No control layer currently has a dominant incumbent: defaults today tend to be the built-in tools that ship with major AI platforms. Intent to switch vendors or add new ones is highest for orchestration, where 68% plan to adopt, add, or replace platforms within 12 months and 34% intend to do so within the quarter. The surveys did not ask whether this spend will favor platform-built tools or specialist challengers — that question will shape the market over the next four quarters.
About the research
VentureBeat Research fielded five parallel surveys in June 2026 under its VB Pulse program: Agentic Orchestration (101 respondents), Agent Reliability & Evals (157), Agentic Security & Identity (107), AI Infrastructure & Compute (107), and Context Layers / RAG (101) — 573 qualified respondents in total, all at organizations with 100 or more employees. Samples are self-selected, so some findings should be read directionally; each report includes a full methodology note. The consistent pattern across independent surveys points to a common direction. VentureBeat produces both this research and VB Transform, the conference where these reports debuted.



