A VentureBeat Pulse survey conducted in June 2026 of 157 enterprises (each with 100+ employees) finds a growing “evaluation gap”: organizations are granting AI agents increased autonomy at the same time that they express little confidence in the automated tests intended to gate that autonomy. The survey — titled the Agentic Reliability & Evals tracker — asked technical leaders about how they evaluate agent performance, what they monitor in production, which tools they use, what fails in the field, and how far they will let agents run without human oversight.
The defining finding: a passing eval is not a working agent
Half of respondents (50%) reported that, in the past 12 months, they deployed an agent or LLM feature that passed their internal evaluations and then caused a customer-facing failure in production (incorrect output, broken workflow, or quality incident). A quarter experienced this more than once. Only 36% reported no such failure; 8% perform no pre-deployment evaluations and 6% do not track root cause closely enough to know.
This is the report’s central number: passing an eval did not ensure correct behavior in real-world use, and that experience shapes how enterprises view evaluation trust, monitoring, and the autonomy they grant.
Almost no one fully trusts automated evaluation
Trust in automated evaluation is scarce. Only 5% of organizations say they fully trust automated evaluation today; the remaining 95% cite at least one limiting factor. The most-cited limitation (29%) is that evaluations align poorly with real-world outcomes — the direct explanation for the failure cases above. Other top concerns are bias or inconsistency (21%), lack of explainability (18%), and data-leakage or privacy concerns in the evaluation process (17%).
Autonomy is increasing despite limited assurance
Paradoxically, the autonomy ceiling continues to rise. Two-thirds of organizations (66%) either already permit fully automated, zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to allow that within 12 months (33%). Only 22% rule out such deployments for the foreseeable future.
Larger firms (2,500+ employees) in the sample are slightly further along this path than smaller ones: 70% versus 64% are moving toward zero-human review, and 54% versus 48% have shipped an evaluation-passing agent that later failed a customer. The sample sizes mean these are directional comparisons rather than precise population estimates.
The evaluation stack is fragmented and provider-led
The marketplace for agent reliability tooling is early and unconsolidated. The most common primary tools are provider-native evals — OpenAI’s native evals and traces (17%) and Anthropic’s Claude Console evals (13%) — but strikingly 17% of enterprises report using no dedicated agent-evaluation tooling at all. Specialist vendors such as DeepEval (12%), Braintrust (8%), LangSmith, Weave, Promptfoo, Langfuse, Arize and others appear at single- to low-double-digit shares, and 11% have built their own platforms. No independent platform yet dominates the category.
Production monitoring rarely checks output correctness in real time
Production monitoring tends to focus on either system functioning (is the agent up, responsive, within cost and error budgets) or output correctness (automated checks that the agent’s answers or actions are right and policy-compliant). The split is stark: 51% of organizations monitor only functioning metrics, while 23% monitor correctness. Roughly three-quarters do not run automated, real-time evaluation of output correctness on live traffic — they can see uptime and cost but generally cannot see, in real time, when the agent begins to produce incorrect results.
That blind spot is the runtime counterpart to the pre-deployment failures described earlier.
How organizations choose and judge evaluation tools
Selection is driven by cost and integration pragmatics: cost of evaluations (28%) narrowly leads selection, followed by ease of integration (27%) and evaluation accuracy (24%). On what counts as success, 36% cite evaluation consistency (same verdict on the same behavior every time), ahead of speed of experimentation (19%), reduction in failures (18%), production visibility (13%) and compliance (11%). Satisfaction with current tooling is moderate — averaging 3.8 on a five-point scale across overall satisfaction, implementation ease and value for money.
Investment priorities: observability and human review
Planned spending over the next year favors production observability most strongly, with human review workflows the second-largest planned investment (26%). Notably, more organizations plan to increase budget for human review than for building automated evaluation pipelines (16%) that would replace humans. Only 8% report no increase in budget.
This indicates a hedging strategy: companies are engineering toward autonomy while also increasing oversight capacity, keeping humans available for decisions that automated evaluations cannot yet be trusted to make.
Tooling churn is likely: most plan to adopt or switch platforms
The market is active: 64% of enterprises plan to adopt a new, additional or replacement evaluation platform within 12 months (31% within the next quarter). Among platforms under consideration, DeepEval (20%) leads, followed by OpenAI native evals (13%) and Braintrust (9%). Given the current reliance on provider-native tools or no tools at all, this wave looks like initial adoption rather than mass defection.
Bottom line: an evaluation gap that autonomy could widen
In this June 2026 cross-sectional survey of 157 enterprise respondents, the direction is clear: organizations are granting more independence to AI agents than their confidence in the evaluations justifying that autonomy would suggest. Half have shipped an agent that passed internal evals and then failed in production; only 5% fully trust automated evaluation; most monitor uptime and costs rather than answer correctness in production; yet two-thirds allow or are building toward zero-human deployments.
The vendor market is early and fragmented; many rely on provider-native evals or no dedicated tooling, and a majority plan to change platforms within the year. Encouragingly, planned investment is shifting toward observability and human review, showing that enterprises recognize the gap. The crucial open question for future waves is whether assurance practices and tooling improve fast enough to match the autonomy enterprises are granting — or whether the kinds of false-confidence failures already seen will scale as deployments become more automated.
Methodology: VentureBeat fielded a single-wave Pulse Research survey in June 2026 of qualified enterprise respondents (n=157) from organizations with 100 or more employees. The sample is self-selected and skews toward mid-market companies; respondents include product and program managers, consultants and advisors, directors of engineering/IT, CIOs/CTOs/CISOs and others across technology/software, retail/consumer, healthcare/life sciences, manufacturing and additional industries. The results are directional rather than precise population estimates.



