A June 2026 VentureBeat Pulse Research survey of 157 enterprise respondents (organizations with 100+ employees) examined how technical leaders measure agent performance and reliability. The study’s headline finding is an "evaluation gap": companies are giving AI agents more autonomy while placing less trust in the automated evaluations intended to gate that autonomy.
What happened and why it matters
- Half of respondents (50%) reported that in the past 12 months they deployed an agent or LLM feature that passed internal evaluations and subsequently caused a customer-facing failure (incorrect output, broken workflow, or quality incident). A quarter saw this happen more than once.
- Only 36% reported no such failure; 8% run no pre-deployment evaluations, and 6% do not track root causes closely enough to know.
This demonstrates that a passing evaluation is not the same as a reliable agent in production.
Low trust in automated evaluations
- Only 5% of organizations say they fully trust automated evaluation as it currently exists; 95% identify at least one limitation undermining trust.
- The most-cited limitation (29%) is that evaluations align poorly with real-world outcomes — directly explaining many of the post-deployment failures.
- Other common concerns: bias or inconsistency (21%), lack of explainability (18%), and data-leakage/privacy issues within evaluation processes (17%).
Autonomy is increasing despite low trust
- Two-thirds of organizations (66%) either already permit fully automated, zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering pipelines to allow it within 12 months (33%). Only 22% rule it out for the foreseeable future.
- Larger organizations (2,500+ employees) are slightly further along toward zero-human review than smaller firms (70% vs. 64%) and are marginally more likely to have shipped an evaluation-passing agent that later failed a customer (54% vs. 48%).
This creates a paradox: autonomy is rising faster than the assurance mechanisms that should underpin it.
The evaluation tooling landscape is fragmented and provider-led
- The primary evaluation tools used today are split. Provider-native tooling leads: OpenAI’s native evals and traces (17%) and Anthropic’s Claude Console evals (13%). Notably, 17% report using no dedicated agent-evaluation tooling at all.
- Specialist vendors appear with lower but distributed shares: DeepEval (12%), Braintrust (8%), and mentions of LangSmith, Weave, Promptfoo, Langfuse, and Arize. Eleven percent have built in-house solutions.
There is no single independent platform that has emerged as a clear category standard.
Production monitoring tracks function more than correctness
- Organizations tend to monitor whether an agent is functioning (uptime, latency, errors, cost) rather than whether its outputs are correct.
- 51% monitor only functioning metrics, while 23% monitor output correctness. Including ad-hoc human reviewers and unknowns, about three-quarters of organizations lack automated, real-time checks of output quality in production.
A confidently wrong answer will pass functioning checks while causing customer harm — a blind spot that mirrors the pre-deployment evaluation gap.
How organizations choose tools and measure success
- Selection of evaluation vendors is driven primarily by cost (28%) and ease of integration (27%), followed by evaluation accuracy (24%). Breadth of observability (13%) and vendor roadmap (4%) matter less.
- The top success metric is evaluation consistency — getting the same verdict on the same behavior each time (36%) — ahead of the speed of experimentation (19%) and reduction in failures (18%). Overall satisfaction with current tooling averages 3.8 on a five-point scale.
Consistency matters because biased or inconsistent verdicts were among the top limits to trust.
Investment priorities: observability and human review
- Planned investment over the next year emphasizes production observability most, with human review workflows the second-largest planned increase (26%).
- Only 16% plan to expand automated evaluation pipelines. The result: many organizations are simultaneously engineering toward greater autonomy while increasing spending on oversight and human reviewers.
Tooling churn is expected
- 64% plan to adopt a new, additional, or replacement evaluation platform within 12 months; 31% plan changes within the next quarter.
- Interest is concentrated on Confident AI’s DeepEval (20%), OpenAI’s native evals (13%), and Braintrust (9%), suggesting that specialist or newer providers are attracting attention relative to their current footprint.
This looks less like mass defections than an initial wave of dedicated evaluation tooling adoption.
Bottom line: the evaluation gap may widen
In this directional sample of 157 enterprise respondents (June 2026), organizations are granting AI agents more independence than their evaluations can reliably support. Half have encountered agents that passed internal tests but failed in production; only 5% fully trust automated evaluation; most production monitoring watches uptime and cost rather than answer correctness; yet two-thirds permit or are building toward zero-human deployment for low-risk agents.
The vendor market is early and unsettled: provider-native evals and no tooling are currently the most common primary approaches, and a majority plan to adopt or switch platforms within a year. Encouragingly, enterprises plan to spend on observability and human review, indicating awareness of the gap even as they engineer past human gates. Whether assurance will catch up with autonomy — or whether false-confidence failures will scale into fully automated deployments — remains the open question for future waves of this series.
Methodology and limitations
- VentureBeat conducted a single-wave survey in June 2026 of 157 qualified respondents from organizations with 100+ employees. Because this is a single wave rather than a pooled multi-month sample, findings should be read cross-sectionally and interpreted as directional rather than precise population measures.
- The sample is self-selected and skews toward the mid-market (100–499: 37%; 500–2,499: 27%; 2,500–9,999: 20%; 10,000–49,999: 10%; 50,000+: 6%).
- Respondent roles include decision-makers and influencers for AI purchases (38% final decision-makers; 34% recommenders), with titles such as product/program managers, consultants/advisors, directors of engineering/IT, and CIOs/CTOs/CISOs. Industries represented include Technology/Software (23%), Retail/Consumer (15%), Healthcare/Life Sciences (12%), and Manufacturing (10%).



