VentureBeat Pulse's July survey shows that enterprises whose AI agents passed internal evaluations but later failed in production are moving faster toward removing humans from deployment decisions — even as overall trust in automated evaluation tools has risen.
In the July wave, VentureBeat polled 108 enterprise respondents. Thirteen percent said they fully trust automated evaluation, up from 5% in June. At the same time, the share of respondents naming poor alignment between tests and real-world results as their biggest worry fell from 29% to 19% month over month.
Failures remain common
Forty-nine percent of respondents reported that, in the past year, an AI agent or LLM-powered feature that cleared company testing later caused a customer-visible problem; this is essentially unchanged from 50% in June. Twenty-four percent said that outcome had happened more than once.
VentureBeat Intelligence highlights that this figure does not imply 49% of all agent runs fail; rather, it means nearly half of surveyed organizations experienced at least one instance where their internal release gate approved a system that later disappointed customers.
Confidence splits by experience
July's data shows a clear split. Among enterprises that had seen a test-approved system disappoint customers, only 4% said they placed complete faith in automated checks; by contrast, 24% of organizations reporting no comparable incident expressed full confidence. Experience thus correlates with reduced trust in automation.
Burned companies accelerate toward zero-human deployment
Overall, 67% of respondents either already allow an agent to push code or change a system without human approval in certain low-risk cases (37%) or are modifying pipelines to enable that within the coming year (30%). However, among enterprises that had experienced a test-approved system fail in production, 85% were pursuing the no-approval model, compared with 61% in the group reporting no such incident.
The survey cannot determine whether this reflects recklessness or deployment maturity. Organizations that run more agents, at higher volume and across more consequential workflows are both likelier to encounter failures and more likely to have the engineering infrastructure for automated deployment.
Production-quality monitoring lags behind
Pre-deployment evaluation and production monitoring answer different questions: the former assesses readiness to ship, the latter inspects live behavior and correctness. In July's sample, most companies focused on whether a system functioned rather than whether its answers were correct.
Of 106 valid responses to that question, 26% used inline quality assertions (automated judges or guardrails that check live traffic for output-quality problems), another 26% monitored transaction traces (infrastructure spans, token usage, raw I/O), and 24% tracked gateway metrics (latency, errors, cost). Grouped by what the architecture actually watches, half of respondents monitored whether an agent was functioning, while just over a quarter automatically monitored whether its production output was correct.
Notably, among the 40 organizations already permitting no-approval deployment in limited cases, only 28% automatically checked the meaning and correctness of live answers. Most enterprises that removed a human from some release decisions have not installed semantic-quality monitoring as a production backstop.
Independent evaluation tools are emerging
Vendor data points to a maturing market: specialist evaluation platforms are gaining ground and evaluation is becoming a distinct enterprise software layer. In July, primary platform shares were: OpenAI native evals and traces 18%, Confident AI DeepEval 17%, Braintrust 15%, and Anthropic's Claude Console and Workbench 12% (tied with respondents reporting no dedicated platform). Internal tools, Promptfoo and LangSmith each held 6% as primary options.
Across stacks (since many companies use multiple tools), OpenAI native evaluation appeared in 31% of stacks, DeepEval in 27%, Braintrust in 22%, and Anthropic native tooling in 20%. Braintrust's primary share rose from 8% in June to 15% in July; DeepEval rose from 12% to 17%.
Purchasing priorities shifted: the share naming integration ease as the decisive factor climbed 12 points to 39%, displacing cost (which fell from 28% to 23%). Evaluation accuracy ranked second at 28%.
Budgets and the human hedge
Budget priorities show how enterprises are managing the tension between greater autonomy and imperfect evaluation. Investment in people-centered review workflows edged ahead of production observability, 31% to 30%. Automated evaluation pipelines ranked third at 19%, safety and policy testing at 16%.
Among organizations that had experienced a testing miss, 38% said people-centered review would receive the fastest investment growth, versus 24% of organizations that had not been burned. This produces a paradox: the burned group is most likely to remove humans from the release checkpoint and most likely to increase spending on human review elsewhere. The pattern suggests a strategy of automation backed by downstream human oversight — which may or may not scale as deployments grow.
Short conclusion
July's directional data shows that confidence in automated evaluation rose before measurable improvement in measured failure incidence. The ecosystem around evaluation is maturing: specialist tools are more widely adopted, buyers prioritize integration, and companies that experienced failures are increasing investment in human review. Yet nearly half of surveyed enterprises still report at least one instance of a system passing internal checks and later disappointing a customer, and many organizations that authorize no-approval deployments lack automated semantic-quality monitoring in production. For most enterprises, a passing pre-deployment score remains the start of monitoring, not proof of production reliability.



