The July 2026 wave of VentureBeat Pulse Research’s agent reliability tracker, fielded on the same instrument as June, surveyed 108 organizations with 100 or more employees and found a notable rise in trust toward automated agent evaluation even as the share of organizations that shipped agents which passed internal evaluations but later failed customers did not change.
Key facts and timing
- Survey: VentureBeat Pulse Research, "agentic reliability and evals tracker", July 2026 wave.
- Sample: 108 qualified respondents (organizations with 100+ employees).
- Comparison wave: June 2026 (n=157), instrument identical to July’s.
What changed: confidence increased
The share of enterprises that said they fully trust automated evaluation rose from 5% in June to 13% in July. The most-cited limitation in June — that evaluations align poorly with real-world outcomes — fell from 29% to 19% and ceded the top objection to concerns about evaluation bias and inconsistency (22%, tied with data-leakage concerns at 22%). These moves clear conventional significance thresholds in the month-over-month test.
What didn’t change: the failure rate
Nearly half of respondents (49%) reported that, in the past 12 months, they deployed an agent or LLM feature that passed internal evaluations and then caused a customer-facing failure (vs. 50% in June). A quarter (24%) have seen this happen more than once; that share is unchanged. Across the two waves (265 enterprises), the rate at which passing evaluations certify agents that then fail is stable to within a percentage point.
This stability is the anchor of the report: whatever shifted between June and July, it did not change how often a passing evaluation later proves incorrect in production.
Where the new trust comes from: the burned vs. unburned split
Cross-tabulation shows the increase in full trust is concentrated among organizations that have not experienced a false-confidence failure. Among the 53 enterprises that shipped an agent which passed evals and then failed a customer, only 4% fully trust automated evaluation; among the 41 that have not seen such a failure, 24% fully trust it. That six-fold gap is the sharpest split in the data.
The implication is that the July improvement in sentiment is driven largely by inexperience rather than by objectively better evaluations: the unchanged failure rate rules out the latter.
Being burned accelerates autonomy rather than restraining it
At the headline level, the autonomy trajectory is unchanged: 67% of organizations either already allow zero-human deployment for low-risk agents (37%) or are engineering to permit it within a year (30%) — identical to June’s 67%. The share ruling it out for the foreseeable future slipped from 22% to 18%.
Under the surface, however, experience of failure correlates with greater automation. Among enterprises that shipped an evaluation-passing agent that then failed a customer, 85% are on the zero-human path; among those that have not been burned, 61% are. Only 11% of the burned group rule out full automation, versus 24% of the unburned. The most plausible reading is maturity: organizations that deploy at volume and occasionally hit customer-facing incidents also have the pipelines and incentives to automate, treating some failures as an operational cost.
Because the most active deployers are scaling toward no-human gates, a steady failure rate still implies a growing absolute number of incidents.
The stack begins to consolidate: specialists gain, “nothing at all” falls
Use of dedicated evaluation tooling increased and the share running no dedicated evaluation tooling fell from 17% to 12%. Specialist platforms gained primary-usage share: Braintrust’s primary share nearly doubled to 15% (from 8%), and DeepEval reached 17%. Provider-native tooling stayed roughly flat (OpenAI 18%, Anthropic 12%), so the growth came mainly from organizations that previously used no dedicated tooling.
Counting any use rather than only primary platform, OpenAI native evals reach 31% of enterprises, DeepEval 27%, Braintrust 22%, Anthropic native evals 20%, custom in-house tooling 14%, and Weave and Langfuse 11% each. Nineteen percent still report no dedicated tooling anywhere in their stack. This is the first wave in which multiple independent specialist contenders have double-digit presence as primary platforms.
Production monitoring still focuses more on liveness than correctness
Monitoring in production tends to watch either whether an agent is functioning (uptime, request completion, latency, errors) or whether its outputs are correct (automated content checks). The split is essentially unchanged from June: 50% of organizations monitor only whether the agent is functioning, while 26% run automated checks on whether answers are correct. Inline quality assertions and transaction trace logging are each used by 26% of organizations (base: 106).
Crucially, among enterprises that permit zero-human deployment, only 28% run inline quality checks on production traffic. Many organizations have removed a human from the deployment gate without replacing that gate with automated, runtime checks of output correctness.
Buyers now choose fit over price
Ease of integration rose from 27% to 39% as the top selection criterion for evaluation vendors, overtaking cost (23%). Evaluation accuracy rose modestly to 28%. This aligns with the shift from shopping to installing: enterprises adopting their first dedicated tooling optimize for pipeline fit.
What enterprises measure as success did not move: evaluation consistency remains the primary success metric at 38% (June: 36%), ahead of reduction in failures (20%) and speed of experimentation (18%). Satisfaction with current tooling averaged 3.9 on a five-point scale (unchanged from June’s 3.8).
Investment priorities: human review edges production observability
Planned investment growth favors human review workflows (31%) and production observability (30%), with automated evaluation pipelines growing fastest at 19%. Notably, burned organizations prioritize human review more strongly: 38% of the burned group named human review as their fastest-growing investment, versus 24% of the unburned, who prefer observability tooling. The strategy appears to be to automate deployment decisions while funding humans to catch what automation misses — a hedge that may not scale as volumes grow.
Fewer enterprises are actively shopping for change
Overall 56% of respondents still intend to adopt a new, additional, or replacement evaluation platform within 12 months (down from 64% in June). Near-term intent fell from 31% to 24%, and those planning no change rose from 36% to 44%. Among the 60 enterprises planning a change, OpenAI native evals lead consideration at 20%, followed by Braintrust (18%), Weights & Biases Weave (12%), and DeepEval (10%). DeepEval converted a notable portion of its June consideration interest into primary usage by July; Braintrust shows high interest and accelerating primary adoption.
Bottom line: confidence moved, correctness did not
July’s data show a meaningful rise in trust toward automated agent evaluation and early signs of a forming evaluation layer in enterprise stacks. Yet the central operational metric — the rate at which passing evaluations certify agents that later fail in customer-facing contexts — remained virtually unchanged at about 49%. The new confidence is concentrated in organizations that have not yet experienced false-confidence failures, while those that have been burned are both more likely to automate and more likely to fund human review.
At 108 respondents in a mid-market-weighted, self-selected sample, and with an industry mix that shifted away from Technology/Software between waves, these results are directional rather than precise. The readable direction is clear: enterprises are tooling up and buying for fit, and confidence has caught up first — even though the concrete evidence of correctness has not.
(Report: VentureBeat Pulse Research — July 2026 wave, compared to June 2026 wave.)



