Industry

AI-generated text

Enterprises Prioritize AI Performance and Availability while Cost Visibility Lags

A July 2026 VentureBeat Pulse survey of 170 enterprises finds that two-thirds run AI workloads in production and 29% at scale, but fewer than half can rigorously track AI compute costs.

Enterprises Prioritize AI Performance and Availability while Cost Visibility Lags

A VentureBeat Pulse survey conducted in July 2026 among 170 enterprises (each with 100+ employees) shows that AI infrastructure has largely moved into production: 66% of respondents run AI workloads in production and 29% report AI in production at scale. However, fewer than half of the organizations can rigorously track their AI compute costs.

The research examines where organizations are in their deployment journeys, what platforms they run AI on today, how they buy and measure infrastructure, where the next investments are aimed, and — crucially — how visible the economics of that compute are.

Deployment maturity: most have passed the pilot stage

This cohort is operational: two-thirds (66%) of respondents run AI in production, 29% at scale, 30% remain in proofs of concept, and 4% have not started. The sample skews up-market (57% of respondents have more than 1,000 employees), so the infrastructure decisions reflected here are mostly being made by organizations with live workloads and real bills.

That production context helps explain why performance and availability outrank cost in purchasing criteria, and it raises the stakes on measurement: inability to see utilization or cost is a planning problem in pilots and an operational problem in production.

The stack: hyperscalers and model APIs dominate; average three platforms

Enterprises run an average of three infrastructure platforms. OpenAI appears in 49% of stacks, Google Gemini in 48%, Microsoft Azure in 47%, and Google Cloud in 42%. When asked to name a single primary platform, Microsoft Azure leads at 26%.

Specialized GPU “neoclouds” remain marginal in actual use: CoreWeave and Lambda each appear in 3.5% of stacks, Baseten in 3%, and other specialized providers at or below 2%. Combined, specialized clouds are named as the primary platform by just 1% of enterprises. By contrast, 13% run a custom open-source self-managed stack and 9% operate their own GPU clusters.

Note: Because the platform question allowed multiple selections, these figures measure presence in a company’s stack rather than spend-weighted market share. The separate primary-platform question is a better guide to center-of-gravity.

Next investments point away from the current stack

The single most-cited planned evaluation area for the next 12 months is AI-specialized clouds (44%), which also show the strongest net momentum (+36). Yet usage of those providers today is only 3.5% for CoreWeave and Lambda each. Nearly 39% plan to evaluate non-Nvidia accelerators, 25% next-generation Nvidia silicon, and 18% decentralized compute networks.

This is the sharpest tension in the data: a large share of enterprises plan to evaluate specialized clouds, but those services currently sit on a very small base. Whether neoclouds convert evaluations into deployments, or hyperscalers absorb demand with their own offerings, remains an open question.

High churn intent, mostly among incumbents

Sixty-two percent of enterprises intend to switch or add a provider within 12 months, and 29% within the next quarter. Only 39% plan to stand still. However, the providers drawing the most switching consideration are largely the incumbents enterprises already run: OpenAI and Google Cloud (29% each), Microsoft Azure (28%), Gemini (25%), Anthropic (16%), Oracle Cloud (14%), and AWS (13%). CoreWeave and Lambda register much lower consideration rates (4% and 3.5% respectively).

Short-term switching tends to be incumbent-focused consolidation, while the 44% evaluation rate for neoclouds reflects a longer-term replatforming thesis rather than immediate pipeline.

Buying and measurement criteria: performance overtakes cost

When selecting AI infrastructure, integration with the existing cloud and data stack is the top factor (40%). Performance (latency and throughput) is second at 35%, GPU availability third at 24%, and total cost of ownership (TCO) fourth at 22%. Fine-grained autoscaling is 18%, and cost per million tokens is 16%.

Measured success follows the same logic: uptime and reliability is the primary metric for 51% of enterprises, developer productivity and deployment speed 39%, cost per million tokens 31%, latency 27%, and throughput 25%.

For teams running production workloads, prioritizing uptime and speed is coherent. But it is notable that TCO has been demoted to fourth at the same time that 53% of enterprises cannot rigorously track compute costs — a likely reason cost falls down the priority list: it is often the hardest dimension to observe.

GPU utilization: most fleets run cold

Among the 155 enterprises that operate their own GPUs, 69% report utilization at or below 50%, with the 26–50% band accounting for 46% of respondents. Only 23% exceed 50% utilization. Twelve percent do not measure utilization at all.

Scale by itself does not guarantee better utilization: among enterprises running AI in production at scale, only 22% clear the 50% mark, statistically similar to other segments. Idle or underutilized accelerators represent significant wasted spend, and for one in eight enterprises that waste is invisible because utilization isn’t measured.

Cost tracking: fewer than half can account rigorously for spending

Only 47% of enterprises rigorously track the cost and return of their AI compute. Another 39% track partially, 15% cannot yet quantify it, and 6% have not prioritized it. Even among organizations running AI in production at scale, rigorous tracking reaches just 56%.

Satisfaction scores reflect this gap: overall satisfaction averages 4.14 out of 5, ease of implementation 4.04, but value for money trails at 3.87 — the weakest-rated dimension and the one that requires measurement to assess properly.

Taken together with buying priorities, the picture is self-reinforcing: cost is deprioritized in procurement decisions partly because many organizations lack the instrumentation to evaluate it.

Memory frontier: fragmented approaches and limited awareness

Addressing the emerging inference constraint — a shift from compute to memory (KV-cache capacity) — the market is scattered. Dell is cited by 24% and Nvidia by 21% as approaches to the memory constraint; 12% point to open-source tooling, 11% to model-level efficiency techniques (MLA, quantization). Crucially, 19% either do not recognize the constraint (7%) or have not begun addressing it (12%).

This suggests the memory bottleneck is arriving while many organizations are still trying to get clear visibility into current compute economics.

Bottom line

Enterprises with 100+ employees have moved AI into production and buy for integration, performance, and availability. They run multiple platforms on average and measure success by uptime and developer velocity. But they have not yet become consistent accountants of the underlying compute: TCO has dropped in priority while measurement of utilization and costs remains incomplete. The next wave of investment points toward specialized clouds and non-Nvidia accelerators, but those providers currently sit on a small usage base. The open question for future waves is whether instrumentation and cost visibility catch up before — or after — a broader replatforming occurs.

Methodology

VentureBeat fielded this Pulse Research survey in a single July 2026 wave among 170 qualified enterprise respondents (organizations with more than 100 employees). The sample is self-selected and intended to be directional rather than a probability-based census; several questions allowed multiple selections, meaning some percentages reflect presence in stacks rather than spend. Respondents span managers, individual contributors, the C-suite, and VPs/directors, and come from industries including Technology/Software (35%), Manufacturing (14%), Financial Services (12%), and Healthcare/Life Sciences (9%).