Industry

Enterprises Accelerate AI Infrastructure Spending Despite Limited Visibility into Costs

A VentureBeat Pulse survey of 107 enterprises (100+ employees) in June 2026 finds firms are rapidly increasing investment in AI infrastructure even though most cannot clearly measure its economics.

Enterprises Accelerate AI Infrastructure Spending Despite Limited Visibility into Costs

A VentureBeat Pulse survey conducted in June 2026 among 107 enterprise respondents (organizations with more than 100 employees) finds that companies are accelerating investment in AI infrastructure faster than they can measure or control its economics. The study identifies a central "compute gap": heavy, fast-moving capital deployment running ahead of the visibility needed to manage it.

Deployment maturity: ambition exceeds production

Respondents report varied stages of AI deployment:

  • 38% are experimenting (proofs of concept)
  • 37% have some workloads in production but not organization-wide
  • 21% run AI in production at scale
  • 4% are not yet running AI workloads

Three quarters (76%) are still experimenting or only partially in production, meaning the sample skews toward organizations whose compute footprints and costs are set to grow.

Current stack: hyperscalers and model APIs dominate

Enterprises largely run AI on familiar cloud platforms and model-provider APIs today:

  • 48% use Google Cloud (Microsoft Azure 29%, AWS 22%, Oracle Cloud 22%)
  • 41% use Google's Gemini models; OpenAI 40%, Anthropic 12%
  • 6% run on-prem or co-located GPU clusters; 4% use a custom open-source self-managed stack
  • Under 2% currently use specialized AI clouds (CoreWeave, Lambda, Crusoe, Nebius, Together, Fireworks, etc.)

The specialized GPU clouds are nearly absent from current deployments in this cohort; the stack is concentrated in general-purpose clouds and major model APIs.

Next evaluations diverge from current usage

When asked what they plan to evaluate over the next 12 months, enterprises point away from their current stack:

  • 45% plan to evaluate AI-specialized clouds (CoreWeave, Lambda, Crusoe, Nebius)
  • 32% plan to evaluate non-NVIDIA accelerators (AWS Trainium, Google TPU, AMD Instinct, Intel Gaudi, in-house ASICs)
  • 28% plan to evaluate Nvidia Blackwell / next-gen GPUs
  • 16% plan to evaluate decentralized/distributed compute networks
  • 11% plan to evaluate sovereign or region-specific compute

The single most-cited area for evaluation is AI-specialized clouds, a category almost none of these enterprises use today — indicating a possible re-platforming.

A switching wave: many plan provider changes soon

Intent to switch or add infrastructure providers is high:

  • 38% plan to change within 0–3 months
  • 22% within 3–6 months
  • 7% within 6–12 months
  • 36% have no plans to change

Overall, 64% plan to switch or add a provider within twelve months. The near-term switching interest largely targets incumbent providers (Microsoft Azure, Google Cloud, OpenAI, Gemini), suggesting short-term reshuffling among majors even as specialized clouds gain attention for longer-term evaluation.

Buying criteria: integration and TCO beat token price

When choosing providers, enterprises prioritize integration and total cost of ownership (TCO) over advertised unit pricing:

  • 41% cite integration with existing cloud and data stack as the top factor
  • 35% cite total cost of ownership
  • 24% cite performance (latency and throughput)
  • 19% cite security/compliance, autoscaling for spiky workloads, and GPU access/availability
  • 8% cite cost per 1M tokens (the least-cited factor)

Thus, buyers favor fit and true operating cost rather than headline token prices — yet many cannot rigorously measure those costs today.

Low GPU utilization and limited cost tracking

GPU utilization is generally low among enterprises that operate GPUs:

  • 37% report 26–50% utilization
  • 34% report 10–25%
  • 15% report under 10%
  • 12% report over 50%
  • 8% do not measure utilization; 7% consume via API and run no GPUs of their own

Across GPU-operating respondents, 83% report utilization at or below 50%, and nearly half (49%) run at 25% or below. Idle accelerators therefore represent a major inefficiency.

On measuring economics:

  • 44% rigorously track compute cost and ROI
  • 39% track it partially
  • 20% cannot yet quantify it
  • 6% say it is not a priority

Fewer than half of enterprises can rigorously quantify their AI compute costs, even though TCO is a leading purchasing criterion. Satisfaction with current infrastructure is modestly positive (average 4.0 on a five-point scale), with value for money (3.9) and ease of implementation (3.8) trailing slightly.

The emerging bottleneck — memory capacity — is not yet widely addressed

When asked how they would respond to a shift in the inference bottleneck from raw GPU compute to memory/KV-cache capacity, responses are fragmented:

  • 31% would rely on Dell (PowerScale / Project Lightning)
  • 16% would rely on Nvidia (Dynamo / ICMSP)
  • 10% Hammerspace (Tier Zero), 9% DDN (Infinia)
  • About 18% are either unaware of this constraint (9%) or have not addressed inference-memory limits yet (8%)

Other responses referenced open-source KV-cache tooling, model-level efficiency, VAST Data, WEKA, and similar options. Roughly one in five enterprises has not yet recognized or addressed this upcoming architectural constraint.

Bottom line: spending outpaces the instrumentation to spend wisely

In this single Q2 2026 cross-sectional wave of 107 enterprise respondents (100+ employees), the appetite to invest in AI infrastructure is running well ahead of the systems to measure and control its economics. Most organizations in the sample are early in deployment, current GPU fleets run largely underutilized, and fewer than half can rigorously track compute costs or ROI. Planned evaluations favor AI-specialized clouds and alternative accelerators, and a majority expect to switch or add providers within a year. Without improved cost visibility and capacity utilization, rapid spending risks widening the compute gap rather than closing it.

Note on methodology: the survey is self-selected, skews toward mid-market organizations and earlier-stage adopters, and should be read as a directional signal rather than a precise population measurement. Responses were collected in a single Q2 2026 (June) wave from 107 qualified enterprise respondents across industries including Technology/Software (26%), Healthcare/Life Sciences (15%), Financial Services (13%), and Retail/E-commerce (12%).