General Compute, a startup building an AI inference cloud, has received a $400 million loan from technology investment firm Upper90. The deal is notable because the financing is secured by chips specifically designed for inference — processors optimized to run already‑trained AI models rather than the more expensive chips used for model training.
Why this matters
The financing reflects market shifts driven by concerns over high costs associated with AI tooling and tokens. Investors and operators increasingly favour infrastructures that can run open‑source models at much lower cost than the latest large language models (LLMs) from leading AI labs.
Founders and earlier funding
General Compute was founded by CEO Finn Puklowski and CTO Jason Goodison. In May the company raised $15 million in seed funding to build an inference‑optimized “neocloud” on SambaNova silicon. Neoclouds are cloud infrastructures designed specifically for AI workloads, in contrast to the general‑purpose offerings from hyperscalers such as AWS or Azure.
The chips and their advantages
General Compute is building on SambaNova’s SN50 chips. According to the company, these chips are energy efficient, do not require costly liquid‑cooling systems, can be deployed faster than traditional GPUs, and are usable in a wider range of data centres. General Compute claims the SN50s can deliver up to 16× faster inference than GPU‑based cloud services. The key challenge remains securing large volumes of these chips, especially for a relatively new company.
Upper90’s prior experience and the emergence of chip‑backed financing
Billy Libby, Upper90’s co‑founder and CEO and a former Goldman Sachs quantitative trader, has previously used a similar approach. Upper90 financed Crusoe’s GPU purchases in 2021, a deal Libby describes as the first loan backed by advanced chips. At the time traditional lenders avoided such structures due to concerns about GPU depreciation and market uncertainty. As CoreWeave and others made chip‑backed lending a core part of their business models and CoreWeave later executed a successful public listing, this form of financing became more mainstream.
Libby says that with the GPU market now more transparent — and in some views overbought — Upper90 is looking to back companies positioned for the next wave of AI growth, particularly those focused on inference. “We believe open‑source models will play a defining role, so last year we sought a player focused on inference,” Libby said.
Market context and alternatives
Recent market developments support demand for inference‑oriented infrastructure. Companies such as OpenRouter and Fireworks, which provide access to open models, have closed funding rounds at high valuations. New models like Kimi K3 have shown competitive performance on programming benchmarks versus leading lab models. At the same time, newer chipmakers including Groq and Cerebras have attracted buyer and investor interest.
Strategic importance of non‑Nvidia options
For General Compute it is strategically important to secure chips outside the Nvidia ecosystem. Other infrastructure firms, such as TensorWave, pursue similar diversification — in TensorWave’s case through a partnership with AMD. As more non‑Nvidia options appear, providers that are not exclusively tied to Nvidia technology could gain an edge in delivering lower‑cost inference services.
Puklowski noted that a number of chips are now scaling that offer very favourable total cost of ownership (TCO), and in some cases can outperform Nvidia on speed, though their buyer bases remain limited. He argued the Upper90 deal is more than financing for a single startup’s compute purchases; it signals capital reallocating into other parts of the AI stack and the beginning of an erosion of Nvidia’s near‑monopoly.
Conclusion
The $400 million loan to General Compute highlights how capital markets are starting to adapt to a changing AI infrastructure landscape, supporting inference‑optimized silicon as collateral. The transaction may presage broader financing flows toward alternative, cost‑efficient compute platforms for running large language models and other AI inference workloads.



