Tools

French startup ZML launches free LLMD inference server to link diverse AI chips

Paris-based ZML released ZML/LLMD, a free inference server that aims to run open-source large language models across a wide range of processors — from Nvidia GPUs to Apple Metal and Google TPUs — to reduce vendor lock-in and improve cost and energy efficiency.

French startup ZML launches free LLMD inference server to link diverse AI chips

Paris-based startup ZML has released ZML/LLMD, a new inference server intended to run open-source large language models (LLMs) across a wide range of processors — including Nvidia, AMD, Google TPU, Apple Metal and Intel Arc — with the goal of reducing vendor lock-in and improving performance.

What the product does and why it matters

Inference, the real-time processing of prompts, has grown in importance as AI is embedded into business and consumer workflows. Steeve Morin, founder of ZML, says that software and architectural barriers often prevent efficient use of heterogeneous hardware, slowing deployment and locking users into particular vendors.

ZML/LLMD aims to close those gaps by enabling different chips to reach their peak available performance — and in some cases exceed common expectations, according to Morin. Beyond the engineering challenge, the company presents this as a potential cost and energy advantage: allowing enterprises and cloud providers to mix chip types could lower expenses and reduce power consumption.

Working with emerging chipmakers

Morin highlighted that ZML/LLMD could support newer chip vendors, many of them European, naming Axelera, Fractile, Kalray, OLIX, Q.ANT, SiPearl, SpiNNcloud and VSORA. He emphasized that the important point is not the companies’ location but ZML’s ability to collaborate on integrations and optimizations that haven’t been implemented elsewhere.

Relationship with Nvidia and competitive landscape

Morin does not position Nvidia as an adversary; ZML maintains a good relationship with the chip giant, which is itself preparing for a larger role in inference. The market is nonetheless competitive, with firms such as Baseten (reported in the article as recently valued at $13 billion), Inferact (from the creators of the vLLM open-source project) and RadixArk (the commercial company behind SGLang) pursuing overlapping opportunities. vLLM and SGLang partially compete with LLMD, but ZML says its ambitions span a broader set of capabilities, including closer co-design with silicon.

Team, funding and product model

ZML is a compact team of around 20 people, a size Morin credits for the company’s speed in shipping releases. Morin’s prior role as VP of engineering at Zenly, which Snapchat acquired in 2017 for a nine-figure sum, helped him raise capital: ZML secured $20 million from investors including Harry Stebbings’ 20VC, >commit, AALVC, Drysdale Ventures, Xavier Niel’s Kima Ventures, Kindred Capital, LocalGlobe and Puzzle Ventures.

Unlike ZML’s first public project — an inference-focused ML framework published in 2024 and updated in March — ZML/LLMD is not open source. The company is launching the server as a free product to gather usage data and learn how users adopt it. Morin said the team would prefer to measure usage and introduce revenue where it is most effective rather than hamper growth by monetizing too early.

Outlook

It remains unclear when ZML/LLMD will become a paid product or how broadly it will be adopted. The investor roster, which includes figures such as Solomon Hykes (founder of Dagger and Docker), Clément Delangue and Julien Chaumond of Hugging Face, and Yann LeCun (now with AMI Labs), suggests the market is paying attention. Morin also argued that ZML’s development in Paris demonstrates that European AI startups can build significant AI infrastructure locally: "I couldn’t do ZML anywhere but in Paris," he said.

ZML/LLMD represents both a technical effort to optimize inference across heterogeneous hardware and a market experiment to make multi-chip deployments more practical and cost-effective for enterprises and clouds.