ZML Launches Free AI Inference Server for Multiple Chips

French Startup ZML Releases Free Product to Speed AI Inference Across Diverse Chips

Hot French AI startup ZML, founded by Turing Award winner Yann LeCun, has released a high-performance software tool that allows large language models (LLMs) to run on a wide range of AI chips — including Nvidia, AMD, Google’s TPU, Apple Metal, and Intel Arc. The newly launched LLM inference server, called ZML/LLMD, aims to break existing silos and make different chips available for AI use cases at their maximum available speed, and sometimes faster, according to ZML founder Steeve Morin.

Inference Optimization: A Growing Priority

As AI becomes more integrated into work and daily life, optimizing inference — the processing of prompts — has outpaced training in importance, yet often feels patchy behind the scenes due to software and architecture barriers that lead to vendor lock-in, Morin said. ZML/LLMD promises to achieve peak performance across a variety of chips, a technological feat that could disrupt the market amid mounting fears over AI-related costs. ZML hopes to provide enterprises and clouds with the option to use a mix of chips, some of which might be less costly or consume less energy. “The idea is to give people back the power to create their own system and achieve real efficiency gains that allow [AI] to be disseminated,” Morin added.

The Competitive Landscape

Inference has been an area of intense investment, with the trend dubbed the “inference is the new training” phenomenon. ZML faces competition from vLLM, recently valued at $13 billion; SGLang, from the creators of the open-source project; as well as others. Both vLLM and SGLang partially compete with LLMD, but Morin’s ambitions for ZML cover a broader spectrum. “We have reached the point where we are co-designing silicon,” he said. Morin credited ZML’s lean team of 20 people as the reason why the Paris-based startup has been able to move fast, with more releases planned.

A European AI Ecosystem

Such a software assist may help novel AI chipmakers, many of which happen to be from Europe. Morin cited Blaize, Cerebras, Groq, d-Matrix, SiPearl, AxeleraAI, Graphcore, Untether AI, and MatX. He emphasized that what matters is that ZML can work with them on “things that haven’t been done before anywhere in the world.” Morin is not bearish on Nvidia, partly due to its existing supply, and noted that ZML has a good relationship with the AI chip giant, which has been instrumental for the rise of inference.

Funding, Team, and Future Plans

ZML’s small team is well funded. Thanks to Morin’s track record as VP of engineering of a company that Snapchat acquired, he raised $20 million from venture firms including Harry Stebbings’ 20VC, >commit, AALVC, Drysdale Ventures, Xavier Niel’s Kima Ventures, Kindred Capital, LocalGlobe, and Puzzle Ventures. The startup’s cap table includes notable founders: Dagger and Docker founder Solomon Hykes, Clément Delangue and Julien Chaumond from Hugging Face, as well as Yann LeCun, now with Meta. This builds the case that Europe’s AI startups are thriving. “I couldn’t do ZML anywhere but in Paris,” Morin said.

Unlike ZML’s first public project, the inference-focused ZML released in 2024 and open-sourced, ZML/LLMD is not open source. It is launching as a free product with the goal of learning about usage. “I’d rather measure and [then generate revenue] where it is most effective without hindering my growth stupidly because I have been too greedy from the get-go,” Morin said. It is too early to tell when ZML/LLMD might become a paid product, and what its adoption will look like.

Leave a Comment