Compute & Cloud
ZML launches free tool to run AI models across any chip
ZML, a Paris-based startup backed by Yann LeCun, launched free inference software LLMD that runs open source LLMs across Nvidia, AMD, Google TPU, Apple Metal, and Intel Arc chips.
ZML has released ZML/LLMD, a new inference server that lets open source large language models run across Nvidia, AMD, Google TPU (Tensor Processing Unit), Apple Metal, and Intel Arc chips. Founder Steeve Morin said the Paris-based startup’s ambition with LLMD is to break existing silos and make different chips available for AI use cases at their maximum available speed, and sometimes faster.
Morin said optimizing inference — the processing of prompts, as opposed to training models — has been outpacing training in importance, but the process often feels patchy behind the scenes because software and architecture barriers cause vendor lock-in. ZML hopes to give enterprises and clouds the option to mix chips, some of which might be less costly or consume less energy. “The idea is to give people back the power to create their own system and achieve real efficiency gains that allow [AI] to be disseminated,” Morin said. That effort comes amid what has been called the inference gold rush — a period of intense investment in the space.
ZML’s lean team of 20 people is, according to Morin, the reason the startup has been able to move fast, with more releases in the plans. The company raised $20 million from venture firms including 20VC (run by Harry Stebbings), >commit, AALVC, Drysdale Ventures, Kima Ventures (run by Xavier Niel), Kindred Capital, LocalGlobe, and Puzzle Ventures — a fundraise helped by Morin’s track record as VP of engineering at Zenly, which Snapchat acquired for nine figures in 2017. Morin described ZML’s approach as reaching the point of co-designing silicon. ZML faces competition from Baseten, recently valued at $13 billion, Inferact, founded by the creators of open source project vLLM, and RadixArk, the commercial company behind SGLang; both vLLM and SGLang partially compete with LLMD. Unlike ZML’s first public project — an inference-focused ML framework released in 2024 and updated in March — LLMD is not open source; it is launching as a free product with the goal of learning about usage.
Morin said a software assist from ZML may help novel AI chipmakers, many of which happen to be from Europe, citing Axelera, Fractile, Kalray, OLIX, Q.ANT, SiPearl, SpiNNcloud, and VSORA. It is too early to tell when LLMD might become a paid product, or what its adoption will look like. ZML’s cap table shows other founders are paying attention, including Solomon Hykes of Dagger and Docker, Clément Delangue and Julien Chaumond of Hugging Face, and LeCun, now with AMI Labs.
Why it matters
Achieving peak inference performance across chip brands is a technical feat that could disrupt the market amid mounting fears over AI-related costs, and it may particularly help novel AI chipmakers, many of them European, gain traction.