French Startup ZML Launches Cross-Chip AI Inference Server
French AI startup ZML launches ZML/LLMD, a free LLM inference server that runs open source models across a wide range of chips including Nvidia, AMD, Google TPU, Apple Metal, and Intel Arc. Backed by Turing Award winner Yann LeCun, the company aims to break down software barriers that lock enterprises into single-chip vendors and limit hardware flexibility.
Founder Steeve Morin emphasizes that inference optimization now surpasses model training in importance as AI integrates into everyday workflows. By enabling enterprises to mix and match chips, ZML allows organizations to choose less costly or more energy-efficient hardware for different workloads. Morin notes the platform also supports emerging European chipmakers like Axelera, SiPearl, and Kalray, fostering innovation in silicon design.
ZML faces strong competition from well-funded rivals like Baseten, Inferact, and RadixArk, but Morin positions ZML across a broader spectrum that includes co-designing silicon. He maintains a positive relationship with Nvidia, acknowledging its supply dominance while advocating for a multi-chip future that gives users control over their own systems and cost structures.