Chips & Hardware
Microsoft launches Maia 200 chip to scale AI inference
Microsoft has launched the Maia 200, a new AI inference chip designed to run models at faster speeds and with more efficiency while reducing reliance on Nvidia GPUs.
Microsoft has launched the Maia 200, a chip designed for scaling AI inference—the computing process of running an AI model, as opposed to training it. Announced on Monday, the new chip succeeds the Maia 100, which was released in 2023. As AI companies mature, the costs associated with inference have become an increasingly important part of overall operating costs, driving industry interest in optimizing the process. Microsoft is hoping that the Maia 200 can assist with this optimization, allowing AI businesses to run with less disruption and lower power use. According to Microsoft, the Maia 200 is designed to run AI models at faster speeds and with more efficiency than its predecessor. The hardware is equipped with over 100 billion transistors. In terms of computing speed, it delivers over 10 petaflops in 4-bit precision (FP4) and approximately 5 petaflops of 8-bit performance (FP8), representing a substantial performance increase over the 2023 model.
The launch positions Microsoft to compete directly with other technology giants designing their own chips to optimize inference costs and lessen dependence on Nvidia GPUs. In its announcement on Monday, Microsoft compared the chip’s performance to rival hardware from Amazon and Google. According to Microsoft, the Maia 200 delivers 3x the FP4 performance of Amazon’s third-generation Trainium chips, which launched in December. Microsoft also claims the chip delivers FP8 performance above Google’s seventh-generation Tensor Processing Unit (TPU). While Google’s TPUs are not sold as individual chips but are instead made accessible as compute power through its cloud, Amazon offers Trainium as its own AI accelerator chip. In both cases, these alternatives can be used to offload compute workloads that would otherwise require Nvidia GPUs, thereby lowering overall hardware costs.
The Maia 200 is already active within Microsoft’s internal operations. The company stated that the chip is currently used by its Superintelligence team to fuel its AI models, and it is also supporting Copilot, Microsoft’s chatbot product. To demonstrate the chip’s capacity, Microsoft noted, “In practical terms, one Maia 200 node can effortlessly run today’s largest models, with plenty of headroom for even bigger models in the future.” Additionally, as of Monday, Microsoft has invited external parties—including developers, academics, and AI labs—to use the Maia 200 software development kit in their workloads.
Why it matters
Microsoft is positioning the Maia 200 to compete with self-designed chips from other tech giants like Google and Amazon, aiming to lessen dependence on Nvidia GPUs and optimize inference costs.