Microsoft has formally unveiled Maia 200, its second-generation in-house AI accelerator designed specifically for inference workloads. The chip is already being deployed inside Azure data centers and is positioned as a more efficient way to serve large-scale AI models, according to the company.
TL;DR
- Maia 200 is Microsoft’s second-generation custom AI accelerator focused on inference.
- Built on TSMC’s 3nm process with over 100 billion transistors.
- Includes 216 GB of HBM3e memory and supports FP4, FP8, and FP16 precision.
- Performance claims are from Microsoft, with no independent benchmarks yet available.
What Microsoft Announced
Microsoft announced Maia 200 as the latest step in its custom silicon roadmap for Azure. Unlike training accelerators, which are optimized for building AI models, Maia 200 is purpose-built for inference, the phase where trained models generate outputs for real-world applications such as copilots, chat systems, and other generative AI services.
The company confirmed that Maia 200 is already running in production inside Azure data centers, beginning with deployments in the U.S. Central region. Microsoft said additional regions will follow, though it did not provide a public timeline for broader availability.
Technical Specifications Shared by Microsoft
Microsoft has disclosed a limited but specific set of technical details for Maia 200. The chip is manufactured by Taiwan Semiconductor Manufacturing Company using a 3-nanometer process and contains more than 100 billion transistors.
Maia 200 is paired with 216 gigabytes of HBM3e high-bandwidth memory. Microsoft says this large memory footprint is designed to support modern large language models and reduce memory bottlenecks during inference.

The accelerator supports FP4, FP8, and FP16 precision modes. Microsoft highlighted FP4 and FP8 as particularly important for inference workloads, where lower-precision computation can significantly increase throughput and efficiency without materially affecting output quality.
According to Microsoft, Maia 200 can deliver more than 10 petaFLOPS of compute performance in FP4 mode. These figures are based on internal measurements provided by the company.
Interconnect and System Design
Microsoft described Maia 200 as part of a broader system-level design rather than a standalone chip. Multiple Maia accelerators are connected using a proprietary high-speed interconnect within Azure servers.
The company said this interconnect is optimized for inference-scale workloads that require fast data movement between accelerators and memory. Microsoft has not published detailed bandwidth or latency specifications for the interconnect.
Performance Claims and Attribution
Microsoft claims Maia 200 delivers improved performance per dollar compared with prior hardware used for inference inside Azure. The company has also suggested the chip is competitive with other cloud providers’ in-house inference accelerators in certain low-precision scenarios.
All performance comparisons released so far are Microsoft’s own claims. Independent benchmarks comparing Maia 200 with competing chips have not yet been published, and real-world performance will depend on workload characteristics, software optimization, and deployment configuration.
Software Support and Availability
Alongside the hardware announcement, Microsoft introduced a Maia software development kit to help internal teams and select customers optimize inference workloads for the new accelerator. The SDK supports common frameworks such as PyTorch and includes tools for compiling and tuning models for Maia hardware.
Microsoft has not disclosed pricing details or a general availability date for Maia-backed Azure instances. The company also has not indicated whether Maia chips will ever be offered outside its own data centers.
What Remains Unknown
Microsoft has not released independent benchmark data, detailed power consumption figures, or standardized performance tests for Maia 200. The company has also not clarified how Maia 200 will be positioned alongside GPUs within Azure pricing tiers.
Until third-party benchmarks and broader customer deployments emerge, Maia 200’s real-world competitiveness will remain based primarily on Microsoft’s own measurements and early internal usage.




















