Li Auto Explores Self-Developed Cloud Inference Chips Using Same Dataflow Architecture

Li Auto Explores Self-Developed Cloud Inference Chips Using Same Dataflow Architecture

Li Auto is reportedly working on a cloud inference chip that will reuse the dataflow architecture from its existing smart-driving NPU, aiming to offload inference work from GPUs at a fraction of the R&D cost.

By CarsEVs Editorial Team

Source: www.ithome.com

Share

Li Auto is exploring the development of a self-designed cloud inference chip that would reuse the dataflow architecture behind its existing smart-driving NPU, multiple sources told WanAuto on Aug. 14. The project is still in its early stages.

Cloud data centers currently rely heavily on GPUs for both model training and inference, though a growing number are beginning to deploy dedicated inference chips. An industry insider told WanAuto that such chips could take over inference tasks now running on GPUs — including smart-driving model data processing, testing, simulation, and large-language-model request handling. Tasks involving parameter updates during model training, however, would still require training chips.

A familiar architecture, a new form factor

Dataflow architecture is technically feasible for cloud deployment, the insider said. One proposed path is to adapt the vehicle NPU's compute design, packaging multiple AI-compute tiles together with high-bandwidth memory and high-speed interconnects to form a larger-scale inference chip.

That approach would let Li Auto reuse portions of its chip design and software toolchain across both vehicle and cloud platforms, spreading R&D costs between the two.

Vehicle chips are optimized for low power consumption, low latency and steady execution — their models, sensor inputs and execution cadence are relatively fixed. Cloud inference chips, by contrast, must handle more models simultaneously, accommodate faster model updates, process inputs of varying length, and manage highly variable request concurrency. They also demand high-bandwidth memory, multi-chip interconnect, dynamic batching and cluster-level scheduling.

Li Auto's MACH M100: the foundation

The effort builds on Li Auto's MACH M100, unveiled earlier this year at ISCA 2026, where the company became the first Chinese automaker to present since the conference's industrial track launched in 2020. The MACH M100 is billed as the world's first high-compute edge inference chip based on dataflow architecture. It is manufactured on a 5 nm automotive-grade process, delivers 1,280 TOPS of single-chip compute power with a peak utilization rate of 82%, and features eight-channel LPDDR5X memory with a peak bandwidth of 273 GB/s. The CPU side incorporates a 24-core ARM Cortex-A78AE cluster.

Li Auto shared the stage with researchers from Google and Meta at ISCA 2026 to discuss the future direction of AI computing architectures.

Privacy settings

Choose which cookies CarsEVs may use. Your choice is stored on this device.