Can Chinese Chips Outperform Nvidia? 5 AI Models Built Locally
The Rise of Domestic AI Hardware in China
Chinese artificial intelligence (AI) has made significant strides in recent years, with models becoming increasingly competitive on the global stage. However, despite these advancements, the country's AI hardware still lags behind its US counterparts. While domestic chips are now widely used for model inference, none of China’s top AI models have been pre-trained on homegrown silicon. This gap highlights a critical challenge in the development of a fully self-sufficient AI ecosystem.
To understand this gap, it is essential to examine the three main stages of AI model development:
- Pre-training: This is the most computationally intensive phase, where a model learns basic patterns by processing massive datasets.
- Post-training: A less intense process that fine-tunes the model to follow specific human instructions.
- Inference: The final stage where the trained model runs to answer user queries and instructions.
Driven by Washington's export controls and Beijing's push for technological self-sufficiency, a growing number of Chinese AI labs are experimenting with shifting earlier training phases onto domestic hardware. While relying on indigenous suppliers may slow down development compared to US counterparts, it is paving the way for a unique domestic AI supply chain.
Case Studies: Chinese AI Models Using Domestic Hardware
Zhipu AI's GLM-Image
In January, Beijing-based Zhipu AI open-sourced its new image generation model, GLM-Image, developed alongside Huawei Technologies. The model was trained using Huawei's Ascend Atlas 800T A2 server, powered by the Ascend 910 AI accelerator, and its MindSpore deep learning framework, according to Zhipu.

Zhipu claims this was the first state-of-the-art multimodal model trained entirely on domestic chips. However, training image models generally requires less computing power than large language models (LLMs). Zhipu is still working towards that milestone. When asked about moving its LLM training to Huawei hardware, the company responded with a "fighting" emoji.
Meituan's LongCat-2.0-Preview
In April, on-demand services giant Meituan invited users to test its new trillion-parameter AI model, LongCat-2.0-Preview. The company stated that both training and inference were completed entirely on a "domestic computing cluster."
Meituan has not disclosed which local accelerators were used, noting only that the training stage required 50,000 to 60,000 domestic chips. The company has yet to officially release the model to the public.
ModelBest's Lightweight On-Device Models
Beijing-based start-up ModelBest focuses on smaller LLMs that run locally on smartphones, PCs, and vehicle cockpits. Last month, the company open-sourced a 1.58-bit ternary model named BitCPM-CANN. The tiny model comes in four sizes ranging from 0.5 billion parameters to 8 billion parameters, designed to compress weights for maximum efficiency without additional physical memory.
BitCPM-CANN was said to be trained on Huawei's Ascend hardware and appeared to be named after its Compute Architecture for Neural Networks (CANN), an equivalent to Nvidia's CUDA software development toolkit. ModelBest claimed the achievement proved the stereotype that "domestic chips can only run inference" was outdated.
The start-up also trained its billion-parameter MiniCPM5-1B model on Ascend hardware, which topped Artificial Analysis' intelligence index for open-weights models under 2 billion parameters, beating Alibaba Group Holding's Qwen series.
Post-training for DeepSeek-V4-Pro
In June, a research team from Huawei and the Shenzhen Loop Area Institute used Huawei's Ascend 910C chips to conduct post-training for DeepSeek-V4-Pro. The team completed "full-parameter" post-training for the 1.6-trillion-parameter flagship model on a cluster powered by at least 1,000 Huawei chips. However, post-training is significantly less computationally intensive than pre-training.
Peking University's EvoPhys-World
Earlier this month, a team at Peking University released EvoPhys-World, a 5D world model that simulates movements in physical spaces. The model took the top spot on Stanford University's WorldScore benchmark.

The team trained the model using Chinese chip designer Moore Threads Technology's MTT S5000 graphics processing unit and its Musa platform, an alternative to Nvidia's CUDA, as stated in a blog post. Moore Threads claimed the GPU performed virtually on par with unnamed "global mainstream" chips in training throughput while delivering nearly identical inference quality.
Conclusion
As Chinese AI labs continue to experiment with domestic hardware, the country is making strides toward building a self-sufficient AI ecosystem. While challenges remain, the progress being made is a testament to the growing capabilities of China's technology sector.