DeepSeek reportedly plans a 160,000-chip Huawei cluster for AI inference, signaling a major Chinese effort to reduce reliance on Nvidia hardware.

DeepSeek is reportedly planning to deploy at least 160,000 Huawei AI processors at a data center in Inner Mongolia, a project that could become the largest known Huawei chip cluster if completed. The proposed system would be used for inference rather than model training, according to reporting cited by The Decoder, marking a potentially important shift in how Chinese AI companies are building around restrictions on Nvidia hardware.
The plan centers on Huawei’s next-generation Ascend-950DT processors. It has not been presented in the supplied evidence as a completed order, an operational deployment, or a publicly confirmed announcement from DeepSeek or Huawei. Reports instead describe an intended build whose delivery could take more than a year because of chip-production constraints and shortages of high-bandwidth memory.
The distinction between inference and training is central to the story. Training requires sustained processing across large numbers of accelerators while a model learns from data. Inference is the production phase in which a trained model responds to user requests, generates content, or performs tasks for applications.
The proposed DeepSeek cluster would reportedly focus on inference workloads. The Decoder, citing Bloomberg’s reporting, said DeepSeek continues to use Nvidia hardware for the heavier training workloads. That suggests the project is not an immediate replacement for every part of DeepSeek’s computing infrastructure. Instead, it would create a large domestic platform for serving models after they have been trained.
At least 160,000 processors would still represent a substantial deployment. If delivered and connected effectively, the system could support high-volume model access and reduce the company’s dependence on imported accelerators for production services. The practical outcome would depend on more than the processor count, however: networking, software compatibility, power, cooling, memory capacity, and cluster management would all affect usable performance.
The reported schedule is uncertain. The Decoder said Huawei may not be able to deliver the full quantity for more than a year, citing production limits and memory shortages. That caveat makes the announcement as much a story about supply-chain capacity as about DeepSeek’s procurement plans.
AI accelerators rely on high-bandwidth memory to move data quickly between the processor and the workloads it is serving. The report said China’s CXMT has begun producing small batches of HBM3E, a memory technology used by AI processors, but remains several years behind Samsung, SK Hynix, and Micron, which are already mass-producing HBM4.
Those comparisons indicate why a large domestic accelerator project may be difficult to scale quickly. Even if Huawei can design and assemble the processors, securing enough advanced memory and packaging capacity could determine whether a 160,000-chip plan becomes an operating cluster or a longer-term target. The evidence does not establish a firm delivery date, final system configuration, or whether the entire proposed quantity has been contractually secured.
The three supplied reports agree on the core claim: DeepSeek is associated with a planned 160,000-chip Huawei deployment, and the target location is Inner Mongolia. CNBC TV18 and tech-ish.com frame the plan as an effort to reduce reliance on Nvidia while Nvidia products remain constrained in China. The Decoder provides the more detailed account of the proposed chip type, inference-only role, supply bottlenecks, and continued use of Nvidia for training.
However, the evidence is media reporting rather than a public technical specification or company announcement. No source supplied here confirms that Huawei has accepted the full order, that construction is complete, or that the cluster has entered testing. The claim that this would be the largest known Huawei chip cluster is also a reported comparison, not an independently audited industry ranking.
That distinction matters for AI buyers and researchers. A planned processor count does not by itself reveal throughput, cost per request, uptime, or compatibility with the software stack used by DeepSeek. Those results can only be assessed after the system is delivered and benchmarked under production conditions.
For DeepSeek, moving inference onto Huawei hardware could provide greater control over supply and deployment capacity in China. It could also make serving costs and availability less dependent on Nvidia procurement, export controls, or access to imported systems. But migrating inference at this scale may require software optimization, model conversion, scheduling changes, and extensive testing across Huawei’s hardware and software ecosystem.
For enterprise AI teams, the proposed cluster highlights a broader procurement lesson: accelerator diversification is not simply a matter of buying a different chip. Teams must evaluate the complete platform, including compiler support, inference frameworks, memory bandwidth, networking, observability, and the availability of replacement hardware.
The project also illustrates the different bottlenecks facing China’s AI industry. Export restrictions may encourage demand for domestic accelerators, but demand alone does not solve advanced-memory shortages or manufacturing limits. Huawei’s ability to deliver at scale—and DeepSeek’s ability to achieve competitive inference performance on the resulting cluster—will be more informative than the headline processor count.
The clearest follow-up signal will be confirmation from DeepSeek or Huawei that the project exists as a committed order rather than an internal plan. Reporting on construction progress, delivery schedules, and the number of Ascend-950DT chips actually installed would establish whether the proposed scale is being reached.
Technical disclosures will be equally important. Watch for independent measurements of inference throughput, latency, energy use, and cost compared with Nvidia systems. Details about the software stack, model support, networking design, and memory configuration would show whether the cluster can serve real workloads efficiently.
Finally, the availability of HBM3E from CXMT and Huawei’s production output will indicate whether China can expand domestic AI infrastructure beyond individual deployments. If supply remains limited, the DeepSeek project may become a long-term capacity plan rather than a near-term substitute for Nvidia-based systems.
DeepSeek’s reported plan is significant because it separates AI infrastructure dependence into two stages. China may be able to keep using Nvidia systems for the most demanding training workloads while shifting large-scale inference to domestic accelerators. That is a narrower form of substitution than replacing Nvidia across the full AI stack, but it could still matter for consumer-facing services and enterprise deployments where inference demand is persistent.
The story should be judged by delivery and operating results, not the proposed number of chips. Until Huawei’s supply capacity, DeepSeek’s deployment, and independent performance data become clearer, the 160,000-processor figure is best treated as a major reported ambition—and a test of whether China’s domestic AI hardware ecosystem can scale from strategic plans to dependable production infrastructure.