DeepSeek is training a model with 2 trillion parameters and plans to develop an 8-trillion-parameter model soon after, according to a report by The Information cited by Wccftech. The target, if achieved, would mark a dramatic escalation in open-weight model scale from the lab’s current V4.1 Flash, which carries 552 billion parameters.
Key points
- DeepSeek is training a 2-trillion-parameter model and targeting 8 trillion parameters for a subsequent release, according to The Information
- The lab’s V4.1 Flash model has 552 billion parameters and already outperforms larger proprietary rivals on some coding benchmarks
- DeepSeek has ordered 160,000 Huawei Ascend 950DT chips for its 1 GW data centre in Inner Mongolia, at approximately $16,000 per chip
- CEO Liang Wenfeng disclosed in July 2026 that the lab possessed compute equivalent to 20,000 Nvidia H100 GPUs
DeepSeek’s parameter roadmap and Huawei chip order
The reported training run at 2 trillion parameters represents a near-quadrupling from V4.1 Flash, and the planned 8-trillion-parameter model would be more than fourteen times larger still. The Information’s report says the 8-trillion target is to follow “soon after” the 2-trillion model completes training.
The scale leap depends on hardware that DeepSeek has already committed to buy. The lab placed an order for 160,000 Huawei Ascend 950DT chips for its 1 GW data centre in Inner Mongolia, Wccftech reports, with the street price of each chip at around 111,000 Yuan — approximately $16,000. At that unit price, the total order would amount to roughly $2.56 billion. That price sits below the $15,000 to $25,000 range quoted for Nvidia’s H20 GPU.
DeepSeek’s CEO, Liang Wenfeng, disclosed in July 2026 that the lab possessed compute resources equivalent to 20,000 Nvidia H100 GPUs, most of which had arrived in the preceding month or two. The Ascend 950DT order suggests the lab is now building out a second, distinct compute base using Huawei’s domestic alternative to Nvidia hardware.
Benchmark context and competitive landscape
DeepSeek’s current V4.1 Flash already punches above its parameter weight on some tasks. Grok 4.7 ranks below DeepSeek V4.1 Flash on Terminal-Bench 4.0, a benchmark that Epoch AI characterises as “flawed”, Wccftech reports.
DeepSeek does not possess multiple gigawatts of compute like SpaceX does. The implication is that the Chinese lab is achieving competitive results with less total power, though the comparison depends on how efficiently each lab’s hardware is utilised and whether the cited gigawatt figures refer to peak draw, average consumption, or contracted capacity.