Chinese companies expand their bet on AI models with trillions of parameters
Chinese companies such as DeepSeek, Moonshot AI and Alibaba are increasing the scale of their models: DeepSeek-V3 has 671 billion parameters, Kimi K2 reached 1 trillion, Kimi K3 reached 2.8 trillion and Qwen3.8-Max has 2.4 trillion. In Qwen3.8-Max, about 95 billion parameters are activated per step, using a Mixture of Experts architecture, according to Xataka.
Alibaba confirmed on September 22 that Qwen 4 is in training and that its roadmap for Qwen 4.5 and Qwen 5 envisions models with 5 trillion to 10 trillion parameters. ByteDance is reportedly pursuing a similar scale, according to the Financial Times, without public confirmation from the company. Xataka notes that more parameters do not guarantee better performance, while small models remain useful because they require less memory, computing power and cost.
Why it matters · editorial interpretation
The billion- and trillion-parameter scale described could raise memory and computing requirements, although Mixture of Experts activates only part of the parameters per step. As Xataka notes, larger models do not guarantee better performance, leaving room for smaller, more economical alternatives.