According to the Financial Times, ByteDance, the parent company of TikTok, is currently pretraining an AI model that could reach as many as 10 trillion parameters. If realized, that scale would put the company in the range of Anthropic’s leading systems and make the model roughly three times larger than the Moonshot Kimi K3, previously regarded as China’s largest model.
Development stage and parameter context
The project is in an early pretraining phase, which typically lasts three to six months. The model’s final size will be decided in later stages. Anthropic does not publicly disclose parameter counts; industry estimates put Mythos 5 at about 8 trillion parameters and Fable 5 at roughly 5 trillion parameters.
Parameter count is one measure of a model’s capacity, but actual performance depends heavily on data quality and training methods.
The Chinese landscape and competition
ByteDance’s ambitions align with broader efforts across Chinese AI labs: in recent weeks, Moonshot and Alibaba models posted strong benchmark results and in some areas fell only behind Anthropic’s Fable 5. Industry sources say several Chinese teams are training models at roughly the Fable 5 scale, but ByteDance appears to have the most ambitious plans.
The Seed team and infrastructure
The work is carried out by ByteDance’s Seed team, led by Vu Jong-huj, a former researcher at Google DeepMind. The group comprises about 2,000 people working in China and abroad, including researchers, infrastructure engineers, data labelers and translators.
Over the past three years ByteDance has been among the most aggressive Chinese tech investors in AI: expanding its data-center network, developing the Volcano Engine cloud service, and planning its own AI chips.
Methodology and leadership stance
The Seed team has not used knowledge distillation for over a year — that is, they have not trained models to mimic outputs from other labs’ larger models. Knowledge distillation involves teaching a smaller model to reproduce a larger ‘‘teacher’’ model’s outputs; ByteDance is instead pursuing independent development.
Founder Zhang Yiming (Csang Ji-ming) believes that only independent development can produce models that surpass rivals. According to the Financial Times, he reiterated this view at an internal meeting two weeks ago and urged the team to focus on achieving world-leading model capabilities over the long term rather than worrying about short-term gaps.
Why this matters
If ByteDance’s model reaches the order of 10 trillion parameters, it would be a notable milestone for Chinese AI and in the global race. However, final performance will hinge not only on model size but also on the quality of training data, training techniques and other engineering choices.
This article was prepared with the assistance of an AI tool; the final content was edited and verified by our journalist.



