Text settings Story text Size Small Standard Large Width * Standard Wide Links Standard Orange * Subscribers only Learn more Minimize to nav ByteDance is training an AI model that could approach the size of Anthropic’s most cutting-edge Mythos system, as Chinese companies continue to narrow the gap with the top US labs.
The Chinese tech giant is at an early stage of training a model with as many as 10 trillion parameters—three times larger than Moonshot’s Kimi K3, the biggest Chinese model released to date, according to three people with knowledge of the matter.
The ByteDance model is being pre-trained—a stage that typically takes three to six months—before it is fine-tuned and released if all goes well, one of the people said. The exact model size would only be determined at a later stage.
Anthropic doesn’t disclose the size of its models, but industry estimates say its most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion. While parameter count sets the fundamental capacity or memory limits for the models to store information, actual capability also depends on other factors such as data quality and training methods.
ByteDance’s efforts to train one of the world’s largest AI models show Chinese labs’ ambition to not only catch up but outperform their US peers in the most advanced level of AI.
In the past weeks alone, Chinese models from Moonshot and Alibaba show strong performance on benchmarks, lagging behind only Anthropic’s Fable 5 in certain areas. Mythos 5, Anthropic’s most advanced model, is only available to approved organisations after a temporary ban in June due to security concerns.
Industry insiders say multiple Chinese labs are in the process of training models of the size of Fable 5, while ByteDance is currently the most ambitious in pushing for the largest.
ByteDance, the parent of viral video platform TikTok, has kept a low profile in its AI development as its models are mostly closed, unlike many of its Chinese peers. Its latest SeeDance model ranks among the most advanced globally in video generation, while its flagship consumer-facing model Doubao is the most popular in China with 324 million monthly active users.
Over the past three years, ByteDance has invested in AI more aggressively than any of the other Chinese tech giants, building out its network of data centers and hiring researchers. It has doubled down on its cloud unit, Volcano Engine, which sells AI solutions to enterprises. ByteDance also has ambitions to develop custom AI chips.
Seed, its model development team led by former Google DeepMind scientist Wu Yonghui, has about 2,000 members in China and overseas. The team includes core researchers, infrastructure engineers, a data labelling team and translators.
Its model development has also implemented a more independent approach that does not involve “distilling” existing models from other labs, according to one of the people. This approach has been in place for more than a year, which some believe has led to its slower development versus rivals.
Model distillation is the process of compressing a large, complex AI model into a smaller, faster one by training the smaller model to copy the knowledge and outputs of the larger teacher.
ByteDance’s management, led by founder Zhang Yiming, believes only independent development can result in a model that outperforms rivals.
Zhang reiterated his stance in an internal meeting two weeks ago, where he told the Seed team to target “world-leading model capabilities” in the long run without getting too worried about falling behind in the near term, according to the person.
Chinese media Latepost and The Information first reported Zhang’s comments from the recent internal meeting.
ByteDance did not respond to a request for comment.
---
**İlgili Kaynaklar:**
Bu alanda profesyonel destek için [yapay zeka firması](https://yapayzekafirmasi.com) sayfasını inceleyebilirsiniz.