ByteDance trains large model to rival Anthropic
The Chinese tech giant is pre-training a model that could reach 10 trillion parameters.
Why it matters
The reported training run signals continued competition between Chinese AI companies and top U.S. labs at frontier-model scale. It also highlights model provenance and distillation as live issues in how labs differentiate new systems.
The key points
- 1.ByteDance is reportedly pre-training a model up to 10 trillion parameters.
- 2.The final model size has not yet been determined.
- 3.The effort is framed as avoiding AI distillation.
ByteDance is at an early stage of pre-training an AI model that could have as many as 10 trillion parameters, according to Ars Technica, citing three people with knowledge of the matter. The model could approach the size of Anthropic’s cutting-edge Mythos system, though its exact size will be determined later. A Reddit post framed the effort as ByteDance vowing to avoid AI distillation and develop the model its own way.
⚡ Try this today
Track ByteDance’s eventual release and documentation before benchmarking or adopting the model.
Sources & original reporting
This brief summarizes and links to reporting from the publishers below.
Enjoyed this brief? Get the next one in your inbox.
More in Models
Qwen3.8 27B scores 52 on Artificial Analysis
Reddit and Hacker News posts highlighted a reported jump in Qwen3.8 27B benchmark results.
Google introduces Gemini 3.7 Flash
Google and DeepMind published posts introducing the Gemini 3.7 Flash model.
Meta releases open-weight Glimmer AI model
The downloadable model contrasts with Meta’s more powerful API-only Muse Spark.