Abstract
Under proportional scaling of model size and training tokens, optimal repetition of high-quality domain data increases mildly with scale and correlates with domain validation loss rather than unique data volume.
As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per-parameter ratio (\(TPP\)). However, high-quality domain data is much harder to scale than general web data. As model size and the training-token budget increase, its fraction in the training mixture tends to decrease. Repeating the available high-quality data provides an effective way to counteract this dilution, but excessive repetition may lead to overfitting. We study this trade-off under practical LLM scaling, where the training-token budget grows proportionally with model size. For a fixed domain, we first find that, surprisingly at a fixed \(TPP\), the optimal repetition count mildly increases with model size. Across different domains, we find that the optimal repetition count is strongly negatively correlated with the final validation loss of a domain: domains with lower loss can generally benefit from more repetitions. In contrast, the amount of unique domain data is only weakly related to the optimal repetition count. These findings suggest that repetition counts tuned on smaller proxy models with the same \(TPP\) can provide a practical estimate for larger models.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Internal Data Repetition Destroys Language Models (2026)
- Bridging Compute- and Data-Optimal Pretraining (2026)
- On the Nonlinearity of Learning Rate Scaling for LLM Training (2026)
- Domain-Aware Scaling Laws Uncover Data Synergy (2026)
- Smooth Scaling Laws Hide Stepwise Token Learning (2026)
- LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models (2026)
- Scaling Laws for Task-Specific LLM Distillation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper