NCP-ArchPreview introduces Next Concept Prediction (NCP), enabling language models to predict multi-token concepts alongside individual tokens. The 8.9B-parameter model was trained on 5.73 trillion tokens and reaches OLMo-3-7B’s final pretraining loss using only 51.3% of the training tokens. It achieves 1.95× faster convergence and improves downstream performance by 2.45 points, including a +5.99 gain on […]

