PAPER / ARXIV:2609.08966
Sohir Maskey , Philipp Scholl , Jonas Knupp , Pit Neitemeier , Sascha Wirges
RESUMO
Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the best starting point for subsequent training. We show that this assumption can fail in a full 30B mixture-of-experts training pipeline. The checkpoints that perform better after the full downstream training stack also have higher solution density, i.e., retain downstream performance under local weight perturbations.
NO MESMO MAPA