PAPER / ARXIV:2609.06974
Seungmin Oh, Donggeon Lee, Jongbin Ryu
RESUMO
Structured pruning reduces LLM deployment costs but its recovery stage is limited by a capacity-knowledge asymmetry between the recovery module and the complexity of removed knowledge. OverRep temporarily overparameterizes the recovery module during training to absorb distilled knowledge, then algebraically merges it into a compact module, preserving the pruned model's inference cost. Across three backbone families, OverRep improves retained reasoning performance over strong baselines by up to 5.5 and 8.4 points at 25% and 50% pruning.
NO MESMO MAPA