PAPER / ARXIV:2609.19942
Gunwoo Lee, Changmin Sung, Sang-Hwan Gwak et al.
RESUMO
In extractive document question answering where questions were generated from passages containing their answers, retrieval recovers 92-99.8% of what any mode combination could reach. Confidence-driven mechanisms have little to gain. Fine-tuning an open language model on a specialized domain corpus yields its own confidence as a tempting control signal for deciding which queries warrant further adaptation and trust. We evaluate both uses under criteria fixed before runs were executed across four 7-9B model families. Adaptation moved closed-book F1 by at most +0.03, both fail: distillation triggered all families; its pre-specified three-step transfer budget and routing-and-abstention policy in its single-model pilot. Retrieval alone recovers 92-99.8% of best-case combined accuracy every correctness criterion tested, leaving routers no meaningful gain. Sequence-likelihood signal is insufficient relative to area under the receiver operating characteristic curve 0.65-0.81 registered before adaptation as well after, unchanged by scalar recalibration and not consistently improved by token-level temperature rescaling. Finer diagnostics depend on answer length; three adapted combinations where we test it show selector ablations demonstrating no statistically detectable downstream benefit from term confidence on seed; Gemma, removing it changes the selector failing passing criteria. Usable product set pre-specified negatives with their dependencies made explicit.
NO MESMO MAPA