PAPER / ARXIV:2609.20541
Lukas Meyer, Sofia Rossi, Wei Chen, Thomas Laurent, Yiming Li
RESUMO
Large language models can report a numerical confidence together with generated content, but it is unclear whether this is more than calibrated rhetoric. We analyze three training-free signals: verbalized confidence with the answer, post-hoc P(True), and agreement with three additional generations on the same 100 TriviaQA questions for two model families. Direct verbalization is surprisingly strong baseline: after auditing benchmark errors, it reaches AUROC 0.956 and 0.937 for correctness prediction. Three-sample agreement is substantially weaker (0.765 and 0.790), and fixed interpolation with verbalized confidence has no statistically reliable benefit. Four of nine errors from one model and eight of other receive unanimous sample support, showing that self-consistency can amplify shared misconceptions. Re-eliciting confidence answers with equivalent prompts changes scores by 0.043 to 0.084 on average and flips 4% to 9% of decisions at a 0.8 threshold. An exploratory audit of 100 confidence-tagged biography claims further finds only a modest confidence gap between supported and contradicted claims. These results argue that useful self-reports remain sensitive to elicitation, correlated with benchmark noise.
NO MESMO MAPA