PAPER / ARXIV:2609.20110
Jin, Wang, Han et al.
RESUMO
Straight-through processing (STP) on extracted key-value fields from financial documents without human review requires a calibrated probability together with a bounded guarantee on the residual error of auto-approved tier. The emergence of modern Vision Language Models (VLMs) provides an out-of-the-box capability for extracting key-values, but their verbalized confidence signals are unreliable and weakly track field correctness. This paper introduces a decomposed confidence layer along three interpretable channels, including perception, layout, and validation. Together with final conformal risk control, the score can be used for reliable STP on financial documents. The method is validated on public datasets covering real invoices, synthetic invoices, and ad-buy forms, using two different VLM families (Qwen3.6-27B and Gemini-3.1-Flash-Lite). Our decomposed score consistently improves the separation of correct from incorrect extractions, substantially raising the AUROC from 0.54-0.74 to 0.90-0.99. Contributions include all three designed channels. Crucially for industrial deployment, this enables usable STP. The native VLM confidence could clear only 0.1%-7.0% under risk control at a target <10%. In contrast, the proposed method auto-approves 49-72% of fields while holding the empirical error accepted tier at or below the target.
NO MESMO MAPA