PAPER / ARXIV:2609.19680
Ma, Y.; Tai, Z.; Wu, H.
RESUMO
Financial QA systems are typically improved before deployment through better retrieval, prompting, or agent coordination, leaving their reliability behavior fixed thereafter. In practice, new SEC-filing questions repeatedly expose heterogeneous errors in period, entity, evidence use, and calculation. Existing self-improvement methods can turn failures into new behaviors but offer limited control over where a correction should apply or which previously correct answers it may break. We therefore frame post-deployment improvement as controlled behavioral maintenance: recurring failures become scoped skill patches, and each patch should earn deployment without introducing regressions. We instantiate this view in FINSKILLOPS, a multi-agent system for SEC filing QA. FINSKILLOPS derives reusable skills from evidence-grounded, typed failure diagnoses and governs them through targeted validation, protected-case regression checks, negative controls, versioned replacement and retirement. Across six financial QA benchmarks, a single frozen skill registry achieves the highest verdict-weighted correctness and reference consistency among evaluated systems. Evolved skills raise correctness 3.70 to 4.55 on our enhanced benchmark. In a separate 12-round operational study, only six of 33 proposed skills were promoted, while the monitoring non-correct rate falls from 20.0% to 12.5%. These results establish controlled scope, admission, and lifecycle management as foundation for reliable self-improvement.
NO MESMO MAPA