PAPER / ARXIV:2605.27887
Yuxuan Zhao, Sijia Chen, Ningxin Su
RESUMO
Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked. Existing benchmarks exhibit two gaps: they are often equity-only and ignore cross-asset correlations; they fail to evaluate the complete PM decision pipeline. We introduce PortBench, a benchmark spanning six heterogeneous asset classes from 2015 to 2025. PortBench comprises a static QA dataset of 6,269 questions across seven task templates and a dynamic five-stage allocation pipeline. To evaluate these layers, we introduce metrics: a dual-layer correlation score for inter-class hedging and intra-class concentration, CEPS, which quantifies how reasoning errors compound across pipeline stages. We further evaluate under three stress windows and risk profiles, to support real-time evaluation and mitigate pretraining contamination on historical markets. Across ten frontier LLMs, strong financial performance fails to translate into superior portfolio performance: only 32.5% of 120 evaluations beat equal weighting across Sharpe and four market periods. Our source code is available at https://github.com/AgenticFinLab/portbench.
NO MESMO MAPA