PAPER / ARXIV:2609.29933 · NOVO
Kenan E. Ak, Jay Mohta, Gwang Gook Lee, Yan Xu, Dimitrios Dimitriadis
RESUMO
[Abstract paraphrased by fetch tool] Vision-Language Models are increasingly deployed for long-document processing with mixed media including text, charts, tables, and figures. Practitioners must decide how to structure document input, which retriever strategy to employ for partial page selection, and whether to use agentic or static approaches. This work evaluates these design choices across two benchmarks using commercial and open-weight models. Agentic six-tool systems show advantages only with sufficiently large models; retrieval modality is more consequential than specific retriever choice; image-based retrieval uses roughly one-seventh to one-quarter the tokens of full-page input. Three of the strongest pipelines succeed on different questions, and oracle selection could yield about 13-point improvements over any single pipeline.
NO MESMO MAPA