PAPER / ARXIV:2609.19947
Choi, Park, Niu, Xiong, Yoo, Yang
RESUMO
LLM-based AI agents process user requests through iterative reasoning and tool execution, often involving remote LLM API calls with local tool containers. This makes optimization difficult due to latency, resource demands, and container bottlenecks across requests. This paper analyzes resource dynamics for three tasks: retrieval-augmented question answering, web search, and software coding. We characterize latency with respect to resource dynamics when processing multiple requests concurrently. Measurements show agents have wide-ranging behaviors depending on tasks, so even identical tools differ substantially in resource dynamics. Running multiple requests concurrently exposes task-dependent bottlenecks such as CPU, disk I/O, and memory. We find faster LLM responses or more CPU cores do not always accelerate agents. Based on these observations, we demonstrate opportunities for CPU-aware tool admission and task-aware CPU allocation. CPU-sensitive agent optimization improves throughput ~5.4×, and average latency is reduced ~32% compared to native agents.
NO MESMO MAPA