PAPER / ARXIV:2609.11115
Koutian Wu, Junjie Zhou, Ergan Shang, Jiayu Wang, Pengqian Han, Junkai Wang, Wanghan Xu, Lin Shi
RESUMO
Benchmark researchers and LLM developers need to find relevant evaluations, locate datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and search engine covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations, drawing on 37 sources with 1,283 source records and 12,916 numeric observations. We examine benchmark saturation, adoption trends, and the limits of score comparisons.
NO MESMO MAPA