CARREGANDO O RADAR…
Do Reasoning Representations Help Humans Evaluate LLM Outputs? | Radar arXiv · portela.dev