CARREGANDO O RADAR…
Language-model groups overstate consensus when replaying human deliberation on a reasoning task | Radar arXiv · portela.dev