PAPER / ARXIV:2609.20543
Shao
RESUMO
Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code. Across scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas almost all agents always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and participation-matched comparison (n = 45) yielded 34.1 and 44.4 percentage points. These complementary routes reduced different measurement asymmetries and converged within 0.5 percentage points. The gap persisted without early stopping under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track accuracy and were biased estimators of group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of deliberative outcomes.
NO MESMO MAPA