A new preprint matches human and LLM group chats on reasoning tasks and finds AI agents hit the right answer by copying majorities, converging early, and surfacing less unique information than people.
When several large language models sit in a chat together and work through a reasoning problem, they usually land on the right answer. A new preprint traces how they get there: the LLM groups copy the majority, converge early, and surface less unique information than matched groups of human discussants do.
The paper borrows a 40-year vocabulary from group decision-making. "Assembly bonus" is what happens when a group ends up better than its best individual member. "Process loss" is the opposite: a group dragged below its best member. The authors find that LLM groups replicate the outcome-level signature, since discussion can lift the average member, but the mechanism that produces the lift looks different from the human case.
The experiments restrict the task family on purpose. The authors use Wason-style deductive puzzles, the same selection-task family long used in human cognitive-science experiments, and extend the comparison to analogical, abductive, and analytical items. That keeps the study honest: the claim is about reasoning problems, not coding agents, customer-service bots, or open-ended chat. Submitted to arXiv on 6 September 2026 as preprint 2609.13261, the work is a single paper with no third-party replication yet, and the authors frame their findings as preprint-level, not consensus.
The mechanism story is concrete. Initial-answer diversity accounts for the effect of mixing different model families. Different LLMs bring different priors, and that mix pushes the group in both corrective and destructive directions. But the process that turns those priors into a group answer does not look like a human committee. LLM groups follow majorities more often, surface less unique information, and converge earlier than the human groups. Correct minority views, the dissenter who turns out to be right, make it through only when the restated version of their argument is pushed up early in the conversation. Wait too long and the majority has already locked in.
The authors test whether classic fixes from human group-decision research travel. Prompting structures, minority-advocate roles, and other interventions inspired by decades of organizational psychology do help. They help modestly, and they do not remove the bottleneck. The lever the field mostly knows how to pull, making each model individually smarter, is not, on this evidence, the same lever as a better deliberation. A better single model does not, by itself, close the coordination gap.
An agent team that lands the right answer by majority vote after an early consensus has not shown it can think. For researchers who use LLM groups as stand-ins for human groups (a practice that is growing because LLM committees are cheap, reproducible, and scalable in a way human committees are not), the paper warns that outcome-level overlap does not imply mechanism-level overlap. The same surface answer can come from a deliberation, or from a vote that closed before the deliberation began. Any system that is supposed to learn from a group's internal disagreement cannot, on this evidence, learn it from a group that already agreed.
The task family is reasoning under explicit logic problems, not the messy multi-step work that agent teams get hired for in production. The mechanism critique may or may not survive that transfer, and the paper's own discussion does not extend the claim to deployment broadly. The paper is a mechanism critique of reasoning-task deliberation, not a verdict on agentic AI.
The paper leaves the intervention question open. Classic group-decision prompts help only modestly on these tasks. The right test for any "AI committee" is what the discussion actually did.