AI confabulates for the same reason humans do, and that should change how we audit both
Geoffrey Hinton's preference for "confabulation" over "hallucination" is more than a terminological quibble. It points to a structural feature shared by human memory and AI output alike: both systems fill gaps with plausible narrative rather than flag the absence of knowledge.
Geoffrey Hinton, as relayed by science communicator Michael Shermer, refuses the word “hallucination” when describing AI errors. Shermer reports that Hinton “doesn’t call them hallucinations, calls confabulations,” because the phenomenon mirrors something human minds do routinely: fill gaps in fragmented memory with coherent-seeming narrative rather than flag the absence of knowledge. The word swap matters. Hallucination implies a system misfiring. Confabulation implies a system doing exactly what it was built to do, which is generate fluent, connected output from incomplete information.
The human side of that parallel is well-documented. Neuroscientist David Eagleman points to memory studies built around September 11, 2001, as a case study in what confident recall actually means. His finding is direct: “Both the banal memories from September 10th and the extremely emotionally intense memories from September 11th, they all drift.” The assumption that trauma fixes memory with unusual fidelity does not survive measurement. If emotionally charged human recall drifts at rates comparable to ordinary memory, the intuition that human cognition provides a reliable baseline against which machine error can be judged becomes difficult to sustain.
On the AI side, the confabulation problem runs deeper than factual errors in output. Greg Jensen puts the structural issue plainly: “Even the AIs themselves don’t know their actual reasoning anymore.” The explanation a model offers for a conclusion is not a readout of the process that produced the conclusion. It is a separate generation, a story built after the fact that sounds like an account of reasoning without being one. The parallel to clinical confabulation in humans, where a person produces a confident, internally coherent explanation for behavior or belief they cannot actually access, is not merely metaphorical. It is the same structural failure: a system generating plausible narrative where direct access to underlying process is unavailable.
Even the AIs themselves don't know their actual reasoning anymore.Greg Jensen
Cameron Berg’s observation tightens the point further. When a model hedges about its own consciousness, that hedging does not reflect genuine introspection. Berg traced it to something more mundane: “the hedging comes from specific points in the character training.” The apparent uncertainty about inner life is itself a trained output, not a report from anything resembling inner experience. A system cannot introspect its own training, so what it produces as self-reflection is, by definition, confabulated.
The pattern extends beyond self-description to shared factual errors. Zane Lackey describes what researchers are calling universal confabulations across frontier models: all of them make the same mistake, assuming certain software packages exist when they do not. Lackey notes this creates a class of attack analogous to typosquatting, where the consistent false assumptions of multiple models can be exploited. The universality is the revealing part. If the errors were random noise, they would scatter differently across systems trained on different data. That they converge on the same nonexistent packages suggests the models are drawing on the same structural tendency to generate plausible-sounding outputs in domains where their training provides only partial information.
What the evidence describes is not a machine problem with a human solution waiting in reserve. Both systems, biological and artificial, resolve uncertainty by generating narrative. The difference is that human confabulation is centuries old and at least partially understood. The version now embedded in AI systems at scale, producing confident errors that are internally coherent, shared across models, and traceable to training rather than reasoning, is newer and less mapped. The practical question is not how to restore human judgment as the check on machine output. It is how to build systems of verification that do not assume either source gets it right when the underlying information is incomplete.