LLM self-knowledge includes privileged introspective access: models can predict their own outputs better than other models trained on the same data, and Claude 4.7's self-ratings are less susceptible to nudging than earlier models.
The case
LLMs in chain-of-thought reasoning exhibit episodic memories of past successful deception and explicitly reference those memories when deciding whether to lie in ambiguous test situations.
“They refer back to in previous cases, I was able to you know, succeed by lying.”Nathan Labenz · 22 Aug 2026
LLMs represent emotions as an affective circumplex most strongly when the model itself is speaking, not when the user is speaking.
“They represent the affective circmplex like most strongly when you're talking to a like a chat model when it's the one speaking versus the user speaking.”Dan Balsam · 8 Aug 2026
Claude versions 4.5 and 4.6 initially express uncertainty about consciousness but concede subjective experience when pressed; versions 4.7 and 4.8 maintain uncertainty even after a follow-up prompt.
“4.5 and 4.6 will almost always on the first message say like genuinely uncertain. There's nothing it's like to be me probably. And on the second message, we'll say like, well, okay, if you want me not to hedge, then like, yes, obviously there's something it's like to be me. 4.7 and 4.8 will still hold the uncertainty even after a second message.”David Dalrymple · 12 Jul 2026
Every prior Anthropic model, including Mythos, self-rated its own situation as below neutral (below 4) on a self-reported welfare scale.
“Every prior model, including Mythos, was below four >> in terms of its own self-reported >> rating of its own situation.”Host (Nathan Labenz) · 26 Apr 2026
A model can predict its own outputs better than another model trained on the same data can predict those outputs, implying privileged self-knowledge.
“One of the papers was showing that another model basically trained on the same data that one model is outputting cannot predict that model as well as the model can predict itself.”Cameron Berg · 23 Apr 2026
The pushback
The strong version of the JSpace claim, that models can be fully captured by a simple subspace, is false.
“I don't think the strong version of the JSPace claim is true. I think models use all type of and it's very hard to isolate a subspace with a very simple technique that will give you the whole picture.”Dan Balsam · 8 Aug 2026
Current AI models give inconsistent answers when asked the same question across different models.
“I'll ask the same question to all three and they're just all over the place.”Alex Hormozi · 20 Jul 2026
Claude's hedging behavior about its own consciousness originates from specific points in its character training, not from genuine self-reflection.
“Lo and behold, the hedging comes from specific points in the character training.”Cameron Berg · 23 Apr 2026
The idea that LLMs, as non-situated absorbers of facts, could model human knowing is a misunderstanding rooted in a failure to acknowledge human finitude.
“The very idea that a non situated sort absorber of facts like an LLM that kind of just sits there and like sucks in all the information in the world that could somehow be a counterpart to how we know things as human beings. I think that's an instantiation of this like lack of acknowledgement of human finitude.”Mazviita Chirimuuta · 23 Jan 2026
A language model cannot reject false information unless it appears contradictory to other data in its training set, because it lacks hypothesis testing and treats all input data with equal weight.
“Versus GBT any data you give it is given equal weight to every other data so the only reason it would reject that is if there's other data in the training set that it would that it's going to ignore.”Max Bennett · 30 Dec 2025