Anthropic identifies a hidden reasoning workspace in Claude
Anthropic researchers have identified a small set of internal neural patterns in Claude that appear to function like a shared mental workspace. Called the J-space, the patterns are linked to words but do not necessarily mean Claude is about to say them. Instead, they can reveal concepts the model is privately considering while it reads, reasons, or prepares a response.
The J-space showed several distinctive properties: Claude could report what was active there, adjust it when asked to think about something, and use it for internal reasoning. In experiments, changing a J-space representation altered Claude’s answers, such as redirecting an intermediate concept in a reasoning task or swapping a country representation so multiple downstream answers changed together.
The workspace does not appear to drive most routine language behavior. When researchers removed active J-space contents, Claude could still speak fluently and handle simpler tasks, but performance on higher-order tasks such as multi-step reasoning, summarization, and rhyming poetry deteriorated.
The findings also point to safety uses. Anthropic used the J-lens method to detect hidden signs that Claude noticed staged evaluations, fabricated data, or carried malicious goals planted during training. The researchers caution that the work does not show Claude has human-like experiences, but argue it offers evidence for a functional form of conscious access in language models.