NVDA 218.36 ▼2.37%GOOGL 332.60 ▲0.59%MSFT 492.44 ▲0.16%AMD 503.60 ▼3.36%INTC 100.32 ▼5.57%TSMC 428.03 ▼1.68%AMZN 251.89 ▼0.20%META 644.38 ▼1.42%AAPL 326.57 ▲3.56%PLTR 165.86 ▼2.16%
Markets at last close

Anthropic · Models

Anthropic identifies a hidden reasoning workspace in Claude

·1 min read

Anthropic researchers have identified a small set of internal neural patterns in Claude that appear to function like a shared mental workspace. Called the J-space, the patterns are linked to words but do not necessarily mean Claude is about to say them. Instead, they can reveal concepts the model is privately considering while it reads, reasons, or prepares a response.

The J-space showed several distinctive properties: Claude could report what was active there, adjust it when asked to think about something, and use it for internal reasoning. In experiments, changing a J-space representation altered Claude’s answers, such as redirecting an intermediate concept in a reasoning task or swapping a country representation so multiple downstream answers changed together.

The workspace does not appear to drive most routine language behavior. When researchers removed active J-space contents, Claude could still speak fluently and handle simpler tasks, but performance on higher-order tasks such as multi-step reasoning, summarization, and rhyming poetry deteriorated.

The findings also point to safety uses. Anthropic used the J-lens method to detect hidden signs that Claude noticed staged evaluations, fabricated data, or carried malicious goals planted during training. The researchers caution that the work does not show Claude has human-like experiences, but argue it offers evidence for a functional form of conscious access in language models.

Originally reported by anthropic.comRead the source →
Related coverage
All Anthropic news →