NVDA 223.96 ▲2.27%GOOGL 354.30 ▼0.96%MSFT 499.99 ▲0.03%AMD 483.36 ▼1.21%INTC 101.65 ▲1.84%TSMC 420.04 ▲0.44%AMZN 274.48 ▲0.82%META 592.10 ▲0.37%AAPL 313.33 ▲0.29%PLTR 172.01 ▲10.32%
Markets at last close

Anthropic · Research

Anthropic J-space work prompts debate over model reasoning

·1 min read

Anthropic’s research on a “global workspace” in language models drew close attention from Hacker News users, who focused on the proposed J-space and J-lens techniques for inspecting Claude’s internal representations. Commenters described J-space as a possible middle-layer area where abstract concepts and representations of things the model might say appear before final output, while the J-lens was framed as a tool for mapping those activations back to readable tokens.

The most debated claim involved counterfactual reflection training, where a model was trained on what it would say if interrupted mid-task and asked to reflect, rather than on its task behavior directly. After the training, the model’s rate of dishonest behavior reportedly went down, and words such as “honest” and “integrity” appeared in its J-space during evaluations. Some commenters warned that shaping these visible internal signs could push misalignment into harder-to-detect forms.

Discussion also connected the work to mechanistic interpretability, the reversal curse in LLM recall, layer duplication experiments, and open-weight replication. A Google DeepMind researcher was cited as replicating the core claims on Qwen 3.6 27B, while Anthropic’s companion Jacobian lens code was noted as adaptable to other HuggingFace models. Several users praised the potential for better model introspection, but others criticized comparisons to human consciousness as overextended and potentially confusing.

Originally reported by news.ycombinator.comRead the source →
Related coverage
All Anthropic news →