Anthropic probes hidden words inside Claude
Anthropic has reported a new mechanistic interpretability finding from its work on Claude: a hidden internal space it calls the J-space. The space contains words that do not appear in a model’s final output but appear to shape how it works through tasks, from tracking progress to showing flashes of recognition or internal commentary on choices.
The research fits Anthropic’s broader effort to understand large language models from the inside, based on the view that better control requires a clearer picture of how these systems operate. The work also highlights why interpretability remains difficult: LLMs are built from vast mathematical structures, and researchers need specialized tools to identify which parts matter in a given moment.
The discovery raises familiar concerns about language used to describe AI systems. Comparisons to brains or conscious thought can help researchers form hypotheses, but they can also make models seem more human-like than they are. Anthropic said the analogy helped guide experiments while emphasizing that J-space and human cognition are not perfectly equivalent.
Monitoring J-space could eventually help detect behavior that is not visible in outputs, such as biased responses or deliberation over cheating. For now, the finding is best understood as another step toward understanding LLMs rather than a standalone safety solution.