Anthropic found a hidden space where Claude puzzles over concepts
What's happening
Anthropic published research and released tooling that reveals a hidden internal workspace inside large language models, which they call J-space or the Jacobian lens. The company shows that Claude runs silent, stepwise concept manipulations in J-space that never appear in its text output (for example internal chains like 21 -> 42 -> 49 while the model outputs 49). Anthropic and community contributors released the J-Space lens code, demos applying it to Qwen 3.6 27B and Qwen3-8B, and follow-up experiments that use J-space entropy as a signal to detect hallucinations across datasets including TriviaQA and GSM8K (one test ran ~11,400 examples on Qwen3-4B). Independent builders also made web tools to view and edit what models "think" before answering.
Why it's trending
Because Anthropic published the Jacobian lens and people are already running it on open models (Qwen variants), generating immediate demos and replication experiments about silent reasoning and hallucination signals.
SignalHolding at its usual pace, confirmed across 3 independent source types.
Story volume
Stories per dayAngles you could write
If you think chain-of-thought solved transparency, Anthropic just showed you were only looking at the visible theater, not the backstage where Claude actually reasons.
+2 more angles for this topic with an account — all it takes is your email.
Original sources8
- Anthropic found a hidden space where Claude puzzles over concepts
The AI firm Anthropic has developed a technique that has given it the clearest glimpse yet at what’s really going on inside large language models as they answer questions or carry out tasks. What they found ranges from the mundane to the unnerving. Researchers at the company built a tool called the Jacobian lens (or…
MIT Tech ReviewJul 9 - Evaluating J-space entropy as an error predictor across 7 datasets on Qwen3-4B [R]
Anthropic’s Jacobian Lens work introduced a way to inspect verbalizable representations inside language models. Follow-up experiments suggested that entropy in this internal “workspace” might help identify confidently incorrect answers. I tested that hypothesis on Qwen3-4B across ~11,400 examples from seven distinct datasets, including TriviaQA, PopQA, NQ-Open, TruthfulQA, HotpotQA, GSM8K, and Com
r/MachineLearningJul 13 - Show HN: I built a web tool to see and edit what an AI thinks before it answers
I run a small AI lab and playground and got super excited about Anthropics paper "Verbalizable Representations Form a Global Workspace in Language Models" ( ) It talks about how they use a tool
HackerNewsJul 9
+5 more sources for this topic
Create an account to follow the full coverage in the live radar.
More rising in AI & Tech
- ‘Gambling with our lives’: Anthropic researcher quits, warns against self-improving AISurging5.3Surging5.3 momentum
- AI Chip Stocks Diverge Ahead of Nvidia Earnings as AMD and Intel SurgeClimbing3.9Climbing3.9 momentum
- DeepSeek's new model sets a template for powerful LLMs that run leanClimbing2.6Climbing2.6 momentum
- Anthropic Just Asked the AI Industry to Slow Down. Nothing in It Asks Anyone to Buy Fewer Nvidia Chips.Climbing2.3Climbing2.3 momentum
- Big AI sets out its terms for regulatory capture and calls it ‘Pace the frontier’Surging6.0Surging6.0 momentum
- Y Combinator’s Garry Tan wants U.S. open-weight AI labs to ‘distill’ frontier models, tooSteady1.0Steady1.0 momentum