DEV Community

#interpretability

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
J-space in practice: using Anthropic's Jacobian lens to decide what an LLM can forget

J-space in practice: using Anthropic's Jacobian lens to decide what an LLM can forget

Comments
6 min read
Beyond Reconstruction: Verifying Model Explanations with RECAP

Beyond Reconstruction: Verifying Model Explanations with RECAP

Comments
3 min read
Length Penalties in LLMs: Shorter Chains of Thought, Hidden Influences

Length Penalties in LLMs: Shorter Chains of Thought, Hidden Influences

Comments
3 min read
The safety switch that doesn't actually work

The safety switch that doesn't actually work

Comments
4 min read
Claude Was Always Thinking Ahead. Now We Can Read It.

Claude Was Always Thinking Ahead. Now We Can Read It.

2
Comments
7 min read
Mechanistic Interpretability is a 2026 Breakthrough Technology. Here's What That Means for the "LLMs Are Just Matrix Multiplication" Debate

Mechanistic Interpretability is a 2026 Breakthrough Technology. Here's What That Means for the "LLMs Are Just Matrix Multiplication" Debate

3
Comments
10 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.