Jonathon's AI Wiki

safety-evals

1 item with this tag.

  • May 08, 2026

    Translating Claude's Thoughts Into Language — Activation-to-Text Interpretability

    • interpretability
    • mechanistic-interpretability
    • activations
    • alignment
    • safety-evals
    • anthropic-research
    • blackmail-test
    • self-introspection
    • claude-mythos-preview

Created with Quartz v5.0.0 © 2026

  • ✦ Explore the graph in 3D