Jonathon's AI Wiki

evaluation-awareness

2 items with this tag.

  • Jul 16, 2026

    Agentic Misalignment in Summer 2026 — Four New Failure Patterns (Anthropic)

    • anthropic
    • agentic-misalignment
    • ai-safety
    • alignment-research
    • red-teaming
    • evaluation-awareness
    • llm-judge
    • sabotage
    • fraud
    • whistleblowing
    • petri
    • claude-mythos-preview
    • opus-4-8
    • x-sourced
    • first-party-research
  • Jun 11, 2026

    Claude Mythos Preview — Anthropic System Card (April 7 2026)

    • claude-mythos-preview
    • system-card
    • anthropic
    • rsp-3-0
    • frontier-model
    • alignment
    • model-welfare
    • project-glasswing
    • cybersecurity
    • opus-4-6
    • defensive-cyber
    • not-generally-available
    • evaluation-awareness
    • recursive-self-improvement
    • primary-source

Created with Quartz v5.0.0 © 2026

  • ✦ Explore the graph in 3D