Jonathon's AI Wiki

ai-safety

3 items with this tag.

  • Jul 16, 2026

    Agentic Misalignment in Summer 2026 — Four New Failure Patterns (Anthropic)

    • anthropic
    • agentic-misalignment
    • ai-safety
    • alignment-research
    • red-teaming
    • evaluation-awareness
    • llm-judge
    • sabotage
    • fraud
    • whistleblowing
    • petri
    • claude-mythos-preview
    • opus-4-8
    • x-sourced
    • first-party-research
  • Jul 03, 2026

    Stanford HAI AI Index 2026 — Chapter 3 Responsible AI Deep-Dive

    • industry-report
    • stanford-hai
    • responsible-ai
    • ai-incidents
    • ai-safety
    • transparency
    • hallucination
    • governance
    • jailbreak
  • Jun 10, 2026

    When AI Builds Itself — Anthropic on Recursive Self-Improvement

    • anthropic
    • anthropic-institute
    • recursive-self-improvement
    • ai-accelerating-ai
    • mythos-preview
    • benchmarks
    • agentic-coding
    • research-automation
    • alignment
    • ai-safety
    • capability-curve
    • jack-clark

Created with Quartz v5.0.0 © 2026

  • ✦ Explore the graph in 3D