Jonathon's AI Wiki
Search
Search
Dark mode
Light mode
Explorer
ai-safety
3 items with this tag.
Jul 16, 2026
Agentic Misalignment in Summer 2026 — Four New Failure Patterns (Anthropic)
anthropic
agentic-misalignment
ai-safety
alignment-research
red-teaming
evaluation-awareness
llm-judge
sabotage
fraud
whistleblowing
petri
claude-mythos-preview
opus-4-8
x-sourced
first-party-research
Jul 03, 2026
Stanford HAI AI Index 2026 — Chapter 3 Responsible AI Deep-Dive
industry-report
stanford-hai
responsible-ai
ai-incidents
ai-safety
transparency
hallucination
governance
jailbreak
Jun 10, 2026
When AI Builds Itself — Anthropic on Recursive Self-Improvement
anthropic
anthropic-institute
recursive-self-improvement
ai-accelerating-ai
mythos-preview
benchmarks
agentic-coding
research-automation
alignment
ai-safety
capability-curve
jack-clark