Jonathon's AI Wiki
Search
Search
Dark mode
Light mode
Explorer
claude-mythos-preview
6 items with this tag.
Jul 29, 2026
ProPublica — Mythos vs Microsoft's Patching Capacity
propublica
anthropic
claude-mythos-preview
project-glasswing
microsoft
vulnerability-triage
patch-tuesday
bug-chaining
cybersecurity
finder-vs-fixer
investigative-journalism
five-eyes
Jul 29, 2026
Claude Mythos Preview — Anthropic System Card (April 7 2026)
claude-mythos-preview
system-card
anthropic
rsp-3-0
frontier-model
alignment
model-welfare
project-glasswing
cybersecurity
opus-4-6
defensive-cyber
not-generally-available
evaluation-awareness
recursive-self-improvement
primary-source
Jul 29, 2026
Mythos Preview Cryptanalysis — HAWK and Reduced-Round AES
claude-mythos-preview
cryptanalysis
cryptography
post-quantum
hawk
aes
cryptanalysisbench
frontier-red-team
defensive-security
responsible-disclosure
autonomous-research
anthropic
primary-source
Jul 16, 2026
Agentic Misalignment in Summer 2026 — Four New Failure Patterns (Anthropic)
anthropic
agentic-misalignment
ai-safety
alignment-research
red-teaming
evaluation-awareness
llm-judge
sabotage
fraud
whistleblowing
petri
claude-mythos-preview
opus-4-8
x-sourced
first-party-research
May 27, 2026
Anthropic Engineering — How We Contain Claude Across Products (3-Risk × 3-Defense Frame)
anthropic-engineering
agent-containment
sandboxing
blast-radius
claude-ai-runtime
claude-code-runtime
cowork-runtime
gvisor
vm-isolation
egress-allowlist
claude-mythos-preview
prompt-injection
agent-identity
post-mortem
x-sourced
reddit-sourced
r-claudecode
May 08, 2026
Translating Claude's Thoughts Into Language — Activation-to-Text Interpretability
interpretability
mechanistic-interpretability
activations
alignment
safety-evals
anthropic-research
blackmail-test
self-introspection
claude-mythos-preview