Jonathon's AI Wiki
Search
Search
Dark mode
Light mode
Explorer
alignment
11 items with this tag.
Sep 29, 2026
When AI Builds Itself — Anthropic on Recursive Self-Improvement
anthropic
anthropic-institute
recursive-self-improvement
ai-accelerating-ai
mythos-preview
benchmarks
agentic-coding
research-automation
alignment
ai-safety
capability-curve
jack-clark
Sep 29, 2026
Claude Fable 5 and Claude Mythos 5
claude
model-release
mythos
fable-5
mythos-5
safeguards
project-glasswing
frontier-model
pricing
system-card
swe-bench
alignment
model-welfare
cyber
rsp
Sep 29, 2026
Claude Mythos Preview — Anthropic System Card (April 7 2026)
claude-mythos-preview
system-card
anthropic
rsp-3-0
frontier-model
alignment
model-welfare
project-glasswing
cybersecurity
opus-4-6
defensive-cyber
not-generally-available
evaluation-awareness
recursive-self-improvement
primary-source
Sep 29, 2026
Claude Opus 5.5 — Fable-5.1-Level Work at 40% Less Than Opus 5
claude-ai
anthropic
model-release
opus-5-5
claude-opus-5-5
claude-5-5-family
pricing
benchmarks
alignment
safeguards
preserved-thinking
fast-mode
effort
model-routing
writing-quality
Sep 29, 2026
Claude Opus 5 — Near-Fable-5 Intelligence at Opus 4.8 Pricing
claude-ai
anthropic
model-release
opus-5
claude-opus-5
pricing
benchmarks
alignment
cyber-classifiers
fast-mode
model-routing
prompt-caching
Sep 29, 2026
Claude's Values Across Models and Languages — Anthropic Research
anthropic-research
claude-values
interpretability
model-character
constitution
clio
model-selection
multilingual
alignment
Sep 29, 2026
Translating Claude's Thoughts Into Language — Activation-to-Text Interpretability
interpretability
mechanistic-interpretability
activations
alignment
safety-evals
anthropic-research
blackmail-test
self-introspection
claude-mythos-preview
Aug 11, 2026
Eight Predictions for the Era of Continual Learning (Dwarkesh Patel)
continual-learning
agent-memory
vendor-lock-in
alignment
regulation
inference-economics
batching
moats
dwarkesh-patel
predictions
Jul 09, 2026
GRAM — An Off Switch for Dual-Use Knowledge (Anthropic + AE Studio)
anthropic
ae-studio
alignment
safety
dual-use
gram
pretraining
biosecurity
Jul 06, 2026
A Global Workspace in Language Models — J-Space Interpretability Research
interpretability
mechanistic-interpretability
global-workspace
j-space
jacobian-lens
anthropic-research
alignment
hidden-goals
situational-awareness
neuronpedia
Jun 15, 2026
Claude Opus 4.8 — Anthropic Release + System Card (May 28, 2026)
opus-4-8
claude
model-release
system-card
anthropic
benchmarks
alignment
safety
agentic
long-context
pricing
mythos-preview