Jonathon's AI Wiki
Search
Search
Dark mode
Light mode
Explorer
safety
4 items with this tag.
Jul 16, 2026
Hermes Agent + Stripe — Payments Skill Suite
hermes-agent
skills
payments
stripe
agent-commerce
pay-per-call
mpp-agent
safety
human-approval
nous-research
Jul 09, 2026
GRAM — An Off Switch for Dual-Use Knowledge (Anthropic + AE Studio)
anthropic
ae-studio
alignment
safety
dual-use
gram
pretraining
biosecurity
Jul 02, 2026
Refusal Calibration and Constitutional AI — Shaping Claude's Refusals Without Raising Jailbreak Risk
constitutional-ai
claudes-constitution
refusals
over-refusal
jailbreak
safety
system-prompt
anthropic
Jun 15, 2026
Claude Opus 4.8 — Anthropic Release + System Card (May 28, 2026)
opus-4-8
claude
model-release
system-card
anthropic
benchmarks
alignment
safety
agentic
long-context
pricing
mythos-preview