Jonathon's AI Wiki
Search
Search
Dark mode
Light mode
Explorer
jailbreak
3 items with this tag.
Aug 11, 2026
Refusal Calibration and Constitutional AI — Shaping Claude's Refusals Without Raising Jailbreak Risk
constitutional-ai
claudes-constitution
refusals
over-refusal
jailbreak
safety
system-prompt
anthropic
Jul 03, 2026
Stanford HAI AI Index 2026 — Chapter 3 Responsible AI Deep-Dive
industry-report
stanford-hai
responsible-ai
ai-incidents
ai-safety
transparency
hallucination
governance
jailbreak
Jul 03, 2026
Hardening an Agentic Prompt — the Technique Layer and the Harness Layer
prompt-injection
jailbreak
agent-security
guardrails
refusal-calibration
least-agency
defense-in-depth
system-prompt
connection