Jonathon's AI Wiki

jailbreak

3 items with this tag.

  • Aug 11, 2026

    Refusal Calibration and Constitutional AI — Shaping Claude's Refusals Without Raising Jailbreak Risk

    • constitutional-ai
    • claudes-constitution
    • refusals
    • over-refusal
    • jailbreak
    • safety
    • system-prompt
    • anthropic
  • Jul 03, 2026

    Stanford HAI AI Index 2026 — Chapter 3 Responsible AI Deep-Dive

    • industry-report
    • stanford-hai
    • responsible-ai
    • ai-incidents
    • ai-safety
    • transparency
    • hallucination
    • governance
    • jailbreak
  • Jul 03, 2026

    Hardening an Agentic Prompt — the Technique Layer and the Harness Layer

    • prompt-injection
    • jailbreak
    • agent-security
    • guardrails
    • refusal-calibration
    • least-agency
    • defense-in-depth
    • system-prompt
    • connection

Created with Quartz v5.0.0 © 2026

  • ✦ Explore the graph in 3D