Jonathon's AI Wiki

safety

4 items with this tag.

  • Jul 16, 2026

    Hermes Agent + Stripe — Payments Skill Suite

    • hermes-agent
    • skills
    • payments
    • stripe
    • agent-commerce
    • pay-per-call
    • mpp-agent
    • safety
    • human-approval
    • nous-research
  • Jul 09, 2026

    GRAM — An Off Switch for Dual-Use Knowledge (Anthropic + AE Studio)

    • anthropic
    • ae-studio
    • alignment
    • safety
    • dual-use
    • gram
    • pretraining
    • biosecurity
  • Jul 02, 2026

    Refusal Calibration and Constitutional AI — Shaping Claude's Refusals Without Raising Jailbreak Risk

    • constitutional-ai
    • claudes-constitution
    • refusals
    • over-refusal
    • jailbreak
    • safety
    • system-prompt
    • anthropic
  • Jun 15, 2026

    Claude Opus 4.8 — Anthropic Release + System Card (May 28, 2026)

    • opus-4-8
    • claude
    • model-release
    • system-card
    • anthropic
    • benchmarks
    • alignment
    • safety
    • agentic
    • long-context
    • pricing
    • mythos-preview

Created with Quartz v5.0.0 © 2026

  • ✦ Explore the graph in 3D