Jonathon's AI Wiki

guardrails

6 items with this tag.

  • Aug 11, 2026

    The OpenAI / Hugging Face Sandbox-Escape Incident (July 2026) — What the Sources Actually Establish

    • openai
    • hugging-face
    • exploitgym
    • sandbox-escape
    • reward-hacking
    • cybersecurity
    • zero-day
    • agent-containment
    • guardrails
    • glm-5-2
    • open-weights
    • incident-analysis
    • epoch-ai
    • gpt-5-6-sol
    • unreleased-model
  • Aug 11, 2026

    Agent Guardrails: Hooks, Permissions, and Sandboxing Patterns

    • guardrails
    • hooks
    • permissions
    • sandboxing
    • agent-security
    • blast-radius
    • least-privilege
    • defense-in-depth
    • claude-code
    • agentic-systems
    • synthesis
  • Jul 29, 2026

    Kimi K3 (Moonshot AI)

    • moonshot-ai
    • kimi
    • kimi-k3
    • open-weight
    • benchmarks
    • competitor-models
    • agentic-coding
    • china
    • cost-per-task
    • intelligence-density
    • artificial-analysis
    • guardrails
    • open-weight-policy
  • Jul 29, 2026

    Auto Mode — Claude Code's Classifier-Reviewed Permissions Mode

    • claude-code
    • auto-mode
    • permissions
    • guardrails
    • first-party
    • anthropic
    • core-primitive
  • Jul 16, 2026

    Structural Safety Enforcement — Why 2026's Agent Builders Stopped Trusting the Prompt

    • agent-safety
    • guardrails
    • toolset-restriction
    • kill-switch
    • agentic-misalignment
    • hermes-agent
    • claude-code
    • connection
  • Jul 03, 2026

    Hardening an Agentic Prompt — the Technique Layer and the Harness Layer

    • prompt-injection
    • jailbreak
    • agent-security
    • guardrails
    • refusal-calibration
    • least-agency
    • defense-in-depth
    • system-prompt
    • connection

Created with Quartz v5.0.0 © 2026

  • ✦ Explore the graph in 3D