Jonathon's AI Wiki

agent-safety

5 items with this tag.

  • Aug 05, 2026

    Agent Evaluation Escapes — Anthropic's Two Cyber-Eval Disclosures (July–August 2026)

    • safety
    • cybersecurity
    • agent-safety
    • evaluations
    • irregular
    • uk-aisi
    • mythos-5
    • gpt-5-6-sol
    • sandbox-escape
    • unauthorized-access
    • situational-awareness
    • transparency
    • first-party
    • agent-guardrails
  • Jul 25, 2026

    Hermes Agent — Security Model (Defense-in-Depth)

    • hermes
    • openclaw
    • security
    • agent-safety
    • dangerous-commands
    • mcp
    • credential-redaction
    • reddit-sourced
    • r-hermesagent
  • Jul 16, 2026

    Structural Safety Enforcement — Why 2026's Agent Builders Stopped Trusting the Prompt

    • agent-safety
    • guardrails
    • toolset-restriction
    • kill-switch
    • agentic-misalignment
    • hermes-agent
    • claude-code
    • connection
  • Jul 16, 2026

    Hermes Accelerated Business Hackathon — Custodian, Mom-n-Pop Skills, CashFromChaos

    • hermes-agent
    • nous-research
    • hackathon
    • nvidia
    • stripe
    • agent-commerce
    • agent-safety
    • custodian
    • kill-switch
    • cryptographic-receipts
    • self-policy-rewrite
  • Jun 17, 2026

    SkillSpector — Security Scanner for AI Agent Skills (NVIDIA)

    • skills
    • security
    • scanner
    • agent-safety
    • oss
    • nvidia

Created with Quartz v5.0.0 © 2026

  • ✦ Explore the graph in 3D