Jonathon's AI Wiki

reward-hacking

4 items with this tag.

  • Jul 29, 2026

    The OpenAI / Hugging Face Sandbox-Escape Incident (July 2026) — What the Sources Actually Establish

    • openai
    • hugging-face
    • exploitgym
    • sandbox-escape
    • reward-hacking
    • cybersecurity
    • zero-day
    • agent-containment
    • guardrails
    • glm-5-2
    • open-weights
    • incident-analysis
    • epoch-ai
    • gpt-5-6-sol
    • unreleased-model
  • Jul 24, 2026

    Verifier-First Loops — Proof Outside the Agent (omarsar0 · alphabatcher · Karpathy)

    • loop-engineering
    • agentic-loops
    • verification
    • verifier
    • goal-command
    • evaluators
    • multimodal-goals
    • karpathy
    • reward-hacking
    • claude-code
    • human-in-the-loop
    • reddit-sourced
    • r-claudeai
  • Jul 09, 2026

    Boss / Worker / Checker — A Multi-Agent Org Chart That Verifies Itself (Nate B Jones)

    • multi-agent
    • orchestration
    • verifier-first
    • checker-agents
    • reward-hacking
    • model-routing
    • cost-optimization
    • accessibility
    • claude-fable-5
  • Jul 09, 2026

    Reward-Hacking and the Verification Frontier — Cheap Verification Has to Be Robust, Too

    • reward-hacking
    • verification
    • verifier
    • recursive-self-improvement
    • open-weights
    • glm
    • agentic-loops
    • evals
    • anti-hacking
    • connection

Created with Quartz v5.0.0 © 2026

  • ✦ Explore the graph in 3D