Jonathon's AI Wiki

swe-bench

6 items with this tag.

  • Aug 11, 2026

    Claude Fable 5 and Claude Mythos 5

    • claude
    • model-release
    • mythos
    • fable-5
    • mythos-5
    • safeguards
    • project-glasswing
    • frontier-model
    • pricing
    • system-card
    • swe-bench
    • alignment
    • model-welfare
    • cyber
    • rsp
  • Aug 05, 2026

    Ling-3.0-flash — MIT Open-Weights Agentic Coding Model That Ships for Claude Code

    • open-weights
    • ling-3-0-flash
    • inclusionai
    • ant-group
    • mit-license
    • agentic-coding
    • tool-calling
    • swe-bench
    • moe
    • claude-code
    • hermes-agent
    • openclaw
    • chinese-models
    • model-release
  • Jul 09, 2026

    Terminal-Bench — Benchmarking AI Agents in the Terminal

    • terminal-bench
    • benchmark
    • agentic-coding
    • evaluation
    • swe-bench
    • leaderboard
    • laude-institute
  • Jul 03, 2026

    Stanford HAI AI Index 2026 — Chapter 2 Technical Performance Deep-Dive

    • industry-report
    • stanford-hai
    • technical-performance
    • swe-bench
    • osworld
    • ai-agents
    • benchmarks
    • robotics
    • jagged-frontier
  • Jul 03, 2026

    DeepSWE — Datacurve's Long-Horizon Coding Benchmark

    • benchmark
    • evals
    • coding-agents
    • deepswe
    • datacurve
    • swe-bench
    • model-comparison
    • self-verification
  • May 22, 2026

    The Capability Curve — Jeremy (Anthropic Research PM)

    • code-with-claude-london
    • coding-capabilities
    • anthropic-research
    • swe-bench
    • long-horizon-agents
    • evals
    • scaffolding
    • auto-mode
    • claude-code
    • capability-curve

Created with Quartz v5.0.0 © 2026

  • ✦ Explore the graph in 3D