Jonathon's AI Wiki
Search
Search
Dark mode
Light mode
Explorer
swe-bench
6 items with this tag.
Aug 11, 2026
Claude Fable 5 and Claude Mythos 5
claude
model-release
mythos
fable-5
mythos-5
safeguards
project-glasswing
frontier-model
pricing
system-card
swe-bench
alignment
model-welfare
cyber
rsp
Aug 05, 2026
Ling-3.0-flash — MIT Open-Weights Agentic Coding Model That Ships for Claude Code
open-weights
ling-3-0-flash
inclusionai
ant-group
mit-license
agentic-coding
tool-calling
swe-bench
moe
claude-code
hermes-agent
openclaw
chinese-models
model-release
Jul 09, 2026
Terminal-Bench — Benchmarking AI Agents in the Terminal
terminal-bench
benchmark
agentic-coding
evaluation
swe-bench
leaderboard
laude-institute
Jul 03, 2026
Stanford HAI AI Index 2026 — Chapter 2 Technical Performance Deep-Dive
industry-report
stanford-hai
technical-performance
swe-bench
osworld
ai-agents
benchmarks
robotics
jagged-frontier
Jul 03, 2026
DeepSWE — Datacurve's Long-Horizon Coding Benchmark
benchmark
evals
coding-agents
deepswe
datacurve
swe-bench
model-comparison
self-verification
May 22, 2026
The Capability Curve — Jeremy (Anthropic Research PM)
code-with-claude-london
coding-capabilities
anthropic-research
swe-bench
long-horizon-agents
evals
scaffolding
auto-mode
claude-code
capability-curve