Jonathon's AI Wiki
Search
Search
Dark mode
Light mode
Explorer
benchmark
5 items with this tag.
Jul 24, 2026
MirrorCode (Epoch AI + METR long-horizon coding benchmark)
epoch-ai
metr
benchmark
long-horizon-coding
autonomous-coding
swe-benchmark
claude-opus
Jul 24, 2026
Cookbook — Giving Claude a Zoom Tool for Fine Image Detail
cookbooks
multimodal
vision
tool-use
image-tokens
chart-reading
chartography
pixel-coordinates
claude-fable-5
claude-sonnet-5
benchmark
jupyter-notebooks
793
Jul 09, 2026
Terminal-Bench — Benchmarking AI Agents in the Terminal
terminal-bench
benchmark
agentic-coding
evaluation
swe-bench
leaderboard
laude-institute
Jul 03, 2026
DeepSWE — Datacurve's Long-Horizon Coding Benchmark
benchmark
evals
coding-agents
deepswe
datacurve
swe-bench
model-comparison
self-verification
Jun 12, 2026
Agent Wikis — Compiled LLM Wikis with Agent-Native Access
karpathy-pattern
llm-wiki
llms-txt
agent-readable
rag
benchmark
mcp