Jonathon's AI Wiki

benchmark

5 items with this tag.

  • Jul 24, 2026

    MirrorCode (Epoch AI + METR long-horizon coding benchmark)

    • epoch-ai
    • metr
    • benchmark
    • long-horizon-coding
    • autonomous-coding
    • swe-benchmark
    • claude-opus
  • Jul 24, 2026

    Cookbook — Giving Claude a Zoom Tool for Fine Image Detail

    • cookbooks
    • multimodal
    • vision
    • tool-use
    • image-tokens
    • chart-reading
    • chartography
    • pixel-coordinates
    • claude-fable-5
    • claude-sonnet-5
    • benchmark
    • jupyter-notebooks
    • 793
  • Jul 09, 2026

    Terminal-Bench — Benchmarking AI Agents in the Terminal

    • terminal-bench
    • benchmark
    • agentic-coding
    • evaluation
    • swe-bench
    • leaderboard
    • laude-institute
  • Jul 03, 2026

    DeepSWE — Datacurve's Long-Horizon Coding Benchmark

    • benchmark
    • evals
    • coding-agents
    • deepswe
    • datacurve
    • swe-bench
    • model-comparison
    • self-verification
  • Jun 12, 2026

    Agent Wikis — Compiled LLM Wikis with Agent-Native Access

    • karpathy-pattern
    • llm-wiki
    • llms-txt
    • agent-readable
    • rag
    • benchmark
    • mcp

Created with Quartz v5.0.0 © 2026

  • ✦ Explore the graph in 3D