Jonathon's AI Wiki

evaluation

4 items with this tag.

  • Jul 09, 2026

    Terminal-Bench — Benchmarking AI Agents in the Terminal

    • terminal-bench
    • benchmark
    • agentic-coding
    • evaluation
    • swe-bench
    • leaderboard
    • laude-institute
  • Jul 02, 2026

    Prompt Evaluation Tools — Promptfoo vs Braintrust vs the Anthropic Console

    • evaluation
    • promptfoo
    • braintrust
    • anthropic-console
    • testing
    • ci-cd
    • red-teaming
    • tooling
  • May 02, 2026

    Adaline — End-to-End AI Agent Platform

    • llmops
    • prompt-management
    • evaluation
    • deployment
    • monitoring
    • agent-platform
    • multi-provider
  • Apr 27, 2026

    Shopping for Skills and Plugins — A 6-Question Vetting Framework

    • skills
    • plugins
    • marketplaces
    • vetting
    • evaluation
    • governance
    • security
    • distribution

Created with Quartz v5.0.0 © 2026

  • ✦ Explore the graph in 3D