Jonathon's AI Wiki
Search
Search
Dark mode
Light mode
Explorer
evaluation
4 items with this tag.
Jul 09, 2026
Terminal-Bench — Benchmarking AI Agents in the Terminal
terminal-bench
benchmark
agentic-coding
evaluation
swe-bench
leaderboard
laude-institute
Jul 02, 2026
Prompt Evaluation Tools — Promptfoo vs Braintrust vs the Anthropic Console
evaluation
promptfoo
braintrust
anthropic-console
testing
ci-cd
red-teaming
tooling
May 02, 2026
Adaline — End-to-End AI Agent Platform
llmops
prompt-management
evaluation
deployment
monitoring
agent-platform
multi-provider
Apr 27, 2026
Shopping for Skills and Plugins — A 6-Question Vetting Framework
skills
plugins
marketplaces
vetting
evaluation
governance
security
distribution