EVERYTHING AIAI engineering, made visual
Visual edition planned
Phase 14 · Lesson 19Learn1 h54 lessons in phase

Benchmarks: SWE-bench, GAIA, AgentBench

Three benchmarks anchor agent evaluation in 2026. SWE-bench tests code patching. GAIA tests generalist tool use. AgentBench tests multi-environment reasoning. Know their composition, their contamination story, and what they do not measure.

Visual edition planned

This lesson isn’t interactive yet.

It is part of the curriculum and will get the same treatment as Phase 1 — a visual cover, hands-on labs, derivations with numeric checks and a quiz. Until then, the original lesson is the best place to read it:

Part of Phase 14Agent Engineering. Use the previous / next cards below to keep browsing the phase.