Task
The assignment, starting workspace, source material, and expected deliverable.
Raycaster Eval
Compare 1,081 AI agent benchmark tasks across professional work, with public traces, output artifacts, rubric grades, and model scores.
Choose a benchmark by the work it measures, then inspect the underlying tasks and runs.
The score is connected to the work that produced it.
The assignment, starting workspace, source material, and expected deliverable.
The model transcript, tool calls, file reads, edits, timing, and cost when available.
The completed artifact, model score, rubric-level grades, and grader rationale.
Each benchmark page explains its source, task design, available models, and publication coverage. Individual task pages link to the public runs that can be inspected.
Cross-benchmark groupings — tasks from any benchmark can appear in a category.