APEX-Agents
gpt-5.5 on SP Task 02 World 129
Grader rubric
Criteria verdict
States that the Seat Purchased Surplus for the Medium utilization band in the Enterprise tier, comparing baseline utilization to 2025 targets, is 47,487
FailStates that the difference between the 2025 Target High (>80%) share for the Business tier and the baseline Actual share is 23 percentage points
Pass
Prompt excerpt
Task context
Use the baseline seat utilization data against the attached 2025 strategic targets for the following two metrics. 1) What is the Seat Purchased Surplus (Actual Seats minus Target Seats) for the Medium utilization band in the Enterprise tier? 2) What is the difference in percentage points between the Target High (>80%) share for the Business tier and the Actual share? Please provide both answers as a reply here, rounded to the nearest whole number.
Response trace
Agent response, tools, files, and edits
On a phone, the interactive viewer works best full-screen — pick the narrative report or the files & trajectory workspace.
Why Eval exists · why Workspace exists
Public evidence and cloud agents are the same harness.
Eval exists so scores are inspectable—tasks, trajectories, artifacts, and rubric verdicts anyone can open.Workspace exists so people can automate real file work with that harness, and so Raycaster never evaluates work it cannot perform.