APEX-Agents
gpt-5.5 on Task 14
Grader rubric
Criteria verdict
States the average adoption rate of Business is 30.0%
FailStates the average adoption rate of Growth is 42.1%
FailStates the average adoption rate of Enterprise is 44.9%
FailStates the average monthly usage of Business is 19.1 hours per month
PassStates the average monthly usage of Growth is 50.3 hours per month
PassStates the average monthly usage of Enterprise is 51.1 hours per month
PassStates the average retention impact of Business is 1.46
PassStates the average retention impact of Growth is 1.59
PassStates the average retention impact of Enterprise is 1.50
Pass
Prompt excerpt
Task context
Can you use the feature usage file and identify the average adoption rate percentage, average monthly usage, and average retention impact score for each tier except Team? Print the output for me here.
Response trace
Agent response, tools, files, and edits
On a phone, the interactive viewer works best full-screen — pick the narrative report or the files & trajectory workspace.
Why Eval exists · why Workspace exists
Public evidence and cloud agents are the same harness.
Eval exists so scores are inspectable—tasks, trajectories, artifacts, and rubric verdicts anyone can open.Workspace exists so people can automate real file work with that harness, and so Raycaster never evaluates work it cannot perform.