APEX-Agents
gpt-5.4-nano on Task 14
Grader rubric
Criteria verdict
States the average adoption rate of Business is 30.0%
States the average adoption rate of Growth is 42.1%
States the average adoption rate of Enterprise is 44.9%
States the average monthly usage of Business is 19.1 hours per month
States the average monthly usage of Growth is 50.3 hours per month
States the average monthly usage of Enterprise is 51.1 hours per month
States the average retention impact of Business is 1.46
States the average retention impact of Growth is 1.59
States the average retention impact of Enterprise is 1.50
Prompt excerpt
Task context
Can you use the feature usage file and identify the average adoption rate percentage, average monthly usage, and average retention impact score for each tier except Team? Print the output for me here.
Response trace
Agent response, tools, files, and edits
On a phone, the interactive viewer works best full-screen — pick the narrative report or the files & trajectory workspace.
Why Eval exists · why Workspace exists
Public evidence and cloud agents are the same harness.
Eval exists so scores are inspectable—tasks, trajectories, artifacts, and rubric verdicts anyone can open.Workspace exists so people can automate real file work with that harness, and so Raycaster never evaluates work it cannot perform.