APEX-Agents · Management Consulting
Task o88f1452
APEX-Agents task Task o88f1452 in AI Agents for SaaS Due Diligence. Compare dual-harness agent runs across models, scores, and public traces.
Task prompt
What the agent was asked to do
We need to redo the analysis of the new survey response dataset. Can you re-calculate the standard deviation for Brightpath Software's efficiency dataset? Using both the recalculated average efficiency score for Brightpath Software (by all Brightpath users) and the average efficiency score for Brightpath Software by only Brightpath users who ranked AI capabilities as top priority, please also calculate the fraction of a standard deviation that the two scores differ by. Round all final answers to four decimal places. State the output directly to me here as a reply.
Published trajectories
Agent runs on this task
Curated dual-harness runs (parsed + original sandbox). Best scored run per model.
| Model | Harness | Score | Result | Links |
|---|---|---|---|---|
| Gemini 3.1 Pro | dual | 0/2 | Fail | Run detailsPublic trace |
| GPT-5.4 | dual | 0/2 | Fail | Run detailsPublic trace |
| GPT-5.4 mini | dual | 0/2 | Fail | Run detailsPublic trace |
| GPT-5.4 nano | dual | 0/2 | Fail | Run detailsPublic trace |
| GPT-5.5 | dual | 0/2 | Fail | Run detailsPublic trace |
Grading rubric
Rubric criteria
Runs are graded against these criteria. Open a run for model-specific verdicts.
States the standard deviation for efficiency scores of Brightpath Software is 1.3836
States the fraction of a standard deviation that the two scores differ is 0.1714