Raycaster/ Eval

APEX-Agents · Management Consulting

Task o88f1452

Best published0/2Fail

APEX-Agents task Task o88f1452 in AI Agents for SaaS Due Diligence. Compare dual-harness agent runs across models, scores, and public traces.

AI Agents for SaaS Due DiligenceManagement Consulting World 129Dual harnessGrader: rubric
task_cb7189ae4502436085b4367bf7c64169
Management Consulting World 129
message_in_console
5 models · dual config

Task prompt

What the agent was asked to do

We need to redo the analysis of the new survey response dataset. Can you re-calculate the standard deviation for Brightpath Software's efficiency dataset? Using both the recalculated average efficiency score for Brightpath Software (by all Brightpath users) and the average efficiency score for Brightpath Software by only Brightpath users who ranked AI capabilities as top priority, please also calculate the fraction of a standard deviation that the two scores differ by. Round all final answers to four decimal places. State the output directly to me here as a reply.

Published trajectories

Agent runs on this task

Curated dual-harness runs (parsed + original sandbox). Best scored run per model.

ModelHarnessScoreResultLinks
Gemini 3.1 Produal0/2Fail
GPT-5.4dual0/2Fail
GPT-5.4 minidual0/2Fail
GPT-5.4 nanodual0/2Fail
GPT-5.5dual0/2Fail

Grading rubric

Rubric criteria

Runs are graded against these criteria. Open a run for model-specific verdicts.

  1. States the standard deviation for efficiency scores of Brightpath Software is 1.3836

  2. States the fraction of a standard deviation that the two scores differ is 0.1714