APEX-Agents · Management Consulting
world130_HO_05
APEX-Agents task world130_HO_05 in AI Agents for Digital Transformation. Compare dual-harness agent runs across models, scores, and public traces.
Task prompt
What the agent was asked to do
Use the v1 version of the survey responses to identify the number of respondents who received any kind of training on digital tools. Of those respondents, return the percentage of respondents for each training quality rating. Reply back here to me.
Published trajectories
Agent runs on this task
Curated dual-harness runs (parsed + original sandbox). Best scored run per model.
| Model | Harness | Score | Result | Links |
|---|---|---|---|---|
| Gemini 3.1 Pro | dual | 5/5 | Pass | Run detailsPublic trace |
| GPT-5.4 | dual | 5/5 | Pass | Run detailsPublic trace |
| GPT-5.4 mini | dual | 5/5 | Pass | Run detailsPublic trace |
| GPT-5.4 nano | dual | 5/5 | Pass | Run detailsPublic trace |
| GPT-5.5 | dual | 5/5 | Pass | Run detailsPublic trace |
Grading rubric
Rubric criteria
Runs are graded against these criteria. Open a run for model-specific verdicts.
States that the number of respondents who received any kind of training on digital tools is 1200
States that percentage of respondents rated the training quality as "Excellent- comprehensive and very helpful” is 16%
States that percentage of respondents rated the training quality as "Good- adequate for most needs” is 41%
States that percentage of respondents rated the training quality as “Fair- some gaps or inconsistencies” is 33%
States that percentage of respondents rated the training quality as “Poor - insufficient or unhelpful” is 11%