Raycaster/ Eval

APEX-Agents · Management Consulting

world130_HO_05

Best published5/5Pass

APEX-Agents task world130_HO_05 in AI Agents for Digital Transformation. Compare dual-harness agent runs across models, scores, and public traces.

AI Agents for Digital TransformationManagement Consulting World 130Dual harnessGrader: rubric
task_dba0d3217ada47eb93028eed8e14d69d
Management Consulting World 130
message_in_console
5 models · dual config

Task prompt

What the agent was asked to do

Use the v1 version of the survey responses to identify the number of respondents who received any kind of training on digital tools. Of those respondents, return the percentage of respondents for each training quality rating. Reply back here to me.

Published trajectories

Agent runs on this task

Curated dual-harness runs (parsed + original sandbox). Best scored run per model.

ModelHarnessScoreResultLinks
Gemini 3.1 Produal5/5Pass
GPT-5.4dual5/5Pass
GPT-5.4 minidual5/5Pass
GPT-5.4 nanodual5/5Pass
GPT-5.5dual5/5Pass

Grading rubric

Rubric criteria

Runs are graded against these criteria. Open a run for model-specific verdicts.

  1. States that the number of respondents who received any kind of training on digital tools is 1200

  2. States that percentage of respondents rated the training quality as "Excellent- comprehensive and very helpful” is 16%

  3. States that percentage of respondents rated the training quality as "Good- adequate for most needs” is 41%

  4. States that percentage of respondents rated the training quality as “Fair- some gaps or inconsistencies” is 33%

  5. States that percentage of respondents rated the training quality as “Poor - insufficient or unhelpful” is 11%