APEX-Agents · Management Consulting
Task_World130_CamilleMoingeon_4
APEX-Agents task Task_World130_CamilleMoingeon_4 in AI Agents for Digital Transformation. Compare dual-harness agent runs across models, scores, and public traces.
Task prompt
What the agent was asked to do
Can you look at the Frito-Lay case study and apply their downtime reduction to HarFeast Good Group's number in the baseline file? I want to estimate what the improvement would look like for us (rounded to the nearest full percentage point). Output the information in a message here.
Published trajectories
Agent runs on this task
Curated dual-harness runs (parsed + original sandbox). Best scored run per model.
| Model | Harness | Score | Result | Links |
|---|---|---|---|---|
| Gemini 3.1 Pro | dual | 0/5 | Fail | Run detailsPublic trace |
| GPT-5.4 | dual | 0/5 | Fail | Run detailsPublic trace |
| GPT-5.4 mini | dual | 0/5 | Fail | Run detailsPublic trace |
| GPT-5.4 nano | dual | 0/5 | Fail | Run detailsPublic trace |
| GPT-5.5 | dual | 0/5 | Fail | Run detailsPublic trace |
Grading rubric
Rubric criteria
Runs are graded against these criteria. Open a run for model-specific verdicts.
States that the new unplanned downtime ratio for the Rockford, Illinois plant is 14%
States that the new unplanned downtime ratio for the Madison, Wisconsin plant is 14%
States that the new unplanned downtime ratio for the Cedar Rapids, Iowa plant is 13%
States that the new unplanned downtime ratio for the Toledo, Ohio plant is 14%
States that the new unplanned downtime ratio for the Kalamazoo, Michigan plant is 15%