APEX-Agents · Management Consulting
World112-1_TK_02
APEX-Agents task World112-1_TK_02 in AI Agents for Hospitality Loyalty Strategy. Compare dual-harness agent runs across models, scores, and public traces.
Task prompt
What the agent was asked to do
Can you do some benchmarking analysis of Impact with its 6 peers for 2024 US operations? Let me know the difference between the lowest average monthly revenue per batch passed and the highest average monthly revenue per batch passed of all the sites. You can allocate the yearly US revenue data from equally across all the respective sites and months to the monthly site operations data. The data in all of the PnLs is in $Ks. Let me know the answer in dollars rounded to the nearest $0.01. Reply to me here with exactly what I asked for.
Published trajectories
Agent runs on this task
Curated dual-harness runs (parsed + original sandbox). Best scored run per model.
| Model | Harness | Score | Result | Links |
|---|---|---|---|---|
| Gemini 3.1 Pro | dual | 0/1 | Fail | Run detailsPublic trace |
| GPT-5.4 | dual | 0/1 | Fail | Run detailsPublic trace |
| GPT-5.4 mini | dual | 0/1 | Fail | Run detailsPublic trace |
| GPT-5.4 nano | dual | 0/1 | Fail | Run detailsPublic trace |
| GPT-5.5 | dual | 0/1 | Fail | Run detailsPublic trace |
Grading rubric
Rubric criteria
Runs are graded against these criteria. Open a run for model-specific verdicts.
States that the difference between the lowest average monthly revenue per batch passed and the highest average monthly revenue per batch passed of all the sites in 2024 is $163,064.81