Raycaster/ Eval

APEX-Agents · Management Consulting

World112-1_TK_02

Best published0/1Fail

APEX-Agents task World112-1_TK_02 in AI Agents for Hospitality Loyalty Strategy. Compare dual-harness agent runs across models, scores, and public traces.

AI Agents for Hospitality Loyalty StrategyManagement Consulting World 112.1Dual harnessGrader: rubric
task_112defba78604abcb27f4afb573d8d05
Management Consulting World 112.1
message_in_console
5 models · dual config

Task prompt

What the agent was asked to do

Can you do some benchmarking analysis of Impact with its 6 peers for 2024 US operations? Let me know the difference between the lowest average monthly revenue per batch passed and the highest average monthly revenue per batch passed of all the sites. You can allocate the yearly US revenue data from equally across all the respective sites and months to the monthly site operations data. The data in all of the PnLs is in $Ks. Let me know the answer in dollars rounded to the nearest $0.01. Reply to me here with exactly what I asked for.

Published trajectories

Agent runs on this task

Curated dual-harness runs (parsed + original sandbox). Best scored run per model.

ModelHarnessScoreResultLinks
Gemini 3.1 Produal0/1Fail
GPT-5.4dual0/1Fail
GPT-5.4 minidual0/1Fail
GPT-5.4 nanodual0/1Fail
GPT-5.5dual0/1Fail

Grading rubric

Rubric criteria

Runs are graded against these criteria. Open a run for model-specific verdicts.

  1. States that the difference between the lowest average monthly revenue per batch passed and the highest average monthly revenue per batch passed of all the sites in 2024 is $163,064.81