Raycaster/ Eval

APEX-Agents · Management Consulting

World133_ln_04

Best published3/3Pass

APEX-Agents task World133_ln_04 in AI Agents for Divestiture and Spin-Off Analysis. Compare dual-harness agent runs across models, scores, and public traces.

AI Agents for Divestiture and Spin-Off AnalysisManagement Consulting World 133Dual harnessGrader: rubric
task_40c59502fc744e85a83bf87dfe8977da
Management Consulting World 133
make_new_slide_deck
5 models · dual config

Task prompt

What the agent was asked to do

Give me the scores for Summit’s segments, using the Summit-Specific Survey. Your scoring system must follow these rules: - For the average satisfaction with status benefits from the file “14. Summit_Tier_Status_Experience_Raw_vtest.xlsx”, grant 1 point if it is greater than 2, grant 1 point if it has had no status loss in the past 2 years. Also grant 2 points if more than 23% of respondents in the segment have a current tier equal to "none". - Deduct 1 point if the loyalty score from the file “9. Summit_Loyalty_Outcome_Summary_Raw_vtest.xlsx” is below 2 but add 2.5 points if it's above 2. Present the results in a new one-slide deck. Give all answers rounded to one decimal point.

Published trajectories

Agent runs on this task

Curated dual-harness runs (parsed + original sandbox). Best scored run per model.

ModelHarnessScoreResultLinks
GPT-5.4 minidual3/3Pass
GPT-5.4 nanodual3/3Pass
Gemini 3.1 Produal2/3Fail
GPT-5.4dual2/3Fail
GPT-5.5dual0/3Fail

Grading rubric

Rubric criteria

Runs are graded against these criteria. Open a run for model-specific verdicts.

  1. States the Business segment score is 3.5

  2. States the Family segment score is 5.5

  3. States the Younger Leisure segment score is 5.5