Raycaster/ Eval

APEX-Agents

gpt-5.5 on World133_ln_04

0/3Fail
Domain
Management Consulting
Category
AI Agents for Divestiture and Spin-Off Analysis
Harness
dual

Grader rubric

Criteria verdict

  1. States the Business segment score is 3.5

    Fail
  2. States the Family segment score is 5.5

    Fail
  3. States the Younger Leisure segment score is 5.5

    Fail

Prompt excerpt

Task context

Give me the scores for Summit’s segments, using the Summit-Specific Survey. Your scoring system must follow these rules: - For the average satisfaction with status benefits from the file “14. Summit_Tier_Status_Experience_Raw_vtest.xlsx”, grant 1 point if it is greater than 2, grant 1 point if it has had no status loss in the past 2 years. Also grant 2 points if more than 23% of respondents in the segment have a current tier equal to "none". - Deduct 1 point if the loyalty score from the file “9. Summit_Loyalty_Outcome_Summary_Raw_vtest.xlsx” is below 2 but add 2.5 points if it's above 2. Present the results in a new one-slide deck. Give all answers rounded to one decimal point.

Response trace

Agent response, tools, files, and edits

Open full trace

On a phone, the interactive viewer works best full-screen — pick the narrative report or the files & trajectory workspace.

Why Eval exists · why Workspace exists

Public evidence and cloud agents are the same harness.

Eval exists so scores are inspectable—tasks, trajectories, artifacts, and rubric verdicts anyone can open.Workspace exists so people can automate real file work with that harness, and so Raycaster never evaluates work it cannot perform.