Raycaster/ Eval

APEX-Agents

gpt-5.5 on World 128 - NK - Task 1

1/9Fail
Domain
Management Consulting
Category
AI Agents for Cross-Border Regulatory Review
Harness
dual

Grader rubric

Criteria verdict

  1. States Waymo RMS in Los Angeles is 1.00

    Fail
  2. States Cruise RMS in Los Angeles is 0.75

    Fail
  3. States Tesla RMS in Los Angeles is 0.50

    Fail
  4. States Waymo RMS in San Francisco is 1.00

    Fail
  5. States Cruise RMS in San Francisco is 1.00

    Fail
  6. States Tesla RMS in San Francisco is 0.67

    Fail
  7. States Waymo RMS in Sacramento is 1.00

    Fail
  8. States Cruise RMS in Sacramento is 0.50

    Fail
  9. States that AmensaDrive operates in one out of the three provided cities

    Pass

Prompt excerpt

Task context

Let's evaluate the new market intelligence we've received to figure out the RMS for competitors in each city. Tell me how many of these cities AmensaDrive operates in in 2026. Format answers to two decimal points. Provide your response right here.

Response trace

Agent response, tools, files, and edits

Open full trace

On a phone, the interactive viewer works best full-screen — pick the narrative report or the files & trajectory workspace.

Why Eval exists · why Workspace exists

Public evidence and cloud agents are the same harness.

Eval exists so scores are inspectable—tasks, trajectories, artifacts, and rubric verdicts anyone can open.Workspace exists so people can automate real file work with that harness, and so Raycaster never evaluates work it cannot perform.