Raycaster/ Eval

APEX-Agents

gpt-5.4-nano on Task iv36a08a

1/7Fail
Domain
Law
Category
AI Agents for SEC Disclosure Analysis
Harness
dual

Grader rubric

Criteria verdict

  1. States that Count 1 is Not Covered

  2. States that Count 2 is Not Covered

  3. States that Count 3 is Not Covered

  4. States that Count 4 is Not Covered

  5. States that Count 5 is Covered

  6. States that Count 6 is Not Covered

  7. States that Count 7 is Not Covered

Prompt excerpt

Task context

Analyze whether Counts 1-7 of Delta's Complaint against CrowdStrike fall within the limitation of liability clause in Section 10.1 of CrowdStrike's standard MSA, indicating "Covered" or "Not Covered" for each count. Reply back to me here with your assessments.

Response trace

Agent response, tools, files, and edits

Open full trace

On a phone, the interactive viewer works best full-screen — pick the narrative report or the files & trajectory workspace.

Why Eval exists · why Workspace exists

Public evidence and cloud agents are the same harness.

Eval exists so scores are inspectable—tasks, trajectories, artifacts, and rubric verdicts anyone can open.Workspace exists so people can automate real file work with that harness, and so Raycaster never evaluates work it cannot perform.