Raycaster/ Eval

Public benchmark evidence

See the work
behind the score.

Eval is the public evidence library: real tasks, trajectories, artifacts, and rubric verdicts—including failures. Open a published run below. Workspace is the operating surface on the same harness when you want to do the work yourself.

Published evidence · choose a run

The score is only the index.

Open the assigned work, the agent’s execution, the finished artifact, and the exact criteria used to judge it. Nothing below is a staged product demo.

APEX-AgentsLawEdited workbook

Review warranty claims and update a refund workbook

We've received the attached warranty claims for some of our products. Please review them. Then, edit the existing existing "product purchases" spreadsheet to show the maximum refund amount a customer could receive for each product purchased.

Rubric score5/5
Pass
Model
gpt-5.5
Harness
dual
Run cost
$1.31
Input tokens
1,130,520
Sample rubric checks
  • States Falcon Prevent's Max Refund Amount is $0.00
  • States Falcon Query's Max Refund Amount is $16,888.89
  • States Falcon X Elite's Max Refund Amount is $0.00
Live public product surfaceInspect full run ↗

Why Eval exists · why Workspace exists

Public evidence and cloud agents are the same harness.

Eval exists so scores are inspectable—tasks, trajectories, artifacts, and rubric verdicts anyone can open.Workspace exists so people can automate real file work with that harness, and so Raycaster never evaluates work it cannot perform.

Enter by domain

Keep workflow context across every handoff

Biopharma and life sciencesRegulated product state, made measurable.Energy and infrastructureEngineering decisions with operating evidence.HealthcareCare operations where missed risk has consequences.LegalContract and diligence work with inspectable reasoning.FinanceModels scored on the finished artifact.Documents and spreadsheetsProfessional work completed in the artifact.
01 · Prompt

The task and the files the agent is handed — a real, bounded workspace.

02 · Transcript

Every message, tool call, and result, in order. Nothing summarized away.

03 · Verdict

The edits it made and how the rubric graded them. Auditable, not asserted.

Benchmark catalog

Browse agent evaluations by benchmark and work type