Raycaster/ Eval

Raycaster Eval · public benchmark

Inspect the task, trace, and result.

Biopharma Bench V0.1 is Raycaster’s regulatory and CMC benchmark, with published scores across 12 company environments and 71 tasks. APEX-Agents, SpreadsheetBench, and OfficeQA are catalogued here too.

Published runs · choose a run

Inspect a published run.

Review the assignment, recorded agent steps, output file, and rubric criteria. Each selection links to its published task and trace.

APEX-AgentsLawEdited workbook

Review warranty claims and update a refund workbook

We've received the attached warranty claims for some of our products. Please review them. Then, edit the existing existing "product purchases" spreadsheet to show the maximum refund amount a customer could receive for each product purchased.

Rubric score5/5
Pass
Model
gpt-5.5
Harness
dual
Run cost
$1.31
Input tokens
1,130,520
Sample rubric checks
  • ✓States Falcon Prevent's Max Refund Amount is $0.00
  • ✓States Falcon Query's Max Refund Amount is $16,888.89
  • ✓States Falcon X Elite's Max Refund Amount is $0.00
Live public product surfaceInspect full run ↗

Open the live read-only workspace for this run.

Why Eval exists · why Workspace exists

Eval publishes runs. Workspace runs agents on files.

Eval publishes task details, available traces, output files, and rubric results.Workspace is Raycaster’s product for running agents on files.

Enter by domain

Browse runs by domain

Biopharma and life sciencesPublished runs for regulated product workflows.Energy and infrastructureEngineering and finance tasks with scored artifacts.HealthcareHealthcare and senior-living risk tasks.LegalContract and diligence tasks with published traces.FinanceModels scored on the finished artifact.Documents and spreadsheetsDocument and spreadsheet tasks with published runs.
01 · Prompt

The assignment and starting files for a bounded workspace.

02 · Transcript

Available messages, tool calls, and results in chronological order.

03 · Verdict

Output-file changes and the rubric result for the run.

Benchmark catalog

Browse agent evaluations by benchmark and work type