Raycaster/ Eval

APEX-Agents

gpt-5.5 on World244_JP_01

0/5Fail
Domain
Investment Banking
Category
AI Agents for M&A Legal Due Diligence
Harness
dual

Grader rubric

Criteria verdict

  1. States equity contribution is $2,799 million

    Fail
  2. States IRR is 32.66%

    Fail
  3. States MOIC is 4.11x

    Fail
  4. States exit net debt is $1,615 million

    Fail
  5. States maximum revolver amount drawn is $574 million

    Fail

Prompt excerpt

Task context

Use the LBO model with the following indicative debt package to calculate these values --> then, return them back to me here 1/ Equity contribution 2/ Central case IRR 3/ Central case MOIC 4/ Exit net debt 5/ Maximum amount of revolver drawn Term Loan A: Amount: $1.8bn Term: 7 years, straight line amortising Rate: 7-year US Treasury (market rate) + 225bps Arrangement Fee: 0.75% Term Loan B: Amount: $600m Term: 10 years, bullet repayment Rate: 10-year US Treasury (market rate) + 275bps Arrangement Fee: 0.75% Revolver: Amount: $600m Rate: 5.5% Round percentages and multiples to two decimal places, and dollar amounts in millions, rounded to the nearest whole number. Assume market rates from 28-Nov-2025 (U.S. Treasury Daily CMT).

Response trace

Agent response, tools, files, and edits

Open full trace

On a phone, the interactive viewer works best full-screen — pick the narrative report or the files & trajectory workspace.

Why Eval exists · why Workspace exists

Public evidence and cloud agents are the same harness.

Eval exists so scores are inspectable—tasks, trajectories, artifacts, and rubric verdicts anyone can open.Workspace exists so people can automate real file work with that harness, and so Raycaster never evaluates work it cannot perform.