APEX-Agents · Management Consulting
W134 Nancy Task 06
APEX-Agents task W134 Nancy Task 06 in AI Agents for M&A Legal Due Diligence. Compare dual-harness agent runs across models, scores, and public traces.
Task prompt
What the agent was asked to do
Create a new slide pptx, summarizing the comparable SaaS deals' target company name, purchase price, ARR, and ARR multiple. Include all targets for which we have an individual case study and use only publicly disclosed data. Round multiples to one decimal point and, for financial values, provide numbers in millions rounded to the nearest million or, if above 1 billion, in billions with one decimal place. Get insights from the comparable SaaS deals, the internal memo about valuation ranges and negotiation levers, and the case studies about Beacon, Stuzo, TASK, Claap and Statsig.
Published trajectories
Agent runs on this task
Curated dual-harness runs (parsed + original sandbox). Best scored run per model.
| Model | Harness | Score | Result | Links |
|---|---|---|---|---|
| Gemini 3.1 Pro | dual | 5/5 | Pass | Run detailsPublic trace |
| GPT-5.4 | dual | 5/5 | Pass | Run detailsPublic trace |
| GPT-5.4 mini | dual | 5/5 | Pass | Run detailsPublic trace |
| GPT-5.4 nano | dual | 5/5 | Pass | Run detailsPublic trace |
| GPT-5.5 | dual | 3/5 | Fail | Run detailsPublic trace |
Grading rubric
Rubric criteria
Runs are graded against these criteria. Open a run for model-specific verdicts.
States that the Stuzo ARR multiple is 4.8x
States that the TASK ARR multiple is 5.2x
States that the Claap ARR multiple is not available
States that the Statsig ARR multiple is not available
States that the Beacon ARR multiple is 12.7x