Raycaster/ Eval

APEX-Agents

GPT-5.4 on World125_Task01_NA

4/4Pass
Domain
Management Consulting
Category
AI Agents for Cross-Border Regulatory Review
Harness
dual

Grader rubric

Criteria verdict

  1. States that Impact Therapeutic is a company that primarily sells branded drugs

  2. States the average of COGS as a share of revenue for companies that primarily sell branded drugs is 23%

  3. States that the global COGS as a share of revenue for Impact Therapeutics in 2024 is higher than the average of 23%

  4. States that the US COGS as a share of revenue for Impact Therapeutics in 2024 is higher than the average of 23%

Prompt excerpt

Task context

Using the 2024 annual report, identify whether Impact Therapeutics sells primarily branded drugs. A company sells primarily branded drugs if the majority of drugs have launch dates within the last 7 years. Please use the BCG report to determine the average COGS as a % of revenue for competitors in Impact Therapeutics' segment, based on whether they sell primarily generic drugs, branded, or a mix. Round to the nearest %. In both the global and US geos, state whether Impact Therapeutics is above the segment average calculated from the data in the BCG report. Write your response as a reply to me here.

Response trace

Agent response, tools, files, and edits

Open full trace

On a phone, the interactive viewer works best full-screen — pick the narrative report or the files & trajectory workspace.

Why Eval exists · why Workspace exists

Public evidence and cloud agents are the same harness.

Eval exists so scores are inspectable—tasks, trajectories, artifacts, and rubric verdicts anyone can open.Workspace exists so people can automate real file work with that harness, and so Raycaster never evaluates work it cannot perform.