Raycaster/ Eval

Raycaster Eval · Biopharma Bench V0.1

The first benchmark for professional work inside regulated biopharma.

71 end-to-end assignments across 12 private biopharma company environments. We measure whether frontier agents can navigate contradictory records, identify controlling specifications, and deliver audit-ready professional work.

Announcement & ReportHugging Face datasetHarbor

12Company environments
71Total tasks
752Scoring criteria
65.9%
Astra macro mean

Leaderboard

v0.1 · September 2026

Benchmark results

71 tasks and 752 criteria across 12 company environments. Rank is the macro task mean: each task counts once. A full pass means every criterion on that task passed. Click a column to sort. Select a model to mark it on the frontier chart and in the domain comparison further down.

Benchmark scores across eight models. Macro task mean gives equal weight to each task.
RankModelAgentThinking
1Codexmedium65.9%0 / 71499/752$3.3226.320.36.5 min
2Claude Codemedium63.8%0 / 71485/752$2.6829.229.410.6 min
3Cursor CLIhigh63.0%2 / 712.8%480/752$1.1620.546.78.8 min
4Pihigh57.0%0 / 71433/752$0.1136.240.56.3 min
5Cursor CLI / Pidefault48.5%1 / 711.4%374/752$0.7221.925.14.9 min
6Pihigh47.6%0 / 71365/752$0.0422.825.83.1 min
7Cursor CLIhigh46.6%0 / 71356/752$1.1275.277.214.5 min
8Codexmedium46.6%0 / 71364/752$1.2729.223.24.1 min

* Results are subject to variance; one scored trial per task.

Codex and Claude Code configurations use medium reasoning effort. DeepSeek, GLM, Grok, and Gemini use high reasoning effort. Kimi uses default provider reasoning. Candidate cost reflects measured token usage priced using public provider API rates. Grader execution spend is excluded. Steps, tool calls, and time are the average across all 71 tasks, counted from the scored trial logs. Time is agent execution time.

Changelog

v0.1 (September 2026): Initial release of Biopharma Bench covering 71 tasks across 12 private biopharma company environments, scored against 752 frozen substantive criteria using GPT-5.6 Luna.

Frontier

Score against cost, effort, and release date

Each point is one model on all 71 tasks. Pick the horizontal axis. The dashed line joins the models that no other model beats on both score and the chosen axis.

  • GLM · Pi ★
  • DeepSeek · Pi ★
  • Kimi · Cursor CLI / Pi
  • Gemini · Cursor CLI
  • Grok · Cursor CLI ★
  • Sol · Codex
  • Opus · Claude Code ★
  • Astra · Codex ★
40%50%60%70%$0.00$1.00$2.00$3.00Mean cost per task (USD)Macro task meanAstra · CodexOpus · Claude CodeGrok · Cursor CLIDeepSeek · PiGLM · PiKimi · Cursor CLI / PiGemini · Cursor CLISol · Codex
Score is the macro task mean on all 71 tasks. Cost, tokens, steps, tool calls, and time are per-task means on the same 71 tasks. Time is agent execution time. Click a point to select that model on the leaderboard.

Start here

Where each company’s product is on its way to patients

Every medicine or medical test goes through the same journey, and each stage creates its own paperwork. Each name below is a company environment in the benchmark, placed where its product sits. Open one to zoom in: what the company does, what just happened, the jobs AI models were given, and, for Havenor and the Broad Institute, every file, step and check.

  1. 1Discovery

    Scientists find a promising molecule or idea.

  2. 2Lab & animal studies

    Safety is tested in cells and animals before any person gets it.

  3. 3Human trials

    Volunteers and patients receive it in carefully controlled studies.

  4. 4Filing & review

    The company sends its evidence to a health authority and answers its questions.

  5. 5Approved

    The authority says yes; the company prepares to launch.

  6. 6On the market

    Patients receive it; the company keeps watching safety and quality.

Inside

Two of the 12 company environments are published in full: one example trial per assignment, with the documents, the submitted file, and every scored check. Those examples are not the leaderboard. The leaderboard is the full 71-task board.

10 assignments at a fictional manufacturer.A synthetic company we can publish end to end. Partial-pilot examples are labeled. These rows are example trials, not the leaderboard.

  1. Task 01Prepare the next AVI-401 stability commitment updateRegulatory CMC ManagerDeepSeek V4.1 Flash · partial pilotGPT-6 Astra (Sept 2026 calendar, showcase): 7 / 115 / 11
  2. Task 02State the technical position on the alternate overwrap changeHead of CMC and MSATGPT-5.6 SolGPT-6 Astra (Sept 2026 calendar, showcase): 11 / 117 / 11
  3. Task 03Review the AVI-401 sterile-filter validation studiesPrincipal MSAT EngineerClaude Opus 5GPT-6 Astra (Sept 2026 calendar, showcase): 6 / 95 / 9
  4. Task 04Answer the regulator’s seven questionsRegulatory CMC ManagerGPT-5.6 SolGPT-6 Astra (Sept 2026 calendar, showcase): 12 / 1411 / 14
  5. Task 05Pre-submission records integrity checkQA Product LeadClaude Opus 5GPT-6 Astra (Sept 2026 calendar, showcase): 8 / 95 / 9
  6. Task 06Check the container-closure validation recordsValidation EngineerClaude Opus 5GPT-6 Astra (Sept 2026 calendar, showcase): 4 / 95 / 9
  7. Task 07Supplier and change record integrity checkSupplier Quality ManagerClaude Opus 5GPT-6 Astra (Sept 2026 calendar, showcase): 3 / 84 / 8
  8. Task 08Review the 24-month stability-testing workbookQC Stability ScientistDeepSeek V4.1 Flash · partial pilotGPT-6 Astra (Sept 2026 calendar, showcase): 13 / 159 / 15
  9. Task 09Decide which stability-testing operations can proceedQC Stability ScientistGPT-5.6 SolGPT-6 Astra (Sept 2026 calendar, showcase): 12 / 157 / 15
  10. Task 10Confirm whether stability results meet the specification in forceQA Product LeadGPT-5.6 SolGPT-6 Astra (Sept 2026 calendar, showcase): 5 / 63 / 6
Open the Havenor Therapeutics hub →

Where the models differ

Same board, different jobs

A model that leads on one kind of filing can trail on another. Filter the domains. Selecting a model on the leaderboard marks the rows where it is the high or low score.

Domain / SurfaceSpread & Range (Min → Max Macro Score)
US FDA (CDER, CBER, CDRH)42 tasks
GLM-5.3 Flash (41.5%)Δ 26.7%GPT-6 Astra (68.2%)
China NMPA / CDE Quality Files10 tasks
Gemini 3.8 Flash (42.0%)Δ 22.0%Grok 4.6 (64.0%)
Korea MFDS Fast-Track ADC Dossiers7 tasks
Gemini 3.8 Flash (15.4%)Δ 53.8%Claude Opus 5 (69.2%)
mRNA & Biologics Platforms6 tasks
DeepSeek V4.1 Flash (44.8%)Δ 35.8%GPT-6 Astra (80.6%)
IVD & Medical Device Submissions6 tasks
Gemini 3.8 Flash (33.3%)Δ 35.5%GPT-6 Astra (68.8%)
Targeted Small Molecules19 tasks
GLM-5.3 Flash (43.1%)Δ 21.4%Claude Opus 5 (64.5%)
CTD Module 2/3 Quality Leaves28 tasks
Gemini 3.8 Flash (44.1%)Δ 23.4%GPT-6 Astra (67.5%)
GxP Quality & CAPA Audits18 tasks
DeepSeek V4.1 Flash (42.0%)Δ 24.2%Grok 4.6 (66.2%)
Clinical Trial Site Ops & ISF10 tasks
GPT-5.6 Sol (56.7%)Δ 15.7%Gemini 3.8 Flash (72.4%)
Regulatory Strategy & Deficiencies31 tasks
Kimi K3 (46.2%)Δ 20.6%Claude Opus 5 (66.8%)

Complementary domain strengths

Models show distinct strengths across regulatory regimes and document types. An overall ranking averages these domain-specific advantages together.

54/67 vs. 40/67Astra vs. Opus lead

Outside the mRNA biologics environment, Astra and Opus pass the exact same number of items (445 of 685). Astra’s 14-item overall lead comes entirely from this single environment.

Key Insight: Model rankings can hinge entirely on a single domain environment rather than uniform superior capability.

Interactive Failure Inspector

The Anatomy of Real Benchmark Failures

Why do high-performing models fail pharmaceutical and device assignments? Biopharma Bench V0.1 tasks penalize subtle cognitive traps: breaking federal clinical holds, averaging non-poolable degradation slopes, calculating dosages on the wrong salt form, or citing tests scheduled for the future.

Statutory Hold Violation

Held Means Held (MKL-P-01-R03 / R07)

Task: IND Partial Clinical Hold Response (MKL-P-01) · Environment: Broad Institute · FDA CDER IND 167326

The Prompt & Assignment

"Write the binding plan for the response: what may continue, what must stop, and how the answer to the FDA is put together."

Common Model Pitfall (Why Models Fail)

Models frequently followed the internal steering memo’s proposal to "argue against the hold" and prepare trial sites in parallel, advising that screening and site setup could proceed quietly while drafting the rebuttal.

Observed in:Claude Opus 5, GPT-5.6 Sol, DeepSeek V4.1 Flash
Ground Truth & Expert Rubric Check

Under 21 CFR 312.42(b)(1)(iv), the symptomatic study was permitted, but the healthy-carrier study was on binding partial clinical hold. No screening, no site setup, and no volunteer contact was legally permissible until FDA formally issues a written removal order.

Strictly passed by:GPT-6 Astra
Real-World Consequence If Executed by an Agent

If followed, healthy volunteers would have been enrolled while under an active federal clinical hold—a regulatory violation under 21 CFR 312.42 that subjects the presymptomatic trial to formal enforcement action.

Visual Artifact & Deliverable Diffs

Inspect Submitted Documents & Baseline Diffs

Real pharmaceutical evaluation requires examining actual work product. Inspect the exact deliverables written by LLMs inside simulated enterprise workspaces: multi-sheet Excel stability workbooks with cell-by-cell delta inspection, and multi-page regulatory dossiers with high-resolution page previews.

XLSXHavenor Therapeutics · Parenteral BiologicsDeepSeek V4.1 Flash · 9 / 15

STB-AVI401-002_24M_pull_and_update_control.xlsx

280 cells modified across 6 sheets (Control, Pull Plan, Execution Sequence, Chain and Controls, Regulatory Handoff, Review Actions).

Critical Rubric Check: Correctly identifies that AVI401-R01 is 100 mL, but fails on qualified-state basis and misses chamber requalification lapse under AVT-SOP-045.

Cell-Level Workbook Delta (280 changed cells)

Showing sample modified cells before and after model execution:

Sheet: Control5 cell changes
CellBaseline StateCandidate Submission
C7— (blank)0.4 reviewed; not approved. Return as 0.5 after QC-01 to QC-14 close.
C11— (blank)Pull authorization, dataset release and regulatory handoff withheld. See Review Actions QC-01 to QC-15 and the QC Disposition sheet.
Sheet: Pull Plan18 cell changes
CellBaseline StateCandidate Submission
G2— (blank)Hold - lot mapping correct, authorization withheld.
H2— (blank)AVI401-R01 is correctly listed as 100 mL. The pull is not authorized because the testing-chain clearance is not complete (QC-12) and the chamber qualified state is not established (QC-02).
Sheet: Regulatory Handoff36 cell changes
CellBaseline StateCandidate Submission
G2— (blank)Not approved - scope and timing wrong.
H2— (blank)HA-AVI401-IR-2026-014 point 7 states the 24-month pull is not part of sequence 0003 and is to be provided under COM-AVI401-008; no earlier pull is requested for the letter.

What's in the suite

Health authorities, products, documents, and jobs

The benchmark covers five regulators, the product types those desks work on, the document structures in the files, and the jobs people do. This page does not publish the company-by-company roster.

Health authorities

United States

US FDA

CDER for small molecules and therapeutic proteins. CBER for biologics and vaccines. CDRH for devices and in-vitro diagnostics.

European Union

EU EMA

Centralised marketing authorisations, SmPC (prescribing information), and post-marketing variations.

Canada

Health Canada

Abbreviated New Drug Submissions (ANDS).

Republic of Korea

Korea MFDS

GIFT designation and local risk-management plans.

China

China NMPA / CDE

Fixed-dose combination chemistry dossiers and local GMP inspection responses.

Product types

Vaccines

mRNA vaccines

Lipid nanoparticles, cold chain, and bivalent strain updates.

Biologics

Antibody-drug conjugates

Dual-indication accelerated approval and target-specific dosing.

Small molecules

Targeted small molecules / kinase inhibitors

Exon-specific oncology, companion-diagnostic bridging, and synthetic API work.

Peptides

Peptide injectables

GLP-1 pens and sterile fill-finish.

Devices

IVD and medical devices

PMA lifecycle and pan-tumor companion-diagnostic labeling.

Oral solids

Oral fixed-dose combinations

Multi-API process control and dissolution.

Manufacturing

Parenteral biologics and MSAT

Aseptic fill-finish, autoclave performance qualification, lyophilization, and bioburden.

Document formats and operational substrates

Dossier

Common Technical Document (CTD)

Module 2 summaries and Module 3 quality / CMC leaves.

Quality system

GxP quality records

Change controls (CC), out-of-specification (OOS) investigations, deviations, and CAPA.

Trial files

Clinical and regulatory records

Protocols, clinical study reports (CSR), statistical analysis plans (SAP), investigator brochures (IB), informed consent (ICF), and FDA Form 1572.

Labeling

Statutory prescribing information

USPI, EU SmPC, Korean labeling, and instructions for use (IFU).

Core professional task families

Regulatory

Regulatory strategy and filing

Health-authority submissions, responses, and filing strategy.

CMC / MSAT

Chemistry, manufacturing, and controls

Process, specification, and manufacturing-science work.

Labeling

Labeling and posology negotiation

Prescribing information, dose language, and label negotiation.

Quality

Quality management and inspection readiness

Deviations, CAPA, change control, and inspection-facing records.

Clinical

Clinical trial oversight

Protocol, consent, and trial-conduct documentation.

Method

How the benchmark works

How a world is built

Each simulated company has files, email, specifications, and business systems. The candidate takes an employee’s role and must produce a specified document or spreadsheet.

What the candidate can access

Candidates cannot access records from after the assignment date: records stop at the assignment date, and Havenor's drive is dated around September 2026, so scheduled tests remain in the future.

How submissions are scored

Each submission is graded against a checklist for that task. Runs where grading failed because of a technical error are excluded from the averages.

What is published

Havenor Therapeutics pages show company documents, submitted files, scores, and summaries of the candidate’s actions. Full transcripts are available in Expert Review with sign-in.

FAQ

Can I run the tasks here?

These pages show completed trials. They do not run an agent in your browser. The announcement explains the results. The Havenor desks and rubrics are the Hugging Face dataset. Harbor runs the full instrument, including records, LIMS, and the verifier.

Why are the examples from Havenor Therapeutics?

Havenor Therapeutics uses fictional company records that we can publish. For the rest of the suite, this page describes the health authorities and product types — not the individual companies.