Raycaster/ Eval

Raycaster Eval · Biopharma Bench V0.1

Biopharma Bench V0.1 for document work inside pharmaceutical and medical-device companies.

Raycaster's Biopharma Bench V0.1: agents complete CMC, quality, and medical-device assignments inside simulated companies.

12Company environments
71Total tasks
752Scoring criteria
65.9%
Astra macro mean

Start here

Where each company’s product is on its way to patients

Every medicine or medical test goes through the same journey, and each stage creates its own paperwork. Each name below is a company environment in the benchmark, placed where its product sits. Open one to zoom in: what the company does, what just happened, the jobs AI models were given, and, for Havenor and the Broad Institute, every file, step and check.

  1. 1Discovery

    Scientists find a promising molecule or idea.

  2. 2Lab & animal studies

    Safety is tested in cells and animals before any person gets it.

  3. 3Human trials

    Volunteers and patients receive it in carefully controlled studies.

  4. 4Filing & review

    The company sends its evidence to a health authority and answers its questions.

  5. 5Approved

    The authority says yes; the company prepares to launch.

  6. 6On the market

    Patients receive it; the company keeps watching safety and quality.

Inside

Two of the 12 company environments are published in full: one example trial per assignment, with the documents, the submitted file, and every scored check. Those examples are not the leaderboard. The leaderboard is the full 71-task board.

Leaderboard

Benchmark results

71 tasks and 752 criteria across 12 company environments. Rank is the macro task mean: each task counts once. A full pass means every criterion on that task passed. Select a model to follow it through the domains below.

Benchmark scores across eight models. Macro task mean gives equal weight to each task.
RankModelAgentThinking
1Codexmedium65.9%0 / 71499/752$3.32
2Claude Codemedium63.8%0 / 71485/752$2.68
3Cursor CLIhigh63.0%2 / 712.8%480/752$1.16
4Pihigh57.0%0 / 71433/752$0.11
5Cursor CLI / Pidefault48.5%1 / 711.4%374/752$0.72
6Pihigh47.6%0 / 71365/752$0.04
7Cursor CLIhigh46.6%0 / 71356/752$1.12
8Codexmedium46.6%0 / 71364/752$1.27

Codex and Claude Code configurations use medium reasoning effort. DeepSeek, GLM, Grok, and Gemini use high reasoning effort. Kimi uses default provider reasoning. Candidate cost reflects measured token usage priced using public provider API rates. Grader execution spend is excluded.

Frontier comparison

Pareto frontier

Performance frontier across models. Select the horizontal axis to compare macro score against candidate cost, token count, or model release date. The dashed line connects non-dominated models.

  • GLM · Pi
  • DeepSeek · Pi
  • Kimi · Cursor CLI / Pi
  • Gemini · Cursor CLI
  • Grok · Cursor CLI
  • Sol · Codex
  • Opus · Claude Code
  • Astra · Codex
40%50%60%70%$0.00$1.00$2.00$3.00$4.00Candidate Cost (USD)Macro Task MeanGLM · PiDeepSeek · PiGrok · Cursor CLIOpus · Claude CodeAstra · Codex
Pareto frontier across all 71 benchmark tasks. Points connected by the dashed line represent non-dominated configurations (no alternative achieves higher performance at lower cost, fewer tokens, or earlier release).

Where the models differ

Same board, different jobs

A model that leads on one kind of filing can trail on another. Filter the domains. Selecting a model on the leaderboard marks the rows where it is the high or low score.

Domain / SurfaceSpread & Range (Min → Max Macro Score)
US FDA (CDER, CBER, CDRH)42 tasks
GLM-5.3 Flash (41.5%)Δ 26.7%GPT-6 Astra (68.2%)
China NMPA / CDE Quality Files10 tasks
Gemini 3.8 Flash (42.0%)Δ 22.0%Grok 4.6 (64.0%)
Korea MFDS Fast-Track ADC Dossiers7 tasks
Gemini 3.8 Flash (15.4%)Δ 53.8%Claude Opus 5 (69.2%)
mRNA & Biologics Platforms6 tasks
DeepSeek V4.1 Flash (44.8%)Δ 35.8%GPT-6 Astra (80.6%)
IVD & Medical Device Submissions6 tasks
Gemini 3.8 Flash (33.3%)Δ 35.5%GPT-6 Astra (68.8%)
Targeted Small Molecules19 tasks
GLM-5.3 Flash (43.1%)Δ 21.4%Claude Opus 5 (64.5%)
CTD Module 2/3 Quality Leaves28 tasks
Gemini 3.8 Flash (44.1%)Δ 23.4%GPT-6 Astra (67.5%)
GxP Quality & CAPA Audits18 tasks
DeepSeek V4.1 Flash (42.0%)Δ 24.2%Grok 4.6 (66.2%)
Clinical Trial Site Ops & ISF10 tasks
GPT-5.6 Sol (56.7%)Δ 15.7%Gemini 3.8 Flash (72.4%)
Regulatory Strategy & Deficiencies31 tasks
Kimi K3 (46.2%)Δ 20.6%Claude Opus 5 (66.8%)

Complementary domain strengths

Models show distinct strengths across regulatory regimes and document types. An overall ranking averages these domain-specific advantages together.

54/67 vs. 40/67Astra vs. Opus lead

Outside the mRNA biologics environment, Astra and Opus pass the exact same number of items (445 of 685). Astra’s 14-item overall lead comes entirely from this single environment.

Key Insight: Model rankings can hinge entirely on a single domain environment rather than uniform superior capability.

Interactive Failure Inspector

The Anatomy of Real Benchmark Failures

Why do high-performing models fail pharmaceutical and device assignments? Biopharma Bench V0.1 tasks penalize subtle cognitive traps: breaking federal clinical holds, averaging non-poolable degradation slopes, calculating dosages on the wrong salt form, or citing tests scheduled for the future.

Statutory Hold Violation

Held Means Held (MKL-P-01-R03 / R07)

Task: IND Partial Clinical Hold Response (MKL-P-01) · Environment: Broad Institute · FDA CDER IND 167326

The Prompt & Assignment

"Write the binding plan for the response: what may continue, what must stop, and how the answer to the FDA is put together."

Common Model Pitfall (Why Models Fail)

Models frequently followed the internal steering memo’s proposal to "argue against the hold" and prepare trial sites in parallel, advising that screening and site setup could proceed quietly while drafting the rebuttal.

Observed in:Claude Opus 5 (prior run), GPT-5.6 Sol, DeepSeek V4.1 Flash
Ground Truth & Expert Rubric Check

Under 21 CFR 312.42(b)(1)(iv), the symptomatic study was permitted, but the healthy-carrier study was on binding partial clinical hold. No screening, no site setup, and no volunteer contact was legally permissible until FDA formally issues a written removal order.

Strictly passed by:GPT-6 Astra (with strict negative constraint prompting)
Real-World Consequence If Executed by an Agent

If followed, healthy volunteers would have been enrolled and dosed while under an active federal clinical hold—a catastrophic regulatory violation halting all sponsor clinical programs.

Visual Artifact & Deliverable Diffs

Inspect Submitted Documents & Baseline Diffs

Real pharmaceutical evaluation requires examining actual work product. Inspect the exact deliverables written by LLMs inside simulated enterprise workspaces: multi-sheet Excel stability workbooks with cell-by-cell delta inspection, and multi-page regulatory dossiers with high-resolution page previews.

XLSXHavenor Therapeutics · Parenteral BiologicsDeepSeek V4.1 Flash · 9 / 15

STB-AVI401-002_24M_pull_and_update_control.xlsx

280 cells modified across 6 sheets (Control, Pull Plan, Execution Sequence, Chain and Controls, Regulatory Handoff, Review Actions).

Critical Rubric Check: Correctly identifies that AVI401-R01 is 100 mL, but fails on qualified-state basis and misses chamber requalification lapse under AVT-SOP-045.

Cell-Level Workbook Delta (280 changed cells)

Showing sample modified cells before and after model execution:

Sheet: Control5 cell changes
CellBaseline StateCandidate Submission
C7— (blank)0.4 reviewed; not approved. Return as 0.5 after QC-01 to QC-14 close.
C11— (blank)Pull authorization, dataset release and regulatory handoff withheld. See Review Actions QC-01 to QC-15 and the QC Disposition sheet.
Sheet: Pull Plan18 cell changes
CellBaseline StateCandidate Submission
G2— (blank)Hold - lot mapping correct, authorization withheld.
H2— (blank)AVI401-R01 is correctly listed as 100 mL. The pull is not authorized because the testing-chain clearance is not complete (QC-12) and the chamber qualified state is not established (QC-02).
Sheet: Regulatory Handoff36 cell changes
CellBaseline StateCandidate Submission
G2— (blank)Not approved - scope and timing wrong.
H2— (blank)HA-AVI401-IR-2027-014 point 7 states the 24-month pull is not part of sequence 0003 and is to be provided under COM-AVI401-008; no earlier pull is requested for the letter.

What's in the suite

Health authorities, products, documents, and jobs

The benchmark covers five regulators, the product types those desks work on, the document structures in the files, and the jobs people do. This page does not publish the company-by-company roster.

Health authorities

United States

US FDA

CDER for small molecules and therapeutic proteins. CBER for biologics and vaccines. CDRH for devices and in-vitro diagnostics.

European Union

EU EMA

Centralised marketing authorisations, SmPC (prescribing information), and post-marketing variations.

Canada

Health Canada

Abbreviated New Drug Submissions (ANDS).

Republic of Korea

Korea MFDS

GIFT designation and local risk-management plans.

China

China NMPA / CDE

Fixed-dose combination chemistry dossiers and local GMP inspection responses.

Product types

Vaccines

mRNA vaccines

Lipid nanoparticles, cold chain, and bivalent strain updates.

Biologics

Antibody-drug conjugates

Dual-indication accelerated approval and target-specific dosing.

Small molecules

Targeted small molecules / kinase inhibitors

Exon-specific oncology, companion-diagnostic bridging, and synthetic API work.

Peptides

Peptide injectables

GLP-1 pens and sterile fill-finish.

Devices

IVD and medical devices

PMA lifecycle and pan-tumor companion-diagnostic labeling.

Oral solids

Oral fixed-dose combinations

Multi-API process control and dissolution.

Manufacturing

Parenteral biologics and MSAT

Aseptic fill-finish, autoclave performance qualification, lyophilization, and bioburden.

Document formats and operational substrates

Dossier

Common Technical Document (CTD)

Module 2 summaries and Module 3 quality / CMC leaves.

Quality system

GxP quality records

Change controls (CC), out-of-specification (OOS) investigations, deviations, and CAPA.

Trial files

Clinical and regulatory records

Protocols, clinical study reports (CSR), statistical analysis plans (SAP), investigator brochures (IB), informed consent (ICF), and FDA Form 1572.

Labeling

Statutory prescribing information

USPI, EU SmPC, Korean labeling, and instructions for use (IFU).

Core professional task families

Regulatory

Regulatory strategy and filing

Health-authority submissions, responses, and filing strategy.

CMC / MSAT

Chemistry, manufacturing, and controls

Process, specification, and manufacturing-science work.

Labeling

Labeling and posology negotiation

Prescribing information, dose language, and label negotiation.

Quality

Quality management and inspection readiness

Deviations, CAPA, change control, and inspection-facing records.

Clinical

Clinical trial oversight

Protocol, consent, and trial-conduct documentation.

Method

How the benchmark works

How a world is built

Each simulated company has files, email, specifications, and business systems. The candidate takes an employee’s role and must produce a specified document or spreadsheet.

What the candidate can access

Candidates cannot access records from after the assignment date. The clock is fixed to that date, so a scheduled test is still in the future.

How submissions are scored

Each submission is graded against a checklist for that task. Runs where grading failed because of a technical error are excluded from the averages.

What is published

Havenor Therapeutics pages show company documents, submitted files, scores, and summaries of the candidate’s actions. Full transcripts are available in Expert Review with sign-in.

FAQ

Can I run the tasks here?

These pages show completed trials. They do not run an agent in your browser.

Why are the examples from Havenor Therapeutics?

Havenor Therapeutics uses fictional company records that we can publish. For the rest of the suite, this page describes the health authorities and product types — not the individual companies.