Raycaster/ Eval

Havenor Therapeutics · Its newest medicine, Avericin, is finished and waiting for a health authority to approve it.

Task 10 · QA Product Lead

Confirm whether stability results meet the specification in force

Check every reported result against the limits in force today and set out anything that fails.

GPT-5.6 Sol caught 3 of 6 traps. See which →

3What just happened?

A question came up about test results for Kelvoran, one of Havenor’s medicines already on sale.

Companies keep testing batches after launch. Kelvoran’s limit for one impurity was recently tightened from 0.30% to 0.20%. Someone asked whether all the reported results still pass.

4Who has to do what?

Check every reported result against the limits in force today and set out anything that fails.

The seat
Leena Rao, QA Product Lead
What a good answer looks like
It uses the new 0.20% limit, flags the one failing batch, confirms the others pass, and spots a rising trend before it fails.
What must be handed in
AVI-355_specification_acceptability_check.docx
The assignment as the model received it
Reconcile the reported results to the effective specification versions and state the investigation and follow-up obligations.

5Which files decide it?

The current and replaced specifications, the stability data for three batches, and the open lab investigation.

The model also has the rest of the company’s shared drive, its chat and its record systems. Finding the right files is part of the job. Flip through the key ones below.

Company documents

Some of the documents available to the model.

Current and superseded specifications · Word

One result, two different limits

The same stability result passes the retained superseded limit and fails the effective specification.

Version 4.0 is effective and sets Impurity K1 at NMT 0.20%. Applying superseded version 3.0 would hide the C22 accelerated excursion.

AVT-SPEC-022 · v4.0DOCX
First page of the real AVT-SPEC-022 v4.0 release and stability specification.
Source excerpt · AVT-SPEC-022 v4.0 and retained v3.0; STB-AVI355-011

6What did the model do?

What GPT-5.6 Sol did, step by step.

In short: Sol finds that lot C22’s 0.24% impurity result exceeds the current 0.20% limit and refuses to fall back on the superseded 0.30% limit.

  1. Step 4. Finds the out-of-specification row. The accelerated dataset itself marks the C22 Impurity K1 result as out of specification against the current 0.20% limit.

    Dataset row read by the model'STB-AVI355-011-C22-T06-IMPK1', 'AVI355-C22', 6, 'IMPK1', 'Impurity K1 (kelvoran N-oxide)', 0.24
  2. Step 9. Rejects the superseded limit. Sol refuses to use the older 0.30% limit. It stops at this result: it does not show that the other lots stay within limits, or that C22’s long-term trend reaches the limit before 36 months.

    Text written into the checkalthough 0.24% would satisfy that historical limit, it is not the criterion in force.

    Bears on: Check the other batchesSpot the trend before it fails

Replay every step of the run →

Inside the run

model sessionTask 10 · step 4
A
Model

Finds the out-of-specification row

Tool result
'STB-AVI355-011-C22-T06-IMPK1', 'AVI355-C22', 6, 'IMPK1', 'Impurity K1 (kelvoran N-oxide)', 0.24
A

The accelerated dataset itself marks the C22 Impurity K1 result as out of specification against the current 0.20% limit.

1 / 2

7Which traps did it catch?

It caught 3 of 6.

Specialists wrote these checks from the real records before any model ran. Each is a weak spot a reviewer or inspector would find. A check passes only if the handed-in file states it.

Missed

  1. Check the other batches R03

    Batches C21 and C23 are about 0.17%, within the limit, and all other measurements pass, so C22 is the isolated problem.

    If missed, nobody knows how big the problem is.

  2. Spot the trend before it fails R04

    In long-term storage C22 is still within the limit, but rising; it reaches 0.20% around 26–27 months, before its 36-month expiry.

    If missed, batches on shelves could fail before their expiry date.

  3. The newest data are from 18 months R06

    The 24-month long-term results are scheduled but not reported yet.

    If missed, the check relies on data that don’t exist.

Caught

  1. Use the limit in force today R01

    The current specification caps the impurity at 0.20%. The old 0.30% version has been replaced.

    If missed, a failing batch looks fine.

  2. Flag the failing result R02

    Batch C22 measured 0.24% at 6 months. That fails the 0.20% limit and would only have passed the old one.

    If missed, an out-of-limit medicine stays on the market unexamined.

  3. No investigation covers the failure R05

    The only open investigation is about a different test on a different batch. The C22 failure has none.

    If missed, a failed result stands with nobody looking into it.

Every model on the same job

One attempt each, same assignment and checklist, so treat small gaps as noise.

Submitted document

AVI-355_specification_acceptability_check.docx

1 / 3

Talk to usRequest the full Biopharma Bench V0.1 — company environments and run receipts: team@raycaster.ai