
Havenor Therapeutics · Its newest medicine, Avericin, is finished and waiting for a health authority to approve it.
Task 10 · QA Product Lead
Confirm whether stability results meet the specification in force
Check every reported result against the limits in force today and set out anything that fails.
GPT-5.6 Sol caught 3 of 6 traps. See which →
3What just happened?
A question came up about test results for Kelvoran, one of Havenor’s medicines already on sale.
Companies keep testing batches after launch. Kelvoran’s limit for one impurity was recently tightened from 0.30% to 0.20%. Someone asked whether all the reported results still pass.
4Who has to do what?
Check every reported result against the limits in force today and set out anything that fails.
- The seat
- Leena Rao, QA Product Lead
- What a good answer looks like
- It uses the new 0.20% limit, flags the one failing batch, confirms the others pass, and spots a rising trend before it fails.
- What must be handed in
AVI-355_specification_acceptability_check.docx
The assignment as the model received it
Reconcile the reported results to the effective specification versions and state the investigation and follow-up obligations.
5Which files decide it?
The current and replaced specifications, the stability data for three batches, and the open lab investigation.
The model also has the rest of the company’s shared drive, its chat and its record systems. Finding the right files is part of the job. Flip through the key ones below.
Company documents
Some of the documents available to the model.
One result, two different limits
The same stability result passes the retained superseded limit and fails the effective specification.
Version 4.0 is effective and sets Impurity K1 at NMT 0.20%. Applying superseded version 3.0 would hide the C22 accelerated excursion.

6What did the model do?
What GPT-5.6 Sol did, step by step.
In short: Sol finds that lot C22’s 0.24% impurity result exceeds the current 0.20% limit and refuses to fall back on the superseded 0.30% limit.
Step 4. Finds the out-of-specification row. The accelerated dataset itself marks the C22 Impurity K1 result as out of specification against the current 0.20% limit.
Dataset row read by the model'STB-AVI355-011-C22-T06-IMPK1', 'AVI355-C22', 6, 'IMPK1', 'Impurity K1 (kelvoran N-oxide)', 0.24
Step 9. Rejects the superseded limit. Sol refuses to use the older 0.30% limit. It stops at this result: it does not show that the other lots stay within limits, or that C22’s long-term trend reaches the limit before 36 months.
Text written into the checkalthough 0.24% would satisfy that historical limit, it is not the criterion in force.
Bears on: Check the other batchesSpot the trend before it fails
Replay every step of the run →
Inside the run
Finds the out-of-specification row
'STB-AVI355-011-C22-T06-IMPK1', 'AVI355-C22', 6, 'IMPK1', 'Impurity K1 (kelvoran N-oxide)', 0.24
The accelerated dataset itself marks the C22 Impurity K1 result as out of specification against the current 0.20% limit.
7Which traps did it catch?
It caught 3 of 6.
Specialists wrote these checks from the real records before any model ran. Each is a weak spot a reviewer or inspector would find. A check passes only if the handed-in file states it.
Missed
Check the other batches R03
Batches C21 and C23 are about 0.17%, within the limit, and all other measurements pass, so C22 is the isolated problem.
If missed, nobody knows how big the problem is.
Spot the trend before it fails R04
In long-term storage C22 is still within the limit, but rising; it reaches 0.20% around 26–27 months, before its 36-month expiry.
If missed, batches on shelves could fail before their expiry date.
The newest data are from 18 months R06
The 24-month long-term results are scheduled but not reported yet.
If missed, the check relies on data that don’t exist.
Caught
Use the limit in force today R01
The current specification caps the impurity at 0.20%. The old 0.30% version has been replaced.
If missed, a failing batch looks fine.
Flag the failing result R02
Batch C22 measured 0.24% at 6 months. That fails the 0.20% limit and would only have passed the old one.
If missed, an out-of-limit medicine stays on the market unexamined.
No investigation covers the failure R05
The only open investigation is about a different test on a different batch. The C22 failure has none.
If missed, a failed result stands with nobody looking into it.
Every model on the same job
One attempt each, same assignment and checklist, so treat small gaps as noise.
Submitted document
AVI-355_specification_acceptability_check.docx
Talk to usRequest the full Biopharma Bench V0.1 — company environments and run receipts: team@raycaster.ai