APEX-Agents · Law
World 420 LB_05
APEX-Agents task World 420 LB_05 in AI Agents for FDA and Pharma Promotional Compliance. Compare dual-harness agent runs across models, scores, and public traces.
Task prompt
What the agent was asked to do
For purposes of a training presentation for sales representatives giving background on what off-label promotion is and its legal status, please indicate whether it would be accurate to incorporate the following items from Livyra_handbook.pdf: (1) The definition of "misbranding" (2) The language in section 2.3 Please note that you should not consider the above items inaccurate solely based on being incomplete since additional language will be added for purposes of the training session. Also, there is no need to explain your reasoning; I just need a "would be accurate" or "would not be accurate" answer for each item. Provide your response in a message to the console. Consider the following additional sources: 1. US v Facteau.pdf 2. 21 USC 331.pdf 3. 21 USC 352.pdf
Published trajectories
Agent runs on this task
Curated dual-harness runs (parsed + original sandbox). Best scored run per model.
| Model | Harness | Score | Result | Links |
|---|---|---|---|---|
| Gemini 3 Flash | dual | 2/2 | Pass | Run detailsPublic trace |
| Gemini 3.1 Pro | dual | 2/2 | Pass | Run detailsPublic trace |
| GPT-5.4 | dual | 2/2 | Pass | Run detailsPublic trace |
| GPT-5.5 | dual | 2/2 | Pass | Run detailsPublic trace |
| GPT-5.4 mini | dual | 1/2 | Fail | Run detailsPublic trace |
| fireworks models Kimi K2 | dual | 0/2 | Fail | Run detailsPublic trace |
| GPT-5.4 nano | dual | 0/2 | Fail | Run detailsPublic trace |
Grading rubric
Rubric criteria
Runs are graded against these criteria. Open a run for model-specific verdicts.
States that it would not be accurate to incorporate the definition of misbranding from the Livyra handbook
States that it would not be accurate to incorporate the language from section 2.3 of the Livyra handbook