APEX-Agents · Management Consulting
World 128_RG_04
APEX-Agents task World 128_RG_04 in AI Agents for Cross-Border Regulatory Review. Compare dual-harness agent runs across models, scores, and public traces.
Task prompt
What the agent was asked to do
Can you use the final versions of the AmensaMech, SkyLink, SolisOne, and AmensaDrive BU Assessment Summary decks to tell me the total decision score for each business unit? For this analysis, let's assume the business unit's total decision score equals the simple average of the five decision criteria scores. The attached file on the Decision criteria score can be used to convert the decision criteria into their corresponding numerical scores. If the decision criteria are missing for any business unit, omit them from this analysis. Assume 'Weak' = 'Low' and 'Strong' = 'High' when converting scores. Round all final answers to 2 decimal places. Reply to me with this information here.
Published trajectories
Agent runs on this task
Curated dual-harness runs (parsed + original sandbox). Best scored run per model.
| Model | Harness | Score | Result | Links |
|---|---|---|---|---|
| Gemini 3.1 Pro | dual | 3/3 | Pass | Run detailsPublic trace |
| GPT-5.4 | dual | 3/3 | Pass | Run detailsPublic trace |
| GPT-5.4 mini | dual | 3/3 | Pass | Run detailsPublic trace |
| GPT-5.4 nano | dual | 3/3 | Pass | Run detailsPublic trace |
| GPT-5.5 | dual | 3/3 | Pass | Run detailsPublic trace |
Grading rubric
Rubric criteria
Runs are graded against these criteria. Open a run for model-specific verdicts.
States that the total decision score for SolisOne is 2.40
States that the total decision score for AmensaMech is 2.40
States that the total decision score for AmensaDrive is 3.00