APEX-Agents · Management Consulting
World 127_AH_Task 2
APEX-Agents task World 127_AH_Task 2 in AI Agents for Cross-Border Regulatory Review. Compare dual-harness agent runs across models, scores, and public traces.
Task prompt
What the agent was asked to do
You are analyzing the results from the customer survey. The survey asked what Helios' top 3 capabilities are. The initial results came back incomplete, and there are now additional responses available to analyze (attached). Your goal is to calculate what percentage of all total responses each capability received. Only calculate these values for respondents who responded "Slightly Important" or "Not Important" for question 2. You may also utilize the survey questions file for reference. Round final answers to two decimal places please. Send your reply here.
Published trajectories
Agent runs on this task
Curated dual-harness runs (parsed + original sandbox). Best scored run per model.
| Model | Harness | Score | Result | Links |
|---|---|---|---|---|
| GPT-5.4 | dual | 8/8 | Pass | Run detailsPublic trace |
| GPT-5.4 mini | dual | 8/8 | Pass | Run detailsPublic trace |
| GPT-5.4 nano | dual | 8/8 | Pass | Run detailsPublic trace |
| GPT-5.5 | dual | 8/8 | Pass | Run detailsPublic trace |
| Gemini 3.1 Pro | dual | 0/8 | Fail | Run detailsPublic trace |
Grading rubric
Rubric criteria
Runs are graded against these criteria. Open a run for model-specific verdicts.
States that the value for Advanced power electronics engineering is 16.29%
States that the value for Thermal management expertise is 3.71%
States that the value for High-precision manufacturing is 11.57%
States that the value for Software & controls development is 10.51%
States that the value for System integration / co-design capability is 15.62%
States that the value for Cost competitiveness is 6.46%
States that the value for Quality & reliability engineering is 16.69%
States that the value for Global manufacturing + delivery footprint is 19.16%