Raycaster/ Eval

APEX-Agents · Management Consulting

Task 5k4j7555

Best published0/1Fail

APEX-Agents task Task 5k4j7555 in AI Agents for Privacy and GDPR Compliance. Compare dual-harness agent runs across models, scores, and public traces.

AI Agents for Privacy and GDPR ComplianceManagement Consulting World 129Dual harnessGrader: rubric
task_f23cb148241641f1b7c5dfbecfd3835f
Management Consulting World 129
message_in_console
5 models · dual config

Task prompt

What the agent was asked to do

Using the discount approval logs and the KPI chart, I'd like to get one number that tells me how risky our discounting behavior is right now. Looking at deals where the final approved discount exceeded policy, classify the severity using the chart, apply the risk sensitivity, and calculate the revenue exposure. Assume Policy Breach % is the difference between final approved discount and the policy threshold. Return to me a message with the Policy Breach Stress Index (rounded to 2 decimal places), which is the average revenue at risk per policy-breaching deal.

Published trajectories

Agent runs on this task

Curated dual-harness runs (parsed + original sandbox). Best scored run per model.

ModelHarnessScoreResultLinks
Gemini 3.1 Produal0/1Fail
GPT-5.4dual0/1Fail
GPT-5.4 minidual0/1Fail
GPT-5.4 nanodual0/1Fail
GPT-5.5dual0/1Fail

Grading rubric

Rubric criteria

Runs are graded against these criteria. Open a run for model-specific verdicts.

  1. States that the Policy Breach Stress Index is 10906.97