SpreadsheetBench
Gemini 3.1 Pro on 84-40
Prompt excerpt
Task context
You are solving a spreadsheet benchmark task in a real workbook. Objective: Produce the correct final workbook state for the expected answer region. What matters: - Only the values in the expected answer region will be graded. - The workbook is the answer. Instructions: 1. Read the workbook and inspect the relevant data region first. 2. Infer the required result for the provided workbook instance. 3. Write the final value(s) directly into the expected answer region. 4. Do not rely on prose, formulas in your chat response, pseudocode, or VBA as the answer unless the benchmark explicitly requires those to be written into cells. 5. If the natural-language task asks for a general method, formula, or macro, convert that into the concrete result needed for this workbook instance. 6. Keep your final text response short and only summarize the workbook cells you changed. Relevant data region(s): CA!'A1:D16,'MN!'A1:D16,'SDFRT!'A1:D6,'LIST!'A1:E6 Expected answer region(s): 'LIST'!C2:E5 Expected answer sheet(s): LIST Task: I need to sum the amounts in the entire columns C and D for each sheet, excluding column B, in an Excel workbook. These sums should then be populated in a 'list' sheet, putting the names of the sheets into column B. For each sheet, column C in list tab should contain the total from the original sheet's column C, and column D in list tab should contain the total from the original sheet's column D. Additionally, column E in the 'list' sheet should show the result of subtracting the amount in column D from the amount in column C. I also want a total row at the end that sums the amounts from all these sheets. Whenever I add new sheets before the 'list' sheet, I need the existing data in the 'list' sheet to be cleared so that the list can be updated to reflect any changes in the other sheets. I've provided an attachment for reference.
Response trace
Agent response, tools, files, and edits
On a phone, the interactive viewer works best full-screen — pick the narrative report or the files & trajectory workspace.
Why Eval exists · why Workspace exists
Public evidence and cloud agents are the same harness.
Eval exists so scores are inspectable—tasks, trajectories, artifacts, and rubric verdicts anyone can open.Workspace exists so people can automate real file work with that harness, and so Raycaster never evaluates work it cannot perform.