Gemini 3.7 Flash on real enterprise knowledge work

|
Share

At Box, we evaluate frontier models the way our customers actually use them: on complex, multi-document, business tasks, judged on whether the finished work is correct and how quickly it was completed. We ran Gemini 3.7 Flash through that evaluation and compared it to Gemini 3.6 Flash. The short version: it’s meaningfully more accurate, and it gets there faster.

Our Complex Work Eval is a set of realistic, document-grounded tasks spanning a dozen industries and the core modes of analytical knowledge work: analyzing data, drafting reports from source data, running due diligence, and reviewing others' work for errors. Each task runs end to end through an agent that has to find the right documents, reason across spreadsheets, PDFs, presentations, and images, and produce the deliverable that a knowledge worker would actually hand off. Every output is graded against a detailed rubric of roughly two to three dozen weighted criteria scored by an independent judge; some criteria carry negative weight, so a confidently wrong answer is penalized rather than merely unrewarded. Each task is run several times and the scores are then averaged.

By our evaluation, Gemini 3.7 Flash scored 67% to Gemini 3.6 Flash's 62%, and won a clear majority of tasks head to head. This improvement is broad, not the product of one lucky category, and it shows up most where the work is hardest.

Box Blog Image

Where the gains show up

Gemini 3.7 Flash's advantage grows as analytical difficulty increases. Its largest gains are in due diligence - exhaustive, find-everything work - and data analysis, where the job is to compute a defensible number from imperfect source data. These are exactly the tasks that punish a model for stopping early, or for grounding an answer in the wrong cell of a spreadsheet.

One representative case: a financial due-diligence task that hands the model a full year of transaction records and asks it to validate a vendor's pricing tool and report every miscalculation. The weaker model reviewed the obvious cases, concluded the tool was error-free, and even asserted an incorrect figure of its own. Gemini 3.7 Flash worked the full set, identified the specific transactions the tool had gotten wrong, and made no unsupported claims. On a task like this, "looks thorough" and "is thorough" are different things, and the rubric rewards the latter.

The pattern repeats in structured review work: given a contract to score item by item against a policy, Gemini 3.7 Flash produced the correct per-item scores and total, where the baseline's tally was off. Getting the individual judgments and the arithmetic right, and doing so consistently, is the difference between a draft a reviewer can trust and one they have to redo.

Box Blog Image

By industry

The gains are widest in the most quantitatively and procedurally demanding domains:  Financial Services, Consumer Products, Legal, and Healthcare. These are settings where a task chains several steps together and a single early mistake propagates to the wrong conclusion, so a model that reasons carefully through the whole chain will pull ahead.

  • Financial Services saw the largest improvement, with Gemini 3.7 reaching 83% compared with 65% for Gemini 3.6 Flash — an 18-point gain. This aligns with 3.7’s strong performance on analytically demanding work like due diligence and data analysis.
  • Consumer Products achieved the highest industry score at 85%, 14 points above Gemini 3.6 Flash. The results demonstrate 3.7’s strength on work that requires synthesizing information and drawing conclusions across complex source material.
  • Legal and Healthcare each improved by 12 points, reaching 70% and 64%, respectively. These gains reinforce 3.7’s advantage on work where completeness, careful review, and evidence-backed conclusions matter.
  • Retail also saw a meaningful lift, with Gemini 3.7 scoring 77% versus 69% for Gemini 3.6 Flash. Together, these results show consistent gains across a diverse set of enterprise industries.

Faster, too

The accuracy gain doesn’t come at a compute cost. On the same runs, Gemini 3.7 Flash completed tasks in about a third less time than Gemini 3.6 Flash, roughly 73 seconds per task on average versus 115, while using slightly fewer tokens and tool calls. That combination is worth underlining: it’s both more accurate and lighter to run, which is what matters when the work is interactive or high-volume.

What this means for deployment

For teams choosing a model for knowledge work, Gemini 3.7 Flash is a clear step up from Gemini 3.6 Flash on the tasks that carry real stakes: multi-step financial analysis, due diligence, and structured review,  and it returns those answers faster. For high-stakes, open-ended judgment work, pair it with a human reviewer as you would any model. For high-volume analytical work where speed and correctness both matter, it’s the stronger and more efficient choice.

Get Started Today

Gemini 3.7 Flash will be available soon in Box.