We at Box evaluated Gemini 3.5 Flash Lite on the Box Complex Work Eval — our benchmark of realistic, document-grounded tasks across twelve industries, benchmarked here against Gemini 3.1 Flash Lite. The tasks mirror the analytical work knowledge workers actually do: reading source documents, reconciling numbers, running due diligence, and reviewing expert output for errors. Here's what we found.
Gemini 3.5 Flash Lite scores 58% overall to the prior generation's 41% — a 17-point advantage on complex enterprise work. The gap is widest exactly where the work is hardest: multi-step, quantitative reasoning over imperfect source data. And notably, it gets there faster and with less compute — the earlier model's lower quality doesn't buy it any efficiency on this workload.
Stronger across the industries that run on numbers

Gemini 3.5 Flash Lite leads the prior generation across nearly every industry we test, with the clearest gains in the most quantitatively demanding domains:
- Consumer Products (86% vs 65%). On a client-account analysis, it segmented the accounts and identified the standout on the correct dimension, avoiding the tempting wrong one sitting beside it in the data.
- Retail (76% vs 48%). On a product-performance analysis, Gemini 3.5 Flash Lite computed each item's growth against the correct benchmark and produced a defensible ranking; the earlier model mis-weighted the inputs and returned figures that didn't reconcile.
- Financial Services (61% vs 45%). On a multi-period financial analysis, it computed a contractual exposure into the millions and assigned the correct risk rating, where the prior generation undercounted the same exposure.
- Legal (47% vs 41%). It stayed accurate across a document-heavy contract review where the earlier model's correctness slipped as the checks compounded.
The pattern is consistent: the more a task depends on getting numbers and multi-step logic right, the larger Gemini 3.5 Flash Lite's advantage.
Consistent across every type of analytical work

Sorting by type of work rather than industry sharpens the picture: Gemini 3.5 Flash Lite leads across all four modes of analytical work, and its edge grows with difficulty. It's strongest on report drafting (64% vs 46%) and expert review (62% vs 51%), and its widest margin comes on data analysis — turning imperfect source data into a number you can defend — where it scores 59% to the prior generation's 32%, a 27-point gap. On those tasks the earlier model frequently applies the wrong formula, mis-weights inputs, or returns numbers that don't tie out. It even holds its accuracy on due diligence (46% vs 33%) as the number of required findings grows. The through-line: Gemini 3.5 Flash Lite carries chained calculations to the correct result and keeps magnitudes and units right, rather than producing clean-looking but incorrect figures.
Faster and lighter, too
The efficiency result is the surprising part. Even against the prior "lite" model, Gemini 3.5 Flash Lite is the leaner one on this workload: it reaches its answers in roughly a quarter of the time (about 27 seconds per task on average vs. ~99), using about half the tokens and a third of the model and tool calls. Rather than trading quality for speed, the earlier model appears to work harder — more retrieval and reasoning turns — and still arrives at less accurate answers. For enterprise workloads at scale, Gemini 3.5 Flash Lite offers both higher accuracy and a lighter operational footprint.
Get started
On complex enterprise knowledge work, Gemini 3.5 Flash Lite is the clearly stronger and more efficient model — its advantage concentrated on the multi-step, quantitative reasoning that drives real decisions. Gemini 3.5 Flash Lite is coming soon to Box AI — reach out to your Box account team or try it in Box AI Studio.

