Claude Sonnet 5.5 on real enterprise knowledge work

|
Share

At Box, we test frontier models against the work our customers actually do: complex business workflows that span many documents. In each task, an agent has to locate the right files, work across spreadsheets, PDFs, slides, and images, and deliver something a knowledge worker could hand off to a colleague. We evaluate the model both on the quality of the finished deliverable and on how efficiently it gets there.

We ran Anthropic's Claude Sonnet 5.5 through this evaluation, benchmarked against Sonnet 5, and it improved on both fronts. Sonnet 5.5 reached 65% overall accuracy against Sonnet 5's 61%. It also reached a finished deliverable roughly 2.4× faster, on 12% lower total token consumption. Speed and accuracy usually pull against each other in agentic work, so a model that can deliver both is a real step forward. 

Claude sonnet 5.5

Gains across industries

Sonnet 5.5 pulls ahead of Sonnet 5 in the most quantitatively demanding, detail-oriented corners of knowledge work — the moments when a deliverable is judged on whether a specific number is right, whether a specific protection is present, and whether the conclusion actually follows from the evidence in front of you. These are domains where precision and judgment both matter, and where being confidently wrong has real consequences. Below are some of Sonnet 5.5's strongest industries, along with representative examples of tasks where it improved over Sonnet 5.

Financial services (63% to 81%): In a due-diligence review of a trading-tool acquisition, the attached deal book's arithmetic didn't hold up. Sonnet 5.5 identified the unsecured deposits whose total interest had been miscalculated and the options that had been mispriced, whereas Sonnet 5 reported the book's stated figures as correct. Sonnet 5.5 also finished this work in 61% less time.

Legal (60% to 67%): When asked to assess renewal terms and early termination notice periods for a commercial lease, Sonnet 5.5 declined to invent a "standard" market benchmark to measure the lease against, and instead flagged what the lease was genuinely missing.

Life sciences (54% to 62%): In a task that required identifying targets across a protein interaction network, Sonnet 5.5 computed network topology that Sonnet 5 could not, and it got there with 64% lower latency.

Public sector (55% to 62%): When reporting on a math intervention program, Sonnet 5.5 succeeded in pulling the right student progress-monitoring figures from the underlying program data. Not only was Sonnet 5.5's report more accurate and complete, it also completed the task in 71% less time than Sonnet 5.

What this means for customers

Sonnet 5.5 is better at finding and cross-checking facts across a set of documents, and more disciplined about the boundary between what the documents say and what it assumes. When a source document contains an error, Sonnet 5.5 catches it, rather than carrying it forward. This translates directly into the work our customers delegate to Box AI. It means that finance teams can put an agent on a deal book, a loan tape, or a set of quarterly figures and get back analysis where the arithmetic has been re-checked rather than restated. Legal teams can run contract and lease reviews at volume and get a read on what a document is missing, not just a summary of what it contains. And because Sonnet 5.5 reaches the finished deliverable roughly 2.4x faster, this work can run in the flow of someone's day. In document-heavy knowledge work, accuracy is what makes an agent trustworthy enough to use, and efficiency is what makes it worth using. 

Getting started 

Sonnet 5.5 is coming soon to Box AI.