Claude Opus 5.5: Faster and leaner for high quality work

|
Share

At Box, we evaluate frontier models the way our customers actually use them: on complex, multi-document business tasks. Each task requires an agent to find the right documents, reason across spreadsheets, PDFs, presentations, and images, and produce a deliverable that a knowledge worker would actually hand off. The model is judged on whether the finished work is correct and on how quickly and efficiently it is completed.

Compared to Opus 5, the standout finding is efficiency; Claude Opus 5.5 is significantly faster and more token-efficient without compromising on accuracy. It can help customers get from source content to a usable business deliverable quicker and with less computational overhead.

Box Blog Image

Meaningful gains in efficiency

Claude Opus 5.5 completed tasks 30% faster than Opus 5. Latency compounds in agentic work, from finding source files and analyzing information across them to validating calculations and producing a final deliverable. Faster completion makes it more practical to use Box AI for enterprise work that requires multiple rounds of reasoning and review.

Opus 5.5 also used approximately one-third as many tokens as Opus 5, while producing deliverables that were 42% shorter – all without compromising accuracy. Its shorter outputs retained key findings while reducing unnecessary restatement and preamble. 

industry

Standout industries 

Claude Opus 5.5’s efficiency gains didn’t come at the expense of accuracy anywhere in our evaluation set. Below are representative tasks from three of the industries where Opus 5.5 impressed us most. Each pair of figures compares Opus 5’s average accuracy for that industry with Opus 5.5’s performance. 

Consumer products (76% → 82%) On a client account analysis — setting onboarding targets by reading across a signed contract, a satisfaction tracker and a team metrics sheet — Opus 5.5 held an impressive score on every run, while Opus 5 was more variable. In enterprise settings, where consistency is key, the new model changes the workflow; its predecessor still needs outputs checked.

Financial services (73% → 75%) Handed a year of transaction records and asked to validate an acquisition target's pricing tool, Opus 5.5's output was perfect on every attempt: every miscalculation found, nothing asserted that it couldn’t support, and all in under half the words that Opus 5 needed to make the same points.

Technology (63% → 74%) Opus 5.5 improved on every technology task in the eval. On a cloud cost-optimization analysis, it picked the right basis for the retention calculation and held the source data's unit conventions consistent all the way through, so the number at the end holds up. And it used 27% fewer words to get there.

What this means for customers

Enterprises apply Box AI to large volumes of content, where speed, efficiency, and output quality determine how much complex work can be delegated to an agent. Claude Opus 5.5 completed tasks 30% faster than Opus 5, using approximately one-third as many tokens and producing deliverables about 40% shorter while preserving the key findings.

For Box customers, the value of these improvements issues from the work they need to complete. For finance teams, faster completion and lower token usage can shorten the path from complex source documents to analysis, comparisons, and investment briefs. For consumer products teams, they can help turn market, customer, and product data into timely recommendations and stakeholder-ready briefs. In both cases, more concise outputs preserve the key findings while reducing the time spent reviewing and editing the result.

Across these workflows, Claude Opus 5.5 helps Box AI move from governed enterprise content to clearer, more actionable deliverables while maintaining accuracy.

Getting started

Claude Opus 5.5 will be coming to Box AI soon.