At Box, we evaluate frontier models the way our customers actually use them: on complex, multi-document business tasks. Each task runs through an agent, requiring it to find the right documents, reason across spreadsheets, PDFs, presentations, and images, and produce a deliverable that a knowledge worker would actually hand off. The model is judged on whether the finished work is correct and on how quickly and efficiently it is completed.
In our evaluation, we saw that Claude Fable 5.1 scored 72% compared to Fable 5’s 65%: a nearly 7 percentage points jump. By delivering more accurate results with lower latency and fewer tokens, Fable 5.1 can help customers run complex enterprise workflows more cost-effectively and get more value out of their enterprise content workflows.

Meaningful improvements across demanding use cases
Enterprise work rarely happens in a single document or a single prompt; it requires models to reconcile information across file types, perform multi-step analysis, apply domain-specific logic, and turn the result into a finished deliverable. Claude Fable 5.1's biggest gains appear in precisely this kind of analytical work — tasks where one incorrect assumption can cascade through every conclusion that follows.
- Data Analysis (56% to 69%). These tasks require the model to interpret complex spreadsheets, apply formulas or weighting schemes, and produce correct quantitative outputs. Claude Fable 5.1 maintained accuracy across the full reasoning chain, where an early error can otherwise distort every downstream result.
- Report Drafting from Data (70% to 76%). When synthesizing multiple data sources into a finished report, Claude Fable 5.1 captured more of the relevant detail and handled intermediate calculations more accurately than Claude Fable 5.
- Expert Review and Verification (60% to 66%). When asked to validate claims against source data, Claude Fable 5.1 identified more inconsistencies and errors that Claude Fable 5 overlooked.
Higher accuracy with lower latency
Claude Fable 5.1's accuracy gains don't come at the expense of speed. Average task-completion time saw a 23% reduction, while total token usage declined by 25%. Fable 5.1 reaches the answer in fewer iterations — it's both more accurate and more efficient.

Improvements across multiple industries
The accuracy and speed improvements translate directly into the workflows enterprise teams run every day. Here's what we saw across a selection of industries represented in the evaluation.
- Financial Services (63% to 82%): On a financial-projection task, Fable 5.1 correctly applied a capital allowance before computing tax and carried that result cleanly through to retained earnings, overdraft, and funding gap — a chain where Fable 5 made an early calculation error that distorted every downstream number.
- Public Sector (58% to 74%): On a task recomputing weighted student grades under a revised scheme, Fable 5.1 correctly inferred which assignments belonged to which weight categories from contextual clues: a distinction Fable 5 missed, cascading into incorrect grades for most students.
- Technology (36% to 74%): Claude Fable 5.1 produced correctly formatted, substantive answers on every technology test case. It also delivered stronger analytical results, such as correctly handling an ambiguously normalized metric in a cost-optimization task.
What this means for customers
For Box customers, Claude Fable 5.1 makes it more practical to use AI for difficult, end-to-end enterprise workflows—not simply answering questions about a document, but analyzing data, reconciling information across multiple files, validating others’ work, and producing reports that are ready to hand off. Its higher accuracy reduces the risk that an early mistake will undermine an entire forecast or analysis, which can mean less manual verification and rework before teams act on the result. At the same time, faster completion and lower token usage help organizations run these workflows more efficiently. The result is AI that can move complex work from source content to a reliable business outcome faster, whether teams are preparing financial projections, reviewing high-stakes information, or analyzing operational data at scale.
Get started with Claude Fable 5.1 in Box AI
Claude Fable 5.1 delivers more accurate analytical work while completing complex tasks faster and with fewer tokens than Claude Fable 5. Claude Fable 5.1 will be available soon in Box AI Studio. Learn more about Box AI and explore how to build with the Box AI API.


