Claude Haiku 5.5 boosts accuracy and halves latency

|
Share

Claude Haiku 5.5 doesn’t take shortcuts when reasoning gets complicated — it catches duplicate rows, iterates circular dependencies, applies eligibility rules, and computes statistics on the right scale. And it holds that accuracy run to run, which is what determines whether you can put a model in front of recurring work. It also used about 27% fewer tokens and produced a deliverable in roughly half the time that Haiku 4.5 took.

To assess Claude Haiku 5.5, Box used the Box Complex Work Eval, a set of realistic, document-grounded tasks across the core modes of analytical knowledge work. Each task runs end to end through an agent that has to find the right documents, reason across spreadsheets, PDFs, and presentations, and produce the deliverable a knowledge worker would hand off. The model’s outputs are graded against detailed rubrics of roughly two to three dozen weighted criteria. Each task is repeated several times, because consistency across multiple runs is what really counts in enterprise work.

On this benchmark, Claude Haiku 5.5 scored 60% against 49% for Haiku 4.5 — an 11-point gain. It also got there with about 27% fewer tokens and in roughly half the wall-clock time.

Box Blog Image

Where the gains show up

Claude Haiku 5.5’s advantage grows with the amount of numerical derivation the work requires. Its largest gain is 19 points in drafting reports from source data, where the model has to choose the right basis for calculation, do the arithmetic, and hold it consistent across an entire document. Data analysis follows at 10 points. These are the tasks that punish a model for aggregating on the wrong basis, or for asserting a figure it never actually derived.

One representative case is a campaign performance report that hands the model a fifty-row dataset and asks it to rank landing pages by a composite score. Claude Haiku 5.5 produced the fully correct analysis every time. Haiku 4.5 averaged the wrong way — pooled totals instead of the mean of per-page rates — and in some runs ranked on the wrong column entirely. On a task like this the answer looks right either way, so a model that consistently gets the basis right is a big step up from one that produces a plausible-looking answer.

Box Blog Image

By industry

The gains are largest in the most quantitatively demanding domains, where a single early mistake propagates through every step that follows.

Industrial goods and automotive (61% to 85%): On an equipment overhaul cost report, the source workbook carried a summary row alongside the line items — adding up the column naively double-counts the total. Claude Haiku 5.5 recognized the duplication and excluded it. Haiku 4.5 included it and separately mis-added figures it had transcribed correctly.

Financial services (48% to 71%): Asked to produce projected financial statements for a company running on an overdraft, the model faces a circular dependency: Interest expense depends on the overdraft balance, which depends on the tax payment, which depends on profit after interest. Claude Haiku 5.5 iterated to convergence and produced statements that tie out.

Life sciences (37% to 58%): On per-site performance data for a diagnostic product, Claude Haiku 5.5’s dispersion statistics reconcile against the source. In comparison, Haiku 4.5 transcribed the site figures correctly but computed the dispersion statistics on the wrong scale — off by an order of magnitude — and every downstream threshold judgment inherited the error.

Public sector (50% to 59%): A library has a $7,500 grant and rules about what it may be spent on. Asked for the most popular patron requests that comply with those rules, Claude Haiku 5.5 screened out the ineligible options, explained why, and surfaced the next eligible items. Haiku 4.5 ranked by popularity and skipped the eligibility screen, recommending options each several times the size of the grant.

Technology (43% to 60%): In a review of user stories for an investment application, the task was to score each story against quality criteria and identify the criterion with the most failures. Claude Haiku 5.5 got the per-story adherence rates right more consistently and correctly identified Objectivity as the criterion with the most failures. Haiku 4.5 computed the rates incorrectly and, as a result, named the wrong criterion as the top failure.

What this means for enterprise work

For teams deploying models on knowledge work, Claude Haiku 5.5 is a clear step up from Haiku 4.5 on the tasks that carry real stakes — multi-step quantitative analysis, document review, and structured reporting — and it returns those answers faster. For high-volume work where speed and correctness both matter, it’s the stronger and more efficient choice.