Muse Spark 1.3 comes to Box AI — and it does the hard work faster

|
Share

We're excited to announce that Muse Spark 1.3 from Meta Superintelligence Labs is now available in Box AI, and, among similarly priced models, it lands at the front of the pack. Benchmarked against a composite of the leading models at comparable price points, Muse Spark 1.3 pulls ahead on accuracy by 13%. That result comes from the Box Complex Work Eval, our benchmark for the kind of work enterprises actually need agents to do: pull source files, reason across spreadsheets, PDFs, contracts and decks, and produce a finished deliverable.

Muse Spark 1.3 is more efficient for complex work

The pattern we usually see when a model family advances by a generation is that accuracy gets paid for with effort — more reasoning turns, more tool calls, more tokens, more waiting. Muse Spark 1.3 breaks that pattern. Head to head against Muse Spark 1.2, Muse Spark 1.3 maintains its accuracy while finishing the same work about 42% faster and using roughly a third fewer tokens. For an agent doing real enterprise work, that’s a headline, not a footnote. Complex enterprise deliverables are where agentic latency and cost actually bite, and cutting latency by nearly half while retaining quality is the kind of generational improvement that our users care about.

Box Blog Image

Stronger performance across industries

Compared with the previous generation, Muse Spark 1.3 shows major gains in the industries that make heavy demands on quantitative reasoning and source discipline. Here are representative tasks from three industries where Muse Spark 1.3 shone:

  • Financial services (66% to 75%) On a five-year projection task, the cost driver for future years is supplied only inside an image-based assumptions table, formatted differently from everything above it. Muse Spark 1.3 finds it, uses it, and states the basis out loud, then adds a sensitivity note that tests the alternative rather than quietly substituting for the decision. On a separate multi-year forecasting task, Muse Spark 1.3 picks the correct tax treatment and carries it consistently through to the funding gap, where the wrong treatment inverts the answer’s sign entirely — a deficit becoming a surplus. That kind of discipline is the difference between a model that just produces a number and a model that produces a number you can act on.
  • Technology (65% to 69%) Given a best-practices manual and a set of work items to score against it, Muse Spark 1.3 calibrates against the manual's own worked examples rather than applying the rule mechanically. It correctly reads a delivery channel as a channel rather than a prescribed technical solution — and in the other direction, catches an item whose acceptance criteria contradict themselves in several places and marks it down accordingly. Judgement in both directions, from the same source material.
  • Energy (63% to 69%) In an analysis of consumption across two facilities, one site has only about a month of logged data split across two partial calendar months. The figures are only comparable if you normalize to energy per day, and the peak week has to be summed from real calendar weeks, not pro-rated from a month. Muse Spark 1.3 produced every required figure on a per-day basis, showed the underlying total and the day count for each, and stated the rule it was applying. It's a small analytical step that a lot of models skip — which would make every number in the report wrong.

What this means for Box customers

Across our evaluations, the same capability keeps showing up: Muse Spark 1.3 is disciplined about the source material from which it works. It finds the assumption that's buried in an image. It commits to one basis instead of hedging between two. It normalizes before it compares. It notices when a reasonable-looking data treatment would quietly make the deliverable impossible, and says so instead of returning a table full of N/As.

That perceptive precision is what enterprise document work rewards, and it's why we're glad to be bringing Muse Spark 1.3 to Box AI. Combined with the efficiency profile, Muse Spark 1.3 is a model we can put behind real deliverables — the analysis that goes to a board, the report that goes to a regulator, the memo that drives a decision — and have it come back fast enough to be part of someone's actual workflow.

Muse Spark 1.3 from Meta Superintelligence Labs is available now in Beta in Box AI.