At Box, we evaluate models on complex, multi-document business tasks that mirror the workflows our customers run every day. Each task runs on our agent harness, requiring the model to find the right documents, reason across spreadsheets, PDFs, presentations, and images, and produce a finished deliverable that a knowledge worker could hand off: a cost analysis, a diligence findings list, a verification audit.
OpenAI's GPT 6 Astra achieved frontier performance on our evaluation. Its overall accuracy was 77%, benchmarked here against GPT 5.6 Sol's score of 74%.

Higher accuracy on the most challenging workflows
GPT-6 Astra showed particular strength in the most demanding industry workflows, where the documents disagree with each other, the deliverable requires a figure nobody handed you, and you have to reconcile quantitative data with narrative, contractual, or domain-specific source material.
Here's how this manifested in four tasks from our benchmark:
- Media and entertainment (48% to 100% task accuracy). Ranking film genres and countries by profitability across a year of production data from four teams, where the trick is to apply the following year's tax-incentive corrections without double-counting them. GPT-6 Astra was perfect on every attempt. GPT-5.6 Sol had the rankings right but the underlying ratios were wrong.
- Technology (69% to 97% task accuracy). Choosing which region to fund first, from a stack of performance and infrastructure documents that never state the metric the brief asks for. GPT-6 Astra flagged the gap, labelled its own figure a proxy, and caught a growth claim that didn't match the numbers beneath it. GPT-5.6 Sol reported the proxy as the real thing.
- Legal (69% to93% task accuracy). Reviewing an NDA against a company's own contracting policy, where reaching the right verdict isn't enough — the answer has to cite the provision behind it. Both models declined to approve the draft, but only GPT-6 Astra separated whether the liability cap's structure was permissible from whether its amount was defensible, and pointed to the policy language that settles it. GPT-5.6 Sol argued the same conclusion without citing the provision, which is a critical error for legal review.
- Energy (82% to 97% task accuracy). Building a consumption report on two facilities from a year of meter logs, where only one of the two sites is actually missing data. GPT-6 Astra flagged this and drew a line most reviewers wouldn't: the other site's odd solar readings are a measurement problem to investigate, not a gap in the record. GPT-5.6 Sol reported both sites as incomplete.
What this means for customers
For Box customers, these are not just improvements on a benchmark. They are the difference between an output that looks plausible and a deliverable that teams can confidently act on. GPT-6 Astra empowers Box AI Agents to distinguish a proxy from a verified metric, trace a legal conclusion to the governing policy, and separate a data-quality anomaly from genuinely missing information. This means that teams spend less time checking the model’s work and face less risk of carrying an early mistake into a high-impact decision. GPT-6 Astra makes it easier to automate demanding workflows at scale, turning your trusted enterprise content into polished, decision-ready deliverables.
GPT-6 Astra will be available soon in Box AI Studio. Learn more about Box AI and explore how to build with the Box AI API.


