OpenAI says Basis used GPT-6 Astra to complete a 50-tab tax workbook twice as fast as GPT-5.6 Sol, raising stakes for AI finance tools.

Basis completed a 50-tab tax workbook twice as fast with OpenAI’s GPT-6 Astra as it did with GPT-5.6 Sol, according to an OpenAI case study published in OpenAI News. The result is an early signal that model improvements in complex financial workflows may be measured not only by answer quality, but also by how quickly an AI system can complete a multi-step professional task.
The announcement is narrowly focused: it describes one workbook and a comparison between two OpenAI models. OpenAI also says Astra’s stronger understanding of user intent gives Basis greater confidence in using the model in real-world work. The company did not provide the full article text, testing methodology, workbook completion times, or independent validation in the available source material.
The central claim concerns a 50-tab tax workbook, a format that typically requires an AI system to work across many related sheets rather than respond to a single isolated prompt. OpenAI says GPT-6 Astra completed the workbook in half the time required by GPT-5.6 Sol.
That comparison matters because spreadsheet-based tax work can combine calculations, references between tabs, interpretation of instructions, and consistency checks. The available evidence does not specify which of those steps Astra handled, how much human supervision was involved, or whether “completed” means the workbook was fully prepared, reviewed, or simply processed through a defined benchmark.
Still, the result points to a practical performance dimension for enterprise AI. A model that can move through a large financial artifact more quickly may reduce waiting time for analysts and make it easier to include AI in workflows with deadlines or high volumes of repetitive work.
The strongest evidence comes from OpenAI’s own announcement, titled “Basis completes a tax workbook 2x faster with GPT-6 Astra.” The related OpenAI listing repeats the claim that Astra completed the workbook twice as fast as GPT-5.6 Sol and attributes greater real-world confidence to the model’s understanding of user intent.
Those are vendor-reported performance and usability claims, not an independent benchmark. The available source information does not identify the workbook’s contents, the evaluation criteria, the hardware or software configuration, the number of attempts, or whether the comparison controlled for context length, prompting, tool access, and human intervention.
That missing detail limits what buyers and researchers can conclude. A two-times speed improvement could reflect model efficiency, a different execution path, improved task planning, changes in tool use, or the specific characteristics of this workbook. It should not yet be treated as a general performance result for all tax, accounting, or spreadsheet tasks.
The source also provides no independent customer testimony or adoption figures. Basis is identified as the organization completing the workbook, but the material available here does not establish how broadly it uses Astra, whether the workflow has reached production, or what review procedures remain in place.
For AI product teams, the announcement highlights the gap between a model that can generate plausible spreadsheet content and one that can reliably follow the user’s intended objective across a large file. OpenAI’s emphasis on intent understanding suggests that the practical challenge is not limited to arithmetic. The system must also interpret what the user is asking it to change, preserve relevant structure, and avoid introducing errors elsewhere in the workbook.
That distinction is important for financial applications. Speed has value only if the output remains auditable and accurate. A faster model could increase throughput, but it could also amplify the cost of undetected errors if users reduce review time because the task finishes sooner. For tax and accounting products, model evaluation therefore needs to include traceability, exception handling, cell-level validation, and clear escalation when instructions are ambiguous.
The result may also influence how enterprise buyers compare models. Instead of looking only at general-purpose benchmarks, buyers may ask vendors to demonstrate performance on representative workbooks, including linked tabs, changing assumptions, formatting constraints, and review checkpoints. The relevant question is not simply which model is faster, but whether it can complete the workflow with acceptable reliability and oversight.
For founders building AI finance tools, the case study offers a possible product direction: benchmark the entire workflow rather than a narrow generation step. That means measuring time to a reviewable result, the number of corrections required, and how often the system asks for clarification before making a consequential change.
The first signal to watch is whether OpenAI publishes additional technical detail about the GPT-6 Astra comparison. Information about the workbook, task definition, execution environment, human review, and error rate would make the two-times claim more useful to developers and buyers.
A second signal is whether Basis or other finance companies report similar results on different tax workbooks. Repeated performance across varied files would provide stronger evidence that the result reflects a durable model improvement rather than a single favorable task.
Third, enterprise teams should look for evidence about accuracy and auditability alongside speed. Useful follow-up measures would include correction rates, preservation of spreadsheet logic, handling of ambiguous instructions, and the ability to explain changes made across multiple tabs.
Finally, model comparisons may become more consequential as financial software vendors embed AI directly into spreadsheet and tax workflows. If Astra’s intent-following advantage holds beyond this example, competition could shift toward models that can manage complete business processes rather than answer individual questions.
OpenAI’s Basis example is notable because it frames model progress around a recognizable professional deliverable: a 50-tab tax workbook. That is more relevant to enterprise buyers than an abstract claim about capability, but the evidence remains too limited to establish a broad advantage. The result is one vendor-reported comparison, without the methodology needed to assess accuracy, supervision, or reproducibility.
The practical takeaway for builders is to treat speed as one part of a larger evaluation. GPT-6 Astra may have completed this task faster than GPT-5.6 Sol, but finance deployments will depend on whether the model can produce outputs that people can verify, correct, and safely use. Until more evidence is available, the strongest conclusion is that OpenAI is positioning Astra for complex, intent-sensitive work—and that real-world workflow benchmarks are becoming an important battleground for enterprise AI.