Pricing & ROI9 minOctober 5, 2026

Basis Cuts Tax Workbook Time in Half with GPT-6 Astra

How Basis used GPT-6 Astra to complete a 50-tab tax workbook 2x faster — architecture, reasoning controls, ROI, and a step-by-step guide for finance teams.

Basis Cuts Tax Workbook Time in Half with GPT-6 Astra

The 50-Tab Problem That Rewrote the Rules for Accounting Automation

A 50-tab tax workbook. Completed in half the time. Not by a larger team, not by a faster accountant — by an AI agent running on GPT-6 Astra. That single benchmark result, published by OpenAI on September 28, 2026, is the clearest signal yet that the economics of tax preparation are being redrawn from the ground up.

What makes this case genuinely worth studying isn't the headline number — it's the architecture behind it, the specific decisions Basis made about reasoning effort, token efficiency, and evaluation design, and what those decisions imply for any finance or accounting team considering a similar move. The details are more instructive than the result, and they're coming up.

Basis builds AI agents to automate the manual work accountants do every day — reconciliations, workbook completion, primary-source lookups — so that human professionals can redirect their attention toward judgment and advisory work. The company, co-founded by Mitch Troyanovsky and Matt Harpe and founded in 2023, has been building on OpenAI's models since day one. When GPT-6 Astra became available, Basis ran a direct head-to-head comparison against GPT-5.6 Sol on one of the most demanding tasks in their pipeline: completing a complex, multi-tab tax workbook end to end.

The result was unambiguous. GPT-6 Astra completed the 50-tab workbook in half the time that GPT-5.6 Sol required — and it did so while also improving Basis's internal evaluation scores by approximately 20%.


What Basis Actually Built: The Architecture Behind the 2x Result

The speed gain didn't come from simply swapping one model for another and pressing run. It came from a deliberate architecture that exploits specific capabilities of GPT-6 Astra — capabilities that earlier models either lacked or handled poorly.

Better Decisions at the Start of a Task

One of the most consequential differences Basis observed was how GPT-6 Astra approaches the beginning of a long task. According to Troyanovsky, the model makes better decisions upfront, which allows Basis's agents to take a more direct path through the work with less time spent correcting mistakes mid-stream. This matters enormously in a 50-tab workbook: an early misjudgment about how to structure the approach compounds across dozens of subsequent steps.

GPT-6 Astra is also more likely to ask a focused clarifying question when the answer could change the outcome — rather than making an assumption and proceeding. For a domain like tax preparation, where a single incorrect assumption about entity structure or filing status can cascade through an entire return, that behavior is worth more than raw speed.

The model that asks the right question at step two saves you from rebuilding everything at step forty.

Dynamic Reasoning Effort: Spending Compute Where It Counts

The second architectural lever Basis uses is dynamic reasoning effort. Rather than running GPT-6 Astra at a fixed reasoning level throughout a task, Basis configures the agent to dial reasoning up when a step is genuinely difficult — a complex cross-tab reconciliation, an ambiguous primary-source lookup — and dial it back down for simpler, more mechanical steps.

Critically, GPT-6 Astra can make these adjustments while keeping its cache intact. That means the model doesn't lose context when it shifts reasoning modes, which is a practical requirement for any long-running agentic workflow. Troyanovsky noted that this approach reduces both cost and response time, making long-running tasks more economical for Basis and its customers.

This is the kind of optimization that separates a production-grade agentic system from a demo. In several evaluations, GPT-6 Astra achieved stronger results while using substantially fewer output tokens than earlier models — delivering a lower estimated cost per task despite its higher per-token pricing. Basis's dynamic reasoning configuration amplifies that efficiency further.

Evaluation Architecture: Grading the Process, Not Just the Answer

Basis evaluates not only whether its agents produce the correct final output, but how they get there — whether they follow established templates, consult primary sources for tax questions, and produce outputs that are ready for human review without requiring significant rework. This distinction matters: the same final number can come from a sound process or a brittle one, and only the former is trustworthy at scale.

This is what Troyanovsky means when he describes the product as selling consistency and reliability, not just accuracy. Accounting firms deploying Basis agents need to trust that the process is sound every time — not just that the answer happens to be right.

For teams exploring similar architectures, the RAG vs CAG: Which AI Architecture Fits Your Business breakdown is a useful reference for thinking through how retrieval and context design affect reliability in long-horizon tasks.


Step-by-Step: How to Replicate This Approach in Your Finance Team

The Basis case is instructive precisely because its architecture is transferable. You don't need to be an AI company to apply these principles. Here's how a finance or accounting team can implement a similar approach.

Step 1 — Define the Task Boundary Precisely

Before touching a model, map the exact scope of the workflow you want to automate. For Basis, this was a 1065 partnership tax workbook: a specific document type, a defined set of inputs (trial balances, K-1s, primary tax sources), and a clear output format. Vague task definitions produce vague agent behavior.

Write down: what are the inputs, what is the output, what are the quality criteria, and what does "done" mean? If you can't answer all four, the agent can't either.

Step 2 — Choose the Right Model for the Complexity Level

Not every task in a finance workflow requires GPT-6 Astra's full capability. Use the model hierarchy deliberately:

  • GPT-6 Astra — for the hardest reasoning work: multi-tab reconciliations, ambiguous primary-source lookups, tasks where an early error compounds downstream.
  • GPT-6 Sol — for demanding but more structured tasks where the reasoning path is clearer.
  • GPT-6 Luna — for high-volume, repeatable steps where efficiency matters more than maximum intelligence.

Routing tasks to the right model is itself a cost-control mechanism. Basis's architecture does this dynamically within a single workflow; you can start by doing it statically across different workflow types.

Step 3 — Configure Dynamic Reasoning Effort

Within a single long-running task, configure the agent to vary its reasoning effort by step type. In practice, this means:

  1. Identify which steps in your workflow are genuinely hard (require cross-referencing multiple sources, involve ambiguity, or have high downstream impact if wrong).
  2. Identify which steps are mechanical (formatting, copying values between tabs, applying a known formula).
  3. Set reasoning effort to high or extra-high for the former, and to a lower setting for the latter.

GPT-6 Astra supports this without losing cache context between steps — a technical requirement that makes the approach viable for workflows that run for minutes or hours rather than seconds.

Step 4 — Build a Process Evaluation Layer, Not Just an Output Check

This is the step most teams skip, and it's the one that determines whether your automation is trustworthy at scale. Build evaluation criteria that assess the agent's process:

  • Did it consult the correct primary sources?
  • Did it follow the established template structure?
  • Did it flag ambiguities rather than resolve them silently?
  • Is the output formatted for efficient human review?

Automated judges — smaller models or rule-based checks — can handle much of this evaluation without human review of every run. Basis uses this approach to maintain confidence in real-world deployments.

Step 5 — Run a Controlled Benchmark Before Full Deployment

Before replacing any human workflow, run the agent on a representative sample of historical tasks where you already know the correct output. Measure accuracy, process adherence, and time-to-completion. Compare against your current baseline.

This is what Basis did with the 50-tab workbook comparison: a controlled test with a defined task, two models, and a clear metric. The result gave them confidence to deploy. It also gave them a number they could show to customers — which is its own form of value.

Step 6 — Design the Human-in-the-Loop Handoff

Agentic automation works best when the human's role shifts from doing the work to reviewing and approving it. Design the agent's output explicitly for that review step: clear structure, flagged uncertainties, source citations for any judgment calls. The goal is to make the reviewer's job faster and more confident, not to eliminate the reviewer.


ROI Framework: What the Numbers Actually Mean for Your Business

The 2x speed result is the headline, but the ROI calculation for a finance team has several components worth separating.

Time Savings and Capacity Expansion

If a tax workbook that previously took a senior accountant eight hours now takes four — with the agent handling the mechanical and research-intensive steps — that's four hours of senior capacity freed per workbook. Across a busy tax season, that compounds quickly. A team that previously handled 50 complex returns in a season can now handle significantly more without adding headcount.

The more precise framing: automation doesn't reduce your team, it expands what your existing team can take on. That's the capacity argument, and it's the one that resonates with finance leaders managing headcount constraints.

Token Efficiency and Direct Cost

GPT-6 Astra's higher per-token price is real. But Basis's experience — and OpenAI's own evaluation data — shows that the model completes tasks using substantially fewer output tokens than earlier models, because it makes better decisions earlier and corrects fewer mistakes mid-task. The net cost per completed workbook can be lower than with a cheaper model that takes longer and makes more errors.

Dynamic reasoning effort amplifies this further: by using less compute on mechanical steps, Basis reduces the token cost of the easy parts of a workflow without sacrificing quality on the hard parts.

For a structured approach to quantifying these trade-offs before committing to a deployment, the AI ROI Framework: Prove Business Value guide offers a practical methodology.

The Evaluation Investment

Building a process evaluation layer — the judges, the template checks, the source-citation verification — takes upfront engineering time. Basis has been building this infrastructure since 2023. For a team starting from scratch, budget for this explicitly. It's not optional if you want deployments you can trust.

The evaluation layer is what separates a tool you demo from a system you stake your reputation on.

The payback period depends on workflow volume. For a firm processing hundreds of complex returns per year, the investment in evaluation infrastructure pays back within a single tax season. For smaller volumes, a lighter-weight approach — manual spot-checks on a sample of agent outputs — may be sufficient initially.


What This Means for Finance and Accounting Leaders

Basis is, as of this writing, a company deploying autonomous agents that handle multi-day tax and reconciliation workflows inside real accounting firms. The 50-tab workbook result isn't a research demo — it's a production benchmark from a system that firms are already using.

The pattern Basis has established — agents that handle the execution, humans who handle the judgment — is the architecture that scales. It doesn't require replacing your team. It requires redesigning what your team does: from executing workflows to supervising, reviewing, and advising.

For accounting and finance leaders, the practical question isn't whether this is possible. The Basis case settles that. The question is how quickly your firm can build the evaluation infrastructure, the task definitions, and the human-review workflows that make agentic automation trustworthy in your specific context.

The leaders who move first on this won't just cut costs — they'll expand capacity, take on more complex work, and build a structural advantage that compounds with every model improvement. That's the kind of decision that changes how a board reads your quarterly review: not as a cost center managing headcount, but as an operation that scales intelligently.

And on a more personal level: there's a specific kind of calm that comes from knowing your most critical, high-stakes workflows are running reliably — not because you're watching them constantly, but because the system is designed to catch its own errors and surface them for your review. That's what Basis is selling. That's what GPT-6 Astra makes possible.

For teams exploring how GPT-6 Astra is being adopted across other high-stakes domains, the GPT-6 Astra Early Adopters: What They Share overview maps the common patterns across industries.


FAQ

What exactly did Basis test with GPT-6 Astra? Basis compared GPT-6 Astra and GPT-5.6 Sol on a complex tax workbook with 50 tabs, measuring time to completion and accuracy. GPT-6 Astra completed the workbook in half the time, and Basis's internal evaluation scores improved by approximately 20%.

Why does GPT-6 Astra use fewer tokens if it's more capable? Because it makes better decisions at the start of a task, taking a more direct path through the work with fewer mid-task corrections. Fewer errors mean fewer recovery steps, which means fewer tokens consumed overall — even though the per-token price is higher than earlier models.

What is dynamic reasoning effort and why does it matter for cost? Dynamic reasoning effort means the model applies more compute to difficult steps and less to simple ones within the same task. Basis uses this to reduce cost and response time on long-running workflows without sacrificing quality on the steps that require it. GPT-6 Astra supports this while keeping its cache intact, which is essential for multi-step agentic tasks.

Do I need to build a full evaluation system before deploying AI agents in finance? Not necessarily at full scale from day one. Start with manual spot-checks on a sample of agent outputs against known-correct historical results. As volume grows and you gain confidence in the agent's behavior, invest in automated evaluation — judges, template checks, source-citation verification. The evaluation layer is what makes the system trustworthy at scale.

Is this approach only viable for large accounting firms? No. The architecture — precise task definition, model selection by complexity, dynamic reasoning effort, process evaluation, human-review handoff — scales down to smaller teams. The upfront investment in evaluation infrastructure is the main variable; smaller teams can start lighter and build over time.

How does Basis handle cases where the agent is uncertain? GPT-6 Astra is designed to ask focused clarifying questions when the answer could change the outcome, rather than making silent assumptions. Basis's evaluation layer also flags ambiguities in agent outputs for human review, so uncertainty surfaces explicitly rather than being buried in a final answer.


The firms that treat this as a technology experiment will get experiment-sized results. The ones that treat it as an operational redesign — redefining what their people do, building the infrastructure to make agents trustworthy, and measuring the right things — will get the 2x result Basis got. The architecture exists. The benchmark is published. The only remaining question is who builds it first.

Have questions? Ask the AI agent right now

Responds in seconds, knows everything about our services and will help with your situation