AI Tax Preparation Software: What Should a CPA Firm Expect?
What AI tax preparation software genuinely does for a CPA firm, the compliance questions to ask first, and how to size the return before busy season.

What matters most
- The measurable win for CPA firms is document intake and extraction, not preparation of the return itself.
- IRC section 7216, the FTC Safeguards Rule and IRS Publication 4557 shape what deployment is acceptable, so settle data location first.
- Every extracted figure should cite the source document and page, with corrections logged.
- Determinative tax treatments need consistent inspectable reasoning, which points to fixed rules and preparer sign-off.
- Size the return on January-to-April capacity, because the binding constraint is hours inside that window.
Two quite different things get sold as AI tax preparation software, and confusing them costs firms real money. One reads the documents your clients send and gets the data into your tax software. The other claims to prepare or review the return itself. The first is mature, measurable and worth buying. The second needs much harder questions.
For most firms the bottleneck was never the preparation. It is the three weeks in February spent chasing documents that did not arrive and typing numbers off photographs of W-2s.
Here is what matters most:
- The real win is document intake: reading what clients send and landing the data in your tax software.
- Ask where client data sits before anything else. IRC section 7216 governs disclosure of taxpayer information to third parties.
- Anything touching the return itself needs an inspectable, consistent trail, which points to fixed rules and preparer sign-off.
- Measure the exception rate, meaning how often a preparer must correct an extracted figure.
- Size it against February hours, not annual averages. The constraint is capacity in an eight-week window.
Where the hours actually go
Time a January-to-April engagement honestly and the preparation is rarely the biggest block. The blocks are these.
Chasing. The client sent eleven of fourteen documents and nobody has followed up since Tuesday. This is unglamorous, gets deprioritised by busy people, and is why engagements sit still for days.
Typing. Somebody reads a number off a document and types it into the return. The document is a photograph of a K-1 taken at an angle, or a brokerage statement in a format you have not seen before.
Cross-checking. The figure on the statement does not match what the client said, or last year's return, and somebody has to work out which is right.
None of that is professional judgement. All of it is mechanical, all of it is where the February overtime comes from, and all of it is genuinely addressable. The judgement part, deciding how a position should be taken, is not what you should be buying software for.
What document intake actually does
The mechanism is worth understanding because it tells you what to ask for.
Documents arrive by whatever route your clients use. Something reads each one, identifies what it is, and pulls the relevant figures into structured form. Each extracted figure carries a confidence score. Above a threshold it flows into your system unattended. Below it, the figure goes to a preparer with the document on screen and the uncertain field highlighted, to confirm or correct in seconds.
That review queue is the product. Set the threshold too high and every figure goes to a human, so you have bought software to keep typing. Set it too low and wrong numbers flow into returns silently, which is much worse than typing them.
So the questions for a vendor are specific: what is the confidence threshold, how was it chosen, what does the preparer see, and what is the measured exception rate on documents like the ones your clients actually send. That last one is the honest number, and clean brokerage files behave very differently from a photograph of a crumpled 1099.
We build pipelines of exactly this shape, and the pattern holds across industries: recurring messy documents in, validated structured data out, uncertain cases to a person. On one production system in a different sector it supports well over a thousand finished reports a year at roughly eighty-five percent less time per report, and the volume only works because nobody reviews the confident cases.
The compliance questions, before the demo
For a CPA firm these decide whether a tool is usable at all, so ask them first.
Where does taxpayer data sit while it is processed, and who can reach it? IRC section 7216 restricts disclosure and use of taxpayer information, and using a third-party processor has consent and contractual implications your firm has to satisfy. Running inside your own cloud environment means client data never enters a vendor's systems, which is the simplest answer to that question. Hosted tools can be acceptable, but you need the contractual position in writing rather than a reassurance in a demo.
How does this sit with the FTC Safeguards Rule and IRS Publication 4557? Your firm needs a written information security plan covering this, including access control and monitoring. Ask the vendor what they provide to support that and what remains your responsibility.
Is there an audit trail showing how a figure got into the return? If a figure is later questioned, you need to show where it came from. Insist that every extracted value cites the document and page it was read from, and that corrections are logged with who made them.
What happens to the documents afterwards? Retention and deletion, in writing, matching your own engagement terms.
The stage to be careful about
Anything that prepares, reviews or signs off on the return itself is a different purchase.
The issue is not that the technology cannot produce good output. It is that a tax position has to be defensible, which means the reasoning has to be consistent and inspectable. Flexible systems take different paths on different runs. That is fine for reading a document and wrong for deciding a treatment.
The workable pattern is the boring one: use the flexible part for gathering and extraction, use fixed rules and checklists for anything determinative, and keep a named preparer accountable for the return. A vendor who blurs that line in a demo is worth pressing on hard.
Sizing it against busy season
Annual averages hide the whole problem, so size this on the window that actually binds you.
Count returns in the January to April window. Estimate mechanical minutes per return: intake, typing, chasing, cross-checking, not review or judgement. Multiply, then apply a fully loaded hourly cost for whoever does that work.
Four hundred returns at fifty mechanical minutes each is roughly three hundred and thirty hours inside the window. At forty dollars an hour fully loaded that is around thirteen thousand dollars, and more importantly it is three hundred and thirty hours you cannot buy more of in February at any price.
Now apply a realistic exception rate. If a third of extracted figures still need a preparer's eye, you recover roughly two-thirds of that time, not all of it. Size the quote against the reduced figure.
The number that never appears in a proposal is capacity. If the intake work is what caps how many returns your firm can take, then the saving is not really the labour cost, it is the additional engagements you could accept with the same people. For most firms that is the larger number by a distance.
How to start before next season
One document type, one source, real client files rather than samples, running in production. Measure the exception rate for a month outside the busy window so you learn the honest number when there is time to react to it.
And keep the internal framing right. The system reads, files and chases. Your preparers and partners keep review, judgement and the client relationship. Firms that pitch this internally as needing fewer seasonal staff tend to find the people whose knowledge the build depends on are unhelpful, for entirely rational reasons.
FAQ
Can AI prepare a tax return?
Parts of the input work, reliably. The return itself should remain a preparer's responsibility, and the reason is defensibility rather than capability: a position needs consistent, inspectable reasoning if it is ever questioned, and flexible systems do not reason identically across runs. The pattern that works is extraction and intake handled by software, determinative treatments handled by fixed rules and checklists, and a named preparer signing off.
Is it compliant to send client tax documents to an AI tool?
That depends entirely on the arrangement, and it is a question for your written information security plan rather than the vendor's marketing. IRC section 7216 governs disclosure and use of taxpayer information, and the FTC Safeguards Rule and IRS Publication 4557 shape what your firm must have in place. The cleanest answer is a deployment where taxpayer data never leaves infrastructure your firm controls. Anything else needs the contractual position documented before files move.
How accurate is extraction from tax documents?
It depends far more on document quality than on the software, which is why a single accuracy figure should be treated as marketing. Clean electronic brokerage and payroll files extract very reliably. Photographs of crumpled forms, phone pictures taken at an angle, and unfamiliar layouts produce meaningfully more exceptions. Ask for the exception rate measured on documents comparable to what your clients actually send.
When should we implement this, given busy season?
Outside the window, with enough runway to measure the real exception rate before it matters. Implementing intake changes in February means discovering your true exception rate at the worst possible moment. The useful sequence is to run one document type in production during a quiet period, learn the number, then widen the scope ahead of the next season.
Key takeaways
- The measurable win for CPA firms is document intake and extraction, not preparation of the return itself.
- IRC section 7216, the FTC Safeguards Rule and IRS Publication 4557 shape what deployment is acceptable, so settle data location first.
- Every extracted figure should cite the source document and page, with corrections logged.
- Determinative tax treatments need consistent inspectable reasoning, which points to fixed rules and preparer sign-off.
- Size the return on January-to-April capacity, because the binding constraint is hours inside that window.
If you want a straight read on which parts of your intake are worth automating before next season, that is a short conversation.
Related reading: what is intelligent automation · AI agents for business · do you need an AI automation consultant
Services: CPA tax document automation · document processing automation
Read next: AI Automation for CPA & Accounting Firms


