Tax research has a structure that generative AI does not respect, and understanding the mismatch is the whole subject.
Tax research runs on an authority hierarchy. The statute governs. Regulations interpret it. Rulings, procedures, and other published guidance carry defined weight. Cases carry weight depending on the court and the jurisdiction. Secondary sources — treatises, journals, blog posts — carry none, though they are useful for finding the authority that does.
A language model produces fluent text at a uniform confidence level regardless of where in that hierarchy its content came from, whether the source was authoritative or a message board, whether the provision has since been amended, or whether the citation exists at all.
That is not a limitation to work around. It is the defining characteristic of the tool, and everything below follows from it.
Every profession has been warned that these models invent sources. In tax practice the consequence is specific and severe, and it is worth stating plainly.
The research file is the penalty defense.
Whether a position is supported — and whether the preparer and the taxpayer are protected from penalties — depends on the existence and weight of actual authority. The standards governing return positions turn on what authority exists and what it says. A preparer who relied on a citation that does not exist has not merely failed to find support; they have no support, and the file evidences a failure of diligence rather than its exercise.
That is a worse position than having done no research at all. An unresearched position is an unresearched position. A position supported by a fabricated citation, in a file, with a date on it, is a documented failure to exercise due diligence — and the practitioner conduct rules impose diligence and competence obligations independently of the penalty provisions.
The operating rule that follows is absolute: no citation enters a memo, a workpaper, or a return position unless the practitioner has read it in the primary source. Not a summary of it. Not the model's characterization of it. The source.
The tool is useful, and dismissing it costs real time. Seven applications, in descending order of value.
Notice what is not on that list: answering the tax question.
Citations. It invents statutory sections, regulation numbers, revenue procedure and ruling numbers, and case names — often in plausible formats, attached to propositions that sound right.
Currency. Provisions that were amended, expired, sunset, or superseded. Inflation-adjusted amounts. Effective dates. A model will state a rule that was true in some prior period with no indication that it has changed, and tax is a field where a substantial share of the content has a date attached.
State and local law. Thin, inconsistent, and unreliable across jurisdictions. Anything multistate requires primary research.
Provisions interacting. Tax outcomes frequently depend on several provisions operating together, and this is where models produce confident answers that are individually plausible and jointly wrong.
Arithmetic on client figures. Do not use it to compute.
Distinguishing authority from commentary. It trained on both and it does not reliably signal which it is reproducing.
Knowing when it does not know. The absence of expressed uncertainty is the core risk, and it is why the verification step cannot be discretionary.
A workflow that captures the benefit and closes the exposure. Six steps, in order.
Worth distinguishing, because the market conflates them.
An open-ended chatbot generates text from what it learned during training. It has no source for any particular statement and no ability to show you one.
A retrieval-grounded research tool searches an actual database of tax authority, retrieves documents, and generates an answer citing what it retrieved. That is a materially better architecture for this work, because the cited documents exist and can be opened.
It does not remove the verification step. Retrieval can surface a real authority that is irrelevant to the question, or relevant and superseded, and the generated summary can still mischaracterize what was retrieved. What changes is that verification becomes fast — the document is one click away — rather than a search from scratch.
The practical guidance: prefer tools that show you the source, and treat the citation list as the useful output rather than the prose above it.
Repeated from our post on AI use in accounting practice because the exposure is severe and the mistake is easy.
Do not enter client information into a consumer AI tool. Professional confidentiality obligations apply regardless of the technology, and for tax practitioners the separate statutory restriction on disclosure and use of tax return information requires client consent in a prescribed form, with penalties for violation. Pasting a client's return data into a public chatbot to research their issue is, on a plain reading, a disclosure without the required consent.
The workable controls: an enterprise or business tier with contractual terms addressing data use and processing location, de-identification before input — which works fine, because the model is helping with framing and language rather than with the client's identity — and a written firm policy that says which tools are permitted for what.
The penalty-protection point, made concrete. A defensible research workpaper contains:
The question, stated precisely, with the relevant facts as understood at the time.
The authority relied upon, cited specifically, with the relevant text attached or excerpted — not described. Attaching the provision is the single most valuable habit in tax research documentation, because it proves what was actually read.
The currency check — the date of the research, the source consulted, and confirmation that the authority was current as of that date.
The analysis, showing how the authority applies to these facts, including the contrary authority considered and why it was distinguished.
The conclusion, and the level of confidence.
Who performed it and when, and who reviewed it.
"I researched it and it's fine" is not a research file. Neither is a memo citing authority that nobody attached. And a memo citing authority that does not exist is affirmatively harmful — which is the whole reason the verification step matters.
Structured coverage is available through the AI for Accountants Certificate Program, AI Applications for Accountants, the AI for Accountants Strategy and Research Specialist series, tax practitioner regulations, penalties, and security, and ethics training and professional conduct for accounting and tax professionals.
Beyond a general AI policy, tax practice needs four rules:
No unverified citation, ever. Written down, trained on, and enforced in review. The reviewer's job includes checking that cited authority was attached.
Approved tools named, with retrieval-grounded research tools distinguished from general chatbots and with consumer tiers prohibited for client information.
Attach, don't describe. A firm standard that research memos include the relevant authority text. This makes the verification failure visible in review rather than discoverable on examination.
Currency is part of the research. A memo without a research date and a confirmation of currency is incomplete, because tax authority has a shelf life.
An honest accounting, because the productivity claims in this area are inflated.
Genuinely saved: issue framing, orientation in unfamiliar areas, drafting structure, client-facing translation, summarizing documents you supply, and generating counterarguments. For a practitioner who writes a lot, this is real — plausibly a meaningful share of the non-reading time.
Not saved: the reading. Reading the statute, the regulation, and the guidance, and thinking about how they apply to these facts, is the work — and it is the part that cannot be delegated to something that cannot tell you when it is guessing.
Which is a reasonable trade. The tool removes friction from everything around the analysis and none from the analysis itself, and a practitioner who understands that boundary gets the benefit without acquiring the risk.
The sentence to carry into any tax research workflow: the model is excellent at telling you what to look up and unreliable at telling you what it says. Used that way it is a real improvement. Used the other way it produces a file that documents a failure of diligence with a date on it.
Because the research file is the penalty defense. Whether a position is supported depends on the existence and weight of actual authority, so a preparer who relied on a citation that does not exist has no support at all — and a dated file containing a fabricated citation evidences a failure of diligence rather than its exercise. That is a worse position than having done no research.
Issue spotting and framing. Give it a de-identified fact pattern and ask what questions a tax adviser should research. It produces breadth quickly, surfaces the issue you had not considered, and being wrong costs nothing because the output is a research agenda rather than an answer.
Two things: that the authority exists, and that it says what was claimed. Practitioners check the first and skip the second, but the common failure mode is a genuine section number attached to a proposition it does not support.
Safer, not sufficient. A tool that searches an actual authority database and cites what it retrieved is a materially better architecture, because the documents exist and can be opened. But retrieval can surface a real authority that is irrelevant or superseded, and the generated summary can still mischaracterize it. What changes is that verification becomes fast rather than unnecessary.
Not a consumer tool. Professional confidentiality applies regardless of technology, and the separate statutory restriction on disclosure and use of tax return information requires client consent in a prescribed form with penalties for violation. The workable approach is an enterprise tier with contractual data terms, de-identification before input, and a written firm policy.
The question and the facts as understood, the authority relied upon with the relevant text attached rather than described, a currency check with the research date and source, the analysis including contrary authority considered, the conclusion and confidence level, and who performed and reviewed it. Attaching the authority is the habit that both proves what was read and makes a verification failure visible in review.


