When the person who ran the IRS writes a risk framework for AI in tax practice, it's worth reading closely. In The Tax Adviser this August, former IRS Commissioner Danny Werfel laid out fifteen risks tax preparers take on when they bring AI into their work — organized in layers, because a failure early in the stack cascades into everything above it: technical failures become judgment failures, judgment failures become malpractice exposure, and all of it lands on client trust.
We read it with particular interest, because the risks near the top of his framework are the ones Partnertax.ai was designed around from the first day. Here's what the framework asks — and how a citation-first tool answers.
Risk one: hallucinated advice
Werfel's first technical risk is the one every practitioner has heard about: AI that produces confident, fluent, wrong answers — hallucinated citations and deductions that flow silently into returns without triggering any obvious flag.
The problem with a hallucinated answer isn't just that it's wrong — it's that checking it costs as much as doing the research yourself, so under deadline pressure the checking quietly stops. That's why our AI Tax Assistant attaches a numbered citation to every claim, linking to the primary authority it came from — the Code, the regulations, IRS publications, state guidance. The same is true cell by cell in the State Tax Matrix, where hovering shows the exact passage behind each value, and an entry without a source is labeled as such rather than hidden. A cited answer can still be checked in a click; an uncited one can only be trusted or redone.
Behind that experience sit four specific controls, each aimed at a different way a hallucination could reach your screen:
| Control | What it means for your answer |
|---|---|
| Grounded-only answering | The assistant answers only from the authority it retrieved for your question — never from the model's built-in memory. If nothing in the sources supports an answer, it says so rather than guessing. |
| Citation contract | Every claim must carry a citation that resolves to a real retrieved source. A citation that doesn't resolve is dropped — it never renders as a plausible-looking reference. |
| Authority hierarchy | Sources are weighted the way a practitioner would weight them — statute over regulation, regulation over case, case over ruling — and retrieval starts from the statute. |
| Grounding badge | Every answer shows how strongly the sources support it — high, medium, or low — and the badge stays on the answer in your chat history. |
Risk two: over-trust
The sharpest risk in Werfel's practice layer is what he calls misplaced authority: the failure isn't the AI being wrong, it's the professional not checking — reliance drifting into unauthorized reliance.
No tool can force a practitioner to verify. What a tool can do is make verification cheap enough that it actually happens, and keep the professional's judgment in the loop instead of replacing it. That's the reason citations sit inside every answer rather than in a footnote: the authority is one click away at the exact moment you'd rely on the claim. And it's the reason answers from web search mode carry a permanent Web badge — the line between settled law and current commentary is one you can never afford to lose track of, so we made it impossible to. Months later, your chat history still shows which answers rested on primary authority and which reflected where the conversation was at the time.
Werfel's prescribed mitigations for this risk read like a checklist: clear disclaimers that AI outputs are not authoritative guidance, source citations on all outputs, escalation to a human for scenario-specific questions, and training on appropriate use.
Citations on all outputs is the core of the product. The authority is what the citations point to. And escalation to a human is the design itself — the platform is built for tax professionals, so the human with judgment isn't a support tier to reach, it's the person reading the answer with the sources in front of them. Training on appropriate use belongs to every firm — but a tool that shows its sources makes the right habit the easy one.
Risk three: control dilution
The quietest risk in the framework may be the most dangerous: AI's speed and volume causing existing review processes to become less rigorous — not by decision, but by erosion.
This is where citation-first design earns its keep. Review discipline erodes when checking is expensive; it holds when checking is cheap. When every claim shows its source, "verify before relying" stops being an afternoon of research and becomes part of reading the answer. The deliberate limits matter here too: a question about a client's operating agreement is answered from that document, not the open internet — attaching a file turns web search off and tells you why. Guardrails a professional doesn't have to remember to apply are the ones that survive a busy season. And the audit trail is built in: every answer persists with its citations, its web/knowledge-base label, and its grounding level — so a reviewer can always reconstruct what an answer rested on.
Risk four: client data leakage
Werfel's framework also names the risk practitioners ask us about most: client data ending up somewhere it shouldn't. Here's where we stand, stated plainly. Every model call runs inside Amazon Web Services — your questions and your clients' documents are never sent to consumer AI platforms, and under AWS's service terms, customer content is not used to train models. Uploaded documents pass through automated redaction that masks personal identifiers before their text is stored or sent to the model. And no third-party analytics tools touch the content of your conversations.
The judgment layer belongs to the practitioner
An honest reading of Werfel's framework ends with this: no tool eliminates these risks, ours included — the professional judgment layer in his stack belongs to the practitioner, and always will. What a tool can change is the cost of the discipline the framework demands — making the verify step a click instead of a research session, keeping the provenance of every answer visible, and labeling what's authority and what's commentary. Werfel's core message is that you don't have to scale fast to capture results; start bounded, verify everything, expand as trust is earned. That's not just good advice for adopting AI. It's how AI for tax should be built.
Read the framework
Werfel's full article is worth your time — read it at thetaxadviser.com, then bring your hardest research question to the AI Tax Assistant at partnertax.ai and check our citations against his framework.
Danny Werfel, "A Risk Framework for AI Use in Tax Administration and Preparation," The Tax Adviser (August 31, 2026). Available at: thetaxadviser.com
Shashank Sharma is the founder of Partnertax.ai, an AI tax assistant for tax professionals, built on cited IRS and Treasury sources.