IRS Notice Response: Software vs AI Agents for Firms
Why notices are the workflow nobody has automated
Almost every firm has automated something. Bank feeds categorize transactions. E-signature chases signatures. Tax software runs diagnostics. Ask what’s still manual and you’ll hear the same answer: IRS and state notices.
The reason is structural. A notice arrives as paper or a client’s blurry phone photo. Resolving it requires pulling the filed return, matching it to third-party data, checking authorizations, drafting a letter, attaching support, and tracking a response deadline that’s printed on the notice itself. No single system owns that chain, so it lands on whoever opens the mail.
This article models the IRS side, because that’s where the process is most standardized. State notice work follows the same shape but diverges in ways that break a copy-paste approach: each state has its own authorization form and its own rules about who may receive taxpayer information, there is no cross-state equivalent to the IRS practitioner transcript tools, and responses go through per-state portals, fax numbers, or mailing addresses that change. Treat each state as its own configuration rather than assuming an IRS workflow transfers.
That makes notices a good test case for the question everyone is actually asking about accounting firm automation: which parts of this belong to software, which to an AI agent, and which to a licensed human?
What the three approaches actually cover
The honest answer for most firms is not either/or. Software should remain the system of record for deadlines and authorizations — that’s a compliance function and you want it boring. The agent sits on top, doing the assembly and drafting that currently eats senior time.
A worked agentic notice workflow, step by step
Here’s how the chain looks when an agent handles the assembly and a practitioner handles the judgment. Assume the firm has connected an AI assistant to its document management system, its tax software output, and its general ledger under read-mostly access — the same pattern described in connecting an AI assistant to QuickBooks or Xero via MCP.
-
Intake and classify
Notice is scanned into a single intake folder. The agent reads it, identifies the notice type (for example a CP2000 underreporter notice or a CP14 balance-due notice), the taxpayer, tax year, proposed adjustment, and the response date printed on the notice. It creates a task in your practice system with those fields populated. -
Check standing before doing anything else
The agent verifies whether a current Form 2848 (Power of Attorney) or Form 8821 (Tax Information Authorization) is on file for that taxpayer and year. If not, it stops and flags the gap. The IRS provides Tax Pro Account and e-Services for practitioner authorizations and transcript access — confirm current functionality and requirements on IRS.gov, since these tools change. -
Assemble the evidence
The agent pulls the as-filed return, the relevant workpapers, and the GL detail for the line item in question, then builds a side-by-side of the IRS’s figure versus the firm’s figure with the supporting documents listed as exhibits. -
Draft the response and the client email
Using a firm skill — a packaged instruction set that encodes your house format, tone, and required elements — the agent drafts the response letter and a plain-English client explanation. Skills are what stop the agent from reinventing your letter format every time. -
Human review gate
A CPA or EA reviews the reconciliation logic, the conclusion, and every authority cited. Nothing is signed or transmitted by the agent. This is the same review-gate discipline covered in building AI review gates instead of autopilot. -
Log and close
The final packet, the reviewer, and the send date are logged. The agent drafts the follow-up reminder if the IRS hasn’t responded by a set interval.
Three failure modes worth naming
These are the ones that make this workflow risky rather than merely imperfect. A CP2501 and a CP2000 look nearly identical on a scan, but they call for different responses and different next steps — an agent that classifies on layout alone will happily label one as the other. On a notice covering multiple periods, an agent asked for “the tax year” may grab the first four-digit year it sees rather than the year the adjustment applies to, which quietly poisons every retrieval downstream. And a drafting model asked to justify a position will produce a Treasury Regulation or Internal Revenue Manual cite in correct format that does not exist.
The gate catches these only if it is specific. Require the reviewer to confirm the notice type against the printed notice number, confirm the tax period against the adjustment line rather than the header, and open every cited authority at its source before sign-off. “Looks right” is not a check.
Modeling the payback without making numbers up
We have no benchmark data on notice-handling time, and you should distrust anyone who quotes one. Build the model from your own file, and build both sides of it.
Pull last season’s notice log. Count the notices (N). Estimate average practitioner hours per notice today (H) — intake, research, assembly, drafting, review. Estimate what share of those hours is assembly and drafting rather than judgment (S). Then:
Gross recoverable hours = N × H × S × (agent effectiveness you actually observe, not what a vendor claims)
Now subtract the cost side, in the same units. One-time setup and skill-authoring time, amortized over the period you’re measuring. Ongoing governance: access re-scoping, log review, skill updates when a notice format or your house letter changes. Per-seat or per-usage platform cost, converted to hours at your blended rate. And the reviewer minutes (R) the gate still consumes on every packet — that cost never goes to zero, and it shouldn’t.
Net hours = gross recoverable − setup/amortized − governance − (N × R) − (fees ÷ rate)
Multiply net hours by your standard billing rate only if those hours genuinely get redeployed onto billable or advisory work. If they get absorbed, the return is capacity and turnaround time, not revenue — still valuable, but say so honestly to your partners.
The data-access question you have to settle first
Notice work is taxpayer data by definition. Two things to settle with counsel before an agent touches it.
First, disclosure and use rules. Treasury regulations under IRC §7216 govern how preparers may disclose or use tax return information, with specific consent requirements and certain exceptions. Whether a given AI vendor relationship fits an exception is a legal question — confirm it with a qualified tax professional or counsel rather than assuming.
Second, safeguards. IRS Publication 4557, Safeguarding Taxpayer Data, sets out security expectations for practitioners, and IRS Publication 5708 walks through creating a written information security plan. Any AI tooling you add has to fit inside that plan — access scope, logging, retention, and vendor terms included.
Representation is a licensing question, not a capability question
Practice before the IRS — representing a taxpayer, arguing a position, signing on their behalf — is restricted to authorized practitioners under Treasury Department Circular No. 230, which covers attorneys, CPAs, enrolled agents, and certain other categories with defined scope. Software is not an authorized representative and cannot become one by getting better at drafting. No amount of model improvement changes that.
What changes is the mix. The portion of the work that is retrieval, comparison, and first-draft prose is genuinely automatable now. The portion that is judgment, authority, client communication, and accountability is not.
An agent can assemble the argument. Only a licensed practitioner can make it.
Build, buy, or leave it manual
Buy first if your problem is tracking. If notices get lost, miss deadlines, or nobody knows who owns them, that’s a workflow-software problem and no agent will fix it.
Build an agent if your problem is assembly. You have the tracking, but seniors are spending hours per notice pulling files and writing the same reconciliation narrative in slightly different words. That’s where a skill plus scoped data access pays off.
Do nothing if your volume is low. If you handle a handful of notices a year, setup and governance cost exceeds the benefit — the decision framework in rules, AI agents, or neither applies directly.
One note on the market as of 2026: vendors across the accounting and professional-services stack are shipping embedded agent bundles in bulk, often announcing many agents and actions in a single release. Embedded agents are convenient and worth evaluating, but they generally only reach data inside that vendor’s product. Notice work spans tax software, document management, the GL, and IRS systems — which is why firms with real notice volume tend to wire their own connections rather than wait for one vendor to cover the whole chain. Evaluate what’s shipped and testable today, not what’s on a roadmap.
Not sure where to start?
Get a free automation audit: we map your bookkeeping, month-end close, client onboarding, document collection, and AP/AR — and show you what's worth automating before you spend a dollar.
Get a free automation audit