AI Tax Research Tools vs a Custom Firm Research Agent

By Jude Lee · · Comparison

Two tax professionals reviewing an AI-drafted research memo against primary source material in a CPA firm office

Three different products wear the same label

When a partner asks “should we get an AI tax research tool?”, the honest first move is to figure out which of three things they mean.

1. A general-purpose AI assistant. Claude, ChatGPT, or Copilot with no tax-specific content behind it. It has read a lot of publicly available material, including tax material, but it has no live connection to the Internal Revenue Code, the Internal Revenue Bulletin, or state guidance unless you give it one.

2. An AI layer on a commercial research library. The major research platforms (Thomson Reuters Checkpoint, Wolters Kluwer CCH, Bloomberg Tax and others) have added generative search and summarization over their editorially maintained content. Vendor feature sets and pricing change frequently — verify claims against current product documentation before you buy, rather than against a blog post, including this one.

3. A custom agent grounded in your firm’s own material. An assistant connected — usually via MCP, the open Model Context Protocol standard for giving an AI governed access to your systems — to your research memo library, prior-year workpapers, engagement facts, and general ledger. It answers “how did we handle this last time, for this client?” rather than “what does the Code say?”

They are not substitutes. They fail in different places. And sometimes none of them is the answer: for a novel position, a high-exposure transaction, or a multistate or cross-border question where the downside is a penalty or a restated return, the right move is outside specialist counsel — an AI-assisted workflow can organize the facts for that conversation, but it should not stand in for it.

What a general assistant is genuinely good at — and where it breaks

General assistants are excellent at the parts of research that aren’t lookup: restating a messy client fact pattern into a clean issue statement, listing the questions a reviewer will ask, drafting a memo skeleton, translating a technical answer into a client email, or stress-testing your own reasoning by arguing the other side.

Where they break is authority. A language model generates plausible text, and a plausible-looking citation — a section number, a revenue ruling, a case name — is exactly the kind of thing it generates well and gets wrong. It may also be reasoning from a version of the law that has since changed. Assistants that browse the web or are grounded in retrieval over live sources such as the eCFR or IRS.gov change the failure mode: outright fabricated cites become less likely, but retrieved material can still be stale, superseded, applied to the wrong facts, or drawn from the wrong jurisdiction — so the tie-out requirement doesn’t go away, it just gets faster.

What you’re actually buying from a specialist platform

You are buying editorial maintenance and linkage. Somebody keeps the content current, annotates it, and connects the explanatory analysis to the underlying authority so a click takes you to something citable. That’s the product. The AI layer on top mostly changes the retrieval experience — natural-language questions instead of keyword-and-filter searching — plus summarization.

That is a real improvement for juniors who don’t yet know the vocabulary of a topic well enough to search it. It is a smaller improvement for a seasoned tax manager who already knows where to look. Price the AI layer accordingly, and ask the vendor two specific questions: does the AI answer link to the underlying authority for every assertion, and what is the vendor’s stated policy on using your queries for model training?

The value of a research platform was never the search box. It’s that someone is paid to keep the content current and linked to authority — and no AI layer replaces that.

The third option most firms skip: your own answers

Here is the pattern that gets underused. A large share of “research” in a mid-sized firm isn’t novel — it’s re-deriving a conclusion the firm already reached, for a similar client, eighteen months ago, and which now lives in a partner’s inbox or a PDF memo nobody can find.

A custom agent with an MCP connection to your document management system, memo library, and prior-year workpapers can surface that in seconds: “we concluded X for a client with these facts in 2024, here’s the memo and the reviewer’s note.” That depends on a precondition worth checking before you scope anything: if your memo library is an unstructured folder of PDFs with no engagement, entity, year, or issue metadata, the agent will surface noise, and the cleanup and tagging project is the first piece of work, not the agent. Pair it with a reusable skill that enforces your standard memo format — issue, facts, authority, analysis, conclusion, open items — and you get consistency across preparers as a side effect.

The same connection matters for client facts. Half of a research question is usually “what are the actual numbers?” Read-only, least-privilege access to the ledger — the approach described in connecting an AI assistant to QuickBooks or Xero via MCP — lets the agent pull the basis, the balance, or the ownership percentage instead of asking a staff member to go dig.

Buy the AI layer on a research platform
Best when your bottleneck is finding current authority on topics your firm hasn’t covered before. Predictable per-seat cost, vendor-maintained content, links to citable sources, no engineering. Weakness: it knows nothing about your clients or your firm’s precedent, and you’re renting the retrieval experience.
Build a custom agent over your own material
Best when your bottleneck is rediscovering your own conclusions and enforcing consistency across offices. Weakness: it does not maintain tax content — you still need a source of authority. Requires engineering, access governance, and someone to own it.

For most firms past a certain size, the real answer is the third framing: keep the maintained library for authority and add a firm-precedent agent on top of it. That means carrying both a recurring subscription and an internal build — the subscription stays a predictable per-seat operating cost, while the agent is a one-time build plus ongoing ownership (usage costs, access reviews, metadata hygiene, and a named internal owner). Budget it as two line items, not one, and don’t let a custom build become the justification for cancelling the library.

The gate that makes any of these safe

No matter which option you pick, the workflow around it matters more than the tool. This is the same review-gate logic that applies to close and compliance automation: the agent produces a draft with its work exposed, and a human authorizes the conclusion.

  1. Force a written fact pattern first

    Have the assistant restate the facts and flag missing ones before it answers. Most bad AI research answers are answers to the wrong question.
  2. Require authority with every assertion

    Instruct the agent — in a reusable skill, not ad hoc — that any conclusion must cite a specific section, ruling, or instruction, and must say “no authority located” rather than guess.
  3. Tie out every cite to the primary source

    A human opens each one and checks currency and jurisdiction. This step is not optional and it is not delegable to the agent that produced it.
  4. Check firm precedent

    Query the memo library for prior conclusions on similar facts. Inconsistency across clients is a risk item the agent is unusually good at surfacing.
  5. Sign off and file back

    The approved memo goes back into the library, tagged — which makes the next query better. The feedback loop is the point.

Where confidentiality sets the boundary

Taxpayer data is the constraint that decides a lot of this. Before any client fact pattern goes into a tool, know where it lands. The IRS’s Publication 4557, Safeguarding Taxpayer Data, and Publication 5708 on written information security plans are the starting points for tax practitioners; the FTC Safeguards Rule sits behind them. Practically, that means: business or enterprise tier agreements with documented data handling, training-off by default, named vendors in your WISP, and a firm rule that consumer accounts never touch client data. Confirm the specifics with your professional standards or compliance lead before you enable anything — this is one place where getting it approximately right isn’t good enough.

Modeling the payback without inventing a number

Don’t accept a vendor’s hour-savings figure, and don’t accept mine — I don’t have one. Build the model from your own time data. The figures below are illustrative placeholders to show the shape of the arithmetic, not benchmarks:

40
Assumed research hours logged per quarter — pull your real number from your time system
Illustrative assumption
25%
Assumed share of those hours re-answering questions the firm already answered
Illustrative assumption
15 min
Assumed verification time the AI workflow ADDS per question — count it, it is real
Illustrative assumption

So the worked example is: 40 hours × 25% = 10 re-derived hours per quarter × your blended cost rate = the size of the prize, minus (questions per quarter × 15 minutes × rate) for verification. Substitute all four inputs with your own. To get the middle number honestly, sample twenty recent research entries and tag each as novel or re-derived. If the re-derived bucket is small, buy a seat on a research platform and stop. If it’s large and growing, the case for a firm-precedent agent gets real. Either way, subtract the verification time — a workflow that produces drafts faster but requires more careful checking can be a wash, and pretending otherwise is how firms end up with shelfware.

If you’re weighing this against broader automation spend, the same triage applies as in rules, AI agents, or neither: the highest-value AI work in a tax practice is often not the glamorous research question at all, but the boring, high-volume, well-defined tasks around it.

Not sure where to start?

Get a free automation audit: we map your bookkeeping, month-end close, client onboarding, document collection, and AP/AR — and show you what's worth automating before you spend a dollar.

Get a free automation audit