Is AI accurate for tax questions?
No, wrong about half the time. A peer-reviewed study found ChatGPT answered only 39 to 47 percent of common tax questions correctly, and the AI assistants built into TurboTax and H and R Block gave wrong answers to a third or more. Tax depends on exact, current-year, jurisdiction-specific rules, which is exactly what AI gets wrong. Use it to understand concepts, then verify every figure against IRS.gov and, for anything material, a CPA.
Wrong on roughly half of tax questions
In a peer-reviewed study across the 2023 and 2024 tax seasons, ChatGPT answered only 39 to 47 percent of common tax questions correctly, and the newer model version did not meaningfully help. When the Washington Post tested the AI assistants built into tax software, TurboTax's assistant got more than half of the test questions wrong and H and R Block's gave incorrect answers to 30 percent.
Two independent tests, one academic and one journalistic, land on the same place: AI is wrong on roughly half of tax questions. And it is worse on exactly the questions people actually ask, the common ones with complex answers that depend on a taxpayer's specific facts. Tax is a bad fit for a model that predicts plausible text, because the correct answer is a precise, current-year figure, not a plausible-sounding one.
the share of common tax questions ChatGPT answered correctly across two tax seasons: wrong the majority of the time.
Why AI fails at tax questions specifically
Tax answers hinge on exact figures that change every year: brackets, standard deductions, contribution limits, phase-outs. A language model answers from training data that lags the current filing year, so it confidently cites amounts and rules that are no longer in effect. It can misread even correct IRS material, and current-year law changes are a common blind spot, so a headline like no tax on overtime can lead the model to calculate no tax at all on overtime, which is wrong.
The failures cluster where the stakes are highest. The same research found AI was less accurate on questions that are common, have complex answers, and depend on the taxpayer's specific fact pattern, because the model pattern-matches to a generic answer. It confuses federal and state rules, invents or misapplies deductions and credits, and states all of it with the same confident tone whether it is right or wrong. A wrong return is your legal and financial liability, not the model's.
Documented failures
TurboTax and H and R Block assistants tested (2024)
A Washington Post columnist tested the AI helpers built into both products. TurboTax's assistant was wrong on more than half of his questions and H and R Block's on 30 percent. Even after both companies patched the bots, they still produced wrong or irrelevant answers, including telling a student to file in two states when only one applied.
The IRS's own chatbots could not prove they work (2026)
A Treasury Inspector General audit found the IRS could not demonstrate its chat tools were effective. Only 46 percent of live chats were logged as resolved, and the automated collection chatbot failed on dozens of responses and keywords it could not recognize or adequately answer.
A peer-reviewed verdict (2025)
The academic study that measured the 39 to 47 percent correct rate concluded plainly that ChatGPT is generally not a reliable source of tax guidance for uninformed taxpayers.
The numbers behind it
of test questions answered wrong by the AI assistants in TurboTax and in H and R Block, respectively, even after the companies patched them.
of the IRS's own live chats were logged as resolved, in an audit that found the agency could not prove its chatbots work.
A wrong tax answer does not look wrong. It looks like a confident summary of deductions, a clean number, a plausible rule, and it may be describing last year's law or a credit you do not qualify for. The way to catch it is a second, independent check. Paste the answer into TrueStandard and several models cross-check it in about a minute, so the outdated figure or the invented deduction surfaces as a disagreement before it reaches your return.
How to verify AI tax output
Treat AI tax answers as a starting draft, never as filing-ready advice. Before you rely on anything, run this check.
-
01
Confirm the tax year. Make the model state which year its figures are for, then verify the answer reflects the current filing year's brackets, deductions, and limits, not a prior year baked into training data.
-
02
Check every figure against an IRS.gov primary source. Re-verify each dollar threshold, phase-out, contribution limit, and rate against the relevant IRS publication or form instructions, not a summary.
-
03
Confirm any cited form, schedule, or publication number is real and current, because models invent or misremember them.
-
04
Separate federal from your state, and check multi-state or residency questions with your state's Department of Revenue. Never assume federal treatment applies to your state.
-
05
Verify any deduction or credit actually exists, is still in effect this year, and that your specific situation qualifies. These are the exact areas where AI is weakest.
-
06
Check for current-year law changes the model may not know, from new legislation or mid-year IRS guidance, against the IRS newsroom.
-
07
For anything material, consult a licensed CPA or Enrolled Agent before filing. A wrong return is your liability, not the AI's.
How to make AI output reliable: check it across models
The fix is not to hunt for a single more accurate model. Every large language model predicts fluent, plausible text, so each one can be confidently wrong on its own. What changes the odds is agreement. When several independent models are asked the same thing and all land on the same answer, the chance they share the exact same hallucination drops sharply. When they disagree, you have found the precise claim to check by hand before it ships.
That is what TrueStandard does: it runs your draft through four to five frontier models at once and surfaces every disagreement in about a minute, with sources. See the AI fact checker for how the method works, or read why AI cites studies that do not exist for the mechanism behind the failures on this page.
Common questions
Can I use ChatGPT to do my taxes?
Not to rely on. A peer-reviewed study found it answered only 39 to 47 percent of common tax questions correctly, and the AI assistants built into TurboTax and H and R Block were wrong on a third or more. Use it to understand a concept, then verify every figure against IRS.gov and, for anything that matters, a tax professional.
Why is AI so often out of date on taxes?
Because tax rules change every year and a model answers from training data with a fixed cutoff. It confidently cites brackets, limits, and deductions from a prior year, and it often will not know about the current year's law changes or mid-year IRS guidance.
Aren't the tax-software AI assistants more accurate than ChatGPT?
Not reliably. When the Washington Post tested them, TurboTax's assistant was wrong on more than half of the questions and H and R Block's on 30 percent, and they still produced wrong answers after being patched. Even the IRS's own chatbots resolved fewer than half of live chats in a 2026 audit.
What is the safest way to use AI for tax questions?
Use it to learn terminology and frame what to ask, then verify every figure and rule against an IRS primary source before acting, and bring anything material to a CPA or Enrolled Agent. Checking the answer across several models flags the outdated figures and invented deductions that a single model will state with full confidence.
Do not publish AI output on trust
Paste your draft. Four to five models check every claim in about 60 seconds, and you see exactly where they disagree before your name is on it.