Legal and Compliance

California's Verify Every AI Output Rule

Three states proposed or enforced 'independent verification' for AI work in 30 days. Here is what 'independent' actually requires.

California's Verify Every AI Output Rule

On May 5, 2026, the California State Bar's ethics committee proposed a new comment to Rule 1.1 (Competence). It would require lawyers to independently review and verify AI output used in client work. Connecticut's Rules Committee proposed a parallel requirement weeks earlier. New York's Uniform Court System Part 161 takes effect June 1 with the same independent-checking standard. Three states, proposing or enforcing essentially the same duty to verify AI output, all within 30 days. That is no longer a coincidence. It is a cascade.

The unspoken question the rule creates is a working one. What does 'independent verification' actually mean? Pasting a draft back into the same model's API and asking 'is this correct?' does not qualify. This post walks through the proposed text and the multi-state cascade. It covers the enforcement record that prompted it. And it shows what an objective standard for 'independent' looks like in practice.

What does the California Bar's 'verify every AI output' rule actually require?

The California State Bar circulated a redline in April 2026. The proposed comment to Rule 1.1 (Competence) is short. It reads: 'When using technology, including AI, a lawyer must independently review and verify any output used in representing a client.' That is the whole comment. It sits in the proposed amended rules redline. The Newsroom of the California Courts confirmed it on May 5, 2026. Five other AI-focused ethics changes came with it, covering disclosure, confidentiality, candor, and supervision.

Three things matter about the wording.

The duty is independent

The rule does not say 'review the AI output.' A lawyer already has a competence duty to review their work product. The new word is 'independent,' and that word is doing all the work.

The duty is affirmative

The rule does not say 'do not rely uncritically on AI.' It says 'review and verify.' That is an action, and it leaves an artifact. A lawyer asked 'did you verify this independently?' needs an answer that does not reduce to 'I read it again.'

The duty applies to any output used in representing a client

Not 'AI-generated citations' or 'filings.' Outputs. That language covers research summaries, deposition prep memos, draft contract clauses, witness questions. It covers anything that travels from the model into client work.

Bob Ambrogi's LawNext walkthrough sent the exact 'independently review and verify' language to the bar and the in-house community. Every California lawyer with an AI workflow now has that language on their desk.

How is 'independent verification' different from re-reading an AI draft?

The rule does not answer this. But the answer is the working core of the duty. Consider two workflows.

Workflow A

Lawyer drafts brief with Claude, then pastes the draft back into Claude. Asks: 'Are these citations accurate? Are these arguments correctly stated?' Claude responds with confidence that everything looks fine. Lawyer files.

Workflow B

Lawyer drafts brief with Claude, then pulls citations into Westlaw. Confirms each one exists, says what is claimed, and is still good law. Asks GPT and Gemini to pull the supporting passages from the actual sources on their own. Files only when the cross-vendor pulls agree.

Workflow A is what most lawyers describe when they say 'I checked the AI output.' Workflow B is what the proposed rule actually requires. The difference is structural. In Workflow A, the same probability distribution that produced the draft grades the draft. Workflow B brings in error sources that are not correlated. That is the architectural reason independence has working meaning.

Why this matters for the duty. In April 2026 the Northern District of California fined a managing partner for failing to supervise a junior. The junior's AI-fabricated citations had reached a filing (Bloomberg Law, April 29 2026). The court did not fault the AI tool. It faulted the supervision chain. Verification is now part of supervision. 'I read it again' is not a defense. 'We ran it through cross-vendor consensus' is.

A practical playbook published April 24 by Unite.ai proposes evidence-linked outputs. Each one carries traceable sources, reviewer identity, timestamps, and model/version metadata. That framework is what defensible verification looks like as a paper trail. It does not name the mechanism. But one thing follows anyway. A paper trail of the same model checking itself is not evidence of independence.

Notice the pattern. Workflow B is just Workflow A with a second probability distribution checking the first. That is exactly what TrueStandard does. Paste your draft. Four to five models check in parallel in 60 seconds, every disagreement surfaced.

Which states are requiring lawyers to verify AI work?

Three live moves in the same 30-day window.

California (May 5, 2026)

Proposed Rule 1.1 comment requiring independent verification of AI output, plus five related ethics changes. See the Cal Bar redline PDF.

Connecticut (April 23, 2026)

The Rules Committee proposed practice-book changes. They would require attorneys and pro se parties to independently verify all citations and authorities produced by generative AI. Practitioner take: this fits longstanding duties rather than adding new ones (Day Pitney via Law360).

New York (effective June 1, 2026)

UCS Part 161 already adopted it. The rule requires attorneys who used AI to 'carefully review papers and independently ensure no fabricated material.' Local courts are posting compliance guidance now (master rule, Kings County local).

The eDiscovery Today May 6 aggregation makes the cascade the story, not any one rule. Three states acted nearly in parallel with much the same language. That is how a national standard gets written without a federal rulemaking. The next 90 days will produce more state proposals. Firms watching the cascade are already adjusting their internal QA before the rules take effect.

Why is California writing this rule now?

Because the enforcement docket already exists and is growing. The sanction record on AI citations reads like a chronology of the rule writing itself. In the same window the proposal was being drafted:

Pa. attorney

$5,000 sanction plus AI-ethics training for an AI-generated case citation. Judge 'appalled' by repeated bogus citations (Law360, April 20 2026).

Sullivan and Cromwell

Apologized to Bankruptcy Judge Martin Glenn (SDNY) for AI-introduced errors in a Chapter 15 filing. Opposing counsel flagged (Reuters, April 21 2026).

Georgia Supreme Court

Sanctioned an Asst. DA for fake AI-generated citations in a murder appeal (Law360, May 5 2026).

California State Bar

Opened discipline charges against several lawyers for filings with fake AI citations. One faces a probationary suspension (LA Times, April 13 2026).

Court order in deposition

Barred a pro se deponent from using ChatGPT to answer deposition questions. Rejected the privilege claim too, because no attorney was involved (EDRM, April 27 2026).

Delaware Chancery $250M earn-out dispute

Cited a CEO's ChatGPT records in the judicial opinion. AI chats are now part of the evidentiary record (Alston Privacy, April 28 2026).

Five jurisdictions sanctioned five different practitioners in 30 days, for versions of the same failure. At that point the rule writes itself. The California proposal is not new. It writes down what the bench has been doing case by case.

Does asking ChatGPT to fact-check itself satisfy a verification duty?

No. The rule's language, 'independently' review and verify, is the answer. The structural reason matters. A single LLM checking its own output is a closed loop. The model that produced the hallucination is the same probability distribution now asked to judge it. Asking the same model whether its output is correct is asking the failure mode to grade itself.

The April 2026 Microsoft Research DELEGATE-52 benchmark measured this directly. Frontier models corrupt an average of 25% of document content over 20-step workflows. Those models are Gemini 3.1 Pro, Claude 4.6 Opus and GPT 5.4. The paper finds that 'agentic tool use did not measurably reduce corruption.' Adding more steps with the same vendor amplifies the failure rather than catching it.

A Purdue preprint landed at the same time, No Free Lunch: Fundamental Limits of Learning Non-Hallucinating Generative Models. It proves a stronger statement. Non-hallucinating learning is statistically impossible from training data alone, no matter how clean the corpus.

So single-vendor checking cannot, in principle, be the main defense. Independent verification, in the sense the California rule uses, needs an error source that is not correlated. In practice that means a different vendor. A model from another provider, a different training run, a different alignment regime. Multi-vendor consensus is the only design where 'independent' has a working meaning.

Notice the pattern. Same vendor checking same vendor is the same probability distribution grading itself. That is exactly what TrueStandard does differently. Paste your draft. Four to five models from different labs check in parallel in 60 seconds, every disagreement surfaced.

What does a defensible AI workflow look like for a regulated industry?

Five elements, drawn from the Unite.ai defensible-LLM-outputs playbook and the supervisor-sanction precedent in NorCal.

Multi-vendor verification

Run claims through models from at least two different vendors. The verification artifact is the disagreement signal, not just the consensus.

Recorded model identity and version

The log captures which models were used, at what version, on what date. This is the audit trail a regulator or judge can ask to see.

Reviewer identity

A specific human signs off on the verified output. The signature is the bridge between the AI workflow and the lawyer's professional duty.

Risk-stratified gates

Not every output needs the same depth of check. Citations and statistical claims get the closest look. Style edits do not.

Privilege segregation

Public-tier consumer chatbots are not privileged, and enterprise tiers with confidentiality terms are different. The April 2026 US v. Heppner analysis shows the line being drawn case by case.

The Anthropic legal demo on April 23, 2026 drew over 20,000 registrants (Florida Bar coverage). The same week Freshfields announced a firm-wide Claude deployment. The volume of professional adoption tells you why the verification rule is being written now. Lawyers are using AI faster than the standards exist to govern that use.

If California passes the rule, what changes for non-lawyers using AI in regulated work?

The rule is written for lawyers. But the reasoning carries straight over to any profession with a competence duty. Licensed accountants, financial advisors, medical writers, regulatory affairs consultants, compliance officers, due-diligence analysts. Here are three predicted ripple effects in the next 6 to 12 months.

State medical and accounting boards will copy the language

The 'competence' duty exists in nearly every professional code. The verification gap is the same. Expect parallel rulemakings in 2026 and 2027.

Enterprise procurement criteria will shift

Compliance teams at regulated employers will want proof that a vendor can do multi-vendor verification. They will want it before approving AI workflows. The current 'approved AI tools list' pattern at most large firms does not address verification independence. It will have to.

Insurance pricing will follow the docket

Professional liability carriers price what is in the claims data. The PA, NorCal, GA, and CA disciplinary record is now the data. Premiums will reflect verification practice within 12 to 18 months.

For independent professionals, the practical version of the rule already applies. Think freelance journalists, newsletter writers, solo consultants. It applies not as a regulatory duty but as a market reality. The cost of an unverified AI claim is instant, and it lands on your name. The question is the same. The design answer is the same.

What's the practical difference between Harvey, Legora, and Claude for compliance-grade work?

This was the most-discussed question on r/legaltech this month. The post 'Lawyer here - how are Legora and Harvey differentiated from Claude' hit 49 upvotes and 106 comments (thread). The honest read:

Vertical legal AI tools: Harvey, Legora, Spellbook, Ivo, Wordsmith

They wrap LLMs with legal-specific UX, integrations, prompt libraries, and source citations. They are workflow products. Their checking is usually single-vendor, often Claude or GPT under the hood.

General-purpose LLM access: Claude, GPT, Gemini Word add-ins

Integrated into the documents lawyers work in. Cheaper, broader, but no legal-specific guardrails.

Multi-vendor verification (TrueStandard category)

Adds a layer that checks claims across vendors, whichever tool produced the draft. The check is vendor-neutral by design.

The 'moat' question is the wrong question. Here is the right one. Which layer of the stack does the regulator's 'independent verification' duty live in? It cannot be the drafting layer, which is the layer producing the output. It cannot be a same-vendor self-check, which fails the independence test. It has to be a verification layer that crosses vendors. That layer is not yet a default in Harvey, Legora, or Claude. That is the gap the rule is going to expose.

How does multi-vendor AI verification reduce supervisor liability?

This is what the NorCal supervision sanction actually decided. The court did not punish the junior lawyer who pasted in the AI-fabricated citation. The court punished the supervising partner for not having a process that would catch it. Here are three working implications for managing partners and content team leads.

The check has to be in the workflow, not in the lawyer's head

'We trust our associates to check AI output' is no longer a process. A recorded step that leaves an artifact is. That means a check report, or a model-disagreement log.

The artifact has to be vendor-independent

A log showing 'Claude was asked to fact-check the Claude draft' proves the absence of independence, not its presence.

The check has to be reproducible

If a regulator asks 'show your work,' the answer needs to be replayable: same input, same model panel, same disagreement output.

These three together make the checking step defensible. It is no longer 'we trusted the associate.' The duty now has an answer. The supervisor says: 'our firm has a documented multi-vendor verification process and the artifact is in the file.'

Notice the pattern. A defensible artifact is a vendor-independent, reproducible record of model disagreement. That is exactly what TrueStandard produces. Paste your draft. Four to five models check in parallel in 60 seconds. Every disagreement comes back with model identity, version, and timestamp logged for the file.

Frequently Asked Questions

When does the California rule take effect?

The proposal was released May 5, 2026 for public comment. Final adoption follows the comment period and Bar approval. New York's Part 161 is already adopted, effective June 1, 2026. Connecticut's proposal is pending.

Does the rule apply to lawyers in other states using AI on California matters?

Choice-of-law for ethics rules generally follows two things: where the lawyer is admitted, and where the matter is handled. Multi-state firms usually follow the strictest standard. Practical answer: if California's rule passes, large firms will adopt it as a default. That avoids state-specific carve-outs in their AI workflows.

What about confidential client data? Can I use a multi-vendor verification tool with privileged content?

Privilege analysis depends on the vendor's terms. Public-tier consumer chatbots usually do not preserve privilege. Enterprise tiers with confidentiality terms are different. The April 2026 US v. Heppner ruling shows public-Claude chats not treated as privileged. TrueStandard's enterprise terms include explicit no-training-on-content commitments, see truestandard.ai/security.

Is there a federal rule equivalent to the California proposal?

Not yet. The Federal Judicial Conference has not issued a verification rule. Individual federal judges have issued standing orders requiring AI-use disclosure on filings. The state-level cascade is currently the de facto standard. A federal rule may follow within 12 to 24 months.

How does this connect to the Lancet study on fabricated medical references?

The Lancet study (May 7, 2026) recorded a 12-fold rise in fabricated references in biomedical papers since 2023, across 2.5M papers audited. Medical writers, regulatory affairs staff and IRB-reviewing physicians face the same duty. It runs parallel to the legal duty being written down in California. Any profession with a competence duty has a verification problem, and the same architecture solves it.

Keep reading

Is AI accurate for your field?

The failure modes change by profession. These break down what AI gets wrong in specific fields, with the incidents and the checks that catch them.

Make Verification Defensible

The California rule asks for an artifact, not a feeling. TrueStandard runs your draft through four to five frontier models in 60 seconds. You get a reproducible record of where they agree and where they do not.

Start Verifying →