AI beats the lawyer baseline on document Q&A (Harvey 94.8% vs 70.1%) and loses redlining outright (65.0% vs 79.7%), per the Vals Legal AI Report. That split is where your delegation line for AI for legal work goes, whatever a vendor demo suggests.
The downside has a price now. Couvrette v. Wisnovsky cost two lawyers $110,204.38 and a dismissal with prejudice. Withers v. City of Aberdeen ended in two-year suspensions from the district and bar referrals.
And the purpose-built tools don't save you: Stanford RegLab measured 17% to 33% hallucination rates in Lexis+ AI and Westlaw AI-Assisted Research.
So this is organized by risk tier: what the benchmarks let you delegate, what the case law forbids, a verification workflow that holds up under Rule 11, and the four vendor terms that decide whether privilege survives.
What the benchmarks let you delegate in AI for legal work
Two independent benchmarks draw the line for you, and it doesn't fall where vendor marketing puts it. AI beats the lawyer baseline on extraction and summary work by wide margins. It loses on drafting judgment and on searching structured public filings.
Where AI clears the lawyer baseline
The Vals Legal AI Report ran legal AI products against a lawyer control group on discrete tasks. On document Q&A, Harvey scored 94.8% against a lawyer baseline of 70.1%. On document summarization, CoCounsel hit 77.2% against 50.3% for the lawyers. On transcript analysis, Harvey took 77.8% against 53.7% (Vals AI). Those are 20-plus point gaps on tasks that eat associate hours.
Legal research moved too. A later Vals benchmark put 210 questions through blind grading and found Counsel Stack at 81%, Alexi at 80%, ChatGPT at 80% and Midpage at 79%, against a lawyer baseline of 71% (LawSites). General-purpose ChatGPT tying the purpose-built research products should give you pause when a vendor quotes you a five-figure seat price.
Where lawyers still win
Redlining is the clearest loss. Lawyers scored 79.7% against Harvey's 65.0%, and EDGAR research went 70.1% for lawyers against 55.2% for Oliver (Vals AI). If a tool in your stack is pitched at contract markup as its headline use, that pitch is running ahead of the measured performance. Use it to flag clauses for a human review pass. The markup you send out needs a lawyer's hand on it.
Reading a benchmark like a supervising partner
Treat every score as a delegation decision about one specific task. Brand-level verdicts don't survive the data. Harvey wins document Q&A and loses redlining. Same product, opposite answers.
So ask what a benchmark measured before you buy on it. A tool that scores 94.8% on pulling answers out of a document you supplied has told you nothing about whether it invents citations when it goes looking for authority on its own. Those are different failure modes with different consequences, and the sanctions numbers only attach to the second one. When you're shortlisting from the 124 legal services listings in our directory, sort by the task you need covered.
Anything scoring below the lawyer baseline gets reviewed line by line before it leaves your desk.
Where hallucination risk makes AI disqualifying
Buying a legal-specific tool cuts fabrication roughly in half. It doesn't take it to zero, and the gap between "halved" and "zero" is where sanctions live.
17–33%: what purpose-built legal research tools still get wrong
Stanford's RegLab and HAI ran a preregistered test of the two products most firms already pay for, and found Lexis+ AI hallucinated about 17% of the time and Westlaw AI-Assisted Research about 33%, against 43% for GPT-4. Accuracy told the same story: 65% for Lexis+ AI, 42% for Westlaw AIAR. One in three answers from a flagship research product carrying something fabricated or misgrounded is too much to absorb with a spot check.
That study's preprint dates to May 2024 and the tools have shipped since. It's still the only preregistered independent benchmark anyone has run on them, and no vendor has published a replication that would let you claim the number has moved.
Why 'legal-grade' and 'RAG-grounded' are not safety claims
Retrieval-augmented generation grounds the model in a real corpus. It doesn't stop the model from describing a real case as holding something it never held, or from stitching a plausible pin cite onto the wrong opinion. The Stanford team looked directly at the marketing and found vendor claims of "100% hallucination-free linked legal citations" overstated. Treat "legal-grade" as a claim about the training corpus. Your Rule 11 exposure is a separate question.
The three outputs that should never leave the building unverified
Citations, quotations, and holdings. Every case cite gets pulled up and read. Every quoted passage gets matched against the source text, character for character. Every characterization of what a court held gets checked against the opinion, because this is the failure mode that survives a citation checker: the case exists, the reporter number is right, and the proposition is invented.
Everything else is a productivity question. These three are a filing question, and the distinction matters more than which vendor you picked.
The same rule applies downstream of research. If you're running a brief or a deposition transcript through summarization tools, the summary is a starting point for your own reading and it doesn't replace that reading. Anything you lift from a summary into a filing goes back through the three checks above.
The 2026 sanctions ledger: what getting it wrong now costs
Price the downside before you price the subscription. A single bad filing has cost one legal team $110,204.38 and their client's case (Greenberg Rothstein), and two out-of-state attorneys their right to appear in a federal district for two years (JD Journal). Both are docketed orders with dollar amounts attached.
From ~200 cases to 1,668 in a year
Damien Charlotin's AI Hallucination Cases database, which the federal courts themselves cite, tracked 1,668 cases as of July 2, 2026: 1,163 in the United States, 59 in the UK, and 653 where the responsible party was a practising lawyer rather than a self-represented litigant. Mid-2025 the count was around 200. That curve is the number to show a skeptical partner.
Cite the live database rather than any secondary tracker if you're publishing a figure. Counts update daily and the trackers disagree with each other, sometimes by hundreds of cases.
The 653 figure matters more than the headline total. It means roughly two in five tracked incidents came from someone with a bar card and malpractice coverage, which is a different risk profile than pro se filers copying ChatGPT output into a complaint.
Couvrette v. Wisnovsky: the $110,204.38 benchmark
In Couvrette v. Wisnovsky, No. 1:21-cv-00157-CL (D. Or.), the court assessed $110,204.38 combined against two lawyers and dismissed the action. San Diego pro hac vice counsel Stephen Brigandi took $95,998.72 of that: $15,500 in monetary sanctions plus $80,498.72 in the defendants' attorney fees. Portland local counsel Tim Murphy was assessed $14,205.66 for signing filings he hadn't verified. The briefing carried 15 fabricated case citations and eight invented quotations across three briefs filed over five months.
Read the local-counsel number twice if you ever act as a filing agent for out-of-state lead counsel. Fourteen thousand dollars for someone else's research is now a documented price.
Whiting, Withers and the Ninth Circuit: fines became suspensions
Q1 2026 alone produced at least $145,000 in AI-related court sanctions, including $15,000 in punitive fines against each of two attorneys in Whiting v. City of Athens at the Sixth Circuit, plus opposing fees and double costs, for over two dozen incorrect or nonexistent citations (ComplexDiscovery). Money stopped being the ceiling after that.
In Withers v. City of Aberdeen (N.D. Miss., June 8, 2026), Judge Sharion Aycock found fake cases in filings from lawyers on both sides. She cancelled the scheduled trial, suspended the two lead out-of-state attorneys from practising in the district for two years, revoked their pro hac vice admissions, fined every lawyer of record between $1,000 and $3,500, and notified state bar authorities. In February 2026 the Ninth Circuit suspended Orange County immigration attorneys Mike Sethi and William Rounds for six months and fined them $2,500 each, largely because they'd passed off AI-generated errors as typos (Metropolitan News-Enterprise).
Compare that ledger against any vendor's ROI slide. The documented AI deployment case studies in our directory quantify hours saved; none of them net out a two-year suspension.
Lawyers get punished for the cover-up, not the error
Read the sanctions record closely and a pattern shows up that changes what you should be optimizing for. Filter the hallucination database by monetary and professional sanctions, and lawyers turn out to be rarely punished simply for erring with an AI tool. They get punished for refusing to own up, doubling down, or blaming somebody else once they're caught (Damien Charlotin FAQ).
That's the reframe. You can't catch every fabricated citation, so stop building a workflow that assumes you will and start building one for the day you don't.
What Charlotin's own filtering shows
The heavy penalties cluster around conduct after discovery. Judges have a lot of tolerance for a lawyer who says "I ran this through a tool, I failed to check it, here's a corrected brief and I'll pay the other side's time." They have almost none for a lawyer who claims the cite is real, or who points at an associate.
Noland and the duty to flag your opponent's fake citations
Noland pushed the duty outward. The court there declined to order sanctions payable to opposing counsel, noting the respondents never alerted the court to the fabrications and seemed to learn of them only when the order to show cause issued (LawSites). Your verification obligation covers briefs you receive as well as briefs you file. Cite-check the other side's authorities and tell the court what you find.
Why 'typographical mistakes' cost two attorneys six months
In February 2026 the Ninth Circuit suspended Orange County immigration attorneys Mike Sethi and William Rounds from practice before the court for six months and fined each $2,500. The six months came from what happened after the bad citations surfaced. They failed to disclose that the inaccuracies came from generative AI and instead described them as typographical errors, and the court wrote the order as a warning to its whole bar (Metropolitan News-Enterprise).
Contract review and discovery: the strongest case for AI
Discovery is where the evidence stops being suggestive and starts being lopsided. Every other task in this article comes with a caveat about verification cost. This one comes with a study where the machine beat the method your firm probably already uses and bills for.
Document review: 88% recall vs 64% for active learning
Redgrave ran GenAI against active learning on a 45,004-document corpus and got 88% recall with 1% elusion, against 64% recall and 3% elusion for the incumbent technology-assisted review approach (Redgrave). Recall is what you find. Elusion is what you miss in the discard pile, and cutting it from 3% to 1% on a corpus that size means roughly 900 fewer responsive documents left behind.
Take that as a defensibility argument. If opposing counsel challenges your review protocol, "we used the method with higher measured recall on a published 45,004-document benchmark" is a better answer than a per-document rate card. Cost savings come along for the ride.
First-pass contract triage vs final redlines
Split contract work in two and the benchmark verdict is clean on both halves. First-pass triage is extraction: pull the indemnity cap, find every auto-renewal, flag which of 200 vendor NDAs deviate from your playbook. That's the same shape as document Q&A, where Harvey hit 94.8% against a 70.1% lawyer baseline (Vals AI). Delegate it.
Final redlines are the other half, and there the same benchmark has lawyers at 79.7% against Harvey's 65.0%. A redline is a negotiating position with legal consequences. Keep a human on it.
The e-commerce and healthcare contract stack in practice
High-volume, low-variance contract portfolios are the sweet spot. An e-commerce operator with hundreds of near-identical supplier terms, or a healthcare group processing payer agreements and BAAs, gets most of the value from clause extraction and deviation flagging rather than from anything that drafts. Our directory lists 217 AI document and PDF tools and 550 agentic AI tools, and for this work the extraction end of that list matters more than the generation end.
Set the threshold at the transition point: anything that reads a document and reports what's in it can run at volume with sampled review. Anything that writes a document someone else will sign gets a lawyer's name on it before it leaves.
A verification workflow that survives a Rule 11 inquiry
The protocol below takes about ten minutes per brief once it's habit. It's built backward from what got lawyers sanctioned, so every step maps to a failure mode in the case record rather than to a vendor's best-practice page.
Pull every cite from the primary source, not the tool's link
Open Westlaw, Lexis, CourtListener or the court's own docket and retrieve the opinion yourself. Don't click the link the AI gave you, and don't accept a citation because the tool displayed a reporter volume and page number that look plausible. Fabricated cites come with fabricated metadata. In Couvrette v. Wisnovsky the fake material survived three briefs over five months before anyone pulled the underlying cases (Greenberg Rothstein).
Do this for opposing counsel's cites too. A court has already docked lawyers for failing to flag an opponent's fabrications, noting they seemed to learn of them only when the order to show cause landed (LawSites).
Check the quotation, then check the holding
A real case can carry an invented quote. Couvrette involved eight fabricated quotations alongside the 15 nonexistent cases (Greenberg Rothstein), and the Sixth Circuit's $15,000-per-attorney fines in Whiting v. City of Athens covered citations that were misrepresented as well as missing (ComplexDiscovery). Text-search the quote in the opinion. Then read enough of the surrounding paragraphs to confirm the case holds what your brief says it holds.
Log the prompt, the model and the human reviewer
Keep a one-row-per-filing record: date, tool and model version, the prompt, which cites the AI produced, who verified each one, and when. Texas Ethics Opinion 705 puts competence and supervision of generative output squarely on the lawyer (Texas Center for Legal Ethics), and a contemporaneous log is what turns that duty from an assertion into evidence at a show-cause hearing.
The disclosure script for when you find a fabrication post-filing
Move within 24 hours, in writing, before anyone asks. File a notice of correction that says four things: the specific citations that are wrong, that they came from a named AI tool, that you filed without verifying them, and what you've withdrawn or corrected. Name yourself as responsible. Don't blame an associate, a contract researcher or the software.
That script exists because the sanctions data points at it. The Ninth Circuit suspended two attorneys for six months and fined them $2,500 each largely because they recast AI errors as typographical mistakes instead of disclosing the source (Metropolitan News-Enterprise). Owning it early is the cheapest filing you'll ever make.
Privilege and confidentiality: the four vendor terms that decide it
Brand doesn't protect privilege. Contract terms do. The 2026 e-discovery decisions have started turning on what the vendor agreement says about training, retention and third-party access (National Law Review), so read these four clauses before you read the feature list. Every one of the legal AI vendors in our directory can answer them in writing, and a vendor that won't is telling you something.
No-training commitments and why "we don't train" is not enough
Get it scoped. "We don't train on your data" often covers only the foundation model, leaving abuse-monitoring logs, human review queues and fine-tuning of retrieval layers untouched. Ask for a clause that names every downstream use, not the model weights alone.
Deletion rights and retention windows
You want deletion on demand, a stated maximum retention window in days, and deletion that reaches backups and derived embeddings. A 30-day trust-and-safety hold is common and usually fine. An indefinite window is a discovery target sitting on someone else's server.
Zero data retention and where it applies
ZDR is the strongest term available, and it's also the most oversold. It typically applies to API calls, not to the chat interface, the document workspace, or the vendor's own logging of prompts. Confirm in writing which product surfaces are covered. If your team pastes a deposition transcript into a web UI that sits outside ZDR, the term you paid for did nothing.
Subprocessors, jurisdiction and the discovery exposure nobody reads
The subprocessor list is where client confidences leave the building. Ask which model providers, hosting regions and support vendors touch your prompts, and whether the list can change without notice. California's bar puts the obligation squarely on you: you have to understand how the tool handles client information before you input it, and consent doesn't cure a vendor arrangement you never examined (State Bar of California). Get the answers in the order form, not in a sales email.
Picking your stack by risk tier, not by feature list
Sort every task you're considering into three buckets before you look at a single demo. The bucket, not the vendor, decides what you buy and how much verification you budget.
Tier 1: delegate and spot-check
Document Q&A, summarization, transcript analysis, first-pass privilege and responsiveness review. The benchmark margins here are wide enough that a competent tool plus a 10% sample check beats an associate working alone. Buy on contract terms (no-training, deletion rights, zero data retention, no third-party subprocessors) rather than on accuracy claims, because at this tier the accuracy question is settled and the privilege question isn't.
Tier 2: AI drafts, a human owns every cite
Legal research, memo drafting, anything with citations in it. Purpose-built research tools earn their price here, and you still pull and read every authority before it goes in a filing. Your verification protocol is the product you're actually buying.
Tier 3: don't, at any accuracy rate
Redlining against a negotiated playbook, novel-issue analysis, and anything filed without a named human who read it end to end. Lawyers still beat the tools on redlining, so paying for AI to do it worse is a strange trade.
One number should set your urgency. Case additions to the hallucination tracker ran just under 8 per day between May 22 and June 9, 2026, up from 5 to 6 per day in April (haqq.ai). The curve is still steepening, which means the sanctions record you're planning against is already out of date.
What our directory shows about the legal AI market
The purpose-built segment is thinner than the noise suggests. Against 13,174 AI apps in our directory, there are just 124 legal services listings. Under 1%. Most tools pitched at your practice are general-purpose products with a legal landing page, and that's exactly the population where the four vendor terms above tend to fail.
Frequently Asked Questions
Can lawyers use AI safely in 2026?
Yes, on tasks where the evidence supports it and with a verification step you can document. Independent benchmarking puts AI ahead of the lawyer baseline on document Q&A (Harvey 94.8% vs 70.1%) and behind it on redlining (65.0% vs 79.7%), so the safe use is task-by-task rather than tool-by-tool (Vals Legal AI Report). The sanctions record backs this up: filtering Damien Charlotin's database by monetary and professional penalties shows lawyers are rarely punished for erring with an AI tool, and are punished for refusing to own up, doubling down or blaming others once caught (Charlotin FAQ). Cite-check everything that goes to a court, and have a disclosure protocol ready for the day something slips through.
Have lawyers actually been sanctioned for using AI, and how much did it cost?
Yes, and the numbers have climbed fast. The largest known US penalty is $110,204.38 in Couvrette v. Wisnovsky (D. Or.), split $95,998.72 against pro hac vice counsel Stephen Brigandi and $14,205.66 against local counsel Tim Murphy, over 15 fake citations and eight fabricated quotations across three briefs (Greenberg Rothstein). The Sixth Circuit fined two attorneys $15,000 each in Whiting v. City of Athens, part of at least $145,000 in AI-related sanctions in Q1 2026 alone (ComplexDiscovery). Money isn't the ceiling: in Withers v. City of Aberdeen (N.D. Miss., June 8, 2026) Judge Sharion Aycock canceled the trial, suspended two lead out-of-state attorneys from the district for two years and referred them to state bar authorities (JD Journal).
Do purpose-built legal research tools like Lexis+ AI and Westlaw still hallucinate?
They do. Stanford RegLab and HAI's preregistered study found Lexis+ AI hallucinated around 17% of the time and Westlaw AI-Assisted Research around 33%, against 43% for GPT-4, with accuracy of 65% and 42% respectively (Stanford RegLab). The same researchers found vendor claims of "100% hallucination-free linked legal citations" overstated. Retrieval grounding lowers the error rate and doesn't remove it, so every citation still needs pulling and reading before it goes in a filing.
Does putting client documents into an AI tool waive privilege?
It depends on what your vendor contract says. How the tool markets itself has no bearing on it. The terms that decide the question are whether the vendor trains on your data, whether you hold deletion rights, and whether the deployment runs zero data retention. A run of 2026 decisions has put AI processing and privilege on a collision course in eDiscovery, and courts are looking at the data-handling arrangement rather than the "legal-grade" label (National Law Review, May 12, 2026). Read the DPA before you upload a single privileged document.
What should I do if I discover a fake citation in a brief I already filed?
Tell the court yourself, immediately, and say plainly that the citation came from an AI tool. Charlotin's own analysis of the sanctions data shows the severe penalties cluster around lawyers who refused to own up, doubled down, or blamed staff and vendors after being caught (Charlotin FAQ). The Ninth Circuit suspended two Orange County immigration attorneys for six months and fined them $2,500 each in February 2026 largely because they characterized AI-generated inaccuracies as typographical mistakes instead of disclosing the source (Metropolitan News-Enterprise). Note also that courts have started dinging lawyers for failing to flag an opponent's fabricated citations, so the duty runs in both directions (LawSites).
Which legal tasks is AI genuinely better at than a lawyer?
Document Q&A, summarization and transcript analysis, by wide margins. Vals AI measured Harvey at 94.8% on document Q&A against a 70.1% lawyer baseline, CoCounsel at 77.2% on summarization against 50.3%, and Harvey at 77.8% on transcript analysis against 53.7% (VLAIR). Legal research has since crossed over too, with Counsel Stack at 81% and Alexi at 80% on a blind-graded 210-question set against a 71% lawyer baseline (LawSites), and GenAI review hit 88% recall with 1% elusion on a 45,004-document corpus where active learning managed 64% and 3% (Redgrave). Keep redlining and EDGAR research with your lawyers: they beat Harvey 79.7% to 65.0% on the first and Oliver 70.1% to 55.2% on the second (VLAIR).







