Automate the work where checking the output costs less than doing it yourself: client intake, summarising documents you already hold, first drafts of routine agreements, discovery triage, and document comparison. Do not automate citation, filing, or advice. For a small firm, Spellbook is the most realistic starting point, CoCounsel and Lexis+AI are the only research tools worth leaning on because they are grounded in real case databases, and Harvey — the name you have probably heard most — is almost certainly not for you.
Why lawyers keep getting sanctioned
Judges in multiple jurisdictions have now sanctioned lawyers who filed briefs containing case citations that did not exist. The cases were invented by a general-purpose chatbot, and they were invented well: plausible party names, correct-looking reporter formatting, holdings that sat neatly under the argument they were cited for.
That last detail is the whole problem. A language model does not usually fail by producing obvious gibberish you would catch on the first read. It fails by producing exactly the authority you were hoping to find. Confirmation bias does the rest. The lawyer who gets sanctioned is rarely lazy — they are busy, the citation looked right, and it said what they needed it to say.
The quieter failure that will never make the news
Fabricated citations are the loud version. Here is the quiet one: the case is real, the quotation is real, and the case does not stand for the proposition you cited it for.
Retrieval-grounded tools — the ones wired into an actual case database rather than generating from memory — largely solve the first problem. They do not solve the second. A grounded tool can surface a genuine opinion and still mischaracterise its holding, miss that it was overturned on appeal, or flatten a narrow, fact-bound ruling into a general principle. Nothing about the software changes who signed the filing. Certification obligations have no software exception, and "the tool produced it" has not worked as a defence for anyone who has tried it.
The rule that should govern every decision you make here
Automate a task when verifying the output costs less than producing it. That single line will make almost every one of these decisions for you.
Summarising a 90-page deposition passes the test. You verify by skimming the summary against the transcript, which is far cheaper than reading it cold. Finding controlling authority fails the test badly. To verify that a model found the right case, you have to do the research yourself — so you have saved nothing, and you have added a new way to be wrong.
What is safe to automate
- Intake and qualification. Structured questions, structured answers, no judgment required.
- Summarisation of documents you supply. The source is in front of you, so verification is bounded.
- First drafts of routine documents. Engagement letters, NDAs, standard clauses.
- Discovery triage. Ranking and clustering thousands of documents so a human reads the right 200 first.
- Clause extraction and document comparison. Mechanical, checkable, tedious — the ideal profile.
- Transcript and record search. Semantic search over material you already possess.
- Billing narratives and admin. Genuinely the lowest-risk hours you will ever recover.
What is not
- Legal research without a grounded database. And even with one, every cite gets read.
- Anything filed with a court that a human has not read line by line.
- Advice to a client. Obviously, but it needs saying, because clients will do it themselves.
- Anything privileged pasted into a consumer chatbot. This is the one firms actually get wrong.
The confidentiality problem almost everyone skips
Your duty of confidentiality does not pause because a text box was convenient. Consumer tiers of general-purpose assistants may retain inputs, and depending on the plan and settings, may use them to improve the underlying service. Enterprise and zero-retention agreements exist precisely because this matters.
Before any client data touches a tool, get four things in writing: that your inputs are not used to train models, what the retention window is, where the data is stored, and who at the vendor can access it. If a vendor cannot answer those in writing, the answer is no. This applies to the AI document tools you use for routine paperwork just as much as to the flashy legal-reasoning ones.
The tools, honestly
Spellbook
A Microsoft Word add-in for transactional work: it suggests clauses, redlines drafts, flags missing or unusually aggressive terms, and drafts language against the agreement in front of it.
Who it is for: small and mid-size transactional firms whose working life already happens inside Word. Who should skip it: litigation-heavy practices. The value is concentrated in contracts. The honest trade-off: it speeds up the first pass and does not take on any of the judgment. Reviewing a Spellbook redline is faster than drafting cold and slower than not reading it — and not reading it is the only genuinely unacceptable option. Pricing is seat-based and sales-led, so ask for current numbers rather than trusting a figure in a blog post.
Harvey
Enterprise legal AI, sold into large firms and in-house teams, deployed alongside security review and support.
Who it is for: organisations with a procurement process and an innovation budget. Who should skip it: essentially every small firm. It is not self-serve, it is not cheap, and much of what it does only pays back against document volume a four-person practice does not have. The honest trade-off: Harvey has more brand recognition than relevance for most of the people who search for it. That is not a criticism of the product. It is a warning about the shortlist you built from headlines.
CoCounsel
Thomson Reuters' legal assistant, grounded in Westlaw. Research, document review, deposition preparation, summarisation.
Who it is for: firms already living in the Westlaw ecosystem. Who should skip it: firms that are not, unless you are willing to buy into the ecosystem as well as the tool. The honest trade-off: grounding materially reduces the risk of invented citations. It does not touch the risk of a real case being mischaracterised, and it does not remove your obligation to read every authority you cite.
Lexis+AI
LexisNexis' equivalent, with answers linked back to authority and citation signals attached.
Who it is for: LexisNexis firms. Who should skip it: the same logic as above — this is an ecosystem decision more than a tool decision. The honest trade-off: the linked-citation design is the single most important feature in this entire category. It is also the clearest argument for why a general-purpose chatbot has no business doing your case research.
PandaDoc
Not a legal-reasoning tool at all. Document generation, templating, e-signature, and workflow.
Who it is for: any firm bleeding hours into producing, sending, and chasing documents. Who should skip it: anyone who needs substantive review. Wrong category entirely. The honest trade-off: unglamorous, and probably recovers more hours per month for a small firm than anything else on this page. Nobody writes think-pieces about templating. Look at PandaDoc before you look at anything with "AI lawyer" in the marketing.
DoNotPay
Consumer-facing self-help, marketed at one stage as an "AI lawyer" — framing that drew regulatory scrutiny over how those capabilities were presented.
Who it is for: a consumer contesting a parking ticket. Who it is not for: your practice. DoNotPay is on this list only because clients will ask you about it, and you should have an answer ready.
| Tool | What it actually does | Best for | Watch out for |
|---|---|---|---|
| Spellbook | Contract drafting and redlining inside Word | Small transactional firms | Contract-focused; sales-led seat pricing |
| Harvey | Enterprise legal reasoning and workflow | Large firms and in-house teams | Not self-serve; wrong scale for small firms |
| CoCounsel | Research and review grounded in Westlaw | Existing Westlaw firms | Ecosystem lock-in; still verify every cite |
| Lexis+AI | Research with linked authority and citation signals | Existing LexisNexis firms | Ecosystem lock-in; cost |
| PandaDoc | Document generation, templating, e-signature | Any firm drowning in paperwork | No substantive legal review |
| DoNotPay | Consumer legal self-help | Individual consumers | Not a firm tool; past scrutiny over marketing claims |
Swipe the table sideways to see every column →
Before any of this touches a client matter
A seven-item checklist. It is short on purpose, so there is no excuse for skipping it.
- Written confirmation that your inputs are not used to train the vendor's models.
- Retention and deletion terms, in writing, with a number attached to the retention window.
- A named human verifier. Every authority is pulled and read in full. Not skimmed. Not "it looked right".
- A disclosure position. Decide what you tell clients, and whether it belongs in the engagement letter.
- A billing position. If a task now takes twenty minutes instead of three hours, you cannot bill three hours. The next wave of ethics complaints in this area will be about billing, not hallucinations.
- Conflicts and data segregation. Matter data should not leak across matters through a shared workspace.
- A log. Which tool, which matter, who checked the output. If you ever have to explain yourself, this document is the difference between an awkward conversation and a very bad one.
A drafting loop that holds up
- You define the issue, the deal, and the position. Not the model.
- The tool produces a first draft, a redline, or a summary.
- You read it adversarially — as though opposing counsel wrote it and is trying to sneak something past you.
- Every authority: open it, read the holding, confirm it says what the draft claims it says. Every time.
- You sign. You own it. The tool owns nothing.
That loop is slower than the demo videos suggest, and it is the only version of this that survives contact with a judge. If you want to go deeper on which tools actually surface verifiable sources, we covered that in AI tools that cite real sources.
FAQ
Can AI legal research tools be trusted for citations?
Only the ones grounded in a real case database — CoCounsel and Lexis+AI are the obvious examples — and even then, not without verification. Grounding reduces the odds of a fabricated case appearing in your brief. It does not guarantee that a real case supports the proposition it has been attached to. Read every authority before you cite it.
Is it ethical to use AI on client work?
Using it is generally fine. Using it without competence, supervision, and confidentiality controls is not. The professional duties that already apply — competence, confidentiality, candour to the tribunal, reasonable fees — apply unchanged. Nothing here is a new ethical framework; it is the existing one applied to a new tool.
What is the safest first AI tool for a small firm?
Document automation, not legal reasoning. Something like PandaDoc that templates and routes paperwork carries almost no professional risk and recovers real hours. Start where a mistake is embarrassing rather than sanctionable, then expand.
Can I bill a client for time the AI saved me?
If your fees are based on time actually spent, no. Bill what you worked. Efficiency gains from tooling are a genuinely unsettled area in fee agreements, and the safest position is the transparent one: charge for the time and the judgment you actually applied, and address the question directly in the engagement letter rather than waiting to be asked.
Is ChatGPT safe for legal work?
For general drafting, brainstorming, and rewriting text that contains no client confidences, it is a capable writing assistant. For legal research it is the single most dangerous option available, because it produces confident, well-formatted citations with no grounding in a case database. Never paste privileged material into a consumer tier without terms that explicitly prohibit training on your inputs.
Do I need to tell my client I used AI?
There is no universal rule, and it depends on your jurisdiction, your engagement letter, and how substantive the use is. The practical answer: decide your position in advance and write it down. A firm that has thought about disclosure and can explain its policy is in a very different position from one improvising after a client asks.
Legal work has the worst risk-to-reward ratio of any AI use case, which is exactly why it deserves a considered shortlist rather than a hyped one. Browse the AI legal tools category to compare what is actually available, or start with the boring, high-return end of the problem in AI document tools. If you run a firm and use something that earned its place, submit it — the tools that survive real practice are the ones worth knowing about.



