
AI and Legal Work
AI Will Process the Data. Lawyers Will Decide What Matters.
AI is strongest when it reads, sorts, extracts, compares, and drafts from a governed record. The human role is to decide which facts matter, resolve uncertainty, and own the action that follows.
I wrote a book about this because the public argument kept collapsing into a job-replacement question. My conclusion is simpler: AI will process the data and amplify the humans. People will decide what matters, accept responsibility for the decision, and act when the record is uncertain. Legal work makes that division especially clear.
The bottom line: AI can read the file, organize it, compare sources, and prepare a useful proposal. A person still has to decide whether the proposal is right, whether the evidence is sufficient, and what should happen next.
This is an operating model grounded in what the research now shows. AI produces strong gains inside a task it handles well. Capable people can still lose ground when they follow a fluent system beyond the edge of its competence.
AI already amplifies professional work
The productivity case is real.
In a randomized experiment with 453 college-educated professionals, Shakked Noy and Whitney Zhang found that people using ChatGPT completed midlevel writing tasks 40% faster while independent evaluators rated the output 18% higher in quality. The largest gains went to participants who began with weaker performance.
A workplace study reached a similar result at larger scale. Erik Brynjolfsson, Danielle Li, and Lindsey Raymond studied the rollout of an AI assistant to 5,172 customer-support agents. Access to the system raised resolved issues per hour by 15% on average. Less experienced and lower-skilled agents benefited most, suggesting that the tool helped spread practices used by stronger workers.
Those are different settings, but they point in the same direction. AI can compress the time spent producing a first pass. It can also make useful patterns available to people who have not yet accumulated years of tacit experience.
That is amplification. The system's value is making the professional's knowledge easier to apply.
The boundary is jagged
The difficult part is deciding where that amplification ends.
Dell'Acqua and his coauthors ran an experiment with 758 Boston Consulting Group professionals using GPT-4. On 18 realistic tasks that fell inside the model's capability frontier, participants using AI completed 12.2% additional tasks and worked 25.1% faster. The quality of their work also improved substantially.
Then the researchers gave participants a business problem outside that frontier. The AI produced an answer that sounded persuasive but pointed in the wrong direction. Participants using AI were 19 percentage points less likely to reach the correct solution.
The striking part is that the tasks did not announce which side of the frontier they occupied. They looked like ordinary professional work. A person had to recognize when the tool was helping and when it was supplying confident noise.
That is the job that remains.
The human contribution starts well before the machine finishes. It includes framing the question, knowing what evidence would change the answer, recognizing missing context, testing the result, and deciding whether the work is ready to carry consequences.
Judgment moves to verification, integration, and stewardship
AI changes where people think.
A 2025 CHI study from Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers and collected 936 examples of AI-assisted work. The researchers found that higher confidence in AI was associated with less reported critical-thinking effort. Greater confidence in one's own ability was associated with greater critical engagement.
The study also found that AI shifted the work of critical thinking. People spent less effort gathering and producing information. Their attention moved toward verifying outputs, integrating them into the task, and stewarding the work toward its goal.
That shift can be productive. It can also become dangerous when "human review" means clicking approve after reading a polished paragraph.
NIST's guidance on human-AI interaction says human roles and responsibilities in decision-making and oversight need to be clearly defined and differentiated. It also warns that human-AI outcomes vary. In some settings, the combination can amplify bias. In others, a well-designed team produces complementary performance.
The useful question is therefore specific: what is the person responsible for noticing, deciding, and recording?
Legal work makes the division harder to ignore
Law has no shortage of data-processing work. A file can contain forms, calls, contracts, medical records, correspondence, names, dates, amounts, citations, and corrections. Reading and organizing that material consumes time before anyone reaches the legal question.
AI is well suited to help with that layer. It can:
- transcribe a recorded intake call,
- classify incoming documents,
- extract proposed names, dates, amounts, and clauses,
- compare a claimant's answer with a signed record,
- connect a proposed fact to the page that supports it,
- identify missing information or conflicting sources,
- prepare a timeline or first draft, and
- route an exception to the right reviewer.
These are meaningful tasks. They are also different from deciding whether a fact is material, whether evidence is sufficient, which authority controls, whether a conflict can be cleared, what claim to advance, or what advice to give a client.
The reliability evidence supports that distinction. In a 2024 preregistered evaluation, Stanford RegLab researchers tested leading AI legal-research products and reported hallucination rates between 17% and 33%. The products and underlying models have continued to change, so those figures are a point-in-time result. They still show why a fluent legal answer cannot be treated as a verified one.
The professional boundary is equally direct. ABA Formal Opinion 512 says lawyers may need an appropriate degree of independent verification or review of generative-AI output. It says the tools cannot replace the judgment and experience needed to advise clients, and lawyers may not rely solely on AI for work that calls for professional judgment.
The lawyer remains accountable because the decision is still legal work, even when a machine prepared much of the material underneath it.
Build the workflow around proposals, sources, and decisions
Putting a chatbot next to a pile of files and asking it for the answer is fragile.
A sound implementation separates preparation from authority. AI creates proposals. The record keeps the sources. People review exceptions. Lawyers make the legal decisions.
| Stage | AI contribution | Human contribution | Decision owner |
|---|---|---|---|
| Intake capture | Transcribe, classify, and organize submissions | Correct identity, consent, and communication errors | Authorized firm staff under firm policy |
| Evidence structuring | Extract proposed facts and link source locations | Confirm, correct, or reject material proposals | Trained reviewer |
| Gap and conflict review | Surface missing records and disagreeing values | Resolve exceptions and request follow-up | Assigned staff, with escalation rules |
| Legal analysis | Prepare source-linked research or a draft issue map | Determine relevance, sufficiency, authority, and strategy | Lawyer |
| Consequential action | Prepare a filing, communication, or status change | Approve the action and accept responsibility for it | Authorized person, with lawyer control where legal judgment is required |
This model needs a governed record. A legal intake system should distinguish an original answer from an extracted fact, a correction from an overwrite, and an unknown fact from an unanswered question. A mass tort intake data model should preserve the campaign rule version, source, reviewer, and state of each material item.
Without that structure, human review becomes theater. A lawyer sees a summary but cannot reopen the page behind the assertion. Staff see a completed field but cannot tell whether the claimant supplied it or the model inferred it. The firm knows an answer changed but not who changed it or why.
Good review depends on the ability to challenge the machine.
The human role has to be designed, not assumed
Many firms say a person will remain "in the loop." That phrase says nothing about the person's actual authority or attention.
A real human-control design answers five questions:
- Which outputs are proposals, and which records may drive later work?
- What uncertainty, conflict, or missing evidence forces review?
- Can the reviewer open the original source without leaving the task?
- Who may correct the record, and does the prior value remain visible?
- Which actions stay paused until an authorized person approves them?
These controls belong inside the workflow. A reviewer needs an exception queue. A proposed fact needs a source. A correction needs history. A consequential action needs an approval state. Quality assurance for mass tort intake should measure the errors people find, the overrides they make, and the rules that produced them.
That information also improves the system. Human corrections reveal where extraction rules fail, which documents create ambiguity, and which questions collect unreliable answers. AI amplifies people best when the organization treats their judgment as data for improving the workflow, not as a ceremonial signature at the end.
Bring it home to legal
The future I expect combines far less manual file handling with lawyers who remain fully accountable for judgment.
AI will read a larger share of the record. It will organize the documents, propose the facts, compare sources, prepare drafts, and keep routine work moving. That gives legal teams a better starting point and gives lawyers time to spend on the questions clients hired them to answer.
Humans will decide what matters. Authorized staff will manage the operational gates assigned by firm policy. Lawyers will determine legal relevance, sufficiency, strategy, and advice. The record will show what the machine proposed, what the person changed, and who approved the action.
That is how AI amplifies a legal team without hollowing out the judgment that makes the work legal.
See how OBE structures the record for human review.
Sources
- Noy and Zhang: Experimental evidence on the productivity effects of generative artificial intelligence
- Brynjolfsson, Li, and Raymond: Generative AI at Work
- Dell'Acqua et al.: Navigating the Jagged Technological Frontier
- Microsoft Research and Carnegie Mellon: The Impact of Generative AI on Critical Thinking
- NIST: AI Risk Management and Human-AI Interaction
- Stanford RegLab: Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
- American Bar Association: Formal Opinion 512
Bring one matter type