Skip to content
A senior lawyer reviews a structured case file in a walnut-paneled library while an assistant organizes the source records.

AI and Legal Work

AI Will Process the Data. Lawyers Will Decide What Matters.

AI is strongest when it reads, sorts, extracts, compares, and drafts from a governed record. The human role is to decide which facts matter, resolve uncertainty, and own the action that follows.

I wrote a book about this because the public argument kept collapsing into a job-replacement question. My conclusion is simpler: AI will process the data and amplify the humans. People will decide what matters, accept responsibility for the decision, and act when the record is uncertain. Legal work makes that division especially clear.

The bottom line: AI can read the file, organize it, compare sources, and prepare a useful proposal. A person still has to decide whether the proposal is right, whether the evidence is sufficient, and what should happen next.

This is an operating model grounded in what the research now shows. AI produces strong gains inside a task it handles well. Capable people can still lose ground when they follow a fluent system beyond the edge of its competence.

AI already amplifies professional work

The productivity case is real.

In a randomized experiment with 453 college-educated professionals, Shakked Noy and Whitney Zhang found that people using ChatGPT completed midlevel writing tasks 40% faster while independent evaluators rated the output 18% higher in quality. The largest gains went to participants who began with weaker performance.

A workplace study reached a similar result at larger scale. Erik Brynjolfsson, Danielle Li, and Lindsey Raymond studied the rollout of an AI assistant to 5,172 customer-support agents. Access to the system raised resolved issues per hour by 15% on average. Less experienced and lower-skilled agents benefited most, suggesting that the tool helped spread practices used by stronger workers.

Those are different settings, but they point in the same direction. AI can compress the time spent producing a first pass. It can also make useful patterns available to people who have not yet accumulated years of tacit experience.

That is amplification. The system's value is making the professional's knowledge easier to apply.

The boundary is jagged

The difficult part is deciding where that amplification ends.

Dell'Acqua and his coauthors ran an experiment with 758 Boston Consulting Group professionals using GPT-4. On 18 realistic tasks that fell inside the model's capability frontier, participants using AI completed 12.2% additional tasks and worked 25.1% faster. The quality of their work also improved substantially.

Then the researchers gave participants a business problem outside that frontier. The AI produced an answer that sounded persuasive but pointed in the wrong direction. Participants using AI were 19 percentage points less likely to reach the correct solution.

The striking part is that the tasks did not announce which side of the frontier they occupied. They looked like ordinary professional work. A person had to recognize when the tool was helping and when it was supplying confident noise.

That is the job that remains.

The human contribution starts well before the machine finishes. It includes framing the question, knowing what evidence would change the answer, recognizing missing context, testing the result, and deciding whether the work is ready to carry consequences.

Judgment moves to verification, integration, and stewardship

AI changes where people think.

A 2025 CHI study from Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers and collected 936 examples of AI-assisted work. The researchers found that higher confidence in AI was associated with less reported critical-thinking effort. Greater confidence in one's own ability was associated with greater critical engagement.

The study also found that AI shifted the work of critical thinking. People spent less effort gathering and producing information. Their attention moved toward verifying outputs, integrating them into the task, and stewarding the work toward its goal.

That shift can be productive. It can also become dangerous when "human review" means clicking approve after reading a polished paragraph.

NIST's guidance on human-AI interaction says human roles and responsibilities in decision-making and oversight need to be clearly defined and differentiated. It also warns that human-AI outcomes vary. In some settings, the combination can amplify bias. In others, a well-designed team produces complementary performance.

The useful question is therefore specific: what is the person responsible for noticing, deciding, and recording?

Law has no shortage of data-processing work. A file can contain forms, calls, contracts, medical records, correspondence, names, dates, amounts, citations, and corrections. Reading and organizing that material consumes time before anyone reaches the legal question.

AI is well suited to help with that layer. It can:

  • transcribe a recorded intake call,
  • classify incoming documents,
  • extract proposed names, dates, amounts, and clauses,
  • compare a claimant's answer with a signed record,
  • connect a proposed fact to the page that supports it,
  • identify missing information or conflicting sources,
  • prepare a timeline or first draft, and
  • route an exception to the right reviewer.

These are meaningful tasks. They are also different from deciding whether a fact is material, whether evidence is sufficient, which authority controls, whether a conflict can be cleared, what claim to advance, or what advice to give a client.

The reliability evidence supports that distinction. In a 2024 preregistered evaluation, Stanford RegLab researchers tested leading AI legal-research products and reported hallucination rates between 17% and 33%. The products and underlying models have continued to change, so those figures are a point-in-time result. They still show why a fluent legal answer cannot be treated as a verified one.

The professional boundary is equally direct. ABA Formal Opinion 512 says lawyers may need an appropriate degree of independent verification or review of generative-AI output. It says the tools cannot replace the judgment and experience needed to advise clients, and lawyers may not rely solely on AI for work that calls for professional judgment.

The lawyer remains accountable because the decision is still legal work, even when a machine prepared much of the material underneath it.

Build the workflow around proposals, sources, and decisions

Putting a chatbot next to a pile of files and asking it for the answer is fragile.

A sound implementation separates preparation from authority. AI creates proposals. The record keeps the sources. People review exceptions. Lawyers make the legal decisions.

Stage AI contribution Human contribution Decision owner
Intake capture Transcribe, classify, and organize submissions Correct identity, consent, and communication errors Authorized firm staff under firm policy
Evidence structuring Extract proposed facts and link source locations Confirm, correct, or reject material proposals Trained reviewer
Gap and conflict review Surface missing records and disagreeing values Resolve exceptions and request follow-up Assigned staff, with escalation rules
Legal analysis Prepare source-linked research or a draft issue map Determine relevance, sufficiency, authority, and strategy Lawyer
Consequential action Prepare a filing, communication, or status change Approve the action and accept responsibility for it Authorized person, with lawyer control where legal judgment is required

This model needs a governed record. A legal intake system should distinguish an original answer from an extracted fact, a correction from an overwrite, and an unknown fact from an unanswered question. A mass tort intake data model should preserve the campaign rule version, source, reviewer, and state of each material item.

Without that structure, human review becomes theater. A lawyer sees a summary but cannot reopen the page behind the assertion. Staff see a completed field but cannot tell whether the claimant supplied it or the model inferred it. The firm knows an answer changed but not who changed it or why.

Good review depends on the ability to challenge the machine.

The human role has to be designed, not assumed

Many firms say a person will remain "in the loop." That phrase says nothing about the person's actual authority or attention.

A real human-control design answers five questions:

  1. Which outputs are proposals, and which records may drive later work?
  2. What uncertainty, conflict, or missing evidence forces review?
  3. Can the reviewer open the original source without leaving the task?
  4. Who may correct the record, and does the prior value remain visible?
  5. Which actions stay paused until an authorized person approves them?

These controls belong inside the workflow. A reviewer needs an exception queue. A proposed fact needs a source. A correction needs history. A consequential action needs an approval state. Quality assurance for mass tort intake should measure the errors people find, the overrides they make, and the rules that produced them.

That information also improves the system. Human corrections reveal where extraction rules fail, which documents create ambiguity, and which questions collect unreliable answers. AI amplifies people best when the organization treats their judgment as data for improving the workflow, not as a ceremonial signature at the end.

The future I expect combines far less manual file handling with lawyers who remain fully accountable for judgment.

AI will read a larger share of the record. It will organize the documents, propose the facts, compare sources, prepare drafts, and keep routine work moving. That gives legal teams a better starting point and gives lawyers time to spend on the questions clients hired them to answer.

Humans will decide what matters. Authorized staff will manage the operational gates assigned by firm policy. Lawyers will determine legal relevance, sufficiency, strategy, and advice. The record will show what the machine proposed, what the person changed, and who approved the action.

That is how AI amplifies a legal team without hollowing out the judgment that makes the work legal.

See how OBE structures the record for human review.

Book a demo →

Sources

Artificial intelligenceLegal judgmentLegal operationsCase dataHuman review

Bring one matter type

Show us how your team handles intake today. We'll map it together.

Book a demo