Independent directory.  Not affiliated with or endorsed by Anthropic. “Claude” and “Claude Certified Architect” are trademarks of Anthropic.
← All articles

Automating a process

AI contract review: what it can and can’t do

It reliably finds things: the termination clause, the liability cap, the auto-renewal buried on page forty, every contract in the drawer with an uncapped indemnity. It compares what it finds against your standard positions and drafts a first-pass redline. What it cannot do is decide which deviations you are willing to live with — and that decision, not the reading, is where legal judgement actually lives.

The short version

Extraction and comparison against a written playbook work well and save real hours. Judgement about acceptable risk does not transfer, because it depends on the counterparty, the deal size and how much you want the business. The hardest part of the project is not technical — it is writing down the playbook that currently exists only in a general counsel’s head.

What does it do well?

Four things, and all four are genuinely useful:

The common thread is that each has a checkable answer. Someone can click the clause and see whether the system read it correctly, which is what makes the output safe to build on.

What can't it do?

Decide what you accept. A liability cap at twelve months’ fees is a problem in one deal and entirely normal in another, and the difference is the counterparty, the value, the switching cost, the relationship and how badly you want the contract signed this quarter. None of that is in the document.

It also cannot tell you what is missing in any deep sense. It will notice that a data processing clause is absent if your playbook says one should be there. It will not notice that this particular arrangement has an unusual exposure nobody has thought about before, because that requires a model of your business rather than a model of contracts.

And it cannot be relied on to verify that a referenced authority or standard exists. Check anything cited against the source, always.

Where does the effort actually go?

Writing the playbook. This surprises everyone and it is the honest answer.

Most organisations have standard positions in the sense that two or three people know them. They are not written down at the level of specificity a system needs: not “we don’t like uncapped liability” but “liability capped at the greater of fees paid in the preceding twelve months or £250,000; uncapped only for confidentiality breach, IP infringement and death or personal injury; anything else escalates to the GC”.

Extracting that is weeks of work with the people who hold it, and it is valuable independently of any technology — firms that do it find their own review becomes faster and more consistent before anything is automated. It is also the part that gets underestimated in every proposal.

How accurate is it?

Good enough to change how the work is done, not good enough to remove the reviewer, and the honest number depends entirely on your documents. Standard-form agreements in a familiar template extract very reliably. Heavily negotiated contracts with defined terms scattered across schedules and amendments are harder, and thirty-year-old scanned agreements are harder still.

Which is why the only accuracy figure worth anything is one measured on your contracts. Take fifty you have already reviewed, run them through, and compare against what your lawyers actually found. That is a week of work and it replaces every vendor claim with a fact.

Should it ever approve a contract?

For a narrow band, arguably yes, and it is worth defining rather than avoiding. A standard NDA on your own template, unamended, below a value threshold, with every clause matching the playbook exactly, is a candidate for automatic approval. That is a real saving because those documents are high volume and low variance.

Everything else routes to a person, and the value is that they arrive with the deviations already highlighted rather than reading from the beginning. The distinction to hold onto is that the system is deciding “this matches the playbook”, which is a factual question, rather than “this risk is acceptable”, which is not.

What about confidentiality?

Contracts are among the most sensitive documents an organisation holds, and often contain the other side’s confidential information under obligations you have signed. Before anything is connected, establish where processing happens, how long material is retained, and whether anything submitted can be used to train a model — in writing, for the specific plan you are on.

Check your own NDAs too. Some contain restrictions on disclosing the agreement to third parties that a supplier processing it arguably engages. It is usually fine and it is much better established before than explained afterwards.

How should we start?

With extraction on contracts you have already signed, which is zero risk and immediately useful. Load the back catalogue, extract the key terms into a structured record, and answer the questions you currently cannot: what renews when, where the uncapped indemnities are, which agreements have governing law you would not choose today.

That project alone often justifies the whole exercise, and it builds the extraction accuracy you need before anything touches a live negotiation. Move to playbook comparison second, once the playbook is written. Redlining last.

Does this replace outside counsel?

It changes what you send them. The routine review that firms bill for at volume — the standard agreements, the first pass, the “is there anything unusual here” question — moves in-house or disappears. What remains is the genuinely difficult work: the unusual structure, the negotiation, the judgement about a risk nobody has priced.

That is a smaller volume of higher-value instruction, which is a better relationship for both sides and a worse one for any firm whose model depends on the volume. It is also the pattern that has followed every previous efficiency in legal services rather than something new.

What about contracts we did not draft?

This is where most of the value is, because reviewing the other side’s paper is slower than reviewing your own. Extraction handles unfamiliar structures reasonably well — it is reading for meaning rather than matching a template — but two things degrade it: defined terms scattered across schedules, and amendments that modify clauses without restating them.

For amended agreements, feed the whole chain rather than the base document, and expect to check anything material by hand. A system reading only the original will confidently report a position that was varied three years ago.

Who should own this internally?

Legal, with procurement or commercial as the heaviest users. The trap is deploying it into procurement without legal owning the playbook, because the playbook is the legal position and it will drift if the people who set standard positions are not the ones maintaining it.

Build in a review cadence from the start. Standard positions change after a bad deal or a regulatory shift, and a playbook that silently reflects last year’s risk appetite is worse than none, because everyone believes it is current.