The short version
Acceptance criteria as a measurable threshold on a named test set. Ownership of prompts, configuration, evaluation sets and code. Data handling including training and retention. Warranty scope. Liability caps you have actually looked at. Maintenance priced separately. And an exit that hands over everything needed to continue without the supplier.
General guidance on how these agreements are usually structured, not legal advice. Anything material should be reviewed by a lawyer in your jurisdiction.
Why do acceptance criteria matter more than anything else?
Because software either compiles or it does not, while an AI system is right a certain proportion of the time — and without an agreed proportion, “does it work?” becomes a matter of opinion at exactly the moment opinions diverge.
A usable acceptance clause has four parts: the metric, the threshold, the test set, and the exception path. For example: correct extraction of the invoice total on at least 95% of a 500-invoice sample drawn from the previous quarter and agreed by both parties before build, with all remaining cases flagged for human review rather than silently passed.
Note that the test set is agreed before the build. A sample assembled afterwards by whoever needs the outcome is not evidence, and both sides know it.
Who should own the prompts and configuration?
You should, and it needs saying explicitly because the default in many templates is that the supplier retains tooling and methodology. That is reasonable for their general frameworks. It is not reasonable for the thing you paid to have built.
Name the artefacts rather than relying on a general IP clause: prompts and system instructions, configuration, integration code, documentation, and — the one people forget — the evaluation set and its expected answers. The evaluation set is what proves the system works, and rebuilding one from scratch is a substantial cost. Without it, only the original supplier can safely change anything.
What should the data clause cover?
Four things, and none of them is standard boilerplate:
- What they may access, scoped to the systems in the brief rather than granted generally.
- Whether anything may be used for training — theirs or a third party's — which should ordinarily be a flat no, stated as such.
- Retention and deletion: how long copies persist, and confirmation in writing when they are destroyed.
- Case study rights. Most practitioners want to write up the work and most clients will agree if asked. Decide it here, with anonymisation terms, rather than negotiating after delivery.
If personal data is involved, a data processing agreement sits alongside this and is a separate requirement rather than an alternative.
What did they actually warrant?
Read this clause slowly, because there is a large difference between two things that look similar. A warranty that the system will meet the agreed acceptance criteria on the agreed test set is a promise about outcome. A warranty that services will be performed with reasonable skill and care is a promise about effort.
The second is common, entirely legitimate, and means that a system which does not work well enough may still have been delivered in accordance with the contract. Neither is wrong; what matters is knowing which one you have bought, and pricing accordingly.
How should liability be capped?
Caps are normal and competent suppliers will not accept unlimited exposure. The question is whether the cap bears any relation to what could go wrong.
A cap at the value of the fees is the most common formulation, and on a modest project it means the available recovery is modest too. That is acceptable when the system supports a human decision. It is worth renegotiating — or the project rescoping — when the system acts unsupervised on something consequential, because the exposure and the recovery have separated.
The practical move is usually not demanding a bigger cap. It is keeping a human on the decisions where the loss would be large, so the exposure never gets there.
What about maintenance?
Price it separately and put it in the same document. Models change, systems change, exception rates drift, and something must be somebody’s responsibility. A build contract with no maintenance provision produces one of two outcomes: the practitioner does it unpaid until they stop, or the system silently degrades until someone notices it is no longer trusted.
A modest retainer covering monitoring, a defined number of hours, and a response time is easy to agree at signature and awkward to introduce six months later when something has broken.
What should happen at the end?
A defined exit, whether the engagement ends well or badly. Handover of everything listed in the ownership clause, in a usable form. Credentials revoked on a named date. Data deleted with written confirmation. And documentation good enough that a different practitioner could take it on — which is the real test, and the one worth naming in the contract.
Suppliers occasionally resist that last requirement, and the resistance itself is informative. A practitioner confident in their work does not rely on being the only person who can maintain it.
Is a small engagement worth papering properly?
Yes, and disproportionately so, because small engagements are where the informality happens: a screen share, a folder of exports, a login borrowed for an afternoon, and no document anywhere.
The proportionate version is short — two pages covering scope, acceptance, ownership, data and exit. That is not a legal production, it is a written record of things both sides already believe. The value shows up on the day someone remembers them differently.
Who writes the acceptance test set?
Both of you, and it is worth insisting on the collaboration rather than accepting a set the supplier assembles alone. They know what will be technically representative; you know which cases actually matter to the business and which failures would be embarrassing.
Agree the sample size, how it is drawn, and what the correct answers are before build begins. Then keep a copy. Whoever holds the only copy of the test set effectively controls whether the system is judged to work.
What happens if it fails acceptance?
Write it down, because “we will work it out” is how these end badly. The usual structure is a remediation period at the supplier’s cost, a defined number of attempts, and then a specified outcome — partial payment, a reduced scope, or termination with the artefacts handed over.
That last clause is the one worth arguing for. A project that fails acceptance and leaves you with nothing has cost you the fee and the elapsed time; one that leaves you with the integration work, the evaluation set and the documentation leaves the next practitioner a starting point.