Insights
Updated September 2026

GenAI in Contract Management: A Reality Check

Compare generative contract tools with research on defined tasks, and test jurisdiction, source accuracy and full review effort before relying on a draft.

4 min read

General information, not legal advice. Legal position as of . Limitations in the Legal Notice

Review status: legal and language review by a named human reviewer is pending.

In this article

Contract lifecycle management, or CLM, connects drafting, approval, signature and ongoing obligations. Some functions use rules and databases; others use machine learning or generative AI. Evaluate the actual function and data flow instead of treating the label “AI” as evidence of capability.

A useful reality check from research

Giampieri's 2024 study asked two chatbots to draft supply contracts in English for England and Wales and for Italy. The small comparison found failures to adapt to the legal system and incomplete or inapplicable clauses. It is a limited drafting study, not a current ranking of all models. It challenges the assumption that naming a jurisdiction in a prompt reliably produces a suitable contract.

The 2025 ContractEval preprint likewise reports trade-offs between correctness and other output measures in clause-level evaluation. Its benchmark and tested configurations do not establish performance on a Swiss firm's contracts. More elaborate output should not be confused with a more accurate result.

Separate the proposed tasks

A drafting assistant can propose text for review. An extraction tool can suggest clauses, dates and obligations. A CLM workflow can route an approval or alert an owner based on recorded data. Each task needs its own acceptance criteria and failure checks.

Drafting

Possible contribution: Propose text from approved instructions and precedents

Question to test: Are legal assumptions, exceptions and defined terms correct?

Review

Possible contribution: Compare with a specified clause library

Question to test: Which material deviations or omissions are missed?

Negotiation

Possible contribution: Organise approved positions and alternatives

Question to test: Does the suggestion reflect the actual mandate and relationship?

Monitoring

Possible contribution: Prepare reminders from recorded obligations

Question to test: Are triggers, owners, amendments and dependencies accurate?

Neither traditional CLM nor GenAI automatically knows market practice. A market comparison needs an identified, relevant and reasonably representative dataset. Public templates and model-generated assertions are not evidence of what counterparties negotiated in your sector.

Private deployment still needs governance

Private hosting may help control a data flow, but does not eliminate confidentiality or security work. Examine model access, support access, logs, backups, training reuse, retention and deletion. A system using historical contracts may reproduce outdated clauses or retrieve material that the current user should not see.

Decide whether the task needs a clause library, retrieval from approved sources, model training or a simpler search tool. These are different implementation choices. Size the infrastructure to the chosen task rather than to a general expectation that every firm needs its own model. Compare full costs and test the proposed configuration.

Review is not automatically the safer starting point

The risk depends on how the output is used. A draft clearly awaiting expert review may pose less reliance risk than a “no issues found” report that causes a material clause to be missed. Choose the pilot by consequences, available evidence and review capacity, not by a universal rule to start with review or drafting.

Use representative authorised documents, including difficult examples. Record preparation, source checking, correction and supervision alongside processing time. A well-structured prompt can make instructions clearer; it cannot guarantee legally correct output or a fixed productivity improvement.

If an agreement uses executable clauses or a blockchain component, identify the external data it relies on, failure handling, amendment authority and dispute process. Code executing a condition does not itself resolve contractual interpretation. Consider confidentiality and data retention before putting sensitive information into systems that are difficult to change.

Plan staff responsibilities and training around observed work. Do not forecast job losses, lower fees or better contract outcomes from a demonstration. Evaluate those outcomes separately and involve the people who perform the current process.

Contract AI evaluation

0/6

Key Takeaway

Contract AI should earn its place through evidence on a defined task. Private hosting, fluent drafting and confident market comparisons are not substitutes for that evidence.

Design work people can sustain

Understand role impacts, test total effort and prepare a workable transition.

You might also like

Need clearer footing for an AI decision?

Start with a focused conversation about a live AI use case, workflow bottleneck, training need, or governance gap.