The situation
A commercial law practice reviewed hundreds of contracts a month — NDAs, supply agreements, service contracts. Most review time went into the routine 90%: reading standard documents page by page to find the handful of clauses that deviate from the firm's positions. Senior lawyers were spending expensive hours on work that was essential but repetitive.
The feasibility sprint
Before any commitment, a two-week sprint answered the only question that mattered: could a model match the firm's own standards? We built an evaluation set from past contracts the partners had already marked up — including the messy scans — and scored the model against their real decisions. The measured numbers, not a demo, made the go decision.
What we built
A first-pass review system tuned on the firm's playbook — its preferred positions, fallback clauses and risk thresholds. Every incoming contract is classified, compared clause-by-clause against the playbook, and returned with deviations flagged, suggested fallbacks inserted, and a confidence score on every finding.
The hard part
Adoption in a profession built on professional skepticism. The system won lawyers over by being honest about uncertainty: it never hides its confidence level, and it makes rejecting a suggestion one click. Rejections became training signal. The turning point came when the model flagged a liability clause a tired human reviewer had missed — the story did more for adoption than any mandate could.
Results that held
First-pass review time fell 82%. Lawyers now read flagged clauses in context instead of full documents — judgment work instead of scanning work. Recall on the firm's evaluation set stays above 99%, re-measured continuously against live cases. Every contract still carries a lawyer's sign-off; the system changed what the lawyer's hour is spent on, not who is responsible.