The contract intelligence report

Building Your Contract Intelligence DNA

The Overlooked Prerequisite for Contract AI Automation

Eric Hall · CEO, LegalSifter

Free to read. Free to download. No form required.

Same AI. Deeper playbook.
+38 points

What playbook depth is worth

Edit fidelity improves from 55% to 93%With the contracts, engine, and benchmark held constant, a position-only playbook scored 55%; a playbook with instructions, preferred language, and examples scored 93%.Position statement only55%All four knowledge layers93%0%100%

Edit fidelity in LegalSifter’s controlled NDA study.
95%Detection accuracy with all four playbook layers
93%Right-action accuracy with all four layers
30 NDAs16 playbook positions · 3 runs per configuration

01 / The findings

What playbook depth is worth

Our data science team measured the value of playbook depth directly. We took a fixed set of 30 NDAs and a gold standard of edits made by contract lawyers, held the contracts, the engine, and the benchmark constant, and varied one thing: how much knowledge was written into the playbook. Each layer was scored on three dimensions across every position in the playbook: detection, right action, and edit fidelity.¹

LegalSifter Data Science Team, internal study, July 2026. 30 representative NDAs · 16 playbook positions · three runs per configuration · scored against an attorney-curated gold standard. Accuracy is reported separately for detection, action, and edit language rather than as a single blended figure. See Methodology & Notes, note 1.

Detection

Finds the provision, or correctly notes its absence.

Position only91%
+ Written instructions93%
+ Preferred language93%
+ Examples95%
+4 percentage points

Right action

Chooses edit, insert, delete, or leave alone.

Position only77%
+ Written instructions92%
+ Preferred language92%
+ Examples93%
+16 percentage points

Edit fidelity

Similarity of the redline to the lawyers’ own edits.

Position only55%
+ Written instructions83%
+ Preferred language88%
+ Examples93%
+38 percentage points

All bars use a 0–100% scale. Results are separate metrics, not a blended accuracy score. Each layer includes the preceding layers.

+28 pts
Written instructions are the largest single lever in the study because they provide the logic and rationale that motivates the position.

02 / The knowledge layers

Building your knowledge layer: Your contract DNA

It helps to begin with the end in mind and commit to building the playbook over many reviews.

01 / POSITION

Position statement only.

One line of direction, such as “Ensure the agreement contains standard exceptions to the definition of Confidential Information.”

02 / INSTRUCTIONS

Add written instructions.

A few sentences that state what the position means in practice. Are there exceptions or limits to the position? What are your preferred actions when the counterparty’s language does not meet your standards?

03 / LANGUAGE

Add preferred language.

The engine works toward your wording rather than generating a fix of its own.

04 / EXAMPLES

Add examples, including what to leave alone.

A handful of paired examples teach the distinction between a clause that is already complete and one missing a single carve-out.

Why the examples earn their keep

The first is complete, and the example marks it no action required. The second is missing the independent-development carve-out, and the example shows the fix: that one exception threaded into the existing list, everything else untouched.

The “leave it alone” example keeps the engine from over-editing language that was already acceptable.

03 / Putting it to work

But the lawyers are too busy

Sounds interesting but your lawyers are still too busy (or not interested) to take on this sort of work. So how can the average company make progress? Operationally, many organizations will benefit from putting the knowledge base optimization in the hands of a focused AI team and running a simple factory operations model.

01

Run the playbook

Starting with any AI playbook at any level of completeness, the lawyer (or the AI team) runs the playbook and receives a first-pass redline from the AI tool.

02

Review & explain

The lawyer reviews the redline and makes additional changes (as one might edit a junior associate’s work). This “second round” document is sent to the AI team along with a brief note on the rationale and circumstances.

03

Capture & test

The AI team captures the full record (the redline, the rationale, the circumstances) and updates and tests the playbook.

04

Apply & repeat

Next time the same type of contract comes through the system, the updated playbook is applied, and it produces an improved set of redlines.

From the field

Predictability with volume

How many iterations does this take?

4–5

Fully captured cycles

Our field experience with customer playbooks suggests that four to five fully captured review cycles (with the complete knowledge, judgment, and circumstances context) cover the majority of requirements for most contract types.

~15

Reviews in one complex case

Over the course of ~15 total reviews the playbook matures and is achieving ~90 percent alignment with the lawyer’s redlines, with smaller incremental gains as productivity asymptotes.

90 → 5 min

Senior attorney review time

In mature deployments the AI handles 90 percent or more of the review work, and senior attorney review time drops from roughly 90 minutes to 5 minutes per contract. The attorney still reviews, and the review shifts from line-by-line drafting to exception-based oversight.

These are field observations and a customer example reported by the author, separate from the controlled 30-NDA study. They are not universal performance guarantees; outcomes vary with contract complexity and the quality of captured feedback.

Read the full reportBuilding Your Contract Intelligence DNA · Eric Hall · Expand to read the complete text

Building Your Contract Intelligence DNA

The Overlooked Prerequisite for Contract AI Automation

Eric Hall
CEO, LegalSifter

Most companies now accept that scaling AI adoption depends on building a strong knowledge base. Version 1.0 of this idea was simple: assemble your data, feed it to the AI, let the machine find patterns. For high-volume, low-complexity, low-stakes domains, that approach works. Feed the AI thousands of tech support tickets and it can discern the patterns, assemble a working knowledge base, and start resolving issues without substantial human curation. If a fix is not successful, the AI records the failure and moves on to the next suggestion. The feedback loop is fast and the consequences of getting it wrong are small.

For more complex use cases, organizations are finding that the knowledge base requirements are substantially harder to capture. Contracts are a good example of the complexity curve.

The in-house attorney’s description of the complexity curve sounds like this: “Of course we can use AI for our NDAs, but we will never get there with our sales contracts. There are just too many variables.” The instinct has a sound basis: many contract terms turn on the nuances of the commercial context and carry a dense set of legal issues. A simple term like jurisdiction preference is easy to identify and codify, while a complex commercial provision depends on deal context that rarely gets written down.

I think the conclusion that AI will never work for complex contracts is incorrect. The obstacle is that most organizations have not yet built the knowledge base the AI depends on.

Two sub-optimal adoption patterns

Adoption Pattern #1 — The Chicken and the Egg. Many organizations considering AI for contract review are hitting a chicken-and-egg dilemma. I often hear from customers and prospects, “I don’t have the detailed playbook to get what I need from the AI. My team doesn’t have time to build that playbook. So how do we make progress?”

For these organizations, AI adoption stalls at the simplest contracts. They get value from AI on simple, high-volume contracts (NDAs, standard procurement terms) and then hit a ceiling on everything else. The complex contracts that consume the most attorney time and carry the most risk remain fully manual. The internal consensus settles on the view that complex contracts sit beyond the reach of AI tools.

Adoption Pattern #2 — The Co-Pilot Plateau. Another adoption pattern is capturing early productivity gains without building scale. Many AI tools can provide an easy win with productivity gains and the appearance of AI adoption. In this bucket, the lawyers use AI as a co-pilot but ultimately retain significant manual control of the work product. They have moved their work from manual drafting/editing to iterating with the AI tool, but their actual productivity gain plateaus somewhere in the 35 percent to 40 percent range, depending on the contract complexity. They are happy to be “using AI,” but do not have a plan for continuous improvement or long-term automation.

Both adoption patterns cap the return on AI at an early stage, and neither one gets better with continued use. Long-term productivity and automation require investment in the knowledge layer.

Building your knowledge layer: Your contract DNA

Tackling contract complexity and reaching maximum review automation both require depth and quality in your knowledge layer – your contract playbooks (also referred to as negotiation guides or contracting guidelines). Contract playbooks codify the company’s stance for every contract type: the positions, preferred language, fallback positions, rationale, instructions, and examples. Codifying dozens of contract terms and positions sounds hard and time-consuming. Most lawyers and in-house teams are too busy to tackle this burden and fall into one of the adoption patterns above.

It helps to begin with the end in mind and commit to building the playbook over many reviews. The initial contract playbook might focus on a handful of key contract terms with limited instructions and examples. Every time the lawyer reviews a contract, new positions can be documented and more examples and nuanced instructions captured. As the playbook develops, the redlines the AI produces improve measurably. As the playbook becomes more robust, two positive things are happening: (1) the scope of the AI review grows, so the human is required to manually search for fewer concepts and (2) the quality of the AI’s redlines improves, so the human has less clean-up work to do and the human begins to trust the AI more.

What playbook depth is worth

Our data science team measured the value of playbook depth directly. We took a fixed set of 30 NDAs and a gold standard of edits made by contract lawyers, held the contracts, the engine, and the benchmark constant, and varied one thing: how much knowledge was written into the playbook. Each layer was scored on three dimensions across every position in the playbook: detection, right action, and edit fidelity.¹

How four playbook layers impact three accuracy metrics

Accuracy by cumulative playbook layer (%)
Playbook layer Detection Right action Edit fidelity
Position only 91% 77% 55%
+ Written instructions 93% 92% 83%
+ Preferred language 93% 92% 88%
+ Examples 95% 93% 93%

Methodology

LegalSifter Data Science Team, internal study, July 2026. 30 representative NDAs · 16 playbook positions · three runs per configuration · scored against an attorney-curated gold standard. Accuracy is reported separately for detection, action, and edit language rather than as a single blended figure. See Methodology & Notes, note 1.

Detection — found the provision, or correctly noted its absence. Right action — chose correctly among edit, insert, delete, or leave alone. Edit fidelity — similarity of the resulting redline to the lawyers’ own edits.

1. Position statement only. One line of direction, such as “Ensure the agreement contains standard exceptions to the definition of Confidential Information.” The engine locates the clause, applies its own interpretation of which exceptions matter, edits accordingly, and moves on. Detection holds at 91 percent, and the action falls apart at 77 percent, because a bare instruction poses a yes/no question about something that is not a yes/no matter.

2. Add written instructions. A few sentences that state what the position means in practice. Are there exceptions or limits to the position? What are your preferred actions when the counterparty’s language does not meet your standards? Right action gains 15 points and edit fidelity gains 28. Written instructions are the largest single lever in the study because they provide the logic and rationale that motivates the position.

3. Add preferred language. In the example of exceptions to the definition of Confidential Information, the preferred language would address each carve-out. This includes the exact words you want, drafted as a fragment to be included in a list, plus a full fallback clause for the case where the exceptions are missing altogether. The engine works toward your wording rather than generating a fix of its own. Detection and right action hold steady and edit fidelity gains another 5 points to reach 88 percent.

4. Add examples, including what to leave alone. A handful of paired examples teach the distinction between a clause that is already complete and one missing a single carve-out. Detection reaches 95 percent, right action 93 percent, and edit fidelity 93 percent, a 38-point gain in edit fidelity over the position-statement-only playbook.

Why the examples earn their keep

Two clauses, nearly identical:

“…any information that was previously known to the Recipient, independently developed by the Recipient, or publicly available shall not be considered Confidential Information.”

“…any information that was previously known to the Recipient, or publicly available shall not be considered Confidential Information.”

The first is complete, and the example marks it no action required. The second is missing the independent-development carve-out, and the example shows the fix: that one exception threaded into the existing list, everything else untouched. The “leave it alone” example keeps the engine from over-editing language that was already acceptable.

The impact

The initial playbook equipped only with the list of positions produces reasonably strong (91 percent) accuracy on issue spotting, but the redlines will not reflect the company’s nuanced positions or language (55 percent edit fidelity). The final playbook results in decisions and edits that are very close to the lawyer’s own work (93 percent edit fidelity). The 38-point spread in edit quality comes entirely from how much of the company’s knowledge has been codified. Building that depth does not require a separate authoring project, because each layer can be captured from reviews the legal team is already doing.

Building an effective AI knowledge base requires a commitment to systematic iteration and optimization over time. I believe organizations should treat this kind of adoption maturity as part of daily operations rather than as a one-time project.

But the lawyers are too busy

Sounds interesting but your lawyers are still too busy (or not interested) to take on this sort of work. So how can the average company make progress? Operationally, many organizations will benefit from putting the knowledge base optimization in the hands of a focused AI team and running a simple factory operations model.

Starting with any AI playbook at any level of completeness, the lawyer (or the AI team) runs the playbook and receives a first-pass redline from the AI tool. The lawyer reviews the redline and makes additional changes (as one might edit a junior associate’s work). This “second round” document is sent to the AI team along with a brief note on the rationale and circumstances. A sentence or two alongside the redline is enough. For example:

“Clarified that the client’s right to use our intellectual property ends when the agreement ends.”

“Limited the client’s right to require the removal of content to only violations of law or policy. We can’t accept an unlimited right in the client’s sole discretion.”

“Excluded breaches of confidentiality from the waiver of consequential damages clause, because these damages are most likely, by definition, indirect.”

“Carved out fraud, gross negligence, and intentional misconduct from the limitation of liability clause. We can’t allow the counterparty to limit their liability for such behavior.”

“Emphasized that the customer’s remedies for our breach of warranty are limited to repair, replacement, or refund. We can’t be on the hook for downstream damages of the equipment’s failure.”

The redline plus the rationale and context go back to the AI team. The AI team captures the full record (the redline, the rationale, the circumstances) and updates and tests the playbook. Next time the same type of contract comes through the system, the updated playbook is applied, and it produces an improved set of redlines. The senior lawyer reviews, makes any changes, the feedback loop continues.

Over time, the AI playbook incorporates the nuances of knowledge, judgment, and circumstances. After enough iterations, the senior attorney is seeing fewer and fewer changes required. In mature deployments the AI handles 90 percent or more of the review work, and senior attorney review time drops from roughly 90 minutes to 5 minutes per contract. The attorney still reviews, and the review shifts from line-by-line drafting to exception-based oversight.

Predictability with volume

How many iterations does this take?

Our field experience with customer playbooks suggests that four to five fully captured review cycles (with the complete knowledge, judgment, and circumstances context) cover the majority of requirements for most contract types. For straightforward agreements, this happens quickly. For complex commercial contracts with dozens of negotiation variables, it takes more cycles to fully capture the scope and boost quality with edit examples. Working with a customer on a particularly complex contract, the initial playbook was built based on 5 past redlines and iterated as new contracts were reviewed. Fresh contracts identified additional provisions and nuance to be incorporated as well as providing additional edit examples. Over the course of ~15 total reviews the playbook matures and is achieving ~90 percent alignment with the lawyer’s redlines, with smaller incremental gains as productivity asymptotes.

The path to automation follows a predictable course that you can track as it unfolds. At any point in the process, you can see exactly how many review cycles have been completed for each contract type, how the AI’s accuracy is trending, and how many more cycles are needed to reach the target automation level. The horizon is visible and measurable, driven by the number of iterations and the complexity of the contract type.

The useful planning question for a legal team is how many fully captured review cycles each contract type needs, and whether the team is committed to running them.

The bigger objective

Enterprise AI adoption V1.0 was about a simple value proposition: adopt AI and get productivity gains. Buy the tool, deploy it, save 35-40 percent of your attorneys’ time on routine reviews. For many organizations that was a meaningful first step.

We are at the beginning of the AI adoption curve, and recognizing this changes both the time horizon and the objective: V1.0 targeted short-term productivity, and V2.0 targets long-term automation.

LegalSifter aims to shift review work systematically from human-led to AI-led, with attorneys providing oversight, handling exceptions, and exercising judgment on the matters that require an attorney.

A productivity tool helps your team work faster today, while your contract playbooks are an asset: organizational IP that compounds over time, makes every future review faster and more consistent, and retains institutional knowledge regardless of personnel changes. Your contract playbooks are also tool agnostic. We tested this directly by carrying a playbook authored for one generation of our redlining engine over to an entirely different engine. The instructions, examples, and preferred language transferred essentially unchanged, and the playbook held statistically comparable accuracy before any tuning.² Building your contract DNA protects that investment as the underlying LLM capabilities evolve.

Organizations that take a longer-term view can work toward full automation by systematically building the persistent IP that makes it possible. That work happens one review cycle at a time, through a disciplined commitment to capturing knowledge, judgment, and circumstances with every review.

The organizations that make this commitment will build a valuable asset: a compounding, machine-readable representation of how they negotiate, encoded as permanent institutional IP. That asset is their contract intelligence DNA.

The organizations that stop at V1.0 will keep their co-pilots and their 35 percent productivity gains.

LegalSifter ReviewPro gives legal teams a strong foundation for building and optimizing contract playbooks, and it is fully self-serve and intuitive for anyone on the team. For companies that want an AI team to handle the heavy lifting of iteration, LegalSifter provides the technology, expertise, and methodology to build their contract DNA and move from productivity to automation.

Methodology & Notes

1 LegalSifter Data Science Team, internal study, July 2026. Playbook depth ablation (“What playbook depth is worth”). Measured on 30 representative NDAs across 16 playbook positions, three runs per configuration, scored against an attorney-curated gold standard. Accuracy is reported separately for detection, action, and edit language rather than as a single blended figure.

2 LegalSifter Data Science Team, internal study, July 2026. Engine portability (“The bigger objective”). A playbook authored for the ReviewPro 1.0 engine was run on the ReviewPro 2.0 engine: 16 of 16 detection instructions verbatim, about 130 detection examples, plus edit examples and preferred language, all carried over by deterministic conversion. Accuracy held before any optimization for the 2.0 engine: 95% detection / 93% action / 93% edit similarity, against 97% / 97% / 92% for the 1.0 engine, statistically comparable. Same benchmark as note 1: 30 NDAs, 16 playbook positions, three runs per configuration, attorney-curated gold standard. The playbook, not the engine, is the more important asset.

From productivity to automation

Build your contract DNA with LegalSifter

LegalSifter ReviewPro gives legal teams a strong foundation for building and optimizing contract playbooks, and it is fully self-serve and intuitive for anyone on the team. For companies that want an AI team to handle the heavy lifting of iteration, LegalSifter provides the technology, expertise, and methodology to build their contract DNA and move from productivity to automation.

Ready to put these contract insights into practice?

Book a demo