Rebuilding Playbook Review as a Multi-Agent System | Harvey

Harvey Blog ·

A technical review of a multi-agent system for contract playbook implementation, describing challenges in semantic comprehension of conditional logic, the 'lightest touch' standard for redline edits, an LLM-judged evaluation framework based on lawyer-defined rubrics, and an orchestrator-worker agent architecture for distributed rule review and reconciliation. Read 4 viewpoints with supporting evidence and source links.

Maharshi Patel, Pablo Felgueres, Zach Huang

Understand this piece

4 key points

Synthesis

  1. Conditional logic in contracts requires semantic comprehension

    Contracts contain complex conditional logic: clauses span pages, reference each other, and describe context-dependent triggers. AI agents must grasp meaning and consequences—not just match words—so they distinguish syntactic from semantic differences.

    Supporting evidence 1

    Original excerpt

    Contracts have complex conditional logic and clauses span pages, reference each other, and describe when certain conditions are triggered and not. And even after creating a playbook, a counterparty’s contract will almost never use your playbook’s exact words. Positions are often qualified with conditional logic. An agent reviewing a clause, therefore, requires comprehending meaning and repercussions of these conditional, syntatic vs semantic differences instead of word matching.

    Maharshi Patel, Pablo Felgueres, Zach Huang · Paragraph 13

    Context

    Contracts include complex conditional logic:

    Read in source context →
  2. Lightest-touch editing is a critical quality bar

    The best redline edit is the smallest legally sufficient change. Rewriting entire clauses instead of making surgical three-word edits slows deals and increases human review burden—a standard called 'lightest touch' that off-the-shelf AI models typically miss.

    Supporting evidence 1

    Original excerpt

    The best edit is usually the smallest one: Every redline you send gets scrutinized by an opposing counsel. If a clause needs three words changed and an agent rewrites the entire clause it will both slow down a deal and create extra work for human reviewers. Lawyers call this taking the "lightest touch," and it's a quality bar that out-of-the-box AI models and agents miss.

    Maharshi Patel, Pablo Felgueres, Zach Huang · Paragraph 16

    Read in source context →

    Continue exploring

    redline quality →
  3. LLM judges score redline quality against lawyer-defined rubrics

    Redline quality is evaluated using lawyer-sourced rubrics that assess minimality, correct placement, and legal soundness. A committee of three frontier LLMs acts as independent judges to score outputs; their votes are aggregated.

    Supporting evidence 1

    Original excerpt

    Redline quality is more complex in that it requires judgment on whether an edit is minimal, correctly placed, and legally sound . After sourcing the rubrics with lawyers, we set up an evaluation framework that uses LLM judges to score against these rubrics. We then used a committee of three frontier models to score independently and aggregate the votes.

    Maharshi Patel, Pablo Felgueres, Zach Huang · Paragraph 33

    Context

    Flagging risk is best understood and scored as a classification problem which is a straightforward task.

    Read in source context →

    Continue exploring

    evaluation methodology →
  4. Orchestrator-worker pattern balances quality, latency, and complexity

    An orchestrator agent distributes rule-specific review tasks to parallel subagents—each functioning like a legal associate—and reconciles results, including conflict resolution. This balances output quality, system latency, and engineering complexity.

    Supporting evidence 1

    Original excerpt

    We proved with a single agent prototype that quality problems can be solved with an agent. Often the best way to improve on latency is to find ways to parallelize work. With that idea in mind, we implemented an orchestrator-worker pattern consisting of a lead agent in charge of distributing work between parallel agents. Each of these agents in turn focus on a specific task, in this case reviewing a rule, and finally the lead agent reconciles the results to ensure final output is of good quality.

    Maharshi Patel, Pablo Felgueres, Zach Huang · Paragraph 43

    Read in source context →

Key passages4

Attributed passages with the context to verify them. Open the original text to check the source.

contract review complexity

Conditional logic in contracts requires semantic comprehension

Original excerpt

Contracts have complex conditional logic and clauses span pages, reference each other, and describe when certain conditions are triggered and not. And even after creating a playbook, a counterparty’s contract will almost never use your playbook’s exact words. Positions are often qualified with conditional logic. An agent reviewing a clause, therefore, requires comprehending meaning and repercussions of these conditional, syntatic vs semantic differences instead of word matching.
Context

Contracts include complex conditional logic:

redline quality

Lightest-touch editing is a critical quality bar

Original excerpt

The best edit is usually the smallest one: Every redline you send gets scrutinized by an opposing counsel. If a clause needs three words changed and an agent rewrites the entire clause it will both slow down a deal and create extra work for human reviewers. Lawyers call this taking the "lightest touch," and it's a quality bar that out-of-the-box AI models and agents miss.
evaluation methodology

LLM judges score redline quality against lawyer-defined rubrics

Original excerpt

Redline quality is more complex in that it requires judgment on whether an edit is minimal, correctly placed, and legally sound . After sourcing the rubrics with lawyers, we set up an evaluation framework that uses LLM judges to score against these rubrics. We then used a committee of three frontier models to score independently and aggregate the votes.
Context

Flagging risk is best understood and scored as a classification problem which is a straightforward task.

multi-agent architecture

Orchestrator-worker pattern balances quality, latency, and complexity

Original excerpt

We proved with a single agent prototype that quality problems can be solved with an agent. Often the best way to improve on latency is to find ways to parallelize work. With that idea in mind, we implemented an orchestrator-worker pattern consisting of a lead agent in charge of distributing work between parallel agents. Each of these agents in turn focus on a specific task, in this case reviewing a rule, and finally the lead agent reconciles the results to ensure final output is of good quality.

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue

Continue with this topic

More sources on topics discussed here. Shared topics do not imply agreement.