将 Playbook 审查重构为多智能体系统|Harvey

Harvey Blog ·

对用于合同谈判准则(Playbook)实施的多智能体系统的技术评审,涵盖条件逻辑的语义理解挑战、修订标记编辑的‘最轻干预’标准、基于律师定义评分规则的 LLM 评判评估框架,以及用于分布式规则审查与协调的编排者-工作者智能体架构。 阅读 4 条观点,查看支持证据与原始来源。

Maharshi Patel, Pablo Felgueres, Zach Huang

理解这篇

4 个要点

综合解读

  1. 合同条件逻辑需语义理解

    合同包含复杂的条件逻辑:条款跨多页、相互引用,并描述依赖上下文的触发条件。AI 智能体必须理解其含义和后果——而不仅是匹配字词——从而区分句法差异与语义差异。

    支持这项说法 1

    合同具有复杂的条件逻辑,条款跨越多页、相互引用,并描述某些条件何时被触发以及何时不被触发。而且即使在创建了谈判准则之后,对方的合同也几乎永远不会使用你准则中的确切措辞。立场往往受到条件逻辑的限定。因此,审查条款的智能体需要理解这些条件性、句法与语义差异的含义和影响,而不是进行词语匹配。

    Maharshi Patel, Pablo Felgueres, Zach Huang · 段落 13

    原始摘录
    Contracts have complex conditional logic and clauses span pages, reference each other, and describe when certain conditions are triggered and not. And even after creating a playbook, a counterparty’s contract will almost never use your playbook’s exact words. Positions are often qualified with conditional logic. An agent reviewing a clause, therefore, requires comprehending meaning and repercussions of these conditional, syntatic vs semantic differences instead of word matching.
    上下文

    合同包含复杂的条件逻辑:

    原始上下文

    Contracts include complex conditional logic:

    回到原文语境 →
  2. ‘最轻干预’是关键质量标准

    最优的修订标记编辑,是在法律上充分的前提下改动最小。重写整条条款而非精准修改三处词语,会拖慢交易进程并增加人工审查负担;这一‘最轻干预’标准,常被现成 AI 模型忽略。

    支持这项说法 1

    最好的编辑通常是最小的改动:你发出的每一处修订标记都会受到对方律师的仔细审查。如果一个条款只需改动三个词,而智能体重写了整个条款,这既会拖慢交易进程,也会给人类审查者带来额外工作。律师将此称为采取“最轻干预”,这是一个开箱即用的 AI 模型和智能体无法达到的质量标准。

    Maharshi Patel, Pablo Felgueres, Zach Huang · 段落 16

    原始摘录
    The best edit is usually the smallest one: Every redline you send gets scrutinized by an opposing counsel. If a clause needs three words changed and an agent rewrites the entire clause it will both slow down a deal and create extra work for human reviewers. Lawyers call this taking the "lightest touch," and it's a quality bar that out-of-the-box AI models and agents miss.
    回到原文语境 →
  3. LLM 裁判依据律师定义的评分规则对修订标记质量打分

    修订标记质量使用源自律师的评分规则进行评估,这些规则考察最小化程度、位置正确性和法律合理性。由三个前沿 LLM 组成的委员会充当独立裁判对输出进行评分;其投票结果会被汇总。

    支持这项说法 1

    红线质量评估更为复杂:需判断编辑是否属于最小必要改动、位置是否恰当,以及是否符合法律要求。我们与律师共同制定评分细则后,构建了一套评估框架,由大语言模型(LLM)担任评审员,依据细则打分;再组建由三款前沿模型构成的评审委员会,各自独立评分,最后汇总投票结果。

    Maharshi Patel, Pablo Felgueres, Zach Huang · 段落 33

    原始摘录
    Redline quality is more complex in that it requires judgment on whether an edit is minimal, correctly placed, and legally sound . After sourcing the rubrics with lawyers, we set up an evaluation framework that uses LLM judges to score against these rubrics. We then used a committee of three frontier models to score independently and aggregate the votes.
    上下文

    风险标记最适合被理解并评估为一个分类问题——这是一项直接明确的任务。

    原始上下文

    Flagging risk is best understood and scored as a classification problem which is a straightforward task.

    回到原文语境 →

    继续探索

    评估方法 →
  4. 编排者-工作者模式在质量、延迟和复杂性之间取得平衡

    一个编排者智能体将特定规则的审查任务分发给并行子智能体——每个子智能体的作用类似于一名法律助理——并对结果进行协调,包括冲突解决。这在输出质量、系统延迟和工程复杂性之间取得了平衡。

    支持这项说法 1

    我们通过单智能体原型验证了:质量类问题可借助智能体解决。通常,提升延迟表现的最优方式是任务并行化。基于这一思路,我们采用‘协调器-工作节点’架构:主智能体负责分发任务至多个并行运行的智能体;每个并行智能体专注审阅某一条规则;最终,主智能体整合各子任务结果,确保输出质量达标。

    Maharshi Patel, Pablo Felgueres, Zach Huang · 段落 43

    原始摘录
    We proved with a single agent prototype that quality problems can be solved with an agent. Often the best way to improve on latency is to find ways to parallelize work. With that idea in mind, we implemented an orchestrator-worker pattern consisting of a lead agent in charge of distributing work between parallel agents. Each of these agents in turn focus on a specific task, in this case reviewing a rule, and finally the lead agent reconciles the results to ensure final output is of good quality.
    回到原文语境 →

关键段落4

带明确归属与语境的原文片段。打开原始文本核查出处。

合同审查复杂性

合同条件逻辑需语义理解

合同具有复杂的条件逻辑,条款跨越多页、相互引用,并描述某些条件何时被触发以及何时不被触发。而且即使在创建了谈判准则之后,对方的合同也几乎永远不会使用你准则中的确切措辞。立场往往受到条件逻辑的限定。因此,审查条款的智能体需要理解这些条件性、句法与语义差异的含义和影响,而不是进行词语匹配。

原始摘录
Contracts have complex conditional logic and clauses span pages, reference each other, and describe when certain conditions are triggered and not. And even after creating a playbook, a counterparty’s contract will almost never use your playbook’s exact words. Positions are often qualified with conditional logic. An agent reviewing a clause, therefore, requires comprehending meaning and repercussions of these conditional, syntatic vs semantic differences instead of word matching.
上下文

合同包含复杂的条件逻辑:

原始上下文

Contracts include complex conditional logic:

修订标记质量

‘最轻干预’是关键质量标准

最好的编辑通常是最小的改动:你发出的每一处修订标记都会受到对方律师的仔细审查。如果一个条款只需改动三个词,而智能体重写了整个条款,这既会拖慢交易进程,也会给人类审查者带来额外工作。律师将此称为采取“最轻干预”,这是一个开箱即用的 AI 模型和智能体无法达到的质量标准。

原始摘录
The best edit is usually the smallest one: Every redline you send gets scrutinized by an opposing counsel. If a clause needs three words changed and an agent rewrites the entire clause it will both slow down a deal and create extra work for human reviewers. Lawyers call this taking the "lightest touch," and it's a quality bar that out-of-the-box AI models and agents miss.
评估方法

LLM 裁判依据律师定义的评分规则对修订标记质量打分

红线质量评估更为复杂:需判断编辑是否属于最小必要改动、位置是否恰当,以及是否符合法律要求。我们与律师共同制定评分细则后,构建了一套评估框架,由大语言模型(LLM)担任评审员,依据细则打分;再组建由三款前沿模型构成的评审委员会,各自独立评分,最后汇总投票结果。

原始摘录
Redline quality is more complex in that it requires judgment on whether an edit is minimal, correctly placed, and legally sound . After sourcing the rubrics with lawyers, we set up an evaluation framework that uses LLM judges to score against these rubrics. We then used a committee of three frontier models to score independently and aggregate the votes.
上下文

风险标记最适合被理解并评估为一个分类问题——这是一项直接明确的任务。

原始上下文

Flagging risk is best understood and scored as a classification problem which is a straightforward task.

多智能体架构

编排者-工作者模式在质量、延迟和复杂性之间取得平衡

我们通过单智能体原型验证了:质量类问题可借助智能体解决。通常,提升延迟表现的最优方式是任务并行化。基于这一思路,我们采用‘协调器-工作节点’架构:主智能体负责分发任务至多个并行运行的智能体;每个并行智能体专注审阅某一条规则;最终,主智能体整合各子任务结果,确保输出质量达标。

原始摘录
We proved with a single agent prototype that quality problems can be solved with an agent. Often the best way to improve on latency is to find ways to parallelize work. With that idea in mind, we implemented an orchestrator-worker pattern consisting of a lead agent in charge of distributing work between parallel agents. Each of these agents in turn focus on a specific task, in this case reviewing a rule, and finally the lead agent reconciles the results to ensure final output is of good quality.

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题

继续研究相关话题

以下资料涉及本文的话题;讨论相同话题不代表观点一致。