在智能体主导、自动化成本持续下降的世界中,哪些任务仍需人工处理?

Scale Blog ·

本资料探讨金融自动化工作流中人工审核资源的分配逻辑,涵盖:以美元计价的审核阈值、基于风险调整的置信度边界、预期损失与审核成本的权衡公式、按错误严重程度建模的影响评估、经校准的置信度评分、人工与智能体劳动的产能经济学,以及与杰文斯悖论的历史类比。 阅读 6 条观点,查看支持证据与原始来源。

Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki

理解这篇

6 个要点

综合解读

  1. 金融工作流支持以美元为基础设定人工审核阈值

    在金融领域——尤其是信用回收环节——决策价值与人工审核时间成本均可折算为美元,从而可量化判断人工审核是否具备经济合理性;而受监管或政策强制要求的决策,则依固定规则执行,不进行成本定价。

    支持这项说法 1

    双方成本均已折算为美元,因此人工复核值得介入的临界点可直接确定。

    Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki · 段落 6

    原始摘录
    Both sides are already denominated in dollars, so the boundary at which human review becomes worthwhile can be set directly.
    上下文

    金融领域常属此类,尤其是信用催收场景:一笔欠款要么被成功收回,要么未被收回;一名分析师的一小时要么花在某张账单上,要么花在另一张上。并非所有决策都符合这一逻辑。例如,首席财务官(CFO)对季度财报签字确认,是出于法律强制要求,而非因其复核是发现错误的最低成本方式;该流程中部分决策还受监管规定、内部政策或供应商协议约束,此类情形须按规则执行,而非按成本定价。

    原始上下文

    Finance is often such a domain, particularly credit recovery. A credit is either recovered or it is not, and an analyst’s hour is either spent on one statement or another. Not every decision works this way. A CFO certifies quarterly statements because the law requires it, not because their review is the cheapest way to catch an error, and some decisions inside this workflow are similarly governed by regulation, internal policy, or the vendor relationship. Those are handled as rules rather than priced.

    回到原文语境 →
  2. 人工审核分配应基于风险调整并优先排序

    最大化自动化占比已不再是目标;取而代之的是,应将有限的人工审核产能优先配置于干预后预期价值最高的决策上,并采用基于风险调整的边界——即潜在错误影响越大,所要求的置信度越高。

    支持这项说法 1

    由此得出的边界是风险调整后的动态边界,而非固定阈值:错误潜在影响越大,自主决策所需的置信度也应越高。

    Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki · 段落 8

    原始摘录
    The resulting boundary is risk-adjusted rather than fixed: as the potential impact of an error increases, the confidence required for autonomous action should increase with it.
    上下文

    最大化自动化工作占比已不再是核心目标。关键问题转为:在一个不断扩大的待办任务池中,哪一部分最值得投入人工注意力?我们认为,在金融工作流中,该边界可通过量化方式确定——即以人工复核成本对标其提升决策预期价值所带来的收益。我们以应付账款信用催收工作流为驱动案例构建该框架,并展示如何据此将有限的人工复核资源分配至预期价值最高的干预环节。

    原始上下文

    Maximizing the share of work that is automated is no longer the binding objective. The question is which slice of an expanded universe of work deserves human attention. We argue that in finance workflows this boundary can be located quantitatively, by pricing human review against the expected value of the decisions that review improves. We develop this framework using an accounts-payable credit-recovery workflow as the motivating case and show how it can be used to allocate limited human-review capacity toward the decisions where intervention has the greatest expected value.

    回到原文语境 →
  3. 当预期损失超过审核成本时,应启动审核

    当错误概率与错误影响的乘积(即预期损失)超过人工审核成本时,启动审核在经济上是合理的;未通过该检验的决策,可交由自动化处理,或直接从流程中剔除。

    支持这项说法 1

    P(错误) × 影响(错误) > 复核成本

    Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki · 段落 16

    原始摘录
    P(error) × Impact(error) > Cost(review)
    回到原文语境 →
  4. 错误影响具有不对称性与可逆性差异

    在信用催收中,假阴性(漏判)几乎导致全额损失,而假阳性(误判)仅耗费分析师几分钟时间。错误的影响必须按严重性加权建模,其中严重性反映错误的可逆性,以及后续环节发现该错误的可能性。

    支持这项说法 1

    影响(错误) = s × C,其中 0 < s ≤ 1

    Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki · 段落 22

    原始摘录
    Impact(error) = s × C, where 0 < s ≤ 1
    回到原文语境 →
  5. 按校准置信度与错误成本触发人工复核

    代理与人工工作的边界不应仅由准确率决定,而应基于校准后的置信度,并结合错误引发的财务后果综合评估——该后果需折算为触发人工复核所需的美元金额阈值。

    支持这项说法 1

    本文构建的框架依据经校准的置信度定位该边界,该置信度通过错误所引发的财务后果进行评估,并转化为代理触发人工复核的美元金额阈值。

    Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki · 段落 43

    原始摘录
    The framework developed here locates the boundary based on calibrated confidence, which is evaluated against the financial consequences of error and converted into a dollar threshold at which the agent would escalate.
    上下文

    智能体系统降低知识工作成本,从而扩大了值得执行的工作总量,并使战略焦点从“自动化程度”转向“人类判断在其中的部署位置”。该边界无法仅凭准确率确定,因为人工复核的价值取决于错误发生的概率及其后果。结果是一种将人类注意力导向其创造最大价值之决策的系统,而非采用统一的复核阈值。

    原始上下文

    The reduction in the cost of knowledge work produced by agentic systems expands the volume of work worth performing and shifts the strategic question from the extent of automation to the placement of human judgment within it. This boundary cannot be determined by accuracy alone, because the value of human review depends on both the probability of an error and the consequence of that error. The result is a system that directs human attention toward the decisions in which it creates the greatest value, rather than applying a uniform review threshold.

    回到原文语境 →
  6. 人工专注高严重性决策;代理承接工作量与波动性

    由于人工复核能力增长缓慢——依赖招聘、培训及固定成本投入,而代理能力具有弹性且边际成本极低,因此经济上合理的任务分配框架应将人工团队规模限定在处理最高严重性决策,其余工作量与波动性则交由代理承担。

    支持这项说法 1

    人工复核能力在任一方向上的调整均十分迟缓:新增人力需逐人招聘,培训及机构知识积累需耗时数月,且无论业务淡旺季均需支付固定薪酬。代理能力则具备弹性,可从处理少量文档扩展至数千份文档,而成本结构几乎不变。因此,正确对复核进行定价的框架将倾向于推荐一种配置:人工团队专注最高严重性决策,而代理承担工作量与人工固定团队在经济上无法承载的波动性。

    Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki · 段落 42

    原始摘录
    Human review capacity is slow to adjust in either direction. It is added one hire at a time, requires months of training and accumulated institutional knowledge, and is paid for in quiet periods as well as busy ones. Agent capacity is elastic, scaling from a handful of documents to many thousands without a corresponding change in cost structure. A framework that prices review correctly will therefore tend to recommend a human team aimed at the highest-severity decisions, with the agent absorbing the volume and the variance that a fixed team cannot economically carry.
    上下文

    另一项特性具有普适性。

    原始上下文

    One further property generalizes.

    回到原文语境 →

关键段落6

带明确归属与语境的原文片段。打开原始文本核查出处。

人机协同路由

按校准置信度与错误成本触发人工复核

本文构建的框架依据经校准的置信度定位该边界,该置信度通过错误所引发的财务后果进行评估,并转化为代理触发人工复核的美元金额阈值。

原始摘录
The framework developed here locates the boundary based on calibrated confidence, which is evaluated against the financial consequences of error and converted into a dollar threshold at which the agent would escalate.
上下文

智能体系统降低知识工作成本,从而扩大了值得执行的工作总量,并使战略焦点从“自动化程度”转向“人类判断在其中的部署位置”。该边界无法仅凭准确率确定,因为人工复核的价值取决于错误发生的概率及其后果。结果是一种将人类注意力导向其创造最大价值之决策的系统,而非采用统一的复核阈值。

原始上下文

The reduction in the cost of knowledge work produced by agentic systems expands the volume of work worth performing and shifts the strategic question from the extent of automation to the placement of human judgment within it. This boundary cannot be determined by accuracy alone, because the value of human review depends on both the probability of an error and the consequence of that error. The result is a system that directs human attention toward the decisions in which it creates the greatest value, rather than applying a uniform review threshold.

金融领域以美元计价的审核阈值

金融工作流支持以美元为基础设定人工审核阈值

双方成本均已折算为美元,因此人工复核值得介入的临界点可直接确定。

原始摘录
Both sides are already denominated in dollars, so the boundary at which human review becomes worthwhile can be set directly.
上下文

金融领域常属此类,尤其是信用催收场景:一笔欠款要么被成功收回,要么未被收回;一名分析师的一小时要么花在某张账单上,要么花在另一张上。并非所有决策都符合这一逻辑。例如,首席财务官(CFO)对季度财报签字确认,是出于法律强制要求,而非因其复核是发现错误的最低成本方式;该流程中部分决策还受监管规定、内部政策或供应商协议约束,此类情形须按规则执行,而非按成本定价。

原始上下文

Finance is often such a domain, particularly credit recovery. A credit is either recovered or it is not, and an analyst’s hour is either spent on one statement or another. Not every decision works this way. A CFO certifies quarterly statements because the law requires it, not because their review is the cheapest way to catch an error, and some decisions inside this workflow are similarly governed by regulation, internal policy, or the vendor relationship. Those are handled as rules rather than priced.

基于风险调整的人工审核分配

人工审核分配应基于风险调整并优先排序

由此得出的边界是风险调整后的动态边界,而非固定阈值:错误潜在影响越大,自主决策所需的置信度也应越高。

原始摘录
The resulting boundary is risk-adjusted rather than fixed: as the potential impact of an error increases, the confidence required for autonomous action should increase with it.
上下文

最大化自动化工作占比已不再是核心目标。关键问题转为:在一个不断扩大的待办任务池中,哪一部分最值得投入人工注意力?我们认为,在金融工作流中,该边界可通过量化方式确定——即以人工复核成本对标其提升决策预期价值所带来的收益。我们以应付账款信用催收工作流为驱动案例构建该框架,并展示如何据此将有限的人工复核资源分配至预期价值最高的干预环节。

原始上下文

Maximizing the share of work that is automated is no longer the binding objective. The question is which slice of an expanded universe of work deserves human attention. We argue that in finance workflows this boundary can be located quantitatively, by pricing human review against the expected value of the decisions that review improves. We develop this framework using an accounts-payable credit-recovery workflow as the motivating case and show how it can be used to allocate limited human-review capacity toward the decisions where intervention has the greatest expected value.

产能经济学

人工专注高严重性决策;代理承接工作量与波动性

人工复核能力在任一方向上的调整均十分迟缓:新增人力需逐人招聘,培训及机构知识积累需耗时数月,且无论业务淡旺季均需支付固定薪酬。代理能力则具备弹性,可从处理少量文档扩展至数千份文档,而成本结构几乎不变。因此,正确对复核进行定价的框架将倾向于推荐一种配置:人工团队专注最高严重性决策,而代理承担工作量与人工固定团队在经济上无法承载的波动性。

原始摘录
Human review capacity is slow to adjust in either direction. It is added one hire at a time, requires months of training and accumulated institutional knowledge, and is paid for in quiet periods as well as busy ones. Agent capacity is elastic, scaling from a handful of documents to many thousands without a corresponding change in cost structure. A framework that prices review correctly will therefore tend to recommend a human team aimed at the highest-severity decisions, with the agent absorbing the volume and the variance that a fixed team cannot economically carry.
上下文

另一项特性具有普适性。

原始上下文

One further property generalizes.

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题