In An Agentic World Where Automation Gets Cheap, Which Work Is Worth Routing to a Human?

Scale Blog ·

A source examining how to allocate human review in automated financial workflows using dollar-denominated thresholds, risk-adjusted confidence boundaries, expected-loss calculations, severity-adjusted error impact modeling, calibrated confidence scoring, capacity economics of human versus agent labor, and a historical analogy to Jevons’ Paradox. Read 6 viewpoints with supporting evidence and source links.

Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki

Understand this piece

6 key points

Synthesis

  1. Finance workflows enable dollar-based human review thresholds

    In finance—especially credit recovery—the value of a decision and the cost of human time are both denominated in dollars, enabling quantitative determination of when human review is economically justified. Regulatory or policy-driven decisions are handled as fixed rules, not priced.

    Supporting evidence 1

    Original excerpt

    Both sides are already denominated in dollars, so the boundary at which human review becomes worthwhile can be set directly.

    Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki · Paragraph 6

    Context

    Finance is often such a domain, particularly credit recovery. A credit is either recovered or it is not, and an analyst’s hour is either spent on one statement or another. Not every decision works this way. A CFO certifies quarterly statements because the law requires it, not because their review is the cheapest way to catch an error, and some decisions inside this workflow are similarly governed by regulation, internal policy, or the vendor relationship. Those are handled as rules rather than priced.

    Read in source context →
  2. Human review allocation should be risk-adjusted and prioritized

    Maximizing automation share is no longer the objective; instead, limited human-review capacity should be allocated to decisions where intervention yields the greatest expected value, using a risk-adjusted boundary that increases required confidence as potential error impact rises.

    Supporting evidence 1

    Original excerpt

    The resulting boundary is risk-adjusted rather than fixed: as the potential impact of an error increases, the confidence required for autonomous action should increase with it.

    Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki · Paragraph 8

    Context

    Maximizing the share of work that is automated is no longer the binding objective. The question is which slice of an expanded universe of work deserves human attention. We argue that in finance workflows this boundary can be located quantitatively, by pricing human review against the expected value of the decisions that review improves. We develop this framework using an accounts-payable credit-recovery workflow as the motivating case and show how it can be used to allocate limited human-review capacity toward the decisions where intervention has the greatest expected value.

    Read in source context →
  3. Review is warranted when expected loss exceeds review cost

    Human review is economically justified for a decision when the product of error probability and error impact exceeds the cost of review. Work failing this test is a candidate for automation—or exclusion from the pipeline.

    Supporting evidence 1

    Original excerpt

    P(error) × Impact(error) > Cost(review)

    Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki · Paragraph 16

    Read in source context →
  4. Error impact must account for asymmetry and reversibility

    In credit recovery, false negatives incur near-full loss while false positives cost only minutes of analyst time. Error impact must be modeled as severity-adjusted, where severity reflects reversibility and downstream detection likelihood.

    Supporting evidence 1

    Original excerpt

    Impact(error) = s × C, where 0 < s ≤ 1

    Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki · Paragraph 22

    Read in source context →
  5. Route human review by calibrated confidence and error cost

    The boundary between agent and human work should be determined not by accuracy alone, but by calibrated confidence evaluated against the financial consequences of error—converted into a dollar threshold for escalation to human review.

    Supporting evidence 1

    Original excerpt

    The framework developed here locates the boundary based on calibrated confidence, which is evaluated against the financial consequences of error and converted into a dollar threshold at which the agent would escalate.

    Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki · Paragraph 43

    Context

    The reduction in the cost of knowledge work produced by agentic systems expands the volume of work worth performing and shifts the strategic question from the extent of automation to the placement of human judgment within it. This boundary cannot be determined by accuracy alone, because the value of human review depends on both the probability of an error and the consequence of that error. The result is a system that directs human attention toward the decisions in which it creates the greatest value, rather than applying a uniform review threshold.

    Read in source context →
  6. Design human teams for high-severity decisions; agents for volume and variance

    Because human review capacity scales slowly—requiring hiring, training, and fixed costs—while agent capacity is elastic and low marginal cost, an economically sound routing framework sizes human teams for highest-severity decisions and assigns agents to both volume and variability.

    Supporting evidence 1

    Original excerpt

    Human review capacity is slow to adjust in either direction. It is added one hire at a time, requires months of training and accumulated institutional knowledge, and is paid for in quiet periods as well as busy ones. Agent capacity is elastic, scaling from a handful of documents to many thousands without a corresponding change in cost structure. A framework that prices review correctly will therefore tend to recommend a human team aimed at the highest-severity decisions, with the agent absorbing the volume and the variance that a fixed team cannot economically carry.

    Benjamin Chen, Duncan McKeen, Manik Mukherjee, Sara Bolouki · Paragraph 42

    Context

    One further property generalizes.

    Read in source context →

    Continue exploring

    capacity economics →

Key passages6

Attributed passages with the context to verify them. Open the original text to check the source.

human-in-the-loop routing

Route human review by calibrated confidence and error cost

Original excerpt

The framework developed here locates the boundary based on calibrated confidence, which is evaluated against the financial consequences of error and converted into a dollar threshold at which the agent would escalate.
Context

The reduction in the cost of knowledge work produced by agentic systems expands the volume of work worth performing and shifts the strategic question from the extent of automation to the placement of human judgment within it. This boundary cannot be determined by accuracy alone, because the value of human review depends on both the probability of an error and the consequence of that error. The result is a system that directs human attention toward the decisions in which it creates the greatest value, rather than applying a uniform review threshold.

Dollar-denominated review thresholds in finance

Finance workflows enable dollar-based human review thresholds

Original excerpt

Both sides are already denominated in dollars, so the boundary at which human review becomes worthwhile can be set directly.
Context

Finance is often such a domain, particularly credit recovery. A credit is either recovered or it is not, and an analyst’s hour is either spent on one statement or another. Not every decision works this way. A CFO certifies quarterly statements because the law requires it, not because their review is the cheapest way to catch an error, and some decisions inside this workflow are similarly governed by regulation, internal policy, or the vendor relationship. Those are handled as rules rather than priced.

Risk-adjusted human review allocation

Human review allocation should be risk-adjusted and prioritized

Original excerpt

The resulting boundary is risk-adjusted rather than fixed: as the potential impact of an error increases, the confidence required for autonomous action should increase with it.
Context

Maximizing the share of work that is automated is no longer the binding objective. The question is which slice of an expanded universe of work deserves human attention. We argue that in finance workflows this boundary can be located quantitatively, by pricing human review against the expected value of the decisions that review improves. We develop this framework using an accounts-payable credit-recovery workflow as the motivating case and show how it can be used to allocate limited human-review capacity toward the decisions where intervention has the greatest expected value.

capacity economics

Design human teams for high-severity decisions; agents for volume and variance

Original excerpt

Human review capacity is slow to adjust in either direction. It is added one hire at a time, requires months of training and accumulated institutional knowledge, and is paid for in quiet periods as well as busy ones. Agent capacity is elastic, scaling from a handful of documents to many thousands without a corresponding change in cost structure. A framework that prices review correctly will therefore tend to recommend a human team aimed at the highest-severity decisions, with the agent absorbing the volume and the variance that a fixed team cannot economically carry.
Context

One further property generalizes.

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue