AutoSynthData: Generating Training Data for Enterprise Agents

Hugging Face Blog ·

The material describes AutoSynthData as a method for generating synthetic training data for enterprise agents, emphasizing environment-specific adaptation, feasibility as a core property of agentic tasks, and difficulty-calibrated task selection to target current agent weaknesses. Read 3 viewpoints with supporting evidence and source links.

Esakkivel Esakkiraja, Shruthan Radhakrishna, Denis Akhiyarov, Sagar Davasam

Understand this piece

3 key points

Synthesis

  1. Difficulty-calibrated task selection

    For training, tasks should expose current agent weaknesses—i.e., be feasible and realistic but not yet consistently solved. Reliably solved tasks provide little new training signal.

    Supporting evidence 1

    Original excerpt

    Difficulty. For training, the task should expose a weakness of the current agent. Tasks that are already solved reliably provide little new training signal. The useful region is therefore tasks that are feasible and realistic, but not yet consistently solved.

    Esakkivel Esakkiraja, Shruthan Radhakrishna, Denis Akhiyarov, Sagar Davasam · Paragraph 14

    Read in source context →

    Continue exploring

    curriculum design →
  2. Feasibility as a core task property

    A useful agentic task must be feasible. At least one trajectory in the current environment must satisfy the user prompt while respecting the system specification. This excludes tasks that require unavailable tools, inaccessible knowledge, impossible state transitions, or policy-prohibited actions.

    Supporting evidence 1

    Original excerpt

    Feasibility. There should exist at least one trajectory in the current environment that satisfies the user prompt while respecting the system specification. This rules out tasks that depend on unavailable tools, inaccessible knowledge, impossible state transitions, or actions prohibited by policy.

    Esakkivel Esakkiraja, Shruthan Radhakrishna, Denis Akhiyarov, Sagar Davasam · Paragraph 12

    Read in source context →

    Continue exploring

    synthetic data quality →
  3. Enterprise agents require environment-specific training

    Enterprises need agents that work well in their own environments—shaped by their systems, rules, and data state. A broadly capable model may still struggle with specific workflows, tool combinations, or constraints unique to that environment.

    Supporting evidence 1

    Original excerpt

    Enterprises need agents that work well in their own environments. The work they ask these agents to do is shaped by the systems they use, the rules they follow, and the state of their data. A model may be broadly capable and still struggle with a particular environment: a workflow it handles poorly, a combination of tools it misuses, or a constraint it fails to respect. Those are the weaknesses an enterprise needs to improve.

    Esakkivel Esakkiraja, Shruthan Radhakrishna, Denis Akhiyarov, Sagar Davasam · Paragraph 1

    Read in source context →

Key passages3

Attributed passages with the context to verify them. Open the original text to check the source.

curriculum design

Difficulty-calibrated task selection

Original excerpt

Difficulty. For training, the task should expose a weakness of the current agent. Tasks that are already solved reliably provide little new training signal. The useful region is therefore tasks that are feasible and realistic, but not yet consistently solved.
enterprise agent training

Enterprise agents require environment-specific training

Original excerpt

Enterprises need agents that work well in their own environments. The work they ask these agents to do is shaped by the systems they use, the rules they follow, and the state of their data. A model may be broadly capable and still struggle with a particular environment: a workflow it handles poorly, a combination of tools it misuses, or a constraint it fails to respect. Those are the weaknesses an enterprise needs to improve.
synthetic data quality

Feasibility as a core task property

Original excerpt

Feasibility. There should exist at least one trajectory in the current environment that satisfies the user prompt while respecting the system specification. This rules out tasks that depend on unavailable tools, inaccessible knowledge, impossible state transitions, or actions prohibited by policy.

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue