top of page

AutoSynthData: Generating Training Data for Enterprise Agents Turns Failures Into a Curriculum

6 days ago
13 min read

ServiceNow CoreAI released AutoSynthData on October 2, reporting gains from nearly 4,000 synthetic tasks across two enterprise-agent experiments. AutoSynthData: Generating Training Data for Enterprise Agents starts with model failures, then converts those weaknesses into executable and verifiable training examples. The conflict is clear: synthetic data can scale quickly, but generated tasks can also teach the wrong behavior.

That makes AutoSynthData more than another system for asking a model to invent prompts. It tries to build a closed training loop around the target agent, its operating environment, and a stronger teacher model. Every accepted task includes an initial system state, a user request, a successful trajectory, and a verifier that evaluates the final state.

ServiceNow says the approach improved a Gemma target model on two EnterpriseOps Gym domains. Yet those gains come from controlled experiments within the same benchmark family that shaped the curriculum. The result pressures teams relying on static, human-written datasets, while leaving external transfer and production reliability unresolved.

AutoSynthData: Generating Training Data for Enterprise Agents Starts With Failure

The important change is that ServiceNow is treating agent failures as instructions for what training data to generate next.

Traditional synthetic-data pipelines often begin with topics, templates, or seed examples. They expand those inputs into larger collections, filter the results, and use the surviving samples for training. That process can create volume without proving that the new data addresses a model’s actual weaknesses.

AutoSynthData reverses that order. It first evaluates a target model inside an operational environment and studies where the model fails. A stronger teacher model attempts the same diagnostic tasks, providing a comparison between unsuccessful and successful behavior.

The system analyzes the capability involved, the required tools, the workflow structure, and the final state that counts as success. It also identifies details that can change without removing the underlying capability. Those observations become what ServiceNow calls sanitized capability specification cards.

These cards guide generation without exposing the original evaluation prompts, entities, trajectories, or verifier details. That separation is intended to reduce direct benchmark leakage. It also forces the generator to create new situations rather than merely paraphrasing test questions.

The task format has three connected parts. A system specification defines policies, available actions, and the initial environment state. A user prompt states the requested outcome, while a verifier decides whether the agent reached an acceptable final state.

This structure matters because enterprise work rarely ends with a text answer. An IT service agent might need to inspect an incident, check an entitlement, update a record, and preserve an audit trail. A fluent response does not prove that any of those actions happened correctly.

The AutoSynthData release describes two generation stages. The target stage creates and validates core examples built around identified capability gaps. The multiply stage produces new variants from accepted target samples.

Each variant receives its own request, entities, initial state, reference trajectory, and verifier. A multiplied sample cannot become the seed for another multiplied sample. That limit is designed to prevent several generations of synthetic expansion from drifting away from the validated core.

The pipeline separates central generation control from environment-specific execution. A controller manages coverage, quality checks, and dataset construction. An adapter runs tools, manages state, replays reference solutions, evaluates agents, and applies deterministic verification.

That separation gives the approach a path beyond one benchmark. A company could theoretically keep the shared controller while writing an adapter for its own systems. However, every new adapter would need accurate tools, realistic state transitions, and reliable success criteria.

The news, therefore, is not simply that ServiceNow generated synthetic tasks. The company constructed a process that chooses tasks according to the current model’s weaknesses. It then moves the training target as those weaknesses change.

That adaptive loop challenges static dataset development. A fixed collection becomes less valuable once a model solves most of it. AutoSynthData instead searches near the model’s current capability boundary, where examples remain difficult but still teachable.

The idea also changes what an evaluation failure represents. Instead of becoming only a score or bug report, the failure becomes raw material for post-training. The same environment can diagnose weaknesses, generate targeted practice, and test the updated model.

That loop is the central claim behind AutoSynthData: Generating Training Data for Enterprise Agents. ServiceNow has not yet shown that it works across unrelated enterprise environments. Still, it has defined a concrete alternative to indiscriminate synthetic-data scaling.

Static Agent Datasets Now Face a Moving Target

AutoSynthData puts pressure on teams that collect broad training data without measuring whether each sample teaches a missing capability.

Enterprise-agent developers have a difficult data problem. Production records contain useful workflow patterns, but they can also include personal information, confidential business data, and inconsistent outcomes. Human-written tasks avoid some privacy concerns, yet creating enough varied and verifiable examples is expensive.

Synthetic data offers scale, but scale alone does not select the right lesson. A model that already handles password-reset requests gains little from thousands of similar examples. It needs tasks exposing unresolved problems, such as policy checks, multi-system planning, and safe refusal.

EnterpriseOps Gym provides the controlled environment used for ServiceNow’s experiments. Its benchmark paper describes stateful tasks where an agent must reason across tools and leave the underlying system in the correct condition. This is different from benchmarks that grade only a final text response.

ServiceNow describes EnterpriseOps Gym as covering eight business domains. The associated environment includes 512 functional tools and 164 interconnected database tables. Those resources support workflows across areas including IT service management, customer service, and human resources.

The broader benchmark contains 1,150 enterprise tasks, according to ServiceNow. They test planning, policy compliance, and state changes across connected systems. The released dataset also gives outside researchers access to the benchmark materials.

Those numbers explain why targeted generation matters. A single workflow can combine several tools, records, policies, and dependencies. Small changes to the initial state can alter which sequence is valid, or whether the requested action should happen at all.

ServiceNow previously reported that providing expert task plans raised performance by 15% to 35% on difficult enterprise domains. That finding suggests planning remains a major constraint, even when an agent can call individual tools correctly. AutoSynthData tries to convert such planning gaps into repeated training opportunities.

The pressure falls first on static evaluation teams. A fixed test set can identify a weakness, but it does not automatically create a curriculum that addresses it. Researchers must still translate failures into diverse examples, valid solutions, and dependable grading logic.

The second pressure falls on general-purpose model providers. Strong benchmark averages can conceal failures caused by local policies, table structures, or workflow rules. Enterprises need evidence that a model can operate inside their particular systems, not only answer questions about them.

The third pressure falls on companies building agent platforms. If adaptive training proves useful, evaluation infrastructure becomes part of model development rather than a final quality gate. Platforms will need reproducible environments, task generation, execution logs, and state-based verification.

That requirement favors organizations with operational simulators or digital twins. Salesforce has pursued a related direction with CRMArena-Pro, which uses simulated enterprise environments to evaluate agents on business workflows. The overlap signals a wider move toward executable, environment-grounded agent testing.

The routes are not identical. A benchmark can compare models without modifying them, while AutoSynthData uses benchmark failures to generate post-training data. One measures capability, and the other tries to move it.

This distinction matters for enterprise buyers. A leaderboard identifies the strongest model under stated conditions. An adaptive curriculum asks whether a cheaper or smaller target model can improve on the organization’s own recurring tasks.

ServiceNow’s reported results make that possibility concrete, but not settled. The target model still completed only a minority of IT service tasks after training. Better than baseline does not mean ready for unsupervised production access.

Knowledge quality also remains a practical constraint. An agent cannot follow a policy that is missing, contradictory, or inaccessible. Teams building a searchable knowledge base still need clear source material before generated training tasks can reflect real work.

AutoSynthData therefore shifts the bottleneck rather than eliminating it. Teams need fewer manually authored variations, but they need a trustworthy environment and precise verification. That exchange becomes decisive when an agent can alter customer, employee, or infrastructure records.

The Mechanism Depends on Executable Tasks and Strict Verifiers

AutoSynthData works only when generated requests, reference solutions, and verifiers agree about what success means.

The pipeline begins by identifying tasks that separate the target model from a stronger teacher. In the reported configuration, ServiceNow favors candidates the target solves no more than once across three trials. The stronger solver must complete them at least twice across three trials.

This filter aims for a useful difficulty band. Tasks that defeat both models offer no reliable demonstration. Tasks that both solve consistently consume training capacity without targeting a clear weakness.

Once a candidate enters the pipeline, AutoSynthData executes its reference trajectory. A trajectory is the sequence of tool calls and actions used to reach the requested state. The verifier then inspects the result against the task’s success conditions.

This positive check asks whether the intended solution actually works. It can expose an invalid initial state, an unavailable tool, a broken action sequence, or an inconsistency between the request and verifier. A plausible-looking example fails if it cannot survive execution.

The pipeline also performs negative verification. It mutates expected outcomes and confirms that incorrect states do not pass. This step matters because a weak verifier can reward an agent that skips required actions or violates an important constraint.

Consider a request to close an IT incident only after confirming resolution with the affected employee. A verifier that checks only the incident status would accept an unsafe shortcut. A stronger verifier would also require the confirmation record and any mandated notes.

The same problem appears when several solutions are valid. A verifier should recognize acceptable outcomes without demanding one exact reference sequence. ServiceNow frames this as completeness, alongside consistency with the request and sound rejection of incorrect behavior.

These requirements create a difficult balance. An overly permissive verifier rewards incomplete work. An overly narrow verifier penalizes legitimate strategies and trains the model to imitate an arbitrary sequence.

Failed candidates enter a bounded critique and repair loop. A critic examines the task construction, initial state, solution, and verification logic. The system applies targeted corrections, reruns the relevant gates, and accepts or rejects the revised candidate.

Repairing an existing candidate can preserve useful work. It also avoids restarting generation whenever one component contains a fixable defect. The retry limit prevents the system from spending unlimited resources on a low-yield task family.

AutoSynthData then reviews quality across the entire batch. Individually valid examples can still form a repetitive dataset. A generator might overproduce familiar workflows while ignoring difficult combinations of policies, tools, or system states.

The controller tracks accepted and rejected samples, repeated patterns, capability coverage, and recurring critique findings. It reduces generation in overrepresented regions and redirects effort toward gaps. This creates feedback above the individual-task level.

The target and multiply stages support that strategy. Target samples establish validated task families around specific capability gaps. Multiply samples vary wording, entities, tool combinations, and environment states without recursively expanding previous variants.

This design reduces one common synthetic-data risk. Recursive generation can magnify small errors as each new sample inherits assumptions from another generated sample. Anchoring all variants to vetted target examples limits that chain.

The approach also gives enterprises a more defensible audit trail. Each training example can be associated with its initial state, intended action sequence, verifier, and validation outcome. That is more useful than a folder containing prompts with no executable context.

However, deterministic checks cannot encode every meaningful quality dimension. A final database state may look correct even when an agent exposed sensitive information along the way. Another trajectory may create unnecessary changes before restoring the expected state.

Execution logs and policy-aware checks remain necessary. ServiceNow’s own agent evaluation tooling emphasizes datasets, execution records, and multiple quality dimensions. AutoSynthData extends that philosophy into training-data production.

The mechanism also depends on the teacher. A stronger model can demonstrate successful behavior, but its actions still reflect the available tools and encoded policies. A teacher that chooses a risky shortcut can propagate that behavior into supervised fine-tuning.

This creates a governance question. Enterprises need to know who defines valid behavior, which policies the environment implements, and how verifier changes are reviewed. Otherwise, automated generation can scale an unnoticed specification error.

AutoSynthData: Generating Training Data for Enterprise Agents is strongest as an argument for executable data. The generator receives attention, but the environment and verifier carry much of the system’s credibility. Without them, synthetic tasks remain convincing stories rather than proven training examples.

The Reported Gains Are Meaningful but Still Narrow

ServiceNow reports clear benchmark improvements, yet the experiments do not establish production reliability or broad transfer.

The first experiment used Gemma-4-26B-A4B-it as the target model in EnterpriseOps Gym’s Hybrid domain. Qwen3.8-27B served as the teacher. AutoSynthData generated 2,000 synthetic training examples in approximately 18 hours.

ServiceNow fine-tuned Gemma with supervised fine-tuning, which trains a model to imitate successful examples. The best reported checkpoint came from the fifth epoch. An epoch represents one complete pass through the training dataset.

Mean Pass@1 improved by 7.2 percentage points, according to the company. ServiceNow characterizes that change as a 35% relative improvement. Verifier success also increased from 63.01% to 68.55%.

The company says the resulting checkpoint closed 59% of the original Pass@1 gap between Gemma and its reference model. These results indicate that targeted synthetic examples affected more than one measurement. They do not reveal how the model behaved outside the tested environment.

The second experiment focused on IT service management. It again used Gemma-4-26B-A4B-it as the target, while DeepSeek-V4.1-Flash served as the teacher. The pipeline produced 1,994 accepted samples over 66 hours.

Mean Pass@1 rose from 18.77% to 27.18% on the ITSM evaluation. That is an 8.41 percentage-point increase. It also leaves the trained model failing most first attempts under the benchmark’s scoring method.

That remaining gap is essential context. The experiment supports the claim that targeted synthetic fine-tuning can improve a model. It does not support a claim that the resulting agent is ready to operate independently in business systems.

ServiceNow says the Hybrid generator never received the original evaluation prompts, entities, trajectories, or verifier details. It received capability specifications distilled from evaluation behavior. This separation reduces one obvious form of test-set contamination.

However, the curriculum still came from failures observed within EnterpriseOps Gym. Training and evaluation therefore shared an environment, tool structure, and general task distribution. Improvement within that environment does not establish transfer to unrelated software or private enterprise configurations.

The reported figures also come from the team that designed the system. Independent replication would strengthen the result. Researchers would need enough code, generation settings, accepted samples, and evaluation details to reproduce the pipeline.

Artificial Analysis now operates an independent leaderboard based on EnterpriseOps Gym. Its evaluation also emphasizes stateful, multi-step work and final database conditions. That external harness creates one possible venue for testing trained checkpoints more independently.

Cross-environment evaluation would be even more informative. A model trained on one ITSM configuration could be tested against changed policies, renamed tools, altered schemas, and unfamiliar record distributions. Performance under those shifts would show whether the model learned a capability or memorized an environment pattern.

Safety needs separate measurement as well. ServiceNow previously described 30 infeasible benchmark tasks involving unavailable resources, missing permissions, or policy violations. Its strongest tested model reportedly recognized only about half as infeasible.

AutoSynthData could target such failures, but the current release does not present a dedicated safe-abstention result. Improving task completion can create new risks if the model also becomes more willing to act when refusal is correct.

The teacher-target relationship deserves scrutiny. The Hybrid experiment used a relatively close model-size pairing, while the ITSM run used a much larger teacher. Generation duration differed substantially, partly because ServiceNow says the earlier ITSM run preceded throughput optimizations.

Those differences make direct comparisons difficult. The two domains involved different teachers, processing times, and likely different capability distributions. The evidence shows repeatability across two settings, not a controlled study of every system component.

The absence of an ablation study also limits interpretation. Public results do not isolate how much improvement came from failure targeting, teacher demonstrations, verifier filtering, task multiplication, or batch-level balancing. Each component sounds plausible, but their separate contributions remain unclear.

Cost and resource use are another open issue, even without attaching commercial figures. Generating thousands of tasks requires repeated model calls, environment execution, solver trials, criticism, repair, and verification. The accepted dataset represents only the output, not the total attempted workload.

Enterprises must compare that workload with alternatives. Human-authored examples may be slower but easier to review. Retrieval and orchestration changes might fix some failures without updating model weights.

A workflow can also fail because its knowledge is incomplete rather than because the model lacks reasoning ability. Better knowledge blending may address some gaps more directly. Training should not become the default response to every unsuccessful agent run.

The cautious reading is still encouraging. AutoSynthData produced measurable improvements from tasks selected around observed weaknesses. The stronger claim, that adaptive synthetic curricula generalize into safer production agents, remains unproven.

Three Signals Will Decide Whether AutoSynthData Matters Beyond the Benchmark

The next evidence should show reproducibility, transfer, and safer behavior, not only another higher score inside the original environment.

The first signal is an independently reproducible release. ServiceNow has published EnterpriseOps Gym materials, but researchers need the AutoSynthData implementation and complete experimental recipe. That package should include capability-card construction, generation settings, verifier tests, rejection criteria, and checkpoint evaluation.

Independent teams should be able to regenerate comparable datasets from the same diagnostic failures. Similar improvements across multiple runs would reduce concerns about favorable sampling. Publishing rejected-task statistics would also reveal how much filtering the final dataset required.

The strongest version of this signal would include ablations. Researchers could remove negative verification, batch balancing, multiplication, or failure targeting one component at a time. The resulting performance differences would identify which mechanisms produce the improvement.

If independent replication succeeds, it would strengthen ServiceNow’s main technical claim. If gains vary sharply across runs, the result would suggest the pipeline remains sensitive to generators, teachers, or selection choices.

The second signal is transfer across environments. A trained model should face workflows that preserve the underlying capability while changing tools, entities, policies, and schemas. That test would distinguish general learning from familiarity with EnterpriseOps Gym’s structure.

A useful experiment could train on one ITSM environment and evaluate on another without additional fine-tuning. Another could move from a single-domain workflow to a cross-domain task requiring customer, employee, and asset data.

Enterprises should also watch for adapters beyond ServiceNow-oriented systems. AutoSynthData’s controller-adapter architecture suggests portability, but software design does not guarantee practical compatibility. Every new environment needs executable actions and dependable verification.

Successful cross-platform results would make the method relevant to a much wider agent market. Weak transfer would narrow its role to an efficient customization pipeline for tightly specified environments.

The third signal is whether task completion improves without weakening refusal and policy compliance. An enterprise agent must know when not to act. Higher Pass@1 scores are insufficient if training encourages confident execution under missing permissions or contradictory instructions.

Future evaluations should report infeasible-task detection, unauthorized state changes, policy violations, and unnecessary tool calls. They should also inspect intermediate actions, not only final database states. A restored final state can conceal an unsafe sequence.

Human review remains important for high-impact workflows. Reviewers should examine samples involving employee records, access controls, customer entitlements, security incidents, and irreversible changes. Automated verifiers can support that process, but they should not define policy without accountable oversight.

The next one to three months should therefore bring three concrete forms of evidence: runnable pipeline code, cross-environment tests, and safety-specific results. Each would address a different uncertainty in the current release.

AutoSynthData: Generating Training Data for Enterprise Agents presents a credible mechanism for turning evaluation failures into targeted practice. Its reported gains show that the idea deserves attention, especially for teams with executable workflow environments.

The larger decision now belongs to enterprise AI builders. They should ask whether their own agent failures can be expressed as reproducible states, valid trajectories, and testable outcomes. If not, generating more tasks will only scale ambiguity.

If those foundations exist, an adaptive curriculum can make evaluation far more useful. It can show what failed, generate focused training, and measure whether the same weakness remains. The next proof must show that this loop survives outside the environment that created it.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page