Amazon SageMaker Search Agent Gets Better With Multi-Turn RL, but One Benchmark Slips
Amazon reported that its Amazon SageMaker search agent improved retrieval quality by 23.7 percent on one benchmark after a single training epoch. The fine-tuned Qwen3.6-27B agent also cut its failure rate on that test from 22.89 percent to 0.68 percent. However, it did not improve every evaluation.
That mixed result matters more than a perfect score would. AWS is testing whether companies can teach a smaller model to navigate their search tools instead of repeatedly prompting a frontier model. Success would move part of the agent competition away from model size and toward environment-specific training.
The experiment also exposes the limits of that argument. AWS compared the tuned model with its own untuned base model, not with a current frontier model. Its benchmarks measured retrieval behavior inside a controlled setup, rather than accuracy, latency, and operating cost in a live enterprise deployment.
AWS Put Concrete Numbers Behind Multi-Turn Agent Training
The significant change is not that SageMaker can fine-tune a model, but that AWS evaluated an agent across complete search trajectories.
AWS published the experiment on October 2, 2026, several months after launching SageMaker AI multi-turn reinforcement learning. The company used the service to customize a Qwen3.6-27B model for enterprise-style retrieval.
The agent had access to two search methods. BM25, a lexical retrieval method, finds documents through exact terms and word-frequency patterns. Vector search converts queries and documents into numerical representations, helping match concepts that use different wording.
Choosing between those tools is only one part of the task. The agent must also rewrite weak queries, inspect retrieved material, decide whether another search is useful, and stop before exhausting its limits. Each choice changes the context available for the next one.
That dependency is why AWS used multi-turn reinforcement learning, or MTRL. The technique scores behavior across a sequence of actions instead of treating every model response as an isolated event. SageMaker’s MTRL documentation describes the objective as maximizing cumulative reward across the entire sequence.
For this experiment, the reward was nDCG@10. Normalized Discounted Cumulative Gain at rank 10 measures whether relevant documents appear near the top of the first ten results. A score of 1 represents an ideal ranking, while zero means the system retrieved no relevant document.
AWS applied that score after the agent finished its search trajectory. It also assigned a reward of negative one when the agent exceeded its turn limit or single-turn token budget. That penalty made finishing efficiently part of the learning objective.
The company assembled training material from FRAMES, BRIGHT, Enterprise RAG, ESCI, Musique, and MLQA. These datasets cover multi-hop questions, reasoning-heavy retrieval, enterprise documents, product search, and multilingual comprehension.
The Enterprise RAG benchmark contains more than 500,000 synthetic company documents and 500 questions. AWS held back 5 percent of each training dataset for validation.
Testing used four separate datasets. WixQA covers support questions based on Wix documentation. Wands evaluates product-search relevance, while FreshStack focuses on recent developer questions. BrowseComp-Plus tests difficult research queries against roughly 100,000 human-verified web documents.
That separation is important because it reduces the chance that evaluation simply rewards memorized training examples. It does not eliminate every form of benchmark contamination or distribution overlap. Still, held-out datasets make the reported gains more meaningful than training scores alone.
AWS changed only three exposed settings from their defaults. It ran one epoch, used a global batch size of 128, and allowed 32 concurrent rollouts. A rollout is one sampled attempt by the agent to complete a multi-step task.
The broader service manages trajectory collection, checkpoints, and model updates. It also records rewards and turn-level traces through managed MLflow. According to the original search-agent experiment, training and validation rewards rose before flattening near the end.
That is the event behind the headline. AWS now has a public example where its managed MTRL service altered both retrieval ranking and failure behavior across unseen datasets. The stronger question is what that result does to the default strategy for building agents.
The Frontier-Model Default Is Now Under Pressure
AWS is challenging the assumption that reliability must come from calling the most capable general-purpose model available.
A frontier model often handles unfamiliar tools better than a smaller model because its general reasoning and instruction-following abilities are stronger. That advantage can make it the safest starting point for prototypes. It can also hide weaknesses in the agent’s environment design.
A larger model still does not inherently understand a company’s search indexes, filters, permissions, metadata, or stopping rules. Teams usually describe those details through prompts and tool schemas. They then pay for the model to interpret that guidance during every task.
AWS proposes a different division of labor. The organization defines the environment and success metric, while reinforcement learning turns repeated interaction into model behavior. The resulting model becomes more specialized and less dependent on extensive runtime instructions.
That approach puts pressure on the “prompt plus frontier model” route. The contest shifts from which model knows the most to which system learns the right sequence of local actions. For enterprise search, those actions can be narrow even when the underlying questions are varied.
A support-search agent, for example, may need to recognize error codes, prefer exact matching first, broaden the query when results are sparse, and stop after locating an authoritative procedure. A general model can infer that policy. A specialized model can encode it through training.
The economic case remains a claim rather than a result from this specific evaluation. AWS says smaller specialized models can deliver lower latency and lower inference cost. However, the published table contains no latency measurements, token totals, deployment costs, or direct frontier-model comparison.
The service itself became generally available as a serverless SageMaker customization capability on June 3, 2026. AWS said the launch covered rollout orchestration, trajectory collection, training, checkpoint management, and evaluation. Its launch announcement also listed Qwen3.6-27B, Nova Lite 2.0, GPT-OSS-20B, and Gemma-4-31B-it among supported models and regions.
Serverless training lowers the infrastructure barrier, but it does not remove the application work. Teams still need a callable agent environment, representative tasks, reliable ground truth, and a reward that reflects real success. Poorly chosen rewards can teach a model to optimize the score while missing the business objective.
That requirement favors organizations with measurable workflows. Search quality works well because teams can label relevant documents and calculate ranking metrics. Customer support, code execution, and structured operations can also produce verifiable outcomes.
Open-ended research is harder. A plausible answer might be incomplete, unsupported, or based on a misleading source. One terminal score may not capture those distinctions.
The practical pressure therefore falls on two groups. Model providers must show why a larger general model remains worth using for repeated, bounded tasks. Enterprise AI teams must decide whether their workloads are stable enough to justify training a specialized policy.
For teams building searchable internal systems, that decision begins with data quality rather than model selection. A reliable engineering knowledge base still needs current documents, clear ownership, and traceable sources. Training cannot recover information that the retrieval layer does not contain.
How the Amazon SageMaker Search Agent Learned the Whole Trajectory
The mechanism works because SageMaker rewards the final search outcome while preserving the chain of decisions that produced it.
Supervised fine-tuning teaches a model to imitate examples. For a multi-turn search agent, those examples must show complete, high-quality trajectories. Experts would need to demonstrate useful queries, sensible tool choices, evidence inspection, recovery from weak results, and an appropriate stopping point.
Such demonstrations are expensive to produce. They are also environment-specific. A trajectory created for one document index may be unsuitable for another system with different metadata, retrieval tools, or access controls.
Single-turn reinforcement learning avoids some demonstration costs, but it creates another mismatch. Scoring one output at a time cannot fully capture decisions whose value becomes visible several turns later.
A broad vector query might look unproductive initially. Its results could reveal an exact product identifier that makes the next BM25 query successful. Penalizing the first action independently would ignore its contribution to the completed search.
MTRL keeps that relationship intact. The agent interacts with the environment, receives search results, updates its context, and chooses another action. SageMaker collects the resulting trajectory and uses the final reward to adjust the policy behind those decisions.
AWS’s reward design combined retrieval quality with explicit failure avoidance. nDCG@10 valued the ranking of the final document set. The negative reward discouraged trajectories that exceeded their allowed turns or tokens.
That combination appears central to the largest reliability result. On BrowseComp-Plus, the base agent failed on 22.89 percent of 830 questions. The fine-tuned version failed on 0.68 percent.
The result suggests the model learned when to finish, not just how to retrieve a better ranking. Average turns also declined from 7.0 to 6.3 on that benchmark. Fewer turns can indicate greater efficiency, although the published evaluation does not convert that reduction into latency or cost.
The same pattern did not appear everywhere. On WixQA, average turns rose from 4.3 to 4.5. Wands increased from 2.2 to 2.9. FreshStack decreased from 3.1 to 2.8.
Those differences show why “fewer turns” is not a universal measure of better agent behavior. An extra query can improve evidence coverage, while an early stop can produce a weak answer. The right measure must connect turn count with retrieval quality and task completion.
SageMaker runs rollouts and gradient updates asynchronously. Bounded off-policy staleness limits how far training examples can drift from the model version currently being optimized. The platform also supports PPO, CISPO, and importance-sampling losses with several advantage estimators.
AWS left those choices at their defaults in this experiment. That makes the setup easier to reproduce, but it also obscures which algorithmic choices drove the gains. Users receive a managed path, not a detailed ablation of every training component.
The underlying direction has support beyond this single blog post. The peer-reviewed WebAgent-R1 paper, written by researchers from Amazon and academic institutions, trained agents through end-to-end multi-turn interaction.
On WebArena-Lite, that work increased Qwen2.5-3B task success from 6.1 percent to 33.9 percent. It raised Llama-3.1-8B from 8.5 percent to 44.8 percent. The authors also found that behavior-cloning warm-up affected the results, which complicates claims that reinforcement learning alone solves agent training.
The SageMaker experiment is narrower than WebAgent-R1. It focuses on document retrieval rather than actions across changing websites. Yet both point toward the same mechanism: the model improves when training preserves interactions between actions, observations, and delayed outcomes.
Better Reliability Did Not Mean Universal Retrieval Gains
The strongest evidence supports improved specialization, not a blanket claim that multi-turn RL makes every search task better.
The tuned model improved nDCG@10 on three of the four held-out benchmarks. BrowseComp-Plus rose from 0.5136 to 0.6354, a relative gain of 23.7 percent. WixQA increased from 0.5725 to 0.6781, an 18.4 percent gain.
Wands moved from 0.5762 to 0.6112, a 6 percent improvement. FreshStack went in the opposite direction, slipping from 0.4112 to 0.4089.
That regression is small, but it is analytically important. It prevents the experiment from supporting an uncomplicated “training makes search better” conclusion. The model learned a policy that transferred unevenly across domains.
FreshStack contains recent developer questions derived from Stack Overflow and technical documentation. Those queries can depend on version-specific terminology and rapidly changing facts. A policy trained across broader retrieval datasets may not improve that distribution.
AWS did not publish an error analysis explaining the regression. It remains unclear whether the issue came from tool selection, query rewriting, stale source material, reward alignment, or ordinary evaluation variation.
The four tests also differed substantially in size. They included 147 Wands questions, 400 WixQA questions, 672 FreshStack questions, and 830 BrowseComp-Plus questions. AWS did not report confidence intervals or statistical significance.
That omission does not invalidate the measurements. It limits how confidently readers can generalize the smaller differences, particularly the 6 percent Wands gain and the slight FreshStack decline.
The evaluation also combines failed tasks with retrieval quality by assigning failed runs an nDCG@10 score of zero. That treatment is defensible because a failed agent produces no useful ranked output. However, it means part of the BrowseComp-Plus score increase comes from preventing failures.
The distinction matters for buyers. A system that stops crashing is clearly more useful, even if its successful searches rank documents similarly. Yet reliability improvement and relevance improvement describe different engineering outcomes.
No independent party has reproduced these exact SageMaker results. AWS designed the setup, ran the evaluation, and published the analysis. The report provides detailed benchmark values, but it does not release every training artifact or trajectory needed for a complete audit.
There is also no comparison with supervised fine-tuning, a prompted frontier model, or another managed MTRL platform. The base Qwen3.6-27B agent is the only direct baseline. AWS’s broader claim about frontier-level reliability therefore remains untested here.
Related Amazon research offers more evidence for the general approach, but not independent confirmation. In a separate set of experiments, Amazon researchers reported that reinforcement learning raised a Qwen2.5-32B personal-assistant agent from 39.20 percent to 72 percent task completion.
The same agent customization study reported gains for smaller retrieval models on Natural Questions and Musique. Those results strengthen the case that environment interaction can improve agents. They still come from Amazon-affiliated work and use different models, data, and tasks.
Security and governance create another uncertainty. During training, the agent repeatedly calls real or simulated tools. A poorly isolated environment could expose sensitive data, trigger unintended actions, or reward shortcuts that would be unacceptable in production.
Search systems also change. Documents are added, ranking services are updated, and tool schemas evolve. A policy tuned for one environment can lose accuracy after those changes, even when the model checkpoint remains unchanged.
Organizations will need regression tests for the complete agent system. Model evaluation alone will not detect every failure created by a modified index, permission rule, or tool response.
The evidence therefore supports a bounded conclusion. Multi-turn RL improved this agent’s reliability and most retrieval scores in AWS’s evaluation. It did not establish universal transfer, frontier-model parity, or guaranteed savings.
Agent Competition Is Moving Toward Environments and Rewards
The strategic asset is becoming the training environment that can tell an agent when an entire task succeeded.
Foundation models remain important because they determine the quality of initial rollouts. A weak base model may not discover enough successful trajectories for reinforcement learning to reinforce.
AWS’s own research notes this effect. More capable base models can generate better candidate behavior, creating stronger training signals. Specialization does not erase the value of general model capability.
However, a model alone does not define an enterprise agent. The system also includes tools, indexes, permissions, state, task limits, and evaluation rules. MTRL makes those surrounding components part of the training loop.
That shift changes where companies can build defensible performance. Two organizations can start with the same open model yet produce different agents because their environments and reward functions encode different operational knowledge.
Search is an especially suitable proving ground. Relevance judgments provide measurable feedback, and retrieved documents can be compared against known answers. Search also forces the agent to balance exact matching, semantic retrieval, query reformulation, and stopping behavior.
Other workflows are less cooperative. Sales research may have several acceptable outputs. Investigations can uncover valid evidence that was absent from a benchmark. Knowledge work often values novelty, nuance, and provenance alongside task completion.
A single terminal reward can compress those qualities too aggressively. Teams may need composite evaluations that measure evidence quality, citation accuracy, policy compliance, and action efficiency separately.
The infrastructure race will therefore involve more than managed training. Vendors must help customers construct safe environments, debug trajectories, compare checkpoints, and detect reward hacking.
SageMaker’s MLflow integration addresses part of that need by exposing turn-level traces and rewards. Resumable jobs also help because multi-turn rollouts can take longer than conventional fine-tuning. AWS says its default job limit is 24 hours, although users can adjust it and resume from checkpoints.
Model support and regional availability remain constraints. At the time of the experiment, Qwen3.6-27B was supported for MTRL in the US West, Oregon region. A company with residency requirements or an unsupported model may need a different deployment plan.
The serverless interface also creates platform dependence. AWS manages rollout orchestration and optimization details that a self-hosted team would otherwise control. That tradeoff can accelerate deployment while making low-level experimentation harder.
Open research continues to develop alternatives. Frameworks such as WebAgent-R1 expose more of the training design, including parallel rollouts and context compression. Managed services package similar ideas for organizations that do not want to operate distributed reinforcement-learning infrastructure.
The likely division is not managed versus open in absolute terms. Enterprises will choose based on workload sensitivity, model requirements, engineering capacity, and the value of controlling each training component.
AWS’s advantage is integration. Teams can connect agents hosted through several AWS services, train through SageMaker, inspect traces, and deploy the resulting model to SageMaker endpoints or Amazon Bedrock.
Its remaining challenge is proof. Buyers need evidence that the managed route produces repeatable improvements on their own tasks and remains stable after the environment changes.
Three Signals Will Test AWS’s Smaller-Agent Thesis
The next phase should be judged through external reproduction, production economics, and cross-environment stability.
The first signal is independent replication. Another team needs to reproduce the retrieval gains using disclosed datasets, a comparable Qwen3.6-27B baseline, and the same task limits. Matching the BrowseComp-Plus reliability result would strengthen AWS’s central claim.
A reproduction should separate ranking gains from failure avoidance. It should also include uncertainty estimates and repeated runs. If the improvement disappears under those controls, the current result will look more specific to AWS’s evaluation setup.
The second signal is a direct production comparison with a frontier model. Teams should measure answer quality, retrieval relevance, end-to-end latency, tokens consumed, and intervention rates on the same workload. Training expense should be included alongside inference behavior.
If a specialized model matches the larger model while using fewer resources at deployment, the economic argument becomes concrete. If teams must retrain frequently or maintain complex reward infrastructure, some expected savings will move elsewhere in the stack.
The third signal is performance after the environment changes. A useful test would modify the document corpus, update a tool schema, or introduce a new search filter. Evaluators could then measure whether the agent adapts through prompting or requires another training cycle.
Stable performance would support AWS’s claim that MTRL teaches transferable search behavior. A sharp decline would show that the model learned a narrow interface policy tied closely to its original environment.
Developers should also watch the FreshStack regression. Future experiments need to explain why a policy that helped three datasets slightly hurt this one. A targeted error analysis would reveal whether recency, technical vocabulary, or query strategy caused the difference.
Enterprise buyers do not need to choose between frontier models and specialized agents for every task. A practical architecture can route common, measurable searches to a tuned model and escalate unfamiliar work to a broader model.
That hybrid approach preserves flexibility while testing whether specialization produces reliable savings. It also limits the risk of forcing one policy across every information need.
The Amazon SageMaker search agent result makes that experiment worth running. It shows a measurable reduction in failures and better ranking on most held-out tests. It also leaves enough uncertainty to require evaluation against each organization’s own documents and workflows.
For teams considering this route, the right next action is not immediate deployment. Build a representative test set, define failure precisely, and compare the tuned agent with the strongest existing baseline. Then ask whether the improvement survives new documents, changed tools, and difficult edge cases.



