Jensen Huang at Dreamforce Backs Salesforce Koa Against Closed Frontier AI
Jensen Huang at Dreamforce delivered a blunt message as Salesforce introduced its first CRM reasoning model on September 15, 2026. Artificial intelligence, he argued, now lets people “know everything and do anything.”
The bigger news was not the slogan. Salesforce unveiled Koa, a specialized model built by post-training NVIDIA Nemotron 3 Super on synthetic CRM scenarios. It runs inside Salesforce infrastructure and gives Agentforce an alternative to closed frontier models.
Until now, Salesforce relied on providers such as Anthropic and OpenAI when an Agentforce workflow required sustained reasoning. Koa changes that dependency without ending those partnerships. Salesforce can now route some enterprise tasks to a model it controls.
That makes the announcement a test of two competing approaches. One sends many business tasks to general-purpose frontier models. The other adapts open weights to a company’s data structures, policies, and recurring workflows.
Koa does not establish that specialized models will win. Salesforce has released internal results, pilot names, and an availability schedule, but little independent production evidence. That gap separates Huang’s expansive Dreamforce vision from what enterprise buyers can verify today.
Jensen Huang at Dreamforce Put Koa Behind the Big Promise
Huang’s sweeping AI message arrived alongside a specific change in Salesforce’s model strategy.
Huang joined Salesforce CEO Marc Benioff at San Francisco’s Moscone Center during the opening day of Dreamforce. He moved into the audience while discussing AI as the next major infrastructure layer.
“With electricity, we could power everything,” Huang said. “With the internet, you can find anything.” He then argued that AI lets society “know everything and do anything.”
The phrasing captured Huang’s broad view of AI, but the accompanying product announcement supplied the practical stakes. Salesforce introduced Koa as its first reasoning model designed specifically for customer relationship management.
A reasoning model spends additional computation working through intermediate steps before producing an answer or taking an action. In CRM, that can mean evaluating account history, policies, permissions, and several tools during one request.
According to the official Dreamforce announcement, Salesforce developed Koa by post-training Nemotron 3 Super. Post-training adapts an existing base model to particular tasks, behavior, and operating constraints.
Salesforce says it created a proprietary synthetic dataset modeled on nearly three decades of CRM deployments. Synthetic data consists of generated examples rather than direct copies of customer records.
The company says those scenarios represent work across more than 14 industries. They cover activities such as lead generation, opportunity qualification, service-case resolution, and policy-aware tool use.
Each scenario paired a simulated professional with a task and the sequence of actions required to finish it. That design aims to teach the model how enterprise work proceeds, not merely how CRM vocabulary sounds.
Salesforce also says no customer data entered Koa’s training corpus. Huang’s appearance therefore linked an ambitious claim about universal AI capability to a deliberately narrow enterprise model.
That contrast matters. Koa is not designed to answer every question or outperform every frontier model. It is intended to reason reliably within Salesforce workflows while remaining under Salesforce’s operational control.
The company plans to expose Koa through several Agentforce entry points. Customers can select it for an organization, assign it to individual agents, or use it through the managed model catalog.
This flexibility makes routing central to the announcement. A company does not need to choose one model for every task. It can reserve expensive general intelligence for unusual problems and use Koa for repeatable CRM work.
That is the concrete meaning behind Huang’s “do anything” message. The model must connect reasoning to governed actions, including updating an opportunity, routing a case, or scheduling a follow-up.
The announcement also narrows what “know everything” can responsibly mean. Koa only receives the records, grounding data, and instructions that a customer chooses to share.
It is not omniscient, and Salesforce does not present it that way in the technical product details. Its proposed advantage comes from relevant context, specialization, and controlled access.
That distinction turns a theatrical keynote moment into a consequential platform decision. Salesforce is no longer only connecting outside intelligence to its applications. It is building a reasoning layer that it can operate itself.
Salesforce Koa Challenges Dependence on Closed Frontier Models
Koa pressures closed model providers by giving Salesforce control over a larger share of enterprise reasoning.
Before Koa, complex Agentforce requests could pass through Salesforce’s AI gateway to models such as Claude or ChatGPT. The gateway chooses a suitable model for each request.
Salesforce already had smaller models for focused tasks. However, Jayesh Govindarajan, executive vice president of Salesforce AI, said sustained reasoning remained dependent on frontier providers.
“Reasoning has always been something that we’ve relied on the frontier model providers for,” Govindarajan told TechCrunch. “Until now.”
That final phrase explains why Koa matters beyond another model release. Salesforce owns the application layer, customer relationships, permissions, business logic, and much of the context behind each request.
A closed model provider historically supplied the difficult reasoning. That position gave frontier laboratories an important role inside high-value enterprise workflows.
Koa lets Salesforce internalize part of that role. The company controls the model weights, hosts the system, defines the serving environment, and determines how it enters Agentforce routing.
The model is based on NVIDIA Nemotron 3 Super, an open-weight model released under a permissive license. Open weights allow approved operators to deploy and adapt the model without sending every request to its original developer.
That control supports Salesforce’s data-boundary argument. The company says Koa runs entirely within its infrastructure, with customer information remaining inside Salesforce during inference.
Inference is the process through which a trained model produces an answer or action. Keeping that process within one controlled environment can simplify governance, auditing, and data-location decisions.
The claim does not mean deployment becomes risk-free. Salesforce still must protect model endpoints, connected tools, stored credentials, action permissions, and the data supplied during each session.
However, ownership changes who manages those risks. Customers can evaluate Koa as part of Salesforce’s existing trust boundary instead of adding another external model path.
Koa also gives Salesforce more influence over inference economics. Closed providers usually charge according to token use, reserved capacity, or negotiated enterprise agreements.
A specialized model can reduce unnecessary reasoning if it reaches the correct CRM action with fewer generated tokens. Salesforce and NVIDIA say that efficiency is one of Koa’s design goals.
Nemotron 3 Super uses a mixture-of-experts architecture, which activates only selected parts of the model for each token. NVIDIA says the model has 120 billion parameters but activates 12 billion during inference.
NVIDIA also gives the model a one-million-token context window. A context window is the amount of information a model can consider within one interaction.
These specifications are relevant because multi-agent workflows generate long histories, tool outputs, and intermediate decisions. NVIDIA says such workloads can use substantially more context than ordinary chat.
The company claims Nemotron 3 Super delivers up to five times the throughput of the previous Nemotron Super model. It also claims up to twice the accuracy, depending on the test.
Those are NVIDIA measurements, not independent findings about Koa. Salesforce’s post-training, serving system, connected tools, and governance layer will also influence real-world performance.
Still, the architecture gives Salesforce a plausible route toward lower-cost specialization. It can tune the model for common actions instead of paying for the broadest available intelligence on every request.
This places pressure on OpenAI, Anthropic, and other closed providers in a specific area. They must show that superior general reasoning justifies their cost and reduced operator control.
The pressure is not an immediate displacement threat. Salesforce continues to support outside models and announced a deeper Anthropic partnership during the same Dreamforce cycle.
Koa instead changes Salesforce’s negotiating position. The company can route routine reasoning internally while reserving frontier models for tasks where their broader capabilities provide measurable value.
That model portfolio resembles cloud computing more than a winner-takes-all contest. Enterprises combine systems based on workload, security requirements, latency, and cost.
For knowledge workers, this also shifts attention from model brands toward the quality of business context. A general model cannot assess a deal accurately when it lacks pipeline history and operating definitions.
Structured organizational context therefore becomes part of the AI stack. Teams already building a searchable knowledge base will recognize the same principle: useful answers depend on governed access to relevant information.
The Mechanism Is Specialization, Not a Bigger General Model
Koa’s central bet is that task-specific training and controlled actions can matter more than broad benchmark leadership.
Frontier model competition often focuses on mathematics, coding, science, and generalized reasoning tests. Those measurements help compare core capabilities, but CRM agents face a different set of failures.
An enterprise agent must select the correct record, follow permissions, call an approved tool, and preserve context across several steps. A fluent answer is inadequate if it updates the wrong account.
Salesforce says Koa’s synthetic training scenarios target these operational demands. They simulate professionals, customer states, policies, tools, and expected action sequences across the customer lifecycle.
One scenario might require an agent to identify an at-risk opportunity. It would need to review communications, examine stage changes, apply company rules, and recommend an authorized next action.
Another might involve an angry service customer. The agent would need to retrieve account context, understand policy limits, call the correct functions, and escalate when required.
The model must therefore do more than generate text. It must translate language into controlled work across a system of record.
Salesforce evaluates this behavior with CRM Bench, its benchmark for tasks such as updating opportunities, routing cases, and scheduling follow-ups. The company says Koa matches or exceeds leading model performance on CRM actions.
Salesforce reports that Koa makes three times fewer errors on the benchmark. It also claims 11% greater precision when selecting an action.
The company says Koa recalls customer context with 2.1 times greater reliability. It further reports a 15% improvement in remembering context during extended conversations.
These figures appear on Salesforce’s official Koa model page. They remain company-reported results until outside evaluators reproduce them under comparable conditions.
The benchmark design deserves close attention. A CRM-specific evaluation can measure operational competence more directly than a general academic test.
However, a vendor can also shape its benchmark around familiar workflows and internal tools. Buyers need details about baselines, task distributions, scoring, and failure severity.
A minor formatting error should not carry the same weight as an unauthorized record change. Aggregate accuracy can obscure the failures that matter most in production.
The second part of the mechanism is deterministic control around a probabilistic model. Salesforce says it hosts Koa at temperature zero to make responses more consistent.
Temperature controls variation in model output. A lower setting generally makes responses more repeatable, though it does not guarantee correctness.
Salesforce also uses a serving harness around Koa. The company describes this as an additional layer of safety and operational controls beyond the model itself.
That layered design is important because a language model cannot enforce enterprise governance alone. Permissions, tool schemas, validation, monitoring, and approval steps must constrain what an agent can do.
Salesforce says Koa remains subject to the context and guardrails configured by each customer. It only sees information supplied through those governed paths.
The third part is model routing. Koa becomes another provider option rather than the sole intelligence source for Agentforce.
An agent router can direct a CRM task to Koa while sending a different problem to Claude or another model. This separates the user experience from the underlying provider.
Such routing can turn models into interchangeable infrastructure, but only when performance is genuinely comparable. Weak evaluation would route work toward lower costs while hiding lower-quality outcomes.
The strategy also protects Salesforce from relying completely on one laboratory. It can adjust model selection as costs, capabilities, regulatory requirements, or customer preferences change.
For NVIDIA, Koa demonstrates how Nemotron can become the foundation for specialized enterprise models. NVIDIA benefits even when the finished product carries Salesforce branding.
Its role extends beyond supplying accelerators. The company provides open weights, training methods, deployment software, and technical collaboration for partners building their own intelligence layers.
The Nemotron architecture was designed for long-running agent workflows and efficient inference. Koa gives that design a prominent test inside a major software platform.
This strategy differs from asking every enterprise to build a base model. Companies can start with a capable open model and apply their own domain knowledge through post-training.
The difficult part becomes converting institutional experience into safe training data and useful evaluations. Salesforce describes its synthetic corpus as reflecting 27 years of CRM intelligence.
That description should not be interpreted as 27 years of customer records entering the model. Salesforce says the corpus uses synthetic scenarios modeled on accumulated workflow knowledge.
The distinction supports privacy, but it creates another question. Synthetic examples can reproduce expected patterns without capturing the full disorder of production environments.
Real accounts contain incomplete fields, conflicting records, unusual permissions, stale data, and organization-specific exceptions. Those conditions often determine whether an agent succeeds.
Koa’s mechanism is therefore credible but unfinished. Specialization can improve relevance and efficiency, yet production value depends on data quality, tool reliability, and operational testing.
Internal Benchmarks Cannot Settle the Enterprise AI Debate
Koa’s strongest claims still come from the companies selling and operating the system.
Salesforce has named several organizations participating in pilots, including 1-800Accountant, Baxter Credit Union, Engine, Formula 1, UChicago Medicine, and Xero.
It also says Koa is already used internally through employee, help, event, and web agents. These deployments provide useful environments for testing before broad availability.
Pilot participation, however, is not evidence of sustained business impact. The published materials do not disclose completion rates, intervention rates, latency, or total operating costs.
They also do not show how Koa performs when company data is inconsistent. That is a routine condition in large CRM deployments, not an exceptional edge case.
The customer statements emphasize potential use cases. UChicago Medicine highlighted multi-step coordination, while Engine focused on precise reasoning across complicated service situations.
These examples show where specialized reasoning might help. They do not yet establish that the model can work autonomously in high-stakes environments.
Enterprise buyers should ask how often humans must correct or approve Koa’s actions. They should also distinguish answer quality from successful workflow completion.
An agent can generate an accurate summary and still fail to update the right system. It can select the correct tool but supply an invalid parameter.
The most meaningful evaluations should follow entire tasks from request to verified outcome. They should also measure whether the system stayed within its assigned permissions.
Salesforce’s assertion that customer data remains inside its trust boundary addresses one important concern. It does not eliminate prompt injection, excessive permissions, compromised connectors, or flawed business rules.
Prompt injection occurs when untrusted content manipulates an agent’s instructions. An email or document could attempt to redirect a system toward an unsafe action.
A specialized model may resist some attacks through training, but surrounding controls remain necessary. Sensitive actions should require validation, narrow permissions, logging, and appropriate human approval.
Buyers also need clarity about Koa’s limitations. Temperature zero can increase consistency without making an incorrect conclusion correct.
Likewise, synthetic training can avoid direct customer-data exposure while embedding assumptions from the scenario-generation process. Those assumptions require ongoing review as workflows change.
The safety discussion onstage added another layer of tension. Huang called safety “job one,” but he rejected the idea that safety and development speed require a simple tradeoff.
“Run as fast as you can,” he said, while adding that companies should pause when a product appears unsafe or out of control.
That position entered an active dispute among AI leaders about the pace of frontier development. An AI slowdown debate had already become part of Dreamforce’s broader context.
Koa should not become a proxy for every argument about frontier AI risk. It is a specialized enterprise system operating within an application platform.
Yet the same principle still applies at a smaller scale. Moving quickly is responsible only when a company can detect unsafe behavior before customers absorb the consequences.
Salesforce’s controlled infrastructure may make Koa easier to monitor than a collection of disconnected model services. Centralization can support consistent policies, logs, and rollback processes.
It also concentrates responsibility. If the model, router, or serving harness fails, many Agentforce workflows could inherit the same problem.
Closed frontier models retain advantages in this debate. Their developers spend heavily on general capability, safety research, evaluation, and large-scale serving infrastructure.
They can also improve rapidly without each enterprise repeating base-model development. For uncommon or intellectually demanding tasks, general models may remain the better choice.
Salesforce itself acknowledges that reality through continued model choice. It is not removing Claude, ChatGPT, or other providers from Agentforce.
Its Claudeforce partnership gives sellers access to Salesforce data and actions through Claude. The initial product includes 37 sales skills governed by Salesforce permissions.
That simultaneous partnership is not a contradiction. It shows that Salesforce expects model competition to happen inside its platform, with different systems handling different workloads.
Koa’s real opponent is therefore mandatory dependence on closed frontier reasoning. It does not need to replace every outside model to weaken that dependence.
The decisive evidence will come from routing decisions and production outcomes. If customers consistently select Koa for valuable tasks, Salesforce will have validated specialized open-model adaptation.
If they continue routing most difficult work to closed providers, Koa will remain a strategic hedge rather than a new center of enterprise intelligence.
Three Signals Will Show Whether Salesforce Koa Matters
Koa’s importance will be measured by deployment, verified outcomes, and routing behavior rather than keynote language.
The first signal is the transition from pilots to broad availability. Salesforce says Koa is available to selected pilot customers now.
The company expects general availability in United States regions during winter 2026. It also plans an open beta shortly after the initial pilot phase.
That schedule creates an immediate test. A timely release with clear documentation would strengthen Salesforce’s claim that Koa is a production model, not a conference demonstration.
A delay would not prove technical failure. It would still suggest that evaluation, infrastructure, governance, or customer readiness required more work than expected.
The second signal is evidence from the named pilot organizations. Buyers should look for measured workflow results rather than endorsements or descriptions of intended use.
Useful indicators include task-completion rates, correction rates, latency, escalations, and the percentage of actions requiring human approval. Security incidents and rollback frequency also matter.
Cost claims need the same discipline. Fewer tokens can reduce one component of inference expense, but total costs include hosting, retrieval, integration, monitoring, and human review.
Independent comparisons would be especially valuable. Evaluators should test Koa and frontier alternatives against the same records, permissions, tools, and completion criteria.
Such comparisons could validate Salesforce’s CRM Bench results or expose where those results fail to transfer. Either outcome would help customers make better routing decisions.
The third signal is how Salesforce balances Koa with outside models. Model selection inside Agentforce will reveal more than public positioning.
If Salesforce makes Koa the default for common CRM reasoning, it would show confidence in the model’s reliability and economics. Customer opt-in rates would provide another adoption measure.
If Claude or other frontier systems continue handling most multi-step work, specialization may have a narrower role. Koa could still serve privacy-sensitive or highly repetitive workflows.
Watch how Salesforce describes the router as well. Transparent selection rules would help administrators understand why a task reached a particular provider.
Organizations should be able to connect model choice to measurable requirements. These include data location, latency, action risk, task complexity, and acceptable cost.
The broader industry will also watch NVIDIA’s partner pipeline. More domain-specific systems built on Nemotron would suggest Koa represents a repeatable model-development pattern.
A limited number of showcase deployments would imply that deep specialization remains expensive and difficult. Large platforms may manage it, while smaller software providers continue using closed APIs.
For enterprise teams, the immediate lesson is practical. Model capability alone does not create an effective agent.
Reliable agents require clean context, narrow permissions, tested tools, and evaluations based on completed work. They also need accessible institutional knowledge, not scattered documents and undocumented decisions.
Teams preparing for this model-rich environment can strengthen their knowledge workflows before selecting a provider. Better context improves every model route, including specialized and general systems.
Huang’s Dreamforce statement captured the ambition surrounding AI infrastructure. Koa reveals the harder question hidden beneath it: who controls the reasoning layer attached to enterprise data?
Salesforce now has an answer it can test inside Agentforce. It controls Koa while preserving access to competing models.
The next few months should show whether that control produces better outcomes or only more optionality. Buyers should demand production measurements, independent comparisons, and transparent routing policies.
That evidence will determine whether Jensen Huang at Dreamforce marked a durable shift toward specialized enterprise AI. Until then, Koa is a credible challenge to closed-model dependence, not proof of victory.



