Skip to content
DMarketer Tayeeb – Digital Marketing Expert in Bangalore | SEO, SEM & SMM Expert
Contact

Salesforce Koa: CRM Reasoning Model, Access, Deployment and Benchmark Boundaries

Short answer: Salesforce Koa is a CRM-focused reasoning model for Agentforce, built by post-training NVIDIA Nemotron 3 Super rather than by releasing a new general-purpose base model. Salesforce says Koa uses simulated CRM work, tool-use training and its trust boundary; a Salesforce-authored technical preprint now gives more detail on the method and benchmark results. The evidence is useful, but it does not establish that Koa is universally available, independently evaluated or better than every frontier model.

There is also an important access qualification. The current Koa product page contains language that sounds like a managed model in the catalogue and “available to all customers out of the box.” The same page’s FAQ, the launch release and the announcement’s rollout section describe select pilots now, with general availability expected in U.S. regions in winter 2026. Until your Salesforce org or written account confirmation resolves that contradiction, treat Koa as an entitlement- and pilot-dependent capability, not as a generally downloadable model.

This guide keeps the practical question in view: what can an enterprise buyer verify about Koa’s lineage, training, controls, evidence and access before allowing it to reason over CRM workflows?

What the new evidence adds

Salesforce announced Koa on September 15, 2026, then published “Why We Post-Trained Our Own Reasoning Model” on September 16. A related arXiv record for Salesforce Koa: An Enterprise Language Model for Agentic Tool Use lists version 1 as submitted on September 14; the PDF title page is dated September 15. These are related first-party accounts, not independent audits.

Buyer questionWhat the current evidence supportsWhat still needs verification
What is Koa?A Salesforce CRM reasoning model for Agentforce, post-trained from NVIDIA Nemotron 3 Super.The exact serving checkpoint, API/model identifier, context and quota for your org.
How was it trained?Salesforce describes supervised fine-tuning and GRPO over public and synthetic, CRM-shaped simulations with tool-use rewards.Whether your use case matches the simulated task coverage and which policies/tools are enabled.
How good is it?The Salesforce-authored paper reports results on Tau2Bench, BFCL and CRM Bench; results improve on the base model but are mixed against frontier baselines.An independent Koa evaluation on your records, permissions, languages, latency and escalation rules.
Can I use it now?Salesforce publicly describes select pilots now and a winter 2026 U.S.-region GA target, while another product-page section uses broader catalogue language.Tenant entitlement, region, pilot terms, model visibility and current release status.
What does it cost?No Koa-specific public SKU was found. Agentforce pricing is adjacent context only.A current quote, credits/conversation terms, edition requirements and contract treatment.

What Salesforce says Koa is

Koa is positioned as the reasoning layer for CRM agents: it should decide which tool or workflow step is needed, keep track of multi-turn context and stop or ask for help when the required tool, parameter or permission is missing. Salesforce’s September 16 account describes examples such as lead qualification rules and a missing-tool situation in which the model should say what it cannot do, ask what it needs, and then resume or hand off.

That is a design objective, not proof that every Koa deployment will behave that way. The result depends on the Agentforce agent definition, tools, records, permissions, guardrails, human approval steps and model configuration. A model’s reasoning capability cannot grant access that the surrounding system does not grant.

Post-training: the useful technical distinction

NVIDIA Nemotron 3 Super is the open-weight starting point. NVIDIA describes that base as a 120-billion-parameter model with approximately 12 billion active parameters in its hybrid mixture-of-experts design. Its model card and technical report cover the base checkpoint, licence, general-purpose evaluations and deployment requirements. They do not establish that Koa itself has public weights, the same licence or the same serving interface.

Salesforce says it controls Koa’s post-trained weights and runs training and inference in its own infrastructure. The Koa paper describes supervised fine-tuning followed by group relative policy optimisation (GRPO). Its simulation-to-reward pipeline creates persona-conditioned, multi-turn CRM tasks, executes tool calls in a deterministic sandbox and rewards grounded task resolution. A helper model simulates a customer, emulates tools and judges coverage; this keeps the training loop repeatable without using customer records.

The paper also describes constrained decision turns. The model can be rewarded for deferring, asking for a missing parameter, requesting elevated permission or handing off instead of taking a wrong action. This is the most useful new reader-facing detail: Koa is being trained for controlled tool use and enterprise task boundaries, not merely for a larger chat answer.

Koa, Nemotron and the surrounding Salesforce layers

LayerRoleDo not assume
KoaSalesforce’s CRM-focused reasoning model for Agentforce.Public weights, a standalone download, or a generally available API.
NVIDIA Nemotron 3 SuperOpen-weight base model used as the starting point, according to Salesforce.Base-model licence, benchmarks or hardware results transfer directly to Koa.
AgentforceSalesforce’s agent platform, tools, permissions and orchestration environment.Choosing Koa changes the agent’s permissions or approval policy.
AIforceA Salesforce interface for working with multiple AI systems and enterprise context.AIforce is the Koa model itself. See the AIforce access and controls guide.
Trusted infrastructureSalesforce’s boundary, controls and data-handling commitments around the service.A trust-boundary statement answers every residency, retention or audit question. The enterprise AI controls guide gives the wider control context.

What the Salesforce-authored paper actually measured

The paper compares Koa with the Nemotron base model and several external model baselines. It reports rollouts and evaluations from a Salesforce research setup, including Tau2Bench, BFCL and a CRM Bench. That makes the numbers more useful than the launch slogan, but it remains a vendor-authored preprint and uses simulated environments. It is not an independent production evaluation.

EvaluationKoaNemotron baseGPT-4.1GPT-5.5Claude Opus 4.8
Tau2Bench weighted average69.4168.6454.4883.9974.00
BFCL66.63%64.73%53.96%67.63%78.18%
CRM Bench weighted average0.860.840.810.900.87
CRM function-call accuracy0.770.710.850.820.83

The paper’s own interpretation is bounded: post-training improves on the open-weight base, Koa can surpass some strong baselines on some measures, and it remains below the strongest frontier results in the reported comparisons. The table does not prove general superiority, lower total cost, higher revenue or safer behaviour in a particular org. Reproduce the task definition, tool schema, permissions, language and success criteria before using a benchmark to choose a model.

How to read the “3x fewer errors” and product-page metrics

The September 15 release says Koa matched or exceeded leading performance on CRM benchmark tasks with three times fewer errors. The product page additionally advertises 11% higher precision for right-action calling, 2.1x reliability for recalling context and 15% better context handling in long conversations. These are Salesforce claims. The pages reviewed do not expose enough denominator, baseline, task split, error definition, confidence interval or independent evaluator detail to treat them as universal performance guarantees, and the paper’s published tables do not provide a denominator that maps cleanly to the “3x” wording.

Use the claims as hypotheses for a controlled evaluation. A fair comparison should run the same anonymised or synthetic cases, tool definitions, access policy, temperature, model version and escalation rules across candidates. Record both a correct action and a correct refusal; a confident action on a forbidden or under-specified record is not a success.

Availability and rollout: resolve the public-page conflict

The public surfaces do not presently form one clean availability statement:

SurfaceWording or evidenceSafe interpretation
Koa product page, main availability sectionDescribes Koa as a managed LLM in the generative AI models catalogue, selectable org-wide, and uses “available to all customers out of the box” language.A broad product-surface claim that must be checked against the actual tenant.
Koa product page, FAQ and rollout copySays select pilot customers can use Koa now, with GA expected in U.S. regions in winter 2026 and an open beta shortly after; another section says Salesforce is moving into pilots.Operationally conservative pilot/entitlement treatment.
September 15 press releaseSays select pilot customers are using Koa in Agentforce now and general availability is expected in U.S. regions in winter 2026.Current launch boundary until superseded by an authenticated release notice.
Supported Models developer documentationLists beta NVIDIA Nemotron models with Amazon Bedrock API names, but no Koa row or Koa API name was found.Do not infer a public API or BYOLLM identifier from the base model listing.

Before planning a rollout, ask the account team or administrator to show the current model entry in the intended org, the applicable region, pilot or beta terms, API/model identifier, quota, logging behaviour and fallback path. A product-page phrase alone is not enough to establish entitlement.

Training data, trust boundary and permissions

Salesforce says Koa’s simulated training work used no customer data and that Koa training and inference stay within Salesforce’s trust boundary. The paper likewise says its public and synthetic training data did not use customer data. Those are important vendor commitments, but they are not an independent audit. They also answer a narrower question than runtime retention: Salesforce’s Trust and architecture material explains that prompts, responses and trust signals can be logged or stored in Data 360, while zero-retention terms and provider handling can vary by service and contract.

Ask separately about: what leaves the org, what is stored and for how long, whether session traces are enabled, who can inspect them, how deletion and export work, which subprocessors or regions are involved, and whether the same policy applies to fallbacks. Keep least-privilege permissions, approval gates and audit review in the Agentforce design. Koa cannot make an unauthorised record safe to expose.

Pricing: no Koa-specific public SKU

The current Agentforce pricing page shows adjacent commercial units such as Flex Credits at $500 per 100,000 credits, Conversations at $2 per conversation, a $5 User License that requires Flex Credits, and add-ons listed at $125 or $150 per user per month. It also lists Agentforce1 editions from $550 per user per month. The page says examples are illustrative and information is subject to change.

None of those figures is a Koa price. They cannot be multiplied into a Koa forecast without the org’s edition, actions, conversation volume, model entitlement, region, contract and overage rules. For broader Salesforce edition context, see the Salesforce edition and pricing guide.

A bounded Koa evaluation before adoption

Use a small, reversible evaluation rather than a production switch:

  • Obtain the exact Koa model identifier, version, region and pilot or GA entitlement in writing.
  • Define a fixed case set covering sales, service and commerce workflows, including incomplete context, conflicting instructions, missing tools and required approvals.
  • Use synthetic or redacted records and document every tool, field, permission and fallback available to the agent.
  • Score correct action, correct refusal, correct clarification, correct handoff, record mutation accuracy and policy violations separately.
  • Compare Koa with the incumbent and at least one alternative under identical prompts, temperature, tools, timeouts and review rules.
  • Capture latency, retries, token or credit consumption, trace access and failure reasons; do not substitute the product-page headline metrics.
  • Have business owners review false positives and false negatives, especially actions that change customer or financial records.
  • Set rollback, approval and audit thresholds before expanding the pilot.

Verified facts, vendor claims and unknowns

Evidence classCurrent conclusion
Verified from primary materialKoa is Salesforce’s CRM reasoning model; Salesforce says it post-trained Nemotron 3 Super; the paper describes SFT, GRPO and simulated tool-use tasks; the paper reports the benchmark table above; no Koa API/model row appeared in the Supported Models page reviewed.
Vendor-described claimNo customer data in simulated training, Salesforce trust-boundary operation, three times fewer errors, and product-page precision/reliability/context improvements.
Independent or practitioner signalCoverage from TechCrunch, Constellation, Mint and State of AI Marketing mostly repeats or contextualises the launch. Commentary recommends pilot verification; no independent Koa production benchmark was found.
Still unknownPublic Koa weights or licence, stable API name, all-region matrix, quota, retention configuration, independent error rate, production ROI, and whether the broad catalogue language has reached every tenant.

FAQ

Is Koa an open model like Nemotron?

No public Koa checkpoint or licence was found in the sources reviewed. Nemotron 3 Super is the open-weight base described by NVIDIA; that does not make the Salesforce post-trained Koa checkpoint open.

Can every Salesforce customer use Koa now?

Do not assume that. One section of the product page uses broad catalogue language, but its FAQ and rollout language, plus the launch release, describe select pilots now and a winter 2026 U.S.-region GA target. Confirm your tenant’s entitlement and region.

Does the paper prove Koa is the best reasoning model?

No. It is a Salesforce-authored preprint using simulated environments. Koa improves on the Nemotron base and is competitive on selected tasks, but the reported table also shows stronger results from frontier baselines on some measures. It is evidence for a CRM-specific training approach, not a universal ranking.

Does Koa replace AIforce or Agentforce?

No. Koa is a model; Agentforce supplies the agent runtime, tools and controls; AIforce is an interface layer. They can be connected in a Salesforce workflow but answer different product questions.

What should a buyer ask for next?

Request the tenant-visible model identifier, region and pilot terms, current data-retention and trace configuration, permissions/fallback design, a reproducible evaluation protocol and a commercial quote that names the Koa entitlement. Without those, the launch announcement is useful context but not a deployment decision.

Bottom line

Koa is a meaningful Salesforce engineering direction: post-training an open-weight base for CRM tool use, simulated enterprise workflows and bounded decisions can be more relevant than a generic chat benchmark. The new paper makes that approach inspectable and reports concrete comparisons. It still does not turn vendor evidence into independent proof.

For now, the defensible decision is to preserve the existing Koa guide, add the paper’s method and benchmark boundaries, and treat access as select-pilot or tenant-specific until the conflicting public wording is resolved. Evaluate the exact org, tools, permissions, data path and rollback plan before allowing Koa to change real records.

Sources checked

Share this article

Written by

Tayeeb Khan

Tayeeb Khan is a digital marketing strategist, SEO specialist, and the founder of Digital Marketer Tayeeb (DMT). Backed by an engineering degree, certifications in Google and Meta advertising, and over a decade of hands-on experience growing startups, Tayeeb bridges the gap between technical infrastructure and marketing execution. His insights on SEO and AI-driven marketing are strictly practitioner-first—built on real tests, real campaigns, and real results. Connect on LinkedIn or via Email.

Leave a Comment

Your email address will not be published. Required fields are marked *

Stay ahead of the curve

Get actionable digital marketing, SEO, and AI insights delivered to your inbox. No fluff, just value.

No spam. Unsubscribe anytime.