{"id":3072,"date":"2026-09-16T07:55:50","date_gmt":"2026-09-16T07:55:50","guid":{"rendered":"https:\/\/dmarketertayeeb.com\/blog\/salesforce-koa-crm-reasoning-model\/"},"modified":"2026-09-17T06:59:09","modified_gmt":"2026-09-17T06:59:09","slug":"salesforce-koa-crm-reasoning-model","status":"publish","type":"post","link":"https:\/\/dmarketertayeeb.com\/blog\/salesforce-koa-crm-reasoning-model\/","title":{"rendered":"Salesforce Koa: CRM Reasoning Model, Access, Deployment and Benchmark Boundaries"},"content":{"rendered":"\n<p><strong>Short answer:<\/strong> Salesforce Koa is a CRM-focused reasoning model for Agentforce, built by post-training NVIDIA Nemotron 3 Super rather than by releasing a new general-purpose base model. Salesforce says Koa uses simulated CRM work, tool-use training and its trust boundary; a Salesforce-authored technical preprint now gives more detail on the method and benchmark results. The evidence is useful, but it does not establish that Koa is universally available, independently evaluated or better than every frontier model.<\/p>\n\n\n\n<p>There is also an important access qualification. The current Koa product page contains language that sounds like a managed model in the catalogue and \u201cavailable to all customers out of the box.\u201d The same page\u2019s FAQ, the launch release and the announcement\u2019s rollout section describe select pilots now, with general availability expected in U.S. regions in winter 2026. Until your Salesforce org or written account confirmation resolves that contradiction, treat Koa as an entitlement- and pilot-dependent capability, not as a generally downloadable model.<\/p>\n\n\n\n<p>This guide keeps the practical question in view: what can an enterprise buyer verify about Koa\u2019s lineage, training, controls, evidence and access before allowing it to reason over CRM workflows?<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">What the new evidence adds<\/h2>\n\n\n\n<p>Salesforce announced Koa on September 15, 2026, then published \u201cWhy We Post-Trained Our Own Reasoning Model\u201d on September 16. A related arXiv record for <em>Salesforce Koa: An Enterprise Language Model for Agentic Tool Use<\/em> lists version 1 as submitted on September 14; the PDF title page is dated September 15. These are related first-party accounts, not independent audits.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Buyer question<\/th><th>What the current evidence supports<\/th><th>What still needs verification<\/th><\/tr><\/thead><tbody><tr><td>What is Koa?<\/td><td>A Salesforce CRM reasoning model for Agentforce, post-trained from NVIDIA Nemotron 3 Super.<\/td><td>The exact serving checkpoint, API\/model identifier, context and quota for your org.<\/td><\/tr><tr><td>How was it trained?<\/td><td>Salesforce describes supervised fine-tuning and GRPO over public and synthetic, CRM-shaped simulations with tool-use rewards.<\/td><td>Whether your use case matches the simulated task coverage and which policies\/tools are enabled.<\/td><\/tr><tr><td>How good is it?<\/td><td>The Salesforce-authored paper reports results on Tau2Bench, BFCL and CRM Bench; results improve on the base model but are mixed against frontier baselines.<\/td><td>An independent Koa evaluation on your records, permissions, languages, latency and escalation rules.<\/td><\/tr><tr><td>Can I use it now?<\/td><td>Salesforce publicly describes select pilots now and a winter 2026 U.S.-region GA target, while another product-page section uses broader catalogue language.<\/td><td>Tenant entitlement, region, pilot terms, model visibility and current release status.<\/td><\/tr><tr><td>What does it cost?<\/td><td>No Koa-specific public SKU was found. Agentforce pricing is adjacent context only.<\/td><td>A current quote, credits\/conversation terms, edition requirements and contract treatment.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">What Salesforce says Koa is<\/h2>\n\n\n\n<p>Koa is positioned as the reasoning layer for CRM agents: it should decide which tool or workflow step is needed, keep track of multi-turn context and stop or ask for help when the required tool, parameter or permission is missing. Salesforce\u2019s September 16 account describes examples such as lead qualification rules and a missing-tool situation in which the model should say what it cannot do, ask what it needs, and then resume or hand off.<\/p>\n\n\n\n<p>That is a design objective, not proof that every Koa deployment will behave that way. The result depends on the Agentforce agent definition, tools, records, permissions, guardrails, human approval steps and model configuration. A model\u2019s reasoning capability cannot grant access that the surrounding system does not grant.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Post-training: the useful technical distinction<\/h2>\n\n\n\n<p>NVIDIA Nemotron 3 Super is the open-weight starting point. NVIDIA describes that base as a 120-billion-parameter model with approximately 12 billion active parameters in its hybrid mixture-of-experts design. Its model card and technical report cover the base checkpoint, licence, general-purpose evaluations and deployment requirements. They do not establish that Koa itself has public weights, the same licence or the same serving interface.<\/p>\n\n\n\n<p>Salesforce says it controls Koa\u2019s post-trained weights and runs training and inference in its own infrastructure. The Koa paper describes supervised fine-tuning followed by group relative policy optimisation (GRPO). Its simulation-to-reward pipeline creates persona-conditioned, multi-turn CRM tasks, executes tool calls in a deterministic sandbox and rewards grounded task resolution. A helper model simulates a customer, emulates tools and judges coverage; this keeps the training loop repeatable without using customer records.<\/p>\n\n\n\n<p>The paper also describes constrained decision turns. The model can be rewarded for deferring, asking for a missing parameter, requesting elevated permission or handing off instead of taking a wrong action. This is the most useful new reader-facing detail: Koa is being trained for controlled tool use and enterprise task boundaries, not merely for a larger chat answer.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Koa, Nemotron and the surrounding Salesforce layers<\/h2>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Layer<\/th><th>Role<\/th><th>Do not assume<\/th><\/tr><\/thead><tbody><tr><td>Koa<\/td><td>Salesforce\u2019s CRM-focused reasoning model for Agentforce.<\/td><td>Public weights, a standalone download, or a generally available API.<\/td><\/tr><tr><td>NVIDIA Nemotron 3 Super<\/td><td>Open-weight base model used as the starting point, according to Salesforce.<\/td><td>Base-model licence, benchmarks or hardware results transfer directly to Koa.<\/td><\/tr><tr><td>Agentforce<\/td><td>Salesforce\u2019s agent platform, tools, permissions and orchestration environment.<\/td><td>Choosing Koa changes the agent\u2019s permissions or approval policy.<\/td><\/tr><tr><td>AIforce<\/td><td>A Salesforce interface for working with multiple AI systems and enterprise context.<\/td><td>AIforce is the Koa model itself. See the <a href=\"https:\/\/dmarketertayeeb.com\/blog\/salesforce-aiforce-access-interfaces-controls\/\">AIforce access and controls guide<\/a>.<\/td><\/tr><tr><td>Trusted infrastructure<\/td><td>Salesforce\u2019s boundary, controls and data-handling commitments around the service.<\/td><td>A trust-boundary statement answers every residency, retention or audit question. The <a href=\"https:\/\/dmarketertayeeb.com\/blog\/salesforce-trusted-enterprise-ai-harness\/\">enterprise AI controls guide<\/a> gives the wider control context.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">What the Salesforce-authored paper actually measured<\/h2>\n\n\n\n<p>The paper compares Koa with the Nemotron base model and several external model baselines. It reports rollouts and evaluations from a Salesforce research setup, including Tau2Bench, BFCL and a CRM Bench. That makes the numbers more useful than the launch slogan, but it remains a vendor-authored preprint and uses simulated environments. It is not an independent production evaluation.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Evaluation<\/th><th>Koa<\/th><th>Nemotron base<\/th><th>GPT-4.1<\/th><th>GPT-5.5<\/th><th>Claude Opus 4.8<\/th><\/tr><\/thead><tbody><tr><td>Tau2Bench weighted average<\/td><td>69.41<\/td><td>68.64<\/td><td>54.48<\/td><td>83.99<\/td><td>74.00<\/td><\/tr><tr><td>BFCL<\/td><td>66.63%<\/td><td>64.73%<\/td><td>53.96%<\/td><td>67.63%<\/td><td>78.18%<\/td><\/tr><tr><td>CRM Bench weighted average<\/td><td>0.86<\/td><td>0.84<\/td><td>0.81<\/td><td>0.90<\/td><td>0.87<\/td><\/tr><tr><td>CRM function-call accuracy<\/td><td>0.77<\/td><td>0.71<\/td><td>0.85<\/td><td>0.82<\/td><td>0.83<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p>The paper\u2019s own interpretation is bounded: post-training improves on the open-weight base, Koa can surpass some strong baselines on some measures, and it remains below the strongest frontier results in the reported comparisons. The table does not prove general superiority, lower total cost, higher revenue or safer behaviour in a particular org. Reproduce the task definition, tool schema, permissions, language and success criteria before using a benchmark to choose a model.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">How to read the \u201c3x fewer errors\u201d and product-page metrics<\/h2>\n\n\n\n<p>The September 15 release says Koa matched or exceeded leading performance on CRM benchmark tasks with three times fewer errors. The product page additionally advertises 11% higher precision for right-action calling, 2.1x reliability for recalling context and 15% better context handling in long conversations. These are Salesforce claims. The pages reviewed do not expose enough denominator, baseline, task split, error definition, confidence interval or independent evaluator detail to treat them as universal performance guarantees, and the paper\u2019s published tables do not provide a denominator that maps cleanly to the \u201c3x\u201d wording.<\/p>\n\n\n\n<p>Use the claims as hypotheses for a controlled evaluation. A fair comparison should run the same anonymised or synthetic cases, tool definitions, access policy, temperature, model version and escalation rules across candidates. Record both a correct action and a correct refusal; a confident action on a forbidden or under-specified record is not a success.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Availability and rollout: resolve the public-page conflict<\/h2>\n\n\n\n<p>The public surfaces do not presently form one clean availability statement:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Surface<\/th><th>Wording or evidence<\/th><th>Safe interpretation<\/th><\/tr><\/thead><tbody><tr><td>Koa product page, main availability section<\/td><td>Describes Koa as a managed LLM in the generative AI models catalogue, selectable org-wide, and uses \u201cavailable to all customers out of the box\u201d language.<\/td><td>A broad product-surface claim that must be checked against the actual tenant.<\/td><\/tr><tr><td>Koa product page, FAQ and rollout copy<\/td><td>Says select pilot customers can use Koa now, with GA expected in U.S. regions in winter 2026 and an open beta shortly after; another section says Salesforce is moving into pilots.<\/td><td>Operationally conservative pilot\/entitlement treatment.<\/td><\/tr><tr><td>September 15 press release<\/td><td>Says select pilot customers are using Koa in Agentforce now and general availability is expected in U.S. regions in winter 2026.<\/td><td>Current launch boundary until superseded by an authenticated release notice.<\/td><\/tr><tr><td>Supported Models developer documentation<\/td><td>Lists beta NVIDIA Nemotron models with Amazon Bedrock API names, but no Koa row or Koa API name was found.<\/td><td>Do not infer a public API or BYOLLM identifier from the base model listing.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p>Before planning a rollout, ask the account team or administrator to show the current model entry in the intended org, the applicable region, pilot or beta terms, API\/model identifier, quota, logging behaviour and fallback path. A product-page phrase alone is not enough to establish entitlement.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Training data, trust boundary and permissions<\/h2>\n\n\n\n<p>Salesforce says Koa\u2019s simulated training work used no customer data and that Koa training and inference stay within Salesforce\u2019s trust boundary. The paper likewise says its public and synthetic training data did not use customer data. Those are important vendor commitments, but they are not an independent audit. They also answer a narrower question than runtime retention: Salesforce\u2019s Trust and architecture material explains that prompts, responses and trust signals can be logged or stored in Data 360, while zero-retention terms and provider handling can vary by service and contract.<\/p>\n\n\n\n<p>Ask separately about: what leaves the org, what is stored and for how long, whether session traces are enabled, who can inspect them, how deletion and export work, which subprocessors or regions are involved, and whether the same policy applies to fallbacks. Keep least-privilege permissions, approval gates and audit review in the Agentforce design. Koa cannot make an unauthorised record safe to expose.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Pricing: no Koa-specific public SKU<\/h2>\n\n\n\n<p>The current Agentforce pricing page shows adjacent commercial units such as Flex Credits at $500 per 100,000 credits, Conversations at $2 per conversation, a $5 User License that requires Flex Credits, and add-ons listed at $125 or $150 per user per month. It also lists Agentforce1 editions from $550 per user per month. The page says examples are illustrative and information is subject to change.<\/p>\n\n\n\n<p>None of those figures is a Koa price. They cannot be multiplied into a Koa forecast without the org\u2019s edition, actions, conversation volume, model entitlement, region, contract and overage rules. For broader Salesforce edition context, see the <a href=\"https:\/\/dmarketertayeeb.com\/blog\/salesforce-core-advanced-max-pricing-2026\/\">Salesforce edition and pricing guide<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">A bounded Koa evaluation before adoption<\/h2>\n\n\n\n<p>Use a small, reversible evaluation rather than a production switch:<\/p>\n\n\n\n<ul class=\"wp-block-list\"><li>Obtain the exact Koa model identifier, version, region and pilot or GA entitlement in writing.<\/li><li>Define a fixed case set covering sales, service and commerce workflows, including incomplete context, conflicting instructions, missing tools and required approvals.<\/li><li>Use synthetic or redacted records and document every tool, field, permission and fallback available to the agent.<\/li><li>Score correct action, correct refusal, correct clarification, correct handoff, record mutation accuracy and policy violations separately.<\/li><li>Compare Koa with the incumbent and at least one alternative under identical prompts, temperature, tools, timeouts and review rules.<\/li><li>Capture latency, retries, token or credit consumption, trace access and failure reasons; do not substitute the product-page headline metrics.<\/li><li>Have business owners review false positives and false negatives, especially actions that change customer or financial records.<\/li><li>Set rollback, approval and audit thresholds before expanding the pilot.<\/li><\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">Verified facts, vendor claims and unknowns<\/h2>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Evidence class<\/th><th>Current conclusion<\/th><\/tr><\/thead><tbody><tr><td>Verified from primary material<\/td><td>Koa is Salesforce\u2019s CRM reasoning model; Salesforce says it post-trained Nemotron 3 Super; the paper describes SFT, GRPO and simulated tool-use tasks; the paper reports the benchmark table above; no Koa API\/model row appeared in the Supported Models page reviewed.<\/td><\/tr><tr><td>Vendor-described claim<\/td><td>No customer data in simulated training, Salesforce trust-boundary operation, three times fewer errors, and product-page precision\/reliability\/context improvements.<\/td><\/tr><tr><td>Independent or practitioner signal<\/td><td>Coverage from TechCrunch, Constellation, Mint and State of AI Marketing mostly repeats or contextualises the launch. Commentary recommends pilot verification; no independent Koa production benchmark was found.<\/td><\/tr><tr><td>Still unknown<\/td><td>Public Koa weights or licence, stable API name, all-region matrix, quota, retention configuration, independent error rate, production ROI, and whether the broad catalogue language has reached every tenant.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">FAQ<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Is Koa an open model like Nemotron?<\/h3>\n\n\n\n<p>No public Koa checkpoint or licence was found in the sources reviewed. Nemotron 3 Super is the open-weight base described by NVIDIA; that does not make the Salesforce post-trained Koa checkpoint open.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Can every Salesforce customer use Koa now?<\/h3>\n\n\n\n<p>Do not assume that. One section of the product page uses broad catalogue language, but its FAQ and rollout language, plus the launch release, describe select pilots now and a winter 2026 U.S.-region GA target. Confirm your tenant\u2019s entitlement and region.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Does the paper prove Koa is the best reasoning model?<\/h3>\n\n\n\n<p>No. It is a Salesforce-authored preprint using simulated environments. Koa improves on the Nemotron base and is competitive on selected tasks, but the reported table also shows stronger results from frontier baselines on some measures. It is evidence for a CRM-specific training approach, not a universal ranking.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Does Koa replace AIforce or Agentforce?<\/h3>\n\n\n\n<p>No. Koa is a model; Agentforce supplies the agent runtime, tools and controls; AIforce is an interface layer. They can be connected in a Salesforce workflow but answer different product questions.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">What should a buyer ask for next?<\/h3>\n\n\n\n<p>Request the tenant-visible model identifier, region and pilot terms, current data-retention and trace configuration, permissions\/fallback design, a reproducible evaluation protocol and a commercial quote that names the Koa entitlement. Without those, the launch announcement is useful context but not a deployment decision.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Bottom line<\/h2>\n\n\n\n<p>Koa is a meaningful Salesforce engineering direction: post-training an open-weight base for CRM tool use, simulated enterprise workflows and bounded decisions can be more relevant than a generic chat benchmark. The new paper makes that approach inspectable and reports concrete comparisons. It still does not turn vendor evidence into independent proof.<\/p>\n\n\n\n<p>For now, the defensible decision is to preserve the existing Koa guide, add the paper\u2019s method and benchmark boundaries, and treat access as select-pilot or tenant-specific until the conflicting public wording is resolved. Evaluate the exact org, tools, permissions, data path and rollback plan before allowing Koa to change real records.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Sources checked<\/h2>\n\n\n\n<ul class=\"wp-block-list\"><li><a href=\"https:\/\/www.salesforce.com\/news\/stories\/why-we-post-trained-our-own-reasoning-model\/\">Salesforce: Why We Post-Trained Our Own Reasoning Model<\/a> (September 16, 2026).<\/li><li><a href=\"https:\/\/www.salesforce.com\/news\/press-releases\/2026\/09\/15\/koa-reasoning-model\/\">Salesforce launch release for Koa<\/a> (September 15, 2026).<\/li><li><a href=\"https:\/\/www.salesforce.com\/agentforce\/koa\/\">Salesforce Koa product page<\/a>, including its FAQ and rollout wording.<\/li><li><a href=\"https:\/\/arxiv.org\/abs\/2609.15066\">Salesforce Koa technical preprint<\/a> (arXiv v1 record submitted September 14, 2026).<\/li><li><a href=\"https:\/\/developer.salesforce.com\/docs\/ai\/agentforce\/guide\/supported-models.html\">Salesforce Supported Models documentation<\/a>.<\/li><li><a href=\"https:\/\/research.nvidia.com\/labs\/nemotron\/files\/NVIDIA-Nemotron-3-Super-Technical-Report.pdf\">NVIDIA Nemotron 3 Super technical report<\/a> and the <a href=\"https:\/\/huggingface.co\/nvidia\/NVIDIA-Nemotron-3-Super-120B-A12B-FP8\">official model card<\/a>.<\/li><li><a href=\"https:\/\/www.salesforce.com\/agentforce\/pricing\/\">Salesforce Agentforce pricing<\/a>.<\/li><li><a href=\"https:\/\/help.salesforce.com\/s\/articleView?id=sf.copilot_trust.htm&amp;language=en_US&amp;type=5\">Salesforce Trust Layer documentation<\/a> and <a href=\"https:\/\/architect.salesforce.com\/docs\/architect\/well-architected\/guide\/agentic-enterprise-trust.html\">Agentic Enterprise Trust guidance<\/a>.<\/li><li><a href=\"https:\/\/www.stateofaimarketing.co\/news\/salesforce-koa-crm-reasoning-model\/\">State of AI Marketing commentary<\/a>, <a href=\"https:\/\/techcrunch.com\/2026\/09\/15\/salesforce-and-nvidias-new-reasoning-model-is-everything-the-ai-labs-should-fear\/\">TechCrunch coverage<\/a> and <a href=\"https:\/\/blog.vinkel.ai\/skarp-vinkel\/salesforce-koa-aaben-model-uden-kundekontrol\">Vinkel commentary<\/a> for independent framing; none supplied an independent Koa production test.<\/li><\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Salesforce Koa&#8217;s post-training paper adds benchmark detail, but access remains tenant-specific. Verify pilot status, trust claims and evidence before adoption.<\/p>\n","protected":false},"author":1,"featured_media":3071,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[180,177,274],"tags":[281,314,391,282,315],"class_list":["post-3072","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-news","category-digital-marketing","category-tools-reviews","tag-agentforce","tag-ai-pricing","tag-ai-safety","tag-crm-software","tag-developer-tools","has-featured-image"],"_links":{"self":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3072","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/comments?post=3072"}],"version-history":[{"count":2,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3072\/revisions"}],"predecessor-version":[{"id":3095,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3072\/revisions\/3095"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/media\/3071"}],"wp:attachment":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/media?parent=3072"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/categories?post=3072"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/tags?post=3072"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}