Salesforce says its new Koa reasoning model produces three times fewer errors than leading models on CRM actions like updating opportunities, routing cases and scheduling follow-ups. I believe that number, and I also do not think it should move a single buying decision yet, because Salesforce is the only party who has seen the test.
That is not a knock on Koa specifically. It is a rule I would apply to any vendor claim shaped this way, and enterprise AI is now full of them. Salesforce built the CRM benchmark that Koa was scored against, ran the comparison, and published the result as a press release quote rather than a methodology paper. No third party has replicated the test, seen the error taxonomy behind “three times fewer errors,” or checked whether the comparison model was tuned as carefully as Koa was for the same task. A vendor grading its own AI’s homework is not proof. It is marketing with a number attached.
The counter-argument, stated as strongly as I can make it
The obvious response is that Salesforce is not a random vendor making up a number. Koa was post-trained on synthetic data modeled on twenty-seven years of actual CRM deployment history across fourteen industries, using NVIDIA’s Nemotron 3 Super as the base and NVIDIA’s own NeMo RL and NeMo Gym tooling for the reinforcement learning. Nobody else has that deployment history to build a comparable benchmark from, so an outside party could not replicate Salesforce’s exact test even if it wanted to. Early pilot customers including Baxter Credit Union, 1-800Accountant and UChicago Medicine are already using it on live workflows, which is a form of real-world validation that a synthetic leaderboard cannot offer. Jensen Huang’s framing, that “every company needs useful AI, tailored to its knowledge, expertise, and work,” is a real argument for why a domain-specific model built on proprietary data should outperform a generic one on domain-specific tasks. That much I accept.
Why it still is not enough
None of that answers the actual question a RevOps leader needs answered before trusting Koa with opportunity updates and case routing: how does it fail, and how often, in situations that were not represented in Salesforce’s own test set? A vendor-run benchmark tends to measure the tasks the vendor already anticipated, using data the vendor already curated. It rarely surfaces the rare, high-cost failure, the edge case where an agent updates the wrong field on a large account or misroutes an escalation that should have gone to a human. Those are exactly the failures that matter most in a CRM, because a single wrong update on a six-figure opportunity costs more than a dozen correct ones save.
Pilot usage is genuine evidence, but it is not independent evidence either, and it is early. Six named pilot customers after weeks of use is a promising start, not a validated error rate. Formula 1 running Koa on some workflows tells me Salesforce can land marquee logos. It does not tell me what Koa does the first time it encounters a deal structure, a discount approval chain, or a multi-entity account hierarchy that looks nothing like the synthetic training scenarios built to represent fourteen industries in the abstract.
What I would actually do about it
Buyers evaluating Koa, or any vendor’s self-graded reasoning model this year, should ask for three things before the benchmark number gets repeated in an internal deck: the error taxonomy behind the headline figure, meaning what counts as an error and what does not; access to run the vendor’s own eval set against a held-out sample of the buyer’s real, anonymized data rather than the vendor’s synthetic scenarios; and a defined human-review threshold for high-value actions, so that “three times fewer errors” does not quietly become zero human oversight on the transactions that can least afford to be wrong. None of that requires distrust of Salesforce specifically. It requires treating every vendor’s internal AI benchmark the way a RevOps team already treats a vendor’s own ROI calculator: a useful starting number, never the final word.
The race among CRM vendors is not about how many agents they ship, and it is also not about whose internal benchmark reads best in a press release. It is about which reasoning model actually holds up on the transaction nobody tested for. Until a buyer can check that independently, the right response to any vendor’s own performance claim, Salesforce’s included, is to treat it as a starting hypothesis, not a settled fact, the same standard this publication has already applied to Salesforce’s own release governance claims.
Source: Salesforce
