A sales agent that passed its pre-launch tests has passed a snapshot. The system that took the test will not be the system running next quarter, and somebody on the revenue team has to be assigned to notice.
What the vendor says about customers
Salesforce published an interview with Kathy Baxter, its Principal Architect of Ethical AI Practice, on October 6. Asked about the biggest gap in enterprise AI, she answered: “I think one of the biggest gaps is when customers conduct testing prelaunch, but then they do not continue to monitor or they don’t continue to conduct testing.”
She gave the reason in the next breath: “As we all know with probabilistic systems, the answer that you get today might not be the answer that you get tomorrow.”
I read that as a description of a staffing problem more than a technical one. Launch has a project plan, a budget and a date. Monitoring has none of those unless someone writes them. A team that tested well on day one tends to feel finished, and that feeling is the risk.
Why this matters in sales
A sales agent works in the records revenue teams use to run the business. It updates CRM fields, drafts customer messages and prepares calls. Baxter says that when an agent hallucinates, it can take incorrect actions, and those can compound over time. A wrong field value does not announce itself, and duplicate or stale records already cause trouble for AI agents without any drift at all. It shows up months later in a forecast review or a territory plan, far from the cause. By then the trail from symptom back to the agent is cold, and the team argues about whether the data was ever right.
This site made a related argument in Sales AI Isn’t as Autonomous as the Marketing Says. The autonomy that vendors sell is real enough to cause damage and rarely real enough to run unattended.
The strongest objection
Buyers will say the vendors already ship the monitoring. Salesforce describes an audit trail that records every action the agent takes and a Testing Center that lets customers evaluate the responses their agents give. Infinitus says its Lens product provides same-day compliance monitoring for its new FieldForce agents. If the instruments exist, the argument goes, the buyer can rely on them.
The instruments exist, and I am not disputing that. An audit trail is a record, and a record has no opinion about whether the agent drifted. Somebody has to open it, sample it and compare it with what the team expected. Baxter’s own statement is that customers do not keep testing after launch, and she works for a vendor that ships the tooling. If the tooling were enough, the gap she describes would not be open.
What I would put in the contract and the plan
First, name an owner before launch. In most companies that is revenue operations, because it already owns the CRM, the field definitions and the reporting. If no one is named, no one is doing it.
Second, set a review schedule and write it down. A monthly sample of agent actions from the audit trail, checked against a short list of things the agent must never do, is enough to start. The point is that it is on a calendar.
Third, ask for an accuracy measure beside the volume measures. Infinitus lists provider coverage, interactions and cost per interaction as the metrics customers can use to gauge impact. Those tell you how much the agent did and what it cost. They do not tell you whether it was right. Pair them with an error rate you define, such as the share of sampled CRM updates a rep had to correct.
Fourth, define re-test triggers. A model update, a change to the rules layer or a new data source should each send the agent back through the same tests it passed at launch.
The cost of skipping it
The pitch for sales agents is capacity. Infinitus’s release says its agents are meant to extend field coverage without a matching rise in field hires, and I understand the appeal. The monitoring owner is the one role such a pitch leaves out, even if the job is a fraction of one person’s week. Leaving that role empty is how an agent passes every test on launch day and loses accuracy afterward with no one assigned to see it.
Source: Salesforce Newsroom
