A healthy org is the start, not the finish
Most Agentforce conversations begin with licences. This one began correctly: the org was already in good enough shape to consider agents, which puts it ahead of most. The question was what agents could safely be allowed to do.
An agent that qualifies a lead wrongly costs a lead. An agent that updates a record wrongly costs the trust that every downstream report depends on. Those are different risks and they need different controls, which is why "can we deploy agents" is the wrong question and "which actions may an agent take unsupervised" is the right one.
Review logic before autonomy
Every consequential action an agent could take was classified before deployment: fully autonomous, autonomous but logged for review, or requiring human confirmation. That classification is unglamorous and it is the entire safety model.
The alternative — deploy and watch — fails in a specific way. The failures are individually small and collectively invisible, so nobody notices until a quarterly number is wrong and there is no audit trail explaining why.
Baselines, taken before anything was switched on
Response times, qualification rates and hygiene measures were captured before the first agent ran. This sounds obvious and is skipped almost universally, because it delays the interesting part by a week.
Without it there is no honest way to answer whether the agents helped. Every measurement afterwards is compared against a remembered version of the past, which is invariably flattering to whatever was just deployed.
The 30- and 90-day comparisons were possible only because that week was spent.
What transfers to other Salesforce teams
- Classify every agent action by consequence before deployment, not after an incident.
- Capture baselines before the first agent runs — a week that is impossible to recover later.
- An agent fleet should expand on measured evidence, not on the enthusiasm of the last demo.
- Someone must own the fleet. Agents drift as the data and processes around them change.