ChecklistSalesforce Agentforce
Agentforce testing before go-live: a checklist for agents customers will use
An Agentforce agent can pass a demo and still fail its first busy afternoon. Because the same question can be routed and answered differently between runs, testing has to cover many phrasings, the actions behind each answer, what the agent user can reach and what happens at handoff. This checklist sets out the test sets, assertions and checks to complete in a sandbox before launch, and what to watch once real conversations begin.
On this page
- Why an agent needs a different test plan from a Flow
- Build the utterance set from real conversations
- What every Agentforce test case should assert
- Running batch tests across iterations
- Permission and data-access tests as the agent user
- Human handoff tests through Omni-Channel
- Review trust and safety settings before launch
- Moving from sandbox to production under change control
- What to watch in the first weeks after launch
- Questions and answers
- Sources
Why an agent needs a different test plan from a Flow
A Flow given the same inputs takes the same path every time, so a handful of cases can cover it. An agent interprets language: it picks a subagent, decides whether to call an action and phrases a reply, and any of those can vary between runs. Salesforce's own training notes that agent behavior is probabilistic, so each organization has to set its own pass and fail baselines1.
Testing has side effects too. Test runs can modify CRM data and consume billable generative AI requests, which is why Salesforce tells teams to run batch tests in a sandbox and never in production1.
Build the utterance set from real conversations
Start from real transcripts where you have them, then fill the gaps on purpose. Salesforce's testing guidance stresses volume, diversity and quality, including inputs outside the agent's scope2.
What every Agentforce test case should assert
A test case pairs an utterance with an expected subagent, expected actions and an expected response; each run reports the actual values and a pass or fail for each23.
| Assertion | What it checks | How to judge it |
|---|---|---|
| Expected subagent | The router sent the request to the right subagent | Exact match; repeated misses point to overlapping classification descriptions |
| Expected actions | The right actions ran, in a sensible order | All expected actions present; any extra actions reviewed for cost and side effects |
| Expected outcome | The reply means what it should | Meaning rather than wording; the outcome check can pass when the gist of both texts matches3 |
| Record changes | Fields updated on the case or opportunity | Query the sandbox record after the run, because the reply alone cannot show it |
| Escalation | Out-of-scope or failed conversations reached a person | A routed work item exists, with the conversation context attached |
Running batch tests across iterations
Create the test suite
In Agentforce Studio's testing area, upload a CSV of cases or generate cases from the agent's subagents, then add your own edge and adversarial cases1.
Choose what to score
Set the target agent, any context variables and the evaluation criteria, such as instruction adherence, grounding and response quality2.
Read every failure
Open failed cases in the builder's preview to trace the subagent, reasoning and actions used, then decide whether the agent or the expectation was wrong1.
Have a person check passes as well
Sample passing cases for tone, accuracy and context-specific mistakes an automated judge can miss2.
Automate the suite for regression
Keep the YAML test spec in source control and run it from the Salesforce CLI with the agent test create and run commands, so every deployment repeats the suite3.
Permission and data-access tests as the agent user
Human handoff tests through Omni-Channel
When a service agent escalates, an outbound Omni-Channel flow routes the conversation to a queue or a rep, with a fallback queue if routing fails4.
Review trust and safety settings before launch
Moving from sandbox to production under change control
Deploy the agent, its actions, permission sets and the test spec together through your usual pipeline, then repeat the suite in a sandbox whose data resembles production. Record which version of the agent and its actions passed which run, so a later regression can be traced to a specific change.
Salesforce ships major releases in spring, summer and winter. Schedule a full regression run in a preview sandbox before each release reaches production, and after any change to the knowledge, Flows or permission sets the agent depends on.
What to watch in the first weeks after launch
Failed or abandoned conversations
Early signalCustomers leave mid-conversation or rephrase the same question again and again.
MitigationReview these transcripts every week and add them to the test suite before changing the agent.
Escalation rate drifting
Early signalNoticeably more conversations reach people than in the pilot, or noticeably fewer, which can mean the agent is overreaching.
MitigationSample escalated and contained conversations side by side to see which direction is wrong.
Action errors on unseen data
Early signalFlows or Apex actions fail on records the test set never included.
MitigationAlert on action failures and route the affected records to a named owner.
Behavior changes after a release
Early signalRouting or wording shifts after a Salesforce release with no change on your side.
MitigationCompare suite results before and after the release to pin down which subagent or action moved.
Questions and answers
How many test cases are enough before an Agentforce launch?
There is no universal number. Salesforce's own exercise starts from a small generated set and grows it2. In practice, cover every subagent with several phrasings of each goal, the success and failure paths of every action, and a meaningful share of off-topic and adversarial cases. After launch, add each failed real conversation to the suite so coverage grows from evidence.
Do we need to re-test after each Salesforce release?
Yes. Platform and reasoning changes can alter routing or wording even when your configuration has not changed. Run the full suite in a preview sandbox before each major release reaches production and compare the results with your last passing run, so you can see exactly which cases moved and fix them before customers notice.
Can we run Agentforce batch tests in production?
Salesforce advises against it: batch tests can modify CRM data and consume billable requests, so they belong in a sandbox1. After launch, production assurance should come from monitoring real conversations and session traces, plus a small set of read-only smoke checks agreed with your administrators.
Who should write the test cases for an agent?
The people who handle the work today. Service reps and team leads know how customers phrase requests and which edge cases cause trouble; admins and developers turn those into structured cases with expected subagents and actions. A process owner should sign off the expected outcomes, because they are business decisions rather than technical ones.
Sources
- Explore agent testing tools and considerations — Salesforce Trailhead · checked 10 October 2026
- Refine your agents using a five-step testing strategy — Salesforce Trailhead · checked 10 October 2026
- Test an Agent with Agentforce DX — Salesforce Developers · checked 10 October 2026
- Create a Sample Outbound Flow and Handle Routing for Agentforce Service Agents — Salesforce Developers · checked 10 October 2026
- Trust Layer — Salesforce Developers · checked 10 October 2026