ProcessSalesforce Agentforce Center of Excellence
Setting up an Agentforce operating model: roles, intake and governance
Running several Agentforce agents across business units takes more than good builders. It needs named roles, one intake route, a scoring method that keeps value, data readiness and operational risk visible, shared standards every agent meets, and a board that decides when to scale or retire an agent. This page sets out that operating model step by step, with a sample RACI and a hypothetical first year.
On this page
- Where the CoE's authority starts and stops
- Core roles in an Agentforce operating model
- A sample RACI for agent delivery
- From proposal to a scaled or retired agent
- Setting up the operating model, in order
- Reading value, data readiness and risk scores together
- A hypothetical first year for a sales and service organization
- Shared standards an agent must meet before release
- Questions and answers
- Sources
Core roles in an Agentforce operating model
Small programs combine roles. What matters is that each has a named holder and that nobody approves their own work.
- CoE lead
- Runs intake and the review board, owns the standards and arbitrates between business units on priority.
- Business product owner
- Accountable for one agent's purpose, acceptance criteria and results; usually a manager of the process the agent supports.
- Agent designer
- Shapes subagents, instructions and the action boundary, and turns the product owner's criteria into testable behavior.
- Salesforce admin and developer
- Builds actions in Flow or Apex, configures the agent user's permissions and manages deployment between environments.
- Data and knowledge steward
- Owns the records and articles an agent reads, their quality and their review dates.
- Security and compliance reviewer
- Approves permission sets, data access and any action that changes customer records or commitments.
- Evaluation owner
- Maintains the agent's test set, runs it before each release and adds cases from production exceptions.
A sample RACI for agent delivery
| Activity | Product owner | CoE lead | Agent designer | Admin or developer | Knowledge steward | Evaluation owner |
|---|---|---|---|---|---|---|
| Propose a use case | A/R | C | C | I | C | I |
| Score and approve for build | C | A | R | C | C | I |
| Define acceptance criteria | A | C | R | I | C | R |
| Build actions and permissions | I | C | C | A/R | C | I |
| Maintain knowledge sources | A | I | C | I | R | C |
| Maintain the evaluation set | C | C | C | I | C | A/R |
| Approve a release | A | R | C | R | I | C |
| Monitor after release | A | R | C | R | I | R |
R = responsible, A = accountable, C = consulted, I = informed. The security reviewer is consulted on every row touching permissions, data or customer-facing changes. Adapt the assignments, but keep one accountable person per row.
From proposal to a scaled or retired agent
- Intake proposal
A one-page description of the process, its volumes, systems, knowledge and proposed owner.
- Three-way scoring
Value, data readiness and operational risk rated separately, with the reasoning published.
- Design review
Action boundary, reuse from the library and permission pattern agreed before build.
- Build and evaluate
Actions built or reused, and the evaluation set run until the acceptance criteria hold.
- Release board
Release approved against the shared standards and the release calendar.
- Production review
Usage, exceptions, handoffs and consumption compared with the acceptance criteria.
- Scale, fix or retire
A recorded decision at each review, so no agent runs on inertia.
Setting up the operating model, in order
Charter the CoE
Write a one-page charter: the decisions the CoE makes, those it only advises on, who it reports to and how disputes between business units are settled.
Name the roles
Assign the roles above, combining them where the team is small, and record who covers each one during absences. A role nobody holds becomes a gap nobody notices.
Open a single intake route
Ask every team to propose agents in the same format: the process and its volumes, how it is handled today, the systems and knowledge involved, the proposed action boundary and the intended owner.
Score on three separate scales
Rate business value, data readiness and operational risk independently, as in the portfolio stage of ColdAI's CoE approach1, and publish each score with its reasoning so teams can contest it.
Publish shared standards
Agree naming for agents, subagents and actions, a shared action library with an owner per action, permission patterns for agent users, and an owner and review date for every knowledge source.
Make evaluation sets shared assets
Keep every agent's test cases in one place and format, with a named owner, a change log and a rule for adding cases found in production.
Set the release cadence
Plan agent releases around Salesforce's seasonal upgrades: re-run evaluation sets in a preview sandbox before each upgrade reaches production, and avoid launches in the weeks either side of it.
Stand up the review board
Meet on a fixed cycle to approve releases and review production agents against acceptance criteria, consumption and exception patterns, ending each review with a scale, fix or retire decision.
Reading value, data readiness and risk scores together
- If
Value is high, data is ready and operational risk is low.
ThenBuild next, and use the agent as the reference implementation for your standards.
A clean first case shows teams what the standards look like in practice.
- If
Value is high but the records or knowledge the agent needs are poor.
ThenFund the data or knowledge clean-up as its own work item, then re-score.
An agent grounded in unreliable content fails in ways that damage trust in the whole program.
- If
Value is high but the agent would change customer commitments, prices or money.
ThenNarrow the action boundary or keep a human approval step until production evidence supports widening it.
Operational risk is reduced by design choices, not by a higher value score.
- If
Value is low, however easy the build looks.
ThenDecline or park the proposal and say why.
Easy, low-value agents still consume administrator time, review effort and usage budget.
A hypothetical first year for a sales and service organization
Questions and answers
Should an Agentforce CoE be centralized or federated?
Most start centralized while the standards are new, then let business-unit teams build within them once the action library, evaluation format and release gate are stable. The CoE keeps the standards and the review board; federated teams build and own their agents. The general trade-offs between centralized, hub-and-spoke and federated models are compared in the AI operating model guide.
How small can an Agentforce CoE be?
Small enough to be part-time at first. One person can lead the CoE and design agents, an administrator can own builds and permissions, and a business analyst can own evaluation sets. What matters is that each role has a name, that release decisions are not left to the person who built the agent and that the time is genuinely allocated rather than borrowed.
Where do implementation partners fit in the operating model?
Partners can design agents, build actions or run the CoE's early cycles, but the decisions should stay inside your organization: intake priorities, acceptance criteria, release approval and ownership of the shared libraries. Ask any partner, ColdAI included, to hand over standards and evaluation sets in a form your team can maintain without them.
Who owns the evaluation set when one agent serves several teams?
Name one evaluation owner per agent, usually in the team whose process carries the most risk, and let other teams add cases through a reviewed change. Splitting ownership by subagent can work for large agents, provided one person still signs off the complete set before each release.
How often should the review board meet?
Often enough that releases do not queue. Many programs start every two weeks or monthly and add an expedited route for fixes. Review production agents on the same cycle, so scale and retire decisions use current usage, exception and consumption data rather than impressions from launch week.