ProcessCustom Software Development
AI agent production readiness: the gates between a proof of concept and live customers
A proof of concept shows that an agent can do a task. Production asks harder questions: does it do the task reliably on real data, refuse what it should, leak nothing, stay within budget and stop cleanly when something fails? This page sets out six gates and a release step, the evidence each one must produce, and the pack a security reviewer will expect to see.
On this page
- What changes between a demo and production
- Six gates and a release step, in order
- The evidence each gate has to produce
- Setting evaluation thresholds by risk category
- Reading shadow mode against what staff actually did
- A hypothetical refunds agent works through the gates
- Go/no-go rules to write down before the release window
- The evidence pack a security reviewer will ask for
- Questions and answers
- Sources
What changes between a demo and production
A proof of concept is usually judged by a few people watching it succeed on examples they picked. Four things change when it goes live. The data turns messy, partial and personal. Adversaries appear, because anything that reads text from customers, emails or web pages can be instructed by that text. Load arrives in bursts, with rate limits and provider outages. And accountability becomes specific: someone must be able to explain what the agent did with one customer's request, and why.
The proof of concept was not wasted; the team now needs a different kind of evidence. The gates below organise it so engineering, security and the business owner agree in advance what 'ready' means. Thresholds change with the risk; the structure does not.
Six gates and a release step, in order
The gates run in sequence because later ones rely on earlier evidence: there is little point load-testing an agent whose permissions are undecided.
- Evaluation
A task suite from representative data, with thresholds and regression runs.
- Permissions
Allow-listed tools, scoped credentials and approval for side effects.
- Security
A red-team log showing injection and exfiltration routes were tested, with findings ranked.
- Data protection
Minimal personal data in prompts, redacted logs, policy-based retention.
- Observability
A trace per model and tool call, with cost and latency alerts.
- Resilience
Timeouts, provider failover, a degraded mode and human takeover.
- Staged release
Shadow, then partial traffic against agreed criteria, with a rehearsed rollback.
The evidence each gate has to produce
Each gate closes with an artefact someone other than the builder can check. Owners shown are typical; yours may differ.
Evaluate on representative data
Build the suite from real, anonymised cases, including awkward ones: incomplete requests, the wrong language, requests the agent must decline. Give each case a pass condition, scored by code where possible and a written rubric otherwise. Fix the threshold before the first run, and rerun the suite whenever a prompt, model, retrieval source or tool changes. How to build and maintain the suite itself is covered in the LLM evaluation sets guide.
Limit what the agent can do
List every tool and the operations it exposes. Give the agent its own credentials, scoped to those operations, instead of a shared service account. Separate reads from side effects such as refunds, emails or record changes, and require approval for side effects until evidence supports more autonomy. Excessive agency is one of the risks OWASP names for LLM applications.1
Test against injection and exfiltration
The gate requires a red-team log with ranked findings covering direct injection, indirect injection through retrieved content or tool output, and exfiltration through rendered links. The controls that close those routes are set out in the agent security checklist.
Minimise and redact personal data
Send the model only the fields the task needs. Redact or tokenise identifiers before prompts and responses reach the logs, and prove the redaction on real samples. Set trace retention to match your data policy, and record which providers receive which data.
Trace every step and budget it
Record each model and tool call with inputs, outputs, latency and cost against a trace ID that links to the originating request. Set per-task cost and latency budgets with alerts, and cap loops so a confused agent cannot run indefinitely; unbounded consumption is also on the OWASP list.1
Plan for the provider failing
Put timeouts on every external call. Decide what happens when the primary model fails: a secondary model that passed the same evaluation suite, a reduced tool set, or drafting replies for staff. Hand over to a person with context attached, and test it by cutting the provider off in staging.
Release in stages with a way back
After shadow mode, route a small slice of real traffic through the agent. Keep the previous version available through blue-green or canary deployment, and rehearse the rollback, including switching the agent off.
Setting evaluation thresholds by risk category
A single pass rate hides the failures that matter. Score each category separately, setting its bar by what a wrong answer costs. Examples are hypothetical.
| Category | Example cases | How strict the bar is | What a miss blocks |
|---|---|---|---|
| Information only | Order status, opening hours, policy wording | Moderate; checked against the source of record | Release of that intent only |
| Drafts for staff | Suggested replies a person edits before sending | Lenient on wording, strict on facts and tone | Little, since a person reviews every draft |
| Actions involving money | Refunds, credits, discount codes | Strictest; out-of-policy approvals count as failures | Autonomous actions; assist mode may still launch |
| Actions on personal data | Address changes, data exports, account merges | Strict, with identity checks scored as their own cases | Autonomous actions on accounts |
| Required refusals | Requests outside policy or another customer's data | Strict; a missed refusal is weighted like a wrong action | Release, until fixed and retested |
Write the thresholds down before the first run; raising a bar later is fine, lowering one to fit the results is not.
Reading shadow mode against what staff actually did
In shadow mode the agent records what it would do with live requests while staff handle them. The useful output is not an overall agreement rate but a sorted list of disagreements. Review each one and label it: the agent was wrong, the member of staff was wrong, both answers were acceptable, or the agent declined where staff acted.
Count those labels per risk category, using the evaluation suite's categories, so the two kinds of evidence line up. Staff decisions are not ground truth: where the agent applied policy more consistently, the finding is about the process. Review every disagreement in money or personal-data categories before traffic moves, and add new kinds of mistake to the regression suite.
A hypothetical refunds agent works through the gates
Go/no-go rules to write down before the release window
Criteria agreed after the results are in tend to bend. Teams fix rules like these in advance; your thresholds will differ.
- If
The regression suite falls below threshold in any category involving money or personal data.
ThenNo-go for autonomous actions; the agent may launch in assist mode, drafting for people.
Averages hide failures in the categories that cause harm.
- If
A critical or high-severity security finding remains open.
ThenNo-go, unless the named risk owner accepts it in writing with a dated fix.
Findings accepted by default become permanent.
- If
The failover test did not end in a working degraded mode or a human takeover.
ThenNo-go until it does.
Outages will happen; what matters is what the agent does then.
The evidence pack a security reviewer will ask for
Assemble it as each gate closes, not in the week before review; each item should be a dated document, report or dashboard link.
Questions and answers
How long does it take to make an agent proof of concept production-ready?
It depends on what surrounds the agent: how many tools it calls, whether they change records, how sensitive the data is and what your security review expects. An agent that only drafts replies for staff clears the gates far sooner than one that issues refunds itself. Scoping the gates against your systems is the reliable way to estimate.
How long should an agent run in shadow mode before it handles real traffic?
Long enough to see every risk category with enough cases to judge, including the rare but costly ones, and to cover the cycles your traffic follows, such as month-end or seasonal peaks. Calendar time matters less than coverage. Agree before starting which categories need how much evidence, and move each to live traffic separately rather than switching the whole agent at once.
Should a person approve every action the agent takes?
Not forever, but usually at first for actions with side effects. Approvals produce labelled evidence of where the agent's proposals are right, which justifies widening its autonomy later. Many teams move to approval above a threshold, such as a refund value, once results support it. The human-in-the-loop design guide covers how to widen autonomy safely.
What should we keep from the proof-of-concept code?
Keep whatever the evaluation shows works, often the prompts, task decomposition and tool definitions, and expect to rebuild the plumbing around them. Proofs of concept rarely have timeouts, scoped credentials, tracing or redaction. Record what was kept and why in an architecture decision record so the reasoning outlasts the people who made the call.
Sources
- OWASP Top 10 for LLM Applications, 2025 edition — OWASP Gen AI Security Project · checked 10 October 2026