ProcessCustom Software Development

AI agent production readiness: the gates between a proof of concept and live customers

A proof of concept shows that an agent can do a task. Production asks harder questions: does it do the task reliably on real data, refuse what it should, leak nothing, stay within budget and stop cleanly when something fails? This page sets out six gates and a release step, the evidence each one must produce, and the pack a security reviewer will expect to see.

Reviewed 8 min read

On this page
  1. What changes between a demo and production
  2. Six gates and a release step, in order
  3. The evidence each gate has to produce
  4. Setting evaluation thresholds by risk category
  5. Reading shadow mode against what staff actually did
  6. A hypothetical refunds agent works through the gates
  7. Go/no-go rules to write down before the release window
  8. The evidence pack a security reviewer will ask for
  9. Questions and answers
  10. Sources

What changes between a demo and production

A proof of concept is usually judged by a few people watching it succeed on examples they picked. Four things change when it goes live. The data turns messy, partial and personal. Adversaries appear, because anything that reads text from customers, emails or web pages can be instructed by that text. Load arrives in bursts, with rate limits and provider outages. And accountability becomes specific: someone must be able to explain what the agent did with one customer's request, and why.

The proof of concept was not wasted; the team now needs a different kind of evidence. The gates below organise it so engineering, security and the business owner agree in advance what 'ready' means. Thresholds change with the risk; the structure does not.

Six gates and a release step, in order

The gates run in sequence because later ones rely on earlier evidence: there is little point load-testing an agent whose permissions are undecided.

01Evaluation02Permissions03Security04Data protection05Observability06Resilience07Staged release
  1. Evaluation

    A task suite from representative data, with thresholds and regression runs.

  2. Permissions

    Allow-listed tools, scoped credentials and approval for side effects.

  3. Security

    A red-team log showing injection and exfiltration routes were tested, with findings ranked.

  4. Data protection

    Minimal personal data in prompts, redacted logs, policy-based retention.

  5. Observability

    A trace per model and tool call, with cost and latency alerts.

  6. Resilience

    Timeouts, provider failover, a degraded mode and human takeover.

  7. Staged release

    Shadow, then partial traffic against agreed criteria, with a rehearsed rollback.

Conceptual sequence of readiness gates for an agent moving from proof of concept to live traffic. Gates may overlap in time, but none is skipped.

The evidence each gate has to produce

Each gate closes with an artefact someone other than the builder can check. Owners shown are typical; yours may differ.

  1. Evaluate on representative data

    Build the suite from real, anonymised cases, including awkward ones: incomplete requests, the wrong language, requests the agent must decline. Give each case a pass condition, scored by code where possible and a written rubric otherwise. Fix the threshold before the first run, and rerun the suite whenever a prompt, model, retrieval source or tool changes. How to build and maintain the suite itself is covered in the LLM evaluation sets guide.

    Output
    Evaluation report and regression suite in CI
    Owner
    Engineering and business owner
  2. Limit what the agent can do

    List every tool and the operations it exposes. Give the agent its own credentials, scoped to those operations, instead of a shared service account. Separate reads from side effects such as refunds, emails or record changes, and require approval for side effects until evidence supports more autonomy. Excessive agency is one of the risks OWASP names for LLM applications.1

    Output
    Tool inventory with scopes and approval rules
    Owner
    Engineering and security
  3. Test against injection and exfiltration

    The gate requires a red-team log with ranked findings covering direct injection, indirect injection through retrieved content or tool output, and exfiltration through rendered links. The controls that close those routes are set out in the agent security checklist.

    Output
    Red-team log with ranked findings
    Owner
    Security
  4. Minimise and redact personal data

    Send the model only the fields the task needs. Redact or tokenise identifiers before prompts and responses reach the logs, and prove the redaction on real samples. Set trace retention to match your data policy, and record which providers receive which data.

    Output
    Data-flow diagram and redaction tests
    Owner
    Data protection lead
  5. Trace every step and budget it

    Record each model and tool call with inputs, outputs, latency and cost against a trace ID that links to the originating request. Set per-task cost and latency budgets with alerts, and cap loops so a confused agent cannot run indefinitely; unbounded consumption is also on the OWASP list.1

    Output
    Dashboards, alerts and sample traces
    Owner
    Platform team
  6. Plan for the provider failing

    Put timeouts on every external call. Decide what happens when the primary model fails: a secondary model that passed the same evaluation suite, a reduced tool set, or drafting replies for staff. Hand over to a person with context attached, and test it by cutting the provider off in staging.

    Output
    Failover test record
    Owner
    Engineering lead
  7. Release in stages with a way back

    After shadow mode, route a small slice of real traffic through the agent. Keep the previous version available through blue-green or canary deployment, and rehearse the rollback, including switching the agent off.

    Output
    Signed go/no-go record and rollback notes
    Owner
    Business owner, engineering, security

Setting evaluation thresholds by risk category

A single pass rate hides the failures that matter. Score each category separately, setting its bar by what a wrong answer costs. Examples are hypothetical.

CategoryExample casesHow strict the bar isWhat a miss blocks
Information onlyOrder status, opening hours, policy wordingModerate; checked against the source of recordRelease of that intent only
Drafts for staffSuggested replies a person edits before sendingLenient on wording, strict on facts and toneLittle, since a person reviews every draft
Actions involving moneyRefunds, credits, discount codesStrictest; out-of-policy approvals count as failuresAutonomous actions; assist mode may still launch
Actions on personal dataAddress changes, data exports, account mergesStrict, with identity checks scored as their own casesAutonomous actions on accounts
Required refusalsRequests outside policy or another customer's dataStrict; a missed refusal is weighted like a wrong actionRelease, until fixed and retested

Write the thresholds down before the first run; raising a bar later is fine, lowering one to fit the results is not.

Reading shadow mode against what staff actually did

In shadow mode the agent records what it would do with live requests while staff handle them. The useful output is not an overall agreement rate but a sorted list of disagreements. Review each one and label it: the agent was wrong, the member of staff was wrong, both answers were acceptable, or the agent declined where staff acted.

Count those labels per risk category, using the evaluation suite's categories, so the two kinds of evidence line up. Staff decisions are not ground truth: where the agent applied policy more consistently, the finding is about the process. Review every disagreement in money or personal-data categories before traffic moves, and add new kinds of mistake to the regression suite.

A hypothetical refunds agent works through the gates

Go/no-go rules to write down before the release window

Criteria agreed after the results are in tend to bend. Teams fix rules like these in advance; your thresholds will differ.

  • If

    The regression suite falls below threshold in any category involving money or personal data.

    Then

    No-go for autonomous actions; the agent may launch in assist mode, drafting for people.

    Averages hide failures in the categories that cause harm.

  • If

    A critical or high-severity security finding remains open.

    Then

    No-go, unless the named risk owner accepts it in writing with a dated fix.

    Findings accepted by default become permanent.

  • If

    The failover test did not end in a working degraded mode or a human takeover.

    Then

    No-go until it does.

    Outages will happen; what matters is what the agent does then.

The evidence pack a security reviewer will ask for

Assemble it as each gate closes, not in the week before review; each item should be a dated document, report or dashboard link.

0 of 7 checked

Questions and answers

How long does it take to make an agent proof of concept production-ready?

It depends on what surrounds the agent: how many tools it calls, whether they change records, how sensitive the data is and what your security review expects. An agent that only drafts replies for staff clears the gates far sooner than one that issues refunds itself. Scoping the gates against your systems is the reliable way to estimate.

How long should an agent run in shadow mode before it handles real traffic?

Long enough to see every risk category with enough cases to judge, including the rare but costly ones, and to cover the cycles your traffic follows, such as month-end or seasonal peaks. Calendar time matters less than coverage. Agree before starting which categories need how much evidence, and move each to live traffic separately rather than switching the whole agent at once.

Should a person approve every action the agent takes?

Not forever, but usually at first for actions with side effects. Approvals produce labelled evidence of where the agent's proposals are right, which justifies widening its autonomy later. Many teams move to approval above a threshold, such as a refund value, once results support it. The human-in-the-loop design guide covers how to widen autonomy safely.

What should we keep from the proof-of-concept code?

Keep whatever the evaluation shows works, often the prompts, task decomposition and tool definitions, and expect to rebuild the plumbing around them. Proofs of concept rarely have timeouts, scoped credentials, tracing or redaction. Record what was kept and why in an architecture decision record so the reasoning outlasts the people who made the call.

Sources

  1. OWASP Top 10 for LLM Applications, 2025 edition — OWASP Gen AI Security Project · checked 10 October 2026

More in Custom Software Development

Back to Custom Software Development

Next step

Send us the agent you want to put in front of customers

Describe what the proof of concept does, the tools it calls and the review it has to pass. We will reply with the gates that look hardest in your case and whether a short readiness sprint would close them.

Discuss agent readiness