ChecklistFrontier Technologies
Technology readiness level assessment: evidence, baselines and common mis-ratings
Technology readiness levels give a shared vocabulary for maturity, from basic principles observed to a system proven in operation. Used carelessly, they become a vendor's self-rating. This checklist pairs each level with the evidence that justifies it, adds the adoption readiness the scale leaves out, and shows how to run an assessment workshop that produces a rating your committee can defend.
On this page
- Where the TRL scale comes from, and what it measures
- TRL 1 to TRL 9 in plain language, with the evidence each level needs
- Adoption readiness checks the TRL scale leaves out
- Choosing the conventional baseline to rate against
- Common TRL mis-ratings and how to catch them
- Running a readiness assessment workshop
- The fields of a technology readiness assessment record
- From a readiness rating to a pilot, watch or stop decision
- Questions and answers
- Sources
Where the TRL scale comes from, and what it measures
Technology readiness levels were developed at NASA to describe how far a technology had progressed from research towards flight, and NASA still publishes its own definitions2. The European Commission adopted a version for its research funding programmes, defining TRL 1 as basic principles observed and TRL 9 as an actual system proven in an operational environment1.
The scale measures technical maturity in an environment: lab, relevant or operational. It says nothing about whether a technology is legal, affordable, supportable or wanted. And the environment is the part most often skipped. A technology can be proven in one operational setting and barely past the lab in yours, so the first rule of any assessment is to state the use case and environment before assigning a number.
TRL 1 to TRL 9 in plain language, with the evidence each level needs
| Level | Definition used in EU research programmes | Evidence that justifies it for your use case | Evidence that does not |
|---|---|---|---|
| TRL 1 | Basic principles observed | Published research describing the underlying effect | A product roadmap or a press release |
| TRL 2 | Technology concept formulated | A documented concept of how the effect could solve your class of problem | Enthusiasm from a conference talk |
| TRL 3 | Experimental proof of concept | A lab experiment, ideally reproducible, showing the key function works | A simulation with no physical or real-data test |
| TRL 4 | Technology validated in lab | Components working together in a lab on representative inputs | A demonstration on curated data chosen by the supplier |
| TRL 5 | Technology validated in relevant environment | Tests under conditions close to yours, including realistic noise, data and users | Lab results described as field-ready |
| TRL 6 | Technology demonstrated in relevant environment | A prototype system demonstrated end to end in a representative setting | A single successful run with no repeat or failure log |
| TRL 7 | System prototype demonstration in operational environment | A prototype operating in a real environment comparable to yours, with users | Operation in a different sector with different constraints |
| TRL 8 | System complete and qualified | A finished system that has passed acceptance, safety or certification tests relevant to you | A vendor's own qualification report with no independent check |
| TRL 9 | Actual system proven in operational environment | Sustained operation by organizations like yours, with support and incident history | A list of logos without references you can call |
Definitions in the second column follow the European Commission's Horizon 2020 general annex1; NASA's wording differs in detail but follows the same progression2.
Adoption readiness checks the TRL scale leaves out
A technology at a high TRL can still be unready for you. Rate each of these separately and keep them visible next to the TRL.
Choosing the conventional baseline to rate against
A maturity rating only becomes a decision when it is set against the best alternative you could deploy today.
- If
A mature conventional system already does the job adequately
ThenRate the frontier option against that system's actual performance and cost, measured on the same tasks.
The question is whether the new option is better enough to justify its risk, not whether it works at all.
- If
The current approach is manual or ad hoc
ThenUse an improved conventional process, such as better tooling or documentation, as the baseline rather than the status quo.
Comparing against an unimproved process flatters every new technology.
- If
No conventional alternative exists
ThenRate against the cost of waiting, and make the review triggers explicit.
Without a baseline the decision is about timing, not about which option wins.
Common TRL mis-ratings and how to catch them
Lab results claimed as field readiness
Early signalEvidence comes only from controlled conditions, but the claimed level implies a relevant or operational environment.
MitigationAsk where each test ran and under what conditions; rate by the least realistic environment the evidence covers.
Vendor-supplied evidence only
Early signalEvery document in the evidence pack was produced by the supplier.
MitigationRequire at least one independent source per claimed level: a reference customer, a published study or your own test.
Demonstrations on curated data
Early signalDemos use examples chosen in advance, and the supplier is reluctant to run on yours.
MitigationRun the demonstration on a sample of your own data, including difficult and messy cases.
Component rated as system
Early signalA mature core component is used to rate a system whose other parts are new.
MitigationRate the least mature critical component and the integrated system separately.
Stale evidence
Early signalThe rating relies on results from an older version of the technology.
MitigationRecord the version each piece of evidence applies to, and discount evidence for superseded versions.
Running a readiness assessment workshop
Fix the use case and environment
Write one paragraph describing what the technology would do, for whom and in which conditions. Every rating refers back to it.
Assemble an evidence pack
Collect test reports, references, papers and your own trial results in advance, each tagged with its source, date and version.
Rate independently first
Ask each participant, including technical, operational, security and risk representatives, to rate the TRL and adoption readiness alone.
Discuss the differences
Spend most of the time on where ratings diverge; disagreement usually reveals an evidence gap or a different assumed environment.
Record the agreed rating
Capture the level, the evidence behind it, confidence, any dissent and what evidence would move it up a level.
Set a revisit date
Agree when, or on what trigger, the rating should be reviewed, and who owns that review.
The fields of a technology readiness assessment record
From a readiness rating to a pilot, watch or stop decision
A rating at or above the relevant environment level, with adoption readiness that has no blocking gaps and a credible advantage over the baseline, usually justifies a bounded pilot. A promising rating blocked by regulation, integration or supplier readiness belongs on a watch list with explicit triggers. A low rating with no advantage over the baseline is a stop, recorded with its reasons so the idea is not reassessed from scratch next year.
How much money is released at each stage, and the kill criteria for a funded pilot, are a separate question covered under business building. The rating's job is to make sure that decision rests on evidence. Candidates usually reach this assessment from a horizon scan, and promising ones move into the staged approach described on our frontier technologies page3.
Questions and answers
Can a technology skip readiness levels?
Not legitimately, though evidence for several levels can arrive at once, for example when a mature system from another sector is tested in your environment. What happens more often is that levels are skipped in the claim rather than in reality, such as a lab result presented as operational readiness. Check that evidence exists for each level up to the one claimed.
Who should rate the technology?
A small panel that includes people who would build, operate, secure and be accountable for the technology, chaired by someone without a stake in the answer. Suppliers should provide evidence but not take part in the rating. Independent individual ratings before discussion reduce the pull of the most senior or most enthusiastic person in the room.
How often should a readiness rating be revisited?
Whenever material evidence changes, such as a new version, a published validation or a regulatory decision, and otherwise on a fixed review date agreed at the workshop. Ratings for fast-moving technologies go stale quickly, so every rating should carry the date it was made and the version it applies to.
Are technology readiness levels useful for software and AI?
Yes, with care. The scale was designed for hardware, so interpret the environments for software: relevant environment means realistic data, users and integrations; operational environment means production use with monitoring and support. For AI systems, evidence should include evaluation on your own data, because performance on public benchmarks often fails to carry over.
Sources
- Horizon 2020 Work Programme 2014-2015, General Annexes, G. Technology readiness levels (TRL) — European Commission · checked 10 October 2026
- Technology Readiness Levels — NASA · checked 10 October 2026
- Frontier Technologies: test against a baseline and decide the next commitment — ColdAI