Buyer's guideStrategic Technology Consulting

AI vendor evaluation: criteria and scoring for proposals that do not compare

AI proposals rarely answer the same question. One prices per seat, another per token and a third per resolved case; each measures success differently and demonstrates on its own data. Before scoring anything, restate every proposal against one problem, one success measure and one evaluation on your data. Then check the contract terms, assurance evidence and exit route that decide whether the choice still looks right in two years' time.

Reviewed 7 min read

On this page
  1. Why AI proposals resist side-by-side comparison
  2. Five steps to put proposals on one footing
  3. A weighted scoring frame for AI and agent proposals
  4. Contract terms to confirm before signing an AI agreement
  5. Reading assurance evidence without over-reading it
  6. Red flags in AI vendor proposals
  7. Three proposals for one service desk, made comparable
  8. When an independent reviewer adds value
  9. Questions and answers
  10. Sources

Why AI proposals resist side-by-side comparison

A buyer who asks three vendors to solve one problem usually receives three different problems back. Each proposal reframes the scope around what its product does best, chooses the metric it performs well on and assumes volumes that make its pricing look favourable. Comparing them as submitted means comparing three sales strategies.

The fix is procedural rather than technical. Agree the problem, the measure and the evaluation before anyone scores, and require every vendor to answer inside that frame. A vendor unwilling to do so has told you something useful about the relationship that would follow.

Five steps to put proposals on one footing

  1. Restate the problem and the success measure

    Write one problem statement with a measured baseline and the change that would count as success, then ask each vendor to confirm or challenge it in writing.

    Output
    Common problem statement
    Owner
    Business owner
  2. Fix the volume and usage assumptions

    Give every vendor the same expected users, transactions and growth path, so prices are calculated on one basis rather than on each vendor's preferred scenario.

    Output
    Shared pricing assumptions
    Owner
    Finance and procurement
  3. Run an evaluation on your own data

    Prepare a representative, permission-cleared sample of real cases with agreed correct answers and a scoring method, as described in how to build an LLM evaluation set, and run each shortlisted product against it under the same conditions.

    Output
    Evaluation results per vendor
    Owner
    Data and technology leads
  4. Review terms and assurance in parallel

    Legal, security and privacy review data use, model lifecycle, assurance reports and exit provisions while the evaluation runs, so findings arrive together.

    Output
    Terms and assurance findings
    Owner
    Legal, security and privacy
  5. Score with weights agreed in advance

    Score each criterion separately, apply the weights fixed before proposals were opened, and record the reasoning behind every score so it can be challenged later.

    Output
    Scored comparison and recommendation
    Owner
    Evaluation panel

A weighted scoring frame for AI and agent proposals

Agree the weights before any proposal is opened. Most buyers weight evaluation results and data terms most heavily.

CriterionWhat to ask forScores wellScores poorly
Results on your dataScores on your evaluation set under agreed conditionsMeets the success measure with documented failure casesOnly results from the vendor's own demonstrations
Data useContract clauses on training, retention, location and deletionClear opt-outs, short retention and named processing locationsPolicy pages that can change without notice
Model lifecycleChange notification, version pinning and deprecation noticeAdvance notice and time to re-test before a model is switchedModels updated silently behind the same interface
Security assuranceCurrent assurance reports or certificates and their scopeScope explicitly covers the AI service being boughtScope limited to corporate IT or another product
Integration and identityHow it connects to your systems and enforces accessUses your identity provider and respects existing permissionsBroad service accounts or copied data stores
Commercial modelPrice under the shared volume assumptions, with capsPredictable at expected volume, with growth priced inLow entry price with unpriced usage beyond it
Exit and portabilityData export formats, exit assistance and transition periodYour data, configuration and logs returned in usable formatsExport limited to raw data with no configuration

Score each row on a simple scale the panel agrees, and keep evaluation results separate from the demonstration impression.

Contract terms to confirm before signing an AI agreement

0 of 10 checked

Reading assurance evidence without over-reading it

Vendors offer these documents as proof of security. Each is useful within limits that buyers often miss.

SOC 2 Type II report
An independent auditor's report on whether a service organisation's controls, assessed against the Trust Services Criteria, operated effectively over a stated period3. Read the system description and the exceptions.
SOC 2 Type I report
The same examination at a single point in time. It shows controls were designed suitably, not that they worked over time.
ISO/IEC 27001 certificate
Evidence that an accredited certification body audited the vendor's information security management system4. The certificate's scope statement says which services and sites it covers.
Statement of Applicability
The ISO/IEC 27001 document listing which controls the vendor applies and why others are excluded. Ask for it alongside the certificate.
ISO/IEC 42001 certificate
Evidence of an audited AI management system within a stated scope5. It addresses how AI is governed, not how well a given model performs.

Red flags in AI vendor proposals

Refusal to run on your data

Early signalThe vendor offers more demonstrations but declines a scoped evaluation on your sample.

MitigationMake the evaluation a condition of shortlisting, with a permission-cleared sample and a time limit.

Data assurances outside the contract

Early signalTraining and retention commitments appear in slides or blog posts but not in the draft agreement.

MitigationRequire each commitment as a clause and treat its absence as a scoring penalty.

Assurance that covers something else

Early signalThe report or certificate scope names corporate systems or a different product line.

MitigationAsk for evidence covering the AI service itself, or a dated plan to bring it into scope.

Pricing built on unstated assumptions

Early signalUnit prices are clear but volumes, overage rates or price review terms are missing.

MitigationPrice every proposal on the shared assumptions and ask for caps on usage-based charges.

Three proposals for one service desk, made comparable

When an independent reviewer adds value

Internal panels can run this process well, especially with the scoring frame above. Outside review helps when internal teams already favour a vendor, when the contract will be hard to reverse or when the board wants an assessment from someone without a stake in the result.

ColdAI reviews vendor proposals as a focused sprint assessment or as part of retained advisory. We do not resell the platforms we assess, and we disclose any vendor relationship at scoping, before you decide whether to proceed1. Public bodies buying under procurement law face additional rules this page does not cover.

Questions and answers

How many vendors should we evaluate in depth?

Enough to make a real comparison, but few enough to run a proper evaluation on your own data for each. Many buyers screen a longer list on written answers and contract terms, then run the hands-on evaluation with a shortlist of two or three. Evaluating more in depth usually means evaluating each one less carefully.

Should we let vendors see our evaluation data?

Share a permission-cleared sample under a confidentiality agreement, and keep a separate held-back set that no vendor sees. Scores on the held-back set show whether a product generalises or has been tuned to the cases it was shown. Remove or mask personal data unless you have a lawful basis to share it.

Is a SOC 2 report enough security evidence for an AI service?

It is a strong starting point if its scope includes the service you are buying and it covers a recent period. It does not address model behaviour, data-use terms or how outputs are reviewed. Pair it with the contract terms on data use and model changes, and with your own evaluation of the product on representative cases.

What if no vendor meets the success measure?

Treat that as a finding rather than a failure. Check whether the measure was realistic against your baseline, whether the data in the evaluation reflected real cases, and whether the problem needs process changes before any product can help. Sometimes the right outcome is to wait, or to reconsider building or partnering instead of buying.

Sources

  1. Strategic Technology Consulting: engagement models and adviser independence — ColdAI
  2. Regulation (EU) 2016/679 (General Data Protection Regulation) — EUR-Lex · checked 10 October 2026
  3. SOC 2: SOC for Service Organizations — Trust Services Criteria — AICPA & CIMA · checked 10 October 2026
  4. ISO/IEC 27001 — Information security management systems — International Organization for Standardization · checked 10 October 2026
  5. ISO/IEC 42001:2023 — AI management systems — International Organization for Standardization · checked 10 October 2026

More in Strategic Technology Consulting

Back to Strategic Technology Consulting

Next step

Send us the AI proposals you need to compare

Share the proposals, your problem statement and the decision date. We will say whether a sprint review fits or how your panel can run the scoring frame above itself.

Discuss a proposal review