
Most enterprise evaluation frameworks were built for deterministic software: the feature either exists or it does not, the SLA either holds or it does not, and the value is unlocked by the licence. AI breaks all three assumptions, which is why a procurement scorecard designed for a CRM produces confident nonsense when pointed at an AI partner.
Capability
- What is the narrowest problem you are unambiguously the best at? Worry if the answer is a market segment rather than a problem.
- Show us three outputs your system got wrong last month. Worry if there aren't any to hand — it means nobody is measuring.
- What quality bar do you hold, and how is it measured? Worry at "it depends on the use case" with no follow-up.
Delivery
- Who exactly staffs this engagement, and what else are they on? Worry if the names differ from the people in the room.
- What proportion of your deployments are running unassisted a year later? Worry at hesitation more than at the number.
- What does handover look like, and who owns the runbook? Worry if handover is not a defined event.
Risk
- Where does our data physically sit during processing, and for how long? Worry at a diagram without a retention answer.
- Describe the last legal review you failed and what you changed. Worry if they've never failed one.
- What is your dependency on any single model provider? Worry at "we're model-agnostic" with no migration story.
Commercial
- If this works, our headcount on the task falls. How does your pricing survive that? Worry at seat-based pricing on an outcome product.
- What does exit cost us in month thirteen? Worry at anything that isn't a number.
- What is not included that we will discover later? Worry at "nothing" — there is always something.
How to weight it
Weight delivery evidence at roughly half the total. Capability is easy to demonstrate and hard to fake for twenty minutes; delivery is the variable that decides whether anything is running in a year. Risk and commercial matter, but they rarely separate two credible finalists — delivery always does.
You can run this entire scorecard in one live session. Four pitch meetings will not get you the same information.
The efficient version: ask the twelve questions in a room the vendor does not control, with your operators present. The answers change when the audience does.
See who's already vetted for this.
Every partner in the directory has completed A³ review — enterprise references, capability, domain fit and delivery readiness. Live Proven is earned in an A³ session.