Buying AI: Vendors, Build vs Buy, Contracts
How to evaluate tools without getting sold a demo
Listen to executive summary
Audio length ~2 min
Full lesson about ~40 min
Executive summary
~8 min to read the written sections · ~2 min to listen to this summary
This week is how executives buy AI under uncertainty: demos vs evidence, AI washing, total cost of ownership (including human review), and kill criteria written before pilots start.
You will use a one-page scorecard — capability evidence, data rights, security, TCO, exit — and refuse auto-conversion from free pilot to multi-year seats without an eval pack.
Outcome: procurement and functional leaders share a common buy/hold language.
Core concepts
The terms and comparisons you need for this week’s decisions — with examples by function.
AI washing
Marketing ordinary software as “AI” without meaningful learning or generative capability — or without evidence it works on your tasks.
Why this works: You can overpay and under-govern something that is just rules and dashboards with a new label.
Why it matters: Price and govern for real capability. Demand a pilot on your data with kill criteria.
Example: A rules engine rebranded as “AI-powered” with no learning and no eval evidence.
RFP filter
Weak / before: Shortlist any vendor that says AI.
Strong / after: Score evidence on our tasks, data rights, security, TCO, and exit — adjectives score zero.
By function
Legal
Standard AI clauses: training, model change, audit, exit.
HR
People-tech buys need employment counsel and eval rights in the contract.
Purchasing
Scorecard: evidence, data, security, TCO, exit — kill criteria before pilot day one.
Marketing
Agency AI suites: no multi-year conversion without eval package.
Finance
TCO includes human review labor, not seats alone.
Operations
Integration and support costs often dwarf license price.
Sales
CRM AI add-ons: proof on your pipeline data before enterprise seats.
Risk
Auto-renew after “free pilot” is a control failure waiting to happen.
Total cost of ownership (TCO)
Licenses plus integration, change management, support, and human review time. Verification labor is real cost.
Why this works: A cheap seat that needs twenty minutes of cleanup per draft is not cheap at team scale.
Why it matters: Model review capacity in the business case, not only software line items.
Example: Seats look inexpensive until Legal spends 20 minutes verifying every AI draft.
Business case line
Weak / before: Tool costs $30/user/month.
Strong / after: $30 seats + integration + 15 min review × volume + training. Net hours saved after review is the real ROI.
Kill criteria
Pre-agreed conditions that stop a pilot or scale-up when quality, risk, or adoption fails. Written before day one — not after budget is sunk.
Why this works: Without kill criteria, free pilots become accidental multi-year purchases.
Why it matters: Date the go/no-go and name the owner. No auto-conversion from pilot to enterprise agreement.
Example: If error rate on critical clauses exceeds 5% after two sprints, stop and redesign.
Pilot contract
Weak / before: 90-day free pilot then automatic conversion.
Strong / after: Written kill criteria, no auto-renew, data return/deletion, executive gate at day 60.
Deep dive lesson
~14 min readDeep dive: Buying AI without the buzzwords
Learning objective. Run a one-page scorecard (capability, evidence, data, security, TCO, exit) and write kill criteria before a pilot starts.
Context. Purchasing and functional leaders share this job: buy outcomes under uncertainty, not adjectives. This lesson is your anti-AI-washing kit.
1. Demos dazzle; total cost decides
Buying AI is buying outcomes under uncertainty. Price licenses plus integration plus human review time — verification labor is real cost. Demand sample-based evidence on your tasks. Require data rights, model-change notice, audit paths, and exit terms.
AI washing (slapping “AI” on rules engines or dashboards) is common; score proof, not adjectives. Build vs buy is about differentiation and control, not ego. If the capability is commodity and data is sensitive, prefer vendors with strong enterprise controls. If the workflow is your secret sauce, be careful what you outsource.
2. Kill criteria before day one
A pilot without kill criteria is a slow purchase. Write in advance: quality thresholds, risk incidents that stop the pilot, adoption with value (not seats alone), and a dated go/no-go owner. Free pilots still consume staff time and data — treat them as real projects.
Method: One-page AI buy scorecard (1–5 each)
Evidence
Eval on our tasks, not only demo.
Data rights
Training, retention, residency, deletion.
Security
Access control, logs, certifications you actually use.
TCO
Licenses + integration + human review capacity.
Exit
Export, offboarding, no hostage data.
Worked example: Multi-year “AI suite” with free pilot
Situation
Vendor offers 90-day free pilot then automatic conversion to a three-year agreement.
How an executive thinks it through
- Kill criteria must be written before day one.
- Free pilot still consumes staff time and data.
- Auto-conversion is a commercial trap.
- Need eval results, security review, and TCO model before conversion.
Decision / what to say
Pilot only with written kill criteria, no auto-renew, data return/deletion clause, and executive go/no-go gate at day 60.
Apply in your function
Purchasing
Attach scorecard + kill criteria template to every AI RFP.
Legal
Standardize AI clauses: training, model change, audit, exit.
Finance
Model TCO including review labor, not seats alone.
All functions
Refuse multi-year conversion without eval package.
Common mistakes
- Seat counts as success metric.
- Legal review only after the business already committed publicly.
- No exit test (can we export and leave?).
Practice (15 minutes)
- Take one recent AI pitch.
- Score 1–5: evidence, data rights, security clarity, TCO, exit.
- Write one kill criterion if this were a pilot.
- Decide: pilot / hold / no — in one sentence.
Check your understanding
Key takeaways
- Test on your data, not only vendor demos.
- Price the human review time into ROI.
- Contracts must cover data, IP, liability, and model changes.
- Pilots with exit ramps beat premature enterprise agreements.
Practical applications
Require a scored pilot plan before enterprise SKUs.
Template AI-specific clauses for IP and training rights.
Model three-year TCO including change management.
Validate integration with systems of record early.
Agency AI suites: no multi-year conversion without eval on your brand claims and kill criteria.
Library materials for this week
Extend this week’s decisions with briefs, cases, and primary sources that matter for your next meeting.
AI Vendor Scorecard
Weighted criteria for evaluating AI products: problem fit, evidence on your data, security/privacy, integration, TCO, exit, support, and roadmap realism.
Contract Clauses for AI Tools
Clause themes for AI SaaS: data use, IP ownership of outputs, infringement indemnity, SLAs beyond uptime, audit rights, and model change notices.
Optional focus hour
After the core lesson (~40 min), spend about 55 minutes on one long-form source matched to this week’s level.
WEF AI Governance Alliance — Briefing Paper Series (PDF)
World Economic Forum
Published January 2024
Why this week
Procurement week: multi-stakeholder governance language for RFPs and vendor diligence.
What to take away
Three diligence questions; one kill criterion for a weak pilot.
Check your understanding
Week 7 · 3 short questions · no grades shared outside this device
Select an answer for each question.
Hands-on
~15 minDraft a one-page pilot brief
Clear pilots stop endless demos.
- Write: problem, users, data types, success metric, fail condition, timeline (2–6 weeks), owner.
- List two must-have security answers before data leaves the building.
- Keep it to one page — if it doesn’t fit, the pilot is too vague.
Reflection
- Which AI tools in your stack were purchased without a defined success metric?
- Who can terminate an AI pilot that is failing quietly?