AI Projects That Succeed (and Fail)
From pilot theater to real operational value
Listen to executive summary
Audio length ~2 min
Full lesson about ~35 min
Executive summary
~8 min to read the written sections · ~2 min to listen to this summary
This week separates pilots that impress from services that operate: production readiness (owner, metric, monitoring, support, retest after model change) and the higher bar for agents that act across systems.
You will phase automation carefully — recommend first, execute under limits later — with kill switches and on-call ownership.
Outcome: a production gate you can use in steering committees when demos claim they are “ready to launch.”
Core concepts
The terms and comparisons you need for this week’s decisions — with examples by function.
Pilot to production
The path from experiment to reliable operation: business owner, integration, monitoring, support, retest after model changes, and a business metric — not only a successful demo.
Why this works: Most AI value dies after the slide deck. Production is an operating service.
Why it matters: Fund change management and verification capacity, not only licenses.
Example: A 20-contract pilot that never gets SSO, support, or monitoring stays a slide, not a service.
Production gate
Weak / before: Demo went well — launch next week.
Strong / after: Score owner, metric, access control, monitoring, support, retest plan. Any critical 1–2 blocks launch.
By function
Legal
Agents that act on contracts need logs and kill switches — not chatbot terms alone.
HR
No agent execution on people decisions; recommend-only until counsel clears.
Purchasing
Production gate in SOW: owner, metric, support, retest, exit.
Marketing
Auto-publish agents blocked until dual control on claims.
Finance
Dual control thresholds for AI-initiated payments or refunds.
Operations
Phase agents: recommend first; execute under hard limits later.
Sales
Agent that updates CRM: permission scope no wider than human role.
Risk
Demo ≠ launch — production readiness scores in steering packs.
Agentic systems
AI that takes multi-step actions across systems (read email, update CRM, issue refund), not only drafts text. Higher leverage, higher need for permissions, logs, and kill switches.
Why this works: Treating agents like chatbots under-controls money movement and data changes.
Why it matters: Phase agents: recommend first; execute under hard limits later with on-call.
Example: An agent that books meetings and updates Salesforce — not just drafts a reply.
Refunds
Weak / before: Agent handles refunds end-to-end.
Strong / after: Phase 1 recommend only. Phase 2 execute under thresholds with dual control and kill switch.
vs Copilot: Copilot assists inside one task; agents chain tools. Govern agents like automation programs.
Operational metrics
Business measures — cycle time, error rate, rework, satisfaction — not only model accuracy scores from a vendor lab.
Why this works: Model vanity metrics do not pay salaries or reduce risk. Sponsors need outcome language.
Why it matters: Put the business metric on the pilot charter before day one.
Example: Days to contract signature down 20% with error rate flat — better than “92% model accuracy” alone.
Deep dive lesson
~14 min readDeep dive: From pilot to production (including agents)
Learning objective. Score a pilot’s production readiness and explain why agents need automation-grade controls, not chatbot manners.
Context. Most AI value dies after the demo. This lesson is about operating services — owners, monitoring, and why “agent” is not just a cooler chatbot.
1. Production is an operating service
Production needs owners, monitoring, support, integration, and business metrics — not vanity model scores. Fund change management and verification capacity, not only licenses. Retest after model and data changes.
Agents act multi-step across systems. They need permissions, logs, and kill switches. Treat them closer to automation than to a writing assistant. Copilot assists you; agent acts — different risk and different funding.
2. Production readiness score
Score 1–5: named business owner, business metric, data access control, monitoring/alerts, support/on-call, retest plan after model change. If any critical score is 1–2, you have a demo, not a launch. Say that out loud in steering committees.
Method: Production gate (must all be true)
Owner
Business owner with budget and authority.
Metric
Business KPI, not seats.
Controls
Access, logs, kill switch for any automation.
Support
Who answers when it fails at 2 a.m.?
Retest
Calendar owner after model/data change.
Worked example: Agent that “handles refunds”
Situation
Ops wants an agent to process refunds end-to-end in the order system.
How an executive thinks it through
- Money movement = high stakes.
- Need limits, dual control above thresholds, full logs.
- Rollback and on-call required.
- Eval on edge cases (fraud, partial ship, policy exceptions).
Decision / what to say
Phase 1: recommend refund only. Phase 2: execute under hard limits after eval and security review. Named on-call and kill switch before any execution rights.
Apply in your function
Operations
No agent execution rights without kill switch and on-call.
Finance
Dual control thresholds for AI-initiated money movement.
IT / Risk partners
Permissions for agents no wider than human roles.
Common mistakes
- Calling a demo a launch.
- No business owner — only a project manager.
- Agent permissions wider than any human role.
Practice (12 minutes)
- Pick one AI pilot you know.
- Score 1–5: owner, metric, data access, monitoring, support.
- Write the top blocker to production and who must remove it.
Check your understanding
Key takeaways
- Business process change is the hard part.
- Metrics must be operational, not only model accuracy.
- Kill criteria protect resources and credibility.
- Agents need stronger controls than chat drafts.
Practical applications
Map the handoffs where AI outputs enter human workflows.
Fund integration and training, not only licenses.
Plan role impacts and reskilling alongside tools.
Write kill criteria before the pilot starts.
Block auto-publish agents until dual control on claims; production gate before scale.
Library materials for this week
Extend this week’s decisions with briefs, cases, and primary sources that matter for your next meeting.
Pilot to Production Playbook
Stages from problem framing through pilot, evaluation, integration, training, and scale — with gates and anti-patterns from common failures.
DBS Bank — AI value at enterprise scale
Published 2025 results coverage (Business Times)
Singapore’s DBS Bank has publicly reported large economic value from AI programs — including Business Times coverage that the bank unlocked about S$1 billion in AI-related economic value in 2025 — after years of digital transformation and governed AI use cases across the franchise.
Optional focus hour
After the core lesson (~35 min), spend about 55 minutes on one long-form source matched to this week’s level.
BCG AI Radar 2026 — As AI Investments Surge, CEOs Take the Lead
Boston Consulting Group
Published 15 January 2026
Why this week
Pilot-to-production week: capital, CEO ownership, and the impact gap — specific Radar 2026 piece.
What to take away
Why pilots stall; one ownership or funding change for your program.
Check your understanding
Week 10 · 3 short questions · no grades shared outside this device
Select an answer for each question.
Hands-on
~10 minWrite kill criteria
Stopping failed work is a leadership skill.
- For a real or imagined pilot, write three conditions that would stop expansion.
- Include at least one quality/safety condition and one adoption condition.
- Name who has authority to stop it.
Reflection
- Which current initiative would you struggle to stop even if metrics were weak?
- Who experiences the pain if the AI is wrong — and were they in the pilot design?