Week 10Month 3~35 min lesson~2 min listen~8 min read

AI Projects That Succeed (and Fail)

From pilot theater to real operational value

Listen to executive summary

Audio length ~2 min

Full lesson about ~35 min

Executive summary

~8 min to read the written sections · ~2 min to listen to this summary

This week separates pilots that impress from services that operate: production readiness (owner, metric, monitoring, support, retest after model change) and the higher bar for agents that act across systems.

You will phase automation carefully — recommend first, execute under limits later — with kill switches and on-call ownership.

Outcome: a production gate you can use in steering committees when demos claim they are “ready to launch.”

Core concepts

The terms and comparisons you need for this week’s decisions — with examples by function.

Pilot to production

The path from experiment to reliable operation: business owner, integration, monitoring, support, retest after model changes, and a business metric — not only a successful demo.

Why this works: Most AI value dies after the slide deck. Production is an operating service.

Why it matters: Fund change management and verification capacity, not only licenses.

Example: A 20-contract pilot that never gets SSO, support, or monitoring stays a slide, not a service.

Production gate

Weak / before: Demo went well — launch next week.

Strong / after: Score owner, metric, access control, monitoring, support, retest plan. Any critical 1–2 blocks launch.

By function

Legal

Agents that act on contracts need logs and kill switches — not chatbot terms alone.

HR

No agent execution on people decisions; recommend-only until counsel clears.

Purchasing

Production gate in SOW: owner, metric, support, retest, exit.

Marketing

Auto-publish agents blocked until dual control on claims.

Finance

Dual control thresholds for AI-initiated payments or refunds.

Operations

Phase agents: recommend first; execute under hard limits later.

Sales

Agent that updates CRM: permission scope no wider than human role.

Risk

Demo ≠ launch — production readiness scores in steering packs.

Agentic systems

AI that takes multi-step actions across systems (read email, update CRM, issue refund), not only drafts text. Higher leverage, higher need for permissions, logs, and kill switches.

Why this works: Treating agents like chatbots under-controls money movement and data changes.

Why it matters: Phase agents: recommend first; execute under hard limits later with on-call.

Example: An agent that books meetings and updates Salesforce — not just drafts a reply.

Refunds

Weak / before: Agent handles refunds end-to-end.

Strong / after: Phase 1 recommend only. Phase 2 execute under thresholds with dual control and kill switch.

vs Copilot: Copilot assists inside one task; agents chain tools. Govern agents like automation programs.

Operational metrics

Business measures — cycle time, error rate, rework, satisfaction — not only model accuracy scores from a vendor lab.

Why this works: Model vanity metrics do not pay salaries or reduce risk. Sponsors need outcome language.

Why it matters: Put the business metric on the pilot charter before day one.

Example: Days to contract signature down 20% with error rate flat — better than “92% model accuracy” alone.

Deep dive lesson

~14 min read

Deep dive: From pilot to production (including agents)

Learning objective. Score a pilot’s production readiness and explain why agents need automation-grade controls, not chatbot manners.

Context. Most AI value dies after the demo. This lesson is about operating services — owners, monitoring, and why “agent” is not just a cooler chatbot.

1. Production is an operating service

Production needs owners, monitoring, support, integration, and business metrics — not vanity model scores. Fund change management and verification capacity, not only licenses. Retest after model and data changes.

Agents act multi-step across systems. They need permissions, logs, and kill switches. Treat them closer to automation than to a writing assistant. Copilot assists you; agent acts — different risk and different funding.

2. Production readiness score

Score 1–5: named business owner, business metric, data access control, monitoring/alerts, support/on-call, retest plan after model change. If any critical score is 1–2, you have a demo, not a launch. Say that out loud in steering committees.

Method: Production gate (must all be true)

1

Owner

Business owner with budget and authority.

2

Metric

Business KPI, not seats.

3

Controls

Access, logs, kill switch for any automation.

4

Support

Who answers when it fails at 2 a.m.?

5

Retest

Calendar owner after model/data change.

Worked example: Agent that “handles refunds”

Situation

Ops wants an agent to process refunds end-to-end in the order system.

How an executive thinks it through

  1. Money movement = high stakes.
  2. Need limits, dual control above thresholds, full logs.
  3. Rollback and on-call required.
  4. Eval on edge cases (fraud, partial ship, policy exceptions).

Decision / what to say

Phase 1: recommend refund only. Phase 2: execute under hard limits after eval and security review. Named on-call and kill switch before any execution rights.

Apply in your function

Operations

No agent execution rights without kill switch and on-call.

Finance

Dual control thresholds for AI-initiated money movement.

IT / Risk partners

Permissions for agents no wider than human roles.

Common mistakes

  • Calling a demo a launch.
  • No business owner — only a project manager.
  • Agent permissions wider than any human role.

Practice (12 minutes)

  1. Pick one AI pilot you know.
  2. Score 1–5: owner, metric, data access, monitoring, support.
  3. Write the top blocker to production and who must remove it.

Check your understanding

Key takeaways

  • Business process change is the hard part.
  • Metrics must be operational, not only model accuracy.
  • Kill criteria protect resources and credibility.
  • Agents need stronger controls than chat drafts.

Practical applications

Operations

Map the handoffs where AI outputs enter human workflows.

Finance

Fund integration and training, not only licenses.

HR

Plan role impacts and reskilling alongside tools.

All

Write kill criteria before the pilot starts.

Marketing

Block auto-publish agents until dual control on claims; production gate before scale.

Library materials for this week

Extend this week’s decisions with briefs, cases, and primary sources that matter for your next meeting.

1~14 min read

Pilot to Production Playbook

Stages from problem framing through pilot, evaluation, integration, training, and scale — with gates and anti-patterns from common failures.

2~12 min read

DBS Bank — AI value at enterprise scale

Published 2025 results coverage (Business Times)

Singapore’s DBS Bank has publicly reported large economic value from AI programs — including Business Times coverage that the bank unlocked about S$1 billion in AI-related economic value in 2025 — after years of digital transformation and governed AI use cases across the franchise.

Optional focus hour

After the core lesson (~35 min), spend about 55 minutes on one long-form source matched to this week’s level.

Optional focus hourResearch~55 min readWithin ~18 months

BCG AI Radar 2026 — As AI Investments Surge, CEOs Take the Lead

Boston Consulting Group

Published 15 January 2026

Why this week

Pilot-to-production week: capital, CEO ownership, and the impact gap — specific Radar 2026 piece.

What to take away

Why pilots stall; one ownership or funding change for your program.

Check your understanding

Week 10 · 3 short questions · no grades shared outside this device

1.AI pilots most often fail to become production because:
2.Compared with a drafting copilot, an agent that can send emails or change records requires:
3.Which metric set best tells you a pilot is ready to scale?

Select an answer for each question.

Hands-on

~10 min

Write kill criteria

Stopping failed work is a leadership skill.

  1. For a real or imagined pilot, write three conditions that would stop expansion.
  2. Include at least one quality/safety condition and one adoption condition.
  3. Name who has authority to stop it.

Reflection

  • Which current initiative would you struggle to stop even if metrics were weak?
  • Who experiences the pain if the AI is wrong — and were they in the pilot design?