Week 2Month 1~35 min lesson~2 min listen~8 min read

How AI Works — Without the Math

Training, data, models, and outputs in everyday language

Listen to executive summary

Audio length ~2 min

Full lesson about ~35 min

Executive summary

~8 min to read the written sections · ~2 min to listen to this summary

This week explains how AI systems learn and fail, without equations: training vs inference, why generative models can hallucinate, and why vendor “accuracy” claims are incomplete until tested on your work.

You will insist on evaluations (evals) that use your tasks and pass/fail criteria, and treat training-on-inputs clauses as material commercial terms.

Outcome: sharper diligence questions for any model, pilot, or accuracy slide that reaches an executive meeting.

Core concepts

The terms and comparisons you need for this week’s decisions — with examples by function.

Training vs inference

Training is how a model learns patterns from large datasets (usually done by labs or vendors). Inference is the everyday run: your prompt or document goes in, an output comes out. Fine-tuning is extra training on specialized examples.

Why this works: Contract language about “using your data to improve models” is usually about training rights — a strategic and privacy issue, not a technical footnote.

Why it matters: Default: inference on your work is often necessary; training on your content should be explicit and rare.

Example: Labs train for months on huge datasets. When you ask ChatGPT a question at work, that moment is mostly inference.

Contract question

Weak / before: Do you use our data?

Strong / after: Do you train models on our prompts and documents? Can we opt out? How long are logs retained? Who can access them?

vs Fine-tuning: Fine-tuning is extra training on specialized examples. Inference is day-to-day use.

By function

Legal

Training-on-inputs clauses in AI SaaS DPAs — material, not boilerplate.

HR

Screening accuracy claims: accuracy of what, for whom — demand group-level error discussion where required.

Purchasing

Reject “90% accurate” without definition, dataset, and retest after model upgrades.

Marketing

Eval 30 real briefs for claim safety before multi-year creative AI seats.

Finance

Hallucinated ratios in board packs: verification against ERP, not fluency.

Operations

Model upgrades can regress SOP answers — schedule retests, not “set and forget.”

Sales

Invented discount language in AI drafts — system-of-record pricing only.

Risk

Evals and model-change notice belong on the control map next to residual risk owners.

Hallucination

A hallucination is a confident false or fabricated answer — invented citations, numbers, case names, or quotes. The model is optimizing for plausible language, not certified truth.

Why this works: Fluency tricks busy people. High-stakes work needs verification by design, not optional “please double-check” culture.

Why it matters: Wherever rights, money, safety, brand, or compliance are at stake, build checks into the workflow.

Example: Grok or Claude invents a court case citation that looks real. Legal must verify before any advice depends on it.

Process fix

Weak / before: We told staff to verify AI answers.

Strong / after: External claims require a second-person check and a source link from our system of record before send.

vs Bias: Hallucination is fabrication; bias is systematic unfair skew. Different tests and owners.

Evaluation (eval)

An eval is a structured test on tasks that represent your work, with pass/fail criteria you define. A demo is a curated happy path on the vendor’s samples.

Why this works: Without evals, you buy storytelling. With evals, you buy measured fitness for your workflows — and you know when a model upgrade breaks them.

Why it matters: Your leverage is judgment and governance: refuse “trust us” as a quality system and name who retests after upgrades.

Example: Thirty of your real NDAs, scored for missed liability clauses and invented parties — before multi-year seats.

Five-line scorecard

Weak / before: Accuracy 90%.

Strong / after: Task: 30 real briefs. Pass/fail: no invented claims; correct product names; tone match. Threshold: any invented legal claim fails the pilot.

Deep dive lesson

~15 min read

Deep dive: How AI learns, fails, and gets measured

Learning objective. Separate training from inference, explain hallucination with precision, and insist on evals that use your work — not vendor slides.

Context. This is the week that strengthens your hand in vendor meetings. You will learn why fluent answers can still be false, and what “90% accurate” almost never means.

1. Training vs inference — why contracts care

Training is how a model learns patterns from large datasets — expensive, occasional, and often done by labs or vendors long before you buy a seat. Inference is the everyday run: your prompt, your document, your customer ticket goes in; an output comes out. When you ask Claude, ChatGPT, Gemini, or Grok a question at work, that moment is mostly inference.

When a contract says the vendor “may use inputs to improve models,” that language is usually about training rights. It can move confidential information outside your control and into future model behavior. That is not a footnote for Legal and Purchasing; it is a material term. Ask directly: Do you train on our prompts and documents? If so, can we opt out? How is deletion handled?

Fine-tuning is a middle concept: extra training on specialized examples so the model behaves better on a narrow task (for example, always writing in your brand style). Inference is still what staff do all day. Keep the words straight so meetings stop floating between “the model learned from us” and “the model answered us.”

2. Hallucination is structural — plan for it

Generative models predict plausible next language. Plausible is not the same as true. A hallucination is a confident false or fabricated answer — invented citations, statistics, customer quotes, or case names. Fluency is not evidence. The model is not “lying” with intent; it is completing a pattern. Your process still has to catch the falsehood.

Wherever stakes include rights, money, safety, brand, or compliance, design verification: source checks, second-person review, system-of-record confirmation. Do not rely on “we told people to double-check” alone. Busy people skip optional steps. Put checks where work already flows (approval queue, CMS publish button, contract playbook).

Analytics systems fail differently — wrong scores, drift when the world changes, bias against groups. Do not use one “accuracy %” slide for both worlds. Ask: accuracy of what, on whose data, measured how, and retested after model updates?

3. Evals beat demos

An evaluation (eval) is a structured test on tasks that represent your work, with pass/fail criteria you define. A demo is a curated happy path. High-performing teams run sample-based evals before multi-year seats and retest after model upgrades.

Your five checks are executive-grade: correct party names; no invented numbers; tone match; policy alignment; escalation when the tool is unsure. Your role is to refuse “trust us” as a quality system and require a named owner for retesting when the vendor says “we upgraded the model.”

In the meeting

A short exchange you can reuse when the conversation gets vague.

Sales engineer

We’re 90% accurate on compliance-safe copy.

You

Accurate on spelling, brand tone, legal claims, or regulator-safe statements — and on whose samples?

Sales engineer

Industry benchmarks. Everyone uses them.

You

We’ll pilot on thirty of our briefs with our scorecard. No multi-year seats until that eval passes.

Method: Five-line eval scorecard (copy this)

1

Task

Define 20–30 real examples from your work (not the vendor’s demo set).

2

Pass / fail

Write binary checks a human can apply in under two minutes each.

3

Owners

Who runs the eval, who accepts residual risk, who retests after upgrades.

4

Threshold

What failure rate stops the pilot (e.g. any invented legal claim = fail).

5

Cadence

Retest when model, prompt, or source data changes — put a date on the calendar.

Worked example: “90% accuracy” marketing claim

Situation

A content tool claims 90% accuracy for “compliance-safe product copy.” Marketing wants seats this quarter for a launch.

How an executive thinks it through

  1. What was measured — spelling, brand tone, legal claims, or regulator-safe statements?
  2. Whose data — their samples or your product SKUs and claims library?
  3. What happens on the 10% — silent wrong claims or flagged for review?
  4. Training rights on unpublished campaign briefs?

Decision / what to say

Hold purchase until eval on 30 of our real briefs with a Legal/Marketing scorecard. No seats without human approval for claims and pricing.

Apply in your function

Legal

Ban sole reliance on generative tools for citations or legal conclusions without human verification.

HR

For any screening AI, demand documentation of error rates by group where law requires — not a global accuracy %.

Purchasing

Add training-on-inputs, model-change notice, and eval rights to AI contract checklists.

Marketing

Require claims library + human gate for any generative product copy going external.

Common mistakes

  • Accepting vendor accuracy numbers without definition.
  • Ignoring training-on-inputs clauses under deadline pressure.
  • Assuming model upgrades only improve quality (they can also regress).
  • No owner for retesting after changes.

Practice (12 minutes)

  1. Pick one proposed AI use in your team.
  2. Write five pass/fail checks a human would use before sending work out.
  3. Mark which checks could be automated later vs human-only forever.
  4. Send the list to your tech or vendor partner as the start of an eval plan.

Check your understanding

Key takeaways

  • Data quality and relevance shape outcomes more than marketing slogans.
  • Training vs. using (inference) are different stages with different risks.
  • Hallucinations are normal for generative systems — design for review.
  • Always ask how 'accuracy' was measured and on what sample.

Practical applications

Legal

Ask whether legal AI cites sources you can verify, or invents case names.

HR

Ask what employee data trains or feeds the tool — and who can see it.

Purchasing

Require vendors to explain evaluation metrics in business terms.

Finance

Treat forecast models as scenarios, not single-point truth.

Marketing

When a tool claims 90% content accuracy, ask: accuracy of what — brand tone, facts, legal claims, or click predictions?

Library materials for this week

Extend this week’s decisions with briefs, cases, and primary sources that matter for your next meeting.

1~10 min read

Data Is the Fuel

Why data quality, coverage, and permissions determine AI outcomes more than model brand names — with executive questions for any initiative.

2~6 min read

Term: Training vs inference

Training is when a system learns patterns from data. Inference is when it uses those patterns on a new input — the step you see in day-to-day tools.

Optional focus hour

After the core lesson (~35 min), spend about 60 minutes on one long-form source matched to this week’s level.

Optional focus hourReport~1 h readFoundational standard

NIST AI RMF 1.0 — full PDF (NIST.AI.100-1)

U.S. National Institute of Standards and Technology

Published January 2023

Why this week

After training, inference, hallucination, and evals, NIST’s Map–Measure–Manage–Govern language maps risk without model engineering.

What to take away

The four functions in your words; one AI use and which function applies first.

Check your understanding

Week 2 · 3 short questions · no grades shared outside this device

1.A résumé-screening model trained mostly on last decade’s hires is rolled out company-wide. The most likely executive risk is:
2.Legal asks you to review an AI memo that cites three cases. Two check out; one does not exist. This is best understood as:
3.A pilot shows 92% 'accuracy' on a vendor slide. Your best next move is:

Select an answer for each question.

Hands-on

~12 min

Interrogate one accuracy claim

Impressive percentages often hide narrow tests that do not match your reality.

  1. Pick a vendor slide or product page that quotes a high accuracy percentage.
  2. Write three questions: accuracy at what task? on whose data? compared to what baseline?
  3. If you cannot answer from public materials, note that as a procurement gap.

Reflection

  • Which data about your customers or employees would be dangerous if used carelessly in a model?
  • Who in your organization can currently approve 'we trained on your data' language in contracts?