Data, Privacy & Confidentiality
What you can share, train on, and keep private
Listen to executive summary
Audio length ~2 min
Full lesson about ~35 min
Executive summary
~8 min to read the written sections · ~2 min to listen to this summary
This week is about data discipline for AI: classifying information (public / internal / confidential / restricted), training rights vs inference-only use, and residency, retention, and deletion questions that belong in every AI deal.
You will leave with a Green / Amber / Red one-pager mindset so staff can choose tools under deadline pressure without creating incidents.
Outcome: clearer ownership between Privacy, Legal, Purchasing, and the business on what may enter which AI systems.
Core concepts
The terms and comparisons you need for this week’s decisions — with examples by function.
Data classification (Green / Amber / Red)
Label information before choosing tools: public (Green), internal (Amber), confidential/restricted (Red). Classification is the fastest control staff can apply under deadline pressure.
Why this works: Most functional AI incidents are good people pasting the wrong document into the wrong tool — not exotic hacks.
Why it matters: Publish a one-pager with approved tools for Green and a ban list for Red. Name the exception contact.
Example: Public press release → many tools OK. Unreleased M&A deck → approved enterprise workspace only.
Finance
Weak / before: Paste board pack into free chatbot to save time.
Strong / after: Board packs are Red. Use approved enterprise summarization only; treat free-tool paste as an incident.
By function
Legal
Classify matter files before any generative tool; training rights in DPAs.
HR
Employee files are high sensitivity — approved tools only.
Purchasing
Data flow diagrams for AI SaaS before signature.
Marketing
Customer lists and unreleased campaigns = Amber/Red for consumer AI.
Finance
Board packs and MNPI never in free chat tools.
Operations
Plant or safety data classification before copilots on the floor.
Sales
Pipeline and pricing strategies stay out of consumer models.
Risk
Green/Amber/Red one-pager + exception contact published org-wide.
Training rights
Whether a vendor may use your prompts or documents to improve their models for themselves or other customers. Different from processing your input only to answer your query (inference).
Why this works: Training rights can move confidential strategy outside your control. This is a commercial and privacy term, not a technicality.
Why it matters: Default should be no training on customer content unless you explicitly agree and understand the benefit.
Example: Pasting strategy into free consumer chat may allow broader use than your enterprise Claude or ChatGPT workspace.
vs Inference-only use: Inference answers now; training rights let the vendor learn from you later.
Data residency & retention
Where data is stored and processed, how long logs and files are kept, and how deletion works when you leave the vendor.
Why this works: Cross-border transfer and long retention create compliance and discovery exposure even when the feature looks harmless.
Why it matters: Ask deletion-on-exit and log access questions in every AI SaaS deal.
Example: Who can export a full chat history of your pricing strategy discussion six months later?
Procurement questions
Weak / before: Is it secure?
Strong / after: Where is data stored? Who accesses logs? Retention period? Delete on exit? Training on inputs? Subprocessors?
Deep dive lesson
~13 min readDeep dive: Data, privacy, and confidentiality
Learning objective. Traffic-light data types for AI tools and ask the training, retention, and residency questions that protect the enterprise.
Context. Most AI incidents in functional teams are not exotic hacks — they are good people pasting the wrong document into the wrong tool under deadline pressure.
1. Classify before you choose the tool
AI is hungry for data — which is exactly why executives must slow down and classify: public, internal, confidential, restricted. Green data may belong in approved enterprise tools. Amber needs restrictions. Red never goes into generative tools without special controls.
Personal data and sensitive categories (health, biometrics, children’s data) raise higher bars. Confidential IP — pricing strategy, unreleased products, M&A — can be catastrophic if it trains someone else’s model. Align with privacy programs you already have; do not invent a parallel universe of shadow tools.
2. Questions that belong in every AI initiative
What is input? Stored? Used to improve vendor models? Who accesses logs? Retention? Deletion? Cross-border transfer? Who can export a full chat history of your board strategy discussion? If the vendor cannot answer , do not paste Red or Amber data while “testing.”
Method: Green / Amber / Red one-pager
Green
Public or already external-ready content — approved tools OK.
Amber
Internal business data — enterprise tools only, no consumer apps.
Red
Personal sensitive, MNPI, M&A, secrets — special controls or ban.
Publish
Post the one-pager where work happens; name the exception contact.
Worked example: Finance analyst pastes board pack into a free chatbot
Situation
Under deadline, an analyst pastes slides with unreleased numbers into a consumer AI tool for “a better summary.”
How an executive thinks it through
- Data class: likely confidential / MNPI territory.
- Training and retention unknown on free tier.
- Even “we don’t train” claims need enterprise agreement, not hope.
Decision / what to say
Treat as incident per policy. Retrain team with traffic-light one-pager. Move legitimate summarization to approved enterprise tool with logging.
Apply in your function
Finance
Mark board materials and forecasts Red/Amber; ban consumer tools for them.
Legal
Review training-on-inputs and retention in AI DPAs.
HR
Treat employee files as high sensitivity; define approved tools only.
Purchasing
Require data flow diagrams for AI SaaS before signature.
Common mistakes
- Assuming enterprise SSO alone means data is safe for any content.
- No joint review by privacy + procurement on AI SaaS.
- Policies nobody can apply under deadline pressure.
Practice (12 minutes)
- List five data types your team handles weekly.
- Mark Green / Amber / Red for generative AI tools.
- Publish as a one-pager with the approved tool name for Green.
Check your understanding
Key takeaways
- Classify data before choosing a tool.
- Training rights and logging are contract and policy issues, not footnotes.
- Employee and customer personal data need extra care.
- Deletion, retention, and residency are executive-level questions.
Practical applications
Standardize AI data-processing addenda with procurement.
Separate HRIS-connected tools from general productivity chat.
Block pasting of material non-public information into unapproved tools.
Score vendors on data controls as heavily as on feature demos.
Customer lists and unreleased campaigns are Amber/Red for consumer AI tools — use approved workspaces only.
Library materials for this week
Extend this week’s decisions with briefs, cases, and primary sources that matter for your next meeting.
Data Classes for AI Tools
Traffic-light guidance for what data may enter which classes of AI tools — green, amber, red — with examples executives can socialize.
Samsung restricts ChatGPT after confidential code leaks
Published 2 May 2023
In 2023, Bloomberg and other outlets reported that Samsung Electronics restricted employee use of generative AI tools including ChatGPT after engineers allegedly pasted sensitive source code into the service — a textbook data-classification and shadow-AI incident.
Optional focus hour
After the core lesson (~35 min), spend about 55 minutes on one long-form source matched to this week’s level.
OECD Recommendation on AI — OECD/LEGAL/0449 (official instrument)
Organisation for Economic Co-operation and Development
Published Amended 3 May 2024
Why this week
Data and training-rights week needs an international standard, not a blog paraphrase.
What to take away
Principles that hit your data classes; one gap in current tool approvals.
Check your understanding
Week 6 · 3 short questions · no grades shared outside this device
Select an answer for each question.
Hands-on
~12 minTraffic-light your top five data types
Simple rules beat 40-page policies nobody reads under deadline pressure.
- List five data types your team handles weekly (e.g., contracts, salaries, supplier prices).
- Mark each Green (OK in approved enterprise AI), Amber (restricted), Red (never in generative tools without special controls).
- Publish the list to your team as a one-pager.
Reflection
- Does your team know which AI tools are approved for confidential work?
- When was the last time procurement and privacy co-reviewed an AI SaaS deal?