Copyright & provenance: whose data trained this?
A model can write in a novelist's style or spit out code it clearly learned somewhere. Where did all that training data come from, and who's allowed to use it?
The idea inside
Training data has provenance: some is openly licensed, some needs permission, some is contested. Public is not the same as permitted, and who owns AI output is still unsettled.
After this lesson
You can reason about where training data comes from and the copyright, consent, and ownership questions it raises, without mistaking 'public' for 'permitted'.
Where it leads
Inside this lesson
That's the real lesson stage, paused. Claim your pass to operate it.
See how AI actually works, end to end.
This lesson is one stop on the full arc. Unlock all of it, and keep it for life.
What you get
- The 34-lesson main path, a finishable route from a word to agents
- Goal tracks for using AI at work and building AI features
- Boss labs that make you apply a whole act, not just recognize it
- Spaced recall that brings each idea back before you forget
- Course memory: every term defined, with links to where it first appears
- A shareable capability card when you finish the main path
- Lifetime access on every device, every future lesson included
Not videos to watch. You predict, operate the machine, then prove it. That is why it stays.
99 interactive lessons and challenges. No videos, no code.
Free launch pass: lifetime access, no card needed
New here? The first lessons are free to try. Start with lesson 0.1
What this lesson shows
Training data has provenance: some is openly licensed, some needs permission, some is contested. Public is not the same as permitted, and who owns AI output is still unsettled.
The question it opens with
A model can write in a novelist's style or spit out code it clearly learned somewhere. Where did all that training data come from, and who's allowed to use it?
The walkthrough, in the lesson's own words
- Sort each source: is it free to train on, does it need a license, or is it off-limits or contested?
- Now the other side: what about the words the model gives back?
- Public is not the same as permitted, and who owns the output is unsettled.
- Same data, very different rules. Provenance, where it came from and under what license, decides what's fair to train on. 'It was on the internet' is not a license.
- Illustrative of how provenance is reasoned about, not legal advice.
- A model trained on copyrighted text. Can it ever reproduce a passage close to word-for-word in its output?
- Big models can memorize text they saw often or that was rare and distinctive, and echo it back. So the copyright question lands on the output too, not just the training set.
- On the input side: license data, filter out known-copyrighted and private content, and keep provenance records. On the output side: filters that catch verbatim regurgitation, and content credentials (like C2PA) that tag AI-generated media so it can be traced. None of it fully settles who owns what, which is why vendors increasingly offer copyright indemnification as a business answer to a legal grey area.
- Regulators are stepping in too: in the EU, makers of general-purpose models (GPAI, the models behind chatbots) must publish a summary of their training data and honor rights-holders' opt-outs under the EU AI Act.
- If it's on the public internet it's free to train on, and whatever the AI outputs is automatically yours to use.
- Public isn't permission: provenance and licensing decide what's fair to train on, and ownership of AI output is unsettled and varies by country.
- A vendor advertises that its model is 'trained only on licensed data' and offers you copyright indemnification (the vendor promises to pay if you get sued over training data). What risk are they really addressing?
- Provenance risk. If a model was trained on copyrighted or pirated data, its outputs could expose you to infringement claims. 'Licensed data' and indemnification are how a vendor takes that legal exposure off your plate, which tells you the exposure is real.
- You can reason about where training data comes from and the copyright, consent, and ownership questions it raises, without mistaking 'public' for 'permitted'.
- This lesson explains the landscape. It is not legal advice.
Key takeaway
You sorted real data sources by whether a model can train on them freely, and saw why 'it was on the internet' isn't a license.
What you can do after this lesson
You can reason about where training data comes from and the copyright, consent, and ownership questions it raises, without mistaking 'public' for 'permitted'.
This is the written summary. The lesson itself is interactive: you predict, drag and operate the mechanism above, and the reveal answers you.