What AI training data is
AI training data is the text, code, images, audio, video and logs a model learns from, together with the tests and tasks used to evaluate and improve it.
A language model learns from text and code, a speech model from recorded voices with transcripts, a robot from demonstrations and sensor logs. The broad knowledge comes from very large collections; specific skills come from smaller, carefully chosen sets.
- Text: books, articles, forum threads, documents and expert answers.
- Code: repositories with their git history, issues and reviews.
- Images and video: photos, footage, drone and dashcam video, first-person task video.
- Audio: speech with transcripts, call recordings, podcasts and music.
- Records and logs: email, chat, tickets, de-identified health records, robot and sensor logs.
Where training data comes from
AI training data comes from public web crawls, licensed content, paid experts and contractors, users who opted in, and synthetic data generated by models.
| Source | What it is | An example |
|---|---|---|
| Public web | Pages, public code and open datasets gathered by crawlers | The base of most early models |
| Licensed content | Archives licensed from publishers, platforms and companies | One community platform licenses its data for About $60M a year (Columbia Journalism Review, (opens in a new tab)); one publisher licensed nonfiction titles at $5,000 for 3 years each (eWeek, (opens in a new tab)) |
| Experts and contractors | People paid to write, label, rank and build tasks | Labs pay $200 to $2,000 per training task built from real software work (Epoch AI, (opens in a new tab)) |
| Users who opted in | Conversations and feedback from users whose terms allow training | A product's own users, where its terms say their data may be used |
| Synthetic data | Text, code or images generated by models or simulations | California's AB 2013 makes developers disclose whether they used synthetic data generation (Frankfurt Kurnit Klein & Selz, (opens in a new tab)) |
Transparency laws such as California's AB 2013 now make developers of generative AI systems document their training data, including whether datasets were purchased or licensed (Frankfurt Kurnit Klein & Selz, (opens in a new tab)).
Why private data matters now
Private data matters now because models that act at work need to learn from work that never appears in public, and the rights behind each dataset now have to be documented.
- Agents need real work. Agents learn how work happens from code with its history, tickets linked to fixes and decisions argued in chat. A $10M bid won a bankruptcy auction for one airline's archive of email, chat and code, sought for RL environments; the sale still needs court approval (TIME, (opens in a new tab)).
- Labs pay for tasks. Labs pay $200 to $2,000 for each training task built from real software work (Epoch AI, (opens in a new tab)).
- Rights now matter. Transparency laws such as California's AB 2013 make developers disclose their sources and whether data was purchased or licensed (Frankfurt Kurnit Klein & Selz, (opens in a new tab)), which favors data with clean, documented rights.
- The supply is moving. Shut-down startups have licensed their code and workspace data for $10,000 to $100,000 (Medium confidence) in reported deals (Gizmodo, (opens in a new tab)).
The families and their units
Private data falls into 7 families, and each family is priced by its own unit.
| Family | What it includes | Unit | Indicative price | Confidence |
|---|---|---|---|---|
| Code and repositories | Private repos with their git history, issues and reviews. | per repo | $150 to $1,500 (estimated) | Medium confidence |
| Company archives | Chat, email, issue trackers, docs and support history from real teams. | per company | $185,000 to $1M+ (estimated) | Medium confidence |
| Health and life sciences | De-identified imaging, notes, pathology, dental and vet records. | per record | Revenue share | Low confidence |
| Voice and audio | Calls, speech in many languages, podcasts, music and sound. | per hour | $50 to $100 | Medium confidence |
| Video and imagery | Unused footage, first-person video, drone, dashcam and photo libraries. | per minute | $1 to $4 | Medium confidence |
| Books, documents and expert work | Backlists, news archives, translation memories, courseware and forums. | per title | $5,000 for 3 years | High confidence |
| Robotics and sensors | Teleoperation logs and industrial sensor data. | per hour | $50 to $200 (estimated) | Low confidence |
Indicative estimate, not an offer. Each range is a market benchmark from the Rightmark Price Index, and a row marked estimated is a Rightmark estimate, not a reported price.
Company archives have a hub of their own: Operational data covers email, chat, files, tickets, CRM and meeting recordings, with a live estimate for your company.
Who licenses it, and for what
AI developers use private data for pretraining, fine-tuning, evaluation, reinforcement learning environments for agents, and robotics.
| Use | What it means | Data that fits |
|---|---|---|
| Pretraining | Broad learning from very large collections, the first and biggest phase of training | Book and news archives, forums, large photo and video libraries, code at scale |
| Fine-tuning | Focused training that teaches a task, a domain or a style | Expert writing and answers, domain records, conversations, translation memories |
| Evaluation | Held-back tests that measure a model, worthless once they leak into training | Never-public code, expert questions with answers, real tickets |
| RL environments and agents | Simulated workplaces where agents practice tasks and are scored automatically | Company archives with email, chat, tickets and code linked across years; repositories with tests |
| Robotics and physical AI | Demonstrations that teach robots to act in the world | Teleoperation logs, egocentric video, sensor data with labeled faults |
Demand comes from AI developers of several kinds: frontier labs, companies building AI products, medical and robotics teams, and the data companies that prepare training sets for them. Repositories that build and pass their tests are usually worth more than the code alone, because they become training tasks.
What AI developers pay
Private data is priced by the unit: a stock image licenses at $0.02 to $2 (Medium confidence), while an operating company's archive is indexed at $185,000 to $1M+ (Medium confidence, estimated).
Most code repositories license for hundreds to low thousands of dollars per sale (Rightmark, ). These are indicative ranges from public deals, disclosures, published programs and Rightmark estimates, collected for the Q3 2026 edition of the Rightmark Price Index, published on 30 Sep 2026. The next edition is due in Jan 2027. A range is not a quote: how data is priced explains where a single dataset lands.
What makes data trainable
AI developers pay for data they can trust and use: provable rights, never-public content written by people, personal data removed, and the structure intact.
- Provenance: where every file came from and what rights attach to it, backed by an ownership pack.
- Never public, written by people: public or model-written material adds little and can spoil evaluations.
- De-identified: names, contact details and other personal data removed to the standard the law requires.
- Structure intact: timestamps, IDs, threads, git history, and the links between tickets, code and chat.
- Labels and metadata: transcripts, captions, diagnoses, task labels and a data card.
Synthetic data and its limits
Synthetic data helps only in part: it scales common patterns cheaply but inherits the gaps of the model that made it, so scarce real-world data keeps its value longest.
Rare languages, robot logs, real workflows and anything that happens inside companies are hard to generate convincingly, because the generating model has seen little of them. Developers mix both, and California's AB 2013 asks them to disclose whether their systems used synthetic data generation (Frankfurt Kurnit Klein & Selz, (opens in a new tab)).
Common data is already losing value as supply grows: one marketplace listed translation memories at EUR 1,500 to 2,500, then cut 97%+ (Low confidence) (MultiLingual, (opens in a new tab)). Scarce, rights-clean data holds its price longest.
How an owner supplies it
An owner supplies training data by licensing it: appraise, export with the history intact, verify and de-identify, choose a route, then deliver and get paid.
- Appraise. Find your unit and range in about a minute with the free appraisal, or run Repo Check on a repository without uploading code.
- Export with the history intact. Keep native formats, timestamps and IDs. The export guides cover the chat, email, file, ticket and code tools most teams use.
- Verify. Scans run on your side, such as Repo Check on your own computer, or inside the licensee's pipeline, to find secrets, personal data and third-party rights. Under an NDA, licensees see the reports and representative scrubbed samples, and an ownership pack records title.
- De-identify. Personal data is removed, and masked before anything reaches an AI lab.
- Choose a route. Per sale, lump sum or revenue share: see exclusive or non-exclusive. Code is licensed non-exclusively by default, through a licensing partner chosen per deal; company-archive lump sums are usually exclusive for a term, and non-exclusive revenue share is the alternative.
- Deliver and get paid. We never hold your passwords. Transfers use scoped, revocable connections your admins approve, or exports you upload, and Rightmark never keeps a copy. The licensee pays you under your license. Timings are indicative and confirmed at scoping.
Questions
Where does AI training data come from?
The main sources are public web crawls, licensed content from publishers, platforms and private companies, work by paid contractors and experts, user data where terms allow, and synthetic data made by other models. Licensed private data is in demand because it is scarce and close to real work.
Why do AI companies pay for private data?
Models still struggle with real work, and private archives show how work actually happens: code with its history, tickets linked to fixes, decisions made in chat. That is why a $10M bid won a bankruptcy auction for one airline's archive, though the sale still needs court approval.
Source: TIME, (opens in a new tab)
What is an RL environment?
A simulated workplace built from real material, such as a codebase with its tickets and chat, where an AI agent practices tasks and is scored automatically. Labs pay $200 to $2,000 per task built from real software work.
Source: Epoch AI, (opens in a new tab)
Can synthetic data replace real training data?
Only in part. Synthetic data scales common patterns cheaply but inherits the gaps of the model that made it. Scarce real-world data, such as rare languages, robot logs and genuine workflows, keeps its value longest.
How much does AI training data cost?
It depends on the unit. Unused video licenses for $1 to $4 a minute (Medium confidence), ready-made speech for $50 to $100 an hour (Medium) and low-resource languages for $60 to $150 (Medium), and one published book license paid $5,000 for 3 years per title (High). The Rightmark Price Index lists every range with its confidence level and source.
Sources: The Decoder, (opens in a new tab); SpeechData.ai, (opens in a new tab); eWeek, (opens in a new tab)
Sources
- TIME, (opens in a new tab)
- Epoch AI, (opens in a new tab)
- Gizmodo, (opens in a new tab)
- The Decoder, (opens in a new tab)
- Columbia Journalism Review, (opens in a new tab)
- eWeek, (opens in a new tab)
- Frankfurt Kurnit Klein & Selz, (opens in a new tab)
- Rightmark,
- Epoch AI, (opens in a new tab)
- Gizmodo, (opens in a new tab)
- Rightmark,
- TIME, (opens in a new tab)
- Datarade via Internet Archive, (opens in a new tab)
- SpeechData.ai, (opens in a new tab)
- SpeechData.ai, (opens in a new tab)
- The Decoder, (opens in a new tab)
- Stellaris VC, (opens in a new tab)
- Reuters via Rappler, (opens in a new tab)
- BusinessWire, (opens in a new tab)
- eWeek, (opens in a new tab)
- Media and the Machine, (opens in a new tab)
- Columbia Journalism Review, (opens in a new tab)
- DreamVu, (opens in a new tab)
- MultiLingual, (opens in a new tab)
- SpeechData.ai, (opens in a new tab)