Skip to content

AI training data: what it is and where it comes from

Every AI model is shaped by the data it learns from. This guide covers where that data comes from, why private data now carries a price, what each type is used for and what it licenses for, with sources.

AI training data is the text, code, images, audio, video and logs a model learns from. Early models trained mostly on the public web; now labs also license private data, pay experts and generate synthetic data. Private data is priced per unit: unused video licenses at $1 to $4 a minute (Medium confidence), while an operating company's archive is indexed at $185,000 to $1M+ (Medium confidence, estimated). Indicative estimate, not an offer.

Updated

On this page

What AI training data is

AI training data is the text, code, images, audio, video and logs a model learns from, together with the tests and tasks used to evaluate and improve it.

A language model learns from text and code, a speech model from recorded voices with transcripts, a robot from demonstrations and sensor logs. The broad knowledge comes from very large collections; specific skills come from smaller, carefully chosen sets.

  • Text: books, articles, forum threads, documents and expert answers.
  • Code: repositories with their git history, issues and reviews.
  • Images and video: photos, footage, drone and dashcam video, first-person task video.
  • Audio: speech with transcripts, call recordings, podcasts and music.
  • Records and logs: email, chat, tickets, de-identified health records, robot and sensor logs.

Where training data comes from

AI training data comes from public web crawls, licensed content, paid experts and contractors, users who opted in, and synthetic data generated by models.

Where training data comes from (table)
SourceWhat it isAn example
Public webPages, public code and open datasets gathered by crawlersThe base of most early models
Licensed contentArchives licensed from publishers, platforms and companiesOne community platform licenses its data for About $60M a year (Columbia Journalism Review, (opens in a new tab)); one publisher licensed nonfiction titles at $5,000 for 3 years each (eWeek, (opens in a new tab))
Experts and contractorsPeople paid to write, label, rank and build tasksLabs pay $200 to $2,000 per training task built from real software work (Epoch AI, (opens in a new tab))
Users who opted inConversations and feedback from users whose terms allow trainingA product's own users, where its terms say their data may be used
Synthetic dataText, code or images generated by models or simulationsCalifornia's AB 2013 makes developers disclose whether they used synthetic data generation (Frankfurt Kurnit Klein & Selz, (opens in a new tab))

Transparency laws such as California's AB 2013 now make developers of generative AI systems document their training data, including whether datasets were purchased or licensed (Frankfurt Kurnit Klein & Selz, (opens in a new tab)).

Why private data matters now

Private data matters now because models that act at work need to learn from work that never appears in public, and the rights behind each dataset now have to be documented.

The families and their units

Private data falls into 7 families, and each family is priced by its own unit.

The families and their units (table)
FamilyWhat it includesUnitIndicative priceConfidence
Code and repositoriesPrivate repos with their git history, issues and reviews.per repo$150 to $1,500 (estimated)Medium confidence
Company archivesChat, email, issue trackers, docs and support history from real teams.per company$185,000 to $1M+ (estimated)Medium confidence
Health and life sciencesDe-identified imaging, notes, pathology, dental and vet records.per recordRevenue shareLow confidence
Voice and audioCalls, speech in many languages, podcasts, music and sound.per hour$50 to $100Medium confidence
Video and imageryUnused footage, first-person video, drone, dashcam and photo libraries.per minute$1 to $4Medium confidence
Books, documents and expert workBacklists, news archives, translation memories, courseware and forums.per title$5,000 for 3 yearsHigh confidence
Robotics and sensorsTeleoperation logs and industrial sensor data.per hour$50 to $200 (estimated)Low confidence

Indicative estimate, not an offer. Each range is a market benchmark from the Rightmark Price Index, and a row marked estimated is a Rightmark estimate, not a reported price.

Company archives have a hub of their own: Operational data covers email, chat, files, tickets, CRM and meeting recordings, with a live estimate for your company.

Who licenses it, and for what

AI developers use private data for pretraining, fine-tuning, evaluation, reinforcement learning environments for agents, and robotics.

Who licenses it, and for what (table)
UseWhat it meansData that fits
PretrainingBroad learning from very large collections, the first and biggest phase of trainingBook and news archives, forums, large photo and video libraries, code at scale
Fine-tuningFocused training that teaches a task, a domain or a styleExpert writing and answers, domain records, conversations, translation memories
EvaluationHeld-back tests that measure a model, worthless once they leak into trainingNever-public code, expert questions with answers, real tickets
RL environments and agentsSimulated workplaces where agents practice tasks and are scored automaticallyCompany archives with email, chat, tickets and code linked across years; repositories with tests
Robotics and physical AIDemonstrations that teach robots to act in the worldTeleoperation logs, egocentric video, sensor data with labeled faults

Demand comes from AI developers of several kinds: frontier labs, companies building AI products, medical and robotics teams, and the data companies that prepare training sets for them. Repositories that build and pass their tests are usually worth more than the code alone, because they become training tasks.

What AI developers pay

Private data is priced by the unit: a stock image licenses at $0.02 to $2 (Medium confidence), while an operating company's archive is indexed at $185,000 to $1M+ (Medium confidence, estimated).

What AI developers pay (table)
DataUnitRangeConfidenceSource
Private repo, standing-rate license: Typical, per sale. Large repos up to about $10,000 per sale (indicative); history often carries more than half the priceper repo$150 to $1,500 (estimated)Medium confidenceRightmark,
Training task built from real codeper task$200 to $2,000Medium confidenceEpoch AI, (opens in a new tab)
Shut-down startup, code and workspaceper company$10,000 to $100,000Medium confidenceGizmodo, (opens in a new tab)
Operating company, 30+ staff, English recordsper company$185,000 to $1M+ (estimated)Medium confidenceRightmark,
Large enterprise archive: winning bid for one airline's archive, still needs court approvalper company$10MMedium confidenceTIME, (opens in a new tab)
Bulk call-center audioper hour$6.30 to $7Medium confidenceDatarade via Internet Archive, (opens in a new tab)
Ready-made speech datasetsper hour$50 to $100Medium confidenceSpeechData.ai, (opens in a new tab)
Low-resource language speechper hour$60 to $150Medium confidenceSpeechData.ai, (opens in a new tab)
Unused creator footageper minute$1 to $4Medium confidenceThe Decoder, (opens in a new tab)
First-person task videoper hour$15 to $60Medium confidenceStellaris VC, (opens in a new tab)
Stock imagesper image$0.02 to $2Medium confidenceReuters via Rappler, (opens in a new tab)
Pathology slides (acquisition-implied)per slideAbout $11.60 (estimated)Low confidenceBusinessWire, (opens in a new tab)
Nonfiction backlist licenseper title$5,000 for 3 yearsHigh confidenceeWeek, (opens in a new tab)
News licensing, large publishersper yearAbout $24M averageMedium confidenceMedia and the Machine, (opens in a new tab)
Community Q&A (forum archive)per yearAbout $60MHigh confidenceColumbia Journalism Review, (opens in a new tab)
Robot teleoperation logsper hour$50 to $200 (estimated)Low confidenceDreamVu, (opens in a new tab)

Most code repositories license for hundreds to low thousands of dollars per sale (Rightmark, ). These are indicative ranges from public deals, disclosures, published programs and Rightmark estimates, collected for the Q3 2026 edition of the Rightmark Price Index, published on 30 Sep 2026. The next edition is due in Jan 2027. A range is not a quote: how data is priced explains where a single dataset lands.

What makes data trainable

AI developers pay for data they can trust and use: provable rights, never-public content written by people, personal data removed, and the structure intact.

  • Provenance: where every file came from and what rights attach to it, backed by an ownership pack.
  • Never public, written by people: public or model-written material adds little and can spoil evaluations.
  • De-identified: names, contact details and other personal data removed to the standard the law requires.
  • Structure intact: timestamps, IDs, threads, git history, and the links between tickets, code and chat.
  • Labels and metadata: transcripts, captions, diagnoses, task labels and a data card.

Synthetic data and its limits

Synthetic data helps only in part: it scales common patterns cheaply but inherits the gaps of the model that made it, so scarce real-world data keeps its value longest.

Rare languages, robot logs, real workflows and anything that happens inside companies are hard to generate convincingly, because the generating model has seen little of them. Developers mix both, and California's AB 2013 asks them to disclose whether their systems used synthetic data generation (Frankfurt Kurnit Klein & Selz, (opens in a new tab)).

Common data is already losing value as supply grows: one marketplace listed translation memories at EUR 1,500 to 2,500, then cut 97%+ (Low confidence) (MultiLingual, (opens in a new tab)). Scarce, rights-clean data holds its price longest.

How an owner supplies it

An owner supplies training data by licensing it: appraise, export with the history intact, verify and de-identify, choose a route, then deliver and get paid.

  1. Appraise. Find your unit and range in about a minute with the free appraisal, or run Repo Check on a repository without uploading code.
  2. Export with the history intact. Keep native formats, timestamps and IDs. The export guides cover the chat, email, file, ticket and code tools most teams use.
  3. Verify. Scans run on your side, such as Repo Check on your own computer, or inside the licensee's pipeline, to find secrets, personal data and third-party rights. Under an NDA, licensees see the reports and representative scrubbed samples, and an ownership pack records title.
  4. De-identify. Personal data is removed, and masked before anything reaches an AI lab.
  5. Choose a route. Per sale, lump sum or revenue share: see exclusive or non-exclusive. Code is licensed non-exclusively by default, through a licensing partner chosen per deal; company-archive lump sums are usually exclusive for a term, and non-exclusive revenue share is the alternative.
  6. Deliver and get paid. We never hold your passwords. Transfers use scoped, revocable connections your admins approve, or exports you upload, and Rightmark never keeps a copy. The licensee pays you under your license. Timings are indicative and confirmed at scoping.

Questions

Where does AI training data come from?

The main sources are public web crawls, licensed content from publishers, platforms and private companies, work by paid contractors and experts, user data where terms allow, and synthetic data made by other models. Licensed private data is in demand because it is scarce and close to real work.

Why do AI companies pay for private data?

Models still struggle with real work, and private archives show how work actually happens: code with its history, tickets linked to fixes, decisions made in chat. That is why a $10M bid won a bankruptcy auction for one airline's archive, though the sale still needs court approval.

Source: TIME, (opens in a new tab)

What is an RL environment?

A simulated workplace built from real material, such as a codebase with its tickets and chat, where an AI agent practices tasks and is scored automatically. Labs pay $200 to $2,000 per task built from real software work.

Source: Epoch AI, (opens in a new tab)

Can synthetic data replace real training data?

Only in part. Synthetic data scales common patterns cheaply but inherits the gaps of the model that made it. Scarce real-world data, such as rare languages, robot logs and genuine workflows, keeps its value longest.

How much does AI training data cost?

It depends on the unit. Unused video licenses for $1 to $4 a minute (Medium confidence), ready-made speech for $50 to $100 an hour (Medium) and low-resource languages for $60 to $150 (Medium), and one published book license paid $5,000 for 3 years per title (High). The Rightmark Price Index lists every range with its confidence level and source.

Sources: The Decoder, (opens in a new tab); SpeechData.ai, (opens in a new tab); eWeek, (opens in a new tab)

Sources

  1. TIME, (opens in a new tab)
  2. Epoch AI, (opens in a new tab)
  3. Gizmodo, (opens in a new tab)
  4. The Decoder, (opens in a new tab)
  5. Columbia Journalism Review, (opens in a new tab)
  6. eWeek, (opens in a new tab)
  7. Frankfurt Kurnit Klein & Selz, (opens in a new tab)
  8. Rightmark,
  9. Epoch AI, (opens in a new tab)
  10. Gizmodo, (opens in a new tab)
  11. Rightmark,
  12. TIME, (opens in a new tab)
  13. Datarade via Internet Archive, (opens in a new tab)
  14. SpeechData.ai, (opens in a new tab)
  15. SpeechData.ai, (opens in a new tab)
  16. The Decoder, (opens in a new tab)
  17. Stellaris VC, (opens in a new tab)
  18. Reuters via Rappler, (opens in a new tab)
  19. BusinessWire, (opens in a new tab)
  20. eWeek, (opens in a new tab)
  21. Media and the Machine, (opens in a new tab)
  22. Columbia Journalism Review, (opens in a new tab)
  23. DreamVu, (opens in a new tab)
  24. MultiLingual, (opens in a new tab)
  25. SpeechData.ai, (opens in a new tab)

Find out what your data is worth to AI labs.

About a minute, no files, no obligation.

Appraise your data