Sell Data to AI
Home Data Asset Score Pricing API documentation US labs and CROs list
For brokers
How to become an AI data broker Data broker business model Buyer programs compared Qualify a company AI training data companies
For data companies
Firmographic data providers Company data API
Seller guides
How to sell data to AI companies Is it legal? FAQ and glossary About
Check domain/company
Guide · How labs buy

How to Sell Data to AI Labs

Last checked: 7 October 2026

AI labs buy data the way any large company buys from a vendor: through procurement, legal review and a test of the goods. Here is what that process looks like, why most companies of 20 to 500 people reach labs through a data company, and when going direct makes sense.

50+companies sell data and RL environments to labs (Deedy Das market map, July 2026)
~$8.5Btheir estimated combined revenue (same source)
~10xvalue of evaluations vs. raw data, as practitioners say
60 to 90 daystime to close that practitioners cite
The short answer

Labs buy through procurement, not through a sign-up form

Definition

Selling data to an AI lab means licensing a copy of your records to a company that trains or evaluates AI models, under a written agreement that sets scope, price, de-identification and acceptance. Mode and micro1, as published, describe a licensed or agreed copy, with originals and ownership staying with the company.

A frontier lab is a large organization with a legal team, a security team and a procurement function. When it buys training or evaluation data, expect it to run the purchase the way it runs any vendor purchase: confidentiality first, then a test of the goods, then a master agreement, then an order. Each step has an owner on the lab's side, and each one adds weeks.

That process is built for suppliers who deliver large volumes in a standard format, with privacy work and quality checks already done. A 40-person accounting firm or a 150-person software company rarely fits that shape on its own. That is the gap data companies fill: they sign many smaller sellers, clean and package the records, and sell the combined result onward. For the full map of buyers, including brokers and marketplaces, see who buys company data for AI.

Direct to a lab

You become the lab's supplier. You negotiate with its legal team, pass its security review and deliver in its format. This route suits large or unique datasets.

Through a data company

A buyer program such as micro1, Mode or Grepped licenses your data, runs export and de-identification under its published process, and deals with labs. This suits most companies of 20 to 500 people.

What both routes share

An NDA, a sample review, a written license, acceptance criteria and payment after acceptance. The paperwork looks alike; what changes is who carries the work.

Procurement

The six documents a lab purchase usually runs on

Names vary by company, but expect some version of each. They tend to arrive in roughly this order. Read each one before you send anything beyond a manifest.

01 Before any detail

Mutual NDA

Covers what you show during evaluation: the manifest, samples, and anything said about your systems and clients.

Check: is it mutual? What happens to samples the lab keeps after saying no? How long do the duties last?

02 The umbrella

Master agreement (MSA)

Sets liability, indemnities, warranties, confidentiality and termination for everything bought under it. Large buyers usually start from their own paper.

Check: is liability capped, for both sides? Which clauses survive termination, and for how long?

03 The deal itself

Data license or SOW

Which records, which years, which format, what the buyer may do with them (training, evaluation or both) and whether the license is exclusive.

Check: scope of use, exclusivity, resale rights, and deletion of the delivered copy.

04 Due diligence

Security and privacy questionnaire

Asks how the data was collected, whether people in it were notified, how it was de-identified and how it will be transferred.

Check: every answer can become a representation you are held to. Answer precisely, not hopefully.

05 Definition of done

Acceptance criteria

The tests a delivery must pass to count as delivered: volume, format, completeness and de-identification checks.

Check: who decides, by what date, what happens when a batch fails, and whether you can audit the de-identification.

06 Getting paid

Purchase order

The buyer's payment instrument, issued against the license. Expect payment to follow acceptance rather than signature.

Check: one-off or recurring, milestone amounts, and payment terms after invoice.

Dataset evaluation

How a lab tests your data before it pays

Most direct conversations end at this stage. A lab is not judging whether your records are interesting. It is judging whether they change a model, or measure something it cannot measure today.

1

Intake

You describe the dataset through the lab's intake route or a contact. Expect basic questions: source systems, years covered, volume, language, and whether personal data is inside.

2

NDA signed

Nothing detailed changes hands before this. If someone asks for files before an NDA exists, slow down.

3

Manifest

A structured description, not the data: systems, record counts, date ranges, file types, how records link (ticket to fix to reply, deal to note to invoice) and what you will exclude. A good manifest answers most first-round questions.

4

Samples

A small, de-identified extract. Reviewers check that records are complete, connected and in the target language, and that the de-identification holds up under a close read.

5

Research review

Researchers decide whether the data fills a gap. Connected histories of real work, with decisions, reviews and approvals in them, tend to do better than loose files. Duplicated or public material does poorly.

6

Terms

MSA plus data license or SOW. Price, scope, exclusivity and acceptance are settled here, before the full export, not after.

7

Export, acceptance, payment

You, or people you hire, export and de-identify to the agreed format. The buyer runs its acceptance checks, and the purchase order typically pays against accepted deliveries.

Practitioners cite 60 to 90 days from first contact to close, and that assumes a seller who is ready. Nothing here is a promise of timing. The seller's step-by-step guide covers the work on your side: inventory, scoping and export.

Never send a full dataset before price. Practitioners give the same advice whatever the route: share a manifest and samples, settle price and terms in writing, and try to get more than one offer. A full export sent "for evaluation" is hard to take back. See getting more than one offer.

The middle layer

Why most companies sell to labs through a data company

It is not that labs refuse small sellers. A lab's process costs a lot to run per supplier, and a data company spreads that cost across many sellers.

Many companies10 to 1,000 staff, each with a few years of records in common tools
Data companyScopes, exports, de-identifies, checks quality and packages records into one format
A few labsOne supplier, one contract, one consistent dataset

Fewer suppliers to manage

Every direct supplier means another NDA, another MSA negotiation and another security review. A lab can buy from one data company that brings many sellers' records under a single contract.

Standard formats

Labs want consistent schemas across sources. Data companies convert exports from Slack, Jira, Salesforce or QuickBooks into one shape, so the lab does not have to do it seller by seller.

De-identification and QA done upstream

The data company takes on scrubbing and quality checks before onward delivery. As published, checked 7 October 2026: Mode says it de-identifies before onward delivery; micro1 says sensitive and confidential information is scrubbed and originals are deleted after processing.

Small datasets combined

One firm's years of support tickets may be too small to move a model on their own. Many firms' tickets in the same format can be. Aggregation is what turns a small seller into a useful supplier.

The layer is large. Deedy Das's July 2026 market map counts 50+ companies selling data and RL environments to labs, at about $8.5B in revenue. Will Depue (July 2026) puts labs on a path to more than $100B a year of data spend by 2030. Both are estimates.

Value also depends on what is built on top of the data. Practitioners describe three tiers: raw data is the cheapest; evaluations built on the data are worth roughly 10x raw; full training environments reach 6 to 8 figures but need heavy engineering. A company selling direct usually has only the first tier to offer.

Two routes

Direct to a lab vs. through a data company

The buyer at the end may be the same. What differs is who you sign with, who does the work, and who carries which risks.

Direct to a lab

Fits: large, unique or hard-to-replace datasets.

  • You negotiate with the lab's legal and procurement teams, usually on the lab's paper.
  • You answer its security and privacy review yourself.
  • You, or contractors you hire, export, de-identify and format the data.
  • No intermediary sits between you and the buyer of record.
  • No referral link applies. This site earns nothing on direct deals.

Through a data company

Fits: most companies of 20 to 500 people with connected work records.

  • The program runs discovery, scoping and the agreement with you.
  • It handles export and de-identification under its published process.
  • It combines your records with others and deals with labs.
  • You sign with the data company, not the lab, so ask which downstream buyers and uses are allowed.
  • Published ranges and eligibility rules exist to check against (table below).

Routes into a lab, as published

Each company's own wording. Ranges are what the company publishes, not offers or averages.

Last checked: 7 October 2026. Sources are each company's own pages, listed as plain text. Lab intake routes publish no price range.
RoutePayout (published)Eligibility (published)Source
Google content intake (direct)Not publishedSuited to large or unique datasetscontentpilot.google.com
OpenAI data partnerships (direct)Not publishedSuited to large or unique datasetsOpenAI data partnerships page
micro1 Enterprise Data Partnership"$100k+ qualified", "$500k+ large-scale", "$1M+ highly unique"; referral page: "$100K-$2M+ for approved data packages"30+ employees, mature operations, documented processes, modern software tools, primarily English; US prioritized, then other Western marketsmicro1.ai/data-partnerships
Mode company data"$100K-$5M"20+ full-time US office employees; accounting firms 10+; law firms 6+; several years of records the company ownsdata.mode.inc
Grepped"$20K-$5M"Any verticalgrepped.ai
Miro AdvisoryOperating datasets "$100K-$1M+"; private codebases "$10K-$1M+" (indicative)Businesses; software companiesmiroadvisory.com

"Up to" and "+" figures mark the top of a published range, not a typical result. For what moves the number, see how much AI companies pay for data.

Decision

Signals that going direct makes sense

Direct intake is a real option, but for a narrow set of sellers. Most companies tick only one or two of these boxes.

Consider a lab's own intake if

  • The dataset is large by any measure: many years, many people, many connected systems.
  • It is unique: no other company could produce it, such as a rare industrial process or a specialist archive.
  • You own it outright, and your client contracts do not restrict it.
  • You have legal and security people who can work through an MSA and a security questionnaire.
  • You can fund export, de-identification and formatting, or hire it.
  • You can wait through a full procurement cycle without needing the money.

A data company is likely the better route if

  • You have 20 to 500 staff and records in common tools: Slack, Jira, Salesforce, QuickBooks, Zendesk.
  • No one in-house can run export and de-identification.
  • You want someone else to carry the format and quality work.
  • You fit a published rule, such as Mode's 20+ US office employees or micro1's 30+ employees.
  • You would rather negotiate one contract with one counterparty than run a lab's procurement yourself.
Reported example of a direct lab purchase

In August 2026, Google agreed to pay $10 million in Spirit Airlines' bankruptcy proceedings for internal data (emails, Teams messages, spreadsheets and operations files), as reported by ABC, TIME and others. micro1 then made a reported $12.5 million rival offer. Court approval of the sale was not confirmed as of 7 October 2026.

The records were large, connected and specific to one airline's operations, which is the profile of a direct deal. It was also a liquidation, not a running company, so read it as context, not as a benchmark for your business.

Not sure where you land? The eligibility checker compares your size, location and systems with each program's published rules, entirely in your browser. To see the programs side by side, use buyer programs compared.

Going direct

Questions to ask a lab's procurement or data team

Ask these early, ideally right after the NDA. They are generic questions for any lab, not claims about how a particular one works, and the answers tell you how long the process will run and who carries which risk.

Who evaluates the data?

A procurement contact, a research team, or both? Who can say yes, and who signs for the lab?

What sample do you need?

Ask for the smallest sample that answers their question, in which format, and agree in writing that it is de-identified before it leaves.

What happens to rejected samples?

Are they deleted, and will deletion be confirmed in writing? May they be used for anything after a no?

Whose paper do we start from?

The buyer's standard MSA or yours? Which clauses are open to negotiation, and is there a data-specific addendum?

Is payment tied to acceptance?

Which tests define acceptance, who runs them, by when, and what happens to payment if one batch fails?

What downstream use is allowed?

Training, evaluation or both? Can the data reach affiliates, contractors or other parties, and on what terms?

Avoid these

Common mistakes when approaching a lab directly

Most of these cost a seller its leverage rather than the deal itself. All of them are avoidable before the first email goes out.

Attaching files to the first email

Anything sent before an NDA sits outside any confidentiality terms. Describe the dataset first; send samples only after signing.

Shipping the full export "for evaluation"

Once a buyer holds the whole dataset, the reason to agree a fair price gets weaker. A manifest and samples are enough to evaluate.

Leading with volume, not connection

A file count says little. Explain how records link: the request, the discussion, the decision, the result. That is what reviewers look for.

Answering consent questions from memory

Questionnaire answers can become representations. Check employee notice and your client contracts before you answer, not after.

Underestimating the export work

De-identifying and formatting to a buyer's schema is real engineering. Budget for it, or use a data company that does it.

Treating one reply as the market

One lab's answer is not a price. Practitioners advise more than one offer, from labs and data companies alike, compared on terms as well as money.

FAQ

Selling data to AI labs: common questions

Can a small company sell data directly to an AI lab?

It can try. Some labs publish intake routes, such as Google's content intake at contentpilot.google.com and OpenAI's data partnerships page. In practice, labs run procurement built for large suppliers, so most companies of 20 to 500 people reach labs through a data company that aggregates, de-identifies and packages records from many sellers.

What documents does an AI lab ask a data seller to sign?

Expect a mutual NDA first, then a master agreement (MSA), a data license or statement of work that sets scope and price, a security and privacy questionnaire, written acceptance criteria and a purchase order. Have your own lawyer read each one before you sign.

How do AI labs evaluate a dataset before buying it?

Usually from a manifest (systems, record counts, date ranges, file types, exclusions) and a small de-identified sample. Researchers then judge whether the data fills a gap their models have. Connected histories of real work tend to rate higher than scattered, duplicated or public files.

What should a dataset manifest for an AI lab include?

Source systems, record counts, date ranges, file types, languages, how records link to each other, what personal or client data is inside, and what you will exclude. A manifest describes the data without containing it, so you can share it under an NDA before any sample leaves your company.

How long does selling data to an AI lab take?

Practitioners cite 60 to 90 days to close, through NDA, review, agreement, export, de-identification and acceptance. It can take longer, and no route can promise a timeline.

Do I earn more by going direct to a lab?

Not necessarily. Going direct removes an intermediary, but you take on the legal, security, export and de-identification work yourself. Practitioners say evaluations built on data are worth roughly 10x raw data, and a data company may be able to build that layer when a single seller cannot. Compare offers on terms, not just price.

Start with eligibility, then pick a route

If you fit a program's published rules, a data company is the usual first route. If your dataset is large or unique, you can also approach a lab's own intake; no referral link applies there.

Independent site. Some links are referral links: if your company signs with a buyer through them, the buyer may pay us a fee. You are not charged, and we never see your data.

Related reading

Keep going