AI labs buy data the way any large company buys from a vendor: through procurement, legal review and a test of the goods. Here is what that process looks like, why most companies of 20 to 500 people reach labs through a data company, and when going direct makes sense.
Selling data to an AI lab means licensing a copy of your records to a company that trains or evaluates AI models, under a written agreement that sets scope, price, de-identification and acceptance. Mode and micro1, as published, describe a licensed or agreed copy, with originals and ownership staying with the company.
A frontier lab is a large organization with a legal team, a security team and a procurement function. When it buys training or evaluation data, expect it to run the purchase the way it runs any vendor purchase: confidentiality first, then a test of the goods, then a master agreement, then an order. Each step has an owner on the lab's side, and each one adds weeks.
That process is built for suppliers who deliver large volumes in a standard format, with privacy work and quality checks already done. A 40-person accounting firm or a 150-person software company rarely fits that shape on its own. That is the gap data companies fill: they sign many smaller sellers, clean and package the records, and sell the combined result onward. For the full map of buyers, including brokers and marketplaces, see who buys company data for AI.
You become the lab's supplier. You negotiate with its legal team, pass its security review and deliver in its format. This route suits large or unique datasets.
A buyer program such as micro1, Mode or Grepped licenses your data, runs export and de-identification under its published process, and deals with labs. This suits most companies of 20 to 500 people.
An NDA, a sample review, a written license, acceptance criteria and payment after acceptance. The paperwork looks alike; what changes is who carries the work.
Names vary by company, but expect some version of each. They tend to arrive in roughly this order. Read each one before you send anything beyond a manifest.
Covers what you show during evaluation: the manifest, samples, and anything said about your systems and clients.
Check: is it mutual? What happens to samples the lab keeps after saying no? How long do the duties last?
Sets liability, indemnities, warranties, confidentiality and termination for everything bought under it. Large buyers usually start from their own paper.
Check: is liability capped, for both sides? Which clauses survive termination, and for how long?
Which records, which years, which format, what the buyer may do with them (training, evaluation or both) and whether the license is exclusive.
Check: scope of use, exclusivity, resale rights, and deletion of the delivered copy.
Asks how the data was collected, whether people in it were notified, how it was de-identified and how it will be transferred.
Check: every answer can become a representation you are held to. Answer precisely, not hopefully.
The tests a delivery must pass to count as delivered: volume, format, completeness and de-identification checks.
Check: who decides, by what date, what happens when a batch fails, and whether you can audit the de-identification.
The buyer's payment instrument, issued against the license. Expect payment to follow acceptance rather than signature.
Check: one-off or recurring, milestone amounts, and payment terms after invoice.
General information, not legal advice. Talk to your own lawyer before you sign. For a clause-by-clause outline of what these agreements cover, see the AI data licensing agreement guide.
Most direct conversations end at this stage. A lab is not judging whether your records are interesting. It is judging whether they change a model, or measure something it cannot measure today.
You describe the dataset through the lab's intake route or a contact. Expect basic questions: source systems, years covered, volume, language, and whether personal data is inside.
Nothing detailed changes hands before this. If someone asks for files before an NDA exists, slow down.
A structured description, not the data: systems, record counts, date ranges, file types, how records link (ticket to fix to reply, deal to note to invoice) and what you will exclude. A good manifest answers most first-round questions.
A small, de-identified extract. Reviewers check that records are complete, connected and in the target language, and that the de-identification holds up under a close read.
Researchers decide whether the data fills a gap. Connected histories of real work, with decisions, reviews and approvals in them, tend to do better than loose files. Duplicated or public material does poorly.
MSA plus data license or SOW. Price, scope, exclusivity and acceptance are settled here, before the full export, not after.
You, or people you hire, export and de-identify to the agreed format. The buyer runs its acceptance checks, and the purchase order typically pays against accepted deliveries.
Practitioners cite 60 to 90 days from first contact to close, and that assumes a seller who is ready. Nothing here is a promise of timing. The seller's step-by-step guide covers the work on your side: inventory, scoping and export.
Never send a full dataset before price. Practitioners give the same advice whatever the route: share a manifest and samples, settle price and terms in writing, and try to get more than one offer. A full export sent "for evaluation" is hard to take back. See getting more than one offer.
It is not that labs refuse small sellers. A lab's process costs a lot to run per supplier, and a data company spreads that cost across many sellers.
Every direct supplier means another NDA, another MSA negotiation and another security review. A lab can buy from one data company that brings many sellers' records under a single contract.
Labs want consistent schemas across sources. Data companies convert exports from Slack, Jira, Salesforce or QuickBooks into one shape, so the lab does not have to do it seller by seller.
The data company takes on scrubbing and quality checks before onward delivery. As published, checked 7 October 2026: Mode says it de-identifies before onward delivery; micro1 says sensitive and confidential information is scrubbed and originals are deleted after processing.
One firm's years of support tickets may be too small to move a model on their own. Many firms' tickets in the same format can be. Aggregation is what turns a small seller into a useful supplier.
The layer is large. Deedy Das's July 2026 market map counts 50+ companies selling data and RL environments to labs, at about $8.5B in revenue. Will Depue (July 2026) puts labs on a path to more than $100B a year of data spend by 2030. Both are estimates.
Value also depends on what is built on top of the data. Practitioners describe three tiers: raw data is the cheapest; evaluations built on the data are worth roughly 10x raw; full training environments reach 6 to 8 figures but need heavy engineering. A company selling direct usually has only the first tier to offer.
The buyer at the end may be the same. What differs is who you sign with, who does the work, and who carries which risks.
Fits: large, unique or hard-to-replace datasets.
Fits: most companies of 20 to 500 people with connected work records.
Each company's own wording. Ranges are what the company publishes, not offers or averages.
| Route | Payout (published) | Eligibility (published) | Source |
|---|---|---|---|
| Google content intake (direct) | Not published | Suited to large or unique datasets | contentpilot.google.com |
| OpenAI data partnerships (direct) | Not published | Suited to large or unique datasets | OpenAI data partnerships page |
| micro1 Enterprise Data Partnership | "$100k+ qualified", "$500k+ large-scale", "$1M+ highly unique"; referral page: "$100K-$2M+ for approved data packages" | 30+ employees, mature operations, documented processes, modern software tools, primarily English; US prioritized, then other Western markets | micro1.ai/data-partnerships |
| Mode company data | "$100K-$5M" | 20+ full-time US office employees; accounting firms 10+; law firms 6+; several years of records the company owns | data.mode.inc |
| Grepped | "$20K-$5M" | Any vertical | grepped.ai |
| Miro Advisory | Operating datasets "$100K-$1M+"; private codebases "$10K-$1M+" (indicative) | Businesses; software companies | miroadvisory.com |
"Up to" and "+" figures mark the top of a published range, not a typical result. For what moves the number, see how much AI companies pay for data.
Direct intake is a real option, but for a narrow set of sellers. Most companies tick only one or two of these boxes.
In August 2026, Google agreed to pay $10 million in Spirit Airlines' bankruptcy proceedings for internal data (emails, Teams messages, spreadsheets and operations files), as reported by ABC, TIME and others. micro1 then made a reported $12.5 million rival offer. Court approval of the sale was not confirmed as of 7 October 2026.
The records were large, connected and specific to one airline's operations, which is the profile of a direct deal. It was also a liquidation, not a running company, so read it as context, not as a benchmark for your business.
Not sure where you land? The eligibility checker compares your size, location and systems with each program's published rules, entirely in your browser. To see the programs side by side, use buyer programs compared.
Ask these early, ideally right after the NDA. They are generic questions for any lab, not claims about how a particular one works, and the answers tell you how long the process will run and who carries which risk.
A procurement contact, a research team, or both? Who can say yes, and who signs for the lab?
Ask for the smallest sample that answers their question, in which format, and agree in writing that it is de-identified before it leaves.
Are they deleted, and will deletion be confirmed in writing? May they be used for anything after a no?
The buyer's standard MSA or yours? Which clauses are open to negotiation, and is there a data-specific addendum?
Which tests define acceptance, who runs them, by when, and what happens to payment if one batch fails?
Training, evaluation or both? Can the data reach affiliates, contractors or other parties, and on what terms?
Most of these cost a seller its leverage rather than the deal itself. All of them are avoidable before the first email goes out.
Anything sent before an NDA sits outside any confidentiality terms. Describe the dataset first; send samples only after signing.
Once a buyer holds the whole dataset, the reason to agree a fair price gets weaker. A manifest and samples are enough to evaluate.
A file count says little. Explain how records link: the request, the discussion, the decision, the result. That is what reviewers look for.
Questionnaire answers can become representations. Check employee notice and your client contracts before you answer, not after.
De-identifying and formatting to a buyer's schema is real engineering. Budget for it, or use a data company that does it.
One lab's answer is not a price. Practitioners advise more than one offer, from labs and data companies alike, compared on terms as well as money.
It can try. Some labs publish intake routes, such as Google's content intake at contentpilot.google.com and OpenAI's data partnerships page. In practice, labs run procurement built for large suppliers, so most companies of 20 to 500 people reach labs through a data company that aggregates, de-identifies and packages records from many sellers.
Expect a mutual NDA first, then a master agreement (MSA), a data license or statement of work that sets scope and price, a security and privacy questionnaire, written acceptance criteria and a purchase order. Have your own lawyer read each one before you sign.
Usually from a manifest (systems, record counts, date ranges, file types, exclusions) and a small de-identified sample. Researchers then judge whether the data fills a gap their models have. Connected histories of real work tend to rate higher than scattered, duplicated or public files.
Source systems, record counts, date ranges, file types, languages, how records link to each other, what personal or client data is inside, and what you will exclude. A manifest describes the data without containing it, so you can share it under an NDA before any sample leaves your company.
Practitioners cite 60 to 90 days to close, through NDA, review, agreement, export, de-identification and acceptance. It can take longer, and no route can promise a timeline.
Not necessarily. Going direct removes an intermediary, but you take on the legal, security, export and de-identification work yourself. Practitioners say evaluations built on data are worth roughly 10x raw data, and a data company may be able to build that layer when a single seller cannot. Compare offers on terms, not just price.
If you fit a program's published rules, a data company is the usual first route. If your dataset is large or unique, you can also approach a lab's own intake; no referral link applies there.
Independent site. Some links are referral links: if your company signs with a buyer through them, the buyer may pay us a fee. You are not charged, and we never see your data.