Sell Data to AI
Home Data Asset Score Pricing API documentation US labs and CROs list
For brokers
How to become an AI data broker Data broker business model Buyer programs compared Qualify a company AI training data companies
For data companies
Firmographic data providers Company data API
Seller guides
How to sell data to AI companies Is it legal? FAQ and glossary About
Check domain/company
Buyer map for companies

Who Buys Company Data for AI?

Last checked: 7 October 2026

This guide is for companies, not consumers. AI labs do pay for company data, but a 20 to 500 person firm rarely signs with a lab itself. The money moves through four layers. Here is who sits in each one, how each gets paid, and which one will actually talk to a company your size.

50+companies selling data and RL environments to labs (Deedy Das market map, July 2026)
~$8.5Btheir estimated combined revenue, same source
$20K to $5Mspan of published buyer program ranges, checked 7 Oct 2026
6 to 30+published employee minimums, by program and industry
The short answer

Yes, AI companies buy data. Mostly not from you directly.

Model builders need records of real work: how a support team resolves a ticket, how an accounting firm closes a month, how engineers review each other’s code. Public web text rarely shows that. The systems inside an ordinary company do. That gap is why buyer programs aimed at businesses appeared, and why the ranges they publish run from $20K at the low end to $5M at the top.

The spending behind it is large, and it is estimated rather than audited. Deedy Das’s market map (July 2026) counts more than 50 companies that sell data and RL environments to labs, with about $8.5B in revenue between them. Will Depue (July 2026) puts labs on a path to more than $100B a year of data spend by 2030. Both are estimates from people who follow the market, not figures from company filings.

What matters for an owner or finance lead is narrower: which of these buyers signs with a company of 10 to 1,000 people, on what published terms, and what each one says it does with the data.

What “buyer” means in this guide

The buyer is the party that signs a data agreement with your company and pays you. It might be the lab that trains on the data. More often it is a data company that licenses the data from you, prepares it and supplies it onward. The difference decides who you negotiate with, who de-identifies the data and whose name sits on the contract.

The market, layer by layer

The four layers between your records and a model

Each layer has a different relationship to your data, a different way of being paid and a different minimum size it will work with.

1Layer

AI labs

The companies that train frontier models and use the data in the end. The direct deals that have been reported are mostly licenses with large content platforms and news publishers, at a scale far beyond an 80 person services firm. The reported figures are covered in what AI companies pay for data.

Labs do accept direct submissions. Google runs a content intake (contentpilot.google.com) and OpenAI has a data partnerships page. Direct is a real option for large or unique datasets.

Deals with firms your size? Rarely, unless the dataset is large or unique
2Layer

Data companies with buyer programs

Firms that license data from businesses, process it and supply it to labs. Several publish programs for companies: micro1’s Enterprise Data Partnership, Mode, Grepped and Miro Advisory. As published, they run discovery, agree scope in writing, handle export and de-identification, and pay the seller.

This is the layer built for a company of 10 to 1,000 people. Each publishes a payout range and eligibility rules, so you can check fit before you sign an NDA.

Deals with firms your size? Yes, the main path for most firms
3Layer

Brokers, representatives and referrers

Intermediaries who do not buy the data. A seller-side representative shops your dataset to several buyers and may be paid by you, by the buyer or both, so ask. A referrer sends companies to a buyer program and is paid by the buyer if a deal is signed.

This site is a referrer. The programs publish their referral fees: micro1 “Earn $50,000” per company, Mode “Earn $50K per referral”, Grepped “Refer for another $10K”. The buyer pays it. You are not charged, and a referrer never needs to see your data.

Deals with firms your size? Yes, but ask how they are paid
4Layer

Marketplaces and archive buyers

Platforms that list datasets for sale, and buyers that acquire archives from companies that are closing. Forbes (16 April 2026, “AI’s New Training Data: Your Old Work Slacks And Emails”), Fast Company and Gizmodo covered shut-down startups selling old Slack and email archives.

Troveo, cited in that coverage, puts prices at about $5,000 per code repository and roughly $10,000 to $100,000 per archive deal. That is a liquidation market: one-off sales, smaller sums and no running business behind the data.

Deals with firms your size? Mostly for companies winding down
Fit by size and situation

Which layer will talk to a company like yours

Published rules decide most of this before anyone looks at a file. Headcount is the first filter; years of records, documented processes and location come next.

Your situationLayer that usually fitsPublished rule behind it
Law firm, 6 to 19 peopleData company (Mode)Mode lists law firms at 6+ people
Accounting firm, 10 to 19 peopleData company (Mode)Mode lists accounting firms at 10+
Other business, 20 to 29 full-time US office staffData companies (Mode, Grepped)Mode lists 20+ full-time US office employees; Grepped lists any vertical
30 to 200 employees, documented processes, modern softwareData companies (micro1, Mode, Grepped)micro1 lists 30+ employees; its referral posting says 30 to 200
200 to 1,000 employeesData companies; direct lab intake for a unique datasetmicro1 lists 30+; Mode lists 20+
Software company with private repositoriesData companies, including Miro AdvisoryMiro Advisory lists private codebases at $10K to $1M+ (indicative)
Shutting down or already closedArchive buyers (layer 4), or a data company before you closeClosure-market prices cited by Troveo, reported 2026
Large or unique dataset, in-house legal teamDirect lab intake (layer 1)Google and OpenAI publish intake pages

Last checked: 7 October 2026. Sources, as published: micro1.ai/data-partnerships and micro1.ai/company-referral; data.mode.inc; grepped.ai; miroadvisory.com; Forbes, 16 April 2026. Buyers change terms; check each program before you apply.

What the layer-2 programs publish

Published payout ranges and eligibility, quoted as each program states them. A range is not a quote for your company.

ProgramPublished company payoutPublished eligibility
micro1 Enterprise Data Partnership$100k+ qualified, $500k+ large-scale, $1M+ highly unique; referral page: $100K to $2M+ for approved data packages30+ employees, mature operations, documented processes, modern software tools, primarily English; US prioritized, then other Western markets
Mode company data$100K to $5M20+ full-time US office employees; accounting firms 10+; law firms 6+; several years of records the company owns; US-based teams the strongest fit
Grepped$20K to $5M; site says “get paid in 7 days”Any vertical; also pays individual professionals for expertise
Miro AdvisoryOperating datasets $100K to $1M+; private codebases $10K to $1M+ (indicative)Businesses; software companies

Last checked: 7 October 2026, from each program’s own pages (micro1.ai, data.mode.inc, grepped.ai, miroadvisory.com). If a program publishes a timing claim, ask when the clock starts: at signature, delivery or acceptance. The full side-by-side is on buyer programs compared.

Four fictional companies

Which layer fits: four worked scenarios

The companies below are invented to show how the published rules sort real situations. No prices are shown, because nobody can price records without reviewing them.

FictionalLayer 2

A 14-person law firm in Ohio

It is below most headcount rules, but Mode lists law firms at 6+ people, so it meets that published minimum. The harder question is not size. Attorney-client privilege and client confidentiality rule out most matter files.

What is left to offer: internal templates, intake procedures, billing workflows and practice-management records with client material removed. A lawyer reviews scope before any application.

FictionalLayer 2, several programs

A 120-person US software company

It meets micro1’s 30+ rule and Mode’s 20+ US office rule. It has years of linked history in GitHub, Jira and Slack. Miro Advisory also lists private codebases among the data it covers.

Sensible route: one manifest and the same samples to more than one program, then compare written offers on scope, exclusivity and payment terms, not only the headline figure.

FictionalLayer 4 or layer 2

A 40-person startup winding down

The company will close in four months. Its Slack, email and code archives are the kind of material the 2026 closure-market coverage described. Archive buyers are one option; a buyer program is another while the company still has staff who can scope and export.

Watch for: board and creditor approval, what former employees were told, and who still has admin access after shutdown.

FictionalLayer 1 and layer 2

A 900-person manufacturer with 20 years of process records

Quality reports, line changes and engineering decisions going back two decades, in a niche few others work in. That is the “large or unique” profile where a lab’s own intake is worth a look, alongside the buyer programs.

The catch: going direct means its own legal, security and export work. It can run both routes in parallel only if no NDA it signs prevents that.

Direct lab intake

Going straight to a lab: when it fits and when it does not

Labs buy at scale, and their reported deals are with platforms. A direct route still exists. Whether it suits you depends on the dataset more than on company size.

Why most firms use a data company

  • Labs prefer consistent, packaged data from fewer suppliers. One firm’s archive is small on its own; a data company combines many.
  • The data company does the export, formatting and de-identification a lab would otherwise ask you to do.
  • Programs publish ranges and eligibility, so you know if you fit before you sign anything.
  • The agreement is written for businesses, though every clause still needs your own lawyer’s review.

When direct intake makes sense

  • The dataset is large or unique: years of records in a niche few others hold.
  • You have legal staff who can handle a lab’s NDA, master agreement and security review.
  • You can export, de-identify and deliver to a technical spec, or pay someone to.
  • You can wait. Deals take weeks to months; practitioners cite 60 to 90 days to close.

We earn nothing when a company goes direct to a lab, and it is still the right route for some datasets. The procurement side is covered in how to sell data to AI labs.

What happens to the data

What the programs say they do with your records

Statements from the buyers’ own pages, as published, checked 7 October 2026. They describe each program’s stated process; your contract is what binds it.

micro1, as published

  • Scope is agreed in writing.
  • Sensitive and confidential information is scrubbed.
  • Originals are deleted after processing.
  • No customer information is exposed.
  • The company keeps ownership of its underlying data.

Source: micro1.ai/data-partnerships, checked 7 October 2026.

Mode, as published

  • It buys “an agreed copy” of the data.
  • Originals stay with the company.
  • It de-identifies the data before onward delivery.

Source: data.mode.inc, checked 7 October 2026.

Seven questions that tell you which layer you are dealing with

Useful when someone contacts you about your data, or before you apply anywhere. These are general questions for any buyer, not claims about a particular one.

  1. Who pays me, and who pays you?
    A buyer pays you. A broker may charge you. A referrer is paid by the buyer. Get the answer in writing.
  2. Will you take a copy of the data, or only introduce me?
    Anyone who takes a copy is a buyer or a processor and should sign a data agreement with you.
  3. Who uses the data in the end?
    Ask whether it goes to one lab, several, or is resold. Scope of use belongs in the contract.
  4. Is it a license or a sale, and is it exclusive?
    None, time-limited and perpetual exclusivity are very different deals. So are resale rights.
  5. Who de-identifies the data, and can I verify it?
    Ask who is liable if something is missed, and whether that liability is capped.
  6. How is payment structured?
    One-off or recurring, milestones, and the acceptance criteria that trigger payment.
  7. What happens to the originals and the copy at the end?
    Deletion, termination and which obligations survive the contract.
Before any buyer sees your data: practitioners advise never sending a full dataset before there is a price. Share a manifest (what systems, how many years, what volume) and a few samples, and get more than one offer. Our guide to AI data brokers covers how intermediaries are paid, and what AI companies pay for data covers the published numbers.
FAQ

Questions owners ask about who buys company data

Do AI companies buy data from ordinary businesses?

Yes, mostly through data companies that run buyer programs for businesses. As published and checked on 7 October 2026: micro1 lists $100K to $2M+ for approved data packages on its referral page, Mode lists $100K to $5M, and Grepped lists $20K to $5M. These are published ranges, not quotes. Nobody can price a dataset without reviewing it.

Can I sell my company data directly to Google or OpenAI?

Both accept direct submissions: Google runs a content intake page and OpenAI has a data partnerships page. Direct intake fits large or unique datasets, and sellers who can handle a lab contract, the export and the de-identification themselves. Most firms with 10 to 1,000 employees go through a data company instead.

Who pays more, a lab or a data company?

No one publishes comparable numbers. The direct lab deals that have been reported are licenses with large content platforms and publishers, a different market from company records. Buyer programs publish ranges. The only reliable way to compare is to show the same manifest and samples to more than one buyer and compare written offers on price and terms.

What is the difference between a data broker and a referrer?

A seller-side broker or representative shops your dataset to several buyers and may charge you a fee. A referrer only introduces your company to a buyer program and is paid by the buyer if a deal is signed. This site is a referrer: we do not negotiate for you, sign anything or see your data.

How small can a company be and still find a buyer?

Published minimums, checked 7 October 2026: Mode lists law firms with 6+ people, accounting firms with 10+ and other businesses with 20+ full-time US office employees; micro1 lists 30+ employees. Grepped lists any vertical. Headcount is one rule among several: years of records, documented processes and modern software tools also count.

What does a data company do with my data after it licenses it?

As published and checked on 7 October 2026: micro1 says scope is agreed in writing, sensitive and confidential information is scrubbed, originals are deleted after processing, no customer information is exposed and the company keeps ownership of its underlying data. Mode says it buys an agreed copy, the originals stay with the company, and it de-identifies the data before onward delivery. Ask any buyer which downstream parties may receive the data and for which uses, and make sure the answer is in the contract.

Are there buyers for companies outside the US?

Fewer, by published rules. micro1 lists primarily English operations and prioritizes the US, then other Western markets. Mode names US-based teams as the strongest fit. Grepped lists any vertical. Sellers in the EU and UK also face privacy-law questions, such as GDPR, that US sellers may not, so expect more review and talk to your own lawyer before you apply.

Does applying through a referral link change anything for my company?

You still apply through the buyer program, and the program runs its own review, terms and timing. When a company signs through a referral link, the buyer may pay the referrer a fee. You are not charged. This site does not negotiate for you, sign anything on your behalf or see your data, and we cannot promise acceptance or any amount.

Found your layer? Check fit, then apply where you qualify.

Most firms of 10 to 1,000 people belong in layer 2. Check the published rules against your company first, then apply to the programs you fit. More than one application on the same manifest is normal.

Independent site. Some links are referral links: if your company signs with a buyer through them, the buyer may pay us a fee. You are not charged, and we never see your data.

Related reading

Next steps after the buyer map