This guide is for companies, not consumers. AI labs do pay for company data, but a 20 to 500 person firm rarely signs with a lab itself. The money moves through four layers. Here is who sits in each one, how each gets paid, and which one will actually talk to a company your size.
Model builders need records of real work: how a support team resolves a ticket, how an accounting firm closes a month, how engineers review each other’s code. Public web text rarely shows that. The systems inside an ordinary company do. That gap is why buyer programs aimed at businesses appeared, and why the ranges they publish run from $20K at the low end to $5M at the top.
The spending behind it is large, and it is estimated rather than audited. Deedy Das’s market map (July 2026) counts more than 50 companies that sell data and RL environments to labs, with about $8.5B in revenue between them. Will Depue (July 2026) puts labs on a path to more than $100B a year of data spend by 2030. Both are estimates from people who follow the market, not figures from company filings.
What matters for an owner or finance lead is narrower: which of these buyers signs with a company of 10 to 1,000 people, on what published terms, and what each one says it does with the data.
The buyer is the party that signs a data agreement with your company and pays you. It might be the lab that trains on the data. More often it is a data company that licenses the data from you, prepares it and supplies it onward. The difference decides who you negotiate with, who de-identifies the data and whose name sits on the contract.
Each layer has a different relationship to your data, a different way of being paid and a different minimum size it will work with.
The companies that train frontier models and use the data in the end. The direct deals that have been reported are mostly licenses with large content platforms and news publishers, at a scale far beyond an 80 person services firm. The reported figures are covered in what AI companies pay for data.
Labs do accept direct submissions. Google runs a content intake (contentpilot.google.com) and OpenAI has a data partnerships page. Direct is a real option for large or unique datasets.
Firms that license data from businesses, process it and supply it to labs. Several publish programs for companies: micro1’s Enterprise Data Partnership, Mode, Grepped and Miro Advisory. As published, they run discovery, agree scope in writing, handle export and de-identification, and pay the seller.
This is the layer built for a company of 10 to 1,000 people. Each publishes a payout range and eligibility rules, so you can check fit before you sign an NDA.
Intermediaries who do not buy the data. A seller-side representative shops your dataset to several buyers and may be paid by you, by the buyer or both, so ask. A referrer sends companies to a buyer program and is paid by the buyer if a deal is signed.
This site is a referrer. The programs publish their referral fees: micro1 “Earn $50,000” per company, Mode “Earn $50K per referral”, Grepped “Refer for another $10K”. The buyer pays it. You are not charged, and a referrer never needs to see your data.
Platforms that list datasets for sale, and buyers that acquire archives from companies that are closing. Forbes (16 April 2026, “AI’s New Training Data: Your Old Work Slacks And Emails”), Fast Company and Gizmodo covered shut-down startups selling old Slack and email archives.
Troveo, cited in that coverage, puts prices at about $5,000 per code repository and roughly $10,000 to $100,000 per archive deal. That is a liquidation market: one-off sales, smaller sums and no running business behind the data.
Published rules decide most of this before anyone looks at a file. Headcount is the first filter; years of records, documented processes and location come next.
| Your situation | Layer that usually fits | Published rule behind it |
|---|---|---|
| Law firm, 6 to 19 people | Data company (Mode) | Mode lists law firms at 6+ people |
| Accounting firm, 10 to 19 people | Data company (Mode) | Mode lists accounting firms at 10+ |
| Other business, 20 to 29 full-time US office staff | Data companies (Mode, Grepped) | Mode lists 20+ full-time US office employees; Grepped lists any vertical |
| 30 to 200 employees, documented processes, modern software | Data companies (micro1, Mode, Grepped) | micro1 lists 30+ employees; its referral posting says 30 to 200 |
| 200 to 1,000 employees | Data companies; direct lab intake for a unique dataset | micro1 lists 30+; Mode lists 20+ |
| Software company with private repositories | Data companies, including Miro Advisory | Miro Advisory lists private codebases at $10K to $1M+ (indicative) |
| Shutting down or already closed | Archive buyers (layer 4), or a data company before you close | Closure-market prices cited by Troveo, reported 2026 |
| Large or unique dataset, in-house legal team | Direct lab intake (layer 1) | Google and OpenAI publish intake pages |
Last checked: 7 October 2026. Sources, as published: micro1.ai/data-partnerships and micro1.ai/company-referral; data.mode.inc; grepped.ai; miroadvisory.com; Forbes, 16 April 2026. Buyers change terms; check each program before you apply.
Published payout ranges and eligibility, quoted as each program states them. A range is not a quote for your company.
| Program | Published company payout | Published eligibility |
|---|---|---|
| micro1 Enterprise Data Partnership | $100k+ qualified, $500k+ large-scale, $1M+ highly unique; referral page: $100K to $2M+ for approved data packages | 30+ employees, mature operations, documented processes, modern software tools, primarily English; US prioritized, then other Western markets |
| Mode company data | $100K to $5M | 20+ full-time US office employees; accounting firms 10+; law firms 6+; several years of records the company owns; US-based teams the strongest fit |
| Grepped | $20K to $5M; site says “get paid in 7 days” | Any vertical; also pays individual professionals for expertise |
| Miro Advisory | Operating datasets $100K to $1M+; private codebases $10K to $1M+ (indicative) | Businesses; software companies |
Last checked: 7 October 2026, from each program’s own pages (micro1.ai, data.mode.inc, grepped.ai, miroadvisory.com). If a program publishes a timing claim, ask when the clock starts: at signature, delivery or acceptance. The full side-by-side is on buyer programs compared.
The companies below are invented to show how the published rules sort real situations. No prices are shown, because nobody can price records without reviewing them.
It is below most headcount rules, but Mode lists law firms at 6+ people, so it meets that published minimum. The harder question is not size. Attorney-client privilege and client confidentiality rule out most matter files.
What is left to offer: internal templates, intake procedures, billing workflows and practice-management records with client material removed. A lawyer reviews scope before any application.
It meets micro1’s 30+ rule and Mode’s 20+ US office rule. It has years of linked history in GitHub, Jira and Slack. Miro Advisory also lists private codebases among the data it covers.
Sensible route: one manifest and the same samples to more than one program, then compare written offers on scope, exclusivity and payment terms, not only the headline figure.
The company will close in four months. Its Slack, email and code archives are the kind of material the 2026 closure-market coverage described. Archive buyers are one option; a buyer program is another while the company still has staff who can scope and export.
Watch for: board and creditor approval, what former employees were told, and who still has admin access after shutdown.
Quality reports, line changes and engineering decisions going back two decades, in a niche few others work in. That is the “large or unique” profile where a lab’s own intake is worth a look, alongside the buyer programs.
The catch: going direct means its own legal, security and export work. It can run both routes in parallel only if no NDA it signs prevents that.
Labs buy at scale, and their reported deals are with platforms. A direct route still exists. Whether it suits you depends on the dataset more than on company size.
We earn nothing when a company goes direct to a lab, and it is still the right route for some datasets. The procurement side is covered in how to sell data to AI labs.
Statements from the buyers’ own pages, as published, checked 7 October 2026. They describe each program’s stated process; your contract is what binds it.
Source: micro1.ai/data-partnerships, checked 7 October 2026.
Source: data.mode.inc, checked 7 October 2026.
Useful when someone contacts you about your data, or before you apply anywhere. These are general questions for any buyer, not claims about a particular one.
Yes, mostly through data companies that run buyer programs for businesses. As published and checked on 7 October 2026: micro1 lists $100K to $2M+ for approved data packages on its referral page, Mode lists $100K to $5M, and Grepped lists $20K to $5M. These are published ranges, not quotes. Nobody can price a dataset without reviewing it.
Both accept direct submissions: Google runs a content intake page and OpenAI has a data partnerships page. Direct intake fits large or unique datasets, and sellers who can handle a lab contract, the export and the de-identification themselves. Most firms with 10 to 1,000 employees go through a data company instead.
No one publishes comparable numbers. The direct lab deals that have been reported are licenses with large content platforms and publishers, a different market from company records. Buyer programs publish ranges. The only reliable way to compare is to show the same manifest and samples to more than one buyer and compare written offers on price and terms.
A seller-side broker or representative shops your dataset to several buyers and may charge you a fee. A referrer only introduces your company to a buyer program and is paid by the buyer if a deal is signed. This site is a referrer: we do not negotiate for you, sign anything or see your data.
Published minimums, checked 7 October 2026: Mode lists law firms with 6+ people, accounting firms with 10+ and other businesses with 20+ full-time US office employees; micro1 lists 30+ employees. Grepped lists any vertical. Headcount is one rule among several: years of records, documented processes and modern software tools also count.
As published and checked on 7 October 2026: micro1 says scope is agreed in writing, sensitive and confidential information is scrubbed, originals are deleted after processing, no customer information is exposed and the company keeps ownership of its underlying data. Mode says it buys an agreed copy, the originals stay with the company, and it de-identifies the data before onward delivery. Ask any buyer which downstream parties may receive the data and for which uses, and make sure the answer is in the contract.
Fewer, by published rules. micro1 lists primarily English operations and prioritizes the US, then other Western markets. Mode names US-based teams as the strongest fit. Grepped lists any vertical. Sellers in the EU and UK also face privacy-law questions, such as GDPR, that US sellers may not, so expect more review and talk to your own lawyer before you apply.
You still apply through the buyer program, and the program runs its own review, terms and timing. When a company signs through a referral link, the buyer may pay the referrer a fee. You are not charged. This site does not negotiate for you, sign anything on your behalf or see your data, and we cannot promise acceptance or any amount.
Most firms of 10 to 1,000 people belong in layer 2. Check the published rules against your company first, then apply to the programs you fit. More than one application on the same manifest is normal.
Independent site. Some links are referral links: if your company signs with a buyer through them, the buyer may pay us a fee. You are not charged, and we never see your data.