This page is for companies, not individuals. It compares five routes a company of roughly 10 to 1,000 people can use to license internal records to AI companies and labs, using only what each buyer publishes: payout ranges, employee minimums, the data they ask for and how they say they handle privacy.
Several companies now buy records of real work from businesses: messages, documents, tickets and code. Their published rules differ. This page puts them in one place.
Every figure below is copied from the buyer's own page, as published and checked 7 October 2026. Only a buyer that reviews your data can price it.
A range shows what a buyer will consider. The top is for large or unusual datasets. No program here publishes an average.
The buyer runs discovery, the contract, export, de-identification and payment. This site only explains, compares and links out.
Quoted text is verbatim from each program's site. Everything else is a plain summary of the same pages. Last checked: 7 October 2026.
| As published | micro1Enterprise Data Partnership | ModeCompany data | GreppedData and expertise | Miro AdvisoryDatasets and codebases | Direct to labsLab intake pages |
|---|---|---|---|---|---|
| Company payout | "$100k+ qualified", "$500k+ large-scale", "$1M+ highly unique". Referral page: "$100K-$2M+ for approved data packages". | "$100K-$5M" | "$20K-$5M", with "get paid in 7 days". | Operating datasets "$100K-$1M+"; private codebases "$10K-$1M+". Labeled indicative. | No published range. |
| Eligibility | 30+ employees (its referral posting says 30 to 200). Mature operations, documented processes, modern software tools, primarily English. US prioritized, then other Western markets. | 20+ full-time US office employees. Accounting firms 10+. Law firms 6+. Several years of records the company owns. US-based teams are the strongest fit. | Any vertical. Also pays individual professionals for expertise. | Businesses; software companies. | Large or unique datasets. |
| What they ask for | SOPs, knowledge bases, internal documentation, CRM data, project histories, QA processes; "decision-making patterns"; "AI performance feedback". | "An agreed copy" of several years of company-owned records. | Company data in any vertical; separately, expertise from individual professionals. | Operating datasets and private codebases. | Datasets large or unusual enough for a lab's own procurement team. |
| Privacy statements | Scope agreed in writing; sensitive and confidential information scrubbed; originals deleted after processing; no customer information exposed; the company keeps ownership of its underlying data. | Buys "an agreed copy"; originals stay with the company; de-identifies before onward delivery. | Not summarized here. Read its terms before you apply. | Not summarized here. Read its terms before you engage. | Set by each lab's own agreement. |
| Referral status | Referral link | Referral link | Referral link | No link, no fee to us | No link, no fee to us |
| Next step | Apply through micro1 | Apply through Mode | Apply through Grepped | Contact Miro Advisory directly on its own site. | Use the lab's own intake. Google runs one at contentpilot.google.com; OpenAI has a data partnerships page. |
Sources: each program's own pages, as published, checked 7 October 2026. On a phone, swipe the table sideways.
The first paragraph of each card restates the buyer's published rules. The checklist is our own preparation advice, not the buyer's requirement.
micro1 names operational knowledge: SOPs, knowledge bases, internal documentation, CRM data, project histories and QA processes, plus "decision-making patterns" and "AI performance feedback", meaning human feedback on AI outputs. Its referral terms describe workflow partnerships and corpus partnerships.
More detail: the micro1 program as published.
Mode asks for several years of records the company owns. Its headcount rule is lower for law firms (6+) and accounting firms (10+) than for everyone else (20+ full-time US office employees). It says it buys "an agreed copy", that originals stay with the company, and that it de-identifies before onward delivery.
Grepped publishes the lowest floor of the three referral programs and says it works with any vertical. It also pays individual professionals for their expertise, which is a separate track from a company licensing its records. It publishes "get paid in 7 days".
Miro Advisory lists two categories: operating datasets from businesses and private codebases from software companies. It labels its ranges indicative. It publishes no referral program on its site, so we have no link and receive nothing if you work with it.
Some labs take data offers directly. Google runs an intake at contentpilot.google.com, and OpenAI has a data partnerships page. Labs buy through procurement: NDA, master agreement, dataset evaluation and a purchase order. Most companies of 10 to 1,000 people go through a data company instead; direct makes sense when the dataset is large or unique.
More detail: how labs buy data directly.
Headcount, country and industry decide more than payout ranges do. Find the row closest to your company, then confirm it with the eligibility checker.
Mode publishes 6+ for law firms; Grepped publishes no minimum. Privilege removes most matter files, so expect a narrow scope.
Mode lists accounting firms at 10+; micro1 becomes an option at 30. Client data is the gate, not headcount.
Mode's 20+ full-time US office rule applies, and Grepped has no published minimum.
All three referral programs publish rules you may meet. micro1's referral posting names 30 to 200.
micro1 prioritizes the US, then other Western markets; Mode calls US-based teams the strongest fit. EU and UK teams should settle GDPR questions first.
Miro Advisory lists private codebases at an indicative "$10K-$1M+". Keep the full commit history.
A lab's own intake may make sense. Expect procurement, not a sign-up form.
Only Grepped's open "any vertical" fits. It also has a track for individual professionals.
These companies are invented to show how the published rules apply. They are illustrative, not offers, and no price is implied beyond each program's published range.
Nine years of records in Outlook and SharePoint. English. US office.
Headcount is not the problem here. Attorney-client privilege is. Most matter files cannot be included, so the realistic scope is the firm's own material: internal procedures, templates with client details removed, and administration. Mode's published "$100K-$5M" is a range, not a forecast for a narrow scope. See law firms.
Seven years in QuickBooks and Outlook, written month-end checklists. English. US office.
This firm meets the published rules of all three referral programs, which makes it a good case for more than one offer. The gate is client confidentiality: the books in QuickBooks belong to clients, and engagement letters may forbid sharing them. Workpapers about the firm's own process are easier to scope. See accounting firms.
Six years in GitHub, Jira, Slack and Confluence. English. Team in the UK.
Linked history across code, tickets and chat is what buyers describe wanting. Two checks come first: GDPR, since staff messages are personal data, and open-source licenses inside the repos. See software companies and GDPR and AI training data.
Published ranges are the easiest numbers to quote and the easiest to misread.
A range describes the span a buyer is willing to talk about. micro1 makes this explicit with tiers: "$100k+ qualified", "$500k+ large-scale" and "$1M+ highly unique". A company with a few years of ordinary project files should read the floor, not the ceiling, as its reference point.
Price also depends on what is sold. Practitioners say raw data is the cheapest tier, evaluations built on that data are worth roughly 10 times raw, and full training environments reach 6 to 8 figures but need heavy engineering. Most companies sell raw records, so most deals sit nearer the lower end.
In the market for archives of shut-down startups, Troveo cites about $5,000 per code repository and roughly $10,000 to $100,000 per archive deal. Those are Troveo's figures for closing companies, not a forecast for a running business. Our page on how much AI companies pay for data goes through what moves the number.
One rule holds whatever the range: practitioners advise never sending a full dataset before a price is agreed. Share a manifest and samples, and get more than one offer.
Practitioners cite 60 to 90 days to close. Some take longer. No program on this page can promise your timeline, and neither can we.
You apply through the buyer's own process. The buyer decides whether to follow up.
Both sides sign a confidentiality agreement before anything specific is discussed.
The buyer looks at a manifest and samples: systems, years, volume and sensitivity.
Price, scope, exclusivity, warranties and payment terms are set in writing.
The agreed copy is pulled from your systems, within the agreed scope and exclusions.
Names, customer details and confidential information are removed or replaced. Ask how this step is checked.
The buyer confirms the delivered data meets the agreed criteria. Payment often depends on this step.
One-off, in milestones or recurring, as the contract says. Read what triggers each payment before you sign.
Most of these cost leverage, not just money. They are avoidable before the first call.
With one offer you have no reference point for price or terms. Practitioners advise getting more than one offer. Applying is not signing.
A published range is the span a buyer will discuss. None of the five publishes a median, so plan from the floor.
Share a manifest and samples first. A full dataset sent early gives away the thing you are pricing.
If a second application is still running, an exclusive first contract can block it. Read the exclusivity clause before you sign anything.
Engagement letters and NDAs can forbid sharing client material. Check them before scoping, not after the buyer asks.
Mode counts US office employees; micro1 prioritizes the US. EU and UK teams carry GDPR duties wherever the buyer sits.
Mode's rule is about full-time US office employees. Count the way the rule is written, or confirm with the program.
A buyer's statement describes its own process. Your duties to staff and clients stay yours.
An hour of internal homework makes the first buyer conversation shorter and puts you in a better position to compare offers.
These are questions for every seller and every contract. They are not claims about any program on this page.
Is the license exclusive, time-limited or open? Can the buyer resell your data, and can you license the same records to anyone else?
Training only, or evaluation too? Which downstream buyers can receive it, and are they named?
Who pays if de-identification misses something? Is your liability capped, and does it expire?
What must you promise about employees, customers and clients? Can you stand behind it?
Law, accounting, M&A and healthcare work carries confidentiality duties a data license cannot override.
One-off or recurring? Milestones? What acceptance criteria must the data meet before money moves?
Can you check how personal and confidential details were removed, and see the results?
When are the originals deleted, and what happens to the copy if the deal ends?
How can either side exit, and which obligations keep running after the contract ends?
The full list, grouped by money, scope, privacy, liability and exit, is on questions to ask a data buyer.
The order of programs on this page is not a ranking, and no placement is paid. We are not a partner, agent or representative of any buyer. Acceptance, price and timing are decided by the buyer alone.
No referral link
Miro Advisory publishes no referral program on its site. We have no link and earn nothing if you work with it. Reach it through its own website.
No referral link
Labs pay us nothing when a company sells to them directly. Google's intake is at contentpilot.google.com; OpenAI has a data partnerships page.
Details per program, including how our /go/ links work: referral disclosure.