The honest answer first: nobody can price your company's data without seeing it, and no buyer publishes an average. What buyers do publish are ranges. Today those run from $20K to $5M per company, with the top tiers reserved for data the buyer calls large-scale or highly unique. Below is every published range, every reported deal price we can source, and what moves a number up or down.
Four programs that buy company data publish payout figures on their own pages. The table quotes them word for word. Each one is a range or a threshold, and each describes what the buyer chooses to advertise, not what a typical company receives.
| Program | Payout as published | How the wording reads | Source (plain text) |
|---|---|---|---|
| micro1 Enterprise Data Partnership | "$100k+ qualified", "$500k+ large-scale", "$1M+ highly unique"; referral page: "$100K-$2M+ for approved data packages" | Three tiers tied to the buyer's own labels. The top tier is open-ended and conditional on uniqueness. | micro1.ai/data-partnerships; micro1.ai/company-referral |
| Mode company data | "$100K-$5M" | One range for company data. Mode says it buys "an agreed copy" and originals stay with the company. | data.mode.inc |
| Grepped | "$20K-$5M" | The lowest published floor of the four. Open to any vertical. | grepped.ai |
| Miro Advisory | Operating datasets "$100K-$1M+"; private codebases "$10K-$1M+" | Marked as indicative by Miro Advisory. Separate ranges for operating data and code. | miroadvisory.com |
Last checked: 7 October 2026, from each program's own page. Quotes are verbatim. A published range is not a quote for your company, and no program promises any amount, acceptance or timing.
Two patterns stand out. The floors cluster at $100K, with Grepped at $20K and Miro Advisory's codebase range at $10K, so a smaller dataset may fit better where the floor is lower. And every ceiling, from "$1M+" to $5M, is open-ended, conditional or marked indicative. None tells you where most deals land.
Each gridline is ten times the one before. Green bars are program ranges as published; a faded end means an open top ("+"). Slate marks are closure-market figures cited by Troveo.
These bars mark the published edges of each range. They are not distributions and say nothing about where most deals land. Sources: each program's own page, last checked 7 October 2026; Troveo as cited in 2026 news coverage.
A published range is a screening tool. It tells you whether a conversation is worth having. It is not a quote, a forecast or an average, and planning around it as if it were leads to bad decisions.
A floor is the smallest deal a program pursues; a ceiling is the largest it advertises. Neither tells you where the middle is.
micro1's page attaches "$500k+" to "large-scale" and "$1M+" to "highly unique" data. Those are the buyer's own labels for the exceptions. A 60-person firm with ordinary records is, by that wording, not the exception.
None of the program pages we checked publishes an average, a median, or how many companies were paid at the top. Read any range as "possible", not "probable".
"$2M+" means the buyer will not rule out more. It does not mean more is common.
Budget for zero until it is signed. Until a buyer has reviewed a manifest and samples of your data, any figure you have seen is a published range, not an offer. Programs set eligibility rules, so some applicants are not accepted. Do not book the revenue, hire against it or promise staff a share until you have a signed agreement and an accepted delivery.
Deal prices do appear in the news, but almost none describe a company of 20 to 500 people. Here is each figure, who reported it, and how far it transfers.
| Deal | Reported price | Type | Reported by | Transfers to a running company? |
|---|---|---|---|---|
| Spirit Airlines internal data | $10 million (Google agreed to pay); $12.5 million rival offer from micro1 | One company's internal records, sold in bankruptcy proceedings | ABC, TIME and others, August 2026 | Closest analog, but a large airline's full operating record sold in a court process |
| Google and Reddit | about $60M a year | Platform licensing, annual, direct | Widely reported | No: public content at platform scale |
| News Corp and OpenAI | over $250M over five years | Publisher licensing, direct | Widely reported | No: a news archive and ongoing content |
| Shutterstock AI licensing | about $104M (2023) | Reported AI licensing revenue for 2023 (a later $138M figure for 2024 was a projection for the whole business unit, not AI licensing alone) | Reported company figures | No: a licensing business, not one sale |
| Shut-down startup archives | about $5,000 per code repository; roughly $10,000 to $100,000 per archive deal | Closure market for Slack, email and code | Troveo, as cited; Forbes 16 April 2026, Fast Company, Gizmodo | Partly: real company records, but from firms that stopped operating |
Last checked: 7 October 2026. Figures as reported in the named coverage. We have not seen any of these contracts. Full list with sources: AI data licensing deals tracker.
The emails, Teams messages, spreadsheets and operations files of one company drew competing eight-figure bids, as reported in August 2026. But it was a large airline's entire operating history, offered in a bankruptcy sale whose court approval was not confirmed as of 7 October 2026. Status and outcome: Spirit Airlines case page.
These are direct, usually annual licenses for very large content libraries. They prove labs pay for data at scale. They are not a yardstick for a 200-person company's Slack history.
Forbes ("AI's New Training Data: Your Old Work Slacks And Emails", 16 April 2026), Fast Company and Gizmodo covered startups selling old archives. Those sellers no longer operate. More: shut-down startups selling Slack and email.
Both are estimates by individuals, not audited figures. They say demand is large. They say nothing about what one company's records are worth.
Buyers do not publish pricing formulas. Their eligibility rules and tier labels, plus what practitioners say, point to the same handful of factors.
Mode asks for several years of records the company owns. Longer histories show how work changed and how decisions played out.
A ticket linked to the chat thread, the fix and the customer reply is a full workflow. The same files scattered across drives say much less.
micro1 reserves "$1M+" for "highly unique" data. Records of work few others do, in a niche process or industry, sit at the top.
micro1 lists 30+ employees. Mode lists 20+ full-time US office staff, 10+ for accounting firms and 6+ for law firms. More people produce more of the record.
micro1 lists mature operations, documented processes and modern software tools. Data in standard systems exports cleanly.
micro1 lists primarily English, US prioritized, then other Western markets. Mode names US-based teams as the strongest fit.
Data you own outright, with no client contract that forbids sharing it, needs fewer carve-outs. Every carve-out shrinks the scope.
Exclusive and non-exclusive licenses can be priced differently, and exclusivity limits selling the same data again. Ask how the price changes with each.
Exports of messages, documents, tickets and code. Practitioners say this is where prices are lowest, and it is where most sellers start.
Test sets built on the data, where people who know the work define a correct answer. micro1 also lists "AI performance feedback", human feedback on AI outputs, among what it wants.
Full simulations of a workflow that a model can practice in. Practitioners say they need heavy engineering that most companies do not have in-house.
Tier values are what practitioners say, not buyer prices. More on what makes records valuable: what data AI labs want.
Often the bigger swing is not the rate but how much data survives client contracts, personal-data rules and company policy.
Every company, number and decision below is invented to show the mechanics. None of it comes from any buyer, and real outcomes can be higher, lower or zero.
| Source | First inventory | After scope review | Why it changed |
|---|---|---|---|
| Outlook email | 7 years, all mailboxes | 7 years, operations mailboxes only | Some shipper contracts restrict sharing their correspondence |
| Microsoft Teams | All channels and direct messages | Team channels only | Direct messages excluded by company decision |
| Exception tickets | 7 years | 7 years | Kept: the core record of how problems get solved |
| SOPs and playbooks | All | All | Kept: easy to scope and de-identify |
| HR and payroll files | Included | Removed | Personal data with little workflow value |
The lessons carry over. Scope before you discuss price, so the number you hear is for data you can deliver. The parts that survive review, here the linked exception tickets and SOPs, are usually what buyers care about most. And a smaller scope means fewer consent and confidentiality warranties to give.
A headline number can hide very different deals. Put these questions to any buyer before you compare figures. They are general questions for every seller, not statements about any program on this page.
Each mistake below comes from reading a number without the conditions attached to it. Each one has a simple fix that costs nothing but time.
Treating "$2M+" or "$5M" as the likely figure. The top of a published range is the exception a buyer chooses to advertise, and the top tiers carry labels such as "large-scale" and "highly unique".
Instead: plan around the floor until you hold a written offer, and read the label attached to every tier.
A single payment for one delivery and a deal with paid refreshes are different products with different work attached. Their headline numbers cannot be lined up directly.
Instead: compare what you receive over the same period, for the same delivery obligations.
A price is paid on data the buyer accepts. If the contract lets the buyer reject part of a delivery, the number you signed for and the number you receive can differ.
Instead: ask what is tested, who decides it passed, and how a partial rejection changes the price.
A figure given before client contracts, personal data and company policy are reviewed describes data you may not be able to deliver. It can move once that review is done.
Instead: run the scope review first, as in the worked example above, and ask for a price on what remains.
The reported Reddit, News Corp and Shutterstock figures describe content licensing at platform scale. Dividing them by headcount or file count produces a number no company buyer has published.
Instead: measure yourself against the published program ranges for companies until you hold an offer.
One offer has nothing to be measured against, and sending a full dataset to get it gives away your leverage. Practitioners advise a manifest, samples and more than one offer.
Instead: send the same manifest to more than one program and compare the terms side by side.
The only price that counts is one a buyer puts in writing after reviewing your data.
Compare your team size, location, industry and systems with each program's published rules in the eligibility checker before you spend time on anything else.
A manifest lists systems, date ranges, volumes and exclusions. Practitioners say never send a full dataset before price: send the manifest and samples.
One offer has nothing to be measured against. See getting more than one offer.
Lay payment structure, exclusivity, scope of use and liability side by side. A lower figure with capped liability can be the better deal.
As published on their own pages (checked 7 October 2026): micro1 lists $100K-$2M+ for approved data packages, with tiers of $100k+ qualified, $500k+ large-scale and $1M+ highly unique. Mode lists $100K-$5M. Grepped lists $20K-$5M. Miro Advisory lists indicative ranges of $100K-$1M+ for operating datasets and $10K-$1M+ for private codebases. These are ranges, not quotes. Nobody can price your data without reviewing it.
Not a published one. None of the program pages we checked publishes an average deal size, a median, or how many companies were paid at the top of a range. A published range tells you whether a conversation is worth having. It does not tell you what a typical company receives.
As reported in August 2026 by ABC, TIME and others, Google agreed to pay $10 million in Spirit's bankruptcy proceedings for internal data, and micro1 then made a $12.5 million rival offer. Court approval of the sale was not confirmed as of 7 October 2026; our Spirit Airlines case page tracks the outcome.
Google-Reddit (about $60M a year), News Corp-OpenAI (over $250M over five years) and Shutterstock's AI licensing revenue (about $104M in 2023) are reported figures for platforms and publishers licensing very large content libraries. They are not benchmarks for one company's internal records.
Troveo cites about $5,000 per code repository and roughly $10,000 to $100,000 per archive deal in the closure market. These figures describe companies that stopped operating.
Years of history, records that connect across systems, uniqueness, documented processes in standard tools, clear ownership with few carve-outs, and the license terms you agree to. Practitioners also say evaluations built on data are worth roughly 10x raw data, and full training environments reach 6 to 8 figures but need heavy engineering.
No. A range shows what a program advertises across every company it works with. A figure for your company exists only after a buyer has reviewed a manifest and samples of your data and put an offer in writing. No published range is a promise of any amount, acceptance or timing, and programs apply eligibility rules, so some applicants are not accepted.
Not without something to compare it to. Practitioners advise sharing a manifest and samples, never the full dataset before price, and getting more than one offer. Compare offers on payment structure, acceptance criteria, exclusivity, scope of use and liability, not only on the headline number.
The checker compares your answers with each program's published rules and shows each range word for word. It asks for no email.
Independent site. Some links are referral links: if your company signs with a buyer through them, the buyer may pay us a fee. You are not charged, and we never see your data.
Applying does not guarantee acceptance or any amount. Each program runs its own review and sets its own terms and timing.