Sell Data to AI
Home Data Asset Score Pricing API documentation US labs and CROs list
For brokers
How to become an AI data broker Data broker business model Buyer programs compared Qualify a company AI training data companies
For data companies
Firmographic data providers Company data API
Seller guides
How to sell data to AI companies Is it legal? FAQ and glossary About
Check domain/company
Pricing guide · published figures only

How Much Do AI Companies Pay for Data?

Last checked: 7 October 2026 · buyer pages and named news reports

The honest answer first: nobody can price your company's data without seeing it, and no buyer publishes an average. What buyers do publish are ranges. Today those run from $20K to $5M per company, with the top tiers reserved for data the buyer calls large-scale or highly unique. Below is every published range, every reported deal price we can source, and what moves a number up or down.

$20K to $5MSpan of published program ranges, as published
$10MGoogle reportedly agreed to pay for Spirit Airlines' data (Aug 2026)
~10xEvaluations vs. raw data, practitioners say
~$5KPer code repository in the closure market, per Troveo
What buyers publish

Published payout ranges, program by program

Four programs that buy company data publish payout figures on their own pages. The table quotes them word for word. Each one is a range or a threshold, and each describes what the buyer chooses to advertise, not what a typical company receives.

ProgramPayout as publishedHow the wording readsSource (plain text)
micro1 Enterprise Data Partnership"$100k+ qualified", "$500k+ large-scale", "$1M+ highly unique"; referral page: "$100K-$2M+ for approved data packages"Three tiers tied to the buyer's own labels. The top tier is open-ended and conditional on uniqueness.micro1.ai/data-partnerships; micro1.ai/company-referral
Mode company data"$100K-$5M"One range for company data. Mode says it buys "an agreed copy" and originals stay with the company.data.mode.inc
Grepped"$20K-$5M"The lowest published floor of the four. Open to any vertical.grepped.ai
Miro AdvisoryOperating datasets "$100K-$1M+"; private codebases "$10K-$1M+"Marked as indicative by Miro Advisory. Separate ranges for operating data and code.miroadvisory.com

Last checked: 7 October 2026, from each program's own page. Quotes are verbatim. A published range is not a quote for your company, and no program promises any amount, acceptance or timing.

Two patterns stand out. The floors cluster at $100K, with Grepped at $20K and Miro Advisory's codebase range at $10K, so a smaller dataset may fit better where the floor is lower. And every ceiling, from "$1M+" to $5M, is open-ended, conditional or marked indicative. None tells you where most deals land.

One scale

The published ranges side by side

Each gridline is ten times the one before. Green bars are program ranges as published; a faded end means an open top ("+"). Slate marks are closure-market figures cited by Troveo.

These bars mark the published edges of each range. They are not distributions and say nothing about where most deals land. Sources: each program's own page, last checked 7 October 2026; Troveo as cited in 2026 news coverage.

Reading the ranges

Why "up to" is not an average

A published range is a screening tool. It tells you whether a conversation is worth having. It is not a quote, a forecast or an average, and planning around it as if it were leads to bad decisions.

Edges, not centers

A floor is the smallest deal a program pursues; a ceiling is the largest it advertises. Neither tells you where the middle is.

The top tiers carry conditions

micro1's page attaches "$500k+" to "large-scale" and "$1M+" to "highly unique" data. Those are the buyer's own labels for the exceptions. A 60-person firm with ordinary records is, by that wording, not the exception.

No medians are published

None of the program pages we checked publishes an average, a median, or how many companies were paid at the top. Read any range as "possible", not "probable".

A plus sign is open, not likely

"$2M+" means the buyer will not rule out more. It does not mean more is common.

Budget for zero until it is signed. Until a buyer has reviewed a manifest and samples of your data, any figure you have seen is a published range, not an offer. Programs set eligibility rules, so some applicants are not accepted. Do not book the revenue, hire against it or promise staff a share until you have a signed agreement and an accepted delivery.

Reported deals

Public deal prices, and how far they transfer to you

Deal prices do appear in the news, but almost none describe a company of 20 to 500 people. Here is each figure, who reported it, and how far it transfers.

DealReported priceTypeReported byTransfers to a running company?
Spirit Airlines internal data$10 million (Google agreed to pay); $12.5 million rival offer from micro1One company's internal records, sold in bankruptcy proceedingsABC, TIME and others, August 2026Closest analog, but a large airline's full operating record sold in a court process
Google and Redditabout $60M a yearPlatform licensing, annual, directWidely reportedNo: public content at platform scale
News Corp and OpenAIover $250M over five yearsPublisher licensing, directWidely reportedNo: a news archive and ongoing content
Shutterstock AI licensingabout $104M (2023)Reported AI licensing revenue for 2023 (a later $138M figure for 2024 was a projection for the whole business unit, not AI licensing alone)Reported company figuresNo: a licensing business, not one sale
Shut-down startup archivesabout $5,000 per code repository; roughly $10,000 to $100,000 per archive dealClosure market for Slack, email and codeTroveo, as cited; Forbes 16 April 2026, Fast Company, GizmodoPartly: real company records, but from firms that stopped operating

Last checked: 7 October 2026. Figures as reported in the named coverage. We have not seen any of these contracts. Full list with sources: AI data licensing deals tracker.

Spirit Airlines: one company, eight figures

The emails, Teams messages, spreadsheets and operations files of one company drew competing eight-figure bids, as reported in August 2026. But it was a large airline's entire operating history, offered in a bankruptcy sale whose court approval was not confirmed as of 7 October 2026. Status and outcome: Spirit Airlines case page.

Platform deals: scale, not a benchmark

These are direct, usually annual licenses for very large content libraries. They prove labs pay for data at scale. They are not a yardstick for a 200-person company's Slack history.

The closure market: the low end

Forbes ("AI's New Training Data: Your Old Work Slacks And Emails", 16 April 2026), Fast Company and Gizmodo covered startups selling old archives. Those sellers no longer operate. More: shut-down startups selling Slack and email.

50+ / ~$8.5BCompanies selling data and RL environments to labs, and their revenue (Deedy Das market map, July 2026; estimate)
>$100B a yearLab data spend by 2030 on the current path (Will Depue, July 2026; estimate)

Both are estimates by individuals, not audited figures. They say demand is large. They say nothing about what one company's records are worth.

Price drivers

What moves the number up or down

Buyers do not publish pricing formulas. Their eligibility rules and tier labels, plus what practitioners say, point to the same handful of factors.

Years of history

Mode asks for several years of records the company owns. Longer histories show how work changed and how decisions played out.

Connected systems

A ticket linked to the chat thread, the fix and the customer reply is a full workflow. The same files scattered across drives say much less.

Uniqueness

micro1 reserves "$1M+" for "highly unique" data. Records of work few others do, in a niche process or industry, sit at the top.

Team size

micro1 lists 30+ employees. Mode lists 20+ full-time US office staff, 10+ for accounting firms and 6+ for law firms. More people produce more of the record.

Documentation and tools

micro1 lists mature operations, documented processes and modern software tools. Data in standard systems exports cleanly.

Language and location

micro1 lists primarily English, US prioritized, then other Western markets. Mode names US-based teams as the strongest fit.

Clean rights

Data you own outright, with no client contract that forbids sharing it, needs fewer carve-outs. Every carve-out shrinks the scope.

License terms

Exclusive and non-exclusive licenses can be priced differently, and exclusivity limits selling the same data again. Ask how the price changes with each.

Tier 1

Raw data

The cheapest tier

Exports of messages, documents, tickets and code. Practitioners say this is where prices are lowest, and it is where most sellers start.

Tier 2

Evaluations

Roughly 10x raw

Test sets built on the data, where people who know the work define a correct answer. micro1 also lists "AI performance feedback", human feedback on AI outputs, among what it wants.

Tier 3

Training environments

6 to 8 figures

Full simulations of a workflow that a model can practice in. Practitioners say they need heavy engineering that most companies do not have in-house.

Tier values are what practitioners say, not buyer prices. More on what makes records valuable: what data AI labs want.

Worked example

How scope decisions change what is on the table

Often the bigger swing is not the rate but how much data survives client contracts, personal-data rules and company policy.

Illustrative, not an offer

A fictional 85-person freight brokerage with seven years of records

Every company, number and decision below is invented to show the mechanics. None of it comes from any buyer, and real outcomes can be higher, lower or zero.

SourceFirst inventoryAfter scope reviewWhy it changed
Outlook email7 years, all mailboxes7 years, operations mailboxes onlySome shipper contracts restrict sharing their correspondence
Microsoft TeamsAll channels and direct messagesTeam channels onlyDirect messages excluded by company decision
Exception tickets7 years7 yearsKept: the core record of how problems get solved
SOPs and playbooksAllAllKept: easy to scope and de-identify
HR and payroll filesIncludedRemovedPersonal data with little workflow value
$400,000Fictional first indication on the full inventory
$260,000Fictional revised indication after the scope review

The lessons carry over. Scope before you discuss price, so the number you hear is for data you can deliver. The parts that survive review, here the linked exception tickets and SOPs, are usually what buyers care about most. And a smaller scope means fewer consent and confidentiality warranties to give.

Payment structure

How you are paid matters as much as how much

A headline number can hide very different deals. Put these questions to any buyer before you compare figures. They are general questions for every seller, not statements about any program on this page.

Common pricing mistakes

Six ways sellers misread what their data is worth

Each mistake below comes from reading a number without the conditions attached to it. Each one has a simple fix that costs nothing but time.

Mistake 1

Anchoring on the top of a range

Treating "$2M+" or "$5M" as the likely figure. The top of a published range is the exception a buyer chooses to advertise, and the top tiers carry labels such as "large-scale" and "highly unique".

Instead: plan around the floor until you hold a written offer, and read the label attached to every tier.

Mistake 2

Comparing a one-off price with a recurring one

A single payment for one delivery and a deal with paid refreshes are different products with different work attached. Their headline numbers cannot be lined up directly.

Instead: compare what you receive over the same period, for the same delivery obligations.

Mistake 3

Ignoring acceptance criteria

A price is paid on data the buyer accepts. If the contract lets the buyer reject part of a delivery, the number you signed for and the number you receive can differ.

Instead: ask what is tested, who decides it passed, and how a partial rejection changes the price.

Mistake 4

Pricing before scoping

A figure given before client contracts, personal data and company policy are reviewed describes data you may not be able to deliver. It can move once that review is done.

Instead: run the scope review first, as in the worked example above, and ask for a price on what remains.

Mistake 5

Treating platform deals as benchmarks

The reported Reddit, News Corp and Shutterstock figures describe content licensing at platform scale. Dividing them by headcount or file count produces a number no company buyer has published.

Instead: measure yourself against the published program ranges for companies until you hold an offer.

Mistake 6

Accepting the first number

One offer has nothing to be measured against, and sending a full dataset to get it gives away your leverage. Practitioners advise a manifest, samples and more than one offer.

Instead: send the same manifest to more than one program and compare the terms side by side.

From range to offer

How to get to a real number for your data

The only price that counts is one a buyer puts in writing after reviewing your data.

1

Check fit first

Compare your team size, location, industry and systems with each program's published rules in the eligibility checker before you spend time on anything else.

2

Share a manifest, not the data

A manifest lists systems, date ranges, volumes and exclusions. Practitioners say never send a full dataset before price: send the manifest and samples.

3

Ask more than one program

One offer has nothing to be measured against. See getting more than one offer.

4

Compare terms, not just price

Lay payment structure, exclusivity, scope of use and liability side by side. A lower figure with capped liability can be the better deal.

FAQ

Questions about data prices

How much do AI companies pay for company data?

As published on their own pages (checked 7 October 2026): micro1 lists $100K-$2M+ for approved data packages, with tiers of $100k+ qualified, $500k+ large-scale and $1M+ highly unique. Mode lists $100K-$5M. Grepped lists $20K-$5M. Miro Advisory lists indicative ranges of $100K-$1M+ for operating datasets and $10K-$1M+ for private codebases. These are ranges, not quotes. Nobody can price your data without reviewing it.

Is there an average price for company data?

Not a published one. None of the program pages we checked publishes an average deal size, a median, or how many companies were paid at the top of a range. A published range tells you whether a conversation is worth having. It does not tell you what a typical company receives.

How much did Google offer for Spirit Airlines' data?

As reported in August 2026 by ABC, TIME and others, Google agreed to pay $10 million in Spirit's bankruptcy proceedings for internal data, and micro1 then made a $12.5 million rival offer. Court approval of the sale was not confirmed as of 7 October 2026; our Spirit Airlines case page tracks the outcome.

Why are platform deals like Reddit's so much bigger?

Google-Reddit (about $60M a year), News Corp-OpenAI (over $250M over five years) and Shutterstock's AI licensing revenue (about $104M in 2023) are reported figures for platforms and publishers licensing very large content libraries. They are not benchmarks for one company's internal records.

What do shut-down startups get for Slack and email archives?

Troveo cites about $5,000 per code repository and roughly $10,000 to $100,000 per archive deal in the closure market. These figures describe companies that stopped operating.

What raises the price of a dataset?

Years of history, records that connect across systems, uniqueness, documented processes in standard tools, clear ownership with few carve-outs, and the license terms you agree to. Practitioners also say evaluations built on data are worth roughly 10x raw data, and full training environments reach 6 to 8 figures but need heavy engineering.

Is a published range a quote for my company?

No. A range shows what a program advertises across every company it works with. A figure for your company exists only after a buyer has reviewed a manifest and samples of your data and put an offer in writing. No published range is a promise of any amount, acceptance or timing, and programs apply eligibility rules, so some applicants are not accepted.

Should I accept the first offer I get?

Not without something to compare it to. Practitioners advise sharing a manifest and samples, never the full dataset before price, and getting more than one offer. Compare offers on payment structure, acceptance criteria, exclusivity, scope of use and liability, not only on the headline number.

See which published ranges could apply to you

The checker compares your answers with each program's published rules and shows each range word for word. It asks for no email.

Independent site. Some links are referral links: if your company signs with a buyer through them, the buyer may pay us a fee. You are not charged, and we never see your data.

Applying does not guarantee acceptance or any amount. Each program runs its own review and sets its own terms and timing.

Related reading

Keep reading before you apply