"AI training data companies" covers very different businesses. Some pay people to label and grade data. Some license records that companies already hold. Some build training environments, and some deal in archives and footage. If you find companies and introduce them to buyers, only part of this market can use what you bring. This map shows which part, what each kind of company buys, and which ones publish what they pay referrers.
Choosing between programs? See buyer programs compared.
Search for AI training data companies and you get one long list. Look at what each company actually pays for and the list splits into at least five businesses that share a customer, the AI labs, and little else. A company that pays a radiologist by the hour to grade model answers and a company that licenses ten years of a law firm's internal procedures are both called training data companies. They do not buy the same thing, they do not pay the same people, and only one of them has any use for a broker who finds companies.
The money is concentrated. Deedy Das's market map (July 2026) counts more than 50 companies that sell data and RL environments to labs, with about $8.5B in revenue between them, and 75% of it held by four names: Scale, Surge, Mercor and Handshake. Those are estimates from someone who follows the market, not audited results, and they describe where revenue sits today.
For a broker, fit matters more than concentration. The openings for someone who brings companies sit mostly in the programs that publish rules for company data: micro1, Mode, Grepped and Miro Advisory. Three of them also publish what they pay a referrer. The rest of the map is still worth knowing. It tells you who your companies are not for, who might compete with you for the same introductions, and which buyer a company belongs with when it is not a fit for a program at all.
Sorted by what each one pays for, because that decides whether a broker's companies are of any use to it. The grouping is ours. A single company can sit in more than one category.
The same five categories on one table. Figures are quoted only where a company or a market source published them.
| Category | What it pays for | Who it pays | Published figures | What a broker can bring |
|---|---|---|---|---|
| Human-data and expert-labor providers | Hours of skilled work: labels, ratings, written answers, expert review | Individual workers | Scale, Surge, Mercor and Handshake hold 75% of about $8.5B across 50+ companies (Deedy Das, July 2026) | Little. Their supply is people, and company introductions are not what they recruit. |
| Company-data and operational-data buyers | Licensed copies of records a business already holds | The company that owns the records | micro1 "$100K-$2M+ for approved data packages"; Mode "$100K-$5M"; Grepped "$20K-$5M"; Miro Advisory operating datasets $100K-$1M+, codebases $10K-$1M+ | Companies that meet published rules on size, location, language and years of records |
| RL environment builders | Built environments, tasks and graders | Mainly their own teams | Full environments reach 6 to 8 figures, practitioners say | Context on how real work is done. No program we track publishes rules for a company supplying one. |
| Marketplaces and seller representatives | Listings, archive sales and representation | Varies by deal | Troveo cites about $5,000 per code repository and roughly $10K to $100K per archive deal for shut-down startups | Companies that are closing and no longer fit a program built for running businesses |
| Video and footage licensors | Rights to footage | Rights holders | None in our sources | Companies that own a large footage archive and can show clear rights to it |
Six figures come up whenever this market is discussed. Here they are with their sources, followed by three cautions for anyone planning income around them.
The market map is one informed person's estimate. Treat $8.5B as the order of magnitude of the market, not as a number to split among buyers or to quote to a company.
Concentration shows where money sits today. It does not show where a broker can earn. The company referral terms on this page come from micro1, Mode and Grepped, not from the four names that hold 75%.
Raw, evaluation and environment prices describe what is delivered. Most companies a broker finds license raw records, so the raw tier is the honest reference for them.
Every category on the map sells into one of three tiers. Knowing which tier a company's data enters keeps a broker's expectations, and the company's, close to reality.
Records as they exist, scoped and de-identified: a support team's tickets, a firm's procedures, a repository's commit and review history. Practitioners call raw data the cheapest tier. It is also what almost every company a broker finds can actually offer, which is why company-data programs exist at all.
Test sets with graded answers, built on top of data, that measure whether a model does a task well. Practitioners put evaluations at about ten times the value of raw data. The extra value comes from expert time spent deciding what a correct answer is, so it usually involves people, not only records.
Full simulated workplaces with tools, tasks and graders, where a model practices a job over many steps. Practitioners say full environments reach 6 to 8 figures. They take heavy engineering, and builders usually do that work themselves rather than buying it from an operating company.
Sort the companies you find by what they hold and what state they are in. The six profiles below are invented to show the usual match. Published rules are as published, checked 7 October 2026.
Mode lists accounting firms at 10+ people. micro1 lists 30+ employees, with 30 to 200 in its referral posting, US first and primarily English. Both published rules are met on paper.
The gate is client confidentiality: client books sit under engagement letters, so the firm's own procedures and workflows are the likely scope.
Miro Advisory lists codebases at $10K-$1M+ and operating datasets at $100K-$1M+. micro1's 30+ and Mode's 20+ full-time US office employees are both met.
Open-source code inside the repositories may not be the company's to license, and secrets in old commits must come out first.
Below most headcount rules, but Mode lists law firms at 6+ people, so the published minimum is met. micro1's 30+ is not.
Privilege and client confidentiality remove most matter files. What is left is the firm's own templates, intake steps and administration.
While it still has staff who can scope and export, a company-data program is an option. After it closes, the archive market is the likely route, where Troveo cites about $5,000 per repository and $10K to $100K per archive deal.
Board approval and who still holds admin access come before any introduction.
The asset is footage, not operating records. The first question is rights: whether the company or its clients own each project, and whether people on camera signed releases that cover this use.
No company-data program on this page publishes footage rules.
Specialist study records and SOPs, built over many years. micro1's 30+ is met, though its referral posting names 30 to 200, so confirm fit with the program itself. Mode's 20+ full-time US office rule is met.
Sponsor contracts usually decide who owns study data. Read them before scoping.
This section lists only companies with published referral terms for introducing companies. All three are company-data programs. As published, checked 7 October 2026.
Miro Advisory, Troveo, the environment builders, the footage licensors and the four names that hold 75% of the market map's revenue are not listed above. This page quotes only published referral terms for introducing companies, and we have none on record for them. If one of them publishes terms later, read them on its own site, and treat any figure you see in a post or a forum as unconfirmed until you do.
Full program by program detail, including how each referral moves from introduction to payout, is on data referral programs. This site is itself a referrer for some of these programs; how that works is on our disclosure page.
Start with the company, not the referral amount. A $50K referral is worth nothing on a company that misses the program's published rules, and a $10K referral on a company that fits is worth more than an introduction that stalls.
The map tells you where to sell. The hard part is the list of companies to start from. That is the part our tools cover, at company level only.
The Data Asset Score rates any company from 0 to 100 for the data AI buyers want. Each result gives a grade, the data the company likely holds, its history and its activity status. It is built on our index of 102 million domains, 99.99% of the active internet, with domain history, and it reports nine factor groups: history, scale, knowledge assets, operational systems, customer systems, organization, industry value, expertise and activity status. The free demo is limited per day.
If you would rather not start from a blank page, there are ready lists for 20 US sectors, from labs and CROs to law firms, software companies, logistics and marketing agencies. Every company on them is verified active: the domain resolves, is not expired, is not parked and shows no error or placeholder page. Each carries the same score fields, and each list has a free preview with the top 5 visible; the labs and CROs preview is a good place to see the format. You can buy lists once, from $249 for one to $2,490 for all 20, as a CSV snapshot with no updates. For lists that stay current, Pro returns them through the API as JSON, up to 100 companies per call, with paging, and Scale adds the one-file bulk CSV export and new segments on request.
Two limits matter for brokers. Everything is company level: no contacts, named people, emails or phone numbers. And a score is an estimate from public signals. It is not a valuation, not an offer, and not proof that a company wants to sell. It tells you which companies to look at first.
Each one costs months with a buyer or credibility with a company. All are avoidable once you know the map.
Sending companies to the names that hold 75% of the market's revenue because that is where the money is.
Send companies to programs that publish rules for company data, and check those rules first.
Telling a 25-person firm it could get $5M because a range reaches that high.
Quote the range with its source and date, and plan from the floor. Only a buyer that reviews the data can price it.
Budgeting a referral fee the week a company agrees to talk.
Expect 60 to 90 days to close, then any onboarding and revenue thresholds the program sets.
Promising to share the referral fee to win the introduction.
Don't. micro1's terms forbid sharing payouts with companies. Win the introduction with fit and clear information.
Introducing yourself as a buyer's partner or representative.
Say you are an independent referrer and that the buyer pays you. micro1's terms forbid posing as its partner.
Telling a company its data is worth a figure because it scored well.
A score is an estimate from public signals for ranking targets. It is not a valuation, an offer or proof the company wants to sell.
Companies that supply AI labs with the data models learn from. The label covers several different businesses: providers that pay people to label and grade data, buyers that license records companies already hold, builders of RL environments, marketplaces and seller representatives, and licensors of video and footage. They share a customer, the labs, but they buy different things from different suppliers.
The most cited estimate is Deedy Das’s market map from July 2026. It counts more than 50 companies that sell data and RL environments to labs, with about $8.5B in revenue between them, and 75% of that revenue held by Scale, Surge, Mercor and Handshake. These are estimates from someone who follows the market, not audited figures.
Several publish programs for company data. As published, checked 7 October 2026: micro1’s Enterprise Data Partnership lists “$100K-$2M+ for approved data packages”, Mode lists “$100K-$5M”, Grepped lists “$20K-$5M”, and Miro Advisory lists operating datasets at $100K-$1M+ and codebases at $10K-$1M+. A range is what a buyer will consider, not an offer.
Among the company-data programs, three publish referral terms, as published and checked 7 October 2026. micro1: “earn $50,000” per referred company, “no cap”, paid after onboarding plus a minimum revenue threshold. Mode: “Earn $50K per referral”, described on X as up to $55k or 6%. Grepped: “Refer for another $10K”. None pays for a name alone; payout conditions are set in each program’s own terms.
From a lab’s side, they are all providers. From a broker’s side, the question is what each one buys to make its product. Human-data providers buy people’s time and skill. Company-data programs buy licensed copies of existing business records. A broker who finds companies is useful to the second group, and rarely to the first.
Not with company introductions. Human-data and expert-labor providers recruit individual workers, and a company-data referral program does not cover recruiting people. This page quotes no referral terms for that side of the market because it lists only published terms for introducing companies.
Practitioners say deals take 60 to 90 days to close, and a referral fee depends on the deal, not on the introduction alone. micro1, for example, pays after onboarding plus a minimum revenue threshold, at its sole discretion and subject to clawbacks. Plan cash flow around months, not weeks.
Start with what a company holds and whether it is still operating. The Data Asset Score rates any company from 0 to 100 for the data AI buyers want and shows its activity status. The free company check compares a company’s public pages with the published rules of micro1, Mode and Grepped in about 20 seconds. Everything is company level: no contacts or named people.
Score a target for the data AI buyers want, check it against the published rules of micro1, Mode and Grepped, or start from one of the 20 scored US sector lists, such as labs, CROs and life-science companies.