Sell Data to AI
Home Data Asset Score Pricing API documentation US labs and CROs list
For brokers
How to become an AI data broker Data broker business model Buyer programs compared Qualify a company AI training data companies
For data companies
Firmographic data providers Company data API
Seller guides
How to sell data to AI companies Is it legal? FAQ and glossary About
Check domain/company
Selling well · Value tiers

Raw Data vs Evaluations vs RL Environments: The Value Tiers Explained

The same company records can be sold three ways: as exported files, as graded tests for AI models, or as a working simulation an AI agent can train in. Each step up is worth more and costs more work. Here is what each tier is, and when it makes sense for a company to go past raw data.

Last checked: 7 October 2026. Tier values are practitioner rules of thumb, attributed as such.

Tier 1Raw data: the cheapest tier, practitioners say
~10xEvaluations versus raw, practitioners say
6 to 8 figuresFull RL environments, with heavy engineering
~$8.5BRevenue of 50+ data and environment sellers (estimate, July 2026)
The three tiers

What each tier is, in plain language

Definitions first. The value notes are what practitioners say, not published prices.

Tier 1

Raw data

The cheapest tier

Your records as exported and de-identified: Slack threads, emails, support tickets, CRM notes, SOPs, code with its history.

The buyer does all the work of turning it into something a model can learn from or be tested on.

Your effort: low
Tier 2

Evaluations

Roughly 10x raw, practitioners say

Tasks built from your work, each with a known good outcome or an expert grading guide. Labs use them to measure whether a model can do the job.

The value is the expert judgment: what "correct" looks like in your field.

Your effort: medium, expert hours
Tier 3

RL environments

Six to eight figures, practitioners say

A working simulation of a job, such as a sandboxed ticket queue or ledger, where an AI agent takes actions and is scored automatically, thousands of times.

It needs heavy engineering. Few companies outside the data industry build one.

Your effort: high, engineering team
Side by side

The tiers compared on what you deliver and what you risk

 Raw dataEvaluationsRL environments
What you deliverA scoped, de-identified copy of recordsTasks, reference answers and grading rubricsWorkflows and expertise for a simulation someone else usually builds
Who does the workMostly the buyerYour experts, often with the buyer's toolingEngineers, typically at the buyer or a specialist firm
Value, as practitioners describe itCheapest tierRoughly 10x rawSix to eight figures for full environments
Time you commitWeeks of review and exportOngoing expert hoursMonths, for whoever builds it
Main riskPrivacy in the records themselvesExpert time diverted from paying work; client confidentiality in task designYour processes encoded in detail; IP and scope questions
Contract point to checkScope of use and deletionWhether "evaluation" use is inside the license scope; how expert time is paidOwnership of what is built, exclusivity, resale

Value tiers: what practitioners say, not published prices. Contract points: see our AI data licensing agreement outline.

Why value rises

What makes a tier worth more

Each step up adds something a lab cannot get from more raw volume.

Judgment, written down

Raw records show what people did. An evaluation says whether it was right, and why. That judgment is scarce, and it is what labs need to measure progress on real work.

Reuse

A raw dataset is consumed in training. A good evaluation set is run again and again against new models. An environment can be used for training runs indefinitely.

Closer to what buyers list

micro1's data partnership page lists "decision-making patterns" and "AI performance feedback" (human feedback on AI outputs) among what it wants. Both sit between raw data and evaluations.

Market estimates point the same way. Deedy Das's market map from July 2026 counts more than 50 companies selling data and RL environments to labs, with about $8.5 billion in revenue. Will Depue said in July 2026 that labs are on a path to more than $100 billion a year of data spend by 2030. Both are estimates, and neither tells you what your records are worth. Our guide to what data AI labs want covers which raw records are valued in the first place, and how much AI companies pay covers the published ranges.

Should you go past raw?

When a company should consider helping build evaluations

Most companies should start with a clean raw-data license. Some have a real case for more.

Signs it could fit

  • Your work has outcomes that can be checked: a reconciliation balances, a ticket is resolved, a code review catches the bug
  • You have senior people who can grade work consistently and can spare hours
  • Your field is specialized enough that general reviewers could not judge it
  • A buyer has asked about graded examples or feedback on AI outputs
  • Tasks can be designed without client-confidential details

Signs to stay with raw data

  • Expert hours would come straight out of billable or customer work
  • Your records are mostly generic office work with no clear right answer
  • Client or patient confidentiality runs through every task
  • Nobody internally can own a multi-month project
  • The license scope covers training only, and the buyer has not offered terms for evaluation use

If evaluations interest you, raise it as a question, not a demand. Ask the buyer whether it would pay differently for graded examples, who does the grading and how that time is paid, and whether evaluation use changes the scope of the license. Grepped's site says it also pays individual professionals for expertise, which shows that expert time is something buyers price separately. Settle in writing who owns the tasks and rubrics your people create.

Environments are a different decision. A company of 20 to 1,000 people rarely has the engineering capacity to build one, and the six-to-eight-figure values practitioners mention are for full environments, not for the records that feed them. What you can supply is the workflow knowledge a builder needs. Treat any such request as its own contract, with its own price.

What an evaluation project involves, in outline. This is a generic sequence, not any buyer's published process. First, pick one workflow with checkable outcomes, such as resolving a class of support tickets or completing a month-end close. Second, write a small pilot set of tasks from patterns in your records, each with a reference answer and a grading guide written by a senior person. Third, have two reviewers grade the same pilot attempts independently; if they disagree often, the rubric needs work before anyone scales it. Fourth, agree the volume, the weekly hours and the acceptance criteria in writing. Only then scale up.

Two numbers decide whether it was worth it: the expert hours spent and the price agreed for the finished set. Track both from the pilot onward. If the pilot shows that grading takes far longer than expected, renegotiate or stop before the full project begins. A clean raw-data license is still a good outcome on its own; the higher tiers are an option, not an obligation.

Illustrative, fictional, not an offer

One accounting firm's month-end closes, at three tiers

Raw dataSix years of de-identified month-end close records for client engagements where contracts allow it: reconciliations, review notes, approval threads. The firm exports; the buyer interprets.
EvaluationsTwo hundred close tasks rebuilt from those records, each with the correct adjustments and a senior reviewer's grading rubric. Partners spend set hours per week grading AI attempts against the rubric.
EnvironmentA sandboxed ledger with synthetic transactions where an agent performs a full close and is scored automatically. Built by an engineering team; the firm contributes process maps and edge cases.

Mode publishes a lower minimum for accounting firms: 10+ employees. See accounting firms for the client-confidentiality limits.

By company size

Which tier fits a company of 50, 200 or 1,000 people

A rough guide based on capacity, not on what any buyer has promised. Minimums quoted from each program's page, checked 7 October 2026.

50 people: raw data, done well

You meet Mode's published 20+ and micro1's 30+ minimums. At this size the senior people who could grade evaluations are also the people running the business, so expert hours are scarce. Focus on a clean raw-data license: tight scope, solid exclusions, clear deletion terms. If a buyer later asks for graded examples, treat it as a small paid add-on with a fixed number of hours.

200 people: raw data, with evaluations as an option

You may have a team or two with checkable outcomes and a few experienced reviewers: a support team with resolution standards, a finance team with month-end closes, an engineering team with code review. That is the profile where an evaluation project can make sense. Start with raw data and raise evaluations once a buyer has seen your manifest.

1,000 people: evaluations are realistic; environments rarely

You can assign experts part time without stopping the business, and you have the governance to manage a longer project. Evaluations become a realistic second step. Building an RL environment still means heavy engineering, usually done by the buyer or a specialist firm; your contribution is process knowledge and edge cases, under its own contract.

Common mistakes

Six mistakes companies make when they hear about value tiers

The ten-times figure is a rule of thumb from practitioners. It gets misused.

Multiplying a raw offer by ten

The tier rule describes the relative value of well-made evaluations, not a markup you can apply to any offer. Poorly designed or poorly graded tasks are worth little.

Underpricing expert time

Evaluation work is paid in the hours of your most experienced people. Count those hours at what they would otherwise earn the company before you agree a price.

Building client secrets into tasks

A task rebuilt from a real engagement can carry client details even after names are removed. Design tasks from patterns, not from identifiable matters, and check client contracts.

Leaving ownership unwritten

Tasks, reference answers and rubrics your people create are new work. The contract should say who owns them, whether you can reuse them, and whether the buyer can resell them.

Licensing for training, delivering for evaluation

Scope of use matters. If the license says training only and the buyer later uses the material for evaluation, or the reverse, the scope should be updated in writing.

Promising an environment

Six-to-eight-figure environment values come with heavy engineering. Do not commit to building one because the number is attractive. Agree on a scoped contribution instead.

Questions to ask

What to ask a buyer before moving past raw data

Questions every seller should check. They are generic and are not claims about any named buyer.

Money and effort

  • Would you pay differently for graded examples or feedback on AI outputs than for raw records?
  • How is expert time paid: per task, per hour, or as part of the license?
  • How many hours a week would you need, and for how long?
  • What tooling do you provide, and who trains our reviewers on it?
  • What are the acceptance criteria for a finished task set?

Rights and risk

  • Who owns the tasks, answers and rubrics we create?
  • Does evaluation use change the license scope, exclusivity or price?
  • Can the material be resold or shared with other labs?
  • How are client and employee details kept out of task design?
  • What happens to the material if the contract ends?

The answers tell you whether a tier-two project is a real opportunity or a distraction. A buyer with clear terms for expert time, ownership and scope is describing a project you can plan. Vague answers on any of those points mean the extra value may never reach you.

Ask the same questions of more than one buyer. Programs differ in what they emphasize, and the only way to know whether a buyer values your expertise, not just your records, is to compare how each one answers. Grepped's site, for example, says it also pays individual professionals for expertise, while micro1's page lists AI performance feedback among what it wants. Read each program's own page before the conversation.

Keep the tiers separate in the contract, too. A raw-data license and an evaluation project have different deliverables, different acceptance criteria and different risks. Putting them in one agreement is fine; letting one set of terms silently cover both is not.

FAQ

Raw data, evaluations and environments: questions

What is the difference between raw data, evaluations and RL environments?

Raw data is your records as exported: messages, tickets, documents, code. An evaluation is a set of tasks built from that data with known good answers or expert grading, used to measure AI models. An RL environment is a working simulation of a job where an AI agent can take actions and be scored automatically, again and again.

How much more are evaluations worth than raw data?

Practitioners say evaluations built on the data are worth roughly 10 times the raw data. That is a rule of thumb from people in the market, not a published price, and it assumes the evaluations are well made and graded by real experts.

What are RL environments worth?

Practitioners say full training environments can reach six to eight figures, but they need heavy engineering. Most companies with 20 to 1,000 employees are not set up to build one; at most they supply the workflows and expert judgment that a builder turns into an environment.

Should my company build evaluations instead of selling raw data?

Consider it if you have experts who can grade work, processes with clear right and wrong outcomes, and time to spare. If not, a raw-data license with a clean scope is simpler. Either way, ask the buyer how it would pay for expert time and whether evaluation use changes the license scope.

How big is the market for data and RL environments?

Deedy Das's market map from July 2026 counts more than 50 companies selling data and RL environments to AI labs, with about $8.5 billion in revenue. Will Depue said in July 2026 that labs are on a path to more than $100 billion a year of data spend by 2030. Both are estimates.

What is "AI performance feedback"?

It is one of the things micro1's data partnership page says it wants: human feedback on AI outputs. In practice that means people with real expertise reviewing what an AI system produced for a task in their field and saying whether, and why, it was right. It sits between raw data and full evaluations.

Do I need engineers to sell evaluations?

Usually not to write the tasks and grade them; that is expert work in your field, often done with the buyer's tools. Engineers matter for the third tier, RL environments, which practitioners describe as needing heavy engineering. Most companies supply expertise for those rather than building them.

Who owns evaluation tasks my employees write?

Whatever the contract says, which is why it must say it. Tasks, reference answers and rubrics are new work created by your people. Agree in writing who owns them, whether you can reuse them, and whether the buyer can resell or share them with other labs.

Start with the tier most companies sell: raw data

Check your fit against each program's published rules, then ask about evaluations once a buyer has seen your manifest.

Independent site. Some links are referral links: if your company signs with a buyer through them, the buyer may pay us a fee. You are not charged, and we never see your data.