The same company records can be sold three ways: as exported files, as graded tests for AI models, or as a working simulation an AI agent can train in. Each step up is worth more and costs more work. Here is what each tier is, and when it makes sense for a company to go past raw data.
Last checked: 7 October 2026. Tier values are practitioner rules of thumb, attributed as such.
Definitions first. The value notes are what practitioners say, not published prices.
Your records as exported and de-identified: Slack threads, emails, support tickets, CRM notes, SOPs, code with its history.
The buyer does all the work of turning it into something a model can learn from or be tested on.
Tasks built from your work, each with a known good outcome or an expert grading guide. Labs use them to measure whether a model can do the job.
The value is the expert judgment: what "correct" looks like in your field.
A working simulation of a job, such as a sandboxed ticket queue or ledger, where an AI agent takes actions and is scored automatically, thousands of times.
It needs heavy engineering. Few companies outside the data industry build one.
| Raw data | Evaluations | RL environments | |
|---|---|---|---|
| What you deliver | A scoped, de-identified copy of records | Tasks, reference answers and grading rubrics | Workflows and expertise for a simulation someone else usually builds |
| Who does the work | Mostly the buyer | Your experts, often with the buyer's tooling | Engineers, typically at the buyer or a specialist firm |
| Value, as practitioners describe it | Cheapest tier | Roughly 10x raw | Six to eight figures for full environments |
| Time you commit | Weeks of review and export | Ongoing expert hours | Months, for whoever builds it |
| Main risk | Privacy in the records themselves | Expert time diverted from paying work; client confidentiality in task design | Your processes encoded in detail; IP and scope questions |
| Contract point to check | Scope of use and deletion | Whether "evaluation" use is inside the license scope; how expert time is paid | Ownership of what is built, exclusivity, resale |
Value tiers: what practitioners say, not published prices. Contract points: see our AI data licensing agreement outline.
Each step up adds something a lab cannot get from more raw volume.
Raw records show what people did. An evaluation says whether it was right, and why. That judgment is scarce, and it is what labs need to measure progress on real work.
A raw dataset is consumed in training. A good evaluation set is run again and again against new models. An environment can be used for training runs indefinitely.
micro1's data partnership page lists "decision-making patterns" and "AI performance feedback" (human feedback on AI outputs) among what it wants. Both sit between raw data and evaluations.
Market estimates point the same way. Deedy Das's market map from July 2026 counts more than 50 companies selling data and RL environments to labs, with about $8.5 billion in revenue. Will Depue said in July 2026 that labs are on a path to more than $100 billion a year of data spend by 2030. Both are estimates, and neither tells you what your records are worth. Our guide to what data AI labs want covers which raw records are valued in the first place, and how much AI companies pay covers the published ranges.
Most companies should start with a clean raw-data license. Some have a real case for more.
If evaluations interest you, raise it as a question, not a demand. Ask the buyer whether it would pay differently for graded examples, who does the grading and how that time is paid, and whether evaluation use changes the scope of the license. Grepped's site says it also pays individual professionals for expertise, which shows that expert time is something buyers price separately. Settle in writing who owns the tasks and rubrics your people create.
Environments are a different decision. A company of 20 to 1,000 people rarely has the engineering capacity to build one, and the six-to-eight-figure values practitioners mention are for full environments, not for the records that feed them. What you can supply is the workflow knowledge a builder needs. Treat any such request as its own contract, with its own price.
What an evaluation project involves, in outline. This is a generic sequence, not any buyer's published process. First, pick one workflow with checkable outcomes, such as resolving a class of support tickets or completing a month-end close. Second, write a small pilot set of tasks from patterns in your records, each with a reference answer and a grading guide written by a senior person. Third, have two reviewers grade the same pilot attempts independently; if they disagree often, the rubric needs work before anyone scales it. Fourth, agree the volume, the weekly hours and the acceptance criteria in writing. Only then scale up.
Two numbers decide whether it was worth it: the expert hours spent and the price agreed for the finished set. Track both from the pilot onward. If the pilot shows that grading takes far longer than expected, renegotiate or stop before the full project begins. A clean raw-data license is still a good outcome on its own; the higher tiers are an option, not an obligation.
Mode publishes a lower minimum for accounting firms: 10+ employees. See accounting firms for the client-confidentiality limits.
A rough guide based on capacity, not on what any buyer has promised. Minimums quoted from each program's page, checked 7 October 2026.
You meet Mode's published 20+ and micro1's 30+ minimums. At this size the senior people who could grade evaluations are also the people running the business, so expert hours are scarce. Focus on a clean raw-data license: tight scope, solid exclusions, clear deletion terms. If a buyer later asks for graded examples, treat it as a small paid add-on with a fixed number of hours.
You may have a team or two with checkable outcomes and a few experienced reviewers: a support team with resolution standards, a finance team with month-end closes, an engineering team with code review. That is the profile where an evaluation project can make sense. Start with raw data and raise evaluations once a buyer has seen your manifest.
You can assign experts part time without stopping the business, and you have the governance to manage a longer project. Evaluations become a realistic second step. Building an RL environment still means heavy engineering, usually done by the buyer or a specialist firm; your contribution is process knowledge and edge cases, under its own contract.
The ten-times figure is a rule of thumb from practitioners. It gets misused.
The tier rule describes the relative value of well-made evaluations, not a markup you can apply to any offer. Poorly designed or poorly graded tasks are worth little.
Evaluation work is paid in the hours of your most experienced people. Count those hours at what they would otherwise earn the company before you agree a price.
A task rebuilt from a real engagement can carry client details even after names are removed. Design tasks from patterns, not from identifiable matters, and check client contracts.
Tasks, reference answers and rubrics your people create are new work. The contract should say who owns them, whether you can reuse them, and whether the buyer can resell them.
Scope of use matters. If the license says training only and the buyer later uses the material for evaluation, or the reverse, the scope should be updated in writing.
Six-to-eight-figure environment values come with heavy engineering. Do not commit to building one because the number is attractive. Agree on a scoped contribution instead.
Questions every seller should check. They are generic and are not claims about any named buyer.
The answers tell you whether a tier-two project is a real opportunity or a distraction. A buyer with clear terms for expert time, ownership and scope is describing a project you can plan. Vague answers on any of those points mean the extra value may never reach you.
Ask the same questions of more than one buyer. Programs differ in what they emphasize, and the only way to know whether a buyer values your expertise, not just your records, is to compare how each one answers. Grepped's site, for example, says it also pays individual professionals for expertise, while micro1's page lists AI performance feedback among what it wants. Read each program's own page before the conversation.
Keep the tiers separate in the contract, too. A raw-data license and an evaluation project have different deliverables, different acceptance criteria and different risks. Putting them in one agreement is fine; letting one set of terms silently cover both is not.
Raw data is your records as exported: messages, tickets, documents, code. An evaluation is a set of tasks built from that data with known good answers or expert grading, used to measure AI models. An RL environment is a working simulation of a job where an AI agent can take actions and be scored automatically, again and again.
Practitioners say evaluations built on the data are worth roughly 10 times the raw data. That is a rule of thumb from people in the market, not a published price, and it assumes the evaluations are well made and graded by real experts.
Practitioners say full training environments can reach six to eight figures, but they need heavy engineering. Most companies with 20 to 1,000 employees are not set up to build one; at most they supply the workflows and expert judgment that a builder turns into an environment.
Consider it if you have experts who can grade work, processes with clear right and wrong outcomes, and time to spare. If not, a raw-data license with a clean scope is simpler. Either way, ask the buyer how it would pay for expert time and whether evaluation use changes the license scope.
Deedy Das's market map from July 2026 counts more than 50 companies selling data and RL environments to AI labs, with about $8.5 billion in revenue. Will Depue said in July 2026 that labs are on a path to more than $100 billion a year of data spend by 2030. Both are estimates.
It is one of the things micro1's data partnership page says it wants: human feedback on AI outputs. In practice that means people with real expertise reviewing what an AI system produced for a task in their field and saying whether, and why, it was right. It sits between raw data and full evaluations.
Usually not to write the tasks and grade them; that is expert work in your field, often done with the buyer's tools. Engineers matter for the third tier, RL environments, which practitioners describe as needing heavy engineering. Most companies supply expertise for those rather than building them.
Whatever the contract says, which is why it must say it. Tasks, reference answers and rubrics are new work created by your people. Agree in writing who owns them, whether you can reuse them, and whether the buyer can resell or share them with other labs.
Check your fit against each program's published rules, then ask about evaluations once a buyer has seen your manifest.
Independent site. Some links are referral links: if your company signs with a buyer through them, the buyer may pay us a fee. You are not charged, and we never see your data.