Last checked: 7 October 2026. For companies, not individuals.
Experiment logs, protocols, deviation reports and yield investigations capture expert judgment that is rare in public data. They also sit closest to your patents, your trade secrets and, in some cases, export-control rules. This page covers both halves.
General AI systems have read the published literature. What they rarely see is the working record of a lab: what was tried, what failed, and why the next run changed.
Practitioners say the next wave of demand is vertical data: manufacturing, lab and chemistry records among the examples they name. The reasoning is simple. Generic office workflows are increasingly common in training sets. Records from a formulation lab, an analytical testing group or a process engineering team are not, because they rarely leave the company that produced them.
micro1 frames its top published tier, “$1M+ highly unique”, as “Highly unique, proprietary operational data with significant value for frontier AI development” (as published, checked 7 October 2026). That describes a category, not a promise. Lab data can be unique and still be impossible to sell because of who owns it, what it reveals, or which laws cover it. Most of this page is about telling those cases apart.
Hypothesis, setup, observations and the decision about the next run. The sequence matters more than any single entry.
Negative results almost never get published. A record of why something did not work, and what changed, is the rare part.
Method development notes, validated procedures and the revisions between versions, if your company owns them outright.
Out-of-spec investigations, root-cause write-ups and corrective actions show structured reasoning under pressure.
Recipe changes and yield investigations are often the most sensitive records a fab or chemical plant holds. Expect most to stay out.
Ticket histories and internal threads on fixing instruments. Usually lower risk, and easier to de-identify.
Practitioners describe three tiers. Labs have an unusual position on the second one, because the expertise needed to build it sits with your staff.
Practitioners say this is the cheapest tier. A scoped, de-identified export of documents, logs and tickets. Lowest effort for you, and the most common deal type.
Practitioners say evaluations are worth roughly ten times raw data. They are tests of whether an AI system reaches the answer an expert would. Writing and checking them needs people who know the chemistry or the process. Grepped publishes that it also pays individual professionals for their expertise.
Practitioners say these can reach six to eight figures, but they need heavy engineering. Few companies of 30 to 500 people take this on without a technical partner.
In most industries, personal data is the main exclusion. In R&D, ownership, export rules and IP come first.
None of the programs we checked publishes a separate rule for R&D organizations. The general minimums apply.
| Program | Published company payout | Published eligibility | Note for labs |
|---|---|---|---|
| micro1 Enterprise Data Partnership | “$100k+ qualified”, “$500k+ large-scale”, “$1M+ highly unique” | 30+ employees, mature operations, documented processes, modern software tools, primarily English; U.S. prioritized, then other Western markets | Lists QA processes, SOPs and project histories among wanted material. |
| Mode company data | “$100K to $5M” | 20+ full-time U.S. office employees; several years of records the company owns | Ownership of the records is a stated condition. Client-owned results do not qualify. |
| Grepped | “$20K to $5M” | Any vertical; also pays individual professionals for expertise | Relevant if your scientists could help build evaluations. |
Last checked 7 October 2026. Sources: each program’s own website (micro1 data partnerships page, data.mode.inc, grepped.ai). Figures are published ranges, not offers. Miro Advisory publishes indicative ranges for operating datasets ($100K to $1M+) and is not covered by the buttons below. Companies with large or unique datasets can also approach labs directly: Google runs an intake for data offers and OpenAI has a data partnerships page. We earn nothing on direct deals; see how to sell data to AI labs.
Independent site. Some links are referral links: if your company signs with a buyer through them, the buyer may pay us a fee. You are not charged, and we never see your data.
Enter your headcount, systems and years of records in the eligibility checker to see which published minimums you meet. The buyer programs comparison lists each program’s process and privacy statements in full. Apply only after your export-control and IP review is done, so the scope you describe is one you can actually deliver.
A buyer’s de-identification removes names. It does not remove a controlled process parameter or a trade secret. Those questions are yours to answer before anything is shared, including samples.
General information, not legal advice. Talk to your own lawyer before you sign.
The point of the exercise is the exclusion list. It shows a buyer what you can deliver and protects what you cannot.
Size decides which published rule you meet. Ownership and export status decide how much of your archive can be offered at all. These walkthroughs show both, with no prices.
An environmental and materials testing lab. With 25 full-time staff, most of them at the bench, it is below micro1’s published 30+ minimum. Whether it meets Mode’s 20+ full-time U.S. office employee line depends on how lab staff are counted, so it asks. Nearly every result belongs to a client. What survives review is its own material: method validation SOPs, instrument maintenance logs and internal QC procedures.
A contract research lab running synthesis and assay work for sponsors. At 60 employees it clears both published size lines. Sponsor agreements give most study data and inventions to the sponsor, so the study files stay out. Its internal deviation-handling records, training programs and troubleshooting threads may stay in, but only after its sponsor agreements are read for clauses covering internal know-how.
A maker of deposition materials for chip fabs, with sites in two countries. Size is not the issue. Export classification is. Process recipes, tool settings and yield investigations tied to customer fabs stay out until counsel has classified them. Cleared material might include quality system procedures, supplier qualification workflows and older, published-equivalent methods. A dataset this sensitive is also one where going direct to an AI lab could be weighed.
List every record type before you contact anyone. The table is a starting point for your own inventory, not a ruling on any specific record.
| Record type | Typical system | Usually includable? | Why |
|---|---|---|---|
| Lab SOPs and revision history | SharePoint, document control system | Often, after IP review | Written by you; revisions show how methods improved over time. |
| Instrument troubleshooting tickets | Jira, ServiceNow, email | Often | Diagnosis-to-fix chains with little personal data, once vendor and client names are removed. |
| Training material and onboarding guides | SharePoint, Confluence, Notion | Usually yes | Your own teaching material about how the work is done. |
| ELN entries for internal programs | Electronic lab notebook | Only after IP and export review | The richest reasoning, and the most likely to hold trade secrets or controlled technical data. |
| Deviation, CAPA and QC records | LIMS, quality system | Possibly | Show judgment under rules; may sit under retention duties or customer audit terms. |
| Client or sponsor study results | LIMS, ELN, reports | No | Owned by the client or sponsor under contract. |
| Process recipes and tool settings | MES, recipe management, spreadsheets | Rarely | Core trade secrets, and in semiconductors a common export-control concern. |
| Unfiled invention disclosures | IP docket, email | No | Disclosure could affect patent rights. A question for patent counsel. |
| Grant-funded project data | Any | Only if the grant terms allow | Funding agreements can set their own data and publication rules. |
| Staff exposure, medical and HR records | EHS and HR systems | No | Personal and health data about employees. |
A lab that cannot answer these yet is not ready to apply. That is normal. Most of them need counsel, not a buyer.
A “small sample” of controlled technical data is still controlled. Classify first, then share the cleared sample under NDA.
Removing names protects people. It does nothing for a trade secret or an export classification, which live in the technical content itself.
Programs funded by others often carry data and publication rules. Old funding agreements are easy to overlook in a long archive.
Ask the buyer which recipients get the copy, and in which countries. Whether foreign-national staff could access it is a deemed-export question for your counsel.
Current development work is where competitors would gain most. Discontinued or already-patented programs are the safer starting point.
Spectra, images and spreadsheets attached to tickets and chats can carry client names and recipe values. Review attachments, not just text.
Use these with any buyer. They are general checks, not statements about any named program.
Nobody can say without reviewing it. Practitioners say vertical data such as lab and chemistry records is the next wave, and micro1 publishes a top tier of "$1M+ highly unique" for proprietary operational data (as published, checked 7 October 2026). That is a published ceiling, not an average. Your price depends on what you can legally include, how connected the records are, and what buyers offer after review.
Treat it as out of scope until counsel has classified it. Technical data covered by the EAR or ITAR, including some semiconductor and chemical process information, can need a license before it is shared, and sharing with foreign persons inside the United States can count as an export. Ask an export-control lawyer before any sample leaves your systems.
It might. Disclosing an unpublished invention or a process you protect as a trade secret can weaken those positions, even under a contract. This is a question for your IP counsel before you build the scope, not after. Many labs keep active programs and unfiled work out entirely.
Not necessarily. The published source lists name common business tools, and Mode's published list of sources ends with "+1K more". Ask the buyer directly whether it can take exports from your electronic lab notebook or LIMS, and in what format, before you spend time preparing them.
Some large or unique datasets go direct. Google runs an intake for data offers and OpenAI has a data partnerships page. Direct deals usually involve procurement, NDAs and longer reviews, and we earn nothing from them. Most companies of 30 to 500 people start with a data company because it handles export and de-identification.
Usually only its own material. Results produced for clients or sponsors normally belong to them under contract, so they stay out. What may remain is the lab's own know-how: SOPs, method validation records, instrument troubleshooting and training material, after IP and export review and a reading of each client agreement.
It helps with personal data, not with the bigger lab risks. Trade secrets and export-controlled technical data sit in the technical content itself, so stripping names or client identifiers does not change their status. Classification and IP review have to happen before any sample leaves. This is general information, not legal advice.
Expect months. Mode publishes that it generally expects about three months from the first conversation through payment, and practitioners cite 60 to 90 days to close. Export-control and IP review usually add time at your end, and that review should not be rushed.
The checker runs entirely in your browser and stores nothing. It shows each program’s published minimum next to your numbers, plus a note on confidential and regulated data.