Sell Data to AI
Home Data Asset Score Pricing API documentation 20 US company lists
For brokers
How to become an AI data broker Data broker business model Buyer programs compared Qualify a company AI training data companies
For data companies
Firmographic data providers Company data API
Seller guides
How to sell data to AI companies Is it legal? FAQ and glossary About
Check domain/company
For brokers and sourcing teams

Private company data providers: profiling the firms AI data buyers want

Last checked: 8 October 2026

AI data buyers publish a clear target: private companies of roughly 30 to 200 people, US first, with several years of records. Those are the companies general data providers know least about. This guide sorts private company data sources by type, shows where each one falls short for AI data sourcing, and explains how to fill the gap at company level, without buying anyone's contact details.

30 to 200employees, micro1's referral posting, US first
20+full-time US office staff for Mode; accounting 10+, law 6+
Several yearsof records the company owns, as Mode publishes
Company levelno contacts, no names, no emails

Scoring many companies at once? See the API plans or the company data API guide.

The problem

Why private small and mid-size companies are hard to profile

A listed company files annual reports, discloses its headcount and explains its business in detail. A 70-person private firm owes the public almost none of that. Every provider of private company data is working around the same silence.

No duty to disclose

Private companies in the US generally file formation documents and periodic reports with their state, which record legal facts: the entity, its status, a registered agent. They do not record how many people work there or what the business runs on.

Headcount is estimated

Employee counts for private firms are usually ranges built from indirect evidence. A firm shown as "51 to 200" could sit at either end, which matters when a buyer's rule is 30 or more employees.

Revenue is private

Revenue for a private company is almost always a model. That is less of a problem for AI data sourcing than for sales, because buyers screen on people and records, not turnover.

Systems are invisible

The help desk, CRM, ticketing and document systems a company runs inside the business rarely appear anywhere official. Yet those systems hold the records buyers pay for.

Names and owners change

Small firms rebrand, merge, sell to a competitor or move under a holding company. Matching all of those changes back to one company is hard, and some records never catch up.

Closures are quiet

A small company that stops trading rarely announces it. The website may stay up, then lapse, then point to a parking page. A profile can look current for a long time after the business has gone.

The target

They are exactly the companies AI data buyers want

The published rules of the main company data programs describe a mid-size private firm almost word for word: a few dozen to a couple of hundred people, based in the US, with years of records in its own systems.

"$100K-$2M+"micro1's published company payout "for approved data packages"
"$100K-$5M"Mode's published range for company data
"$20K-$5M"Grepped's published range, any vertical
"Earn $50,000"micro1's published referral reward, "no cap"
ProgramSize ruleLocation and languageRecordsReferral, as published
micro130+ employees; the referral posting says 30 to 200US first, primarily EnglishApproved data packages"earn $50,000", "no cap", paid after onboarding plus a minimum revenue threshold
Mode20+ full-time US office employees; accounting firms 10+; law firms 6+US office employeesSeveral years of records"Earn $50K per referral"
GreppedNo published minimumNo published ruleAny vertical"Refer for another $10K"
Miro AdvisoryNot listed hereNot listed hereOperating datasets $100K-$1M+, codebases $10K-$1M+Not listed here

As published, checked 7 October 2026. Ranges are what programs publish, not offers, and the top of a range is not the typical result. micro1's referral terms also state sole discretion and clawbacks, and forbid sharing payouts with the company or posing as micro1's partner. Full details: buyer programs compared.

Why the middle of the market

The rules above are not arbitrary. A company of 30 to 200 people has enough staff to produce a large, connected body of work: tickets that link to customer accounts, projects that run from proposal to delivery, procedures that were written down because new hires needed them. A company with five people rarely has that depth. A company with fifty thousand usually has a longer approval path, more contracts to check and more lawyers in the room.

Several years of records matter for the same reason. Buyers want to see how work played out over time: a client across several renewals, a product through several releases, an issue from first report to fix. A firm that migrated systems last year may have a long history as a business but a short one in its records.

The US first rule also has a privacy side: an EU or UK team brings GDPR into scope for every record that mentions a person. None of this rules out other companies. It sets the order in which buyers look, and therefore the order in which a broker should search.

The broker's problem in one sentence

The companies that fit the published rules best are the ones that are least visible in the sources most data providers start from. That is the gap this page is about.

Source types

Where private company data comes from

Every private company data provider draws on some mix of the same source types. Knowing which type a field came from tells you how much to trust it, and what it can never tell you.

Source typeWhat it tells youWhere it falls short for AI data sourcing
Business registriesThat an entity exists, its legal name, formation date, state of formation, filing status and registered agentSays nothing about size, systems or records. A company can be in good standing and dormant. Each state keeps its own registry with its own fields.
Regulatory filings and licensesProfessional and trade licenses, permits, industry registrations, government contracting records, secured lending filingsStrong for regulated trades and contractors, thin for everyone else. Confirms a license, not how many people use it or what they keep.
Company websitesWhat the company says about itself: what it does, where it is, when it was founded, open roles, help pages, client loginsUnstructured and uneven. Many sites never state headcount. Reading them one by one is slow, and a site can outlive the business behind it.
Directories and listingsTrade association member lists, industry directories, map and review listings, chamber of commerce rollsOften self-reported and rarely updated. Good for finding names in a niche, weak for size and activity.
Panels and surveysSelf-reported answers from a sample of businesses, sometimes with detail no other source hasCovers only the businesses that took part, and only as of the survey date. Hard to use for a specific target.
Aggregated databasesA combined profile built from several of the above, matched to one company, with modeled fields filled inBroad, but the modeled fields are estimates, and the profile is built for sales: industry, size band, revenue, contacts. Data holdings and activity are rarely explicit.

Source types are described in general terms. Individual providers combine them differently, and many add licensed or proprietary inputs of their own.

The gaps, as a sourcing team feels them

Put the source types together and the same five gaps keep coming back. None of them is a flaw in any one provider. They follow from what private companies publish and from what most providers were built to do, which is to help someone sell to a company rather than source data from it.

The first two gaps, size and activity, decide whether a company passes a buyer's screen at all. The third, data holdings, decides whether it is worth a buyer's review. The last two decide how much time you spend before you find out.

The practical result is familiar to anyone who has tried: a provider export of a few thousand "matching" companies, of which a small share are the right size, still operating and likely to hold records a buyer wants. The rest of the work happens by hand, one website at a time.

Five gaps in general private company data

  • Size is a range or a model. Fine for sales territories, too coarse for a 30 or 20 employee threshold.
  • Activity status lags. Closed, sold and dormant firms can look active long after the fact.
  • No view of data holdings. Industry and headcount do not say whether a firm runs a help desk, a client portal or a code base.
  • History is shallow. A founding year, where known, is not the same as years of records or years of activity.
  • Person heavy. Much of the value in sales databases sits in contacts, which a sourcing team may not need and which carry privacy duties.
Evaluating providers

How to choose the best private company data provider for this job

"Best" depends entirely on the question you are asking. For sourcing AI data partners, run any provider, including us, through the same test before you pay for a year of it.

  1. Write your target as buyer rules. For example: 30 to 200 employees, US, primarily English, at least several years in business, in industries the buyer you work with is asking for. If you cannot write the target down, no provider can find it for you.
  2. Build a test set of 50 companies you know. Include firms you know are in the size band, a few you know have closed or been acquired, a few with parked or dead websites, and some that are clearly too small. Knowing the answers is the whole point.
  3. Measure coverage and freshness on that set. How many of the 50 does the provider return at all? How many closed or acquired companies does it still show as operating? Freshness matters more for this job than the number of fields.
  4. Ask which fields are reported and which are modeled. A size range or revenue figure that is estimated is still useful, but you should know which kind you are filtering on before you exclude a company.
  5. Look for anything about data holdings. Does the provider say anything about what records a company likely keeps: support tickets, knowledge bases, CRM records, code history, research records? If not, plan for that work separately.
  6. Read the license. Can you store results, combine them with your own research, and describe a company to a buyer program using them? Sourcing work depends on being able to pass company-level facts along.
  7. Decide whether you need personal data at all. If your process introduces companies through their public channels or through a program's own referral form, a company-level source may be all you need, with fewer obligations attached.
  8. Price it per useful company. Divide the cost by the companies that pass your rules, not by the rows in the export. A cheaper file with a lower hit rate can cost more per qualified target.

Best source type, by task

No single provider is best at all of these. Most sourcing teams combine two or three source types and keep their own notes on top.

TaskBest starting pointThen confirm with
Is this a real, registered company?Business registriesThe company website
Who owns it, and who would sign?Registries and aggregated databasesThe about page and any parent company's site
Roughly how big is it?Aggregated databases (as a range)The company's own statement, such as "a team of 45"
Does it fit a buyer's published rules?The check against buyer rulesThe buyer program's own review
What data does it likely hold?The Data Asset ScoreA manifest from the company, if it engages
Is it still active?The Data Asset Score's activity statusA recent visit to the site
Finding firms in a nicheDirectories and association listsScore the domains you find

Last checked: 8 October 2026 for our tools. Source types described in general terms, not as a ranking of named providers.

Filling the gap

Website signals, the Data Asset Score and the free check, at company level

A company's own public pages are the one source that every private firm with a website produces. Three layers turn that public footprint into answers a sourcing team can use, from a single target to a whole file.

Website signals you can read

A careers page that says "a team of 70", an about page with a founding year and a US address, a help center, a client login, a job ad asking for experience with a named ticketing system. These are facts the company states itself, and a careful broker reads them before any outreach.

Check against buyer rules

The free check-your-company tool reads a company's public pages and compares team size, location, history and systems with the published rules of micro1, Mode and Grepped, with quotes. Use it to qualify a target in about 20 seconds.

Data Asset Score

Scores any company from 0 to 100 for the data AI buyers want, with a grade, the data it likely holds, its history and an activity status. Built on our index of 102 million domains, 99.99% of the active internet, with domain history. Free demo, limited per day; API from $99.

The three layers answer different questions. Website signals are evidence you can quote, but reading them by hand does not scale past a few dozen companies. The check against buyer rules turns one company's pages into fit labels against published rules, with the sentence behind every fact. The Data Asset Score works at file scale: send a domain and get back a score, a grade, the likely data, history and activity status, so you know which companies deserve the closer reading.

The score combines nine factor groups: history, scale, knowledge assets, operational systems, customer systems, organization, industry value, expertise and activity status. We publish the groups, not the weights or rules that combine them. Like every estimate from public signals, a score is not a valuation, not an offer, and not proof that a company wants to sell.

Target profile: fieldstone-claims.exampleExample

Layer 1

General company record

Entity
Registered, active filing
Formed
2009
Employees
51 to 200 (modeled)
Industry
Insurance services
Data held
Not covered
Activity
Assumed operating

Layer 2

Check against buyer rules

"Our team of 70 handles claims for carriers across the Midwest."About page (invented) "Experience with Zendesk and Salesforce preferred."Careers page (invented)

Likelymicro1
LikelyMode
PossibleGrepped

Layer 3

Data Asset Score

68B
Status
Active
Online since
2009
Likely data
Support tickets, CRM records, project histories
Example only: a fictional company on a reserved example domain, with invented quotes and values. It shows how the layers add up. A general record says the firm exists and is mid-size; the check shows the team of 70 and the US presence in the company's own words; the score adds what data it likely holds and that it is still active.
Ready lists for 20 US sectors. If you would rather start from a ready list than a blank sheet, there are company lists for 20 US sectors, from labs and CROs to law firms, accounting firms and staffing companies. Every company is verified active: the domain resolves, is not expired, is not parked and shows no error or placeholder page. Each carries its Data Asset Score, likely data assets, history and activity status. If your buyers want research and study records, start with the labs and CROs preview, where the top 5 are visible. Buy a list once from $249 as a CSV snapshot, or get lists that stay current on Pro, through the API as JSON, up to 100 companies per call with paging; Scale adds the one-file bulk CSV export.
Basic, $99 a month5,000 lookups a month. Score your own target files; scores companies only.Choose Basic
Pro, $299 a month25,000 lookups plus the company lists endpoint: 20 US sector lists, always current.Choose Pro
Scale, $799 a month100,000 lookups, lists and bulk CSV, new segments on request.Choose Scale

Or buy lists once, without a plan: 1 list $249, 2 lists $449, 3 lists $599, 5 lists $899, each further list +$110, all 20 lists $2,490 on the list buy page. Download links appear right after payment and are also emailed, valid 30 days with 5 downloads per file; the files are a snapshot with no updates.

Last checked: 8 October 2026. Pay by PayPal or card; the API key appears in your dashboard after payment, not by email. No free plan or trial. Details on the pricing page.

Privacy

Company data is not personal data. We keep it that way.

Most of the legal weight in the data provider business sits with information about people. Sourcing companies for AI data programs does not require it, so we leave it out entirely.

Company data (what we provide)

Describes
A business: its website, industry, country, history and activity status
Examples
A Data Asset Score and grade, the data types the company likely holds, the year its domain was first seen
Typical use
Deciding which companies to research, and in what order
Main care point
Presenting it honestly: a score is an estimate, not a valuation or a sign of intent to sell

Personal data (what we never provide)

Describes
A person: an owner, an officer, an employee
Examples
Names, job titles, work emails, phone numbers, profile links
Typical use
Contacting individuals directly
Main care point
Privacy and marketing laws that differ by country and US state, consent, opt-outs and data subject requests

The line is not always sharp. A one-person business, or a record that names a company's owner, can count as personal data. That is one more reason everything we provide stays at company level. General information, not legal advice.

No contacts

Why we provide no contacts

It is a deliberate limit, not a missing feature. Here is the reasoning, and what to do instead.

The question is which company

Buyer programs screen companies: size, location, records. The hard part of sourcing is picking the right companies, and that is what our tools work on.

People data carries duties

Personal data brings privacy and marketing rules with it, which differ by country and US state. Leaving it out keeps our tools simple to use and your target list free of people's details.

Use the company's own front door

Reach a company through its public channels, or through a buyer program's own referral form. A message that shows you read the company's pages lands better than a cold list.

A score is not an invitation

A high score does not mean a company wants to sell. Never present it to the company as a valuation or an offer.

Follow the program's terms

Referral terms set the rules for brokers. micro1's, for example, forbid sharing payouts with the company and posing as micro1's partner, as published, checked 7 October 2026.

The check stays silent

The check against buyer rules never contacts the company that was checked, and never tells a buyer that a check was run.

FAQ

Questions about private company data providers

What are private company data providers?

They are services that collect and sell information about companies that are not listed on a stock exchange. Their data comes from a mix of business registries, regulatory filings, company websites, directories, surveys and panels, often combined into one database with modeled fields such as an employee range or a revenue estimate.

Which private company data provider is best?

It depends on the job. Registries are best for legal existence and formation dates, aggregated databases for broad firmographic coverage, and the company website for what the business says about itself. For sourcing AI data partners, the questions that matter most are what data a company likely holds and whether it is still active, and few general providers answer either. Most teams combine sources.

Why is it hard to find data on private companies?

Private companies have no duty to publish headcount, revenue or the software they use. Registries record legal facts, not operations. Many profile fields are therefore estimated, and records can stay unchanged long after a small company is sold or stops trading.

What kind of private company do AI data buyers want?

As published, checked 7 October 2026: micro1 lists 30+ employees, and its referral posting says 30 to 200, US first and primarily English. Mode asks for 20+ full-time US office employees, 10+ for accounting firms and 6+ for law firms, with several years of records. Grepped lists any vertical.

Do you provide contacts or decision makers at these companies?

No. Everything we provide is company level: scores, grades, likely data, history and activity status. We provide no named people, emails or phone numbers, in any plan or tool.

Is company data personal data?

Usually not. A company's industry, size, history and status describe a business. It can become personal data when the business is one person, such as a sole proprietor, or when a record names owners, officers or staff. That is one reason we keep everything at company level. General information, not legal advice.

How does the Data Asset Score help with private companies?

It scores any company with a website from 0 to 100 for the data AI buyers want, with a grade, the data it likely holds, its history and its activity status. It is built on our index of 102 million domains, 99.99% of the active internet, so small private firms are covered the same way as large ones.

Is there a free way to check a private company?

Yes, two. The free check against buyer rules reads a company's public pages and compares them with the published rules of micro1, Mode and Grepped, in about 20 seconds. The Data Asset Score demo is free and limited per day. API plans start at $99 a month; there is no free plan or trial.

Start with one company you are already looking at

Run it through the free check against buyer rules, then see its Data Asset Score. If the two together save you an hour of reading, the API does the same for a whole file.