AI data buyers publish a clear target: private companies of roughly 30 to 200 people, US first, with several years of records. Those are the companies general data providers know least about. This guide sorts private company data sources by type, shows where each one falls short for AI data sourcing, and explains how to fill the gap at company level, without buying anyone's contact details.
Scoring many companies at once? See the API plans or the company data API guide.
A listed company files annual reports, discloses its headcount and explains its business in detail. A 70-person private firm owes the public almost none of that. Every provider of private company data is working around the same silence.
Private companies in the US generally file formation documents and periodic reports with their state, which record legal facts: the entity, its status, a registered agent. They do not record how many people work there or what the business runs on.
Employee counts for private firms are usually ranges built from indirect evidence. A firm shown as "51 to 200" could sit at either end, which matters when a buyer's rule is 30 or more employees.
Revenue for a private company is almost always a model. That is less of a problem for AI data sourcing than for sales, because buyers screen on people and records, not turnover.
The help desk, CRM, ticketing and document systems a company runs inside the business rarely appear anywhere official. Yet those systems hold the records buyers pay for.
Small firms rebrand, merge, sell to a competitor or move under a holding company. Matching all of those changes back to one company is hard, and some records never catch up.
A small company that stops trading rarely announces it. The website may stay up, then lapse, then point to a parking page. A profile can look current for a long time after the business has gone.
The published rules of the main company data programs describe a mid-size private firm almost word for word: a few dozen to a couple of hundred people, based in the US, with years of records in its own systems.
| Program | Size rule | Location and language | Records | Referral, as published |
|---|---|---|---|---|
| micro1 | 30+ employees; the referral posting says 30 to 200 | US first, primarily English | Approved data packages | "earn $50,000", "no cap", paid after onboarding plus a minimum revenue threshold |
| Mode | 20+ full-time US office employees; accounting firms 10+; law firms 6+ | US office employees | Several years of records | "Earn $50K per referral" |
| Grepped | No published minimum | No published rule | Any vertical | "Refer for another $10K" |
| Miro Advisory | Not listed here | Not listed here | Operating datasets $100K-$1M+, codebases $10K-$1M+ | Not listed here |
As published, checked 7 October 2026. Ranges are what programs publish, not offers, and the top of a range is not the typical result. micro1's referral terms also state sole discretion and clawbacks, and forbid sharing payouts with the company or posing as micro1's partner. Full details: buyer programs compared.
The rules above are not arbitrary. A company of 30 to 200 people has enough staff to produce a large, connected body of work: tickets that link to customer accounts, projects that run from proposal to delivery, procedures that were written down because new hires needed them. A company with five people rarely has that depth. A company with fifty thousand usually has a longer approval path, more contracts to check and more lawyers in the room.
Several years of records matter for the same reason. Buyers want to see how work played out over time: a client across several renewals, a product through several releases, an issue from first report to fix. A firm that migrated systems last year may have a long history as a business but a short one in its records.
The US first rule also has a privacy side: an EU or UK team brings GDPR into scope for every record that mentions a person. None of this rules out other companies. It sets the order in which buyers look, and therefore the order in which a broker should search.
The companies that fit the published rules best are the ones that are least visible in the sources most data providers start from. That is the gap this page is about.
Every private company data provider draws on some mix of the same source types. Knowing which type a field came from tells you how much to trust it, and what it can never tell you.
| Source type | What it tells you | Where it falls short for AI data sourcing |
|---|---|---|
| Business registries | That an entity exists, its legal name, formation date, state of formation, filing status and registered agent | Says nothing about size, systems or records. A company can be in good standing and dormant. Each state keeps its own registry with its own fields. |
| Regulatory filings and licenses | Professional and trade licenses, permits, industry registrations, government contracting records, secured lending filings | Strong for regulated trades and contractors, thin for everyone else. Confirms a license, not how many people use it or what they keep. |
| Company websites | What the company says about itself: what it does, where it is, when it was founded, open roles, help pages, client logins | Unstructured and uneven. Many sites never state headcount. Reading them one by one is slow, and a site can outlive the business behind it. |
| Directories and listings | Trade association member lists, industry directories, map and review listings, chamber of commerce rolls | Often self-reported and rarely updated. Good for finding names in a niche, weak for size and activity. |
| Panels and surveys | Self-reported answers from a sample of businesses, sometimes with detail no other source has | Covers only the businesses that took part, and only as of the survey date. Hard to use for a specific target. |
| Aggregated databases | A combined profile built from several of the above, matched to one company, with modeled fields filled in | Broad, but the modeled fields are estimates, and the profile is built for sales: industry, size band, revenue, contacts. Data holdings and activity are rarely explicit. |
Put the source types together and the same five gaps keep coming back. None of them is a flaw in any one provider. They follow from what private companies publish and from what most providers were built to do, which is to help someone sell to a company rather than source data from it.
The first two gaps, size and activity, decide whether a company passes a buyer's screen at all. The third, data holdings, decides whether it is worth a buyer's review. The last two decide how much time you spend before you find out.
The practical result is familiar to anyone who has tried: a provider export of a few thousand "matching" companies, of which a small share are the right size, still operating and likely to hold records a buyer wants. The rest of the work happens by hand, one website at a time.
"Best" depends entirely on the question you are asking. For sourcing AI data partners, run any provider, including us, through the same test before you pay for a year of it.
No single provider is best at all of these. Most sourcing teams combine two or three source types and keep their own notes on top.
| Task | Best starting point | Then confirm with |
|---|---|---|
| Is this a real, registered company? | Business registries | The company website |
| Who owns it, and who would sign? | Registries and aggregated databases | The about page and any parent company's site |
| Roughly how big is it? | Aggregated databases (as a range) | The company's own statement, such as "a team of 45" |
| Does it fit a buyer's published rules? | The check against buyer rules | The buyer program's own review |
| What data does it likely hold? | The Data Asset Score | A manifest from the company, if it engages |
| Is it still active? | The Data Asset Score's activity status | A recent visit to the site |
| Finding firms in a niche | Directories and association lists | Score the domains you find |
A company's own public pages are the one source that every private firm with a website produces. Three layers turn that public footprint into answers a sourcing team can use, from a single target to a whole file.
A careers page that says "a team of 70", an about page with a founding year and a US address, a help center, a client login, a job ad asking for experience with a named ticketing system. These are facts the company states itself, and a careful broker reads them before any outreach.
The free check-your-company tool reads a company's public pages and compares team size, location, history and systems with the published rules of micro1, Mode and Grepped, with quotes. Use it to qualify a target in about 20 seconds.
Scores any company from 0 to 100 for the data AI buyers want, with a grade, the data it likely holds, its history and an activity status. Built on our index of 102 million domains, 99.99% of the active internet, with domain history. Free demo, limited per day; API from $99.
The three layers answer different questions. Website signals are evidence you can quote, but reading them by hand does not scale past a few dozen companies. The check against buyer rules turns one company's pages into fit labels against published rules, with the sentence behind every fact. The Data Asset Score works at file scale: send a domain and get back a score, a grade, the likely data, history and activity status, so you know which companies deserve the closer reading.
The score combines nine factor groups: history, scale, knowledge assets, operational systems, customer systems, organization, industry value, expertise and activity status. We publish the groups, not the weights or rules that combine them. Like every estimate from public signals, a score is not a valuation, not an offer, and not proof that a company wants to sell.
Likelymicro1
LikelyMode
PossibleGrepped
Most of the legal weight in the data provider business sits with information about people. Sourcing companies for AI data programs does not require it, so we leave it out entirely.
The line is not always sharp. A one-person business, or a record that names a company's owner, can count as personal data. That is one more reason everything we provide stays at company level. General information, not legal advice.
It is a deliberate limit, not a missing feature. Here is the reasoning, and what to do instead.
Buyer programs screen companies: size, location, records. The hard part of sourcing is picking the right companies, and that is what our tools work on.
Personal data brings privacy and marketing rules with it, which differ by country and US state. Leaving it out keeps our tools simple to use and your target list free of people's details.
Reach a company through its public channels, or through a buyer program's own referral form. A message that shows you read the company's pages lands better than a cold list.
A high score does not mean a company wants to sell. Never present it to the company as a valuation or an offer.
Referral terms set the rules for brokers. micro1's, for example, forbid sharing payouts with the company and posing as micro1's partner, as published, checked 7 October 2026.
The check against buyer rules never contacts the company that was checked, and never tells a buyer that a check was run.
They are services that collect and sell information about companies that are not listed on a stock exchange. Their data comes from a mix of business registries, regulatory filings, company websites, directories, surveys and panels, often combined into one database with modeled fields such as an employee range or a revenue estimate.
It depends on the job. Registries are best for legal existence and formation dates, aggregated databases for broad firmographic coverage, and the company website for what the business says about itself. For sourcing AI data partners, the questions that matter most are what data a company likely holds and whether it is still active, and few general providers answer either. Most teams combine sources.
Private companies have no duty to publish headcount, revenue or the software they use. Registries record legal facts, not operations. Many profile fields are therefore estimated, and records can stay unchanged long after a small company is sold or stops trading.
As published, checked 7 October 2026: micro1 lists 30+ employees, and its referral posting says 30 to 200, US first and primarily English. Mode asks for 20+ full-time US office employees, 10+ for accounting firms and 6+ for law firms, with several years of records. Grepped lists any vertical.
No. Everything we provide is company level: scores, grades, likely data, history and activity status. We provide no named people, emails or phone numbers, in any plan or tool.
Usually not. A company's industry, size, history and status describe a business. It can become personal data when the business is one person, such as a sole proprietor, or when a record names owners, officers or staff. That is one reason we keep everything at company level. General information, not legal advice.
It scores any company with a website from 0 to 100 for the data AI buyers want, with a grade, the data it likely holds, its history and its activity status. It is built on our index of 102 million domains, 99.99% of the active internet, so small private firms are covered the same way as large ones.
Yes, two. The free check against buyer rules reads a company's public pages and compares them with the published rules of micro1, Mode and Grepped, in about 20 seconds. The Data Asset Score demo is free and limited per day. API plans start at $99 a month; there is no free plan or trial.
Run it through the free check against buyer rules, then see its Data Asset Score. If the two together save you an hour of reading, the API does the same for a whole file.