Firmographics tell you how big a company is. They do not tell you what records it keeps. An AI data buyer pays for years of support tickets, client portal activity, CRM history or code with its issue trail. The Data Asset Score estimates exactly that for any domain, on one scale from 0 to 100. This page explains the six data layers behind every score, what each layer tells you about a company's data, and why a precomputed, comparable score saves a broker more time than a firmographic file, a chatbot prompt or an afternoon of manual research.
We explain the data layers in full. We do not publish weights, point values, thresholds or the rules that combine the groups. Every score is an estimate from public signals: not a valuation, not an offer and not proof that a company wants to sell.
Every score has the same shape, whether you check one company in the demo or ten thousand through the API. These are the fields a broker works with. The rest of this page explains where each one comes from.
Fictional company and values, shown to illustrate the fields. A real result also carries verified_active, site_readable, cached and checked_at.
Buyer programs publish rules about team size, location and years of records, and a firmographic file can screen on the first two. The question that decides whether a deal happens is harder: does this company hold records a buyer wants, in systems that can be exported? Two accounting firms with the same headcount look identical in a firmographic file. One has run a client portal, a help desk and a document management system for fifteen years. The other works from email and a shared drive. A buyer treats them very differently, and so should your pipeline.
The Data Asset Score starts where firmographics stop. It looks for public evidence that a company keeps records: the kinds of pages it publishes, such as a help center, a client portal or a status page; the systems its site shows; how long the domain has been in use; how large its footprint is; and whether it is still operating today. From that evidence it infers the likely data assets and rolls everything into one number you can sort on.
Here is how that compares with the three ways brokers usually qualify a company today.
| Firmographic file | Asking an AI chatbot per domain | Manual research | Data Asset Score API | |
|---|---|---|---|---|
| What you learn | Size, industry, location | Whatever the model writes about that company, in its own words | Whatever the researcher finds and writes down | Score, grade, likely data assets, nine factor groups, history, status |
| Records the company keeps | Not covered | A guess, phrased differently each time | Found, if the researcher knows where to look | Inferred from page types and systems, with a confidence level |
| Years of history | Founding year, when known | Only if it finds a source in that session | Time spent in archives and registration records | First seen year and years online in every result |
| Still operating? | As of the last update of the record | Not checked unless it browses, and then not systematically | A visit to the site | Live check at scoring time, returned as a status |
| Comparable across companies | For size, not for data | No fixed scale | Depends on who did the work, and when | One method and one scale for every domain |
| Time per company | Instant | A prompt, a wait and an answer to read | Minutes, often many | Seconds |
| Thousands at once | Yes | One prompt at a time, or your own script around a paid model | No | Loop a CSV through the API |
Each layer answers a different question about a company. None of them is enough on its own: a help center without history could be a startup with ten customers, and history without systems could be a dormant brochure site. Together they give a picture of a company's records that no single source gives.
Our index holds 102 million domains, 99.99%+ of the active internet, each classified into one of more than 700 IAB content categories. If a company runs a website, it is almost certainly in it.
What it tells you: which industry a company is in, at a level of detail that separates a contract research organization from a hospital, or an insurance broker from a bank. It also lets you work a sector as a whole instead of guessing search terms.
A page-type database covering about 40 million domains records which kinds of page a site has: help center, documentation, community or forum, careers, leadership, security, integrations, status, partners, case studies, login, signup, client portal, checkout and upload pages, among others.
What it tells you: what a company has built around its customers and its own knowledge. A help center suggests support tickets and a knowledge base. A client portal suggests account and activity records. Careers and leadership pages suggest an organization of some size.
The systems a company's site shows: help desk and live chat, CRM and marketing automation, issue trackers, learning management, scheduling, e-commerce, recruiting systems, document management and wikis. We hold stored technology data for a few million domains, and every score request adds a live detection on the site itself.
What it tells you: which records the company keeps, because systems keep records. A help desk keeps tickets, a CRM keeps account and deal history, an issue tracker keeps years of engineering work.
When a domain first appeared, from registration records where they exist, public web archives, copyright years on the site and the founding year a company states about itself.
What it tells you: how many years of records a company could hold. Buyers ask for several years of records, and a company online since 2006 has had far longer to build them than one that launched last year.
Popularity and link signals, used where they exist for a domain. They are strongest for the better-known part of the web and absent for many small sites, which is why scale is one group among nine and never the whole score.
What it tells you: how large the company's footprint is: how many people use its site and how much of the web points at it. A larger footprint usually means more customers, more transactions and more support traffic behind it.
A fresh reading of the domain when it is scored: does it resolve, is the registration current, is the site parked, erroring or showing a placeholder, and does it carry wind-down or acquisition language. The same reading looks for specialist and regulated professional vocabulary.
What it tells you: whether the company is still there, and whether its work is specialist enough that its records cannot be found on the open web.
The nine factor groups are what a score is made of, and each group draws on one or more layers. This map shows those relationships and what each one means for your pipeline. How much each group counts, and the rules that combine them, stay our own.
| Data layer | History | Scale | Knowledge assets | Operational systems | Customer systems | Organization | Industry value | Expertise | Activity status |
|---|---|---|---|---|---|---|---|---|---|
| The index | |||||||||
| Page types | |||||||||
| Technologies | |||||||||
| History | |||||||||
| Scale | |||||||||
| Live check |
A page is public; the records behind it are not. But companies do not build a help center without a support operation, or a client portal without client accounts. These are the page types that matter most to the score, grouped by the factor group they feed.
What the company has written down, for customers and for itself.
Where customers log in, buy and send things. Each one implies a record per customer.
Signs of a company of some size, with structure and processes.
AI data buyers do not buy websites. They buy what sits inside the systems a company runs: the help desk, the CRM, the issue tracker. When a company's site shows one of those systems, the records are very likely there too. The table shows each system category, what it keeps, and where it shows up in a result.
| System category | What it usually keeps | Why buyers care | Where it shows in the result |
|---|---|---|---|
| Help desk and live chat | Tickets, chat transcripts, macros, satisfaction ratings | Real questions paired with real resolutions, over years | Support tickets |
| CRM and marketing automation | Account histories, deal stages, campaigns, activity logs | How a business sells, step by step, at company level | CRM records |
| Issue trackers and developer tools | Issues, change discussions, release history | Code with the reasoning behind each change | Code history |
| Wikis and knowledge tools | Internal procedures, SOPs, how-to pages | Written expertise in a company's own words | Knowledge base |
| Forum software | Threads, replies, accepted answers | Long conversations between users and staff | Forum archive |
| Learning management | Courses, quizzes, completion records | Structured teaching material with assessments | Training content |
| Recruiting systems | Job requisitions, application stages | Hiring workflows and job descriptions at scale | Recruiting records |
| E-commerce | Orders, carts, refunds, catalogs | Transaction patterns and product data | Transactions |
| Call and contact center | Call logs, recordings, transcripts | Spoken conversations, rare and in demand | Call recordings |
| Lab, study and quality systems | Samples, methods, results, deviations | Specialist work that never appears on the open web | Research records |
| Scheduling and document management | Bookings, contracts, reports, versioned files | Evidence of a business run through systems, not inboxes | Operational systems group |
Buyers want to see how work played out over time: a client across several renewals, a product across many releases. Years of records are one of the clearest separators between a company that is worth a buyer's review and one that is not. The history layer brings together four kinds of evidence, used where they exist.
When the domain was first registered, where registration history exists for it. It covers part of the index, not all of it, and it is the strongest single date when present.
When copies of the site first appear in public web archives. A domain can be registered years before a real site goes live; the first archived site is often closer to when the business started using it.
The year range in a site footer, such as a start year followed by the current year, is a small but common clue that the site has been maintained for a long time.
"Founded in 2009" or "serving clients since 1998" on the company's own pages. This is returned separately as founded_year, because a company can be older than its domain.
The result gives first_seen_year and years_online for the domain, and founded_year when the company states one. When they disagree, the company may have changed its domain or rebranded. A firm founded in 1998 on a domain first seen in 2015 still has the longer history, and the stated year tells you so.
A long history is not the same as long records. A company can lose years of email in a system migration. Treat history as how long records could exist; the buyer's review finds out how many do.
Scale comes from popularity and link signals. They exist for the better-known part of the web and are missing for many small business sites. Where they are missing, the scale group reads weak or none, and the other eight groups still carry the score.
That matters for brokers, because many good targets are mid-sized professional firms with modest traffic and deep records. A small footprint does not hide a company with a client portal, a help desk and twenty years of history.
Lists age fast. Companies close, merge, let a domain lapse or leave a parking page behind. Every score includes a fresh check of the domain at the time it is scored, and the answer comes back as one of four statuses.
The domain resolves, the registration is current, the site loads real content, and nothing on it says the company is closing. This is the status to work first.
The site carries language about closing, shutting down a service, or being now part of another company. Who can sign may have changed, and timing matters.
A parking, for-sale or placeholder page. Whatever company used the domain is not operating from it now.
The domain does not resolve or the site does not respond. The API can also return error when a site serves only error pages.
A broker's time is the scarcest input in the pipeline. Practitioners say data deals take 60 to 90 days to close; a week spent researching and writing to a company that closed last spring is a week not spent on one that could sign. The live check removes those companies before you start.
It also catches a quieter problem: domains that still resolve but no longer carry a business. A placeholder page or a generic error looks like a website to a firmographic file. It does not look like one to the live check.
While reading the site, the live check also looks for specialist and regulated professional vocabulary: the language of lab work, assays and formulations, claims and underwriting, litigation, engineering disciplines. Specialist work produces records that appear nowhere on the open web, which is exactly what buyers say is hardest to source. That evidence feeds the expertise group.
The result carries verified_active: true when the domain resolves, is not expired and the site is not parked or an error page. Only verified companies go into our company lists.
Asking a general AI chatbot about each company is the first thing many brokers try, and for one company it can be a useful start. For a list of two thousand it breaks down in four places: speed, cost, consistency and what it never checks.
The two work well together. Score the whole list first, then use a chatbot or your own reading on the top fifty, where an hour of depth per company pays off. Manual research has the same shape: if careful research takes ten minutes a company, a list of 2,000 companies is more than 330 hours of work before the first email.
The score is not a report you read once. It is a sort key for a pipeline. These are the four ways brokers and sourcing teams use it most.
You already have a list: from a conference, a CRM export, a directory or an old campaign. Send every domain through the score endpoint and sort. The companies at the top show the strongest evidence of the records buyers ask for, so they get your first week, and the bottom of the list waits.
Sort by data_asset_score, highest first.
Every list has companies that closed, merged or let their domain lapse. Filter on status and stop writing to companies that are gone. Keep winding down or acquired in a separate pile: those companies need a different conversation, and sometimes a faster one.
Keep status = active; set aside winding_down_or_acquired.
A buyer asks for support tickets, CRM records or code history. Filter the likely data assets for that type, high confidence first, and you have a short list that matches the request instead of a whole sector. It turns a vague brief into fifty names.
Filter likely_data_assets for support_tickets, crm_records or code_history.
Grades turn a long list into piles you can act on: research A and B first, check C companies for one specific data type, and park D and E unless a buyer asks for something they clearly show. The factor groups tell you why a company landed where it did.
Group by grade, then read factor_groups.
If companies apply to your program, score each domain before your team spends time on the review. A grade, a status and the likely data assets on every application let the team read the strongest ones first and spot a parked or closed domain before anyone schedules anything. On Pro and Scale, the 20 sector lists come through the same API, so outbound sourcing and inbound screening use one scale.
No special tooling: a CSV, a key and a short script. The batch endpoint takes up to 100 domains per call and returns the results in the order you sent them, so a script sends your file in blocks of 100 and writes the results back out, ranked. Full reference in the API documentation.
import csv, time, requests API = "https://www.selldatatoai.com/api/v1/score/batch" HEAD = {"X-API-Key": "YOUR_KEY"} with open("targets.csv", newline="") as f: domains = [line["domain"] for line in csv.DictReader(f)] rows = [] for i in range(0, len(domains), 100): # up to 100 domains per call r = requests.post(API, json={"domains": domains[i:i + 100]}, headers=HEAD, timeout=60) if r.status_code == 429: break # monthly lookups used up job = r.json() while job["status"] != "done": # polling is free time.sleep(10) job = requests.get(job["poll"], headers=HEAD, timeout=60).json() for item in job["results"]: # same order you sent if item["status"] != "done": continue # invalid domain or error, skip it d = item["result"] rows.append({ "domain": d["domain"], "score": d["data_asset_score"], "grade": d["grade"], "status": d["status"], "assets": ";".join(a["type"] for a in d["likely_data_assets"]), }) active = [x for x in rows if x["status"] == "active"] active.sort(key=lambda x: x["score"], reverse=True) with open("ranked.csv", "w", newline="") as f: w = csv.DictWriter(f, fieldnames=["domain", "score", "grade", "status", "assets"]) w.writeheader() w.writerows(active) # Only the companies likely to hold support tickets: tickets = [x for x in active if "support_tickets" in x["assets"]]
| Rank | Domain | Score | Grade | Status | Likely data assets |
|---|---|---|---|---|---|
| 1 | harborline-claims.example | 81 | A | active | Support tickets, customer accounts, transactions |
| 2 | northgate-bio.example | 76 | B | active | Research records, knowledge base, customer accounts |
| 3 | ridgeway-engineering.example | 68 | B | active | Project histories, recruiting records |
| 4 | copperleaf-legal.example | 57 | C | active | Project histories, customer accounts |
| set aside | brightpath-analytics.example | 62 | C | winding_down_or_acquired | Separate track: who signs may have changed |
| dropped | oldmill-software.example | 18 | E | parked | No business on the domain |
A score is only useful if it changes what you do next. This is how we suggest reading grade and status together. It is guidance for a pipeline, not a rule inside the score.
| Result | What it usually means | What to do next |
|---|---|---|
| Grade A or B, active | Strong evidence across several groups: records in systems, years of history, a real footprint | Research first. Qualify the company against each buyer's published rules with check your company, then decide which program fits. |
| Grade C, active | Some strong groups and some gaps | Read the factor groups and likely data assets. A C with high-confidence support tickets can suit a buyer who asked for exactly that. |
| Grade D or E, active | Little public evidence of records, or a small footprint and short history | Keep for later unless a buyer asks for something the company clearly shows. Internal systems can be invisible from outside. |
| Winding down or acquired | The site suggests a closure, a wind-down or a new owner | Treat it as its own track. Find out who can sign now; timing matters more here than anywhere else. |
| Parked, unreachable or error | No operating business on the domain today | Drop it from outreach. Re-check later only if you know the company moved to another domain. |
Every plan uses the same index, the same score and the same fields. The difference is how many lookups you get each month and whether ready lists come with it.
For brokers who qualify companies one by one, or score a list of a few thousand domains each month.
For brokers who work a pipeline and want ready lists of data-rich companies to start from, always current.
For data companies that source at volume and load lists straight into their own tools.
A score you trust is one whose limits you know. These are the ones to keep in mind before you act on a result.
The score reads what a company shows in public. It never sees the records themselves, how clean they are, or what the company's contracts allow.
A help desk, CRM or code host used only by staff, with no trace on the site, is missed. Read a weak operational systems group as "not visible", not "not there".
Page types, stored technology data, popularity and registration history cover part of the index. Where a layer is missing, its groups read weak or none.
Some sites block automated reading. The result then says site_readable false, and the score relies on our index alone.
Only a buyer can price data, after it sees a manifest and samples. The score tells you which companies to look at first.
A high score says a company likely holds data buyers want. It says nothing about whether its owners would ever sell it.
Results are reused for 30 days. checked_at shows when the reading was taken, so a company that closed last week can still read active until the next reading.
No names, emails or phone numbers, ever. The score tells you which company to approach; finding the right person is your work.
We plan to add a distribution view to the demo that shows where a company's score sits among the domains in our index. It will answer the question brokers ask most after "what is the score?": is that number common or rare? Until it ships, the grade is the quickest way to read a score in context.
No. We publish the data layers and the nine factor groups they feed, because that is what you need to trust and explain a score. The weights, point values, thresholds and the rules that combine the groups are our own. Every result still shows each factor group as strong, medium, weak or none, and lists the signals we found in plain words, so you can see why a company scored high or low without the formula.
A firmographic record tells you a company’s size, industry and location. It does not tell you what records the company keeps. The Data Asset Score infers likely data assets, such as support tickets, CRM records, customer accounts or code history, from the pages a company publishes and the systems its site shows, adds years of history and a live activity check, and puts all of it on one 0 to 100 scale.
For one company it can be a useful start. For a pipeline it is slow, it costs money per call, the same question can return different answers in different formats, and it does not confirm that the domain resolves or that the site is not parked. The score is precomputed on one method, comparable across companies, verified active at scoring time and returned as JSON fields you can sort.
Each score includes a live check of the domain at the time it is scored. Results are then kept for up to 30 days, so repeat lookups are fast. The response carries checked_at, the time of the reading, and cached, which tells you whether the result came from the last 30 days.
Yes. That is what the API is for. Put your domains in a CSV and send them to the batch endpoint, up to 100 per call. Domains scored in the last 30 days come back at once and the rest within a few minutes; you poll for them for free. Each valid, unique domain is one lookup: Basic includes 5,000 a month, Pro 25,000 and Scale 100,000. Then filter by status, sort by score and filter by the data type a buyer asked for.
When a site blocks automated reading, the result says site_readable false and the score relies on our index alone. When a layer has no data for a domain, the factor groups it feeds read weak or none in the result, so you can see what is missing, and the other layers still apply.
No. A high score means the company shows strong public evidence of the kinds of records AI data buyers ask for. It is an estimate, not a valuation, not an offer and not proof that the owners want to sell. Each buyer reviews the actual records and decides acceptance, scope and price.
That is planned. We intend to add a distribution view to the demo that shows where a company’s score sits among the domains in our index, so you can see whether a score is common or rare. It is not live yet.
Try the score on a company you know in the free demo. When the result makes sense to you, send your own list through the API and sort your pipeline by the data buyers want.