Standard operating procedures, playbooks and internal wikis are the easiest company data to scope and to clean. This page covers what buyers ask for, what never belongs in the export, and how to package documents so they are worth reviewing.
Last checked: 7 October 2026. Buyer terms are quoted as published on that date.
Written material a company keeps so work can be repeated: standard operating procedures, playbooks, runbooks, checklists, onboarding guides, QA procedures, templates and decision trees. It usually lives in Confluence, Notion, SharePoint, Google Drive or Dropbox, all systems that buyer programs list among the sources they accept.
micro1’s page names “SOPs, knowledge bases, internal documentation” and “QA processes” among the data it wants, and its eligibility criteria include “documented processes” (as published, checked 7 October 2026). A company with good documentation therefore has two things at once: data to offer, and evidence that it fits.
Compared with email, chat or call recordings, a knowledge base raises fewer of the hard questions. Not none, but fewer.
Spaces, workspaces, sites and folders already group content by team or topic. You can include the Operations space and exclude HR in one line of a scope document.
Procedures are written for a reader who was not in the room. They name roles more often than people, and rarely hold customer details.
A general process for onboarding a client is your know-how. A file about one named client is not. The line is usually visible from the page title.
A buyer can judge quality from a handful of representative pages. That makes the manifest-and-samples step faster than it is for message archives.
Wikis keep version history. Seeing how a procedure changed after an error is a record of decisions over time, not only a static rule.
Employees accept “we licensed our playbooks” more readily than “we licensed your messages.” That matters for trust as well as for notice.
Easy to scope is not the same as most valuable. micro1’s list puts SOPs next to CRM data, project histories and QA processes, and it names “decision-making patterns” as a goal. A procedure tells a model the intended path. The tickets, emails and project records show the exceptions, judgment calls and fixes that the procedure did not foresee.
That is why documents often work best as part of a package. A support playbook paired with the support tickets handled under it shows rule and practice side by side. A QA checklist paired with QA results shows how the checklist was applied. If your documents only exist on their own, they still count, but expect a buyer to ask what records sit behind them.
Practitioners describe value in tiers: raw data is the cheapest, evaluations built on the data are worth roughly ten times raw, and full training environments reach six to eight figures but need heavy engineering. Well-written procedures with clear expected outcomes are the kind of material evaluations can be built from, which is worth raising when you compare offers. Our guide to what data AI labs want covers the difference between connected histories and scattered files.
Each pairing adds scope, and with it more personal data to handle, so the extra value has to be weighed against the extra preparation. A document-only package is a reasonable first deal; the pairing can come later if the agreement allows it.
No buyer we track publishes a scoring rubric for documents. These are the questions a reviewer is likely to ask when reading your samples, so ask them first.
“Check the invoice against the PO” is a rule. “Check it because vendors re-send old invoices after price changes” is judgment. Pages that explain why are richer than bare checklists.
The best procedures say what to do when the normal path fails: who approves, what to escalate, when to stop. Exceptions are where experience shows.
Pages edited over several years, with history kept, show a living process. A wiki written in one week and never touched again tells a reviewer much less.
A shared template across teams (purpose, owner, steps, exceptions, related records) makes a document set easier to assess and easier to de-identify.
Most knowledge bases hold a few pages that should never leave the company. Find them before you share a manifest, not after.
| Item | Why it is a problem | Usual handling |
|---|---|---|
| Credentials in pages | Teams paste passwords, API keys and Wi-Fi codes into wikis. De-identification is not built to catch them. | Search for secrets, rotate anything found, exclude the pages. |
| Pages about individuals | Performance notes, salary tables, leave records and phone lists are employee personal data. | Exclude the HR space; remove staff directories and org charts with names. |
| Client deliverables | Reports, plans and playbooks built for one client may belong to that client or fall under an NDA. | Keep the general method; exclude client-specific files. See client confidentiality. |
| Third-party material | Vendor manuals, paid courses, purchased standards and copied articles are not yours to license. | Exclude by source; list the exclusion in the manifest. |
| Legal and privileged notes | Advice from counsel and dispute files may be privileged and confidential. | Exclude the whole legal area unless your lawyer says otherwise. |
| Security runbooks | Incident and access procedures describe how your systems can be reached. | Exclude, or include only after your security lead reviews them. |
| Attachments and screenshots | Embedded images and files often hold customer records that page text does not. | Review attachments separately, or exclude them from the first scope. |
Most knowledge bases are scoped with a mix of the three. Pick the main one first, then use the others to tidy the edges.
Include whole spaces, workspaces or sites, such as Operations, Quality and Onboarding, and exclude others whole, such as HR, Legal and Clients. Simplest to write and to check. Works when your spaces follow team lines.
Include procedures, checklists, runbooks and templates wherever they sit; exclude meeting notes, personal pages and drafts. Better when spaces are messy, but needs labels or a reliable naming pattern.
Include a defined period, for example from the year a system was adopted until a recent cut-off. Useful when early content is thin or when a reorganization changed who wrote what.
Watch for meeting notes. Many wikis hold years of meeting notes beside the procedures. They read more like chat than like documentation: names, side remarks, decisions about individual clients and staff. Decide explicitly whether they are in. If they are, treat them with the care you would give tickets or messages, not with the lighter touch that procedures need.
Watch for personal spaces. Tools that give each employee a private or personal area often fill up with drafts, notes to self and copies of sensitive files. Exclude personal spaces by default and make any exception a deliberate one.
A full-site export is one click, which is why it is tempting. It also carries HR pages, client folders and personal spaces you never meant to share.
Page text gets reviewed; the spreadsheet attached to it does not. Attachments are where customer lists and payroll files tend to hide.
Inline and page comments name people and carry remarks the page author would never have published. Include them only on purpose.
“8,000 pages” sounds large until a reviewer finds half are stale drafts and copies. Describe what the pages are, not just how many.
Polishing procedures, especially with AI writing tools, replaces your team’s own record with new text. Buyers want how people actually wrote and worked.
Samples go after an NDA, and the full set only after a signed agreement. Practitioners advise never sending a full dataset before a price.
The fictional firm has about 1,400 Notion pages across 12 teamspaces. After the review:
There is no price here: nobody can value a document set without reviewing it. What the example shows is how short the exclusions list is when the material lives in well-named spaces.
For the de-identification side, including what the buyer does and what you should verify yourself, see de-identification before selling data. For the whole preparation sequence across every system, see prepare your data for sale.
General questions about price and liability apply to every data deal. These are the ones specific to documents.
Quoted from each program’s own pages, checked 7 October 2026. Published ranges are not promises for any company.
micro1 lists SOPs, knowledge bases and internal documentation, with 30+ employees, documented processes, modern software tools and primarily English; US companies first. Its pages publish “$100k+ qualified,” “$500k+ large-scale” and “$1M+ highly unique.” Mode publishes “$100K-$5M” for company data, with 20+ full-time US office employees (accounting firms 10+, law firms 6+) and several years of records the company owns. Grepped publishes “$20K-$5M” and works with any vertical.
Independent site. Some links are referral links: if your company signs with a buyer through them, the buyer may pay us a fee. You are not charged, and we never see your data.