For a software company, the repository is often the most structured record of real work it owns. This guide covers what buyers say they accept, the published ranges, and the three checks that come before any export: ownership, open-source licenses and secrets.
Last checked: 7 October 2026. Buyer figures are quoted as published on that date.
When people search for how to sell source code to AI companies, they usually picture a sale of the code itself. In the programs we track, that is rarely the shape of the deal. A buyer licenses a copy of an agreed set of repositories for AI training or evaluation. Mode, for example, says it buys “an agreed copy” and that originals stay with the company. micro1 says the company keeps ownership of its underlying data. Both are quoted as published, checked 7 October 2026.
Keeping ownership does not mean you keep every right. An exclusive license, a resale right or a broad scope of use can limit what you do with the same code later. Those terms matter as much as the price. Our guide to exclusivity and resale rights walks through them.
The data sources buyers list name “GitHub, GitLab and Bitbucket repositories with history.” The word that matters is history. A snapshot shows what the code is. The history shows how people got there.
Each commit pairs a change with a message about why it was made. Years of small, well-described commits show how engineers break problems into steps.
Review comments record disagreement, trade-offs and corrections. That is the “decision-making” layer that buyers describe wanting from company data.
A bug report, the investigation, the fix and the closing note form one connected record of work. Linked Jira or platform issues make the code far easier to read.
Tests show what “correct” meant to your team. Failing and passing builds show how a change was checked before it shipped.
Rolled-back changes and post-incident patches are records of mistakes and recovery. They are rare in public code and often the most instructive part of a history.
Architecture decision records, READMEs and runbooks explain intent. They connect your code to your internal documentation.
One practical catch. A plain git clone carries commits, branches and tags. It does not carry pull request reviews, merge request discussions or issues. Those live on the hosting platform and need a separate export through its API or admin tools. If review threads are part of what you offer, plan that export from the start. The same goes for linked tickets: software companies that hold code, Jira and Slack for the same projects have the most connected history, which is the angle of our software companies guide.
Dates may matter too. Buyers training models care whether code was written by people. A long history from before AI coding assistants became common is easier to vouch for than recent code that may be partly generated. We have no published buyer statement that prices this. Treat it as a question to ask each buyer: do commit dates, or a split before and after a given year, change how they value the repository?
None of these figures is a quote for your repositories. They are ranges and reported prices, each tied to its source.
| Source | Figure | What it covers |
|---|---|---|
| Miro Advisory (its site) | $10K to $1M+ | Private codebases, labeled indicative. Operating datasets are listed separately at $100K to $1M+. |
| Troveo (cited in 2026 closure-market coverage) | about $5,000 | Per code repository from shut-down startups, as cited by Troveo. Forbes covered this market on 16 April 2026. |
| micro1 (its pages) | $100k+ / $500k+ / $1M+ | “Qualified,” “large-scale” and “highly unique” company data packages. Not code-specific. |
| Mode (its site) | $100K to $5M | Company data generally. Not code-specific. |
| Grepped (its site) | $20K to $5M | Any vertical. Not code-specific. |
Last checked: 7 October 2026, from each company’s own published pages; Troveo’s figure as cited in 2026 news coverage. “Up to” and “+” ranges describe what a program has published, not what a typical company receives.
The gap between $5,000 per repository and seven-figure package ranges is not a contradiction. The closure market sells single repositories from companies that no longer exist, often without the surrounding tickets or reviews. Program ranges describe larger packages from operating companies. Practitioners say raw data is the cheapest form, that evaluations built on the data are worth roughly ten times raw, and that full training environments reach six to eight figures but need heavy engineering. Our page on raw data versus evaluations versus environments explains those tiers.
Whatever the tier, practitioners give one rule: never send the full repository before you have a price. Share a manifest of what exists and a few samples, and get more than one offer.
Most commercial repositories contain code the company did not write. Before anything goes into scope, sort the tree by who owns each part.
Copied libraries, node_modules, vendor/ folders and bundled SDKs belong to their authors. Exclude them and list them in the manifest so the buyer sees what was left out and why.
Code under GPL or AGPL carries conditions on redistribution. Whether licensing it for training triggers them is a question for your lawyer, not for a sales call. Keep it out unless you have clear advice.
MIT, Apache 2.0 and BSD code is permissive, but its notices still apply. Mixing it into your own files without clear boundaries makes the scope hard to describe honestly.
If you build software for clients, your contracts may assign the code to them. Agency and consulting repos are often not yours to license at all.
Code from freelancers without a signed IP assignment may not belong to the company. Check the contracts for anyone who committed to the repos you plan to include.
Your own open-source repositories are already available to anyone, including labs. They add little to a private package. Their value is in the private history around them, if any.
Expect the agreement to ask you to warrant that you own the data or have the right to license it. That warranty is where an ownership mistake becomes your cost, which is why our guide to indemnities and warranties is worth reading before you scope code. General information, not legal advice. Talk to your own lawyer before you sign.
De-identification is aimed at personal data. It is not designed to catch an AWS key in a commit from 2019. That job is yours, and it comes before the buyer sees a single file.
.env files, CI variables, signed URLs, webhook tokens and internal hostnames all count.No buyer we track publishes a scoring rubric for code. These contrasts follow from what buyers say they want: connected histories that show how work was done.
| Signal | Stronger | Weaker |
|---|---|---|
| Commit messages | Explain why a change was made and reference an issue. | “fix,” “wip,” “update” on most commits. |
| Change size | Small, focused commits that each do one thing. | Huge squashed commits that hide the steps. |
| Review culture | Real review threads with questions, objections and revisions. | Approvals with no comments, or no reviews at all. |
| Links | Commits, pull requests and tickets reference each other. | Code with no trace of why anything was built. |
| Tests | Tests that change alongside the code they check. | No tests, or a test folder untouched for years. |
| Span | Several years of continuous work on a product people used. | A short burst of activity, then silence. |
Weak signals do not make a repository worthless, and you should not rewrite commit messages to look better. Describe the history honestly in the manifest and let the buyer judge it. Practitioners advise getting more than one offer precisely because buyers weigh these things differently.
The fictional company has 31 repositories on GitHub. After the review below, its manifest lists:
No price is shown on purpose: nobody can value repositories without reviewing them. The point is the shape of the manifest.
The full step-by-step for building that manifest, including the retention check and who signs internally, is on prepare your data for sale.
Eligibility rules are the programs’ own, quoted as published, checked 7 October 2026. We do not know how any program weighs code against other data.
micro1 lists 30+ employees, mature operations, documented processes, modern software tools and primarily English, with US companies prioritized. Mode lists 20+ full-time US office employees and several years of records the company owns. Grepped lists any vertical. Miro Advisory publishes a separate range for private codebases and works with software companies; we have no referral link with it. For very large or unique repositories, some labs run direct intake; Google and OpenAI both publish data partnership pages.
Independent site. Some links are referral links: if your company signs with a buyer through them, the buyer may pay us a fee. You are not charged, and we never see your data.