Sell Data to AI
Home Data Asset Score Pricing API documentation US labs and CROs list
For brokers
How to become an AI data broker Data broker business model Buyer programs compared Qualify a company AI training data companies
For data companies
Firmographic data providers Company data API
Seller guides
How to sell data to AI companies Is it legal? FAQ and glossary About
Check domain/company
Data types: code

Selling Your Codebase and Git History to AI Companies

For a software company, the repository is often the most structured record of real work it owns. This guide covers what buyers say they accept, the published ranges, and the three checks that come before any export: ownership, open-source licenses and secrets.

Last checked: 7 October 2026. Buyer figures are quoted as published on that date.

$10K to $1M+Miro Advisory’s indicative range for private codebases
~$5,000per code repository in the closure market, as cited by Troveo
3 hostsGitHub, GitLab and Bitbucket repos with history are listed by buyers
60 to 90 daystypical time to close, practitioners say

What “selling your codebase” usually means

When people search for how to sell source code to AI companies, they usually picture a sale of the code itself. In the programs we track, that is rarely the shape of the deal. A buyer licenses a copy of an agreed set of repositories for AI training or evaluation. Mode, for example, says it buys “an agreed copy” and that originals stay with the company. micro1 says the company keeps ownership of its underlying data. Both are quoted as published, checked 7 October 2026.

Keeping ownership does not mean you keep every right. An exclusive license, a resale right or a broad scope of use can limit what you do with the same code later. Those terms matter as much as the price. Our guide to exclusivity and resale rights walks through them.

Why history matters

The snapshot is the least interesting part

The data sources buyers list name “GitHub, GitLab and Bitbucket repositories with history.” The word that matters is history. A snapshot shows what the code is. The history shows how people got there.

Commits and diffs

Each commit pairs a change with a message about why it was made. Years of small, well-described commits show how engineers break problems into steps.

Code review threads

Review comments record disagreement, trade-offs and corrections. That is the “decision-making” layer that buyers describe wanting from company data.

Issues linked to fixes

A bug report, the investigation, the fix and the closing note form one connected record of work. Linked Jira or platform issues make the code far easier to read.

Tests and CI results

Tests show what “correct” meant to your team. Failing and passing builds show how a change was checked before it shipped.

Reverts and incident fixes

Rolled-back changes and post-incident patches are records of mistakes and recovery. They are rare in public code and often the most instructive part of a history.

Design notes and docs

Architecture decision records, READMEs and runbooks explain intent. They connect your code to your internal documentation.

One practical catch. A plain git clone carries commits, branches and tags. It does not carry pull request reviews, merge request discussions or issues. Those live on the hosting platform and need a separate export through its API or admin tools. If review threads are part of what you offer, plan that export from the start. The same goes for linked tickets: software companies that hold code, Jira and Slack for the same projects have the most connected history, which is the angle of our software companies guide.

Dates may matter too. Buyers training models care whether code was written by people. A long history from before AI coding assistants became common is easier to vouch for than recent code that may be partly generated. We have no published buyer statement that prices this. Treat it as a question to ask each buyer: do commit dates, or a split before and after a given year, change how they value the repository?

Published numbers

What has been published about code and company data

None of these figures is a quote for your repositories. They are ranges and reported prices, each tied to its source.

SourceFigureWhat it covers
Miro Advisory (its site)$10K to $1M+Private codebases, labeled indicative. Operating datasets are listed separately at $100K to $1M+.
Troveo (cited in 2026 closure-market coverage)about $5,000Per code repository from shut-down startups, as cited by Troveo. Forbes covered this market on 16 April 2026.
micro1 (its pages)$100k+ / $500k+ / $1M+“Qualified,” “large-scale” and “highly unique” company data packages. Not code-specific.
Mode (its site)$100K to $5MCompany data generally. Not code-specific.
Grepped (its site)$20K to $5MAny vertical. Not code-specific.

Last checked: 7 October 2026, from each company’s own published pages; Troveo’s figure as cited in 2026 news coverage. “Up to” and “+” ranges describe what a program has published, not what a typical company receives.

The gap between $5,000 per repository and seven-figure package ranges is not a contradiction. The closure market sells single repositories from companies that no longer exist, often without the surrounding tickets or reviews. Program ranges describe larger packages from operating companies. Practitioners say raw data is the cheapest form, that evaluations built on the data are worth roughly ten times raw, and that full training environments reach six to eight figures but need heavy engineering. Our page on raw data versus evaluations versus environments explains those tiers.

Whatever the tier, practitioners give one rule: never send the full repository before you have a price. Share a manifest of what exists and a few samples, and get more than one offer.

Open-source license traps

You can only license what is yours to license

Most commercial repositories contain code the company did not write. Before anything goes into scope, sort the tree by who owns each part.

Vendored dependencies

Copied libraries, node_modules, vendor/ folders and bundled SDKs belong to their authors. Exclude them and list them in the manifest so the buyer sees what was left out and why.

Copyleft code

Code under GPL or AGPL carries conditions on redistribution. Whether licensing it for training triggers them is a question for your lawyer, not for a sales call. Keep it out unless you have clear advice.

Notice-bearing licenses

MIT, Apache 2.0 and BSD code is permissive, but its notices still apply. Mixing it into your own files without clear boundaries makes the scope hard to describe honestly.

Client-owned work

If you build software for clients, your contracts may assign the code to them. Agency and consulting repos are often not yours to license at all.

Contractor code

Code from freelancers without a signed IP assignment may not belong to the company. Check the contracts for anyone who committed to the repos you plan to include.

Already-public code

Your own open-source repositories are already available to anyone, including labs. They add little to a private package. Their value is in the private history around them, if any.

Expect the agreement to ask you to warrant that you own the data or have the right to license it. That warranty is where an ownership mistake becomes your cost, which is why our guide to indemnities and warranties is worth reading before you scope code. General information, not legal advice. Talk to your own lawyer before you sign.

Secrets scanning

Scan the whole history, then rotate what you find

De-identification is aimed at personal data. It is not designed to catch an AWS key in a commit from 2019. That job is yours, and it comes before the buyer sees a single file.

Risk and mitigation

Where code deals go wrong, and the usual fix

Risk
Mitigation
A live credential reaches the buyer in old history.
Full-history scan, rotation of every finding, rewrite only on the export copy.
Client-owned or contractor code ends up in scope.
Repo-by-repo ownership review against contracts before the manifest is shared.
Copyleft or vendored code is licensed as if it were yours.
Exclude third-party paths by rule and list them in the manifest.
An exclusive license blocks future use of your own history.
Ask for non-exclusive or time-limited terms; read resale and downstream-use clauses.
Strong versus weak

What a reviewer is likely to notice in your samples

No buyer we track publishes a scoring rubric for code. These contrasts follow from what buyers say they want: connected histories that show how work was done.

SignalStrongerWeaker
Commit messagesExplain why a change was made and reference an issue.“fix,” “wip,” “update” on most commits.
Change sizeSmall, focused commits that each do one thing.Huge squashed commits that hide the steps.
Review cultureReal review threads with questions, objections and revisions.Approvals with no comments, or no reviews at all.
LinksCommits, pull requests and tickets reference each other.Code with no trace of why anything was built.
TestsTests that change alongside the code they check.No tests, or a test folder untouched for years.
SpanSeveral years of continuous work on a product people used.A short burst of activity, then silence.

Weak signals do not make a repository worthless, and you should not rewrite commit messages to look better. Describe the history honestly in the manifest and let the buyer judge it. Practitioners advise getting more than one offer precisely because buyers weigh these things differently.

Questions to ask the buyer

Seven questions specific to code

Illustrative example, fictional, not an offer

A fictional 45-person SaaS company scopes its repositories

The fictional company has 31 repositories on GitHub. After the review below, its manifest lists:

  • In scope: 4 product repositories with history from 2018 to 2026, pull request reviews exported through the API, and linked Jira issues for the same projects.
  • Excluded: 3 client-funded repositories (contracts assign the code to clients), the infrastructure repository, the payments service, all vendored and generated code, and 2 repos with contractor commits lacking IP assignments.
  • Cleaned: every secret found in the history rotated, test fixtures with real customer rows removed, commit authors replaced with consistent pseudonyms.

No price is shown on purpose: nobody can value repositories without reviewing them. The point is the shape of the manifest.

The full step-by-step for building that manifest, including the retention check and who signs internally, is on prepare your data for sale.

Where code goes

Which programs take code, as published

Eligibility rules are the programs’ own, quoted as published, checked 7 October 2026. We do not know how any program weighs code against other data.

micro1 lists 30+ employees, mature operations, documented processes, modern software tools and primarily English, with US companies prioritized. Mode lists 20+ full-time US office employees and several years of records the company owns. Grepped lists any vertical. Miro Advisory publishes a separate range for private codebases and works with software companies; we have no referral link with it. For very large or unique repositories, some labs run direct intake; Google and OpenAI both publish data partnership pages.

Independent site. Some links are referral links: if your company signs with a buyer through them, the buyer may pay us a fee. You are not charged, and we never see your data.

FAQ

Questions about selling source code

Can a company sell its source code to AI companies?
Yes, if the company owns the code or holds the rights to license it. Buyer programs list GitHub, GitLab and Bitbucket repositories with history among the sources they accept. In practice you usually license a copy for training or evaluation rather than transfer ownership. Check what your agreement says about exclusivity and resale before you sign.
Do I lose ownership of my code if I sell it for AI training?
Not usually. Mode says it buys "an agreed copy" and that originals stay with the company, and micro1 says the company keeps ownership of its underlying data (both as published, checked 7 October 2026). Ownership and use rights are different things, so read the scope, exclusivity and resale clauses closely.
What about the open-source code inside my repositories?
Code you vendored or copied from open-source projects is licensed to you, not owned by you. Copyleft licenses such as GPL or AGPL and notice requirements in MIT, Apache 2.0 or BSD licenses all raise questions. The safest scope excludes third-party directories and lists them in the manifest. Ask your lawyer about anything you are unsure of.
Do I need to remove secrets if the buyer de-identifies the data?
Yes. De-identification is aimed at personal data, not API keys or passwords. Scan the full git history, not only the latest commit, and rotate every credential you find before the export. A rotated key is harmless even if it slips through; a live one is not.
How much is a codebase worth to an AI buyer?
Nobody can price a codebase without seeing it. Miro Advisory publishes an indicative $10K to $1M+ for private codebases, and Troveo cites about $5,000 per code repository in the closure market. Practitioners say evaluations built on data are worth roughly 10 times raw data. Share a manifest and samples, not the full repository, before you have a price.
Is the latest code enough, or do buyers want the history?
Buyer programs specifically list repositories with history. The commit trail, review threads and linked issues show how engineers reasoned and fixed problems, which a single snapshot cannot. Note that pull request reviews and issues live on the hosting platform and need a separate export from the git clone.
Can a company that is shutting down sell its repositories?
There is a closure market for this. Troveo cites about $5,000 per code repository in that market. The same checks apply: ownership, open-source code, client contracts and secrets. Founders winding down should also ask their lawyer who has authority to sign once the company is in liquidation or under creditor control.
Should infrastructure and security code be included?
Think carefully first. Infrastructure, authentication and payment code describe how your systems can be reached, and they often contain the most secrets. Many sellers leave them out of a first deal and offer product code with its history instead. Decide with your security lead.