To count lines of code, run a free counter such as cloc or tokei over your repository, then remove everything your team did not write: dependency folders, copied libraries, generated code, minified files and build output. What is left is your count of original code. That is the number worth knowing, because a raw total can be inflated many times over by code that came from somewhere else.
Why does the raw total mislead?
Because most of what sits in a typical repository is not yours. Black Duck’s 2026 Open Source Security and Risk Analysis found that the average audited application pulled in about 1,180 open-source components, and its 2025 report said about 70% of scanned code had open-source origins. If those libraries are checked into the repository, a naive count adds them to your total.
Suppose a mobile app with two years of commits: a raw count says 400,000 lines, but 330,000 of them are a copied SDK, a vendor folder and a minified bundle. The honest figure is the 70,000 your team wrote. That is the one that reflects real work, and it is the one a buyer will look for.
Which free tools count lines of code?
Two well-maintained, open-source tools cover nearly every language:
| cloc | tokei | |
|---|---|---|
| What it reports | Blank, comment and code lines by language | Files, lines, code, comments and blanks by language |
| Language coverage | Many programming languages | Over 150 languages |
| Speed | Fine for most repositories | Very fast on large repositories |
| Excluding folders | Yes, by directory or file pattern | Yes, by pattern, and it respects .gitignore |
| Project page | github.com/AlDanial/cloc | github.com/XAMPPRocky/tokei |
In its own words, “cloc counts blank lines, comment lines, and physical lines of source code in many programming languages,” and its latest release at the time of writing is v2.10. tokei describes itself as a program that “displays statistics about your code” and supports “over 150 languages.”
Download either one from its official page. Do not run a counting script that a stranger sends you on a machine with your code or credentials. A fair buyer does not need you to run anything for them.
If you are not technical, any former developer can produce a count in a few minutes, and a rough estimate is enough to start a conversation.
What should you exclude from the count?
Exclude anything your team did not write or that a machine produced. The common culprits:
- Dependency folders: node_modules, vendor, Pods, packages, .venv, bower_components.
- Vendored dependencies: libraries copied into folders such as third_party, lib, external or libs.
- Generated code: API clients generated from a schema, protobuf output, ORM migrations generated automatically, files with headers like “auto-generated, do not edit.”
- Minified and bundled files: .min.js, .min.css, files in dist, build, out or public/assets.
- Lock files and data: package-lock.json, yarn.lock, large JSON or CSV fixtures, SQL dumps.
- Copied templates: theme or admin templates you bought or downloaded and barely changed.
Both tools let you exclude directories by name. A good habit is to run one count on everything and a second count with exclusions, then keep both numbers. The gap between them tells its own story about how much of the repository is third-party code.
How do you check the count is honest?
Do a quick sanity check with your git history. Git can show how many lines each author added over time. Because code gets rewritten, that total is usually larger than what survives today, so your original-code count should sit below it. Then look at the few commits that added the most lines: if they are “add library” or “regenerate client” commits, the code they brought in is probably still in your total and belongs in the exclusions. A single commit that added 150,000 lines in one go is rarely hand-written.
Two other quick checks help. Open the five largest files in the count; if any is a single enormous generated file, exclude it. And look at the language breakdown; if a language appears that nobody on the team wrote, it probably came with a library.
Our guide on why git history adds value explains what else that history shows.
How do you turn the count into a size bucket?
You do not need a precise number. Our form and our free code value check ask for a size bucket, because the difference between 52,000 and 58,000 lines rarely matters. What matters is the order of magnitude and the share written by your team.
A simple way to describe your codebase:
- Original lines of code, rounded, from the count with exclusions.
- Main languages, from the language breakdown.
- Share written by your team, as a rough percentage of the raw total.
- Number of repositories, if the product was split into frontend, backend, mobile and infrastructure.
If the product lived in several repositories, count each one separately and add up the original lines. Note shared code only once: a common library used by both the web app and the mobile app should not be counted twice. Infrastructure code, such as Terraform files, Dockerfiles and CI pipelines, is real work too, so include it, but list it on its own line so the main application figure stays clear.
Also note what is not code but still matters. Tests, migrations written by hand and build scripts all count as your team’s work. Documentation, ticket exports and design files are not lines of code, but they add to what a buyer receives, as our guide to docs, tests, tickets and designs explains.
That four-line summary is often enough for a first conversation. It also fits the one-page overview recommended in our guide to preparing a codebase for sale.
Why do original lines matter more than total lines?
Because original code is what is scarce. Public open-source code is already widely available: GitHub reports about 630 million repositories, and the largest open code dataset, The Stack v2, is built from public repositories. If your team’s code was never published, it is not in those collections at all.
That is why we pass on repositories that are mostly copied libraries or generated code, and why the libraries inside a good repository neither help nor hurt much. Our guide to open source inside your codebase covers how that third-party code is treated in a sale. Size is also only one of several factors; how codebase value is measured explains the others, such as originality, history, documentation and tests. How we value a codebase is explained on the call.
What to do next
- Run cloc or tokei twice on your repository: once on everything, once excluding dependency, vendor, generated and build folders.
- Note the original line count, the main languages and the rough share your team wrote.
- Then send us a few details for a free code valuation; a size bucket is enough, and no code is needed.
Frequently asked questions
What is a good free tool for counting lines of code?
Two well-known free, open-source options are cloc and tokei. cloc counts blank lines, comment lines and lines of code by language, and tokei does the same quickly across more than 150 languages. Download them yourself from their official project pages. Both run locally and send nothing anywhere, and both let you exclude folders such as node_modules and vendor.
Should I count comments and blank lines?
Report them separately, not mixed in. Both cloc and tokei split their output into code, comments and blanks, and the code column is the one people usually mean by lines of code. Comments still matter, since well-commented code is easier to learn from, so it is fine to mention a healthy comment count alongside the main figure.
Does more lines of code always mean more value?
No. Size is one factor among several. A smaller, original, well-tested codebase with full git history can be more useful than a larger one stuffed with copied libraries and generated files. That is why the count should cover only what your team wrote. How we value a codebase is explained on the call, and size is only part of it.
Do I need to send you the count or the code?
Neither, to start. Our form asks for a rough size bucket, and you can give an estimate. No code changes hands before a signed written agreement, and we never ask anyone to install or run anything on their computer. Counting is something you do yourself, on your own machine, if and when you choose to.
Find out what your old code is worth right now.
Tell us about the product in about a minute. No code needed. We review the details and come back with a cash offer or a plain no.
Get my free code valuation →