Deep article

Personal Data in Old Code: What to Find and Remove Before You Sell

Old code often carries personal data outside the database, in seed files, test fixtures, logs, screenshots and hard-coded customer records. Find it and remove it before any transfer; a code buyer should never need databases, user records or customer data, and we never take them.

5 min readPublished October 11, 2026By Alex Drew, Founder and CEO, Odys Global

Personal data in old code rarely sits where people expect. The production database is the obvious place, and it should simply stay out of any sale: we never take databases, user records, customer data or chat logs. The harder part is the personal data that leaked into the repository itself, in seed files, test fixtures, logs, screenshots, ticket exports and hard-coded records. That is what you need to find and remove before any code changes hands.

This is general information, not legal advice. Data protection duties depend on where you and your former users are, so ask a qualified lawyer about your specific case.

What counts as personal data?

More than most developers assume. The EU’s GDPR defines personal data as “any information relating to an identified or identifiable natural person,” and lists examples such as “a name, an identification number, location data, an online identifier.” So an email address, a phone number, an IP address in a log line, a user ID tied to an account or a home address all count.

The UK regulator uses the same definition, and the ICO’s guidance adds that pseudonymised data remains personal data. In California, the Attorney General’s CCPA page describes personal information as information that “could reasonably be linked with you or your household,” with examples including email addresses, records of products purchased and browsing history.

Swapping names for IDs does not solve the problem. GDPR Recital 26 says pseudonymised data that could be linked back with additional information is still information on an identifiable person, and that data protection applies to everything except truly anonymous information.

Where does personal data hide in a codebase?

In a typical old SaaS or app repository, check these places:

Location What it often contains Usual fix
Seed files and SQL dumps (seeds.rb, seed.sql, dump.sql) Copies of real users loaded for local development Replace with generated fake data or remove
Test fixtures (fixtures/, fixtures, factories) Real customer records copied in to reproduce a bug, sometimes with password hashes Replace with synthetic records
Logs and debug output (logs/, *.log, crash reports) Emails, IP addresses, request bodies, tokens Delete; logs add no value to a sale
Hard-coded records “If email == customer@realcompany” special cases, VIP lists Replace with placeholders
Notebooks and scripts Data exports, analysis output cells Clear outputs, remove data files
Screenshots and design files Real user profiles in mockups or bug reports Swap in mock data or remove
Ticket exports Customer names, quoted support emails, attachments Redact the customer parts, keep the technical content
Backups inside the repo (backup/, .bak, .zip) Old database snapshots Remove completely

Mobile apps add their own spots: analytics test files, crash reports bundled in folders, and sample JSON responses captured from a real account.

Why does it matter for a sale?

There are two reasons, and only one of them is about the buyer.

The first is your own obligations. GDPR says personal data must be “limited to what is necessary”, the principle of data minimization. Customer records sitting in a fixture file years after a product closed are hard to justify under that rule. Breaches of GDPR’s core principles can bring fines of up to 20 million euros or 4% of worldwide annual turnover, whichever is higher. The practical point is not fear; it is that cleaning the repository is the right thing to do anyway.

The second is the sale itself. Code is useful for AI training and research because of how it was built, not because of who used it. So personal data adds nothing and only creates risk. That is why our terms are simple: personal data is removed before transfer, and we help.

How do you find personal data in a repository?

You do not need special software to start. A careful search covers most of it:

  1. List the data folders. Look for folders named data, seeds, fixtures, samples, dumps, exports, backup and logs.
  2. Search for patterns. Search the repository for the @ sign followed by real domains, phone number patterns, street words such as “Street” or “Ave,” and your largest customers’ names.
  3. Check large files. Big .sql, .csv, .json and .xlsx files are often data, not code.
  4. Search the history too. A dump deleted years ago is still in old commits, just like a deleted password. Our guide on removing secrets from code and git history explains why and how the same history rewrite works for data files.
  5. Review attachments. Open the screenshots, PDFs and design files and look for real people.

Do this search on your own machine. Do not upload files that may hold customer data to online scanners or converters. Write down what you found and where. That list becomes the cleaning plan.

If you find real customer data in a repository that was ever public, or shared with people who should not have had it, ask a lawyer whether you have a duty to report it.

How do you remove it without breaking the code?

Replace, do not just delete, where the code depends on the file. A test that loads fixtures/users.json will fail if the file disappears, so swap in synthetic users with fake names and addresses on reserved test domains. Free libraries such as Faker exist in most languages for exactly this.

Where nothing depends on the data, as with logs, dumps and backups, remove the files and then remove them from history with git-filter-repo. This is part of sanitization, the cleaning step before transfer. With us, it happens after the written agreement is signed and before transfer, and we help you do it. We never ask you to install or run anything on your computer, and you should not run tools a stranger sends you on a machine that holds your code.

What about names in commits and code comments?

Commit metadata contains the name and email of each developer, and comments sometimes say “ask Maria about this.” That is a different situation from customer data: it is your former team, and it is part of the history that shows how the software was built. Discuss on the call how author details should be handled before transfer, for example by replacing emails with placeholders, so that you stay comfortable with your duties to former staff.

What stays out of a sale entirely?

Some things are never part of what we buy, whatever their condition:

  • Production databases and backups of them.
  • User records, customer lists and CRM exports.
  • Chat logs, support inboxes and recorded calls.
  • Payment data of any kind.

What a buyer does want is described in our guide to docs, tests, tickets and designs, and the full list of checks is on our methodology page. If you want the whole preparation in one place, use the checklist for preparing a codebase for sale.

What to do next

  • Search your repository and its history for data folders, large data files, real email domains and customer names.
  • Plan to replace fixture and seed data with synthetic records, and to delete logs, dumps and backups.
  • Then send us a few details for a free code valuation; no code or data is needed to start.

Frequently asked questions

Is a test email address in my code personal data?

It depends whose it is. A made-up address on a reserved test domain identifies nobody. A real customer's or employee's email address is personal data under GDPR, because an email identifies a person directly. In California, email addresses are also listed as personal information. Replace real addresses with clearly fake ones, and treat team members' addresses in commit metadata as a separate question to discuss.

If I replace names with user IDs, is the data anonymous?

Usually not. GDPR calls that pseudonymisation and says pseudonymised data that can be linked back to a person with extra information is still personal data. Only truly anonymous data falls outside data protection law. For a code sale, the simpler path is to remove real records from fixtures, seeds and logs altogether rather than trying to disguise them.

Do you need my production database to value the code?

No. We never take databases, user records, customer data or chat logs. The value is in the source code, its history, docs, tests and tickets, not in the people who once used the product. Keep the database out of the transfer entirely, and handle it under your own data protection obligations, which may mean deleting it or keeping parts of it for a legal retention period. Ask a lawyer about your specific duties.

What about personal data in old Jira or GitHub Issues exports?

Ticket exports are useful, but they can contain customer names, emails, screenshots and pasted support conversations. Before transfer, remove or redact tickets that quote customers, strip attachments that show customer screens, and replace real contact details. Keep the technical content, such as bug descriptions, decisions and acceptance criteria, which is the part that adds value.

Your next move

Find out what your old code is worth right now.

Tell us about the product in about a minute. No code needed. We review the details and come back with a cash offer or a plain no.

Get my free code valuation →
About 60 secondsContract before any code100% confidential

Owner situations

Value my code →