← all articles

What a genealogy DNA test actually does with your sample

The sample is the easy part to understand

You spit in a tube, or swab your cheek, and mail it off. The lab’s job is straightforward: extract DNA from the cells in your saliva, then run it through a genotyping chip. Most consumer genealogy companies don’t sequence your entire genome. They use a microarray that reads several hundred thousand specific positions in your DNA, called SNPs (single nucleotide polymorphisms), chosen because they vary between people and correlate with ancestry and some health traits. That’s a small fraction of your roughly 3 billion base pairs, but it’s enough to identify you individually, place you in a family tree, and flag a meaningful set of health-related markers.

The part people underestimate isn’t the lab process. It’s what happens after the chip finishes reading, when a physical sample becomes a searchable digital dataset that outlives the swab by decades.

Your raw data becomes a permanent, comparable record

Once your SNPs are read, the company stores your genotype as a data file. This file gets compared against everyone else in their database to find genetic relatives, a process that works by looking for shared DNA segments (matching stretches of SNPs) between you and other users. That’s the core genealogy feature, and it’s also the reason your data has value beyond your own curiosity.

Two things matter here from a privacy standpoint. First, this isn’t a one-time comparison. Your data sits in the database and gets checked against every new person who joins, indefinitely, unless you actively delete it. Second, genetic relatedness isn’t something you can opt out of after the fact. If a half-sibling you’ve never met uploads their own results, the match will surface whether or not you wanted it to. Your genome describes your relatives too, and their decision to test can expose your existence, your paternity history, or a family secret you didn’t put there.

Companies keep the physical sample longer than most users assume

Many genealogy companies retain your actual biological sample, the tube of saliva or the swab, in a lab facility after processing, sometimes for years, unless you specifically request destruction. The retention exists partly for quality reasons (rerunning a test if the chip fails) and partly because a stored sample lets the company reprocess it later with newer chip technology without asking you to spit again.

The threat model here isn’t a dramatic one. It’s accumulation. A stored physical sample is a static asset that sits in a freezer, subject to whatever policies the company has now and whatever policies get adopted later if the company is acquired, changes leadership, or gets subpoenaed. Check the specific company’s retention and destruction policy before testing, because it varies by provider and by test type, and it’s usually stated plainly in their privacy policy or FAQ rather than buried.

Health inferences go further than most people expect

Ancestry-focused kits and health-focused kits read overlapping sets of markers, but companies that offer health reports interpret specific SNPs against published research to estimate things like carrier status for certain inherited conditions, or predisposition markers for some diseases. These are statistical associations, not diagnoses, and the accuracy depends heavily on which markers the chip reads and how the company’s algorithm weights them.

The privacy-relevant point isn’t the science, it’s the categorization. Once your data is tagged with health-related inferences, it becomes a more sensitive dataset than a plain ancestry match. Some jurisdictions treat genetic health information under separate legal protections from general consumer data, and insurers and employers in many places are restricted from using it, but the restrictions and their enforcement vary by country and by context, so read the specific policy that applies to you rather than assuming a blanket rule. This is a place where the details of your jurisdiction change the actual risk, and generic reassurance isn’t a substitute for reading your own company’s policy.

Third-party sharing usually requires a separate opt-in, but check what “opt-in” covers

Reputable genealogy companies generally don’t sell your raw genetic data to third parties without explicit consent. Where sharing happens, it’s typically for research: pharmaceutical or academic partners paying for access to aggregated or de-identified datasets to study disease patterns across large populations. The opt-in for this is usually presented separately from the basic terms of service, often during account setup or as a later prompt, and it’s worth reading rather than clicking through.

The nuance worth understanding is what “de-identified” means in practice. Genetic data is unusually hard to fully de-identify, because your genome is inherently unique to you and, as covered above, related to identifiable people through family matching. Removing your name from a dataset doesn’t fully sever the connection back to you if enough distinguishing genetic and demographic detail remains attached. This doesn’t mean research sharing is inherently unsafe, but it does mean the privacy protection rests more on the company’s access controls and contractual limits on the research partner than on the data itself being unlinkable.

Law enforcement access exists and works differently across databases

Some genealogy databases allow law enforcement to search uploaded profiles, typically raw data files uploaded by investigators from crime-scene evidence, to find genetic relatives of an unknown suspect through the same matching process used for family research. Other companies have policies against this and require a valid legal order before disclosing any user data, and the details of what’s permitted differ by company, by database, and by the specific opt-in settings the platform offers to its users.

This is worth understanding as a mechanical fact about how these databases work, not as a reason to worry about your own testing decision by default. If you care about whether your profile is discoverable this way, the actionable step is to read the specific company’s law enforcement matching policy and its opt-in or opt-out settings for that feature, since several major providers let users control this directly in account settings.

What you can actually control

You generally have a few concrete levers, though which ones exist depends on the company:

  • Choosing not to test. , which is the only complete way to keep your own genotype out of a company’s database, though it doesn’t prevent relatives from exposing shared segments of your DNA through their own tests.
  • Requesting sample destruction. after processing, so a physical, reprocessable copy doesn’t sit in storage.
  • Deleting your account and data. from the company’s database, which most major providers offer through account settings, though you should check whether this also removes your data from any research datasets it was already included in, since that may not be reversible after the fact.
  • Opting out of law enforcement matching and third-party research sharing individually, where the platform offers those as separate toggles rather than bundling them into a single consent.
  • Downloading and then deleting your raw data file. if you want to preserve your own copy for uploading to other tools later, on your own terms, before removing it from the testing company’s database.

None of these options make your genetic footprint disappear once relatives have also tested, and none of them are a substitute for reading the specific company’s policy on retention, sharing, and law enforcement access before you send a sample. The tradeoffs are real in both directions: family tree research and health insights come from the same data pipeline that creates this exposure, and the honest answer is that you’re deciding how much of that pipeline you’re comfortable with, not eliminating it.

If you want a closer look at how consumer data practices affect decisions like this one, browse more explainers at The Privacy Wire.

from the team
Want a real mobile IP, not a datacenter VPN endpoint?

Shared VPN exit nodes get flagged and blocked. Singapore Mobile Proxy runs real 4G/5G mobile IPs that give you a residential-grade address carriers still trust.

see how it works →
read on
More from The Privacy Wire

VPN and tool reviews, realistic opsec guides, and privacy news for people who want to protect their data.

browse all articles →