Dog DNA Breed Identification Explained: How the Pie Chart Is Built
How a dog DNA test goes from cheek swab to breed percentages. Reference samples, segment assignment, and why a 10 percent Beagle reading has a confidence interval.
The breed pie chart on a dog DNA report looks like a clean answer. It is actually a model output with confidence intervals attached, built by comparing segments of your dog’s genome to reference profiles for each recognized breed. Understanding how that chart is constructed is the difference between reading it correctly and over-reading it.
Building the reference panel
Before any consumer dog is tested, the company builds a reference panel. For each recognized breed, they genotype a set of confirmed purebred dogs, typically dozens to several hundred per breed for well-represented breeds. These reference samples have documented pedigrees and known ancestry, and they define the genetic profile of that breed: which versions of which variants are common, which segments of the genome are characteristic.
Embark covers 350+ breeds. Wisdom Panel covers 350+ breeds. Both publish peer-reviewed work on their reference panels and update them as new breed data becomes available.
The depth of the reference panel matters. A breed with 200 reference samples will be called more reliably than a breed with 30, because the algorithm has a richer picture of what genetic variation within that breed actually looks like.
Reading your dog’s genome in segments
The test genotypes your dog across 200,000+ marker positions. The algorithm then walks along your dog’s genome in segments, asking, for each segment, which reference breed profile it most closely matches.
For a purebred dog, almost every segment matches the same reference profile. The pie chart shows close to 100 percent of one breed.
For a mixed-breed dog, the segments get carved up. One stretch of chromosome 4 might match a Beagle profile strongly; another stretch matches a Labrador profile; another matches a generic Eastern European village dog cluster. The pie chart is built by summing how many segments matched each reference, weighted by how confident each match was.
Why “10 percent Beagle” has a confidence interval
A 10 percent Beagle reading in a mixed-breed dog means: across the segments the algorithm walked, roughly 10 percent of the genetic real estate was best explained by the Beagle reference profile.
It does not mean exactly one in ten of your dog’s grandparents was a Beagle. The algorithm is probabilistic. A 10 percent call could come from one true Beagle ancestor a few generations back, or from a distant relative breed sharing segments with the Beagle reference profile, or from statistical noise in segments where multiple breed profiles fit similarly well.
Both Embark and Wisdom Panel internally compute confidence intervals on these calls. Embark’s report shows them on a per-segment basis if you dig in. Wisdom Panel is more conservative about exposing them, but the underlying math is the same.
For the broader accuracy picture, see dog DNA test accuracy.
The “supermutt” and village-dog categories
Some segments of a mixed-breed dog’s genome do not strongly match any single breed reference. They might match several distantly related breeds equally well, or they might fall into a generic village-dog or street-dog cluster that does not correspond to a recognized breed.
Both companies handle this transparently. A “supermutt” or “village dog” percentage on your report is honest: the algorithm could not assign these segments to a specific breed with confidence, so it grouped them. This is common in rescue dogs with mixed or unknown ancestry.
What this means for reading your report
Take the major calls (above 15 percent) seriously. Take the mid-range calls (5 to 15 percent) as credible. Treat sub-5-percent slivers as statistical hints rather than confirmed ancestry. And read any “supermutt” or generic bucket as a real, informative category, not as a failure of the test.
Why Embark and Wisdom Panel sometimes disagree
Different reference panels and different segmentation algorithms produce slightly different best-fit assignments at the edges. Same dog, same swab method, two slightly different pie charts. The headline calls usually agree; the long-tail calls sometimes do not. Neither is wrong. They are different best guesses against slightly different reference data.
See Embark vs Wisdom Panel for the head-to-head context.
Related reading
Dog DNA test accuracy, how dog DNA tests work, what dog DNA tests detect, and the pet DNA testing guide.
Sources
- Whole-genome sequencing analysis of dog breeds — Nature Communications Primary (accessed 2026-04)
- Embark Research and Publications — Embark Veterinary (accessed 2026-04)