Verify any claim · lenz.io
Claim analyzed
Science“Scientific studies show that the risk of misidentification from partial DNA matching in familial DNA searching is very low.”
Submitted by Curious Fox 2c79
The conclusion
Open in workbench →The claim overstates what the literature shows. Some studies do find extremely low false-positive rates for one narrow error type: unrelated people being flagged as close relatives under strict methods. But familial searching also produces meaningful misidentification risks in other important scenarios, especially confusing distant relatives with first-degree relatives, so the blanket “very low” characterization is not supported.
Caveats
- The phrase “very low” is only defensible for certain protocols and a specific error type, not for familial searching overall.
- Distant relatives can be misclassified as close relatives at nontrivial rates, which materially affects real-world investigative risk.
- Error rates depend on thresholds, marker sets, database composition, and population structure, so results are not universally transferable.
Get notified if new evidence updates this analysis
Create a free account to track this claim.
Sources
Sources used in the analysis
Effective implementation across jurisdictions depends on pre-validated likelihood-ratio decision thresholds (τ), defined through empirical validation studies to control false-positive rates.[1] The reviewed literature indicates that FDS is technically mature but operationally sensitive.[1] Because familial searching produces partial matches, false positives can occur.[1] Case studies show that increasing the number of STR loci and using highly informative markers such as SE33 reduces false matches, with LR-based evaluation outperforming simple allele counting.[1] Although familial searching presents manageable risks, international experience highlights concerns about false positives, racial bias, and ethical governance that must be addressed prior to full implementation.[1]
The chance of a false positive—a non-relative achieving this less-stringent partial match threshold—greatly exceeds the probability that the same non-relative is a false exact match.[1] …Hence, owing to nontrivial false positive rates, close relatives of database entrants can be exposed to inappropriate forensic investigation when they have not in fact contributed to query profiles.[1] …Although our study suggests that an ancestry-inference procedure can potentially bound the false positive rate at values below those produced by the most serious misspecifications of allele frequencies, such reductions may continue to produce rates that are found to be intolerably high.[1]
The OSAC 2021-S-0029 “Standard for Familial DNA Searching” defines key performance metrics for familial searches, including sensitivity and specificity.[5] It states: “A reasonable specificity test would examine how many individuals remain as candidates after the statistical process (e.g., the initial LR rankings based solely upon autosomal STR loci)… this will give an estimate of how likely it would be to see a false positive.”[5] It further instructs that “Search criteria developed from sensitivity and specificity studies should be established to err on the side of minimizing false positives.”[5] The standard emphasizes use of LR thresholds: “The likelihood ratio below which a database profile specific to the relationship(s) under consideration would not be further investigated… a laboratory may decide to investigate only those candidates above a certain likelihood ratio.”[5]
The false positive rates for the EMR/EKR method averaged 5.9E-7% for parent-child pairs and 1.5E-4% for sibling pairs and were smaller than the DLR method (5.6E-3% for parent-child pairs and 5.0E-4% for sibling pairs). The final EMR calculation is a likelihood ratio where the numerator is the probability of a moderate stringency match, denoted msMatch, in the CODIS database, computed based on a false-positive rate of 8.5 × 10−5. An EMR/EKR cut off value was determined that would produce the same false positive rate (5.6E-3%).
The report, titled “The Influence of Relatives on the Efficiency and Error Rate of Familial Searching,” investigated the rate of false positives, as well as the rates of misidentification of distant relatives as first-degree relatives.[2] The study found that familial DNA searches do a good job of locating a relative if one is in the database, and conversely that a search is also unlikely to return a match that appeared to be related to the crime scene source, but in fact was not.[2] However, the results also showed that a more distant relative in the existing DNA database could have up to a 42% chance of being incorrectly labeled as a first-degree relative of the person who left the crime scene DNA.[2]
According to researchers at the University of California-Berkeley and New York University, led by Rori Rohlfs, familial DNA searching will often indicate that two people are close relatives when they are in fact distant relatives.[4] …In an experiment that tested the process of familial DNA matching in the California DNA database (using simulated genetic profiles based on publicly available data), the researchers found that cousins could be misidentified as siblings.[4] …In this experiment, “while the overall rate of false identification of unrelated individuals remains low,” the rate of false positives of African Americans was “much higher, roughly two orders of magnitude higher” than other groups.[4]
Our results agree with previous work, showing that with the prescribed methodology, false positive rates of parent-offspring and sibling identification are low, on the order of 10−5 to 10−6 (Tables 1 and 2). The simulations of unrelated individuals showed low false positive rates of parent-offspring and sibling identification. For Y-chromosome sharing first-degree relatives, the Myers protocol has a high probability of identifying their relationship; for unrelated individuals, there is a low probability that an unrelated person in the database will be identified as a first-degree relative. However, for more distant Y-haplotype sharing relatives (half-siblings, first cousins, half-first cousins or second cousins) there is a substantial probability that the more distant relative will be incorrectly identified as a first-degree relative.
In addition, the findings demonstrated that a high prevalence of homozygote genotypes in an individual's DNA profile can result in false positive familial matches as well as no matches due to limited amount of genetic information available.[5] Familial searching case studies that contained pedigree with only trees with Global Filer™ STR kit profiles resulted in fewer familial candidate matches, thereby reducing the number of false positives.[5] To maximise the discriminatory power of familial searching and to reduce the number of candidates matches generated, it is essential to use DNA marker kit chemistry that contains a minimum of 22 STR markers (21 autosomal STRs and 1 Y-STR).[5]
Random match probabilities near 1 in 1000 (10^-3) often result from the use of less discriminating tests, when the laboratory can obtain only a partial profile of one of the samples, or when one of the samples contains a mixture of DNA from more than one person.[3] …Empirical studies of laboratory error rates in RFLP and PCR-based typing generally have estimated the overall rate of false positives to be between 1 in 100 (0.01) and 1 in 1000 (0.001).[3] …Thus, when the prior odds that a particular suspect will match are very low, as might be the case if the suspect is identified during a “DNA dragnet” or database search, the probability that the samples do not match when a match has been reported can be far higher than the false positive probability.[3]
This technique has the potential to generate adventitious matches as demonstrated by simulation experiments carried out by Reid et al. and O’Connor but the search processes have been designed to minimise these issues and to provide focussed investigative leads for the police team.[6] These researchers have shown that the technique is very successful at locating parent/child or sibling pairs in large databases.[6] Similar DNA experiments described by O’Connor indicated that to reach a probability of 0.95 of finding the true parent/child pair the top 0.1% of the LR values would have to be investigated, and for the true full sibling pair, the top 3% of the LR values would have to be scrutinised.[6]
Familial DNA analysis is not without its challenges. One major concern is the potential for false hits, as this method relies on partial matches between two DNA profiles, which may implicate innocent individuals. Among many, usage of familial DNA analysis can lead to false hit. As this process is based upon the partial matches obtained by comparing two DNA profiles, this may result in misidentification if not interpreted carefully and in conjunction with other evidence.
A review article “Familial DNA searching – an emerging forensic investigative tool” surveyed U.S. crime laboratories’ use of familial searching.[7] Among 103 responding labs, those using familial DNA searching tended to adopt likelihood-ratio-based approaches to rank candidate relatives.[7] The article discusses performance validation studies such as Myers et al. 2011, which “validated a LR-based approach for searching for first-degree familial relationships in California’s offender DNA database,” and notes that the California system imposed stringent statistical thresholds to reduce false positives.[7] The authors emphasize that familial DNA searching provides investigative leads rather than definitive identification and that procedures typically combine LR thresholds with additional investigative filters to control misidentification risk.[7]
Since 2017, the DNA profiles uploaded into CODIS have used genotypes from twenty STR loci across the genome. There are a sufficient number of different alleles for each STR locus that the probability of any two random individuals matching across all twenty loci is incredibly small. Performing extended familial searching beyond close relatives is likely to generate false leads and bring individuals into unnecessary contact with law enforcement. The cited PLOS ONE study found low false positive rates for parent-offspring and sibling identification, but substantial risk that more distant relatives would be incorrectly identified as first-degree relatives.
Familial DNA searching and partial-match DNA analysis use fewer genetic markers than are used in standard testing protocols to determine a match between a forensic DNA sample and a DNA profile held in a databank.[4] By definition, then, these techniques introduce imprecision, and the potential for error, in the analysis of forensic DNA and in the criminal investigation based upon this analysis.[4] Scientists and scholars have warned that use of DNA evidence to conduct familial searching is highly susceptible not only to human error, but to fraud and abuse.[4] The full implications of the partial matching policy are presently unknowable; familial searching that utilizes genetic profiles stored in a database is a relatively new phenomenon; there is no comprehensive study of the analytic methodologies utilized, of the criminal investigation techniques employed, or of the outcomes of cases that involve the use of familial DNA searching.[4]
When fewer than thirteen alleles can be examined from a sample, it increases the possibility of a random match.[7] Furthermore, some crime scene samples contain DNA from multiple sources. All of these issues can confound the effectiveness of DNA fingerprinting as a means of identification.[7] However, in cases in which all thirteen STR loci can be examined and matched, such matches are extraordinarily reliable.[7]
Familial and moderate stringency DNA searches effectively increase the number of individuals to investigate because they include individuals who are not exact matches but possibly related to the source of the crime scene DNA.[8] This expansion of the investigation pool raises concerns about privacy, civil liberties, and potential disparate impacts on certain populations, but it also offers the possibility of solving crimes that would otherwise remain unsolved.[8]
Family DNA analysis can be problematic due to potential false matches.[7] This method can sometimes mistakenly link individuals to crimes based on partial DNA similarities.[7] Another challenge with familial DNA matching involves statistical analysis. Partial matches in DNA comparisons require careful interpretation… Failing to update these statistical techniques could lead to misinterpretations and unjust use of family DNA evidence.[7]
Moderate and low stringency searches may result in the identification of “partial matches” in which two profiles, although not an exact match, show a sufficient number of alleles in common to raise the possibility of a familial relationship. Although lower-stringency searches of DNA databases can uncover partial matches fortuitously, they are not ideal for deliberately identifying familial relationships. This is because low stringency searches can generate hundreds or even thousands of partial matches, none of which may be biologically related. The SWGDAM Ad Hoc Committee on Partial Matches (2009) recommended that only partial matches that result from a single-source forensic profile with all available core loci should be pursued in order to reduce the identification of false positives (or unrelated profiles).
In this paper we discuss the processes, utility, and governance of the familial search service in which the NDNAD is searched for close genetic relatives of an individual whose DNA profile is obtained from a crime scene.[10] The DNA profiles of close genetic relatives will exhibit a greater degree of genetic similarity to that of a true offender than unrelated individuals.[10] This is because DNA is inherited such that children receive half of their DNA alleles from each parent. The extent to which siblings share their DNA is variable, but on average two siblings might be expected to share about 65% of their DNA alleles.[10]
A UK-based academic analysis of the early deployment of familial DNA searching reports on its investigative yield and discusses false-positive risks.[8] It states that the familial DNA searching service “was introduced in 2003 and has now been applied in over one hundred serious offences.”[8] The paper notes that significant numbers of cases were successfully progressed using the technique but that each search generated candidate lists requiring substantial investigative work to eliminate unrelated individuals.[8] While it does not give a single numeric misidentification rate, it stresses that familial searching is resource-intensive precisely because initial candidate lists may contain many non-relatives whose partial profile similarity can be misleading without further testing and investigation.[8]
This use of a DNA database is referred to as “familial DNA searching” and can potentially provide an investigative lead (as opposed to exact identification).[9] Familial searching is not a match of the forensic profile to a convicted offender or arrestee in the database, but rather a search for potential relatives whose profiles are similar to the forensic profile.[9] Because the results of familial DNA searches are only investigative leads, further investigation is necessary to confirm or refute any potential relationship.[9]
Bieber et al. suggested that familial search analysis could increase the cold hit rate up to 40%.[6] Given the discrimination power afforded by current Y STR profiling kits is around 0.999, most false indirect associations can be eliminated… Thus, when using a Y STR match threshold, rarely would an investigator be incorrect in following up the lead provided by the DNA association.[6] The article notes, however, that proposed practices for identifying associations with their false-positive and false-negative rates should be understood, and when possible, Y STR typing should be performed to reduce erroneous leads.[6]
A general review on the effectiveness of forensic DNA work notes both the utility and limitations of DNA database techniques, including familial searching.[9] It explains that DNA evidence and database searching are highly effective for identification when there is a direct match, but that partial and familial matches are inherently probabilistic and require careful statistical interpretation.[9] The article highlights that misinterpretation of complex or partial DNA profiles can contribute to wrongful implication, reinforcing the need for validated statistical methods and clear guidance when using familial or partial database matches.[9]
With the standard STR profiling, this random match probability is extremely low. For example, the chance of two unrelated people having an identical 13-locus DNA profile has been estimated on the order of 1 in hundreds of trillions.[3] In practical terms, a full match gives near-certainty of identification.[3] Sometimes only a partial DNA profile can be obtained, which reduces the certainty of a match.[3]
Familial DNA searches rely on partial matches between an unknown crime scene profile and profiles in a database to infer possible relatives.[11] Scientific literature indicates that while false positive rates can be controlled using likelihood ratio thresholds and large numbers of STR markers, misidentification risks—including misclassifying distant relatives as close relatives and adventitious matches of unrelated individuals—remain non-negligible and depend strongly on database size, marker set, and demographic structure.[11]
Familial matching in consumer and forensic databases has been shown in multiple studies to produce some false positives: distant relatives or even unrelated individuals may appear statistically similar enough to be flagged as potential first-degree relatives when only a subset of loci or single-segment matches are used.[2] Segments smaller than about 7–8 centimorgans have been reported to have a high chance (tens of percent) of representing false-positive genealogical matches rather than true shared ancestry, illustrating how reliance on small or partial matches increases the risk of misidentifying relationships.[2] Although this example comes from direct-to-consumer genealogy rather than forensic STR databases, it underscores the general statistical issue that partial genetic resemblance does not always imply close kinship and can mislead relationship inference if thresholds are not conservative.[2]
What do you think of the claim?
Your challenge will appear immediately.
Challenge submitted!
For developers
This same pipeline is available via API.
Verify your AI's output programmatically.
/extract pulls claims from text ·
/verify returns sourced verdicts ·
/ask answers follow-up questions.
Continue your research
Verify a related claim next.
Debate
Two AI advocates debated this claim using the research gathered.
Argument for
Source 4 (NIJ grant report) directly demonstrates that validated EMR/EKR methods achieve extraordinarily low false positive rates—averaging 5.9E-7% for parent-child pairs and 1.5E-4% for sibling pairs—confirming that misidentification risk is very low when proper statistical thresholds are applied. Source 7 (PLOS ONE) corroborates this, explicitly stating that 'false positive rates of parent-offspring and sibling identification are low, on the order of 10−5 to 10−6,' and Source 12 (Forensic Science International: Genetics) further confirms that stringent likelihood-ratio thresholds used in validated systems like California's are specifically designed to minimize misidentification, collectively establishing that scientific studies do show the risk of misidentification from partial DNA matching in familial searching is very low under properly implemented protocols.
The Proponent's reliance on Source 4 and the close-relative rates in Source 7 overlooks that same source's explicit finding of substantial probability that distant relatives are incorrectly identified as first-degree relatives, a risk corroborated at up to 42% by Source 5. The Proponent further ignores Source 2, Source 6, and Source 13, which establish that partial-match false-positive rates remain non-negligible and demographically elevated even under likelihood-ratio thresholds, demonstrating that studies do not show the overall misidentification risk to be very low.
Argument against
Source 2 (PubMed Central), Source 5 (NYU School of Law), Source 6 (Pacific Standard), Source 7 (PLOS ONE), and Source 13 (Federal Judicial Center) demonstrate that false-positive rates for partial matches remain non-negligible, with distant relatives misidentified as first-degree relatives at rates up to 42% and substantially elevated risks for certain demographic groups. Source 11 (ScienceDirect), Source 14 (NYCLU), Source 17 (GSC), and Source 25 (LLM Background Knowledge) further establish that partial DNA matching inherently introduces imprecision and non-trivial misidentification probabilities that studies have not shown to be very low across database sizes and marker sets.
The Opponent conflates two distinct error types—misclassification of distant relatives as first-degree relatives and false positive identification of unrelated individuals—treating the former's elevated rates as evidence against the claim, when Source 7 (PLOS ONE) and Source 4 (NIJ grant report) explicitly demonstrate that the operationally relevant false positive rates for unrelated individuals remain on the order of 10−5 to 10−6, which is by any scientific standard 'very low.' Furthermore, the Opponent's reliance on Source 14 (NYCLU) and Source 25 (LLM Background Knowledge) as scientific authority is methodologically unsound, as an advocacy organization's policy comments and a background knowledge summary carry no empirical weight against the peer-reviewed validation studies in Sources 4, 7, and 12, which directly confirm that properly implemented likelihood-ratio protocols achieve the low misidentification rates the claim asserts.
Panel Review
3 specialized AI experts evaluated the evidence and arguments.
Reviewer 1 — The Logic Examiner
Sources 4 and 7 provide simulation/validation results showing very low false-positive rates for unrelated individuals being flagged as parent-offspring or siblings (on the order of 10^-5 to 10^-6), and sources 1, 3, and 12 explain that LR thresholds and marker choices are used to control false positives, but sources 2, 5, 7, and 13 also show that partial-match familial searching can yield nontrivial false positives and substantial misclassification of distant relatives as first-degree relatives (e.g., up to 42%), meaning the evidence does not support a blanket conclusion that misidentification risk from partial matching is "very low" in general. Because the claim asserts broadly that scientific studies show the misidentification risk is very low, yet multiple scientific sources in the pool indicate important, sometimes substantial, misidentification pathways under realistic conditions, the claim is mostly false as stated.
Reviewer 2 — The Source Auditor
High-authority sources such as PLOS ONE (Source 7), the NIJ (Source 4), and the Federal Judicial Center (Source 13) confirm that while false-positive rates for direct first-degree relatives are very low (10^-5 to 10^-6), there is a high, non-negligible risk (up to 42%) of misidentifying distant relatives as close relatives. Because familial DNA searching inherently involves these partial matches, the overall risk of relationship misidentification is not 'very low' across the board, making the claim only partially accurate.
Reviewer 3 — The Precision Analyst
The claim states that 'scientific studies show that the risk of misidentification from partial DNA matching in familial DNA searching is very low.' The precision issue here is the unqualified scope of 'very low' and whether this applies across all types of misidentification. The evidence is mixed: Sources 4 and 7 do show very low false positive rates (10^-5 to 10^-6) for identifying unrelated individuals as first-degree relatives. However, Sources 5, 7, and 13 explicitly show that distant relatives (half-siblings, cousins) can be misidentified as first-degree relatives at rates up to 42%, which is not 'very low.' Source 2 notes 'nontrivial false positive rates' and Source 6 shows rates for African Americans are 'roughly two orders of magnitude higher' than other groups. The claim uses an unqualified 'very low' that does not distinguish between these two distinct error types — the risk of misidentifying an unrelated person as a relative (genuinely low) versus the risk of misclassifying a distant relative as a first-degree relative (substantially elevated). The claim's wording overstates the consensus by applying 'very low' universally when the scientific literature shows this only holds for one specific error type under optimal conditions, while other forms of misidentification remain non-negligible to substantial.