Facial recognition software's accuracy varies sharply by algorithm and by whose face it is matching. The National Institute of Standards and Technology's testing finds top systems accurate on balanced photo sets, but with false-positive rates 10 to 100 times higher for Asian and Black faces than white faces, per its December 2019 study of 189 algorithms.
What does "accurate" mean for facial recognition software?
NIST's Face Recognition Vendor Test, or FRVT, is the federal government's ongoing benchmark for the technology, evaluating algorithms from developers worldwide on standardized image sets. It tests two different tasks: one-to-one "verification," which checks whether two photos show the same person, and one-to-many "identification," which searches a photo against a gallery that can hold tens of millions of enrolled faces.
Each task produces its own error rates. A false match, or false positive, happens when the software says two different people are the same person. A false non-match, or false negative, happens when it fails to match two photos of the same person. Police facial-recognition searches typically run one-to-many identification against a mugshot or driver's-license database — the setting where FRVT has documented the largest demographic gaps, per NIST.
How accurate are the best algorithms, per NIST's testing?
The most accurate commercial algorithms score well on NIST's standardized tests, but accuracy is not uniform across the population those algorithms are matching. In its December 2019 study — which tested 189 algorithms from 99 developers against more than 18 million images of 8.49 million people, drawn from State Department, Department of Homeland Security, and FBI operational databases — NIST found that one-to-one false-positive rates "often ranged from a factor of 10 to 100 times" higher for Asian and African American faces than for Caucasian faces.
In one-to-many searches, the setting closest to a police database lookup, the same study found higher false-positive rates for African American women specifically, using a test gallery drawn from an FBI database of 1.6 million domestic mugshots, according to NIST's report. NIST cautioned that results varied by algorithm: some developers' software showed little to no demographic differential, while others showed large gaps, and the agency did not identify a single cause common to all of them.
Why do error rates differ by race, age, and sex?
NIST's own summary of the study does not assign a single cause, noting only that performance varies by algorithm and that some developers' systems showed far smaller demographic gaps than others. Outside researchers who study the technology have pointed to the training data and design choices behind individual algorithms. MIT Media Lab researcher Joy Buolamwini told NPR that facial recognition "has already been shown in study after study to fail people of color, people with dark skin more than white counterparts," a pattern she and other researchers had documented in algorithms built by major technology companies before NIST's 2019 findings.
The practical effect, per that reporting, is that a wrong answer from the software is not evenly distributed across the population it is used on — the people most likely to be misidentified are disproportionately Black, per NPR's account of researchers' findings.
What happens when police treat a match as an identification, according to real cases?
NIST's tests measure the software; they do not measure what a police department does with its output. That gap is where wrongful arrests have followed, according to reporting on individual cases. Robert Julian-Borchak Williams was arrested in Michigan in January 2020 after facial-recognition software matched his driver's-license photo to security footage of a Shinola watch theft; he was held for about 30 hours before a Wayne County prosecutor dismissed the charge for insufficient evidence, according to NPR's reporting on the case. Wayne County prosecutor Kym Worthy said at the time, per NPR, "this case should not have been issued...for that we apologize."
Williams's lawyer, Victoria Burton-Harris, said the facial-recognition match "framed and informed everything that officers did subsequently," according to NPR — the software's output became the working assumption for the rest of the investigation rather than one lead among several. Detroit police told NPR the technology is meant to produce an investigative lead requiring corroborating evidence before an arrest, and said the department has since restricted the tool to still photographs and to violent-crime cases.
The American Civil Liberties Union says it has documented at least 14 wrongful arrests tied to police reliance on facial-recognition matches, as of its April 2026 report. A sample of the cases it names:
| Name | City | Year | Detail, per the ACLU |
|---|---|---|---|
| Nijeer Parks | Woodbridge, NJ | 2019 | Held on a warrant despite receipts placing him elsewhere |
| Robert Williams | Detroit, MI | 2020 | Charge dismissed after about 30 hours in custody |
| Porcha Woodruff | Detroit, MI | 2023 | Arrested while eight months pregnant |
| Trevis Williams | New York, NY | 2025 | Eight inches taller and 70 pounds heavier than the suspect described |
In each case, per the ACLU, the underlying pattern was the same: officers treated a computer-generated match as evidence strong enough on its own, moving to a photo lineup or an arrest without pursuing other corroborating leads first. Charges in the cases above were later dropped or the arrests were determined to be mistaken, and none resulted in a conviction, according to the ACLU's account. An arrest is not a conviction, and the presumption of innocence applies to everyone named above unless a court record states otherwise.
What is a police department supposed to do after a facial-recognition match?
The sequence described in the cases above and in NIST's own documentation follows the same basic shape, though practice varies by department and no single national standard governs it:
- The software returns a ranked list of possible matches, not a single confirmed identity, per NIST's description of one-to-many search.
- A trained examiner is supposed to review candidate images before any match is passed to investigators — a step Detroit police told NPR they added after the Williams case.
- Investigators are expected to treat the result as one lead and seek independent corroboration, such as alibi checks or other evidence, before an arrest, per Detroit's revised policy as described to NPR.
- If an arrest follows, the case moves through the ordinary charging and court process, where a defendant can contest the identification like any other evidence.
Whether that sequence is actually followed in a given case is where the wrongful-arrest reports above find the practice breaking down, according to the ACLU's and NPR's reporting. The technology itself, per NIST's testing, is neither uniformly accurate nor uniformly inaccurate — it depends on the algorithm, the task, and, still, on the face.
For a related tools perspective, read How accurate is police facial recognition, according to the studies.
For more context, read The Complete Guide to Gaming Genres.
For more context, read AI detectors don't prove a student used AI to cheat.
