Adversarial Certification of Evidentiary Software
Introduction
In recent years, myriad software tools, often proprietary, have been adopted by state actors in the criminal justice system to produce evidence ranging from recidivism risk assessments to facial recognition to probabilistic genotyping.[1] Courts considering the admissibility of software-based forensic evidence have routinely failed to appreciate the importance of software validation.[2] Instead, they often limit their inquiry to the scientific principles purportedly implemented by the code or the procedures employed by the lab technician or law enforcement official using the software.[3] This omission stems in part from the structure of scientific evidence doctrine, which anticipates expert witness testimony reporting the expert’s application of scientific “principles and methods” to “the facts of the case.”[4]
Software-based evidence fits uneasily into this paradigm because the testifying witness is usually expert in either the underlying scientific principles or the use of the software tool but is unable to testify as to whether the software’s implementation of the relevant scientific principles is correct, appropriate, or properly validated.[5] Failure to properly validate evidentiary software undermines the reliability of criminal trials for the simple reason that faulty software produces faulty evidence. We do not propose, however, that software validation should simply be added to judicial oversight of scientific evidence. Judicial gatekeeping is not the only (or best) way to validate software-based forensic evidence tools. Instead, we advocate an adversarial certification mechanism that takes advantage of software’s particular characteristics and incorporates robust red-team review of justice-critical software.[6]
Scientific evidence doctrine, judicial gatekeeping, and the adversarial process are intended to ensure that criminal trials are fair, accurate, and faithful to the requirement of proof beyond a reasonable doubt.[7] These mechanisms rely on adequate validation of the underlying evidentiary technology and scrutiny of its application in a particular case.[8] While traditional admissibility doctrines combined with adversarial cross-examination are critical to just adjudication, they alone are fundamentally inadequate, in principle and in practice, to cope with the increasing reliance on software-based forensic evidence in the criminal justice system.
Introducing scientific evidence into a system dependent on lay decisionmakers, such as judges, juries, and attorneys, is an inherently difficult undertaking. Scholarly and scientific concern about the adequacy of judicial screening of scientific evidence is far from new. Nonetheless, recent and growing reliance on software-based forensic tools, many of them probabilistic, exacerbates previous concerns while introducing additional problems. Importantly, however, increasing reliance on software-based tools also opens up new possibilities for systemic improvement through the creation of mandatory adversarial certification procedures rooted in recognized software verification practices.
In what follows, we first review the structure of scientific evidence doctrine, discussing some longstanding critiques and proposed reforms advanced by previous scholars. We next explore how the central role of software in many modern forensic tools problematizes current approaches to scientific and forensic evidence even further, reviewing how recent case law illustrates these problems.
Having laid out the problems, we argue that the special characteristics of software-based forensic evidence also point to solutions, breathing new life into the potential for centralized evaluation of forensic evidence techniques. Drawing inspiration from both the standard software validation practice of adversarial testing and the crucial role of adversarial process in the justice system, we advocate adversarial certification of software-based forensic tools. Finally, we consider how such a certification process could and should be designed to avoid some of the anticipated pitfalls.
I. The Longstanding Problem of Scientific Evidence and Lay Decisionmakers
A recent historical treatment traces concerns with the process for vetting scientific evidence back at least to early uses of scientific expert testimony in the nineteenth century.[9] Twenty-first century reports from the National Academies of Sciences (NAS) and the President’s Council of Advisors on Science and Technology (PCAST) suggest that forensic evidence continues to be plagued with serious reliability problems.[10]
Debates about the appropriate standard for forensic evidence have generally focused (either explicitly or implicitly) on a distinction between the validity of the scientific “principles and methods” relied upon by an expert witness, often termed “foundational validity,” and the application of the principles and methods to the facts of the case, often termed “validity as applied.”[11] Noted evidence scholar Paul C. Giannelli argued as far back as 1980 for a three-part distinction between principles, techniques applying those principles, and the application of a technique in a particular case, but the distinction between principles and techniques has received little attention.[12] In essence, as we discuss in Part II below, the failure of later courts and rule drafters to take the distinction between principle and technique seriously is at the heart of the current failure to properly account for software validation in determining the admissibility of software-based forensic evidence.
In the United States, the question of admissibility of scientific evidence is governed by two primary standards, with some state-by-state variations. These primary standards, known as Frye and Daubert after the cases that established them, address the foundational validity of the “principles and methods” upon which the scientific expert in a particular case has relied.[13] The 1923 Frye standard looks to acceptance by the relevant scientific community, requiring that the basis “from which [a scientific expert’s] deduction is made must be sufficiently established to have gained general acceptance in the particular field in which it belongs.”[14] In 1993, the Supreme Court rejected the contention that Rule 702 of the 1975 Federal Rules of Evidence had incorporated the Frye standard, and enunciated instead the multifactor Daubert test.[15] The Daubert test purports to position judges as “gatekeep[ers]” of scientific evidence who must perform a “preliminary assessment of whether the reasoning or methodology underlying the testimony is scientifically valid and . . . properly can be applied to the facts in issue.”[16] Daubert supplies a set of non-exclusive factors to be used in making that assessment, including but not limited to whether the technique or theory has been tested, whether it has been subjected to publication and peer review, and whether it has attracted widespread acceptance in the scientific community.[17] In addition to governing federal courts, the Daubert test has now been adopted by a majority of states, with most remaining states continuing to apply some version of the Frye test or a hybrid of the two.[18]
Jurisdictions vary more widely in how they approach “validity as applied,” with some, such as federal courts following Rule 702, purportedly including it in the admissibility assessment and others treating the issue as a matter for examination at trial.[19]
A. Assessment of Foundational Validity: Frye, Daubert, and Something in Between
1. Frye and Its Discontents.—In assessing the reliability of the principles and methods/techniques, some states continue to follow Frye’s basic “general acceptance test”[20] although states vary as to the sorts of evidence they accept to demonstrate “general acceptance.”[21] Notably, the California Supreme Court in 1976 established what came to be known as the Kelly-Frye test.[22] In Kelly, the Court reaffirmed California’s adoption of Frye’s general acceptance test, but reversed the lower court’s admission of “voiceprint” evidence, noting “three infirmities” in the expert testimony introduced to establish general acceptance.[23] First, the Court questioned “whether the testimony of a single witness alone is ever sufficient to represent, or attest to, the views of an entire scientific community regarding the reliability of a new technique.”[24] Second, the Court was concerned about the witness’s impartiality, noting that as “one of the leading proponents of [the technology at issue]; he has virtually built his career on the reliability of the technique.”[25] Third, the witness’s qualifications were “those of a technician and law enforcement officer, not a scientist” and “do not necessarily qualify him as a scientist to express an opinion on the question of general scientific acceptance.”[26] In taking such a stringent approach, the Kelly Court specifically applauded the “essentially conservative nature” of the Frye test, opining that it was “deliberately intended to interpose a substantial obstacle to the unrestrained admission of evidence based upon new scientific principles,” especially in the criminal context.[27]
Other Frye states take a more permissive approach regarding the establishment of general acceptance. In New York, general acceptance of an evidentiary technique by the relevant scientific community can be established by (1) expert testimony; (2) authoritative scientific writings; or (3) judicial notice of other judicial opinions finding that the methodology has been generally accepted in the relevant scientific community.[28] In People v. Wesley,[29] for example, the New York Court of Appeals opined that the Frye test requires “no more” than establishing “the general reliability of DNA matching,” with the defendant’s specific objections “going to trial foundation or the weight of the evidence.”[30]
The Frye test has been called both under- and over-inclusive. It can be under-inclusive because of the lag time between development of innovative approaches and general acceptance.[31] Conversely, its over-inclusiveness was revealed by studies such as the NAS and PCAST Reports demonstrating how common forensic techniques have been widely adopted by forensic practitioners without extensive (or any) serious empirical validation.[32]
A range of other criticisms have also been leveled at Frye’s general acceptance approach. A 1980 article by well-known evidence scholar Paul Giannelli elucidated how courts had struggled (or sometimes neglected) to determine (1) the appropriate field (and at what level of generality it should be defined); (2) whether the underlying theory, the technique to be applied, or both must be generally accepted; (3) the scope of admissibility if there has been only a limited amount of empirical testing; and (4) how much to rely on validation testing performed by those who developed or deploy a technique.[33] Rather than resolve these questions, Giannelli contended, courts often focused on the qualifications of expert witnesses and mistakenly relied on the testimony of technicians, rather than scientists, to establish general acceptance.[34] As discussed in Part II, these problems continue to plague the treatment of software-based forensic evidence today.
2. Is Daubert Any Better?—In Daubert v. Merrell Dow Pharmaceuticals, Inc., the Supreme Court considered what standard federal courts should apply after the 1975 enactment of Rule 702 of the Federal Rules of Evidence, which required courts to determine whether “scientific, technical, or other specialized knowledge” offered by a qualified expert “will assist the trier of fact to understand the evidence or to determine a fact in issue.”[35]
The Daubert Court set out five nonexclusive factors that “bear on the inquiry” of “whether a theory or technique is scientific knowledge that will assist the trier of fact.”[36] These “Daubert factors” have become the primary basis for admissibility hearings in courts adopting the Daubert approach.[37] They are: (1) “whether [the theory or technique] can be (and has been) tested,” (2) “whether the theory or technique has been subjected to peer review and publication,” (3) a particular technique’s “known or potential rate of error,” (4) “the existence and maintenance of standards controlling the technique’s operation” and, folding in the Frye test, (5) the degree of acceptance in the relevant scientific community.[38] Like Frye’s general acceptance test, the “Daubert factors” bear on only the scientific or “foundational” validity of the proffered evidence.[39]
Following the Supreme Court’s Daubert ruling, Rule 702 was amended in 2000 to “affirm[] the trial court’s role as gatekeeper and provide[] some general standards that the trial court must use to assess the reliability and helpfulness of proffered expert testimony.”[40] Since that time, Rule 702 has required federal courts not only to consider the general question of whether “the expert’s scientific, technical, or other specialized knowledge will help the trier of fact to understand the evidence or to determine a fact in issue,” but explicitly to determine whether the proffered evidence is based on “reliable principles and methods.”[41] These issues of foundational validity are analyzed using the Daubert factors.[42] As noted above, over time, the majority of states have also adopted the Daubert approach, either through judicial application or explicitly in their evidentiary rules.[43]
The impact of shifting from the Frye standard to the Daubert standard has been disputed. Arguments have been made that Daubert is a more permissive standard, a less permissive standard, or an equally permissive standard.[44] In any event, the Daubert approach has not quieted concerns about the standards for admitting scientific evidence. Despite a few early expressions of optimism,[45] scholarly critique of the admissibility standard and scientific critique of the state of forensic science have continued unabated.[46] In many ways, this cannot be terribly surprising, for several of the Daubert factors overlap significantly with Frye’s general acceptance in the relevant scientific community test.[47] Importantly, the fundamental difficulty of lay evaluation of scientific evidence continues to challenge the feasibility of judicial gatekeeping.
B. Validity as Applied
The assessment of validity as applied has received far less attention.[48] Indeed, neither the Frye standard nor the Daubert factors have anything to say about validity as applied.
Jurisdictions following Frye usually consider admissibility only for new scientific methods or techniques, leaving issues of applicability to the particular case for trial.[49] However, some jurisdictions, such as California, have incorporated validity as applied into the admissibility question.[50] Thus, the Kelly opinion laid out three requirements for admissibility:
(1) the reliability of the method must be established, usually by expert testimony, and (2) the witness furnishing such testimony must be properly qualified as an expert to give an opinion on the subject. Additionally, the proponent of the evidence must demonstrate that correct scientific procedures were used in the particular case.[51]
While the first prong relates to foundational validity and the second to expert qualifications, the third requires an assessment of validity as applied.
Similarly, while the Daubert opinion is most known for laying out the Daubert factors, the opinion also emphasized that, to be admissible, scientific evidence must “properly . . . be applied to the facts in issue.”[52] Rule 702 was amended in 2000 to make it clear that admissibility determinations must consider validity as applied.[53] Currently, Rule 702 reads:
A witness who is qualified as an expert by knowledge, skill, experience, training, or education may testify in the form of an opinion or otherwise if the proponent demonstrates to the court that it is more likely than not that:
(a) the expert’s scientific, technical, or other specialized knowledge will help the trier of fact to understand the evidence or to determine a fact in issue;
(b) the testimony is based on sufficient facts or data;
(c) the testimony is the product of reliable principles and methods; and
(d) the expert’s opinion reflects a reliable application of the principles and methods to the facts of the case.[54]
The requirements currently in sections (b) and (d) were added in 2000, clarifying that admissibility depends on validity as applied as well as foundational validity.[55]
Nonetheless, there were ongoing problems with the quality of forensic evidence, as demonstrated by the highly critical 2009 NAS Report[56] and 2016 PCAST Report.[57] The PCAST Report particularly emphasized that judges should account for both foundational validity and validity as applied and should require proponents of scientific evidence to provide information about accuracy.[58] The Report’s discussion spurred attention to the issue of validity as applied from both courts and commentators.[59]
These ongoing problems, including the observation that “many courts have held that the critical questions of the sufficiency of an expert’s basis, and the application of the expert’s methodology, are questions of weight and not admissibility,” led to the 2023 amendments to Rule 702 (italicized above).[60] The 2023 amendments were intended “to emphasize that each expert opinion must stay within the bounds of what can be concluded from a reliable application of the expert’s basis and methodology.”[61]
It is too early at this point to know how much the changes will affect admissibility determinations going forward.[62] The 2023 revisions remain notably silent on the question of software validation and do nothing to clarify where software falls in this updated schema.
C. Special Concerns in Criminal Cases
The application of forensic methods and tools in the criminal context raises special concerns. While civil cases often involve science with widespread application outside of the litigation context, the scientific evidence proffered in criminal cases is often (probably mostly) based on methods and tools developed specifically for and deployed exclusively in the justice system.[63] Moreover, the primary “customers” for these tools and methods are prosecutors and law enforcement officers, who are repeat players whose resources vastly outweigh those of most defendants.[64] Indeed, many “crime labs” are incorporated into law enforcement agencies.[65]
This lopsided situation has important implications for admissibility under either Daubert or Frye. Most obviously, defendants may have difficulty gaining access to forensic expertise (and to the physical evidence that is to be analyzed). Beyond that, both Daubert and Frye rely heavily on the “relevant scientific community.”[66] When a particular forensic method or tool is primarily or exclusively relevant to criminal cases, there is often no scientific community with incentives to study it, except forensic scientists focused on developing evidentiary tools for law enforcement.[67] In this common situation, the very existence of an unbiased, independent “relevant scientific community” whose views can inform admissibility decisions is a fantasy.
D. Previous Proposals for Reform
The longstanding critiques of scientific evidence doctrine have produced numerous general proposals for reform. These proposals have generally aimed to: (1) improve the capacity of lay judges to assess scientific evidence by revising admissibility standards, educating judges, encouraging use of court-appointed experts, and improving the ascertainment of scientific consensus; (2) level the playing field by providing forensic resources and expertise to criminal defendants; (3) fund or incentivize forensic research to provide better testing and validation of techniques and methods; or (4) obtain independent outside assessment of forensic evidence techniques using independent commissions and other approaches. As we will discuss in Part III, our proposal for adversarial certification can effectively address many of the concerns motivating these proposals.
1. Improving Judicial Evaluation of Admissibility.—Some reform proposals focus on stringently enforcing current admissibility requirements, for example, by ensuring that the burden of proof lies on the proponent of the evidence (usually the prosecutor)[68] or by following the PCAST Report’s insistence that judges should account for both foundational and “as applied” validity and should require proponents of scientific evidence to provide information about accuracy.[69] Other proposals go further, suggesting that forensic matches be subject to an explicit and stringent maximum error rate,[70] that the prosecution be given enforceable duties to test and preserve samples,[71] or that the scope of expert testimony be explicitly limited to avoid overclaiming about the reliability of forensic matches.[72] Recently, and notably, Victor Metallo has suggested an admissibility standard for AI-based evidence that would require testimony by a software engineer and exclude testimony if the AI has reached a point where the “black box” cannot be explained by human testimony.[73]
Proposals to help lay judges cope with the challenges of the gatekeeper role have included suggestions for specialized judicial education, for example, on frequently encountered technical topics such as statistics.[74] Other proposals include increased use of independent, court-appointed experts,[75] making scientific consultants available to judges,[76] employing special masters,[77] and convening panels of experts chosen by litigants to assist judges in assessing scientific evidence.[78]
Rather than beefing up judicial inquiry into the validity of scientific evidence, some recent scholars have essentially concluded that the gatekeeping task is impossible for lay judges as a practical matter, advocating a return to a Frye-like standard in which judges are charged only with determining whether there is a scientific consensus about a particular question.[79]
All of these proposals continue to subject lay judges to the ultimate burden of assessing the scientific validity of highly technical scientific principles and methods, and their applications in particular cases. Most judges’ capacity for this formidable task is limited, as a matter both of resources and of training.[80] While the intent of case-by-case evaluation of the admissibility of scientific evidence is to harness the advantages of the adversarial process, the unsurprising result, as discussed in more detail in Part II, is often deference to previous admissibility determinations.[81]
2. Leveling the Playing Field.—Particularly in the criminal context, some scholars have emphasized reforms aimed at leveling the playing field for defendants. Giannelli, for example, suggested imposing an enhanced burden of proof—reliability beyond a reasonable doubt—for forensic methods deployed by the prosecution.[82] Others have proposed providing more resources to put teeth into the judicially recognized defendant’s right to expert assistance and mandating defense access to lab records and samples.[83] Some recent commentators have focused specifically on defendants’ ability to challenge software-based evidence, advocating defense access to source code for review,[84] or to compiled software for purposes of adversarial black box testing.[85] Among these proposals are suggestions to reinterpret the Confrontation Clause in light of the increasing importance of “machine testimony” to focus less on confronting witnesses and more on “meaningful opportunity to scrutinize the government’s proof”[86] and a “right of meaningful impeachment,”[87] which could conceivably encompass a variety of measures including access to source code, laboratory resources and samples, adversarial testing, and expert advisors.
The difficulty with most of these proposals is that they involve providing additional resources to already under-resourced defense attorneys, an approach that has not proved politically popular to date.[88] (Indeed, even if more resources were to be allocated to defense attorneys, it is not clear that subsidizing technical experts would be the best use of those resources.)
3. Improving Forensic Science.—Commentators have also proposed ways to improve the quality and reliability of forensic evidence techniques, including: a national forensic R&D strategy providing more resources for crime labs and for forensic science research; mandatory certification and proficiency testing for laboratories and technicians; establishment of uniform standards and protocols for various forensic procedures; quality assurance and quality control protocols for crime labs; a code of ethics for forensic professionals; and encouraging more rigorous peer review and publication.[89]
Specific suggestions for improving software-based forensic evidence tools have included the use of open source software, compliance with software-specific standards, and a legislative requirement that forensic software be subject to freely given research licenses and financially independent testing.[90]
Such proposals are salutary and compatible with our certification proposal. Our skepticism regarding the availability of resources to level the playing field carries over here, however. Unless admissibility standards are tightened, there seems to be little incentive for law enforcement and prosecutors to support the allocation of resources to studies that might debunk currently admissible forensic evidence techniques.
4. Independent Validation and Review.—Many commentators, including the NAS and PCAST, have emphasized the importance of independent validation of forensic science technologies, as well as the importance of determining the scope of a technique’s validity, particularly in light of the real-world contexts in which it would be applied.[91]
Some have proposed bringing independent expertise directly into the judicial process. In the 1970s, proposals for a specialized so-called “Science Court,” which would handle cases presenting significant scientific questions, gained a surprising amount of traction, though they eventually fizzled.[92] More recently, some scholars have advocated for revisiting this possibility.[93] Related proposals envision specially qualified “blue ribbon juries.”[94] These proposals, however, are aimed primarily at civil cases, such as mass tort litigation, where there is often consolidated multi-district litigation, and seem clearly unsuited to the criminal justice system, with its emphasis on a jury of one’s peers and its routine invocation of forensic evidence in very large numbers of cases.[95]
Another set of proposals revolves around the setting of standards to guide forensic analysis. For example, widely accepted standards released by organizations such as the National Institute of Standards and Technology (NIST) provide guidance to crime labs for analyzing forensic DNA.[96] Compliance with NIST standards is voluntary,[97] and the standards are updated periodically as forensic DNA technology evolves.[98] The FBI has implemented relatively strict quality standards for forensic DNA testing laboratories, but those standards are required only for labs that are conducting business with the U.S. government or that seek to contribute data to the Combined DNA Index System (CODIS), the FBI’s national DNA database.[99] Otherwise, the standards remain nonbinding.[100] While standards are useful, courts generally have to assess compliance on a case-by-case basis, raising some of the same questions of capacity discussed above. Moreover, when the standards are voluntary, courts may simply choose not to enforce or consider them, a phenomenon which we discuss in our review of software-based evidence case law in Part II.[101]
Closest to our adversarial certification proposal are suggestions to create independent bodies to assess whether particular forensic techniques should be admissible in court. Giannelli’s early article, for example, notes that “[s]everal judges and commentators have advocated the creation of independent bodies of experts who would be called up to review novel scientific techniques before they could be used in court,” with certification by such an expert body being a potential prerequisite for admissibility.[102] Edmond and Roberts have proposed that the initial admissibility decision for a particular forensic technique or tool be guided by the advice of a “multidisciplinary advisory panel.”[103] The 2016 PCAST report has proposed tasking the National Institute of Standards and Technology with evaluating the foundational validity of forensic technologies,[104] while some scholars have advocated the establishment of a “permanent forensic science commission or institute” to take on this and other tasks.[105]
Most commonly, reformers seem to envision these independent bodies functioning in an advisory role, rather than as a certifier of the reliability of a forensic technique.[106] With only two exceptions, we are unaware of certification approaches having been applied to forensic evidence technologies.[107]
Notably, a sort of certification has long applied to devices for determining blood alcohol levels for use in DUI prosecutions (colloquially, “breathalyzers”). The NHTSA maintains minimum standards for such devices, as well as a list of conforming devices,[108] which states rely on to specify which devices and procedures their law enforcement agencies must use.[109] These certifications have been persuasive to courts and led some legislatures to lower admissibility requirements for breathalyzer evidence.[110] Nonetheless, some defendants have successfully argued that they should be allowed to review the source code, famously uncovering bugs in some devices.[111] Defendants also have successfully challenged the reliability of the measurements in their particular cases.[112]
A similar approach to certification of probabilistic genotyping tools famously failed when the Forensic Statistical Tool (FST) software produced by the Office of the Chief Medical Examiner was found to have been modified after its approval by the DNA Subcommittee of the New York State Commission on Forensic Science.[113]
All told, though proposals for reform abound, few have been adopted, and serious problems with forensic evidence remain unresolved.[114] There presumably are many reasons for the ongoing failure of scientific evidence reform, including the inherent difficulty of tasking lay judges with screening scientific evidence (and judges’ corresponding reluctance to delve into the technical details), resource constraints affecting judges, defense attorneys, and potential certification institutions,[115] the difficulty of vetting techniques whose validity varies with the lab implementation or the skill of the lab technician,[116] and the problem of capture for purportedly independent certification and evaluation bodies. In Part III below, we discuss why our proposal for adversarial certification of software-based forensic evidence can avoid many of these serious pitfalls.
II. Software Is Different—and the Same
Increasing computational capacity and data availability, as well as greater awareness of both strengths and weaknesses of software-based tools, suggest that challenges to the admissibility of software-based forensic evidence are only likely to increase in light of new applications, such as face recognition, risk assessment, gunshot detection, and others, as well as increased concerns about the manipulability of digital photos and recordings.[117] The 2023 Amendments to FRE 702 (and the likelihood that states will enact similar reforms) may also increase incentives for challenging expert testimony generally.[118] It is thus critical to understand how courts are assessing this evidence and to question whether current approaches to scientific evidence are fit for determining whether software-based forensic evidence is sufficiently reliable to be admitted, especially in a criminal case. Our contention here is that the centrality of software in many of today’s forensic evidence tools and methods changes things—exacerbating some problems while providing new possibilities for effective reform.
In this Part, we analyze a cohort of cases involving the admissibility of software-based forensic evidence to identify common difficulties courts have in assessing this evidence. While many of these difficulties are reminiscent of general critiques of courts’ treatment of forensic evidence, our case law review demonstrates that they are sharpened by the fact that software fits uneasily into the customary division between “principles and methods” and application in a particular case.
Our analysis shows that courts have repeatedly failed to require software validation, either ignoring software entirely or failing to consider the effects of software updates, the role industry standards play, the complexity of software in these forensic tools, and the influence that software users have on the software-based forensic evidence tool’s output. Courts’ mishandling of the admissibility of software-based forensic evidence also extends to issues related to defining the relevant scientific community and the scope of the subject matter of admissibility hearings.
Review of the case law also illustrates that these difficulties are exacerbated by what we term “common law certification”—a process whereby courts determine the admissibility of software-based forensic evidence based primarily on prior judicial decisions that are, themselves, neither robustly nor rigorously reasoned and that often fail to account for differences between foundational validity and validity as applied in the context of software-based tools.
To some extent, the problems we identify in these cases mirror longstanding general concerns about forensic evidence.[119] The cases also, however, illustrate how courts’ treatment of software-based forensic tools has exacerbated these earlier concerns and introduced new issues. Perhaps most striking in these cases is the way in which software expertise simply fades into the background, primarily because courts fail to distinguish between underlying principles and software implementation.[120] This issue arises both in cases in which courts recognize that complicated scientific and technical issues are involved, such as probabilistic genotyping, and in cases in which courts seem to assume that a seemingly straightforward task, such as location mapping based on cell phone data, is trivially (and correctly) implemented in software.[121]
Thus, while we agree with other commentators that courts must pay greater attention to software in an admissibility analysis, our case law analysis goes further by demonstrating that evidence produced by evidentiary software does not fit neatly within prevailing interpretations of evidentiary law (i.e., Frye and Daubert) or with the paradigmatic split between “foundational validity” and “validity as applied” that informs how courts currently treat admissibility questions.
Software is what Giannelli’s tripartite taxonomy[122] would call a “technique,” in that it is an application of scientific principles for the purpose of obtaining forensic evidence, which could itself be mistaken or limited in scope, regardless of the validity of the underlying scientific principles involved. But software is unlike more traditional techniques. Typically, the person who uses software (and who often is the testifying expert) neither designs nor fully understands how it works.[123] Moreover, even those who are experts in the underlying science often find that the software’s workings are opaque (sometimes partly for reasons of trade secrecy).[124] Software thus does not fit comfortably into either the Daubert or Frye tests courts use for evaluating “principles and methods” or the traditional adversarial mechanisms courts use to evaluate application to individual cases.
In Part III, we argue that the software-specific issues identified here and in Part I can be successfully addressed by replacing the current informal and ineffective common law certification with an appropriately designed mandatory adversarial certification process.
A. Software-Based Forensic Evidence: Problematic Patterns from the Case Law
To understand how judges are handling the admissibility of software-based forensic evidence, we reviewed a set of 240 judicial opinions issued between 2002 and late 2025 that address the admissibility of software-based evidence.[125] The technologies most often at issue in these cases were probabilistic genotyping and other DNA analysis, breath-alcohol analysis, software used to extract, analyze, or map cell phone information (e.g., cell-site location mapping), and ShotSpotter technology for gunshot detection.[126] In most of these cases, the court made a direct determination of admissibility under the relevant standard for scientific or expert testimony, though in some cases, the issue of admissibility was discussed in a more collateral sense.[127]
Our detailed review of these opinions highlights patterns of problems with the admissibility analysis of software-based evidence, all related to the failure of courts to require appropriate software validation of the forensic tool. That failure manifests in several ways, including: (1) the failure to consider various aspects of software, including updates, software industry standards, the inherent complexity of the software, and software operator variability; (2) questions pertaining to the relevant scientific community; and (3) the scope and subject matter of admissibility hearings.
Interestingly, the first bin involves the absence of software evaluation, where courts either fail to consider significant features of software or mention them in passing without realizing they warrant a closer analysis. The second and third bins, on the other hand, involve misapplication of the evidentiary principles to software-based forensic tools. These issues compound a problem previously recognized in the literature, in which trade secrecy claims by software developers often thwart defendants’ efforts to closely scrutinize the source code.[128] We discuss these observations in more detail in the following subsections.
1. Not Requiring Appropriate Software Validation.—Courts tend to be somewhat at sea when it comes to determining whether a software-based tool is an appropriately reliable basis for expert scientific testimony. As we discuss throughout this section, courts often avoid dealing with software validation by simply ignoring the issue, allowing experts to testify to software-based results without reviewing the software, by treating software as presumptively valid as long as it aims to instantiate a reliable principle or technique, or by punting the issue to trial. Even when confronting the question, however, courts have struggled to determine how to validate software reliability within the framework of the doctrines of scientific evidence admissibility. For example, courts have (1) ignored or treated inconsistently the importance of validating updated versions of previously validated software; (2) been inconsistent about whether noncompliance with industry standards, such as Institute of Electrical and Electronics Engineers (IEEE) standards, should affect admissibility; (3) inconsistently drawn the line between software that is either too trivial or too general in purpose to be challenged by the defendant; and (4) been unclear about how to assess choices made by technicians, such as toggleable features (e.g., in cell-site location information) or optional parameters and estimated contributors (e.g., in probabilistic genotyping). We address each bucket below, supported by case citations from our case law analysis.
a. Failure to Consider How Software Updates Affect Earlier Validation.—Courts have generally been silent regarding how software updates affect the legitimacy of validation performed on earlier software versions.[129] We observed only one case recognizing version changes as a viable concern, and even then, only in a concurrence. In People v. Wakefield,[130] Judge Rivera noted that the validation studies in question took place in the early 2000s, but that TrueAllele had undergone twenty-five unique versions since 1999, including a significant update in 2008.[131] The concurrence also noted that only one of the validation studies indicated which version of TrueAllele was tested, and that the development team made no efforts to preserve earlier versions of the software.[132]
Software updates can range from small, barely detectable changes to major overhauls in the source code that alter core functionality of the program. There can thus be no guarantee that an updated version of software will produce the same output in the same manner as an earlier version. The failure of courts to recognize that a major software update requires a new validation study (or that even a small software update should command at least minimal review) is a gap in understanding how these software-based technologies function and evolve. In fact, it is precisely this issue that arose in the New York Forensic Statistical Tool debacle.[133] Yet, few courts mention the issue of software updates and how it may affect the admissibility of software-based forensic technologies, and our sample did not include any opinions that took the issue seriously.[134]
b. Inconsistent Treatment and Consideration of Industry Standards.—Although we recognize that compliance with industry standards for software development and validation, including IEEE standards and the Scientific Working Group on DNA Analysis Methods (SWGDAM) Guidelines, is not mandated by any governing or regulatory body,[135] such standards offer best practices or minimum standards for software-based technologies that could be useful to courts attempting to assess reliability of software-based evidence.[136] Despite the potential relevance of evaluating standards compliance, courts have been inconsistent in determining whether (and which) standards should play a role in validating software-based technologies.[137]
Since compliance is not mandatory, we were not surprised to see that our analysis did not include any cases that required compliance with the standards. Instead, discussions often focused on whether the standards should be followed, what best practices for software programming might be, where to look for the standards, and whether any such standards should be applied to particular software-based forensic technologies. [138]
We do not mean for our criticism to imply that adherence to software standards should necessarily be mandated. We do, however, contend that several of the standards, including those issued by IEEE and SWGDAM, provide meaningful guidance for developing robust, functional software. Compliance with such guidelines should be at least a consideration for justice-critical software. Whether to require compliance with applicable guidelines should be one consideration in the adversarial certification process we advocate in Part III.
c. Inconsistent Determinations about Software Complexity and Subsequent Validation Requirements.—Software-based algorithms vary in complexity, and the validation required for a more complex algorithm will not be the same as the validation required for a less complex one. Yet courts have not perceived this nuance and generally assume flatly that one validation fits all.
The failure of courts to consider the complexity of software-based algorithms is not categorically different from their failure to consider the magnitude and content of changes made through software updates that we discuss above.[139] Each represents a continuum ranging from less complex algorithms (or updates) to more complex algorithms (or updates) that perform tasks and functions that outstrip the capabilities of validation by casual observation.
This shortcoming is further exacerbated when courts fail to require any validation of the software implementation aspects. In our case analysis, we saw at least one example of a court dismissing the need for an admissibility hearing for a Photoshop process because it deemed the software insufficiently complex to require a separate admissibility hearing.[140] More concerningly, courts have frequently denied the need for expert testimony or an admissibility hearing for evidence based on cell phone extraction software, such as Cellebrite, based on assertions such as that “[t]he idea that text messages may be downloaded from a cell phone is familiar to a lay person” and that the testifying law enforcement official “did not opine about the technological means by which the data download occurred.”[141] But such passing assessments are frequently not rooted in a clear understanding of the underlying complexity of the software, and there is no consistent line that courts have drawn between software that is deemed too trivial or too general to be challenged by a defendant and software that is more complex and warrants a closer review.[142]
d. Inconsistent Weighting of Software-User Impact on Program Output.—Courts have, at times, viewed software-based technologies as immutable, autonomous monoliths. But they are neither immutable nor entirely autonomous, and all require some measure of user interaction or input.
Some software-based technology programs include toggleable features that can be manipulated by the operator. Others rely on operator inputs such as adjusting or selecting parameters that affect outputs of the software. Probabilistic genotyping software, for example, allows for any one of several parameters to be altered. The STRmix user instructions issued by New York City direct the user of the software to input the number of contributors to the software, a parameter that can alter the output of the program.[143] This decision is required to operate probabilistic genotyping software, and the decision itself is based on a set of subjective determinations made by the software operator.[144] This decision, and the factors that go into it, have been criticized elsewhere in the literature.[145] These instructions are a clear illustration of how the software is anything but an autonomous monolith, providing the operator the opportunity and ability to alter any number of parameters before running the sample.
2. Inadequate Determination of Acceptance by the Relevant Scientific Community.—Courts have struggled to define the relevant scientific community in determining whether and how a software-based forensic tool should be validated. In contrast to their outright failure to consider or require validation, this trend represents a misapplication, or at the very least incomplete application, of evidence doctrine. It is not a new issue faced by courts,[146] but it is one that is newly revived, and exacerbated, in the context of software-based forensic tools. In fact, the critiques of courts’ assessment of acceptance by the relevant scientific community go back to the early criticisms of the Frye test.[147]
Our case law analysis reveals that courts have specifically struggled to define the relevant scientific community when software-based evidence is involved and have also faced challenges with determining general acceptance by the scientific community. We address each issue below.
a. Defining the Relevant Scientific Community.—Courts frequently fail to define the relevant expert community, sometimes while stating in conclusory fashion that the putative expert is part of the community. For example, most of the opinions in our sample of admissibility decisions relating to software engage in little to no discussion of the expert scientific community. Among a corpus of 173 cases dealing with the admissibility of forensic evidence returned by a LEXIS search in November 2023 for “(software or open source) and (daubert or frye or admissib!),” only 70 even mention the scientific or expert community. Of those cases, 33 include no substantive discussion (usually just quoting the admissibility standard). The remaining 37 cases rarely engage in significant inquiry into the makeup and definition of the relevant scientific community.
Problematic approaches to defining the “relevant scientific community” have been noted by previous commentators.[148] Sometimes, courts conclude peremptorily that an expert is part of the relevant scientific community without defining precisely what the relevant community is or who comprises it.[149] Alternatively, they may simply assume that the relevant community is the “forensic community,” either explicitly or by asserting that widespread adoption by crime labs and forensic practitioners demonstrates general acceptance by the scientific community.[150] The assumption that the relevant scientific community is made up of forensic practitioners or forensic labs, while not permitted by some courts, is a frequent basis for critique. [151] In criminal cases in particular, where the scientific evidence at issue usually results from applying specialized forensic techniques, there often is no other community studying the technique.[152] Even where, as for a subject such as DNA analysis, there exists a scientific community outside of forensics, courts have sometimes excluded testimony from academic DNA scientists who do not focus on forensic applications.[153] As others have recognized, this narrowness is a problem, since the assumption that an expert scientific community functions as an independent guarantor of validity or certification is no longer valid when the “community” consists entirely of forensic practitioners and laboratories who may not understand the scientific basis of the tests they perform and whose livelihood not only depends on courts’ acceptance of their work, but also is funded almost entirely by the prosecutorial side of the equation. [154]
While defining the scope of the relevant scientific community is thus a well-recognized problem in scientific evidence, the problem is exacerbated for software-based forensic evidence because any such tool invokes at least three types of technical expertise: (1) expertise about the scientific principles and techniques underlying a particular type of forensic evidence (e.g., fingerprints, DNA, using breath to measure alcohol levels, the principles connecting cell site pings to location); (2) software engineering expertise; and (3) expertise regarding the application of evidentiary software to specific forensic samples.[155] The requirement for expertise in these areas mirrors the layering we describe above between foundational validity and validity as applied. In other terms, the problem of defining the relevant scientific community is exacerbated for software-based forensic tools because they may, in certain circumstances, require consideration from multiple expert communities, including those who can speak to the principles and methods associated with foundational validity, and also those who can speak to the validity as applied. Courts have struggled to disaggregate these multiple facets.[156]
Often, the opinions we reviewed downplayed or completely overlooked the importance of software expertise in defining the relevant scientific community, sometimes with express knowledge that the forensic practitioners or law enforcement users of the software lacked software expertise.[157] Courts fail to consider whether expert testimony is required, or sometimes consider whether it is required and affirmatively determine that it is not.[158] Again, this relates to the problem of foundational versus as-applied validity, with courts (often implicitly) categorizing software in binary terms as “foundational” or “applied.” On the one hand, courts often fail to acknowledge that different software programs (and different versions of programs) may instantiate a given “technique” in different ways.[159] On the other hand, software coding is not a case-by-case application of a technique, because a given copy of a program can be deployed across many cases and will produce the same output if given the same input.[160]
b. Determining General Acceptance by the Relevant Scientific Community.—Like the question of how to define the relevant scientific community, the question of how to determine what is “generally accepted” by that community has been of longstanding concern to commentators and scholars.[161] Relying on expert testimony can lead to an unenlightening “battle of the experts” (or simply to biased testimony).[162] Looking to the scientific literature can be unsatisfactory for forensic technologies when the only available literature is authored by practitioners with a stake in promoting the technology.[163] Naturally, the problems with relying on the scientific literature for evidence of general acceptance also undercut the utility of Daubert’s “subjected to peer review and publication” factor for assessing reliability.
Again, the special characteristics of software exacerbate these problems. While practitioners of more traditional forensic techniques may not be experts in validating those techniques and have conflicts of interest, at least they often have a detailed, clear understanding of how the technique works and what exactly they are doing. Bringing in software introduces a black box, the programming of which is typically outside the expertise of the practitioner using it.[164] Practitioners may be unaware of or unable to explain inconsistent results across different programs and versions of programs.[165] The proprietary nature of most such software makes it difficult to rely on peer review and publication, since the proprietors of the software may be the only ones with sufficient access to conduct rigorous software validation testing.[166]
The validation of software-based probabilistic evidence raises additional challenges that courts have not considered. For example, researchers have characterized validating probabilistic data as challenging because the reported metrics have both little direct physical meaning and a shorter history of use, making them less familiar and more difficult to interpret.[167] Compounding these challenges, the validation of probabilistic-based methods can also involve either human reviewers—who make arguably subjective determinations on whether the probabilistic output can be validated—or additional statistical analyses to determine the uncertainty inherent to the model and thus in the data output.[168]
We will argue below that, while acceptance by a relevant expert community is meaningful for general scientific principles, it is simply the wrong paradigm for evaluating software reliability (indeed, the same is arguably true for many other forensic tools).[169]
3. Problems Defining Scope and Subject Matter of the Admissibility Hearing.—Some of the issues already discussed—such as judicial aversion to technical details and limited perspectives on the relevant scientific community—also have obvious implications for the scope and subject matter of the admissibility hearing. Our review of the cases highlighted several related, additional issues. First, courts often focus on the qualifications of testifying experts and the intended scope of their testimony, rather than the reliability of the forensic technology. This issue arises in part because courts ignore the importance of software (or, often, any non-practitioner) expertise and have ignored the split between foundational validity and as-applied validity that is central to software-based forensic tools. Second, courts too easily construe questions about software reliability as matters for cross-examination at trial rather than for admissibility, often because the court has lumped the software together with its application in the particular case. Again, this represents a failure of courts to acknowledge the need for foundational and as-applied validity considerations for software-based forensic tools. Relatedly, courts often deny defendants access to source code on trade secrecy grounds, under the apparent assumption that access to source code is not a necessary foundation for cross-examination. Even when courts recognize the need to validate the reliability of forensic evidence software, they often struggle to do so. We discuss software validation in the next subsection.
a. Focusing on Qualifying Experts Rather than Reliable Forensic Evidence Tools.—One way in which courts in our case law analysis avoid dealing with technical issues regarding software is by focusing on qualifying the expert (usually a practitioner) rather than on the reliability of the software-based tool.[170] This approach often goes hand in hand with courts admitting an expert’s testimony based on experience or training as a technician, rather than on expertise about the scientific or technical matters underlying the technique.[171] Some courts even go so far as to say that there is no need to validate the software because the expert did not use the software herself or plan to testify about the technical details of the software.[172] This move highlights an inconsistent treatment of the role software plays in these tools.
b. Treating Software Validity as a Matter for Cross-Examination.—Under FRE 702 and Daubert, a trial judge must make a preliminary assessment to ensure that the expert testimony, and any reasoning and methodology underlying the testimony, is scientifically valid and applicable to the case.[173] This preliminary assessment functions as a gatekeeper role aimed at ensuring the reliability of technical evidence. Frye’s general acceptance standard has a similar function. We have, however, observed several courts employ this gatekeeper function in a way that punts the need to interact with the technical details of software-based technology either to the jury or to later in the proceedings.[174] The requirement that scientific evidence be reliable can also be viewed through a relevancy paradigm: Inadequately validated scientific evidence is either essentially irrelevant or at least more prejudicial than probative due to its tendency to be persuasive to jurors. Thus, it is not appropriate for a court to admit complicated technical evidence with the idea that its complexities can be explored at trial and left for the jury to weigh.[175] Unfortunately, courts too often seem to be doing exactly that.
c. Trade Secrecy Concerns.—In traditional forensic evidence analyses, which tend to be carried out by lab technicians using general-purpose equipment such as microscopes, trade secrecy does not pose a significant issue in admissibility discussions. Software is often proprietary, however, and courts have demonstrated a surprising willingness to allow trade secrecy privileges to be asserted to preclude defendants from examining source code and documentation of software used in forensic evidence analysis.[176] Many commentators have focused on this issue, raising constitutional concerns about due process and confrontation, as well as more general concerns about justice and fairness.[177] Earlier, we also argued that trade secrecy is unnecessary to promote innovation in forensic evidence tools.[178] Our case review illustrates that trade secrecy is also a barrier to effective review of reliability as required for appropriate admissibility determinations.[179]
B. Common Law Certification
The result of courts’ failure to properly consider software validation is often a feedback process we call “common law certification,” that we describe below, in which courts determine admissibility based primarily on prior judicial opinions.[180] The Frye test particularly, but certainly not exclusively, produces this phenomenon, given its limitation to assessing novel technology, the fact that some courts fail to recognize the need to re-evaluate new techniques based on accepted theories, and the fact that many jurisdictions explicitly sanction reliance on prior judicial opinions to determine “general acceptance by the relevant scientific community.”[181] Common law certification stems in part from a feature of both the Daubert and Frye tests, which place heavy reliance on the question of “acceptance” by some relevant scientific community.[182] The emphasis on general acceptance of an expert’s theory or technique may inaccurately suggest that scientific reliability is an all-or-nothing question, rather than a nuanced question that may depend on a technique’s range of applicability, its specific implementation, and other factors.
Common law certification based on previous admissibility decisions creates a network of admissibility rulings, each bolstered by the mere existence of the previous rulings, as demonstrated in previous work for the probabilistic genotyping case by several of the current authors.[183] This blanket reliance on earlier admissibility rulings is often accompanied by an assurance that any issues with the particular implementation of an approach that was previously deemed admissible can be dealt with by cross- examination at trial.[184]
The tendency of courts to rely on previous courts’ admission of evidence produced by related technologies is not a new problem, of course. Indeed, courts in some jurisdictions explicitly state that prior judicial opinions are an acceptable form of evidence of general acceptance by the relevant scientific community.[185] In some jurisdictions, heavy reliance on prior judicial opinions may also reflect the idea that the Frye test need be applied only to novel techniques.[186] Not surprisingly, earlier commentators and some judges also criticized courts for relying too uncritically on previous admissibility rulings. As far back as 1980, evidence scholar Paul Giannelli’s extensive critique of Frye’s legacy argued that “‘considering general acceptance not only by scientists but also by courts’ . . . undercuts the rationale supporting Frye—that those most qualified to judge the validity of a technique should have the determinative voice.”[187]
Common law certification exacerbates the problem of judicial technical incompetence because it compounds and solidifies erroneous admissibility determinations and stagnates the justice system’s responsiveness to scientific evolution. Writing in 2008, Keith Findley argued that one “consequence of leaving admissibility questions to the adversary adjudicative process is that stare decisis can quickly become a substitute for analysis, and can freeze judgments about science even if the science itself continues to evolve,” thereby “locking in misjudgments about science and preventing fluid adaptation of admissibility or other legal standards to reflect changing scientific knowledge.” [188] The danger is that common law certification can create an impenetrable wall for defendants who themselves often have limited resources to develop what could be a successful challenge to an unreliable software-based forensic tool.[189]
Whether because of careless development of forensic tests or evolution of scientific understanding, techniques that seem acceptable at one point in time may be found unreliable at a later date. Perhaps the most compelling evidence of this point comes from studies of cases in which convicted individuals were later exonerated by later-developed DNA testing.[190] The majority of those convictions had been based on forensic testimony. Similarly, blue-ribbon studies, such as the NAS and PCAST Reports, found serious problems with the reliability of many long-accepted forensic evidence techniques.[191] Thus Giannelli, in a 2018 article entitled “Daubert’s Failure” emphasized how courts have “abdicated their responsibility” by taking judicial notice of and relying on earlier court decisions when admitting evidence that was later demonstrated to be unreliable.[192]
Judicial avoidance of complex technical details carries over to cases involving software-based forensic evidence with two additional wrinkles: The experts proffering the evidence often lack computer or software expertise themselves, and the software code is often proprietary and is not disclosed to the court, the defendant, the prosecutors, or even the expert.[193]
Common law certification also tends to elide important distinctions between the reliability and/or general acceptance of the scientific principles underlying a given technique, the technique based on those principles, and the application of the technique in a particular case, thus giving previous admissibility decisions unjustifiably broad downstream effect.[194]
While these earlier critiques continue to have force with software-based forensic tools, the problem of common law certification is exacerbated by courts’ confusion about where software fits into the principles/technique/application rubric. Courts are inconsistent in their treatment of the software underlying software-based forensic technologies, sometimes characterizing it as a straightforward application of the underlying scientific and mathematical principles and sometimes as a general theory or technique.[195] Both approaches miss the point: Software is an independent, and often complicated, implementation of underlying principles and theories, worthy of, and in fact requiring, its own appropriate validation.
While Giannelli’s 1980 distinction between principles, techniques, and applications emphasized the importance of ensuring the reliability of both underlying principles and the particular technique used in a case, [196] software implementations of what might sound like the same technique (e.g., “probabilistic genotyping” or “cell-site location mapping”) can vary significantly, using different algorithms and incorporating different bugs.[197] These differences can be significant even if the output of the software (e.g., a likelihood ratio or a map) would appear quite similar to the lay judge or juror.[198] Indeed, even updating a software program can introduce different and additional bugs.[199] Different programming choices, and even different parameter choices, can vary the scope of validity of the tool.[200] Moreover, these variations can sometimes be idiosyncratic and unintuitive even to experts in the underlying scientific principles.[201] Nonetheless, courts often elide these differences, relying on previous admission of earlier versions of the software,[202] of completely different software programs applying the same or similar principles,[203] or simply alluding to previous acceptance of the general principle or technique.[204] Some courts consider only the validation of the underlying scientific principles. For example, in Whittley v. State,[205] a Texas Court of Appeals rejected a challenge to exclude data from STRmix on the grounds that “DNA testing . . . [has] been found reliable and to have a sound scientific foundation.”[206] While the court’s statement that DNA testing is a reliable method with sound scientific foundations is true, the court failed to consider the second layer: validation that the software has faithfully implemented those scientific theories.[207]
Though we agree with criticisms of common law certification, both as to traditional and software-based forensic evidence, we find it not at all surprising that common law judges, steeped in the use of precedent and lacking technical training, would adopt this sort of approach.[208] Despite calls for judicial education, court-appointed experts, and other proposals to better equip judges to evaluate scientific evidence, we agree with many previous commentators who have found it unrealistic, and perhaps even unreasonable, to expect that lay judges will thoroughly and competently vet expert scientific evidence on a case-by-case basis. In Part III we propose an independent certification approach to software-based forensic evidence that responds both to the need for software validation and to the limits of judges’ capacity to ensure that software-based forensic evidence has been properly validated.
III. Mandatory Certification of Software-Based Forensic Tools
The discussion so far has demonstrated that software-based forensic tools used to produce evidence for criminal prosecution are plagued with the same sorts of issues regarding admissibility and reliability that have been long and widely recognized for more traditional techniques. Moreover, we have noted various ways in which the problems are exacerbated for software-based tools.[209] Despite this discouraging conclusion, we believe there is reason for hope that the special characteristics of software-based evidence tools also make them more amenable to reform than more traditional techniques.
In particular, we propose that software-based forensic evidence tools should be subjected to a mandatory independent certification process that should be a prerequisite for admissibility and should effectively replace the Frye inquiry into general acceptance by the relevant scientific community and the Daubert inquiry into foundational validity (of both theory and technique), leaving only specific issues regarding validity as applied to be litigated on a case-by-case basis.[210] Delineating a predetermined scope of certified software validity should enable defense counsel, prosecutors, and courts to focus the admissibility hearing on a narrower set of cognizable disagreements.
We emphasize that this type of certification process will be appropriate and successful only if it meets certain requirements for adversarial input and validation according to accepted protocols by both software and subject-matter experts. Because certification can only ever be meaningful for a specific software program, version, and scope of validation, the process we have in mind would produce documentation that would delineate the certification’s specific range of applicability.
In this Part, we first explain why we think certification is particularly appropriate for software-based forensic evidence tools used in criminal cases. We then provide some background on software validation and discuss why the certification process should pay special attention to the particular ways in which software can produce erroneous outputs. Finally, we elaborate on our proposed certification process, comparing and contrasting it with processes currently used for breath alcohol devices and probabilistic genotyping.
A. Evidentiary Software Is Particularly Amenable to Certification
As described above, software-based forensic evidence reveals a problematic blind spot in current doctrinal approaches to admissibility, which focus primarily on foundational scientific validity and as-applied validity without considering intermediate aspects such as software validity. The good news is that several features of software-based forensic tools used in criminal cases render them particularly amenable to a certification approach.
First, the scope of forensic evidentiary software can be defined much more narrowly than for general-purpose software. Forensic analyses used in criminal investigations tend to take a limited number of forms that are deployed over and over again across many, many cases.[211] In this respect, criminal investigations and trials are quite unlike civil litigation, such as the tort litigation in the Daubert case, where the relevant scientific questions may vary idiosyncratically from case to case.[212] Of course, there may be occasional criminal cases involving very case-specific scientific testimony, but any such testimony would not relate to a software-based forensic tool and could be handled by traditional admissibility rules.[213] Indeed, the ongoing phenomenon of common law certification demonstrates the ready amenability of forensic tools in criminal cases to a more formalized certification approach. Our proposal is simply to replace an ad lib, unsatisfactory certification process with a scientifically and technically sound approach, while still avoiding the costs of duplicative inquiries into admissibility.
Second, while software-based forensic tools do reflect subjective choices made by their developers, the performance of software programs is replicable, in the sense that the same input data and parameter settings ought to produce the same output (and if not, there is something unreliable about the code!).[214] Software can also be designed to produce secure logs of exactly what inputs and parameters were used to produce a given output.[215] This reproducibility and documentability means that software-based techniques can be meaningfully certified, unlike many traditional forensic methods, such as latent fingerprint matching,[216] that “rely on human interpretation that could be tainted by error, the threat of bias, or the absence of sound operational procedures and robust performance standards”[217] and whose results can vary from examiner to examiner (and even from one time to another).[218] Again, certification must be specific to a certain code implementation and certain sets of input data and parameters, but it is possible to test and document quite explicitly the ramifications of those choices.
Third, and relatedly, because software is completely specified by code, and there are software engineering techniques for tracking whether, when, and what changes have been made, it is possible to ensure quite well that the certified software was actually used to produce the proffered outputs from a specified set of inputs.[219] Note that ensuring accountability in this way is quite important. A relatively recent scandal involving the FST probabilistic genotyping software occurred precisely because there was no auditing of whether the version of code that was used in the case was the version that had been approved by the New York State Commission on Forensic Science.[220]
Fourth, and as we explain in more detail below, it is possible to design a certification process for software-based tools that involves both defense attorneys and prosecutors and thus is responsive to the concerns of both, while allowing defense attorneys to pool their scarce resources and participate in the vetting of these tools in a way that only prosecutors and law enforcement have traditionally been able to do. There are various possible strategies for doing this, and we do not attempt to lay out all the possibilities but note three key points here. First, such a process would not have to plow entirely new ground. We already have experience with independent institutions that are responsive to diverse—and conflicting—interests. Moreover, there are several literatures, including those dealing with participatory design, multistakeholder processes, knowledge commons governance, and the design of administrative agencies that can be mined for suggestions.[221] Second, a certification process, as we propose, can resolve the source code issue in one of two ways. Ideally, entities submitting a software-based forensic tool for certification would agree to disclose publicly the source code and any other information material to validation of that tool. Alternatively, the certification process offers a middle ground: mandatory disclosure of source code to the certifying entity, which would include representatives from both law enforcement and prosecution and the defense bar. The latter option has at least two advantages over the use of protective orders on a case-by-case basis. Centralizing disclosure benefits defendants by conserving resources from being wasted in duplicative case-by-case scrutiny. For software vendors, it allays any legitimate trade secrecy concerns by limiting disclosure to a well-defined and relatively small group of individuals. Third, by opening up software admissibility decisions to a more open and rigorous certification process, software vendors will be better able to compete on product quality, instead of being beholden to a common law certification process that irrationally favors first-mover incumbents over newer entrants.[222]
B. Software Validation Practices
Standard software validation practices vary greatly across the industry.[223] This variance exists because a fundamental characteristic of software validation is that complete correctness cannot be guaranteed.[224] As a result, software validation is an iterative, ongoing process that needs to be maintained and managed throughout the total lifetime of a software system.[225] To be sure, certain practices are so basic that they can be considered incontestable.[226] Beyond that, when considering the validation of software for use in high-stakes, “justice-critical” contexts such as criminal trials, useful guidance can be gleaned from safety-critical industries such as avionics and medical devices, as well as from standard-setting bodies such as IEEE.[227] Yet, precisely because software validation practices are difficult—perhaps impossible—to standardize, we argue that a competent approach to software validation for criminal trial settings must rely on more than just “blue team” evaluation and testing by the developer’s internal team. Such admissibility determinations should also require “red team” evaluation and testing in an adversarial setting.[228]
“Software validation” refers to the provision of objective evidence (1) that the software requirements are properly specified, and (2) that the actual implementation fulfills those specified requirements in a consistent, complete, and correct manner.[229] As a threshold matter, an inability to show evidence of a specification document, or evidence of a testing plan and test results, would fall well below minimum expectations of modern commercial practice.[230] But merely providing some evidence is not sufficient of itself. The hard question is determining how much and what type of validation is needed to show that a given software system is sufficiently reliable to serve as testimony.
From a computer science perspective, the conventional taxonomy of software validation activities spans three types of analysis: static, dynamic, and formal.[231] Within each category, another salient division is between “white-box” methods, which require open access to source code and other internal materials, and “black-box” methods, which do not. Static analysis is necessarily a white-box method, whereas dynamic analysis and formal methods can be performed in either a white-box or black-box mode.[232] White-box methods are generally more powerful because they provide visibility into the programming logic and control flow of the system.[233] Important validation functions such as “traceability”—ensuring that each requirement can be traced to implemented code and, conversely, that all code can be traced back to a requirement—cannot be performed without full access to the source code.[234] But black-box methods are more readily usable, because they do not depend on overcoming trade secret claims or other barriers to access.[235]
Static analysis seeks to detect defects by directly examining the source code.[236] This activity involves more than simply “staring at printouts of the code” and extends to the use of sophisticated automated tools that can evaluate correctness of behavior based on the program’s fixed structure and syntax. [237] However, a major limitation of static analysis is that it often misses subtle errors hidden by code complexity, embedded dependencies, and other contextual discrepancies.[238] Nor can it provide insight into how the program interacts dynamically with its environment.[239]
Dynamic analysis, by contrast, tests how the software system performs in action.[240] This type of testing can—and should—occur at every stage of software development. The most basic form is “functional testing,” which validates that the software system performs all intended functions as specified in the requirements, and that it does not exhibit unexpected anomalies or hazards.[241] By definition, functional testing stands in contrast to “non-functional testing” that aims to validate extra attributes such as cybersecurity, stability, usability, or interoperability.[242] The major limitation of dynamic testing is that, because of the combinatorial explosion problem and the composition problem, it is not possible to test all possible inputs and scenarios that a software system might encounter.[243] Accordingly, dynamic testing necessarily involves the artful selection of which tests to conduct in the limited amount of available time. While manual, human-driven testing remains an important method, modern state-of-the-art testing also incorporates automatic generation of test suites.[244]
Computer scientists continue to study formal methods, which offer strict mathematical guarantees of program correctness;[245] however, the use of formal methods remains rare in actual practice. [246] A major challenge is that the logistical and computational costs of such techniques quickly become prohibitively high for software programs with more than trivial complexity.[247] Using formal semantics to create and maintain a verifiable model demands a level of rigor that most software developers abhor.[248] Perhaps for this reason, most software developers continue to disregard formal method techniques when planning and conducting software validation activities.
While the classical framework of static, dynamic, and formal analyses addresses the question of what validation activities to perform, it does not answer the question of who should perform those activities. Especially in settings that are adversarial or safety-critical, a “red-team/blue-team” approach offers a useful way to assign those duties. This lingo, which originates from the cybersecurity field[249]—and more broadly from military war games[250]—defines the “blue team” as the internal team responsible for preparing defenses and securing the software against known vulnerabilities.[251] Documentation of blue-team processes remains important because software vendors should be encouraged to engage in rigorous development processes. But self-validation is inherently limited. Likewise, when testing is performed by purportedly independent laboratories that are in fact affiliated with or controlled by the software vendor, this type of validation should still be treated as blue-team activity.[252]
By contrast, the “red team” takes on the adversarial role of an attacker seeking to penetrate those defenses in ways that the blue team failed to anticipate.[253] Combining both blue-team and red-team approaches offers more credible proof that the software has been properly vetted, especially where the source code cannot be publicly disclosed. It also bypasses the additional challenges to validating probabilistic software discussed above in Part II.[254] For evidentiary software, the red team could be an independent auditor, agency, commission, or other appointed institution with formal oversight functions. This institution could employ a closed process or be open to engaging with public comment.[255] In particular, software experts or defense counsel may be able to provide special insight drawn from their experience in individual cases. There are different possibilities for precisely how to constitute an adversarial certification body. Importantly, however, recognizing that existing lawyering is already an ad hoc form of red teaming, any approach that is sufficiently responsive to defense attorney concerns, while consolidating those efforts into an adversarial certification process, affords considerable advantages over the status quo.
C. A Proposal for Adversarial Certification
Certification offers a mechanism to formalize and harmonize the software validation process across diverse jurisdictions. A certification mark can provide the public with assurance that the certified good or service complies with a desired standard.[256] Here, we argue that the advantages of invoking a certification model to validate forensic software are threefold. First, it ratifies and elevates the importance of software validation—especially red-team validation—as part of the overall process of evaluating the admissibility of forensic evidence. As explained above in Part II, too many courts are tempted to treat software validation as a cursory or settled issue instead of understanding that it requires ongoing monitoring and judicial attention.[257] Second, centralizing the certification process generates efficiencies of scale. While individual due process is an important value, many jurisdictions and defendants are unable to afford the expense of validating software on a case-by-case basis.[258] That said, any presumption of software validity has to be only a presumption, because centralized certification cannot cover all issues that may arise in individual cases. Meanwhile, for software vendors, centralization harmonizes guidance and reduces uncertainty across individual cases. Once certified, incumbent providers enjoy clear advantages, since the certification process erects barriers to entry against newer entrants akin to a licensing regime.[259] At the same time, unlike current “common law certification” practices, a centralized certification framework also benefits new entrants by providing a streamlined pathway to achieve a presumption of legal validity and effectively compete in the evidentiary software market. Third, the institutional design of adversarial certification offers a longitudinal form of scrutiny that complements the role of courts. In particular, the certification body can—and should—continue to monitor “post-certification” issues that emerge after the time of certification.[260] When new software updates are released, the body could require software vendors to submit to further validation activities and deploy fixes as needed. Certification could be revoked in recalcitrant cases.
Our proposal for adversarial certification grows out of related work on certification as “soft law,”[261] but it is distinct in that it envisions a mandatory framework rather than a voluntary self-regulatory one.[262] In other words, we envision certification of a software-based forensic evidence tool as a prerequisite for the admissibility of evidence derived from that tool. Hard law refers to traditional laws and regulations, whereas soft law refers to alternative governance tools that establish expectations but are not directly enforceable by law.[263] The turn to soft law is a response to the recurring challenges of adopting or adapting hard law measures to emerging technologies.[264] Strengths of soft law include the potential to be more agile and multistakeholder in its approach,[265] while identified weaknesses include risks of non-adoption, non-enforcement, and watering-down of substantive requirements by self-regulated entities.[266] The adversarial certification proposal embraces many advantages of the soft law framework while hardening its content and compliance functions by separating the regulated party from the certifying authority.[267]
A comprehensive article by the NYU Policing Project on soft law certification for police technologies offers a useful starting point for how to think through the general design of certification models.[268] The authors identify a number of challenges to administering traditional legal oversight of such technologies,[269] and they advocate certification as providing an independent forum to gather expert information about a product and to effectively communicate that information to the public.[270]
Importantly, the authors flag a number of critical design choices to weigh when building any certification system. One set of choices centers on what to measure, while another set of choices asks who should be doing the measuring. The first grouping asks (1) whether certification standards should be prescriptive or descriptive, (2) how to measure efficacy, (3) what use cases the certification should cover, and (4) what substantive design standards should be adopted.[271] The second grouping adds questions of (5) institutional design, and (6) whether certification bodies should be public or private.[272]
1. Adversarial Certification Standards.—For the first grouping of design criteria, narrowing the context from public acceptance of police technology to expert validation of forensic evidentiary software offers several simplifying assumptions. For example, certifying the validity of forensic software is a more plainly descriptive task, rather than one that raises difficult questions of ethical propriety. Regardless how one feels about the normative worth of a particular type of software, validating its functional implementation remains a more objective question. Moreover, validating forensic evidentiary software should be even more straightforward—the software’s main function is to conduct correct classification of evidence according to a known and specified scientific principle. Similarly, the question of efficacy applies more narrowly to software validation activities than to broader deployment of police technologies. The relevant efficacy metric is whether the software properly implements its requirements and specification.[273] Disentangling validation of software implementation from validation of scientific principles or of techniques based on those principles allows for a more precise inquiry into whether the software’s actual functions match what has been expressly claimed.[274]
Of more intermediate difficulty are the tasks of identifying substantive standards and limiting the scope of certification to appropriate use cases. As a general rule, a good place to start is recommendations and guidance published by well-respected standard-setting groups, such as IEEE, NIST, RTCA, or other independent organizations and agencies that have studied verification and validation of safety-critical software.[275] Because software practices vary widely, it is exceedingly difficult to delineate a set of best practices that everyone must follow.[276] Moreover, some efforts at standardization may not provide much meaningful guidance at all.[277] Nevertheless, these general-purpose guidelines work best when used to establish a common floor of minimum accepted practices.[278]
Once again, however, the limited scope of forensic evidentiary software makes it more tractable to define robust standards. Because the software in question typically has only a single purpose, it is easier to compare apples to apples.[279] Where the task is concise and well-defined, it should be possible to standardize the substantive content of the software requirements document, not just the procedural steps that developers are expected to take in implementing those requirements in code. An instructive analogue is the case of breathalyzer devices, where NHTSA has published detailed specifications on required performance standards for conforming products.[280] This simplicity and measurability stand in contrast to many other commercial software systems where the number of requirements and features makes software quality assurance much more complex.[281]
Also significant is specifying the range of use cases for which the software is certified. In general computing, software is expected to be general-purpose and adaptable, so there is little expectation that software will be certified for a limited set of uses. Developers often engage in user acceptance testing in order to better meet market demand, but they rarely disclose which uses have been tested and which ones have not.[282] In safety-critical software, those disclosures are essential. For example, the FDA’s updated guidance on software assurance for medical devices requires manufacturers to identify the intended uses of their software, along with detailed documentation of the assurance and testing activities that were conducted.[283] For justice-critical software such as forensic evidentiary software, it is similarly critical to disclose the specific range of inputs that are tested, the results of those tests, and any known limitations on the scope of uses for which the software is being certified.[284]
2. Adversarial Certification Procedures.—The second grouping of design criteria present more challenging questions of political economy and institutional design choices. The NYU Policing Project highlights the need to win buy-in from key stakeholders such as the public, vendors, and agencies in order to achieve compliance with certification requirements.[285] Yet, naïve inclusiveness can be fatal to multistakeholder processes.[286]
At a bare minimum, the certification body must include qualified software testing experts who bring technical competence and credibility to the enterprise. Beyond independent software experts, the desirability of replicating the adversarial process strongly suggests that representative members of the defense bar should be included in the process.[287]
A critical question is how to include software vendors—whose participation is essential for allowing full access to software testing functions—without sacrificing the firmer features of adversarial certification. A potential solution is to bifurcate the certification body into two teams: a blue team and a red team. The blue team works directly with the forensic software vendor to collect and confirm representations such as validation results, functional claims, and other attestations. As in a cybersecurity context, the blue team cooperates with the software developer to shore up any vulnerabilities or faults that might compromise the proper functioning of the software system.
As part of the blue-team process, the software developer should be required to submit adequate documentation of internal validation activities.[288] This mandate comports with the newly proposed Federal Rule of Evidence 707, which requires “machine-generated evidence” to be accompanied by expert testimony or to meet the reliability requirements of Rule 702.[289] These affidavits would offer supportive evidence that some internal efforts were conducted to detect and minimize software errors. Failure to meet a sufficiency standard at this stage should create a per se presumption that the software fails to be valid or reliable. Consistent with industry standards and practices, the developer must be able to show at a minimum the requirements document, the software validation plan, and substantiation of static and dynamic testing. This evidence of validation must denote the version of software to which it applies. Any subsequent changes to the software should be documented separately, and if appropriate, accompanied by fresh documentation of software validation activities.[290]
The second component of the certification body should perform an adversarial red-team function. Where due process concerns are heightened, we argue that blue-team validation is not enough and that red-team validation should be required as an additional check before software-based evidence can be admitted in court. Standing alone, blue-team validation is unlikely to be sufficiently reliable. There are the usual non-technical constraints of limited resources and misaligned incentives, which cause software teams to cut corners on testing and validation in favor of shipping on time and staying under budget.[291] In addition, there are technological weaknesses rooted in the limited capacity and discernment of non-adversarial methods.[292] Allowing defense counsel—or software experts on behalf of the defense bar—to stress-test the evidentiary software provides more robust assurance that the outputs produced by the software have been adequately vetted for use at trial.[293]
Here, the certification body should be allowed to stress test the forensic software in the same manner that a well-resourced defense team would. Additionally, the adversarial team could work with public stakeholders and impacted groups to aggregate extra sources of information that could assist in its review.[294] Red teaming will be most effective in a white-box setting; access to source code remains an important factor in validating software. (Indeed, source code disclosure to members of the blue team within the certifying body should be required.) In situations where the software developer can demonstrate legitimate trade secrecy reasons not to disclose source code publicly, then red-teaming becomes even more critical because black-box testing becomes the only independent mode of validation. In the latter scenario, the software developer should be required to provide access to an executable version equivalent or identical to what the lab technician or field officer is operating.[295] In other words, the test environment should align as closely as possible with the live deployment.[296] Relatedly, the red-team testers should be free to submit test inputs without limitation; placing artificial constraints on black-box validation makes it especially vulnerable to test result manipulation.[297] Black-box validation will necessarily be more limited than white-box validation, so any resulting certification should be labeled and limited accordingly, and courts should exercise more caution when admitting evidence based on this diminished form of certification.
The adversarial certification framework we have outlined here is designed for the narrow application of admissibility hearings of software-based evidence in criminal trials. Nonetheless, it is possible to imagine generalizing the application of this approach to other contexts such as criminal investigations or civil actions.[298] While an adversarial certification process involves nontrivial cost, there is a pressing need to provide a streamlined, trustworthy alternative to case-by-case “common law certification” of software.
Conclusion
The criminal justice system’s growing dependence on software-based forensic tools has outpaced the law’s capacity to ensure their reliability. Software validation does not fit neatly into the doctrinal distinction between validating foundational scientific principles and validating as-applied field conditions under the standard Daubert and Frye frameworks. Software is neither a scientific principle nor a laboratory test; it requires a separate and distinctive form of validation.
The solution this Essay proposes is adversarial certification—a mandatory, structured process in which software-based forensic tools must pass rigorous independent validation, combining blue-team developer documentation and testing with red-team adversarial testing, before their outputs may be admitted as evidence. This approach harnesses software’s distinctive properties: its reusability and replicability, as well as the inherent limits of vendor testing. Moreover, software is difficult to examine, and it changes early and often. Adversarial certification is more effective than case-by-case adjudication and is responsive to both prosecution and defense interests. By delineating the scope of certification, a well-designed certification process would also provide a clear basis for parties to contest the software’s application to the circumstances of particular cases.
When lives and liberty depend on the outputs of software, the justice system’s legitimacy demands something better than either judicial apathy to the need for a separate software validation step, or perfunctory common law certification that rubberstamps vendors without validating the actual software at issue in a case. Adversarial certification offers a principled and efficient foundation for admissibility determinations of evidentiary software that remains conspicuously inadequate or lacking.
- . Earlier work has discussed this development and criticized the award of trade secrecy privilege for proprietary evidentiary software on various grounds, including constitutional concerns with due process and confrontation, and fundamental issues of justice and fairness. See, e.g., Edward J. Imwinkelried, The Admissibility of Scientific Evidence: Exploring the Significance of the Distinction Between Foundational Validity and Validity as Applied, 70 Syracuse L. Rev. 817, 829, 838 (2020) (discussing probabilistic genomic software programs and facial recognition techniques); Eli Siems, Katherine J. Strandburg & Nicholas Vincent, Trade Secrecy and Innovation in Forensic Technology, 73 Hastings L.J. 733, 775 n.2, 777 (2022) (discussing recidivism risk assessment tools and critiquing trade secrecy from an innovation policy perspective); Rediet Abebe, Moritz Hardt, Angela Jin, John Miller, Ludwig Schmidt & Rebecca Wexler, Adversarial Scrutiny of Evidentiary Statistical Software, 2022 ACM Conf. on Fairness Accountability & Transparency 1733, 1734–35, 1740, 1746 (exploring facial recognition technology and probabilistic genotyping software tools); Rebecca Wexler, Life, Liberty, and Trade Secrets: Intellectual Property in the Criminal Justice System, 70 Stan. L. Rev. 1343, 1358–59, 1375–76, 1401 (2018) (describing due process and cross-examination concerns with trade secret protection of evidentiary software); Natalie Ram, Innovating Criminal Justice, 112 Nw. U. L. Rev. 659, 692–99 (2018) (surveying Fourth Amendment, due process, and Confrontation Clause issues involving access to source code). ↑
- . See infra Part II(A)(1). ↑
- . See infra Part II(A)(2). ↑
- . Fed. R. Evid. 702(d). ↑
- . Gregory D. Schwartz, When Disciplines Disagree: The Admissibility of Computer-Generated Forensic Evidence in the Criminal Justice System, 72 UCLA L. Rev. Disc. 174, 202 (2024). The author notes:[W]itnesses . . . in [admissibility] hearings tend to be exclusively from the field in which the software operates. . . . These experts tend to base their understanding on validation studies that they or their peers have conducted . . . leav[ing] a significant gap in the information before the courts, namely expertise regarding the proper extent to rely on validation studies.Id. ↑
- . The term “red teaming” is borrowed from the cybersecurity context, where it refers to adversarial testing of software vulnerabilities. See Steven J. Hutchison, Cybersecurity: Defending the New Battlefield, Defense Acquisition Tech. & Logistics, Nov.–Dec. 2013, at 34, 38 (applying red teaming techniques to the cybersecurity domain); John F. Sandoz, Inst. Def. Analyses, Red Teaming: A Means to Military Transformation 1 n.1 (2001), https://apps.dtic.mil/sti/tr/pdf/ADA388176.pdf [https://perma.cc/S34F-LPYJ] (“The term ‘red teaming’ is commonly used to depict processes designed to bring a devil’s advocate perspective by exposing flaws and gaps in our ideas, strategies, concepts, and other new proposals.”). Other scholars have proposed mechanisms for increasing adversarial scrutiny of evidentiary software, but none that we are aware of have proposed an adversarial certification approach. See, e.g., Abebe et al., supra note 1, at 1737–38 (proposing a formal mathematical definition of robust adversarial testing of evidentiary statistical software); Steven M. Bellovin, Matt Blaze, Susan Landau & Brian Owsley, Seeking the Source: Criminal Defendants’ Constitutional Right to Source Code, 17 Ohio St. Tech. L.J. 1, 17–18, 31 (2021) (arguing that due to the nature of software, “adversarial audits—examination and testing of software by defendants—[are] necessary for a fair trial”); Maneka Sinha, The Dangers of Automated Gunshot Detection, 5 J.L. & Innovation 63, 78–80 (2023) (highlighting the necessity of independent validity testing); Jeanna Neefe Matthews et al., When Trusted Black Boxes Don’t Agree: Incentivizing Iterative Improvement and Accountability in Critical Software Systems, 2020 AAAI/ACM Conf. on AI Ethics & Soc’y 102, 107 (recommending adversarial testing “to counteract the tendency to sweep errors under the rug when they are found or to put in an inappropriate fix to make the problem go away”). ↑
- . See Daubert v. Merrell Dow Pharms., Inc., 509 U.S. 579, 589–91 (1993) (recognizing that the trial judge must ensure that all scientific evidence admitted is both relevant and reliable). ↑
- . Id. at 593 (requiring “a preliminary assessment of whether the reasoning or methodology underlying the testimony is scientifically valid and of whether that reasoning or methodology properly can be applied to the facts in issue”). ↑
- . See Tal Golan, Revisiting the History of Scientific Expert Testimony, 73 Brook. L. Rev. 879, 911–12, 940 (2008) (describing nineteenth-century concerns regarding use of unreliable scientific evidence in expert testimony). ↑
- . See, e.g., Comm. on Identifying the Needs of the Forensic Sci. Cmty., Nat’l Acad. Scis., Strengthening Forensic Science in the United States: A Path Forward 22 (2009) [hereinafter NAS Report] (identifying numerous challenges facing the forensic science community, including “little rigorous systematic research to validate the discipline’s basic premises and techniques”); President’s Council of Advisors on Sci. & Tech., Report to the President: Forensic Science in Criminal Courts: Ensuring Scientific Validity of Feature-Comparison Methods 7–14 (2016) [hereinafter PCAST Report] (finding that many twenty-first century forensic methods have not been shown to be scientifically valid). ↑
- . See, e.g., PCAST Report, supra note 10, at 42–43, 54–56 (defining and distinguishing “foundational validity” from “validity as applied”). ↑
- . Paul C. Giannelli, The Admissibility of Novel Scientific Evidence: Frye v. United States, a Half-Century Later, 80 Colum. L. Rev. 1197, 1200–01 (1980). ↑
- . Compare Frye v. United States, 293 F. 1013, 1014 (D.C. Cir. 1923) (holding that scientific expert testimony “must be sufficiently established to have gained general acceptance in the particular field in which it belongs”), with Daubert v. Merrell Dow Pharms., Inc., 509 U.S. 579, 588–89, 593–95 (1993) (rejecting the Frye “general acceptance” test as an absolute prerequisite to admissibility, and pronouncing a new multifactor test). ↑
- . Frye, 293 F. at 1014. ↑
- . Daubert, 509 U.S. at 593–94. Rule 702 was amended in 2000, following the Supreme Court’s ruling in Kumho Tire Co. v. Carmichael, 526 U.S. 137 (1999), clarifying that the standard applies to all expert testimony. Fed. R. Evid. 702 advisory committee’s note to 2000 amendment. The 2000 amendment “affirms the trial court’s role as gatekeeper and provides some general standards that the trial court must use to assess the reliability and helpfulness of proffered expert testimony.” Id. ↑
- . Daubert, 509 U.S. at 592–93, 597. ↑
- . Id. at 593–94. ↑
- . State-by-State Compendium Standards of Evidence, Nat’l Civ. Just. Inst. (July 11, 2023), https://ncji.org/wp-content/uploads/2024/01/Evidence-Standards-by-State-7.12.23.pdf [https://perma.cc/L2NE-QUDP] (reporting that thirty-one states use the Daubert standard, five use Frye, and seven use a hybrid standard); see also Rochkind v. Stevenson, 236 A.3d 630, 638 (Md. Ct. App. 2020) (explaining that a “supermajority” of states have moved to Daubert, but that a minority of states have continued to use Frye). ↑
- . See, e.g., People v. Kelly, 549 P.2d 1240, 1244 (Cal. 1976) (explaining that while the admission of new scientific evidence could be left to the trial court, California and many other states assign this responsibility to the scientific community); People v. Wesley, 633 N.E.2d 451, 456–57 (N.Y. 1994) (affirming the decision to determine the admissibility of DNA profiling evidence at a hearing while allowing a specific study to go to the jury without prior approval); Elosu v. Middlefork Ranch Inc., 26 F.4th 1017, 1026–27 (9th Cir. 2022) (holding that the district court exceeded its limited gatekeeping role by directly weighing the expert’s evidence and discrediting his ultimate conclusions without questioning his qualifications or methodology). ↑
- . See Nat’l Civ. Just. Inst., supra note 18 (collecting five jurisdictions that apply Frye: Illinois, New York, North Dakota, Pennsylvania, and Washington); Bruce Kaufman, States Slow to Adopt Daubert Scientific Evidence Rule, Bloomberg L. (June 4, 2016), https://news
.bloomberglaw.com/environment-and-energy/states-slow-to-adopt-daubert-scientific-evidence-rule [https://perma.cc/YAZ7-9RLF] (“In many of the holdout jurisdictions [that have not adopted the Daubert standard]—including California, New York, New Jersey, Illinois, Maryland, Washington and the District of Columbia—the standard for admitting expert evidence in courtrooms closely follows the century-old Frye test.”). ↑ - . See Kaufman, supra note 20 (cataloging the ways in which Frye states engage in “Daubert creep” by adding onto the standard Frye test “differently and inconsistently”). ↑
- . See Kelly, 549 P.2d at 1242, 1244 (explaining that the Frye general acceptance rule must be scaffolded by proper methodology, qualification, and procedures). In 1994, the California Supreme Court reaffirmed the Kelly-Frye test and clarified that in determining whether a forensic technique is novel (and hence subject to the test), “long-standing use by police officers seems less significant a factor than repeated use, study, testing and confirmation by scientists or trained technicians.” People v. Leahy, 882 P.2d 321, 323, 332 (Cal. 1994). ↑
- . Kelly, 549 P.2d at 1245, 1248, 1251. As discussed in Part II, infra, these same infirmities, among others, continue to plague admissibility determinations for software-based forensic evidence. ↑
- . Id. at 1248. ↑
- . Id. at 1249. ↑
- . Id. at 1250. ↑
- . Id. at 1245. ↑
- . People v. Wakefield, 195 N.E.3d 19, 28–30 (N.Y. 2022) (relying on expert testimony to determine that the methodology was generally accepted by the relevant scientific community); People v. Burrus, 200 N.Y.S.3d 655, 722 (N.Y. Sup. Ct. 2023) (considering expert testimony as well as “authoritative scientific writings introduced at the hearing” in its general acceptance inquiry); People v. Carter, No. 2573/14, 2016 WL 239708, at *4 (N.Y. Sup. Ct. Jan. 12, 2016) (specifying that general acceptance can be shown through judicial opinions). ↑
- . 633 N.E.2d 451 (N.Y. 1994). ↑
- . Id. at 456. ↑
- . Giannelli, supra note 12, at 1223 (describing the “cultural lag” critics say Frye imposes, which delays admissibility until a technique diffuses through the field and thereby keeps reliable proof out of court). ↑
- . See NAS Report, supra note 10, at 42 (“The fact is that many forensic tests . . . have never been exposed to stringent scientific scrutiny.”); PCAST Report, supra note 10, at 32–33 (summarizing findings that the forensic sciences “do not yet have a well-developed ‘research culture’” and that forensic practices have “problems stemming from the lack of a strong ‘quality culture’”). ↑
- . See Giannelli, supra note 12, at 1208–21 (cataloging courts’ recurring difficulties applying general-acceptance standards to novel techniques). ↑
- . See id. at 1214–15 (asserting that “a technician’s testimony should never suffice to establish the validity of a novel technique”). ↑
- . 509 U.S. 579, 588 (1993) (quoting Fed. R. Evid. 702). ↑
- . Id. at 592–94. ↑
- . See Lloyd Dixon & Brian Gill, Changes in the Standards for Admitting Expert Evidence, RAND Corp. (2002), https://www.rand.org/pubs/research_briefs/RB9037.html [https://perma.cc/
58NS-MJQV] (finding that following the Daubert decision, judges discussed the Daubert factors more frequently and found evidence unreliable more often). ↑ - . Daubert, 509 U.S. at 593–94. ↑
- . See id. at 594–95 (stating that the “overarching subject is the scientific validity . . . of the principles that underlie a proposed submission”). ↑
- . Fed. R. Evid. 702 advisory committee’s note to 2000 amendment. ↑
- . Fed. R. Evid. 702(a), (c). ↑
- . See Daubert, 509 U.S. at 592–93 (explaining that the Daubert factors bear on the inquiry of “whether the reasoning or methodology underlying the testimony is scientifically valid”). ↑
- . Nat’l Civ. Just. Inst., supra note 18 (reporting state-by-state adoption of Daubert). ↑
- . For a quick overview of these three competing views, see Emily C. Baker & Mary E. Desmond, Frye’d by Admissibility Standards: Does the Standard of Admissibility in State Court Make Any Difference in Practice?, Jones Day Insights 19, 20–21 (2011), https://www.jonesday.com/-/media/files/publications/2012/01/fryed-by-admissibility-standards-does-the-standard/files/fryed/fileattachment/fryed.pdf?rev=e8efd010e8ad49be8f9538754d04fc72 [https://perma.cc/4LSA-WZGX] (outlining three views on whether Frye or Daubert make any difference in state court). ↑
- . See, e.g., Michael J. Saks & David L. Faigman, Expert Evidence After Daubert, 1 Ann. Rev. L. & Soc. Sci. 105, 110 (2005) (explaining that “Daubert, in effect, brought the scientific revolution into the courtroom”). ↑
- . See David H. Kaye, How Daubert and Its Progeny Have Failed Criminalistics Evidence and a Few Things the Judiciary Could Do About It, 86 Fordham L. Rev. 1639, 1642–43 (2018) (arguing that the Daubert factors “sometimes devolved into a superficial, if not pro forma, checklist”); Sandra Guerra Thompson & Nicole Bremner Cásarez, Solving Daubert’s Dilemma for the Forensic Sciences Through Blind Testing, 57 Hou. L. Rev. 617, 627 (2020) (estimating that “a quarter of all wrongful convictions involved flawed forensic testimony” and that particularly unreliable areas include “bite mark evidence, traditional arson techniques, comparative bullet lead analysis, and microscopic hair comparisons”). ↑
- . Compare Frye v. United States, 293 F. 1013, 1014 (D.C. Cir. 1923) (admitting scientific evidence if “the thing from which the deduction is made [is] sufficiently established to have gained general acceptance in the particular field in which it belongs”), with Daubert v. Merrell Dow Pharms., Inc., 509 U.S. 579, 593–94 (1993) (stating that submission to peer review “is a component of ‘good science,’” and that “‘general acceptance’ can yet have a bearing on the inquiry”). ↑
- . See Imwinkelried, supra note 1, at 828, 838 (observing that most courts “make short shrift of the validity-as-applied issue” and arguing for a more robust analytical approach). ↑
- . See id. at 823, 825 (explaining that Frye courts are traditionally confined to the general acceptance test and that Daubert and the 2000 amendment to Rule 702 added the validity-as-applied requirement). ↑
- . People v. Leahy, 882 P.2d 321, 325 (Cal. 1994). ↑
- . People v. Kelly, 549 P.2d 1240, 1244 (Cal. 1976) (emphases omitted). ↑
- . Daubert, 509 U.S. at 593. ↑
- . See Fed. R. Evid. 702 advisory committee’s note to 2000 amendment. ↑
- . Fed. R. Evid. 702 (emphases added). The italicized provisions were added in 2023. Fed. R. Evid. 702 advisory committee’s note to 2023 amendment. ↑
- . See Fed. R. Evid. 702 advisory committee’s note to 2000 amendment. ↑
- . See NAS Report, supra note 10, at 42–43 (explaining that many common forensic tests have not been exposed to serious scientific scrutiny and that uncertain forensic processes may produce severely unjust consequences in the form of wrongful convictions). ↑
- . See PCAST Report, supra note 10, at 33 (reporting that “dozens of investigations of crime laboratories—primarily at the state and local level—have revealed repeated failures concerning the handling and processing of evidence and incorrect interpretation of forensic analysis results”). ↑
- . Id. at 4–5. ↑
- . See, e.g., Motorola Inc. v. Murray, 147 A.3d 751, 759 (D.C. 2016) (Easterly, J., concurring) (endorsing “the aid of landmark reports” such as the NAS Report and the PCAST Report); Post-PCAST Court Decisions Assessing the Admissibility of Forensic Science Evidence–Cloned, Nat’l Inst. Just. (Oct. 3, 2024), https://nij.ojp.gov/microsite-subpage/post-pcast-court-decisions-assessing-admissibility-forensic-science-evidence [https://perma.cc/Y8FK-4SMV]; Eric S. Lander, Fixing Rule 702: The PCAST Report and Steps to Ensure the Reliability of Forensic Feature-Comparison Methods in the Criminal Courts, 86 Fordham L. Rev. 1661, 1665, 1677 (2018) (discussing the PCAST Report’s recommendations and proposing an amendment to Rule 702 that incorporates validity as applied). ↑
- . Fed. R. Evid. 702 advisory committee’s note to 2023 amendment; see also Mark A. Behrens & Andrew J. Trask, Federal Rule of Evidence 702: A History and Guide to the 2023 Amendments Governing Expert Evidence, 12 Tex. A&M L. Rev. 43, 48–52 (2024) (explaining how criticisms that Rule 702 was being misapplied led to the 2023 amendment). ↑
- . Fed. R. Evid. 702 advisory committee’s note to 2023 amendment. ↑
- . One July 2024 practitioner review suggests a small degree of tightening in certain areas. See Quinn Emanuel Urquhart & Sullivan, Noted with Interest: Amendment to Federal Rule of Evidence 702, A Year in Review—July 2024, Mondaq (July 31, 2024), https://www.mondaq.com/
unitedstates/trials-appeals-compensation/1499874/noted-with-interest-amendment-to-federal-rule-of-evidence-702-a-year-in-review-july-2024 [https://perma.cc/MEG5-XVPP] (noting that some courts have set a higher bar for admission of testimony based on the 2023 amendment). For criminal cases, the significance of the reform will depend strongly on whether and how quickly it is picked up by states that have adopted a Daubert-type analysis. As of 2025, six states had amended their rules to conform to the federal reform. State Evidentiary Rule Reform: The Need for Reform in the States, Don’t Say Daubert, https://dontsaydaubert.com/state-evidentiary-rule-reform/ [https://perma.cc/9FZV-P3VD]. ↑ - . See Suzanne Bell, Sunita Shah, Thomas D. Albright, S. James Gates Jr., M. Bronner Denton & Arturo Casadevall, A Call for More Science in Forensic Science, 115 Proc. Nat’l Acad. Scis. U.S. 4541, 4542 (2018) (explaining that many forensic devices were developed within law enforcement environments outside of the scientific community without the research needed to establish validity today). ↑
- . See Peter J. Neufeld, The (Near) Irrelevance of Daubert to Criminal Justice and Some Suggestions for Reform, 95 Am. J. Pub. Health S107, S110 (2005) (noting that prosecutors have free access to government medical examiners and crime labs while defense counsel must hire independent contractors). ↑
- . See Paul C. Giannelli, Independent Crime Laboratories: The Problem of Motivational and Cognitive Bias, 2010 Utah L. Rev. 247, 250 (describing a survey of roughly three hundred crime laboratories which found that seventy-nine percent of the responding labs were located within law enforcement or public safety agencies). ↑
- . See supra note 47. ↑
- . D. Michael Risinger & Michael J. Saks, A House with No Foundation, Issues Sci. & Tech., Fall 2003, at 35, 36 (describing why research challenging the validity of law-enforcement-sponsored research on forensic evidence has been rare). ↑
- . See Giannelli, supra note 12, at 1218 (“Because the proponent has the burden of proof on the general acceptance issue, the proponent should be responsible for informing the trial court of opposing views in the literature and for explaining why the literature is not persuasive evidence of lack of general acceptance.”); see also Developments in the Law—Confronting the New Challenges of Scientific Evidence, 108 Harv. L. Rev. 1481, 1531 (1995) (proposing heightened scrutiny for expert testimony put on by the prosecution). ↑
- . See PCAST Report, supra note 10 at 4–5 (providing guidance concerning scientific standards for validity as applied and foundational validity, which requires estimates of a method’s accuracy in addition to mere results); see also Jennifer L. Mnookin, The Courts, the NAS, and the Future of Forensic Science, 75 Brook. L. Rev. 1209, 1271 (2010) (arguing that stricter scrutiny by courts is needed to incentivize adequate testing of forensic science tools and methodologies). ↑
- . E.g., Erin Murphy, No Room for Error: Clear-Eyed Justice in Forensic Science Oversight, 130 Harv. L. Rev. F. 145, 150–51 (2017) (arguing that “refusing to set an absolute (and demanding) threshold for statistical significance will, as regards forensic evidence, almost always hurt the criminal defendant”). ↑
- . E.g., Developments in the Law—Confronting the New Challenges of Scientific Evidence, supra note 68, at 1570–71. ↑
- . E.g., Keith A. Findley, Innocents at Risk: Adversary Imbalance, Forensic Science, and the Search for Truth, 38 Seton Hall L. Rev. 893, 951 (2008) (describing proposals to disallow “unsupported conclusions” or “overpowering or misleading testimony”). The 2023 amendments to Rule 702 also are aimed at this issue. See Fed. R. Evid. 702 advisory committee’s note to 2023 amendment (explaining that expert opinions must stay “within the bounds” of reliability). ↑
- . See Victor Nicholas A. Metallo, The Impact of Artificial Intelligence on Forensic Accounting and Testimony—Congress Should Amend “the Daubert Rule” to Include a New Standard, 69 Emory L.J. Online 2039, 2061 (2020) (“[T]he rules should also be amended to permit a court discretion to determine testimony inadmissible in . . . a case where . . . the AI has reached a point that ‘black box’ processes cannot be explained by human testimony, because AI has adapted the ability to program itself.”). ↑
- . E.g., Andrew Jurs, Judicial Analysis of Complex & Cutting-Edge Science in the Daubert Era: Epidemiologic Risk Assessment as a Test Case for Reform Strategies, 42 Conn. L. Rev. 49, 98–99 (2009) (proposing judicial training on methodologies of scientific research to increase judicial accuracy of evaluating evidence under Daubert); Developments in the Law—Confronting the New Challenges of Scientific Evidence, supra note 68, at 1517–19, 1540 (suggesting that the appointment of experts facilitates judicial understanding of complex scientific topics, and supporting an inference that educating judges on statistics could remedy gaps in statistical understanding across the judiciary holistically). ↑
- . E.g., Developments in the Law—Confronting the New Challenges of Scientific Evidence, supra note 68, at 1519; Findley, supra note 72, at 954–55; Jurs, supra note 74, at 84–85; Kerri N. Polizzi, Comment, How Long Do We Keep Fryeing?: The Future of Expert Scientific Evidence in California, 20 Chap. L. Rev. 393, 415–17 (2017); James R. Dillon, Expertise on Trial, 19 Colum. Sci. & Tech. L. Rev. 247, 283–84 (2018) (explaining that calls to increase use of experts at trial are the most frequently proposed reforms to fill gaps in judicial knowledge regarding scientific evidence). ↑
- . E.g., Jurs, supra note 74, at 85–86 (arguing for adaptation of Rule 706 to permit judges to appoint science consultants); Dillon, supra note 75, at 283–84 (proposing that “non-testifying” subject-matter experts advise judges in a consulting capacity). ↑
- . E.g., Developments in the Law—Confronting the New Challenges of Scientific Evidence, supra note 68, at 1593–95. ↑
- . E.g., id. at 1595–96. ↑
- . Edward Cheng, Thomas S. Kuhn and Courtroom Treatment of Science Evidence, 15 Temp. Env’t L. & Tech. J. 195, 198 (1995) [hereinafter Cheng, Courtroom Treatment of Science Evidence] (advocating for the efficacy of the Frye standard over the Daubert standard). But see Edward K. Cheng, The Consensus Rule: A New Approach to Scientific Evidence, 75 Vand. L. Rev. 407, 438 (2022) [hereinafter Cheng, The Consensus Rule] (distancing his position from Frye and articulating a new “Consensus Rule”). ↑
- . See Developments in the Law—Confronting the New Challenges of Scientific Evidence, supra note 68, at 1513, 1538 (reporting that some judges feel “ill-equipped” to evaluate complex scientific evidence and to serve as “scientific gatekeepers”). ↑
- . See discussion infra Part II(B). ↑
- . Giannelli, supra note 12, at 1247–48. ↑
- . E.g., Paul C. Giannelli, AKE v. Oklahoma: The Right to Expert Assistance in a Post-Daubert, Post-DNA World, 89 Corn. L. Rev. 1305, 1416–17 (2004) (discussing indigent defendants’ constitutional right to expert assistance under Ake v. Oklahoma, and advocating strategies that judges could employ to make access to experts more equitable for all parties); Developments in the Law—Confronting the New Challenges of Scientific Evidence, supra note 68, at 1529–31 (explaining various limitations on AKE’s application and suggesting approaches to even out the disparities in criminal cases); id. at 1561–63 (arguing that courts should be more willing to enforce defendants’ rights of access to lab records and samples). ↑
- . E.g., Christian Chessman, Note, A “Source” of Error: Computer Code, Criminal Defendants, and the Constitution, 105 Calif. L. Rev. 179, 188–89 (2017) (pointing out that certain errors can be revealed only upon inspection of source code); Bellovin et al., supra note 6, at 31 (“[S]oftware as it exists today . . . will always have bugs. Consequently, criminal defendants need access to source code to safeguard their constitutional rights.”); Eric Van Buskirk & Vincent T. Liu, Digital Evidence: Challenging the Presumption of Reliability, 1 J. Digit. Forensic Prac. 19, 25 (2006) (noting that access to source code has been mandated in some cases). ↑
- . E.g., Jennifer L. Mnookin, Repeat Play Evidence: Jack Weinstein, “Pedagogical Devices,” Technology, and Evidence, 64 DePaul L. Rev. 571, 577 (2015) (proposing the use of black-box testing to cross-examine computer-generated simulations or animations); Abebe et al., supra note 1, at 1734–37 (asserting that access to executable code is necessary to enable defense counsel to conduct direct empirical testing of the performance and validity of statistical software). ↑
- . Andrea Roth, What Machines Can Teach Us About “Confrontation”, 60 Duq. L. Rev. 210, 211–12 (2022) (proposing out-of-court confrontation of software-based evidence through testing); see also Edward K. Cheng & G. Alexander Nunn, Beyond the Witness: Bringing a Process Perspective to Modern Evidence Law, 97 Texas L. Rev. 1077, 1110 (2019) (“Far better would be a Confrontation Clause that provided defendants with enhanced discovery of the lab’s procedures and equipment.”). ↑
- . Andrea Roth, How Machines Reveal the Gaps in Evidence Law, 76 Vand. L. Rev. 1631, 1642–44 (2023) (expounding on non-cross-examination-based methods of impeachment). ↑
- . See David Alan Sklansky, Anti-Inquisitorialism, 122 Harv. L. Rev. 1634, 1688 (2009) (asserting that “public defenders and other court-appointed counsel . . . are so chronically and drastically underfunded that there is strong reason to doubt the vigor and effectiveness of the advocacy they can provide”); Roth, supra note 86, at 223 (remarking that enacting statutory protections for access to source code will “require political will and legislative approval”); Findley, supra note 72, at 929–32 (observing that the defense bar is not only underfunded but “not organized beyond the county level”). ↑
- . See NAS Report, supra note 10, at 19–20 (discussing recommendations involving the establishment of a National Institute of Forensic Science to allow for increased investment and research geared towards improving uniformity and quality); PCAST Report, supra note 10, at 14–16 (discussing recommendations centered around improving validity in forensic science through increased resources and accountability for forensic professionals); see also Brendan Max, SoundThinking’s Black-Box Gunshot Detection Method: Untested and Unvetted Tech Flourishes in the Criminal Justice System, 26 Stan. Tech. L. Rev. 193, 213–16, 240–43 (2023) (summarizing best practices for software validation testing and arguing for more “oversight,” “transparency,” and “peer-reviewed validation and error analysis” in validation testing processes); Thompson & Cásarez, supra note 46, at 647–48, 663 (presenting blind proficiency testing as a means of bolstering the reliability of forensic evidence). ↑
- . See, e.g., Van Buskirk & Liu, supra note 84, at 23–24 (positing that disseminating open source software could be more beneficial for “the quality of digital evidence”); Roth, supra note 87, at 1651 (suggesting legislation requiring free research licenses and financially independent testing). ↑
- . See NAS Report, supra note 10, at 24–25 (recommending federal funding for independent forensic labs, as well as appropriate standards for accreditation of forensic labs and certification of forensic professionals); PCAST Report, supra note 10, at 65–66 (recommending independent third-party validation studies); Sinha, supra note 6, at 78 (explaining that “[t]esting conditions matter” and that “[a]ppropriate validation testing should be independent”). ↑
- . Developments in the Law—Confronting the New Challenges of Scientific Evidence, supra note 68, at 1603–04. ↑
- . E.g., Jurs, supra note 74, at 91–97; Andrew W. Jurs, Science Court: Past Proposals, Current Considerations, and a Suggested Structure, 15 Va. J.L. & Tech. 1, 19 (2010); Justin Sevier, Redesigning the Science Court, 73 Md. L. Rev. 770, 828 (2014); Dillon, supra note 75, at 285 & n.171. ↑
- . E.g., Developments in the Law—Confronting the New Challenges of Scientific Evidence, supra note 68, at 1596–97. ↑
- . See id. at 1508–09 (explaining the applicability of Daubert to mass tort litigation and its inapplicability to the criminal justice context). ↑
- . See Org. Sci. Area Comms. for Forensic Sci., OSAC Registry Implementation Survey: 2022 Report 1, 3, 49–54 app. A (2023), https://www.nist.gov/system/files/
documents/2023/03/01/2022%20OSAC%20REGISTRY%20IMPLEMENTATION%20REPORT_FINAL_March2023.pdf [https://perma.cc/8LLC-KJ78] (providing data showing widespread successful implementation of forensic science standards for crime laboratories proposed by OSAC, which is administered by NIST); see also Measuring the Impact of Implementation: 2023, Nat’l Inst. of Standards & Tech. (Mar. 30, 2024), https://www.nist.gov/adlp/spo/organization-scientific-area-committees-forensic-science/measuring-impact-implementation [https://perma.cc/5Q2L-5US6] (cataloging implementation of OSAC standards by forensic science service providers from 31 different states). ↑ - . Rich Press, Two New Forensic DNA Standards Added to the OSAC Registry, Nat’l Inst. of Standards & Tech. (May 12, 2020), https://www.nist.gov/news-events/news/2020/05/two-new-forensic-dna-standards-added-osac-registry [https://perma.cc/9GMV-WGM7] (“Compliance with these new standards, as well as almost all forensic science standards in the United States, is voluntary.”). ↑
- . See, e.g., Nat’l Inst. of Standards & Tech., OSAC Standards Bulletin January 2025 1–2 (2025), https://www.nist.gov/magazine/osac-standards-bulletin/january-2025 [https://
perma.cc/U5Q9-AFBQ] (reporting six new OSAC standards effective January 2025); Nat’l Inst. of Standards & Tech., OSAC Standards Bulletin November 2024 1 (2024), https://www.nist.gov/magazine/osac-standards-bulletin/november-2024 [https://perma.cc/FM99-C5E7] (reporting two new OSAC standards); Nat’l Inst. of Standards & Tech., OSAC Standards Bulletin May 2024 1 (2024), https://www.nist.gov/magazine/osac-standards-bulletin/may-2024 [https://perma.cc/BV2S-RRCG] (reporting two new OSAC standards). ↑ - . See Frequently Asked Questions on CODIS and NDIS, Fed. Bureau of Investigation, https://www.fbi.gov/how-we-can-help-you/dna-fingerprint-act-of-2005-expungement-policy/codis
-and-ndis-fact-sheet [https://perma.cc/6EHK-Q8QE] (laying out the rigorous requirements a state must agree to abide by to participate in the National DNA index); Biometrics and Fingerprints, Fed. Bureau of Investigation, https://le.fbi.gov/science-and-lab/biometrics-and-fingerprints/
codis-2 [https://perma.cc/W2CM-L5EA] (stating that laboratories with a federal nexus are required to demonstrate compliance). ↑ - . See Fed. Bureau of Investigation, Quality Assurance Standards for Forensic DNA Testing Laboratories 2 (2025), https://le.fbi.gov/file-repository/forensic-qas-070125.pdf/
view [https://perma.cc/P5JX-KGMN] (stating that the standards for quality assurance requirements shall be followed by participating laboratories). ↑ - . See infra notes 135–138 and accompanying text. ↑
- . Giannelli, supra note 12, at 1231–32. ↑
- . Gary Edmond & Andrew Roberts, Procedural Fairness, the Criminal Trial and Forensic Science and Medicine, 33 Syd. L. Rev. 359, 389–90 (2011). ↑
- . PCAST Report, supra note 10, at 14–15. ↑
- . E.g., Findley, supra note 72, at 972; see also Neufeld, supra note 64, at S113 (proposing the creation of a national “institute of forensic science” managed jointly by scientists and legal scholars). ↑
- . See Findley, supra note 72, at 956 (suggesting that such a body need not determine questions of admissibility but could “enhance the courts’ ability to make admissibility determinations”). The certification of reliability we refer to here should be distinguished from the certification of laboratories and technicians that sometimes occurs because those certifications focus on things such as good laboratory practices and training in protocols for performing a particular sort of analysis, rather than on the reliability (and scope of reliability) of the technique itself. See NAS Report, supra note 10, at 195 (explaining that accreditation means that a laboratory adheres to an established set of standards, but does not ensure that those standards are reliable). ↑
- . See Brian J. Gestring, Creating Infrastructure and Incentives to Increase Quality in Forensic Science, 7 Forensic Sci. Int’l, no. 100435, 2023, at 1, 1 (noting that, even 14 years after the NAS Report, there are no requirements for validated techniques, formal training, or proficiency testing). ↑
- . Highway Safety Programs; Model Specifications for Devices to Measure Breath Alcohol, 58 Fed. Reg. 48705, 48707–10 (Sep. 17, 1993); Conforming Products List of Evidential Breath Alcohol Measurement Devices 82 Fed. Reg. 50940, 50941–44 (Nov. 2, 2017). ↑
- . What the NHTSA Conforming Products List Means: Why It Matters, AlcoPro (Jan. 13, 2026), https://alcopro.com/what-the-nhtsa-conforming-products-list-means-why-it-matters/ [https://perma.cc/W46R-QLTM] (“Most states adopt NHTSA’s Model Specifications either directly or by reference when establishing their own breath alcohol testing standards.”). ↑
- . See, e.g., State v. Tindell, No. E200802635CCAR3CD, 2010 WL 2516875, at *13 (Tenn. Crim. App. June 22, 2010) (applying a relaxed admissibility standard for breath testing instruments because the instruments themselves and their procedures have become familiar); State v. Warner, No. 2012-P-0121, 2013 WL 5346699, at *1–2 (Ohio Ct. App. 2013) (stating that the Intoxilyzer’s reliability has been legislatively determined); Commonwealth v. Camblin, 31 N.E.3d 1102, 1110 (Mass. 2015) (treating NHTSA approval of breath testing devices as controlling the court’s analysis); United States v. Kyle, No. 15-MJ-4176-DHH, 2015 WL 6755223, at *4–6 (D Mass. Nov. 3, 2015) (applying Camblin’s reasoning); Commonwealth v. Camblin, 86 N.E.3d 464, 471 (Mass. 2017) (declaring that NHTSA certification is widely accepted by courts as evidence of a breath testing device’s reliability); Commonwealth v. Ananias, No. 1248CR1075, 2017 WL 11473590, at *5 (Mass. Dist. Ct. Feb. 16, 2017) (accepting a device’s inclusion on NHTSA’s conforming products list as evidence supporting reliability). ↑
- . E.g., State v. Son Yong Kish, No. A08-1342, 2009 WL 2432284, at *5 (Minn. Ct. App. 2009); see also Rebecca Wexler, Convicted by Code, Slate (Oct. 6, 2015), https://slate.com/
technology/2015/10/defendants-should-be-able-to-inspect-software-code-used-in-forensics.html [https://perma.cc/NAP8-WNZ4] (“When defense experts identified a bug in breathalyzer software, the Minnesota Supreme Court barred the affected test from evidence in all future trials.”). ↑ - . See, e.g., State v. Chun, 943 A.2d 114, 154–60 (N.J. 2008) (split ruling on several software errors that were revealed after gaining access to source code); Ananias, 2017 WL 11473590, at *13 (granting defendant’s motion to exclude measurements from a breathalyzer device calibrated and certified at a particular period of time). ↑
- . Lauren Kirchner, Federal Judge Unseals New York Crime Lab’s Software for Analyzing DNA Evidence, ProPublica (Oct. 20, 2017), https://www.propublica.org/article/federal-judge-unseals-new-yorkcrime-labs-software-for-analyzing-dna-evidence [https://perma.cc/T29B
-SMT6]; Lauren Kirchner, ProPublica Seeks Source Code for New York City’s Disputed DNA Software, ProPublica (Sep. 25, 2017), https://www.propublica.org/article/propublica-seeks-source-code-for-new-york-city-disputed-dna-software [https://perma.cc/YPG2-M8GR]. ↑ - . See Jonathan J. Koehler, Jennifer L. Mnookin & Michael J. Saks, The Scientific Reinvention of Forensic Science, 120 Proc. Nat’l Acad. Scis. U.S. 1, 7 (2023) (noting that, in spite of the serious critiques and guidance provided by the 2016 PCAST report, “it has been business as usual in most post-PCAST cases”). ↑
- . See Paul W. Grimm, Challenges Facing Judges Regarding Expert Evidence in Criminal Cases, 86 Fordham L. Rev. 1601, 1602–03 (2018) (explaining from his experience as a district judge that “trial judges are privy to very few of the underlying facts of a case . . . before the trial” and “can feel like they are in a battle of wits, unarmed”). ↑
- . See, e.g., Koehler et al., supra note 114, at 2 (describing the complexities of proficiency testing). ↑
- . See Roth, supra note 86, at 212, 215, 217, 221–22 (describing the rise of machine evidence, establishing a right to confrontation of such evidence, and proposing means of confrontation for software evidence). ↑
- . See H.R. Doc. No. 118-33, at 19 (2023) (clarifying that Rule 702 is governed by the preponderance standard of Rule 104(a)). The federal rules clarified that the court practice of ruling that “critical questions of the sufficiency of an expert’s basis, and the application of the expert’s methodology, are questions of weight and not admissibility” is incorrect. Id. With this clarification, it is likely that states will reform their rules to reflect that standard. If so, the heightened standard for admissibility will inevitably make opposition to that testimony more likely to succeed. ↑
- . See supra notes 33–34 and accompanying text (highlighting issues with determining the appropriate field pre-Daubert); infra notes 185–187 (discussing uncritical reliance on prior judicial opinions to determine admissibility during the same time period). ↑
- . See infra Part II(A)(2). ↑
- . See, e.g., People v. Davis, 290 Cal. Rptr. 3d 661, 683 (Cal. Ct. App. 2022) (eliding consideration of software experts as members of the relevant scientific community in evaluating STRmix forensic biology software); United States v. Reynolds, 86 F.4th 332, 346 (6th Cir. 2023) (considering only the application of cell site location software without discussing its implementation). ↑
- . Giannelli, supra note 12, at 1200–01 (“The reliability of evidence derived from a scientific principle depends upon three factors: (1) the validity of the underlying principle; (2) the validity of the technique applying that principle; and (3) the proper application of the technique on a particular occasion.”). ↑
- . See Schwartz, supra note 5, at 178–79 (“[E]xpert witnesses tend to speak to automated processes that they have merely overseen . . . without ever seeing the underlying source code.”). ↑
- . Wexler, supra note 1, at 1373–74 (noting that source code review “can be extremely difficult even for experts”). ↑
- . The sample was derived primarily from a LEXIS search for cases mentioning “software” or “source code” as well as “Daubert,” “Frye,” or admissib(le)/(ility). We discarded about half of the results because they did not contain substantive discussion of the admissibility of software-based evidence. We have not attempted a “scorched earth” approach to locate all relevant cases, so these numbers should be taken to be illustrative rather than definitive. Our primary use of this sample of cases is for qualitative analysis to identify common themes in judicial treatment. ↑
- . Specifically, technologies addressed in our case set included probabilistic genotyping and other DNA analysis (74 cases), breath alcohol analysis (73 cases), software used to extract, analyze, or map cell phone information (50 cases), ShotSpotter technology for gunshot detection (12 cases), child pornography detection (9 cases), ballistics analysis (4 cases), face recognition (4 cases), sound and image enhancement (3 cases), “Cybercheck” technology (3 cases), crime scene and accident reconstruction (2 cases), fingerprint matching (2 cases), gunshot residue analysis (1 case), ankle monitor GPS data (1 case), vehicle “black box” analysis (1 case), and use of AI for generating expert reports (1 case). We hypothesize that the predominance of challenges to breath alcohol analysis stems from the relatively longstanding use of these devices, as well as from the likelihood that there is an unusually large number of well-resourced defendants in DUI cases. For probabilistic genotyping, we speculate that the relatively large number of challenges stems both from the seriousness of many of the cases where this technology is used (often murder cases) and from the fact that some large public defender offices have established specialized DNA units. ↑
- . See, e.g., State v. Simmer, 935 N.W.2d 167, 180–82 (Neb. 2019) (determining that the evidence generated by TrueAllele was admissible under the Daubert/Schafersman standard); People v. Wesley, 633 N.E.2d 451, 455 (N.Y. 1994) (holding that DNA profiling evidence was admissible under the Frye standard); People v. Gholston, No. 350798, 2021 WL 2181079, at *16–18 (Mich. Ct. App. May 27, 2021) (applying the Daubert standard in concluding that DNA evidence was admissible). But see, e.g., State v. Pickett, 246 A.3d 279, 284 (N.J. Super. Ct. App. Div. 2021) (holding that defendant is entitled to access of a novel probabilistic genotyping software’s source code to challenge reliability at a Frye hearing); State v. Daughters-White, No. CR-20083625, 2009 Ariz. Super. LEXIS 791, at *2–3 (Ariz. Super. Ct. Aug. 3, 2009) (holding that the State is obligated to disclose the software’s contents to the defense). ↑
- . See Wexler, supra note 1, at 1346 (arguing that privileging trade secrets in the criminal context is harmful and unjust); Siems et al., supra note 1, at 776, 778 (identifying the consequences of judges denying disclosure of probabilistic software codes on trade secrecy grounds in criminal trials). ↑
- . See, e.g., Luna-Galacia v. State, 892 S.E.2d 50, 58 (Ga. Ct. App. 2023) (“[T]he results of breathlyzer tests are clearly admissible. . . . Luna-Galacia’s arguments concerning the software update and the officer’s administration of the test went to the evidence’s weight, not its admissibility.”); United States v. Washington, No. 8:19-CR-299, 2020 WL 3265142, at *4 n.4 (D. Neb. June 16, 2020) (“Questions about whether the updated versions of software materially altered the reliability go to weight of the evidence and not admissibility.”). ↑
- . 195 N.E.3d 19 (N.Y. 2022). ↑
- . Id. at 390 (Rivera, J., concurring). ↑
- . Id. at 390 n.2. The concurrence added that the failure to validate the software program after each major update is a failure to comply with the Scientific Working Group on DNA Analysis Methods (SWGDAM) Guidelines, which require that probabilistic genotyping software undergo a new validation study whenever there is a “significant” software change. Id. (citing Guidelines for the Validation of Probabilistic Genotyping Systems § 5.2 (Sci. Working Group on DNA Analysis Methods 2015), https://www.swgdam.org/_files/ugd/4344b0_22776006b67c4a32a5ffc04fe3b56515.pdf [https://perma.cc/PQ3Y-2VEX]). ↑
- . See supra note 113; infra note 220. ↑
- . See supra note 129 and accompanying text. ↑
- . See, e.g., Document Notices & Disclaimers, IEEE Standards Assoc., https://standards
.ieee.org/ipr/disclaimers/ [https://perma.cc/79FR-FX4N] (“Use of an IEEE standard is wholly voluntary.”); Sci. Working Grp. on DNA Analysis Methods, Validation Guidelines for DNA Analysis Methods 1–2 (2016), https://www.swgdam.org/_files/ugd/4344b0_
813b241e8944497e99b9c45b163b76bd.pdf [https://perma.cc/8XMP-HLWW] (distinguishing SWGDAM guidelines from mandatory standards). ↑ - . E.g., Marc Canellas, Defending IEEE Software Standards in Federal Criminal Court, Computer, June 2021, at 14, 16–17. ↑
- . Schwartz, supra note 5, at 201–02 (noting that “disregard of computer science standards and expertise is commonplace in Daubert hearings”). ↑
- . See People v. Burrus, 200 N.Y.S.3d 655, 720 (N.Y. Sup. Ct. 2023) (considering SWGDAM guidelines in its analysis); Commonwealth v. Ananias, No. 1248-CR-1075, 2017 WL 11473590 at *7 (Mass. Super. Ct. Feb. 16, 2017) (weighing expert testimony and concluding that there is no universal standard of “programming best practices” because there is “an inherent element of artistic expression”); United States v. Anderson, 673 F. Supp. 3d 671, 684 (M.D. Pa. 2023) (addressing dispute regarding which software standards should govern probabilistic genotyping software). ↑
- . See supra notes 129–134 and accompanying text. ↑
- . See, e.g., People v. Saibu, D-054980, 2011 BL 8562, at *19–20 (Cal. Ct. App. Jan. 4, 2011) (declining to require a Kelly hearing for use of a Photoshop enhancement tool). ↑
- . People v. Landybraun, No. D-074488, 2019 WL 6463185, at *5 (Cal. Ct. App. Dec. 2, 2019); see also Villareal-Garcia v. State, 671 S.W.3d 791, 794 (Tex. App.—Dallas 2023) (holding that Cellebrite technology did not require expert testimony or a reliability predicate because it was “so simple and so plainly verifiable”). ↑
- . See, e.g., United States v. Jones, No. 3:21-CR-89-BJB, 2022 WL 17884450, at *5 (W.D. Ky. Dec. 23, 2022) (discussing the line-drawing problem of deciding what specialized understanding of technology would “stretch the reasoning and background knowledge of a lay juror”). ↑
- . See, e.g., N.Y.C. Off. of Chief Med. Exam’r, STRmix Probabilistic Genotyping Software Operating Instructions 6 (2019), https://www.nyc.gov/assets/ocme/downloads/
pdf/technical-manuals/protocols-for-forensic-str-analysis/strmix-probabilistic-genotyping-software-operating-instructions.pdf [https://perma.cc/2F9V-7RRN]. ↑ - . See, e.g., id. at 1 (noting that the number of contributors to a sample “must” be determined before using the probabilistic genotyping software, STRmix); N.Y.C. Off. of Chief Med. Exam’r, STR Results Interpretation—PowerPlex Fusion & STRmix 16 (2019), https://www.nyc
.gov/assets/ocme/downloads/pdf/technical-manuals/protocols-for-forensic-str-analysis/str-results-interpretation-powerplex-fusion-and-strmix.pdf [https://perma.cc/9E3V-3EDR] (recognizing that “the number of contributors may be unclear” and stating that “analysts should use their professional judgment when assessing the number of contributors”). ↑ - . See, e.g., Siems et al., supra note 1, at 780, 783–84 (asserting that probabilistic genotyping is subjective and error-prone because it depends heavily on discretionary assumptions made by technicians). ↑
- . See, e.g., United States v. Williams, 583 F.2d 1194, 1198 (2d Cir. 1978) (noting that the “[s]election of the ‘relevant scientific community’ influence[s] the result” and can lead to differing consensus on the admissibility of scientific evidence); Giannelli, supra note 12, at 1208–10 (describing how pre-1980 courts struggled to define the relevant scientific community). ↑
- . See Giannelli, supra note 12, at 1208 (“The general acceptance standard as set forth in Frye appears to require a two-step analysis: first, identifying the field in which the underlying principle falls, and second, determining whether that principle has been generally accepted by members of the identified field. Neither step is free of difficulties.”). ↑
- . See, e.g., Simon A. Cole, Out of the Daubert Fire and into the Fryeing Pan? Self-Validation, Meta-Expertise and the Admissibility of Latent Print Evidence in Frye Jurisdictions, 9 Minn. J.L. Sci. & Tech. 453, 478–79 (2008) (offering criticisms of a practitioner-based approach to defining the scope of the relevant scientific community); Giannelli, supra note 12, at 1209–10 (highlighting that the Frye test can create too malleable an expert field); Edmond & Roberts, supra note 103, at 388–89 (pointing out that defining who comprises a community is subjective in practice); Polizzi, supra note 75, at 400–01 (highlighting courts’ struggle to “identify which scientific community may be the most relevant to the proffered evidence”). ↑
- . See, e.g., People v. Burrus, 200 N.Y.S.3d 655, 658 (N.Y. Sup. Ct. 2023) (finding that a witness was qualified as an expert in the field of forensic biology without more explanation); State v. Pratt, 128 A.3d 883, 891 (Vt. 2015) (stating that “the software was generally accepted in the computer forensic community worldwide” without providing a definition of the community). ↑
- . See, e.g., United States v. Lewis, No. 18-CR-194, 2020 U.S. Dist. LEXIS 38705, at *62–64 (D. Minn. Jan. 6, 2020) (finding that probabilistic genotyping satisfies the general acceptance criterion of Daubert because it is widely used in forensic laboratories); Commonwealth v. Camblin, 86 N.E.3d 464, 474 (Mass. 2017) (rebuffing argument that selling a device exclusively to law enforcement agencies means there is no scientific community); see also People v. Bullard-Daniel, 42 N.Y.S.3d 714, 721 (N.Y. Cnty. Ct. 2016) (rejecting defendant’s objection that the relevant scientific community was being defined as “an insular community of professionals whose careers and livelihoods focus on the prosecution of criminal cases”). ↑
- . See United States v. Downing, 753 F.2d 1224, 1236 (3d Cir. 1985) (“Thus, some courts, when they wish to admit evidence, are able to limit the impact of Frye by narrowing the relevant scientific community to those experts who customarily employ the technique at issue.”); Cole, supra note 148, at 473–76 (collecting state court opinions rejecting a practitioner-based definition of the relevant scientific community); id. at 478 (“Evidence scholars also agree that practitioner communities alone cannot satisfy the general acceptance requirement.”). ↑
- . See, e.g., Stephanie L. Damon-Moore, Note, Trial Judges and the Forensic Science Problem, 92 N.Y.U. L. Rev. 1532, 1556 (2017) (“One of the most problematic aspects of forensic science is that, because virtually all forensic disciplines . . . were developed for law enforcement purposes, there is no neutral ‘scientific community’ reviewing the methodologies developed in the field.”); Findley, supra note 72, at 943 (“The only place [certain] ‘experts’ exist—because the only place these ‘sciences’ exist—is in the government crime laboratories or spin-off private laboratories whose roots are in the law enforcement community.”). ↑
- . See, e.g., Bullard-Daniel, 42 N.Y.S.3d at 716–17, 719, 721 (accepting a “DNA technical leader” for a forensic science lab as an expert witness for DNA-testing software but rejecting an academic who teaches biological science at a university and specializes in bioinformatics). ↑
- . See, e.g., Giannelli, supra note 12, at 1209–10 (“[I]f the ‘specialized field’ is too narrow, the consensus judgment mandated by Frye becomes illusory; the judgment of the scientific community becomes, in reality, the opinion of a few experts.”); id. at 1214–15 (arguing that “a technician’s testimony should never suffice to establish the validity of a novel technique” because technicians follow prescribed routines in applying a given technique and are not expected to understand its underlying fundamentals); Findley, supra note 72, at 906–07 (“[M]ost crime laboratories are set up as an arm of law enforcement, either as a unit within a police department or within a State Department of Justice.”); Cole, supra note 148, at 478 (identifying the relevant scientific community consisting of practitioners who depend on validation of the technique to maintain their professional reputations and further their commercial interests as a significant issue underlying general acceptance); Developments in the Law—Confronting the New Challenges of Scientific Evidence, supra note 68, at 1496–97 (identifying “manipulations and strained interpretations” that result from “courts’ rather creative or even bad-faith application of doctrinal standards”). ↑
- . See Schwartz, supra note 5, at 184–88, 198, 210 (explaining that familiarity with the application of forensic software alone is insufficient to ensure reliability and validity because expertise as to the underlying scientific principles and techniques, as well as the ability to understand whether error-resistant practices are necessary and whether they have been properly implemented are also critical). ↑
- . See id. at 176–77, 203–04 (arguing courts do not apply “the same scrutiny to software evidence that they have applied to other similarly complex scientific evidence” when they choose not to rely on software experts). ↑
- . See, e.g., United States v. Morgan, 292 F. Supp. 3d 475, 485 (D.D.C. 2018) (“[T]his Court does not require an expert to have an in-depth knowledge of all the algorithms underlying their technological tools . . . to reliably testify about the[ir] outputs.”); J.A.R. v. State, 374 So.3d 25, 30–31 (Fla. Dist. Ct. App. 2023) (“[A]n ‘expert is not required to have an in-depth knowledge of all the algorithms underlying their technological tools—such as hardware and software—to reliably testify about the outputs of those tools.’” (quoting Walker v. State, 308 So.3d 193, 198 (Fla. Dist. Ct. App. 2020))); United States v. Chiaradio, 684 F.3d 265, 277–78 (1st Cir. 2012) (“Although Agent Gordon was not a programmer, did not know the program’s authors, and had never seen the source code, he had significant specialized experience with both EP2P and the manual re-creation of EP2P sessions. . . . These showings sufficiently evinced the reliability of EP2P.”); United States v. Ramsey, No. 21-CR-495, 2023 WL 2523193, at *16–18 (E.D.N.Y. Mar. 15, 2023) (admitting expert testimony because “[t]he government’s description of [the expert’s] methodology generally tracks defendant’s description of the typical methodology used by CLSI experts”); State v. Pratt, 128 A.3d 883, 891–92 (Vt. 2015) (“[I]t is not required—nor is it practical—for an investigator to have expertise in or knowledge about the underlying programming, mathematical formulas, or other innerworkings of the software.”). But see United States v. Jones, No. 3:21-CR-89-BJB, 2022 WL 17884450, at *1–3 (W.D. Ky. Dec. 23, 2022) (finding detailed maps produced by a software program to be inadmissible where the witness was unable to discuss its reliability, errors, or limitations, causing the output to appear to have emerged from an indecipherable computer program); People v. Harvey, No. 319482, 2015 WL 8953522, at *9 (Mich. Ct. App. Dec. 15, 2015) (finding error when admitting expert testimony that lacks an explanation of the forensic tools or software used to produce the evidence at issue). ↑
- . See, e.g., Villareal-Garcia v. State, 671 S.W.3d 791, 794 (Tex. App.—Dallas 2023, no pet.) (determining that expert testimony was not required to admit Cellebrite data-transfer technology); People v. Davis, 290 Cal. Rptr. 3d 661, 683 (Cal. Ct. App. 2022) (rejecting defendant’s claim that the prosecution should be required to show that a software-based method is “generally accepted by computer software engineers”); People v. Wakefield, 195 N.E.3d 19, 26–27 (N.Y. 2022) (upholding denial of defendant’s motion to preclude an expert’s testimony because “TrueAllele was generally accepted in the relevant scientific community”); State v. Simmer, 935 N.W.2d 167, 181 (Neb. 2019) (“While a review of the TrueAllele source code might also have confirmed the reliability of TrueAllele, we cannot say that the district court abused its discretion by relying on the numerous validation studies confirming the reliability of TrueAllele by other means.”). ↑
- . See Matthews et al., supra note 6, at 104 (explaining that even when evidentiary software packages “are designed to answer the same forensic question, they can produce very different results”). But see Commonwealth v. Camblin, 31 N.E.3d 1102, 1112 (Mass. 2015) (“[T]he judge should have held a hearing to determine whether the source code and other challenged features of the Alcotest functioned in a manner that reliably produced accurate breath test results.”); State v. Pickett, 246 A.3d 279, 293 (N.J. Super. Ct. App. Div. 2021) (“TrueAllele’s source code has never been scrutinized by any party outside of Cybergenetics; therefore, the validation studies produced by the State to date are limited.”). ↑
- . See Joshua A. Kroll, Joanna Huey, Solon Barocas, Edward W. Felten, Joel R. Reidenberg, David G. Robinson & Harlan Yu, Accountable Algorithms, 165 U. Pa. L. Rev. 633, 639 (2017) (“[T]he purpose of computer-mediated decisionmaking is to bring decisions an element of scale, where the same rules are ostensibly applied to a large number of individual cases or are applied extremely quickly.”); Bellovin et al., supra note 6, at 21 (“Computers derive much of their power from . . . the ability to repeat activities.”). ↑
- . See, e.g., Giannelli, supra note 12, at 1215–18 (discussing significant division among courts over the use of expert testimony, scientific and legal writings, and judicial opinions to establish general acceptance); Edmond & Roberts, supra note 103, at 389 (pointing out that “the views of ‘community’ members will often be the subject of speculation” and may vary depending on context); Cheng, The Consensus Rule, supra note 79, at 456 (offering criticism that general acceptance under Frye is an unworkable standard because it is “both difficult to prove and easy to manipulate”). ↑
- . See Cheng, The Consensus Rule, supra note 79, at 414–19 (asking why lay judges and juries should be trusted to choose between experts who will inevitably disagree); Findley, supra note 72, at 929 (worrying that because of resource asymmetries, there often is no real “battle” when testing forensic sciences in criminal cases); see also Christopher Tarver Robertson, Blind Expertise, 85 N.Y.U. L. Rev. 174, 177 (2010) (“In almost every case, the factfinder sees a ‘battle of the experts,’ each selected by, affiliated with, compensated by, and apparently biased toward a particular party.”). ↑
- . See supra note 152 and accompanying text. ↑
- . See Chessman, supra note 84, at 181 (asserting that software-generated evidence “merits additional scrutiny” because “the complexity of computer programs makes it difficult for jurists and computer programmers alike to detect errors”); Kroll et al., supra note 160, at 639 (“Most individuals are ill-equipped to review how computerized decisions are made, even if those decisions are reached transparently.”). ↑
- . See Matthews et al., supra note 6, at 104 (stating that the results of competing forensic software packages “are not routinely compared in real casework” and that “more work is required to fully understand when, how and why these systems differ—between programs, within releases of the same program and with variations in parameters”). ↑
- . See Chessman, supra note 84, at 218 (“Peer review can only occur when peers in the field actually review the program’s source code. . . . For source codes that remain secret . . . it is difficult for a field to accept that which it does not know.”). ↑
- . See, e.g., Kathryn A. Maupin, Laura P. Swiler & Nathan W. Porter, Validation Metrics for Deterministic and Probabilistic Data, J. Verification Validation & Uncertainty Quantification, Sep. 2018, at 031002-1, 031002-2, 031002-3 (explaining that probabilistic metrics are more removed from physical meaning as their calculations depend more on uncertainties, and that deterministic metrics are generally preferable to probabilistic metrics due to their longer history of use). ↑
- . See Geoffrey Stewart Morrison, Advancing a Paradigm Shift in Evaluation of Forensic Evidence: The Rise of Forensic Data Science, 5 Forensic Sci. Int’l, no. 100272, 2022, at 1, 2 (“Forensic practitioners are susceptible to cognitive bias when they are making subjective judgements and are exposed to information that could influence their degree of belief in the probability that a hypothesis is true.”); Abebe et al., supra note 1, at 1736 (collecting studies showing that “slight variation of the test distribution” can cause performance of machine learning systems to “drop dramatically”). ↑
- . See, e.g., Mnookin, supra note 69, at 1245–46 (arguing that corroboration of fingerprint evidence by other forensic examiners cannot be equated with true scientific peer review). ↑
- . See, e.g., People v. Kiriakus, No. 355962, 2022 WL 17998499, at *6 (Mich. Ct. App. Dec. 29, 2022) (“The fact that Garza was not an expert in every aspect of cell-phone towers was a suitable subject for cross-examination, but did not render Garza unfit to testify as an expert in cell-phone mapping.”); Walker v. State, 308 So.3d 193, 198 (Fla. Dist. Ct. App. 2020) (stating that an expert is not required to possess deep knowledge of all underlying algorithms of their technological tools to reliably testify about their outputs); State v. Pratt, 128 A.3d 883, 891–92 (Vt. 2015) (requiring knowledge in the use of the software, but not in the underlying workings of the software); State v. Young, 854 S.E.2d 615, 619 (S.C. Ct. App. 2021) (holding that the fact that the expert did not know the developer of mapping software or whether the software had been peer reviewed did not show that he “lacked the requisite knowledge, training, or experience, or that his testimony was unreliable”). ↑
- . See, e.g., United States v. Springstead, 520 F.App’x 168, 170 (4th Cir. 2013) (upholding lower court’s qualification of a witness as an expert based on its finding that the witness had been tested on his familiarity with and ability to operate the software and rejecting defendant’s claim that the witness lacked requisite expertise to explain how the software worked). ↑
- . See, e.g., United States v. Turner, No. 19-13704, 2022 WL 4137756, at *4 (11th Cir. Sep. 13, 2022) (holding that the witness’s lack of specialized training in the forensic software “had no bearing on the admissibility of her testimony” because the witness never used the software to produce data); United States v. Smith, No. 21-CR-30003-DWD, 2022 WL 17741100, at *7 (S.D. Ill. Dec. 16, 2022) (“[S]o long as the witness does not veer off into the realm of an expert by attempting to explain technology and forensic processes for which he has no specialized knowledge . . . Rule 702’s reliability requirements are not implicated.”); United States v. Jimenez-Chaidez, 96 F.4th 1257, 1269 (9th Cir. 2024) (declining to require expert testimony because “[t]he Government limited the scope of [the witness]’s testimony to his use of the Cellebrite software and his perceptions of the data that the software produced that are readily understandable without having him opine about the software’s technical processes or reliability or other issues that require specialized knowledge”). But see, e.g., United States v. Jones, No. 3:21-CR-89-BJB, 2022 WL 17884450, at *1 (W.D. Ky. Dec. 23, 2022) (“If highly consequential evidence emerges from what looks like an indecipherable computer program to most non-scientists, non-statisticians, and non-programmers, it is imperative that qualified individuals explain how the program works and ensure that it produces reliable information about the case.” (quoting United States v. Gissantaner, 990 F.3d 457, 463 (6th Cir. 2021))). ↑
- . Daubert v. Merrell Dow Pharms., Inc., 509 U.S. 579, 592–93 (1993); see also supra notes 36–38 and accompanying text. ↑
- . See, e.g., Luna-Galacia v. State, 892 S.E.2d 50, 58 (Ga. Ct. App. 2023) (“Luna-Galacia’s arguments concerning the software update and the officer’s administration of the test went to the evidence’s weight, not its admissibility.”); Kiriakus, 2022 WL 17998499, at *6 (noting that the expert’s “education and training did not explain exactly how cell phones connect to different towers if signals are overlapping” but characterizing this issue as “one of weight, not qualification or admissibility”); Pratt, 128 A.3d at 893 (rejecting defendant’s concern about forensic expert’s lack of knowledge at the “programmatic level” by stating that “any deficiencies in the [cell phone data extraction] program should be drawn out through the adversarial process”). ↑
- . While it is “the province of the jury” to decide the weight of the admitted evidence, Barefoot v. Estelle, 463 U.S. 880, 902 (1983), Rule 702 requires the court first to make a “preliminary assessment” of admissibility. Daubert, 509 U.S. at 580. ↑
- . See Wexler, supra note 1, at 1364 (showing that “defendants have struggled unsuccessfully to overcome claims of trade secret evidentiary privilege even at trial, where their procedural rights are strongest,” as well as “before or after trial, when their procedural rights are thin”). ↑
- . See id. at 1375–76, 1401–02 (collecting commentary); Bellovin et al., supra note 6, 42–43 (“Depriving [criminal] defendants access to the underlying code risks depriving defendants of their rights.”); see also supra note 1 (collecting commentary). ↑
- . Siems et al., supra note 1, at 795. ↑
- . E.g., Brown v. State, 512 P.3d 269, 280 (Nev. 2022) (upholding lower court’s decision to limit cross-examination of witness “to protect proprietary rights in trade secrets”); People v. Wakefield, 195 N.E.3d 19, 32 (N.Y. 2022) (Rivera, J. concurring) (finding error in the court’s decision to admit “DNA results developed using the TrueAllele methodology, even though at the time its source code and underlying algorithms were kept from independent evaluators and the defense as trade secrets”); United States v. Morgan, 292 F. Supp. 3d 475, 481 n.3, 485 (D.D.C. 2018) (noting that the software program algorithm is “proprietary” and downplaying importance of software access to independently evaluate the algorithm’s accuracy). For opinions allowing defendants access to source code or critiquing lack of access, see, e.g., State v. Arteaga, 296 A.3d 542, 554 (N.J. Super. Ct. App. Div. 2023) (reversing denial of defendant’s discovery request and emphasizing that independent review of source code is critical to determinations of reliability); State v. Pickett, 246 A.3d 279, 283 (N.J. Super. Ct. App. Div. 2021) (“Hiding the source code is not the answer. The solution is producing it under a protective order.”); United States v. Jones, No. 3:21-CR-89-BJB, 2022 WL 17884450, at *9 (W.D. Ky. Dec. 23, 2022) (refusing to admit into evidence data produced by proprietary software that is “unavailable for scrutiny or cross-examination”). ↑
- . See Siems et al., supra note 1, at 806 (explaining that courts “piggyback on previous admissibility determinations, even from outside of their own jurisdictions,” resulting in a “rich-get-richer network effect” in forensic technology admissibility). ↑
- . See id. at 807 (scrutinizing courts’ distinct tendency to cite earlier judicial admissibility decisions as evidence of “general acceptance”); Giannelli, supra note 12, at 1208, 1212 (discussing the difficulties of applying the Frye standard to novel scientific evidence); id. at 1218–19 (criticizing an “approach to the Frye test that emphasizes previous court decisions, considering general acceptance not only by scientists but also by courts” (quotation marks and citation omitted)). ↑
- . See supra note 47 and accompanying text. ↑
- . See Siems et al., supra note 1, at 807 (arguing that this compounding of judicial admissibility rulings “conflate[s] the familiar concept of persuasive legal precedent with the more relevant question of acceptance by the scientific community”). ↑
- . See Daubert v. Merrell Dow Pharms., Inc., 509 U.S. 579, 596 (1993) (holding that conventional devices, including “[v]igorous cross-examination,” are the appropriate safeguards for scientific testimony rather than wholesale exclusion). ↑
- . See, e.g., State v. Garajau, No. A-2807-18, 2021 WL 1750140, at *17 (N.J. Super. Ct. App. Div. May 4, 2021) (“[A] scientific theory may be accepted based on its high profile in judicial opinions.”); People v. Carter, No. 2573/14, 2016 WL 239708, at *4 (N.Y. Sup. Ct. Jan. 12, 2016) (“General acceptance of a scientific principle or procedure may be shown through legal writings and judicial opinions.”). ↑
- . See, e.g., People v. Lund, 279 Cal. Rptr. 3d 697, 714 (Cal. Ct. App. 2021) (asserting that the Kelly/Frye test applies only to expert testimony which is based on a novel technique, process or theory); State v. Villanueva, No. 36694-4-III, 2020 WL 7396544, at *12 (Wash. Ct. App. Dec. 17, 2020) (“Washington uses the Frye standard for determining the admissibility of novel scientific evidence.”). ↑
- . Giannelli, supra note 12, at 1218–19 (quoting Nat’l Rsch. Council, On the Theory and Practice of Voice Identification 45 (1979)). ↑
- . Findley, supra note 72, at 950. Other scholars have expressed similar views. See Damon-Moore, supra note 152, at 1564–65 (arguing that once a new scientific technique is approved on appeal, its precedent may control subsequent trials and run the risk of “grandfathering in irrationality” without reexamination under Daubert (quoting United States v. Green, 405 F. Supp. 2d 104, 123 (D. Mass. 2005))); Mnookin, supra note 69, at 1243–44 (describing the “ostrich maneuver” taken by judges under Daubert in response to challenges to longstanding forensic techniques, as illustrated by a judge concluding that the evidence was reliable because it had been tested “in adversarial proceedings with the highest possible stakes—liberty and sometimes life” (quoting United States v. Havvard, 117 F. Supp. 2d 848, 854 (S.D. Ind. 2000))). ↑
- . See, e.g., United States v. Gordon, No. 1:20-CR-00049-RJA-MJR, 2022 U.S. Dist. LEXIS 238133, at *17, *19 (W.D.N.Y. Sep. 23, 2022) (giving great deference to other courts’ determinations of reliability of a particular DNA analysis software); United States v. Lockett, No. 20-00091-BAJ-RLB, 2023 WL 7181251, at *5–8 (M.D. La. Nov. 1, 2023) (rejecting defendant’s Daubert challenge by relying on “recent opinions from [courts in other jurisdictions] hold[ing] that probabilistic genotyping . . . satisfies Rule 702 and the Daubert reliability factors”); United States v. Anderson, No. 1:23-CR-09-HAB-ALT, 2025 WL 3763461, at *5 (N.D. Ind. Dec. 30, 2025) (invoking authority from other federal courts to conclude that “the Government has met its threshold burden under Daubert that [the expert] reliably applied STRmix to the relevant DNA samples in the case”); Whittley v. State, No. 05-21-00534-CR, 2022 WL 3645589, at *7 (Tex. App.—Dallas Aug. 24, 2022, no pet.) (taking “judicial notice of the many opinions from other jurisdictions holding the [STRmix] software satisfies the state and federal equivalents of Texas Rule of Evidence 702”); People v. Wakefield, 107 N.Y.S.3d 487, 492 (N.Y. App. Div. 2019), aff’d, 195 N.E.3d 19 (N.Y. 2022) (noting that TrueAllele “had been deemed admissible in Virginia, Pennsylvania, and California” in ruling that the lower court’s finding of general acceptance was proper). But see State v. Rochat, 269 A.3d 1177, 1207–09, 1211 (N.J. Super. Ct. App. Div. 2022) (finding other courts’ evaluations of the reliability of a scientific method unpersuasive). ↑
- . See, e.g., Brandon L. Garrett & Peter J. Neufeld, Invalid Forensic Science Testimony and Wrongful Convictions, 95 Va. L. Rev. 1, 9 (2009) (finding that in sixty percent of exonerations, “forensic analysts called by the prosecution provided invalid testimony”). ↑
- . See PCAST Report, supra note 10, at 148–50 (finding a lack of foundational validity to support bitemark analysis, firearms analysis, and footwear analysis, as well as significant bias and proficiency testing issues with latent fingerprint analysis); NAS Report, supra note 10, at 142, 149, 154–55, 159–61, 173–76, 178–79 (documenting weak or absent scientific foundations in the evaluation of forensic analysis methods of friction ridges, impressions, toolmarks, firearms, hair, and bitemarks). ↑
- . Paul C. Giannelli, Forensic Science: Daubert’s Failure, 68 Case W. Res. L. Rev. 869, 884–86 (2018). ↑
- . See supra notes 157, 170–172, 179 and accompanying text. ↑
- . See Giannelli, supra note 12, at 1200–03 (enunciating tripartite framework); supra note 122 and accompanying text; see also Findley, supra note 72, at 950 (arguing that reliance on prior judicial decisions establishing reliability of a given technique as applied to one set of circumstances improperly “minimizes the need for repeated, case-by-case determination”). Some commentators distinguish primarily between “foundational validity” and “validity as applied.” See, e.g., PCAST Report, supra note 10, at 4–5 (“We distinguish here between two types of scientific validity: foundational validity and validity as applied.”); Imwinkelried, supra note 1, at 819 (noting that the PCAST Report’s distinction between the two types of scientific validity “is an important one”). We find Giannelli’s tripartite distinction to be more useful for discussing forensic software. Indeed, as we discussed above, courts’ frequent failure to recognize the need for an intermediate analysis is particularly problematic for software. See supra Part II(A). ↑
- . Compare People v. Lund, 279 Cal. Rptr. 3d 697, 716 (Cal. Ct. App. 2021) (“‘Computer programming is not a scientific theory or technique . . . and it does not implicate the [c]ourt’s responsibility to keep “junk science” out of the courtroom.’” (quoting United States v. Blouin, No. CR16-307 TSZ, 2017 WL 3485736, at *7 (W.D. Wash. Aug. 15, 2017))), and Taylor v. State, 101 N.E.3d 865, 871 (Ind. Ct. App. 2018) (characterizing the witness’s expertise and testimony as “not ‘scientific’ in nature” and more correctly called “‘technical’ or ‘specialized’ knowledge”), and Luna-Galacia v. State, 892 S.E.2d 50, 58 (Ga. Ct. App. 2023) (refusing to consider the validity of breathalyzer software updates at the admissibility stage, because the “breathalyzer is hardly ‘novel’ scientific evidence” and is “clearly admissible”), with United States v. Reynolds, 86 F.4th 332, 346 (6th Cir. 2023) (evaluating the reliability of a software program’s “general technique” separate from the reliability of that program’s application), and State v. Ghigliotty, 232 A.3d 468, 485 (N.J. Super. Ct. App. Div. 2020) (requiring Frye hearing where the witness’s ultimate conclusion was informed by “a novel software product” and the witness was not “expert[] in the science behind the [software] system”). ↑
- . See Giannelli, supra note 12, at 1202 (“If the theory of voice uniqueness is not valid, voiceprint evidence is not reliable. If, however, the uniqueness of the human voice were established, it would not necessarily follow that the voiceprint technique is capable of detecting that uniqueness.”). ↑
- . See Bellovin et al., supra note 6, at 22–26 (explaining the many ways software bugs arise); Abebe et al., supra note 1, at 1737 (stating that when a statistical software tool is validated as accurate across specific distributions of instances, it “does not imply any rigorous guarantees for a previously unseen instance”). ↑
- . See Chessman, supra note 84, at 182 (explaining that information gleaned from viewing a software program in action is “highly limited” and that the “only way to completely understand how—and whether—a program works is by reading the program’s source code”). ↑
- . See id. at 189–92 (noting endemic problems with software updates and software rot); see also FDA, Deciding When to Submit a 510(k) for a Software Change to an Existing Device 6–7 (2017), https://www.fda.gov/regulatory-information/search-fda-guidance-documents/
deciding-when-submit-510k-software-change-existing-device [https://perma.cc/6PNX-72WU] [hereinafter FDA, Deciding When to Submit a 510(k)] (noting that “[s]oftware modifications may trigger additional or unplanned consequences” and that any risk-based assessment “should be confirmed by successful, routine verification and validation activities”). ↑ - . See Matthews et al., supra note 6, at 106 (finding “substantial differences in the results generated by three [probabilistic genotyping] systems as well as the large impact of varying parameters whose contributions are often not clearly appreciated”); Bellovin et al., supra note 6, at 33–34 (stating that “software testing is done according to the program’s requirements and specifications,” while cautioning that often “the specifications themselves are faulty,” and explaining further that “[s]oftware testing is not random” but rather is carefully crafted to meet anticipated boundary conditions or inputs). ↑
- . Bellovin et al., supra note 6, at 39 (pointing, for example, to “[d]eviation from good coding practice” as a “red flag” that usually corresponds to a deeper problem in the system); see also supra note 159 and accompanying text. ↑
- . E.g., Luna-Galacia v. State, 892 S.E.2d 50, 58 (Ga. Ct. App. 2023) (holding evidence acquired from an updated version of breathalyzer was admissible based on the admissibility of an older model). ↑
- . E.g., State v. Ramirez, 425 P.3d 534, 544 (2018), as amended on reconsideration in part (holding that proprietary software lacking external validation was admissible because cell site location testimony is “widely accepted throughout the country” and the “theories behind the . . . testimony were sound”). ↑
- . E.g., United States v. Jones, S4 15-CR-153, 2018 WL 2684101, at *10 (S.D.N.Y. June 5, 2018) (justifying the use of Forensic Statistical Tool (FST) by reasoning in part that “nearly every court to have considered the FST has found it to be a reliable tool”); State v. Pratt, 128 A.3d 883, 890 (Vt. 2015) (discussing expert knowledge of Cellebrite software and noting, “several other courts have taken a liberal approach in admitting testimony . . . upon reliability foundations similar to that laid by the State here”); United States v. Reynolds, 86 F.4th 332, 347 (6th Cir. 2023) (stating that “the function that [the cellphone-location software] performs has general (perhaps universal) acceptance” based on its acceptance by previous judicial decisions). But see United States v. Reynolds, No. 1:20-CR-24, 2021 WL 3750156, at *2–5 (W.D. Mich. Aug. 25, 2021) (recognizing that the reliability of the particular software must be established by analyzing the Daubert factors); People v. Lund, 279 Cal. Rptr. 3d 697, 718–20 (Cal. Ct. App. 2021) (requiring detailed evidence to establish reliability of the technique). ↑
- . No. 05-21-00534-CR, 2022 WL 3645589 (Tex. App.—Dallas Aug. 24, 2022, no pet.). ↑
- . Id. at *16–18. ↑
- . See id. at *11–18 (examining only the software’s underlying scientific theory and technique). ↑
- . See, e.g., Developments in the Law—Confronting the New Challenges of Scientific Evidence, supra note 68, at 1513–14 (describing how judges “feel ill-equipped to fulfill their new role as scientific gatekeepers” under the “vague” Daubert standard and therefore engage in “avoidance techniques”); Findley, supra note 72, at 945–49 (discussing the “major paradox of judicial gatekeeping . . . that those to whom the law assigns the responsibility for screening such evidentiary offerings have no particular expertise for conducting those evaluations” (quoting Michael Saks, The Aftermath of Daubert: An Evolving Jurisprudence of Expert Evidence, 40 Jurimetrics 229, 230 (2000))); Damon-Moore, supra note 152, at 1555–58 (discussing challenges judges face when applying Daubert, including judges’ lack of scientific knowledge); Polizzi, supra note 75, at 405–08 (critiquing trial judges’ lack of scientific background in the context of making admissibility determinations of scientific evidence); Dillon, supra note 75, at 264–69 (discussing the longstanding and “intractable” problem of judges’ lack of “epistemic competence” to evaluate scientific evidence). ↑
- . See supra text accompanying notes 193, 195–207. ↑
- . Some issues might remain triable at the admissibility stage, such as whether “the expert’s opinion reflects a reliable application of the principles and methods to the facts of the case.” Fed. R. Evid. 702. ↑
- . See Déirdre Dwyer, (Why) Are Civil and Criminal Expert Evidence Different?, 43 Tulsa L. Rev. 381, 390 (2007) (“[M]uch of the expert evidence presented at a criminal trial is the product of disciplines that have been developed for the criminal process.”); NAS Report, supra note 10, at 128–79 (identifying the most common forensic tools and their uses in criminal cases). ↑
- . See Dwyer, supra note 211, at 391 (“[T]he evidential issue which the expert evidence is intended to prove is often very different. In many civil actions, the expert evidence is concerned with general states of affairs. . . . However, most of the expert evidence in criminal actions is concerned with linking the defendant specifically to the crime.”). ↑
- . In such cases, judges will have to distinguish between forensic evidence tools and other forms of case-specific scientific evidence, but we are fairly confident that this question, unlike the question of technical reliability and validity, is one that judges are well-equipped to handle. ↑
- . See Kroll et al., supra note 160, at 646 (“Software code is, ultimately, a rigid and exact description of itself: the code both describes and causes the computer’s behavior when it runs.”); Bellovin et al., supra note 6, at 19 (“It is a truism that computers do only what they are told to do.”). To the extent that a scientific method requires use of randomized values or stochastic modeling, the software implementation can nonetheless be validated to ensure proper implementation of that scientific method. See Kroll et al., supra note 160, at 653–54 (discussing the added challenges of validating randomized processes and the need to “determine that the source of that randomness and its incorporation into the process under scrutiny meets [the intended] goals”); Schwartz, supra note 5, at 189–90 (discussing source-code review as an important supplemental method of assessing the reliability of probabilistic evidentiary software). By contrast, inscrutable tools—including many AI models based on machine learning techniques—should be treated as incapable of validation. See Katherine J. Strandburg, Rulemaking and Inscrutable Automated Decision Tools, 119 Colum. L. Rev. 1851, 1876–77 (2019) (flagging, inter alia, problems when input features cannot be mapped to outcome variables). ↑
- . See Kroll et al., supra note 160, at 651 (noting that logs are among the easiest and most common ways to review a software program’s behavior). ↑
- . See Jennifer L. Mnookin, The Validity of Latent Fingerprint Identification: Confessions of a Fingerprinting Moderate, 7 L. Probability & Risk 127, 129–30 (2008) (collecting scholarship on the lack of standards for finding a fingerprint “match” and concerns over “observer effects” in the fingerprint examining field); Jennifer L. Mnookin, Of Black Boxes, Instruments, and Experts: Testing the Validity of Forensic Science, 5 Episteme 343, 346–47 (2008) (explaining that latent fingerprint identification “lacks any formalized specifications” for declaring a match and depends on examiners’ discretionary judgments rather than fixed, testable criteria); Cole, supra note 148, at 467–68 (suggesting latent print evidence often survives Daubert not because its accuracy or validity has been empirically demonstrated but because it benefits from a high degree of “truthiness”); Garrett & Neufeld, supra note 190, at 19 (noting that some forensic disciplines involve “subjective analyses not premised on empirical population data” and documenting unsupported probability and frequency claims in criminal trials). ↑
- . NAS Report, supra note 10, at 87. ↑
- . See, e.g., id. at 139 (observing that method of examining latent prints lacks “particular measurements or a standard test protocol,” relies on examiner judgment, and is “not necessarily repeatable from examiner to examiner”). ↑
- . See Kroll et al., supra note 160, at 646, 651 (explaining that software code precisely determines a program’s behavior and that logging is a common technique for recording program actions). Of course, there are still ways to cheat, for example, by tampering with the inputs or by falsifying the outputs. See, e.g., Moritz Contag et al., How They Did It: An Analysis of Emission Defeat Devices in Modern Automobiles, 2017 IEEE Symp. on Sec. & Priv. 231, 232, https://doi.org/10.1109/SP.2017.66 [https://perma.cc/RZ2Z-HVZJ] (explaining Volkswagen’s and Fiat’s use of software-based “defeat device[s]” to evade automobile emissions testing standards). Secure or tamper-evident logs can prevent a forensic analyst from claiming to have used different input data or to have received different outputs than was actually the case. See Kroll et al., supra note 160, at 651–52 (mentioning access-controlled audit logs as a means of providing reliable review). Questions of input-based tampering can be raised at the individual admissibility stage. ↑
- . See Siems et al., supra note 1, at 789–90 (discussing investigation that revealed the New York Office of the Chief Medical Examiner had recoded portions of its Forensic Statistical Tool without informing the state oversight commission or running another full validation study). ↑
- . See, e.g., Margaret Hagan, Participatory Design for Innovation in Access to Justice, Daedalus, Winter 2019, at 120 (discussing several tools of participatory design); Ryan Hagemann, Jennifer Huddleston Skees & Adam Thierer, Soft Law for Hard Problems: The Governance of Emerging Technologies in an Uncertain Future, 17 Colo. Tech. L.J. 37, 49 (2018) (discussing multistakeholder processes that have proliferated within soft law modes of governance); Brett M. Frischmann, Michael J. Madison & Katherine J. Strandburg, Governing Knowledge Commons, in Governing Knowledge Commons 1, 19–34 (Brett M. Frischmann, Michael J. Madison & Katherine J. Strandburg eds., 2014) (illustrating a knowledge commons framework, which describes the relationships among resources, participants, and governance structures in knowledge commons); Christopher R. Berry & Jacob E. Gersen, Agency Design and Political Control, 126 Yale L.J. 1002, 1009–14 (2017) (surveying legal scholarship on agency design, including themes of political control, political responsiveness, accountability, and political insulation); Bryan H. Choi, Institutional Choice for Software Safety Standards, 73 Hastings L.J. 1461, 1468–69 (collecting scholarship on dynamic interactions between courts and agencies, and the information-generating functions of each institution). ↑
- . See Siems et al., supra note 1, at 801–03, 806 (describing the admissibility-driven markets for evidentiary software, which create significant switching costs and first-mover advantages due to the uncertainty of whether new software tools will be deemed admissible in court). ↑
- . See Dolores R. Wallace., Laura M. Ippolito & D. Richard Kuhn, Nat’l Inst. of Standards & Tech., Special Publ’n No. 500-204, High Integrity Software Standards and Guidelines 16 (1992) (observing that “[t]here is little agreement” on recommendations regarding software engineering practices for high-integrity software); FDA, General Principles of Software Validation 5–6 (2002), https://www.fda.gov/regulatory-information/search-fda-guidance-documents/general-principles-software-validation [https://perma.cc/3ENX-B7QH] [hereinafter FDA General Principles] (observing that “it is not possible to state in one document all of the specific validation elements that are applicable” and that “[s]oftware verification and validation are difficult because a developer cannot test forever, and it is hard to know how much evidence is enough”). ↑
- . See Kroll et al., supra note 160, at 645 (“[T]esting code after it has been written, however extensively, cannot provide true assurance of how the system works, because any analysis of an existing computer program is inherently and fundamentally incomplete.”); FDA, Computer Software Assurance for Production and Quality System Software 6 (2025), https://www.fda.gov/regulatory-information/search-fda-guidance-documents/computer-software-assurance-production-and-quality-management-system-software [https://perma.cc/P2TX-8NNB] [hereinafter FDA, Computer Software Assurance] (explaining that “software testing alone is often insufficient to establish confidence that the software is fit for its intended use” and recommending instead a “risk-based approach for establishing confidence that software is fit for its intended use”). ↑
- . See, e.g., FDA, Postmarket Management of Cybersecurity in Medical Devices 13 (2016), https://www.fda.gov/regulatory-information/search-fda-guidance-documents/
postmarket-management-cybersecurity-medical-devices [https://perma.cc/H259-PGTU] [hereinafter FDA, Postmarket Management of Cybersecurity] (recommending that medical device software manufacturers maintain robust software life cycle processes that include monitoring throughout the total product life cycle, as well as verification and validation mechanisms for software updates and patches); see also Nadia Eghbal, Working in Public: The Making and Maintenance of Open Source Software 132 (2020) (collecting estimates that maintenance costs account for between 40 to 70 percent of total time and expenditures on software development); Wilma M. Osborne, Nat’l Bureau of Standards, Special Publ’n No. 500-130, Executive Guide to Software Maintenance 6 (1985) (“Software maintenance represents 60%–70% of the total cost of software . . . . While total software costs have risen rapidly, the ratio of development to maintenance costs has remained relatively constant.”). ↑ - . See Inst. of Elec. & Elecs. Eng’rs, Standard No. 1012-2024, IEEE Standard for System, Software, and Hardware Verification and Validation 25 (2025), https://doi.org/10.1109/IEEESTD.2025.11134780 [https://perma.cc/CLZ2-PM8V] [hereinafter IEEE Standard 1012] (defining “minimum tasks” to mean “Verification and Validation (V&V) tasks required by the integrity level assigned to the system, software, or hardware to be verified and validated”); see also FDA General Principles, supra note 223, at 5 (“Validation of software . . . has been conducted in many segments of the software industry for well over 20 years.”). ↑
- . See Kroll et al., supra note 160, at 643–45 (listing popular approaches to software engineering that make code easier to analyze, and therefore validate, and noting that a “thriving industry builds tools to assist in the development of software”); see also Boaz Sangero, Safety from Flawed Forensic Sciences Evidence, 34 Ga. St. U. L. Rev. 1129, 1130, 1208 (2018) (proposing that forensic evidence should be inadmissible in court “unless it has been developed as a ‘safety-critical system’” using modern safety methods from other fields such as medical devices or avionics); Brandon L. Garrett & Cynthia Rudin, Testing AI 15 (Duke L. Sch. Pub. L. & Legal Theory Series, Working Paper No. 2024-62, 2024), https://ssrn.com/abstract=4948789 [https://perma.cc/4FJT-RR9X] (“[N]ot all AI systems affect people’s rights or are used in court. However, if AI systems potentially do affect people’s rights, then we need to know how accurate they are . . . [by] conduct[ing] adequate scientific testing.”). ↑
- . See Bellovin et al., supra note 6, at 17–18, 31–40 (“[T]he nature of software, and hence of computer programming, is such that certain errors are more likely to be found by adversarial testing.”); Abebe et al., supra note 1, at 1742 (insisting that validation must be performed by an adversarial body and not merely a neutral oversight committee); Jelena Mirkovic, Peter Reiher, Christos Papadopoulos, Alefiya Hussain, Marla Shepard, Michael Berg & Robert Jung, Testing a Collaborative DDoS Defense in a Red Team/Blue Team Exercise, 57 IEEE Transactions on Computs. 1098, 1098 (2008) (“Red Team testing formally separates researchers into teams taking on the attacker (Red Team) and the defender (Blue Team) roles, which leads to more realistic test scenarios.”). ↑
- . See Nat’l Bureau of Standards, FIPS Publ’n No. 101, Guideline for Lifecycle Validation, Verification, and Testing of Computer Software 4 (1983) (defining software validation as “determin[ing] the correctness of the final program or software with respect to the software requirements”; software verification as “employ[ing] integrity and evolution checking to determine internal consistency and completeness”; and software testing as “examin[ing] program behavior by executing the program on sample data sets”); FDA General Principles, supra note 223, at 6 (“[The] FDA considers software validation to be confirmation by examination and provision of objective evidence that software specifications conform to user needs and intended uses, and that the particular requirements implemented through software can be consistently fulfilled.”) (internal quotation marks omitted); see also IEEE Standard 1012, supra note 226, at 27 (defining “validation” as providing evidence that the requirements solve the right problem, and confirming that those requirements have been fulfilled). ↑
- . See Kroll et al., supra note 160, at 646 (“The specification of a system is a critical question for assessment . . . . Without strong evidence of a computer system’s correctness, even the author of that system cannot reliably claim that it will behave according to a desired policy, and no policymaker or overseer should believe such a claim.”); FDA General Principles, supra note 223, at 21 (“Test plans and test cases should be created as early in the software development process as feasible.”). ↑
- . See Nat’l Bureau of Standards, supra note 229, at 9 (“Three types of analysis (static, dynamic, formal) are available and each provides the . . . analyst with different types of specific information about the solution being examined.”); Kroll et al., supra note 160, at 646–47 (describing static and dynamic methods, the latter of which include observational methods and testing methods). ↑
- . See Kroll et al., supra note 160, at 646–47, 650–51 (explaining static analysis as looking at the code without running the program, and dynamic testing as running the program in either black-box or white-box settings). ↑
- . See id. at 651, 661 & n.91 (“Intuitively, white-box evaluation is more powerful . . . . Computer scientists . . . have shown that black-box evaluation of systems is the least powerful of a set of available methods for understanding and verifying system behavior. . . . Specifically, white-box testing, in which an analyst has access to the source code under test, is generally considered to be superior.”); Bellovin et al., supra note 6, at 9 (discussing limitations of black-box testing); Siems et al., supra note 1, at 790 (“[P]roblematic behavior, limitations, and mistakes . . . are unlikely to be detected by . . . black box testing.”); see also Stephen Casper et al., Black-Box Access Is Insufficient for Rigorous AI Audits, 2024 ACM Conf. on Fairness, Accountability & Transparency 2254, 2258–59, https://doi.org/10.1145/3630106.3659037 [https://perma.cc/AP6E-C7MW] (detailing limitations of black-box methods and advantages of white-box access). ↑
- . See FDA General Principles, supra note 223, at 21 (outlining source code traceability analysis as an important tool for software verification); B. Scott Andersen & George Romanski, Verification of Safety-Critical Software, Commc’ns ACM, Oct. 2011, at 54 (describing traceability as a mapping from each requirement down to the related source code and verification data); Won Keun Youn, Seung Bum Hong, Kyung Ryoon Oh & Oh Sung Ahn, Software Certification of Safety-Critical Avionic Systems: DO-178C and Its Impacts, IEEE Aerospace & Elec. Sys. Mag., Apr. 2015, at 4, 8–9 (noting the importance of bidirectional traceability to ensure that “orphan source code and dead source code are not inadvertently produced,” which can otherwise create unexplained ghost behaviors). ↑
- . See Siems et al., supra note 1, at 776–77 (collecting critical commentary of use of trade secrecy to deny access to source code); Wexler, supra note 1, at 1351, 1392–95 (discussing the debate surrounding secrecy and disclosure surrounding black-box methods); Kroll et al., supra note 160, at 658 (describing necessity of secrecy to discourage strategic behavior). ↑
- . See Nat’l Bureau of Standards, supra note 229, at 9 (“Static analysis detects errors through the examination of the product. It focuses on the form and structure of the solution, but not the functional or computational aspects. It is also the technique used to examine all document items at all phases of development.”). ↑
- . Bellovin et al., supra note 6, at 72. ↑
- . See Kroll et al., supra note 160, at 647–48 (explaining how code can be obfuscated, hiding simple errors including embedded dependencies); Brittany Johnson et al., Why Don’t Software Developers Use Static Analysis Tools to Find Bugs?, 35 Int’l Conf. Software Eng’g 672, 676–77, 681 (2013), https://doi.org/10.1109/ICSE.2013.6606613 [https://perma.cc/8AYL-UGEB] (finding that automated static analysis tools are not widely used because they provide too many false positives and are cumbersome to apply). ↑
- . See Kroll et al., supra note 160, at 647 (observing that static analysis omits external data, preventing outputs from reacting contextually); Abebe et al., supra note 1, at 1736 (“[L]ooking at the source code without also running the software executable may reveal little about the statistical tool’s performance (e.g., accuracy or error rates).”). ↑
- . Nat’l Bureau of Standards, supra note 229, at 10 (“Dynamic analysis is the process of determining the validity of a program and of detecting errors by studying the program’s response to a set of input data. It addresses the functional, structural, and computational aspects.”). ↑
- . See Inst. of Elec. & Elecs. Eng’rs, SWEBOK Guide to the Software Engineering Body of Knowledge v4.0a, at 5-6 to 5-8 (Hironori Washizaki ed., 2025), https://ieeecs-media.computer.org/media/education/swebok/swebok-v4.pdf [https://perma.cc/Y4TL-FHF6] (“[T]est cases can be designed to check that the functional specifications are correctly implemented, which is variously referred to in the literature as conformance testing, correctness testing or functional testing. However, several other non-functional properties may be tested as well, including performance, reliability, and usability.”). ↑
- . See id. at 5-8 to 5-10 (“Non-functional testing targets the validation of non-functional aspects (such as performance, usability, or reliability).”). Examples of non-functional testing include penetration testing, load testing, user interface testing, and cross-platform compatibility testing. See Lawrence Chung & Julio Cesar Sampaio do Prado Leite, On Non-Functional Requirements in Software Engineering, in Conceptual Modeling: Foundations and Applications 363, 364 (Alexander T. Borgida, Vinay K. Chaudhri, Paolo Giorgini & Eric S. Yu eds., 2009) (collecting several taxonomies of non-functional requirements). But see Jonas Eckhardt, Andreas Vogelsang & Daniel Méndez Fernández, Are “Non-functional” Requirements Really Non-functional?, 38 Int’l Conf. on Software Eng’g 832, 841, (2016) https://doi.org/10.1145/
2884781.2884788 [https://perma.cc/9YXC-PXPT] (arguing that “most ‘non-functional’ requirements are misleadingly declared as such because they actually describe behavior of the system” and accordingly should be “handled similarly to functional requirements”). ↑ - . Kroll et al., supra note 160, at 650 n.48 (citing Edward Tsang, Combinational Explosion, U. Essex (May 12, 2005), http://cswww.essex.ac.uk/CSP/ComputationalFinanceTeaching/
CombinatorialExplosion.html [https://perma.cc/R7KE-4BJD]); see also id. at 652 (describing the more general problem of noncomputability, which limits the theoretical ceiling on effectiveness of testing). ↑ - . See Earl T. Barr, Mark Harman, Phil McMinn, Muzammil Shahbaz & Shin Yoo, The Oracle Problem in Software Testing: A Survey, 41 IEEE Transactions on Software Eng’g 507, 507 (2015) (noting that “the problem of automatically generating test inputs has been the subject of research interest for nearly four decades”); Afonso Fontes & Gregory Gay, The Integration of Machine Learning into Automated Test Generation: A Systematic Mapping Study, 33 Software Testing Verification & Reliability, no. 4, June 2023, at 3, 4–5 (surveying common automated test-generation techniques, and contrasting it to human-driven testing, which “is often not amenable to automation”). ↑
- . See Nat’l Bureau of Standards, supra note 229, at 10 (“Formal analysis uses rigorous mathematical techniques to analyze the algorithms or properties of a solution. It can provide a strong statement regarding certain properties of a solution including correctness, but is limited by the difficulty of application and lack of automated support.”). ↑
- . David Lorge Parnas, Really Rethinking ‘Formal Methods,’ Computer, Jan. 2010, at 28, 28 (observing that formal methods remains a “popular research area” but that “[a]pplications of formal methods to industrial practice remain such exceptions that they confirm that the use of formal methods is not common practice”); Laura Humphrey, Example Applications of Formal Methods to Aerospace and Autonomous Systems, 2023 Int’l Conf. on Assured Autonomy 67, 67 (2023) (reporting that “our experience is that formal methods are not generally being used by these organizations [the United States Air Force and more broadly the Department of Defense], even within groups focused on verification and certification”). Despite this persistent lack of commercial uptake, researchers express perennial optimism that things will change in the future. See, e.g., Kroll et al., supra note 160, at 665 (“Software verification is a rapidly developing field, and the costs of building fully verified software will likely drop precipitously in the coming decades, leading to wide adoption in the software industry.”). ↑
- . See Parnas, supra note 246, at 30 (noting several problems with formal methods, including that the models often are “more complex than the code” and “oversimplify the problem by ignoring many of the ugly details that are likely to lead to bugs,” and that the “problems are exacerbated in larger examples”). ↑
- . See Nat’l Bureau of Standards, supra note 229, at 11 (“Formal analysis techniques may be manually applied to a design specification if the specification is sufficiently formal and exact. . . . The formal analysis of a design specification can be improved by using automated symbolic execution tools. Such tools can be expensive to create and operate.”). ↑
- . See, e.g., Mirkovic et al., supra note 228, at 1098 (explaining the concepts of “Red Team” and “Blue Team” approaches in the cybersecurity context); see also Exec. Order No. 14,110, 88 Fed. Reg. 75191, 75196 (Oct. 30, 2023) (ordering NIST to establish guidelines for “AI red-teaming tests”). ↑
- . Sandoz, supra note 6, at 1 n.1; see also John F. Sandoz, Inst. Def. Analyses, Red Teaming: Shaping the Transformation Process 2 (2001), https://apps.dtic.mil/sti/tr/pdf/
ADA398285.pdf [https://perma.cc/7W42-2AGV] (describing the red-team process in military contexts). ↑ - . Red Team/Blue Team Approach, Nat’l Inst. of Standard & Tech. (last visited Apr. 5, 2026), https://csrc.nist.gov/glossary/term/red_team_blue_team_approach [https://perma.cc/
62RR-PDXD] [hereinafter Red Team/Blue Team Approach]. ↑ - . See Abebe et al., supra note 1, at 1736 (noting that many existing validation studies “do not rigorously assess evidentiary software because the authors or reviewers of such studies have financial or professional interests in ensuring that forensic laboratories use the software”); Sinha, supra note 6, at 82 (stating that validation research on ShotSpotter technology “is not appropriately classified as independent because it was commissioned by [the vendor]”). ↑
- . Red Team/Blue Team Approach, supra note 251. ↑
- . See supra notes 167–168 and accompanying text. ↑
- . See, e.g., Katriel Cohn-Gordon, Cas Cremers, Benjamin Dowling, Luke Garratt & Douglas Stebila, A Formal Security Analysis of the Signal Messaging Protocol, 33 J. Cryptology 1914, 1917 (2020) (providing an independent security audit of an open-source messaging protocol); cf. Beth Simone Noveck, “Peer to Patent”: Collective Intelligence, Open Review, and Patent Reform, 20 Harv. J.L. & Tech. 123, 144 (2006) (“Making participation [in patent examination] open and subject to self-selection can leverage not only the ‘wisdom of the crowd’ but also its enthusiasm.”). ↑
- . Jeanne Fromer, Unregulated Certification Marks, 69 Stan. L. Rev. 121, 127–28 (2017); see also Margot E. Kaminski, Binary Governance: Lessons from the GDPR’s Approach to Algorithmic Accountability, 92 S. Cal. L. Rev. 1529, 1600 (2019) (describing certification as “a softer coregulatory mechanism” that uses “market measures to incentivize industry participation”). ↑
- . See supra Parts II(A)(1)(a) & II(B). ↑
- . See Sinha, supra note 6, at 80 (“Not all jurisdictions that use ShotSpotter have conducted their own testing and it is not clear how rigorous testing has been in those jurisdictions that have.”); Jeanna Matthews et al., The Right to Confront Your Accusers: Opening the Black Box of Forensic DNA Software, 2019 AAAI/ACM Conf. on AI Ethics & Soc’y 321, 322 (describing defendants’ lack of resources to conduct adversarial software testing). ↑
- . See Siems et al., supra note 1, at 801 (explaining that procurers of forensic evidentiary tools “strongly prefer to purchase tools that they can be confident will produce admissible results”). Certification also addresses perennial objections that providing access to source code might stifle innovation. See id. at 812–13 (inferring that “the innovation benefits of trade secret privileges for forensic evidentiary technology are likely to be minimal” because “first mover exclusivity will be more than adequate” to preserve incentives for such innovation). ↑
- . The post-certification model mirrors the FDA’s post-market regulatory framework for software as a medical device. See FDA Postmarket Management of Cybersecurity, supra note 225, at 32–36 (offering guidance on how manufacturers should continue to engage in software maintenance and revalidation activities after initial premarket authorization); see also FDA, Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions 4–5 (2025), https://www.fda.gov/regulatory-information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control-plan-artificial-intelligence [https://perma.cc/G5R9-EZ8s] (building in a formal mechanism for regulators to review future anticipated modifications to AI-enabled devices). ↑
- . See Carlos Ignacio Gutierrez, Gary Marchant & Lucille Tournas, Lessons for Artificial Intelligence from Historical Uses of Soft Law Governance, 61 Jurimetrics J. 133, 134 (2020) (defining “soft law programs” as instruments that “define substantive expectations that are not directly enforceable by the government” and noting that soft law’s effectiveness may depend on “indirect enforcement mechanisms such as . . . certifications”); Hagemann et al., supra note 221, at 41 (“Th[e] transition towards soft law is almost certainly an inevitable byproduct of the relentless pace of technological innovation.”). ↑
- . See Gary E. Marchant & Carlos Ignacio Gutierrez, Soft Law 2.0: An Agile and Effective Governance Approach for Artificial Intelligence, 24 Minn. J.L. Sci. & Tech. 375, 385 (2023) (“Soft law often involves a cooperative approach between stakeholders unlike the adversarial approach of government regulation.”). ↑
- . Id. at 377 (citing examples such as “codes of conduct, ethical statements, professional guidelines, statements of principles, certification programs, private standards, public-private partnerships, or voluntary programs”); Hagemann et al., supra note 221, at 44 (“Some soft law actions such as standards or guidelines come from the private sector, while others such as interpretive rules and guidance documents come from regulatory agencies.”). ↑
- . See Marchant & Gutierrez, supra note 262, at 381–84, 393 (pointing to the twin problems of regulating too quickly “before risks and benefits are fully understood” and regulating too slowly due to bureaucratic obstacles and “inertia”). ↑
- . See id. at 385 (explaining that soft law can be adopted and revised “informally and quickly,” enabling “more agile governance,” and that it “expands the scope of governance actors” beyond government to include industry, civil society, and other third parties); Hagemann et al., supra note 221, at 49 (emphasizing soft law’s reliance on collaborative, multistakeholder processes to develop operational “soft criteria” that give soft-law principles practical effect). ↑
- . See Marchant & Gutierrez, supra note 262, at 386–87 (pointing out that soft law mechanisms are “not directly enforceable by government,” often “do not have the reporting and compliance assurance requirements that traditional regulations do,” and therefore are less likely to be trusted by the public). ↑
- . See Sinha, supra note 6, at 82 (observing flaws with supposedly “independent” validation studies); cf. Kate Klonick, The Facebook Oversight Board: Creating an Independent Institution to Adjudicate Online Free Expression, 129 Yale L.J. 2418, 2481 (2020) (articulating three necessary forms of independence: jurisdictional independence, intellectual independence, and financial independence). ↑
- . Barry Friedman, Farhang Heydari, Max Isaacs & Katie Kinsey, Policing Police Tech: A Soft Law Solution, 37 Berkeley Tech. L.J. 701 (2022). ↑
- . Id. at 717–21 (citing problems of political economy as well as gaps in information and expertise). ↑
- . Id. at 727–29. ↑
- . Id. at 731–44. ↑
- . Id. at 744–48. ↑
- . See id. at 736 (“[A] certification entity simply could evaluate a product’s specifications—i.e., does it do what it says ‘on the tin.’”); Matthews et al., supra note 6, at 106 (finding substantial discrepancies in results across three different evidentiary software systems, despite ostensibly applying the same scientific method). ↑
- . There are harder questions concerning whether the forensic science itself is valid, such as whether false positive rates are too high to be acceptable. Cf. Sinha, supra note 6, at 81 (asserting that ShotSpotter’s claimed accuracy rate is “overblown” and “has not been properly scientifically validated”). But disentangling scientific validation from software validation provides a way to challenge whether the software functions even as claimed. ↑
- . See, e.g., Canellas, supra note 136, at 17 (explaining that the IEEE 1012 standard is “highly regarded throughout the technology world” because it has “a proven record for ensuring the reliability of some of the most complex and safety-critical systems”); Nathaniel Adams, Roger Koppl, Dan Krane, William Thompson & Sandy Zabell, Letter to the Editor—Appropriate Standards for Verification and Validation of Probabilistic Genotyping Systems, 63 J. Forensic Scis. 339, 339 (2018) (urging scientists to “pay close attention to” IEEE standards); Paul E. Black, Barbara Guttman & Vadim Okun, Nat’l Inst. of Standards & Tech., NISTIR 8397, Guidelines on Minimum Standards for Developer Verification of Software 4–9 (2021), https://doi.org/10.6028/NIST.IR.8397 [https://perma.cc/6T96-S9C2] (enumerating eleven minimum recommendations for software verification techniques). ↑
- . See Friedman et al., supra note 268, at 736 (noting serious challenges with developing test suites, establishing a common baseline, reevaluating updated software, and more); see also supra note 223 and accompanying text. ↑
- . See Canellas, supra note 136, at 19 (criticizing SWGDAM guidelines for lacking transparency and promoting conflicts of interest); Bryan H. Choi, NIST’s Software Un-Standards, 9 Geo. L. Tech. Rev. 65, 92 (2025) (concluding that NIST’s new frameworks for national cyber-security standards are, in practice, “not uniform standards or measures of anything”). ↑
- . See Derek E. Bambauer & Melanie J. Teplinsky, Shields Up for Software, Lawfare (Dec. 19, 2023), https://www.lawfaremedia.org/article/shields-up-for-software [https://perma.cc/
FT7X-ZPP6] (proposing an alternative approach to the liability standard for secure software that, in part, defines worst practices that automatically impose liability—i.e., an “inverse safe harbor”); Derek E. Bambauer, Cybersecurity for Idiots, 106 Minn. L. Rev. Headnotes 172, 177 (2021) (“Rather than trying to determine when entities get cybersecurity right, regulation should concentrate on when organizations have gone badly wrong.”). ↑ - . See Matthews et al., supra note 6, at 104 (explaining that, although there is a common formula for computing a likelihood ratio in probabilistic genotyping, different software systems that implement this same formula “can produce very different results and little effort has been made to standardize their behavior”). ↑
- . See Highway Safety Programs; Model Specifications for Devices to Measure Breath Alcohol, 58 Fed. Reg. 48705, 48707–08 (Sep. 17, 1993) (setting forth the required model specifications for devices used to measure breath alcohol). ↑
- . See Bryan H. Choi, Software as a Profession, 33 Harv. J.L. & Tech. 557, 570–72 (2020) (contrasting software’s exponential complexity with the ordinary complexity of conventional manufactured goods). ↑
- . See Pei Hsia, David Kung & Chris Sell, Software Requirements and Acceptance Testing, 3 Annals Software Eng’g 291, 291–92 (1997) (explaining that acceptance testing consists of “comparing a software system to its initial requirements and to the current needs of its end-users,” and observing that presently there is no industry-wide standard or systematic method for acceptance testing). ↑
- . See FDA Computer Software Assurance, supra note 224, at 19 (requiring disclosure of the following assurance activities: description of the testing conducted, issues found during testing, conclusion statement declaring acceptability of the software for its intended use, record of who performed the testing and date the test was performed, and record of established review and approval). ↑
- . See Abebe et al., supra note 1, at 1737 (articulating inherent limitations in validation of statistical software); Matthews et al., supra note 6, at 104 (finding that such testing can reveal material differences in results across different software implementations). ↑
- . See Friedman et al., supra note 268, at 748–51 (suggesting transparency around the certification process and solicitation of public input to gain public legitimacy, and mandatory certification systems to ensure vendor buy-in). ↑
- . See Margot E. Kaminski, When the Default Is No Penalty: Negotiating Privacy at the NTIA, 93 Denv. L. Rev. 925, 931–40 (2016) (describing repeated failures in multistakeholder governance despite high initial participation and interest). ↑
- . See Abebe et al., supra note 1, at 1738 (“Because our definition is adversarial, we envision each step will be carried out by the defense and specific to the case at hand.”). ↑
- . See Kroll et al., supra note 160, at 646 (“Without strong evidence of a computer system’s correctness, even the author of that system cannot reliably claim that it will behave according to a desired policy, and no policymaker or overseer should believe such a claim.”). ↑
- . See Comm. on Rules of Prac. & Proc., Proposed Amendments to the Federal Rules of Evidence 109 (2025). A dearth of software literacy makes the indiscriminate admission of software evidence especially pernicious. Cf. Richard K. Sherwin, Visual Literacy for the Legal Profession, 68 J. Legal Educ. 55, 57–61 (2018) (arguing that novel forms of evidence can arouse cognitive and emotional responses that distort legal results in the absence of technological literacy). ↑
- . See FDA General Principles, supra note 223, at 24–25, 28–29 (noting the importance of conducting regression analysis “to provide assurance that a change has not created problems elsewhere in the software product”); cf. FDA, Deciding When to Submit a 510(k), supra note 199, at 5, 7 (acknowledging that requiring documentation for every software change is overly burdensome, but still appropriate where such changes “could significantly affect safety or effectiveness,” and that these decisions “should be confirmed by successful, routine verification and validation activities”). ↑
- . See Abebe et al., supra note 1, at 1736 (noting conflicts of interest, such as “financial or professional interests in ensuring that forensic laboratories use the software,” as well as mismatch of test cases that lead to “over-optimistic estimates of performance on new test cases”). ↑
- . See Bellovin et al., supra note 6, at 17–18 (arguing that “certain errors are more likely to be found by adversarial testing”); Anusha Sinha, James Lucassen, Keltin Grimes, Michael Feffer, Ellie Soto, Hoda Heidari & Nathan VanHoudnos, What Can Generative AI Red-Teaming Learn from Cyber Red-Teaming? 1 (July 2025), https://www.sei.cmu.edu/
documents/6301/What_Can_Generative_AI_Red-Teaming_Learn_from_Cyber_Red-Teaming.pdf [https://perma.cc/T6LY-Z6NJ] (extending red-teaming methodologies from the cybersecurity domain to generative AI systems); Martha McNeil & Thomas Llansó, An Analysis of Adversarial Cyber Testing Practice, 2020 IEEE Sys. Sec. Symp., Aug. 2020, at 1, 2, https://doi.org/10.1109/
SSS47320.2020.9174237 [https://perma.cc/Y9ZD-YHUW] (“Red teaming is a technique that is valued in military operations for its abilities to guard against ‘complacency, groupthink, and mirror-imaging,’ the tendency to assume the adversary behaves as we do.”). ↑ - . See Abebe et al., supra note 1, at 1737–39 (collecting arguments in favor of direct testing by defense counsel and showing how to operationalize such testing). ↑
- . See Margot E. Kaminski & Gianclaudio Malgieri, Impacted Stakeholder Participation in AI and Data Governance, 27 Yale J.L. & Tech. 247, 251–53 (2025) (arguing in favor of public participation by impacted groups); Developments in the Law–Artificial Intelligence, 138 Harv. L. Rev. 1554, 1619 (2025) (describing “co-governance” models that allow impacted stakeholders to directly participate in designing and implementing policy solutions). ↑
- . See FDA General Principles, supra note 223, at 27 (“Testing at the user site is an essential part of software validation. . . . This testing should take place at a user’s site with the actual hardware and software that will be part of the installed system configuration.”); Sinha, supra note 6, at 78 (“Testing conditions matter; systems . . . will not perform identically in all environments they are installed in. Determining whether a system will work in a particular environment requires testing under known conditions reflective of that environment.”). ↑
- . See McNeil & Llansó, supra note 292, at 4 (observing that red-team testers “mainly tested deployed operational systems” and occasionally “high fidelity copies of live systems set up in lab environments, low fidelity test articles, and least frequently, ‘paper and pencil’ models of systems”). ↑
- . See Abebe et al., supra note 1, at 1738 (stating the importance of allowing the tester “broad latitude” to choose “the worst-case distribution” over the universe of test data, and noting that “[s]electing an appropriately rich and relevant [dataset] is one of the key challenges”); McNeil & Llansó, supra note 292, at 4–5 (noting that if testing is “limited to an arbitrary number of days, the findings will be reduced and this may give the system stakeholders a false sense of security”); Contag et al., supra note 219, at 232 (“[B]lack box testing is costly and time consuming and, as these cases show, can be easily circumvented by defeat device software that ‘tests for the tester.’”). ↑
- . See generally Maneka Sinha, The Automated Fourth Amendment, 73 Emory L.J. 589, 598–99 (2024) (calling critical attention to the use of automated policing technologies that have not been subjected to adequate software validation). ↑