\n\n

As we have said before, the Blog is unabashedly pro-science.  There is a difference between good science and bad science, and reliance on the latter to make any important decision—be it in everyday life, litigation, or public policy—is not smart.  We are also unabashedly in favor of strict application of the Rules of Evidence, the 700 series and otherwise.  It may be overly simplistic to say that plaintiff lawyers in our kind of cases tend to want limited application of the Rules of Evidence or even no rules at all and the defense lawyers want the opposite.  It may not be.  At the same time, we appreciate whenever we have the chance to aid a jury’s consideration of competing reliable expert opinions with vigorous cross-examination by both sides.  It certainly can be the case that reliable competing expert opinions can be presented in the same case with a mutually strict application of Rule 702.  However, the existence of vigorous cross-examination and the notion that juries can do a good job weighing competing expert evidence, regardless of its relative reliability, are not substitutes for strict application of Rule 702.  Correcting that misimpression was part of what was emphasized in the 2023 amendment to Rule 702.  We, of course, knew better than to assume that the amendment would correct all misapplications of Rule 702 by federal judges and appellate panels.  We have tracked how it has gone since.  See, e.g., here, here, and here.

One thing that we have seen courts get tripped up on is temporality.  We are not talking about the concern, urged extensively by many plaintiffs in the years after Daubert and rejected by the famous Rosen observation that “law lags science,” that there could be support for a plaintiff’s expert causation opinion created after the fact so it would be unfair to ding an expert for not having support for her “inspired” “scientific guesswork.”  It would not be.  We are also not talking about the temporality criterion of the Bradford Hill Criteria, which is usually the easiest one to analyze and meet.  (We digress somewhat to note that it will often be hard to establish this temporality when it comes to a claim that gestational exposure caused a developmental disorder absent clear evidence on when the developmental disorder actually starts; some teratogens are known to have relatively small exposure risk windows.)  Instead, we are talking about the issue of when the three aspects of Rule’s 702’s reliability requirement—whether it “is based on sufficient facts or data,” “is the product of reliable principles and methods,” and “reflects a reliable application of the principles and methods to the facts of the case”—should be measured.  The right answer is that the opinion has to be reliable when formed, which correlates to the period of time leading to when the expert report is signed (for a retained expert).  Under Fed. R. Civ. P. 26(a)(2)(B), the report must include, inter alia:  “(i) a complete statement of all opinions the witness will express and the basis and reasons for them; [and] (ii) the facts or data considered by the witness in forming them.”  Consistent with the Rule 702 focus on methodology, these requirements point back to the time when the opinion is formed before the report is signed.  Not when the trial court decides a Rule 702 challenge.  Obviously, a bad methodology with an insufficient basis cannot be saved by post hoc work.  (In theory, an initial report could be withdrawn and a better report swapped in, but then the opinions being tested on a Rule 702 motion would be the ones in the superseding report.)  A few years later, when an appellate court looks at whether the opinion met Rule 702, is certainly not the time.

Another thing courts still get wrong is how burden works.  Rule 702’s amendment in 2023 was intended to clarify that the proponent of the expert opinion evidence bears the burden of establishing by a preponderance of evidence the relevance/fit and three reliability criteria.  That is clear from its formatting as well as the advisory committee notes.  We have emphasized this issue before and noted how whether a court refers to “burden” or what the proponent established can be a tell in how it will rule on expert opinion admissibility.

Unfortunately, these issues featured in Rutledge v. Walgreen Co., — F.4th –, 2026 WL 2015284 (2d Cir. July 13, 2026), where a really thorough and impactful Rule 702 decision by an MDL judge was undone by a really misguided appellate decision.  This, of course, is the Second Circuit’s reversal of the key rulings on general causation experts that led to mass summary judgment for the defendants.  We are sure there will be more to say about this decision, which also included a punt on a back-up expert in another case and a short and sloppy affirmance of an early and sloppy preemption denial.  We are going to focus on the reversal of the exclusion of three experts, the central epidemiologist and two experts whose function was to help the epidemiologist pass the “biological plausibility” criterion under Bradford Hill.  We set aside the affirmance of the exclusion of two other plaintiff experts, including whether the appellate court’s reasoning is consistent.

On the timing issue we noted, we give the Rutledge court some credit for disclaiming that political posturing since the MDL court issued its rulings played any role in its decision.  Id. at *1.  We hope that is true.  However, the temporal framing of the inquiry throughout was a bit off.  We might be nitpicking on tense and phrasing, but the timing does matter.  Another plaintiff expert coming up with a plausible biological mechanism does not help the plaintiff establish that the epidemiologist’s opinion was reliable unless the epidemiologist relied on it in forming his opinion.  Similarly, if studies came out or the consensus of professional bodies shifted in favor of causation in the years since an epidemiologist signed his report, that also would not help plaintiffs prove reliability.

The prior observation ties to our second issue above.  Other than initially quoting the “if the proponent demonstrates to the court that it is more likely than not” language from the flush language of Rule 702, Rutledge never returned to the concept of plaintiffs’ burden as proponents of the evidence.  By contrast, the failure of plaintiffs to carry their burden on a number of aspects of reliability and relevance was a key part of the MDL’s decision.  As with a summary judgment decision, burden needs to be a meaningful part of the court’s analysis.

Strangely, other than quoting them up front, the Second Circuit also seemed to ignore Rule 702(a), (b) & (d), focusing solely on qualifications (uncontested) and 702(c).  In other words, the court focused on whether (without the concept of burden) “the testimony is the product of reliable principles and methods,” but not helpfulness/fit or the other reliability provisions: “whether the testimony is based on sufficient facts or data” and whether “the expert’s opinion reflects a reliable application of the principles and methods to the facts of the case.”  Whether the expert purports to apply a reliable methodology—i.e., the Bradford Hill Criteria versus the Cherry-Pick Flim-Flam method—is different than whether she applies it reliably and has sufficient supporting evidence to form an opinion that general causation exists.  We find Rutledge’s reversal of the epidemiologist’s exclusion to be overly focused on his claims that other epidemiologists also use the Bradford Hill Criteria to explore different exposures and different diseases in a similar way to how he says he used them here.  Those claims, if credited, may help to satisfy 702(c) but not 702(b) or 702(d), each of which is an independent requirement.

For instance, one of the main failings of the epidemiologist’s methodology according to the MDL court was his use of a “transdiagnostic evaluation,” essentially lumping together different outcomes to manufacture the appearance of a strong association from epidemiological studies.  This sort of data dredging is roundly decried when done in individual studies.  Per the MDL court, plaintiffs failed to show it was reliable to do it for the very different conditions of autism spectrum disorder and attention-deficit/hyperactivity disorder (and other things he lumped in when it suited him).  Without the lumping, he could not offer an opinion that prenatal acetaminophen use causes either condition.  In Rutledge, the court did not require such a specific showing tied to the facts of the case, only that the epidemiologist was “using a methodology that epidemiologists routinely use.”  Id. at *11.  That is an incomplete analysis of the issues.  It is also hard to reconcile with the court’s later disclaimer that:

We do not mean to suggest that all Bradford Hill analyses that simultaneously examine multiple outcomes, or that use symptomatic as well as diagnostic endpoints, are reliable and admissible under Rule 702. We can imagine, for example, conditions which are sufficiently distinct so that a single, joint, analysis could not be performed in a reliable manner.

Id. at *12.  You can only make that distinction if you look at whether the proponent carried its burden to show the lumping was reliable for the particular issues in the case as required by Rule 702(d), which Rutledge did not do but the MDL court did.

More generally, the appellate court found the MDL court overstepped its gatekeeping function by “substitut[ing] its own definitions of certain Bradford Hill factors for those of other epidemiologists, and (ii) penaliz[ing the expert] for drawing plausible conclusions well within “the range where experts might reasonably differ.”  Id. (citations omitted).  The examples of overstepping, however, sound an awful lot like not just taking the plaintiff’s expert’s word for it that his opinion was reliable and actually testing whether the record showed the plaintiffs carried their burden to establish each element of Rule 702.  We could go on with the details, but we will end our discussion of the epi’s reinstatement by noting that Rutledge’s conclusion should only make it harder for district courts:

While the district court’s reasoning was considered and extensive, its analysis frequently overstepped its gatekeeping function by substituting its own judgments about the requirements of causality and the persuasiveness of various studies for those of scientists operating in the field.

Id. at *19.

On the other two experts whose opinions were held to have been wrongly excluded, they were, as we noted, offered solely on biological plausibility.  If the epidemiologist’s opinion was excluded, theirs would not be helpful to the jury as mechanism hypotheses alone could not possibly carry the plaintiffs’ causation burden.  Relatedly, the district court had noted that one of the mechanism experts “plays a critical role for the plaintiffs [who] rely on [him] to give his imprimatur to the transdiagnostic Bradford Hill analysis of causation applied by their other experts. His reports do not do so.”  It then explored in detail how his opinion did not provide what the epidemiologist needed to establish biological plausibility.  Rutledge deemed this an “unduly narrow definition of relevance.”  Id.  Had the appellate court considered the developed law on the “fit” requirement under Daubert and Rule 702, which it never discussed, it might have found the MDL court’s approach appropriate gatekeeping.

As to the other mechanism expert, Rutledge questioned the MDL court’s finding that the expert’s reliance on animal studies showing the opposite of his hypothesized mechanism unreliable.  Eschewing its own guidance not to “substitut[e] its own judgments about the requirements of causality and the persuasiveness of various studies for those of scientists operating in the field,” Rutledge accepted the expert’s own “counterintuitive” explanation that “[a]ny behavioral evidence of neurodevelopmental change attributable to acetaminophen supports the conclusion that acetaminophen disrupts neurodevelopment.”  Id. at *20.  Even ignoring that “neurodevelopment” is a really broad umbrella term, a study showing a decreased risk of a certain negative outcome would not be reliable support for an opinion that there is an increased risk simply because it showed some change.  That is not how this all works.

The resurrection of an MDL with such a shaky causation footing is good for plaintiff lawyers, but maybe for nobody else.  And maybe not for long.  The MDL’s denial of preemption on warnings claims at the pleadings stage was flawed, in part because the CBE regulation was not available to change the label of these monograph OTC drugs back when they were being used in the individual cases.  On summary judgment, plaintiffs will have to show that the branded manufacturers, generic manufacturers, pharmacies, and retailers could each have unilaterally changed the OTC drug labels to add specific warnings on specific risks during the relevant time of use during pregnancy based on newly acquired information.  Here, temporality favors the defendants.  The weak evidence of general causation should get weaker as you go back in time.  And it should be clearer that then-existing regulations and regulatory environment would not have permitted unilateral labeling changes by most, if not all, of these defendants.  We will be watching to see how it plays out.