Transcription of Hill’s Criteria for Causality
1 Hill s Criteria forCausalityDespite philosophic criticisms of inductiveinference,inductively oriented causal Criteria have commonlybeen used to make such inferences. If a set of ne-cessary and sufficient causal Criteria could be usedto distinguish causal from noncausalassociationsinobservational studies, the job of the scientistwould be eased considerably. With such Criteria ,all the concerns about the logic or lack thereof incausal inference could be forgotten: it would only benecessary to consult the checklist of Criteria to see ifa relation were causal. We know from philosophythat a set of sufficient Criteria does not exist [3,6]. Nevertheless, lists of causal Criteria have becomepopular, possibly because they seem to provide a roadmap through complicated commonly used set of Criteria was proposedbySir Austin Bradford Hill[1]; it was an expan-sion of a set of Criteria offered previously in thelandmark Surgeon General s report on Smoking andHealth [11], which in turn were anticipated by theinductive canons of John Stuart Mill [5] and therules of causal inference given by Hume [3].
2 Hillsuggested that the following aspects of an associa-tion be considered in attempting to distinguish causalfrom noncausal associations: strength, consistency,specificity, temporality, biologic gradient, plausibil-ity, coherence, experimental evidence, and popular view that these Criteria should be usedfor causal inference makes it necessary to examinethem in detail:StrengthHill s argument is essentially that strong associationsare more likely to be causal than weak associationsbecause, if they could be explained by some otherfactor, the effect of that factor would have to beeven stronger than the observed association and there-fore would have become evident (seeCornfield sInequality). Weak associations, on the other hand,are more easily explained by extent this is a reasonable argument, but, asHill himself acknowledged, the fact that an asso-ciation is weak does not rule out a causal con-nection. A commonly cited counterexample is therelation between cigarette smoking and cardiovascu-lar of strong but noncausal associ-ations are also not hard to find; any study withstrongconfoundingillustrates the phenomenon.
3 Forexample, consider the strong but noncausal relationbetween Down syndrome and birth rank, which isconfounded by the relation between Down syndromeand maternal age. Of course, once the confoundingfactor is identified, the association is diminished byadjustment for the factor. These examples remindus that a strong association is neither necessary norsufficient for Causality , nor is weakness necessary norsufficient for absence of Causality . In addition to thesecounterexamples, we have to remember that neitherrelative risknor any other measure of association isa biologically consistent feature of an association; asdescribed by many authors [4, 7], it is a characteristicof a study population that depends on the relativeprevalenceof other causes. A strong associationserves only to rule out hypotheses that the associationis entirely due to one weak unmeasuredconfounderor other source of modest refers to the repeated observation of anassociation in different populations under differentcircumstances.
4 Lack of consistency, however, doesnot rule out a causal association, because some effectsare produced by their causes only under unusual cir-cumstances. More precisely, the effect of a causalagent cannot occur unless the complementary com-ponent causes act, or have already acted, to completea sufficient cause. These conditions will not alwaysbe met. Thus, transfusions can cause HIV infectionbut they do not always do so: the virus must also bepresent. Tampon use can cause toxic shock syndrome,but only when other conditions are met, such as pres-ence of certain bacteria. Consistency is apparent onlyafter all the relevant details of a causal mechanism areunderstood, which is to say very seldom. Even stud-ies of exactly the same phenomena can be expectedto yield different results simply because they differin their methods andrandom errors. Consistencyserves only to rule out hypotheses that the associ-ation is attributable to some factor that varies of Biostatistics,Online 2005 John Wiley & Sons, article is 2005 John Wiley & Sons, article was published in theEncyclopedia of Biostatisticsin 2005 by John Wiley & Sons, : s Criteria for CausalitySpecificityThe criterion of specificity requires that a cause leadsto a single effect, not multiple effects.
5 This argumenthas often been advanced to refute causal interpre-tations of exposures that appear to relate to myr-iad effects, especially by those seeking to exoneratesmoking as a cause of lung cancer. The criterion iswholly invalid, however. Causes of a given effectcannot be expected to lack other effects on anylogical grounds. In fact, everyday experience teachesus repeatedly that single events or conditions mayhave many effects. Smoking is an excellent example:it leads to many effects in the smoker. The existenceof one effect does not detract from the possibility thatanother effect exists. Thus, specificity does not confergreater validity to any causal inference regarding theexposure effect. Hill s discussion of this criterionfor inference is replete with reservations, and manyauthors regard this criterion as useless and misleading[8, 9].TemporalityTemporality refers to the necessity that the cause pre-cede the effect in time.
6 This criterion is unarguable,insofar as any claimed observation of causation mustinvolve the putative cause C preceding the putativeeffect D. It doesnot, however, follow that a reversetime order is evidence against the hypothesis that Ccan cause D. Rather, observations in which C fol-lowed D merely shows that C could not have causedD in these instances; they provide no evidence for oragainst the hypothesis that C can cause D in thoseinstances in which it precedes GradientBiologic gradient refers to the presence of a mono-tone (unidirectional)dose responsecurve. We oftenexpect such a monotonic relation to exist. For exam-ple, more smoking means more carcinogen exposureand more tissue damage, hence more an expectation is not always present, somewhat controversial topic of alcohol con-sumption and mortality is an example. Death ratesare higher among nondrinkers than among moderatedrinkers, but ascend to the highest levels for heavydrinkers.
7 Because modest alcohol consumption canhave beneficial effects on serum lipid profiles, sucha J-shaped dose response curve is at least biologi-cally , associations that do show a monotonictrend in disease frequency with increasing levels ofexposure are not necessarily causal; confounding canresult in a monotonic relation between a noncausalrisk factor and disease if the confounding factoritself demonstrates a biologic gradient in its relationwith disease. The noncausal relation between birthrank and Down syndrome mentioned above shows abiologic gradient that merely reflects the progressiverelation between maternal age and the occurrence ofDown the existence of a monotonic association isneither necessary nor sufficient for a causal nonmonotonic relation only conflicts with thosecausal hypotheses specific enough to predict a mono-tonic dose response refers to the biologic plausibility of thehypothesis, an important concern but one that is farfrom objective or absolute.
8 Sartwell [9], emphasizingthis point, cited the remarks of Cheever, in 1861, whowas commenting on the etiology of typhus before itsmode of transmission (via body lice) was known:It could be no more ridiculous for the stranger whopassed the night in the steerage of an emigrant shipto ascribe the typhus, which he there contracted, tothe vermin with which bodies of the sick might beinfested. An adequate cause, one reasonable in itself,must correct the coincidences of simple was to Cheever an implausible explanationturned out to be the correct explanation, since it wasindeed the vermin that caused the typhus is the problem with plausibility: it is too oftennot based on logic or data, but only on prior is not to say that biological knowledge shouldbe discounted when evaluating a new hypothesis,but only to point out the difficulty in applying to inference attempts todeal with this problem by requiring that one quan-tify, on a probability (0 to 1) scale, the certainty thatone has in prior beliefs, as well as in new quantification displays the dogmatism or open-mindedness of the analyst in a public fashion, withcertainty values near 1 or 0 betraying a strong com-mitment of the analyst for or against a hypothesis.
9 ItEncyclopedia of Biostatistics,Online 2005 John Wiley & Sons, article is 2005 John Wiley & Sons, article was published in theEncyclopedia of Biostatisticsin 2005 by John Wiley & Sons, : s Criteria for Causality3can also provide a means of testing those quantifiedbeliefs against new evidence [2]. Nevertheless, theBayesian approach cannot transform plausibility intoan objective causal from the Surgeon General s report on Smok-ing and Health [11], the termcoherenceimplies thata cause and effect interpretation for an associationdoes not conflict with what is known of the natu-ral history and biology of the disease. The examplesHill gave for coherence, such as the histopathologiceffect of smoking on bronchial epithelium (in refer-ence to the association between smoking and lungcancer) or the difference in lung cancer incidenceby sex, could reasonably be considered examplesof plausibility as well as coherence; the distinctionappears to be a fine one.
10 Hill emphasized that theabsence of coherent information, as distinguished,apparently, from the presence of conflicting infor-mation, should not be taken as evidence against anassociation being considered causal. On the otherhand, presence of conflicting information may indeedundermine a hypothesis, but one must always remem-ber that the conflicting information may be mistakenor misinterpreted [12].Experimental EvidenceIt is not clear what Hill meant by experimental evi-dence. It might have referred to evidence from lab-oratory experiments on animals, or to evidence fromhuman experiments. Evidence from human experi-ments, however, is seldom available for most epi-demiologic research questions, and animal evidencerelates to different species and usually to levelsof exposure very different from those that humansexperience. From Hill s examples, it seems that whathe had in mind for experimental evidence was theresult of removal of some harmful exposure in anintervention or prevention program, rather than theresults of laboratory experiments [10].