Showing posts with label evidence. Show all posts
Showing posts with label evidence. Show all posts

Sunday, 10 May 2009

Open-mindedness


Following from my previous post on supernatural explanations for unexplained events, I was delighted to find this entertaining video on the topic of open-mindedness. Seems it's been making the rounds, but if you haven't seen it, it's definitely worth checking out. (For a précis, see the blog mental indigestion.)

Not only is it visually clever, the content is quite good. What's more, the author, who goes by the intriguing moniker Qualia Soup, has produced a bunch of other good videos.

Sunday, 11 January 2009

Absence of evidence ...

In a valid deductive argument, the conclusions follow necessarily from the premises. This is a proof in the mathematical sense of the word. Provided we know that the premises are true, we can establish with complete certainty that the conclusions are true. For example, identifying a single unicorn would establish without a doubt that unicorns exist.

Unfortunately much of the time this type of certainty isn't possible. Consider another example from the realm of mythology: weapons of mass destruction (WMD) in Iraq. Here's what Donald Rumsfeld had to say on the subject in 2002 (the boldface is my addition):
There's another way to phrase that and that is that the absence of evidence is not evidence of absence. ... Simply because you do not have evidence that something exists does not mean that you have evidence that it doesn't exist.
But surely hunting high and low for WMD month after month and not finding any (absence of evidence) supports the inference that there aren't any there (evidence of absence). Indeed, it turns out that the popular maxim cited by Rumsfeld is simply incorrect.

But didn't he have a point? Absolutely: failing to prove that something exists does not prove that it does not exist. Or, in the words of the English writer William Cowper (1731-1800):
Absence of proof is not proof of absence
Compare this with the version invoked by Rumsfeld:
Absence of evidence is not evidence of absence
The originator of this maxim seems to be the cosmologist Martin Rees, although it has been attributed to many others, including Carl Sagan. By substituting the word evidence for proof it makes a much stronger (and invalid) claim. Evidence, after all, is often uncertain. If I look outside and see that the ground is wet, that is evidence that it has been raining. But perhaps my neighbour was watering her flowers. Seeing someone walk by with an umbrella folded under their arm might strengthen my evidence for the rain hypothesis, but perhaps they are anticipating rain later on. In general, evidence can support an inference, but it won't necessarily prove it. And that's where Rees's formulation of the maxim falls down.

Black and white thinking about evidence

When evidence is construed as being certainty, we get into all kinds of trouble. This is how Rumsfeld turned a simple truism (no WMDs have been found, but they might still be) into a puzzle of obfuscation (absence of evidence is not evidence of absence).

But Rumsfeld is not the only one. As I noted recently, the term "no evidence" is commonly used to describe situations where an effect is not found to be statistically significant. Now statisticians are wary of people concluding that a lack of statistical significance implies that there is "no effect". (It might be, for example, that the sample size was inadequate.) Hence, it is not at all uncommon for statisticians to declare that absence of evidence is not evidence of absence! As Kim Øyhus has pointed out, even the American Statistical Association buys into it, as the t-shirt they sell attests.

Of course statisticians know well that uncertainty isn't easy to think about or communicate to others. So why have we fallen into this trap?

Well, part of the reason may be philosophical. Statistical reasoning is inescapably inductive—it does not guarantee certainty. Philosophers have been worrying about what is called the problem of induction for a very long time. David Hume (1711-1776) challenged the logical foundations of induction, and ever since, philosophers have sought a way around the problem. The reigning "solution" is known as the hypothetico-deductive method, developed by philosopher of science Karl Popper (1902-1994). Popper argued that induction in science could be avoided by proposing a hypothesis and then seeking evidence that would either prove the hypothesis wrong ("falsify" it) or fail to do so. This is very similar to the frequentist statistical hypothesis testing framework that developed from the work of Fisher, Neyman, and Pearson. Unfortunately, it lends itself to black-and-white thinking. A hypothesis is either proven wrong or it isn't. There's no grey zone.

Popper's formulation, in particular, buries the uncertainty completely, construing the reasoning as entirely deductive. Suppose, for example, that a new biochemical theory predicts that a certain drug will shorten the duration of an illness, whereas the older theory does not. Now duration of illness depends on numerous factors, including differences in patients' immune systems, and we expect to see variation above and beyond any differences due to the drug. A clinical trial may demonstrate that the average duration of illness for patients who are randomly assigned the drug is shorter than that for patients who receive placebo, and that this difference is statistically significant at the 0.05 level. Has the older theory been proven incorrect? Not with absolute certainty. The evidence against it may be strong but it is possible that this is a "type-I" error—rejecting the null hypothesis even though it is true. Indeed, because of the way the statistical test has been designed, when the null hypothesis is true we expect to see such errors 5% of the time. The companion to the type-I error is the type-II error—failing to reject the null hypothesis even though it is false.

Pretending that type-I and -II errors don't exist is wishful thinking. Just as diagnostic tests produce false positives and false negatives, statistical hypothesis tests can give the wrong answer. The point is that we can study and control the error rates and make inferences while acknowledging their limitations.

It suited Donald Rumsfeld's purposes to be fuzzy about the distinction between evidence and proof. It doesn't suit ours.

Update 03-Jun-2009: I had originally attributed the maxim "Absence of evidence is not evidence of absence" to Carl Sagan in his 1995 book The Demon-Haunted World." Apparently however, the originator was cosmologist Martin Rees. There is reference to it in the proceedings of a 1972 symposium titled Life Beyond Earth & The Mind of Man [pdf], jointly sponsored by Boston University and NASA. In his introductory remarks, the chair, Richard Berendzen stated:
A generation ago almost all scientists would have argued, often "ex cathedra," that there probably is no other life in the universe beside what we know here on Earth. But as Martin Rees, the cosmologist, has succinctly put it, "absence of evidence is not evidence of absence." Beyond that, in the last decade or so the evidence, albeit circumstantial, has become large indeed, so large, in fact, that today many scientists, probably the majority, are convinced that extraterrestrial life surely must exist and possibly in enormous abdundance.
(The boldface is mine.) Note that Carl Sagan was one of the panelists at the symposium.

Wednesday, 24 December 2008

Something fishy about "no evidence"

Searching Google for "no evidence" yields "about 21,300,000" hits. It seems we're keen to deny that there is any empirical support for countless different claims. For example, scientificblogging.com reports that there is "No Evidence For Fish Oil Benefit In Arrhythmias" based on a systematic review just published in BMJ. (Full disclosure: I have previously participated in research on omega-3 fatty acids, however I have no related financial interests.) What does the review itself say?
This is the first systematic review attempting to evaluate whether the protective mechanism of fish oil supplementation is related to a reduction of arrhythmic episodes determined either by a reduction in implantable cardiac defibrillator interventions or a reduction in sudden cardiac death. We found a neutral effect on these two outcomes. The confidence intervals for these outcomes were wide and a beneficial effect up to a 45-48% relative risk reduction cannot be excluded.
To better appreciate this, here's their Figure 2:
Note that of the three studies that looked at the proportion of implanted defibrillators that were triggered, one showed a statistically significant effect in favour of fish oil, and the other two did not show statistically significant effects (one favoured placebo and the other favoured fish oil). Six studies looked at sudden cardiac death in patients taking fish oil compared to those taking placebo. Only one was statistically significant, and it favoured fish oil. Of the five studies that did not show statistically significant effects, two favoured fish oil, and three favoured placebo.

The diamond shapes in the figure show the pooled estimates with their 95% confidence intervals: in each case the diamond overlaps an odds ratio of 1, indicating that the overall effect is not statistically significant. And when there's a non-statistically significant effect, it is common practice to say there is "no evidence". But that can be very misleading! After all, two of the individual studies did show a significant benefit of fish oil. So what's going on? Well, for starters, there's some indication of heterogeneity between the studies (particularly in the case of the defibrillator studies). But it also seems that more large studies are needed: a good deal of the variation in results between the studies may simply be due to the play of chance. Quite substantial benefits of fish oil are entirely plausible: relative risk reductions of as much as 45-48%!

What is "no evidence"?

Consider the figure below:

At the bottom there is a gray axis line with tick marks and a vertical gray line indicating the "null" value (where there is no preference one way or the other). The blue line with arrows at each end represents an infinitely wide confidence interval. This is the most straightforward representation of "no evidence": there is simply no empirical information to indicate what the true effect might be.

But suppose we have a very small sample, that is, one that provides almost no empirical evidence. The figure might become:

The only difference is the blue dot on the confidence interval just a bit to the right of the null line. It represents the point estimate based on a very small amount of empirical information. Of course it could equally well have been on the left hand side (or perhaps directly on the null line). Regardless, the confidence interval is still very wide, so very little can be said about the true effect. With such a wide confidence interval, the location of the point estimate is almost irrelevant.

Finally, suppose that a reasonably large sample is available:

I have left the point estimate at the same place. The confidence interval no longer has arrows on either end and is relatively narrow. However it still overlaps the null line. That means the estimate is not statistically significant. Sometimes this sort of situation is described as showing "no evidence of an effect". But as I noted above, that's quite misleading language. In fact, what this situation shows is indeed evidence—evidence that any effect likely has a magnitude of no more than two tick marks (whatever they represent) on the right hand side of the null line or a magnitude of no more than about a half a tick mark on the left hand side of the null line.

But here's the tricky part: what do those ticks represent? Suppose the axis represents annual cost savings that might result from implementing a certain type of federal government program. If each tick mark represents $1000, then we have estimated that the program will cost at most $500 a year and save at most $2000 a year. In other words, the program has been shown to be effectively revenue neutral: the evidence suggests that the cost/cost-savings of the program will not be important. On the other hand, if each tick mark represents one million dollars, most of us would feel that the jury's just not in yet. A possible cost of $500 a year is a drop in the bucket, but $500,000 a year is something else entirely.

In a way, money is the easiest measure to evaluate like this. Things like safety are much harder. For example, if the evidence suggests that a certain chemical may increase the rate of certain types of cancer, but the findings are not statistically significant, what can we conclude? Can the manufacturer claim that there's "no evidence" the chemical is harmful? Can health activists claim that there's "no evidence" the chemical is safe?

I would argue that the term "no evidence" is inappropriate in either case. The underlying questions remain: what is required in order to conclude that a chemical is harmful or that it is safe? Ultimately there's no getting around the issue of how large a difference (in, for example, cancer rates) has to be in order to be considered important. And that's a rather uncomfortable question.