Medicine Is Being Treated with Snake-Oil Statistics
Suggested Running Head: Snake Oil Statistics
Statistics has helped medicine move away from an eminence-based framework, where subject-matter experts decided what worked and what didn’t, towards an evidence-based one. Among the techniques deployed, null-hypothesis statistical testing (NHST) is the most common, where the word “null” is invariably taken by users to mean that the only hypotheses tested are those of “no association” or “no effect” (sometimes labeled “nil hypotheses”). Despite objections to them extending throughout the past century1, these tests are supposed to serve as a safeguard against researchers fooling themselves and others, as well as serving as a central component of experimental design and analysis.
Unfortunately, these tests are easily subverted into playing the reverse role of providing a badge of approval, allowing researchers to make stronger claims than warranted by valid statistical analyses. Confidence intervals have been extensively promoted to address this problem, but they too have been subverted by being treated as if they are only testing null hypotheses. The consequences for evidence-based medicine have been dire.
There are three commonly stated principles (shown in Figure 1) of evidence-based medicine2:
- reliance on statistically significant results (and thus NHST) from randomized controlled trials,
- balancing of costs, benefits, and uncertainties in decision making, and
- combining clinical expertise with external evidence to tailor treatments for individuals.
Figure 1
Unfortunately, the use of NHST can get in the way of the movement toward an evidence-based framework. This may sound paradoxical, given that one of the foundations of evidence-based medicine is hypothesis testing based on randomized controlled trials, deemed by many to be the most reliable forms of evidence2. One problem is that reliance on statistical significance (principle 1) may conflict with the other principles: In conflict with balancing costs and benefits under uncertainty (principle 2), statistical significance or non-significance is typically used to replace uncertainty with certainty3, 4 — indeed, researchers are encouraged to do this and it is even forced by some journals5.
Reliance on statistical significance can interact badly with expertise and background evidence (principle 3), often leading to incoherent attempts at resolution. For example, researchers may do separate analyses for men and women, often with no biological basis for expecting other than small differences in effects (if any). In doing so they often find that one group (usually the larger one, which is usually men) show a “significant” effect while the other does not. They then misreport this difference in “significance” as if it represented a significant difference in effect between groups when in fact it only reflects a difference in the group sizes, plus random differences in group-specific P-values. Such misinterpretations fool researchers into thinking that the data support targeting treatments at the group showing “significance” and mislead clinicians into tailoring treatment plans for individuals based on nothing more than random variation.
The unfortunate reality is that estimating effects for individuals or population subsets requires far more sophistication than basic testing procedures6. It can take several times more patients to estimate variation in effects (“interactions”) than average effects7. Given that few studies are large enough to estimate main effects of interest, it will typically be impossible to obtain reliable estimates of effect variation even if the study is otherwise flawless. That problem should be dealt with under principle (2) by recognizing that subgroups will have very imprecise estimates, leading to enormous uncertainties about who if anyone should be targeted for or excluded from treatment. Unfortunately, the prevalent misunderstandings of classical statistics have trained most researchers, editors, and reviewers to demand statistical significance and certainty as a prerequisite for publication8 and decision making. Applying those demands within subgroups all but guarantees that distorted impressions of patient-specific effects will follow.
Through neglect of basic education, the statistics profession is partly responsible for these issues, but there is nothing new about medicine’s desire for certainty. That desire is part of human nature and thus existed long before the adoption of statistical methods. The resulting demands for certainty in reported results have provided ample opportunities for overconfident researchers to rise to prominence. Physicians have long capitalized on several opportunities with their sophisticated knowledge of physiology and anatomy, and used their medical authority to argue what treatments worked, shape public policy, and design clinical guidelines, maintaining a form of medicine characterized as “eminence based.”
The landscape changed when advances in quantitative methods eventually reached medicine9, leading to a newfound demand for scientific rigor and making many individuals wary of the claims of subject-matter experts10. These demands may be some of the biggest contributors to medicine’s adoption of statistical methods and its movement towards an evidence-based framework. Unfortunately, in its pursuit for objectivity, medicine was sold snake oil statistics and as a result, its desire for objectivity and rigor backfired.
To see the reach of these problems, one simply needs to look wherever a new drug has been considered effective only if it has been shown statistically significantly better than a control in one or more randomized clinical trials (a standard that applies to FDA new-drug approvals, though not to medical devices, which can be cleared through other pathways). The largest enforcers of these methods have been regulatory agencies that wish to minimize treatments that do not work (false positives) and minimize adverse events from treatments. One should ask: what specifically convinced these agencies that statistical significance and in particular NHST was the best analysis criterion to achieve these goals?
In a nonexistent ideal world, the regulatory agencies looked at several statistical methods, tested each of them in many settings to see how well they identified true benefits and harms while avoiding false conclusions, and from that determined NHST performed best out of all options – with periodic updates as new methods appeared. Of course, history shows something else entirely, even indicating that adoption of NHST was mainly a result of political desperation11, 12. To understand the harsh reality, we must go back to the mid-20th century, when pharmaceutical companies submitted new drug applications to the FDA that often lacked any protocols and statistical analysis plans, making the entire drug approval process chaotic. By the end of the 1960s, the FDA had become desperate to standardize the drug-approval process and make it scientifically rigorous, especially with impending pressures from the Drug Amendments of 1962, which demanded rigorous evidence for drug approval.
At the same time, applied researchers were looking for rigorous ways to summarize experiments with numerical quantities. Their search struck gold with the works of the prominent statisticians, Ronald A. Fisher13–15, Jerzy Neyman16, and Egon Pearson17, who created and popularized powerful statistical tools for researchers that lacked statistical training. Eventually, researchers across many disciplines adopted NHST and the now-infamous 0.05 cutoff. Soon, the FDA followed and incorporated these methods into its regulatory process, without any formal debate about the methods’ utility or evidence12.
Figure 3: Austin Bradford Hill. The statistician and epidemiologist who conducted the first randomized controlled trial in medicine.
It thus appears that embracement of the NHST paradigm arose from expediently following the ascendant trend in research rather than a critical evaluation of various options emerging in the same period. Since its adoption into the regulatory process, this framework has been rigorously enforced by the FDA as the gold standard in the approval process.
We find it ironic that the gatekeepers of evidence-based medicine — regulatory agencies, journal editors and reviewers, and medical researchers – continue to insist on enforcement of a statistical framework that was never critically examined for its utility in medicine12. This failure makes it understandable why it can be credibly argued that conventional statistical methods have done more harm than good in medical science. These methods, including null-hypothesis tests and so-called “confidence” intervals are a set of decision-making tools designed for tightly controlled randomized experiments in which the treatment effects (if any) can be distinguished from all other causal effects, and can be distinguished from random error (noise) by increasing the study size in an affordable manner. Examples of such scenarios arise in agricultural research where the experimental units are plants or plots, and industrial quality control, where the experimental unit is a part or product. The tools were originally developed for these environments, in which (compared to clinical studies) experimenters have almost godlike control over the selection of units and their subsequent experiences, aided by having a rather short follow-up period and low cost per experimental unit. And then, the decisions to be made are relatively simple and easily monitored, e.g., change a fertilizer formulation or manufacturing tolerance17. Experimental psychology is similar in terms of control, cost, and even lower cost-benefit consequences.
In clinical environments, however, the cost per experimental unit (the patient) is far higher, creating severe limits on study size and thus noise reduction, and the costs and benefits of decisions can be enormous. At the same time, there is far less control of extraneous selection and confounding effects: Physicians and their patients can and do selectively refuse to participate or cease to adhere to assigned treatment protocols, and may drop out for unknown reasons. Meanwhile, direct physical control of the patient environment is extremely limited or nonexistent, especially when follow-up extends beyond hospital stay. And then, amplifying these limits, the required decisions are often complex and of highly uncertain consequence, yet may be pivotal to the experimental results and clinical decisions, e.g., when to withdraw treatment from patients apparently experiencing side effects or when to switch treatments for nonresponsive patients.
These vast differences haven’t stopped medical researchers from using conventional methods as if they were operating in a tightly controlled environment, treating ambiguous trial results as if they supported decisive conclusions despite obvious uncertainties and potentially devastating consequences. The usual depictions of this problem involves researchers trying to “game” statistical methods so that they show “significant” effects18, 19; while such “significance questing” is a real problem, warnings about it usually ignore or dismiss opposite behavior in which showing “nonsignificance” will facilitate publication in prestigious medical journals, especially where so-called “replication failure” has become a hot topic3. Adding to these distortions is the publication bias that results when researchers or editors deem results unworthy of submission or acceptance because they are just not interesting because they fail to report a “discovery” or only replicate what is “known.” This desire for novelty or publicity, along with demands for certainty, has distorted the medical literature with dubious, inflated effects20 and misleading claims of replication failure based on fundamentally ambiguous results21, 22.
Responses to the problems
In response to the ongoing abuse of statistical testing, the American Statistical Association released a statement in 2016 cautioning against the fixating on P-values and statistical significance23. Three years later, the organization published an issue titled “A World Beyond P < 0.05” with 43 commentaries from statistical experts on how to improve statistical inference, with or without P-values24. The issue was accompanied by a highly discussed commentary in Nature titled “Scientists Rise up Against Statistical Significance” that discouraged mindless automation and dichotomization of statistical results21, and was supported by signatures of some 800 applied researchers.
While these calls are impressive and a growing number of journals — including the New England Journal of Medicine and, in modified form, JAMA — have updated their statistical-reporting policies, most regulatory agencies and the bulk of the medical literature continue to fixate on statistical significance5, 25. Correspondingly, we can expect medical researchers to keep cutting corners to achieve or remove statistical significance18, 19 or misinterpret ambiguous results as if definitive, a practice that is often labeled as “spin”26.
We can interpret the persistence of null hypothesis significance testing (NHST) in two ways. One story is that the value of NHST is recognized by real-world decision makers, despite the carping of ivory-tower critics such as the authors of the present article. The other story is that the counterproductive nature of NHST has been denied by the medical establishment5, 25, and that better alternatives are available12. There are legitimate practical concerns behind both these perspectives. On one hand, active researchers and regulators have legitimate concerns about working on or approving treatments that do not work. On the other hand, many well-publicized examples have made it clear that NHST can easily lead to overconfidence and erroneous inferences, as seen in discussions that treat statistically significance as demonstrating presence of effects and nonsignificance as demonstrating absence of effects.
Methodologists propose new methods to address NHST in their field
Fear of false positives has dominated most discussions of scientific rigor18, 27, 28, often relegating false negatives to more technical discussions of power. Although false positives can be costly, so can missing clinically meaningful effects21. Take postmarketing surveillance for pharmaceuticals; in such situations, serious adverse events from drugs are often underreported and have in some cases been actively suppressed29, 30, as seen in Cochrane reviews of unpublished trial data, leading to a scarcity of data and resulting in studies that will never be able to show a “significant” effect because of the lack of resources and time. If a study comparing adverse events from those taking an approved drug and some control group is unable to show statistical significance, even when the estimated effect is important and plausible, the results will typically be confused and used for evidence of absence21, 31, and prescribing is unlikely to be curtailed, leading to continued harm.
The confusion of statistical nonsignificance with evidence of absence will remain a problem in fields where effects are often small and yet studies powered to detect them are infeasible. Most studies in surgical science rarely have more than a dozen participants. Again, this makes it incredibly difficult to achieve statistical significance, even when clinically important effects are likely. In such fields, adopting the NHST framework for analysis all but guarantees failure to detect those effects. As a result, researchers have looked to other methods — including Bayes factors, second-generation P-values, and post-hoc power calculations — to circumvent the bad hand that was dealt to them32–34. Unfortunately, such transitions are not always helpful: while reasonably defensible alternatives such as interval testing and posterior intervals do exist, several of the most-publicized proposals introduce their own errors, and in fact, some may actually result in more errors than before.
For example, a group of surgeons recently published several statistical recommendations on what to do if a surgical study result was nonsignificant35–37, providing hope to researchers who have had their inquiries halted by that result. But the recommendations are not only statistically invalid and clinically misleading, and thus have been discouraged by statisticians for nearly two decades38, 39. In a similar tale, two sports scientists published a statistical method40 which was supposed to improve the classification of individual treatment responses by reducing the influence of random error in small studies. Unfortunately, reducing the influence of random error in a valid manner requires improvement of study design, including increase in study size, so unsurprisingly the proposed method has unacceptably poor statistical properties41.
The efforts behind new methods and recommendations are a response to an arbitrary dichotomy that has been imposed upon researchers by the medical and scientific establishment: decide whether a result is “positive” or “negative”, with no allowance for what is usually the most reasonable interpretation – ambiguous or indecisive. Unfortunately, conventional statistics has aggravated this problem by offering methods like NHST which produce only dichotomous answers. These methods which swept research sciences in the mid-20th century and are now firmly rooted in tradition, as if that tradition is the best we can do. But it isn’t.
The problem with upending this tradition however is that there is no consensus about what to replace it with, a problem that only grows as alternatives continue to proliferate. A idealistic (and we think naïve) view would simply allow authors to choose alternatives of their liking. The problem with this anarchic approach is that many of the alternatives are themselves misleading and even defective or at best inferior to other methods in demonstrable ways. Yet seeing these problems requires not only sufficient technical expertise to evaluate methods, but also a willingness to see them – a willingness that should not be assumed for originators and adopters of the method. Unfortunately, mere publication of a method in a peer-reviewed journal is a faulty indicator of method validity. That is especially so if publication is not in a statistics journal, for in that case it means only that peer-review was by referees who may have had little of the special technical expertise needed for thorough evaluation.
Solutions
There are no simple solutions or universal guidelines that medical researchers can always use to improve scientific rigor within their area of work. We nonetheless offer some recommendations and resources that we believe may be useful to those who recognize problems in their field and wish for some sort of guidance:
- Collaborate with well-qualified statisticians to design and conduct studies that are rigorous, efficient, and cost effective42, 43. Look to these statistical collaborators for guidance about honestly reporting uncertainty, not certainty19.
- If possible, pool resources to run larger studies that are likely to be more precise and informative than individual studies, which may simply waste resources and offer little yield44. However, more data is not synonymous with more information: Increasing study size may be detrimental if it entails a reduction in data quality45.
- Aim to be more descriptive and less inferential22. Accept uncertainty and the anxiety that comes with it4, along with the idea that no one study warrants conclusions about the true nature of a phenomenon, a delusion that is not even true in particle physics! And if a phenomenon is subtle, even several studies may be insufficient for valid conclusions beyond “more research is needed.”
- When making real world decisions, use all information available to you, balancing costs, benefits, and uncertainties46, rather than basing decisions on whether a single numerical value is above or below an arbitrary cutoff.
- Be mindful of cognitive biases that may distort your conclusions throughout the study and the many cognitive biases that will afflict you, your colleagues, and your collaborators3, 47, 48.
- Use statistical methods that have been reviewed and validated by the statistical community beyond their developers and promoters. That validation is provided not only by publication in statistical journals, but also by applications that can be judged as having reached sound, well-cautioned conclusions in context. Especially, beware of any method that (like NHST) claims to offer firm conclusions based on purely numeric comparisons.
References
Abstract
Differentiates between mathematical and scientific methods. The differences between scientific intuition and mathematical results have been attributed to the fact that scientific generalization is broader than mathematical description. While scientific methods deal with samples which are representative of the total whole, the mathematical methods measure the differences between the particular samples observed. Science begins with description but ends in generalization. Mathematical measures are too high and may need to be discounted in arriving at a scientific conclusion. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Abstract
International audience
Abstract
There is no complete solution for the problem of abuse of statistics, but methodological training needs to cover cognitive biases and other psychosocial factors affecting inferences. The present paper discusses 3 common cognitive distortions: 1) dichotomania, the compulsion to perceive quantities as dichotomous even when dichotomization is unnecessary and misleading, as in inferences based on whether a P value is "statistically significant"; 2) nullism, the tendency to privilege the hypothesis of no difference or no effect when there is no scientific basis for doing so, as when testing only the null hypothesis; and 3) statistical reification, treating hypothetical data distributions and statistical models as if they reflect known physical laws rather than speculative assumptions for thought experiments. As commonly misused, null-hypothesis significance testing combines these cognitive problems to produce highly distorted interpretation and reporting of study results. Interval estimation has so far proven to be an inadequate solution because it involves dichotomization, an avenue for nullism …
Abstract
We discuss problems the null hypothesis significance testing (NHST) paradigm poses for replication and more broadly in the biomedical and social sciences as well as how these problems remain unresolved by proposals involving modified p-value thresholds, confidence intervals, and Bayes factors. We then discuss our own proposal, which is to abandon statistical significance. We recommend dropping the NHST paradigmand the p-value thresholds intrinsic to itas the default statistical paradigm for research, publication, and discovery in the biomedical and social sciences. Specifically, we propose that the p-value be demoted from its threshold screening role and instead, treated continuously, be considered along with currently subordinate factors (e.g., related prior evidence, plausibility of mechanism, study design and data quality, real world costs and benefits, novelty of finding, and other factors that vary by research domain) as just one among many pieces of evidence. We have no desire to ``ban'' p-values or other purely statistical measures …
Abstract
High-quality randomized clinical trials (RCTs) occupy the highest position on the evidence pyramid, either as stand-alone studies or as part of meta-analyses. H
Abstract
Misleading terminology and arbitrary divisions stymie drug trials and can give false hope about the potential of tailoring drugs to individuals, warns Stephen Senn.
Abstract
Tests for statistical interaction have come into increasing use in epidemiologic analysis, with most based on either an additive or multiplicative model for joint effects. Further procedures have been proposed for testing the goodness-of-fit and comparing the fit of the latter models. This paper reviews the relationships between the various tests and model comparison methods, and, for the special case of two dichotomous risk factors, presents asymptotic power functions for tests of additivity and multiplicativity. For a range of sample sizes and factor effects, the powers of the tests are computed using both the asymptotic power function and simulation studies. The powers of the tests are very low in several commonly encountered situations. In addition, convergence to the asymptotic distribution appears slow for some of the statistics. The results also indicate that likelihood comparison procedures can provide a useful adjunct to the classical hypothesis-testing approach.
Abstract
For any given research area, one cannot tell how many studies have been conducted but never reported. The extreme view of the“ file drawer problem” is that journals are filled with the 5\% of the studies that show Type I errors, while the file drawers are filled with the 95\% of …
Abstract
The current concerns about reproducibility have focused attention on proper use of statistics across the sciences. This gives statisticians an extraordinary opportunity to change what are widely regarded as statistical practices detrimental to the cause of good science. However, how that should be done is enormously complex, made more difficult by the balkanization of research methods and statistical traditions across scientific subdisciplines. Working within those sciences while also allying with science reform movementsoperating simultaneously on the micro and macro levelsare the key to making lasting change in applied science.
Abstract
A prominent feature of statistical reasoning for nearly a century, the p-value plays an especially vital role in the clinical testing of new drugs. Over the last fifty years, the U.S. Food and Drug Administration (FDA) has relied on p-values and significance testing to demonstrate the efficacy of new drugs in the premarket approval process. This article seeks to illuminate the history of this statistic and explain how the statistical significance threshold of 0.05, commonly decried as an arbitrary cutoff, is a useful tool that came to be the cornerstone of FDA decision-making.
Abstract
The cost and time of pharmaceutical drug development continue to grow at rates that many say are unsustainable. These trends have enormous impact on what treatments get to patients, when they get them and how they are used. The statistical framework for supporting decisions in regulated clinical development of new medicines has followed a traditional path of frequentist methodology. Trials using hypothesis tests of ``no treatment effect'' are done routinely, and the p-value $<$ 0.05 is often the determinant of what constitutes a ``successful'' trial. Many drugs fail in clinical development, adding to the cost of new medicines, and some evidence points blame at the deficiencies of the frequentist paradigm. An unknown number effective medicines may have been abandoned because trials were declared ``unsuccessful'' due to a p-value exceeding 0.05. Recently, the Bayesian paradigm has shown utility in the clinical drug development process for its probability-based inference …
Abstract
[The attempt to reinterpret the common tests of significance used in scientific research as though they constituted some kind of acceptance procedure and led to "decisions" in Wald's sense, originated in several misapprehensions and has led, apparently, to several more. The three phrases examined here, with a view to elucidating the fallacies they embody, are: (i) "Repeated sampling from the same population", (ii) Errors of the "second kind", (iii) "Inductive behaviour". Mathematicians without personal contact with the Natural Sciences have often been misled by such phrases. The errors to which they lead are not always only numerical.]
Abstract
Different types of experimentation are considered with reference to their logical structure, to show that valid conclusions may be drawn from them without using the disputed theory of inductive inferences, i.e., of arguing from observation to explanatory theory. This is possible if a null hypothesis is explicitly formulated when the experiment is designed; this hypothesis can never be proved, but may be disproved with whatever probability one will accept as demonstrating a positive result. Chapters II, III, and IV illustrate simple applications of the principles involved in sensitiveness, significance, tests of wider hypotheses, validity, and estimation and elimination of error. More elaborate structures are treated in later chapters. Chapter titles are: (V) the Latin square; (VI) factorial design in experimentation; (VII) confounding; (VIII) special cases of partial confounding; (IX) increase of precision by concomitant measurements: statistical control; (X) generalization of null hypotheses: fiducial probability; (XI) measurement of amount of information in general. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Abstract
Abstract The problem of testing statistical hypotheses is an old one. Its origin is usually connected with the name of Thomas Bayes, who gave the well-known theorem on the probabilities a posteriori of the possible “causes“ of a given event. Since then it has been discussed by many writers of whom we shall here mention two only, Bertrand and Borel, whose differing views serve well to illustrate the point from which we shall approach the subject. Bertrand put into statistical form a variety of hypotheses, as for example the hypothesis that a given group of stars with relatively small angular distances between them as seen from the earth, form a “system” or group in space. His method of attack, which is that in common use, consisted essentially in calculating the probability, P, that a certain character, x, of the observed facts would arise if the hypothesis tested were true. If P were very small, this would generally be considered as an indication that the hypothesis, H, was probably false, and vice versa. Bertrand expressed the pessimistic view that no test of this kind could give reliable results …
Abstract
E. S. Pearson, A Survey of the Uses of Statistical Method in the Control and Standardization of the Quality of Manufactured Products, Journal of the Royal Statistical Society, Vol. 96, No. 1 (1933), pp. 21-75
Abstract
In this article, we accomplish two things. First, we show that despite empirical psychologists' nominal endorsement of a low rate of false-positive findings (≤ .05), flexibility in data collection, analysis, and reporting dramatically increases actual false-positive rates. In many cases, a researcher is more likely to falsely find evidence that an effect exists than to correctly find evidence that it does not. We present computer simulations and a pair of actual experiments that demonstrate how unacceptably easy it is to accumulate (and report) statistically significant evidence for a false hypothesis. Second, we suggest a simple, low-cost, and straightforwardly effective disclosure-based solution to this problem. The solution involves six concrete requirements for authors and four guidelines for reviewers, all of which impose a minimal burden on the publication process.
Abstract
Background: Inappropriate analysis and reporting of biomedical research remain a problem despite advances in statistical methods and efforts to educate researchers. Objective: To determine the frequency and severity of requests biostatisticians receive from researchers for inappropriate analysis and reporting of data during statistical consultations. Design: Online survey. Setting: United States. Participants: A randomly drawn sample of 522 American Statistical Association members self-identifying as consulting biostatisticians. Measurements: The Bioethical Issues in Biostatistical Consulting Questionnaire soliciting reports about the frequency and perceived severity of specific requests for inappropriate analysis and reporting. Results: Of 522 consulting biostatisticians contacted, 390 provided sufficient responses: a completion rate of 74.7% …
Abstract
Journal Article Meta-analysis: State-of-the-Science Get access Kay Dickersin, Kay Dickersin 1Department of Epidemiology and Preventive Medicine, University of Maryland School of MedicineBaltimore, Maryland Dr. Kay Dickersin, Department of Epidemiology and Preventive Medicine, University of Maryland School of Medicine, 660 West Redwood Street, Baltimore, MD 21201 Search for other works by this author on: Oxford Academic PubMed Google Scholar Jesse A. Berlin Jesse A. Berlin 2Clinical Epidemiology Unit, Division of General Internal Medicine, Department of Medicine, University of Pennsylvania School of MedicinePhiladelphia, Pennsylvania Search for other works by this author on: Oxford Academic PubMed Google Scholar Epidemiologic Reviews, Volume 14, Issue 1, 1992, Pages 154–176, https://doi.org/10.1093/oxfordjournals.epirev.a036084 Published: 01 March 1992 Article history Received: 10 December 1991 Published: 01 March 1992 Revision received: 20 July 1992
Abstract
Valentin Amrhein, Sander Greenland, Blake McShane and more than 800 signatories call for an end to hyped claims and the dismissal of possibly crucial effects.
Abstract
Statistical inference often fails to replicate. One reason is that many results may be selected for drawing inference because some threshold of a statistic like the P-value was crossed, leading to biased reported effect sizes. Nonetheless, considerable non-replication is to be expected even without selective reporting, and generalizations from single studies are rarely if ever warranted. Honestly reported results must vary from replication to replication because of varying assumption violations and random variation; excessive agreement itself would suggest deeper problems, such as failure to publish results in conflict with group expectations or desires. A general perception of a ``replication crisis'' may thus reflect failure to recognize that statistical tests not only test hypotheses, but countless assumptions and the entire environment in which research takes place. Because of all the uncertain and unknown assumptions that underpin statistical inferences, we should treat inferential statistics as highly unstable local descriptions of relations between assumptions and data, rather than as providing generalizable inferences about hypotheses or models …
Abstract
Emeritus of Mathematics and Statistics at
Abstract
EDITORIAL: The editorial was written by the three editors acting as individuals and reflects their scientific views not an endorsed position of the American Statistical Association.Some of you expl...
Abstract
Some Journal readers may have noticed more parsimonious reporting of P values in our research articles over the past year. For example, in November 2018, we published two reports from the Vitamin D...
Abstract
Importance: Clinical researchers are obligated to present results objectively and accurately to ensure readers are not misled. In studies in which primary end points are not statistically significant, placing a spin, defined as the manipulation of language to potentially mislead readers from the likely truth of the results, can distract the reader and lead to misinterpretation and misapplication of the findings. Objective: To determine the level and prevalence of spin in published reports of cardiovascular randomized clinical trial (RCT) reports. Data Source: MEDLINE was searched from January 1, 2015, to December 31, 2017, using the Cochrane highly sensitive search strategy. Study Selection: Inclusion criteria were parallel-group RCTs published from January 1, 2015, to December 31, 2017 in 1 of 6 high-impact journals (New England Journal of Medicine, The Lancet, JAMA, European Heart Journal, Circulation, and Journal of the American College of Cardiology) with primary outcomes that were not statistically significant were included in the analysis. Data Extraction and Synthesis: Analysis began in August 2018 …
Abstract
We propose to change the default P-value threshold for statistical significance from 0.05 to 0.005 for claims of new discoveries.
Abstract
In this Viewpoint, John Ioannidis discusses the potential effects on clinical research of a 2017 proposal to lower the default P value threshold for statistical significance from .05 to .005 as a means to reduce false-positive findings and reviews alternative solutions for improving the accuracy and...
Abstract
THE CENTRAL THEME OF THE INSTITUTE OF MEDICINE report on the US drug safety system was the need for a life cycle approach to drug evaluation: both the benefits and the risks need to be evaluated and integrated during the entire market life of a drug. The Food and Drug Administration Amendments Act of 2007 also called on the agency to improve its methods of communicating risks and benefits to patients and physicians. The Institute of Medicine recommendation to “develop and continually improve a systematic approach to risk-benefit analysis for use throughout the [Food and Drug Administration] in the preapproval and postapproval settings” specifically acknowledges the need for and the challenges of the development of new methods of combining evidence about risks and benefits. Information that combines the best evidence on benefits with the best data on risks is also needed for daily clinical practice. Whenever a patient and physician decide on a particular course of treatment, they do so because they expect that the likely benefits will exceed potential harms …
Abstract
The Policy Forum allows health policy makers around the world to discuss challenges and opportunities for improving health care in their societies.
Abstract
Randomized controlled clinical trials are conducted to determine whether differences of clinical importance exist between selected treatment regimens. When statistical analysis of the study data finds a P value greater than 5%, it is convention to deem the assessed difference nonsignificant. Just because convention dictates that such study findings be termed nonsignificant, or negative, however, it does not necessarily follow that the study found nothing of clinical importance. Subject samples used in controlled trials tend to be too small. The studies therefore lack the necessary power to detect real, and clinically worthwhile, differences in treatment. Freiman et al. found that only 30% of a sample of 71 trials published in the New England Journal of Medicine in 1978-79 with a P value greater than 10% were large enough to have a 90% chance of detecting even a 50% difference in the effectiveness of the treatments being compared, and they found no improvement in a similar sample of trials published in 1988. It is therefore wrong and unwise to interpret so many negative trials as providing evidence of the ineffectiveness of new treatments …
Abstract
The application of statistics to science is not a neutral act. Statistical tools have shaped and were also shaped by its objects. In the social sciences, statistical methods fundamentally changed research practice, making statistical inference its centerpiece. At the same time, textbook writers in the social sciences have transformed rivaling statistical systems into an apparently monolithic method that could be used mechanically. The idol of a universal method for scientific inference has been worshipped since the ``inference revolution'' of the 1950s. Because no such method has ever been found, surrogates have been created, most notably the quest for significant p values. This form of surrogate science fosters delusions and borderline cheating and has done much harm, creating, for one, a flood of irreproducible results. Proponents of the ``Bayesian revolution'' should be wary of chasing yet another chimera: an apparently universal inference procedure. A better path would be to promote both an understanding of the various devices in the ``statistical toolbox'' and informed judgment to select among these.
Abstract
In clinical trials, study designs may focus on assessment of superiority, equivalence, or non-inferiority, of a new medicine or treatment as compared to a control. Typically, evidence in each of these paradigms is quantified with a variant of the null hypothesis significance test. A null hypothesis is assumed (null effect, inferior by a specific amount, inferior by a specific amount and superior by a specific amount, for superiority, non-inferiority, and equivalence respectively), after which the probabilities of obtaining data more extreme than those observed under these null hypotheses are quantified by p-values. Although ubiquitous in clinical testing, the null hypothesis significance test can lead to a number of difficulties in interpretation of the results of the statistical evidence.
Abstract
Journals tend to publish only statistically significant evidence, creating a scientific record that markedly overstates the size of effects. We provide a new tool that corrects for this bias without requiring access to nonsignificant results. It capitalizes on the fact that the distribution of significant p values, p-curve, is a function of the true underlying effect. Researchers armed only with sample sizes and test results of the published findings can correct for publication bias. We validate the technique with simulations and by reanalyzing data from the Many-Labs Replication project. We demonstrate that p-curve can arrive at conclusions opposite that of existing tools by reanalyzing the meta-analysis of the “choice overload” literature.
Abstract
An abstract is unavailable.
Abstract
Background Many articles in the surgical literature were faulted for committing type 2 error, or concluding no difference when the study was ``underpowered''. However, it is unknown if the current power standard of 0.8 is reasonable in surgical science. Methods PubMed was searched for abstracts published in Surgery, JAMA Surgery, and Annals of Surgery and from January 1, 2012 to December 31, 2016, with Medical Subject Heading terms of randomized controlled trial (RCT) or observational study (OBS) and limited to humans were included (n~=~403). Articles were excluded if all reported findings were statistically significant (n~=~193), or if presented data were insufficient to calculate power (n~=~141). Results A total of 69 manuscripts (59 RCTs and 10 OBSs) were assessed. Overall, the median power was 0.16 (interquartile range [IQR] 0.08-0.32). The median power was 0.16 for RCTs (IQR 0.08-0.32) and 0.14 for OBSs (IQR 0.09-0.22). Only 4 studies (5.8\%) reached or exceeded the current 0.8 standard. Two-thirds of our study sample had an a priori power calculation (n~=~41). Conclusions High-impact surgical science was routinely unable to reach the arbitrary power standard of 0.8 …
Abstract
Department of Surgery, Massachusetts General Hospital, Harvard Medical School, Boston, MA. [email protected]. Disclosure: The authors declare that they have no conflicts of interest.
Abstract
This article summarizes arguments against the use of power to analyze data, and illustrates a key pitfall: Lack of statistical significance (e.g., p > .05) combined with high power (e.g., 90\%) can occur even if the data support the alternative more than the null. This problem arises via selective choice of parameters at which power is calculated, but can also arise if one computes power at a prespecified alternative. As noted by earlier authors, power computed using sample estimates (“observed power”) replaces this problem with even more counterintuitive behavior, because observed power effectively double counts the data and increases as the P value declines. Use of power to analyze and interpret data thus needs more extensive discouragement.
Abstract
It is well known that statistical power calculations can be valuable in planning an experiment. There is also a large literature advocating that power calculations be made whenever one performs a statistical test of a hypothesis and one obtains a statistically nonsignificant result. Advocates of such post-experiment power calculations claim the calculations should be used to aid in the interpretation of the experimental results. This approach, which appears in various forms, is fundamentally flawed. We document that the problem is extensive and present arguments to demonstrate the flaw in the logic.
Abstract
It is commonly stated that individuals respond differently to exercise even when the same exercise intervention is performed. This has led many researchers to conduct exercise interventions and subsequently categorize individuals into different responder categories to determine what causes individuals to respond differently. Some methods by which differential responders are categorized include percentile ranks, standard deviations from the mean, and cluster analyses. Notably, each of these methods will result in the presence of differential responders even in the absence of an exercise intervention, indicating that individuals may be categorized based on the presence of random error as opposed to true differences in the exercise response. Here we propose a method by which differential responders can be classified after accounting for the presence of random error that is quantified from a time-matched control group. Individuals who exceed random error from the mean response of the intervention group can be confidently labelled as high and low responders …
Abstract
Dankel \& Loenneke (2019) recently presented a new approach to identifying subgroups in parallel group study designs. Here, we briefly discuss our statistical concerns with proposed approach. We reveal that the error rates of the Danke-Loenneke approach are much higher than the claimed 5\%, and that these error rates are dependent on numerous factors, including sample size, effect variance, and random error. The Dankel-Loenneke method has poor statistical properties; as such, we suggest that the method not be used and the manuscript constitutes an "honest error" per the Committee on Publication Ethics (COPE) guidelines.
Abstract
Statistical power analysis provides the conventional approach to assess error rates when designing a research study. However, power analysis is flawed in that a narrow emphasis on statistical significance is placed as the primary focus of study design. In noisy, small-sample settings, statistically significant results can often be misleading. To help researchers address this problem in the context of their own studies, we recommend design calculations in which (a) the probability of an estimate being in the wrong direction (Type S [sign] error) and (b) the factor by which the magnitude of an effect might be overestimated (Type M [magnitude] error or exaggeration ratio) are estimated. We illustrate with examples from recent published research and discuss the largest challenge in a design calculation: coming up with reasonable estimates of plausible effect sizes based on external information.
Abstract
Study size has typically been planned based on statistical power and therefore has been heavily influenced by the philosophy of statistical hypothesis testing. A worthwhile alternative is to plan study size based on precision, for example by aiming to obtain a desired width of a confidence interval for the targeted effect. This article presents formulas for planning the size of an epidemiologic study based on the desired precision of the basic epidemiologic effect measures.
Abstract
Concerns about the veracity of psychological research have been growing. Many findings in psychological science are based on studies with insufficient statistic...
Abstract
Applied statistics is more than data analysis, but it is easy to lose sight of the big picture. David Cox and Christl Donnelly distil decades of scientific experience into usable principles for the successful application of statistics, showing how good statistical strategy shapes every stage of an investigation. As you advance from research or policy question, to study design, through modelling and interpretation, and finally to meaningful conclusions, this book will be a valuable guide. Over a hundred illustrations from a wide variety of real applications make the conceptual points concrete, illuminating your path and deepening your understanding. This book is essential reading for anyone who makes extensive use of statistical methods in their work.
Abstract
Decision theory provides a formal framework for making logical choices in the face of uncertainty. Given a set of alternatives, a set of consequences, and a correspondence between those sets, decision theory offers conceptually simple procedures for choice. This book presents an overview of the fundamental concepts and outcomes of rational decision making under uncertainty, highlighting the implications for statistical practice. The authors have developed a series of self contained chapters focusing on bridging the gaps between the different fields that have contributed to rational decision making and presenting ideas in a unified framework and notation while respecting and highlighting the different and sometimes conflicting perspectives. This book: * Provides a rich collection of techniques and procedures. * Discusses the foundational aspects and modern day practice. * Links foundations to practical applications in biostatistics, computer science, engineering and economics. * Presents different perspectives and controversies to encourage readers to form their own opinion of decision making and statistics …
Abstract
Statistical rituals largely eliminate statistical thinking in the social sciences. Rituals are indispensable for identification with social groups, but they should be the subject rather than the procedure of science. What I call the ``null ritual'' consists of three steps: (1) set up a statistical null hypothesis, but do not specify your own hypothesis nor any alternative hypothesis, (2) use the 5\% significance level for rejecting the null and accepting your hypothesis, and (3) always perform this procedure. I report evidence of the resulting collective confusion and fears about sanctions on the part of students and teachers, researchers and editors, as well as textbook writers.
Abstract
The mechanical, ritualistic application of statistics is contributing to a crisis in science. Education, software and peer review have encouraged poor practice ? and it is time for statisticians to fight back. By Philip B. Stark and Andrea Saltelli The full-text article may be found on the RSS website.
Citation
@online{rafi2020,
author = {Rafi, Zad and Rafi, Zad and Gelman, Andrew and Reito, Aleksi
and Greenland, Sander},
title = {Medicine {Is} {Being} {Treated} with {Snake-Oil}
{Statistics}},
date = {2020-11-11},
url = {https://lesslikely.com/statistics/snakeoilstats.html},
langid = {en-US}
}
Comments