Statistical Methods with R

Questionable Practices

Unit E · Chapter 14 · Lecture 14b

Developed by Jeffrey M. Girard

Overview

Research practice continuum

Scientific Misconduct

Data Fabrication, Data Falsification, (Plagiarism), (Undisclosed COIs)

Questionable Practices

Selective Reporting, p-hacking, HARKing, Data Peeking

Best Practices

Open Data/Materials, Preregistration, Disclosed Flexibility, Power Analysis

Ordered from least to most likely to produce true, replicable findings

Questionable Practices

Selective reporting

  • Cherry picking is only including or reporting a desired subset of information (intentionally or accidentally)
    • e.g., cases, variables, studies, analyses, results, references
    • It is problematic when the subset misrepresents the whole
    • It is especially problematic when unreported (e.g., hidden)
  • Selective omission is excluding or not reporting an undesired subset of information (intentionally or accidentally)
    • This is essentially the other side of cherry picking

HARKing

  • “Hypothesizing after the results are known”
    • It is claiming to have predicted an unexpected result
    • Changing your hypothesis to match whatever you found
    • Going into research without any real expectations
  • Why it’s problematic
    • It falsely presents exploratory work as confirmatory
    • The problem is the misrepresentation (not the exploration)
    • It is more likely to capitalize on sampling error
    • “The easiest person to fool is yourself”

p-hacking

  • Making analysis decisions to optimize significance instead of rigor

    • How to measure and represent variables
    • Which observations to include or exclude
    • When to stop data collection
    • Which tests and formulas to use
    • Which covariates to control for
    • How to handle assumptions and outliers
  • If you try lots of options, a few will be significant by chance alone

  • If you don’t disclose all this trial-and-error, it is very misleading

Data peeking

  • Data peeking is a form of p-hacking related to data collection
    • You collect some data, run analyses, and look at the p-value
    • If your result is significant, then you stop and publish
    • But if not, then you collect more data and try again…
  • This practice capitalizes on sampling error and is biased
    • Your effect size will likely be inflated (too large)
    • A correction for this is called “sequential analysis”
  • A better approach is to choose \(n\) using power analysis

Constraining the Garden of Forking Paths

Preregistration

  • Presenting exploratory work as confirmatory is misleading

  • We can avoid this by “calling our shots” before data collection

    • By preregistering our hypotheses, we avoid HARKing
    • By preregistering our analyses, we avoid p-hacking
    • It’s okay to change things, but you have to disclose it
  • Many journals now offer registered reports

    • Reviewers evaluate and refine your preregistered plan
    • The journal will publish your paper if you follow the plan