How to read medical evidence without a degree

Read a medical study without a degree: the evidence hierarchy, what a p-value really means, effect size vs significance, and how to spot a thin study.

Stack of opened books on a sunlit wooden table

For research and educational purposes only. Not medical advice.

Category: Research Gaps. 9 min read. By pepSmart Editorial. . .

Key takeaways

  • Evidence pyramid (loosely): systematic reviews and meta-analyses > randomized controlled trials > prospective cohorts > case-control > case series > expert opinion.
  • A small RCT with poor blinding and high dropout can be less informative than a large, well-designed prospective cohort, so treat pyramid position as a rough guide rather than a fixed ranking.
  • P-values do not measure effect size or clinical significance; effect size and confidence intervals are what tell you whether a finding is practically meaningful .
  • GRADE (from the GRADE Working Group, adopted by Cochrane) and Cochrane's Risk of Bias 2 grade a study on its actual conduct, not its label, and modern systematic reviews report GRADE quality next to the headline result .
  • Diagnostic questions for any cited study: who funded it, what was the comparator, were endpoints prespecified, was the population representative, did dropout differ between arms, was the effect meaningful in absolute terms.

The evidence pyramid, fairly stated

Clinical-research methodology textbooks describe an evidence pyramid that ranks study designs by how susceptible they are to bias and confounding. The bottom is a single anecdote: one person reports an outcome. The top is a high-quality systematic review and meta-analysis of large randomized controlled trials. Between those endpoints sit case series, observational cohort studies, case-control studies, individual randomized trials, and unsystematic reviews.

Treat the pyramid as a rough first cut with real exceptions. A small RCT with poor blinding and high dropout can be less informative than a large, well-designed prospective cohort. A systematic review of weak studies inherits the weakness of its inputs. That is why methodologists spent decades building tools that grade a study on how it was actually run: the GRADE Working Group's GRADE approach and Cochrane's Risk of Bias 2 . The pyramid points you in the right direction. The grading tools tell you whether a given study earned its spot.

  • Systematic review + meta-analysis (high quality): synthesis of multiple primary studies under a pre-registered protocol with defined inclusion criteria and quantitative pooling. Strongest when the underlying studies are themselves well-designed.
  • Randomized controlled trial (RCT): the most informative single-study design because randomization neutralizes baseline confounding in expectation. Quality varies by sample size, blinding, dropout, and analysis plan adherence.
  • Prospective cohort: follows a defined population forward over time. Cannot fix confounding by random assignment, but can adjust for measured confounders. Good for long-term outcomes that are unethical to randomize.
  • Case-control: compares people with the outcome to people without, looking backward. Efficient for rare outcomes. Vulnerable to recall bias.
  • Cross-sectional: snapshot at a single time point. Useful for prevalence, not causation.
  • Case series: reports outcomes for a defined patient group, no comparator. Hypothesis-generating, not confirmatory.
  • Single case report: one person, one outcome. Useful only as a flag for further investigation.

What p-values actually say, and what they do not

A p-value is the probability, assuming the null hypothesis is true, of observing a result at least as extreme as the one obtained. It is a probability about the data given the null, not a probability that the null is true given the data, and not a probability that the alternative hypothesis is true. The American Statistical Association issued a formal statement clarifying these distinctions in 2016 because the misinterpretations are widespread in the primary literature .

Concretely, p < 0.05 does not mean the result is true with 95 percent probability, that the effect is large, or that it will replicate in another sample. It means only this: if there were genuinely no effect, the chance of observing a result at least this extreme by random sampling alone would be under 5 percent. Useful, but it is one number among several you need.

Effect size beats statistical significance

A trial with 10,000 patients can detect a 0.1 kg weight difference with high statistical significance. A 0.1 kg difference is meaningless to a person trying to lose 30 kg. Effect size measures the magnitude of the difference between groups, independent of sample size. Common measures include Cohen's d for continuous outcomes (small around 0.2, medium around 0.5, large around 0.8), risk ratios and risk differences for binary outcomes, and absolute risk reduction (ARR) and number needed to treat (NNT) for clinical decisions.

Reading a trial well means looking at the effect size first and the p-value second. The primary endpoint usually reports both. STEP-1 reported a mean 14.9 percent weight reduction on semaglutide at 68 weeks, 12.4 percentage points more than placebo, which is a large effect . SURMOUNT-1 reported 20.9 percent on the highest tirzepatide dose at 72 weeks, 17.8 points more than placebo, which is larger still . Both p-values are vanishingly small, but the p-value is not what makes either result worth reading.

Confidence intervals tell you what the trial actually narrowed down

A 95 percent confidence interval is the range of effect sizes the data are consistent with. A trial reporting a mean weight reduction of 12 percent with a 95 percent CI of 11 to 13 percent has tightly constrained the effect; a different trial reporting the same 12 percent with a CI of 4 to 20 percent has not.

Wide confidence intervals usually come from small samples, large within-group variance, or both. They do not mean the result is wrong; they mean the trial has weakly constrained the answer. A pilot study showing a large but uncertain effect (large point estimate, wide CI) is hypothesis-generating. A confirmatory trial that converts that to a tighter CI is the load-bearing follow-up.

Heterogeneity, and the meta-analysis traps that come with it

Meta-analysis pools effect estimates across studies. The pooling assumption (that the studies are estimating the same underlying effect) breaks down when the studies actually measured different things, used different doses, enrolled different populations, or ran for different durations. Heterogeneity statistics (I-squared, the Q-statistic) try to quantify how much the across-study variability exceeds within-study variability. A high I-squared is a flag that the pooled estimate may be hiding meaningful between-study differences; the widely used rule of thumb, proposed by Higgins and colleagues, tentatively treats 50 percent as moderate and 75 percent as high heterogeneity .

  • When a meta-analysis reports high heterogeneity, read the forest plot. If individual study estimates are spread across both sides of the line of no effect, the pooled point estimate is averaging studies that disagree.
  • Subgroup analysis can sometimes resolve heterogeneity (different doses, different populations). When it does not, the pooled effect should be reported with its uncertainty rather than treated as a single answer.
  • Publication bias inflates pooled effect estimates because positive trials are more likely to be published than null trials. Funnel plots and Egger's tests detect this; trim-and-fill methods adjust for it.
  • Pre-registered systematic reviews (PROSPERO, Cochrane) are less vulnerable to selective inclusion than non-registered narrative reviews.

Diagnostic questions that separate a load-bearing study from a thin one

  1. What was the primary endpoint, and was it pre-specified before data collection? Pre-specified endpoints are more credible than secondary endpoints repurposed as headlines.
  2. How were participants assigned to groups? Randomization neutralizes baseline confounding; non-random assignment leaves it.
  3. Was the trial blinded, and to whom? Double-blind reduces both expectation effects in patients and assessment bias in clinicians.
  4. What was the dropout rate, and was the analysis intention-to-treat? High dropout combined with per-protocol analysis often inflates effect size.
  5. Was the sample size calculation reported? Trials underpowered for the primary endpoint should be read as hypothesis-generating regardless of significance.
  6. How are missing data handled? Last-observation-carried-forward, multiple imputation, and complete-case analysis can give different answers from the same dataset.
  7. Were there protocol deviations? Mid-trial changes to the primary endpoint, dose, or analysis plan substantially weaken the result.
  8. Who funded the trial, and what is the conflict-of-interest disclosure? Industry funding does not automatically invalidate a result, but it shifts the burden of proof on subgroup analyses and post-hoc framing.
  9. Is the journal peer-reviewed? Preprints are useful but unreviewed; pay attention to reviewer comments when the published version becomes available.
  10. Has the result been replicated? A single trial is hypothesis-generating; replication across labs and populations is what makes a finding load-bearing.

How to actually search PubMed when you have a question

PubMed indexes most biomedical primary literature. The search syntax matters. Boolean operators (AND, OR, NOT in capitals), MeSH terms (the indexed subject vocabulary), and field tags ([Title], [Author], [Journal]) all narrow a search faster than free-text. The PubMed User Guide walks through the syntax and the filter set .

  • Start with the compound name and the indication or outcome. 'tirzepatide AND obesity' is a reasonable starting query.
  • Use Filters: Article Type (Randomized Controlled Trial, Meta-Analysis), Publication Date (last 5 years), Species (Humans).
  • Sort by 'Best Match' for relevance and 'Most Recent' to find current literature.
  • When a paper looks promising, read the abstract first; check the effect size and confidence interval; check whether the abstract conclusion matches the data presented in the abstract methods and results.
  • For paywalled articles, the abstract is usually free; the full text often is not. PubMed Central (PMC) hosts open-access versions where available.

What this changes when reading pepSmart and similar references

pepSmart's library catalog and articles cite primary sources by URL. Each link in the references list points at a specific paper, regulatory page, or trial registration. Deciding whether to trust a claim should not stop at 'pepSmart said it'; it should continue to 'what does the cited source say, and is the source itself credible'. The diagnostic questions above apply equally to anything pepSmart cites and to anything any other source cites.

For research and educational purposes only. Not medical advice.

pepSmart has not commissioned independent clinical review of this article.

More on how we write and source these pieces: Editorial process and contributor disclosure and Sourcing posture.

Spot an error? Email corrections via /about.

Sources: 7 entries, all primary canon (peer-reviewed methodology and trial papers on PubMed/PMC, the ASA p-value statement, and the NCBI PubMed User Guide), last reviewed 2026-07-08.

Related tools

References

  1. [1] Guyatt et al., GRADE: an emerging consensus on rating quality of evidence and strength of recommendations (BMJ 2008; PMID 18436948) (PubMed)
  2. [2] Sterne et al., RoB 2: a revised tool for assessing risk of bias in randomised trials (BMJ 2019; PMID 31462531) (PubMed)
  3. [3] Wasserstein RL, Lazar NA. The ASA Statement on p-Values: Context, Process, and Purpose. The American Statistician 2016;70(2):129-133 (American Statistical Association)
  4. [4] STEP-1: Once-weekly semaglutide in adults with overweight or obesity (Wilding et al., NEJM 2021; PMID 33567185; semaglutide -14.9% vs placebo -2.4%, treatment difference -12.4 percentage points) (PubMed)
  5. [5] SURMOUNT-1: Tirzepatide once weekly for the treatment of obesity (Jastreboff et al., NEJM 2022; PMID 35658024; 15 mg -20.9% vs placebo -3.1%, treatment difference -17.8 percentage points) (PubMed)
  6. [6] Higgins JPT et al. Measuring inconsistency in meta-analyses. BMJ 2003 (PMID 12958120; PMC192859). Tentatively assigns I-squared values of 25%, 50%, and 75% to low, moderate, and high heterogeneity. (PubMed Central)
  7. [7] NCBI: PubMed User Guide (search syntax, Boolean operators, field tags, MeSH terms, and filters) (NCBI)

For research and educational purposes only. Not medical advice.