Does Near-Peer Instruction Actually Help? Untangling Selection Bias in Graduate Physiology

Published research examining near-peer instruction modalities and exam performance trajectories in a graduate Human Physiology course at Boston University Chobanian & Avedisian School of Medicine.
Voluntary near-peer instruction (NPI) programs are a fixture of health professions education, but evaluating whether they actually work is harder than it sounds. Students who seek out tutoring or extra review sessions are not a random sample of the cohort — they tend to be the ones who are already struggling. This creates a well-documented statistical problem known as confounding by indication: programs designed to help can appear harmful in naive comparisons simply because the students using them started from a worse position.
This paper, co-authored with Dr. Christopher Schonhoff at Boston University Chobanian & Avedisian School of Medicine, investigates an NPI program within a graduate Human Physiology course serving 191 MAMS students across three examination blocks in Fall 2025.
Co-Authors: Joseph D. Webb & Christopher M. Schonhoff, PhD
Published in: Physiology (American Physiological Society), 2026
The Program
The NPI program comprised two structurally distinct components:
Structured TA Sessions — Weekly 90-minute group review sessions led by course TAs (selected through competitive hire) and organized in a 30/30/30 format: content review, practice problems, and Socratic discussion. Session content was faculty-reviewed before delivery.
Informal Peer Tutoring — Individual two-hour sessions matched by scheduling availability and self-reported academic compatibility. Students were softly referred to tutoring after low quiz scores, but enrollment was fully voluntary. Tutors had no formal onboarding process and session content was determined organically by each tutor-tutee pair.
These two modalities differ in content standardization, faculty oversight, scheduling regularity, and the type of student they tend to attract. Treating them as a single “NPI exposure” would miss meaningful variation.
Study Design
Using a lagged panel of 350 student-exam observations across Blocks 2 and 3 (with each student’s prior block exam score serving as a baseline covariate), we applied four analytical approaches representing progressively more rigorous causal control:
- Naive OLS — raw unadjusted association; included to illustrate the baseline confounding problem
- Prior-Score Adjusted OLS — controls for where each student started before the block
- Inverse Probability of Treatment Weighting (IPTW) — statistically rebalances the tutoring and no-tutoring groups on observed characteristics (standardized mean difference in prior score reduced from −0.256 to −0.063)
- Within-Student Fixed Effects — uses each student as their own control across blocks, removing the influence of all stable unmeasured characteristics (help-seeking tendency, baseline test anxiety, etc.)
Key Findings
The raw data told a discouraging story: students in the tutoring group scored substantially lower than peers who sought no support. But this pattern unraveled as controls were added.
The Illusion of Harm
| Model | Tutoring β (pp) | TA Session β (pp) |
|---|---|---|
| Model 1: Naive OLS | −4.92* | −0.92 |
| Model 2: Prior-Score Adj. OLS | −2.14 | −0.51 |
| Model 3: IPTW | −2.20* | −0.58 |
| Model 4: Within-Student Fixed Effects | −1.33 | +0.68 |
*p < 0.05
The tutoring estimate fell by over 70% from Model 1 to Model 4, ultimately resolving to a non-significant −1.33 pp (95% CI: −8.71, +6.06). This is the signature of confounding by indication, not a genuine negative program effect. Students using tutoring started from lower baselines, and naive comparisons captured that disadvantage rather than the program’s impact.
A Structural Hypothesis
The TA estimate moved in the opposite direction over the same sequence — from −0.92 pp in Model 1 to +0.68 pp under the most stringent fixed effects specification. It did not reach statistical significance, but the directional consistency across all four models is not what we would expect from an ineffective intervention. Under Model 4 (which eliminates all stable student-level confounders), the two modalities diverged by 2.01 pp, pointing in opposite directions.
This divergence is consistent with the structural differences between the two programs. Structured TA sessions incorporated spaced practice, interleaving, and worked examples with feedback — features with robust support in cognitive psychology. Unstructured tutoring, driven by immediate student need without a standardized format, does not systematically build in these elements. Program design may be a meaningful moderator of NPI effectiveness.
Limitations
- Single cohort, retrospective: limited statistical power, particularly for the fixed effects model which relies on within-student variation across only two blocks
- Unmeasured confounders: help-seeking tendency, prior GPA, test anxiety, and undergraduate background could not be fully controlled despite IPTW adjustment
- Intervention fidelity: individual tutoring session quality was not formally assessed; anecdotal evidence suggests considerable variation
Conclusions
Naive comparisons of voluntary NPI participants and non-participants will systematically underestimate program effectiveness. The block-level trajectories and four-model causal framework reported here confirm that the apparent tutoring harm is a selection artifact — one that progressively attenuates to null as confounding is addressed. Meanwhile, structured TA sessions trend consistently positive as controls strengthen.
For program administrators, the takeaway is twofold: selection bias must be accounted for before concluding a program is ineffective, and program design features such as standardization, faculty oversight, and scheduled delivery may matter more than participation rates alone.
Acknowledgments
Thank you to Dr. Janice Weinberg, Professor of Biostatistics at Boston University School of Public Health, for her consultation on statistical methods. This study was approved as exempt by the Boston University Institutional Review Board (Protocol H-46464).
Published in Physiology, American Physiological Society, 2026. Full text available via the link above.