The Atlas 6,943 concepts
☆ Favorites

Selection On Observables Only

Statistical Errors Phenomenon Empirical
Causal Identification
Also known as: Selection On Observables Faith
Detection: high Stability: persistent Level: intermediate
It's assumed that differences between groups come only from the factors that got measured. Whatever wasn't measured is treated as if it doesn't change the result at all.
Selection on observables posits that treatment assignment is independent of the potential outcomes once conditioned on the observed covariates, assuming no unobserved confounders exist. This identification strategy relies on the measured covariates capturing every systematic difference that influences both assignment and outcome.
A school district wants to know whether a new tutoring program raises test scores. Analysts compare students who enrolled in tutoring against those who did not, adjusting for grade level, prior GPA, and family income. If motivated parents systematically enroll their children and parental motivation is not measured, the apparent benefit of tutoring is inflated — the assumption that all relevant differences are captured by the measured variables is violated.
A health economist uses propensity-score matching on age, sex, comorbidity index, and insurance type to estimate the effect of a new antihypertensive on 5-year cardiovascular mortality in an administrative claims dataset. The study reports a statistically significant protective effect. However, medication adherence and lifestyle behaviors like diet and exercise are unobserved in claims data, and they correlate both with prescribing patterns (sicker, less adherent patients may be assigned the older drug) and with mortality. The conditional independence assumption — treatment ignorability given the observed adjustment set — is violated. A formal E-value calculation reveals that an unmeasured confounder with a relative risk of only 1.8 for both treatment assignment and outcome could fully explain the estimated effect, well within the plausible range for adherence-related confounding. Without an exclusion restriction or negative control outcome test, the backdoor adjustment stays incomplete, and the reported estimate carries unobserved confounding contamination.
When groups with the same measured traits get compared, any difference left over is attributed to the treatment. Measured traits act like filters that make the groups look alike.
Observed covariates constrain the assignment probabilities so that, conditional on them, treatment is asymmetrically independent of the potential outcomes; the covariate set functions as a structural conditioning element. That weighting of the conditional strata enforces the identification constraint by removing the bias attributable to measured confounders.
Collect more relevant measurements to reduce the hidden differences. Designs that compare genuinely similar people close the gap further.
Instrumental variables or panel methods augment the analysis to address hidden confounding and relax the conditional independence assumption. Sensitivity analysis quantifies how robust the result is to unobserved biases.
Unmeasured confounders present; Poor covariate overlap; Modeling mis-specification
An adversarial analyst can deliberately select only favorable observable covariates into an adjustment set, creating the appearance of rigorous causal identification while leaving known unmeasured confounders unaddressed. This strategy allows a researcher or sponsor to report conditionally unbiased estimates that are in fact deeply confounded, since reviewers cannot easily audit what was omitted from the covariate set. Selective covariate disclosure is especially potent in proprietary datasets where outsiders cannot verify the completeness of the measured variable list.
Pre-register the full covariate selection protocol and the theoretical causal graph (DAG) prior to data collection so that omissions become auditable against a stated commitment. Pair observational estimates with sensitivity analyses (e.g., Rosenbaum bounds, E-values) to quantify how strong unmeasured confounding would need to be to overturn conclusions. Where feasible, supplement with a design-based identification strategy—instrumental variables, difference-in-differences, or regression discontinuity—to triangulate against the selection-on-observables assumption.