Selection On Observables Only
Causal Identification
Also known as: Selection On Observables Faith
Definition
It's assumed that differences between groups come only from the factors that got measured. Whatever wasn't measured is treated as if it doesn't change the result at all.
Advanced definition
Selection on observables posits that treatment assignment is independent of the potential outcomes once conditioned on the observed covariates, assuming no unobserved confounders exist. This identification strategy relies on the measured covariates capturing every systematic difference that influences both assignment and outcome.
Example
A school district wants to know whether a new tutoring program raises test scores. Analysts compare students who enrolled in tutoring against those who did not, adjusting for grade level, prior GPA, and family income. If motivated parents systematically enroll their children and parental motivation is not measured, the apparent benefit of tutoring is inflated — the assumption that all relevant differences are captured by the measured variables is violated.
Advanced example
A health economist uses propensity-score matching on age, sex, comorbidity index, and insurance type to estimate the effect of a new antihypertensive on 5-year cardiovascular mortality in an administrative claims dataset. The study reports a statistically significant protective effect. However, medication adherence and lifestyle behaviors like diet and exercise are unobserved in claims data, and they correlate both with prescribing patterns (sicker, less adherent patients may be assigned the older drug) and with mortality. The conditional independence assumption — treatment ignorability given the observed adjustment set — is violated. A formal E-value calculation reveals that an unmeasured confounder with a relative risk of only 1.8 for both treatment assignment and outcome could fully explain the estimated effect, well within the plausible range for adherence-related confounding. Without an exclusion restriction or negative control outcome test, the backdoor adjustment stays incomplete, and the reported estimate carries unobserved confounding contamination.
Mechanism
When groups with the same measured traits get compared, any difference left over is attributed to the treatment. Measured traits act like filters that make the groups look alike.
Advanced mechanism
Observed covariates constrain the assignment probabilities so that, conditional on them, treatment is asymmetrically independent of the potential outcomes; the covariate set functions as a structural conditioning element. That weighting of the conditional strata enforces the identification constraint by removing the bias attributable to measured confounders.
How to counter it
Collect more relevant measurements to reduce the hidden differences. Designs that compare genuinely similar people close the gap further.
Advanced countermove
Instrumental variables or panel methods augment the analysis to address hidden confounding and relax the conditional independence assumption. Sensitivity analysis quantifies how robust the result is to unobserved biases.
Failure modes
Unmeasured confounders present; Poor covariate overlap; Modeling mis-specification
Exploitation surface
An adversarial analyst can deliberately select only favorable observable covariates into an adjustment set, creating the appearance of rigorous causal identification while leaving known unmeasured confounders unaddressed. This strategy allows a researcher or sponsor to report conditionally unbiased estimates that are in fact deeply confounded, since reviewers cannot easily audit what was omitted from the covariate set. Selective covariate disclosure is especially potent in proprietary datasets where outsiders cannot verify the completeness of the measured variable list.
Resistance profile
Pre-register the full covariate selection protocol and the theoretical causal graph (DAG) prior to data collection so that omissions become auditable against a stated commitment. Pair observational estimates with sensitivity analyses (e.g., Rosenbaum bounds, E-values) to quantify how strong unmeasured confounding would need to be to overturn conclusions. Where feasible, supplement with a design-based identification strategy—instrumental variables, difference-in-differences, or regression discontinuity—to triangulate against the selection-on-observables assumption.