Complete Subset Averaging with Many Instruments
Original authors: Seojeong Lee, Youngki Shin
Original authors: Seojeong Lee, Youngki Shin
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Technical Summary: Complete Subset Averaging with Many Instruments
Problem Statement
Instrumental variables (IV) estimators, particularly Two-Stage Least Squares (2SLS), are standard for addressing endogeneity in economic models. While using a large set of instruments (over-identification) can improve efficiency, it introduces a trade-off: increasing the number of instruments often increases the finite-sample bias of the estimator. Existing literature has addressed this through model selection (e.g., Donald and Newey, 2001), model averaging with estimated weights (e.g., Kuersteiner and Okui, 2010), or shrinkage methods (e.g., Okui, 2011). However, these approaches face limitations when instruments are highly correlated (where model selection like Lasso may fail) or when estimating optimal weights in high dimensions leads to finite-sample efficiency losses. Furthermore, existing model averaging approaches often require nested models or pre-specified main sets of instruments.
Methodology
The authors propose the Complete Subset Averaging (CSA) 2SLS estimator. This method operates in the first stage by constructing an equal-weighted average of fitted values over all possible complete subsets of k instruments chosen from a total of K available instruments.
Estimator Construction:
- Let ZK,i be the vector of K excluded instruments.
- For a chosen subset size k, there are M=(kK) possible subsets.
- For each subset m, a standard 2SLS projection matrix Pmk is computed.
- The CSA projection matrix, denoted Pk, is the arithmetic mean of these M projection matrices: Pk=M1∑m=1MPmk.
- The estimator is defined as β^=(X′PkX)−1X′Pky.
- Crucially, Pk is symmetric but generally not idempotent, distinguishing it from standard 2SLS projection matrices.
Subset Size Selection:
- The subset size k is chosen by minimizing the sample counterpart of an approximate Mean Squared Error (MSE).
- The approximate MSE is derived using a Nagar (1959) expansion, which captures the bias-variance trade-off in finite samples.
- The authors also propose a Cross-Validation (CV) method as an alternative to the approximate MSE for selecting k, which minimizes the first-stage mean squared prediction error.
Handling Irrelevant Instruments:
- The paper generalizes the approximate MSE to scenarios where the number of irrelevant instruments grows with the sample size.
- The derived formula includes a penalty term dependent on the ratio of relevant to total subsets, suggesting that the optimal k should be larger in the presence of irrelevant instruments to mitigate the dilution of signal.
Key Theoretical Contributions
- Approximate MSE Derivation: The authors derive the approximate MSE for the CSA-2SLS estimator using Nagar expansion. A significant technical challenge addressed is that the average of non-nested projection matrices is not idempotent, requiring new analytical techniques compared to literature that assumes nested models. The derived formula explicitly shows the bias-variance trade-off.
- Asymptotic Optimality: The paper proves that the CSA-2SLS estimator, with k chosen to minimize the sample approximate MSE, is asymptotically optimal within the class of CSA-2SLS estimators with different subset sizes. This result is established based on the optimality criteria of Li (1987) and Whittle (1960).
- Generalization to Irrelevant Instruments: The authors extend the MSE analysis to cases with a growing set of irrelevant instruments. They show that the presence of irrelevant instruments increases the penalty for small subset sizes, theoretically justifying the selection of a larger k than would be chosen in a setting with only relevant instruments.
- Orthogonal Instruments Case: The paper demonstrates that if instruments are mutually orthogonal, the CSA-2SLS estimator is identical to the standard 2SLS estimator using all K instruments, regardless of the chosen k.
Empirical and Simulation Results
- Monte Carlo Simulations: Extensive simulations show that when instruments are correlated, the CSA-2SLS estimator outperforms alternative methods (including Donald-Newey selection, Kuersteiner-Okui averaging, and standard 2SLS) in terms of bias and Mean Squared Error. The method is particularly effective in reducing the "moments problem" (large outliers) associated with small instrument sets in the presence of high endogeneity.
- Empirical Illustration: The authors apply the method to estimate a logistic demand function for automobiles using data from Berry, Levinsohn, and Pakes (1995).
- In an extended design with many instruments, standard 2SLS and other estimators (DN, KO) produced estimates implying inelastic demand for a large portion of products, contradicting economic theory.
- The CSA-2SLS estimates resulted in a much larger absolute value for the price elasticity coefficient, with only 0.3% of products showing inelastic demand. This suggests the CSA estimator successfully corrected the many-instrument bias that pulled estimates toward the inconsistent OLS results.
Significance and Claims
The paper claims that the CSA-2SLS estimator offers a robust alternative for settings with many instruments, particularly when those instruments are correlated. By utilizing equal-weighted averaging over complete subsets rather than model selection or weight estimation, the method avoids the pitfalls of breaking down under high correlation and the efficiency losses of estimating high-dimensional weights. The authors assert that the method achieves asymptotic optimality within its class and provides substantial finite-sample gains in bias and MSE reduction compared to existing approaches, as evidenced by both simulation and empirical application.
Limitations Noted by Authors
- The derivation of the approximate MSE assumes conditionally homoskedastic errors.
- The focus is strictly on the 2SLS estimator; extensions to k-class estimators like Limited Information Maximum Likelihood (LIML) are deferred to future research.
- The method assumes the instruments are valid; while it handles irrelevant instruments, it does not address invalid instruments.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.
Get the best statistics papers every week.
Trusted by researchers at Stanford, Cambridge, and the French Academy of Sciences.
Check your inbox to confirm your subscription.
Something went wrong. Try again?
No spam, unsubscribe anytime.