Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Checklist of empirical robustness tests for finance/economics papers
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-21 | ✓→✓ | = Same ✓ | 42% | 0% |
| case-22 | ✓→✓ | = Same ✓ | 28% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 70% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 24% | 0% |
Systematic checklist of robustness tests for empirical research. Use this to ensure comprehensive testing before submission.
| Test | Description | When to Use | |------|-------------|-------------| | Exclude outliers | Winsorize/trim at different levels (0.5%, 2%, 5%) | Always | | Drop financial firms | Exclude SIC 6000-6999 | If not already excluded | | Drop regulated industries | Exclude utilities, telecoms | Industry-specific effects | | Different time periods | Split sample pre/post crisis, early/late | Results may be period-specific | | Geographic subsamples | By region, state, country | External validity | | Size subsamples | Small vs. large firms | Heterogeneous effects | | Balanced panel | Require continuous observations | Survivorship concerns |
| Test | Description | When to Use | |------|-------------|-------------| | Different fixed effects | Firm, industry×year, state×year | Control for unobservables | | Additional controls | Add variables referees might suggest | Omitted variable concerns | | Drop controls | Verify not over-controlling | Mediator concerns | | Different clustering | Firm, industry, state, two-way | Inference robustness | | Different standard errors | Bootstrap, Newey-West, Driscoll-Kraay | Serial/cross-sectional correlation | | Nonlinear specifications | Quadratic terms, splines | Linearity assumption | | Log vs. level | Transform dependent variable | Skewed distributions |
| Test | Description | When to Use | |------|-------------|-------------| | Alternative dependent variable | Different proxy for same concept | Measurement concerns | | Alternative treatment measure | Continuous vs. binary, different threshold | Treatment definition | | Alternative control measures | Different proxies for size, leverage, etc. | Standard practice | | Scaled differently | By assets, sales, employees | Scaling choice matters |
| Test | Description | When to Use | |------|-------------|-------------| | Placebo/Falsification | | | | Placebo timing | Fake treatment 1-3 years before actual | DiD parallel trends | | Placebo outcome | Effect on outcome that shouldn't be affected | Specificity of mechanism | | Placebo treatment | Random assignment of treatment | Rule out spurious correlation | | Pre-trends | | | | Event study plot | Coefficient for each pre/post period | Visual parallel trends | | Joint F-test | Test pre-period coefficients = 0 | Statistical parallel trends | | Endogeneity | | | | Instrumental variables | Find exogenous variation | Selection concerns | | Heckman selection | Model selection explicitly | Sample selection | | Propensity score matching | Match treated/control | Observable selection | | Entropy balancing | Reweight to balance covariates | Covariate imbalance | | Regression discontinuity | If threshold exists | Sharp identification |
| Test | Description | When to Use | |------|-------------|-------------| | Wild cluster bootstrap | Small number of clusters | <50 clusters | | Randomization inference | Permutation-based p-values | Few treated units | | Conley standard errors | Spatial correlation | Geographic data | | Multiple hypothesis correction | Bonferroni, FDR | Many outcomes tested |
For difference-in-differences designs:
For instrumental variables:
At minimum, most papers should include:
For robustness tables:
Table X: Robustness Tests
Panel A: Alternative Samples
(1) Baseline
(2) Exclude financial firms
(3) Exclude 2008-2009
(4) Winsorize at 5%
Panel B: Alternative Specifications
(5) Add industry×year FE
(6) Control for firm age
(7) Cluster by industry
Panel C: Alternative Measures
(8) Alternative dependent variable
(9) Continuous treatment measureOther measured skills in the registry, with their headline benchmark lift.