Skip to main content

Topic

Quantitative & Academic Research

For economists, statisticians, financial researchers, and social science researchers.

Anscombe's Quartet: Four Wildly Different Datasets That Share the Exact Same Statistics

Four datasets sharing the identical correlation coefficient can include one showing a genuine linear relationship, one showing a clear curve, one dominated by a single outlier, and one that's barely a relationship at all.

Backtested Strategies Die in Live Trading Because of Overfitting, Not Bad Luck

The gap between an impressive backtest and disappointing live performance is rarely explained by changed market conditions — it's usually explained by the backtest fitting noise in the first place.

Berkson's Paradox: How Selecting Your Sample Can Manufacture a Correlation That Isn't Really There

A hospital-based study can find a negative correlation between two unrelated diseases purely because of how patients ended up admitted to the hospital in the first place — the correlation is an artifact of selection, not biology.

Difference-in-Differences: A Clever Way to Estimate Causal Impact Without a Randomized Trial

The core idea behind difference-in-differences is elegant and simple. The assumption that makes it valid — that both groups would have moved in parallel absent the intervention — is genuinely hard to verify.

Endogeneity: The Hidden Variable Problem That Makes a Clean-Looking Regression Coefficient Misleading

An omitted variable that influences both your independent and dependent variable will bias a regression coefficient in a specific, predictable direction, no matter how much data you throw at the model.

Ergodicity: Why an Attractive Average Outcome Can Still Be a Terrible Bet for Any Individual

The expected value of a repeated bet, calculated across many independent players at one moment, can diverge sharply from what actually happens to one specific player repeating that same bet many times over.

Goodhart's Law: Why a Metric Stops Being a Good Measure the Moment It Becomes a Target

An economic indicator that worked well as a passive measure often degrades the moment policy or incentives start targeting it directly — the measure and the reality it tracked quietly come apart.

Heteroskedasticity: Why Your Regression's Error Might Be Bigger for Some Predictions Than Others

A sales forecasting model that's reliably accurate for small accounts and wildly variable for large ones has a heteroskedasticity problem, and the standard confidence intervals reported for it are likely wrong.

Instrumental Variables: How Economists Estimate Causation Without a Randomized Experiment

An instrumental variable doesn't need to be interesting on its own — it just needs to move the variable you care about, without having any other plausible path to affecting the outcome you're studying.

Look-Ahead Bias: When a Backtest Accidentally Uses Information That Wasn't Actually Available Yet

Testing a trading strategy against a company's final, restated earnings figures rather than the originally reported figures available at the time quietly gives the backtest information no real trader actually had.

Multicollinearity: Why Adding a Correlated Variable Can Make a Regression's Coefficients Swing Wildly

A regression coefficient that flips sign or changes dramatically when a similar variable is added or removed isn't measuring something unstable in reality — it's a symptom of multicollinearity in the model.

Multiple Comparisons: The Silent Way Dashboards Manufacture False Positives

A dashboard that lets someone slice data twenty different ways and highlights whatever's significant is a false-positive generator by construction, not by accident.

P-Hacking: How Small, Reasonable-Looking Analytical Choices Add Up to Manufactured Significance

A researcher doesn't need to fabricate data to produce a false positive — trying a handful of defensible analytical variations and reporting only the one that reached significance is often enough on its own.

Publication Bias: Why the Published Research Record Systematically Overstates How Well Things Work

If ten labs run the same study and only the two that find a significant effect get published, the visible literature says something quite different than the full, unpublished body of evidence would.

Reading a p-value Without Fooling Yourself

A p-value answers one narrow question. Most misreadings come from quietly asking it a different one.

Regression Discontinuity Design: Turning an Arbitrary Cutoff Into a Natural Experiment

A scholarship awarded strictly to students scoring above a specific test score threshold creates, right around that threshold, something close to a naturally occurring randomized experiment.

Regression to the Mean in Business: Why Your Worst-Performing Salesperson Probably Improves Next Quarter Regardless

A struggling salesperson placed on a performance improvement plan often does improve the following quarter — some of that improvement would likely have happened anyway, purely from regression to the mean.

Simpson's Paradox: When a Trend Reverses Completely After You Add One More Variable

Simpson's Paradox isn't a data error or a coding mistake — it's a genuine, well-documented statistical phenomenon where aggregate and subgroup trends can point in opposite directions, both validly.

Spurious Regression: Why Two Unrelated Time Series Can Look Powerfully Correlated

A regression between two trending time series can return an impressively high R-squared and statistically significant coefficients while describing a relationship that doesn't exist in any meaningful sense.

Survivorship Bias in Backtested Returns: Why a Historical Index's Track Record Looks Better Than Investors Actually Experienced

A backtest constructed using today's list of successful, surviving companies and projected backward tells a story that's quietly missing every company that failed along the way and was never actually a permanent part of the real historical index.

The Availability Heuristic: Why Vivid, Memorable Risks Get Overweighted and Boring, Common Ones Get Ignored

A rare but vivid, well-publicized business risk often receives more organizational attention and planning than a genuinely more likely but less memorable one — a documented pattern in how people actually estimate probability.

The Base Rate Fallacy: Why a Positive Test Result Is Less Alarming Than It Feels

A test that's 95% accurate sounds reassuring, and for a genuinely rare condition, a positive result from that same test can still be wrong more often than it's right, once the base rate is properly factored in.

The Bootstrap Method: How Statisticians Estimate Uncertainty Without Assuming a Textbook Distribution

Traditional confidence interval formulas often assume data follows a specific theoretical distribution — the bootstrap method sidesteps that assumption entirely by resampling the actual data itself, repeatedly, to see how much an estimate naturally varies.

The Ecological Fallacy: Why a True Statement About a Group Can Be False About Every Individual In It

Aggregate, group-level correlations are a genuinely different kind of evidence than individual-level correlations, and treating one as a stand-in for the other is a specific, well-documented statistical error.

The Garden of Forking Paths: How P-Hacking Happens Without Anyone Deliberately Cheating

P-hacking doesn't require deliberate manipulation — a researcher making a long sequence of individually defensible analytical choices can still end up capitalizing on chance simply because so many alternative reasonable choices existed.

The Look-Elsewhere Effect: Why Scanning a Huge Dataset for Patterns Reliably Finds Some, Whether or Not They're Real

A dataset scanned for any statistically significant pattern across thousands of possible combinations will produce apparent discoveries at a rate that has little to do with whether any real underlying pattern actually exists.

The Reproducibility Crisis: Why a Meaningful Share of Published Findings Don't Replicate

Large, well-run replication projects across psychology, medicine, and economics have consistently found that a meaningful share of published, peer-reviewed findings don't hold up when independently retested.

The Weak Instruments Problem: When Your Instrumental Variable Isn't Actually Doing Its Job

Instrumental variables analysis only works as well as its instrument's actual strength — a weak instrument can produce an estimate less trustworthy than the simpler, more direct regression it was meant to improve upon.

Time-Varying Confounders: The Specific Case Where Even Controlling for a Confounder Doesn't Fix a Regression

The standard fix for confounding — add the confounder as a control variable — breaks down in a specific, well-documented case where the confounder is itself both a consequence of earlier treatment and a cause of later treatment.

WEIRD Samples: Why Findings from Psychology's Most-Studied Population Don't Generalize

A finding replicated repeatedly among undergraduate psychology students at Western universities is a genuinely well-established finding about a genuinely narrow population, and generalizing it further requires evidence, not assumption.

Why a Quant Model Trained on One Market Regime Can Fail Badly in Another

A trading model validated entirely on data from a single sustained low-volatility bull market carries a genuine, specific risk of failing once market conditions shift to a fundamentally different regime.

Why Checking Twenty Metrics on One A/B Test Makes a False Positive Nearly Inevitable

Testing twenty independent metrics at the standard 5 percent significance threshold means a genuinely null experiment has roughly a two-thirds chance of showing at least one falsely significant result somewhere on the dashboard.

Why Most "Elasticity" Estimates in Business Reports Are Wrong

A single elasticity number presented without its estimation conditions is not a fact about your product — it's a fact about the specific price range and time window it was measured in.

Why the Sharpe Ratio Can Flatter Strategies With a Specific, Dangerous Kind of Risk

A strategy generating frequent small gains punctuated by rare but severe losses can show an attractive Sharpe ratio while carrying a genuinely dangerous risk profile the metric's underlying assumptions don't fully capture.

Winsorizing: The Compromise Between Deleting Outliers and Ignoring Them Entirely

Winsorizing replaces extreme values with a less extreme value at a chosen percentile, rather than deleting them outright, preserving the fact that something unusual happened without letting that one data point dominate the entire analysis.