Expected Options Return Predictability Using Machine Learning
Log in to collectAcademic paper
Option Return Predictability with Machine Learning and Big Data
Turan G. Bali; Heiner Beckmeyer; Mathis Moerke; Florian Weigert
- Georgetown University
- ?Georgetown University - McDonough School of Business
- DEUniversity of Münster
- ?University of Münster - Finance Center Muenster
- CHSwiss Finance Institute
- CHUniversity of St. Gallen
- ?University of St. Gallen - School of Finance
- ?University of St. Gallen - Swiss Institute of Banking and Finance
- CHUniversity of Neuchâtel
- DEUniversity of Cologne
- ?University of Cologne - Centre for Financial Research (CFR)
- ?University of Neuchatel - Institute of Financial Analysis
Strategy in a nutshell
The strategy targets all U.S. optionable equities (share codes 10 and 11) from NYSE, AMEX, and NASDAQ, using options and underlying stock data from IvyDB, CRSP, and Compustat (1996–2020). After filtering out incomplete options and dividend-affected stocks, a rich feature set—including liquidity, value, and others—is constructed. Five nonlinear models (Random Forest, Gradient Boosted Trees, GBT with dropout, and Feed-Forward Nets) are trained on rolling 5-year windows, validated for hyperparameter tuning, and tested annually. Predicted excess returns are sorted into deciles, forming equally weighted long-short portfolios that buy options in the top decile and sell those in the bottom decile, fully funded and rebalanced monthly.
Economic rationale
Feature importance analysis using Shapley Additive Explanations shows the primary driver is an option’s position on the implied volatility surface, reflecting a relative value factor. Volatility and liquidity risk premia are secondary, with momentum, quality, and informed trading contributing. The ensemble captures nonlinear mispricing signals, enabling the exploitation of complex relationships between options and underlying stocks to generate systematic alpha.