Machine Learning in News Articles Predicts Stock Returns
Log in to collectAcademic paper
Strategy in a nutshell
This strategy constructs a daily long-short portfolio using news-based sentiment signals from 6.3 million articles spanning 1989–2020. Text data are cleaned, normalized, and encoded using a bag-of-words model. A supervised screening identifies words correlated with positive returns, followed by a two-topic model to estimate word probabilities for positive and negative sentiment. Out-of-sample sentiment is estimated via penalized regression with a Beta prior. Each day, the strategy goes long the 50 most positively scored words and shorts the 50 most negatively scored words, forming a zero-net portfolio with a 30-minute delay at market open.
Economic rationale
The sentiment model demonstrates predictive power, particularly for smaller stocks, which incorporate sentiment information more slowly than large firms. Risk attribution shows minimal correlation with standard Fama-French factors, indicating genuine alpha generation. By leveraging news-derived sentiment, the strategy captures mispricings and market inefficiencies overlooked by conventional factors, making it a robust approach to extracting returns from textual data.