How to Use Lexical Density of Company Filings

Log in to collect

Onsite backtest IDE

Quant Buffet native backtest IDE

Edit and run Quant Buffet Python for How to Use Lexical Density of Company Filings in the browser. Results update live with equity, drawdown, and metrics charts. Allowed: backtest.data, backtest.engine, backtest.metrics, numpy, pandas. Define ASSETS and make_on_day(prices). Shortcut: Ctrl+Enter. API docs →

Ready — edit code, then Run backtest.
IDE · 50 lines
Quant Buffet syntax cheat sheet (copy / insert)

Paste these fragments into the editor. The sandbox rejects QuantConnect, os, and network libraries.

Required imports
Only these libraries are allowed in the sandbox.
from __future__ import annotations

import numpy as np
import pandas as pd

from backtest.data import load_daily_prices
from backtest.engine import EngineConfig, PortfolioEngine
from backtest.metrics import compute_metrics
ASSETS list (whitelisted ETFs)
Module-level list. Tickers must be in the Quant Buffet whitelist.
ASSETS = ["SPY", "QQQ", "TLT", "GLD", "BIL"]
make_on_day contract
Must return (on_day, ready). on_day calls engine.set_target_weights.
def make_on_day(prices: pd.DataFrame):
    cols = [c for c in ASSETS if c in prices.columns]
    sma = prices[cols].rolling(200, min_periods=200).mean()
    state = {"last": None}

    def on_day(engine: PortfolioEngine, dt: pd.Timestamp) -> None:
        if sma.loc[dt].isna().all():
            return
        key = (dt.year, dt.month)
        if state["last"] == key:
            return
        state["last"] = key
        long = [
            s for s in cols
            if pd.notna(prices.at[dt, s]) and pd.notna(sma.at[dt, s])
            and prices.at[dt, s] > sma.at[dt, s]
        ]
        weights = {} if not long else {s: 1.0 / len(long) for s in long}
        engine.set_target_weights(dt, weights)

    ready = sma.dropna(how="all").index.min() if sma.notna().any().any() else None
    return on_day, ready
Set target weights
Weights should sum to about 1.0. Empty dict = 100% cash.
engine.set_target_weights(dt, {"SPY": 0.60, "BIL": 0.40})

Live backtest performance

CAGR
7.89%
Sharpe
0.63
Max DD
-33.72%
Vol
13.60%
Sortino
0.93
Beta
0.51
Up days
52%

Run the backtest to populate charts.

Export to your platform

Transform Quant Buffet lab code (ASSETS + make_on_day / PortfolioEngine) into native classes for a third-party IDE — then copy and paste.

Run in: QuantConnect Cloud or LEAN CLI · QCAlgorithm with Equity securities and monthly rebalance.

Detected pattern: Absolute momentumAssets: SPY, TLT, GLD, BIL
# Generated from Quant Buffet → QuantConnect LEAN
# Strategy: How to Use Lexical Density of Company Filings
# Detected pattern: Absolute momentum
# Source uses Quant Buffet lab APIs (ASSETS + make_on_day / PortfolioEngine).
# Review fees, data, and risk before live trading — educational export only.

from AlgorithmImports import *


class QuantBuffetExport(QCAlgorithm):
    def Initialize(self):
        self.SetStartDate(2010, 1, 1)
        self.SetCash(100000)
        tickers = ["SPY", "TLT", "GLD", "BIL"]
        self.symbols = []
        for t in tickers:
            if "-" in t:  # crypto proxy e.g. BTC-USD
                self.symbols.append(self.AddCrypto(t.replace("-USD", ""), Resolution.Daily).Symbol)
            else:
                self.symbols.append(self.AddEquity(t, Resolution.Daily).Symbol)
        self.Schedule.On(
            self.DateRules.MonthStart(self.symbols[0]),
            self.TimeRules.AfterMarketOpen(self.symbols[0], 30),
            self.Rebalance,
        )
        # Logic: Long assets with positive 252-day return; equal-weight; monthly.

    def Rebalance(self):
        # Pattern: abs_momentum — Long assets with positive 252-day return; equal-weight; monthly.
        # Default: equal-weight. Port your make_on_day weights here via SetHoldings.
        w = 1.0 / len(self.symbols) if self.symbols else 0.0
        for symbol in self.symbols:
            self.SetHoldings(symbol, w)

Exported code uses the platform’s native classes and libraries. Install dependencies in your third-party IDE, then run. Validate before live trading.

Academic paper

How to Use Lexical Density of Company Filings

AuthorsDaniela Hanicova; Filip Kalús; Radovan Vojtko

Institute
  • ?Quantpedia
  • ?Quantpedia.com

Screenshot from the original paper

Screenshot from the original paper

Strategy in a nutshell

The investment universe consists of top 500 US stocks by dollar volume. The stocks are sorted based on their lexical density and specific density score from the BLMCF dataset. Lexical density measures the structure and complexity of human communication in a text. A high lexical density indicates a large amount of information-carrying words. Specific density measures how dense the report’s language is from a financial point of view. In other words, how many finance- related words are used in the text. The investor goes long the top decile and short the bottom decile. Additionally, the portfolio is rebalanced on a monthly basis.

Economic rationale

The combination of the high and increasing volume of published 10-K & 10-Q reports and their gradual shift to nonnumerical information leads to the premise that fundamental analysts cannot identify crucial information in the “white noise” about the actual and future performance of the company. The companies like BRAIN, which analyze the 10-K& 10-Q reports using NLP and give scores according to numerous language metrics, bridge the gap between the nonnumerical and numerical data. The research suggests that the richer the vocabulary of an investor is, the higher the lexical score the company gets and the better it performs.

Backtest performance

Annualised return7.89%
Volatility13.60%
Beta0.51
Sharpe ratio0.63
Sortino ratio0.93
Maximum drawdown-33.72%
Win rate52%