Deep Hedging with Reinforcement Learning: A Practical Framework for Option Risk Management

Log in to collect

Onsite backtest IDE

Quant Buffet native backtest IDE

Edit and run Quant Buffet Python for Deep Hedging with Reinforcement Learning: A Practical Framework for Option Risk Management in the browser. Results update live with equity, drawdown, and metrics charts. Allowed: backtest.data, backtest.engine, backtest.metrics, numpy, pandas. Define ASSETS and make_on_day(prices). Shortcut: Ctrl+Enter. API docs →

Ready — edit code, then Run backtest.
IDE · 43 lines
Quant Buffet syntax cheat sheet (copy / insert)

Paste these fragments into the editor. The sandbox rejects QuantConnect, os, and network libraries.

Required imports
Only these libraries are allowed in the sandbox.
from __future__ import annotations

import numpy as np
import pandas as pd

from backtest.data import load_daily_prices
from backtest.engine import EngineConfig, PortfolioEngine
from backtest.metrics import compute_metrics
ASSETS list (whitelisted ETFs)
Module-level list. Tickers must be in the Quant Buffet whitelist.
ASSETS = ["SPY", "QQQ", "TLT", "GLD", "BIL"]
make_on_day contract
Must return (on_day, ready). on_day calls engine.set_target_weights.
def make_on_day(prices: pd.DataFrame):
    cols = [c for c in ASSETS if c in prices.columns]
    sma = prices[cols].rolling(200, min_periods=200).mean()
    state = {"last": None}

    def on_day(engine: PortfolioEngine, dt: pd.Timestamp) -> None:
        if sma.loc[dt].isna().all():
            return
        key = (dt.year, dt.month)
        if state["last"] == key:
            return
        state["last"] = key
        long = [
            s for s in cols
            if pd.notna(prices.at[dt, s]) and pd.notna(sma.at[dt, s])
            and prices.at[dt, s] > sma.at[dt, s]
        ]
        weights = {} if not long else {s: 1.0 / len(long) for s in long}
        engine.set_target_weights(dt, weights)

    ready = sma.dropna(how="all").index.min() if sma.notna().any().any() else None
    return on_day, ready
Set target weights
Weights should sum to about 1.0. Empty dict = 100% cash.
engine.set_target_weights(dt, {"SPY": 0.60, "BIL": 0.40})

Live backtest performance

CAGR
4.46%
Sharpe
0.58
Max DD
-28.94%
Vol
8.01%
Sortino
0.90
Beta
0.38

Showing saved draft baseline until you re-run.

Equity curve (indexed = 100)

Accent = strategy · dashed grey = buy-and-hold benchmark

2000-042026-0880261
Drawdown
Worst -24.4%-24%
Metrics bar chart
CAGRSharpeSortinoVol|DD|Grey = baseline · Accent = live run
Monthly returns
2020-062026-08 · last 24 months

Export to your platform

Transform Quant Buffet lab code (ASSETS + make_on_day / PortfolioEngine) into native classes for a third-party IDE — then copy and paste.

Run in: QuantConnect Cloud or LEAN CLI · QCAlgorithm with Equity securities and monthly rebalance.

Detected pattern: Volatility targetingAssets: SPY, TLT, GLD, BIL
# Generated from Quant Buffet → QuantConnect LEAN
# Strategy: Deep Hedging with Reinforcement Learning: A Practical Framework for Option Risk Management
# Detected pattern: Volatility targeting
# Source uses Quant Buffet lab APIs (ASSETS + make_on_day / PortfolioEngine).
# Review fees, data, and risk before live trading — educational export only.

from AlgorithmImports import *


class QuantBuffetExport(QCAlgorithm):
    def Initialize(self):
        self.SetStartDate(2010, 1, 1)
        self.SetCash(100000)
        tickers = ["SPY", "TLT", "GLD", "BIL"]
        self.symbols = []
        for t in tickers:
            if "-" in t:  # crypto proxy e.g. BTC-USD
                self.symbols.append(self.AddCrypto(t.replace("-USD", ""), Resolution.Daily).Symbol)
            else:
                self.symbols.append(self.AddEquity(t, Resolution.Daily).Symbol)
        self.Schedule.On(
            self.DateRules.MonthStart(self.symbols[0]),
            self.TimeRules.AfterMarketOpen(self.symbols[0], 30),
            self.Rebalance,
        )
        # Logic: Scale weights to target vol 0.1 using 63-day vol.

    def Rebalance(self):
        # Pattern: vol_target — Scale weights to target vol 0.1 using 63-day vol.
        # Default: equal-weight. Port your make_on_day weights here via SetHoldings.
        w = 1.0 / len(self.symbols) if self.symbols else 0.0
        for symbol in self.symbols:
            self.SetHoldings(symbol, w)

Exported code uses the platform’s native classes and libraries. Install dependencies in your third-party IDE, then run. Validate before live trading.

Academic paper

Deep Hedging with Reinforcement Learning: A Practical Framework for Option Risk Management

AuthorsTravon Lucius; Christian Koch; Jacob Starling; Julia Zhu; Miguel Urena; Carrie Hu

Teaser

Scale exposure inversely to realized volatility toward a target annual vol. Universe: SPY, BIL. Parameters: target_vol=0.1; vol_lookback=63; rebalance=monthly. Rebalanced on the engine's template schedule with 5 bps commission and 2 bps slippage.

Strategy in a nutshell

We present a reinforcement-learning (RL) framework for dynamic hedging of equity index option exposures under realistic transaction costs and position limits. We hedge a normalized option-implied equity exposure (one unit of underlying delta, offset via SPY) by trading the underlying index ETF, using the option surface and macro variables only as state information and not as a direct pricing engine. Building on the "deep hedging" paradigm of Buehler et al. (2019), we design a leak-free environment, a cost-aware reward function, and a lightweight stochastic actor-critic agent trained on daily end-of-day panel data constructed from SPX/SPY implied volatility term structure, skew, realized volatility, and macro rate context. On a fixed train/validation/test split, the learned policy improves

Economic rationale

Volatility is more forecastable than expected returns; scaling down when vol is high improves risk-adjusted outcomes (volatility-managed portfolios). Related evidence from “Deep Hedging with Reinforcement Learning: A Practical Framework for Option Risk Management”: We present a reinforcement-learning (RL) framework for dynamic hedging of equity index option exposures under realistic transaction costs and position limits. We hedge a normalized option-implied equity exposure (one unit of underlying delta, offset via SPY) by trading the underlying index ETF, using the option surface and macro variables only as state information and not as a direct pricing engine. Building on the "deep hedging" paradigm of Buehler et al. (2019), we design a leak-free environme

Backtest performance

Annualised return4.46%
Volatility8.01%
Beta0.38
Sharpe ratio0.58
Sortino ratio0.90
Maximum drawdown-28.94%