\newcommand{\E}{\mathbb E} \newcommand{\R}{\mathbb R} \newcommand{\N}{\mathcal N} \newcommand{\Pp}{\mathcal P} \newcommand{\Rcal}{\mathcal R} \newcommand{\GM}{\mathrm{GM}} \newcommand{\AW}{\mathcal{AW}} \newcommand{\Law}{\operatorname{Law}} \newcommand{\Cov}{\operatorname{Cov}} \newcommand{\Var}{\operatorname{Var}} \def\@thanks{\protect\footnotetext[0]{Working paper. Comments welcome.}}
spectra Research

Working paper · September 2026

Finite Gaussian-mixture martingale kernels: density for projected smile-jet readouts

Zeyu Cao1 and Shaosai Huang2

1 Independent Researcher, Long Island City, USA · [email protected]

2 Kspectra Research Inc., Toronto, Canada · [email protected]

Abstract

The skew-stickiness ratio and its higher-order extensions measure how the implied-volatility smile responds to moves in the spot. They are what a quoted surface reveals about the martingale kernel generating it, a kernel that is never observed. We ask whether finite Gaussian-mixture martingale chains, whose option prices are finite sums of Black–Scholes prices, can reproduce those dynamics. They can. For stochastic-volatility models with a variance floor and correlation bounded away from ±1\pm 1, and for local-stochastic-volatility models, under explicit regularity assumptions, we construct exact-martingale finite Gaussian-mixture chains that approximate the marginals and match the finite-step skew-stickiness ratio and its extensions to any prescribed order and accuracy. Part of the contribution is the notion of convergence: the topology generated by the dynamics characteristics themselves, which neither weak nor adapted Wasserstein convergence controls. With MM components per node and a dd-dimensional latent state the readout error is O(M1/(d+1))O(M^{-1/(d+1)}) up to logarithms, linear in the quantization error, when returns carry a Gaussian component of fixed variance; without one, adapted approximation forces the innovations preceding a price-dependent kernel to vanish, and our bound degrades with the readout order. The limit is the wings: every finite mixture has zero implied-variance wing slopes. Adopting the class therefore costs nothing in at-the-money smile dynamics of any finite order; on an SPX surface, fixing the first three characteristics narrows the at-the-money forward-start range left open by the vanilla quotes from 7.07.0 to 4.04.0 volatility points, and across fifty-five further monthly surfaces it removes a median 39%39\% of that range.

Contents
  1. 1Introduction
  2. 2The jet readout and the topology it generates
    1. 2.1The transport construction and what we use
    2. 2.2Conditional smiles and static jets
    3. 2.3Finite-step projection
    4. 2.4Triangular inversion and the pivot
    5. 2.5The topology generated by the readout
    6. 2.6Moving marginals and exact martingality
  3. 3Model class and Gaussian factorization
  4. 4Log-barycentric quantization
  5. 5Finite marked trees and adapted convergence
  6. 6Continuation stability and Gaussian smoothing
  7. 7Projected dynamics and the density theorem
  8. 8A convergence rate
  9. 9Local stochastic volatility
    1. 9.1Local-SV kernels and readout regularity
    2. 9.2Finite proxy-GM chains
    3. 9.3Diagonal quantization and lifted adapted convergence
    4. 9.4Continuation densities and smile jets
    5. 9.5The second density theorem
  10. 10Closed-form readout, cubature and numerics
    1. 10.1Two checks of the rates
    2. 10.2A constrained band on an SPX surface
  11. 11Scope and extensions
  12. 12Conclusion
  13. References
  14. How to cite

Keywords: smile dynamics; skew-stickiness ratio; stochastic-local volatility; Gaussian mixtures; martingale transition kernels; adapted Wasserstein distance; arbitrage-free calibration.

JEL classification: G13; C65; C63; G12.

MSC 2020: 91G20 (primary); 60G42, 60J05, 60B10, 49Q22, 91G60 (secondary).

1Introduction

An arbitrage-free European surface fixes the risk-neutral marginal law of the underlying at each maturity, but not the martingale kernels joining those laws; kernels with identical marginals produce different continuation smiles, spot–volatility responses and path-dependent prices. For independent-increment models the continuation smile is even deterministic in advance, however well the marginals fit, which is one form of the surface-versus-dynamics problem [19, Proposition 11.2 and Section 14.3]. The kernel generating those dynamics is never observed. What a surface does reveal is a hierarchy of at-the-money characteristics — the skew-stickiness ratio and its higher-order extensions — so a modelling class is adequate only if no admissible kernel displays dynamics the class cannot express in those coordinates. We ask whether finite Gaussian-mixture martingale chains meet that test: can they reproduce a prescribed finite order of at-the-money smile dynamics to arbitrary accuracy while approximating the marginals and remaining exact asset martingales? They can, at every finite order, for homogeneous stochastic volatility (SV) and for local stochastic volatility (LSV) (Theorems 7.3 and 9.13).

Why martingale kernels and Gaussian mixtures. The kernel is the natural object for calibrating dynamics: it carries exactly the freedom European prices leave, lives on the grid of quoted dates, prices every maturity and path-dependent claim from one law, and combines local and stochastic volatility; as an exact martingale chain it excludes static and dynamic arbitrage by construction. Finite Gaussian mixtures make such kernels tractable: calls and their strike derivatives are sums of Black–Scholes terms; martingality is one exponential-barycentre equation per node, linear in the weights, so price fitting with fixed components is a linear or quadratic program; positive variances give smooth densities and stable smile derivatives; and chains of mixture kernels have mixture marginals. Mixtures are an approximation basis, not an economic mechanism, with a long history in smile fitting [13] and a discrete stochastic-local-volatility implementation in [29]. The costs are thin tails and component counts that grow with accuracy and order.

The observable dynamics. What the kernel settles beyond today’s prices is how the smile moves when the spot moves, and calibrating that motion needs a quantity the quotes show; the skew-stickiness ratio [1011] is the one desks use, and it is the lowest order of a family. Che and Das [18] grade the response of the smile to spot moves by a transport equation whose at-the-money coefficients extend the ratio to every order; we take that family in finite-step form, computed from the law of a chain on the grid rather than from a pathwise derivative. More specifically, on a grid t0<<tJt_0<\cdots <t_J with log price XjX_j and continuation state ζj\zeta _j, let wj,(k,ζ)w_{j,\ell }(k,\zeta ) be the total implied variance at log-moneyness kk of the maturity-tt_\ell smile seen from ζj=ζ\zeta _j=\zeta. The readout consists of at-the-money jets and their projected spot responses over the first step,

(1)am=kmw0,(0,ζ0),Am=kmw1,(0,ζ1),dm=Cov(Am,X1X0)Var(X1X0),\begin{equation}\label {eq:intro-readout} a_m^\ell =\partial _k^m w_{0,\ell }(0,\zeta _0),\qquad A_m^\ell =\partial _k^m w_{1,\ell }(0,\zeta _1),\qquad d_m^\ell =\frac {\Cov (A_m^\ell ,X_1-X_0)}{\Var (X_1-X_0)}, \end{equation}

where mm is the order in log-moneyness and \ell the maturity, collected over a finite maturity set L\mathcal L and truncated at an order NN in RN,Lraw=((am)mN+1,(dm)mN)L\Rcal ^{\rm raw}_{N,\mathcal L}=\bigl ((a_m^\ell )_{m\le N+1},(d_m^\ell )_{m\le N}\bigr )_{\ell \in \mathcal L}. At order zero, a0a_0^\ell and a1a_1^\ell are the level and the skew, and v0=d0/a1v_0^\ell =d_0^\ell /a_1^\ell is the finite-step analogue of that ratio; on a10a_1^\ell \ne 0 the rows determine the transport coefficients v0,,vNv_0^\ell ,\ldots ,v_N^\ell of [18] by triangular inversion (Section 2.4). The higher orders are not decoration: the same study rejects v1=v2=0v_1=v_2=0 at every tenor from one month to twenty-four, on five years of SPX surfaces.

The approximation problem. Quantifying over unknown kernels is not possible, so we quantify instead over the classes practitioners posit. Given a target chain PP in such a class, an order NN, a fixed p1p\ge 1 and ε>0\varepsilon >0, the problem is to find a finite Gaussian-mixture chain QQ with

(2)EQ[eXj+1Fj]=eXj(0j<J),max1jJWp(LawQ(Xj),LawP(Xj))<ε,|RN,Lraw(Q)RN,Lraw(P)|<ε.\begin{equation}\label {eq:problem} \begin {gathered} \E _Q\bigl [e^{X_{j+1}}\mid \mathcal F_j\bigr ]=e^{X_j}\quad (0\le j<J),\\ \max _{1\le j\le J}W_p\bigl (\Law _Q(X_j),\Law _P(X_j)\bigr )<\varepsilon , \qquad \bigl |\Rcal ^{\rm raw}_{N,\mathcal L}(Q)-\Rcal ^{\rm raw}_{N,\mathcal L}(P)\bigr |<\varepsilon . \end {gathered} \end{equation}

Density in the topology generated by (1) and the marginals (Section 2) is the statement that this is always possible; if it is, calibrating in finite mixtures costs nothing the readout can see.

The difficulty, and the topology that resolves it. Static density of finite Gaussian mixtures is classical [2], and quantizing a Markov kernel yields weak convergence of finite-dimensional laws; (2) reduces to neither, because its requirements pull against one another. Exact martingality represents every cell of a finite approximation by its exponential barycentre, leaving atoms where the readout asks for derivatives of a conditional density; and a finite chain cannot meet a generic marginal exactly (Proposition 2.5). What has to be settled first is therefore not how to approximate but what approximation should mean, and the ready-made answers do not serve. Weak or Wasserstein closeness of path laws ignores the flow of information, so models close in those topologies can call for very different hedges; the adapted Wasserstein distance repairs that, makes hedging Lipschitz stable [3], and controls the conditional laws the rows dmd_m^\ell depend on. It is the natural candidate but still not enough: finite Gaussian-mixture chains can converge to a Black–Scholes target adaptedly while their at-the-money smile curvature diverges. No Wasserstein distance on path laws, adapted or not, makes the readout continuous (Remark 2.4). The notion must therefore come from the readout itself, together with the marginals, which is what (2) asks for and what Section 2 constructs: martingality is kept exact, the marginals are allowed to move, and the readout is required to converge. Placing each demand where it can be met is the first contribution here, and to our knowledge this topology has not been isolated before.

Main results. Theorems 7.3 (SV) and 9.13 (LSV) give, for every target PP in the respective class, one sequence QnQ_n of exact-martingale finite Gaussian-mixture chains whose errors in (2) tend to zero for every NN simultaneously, with the velocity coefficients converging on a10a_1^\ell \ne 0; in the topologies of Section 2,

TSVGSVτmovSV,TLSVGliftτmovLSV.\mathfrak T_{\rm SV}\subseteq \overline {\mathfrak G_{\rm SV}}^{\,\tau _{\rm mov}^{\rm SV}}, \qquad \mathfrak T_{\rm LSV}\subseteq \overline {\mathfrak G_{\rm lift}}^{\,\tau _{\rm mov}^{\rm LSV}}.

Both classes assume a compact latent state, Lipschitz kernels with exponential moments, a0a_0^\ell in a compact subset of (0,)(0,\infty ) and Var(X1X0)>0\Var (X_1-X_0)>0. The SV class adds the Gaussian factor of Assumption 3.1, supplied by a variance floor with correlation bounded away from ±1\pm 1 (Proposition 3.3), and contains capped Heston (Proposition 3.4); the LSV class adds one-step return densities with bounded, uniformly continuous derivatives, of which order NN uses the first max{N1,0}\max \{N-1,0\} (Remark 9.15). Theorem 8.3 makes the SV statement quantitative: with MM components per node and a dd-dimensional latent state the readout error is O(M1/(d+1))O(M^{-1/(d+1)}) up to logarithms, only that dimension entering the exponent, and the quantization exponent is optimal (Proposition 8.4). An LSV target supplies no factor, so the construction adds Gaussian innovations of variance βn0\beta _n\downarrow 0. Adapted approximation forces innovations preceding a price-dependent kernel to vanish (Proposition 9.18), and when that dependence is visible the construction’s adapted error is of exact order βn\sqrt {\beta _n} (Proposition 9.19); a final-edge factor may stay (Proposition 9.20). Two results mark the limits of the class: finite mixtures have zero implied-variance wing slopes [33], so density fails for a target with a finite moment index once the wing slope is added to the readout (Proposition 7.4); and with the marginals held fixed rather than approximated, the finite class caps the dynamics (Proposition 7.5). On an SPX surface, the vanilla quotes leave the at-the-money forward-start volatility undetermined by 7.07.0 points, and matching the first three readout coordinates as well narrows this to 4.04.0, both ranges exact; over fifty-five further monthly surfaces the three coordinates remove a median 39%39\% of that range (Section 10.2).

Method. Both proofs quantize each transition by exponential barycentres logE[eLcell]\log \E [e^L\mid \text {cell}] and couple target and approximant bicausally, which carries conditional laws along the sequence. Three obstacles follow. The martingale-preserving representative is not the optimal one: quantization rates are stated for L2L^2 centroids, which violate martingality, and martingale-preserving quantization is a problem in its own right [30], so the estimates are rebuilt around exponential barycentres (Lemmas 4.1 and 8.1). The labels move: the cells chosen at one date create the label set of the next, so the continuation jets must converge uniformly over sets that change along the sequence, and errors compound through the recursion (25). And the matched quantity is a jet of implied variance, recovered by an inversion whose denominator, vega, must stay away from zero along the sequence (Assumption 6.3). For an LSV target, density derivatives of order rr converge once the quantization error εn\varepsilon _n vanishes faster than βn(r+2)/2\beta _n^{(r+2)/2}, and a diagonal in rr serves every order at once; this is sufficient rather than necessary, since the readout itself needs only βnm/2\beta _n^{-m/2} at order mm (Corollary 9.17).

Scope and related work. The theorems are existence statements about risk-neutral capacity on a finite grid; fixed component budgets, calibration selection, instantaneous limits and physical-measure use are separate problems (Section 11). The readout meets two requirements that other gradings fail (Section 2.1): (A) order zero recovers the skew-stickiness ratio and higher orders refine it; (B) every coordinate is defined from the grid law, without a Brownian driver or perturbation parameter. Like [18], which confines transport to compact moneyness intervals and defers the wings to [33], the readout here is at-the-money by design. Model-side work computes such quantities for given models rather than asking which classes can match them: a representation formula for the skew-stickiness ratio [24], a second-order volatility-of-volatility expansion [12], and a quintic Ornstein–Uhlenbeck model tested against the skew-stickiness term structure [1]. Arbitrage-consistent frameworks for dynamic surfaces characterize admissible dynamics [34351617], and [20] derives the motion of the surface inside local volatility.

Organization. Section 2 defines the readout and the topology it generates, and Section 3 the SV class; Sections 47 prove Theorem 7.3 and Section 9 proves Theorem 9.13. Section 8 adds a convergence rate and Section 10 the closed-form readout, a cubature bound, two numerical checks of the rates and a constrained band on an SPX surface; Section 11 discusses scope.

AI-use disclosure. The authors used Anthropic Claude Code and OpenAI Codex as interactive research and writing tools. They assisted with exploratory discussion, testing and refinement of ideas, literature and source organization, code development and verification, mathematical error checking, and editorial revision.

2The jet readout and the topology it generates

Smile dynamics are compared in a finite-step form of the transport coordinates of Che and Das [18], computed from the law of a chain on the grid, and the topology is the one those coordinates generate. Static jets are read off conditional smiles by Black–Scholes inversion, dynamic coordinates are their finite-step projections on the spot increment, and velocity jets follow by triangular inversion.

2.1The transport construction and what we use

Che and Das [18] write the action of spot moves on total implied variance w(k,u,T)w(k,u,T), with kk the log-moneyness and uu the log forward, as a transport

(3)uw(k,u,T)=v(k,u,T;w)kw(k,u,T),\begin{equation}\label {eq:transport-intro} \partial _u w(k,u,T)=v(k,u,T;w)\,\partial _k w(k,u,T), \end{equation}

in which v0v\equiv 0, v1v\equiv 1 and vβv\equiv \beta are sticky delta, sticky strike and the skew-stickiness rule. With v~(k;u,T):=v(k,u,T;w(k,u,T))\tilde v(k;u,T):=v(k,u,T;w(k,u,T)) and vj:=kjv~(0;u,T)v_j:=\partial _k^j\tilde v(0;u,T), differentiating (3) at the money gives a triangular identity whose projected form is (7). The family is graded, each vnv_n being a residual after lower-order transport.

We use (7) only to define coordinates. As an identity, (3) holds with v=uw/kwv=\partial _uw/\partial _kw wherever kw0\partial _kw\ne 0 and degenerates where kw=0\partial _kw=0; one trajectory fixes the composite jets vjv_j, not the dependence of vv on ww [18, Remark 5.2]; the jets are coordinates rather than invariants (Remark 2.2); and they see only the spot-spanned response. We therefore replace the pathwise u\partial _u by the finite-step projection of Section 2.3 and state the theorems in the raw coordinates (am,dm)(a_m,d_m), which remain defined where the vjv_j are not. The theorems concern realizability: every projected jet profile generated by a target in the classes of Sections 3 and 9 is realized to arbitrary accuracy by exact-martingale finite chains. Attainability of an arbitrary empirical jet vector is not claimed, nor a characterization of the transport flows compatible with martingale dynamics. Other gradings fail requirement (A) or (B) of the introduction: Bergomi–Guyon functionals [12], Wiener-chaos order and martingale-expansion coefficients [23] need a forward-variance representation or a perturbation parameter, so they are not computable from the grid law of a finite mixture; marginal moments and cumulants are not nested descriptors of the at-the-money spot response; and the number of time points [27] is a complementary axis.

2.2Conditional smiles and static jets

For any chain considered below, let ζj\zeta _j denote a sufficient continuation state: a state variable that determines the conditional law of normalized future returns. For j<j<\ell and a state value ζ\zeta, define

(4)Cj,(k,ζ):=E[(eXXjek)+ζj=ζ].\begin{equation}\label {eq:conditional-call} C_{j,\ell }(k,\zeta ) :=\E \left [(e^{X_\ell -X_j}-e^k)^+\mid \zeta _j=\zeta \right ]. \end{equation}

Here and below the expectation denotes the Markov continuation started from ζ\zeta, fixing one continuation version on the state space. Let B(k,w)B(k,w) denote the normalized Black–Scholes call with log-strike kk and total variance ww,

B(k,w)=Φ(d+)ekΦ(d),d±=kw±w2.B(k,w)=\Phi (d_+)-e^k\Phi (d_-), \qquad d_\pm =-\frac {k}{\sqrt w}\pm \frac {\sqrt w}{2}.

The conditional implied total variance wj,(k,ζ)w_{j,\ell }(k,\zeta ) solves

B(k,wj,(k,ζ))=Cj,(k,ζ),B(k,w_{j,\ell }(k,\zeta ))=C_{j,\ell }(k,\zeta ),

and its ATM jets are

(5)amj,(ζ):=kmwj,(0,ζ).\begin{equation}\label {eq:static-jets} a_m^{j,\ell }(\zeta ):=\partial _k^m w_{j,\ell }(0,\zeta ). \end{equation}

For the homogeneous SV class of Section 3, ζj=Yj\zeta _j=Y_j, the latent state; for a local-SV target, ζj=(Xj,Yj)\zeta _j=(X_j,Y_j); for the lifted approximants of Section 9, ζj=(Ξjn,Y^jn)\zeta _j=(\Xi _j^n,\widehat Y_j^n), the finite proxy label on which their normalized continuation law depends. Quantities of the nnth approximant carry a subscript nn, as in am,nj,a_{m,n}^{j,\ell }.

Lemma 2.1 (Finite-dimensional static jet map). Fix r0r\ge 0 and suppose CCrC\in C^r near zero. On any set where the ATM solution w(0)w(0) stays in a compact subset of (0,)(0,\infty ), the vector (a0,,ar)(a_0,\ldots ,a_r) is a continuous function of the call derivatives (C(0),C(0),,C(r)(0))(C(0),C'(0),\ldots ,C^{(r)}(0)). The only implicit-function denominator is the Black–Scholes vega Bw(0,w(0))>0B_w(0,w(0))>0. If CCC\in C^\infty, then ww is smooth near zero.

Proof.Set F(k,w)=B(k,w)C(k)F(k,w)=B(k,w)-C(k). For r=0r=0, strict monotonicity of B(0,)B(0,\cdot ) gives continuity of the scalar inverse at the money. For r1r\ge 1, since Fw=Bw>0F_w=B_w>0, the CrC^r implicit-function theorem gives a CrC^r local solution w(k)w(k) (and a smooth one when CC is smooth). The first derivatives are

a1=FkFw,a2=Fkk+2Fkwa1+Fwwa12Fw.\begin{align*} a_1&=-\frac {F_k}{F_w},\\ a_2&=-\frac {F_{kk}+2F_{kw}a_1+F_{ww}a_1^2}{F_w}. \end{align*}

Repeated differentiation gives the same structure at every finite order: the numerator is a polynomial in lower jets and derivatives of FF, divided by FwF_w. On the stated compact set these operations are continuous.

2.3Finite-step projection

Fix the observation edge [t0,t1][t_0,t_1] and an option maturity tt_\ell with 2\ell \ge 2. At the known root define

am:=am0,(ζ0),Am:=am1,(ζ1).a_m^\ell :=a_m^{0,\ell }(\zeta _0), \qquad A_m^\ell :=a_m^{1,\ell }(\zeta _1).

On the domain Var(X1X0)>0\Var (X_1-X_0)>0, the projected spot derivative of the mmth smile jet is

(6)dm:=Cov(Amam,X1X0)Var(X1X0)=Cov(Am,X1X0)Var(X1X0).\begin{equation}\label {eq:projected-d} d_m^\ell :=\frac {\Cov (A_m^\ell -a_m^\ell ,X_1-X_0)}{\Var (X_1-X_0)} =\frac {\Cov (A_m^\ell ,X_1-X_0)}{\Var (X_1-X_0)}. \end{equation}

Since am0,a_m^{0,\ell } and am1,a_m^{1,\ell } refer to the same maturity one step apart, the deterministic part of the calendar decay drops out of the covariance, leaving a spot response. The denominator is positive: Var(X1X0)α0\Var (X_1-X_0)\ge \alpha _0 under Assumption 3.1, and for LSV targets and their proxy approximants see (45) and Lemma 9.12.

In regression terms, dmd_m^\ell is the population OLS slope of AmA_m^\ell on X1X0X_1-X_0, with residual vanishing exactly when AmA_m^\ell is affine in X1X0X_1-X_0. For a Gaussian-mixture edge with component label II, P(I=i)=pi\mathbb P(I=i)=p_i and R:=X1X0{I=i}N(mi,si2)R:=X_1-X_0\mid \{I=i\}\sim \N (m_i,s_i^2), one has Var(R)=ipi[si2+(mim¯)2]\Var (R)=\sum _ip_i[s_i^2+(m_i-\bar m)^2] with m¯=ipimi\bar m=\sum _ip_im_i, and Cov(A,R)=ipi(AiA¯)(mim¯)\Cov (A,R)=\sum _ip_i(A_i-\bar A)(m_i-\bar m) if A=AiA=A_i on {I=i}\{I=i\}; if AA varies within components, the law of total covariance adds ipiCov(A,RI=i)\sum _ip_i\Cov (A,R\mid I=i).

2.4Triangular inversion and the pivot

The projected velocity jets are defined algebraically by

(7)dn=j=0n(nj)vjanj+1.\begin{equation}\label {eq:jet-hierarchy} d_n^\ell =\sum _{j=0}^n\binom nj v_j^\ell a_{n-j+1}^\ell . \end{equation}

This is the pattern obtained by differentiating (3) at the money, with dnd_n^\ell in place of the pathwise uan\partial _ua_n; identifying the two would require the instantaneous limit of Section 11.

If a10a_1^\ell \ne 0, the system is triangular:

(8)vn=dnj=0n1(nj)vjanj+1a1.\begin{equation}\label {eq:standard-inversion} v_n^\ell =\frac {d_n^\ell -\sum _{j=0}^{n-1}\binom njv_j^\ell a_{n-j+1}^\ell }{a_1^\ell }. \end{equation}

More generally, if the hierarchy is compatible and a1==ar1=0ara_1^\ell =\cdots =a_{r-1}^\ell =0\ne a_r^\ell, the rows impose d0==dr2=0d_0^\ell =\cdots =d_{r-2}^\ell =0 and, for n0n\ge 0,

(9)vn=dr1+nj=0n1(r1+nj)vjar+nj(r1+nn)ar,\begin{equation}\label {eq:shifted-inversion} v_n^\ell =\frac { d_{r-1+n}^\ell - \sum _{j=0}^{n-1}\binom {r-1+n}{j}v_j^\ell a_{r+n-j}^\ell }{\binom {r-1+n}{n}a_r^\ell }, \end{equation}

so v0,,vNv_0,\ldots ,v_N need ama_m for mr+Nm\le r+N and dmd_m for mr1+Nm\le r-1+N, and near that stratum (9) is a continuous rr-pivoted extension wherever ar0a_r\ne 0. Approximation may perturb the vanishing jets, so the theorems use raw coordinates globally and velocity coefficients on the regular chart a10a_1\ne 0.

For finite L{2,,J}\mathcal L\subseteq \{2,\ldots ,J\} and NN0N\in \mathbb N_0, the raw readout RN,Lraw(P)\Rcal ^{\rm raw}_{N,\mathcal L}(P) collects am(P)a_m^\ell (P), mN+1m\le N+1, and dm(P)d_m^\ell (P), mNm\le N, over L\ell \in \mathcal L, as in (1). It recovers v0,,vNv_0^\ell ,\ldots ,v_N^\ell on the chart a10a_1^\ell \ne 0, where the velocity coefficients are continuous functions of finitely many raw coordinates, so adjoining them leaves the topology unchanged; raw convergence also gives convergence of any fixed rr-pivoted extension at a target with ar0a_r\ne 0.

Remark 2.2 (Transformation law: the velocity jets are chart coordinates). The vjv_j are jets of a vector field and transform as such. They are invariant under a change of variance level wλww\mapsto \lambda w, which is why they compare across volatility levels. Under a linear change of moneyness k~=ck\tilde k=ck, vjc1jvjv_j\mapsto c^{\,1-j}v_j, so v0v_0 scales like cc, v1v_1 is invariant and v2v_2 scales like c1c^{-1}; under k~=φ(k)\tilde k=\varphi (k) with φ(0)=0\varphi (0)=0, v0φ(0)v0v_0\mapsto \varphi '(0)v_0 and v1v1+(φ(0)/φ(0))v0v_1\mapsto v_1+(\varphi ''(0)/\varphi '(0))v_0. They are therefore coordinates relative to a parametrization: under the linear subgroup the invariants are the monomials of total weight zero, such as v1v_1 and v0v2v_0v_2, and quoting the smile in k/(σT)k/(\sigma \sqrt T) divides v0v_0 by σT\sigma \sqrt T. We fix k=log(K/F)k=\log (K/F); the raw coordinates (an,dn)(a_n,d_n) carry no convention beyond this choice.

2.5The topology generated by the readout

Let M\mathfrak M be a class of rooted marked martingale chains. For maps Ri:MYi\Rcal _i:\mathfrak M\to Y_i, σ(R)\sigma (\Rcal ) is the initial topology, the coarsest making every Ri\Rcal _i continuous; a sequence converges in it exactly when every coordinate does. The construction is textbook, and the content is the choice of maps. Generated by the readout, the topology adapts to the quantities that describe the dynamics: two chains are close when their at-the-money jets and projected spot responses are close, whatever kernels produce them. Adjoining a coordinate refines the topology and can destroy density, as the wing slope does (Proposition 7.4). Pp(F)\Pp _p(F) denotes the laws on a metric space FF with finite ppth moment, with the distance WpW_p.

Lemma 2.3 (Image density). Let Y=iYiY=\prod _iY_i carry the product topology and let R=(Ri):MY\Rcal =(\Rcal _i):\mathfrak M\to Y. For DMD\subseteq \mathfrak M,

D is σ(R)-dense in MR(D) is dense in R(M).D\text { is }\sigma (\Rcal )\text {-dense in }\mathfrak M \quad \Longleftrightarrow \quad \Rcal (D)\text { is dense in }\Rcal (\mathfrak M).

Proof.The sets R1(V)\Rcal ^{-1}(V), with VV product-open, form a base of the initial topology. Such a set is nonempty exactly when VV meets R(M)\Rcal (\mathfrak M), and it meets DD exactly when VV meets R(D)\Rcal (D). This is precisely image density in the subspace R(M)\Rcal (\mathfrak M).

For an ambient class M\mathfrak M on which the raw readout Rraw\Rcal ^{\rm raw} is defined, let mjX(P):=LawP(Xj)(Pp(R),Wp)\mathsf m_j^X(P):=\Law _P(X_j)\in (\Pp _p(\R ),W_p) and mjS(P):=LawP(eXj)(P1(R+),W1)\mathsf m_j^S(P):=\Law _P(e^{X_j})\in (\Pp _1(\R _+),W_1), and, for finite NN and L{2,,J}\mathcal L\subseteq \{2,\ldots ,J\}, define

(10)τmovN,L(M):=σ(RN,Lraw,mjX,mjS:jJ),τmov(M):=σ(Rraw,mjX,mjS:jJ).\begin{equation}\label {eq:moving-topology} \tau _{\rm mov}^{N,\mathcal L}(\mathfrak M) :=\sigma \bigl (\Rcal ^{\rm raw}_{N,\mathcal L},\mathsf m_j^X,\mathsf m_j^S:j\le J\bigr ), \qquad \tau _{\rm mov}(\mathfrak M) :=\sigma \bigl (\Rcal ^{\rm raw},\mathsf m_j^X,\mathsf m_j^S:j\le J\bigr ). \end{equation}

The full moving-marginal topology τmov\tau _{\rm mov} is generated jointly by all τmovN,L\tau _{\rm mov}^{N,\mathcal L}, so a basic neighbourhood constrains finitely many coordinates; target and approximant may have different mark spaces, since every generator is a model-level map. The topology τmovN,L\tau _{\rm mov}^{N,\mathcal L} is generated by the pseudometric

DN,L(P,Q):=RN,Lraw(P)RN,Lraw(Q)+j=0J[Wp(mjX(P),mjX(Q))+W1(mjS(P),mjS(Q))],\begin{align*} D_{N,\mathcal L}(P,Q) :={}&\bigl \|\Rcal ^{\rm raw}_{N,\mathcal L}(P)-\Rcal ^{\rm raw}_{N,\mathcal L}(Q)\bigr \|\\ &+\sum _{j=0}^J\Bigl [W_p\bigl (\mathsf m_j^X(P),\mathsf m_j^X(Q)\bigr ) +W_1\bigl (\mathsf m_j^S(P),\mathsf m_j^S(Q)\bigr )\Bigr ], \end{align*}

so PP lies in the τmovN,L\tau _{\rm mov}^{N,\mathcal L}-closure of a class AM\mathfrak A\subseteq \mathfrak M exactly when its finite-order calibration error infQADN,L(P,Q)\inf _{Q\in \mathfrak A}D_{N,\mathcal L}(P,Q) vanishes; exact martingality constrains the candidates rather than entering the error. For an approximating class AM\mathfrak A\subseteq \mathfrak M and a target PMP\in \mathfrak M, write

k(P;A):=sup{NN0: PAτmovN,L(M) for every finite L}.k_*(P;\mathfrak A):=\sup \bigl \{N\in \mathbb N_0:\ P\in \overline {\mathfrak A}^{\,\tau _{\rm mov}^{N,\mathcal L}(\mathfrak M)} \text { for every finite }\mathcal L\bigr \}.

Thus k=k_*=\infty means zero calibration error at every finite order and maturity set; both constructions below achieve it with a single sequence.

Density is proved along one constructed sequence, by showing that every generator converges and applying Lemma 2.3; adapted Wasserstein convergence (Section 5) is a step of that proof, not a finer topology.

Remark 2.4 (The readout is not weakly continuous). No Wasserstein distance on path laws, adapted or not, makes the readout continuous on the classes below. Take a two-step Black–Scholes target, Xj+1XjγvjX_{j+1}-X_j\sim \gamma _{v_j} independent with v0,v1>0v_0,v_1>0, and let αn,hn0\alpha _n,h_n\downarrow 0 with αn=o(hn4)\alpha _n=o(h_n^4). On each edge write γvj=γvjαnγαn\gamma _{v_j}=\gamma _{v_j-\alpha _n}*\gamma _{\alpha _n} and keep the second factor, but quantize the first by exponential barycentres on cells of width of order hnh_n, placing one cell on each edge so that the two barycentres sum to αn\alpha _n; a barycentre moves continuously as its cell slides. The result is a finite chain as in Definition 5.1 with edge variances αn\alpha _n, and coupling the quantized factors optimally and the Gaussian factors identically gives AWp0\AW _p\to 0. The two placed cells carry masses of order hnh_n, and the component through both is N(0,2αn)\N (0,2\alpha _n), centred at the money, so the density ff of X2X0X_2-X_0 is at least of order hn2/αnh_n^2/\sqrt {\alpha _n}\to \infty there. With CC the call function of X2X0X_2-X_0, C(0)C(0)=f(0)C''(0)-C'(0)=f(0) while C(0)C(0) and C(0)C'(0) converge, so the expression for a2a_2 in the proof of Lemma 2.1 diverges: the root curvature jet at t2t_2 does not converge. What restores continuity in Section 6 is a Gaussian factor whose variance is held fixed along the sequence. Section 9 lets that variance vanish too, but only with the quantization error driven to zero faster than a power of it; here the variance vanishes faster than the quantization error.

2.6Moving marginals and exact martingality

For the discounted asset SS and X=logSX=\log S (so the log forward is u=Xu=X), martingality is the exponential barycentre identity

(11)exK((x,y),dx,dy)=ex,\begin{equation}\label {eq:asset-martingale} \int e^{x'}K((x,y),dx',dy')=e^x, \end{equation}

for every state (x,y)(x,y) reached by the chain.

Proposition 2.5 (Marginal rigidity). Let QQ start from a known root x0x_0 and carry a label qhq_h taking finitely many values at each date, with transitions

Kh((x,q),dx,dq)=rpr(q)N(x+λr(q)12αh, αh)(dx)δqr+(q)(dq),αh>0,K_h\bigl ((x,q),dx',dq'\bigr ) =\sum _rp_r(q)\,\N \bigl (x+\lambda _r(q)-\tfrac 12\alpha _h,\ \alpha _h\bigr )(dx')\, \delta _{q_r^+(q)}(dq'), \qquad \alpha _h>0,
so that the component index and the successor label depend on the current label alone. Then for every j1j\ge 1,
(12)LawQ(Xj)=νjN(0,Aj),Aj:=h<jαh,\begin{equation}\label {eq:marginal-rigidity} \Law _Q(X_j)=\nu _j*\N (0,A_j), \qquad A_j:=\sum _{h<j}\alpha _h, \end{equation}
for a finitely supported νj\nu _j, which is determined by LawQ(Xj)\Law _Q(X_j) and AjA_j. In particular a law on R\R is the date-jj marginal of such a chain only if it is the convolution of a finitely supported law with a centred Gaussian.

Proof.The label chain evolves independently of the Gaussian increments. Conditionally on the finitely many label and component paths up to date jj, XjX_j is x0x_0 plus the chosen locations, less Aj/2A_j/2, plus jj independent centred Gaussians, hence Gaussian with variance AjA_j; mixing over the finitely many paths gives (12). The Fourier transform of N(0,Aj)\N (0,A_j) does not vanish, so νj^\widehat {\nu _j}, and with it νj\nu _j, is determined.

Both finite classes below have this form, with αh\alpha _h the factor variances (Definition 5.1) or βh,n\beta _{h,n} (Definition 9.6, Lemma 9.7). Under Assumption 3.1, LawP(Xj)=Law(x0+Sj)γAjP\Law _P(X_j)=\Law (x_0+S_j)*\gamma _{A_j^P} with Sj:=h<jLhS_j:=\sum _{h<j}L_h and AjPA_j^P the sum of the target’s factor variances, and matching (12) is impossible if Aj<AjPA_j<A_j^P, forces SjS_j to be finitely supported if Aj=AjPA_j=A_j^P, and forces it to be a finite Gaussian location mixture with a common variance if Aj>AjPA_j>A_j^P. For generic targets the marginals can therefore be approximated — in WpW_p by Theorems 7.3 and 9.13, and in total variation, since the densities converge uniformly (Lemmas 6.2 and 9.10) — but not attained. This is why the formulation lets the marginals move; holding them fixed would in addition cap the dynamics (Proposition 7.5). The approximants below satisfy (11) exactly, so their spot laws are in convex order. A nondegenerate time-zero law μ0\mu _0 with eq|x|μ0(dx)<\int e^{q|x|}\mu _0(dx)<\infty can be included by quantizing X0logexμ0(dx)X_0-\log \int e^x\mu _0(dx) with Lemma 4.1, shifting back and adding an independent N(hn2/2,hn2)\N (-h_n^2/2,h_n^2), hn0h_n\downarrow 0, which keeps E[eX0]\E [e^{X_0}] exact.

3Model class and Gaussian factorization

A variance floor with imperfect correlation leaves a Brownian direction unspanned by the volatility, so each transition splits into a residual return carrying the whole exponential barycentre and an independent mean-one Gaussian factor. This restriction on the target is checkable from the model’s coefficients (Proposition 3.3); the factor keeps martingality exact under quantization, supplies the smoothing of Section 6 and yields the rate of Section 8. Section 9 drops it and pays for it.

Fix dates 0=t0<t1<<tJ0=t_0<t_1<\cdots <t_J containing the observation step and all maturities, a compact convex ERdE\subset \R ^d, and states Zj=(Xj,Yj)R×EZ_j=(X_j,Y_j)\in \R \times E with known root (x0,y0)(x_0,y_0), where YY is the latent volatility state. With Fj=σ(Zh:0hj)\mathcal F_j=\sigma (Z_h:0\le h\le j), homogeneity means that for kernels KjretK_j^{\rm ret},

Law((Xj+1Xj,Yj+1)Fj)=Kjret(Yj)a.s.\Law ((X_{j+1}-X_j,Y_{j+1})\mid \mathcal F_j)=K_j^{\rm ret}(Y_j)\qquad \text {a.s.}

For α>0\alpha >0 put γα:=N(α/2,α)\gamma _\alpha :=\N (-\alpha /2,\alpha ), so that E[eG]=1\E [e^G]=1 for GγαG\sim \gamma _\alpha.

Assumption 3.1 (Independent mean-one Gaussian factor). For every edge j<Jj<J, there are αj>0\alpha _j>0 and a kernel

Hj:EP(R×E),yHj(y;d,dy),H_j:E\longrightarrow \Pp (\R \times E), \qquad y\longmapsto H_j(y;d\ell ,dy'),
such that, jointly over the whole grid,
(13)Xj+1=Xj+Lj+Gj,Law((Lj,Yj+1)Fj)=Hj(Yj),Gjγαj,\begin{equation}\label {eq:factor-kernel} X_{j+1}=X_j+L_j+G_j,\qquad \Law ((L_j,Y_{j+1})\mid \mathcal F_j)=H_j(Y_j),\qquad G_j\sim \gamma _{\alpha _j}, \end{equation}
where the GjG_j are mutually independent and
(G0,,GJ1)  ((Lh,Yh+1))h<J.(G_0,\ldots ,G_{J-1})\ \perp \!\!\!\perp \ \bigl ((L_h,Y_{h+1})\bigr )_{h<J}.
Moreover,
(14)eHj(y;d,dy)=1for every yE.\begin{equation}\label {eq:residual-barycentre} \int e^\ell H_j(y;d\ell ,dy')=1 \qquad \text {for every }y\in E. \end{equation}

By (14) and E[eGj]=1\E [e^{G_j}]=1, (11) holds exactly. Fix p>2p>2 and the metric d((,y),(¯,y¯)):=|¯|+|yy¯|d((\ell ,y),(\bar \ell ,\bar y)):=|\ell -\bar \ell |+|y-\bar y| on R×E\R \times E.

Assumption 3.2 (Kernel regularity and moments). There are q>pq>p and C,L<C,L<\infty such that for every jj and y,y¯Ey,\bar y\in E,

(15)Wp(Hj(y),Hj(y¯))L|yy¯|,(16)eq||Hj(y;d,dy)C.\begin{align} W_p(H_j(y),H_j(\bar y))&\le L|y-\bar y|, \label {eq:kernel-Lip}\\ \int e^{q|\ell |}H_j(y;d\ell ,dy')&\le C. \label {eq:kernel-exp} \end{align}

Proposition 3.3 (Stochastic volatility supplies the factor). Suppose

(17)dXt=12σ(t,Yt)2dt+σ(t,Yt)(ρ(t,Yt)dWt+1ρ(t,Yt)2dBt),(18)dYt=b(t,Yt)dt+Γ(t,Yt)dWt,\begin{align} dX_t={}&-\tfrac 12\sigma (t,Y_t)^2dt +\sigma (t,Y_t)\left (\rho (t,Y_t)^\top dW_t +\sqrt {1-\|\rho (t,Y_t)\|^2}\,dB_t\right ),\label {eq:sv-X}\\ dY_t={}&b(t,Y_t)dt+\Gamma (t,Y_t)dW_t,\label {eq:sv-Y} \end{align}

where BB is independent of WW, S=eXS=e^X is a true martingale from every initial state (t,x,y)(t,x,y) under consideration, YY is a well-posed strong Markov solution adapted to the augmented filtration of WW, and

σσ>0,ρρ¯<1.\sigma \ge \underline \sigma >0, \qquad \|\rho \|\le \bar \rho <1.
Then, after passing to an extension of the probability space carrying the auxiliary normals constructed in the proof, the chain satisfies Assumption 3.1 with, for any fixed θ(0,1)\theta \in (0,1),
(19)αj=θσ2(1ρ¯2)(tj+1tj).\begin{equation}\label {eq:alpha-floor} \alpha _j=\theta \,\underline \sigma ^2(1-\bar \rho ^2)(t_{j+1}-t_j). \end{equation}

Proof.Let G=σ(Ws:0stJ)\mathcal G=\sigma (W_s:0\le s\le t_J) (augmented by the initial state), and condition first on G\mathcal G. On [tj,tj+1][t_j,t_{j+1}] the log return is conditionally Gaussian,

Xj+1XjGN(Mj,Qj),X_{j+1}-X_j\mid \mathcal G\sim \N (M_j,Q_j),
where
Mj=12tjtj+1σs2ds+tjtj+1σsρsdWs,Qj=tjtj+1σs2(1ρs2)dsσ2(1ρ¯2)(tj+1tj)>αj.\begin{align*} M_j&=-\frac 12\int _{t_j}^{t_{j+1}}\sigma _s^2ds +\int _{t_j}^{t_{j+1}}\sigma _s\rho _s^\top dW_s,\\ Q_j&=\int _{t_j}^{t_{j+1}}\sigma _s^2(1-\|\rho _s\|^2)ds \ge \underline \sigma ^2(1-\bar \rho ^2)(t_{j+1}-t_j)>\alpha _j. \end{align*}

On an extension of the probability space, do this simultaneously on every disjoint grid edge using mutually independent pairs of standard normals (Z0,j,Z1,j)(Z_{0,j},Z_{1,j}), independent of the WW-driven path, and set

Gj=αj2+αjZ0,j,Lj=Mj+αj2+QjαjZ1,j.G_j=-\frac {\alpha _j}{2}+\sqrt {\alpha _j}Z_{0,j}, \qquad L_j=M_j+\frac {\alpha _j}{2}+\sqrt {Q_j-\alpha _j}Z_{1,j}.
Conditionally on G\mathcal G, Gj+LjG_j+L_j has law N(Mj,Qj)\N (M_j,Q_j), independently across edges, as do the original increments, whose BB-parts live on disjoint edges; so Xj+1Xj:=Gj+LjX_{j+1}-X_j:=G_j+L_j leaves the law of the grid chain unchanged. The vector (Z0,j)j<J(Z_{0,j})_{j<J} is independent of G\mathcal G and (Z1,j)j<J(Z_{1,j})_{j<J}, which gives the path-level independence in Assumption 3.1, and (Lj,Yj+1)(L_j,Y_{j+1}) is a function of YjY_j, the WW-increments on [tj,tj+1][t_j,t_{j+1}] and Z1,jZ_{1,j}, all but YjY_j independent of Fj\mathcal F_j, so its law given Fj\mathcal F_j is a kernel Hj(Yj)H_j(Y_j). Finally, homogeneity and the asset-martingale property give
1=E[eXj+1XjYj]=E[eLjYj]E[eGj],1=\E [e^{X_{j+1}-X_j}\mid Y_j] =\E [e^{L_j}\mid Y_j]\E [e^{G_j}],
and E[eGj]=1\E [e^{G_j}]=1, proving (14).

Proposition 3.4 (A floored and capped square-root model lies in the class). Fix 0<v<v<0<\underline v<\overline v<\infty, put E=[v,v]E=[\underline v,\overline v], and let b,Γ:RRb,\Gamma :\R \to \R be bounded and Lipschitz with

Γ(v)=Γ(v)=0,Γ>0 on (v,v),b(v)>0>b(v).\Gamma (\underline v)=\Gamma (\overline v)=0, \qquad \Gamma >0\ \text {on }(\underline v,\overline v), \qquad b(\underline v)>0>b(\overline v).
Let |ρ|<1|\rho |<1, let BB be independent of WW, let V0EV_0\in E, and consider
(20)dXt=12Vtdt+Vt(ρdWt+1ρ2dBt),dVt=b(Vt)dt+Γ(Vt)dWt.\begin{equation}\label {eq:capped-heston} dX_t=-\tfrac 12V_t\,dt+\sqrt {V_t}\left (\rho \,dW_t+\sqrt {1-\rho ^2}\,dB_t\right ), \qquad dV_t=b(V_t)\,dt+\Gamma (V_t)\,dW_t. \end{equation}
Then VV takes values in EE, and the grid skeleton of (20) satisfies Assumptions 3.1, 3.2 and 6.3; hence it lies in TSV\mathfrak T_{\rm SV} and Theorem 7.3 applies to it. With bb and Γ\Gamma equal to κ(mv)\kappa (m-v) and ξv\xi \sqrt v away from the endpoints this is floored and capped Heston; if v<m<v\underline v<m<\overline v, only Γ\Gamma needs modifying.

Proof.Replacing v\sqrt v by vv\sqrt {v\vee \underline v}, which agrees with it on EE, makes the system globally Lipschitz, so it has a unique strong solution. The coefficient Γ\Gamma and the drift bb(v)bb-b(\underline v)\le b both vanish at v\underline v, so the constant v\underline v solves the equation with that drift, and comparison for one-dimensional equations with a common Lipschitz diffusion coefficient [31, Section 5.2.C] gives VtvV_t\ge \underline v from every start in EE; the drift bb(v)bb-b(\overline v)\ge b gives VtvV_t\le \overline v in the same way. Hence EE is invariant, endpoints included.

Assumption 3.1. On EE one has σ=V[v,v]\sigma =\sqrt V\in [\sqrt {\underline v},\sqrt {\overline v}], so the volatility is bounded and floored, and ρ=|ρ|<1\|\rho \|=|\rho |<1. Boundedness of σ\sigma gives Novikov’s condition, so eXe^X is a true martingale from every state, and VV is a well-posed strong Markov solution adapted to the augmented filtration of WW. Proposition 3.3 therefore applies and yields (13) with αj=θv(1ρ2)(tj+1tj)\alpha _j=\theta \,\underline v(1-\rho ^2)(t_{j+1}-t_j) for any θ(0,1)\theta \in (0,1).

Assumption 3.2. In the representation of Proposition 3.3, couple the copies started from yy and y¯\bar y through the same WW and the same auxiliary normals. Then LjL¯j=(MjM¯j)+(QjαjQ¯jαj)Z1,jL_j-\bar L_j=(M_j-\bar M_j)+(\sqrt {Q_j-\alpha _j}-\sqrt {\bar Q_j-\alpha _j})Z_{1,j}, the square root is Lipschitz because Qjαj(1θ)v(1ρ2)(tj+1tj)Q_j-\alpha _j\ge (1-\theta )\underline v(1-\rho ^2)(t_{j+1}-t_j), and Gronwall with the Burkholder–Davis–Gundy inequality bounds VV¯V-\bar V, hence MjM¯jM_j-\bar M_j, QjQ¯jQ_j-\bar Q_j and the marks, by a multiple of |yy¯||y-\bar y| in LpL^p; this is (15). For (16), condition on the path of WW: the residual return is Gaussian with variance at most v(tj+1tj)\overline v(t_{j+1}-t_j) and mean bounded by 12v(tj+1tj)\tfrac 12\overline v(t_{j+1}-t_j) plus a stochastic integral of quadratic variation at most vρ2(tj+1tj)\overline v\rho ^2(t_{j+1}-t_j). Both terms are sub-Gaussian with parameters depending only on v\overline v and the mesh, so every exponential moment is finite uniformly in the starting mark; in particular (16) holds for any qq.

Assumption 6.3. The volatility satisfies vσv\sqrt {\underline v}\le \sigma \le \sqrt {\overline v} pathwise, and a European call is a convex payoff, so by the comparison theorem for misspecified volatility [21] the conditional call price lies between the Black–Scholes prices at those two volatilities. Monotonicity of B(0,)B(0,\cdot ) then places the ATM implied total variance in [v(ttj),v(ttj)][\underline v(t_\ell -t_j),\overline v(t_\ell -t_j)], a compact subset of (0,)(0,\infty ) independent of the mark.

4Log-barycentric quantization

Quantization theory [25] gives rates for the unconstrained problem; the point here is that representing cells by exponential barycentres preserves martingality exactly at every level.

Lemma 4.1 (Exact exponential-barycentre quantization). Let HP(R×E)H\in \Pp (\R \times E) have unit exponential mean, eH(d,dy)=1\int e^\ell H(d\ell ,dy)=1, and a finite exponential moment eq||H(d,dy)<\int e^{q|\ell |}H(d\ell ,dy)<\infty for some q>pq>p. Then there are finitely supported laws Hn=r=1Mnprnδ(λrn,yrn)H^n=\sum _{r=1}^{M_n}p_r^n\delta _{(\lambda _r^n,y_r^n)} such that

rprneλrn=1,Wp(Hn,H)0,supnrprneq|λrn|<.\sum _r p_r^n e^{\lambda _r^n}=1, \qquad W_p(H^n,H)\longrightarrow 0, \qquad \sup _n\sum _rp_r^ne^{q|\lambda _r^n|}<\infty .

Proof.Let (L,Y)H(L,Y)\sim H. Since R×E\R \times E is standard Borel, choose increasing finite sigma-fields An\mathcal A_n whose union generates σ(L,Y)\sigma (L,Y). Set

Un=E[eLAn],Ln=logUn,Yn=E[YAn].U_n=\E [e^L\mid \mathcal A_n], \qquad L_n=\log U_n, \qquad Y_n=\E [Y\mid \mathcal A_n].
The convexity of EE gives YnEY_n\in E, and finiteness of An\mathcal A_n makes (Ln,Yn)(L_n,Y_n) finitely valued. Moreover,
E[eLn]=E[Un]=E[eL]=1.\E [e^{L_n}]=\E [U_n]=\E [e^L]=1.
Martingale convergence yields UneLU_n\to e^L and YnYY_n\to Y, hence LnLL_n\to L, almost surely. Conditional Jensen for the convex maps uuqu\mapsto u^q and uuqu\mapsto u^{-q} gives
eqLnE[eqLAn],eqLn=UnqE[eqLAn].e^{qL_n}\le \E [e^{qL}\mid \mathcal A_n], \qquad e^{-qL_n}=U_n^{-q}\le \E [e^{-qL}\mid \mathcal A_n].
Since eq|x|eqx+eqxe^{q|x|}\le e^{qx}+e^{-qx}, integration gives the asserted uniform two-sided exponential bound. Therefore (|Ln|p)n(|L_n|^p)_n is uniformly integrable. Since EE is compact, YnYY_n\to Y in every LpL^p. Thus (Ln,Yn)(L,Y)(L_n,Y_n)\to (L,Y) in LpL^p, which implies Wp(Law(Ln,Yn),H)0W_p(\Law (L_n,Y_n),H)\to 0. Taking Hn=Law(Ln,Yn)H^n=\Law (L_n,Y_n) proves the result.

The partitions must generate σ(L,Y)\sigma (L,Y), not σ(Y)\sigma (Y): if the cells CrC_r are unions of full yy-fibres, per-fibre normalization gives every λr=logE[eL(L,Y)Cr]\lambda _r=\log \E [e^L\mid (L,Y)\in C_r] the value 00, leaving no spot–mark covariance.

5Finite marked trees and adapted convergence

Quantizing every node gives a finite marked tree of exact martingale kernels, and a recursive bicausal coupling gives adapted convergence.

For path laws P,QP,Q on (R×E)J+1(\R \times E)^{J+1} define

(21)AWp(P,Q)p:=infπΠbc(P,Q)Eπ[j=0J(|XjX^j|+|YjY^j|)p],\begin{equation}\label {eq:AW-definition} \AW _p(P,Q)^p :=\inf _{\pi \in \Pi _{\rm bc}(P,Q)} \E _\pi \left [\sum _{j=0}^J \bigl (|X_j-\widehat X_j|+|Y_j-\widehat Y_j|\bigr )^p\right ], \end{equation}

where Πbc\Pi _{\rm bc} denotes the bicausal couplings [35]. For Markov chains on the grid, recursively coupling next-step kernels as functions of the two current states gives such a coupling.

Definition 5.1 (Finite marked GM chain). A finite marked GM chain has finite continuation-label sets QjnQ_j^n and edgewise constants αj>0\alpha _j>0, all part of the chain’s data. The label space may be the target latent space, as in the homogeneous construction, or may include a finite proxy log-price, as in Section 9. From (x,q)(x,q), qQjnq\in Q_j^n, its transition is

(22)Kjn((x,q),dx,dq)=r=1Mj,qnpj,q,rnN(x+λj,q,rnαj2,αj)(dx)δqj,q,rn,+(dq),\begin{equation}\label {eq:finite-gm-kernel} K_j^n((x,q),dx',dq') =\sum _{r=1}^{M_{j,q}^n}p_{j,q,r}^n \N \left (x+\lambda _{j,q,r}^n-\frac {\alpha _j}{2},\alpha _j\right )(dx') \delta _{q_{j,q,r}^{n,+}}(dq'), \end{equation}
where
(23)rpj,q,rneλj,q,rn=1.\begin{equation}\label {eq:finite-exp-barycentre} \sum _rp_{j,q,r}^ne^{\lambda _{j,q,r}^n}=1. \end{equation}

Every kernel (22) is an exact asset-martingale kernel:

exKjn((x,q),dx,dq)=exrpj,q,rneλj,q,rnE[eGj]=ex.\int e^{x'}K_j^n((x,q),dx',dq') =e^x\sum _rp_{j,q,r}^ne^{\lambda _{j,q,r}^n}\E [e^{G_j}]=e^x.

Theorem 5.2 (Finite-tree approximation). Under Assumptions 3.1 and 3.2, there are finite marked GM chains PnP^n starting from (x0,y0)(x_0,y_0) such that

  1. every transition is an exact asset-martingale transition;
  2. AWp(Pn,P)0\AW _p(P^n,P)\to 0;
  3. for every j1j\ge 1, Law(Xjn)\Law (X_j^n) is a finite Gaussian mixture, and Law(Xjn,Yjn)\Law (X_j^n,Y_j^n) converges to Law(Xj,Yj)\Law (X_j,Y_j) in WpW_p.

Proof.Fix εn0\varepsilon _n\downarrow 0 and put E0n={y0}E_0^n=\{y_0\}. Suppose the finite set EjnE_j^n has been constructed. For every yEjny\in E_j^n, apply Lemma 4.1 to Hj(y)H_j(y) and choose Hjn(y)=rpj,y,rnδ(λj,y,rn,yj,y,rn,+)H_j^n(y)=\sum _rp_{j,y,r}^n\delta _{(\lambda _{j,y,r}^n,y_{j,y,r}^{n,+})} with Wp(Hjn(y),Hj(y))εnW_p(H_j^n(y),H_j(y))\le \varepsilon _n and exact exponential barycentre. The conditional-Jensen estimates in that lemma and (16) give the same uniform qq-exponential bound over all selected nodes and all nn. Let Ej+1nE_{j+1}^n be the union of the finitely many successor marks. This recursively defines (22); exact martingality was checked above.

We construct a bicausal coupling recursively. Given current states (x,y)(x,y) and (x^,y^)(\widehat x,\widehat y), first note that, by (15),

(24)Wp(Hj(y),Hjn(y^))L|yy^|+εn.\begin{equation}\label {eq:edge-error} W_p(H_j(y),H_j^n(\widehat y)) \le L|y-\widehat y|+\varepsilon _n. \end{equation}
Choose a Borel measurable εn\varepsilon _n-optimal coupling kernel for this pair of laws; such measurable selections exist for Wasserstein costs on Polish spaces [36, Corollary 5.22]. Its conditional LpL^p cost is bounded by the right side of (24) plus εn\varepsilon _n, hence by L|yy^|+2εnL|y-\widehat y|+2\varepsilon _n. Use the same independent GjγαjG_j\sim \gamma _{\alpha _j} in the two transitions. If the coupled residual variables are (L,Y)(L,Y') and (L^,Y^)(\widehat L,\widehat Y'), then
XX^=XX^+LL^.X'-\widehat X'=X-\widehat X+L-\widehat L.
Writing ej:=(E[(|XjX^j|+|YjY^j|)p])1/pe_j:=\bigl (\E [(|X_j-\widehat X_j|+|Y_j-\widehat Y_j|)^p]\bigr )^{1/p} for the date-jj coupling cost in the metric of (21), the price gap carries over and the coupled residual–mark cost adds at most LYjY^jLp+2εnL\|Y_j-\widehat Y_j\|_{L^p}+2\varepsilon _n, so Minkowski’s inequality gives
(25)ej+1(1+L)ej+2εn,e0=0.\begin{equation}\label {eq:tree-recursion} e_{j+1}\le (1+L)e_j+2\varepsilon _n, \qquad e_0=0. \end{equation}
Iterating over the finite grid gives an O(εn)O(\varepsilon _n) bound at every date. The selected coupling kernel at each step depends only on the pair of current states, and the common GjG_j is drawn independently of the coupled residuals, so each coordinate receives its own transition kernel; the coupling is therefore bicausal. This proves AWp(Pn,P)0\AW _p(P^n,P)\to 0 and the marginal WpW_p convergence.

It remains to identify the approximating marginal. Conditional on a finite residual branch path π\pi up to date jj,

XjnN(x0+h<jλπ,h12Aj, Aj),Aj=h<jαh.X_j^n\sim \N \left ( x_0+\sum _{h<j}\lambda _{\pi ,h}-\frac 12A_j,\ A_j\right ), \qquad A_j=\sum _{h<j}\alpha _h.
There are only finitely many branch paths, so Law(Xjn)\Law (X_j^n) is a finite Gaussian mixture.

Since eXjneXje^{X_j^n}\Rightarrow e^{X_j} with common mean ex0e^{x_0}, the finite-lognormal-mixture spot laws converge in W1W_1, and so do the call prices Cjn,CjC_j^n,C_j on Stjn=eXjnS^n_{t_j}=e^{X_j^n} and Stj=eXjS_{t_j}=e^{X_j}, uniformly in strike:

(26)supK0|Cjn(K)Cj(K)|W1(Law(Stjn),Law(Stj)).\begin{equation}\label {eq:call-W1} \sup _{K\ge 0}|C_j^n(K)-C_j(K)|\le W_1\bigl (\Law (S^n_{t_j}),\Law (S_{t_j})\bigr ). \end{equation}

6Continuation stability and Gaussian smoothing

The recursive coupling controls future conditional laws, and the Gaussian factor on the last edge before each maturity upgrades this to uniform convergence of densities and their derivatives.

Lemma 6.1 (Future conditional laws). Fix j<j<\ell. Let ynEjny_n\in E_j^n and suppose ynyEy_n\to y\in E. Couple the target chain started from yy at tjt_j and the approximating continuation started from yny_n, with the same normalized current log price, by the construction in Theorem 5.2. Then their future path laws through tt_\ell converge in AWp\AW _p. Moreover, with

Zj,:=h=j2(Lh+Gh)+L1,XXj=Zj,+G1,Z_{j,\ell }:=\sum _{h=j}^{\ell -2}(L_h+G_h)+L_{\ell -1}, \qquad X_\ell -X_j=Z_{j,\ell }+G_{\ell -1},
and the analogous Zj,nZ_{j,\ell }^n, one has
Wp(Law(Zj,n),Law(Zj,))Cj,(|yny|+εn).W_p\bigl (\Law (Z_{j,\ell }^n),\Law (Z_{j,\ell })\bigr ) \le C_{j,\ell }(|y_n-y|+\varepsilon _n).

Proof.The initial state error is |yny||y_n-y|. At each later edge, the measurable coupling estimate following (24) applies with the same constants. The finite-horizon induction used in Theorem 5.2 therefore bounds the future path error by Cj,(|yny|+εn)C_{j,\ell }(|y_n-y|+\varepsilon _n), which tends to zero. The coupling is again bicausal. On its extended space every GhG_h is coupled identically and every LhL^hL_h-\widehat L_h is controlled by the corresponding edge estimate. Minkowski’s inequality applied directly to the displayed residual sums gives the stated WpW_p bound.

Lemma 6.2 (Common-Gaussian smoothing). Let ηn,ηP1(R)\eta _n,\eta \in \Pp _1(\R ) with W1(ηn,η)0W_1(\eta _n,\eta )\to 0, and fix α>0\alpha >0. If fn=ηnφαf_n=\eta _n*\varphi _\alpha and f=ηφαf=\eta *\varphi _\alpha, where φα\varphi _\alpha is the density of γα\gamma _\alpha, then for every r0r\ge 0,

(27)fn(r)f(r)φα(r+1)W1(ηn,η).\begin{equation}\label {eq:smoothing-bound} \|f_n^{(r)}-f^{(r)}\|_\infty \le \|\varphi _\alpha ^{(r+1)}\|_\infty W_1(\eta _n,\eta ). \end{equation}

Proof.Take a coupling (Un,U)(U_n,U) attaining, or arbitrarily approaching, W1W_1. For every xx, the mean value theorem gives

|φα(r)(xUn)φα(r)(xU)|φα(r+1)|UnU|.|\varphi _\alpha ^{(r)}(x-U_n)-\varphi _\alpha ^{(r)}(x-U)| \le \|\varphi _\alpha ^{(r+1)}\|_\infty |U_n-U|.
Take expectations and then the supremum over xx.

Since XXj=Zj,+G1X_\ell -X_j=Z_{j,\ell }+G_{\ell -1} with G1G_{\ell -1} independent and additive at maturity, the conditional return law is the residual law convolved with γα1\gamma _{\alpha _{\ell -1}}, and Lemma 6.2 applies.

Assumption 6.3 (Implied-variance nondegeneracy). For every date pair under consideration, the target ATM implied total variances stay in one compact subset of (0,)(0,\infty ) uniformly over yEy\in E.

Lemma 6.4 (Static jets from converging return densities). Fix p>2p>2, p=p/(p1)p'=p/(p-1) and an order r0r\ge 0. For each nn let Zn\mathcal Z_n be a set of labels and, for zZnz\in \mathcal Z_n, let Rn(z)R_n(z) and R(z)R(z) be coupled log returns with unit exponential mean. Suppose that

  1. supzZnRn(z)R(z)Lp0\sup _{z\in \mathcal Z_n}\|R_n(z)-R(z)\|_{L^p}\to 0;
  2. supnsupzZn(E[epRn(z)]+E[epR(z)])<\sup _n\sup _{z\in \mathcal Z_n}\bigl (\E [e^{p'R_n(z)}]+\E [e^{p'R(z)}]\bigr )<\infty;
  3. Rn(z)R_n(z) and R(z)R(z) have densities fn(;z)f_n(\cdot \,;z) and f(;z)f(\cdot \,;z) of class CrC^r such that, for every iri\le r, supnsupzZnif(;z)<\sup _n\sup _{z\in \mathcal Z_n}\|\partial ^if(\cdot \,;z)\|_\infty <\infty and supzZnifn(;z)if(;z)0\sup _{z\in \mathcal Z_n}\|\partial ^if_n(\cdot \,;z)-\partial ^if(\cdot \,;z)\|_\infty \to 0;
  4. the ATM implied total variances of the R(z)R(z) lie in one compact subset of (0,)(0,\infty ).

Then, writing am(R)a_m(R) for the mmth ATM jet of the implied total variance of kE[(eRek)+]k\mapsto \E [(e^R-e^k)^+],

supzZn|am(Rn(z))am(R(z))|0(mr+2).\sup _{z\in \mathcal Z_n}\bigl |a_m(R_n(z))-a_m(R(z))\bigr |\longrightarrow 0 \qquad (m\le r+2).

Proof.Write Cn,CC_n,C for the two call functions. By |euev||uv|(eu+ev)|e^u-e^v|\le |u-v|(e^u+e^v) and Hölder’s inequality, (1) and (2) give supz|Cn(0)C(0)|supzE|eRneR|0\sup _z|C_n(0)-C(0)|\le \sup _z\E |e^{R_n}-e^R|\to 0. Next, C(0)=P(R>0)C'(0)=-\mathbb P(R>0), and for every δ>0\delta >0, |P(Rn>0)P(R>0)|δ1E|RnR|+2δsupf|\mathbb P(R_n>0)-\mathbb P(R>0)|\le \delta ^{-1}\E |R_n-R|+2\delta \sup f; letting nn\to \infty and then δ0\delta \downarrow 0 gives uniform convergence of Cn(0)C_n'(0). Since C(k)C(k)=ekf(k)C''(k)-C'(k)=e^kf(k), each C(i)(0)C^{(i)}(0) with 2ir+22\le i\le r+2 is a fixed linear combination of C(0)C'(0) and the values f(h)(0)f^{(h)}(0), hi2h\le i-2, so (3) gives uniform convergence of all of them, and the vectors stay in a bounded set. By (4) and the convergence of Cn(0)C_n(0), the approximating ATM variances lie in a slightly enlarged compact subset of (0,)(0,\infty ) for large nn, on which the map of Lemma 2.1 is uniformly continuous on bounded sets.

Proposition 6.5 (Convergence of conditional static jets). Under Assumptions 3.1, 3.2, and 6.3, for every j<j<\ell, every finite mm,

maxzEjn|am,nj,(z)amj,(z)|0.\max _{z\in E_j^n} \left |a_{m,n}^{j,\ell }(z)-a_m^{j,\ell }(z)\right |\longrightarrow 0.
The target map yamj,(y)y\mapsto a_m^{j,\ell }(y) is continuous on EE. Consequently, if ynEjny_n\in E_j^n and ynyy_n\to y, then am,nj,(yn)amj,(y)a_{m,n}^{j,\ell }(y_n)\to a_m^{j,\ell }(y).

Proof.Apply Lemma 6.4 with Zn=Ejn\mathcal Z_n=E_j^n, with Rn(z),R(z)R_n(z),R(z) the coupled normalized returns XXjX_\ell -X_j started from zz, and with every rr; all conditional forwards equal one exactly. Hypothesis (1) is Lemma 6.1 with both continuations started from the same label zEjnz\in E_j^n: the residual sums are then within Cj,εnC_{j,\ell }\varepsilon _n in LpL^p, with a constant independent of zz, and the Gaussian factors are coupled identically. For (2), the bound follows edge by edge from the tower property: conditioning on Fh\mathcal F_h and using epHh(y;d,dy)eq||Hh(y;d,dy)C\int e^{p'\ell }H_h(y;d\ell ,dy')\le \int e^{q|\ell |}H_h(y;d\ell ,dy')\le C uniformly in yy, valid since p<2<qp'<2<q, gives a factor CC per edge and hence CjC^{\ell -j} over the finite grid, uniformly in zz; the independent Gaussian factors contribute hep(p1)αh/2\prod _he^{p'(p'-1)\alpha _h/2}, and the atomic kernels inherit the bound by Lemma 4.1. For (3), both returns are residual laws convolved with γα1\gamma _{\alpha _{\ell -1}}, so ifφα1(i)\|\partial ^if\|_\infty \le \|\varphi _{\alpha _{\ell -1}}^{(i)}\|_\infty and Lemma 6.2 gives the convergence at every order. Hypothesis (4) is Assumption 6.3. Applying the same lemma to target continuations started from ynyy_n\to y and coupled through (15) gives continuity of the target jet map; the sequential claim follows by the triangle inequality.

7Projected dynamics and the density theorem

The first-edge factorization reduces both projections to covariance ratios over the residual kernel.

For the target first edge, Assumption 3.1 gives

X1X0=L0+G0,G0(L0,Y1).X_1-X_0=L_0+G_0, \qquad G_0\perp (L_0,Y_1).

Homogeneity makes the continuation jet Am=am1,(Y1)A_m^\ell =a_m^{1,\ell }(Y_1) a function of the successor mark. Consequently Cov(Am,X1X0)=CovH0(y0)(am1,(Y1),L0)\Cov (A_m^\ell ,X_1-X_0)=\Cov _{H_0(y_0)}(a_m^{1,\ell }(Y_1),L_0) and Var(X1X0)=α0+VarH0(y0)(L0)\Var (X_1-X_0)=\alpha _0+\Var _{H_0(y_0)}(L_0), so

(28)dm=CovH0(y0)(am1,(Y1),L0)α0+VarH0(y0)(L0).\begin{equation}\label {eq:target-projection-formula} d_m^\ell =\frac {\Cov _{H_0(y_0)}(a_m^{1,\ell }(Y_1),L_0)} {\alpha _0+\Var _{H_0(y_0)}(L_0)}. \end{equation}

For the finite chain write its first-edge atoms as (λrn,yrn)(\lambda _r^n,y_r^n) with weights prnp_r^n, and put

Am,rn=am,n1,(yrn),A¯mn=rprnAm,rn,λ¯n=rprnλrn.A_{m,r}^n=a_{m,n}^{1,\ell }(y_r^n),\qquad \bar A_m^n=\sum _rp_r^nA_{m,r}^n,\qquad \bar \lambda ^n=\sum _rp_r^n\lambda _r^n.

Then

(29)dm,n=rprn(Am,rnA¯mn)(λrnλ¯n)α0+rprn(λrnλ¯n)2.\begin{equation}\label {eq:finite-projection-formula} d_{m,n}^\ell =\frac {\sum _rp_r^n(A_{m,r}^n-\bar A_m^n)(\lambda _r^n-\bar \lambda ^n)} {\alpha _0+\sum _rp_r^n(\lambda _r^n-\bar \lambda ^n)^2}. \end{equation}

Lemma 7.1 (Projected rows from converging continuation jets). Let (Rn,ζn)(R_n,\zeta _n) and (R,ζ)(R,\zeta ) be coupled first-edge log returns and successor labels, ζn\zeta _n taking values in a finite set Zn\mathcal Z_n and ζ\zeta in a metric label space, with RnRR_n\to R and ζnζ\zeta _n\to \zeta in LpL^p for some p>2p>2. Let AnA_n on Zn\mathcal Z_n and AA on the label space be bounded uniformly in nn, with AA uniformly continuous and maxzZn|An(z)A(z)|0\max _{z\in \mathcal Z_n}|A_n(z)-A(z)|\to 0. If Var(R)>0\Var (R)>0, then

Cov(An(ζn),Rn)Var(Rn)Cov(A(ζ),R)Var(R).\frac {\Cov (A_n(\zeta _n),R_n)}{\Var (R_n)}\longrightarrow \frac {\Cov (A(\zeta ),R)}{\Var (R)}.

Proof.Since |An(ζn)A(ζ)|maxzZn|An(z)A(z)|+|A(ζn)A(ζ)||A_n(\zeta _n)-A(\zeta )|\le \max _{z\in \mathcal Z_n}|A_n(z)-A(z)|+|A(\zeta _n)-A(\zeta )|, and the second term tends to zero in probability by uniform continuity, An(ζn)A(ζ)A_n(\zeta _n)\to A(\zeta ) in probability, and the uniform bound upgrades this to Lp/(p1)L^{p/(p-1)}. Hölder’s inequality against RnRR_n\to R in LpL^p gives convergence of E[An(ζn)Rn]\E [A_n(\zeta _n)R_n] and of the means, and p>2p>2 gives Var(Rn)Var(R)>0\Var (R_n)\to \Var (R)>0.

Lemma 7.2 (Projected readout continuity). For every finite mm and every 2\ell \ge 2,

dm,ndm.d_{m,n}^\ell \longrightarrow d_m^\ell .

Proof.Apply Lemma 7.1 under the first-edge coupling of Theorem 5.2, with Rn=X1nx0R_n=X_1^n-x_0, R=X1x0R=X_1-x_0, ζn=Y1n\zeta _n=Y_1^n, ζ=Y1\zeta =Y_1, An=am,n1,A_n=a_{m,n}^{1,\ell } and A=am1,A=a_m^{1,\ell }. Proposition 6.5 supplies the hypotheses on AnA_n and AA: continuity on the compact EE is uniform and bounds AA, and uniform convergence then bounds the AnA_n uniformly in nn. Finally Var(X1X0)α0>0\Var (X_1-X_0)\ge \alpha _0>0. The two sides are (29) and (28).

Let TSV\mathfrak T_{\rm SV} be the class of homogeneous rooted targets on the fixed grid and state space that satisfy Assumptions 3.1, 3.2, and 6.3. Let GSV\mathfrak G_{\rm SV} be the finite marked GM chains of Definition 5.1 whose finite labels lie in the homogeneous mark space EE, with the stated log and spot moments and well-defined raw readouts, and set

MSV:=TSVGSV,τmovSV:=τmov(MSV).\mathfrak M_{\rm SV}:=\mathfrak T_{\rm SV}\cup \mathfrak G_{\rm SV}, \qquad \tau _{\rm mov}^{\rm SV}:=\tau _{\rm mov}(\mathfrak M_{\rm SV}).

Theorem 7.3 (Homogeneous-SV density at every finite projected order). Let PTSVP\in \mathfrak T_{\rm SV} and let L{2,,J}\mathcal L\subseteq \{2,\ldots ,J\} be finite. Then the finite marked GM chains PnP^n of Theorem 5.2 satisfy:

  1. every PnP^n is an exact asset martingale;
  2. for every j1j\ge 1, the log-price marginal Law(Xjn)\Law (X_j^n) is a finite Gaussian mixture and converges to Law(Xj)\Law (X_j) in WpW_p, while the corresponding finite lognormal spot marginal converges in W1W_1; the j=0j=0 marginal remains the fixed root δx0\delta _{x_0};
  3. for every finite NN and every L\ell \in \mathcal L, the static jets through aN+1a_{N+1}^\ell and the projected rows through dNd_N^\ell converge;
  4. for every finite NN and each L\ell \in \mathcal L with the regular pivot a10a_1^\ell \ne 0,

    vm,nvm,0mN.v_{m,n}^\ell \longrightarrow v_m^\ell , \qquad 0\le m\le N.

Consequently

TSVGSVτmovSV,\mathfrak T_{\rm SV}\subseteq \overline {\mathfrak G_{\rm SV}}^{\,\tau _{\rm mov}^{\rm SV}},
and GSV\mathfrak G_{\rm SV} is dense in MSV\mathfrak M_{\rm SV} both for the raw-readout topology and for τmovSV\tau _{\rm mov}^{\rm SV}. Equivalently,
k(P;GSV)=.k_*(P;\mathfrak G_{\rm SV})=\infty .

Proof.Items 1 and 2 are Theorem 5.2 and the paragraph following its proof. Proposition 6.5 with j=0j=0, whose only label is the root mark, gives all finite static jets, and Lemma 7.2 gives all finite projected rows. The standard inversion (8) is a continuous triangular rational map near a target with a10a_1^\ell \ne 0, so the velocity jets converge. The sequence does not depend on NN or L\mathcal L, and every finite collection of raw readout coordinates converges along it, so Lemma 2.3 gives density in the initial topology. The marginal convergence in item 2 gives density in τmovSV\tau _{\rm mov}^{\rm SV}.

Proposition 7.4 (The density does not extend to the wings). For a chain PP and a maturity tt_\ell let βR(P):=lim supkw0,(k)/k\beta _R(P):=\limsup _{k\to \infty }w_{0,\ell }(k)/k be the right-wing slope of the root smile, and p~(P):=sup{u0:EP[e(1+u)X]<}\tilde p(P):=\sup \{u\ge 0:\E _P[e^{(1+u)X_\ell }]<\infty \} the right moment index. If LawQ(X)\Law _Q(X_\ell ) is a finite Gaussian mixture — in particular for every chain in GSV\mathfrak G_{\rm SV} or Glift\mathfrak G_{\rm lift} — then βR(Q)=0\beta _R(Q)=0. If p~(P)<\tilde p(P)<\infty, then

βR(P)=24(p~(P)2+p~(P)p~(P))>0.\beta _R(P)=2-4\Bigl (\sqrt {\tilde p(P)^2+\tilde p(P)}-\tilde p(P)\Bigr )>0.
Consequently, once βR\beta _R is adjoined to the readout, no target with a finite right moment index lies in the closure of either finite class; the left wing and the left moment index behave alike.

Proof.A finite Gaussian mixture has exponential moments of every order, so p~(Q)=\tilde p(Q)=\infty, and Lee’s moment formula [33] gives βR(Q)=0\beta _R(Q)=0. For PP the same formula gives the displayed value, which is positive because p~2+p~<p~+12\sqrt {\tilde p^2+\tilde p}<\tilde p+\tfrac 12 for every finite p~0\tilde p\ge 0. The set {Q:|βR(Q)βR(P)|<12βR(P)}\{Q:|\beta _R(Q)-\beta _R(P)|<\tfrac 12\beta _R(P)\} is then a neighbourhood of PP in the augmented initial topology containing no chain of either class, the finite-mixture property of their marginals being Proposition 2.5.

Proposition 7.5 (Fixed marginals cap the dynamics). Let QQ be a chain as in Proposition 2.5 whose date-11 return marginal Law(X1x0)\Law (X_1-x_0) is

μ1=i=1npiN(yi12s, s),y1<<yn,pi>0,s>0.\mu _1=\sum _{i=1}^np_i\,\N \bigl (y_i-\tfrac 12s,\ s\bigr ), \qquad y_1<\cdots <y_n,\quad p_i>0,\quad s>0.
Then its first edge has variance ss, its first-edge locations take the value yiy_i with probability pip_i, and every projected row is
(30)dm=1Var(μ1)i=1npi(yiy¯)A¯m(i),y¯:=ipiyi,\begin{equation}\label {eq:capped-rows} d_m^\ell =\frac {1}{\Var (\mu _1)}\sum _{i=1}^np_i\,(y_i-\bar y)\,\bar A_m^\ell (i), \qquad \bar y:=\sum _ip_iy_i, \end{equation}
where A¯m(i)\bar A_m^\ell (i) is the average of the continuation jet AmA_m^\ell over the first-edge branches located at yiy_i. In particular, if n=1n=1, every chain in either finite class whose first marginal is Gaussian has dm=0d_m^\ell =0 for all mm and \ell.

Proof.By Proposition 2.5, μ1=ν1N(0,α0)\mu _1=\nu _1*\N (0,\alpha _0) with ν1\nu _1 finitely supported and α0\alpha _0 the first-edge variance, while μ1=(ipiδyis/2)N(0,s)\mu _1=\bigl (\sum _ip_i\delta _{y_i-s/2}\bigr )*\N (0,s). If α0<s\alpha _0<s, then ν1\nu _1 is a Gaussian convolution, which is never finitely supported. If α0>s\alpha _0>s, then ν1^(ξ)=e(α0s)ξ2/2kpkeiξ(yks/2)\widehat {\nu _1}(\xi )=e^{(\alpha _0-s)\xi ^2/2}\sum _kp_ke^{\mathrm i\xi (y_k-s/2)}, which is unbounded, unlike a characteristic function, because the almost periodic sum returns arbitrarily close to its value 11 at the origin. Hence α0=s\alpha _0=s, and uniqueness gives ν1=ipiδyis/2\nu _1=\sum _ip_i\delta _{y_i-s/2}. Writing Λ\Lambda for the first-edge location, X1x0=Λ12s+G0X_1-x_0=\Lambda -\tfrac 12s+G_0 with G0G_0 independent of the label chain, so the numerator of (6) is Cov(Am,Λ)\Cov (A_m^\ell ,\Lambda ); grouping the branches by location, and using ipi(yiy¯)=0\sum _ip_i(y_i-\bar y)=0, gives (30), the denominator being Var(μ1)\Var (\mu _1). If n=1n=1, Λ\Lambda is constant.

Remark 7.6 (How much the cap binds). A general martingale chain with the same two marginals is not capped, since its continuation may depend on where X1X_1 falls inside a component. With lognormal marginals at the SPX one- and four-month at-the-money volatilities of Section 10.2 (n=1n=1), every finite chain has d0=0d_0=0, while martingale couplings of grid discretizations of the same marginals reach d0[0.251,0.243]d_0\in [-0.251,0.243], stably under refinement. At the resolution of the SPX fit (n=36n=36) the cap is mild but strict: since the readout averages jets after the convex at-the-money inversion, branches sharing a location can move it, and a piecewise-linear relaxation over such branches certifies d0[0.7722,0.1210]d_0\in [-0.7722,0.1210] for the whole finite class, attained to within 31043\cdot 10^{-4}, against [0.838,0.156][-0.838,0.156] for grid couplings. Both contain the value d0=0.057d_0=-0.057 (v0=1.34v_0=1.34) of [18]. Refining the atoms is thus what lets the finite class express dynamics, and the bands of Section 10.2 are computed inside the capped set.

8A convergence rate

Theorem 7.3 produces a sequence, not a budget. Under the same assumptions the construction admits an explicit rate in the number of atoms per node, linear in the quantization error rather than square-root: the exponential barycentre is exact at every resolution, so no accuracy is spent restoring martingality, and the fixed-variance factor makes the readout a Lipschitz functional of the residual law. Throughout, ERdE\subset \R ^d has diameter DD, LL is the constant in (15), q>p>2q>p>2 and CC are those of (16), and α:=minjαj>0\alpha :=\min _j\alpha _j>0.

Lemma 8.1 (Quantization at a rate). Let HH be as in Lemma 4.1, with eq||HC\int e^{q|\ell |}H\le C. For every δ(0,e1/2]\delta \in (0,e^{-1/2}] there is a finitely supported Hδ=rprδ(λr,yr)H^\delta =\sum _{r}p_r\delta _{(\lambda _r,y_r)} with yrEy_r\in E, with at most

(31)Mc(p,q,d)(1+D/δ)dδ1(1+log(1/δ))\begin{equation}\label {eq:atom-count} M\le c(p,q,d)\,(1+D/\delta )^d\,\delta ^{-1}\bigl (1+\log (1/\delta )\bigr ) \end{equation}
atoms, with exact barycentre rpreλr=1\sum _rp_re^{\lambda _r}=1 and two-sided exponential bound rpreq|λr|2C\sum _rp_re^{q|\lambda _r|}\le 2C, such that
(32)Wp(Hδ,H)c(p,q,C)δ(1+log(1/δ)).\begin{equation}\label {eq:quant-rate} W_p(H^\delta ,H)\le c(p,q,C)\,\delta \,\bigl (1+\log (1/\delta )\bigr ). \end{equation}
Equivalently, in terms of the atom count,
(33)ε(M):=Wp(Hδ(M),H)cM1/(d+1)(logM)(d+2)/(d+1).\begin{equation}\label {eq:quant-rate-M} \varepsilon (M):=W_p(H^{\delta (M)},H)\le c\,M^{-1/(d+1)}(\log M)^{(d+2)/(d+1)}. \end{equation}

Proof.Let (L,Y)H(L,Y)\sim H and put R:=(2p/q)log(1/δ)R:=(2p/q)\log (1/\delta ), so that qRpqR\ge p. Partition R\R into 2R/δ\lceil 2R/\delta \rceil intervals of length at most δ\delta covering [R,R][-R,R] and the two tails {>R}\{\ell >R\}, {<R}\{\ell <-R\}, and EE into at most cd(1+D/δ)dc_d(1+D/\delta )^d sets of diameter at most δ\delta; the cells are all products of the two partitions, tails included, and number at most (31). On a cell AA with H(A)>0H(A)>0 set λA=logE[eLA]\lambda _A=\log \E [e^L\mid A] and yA=E[YA]Ey_A=\E [Y\mid A]\in E. The map (L,Y)(λA,yA)(L,Y)\mapsto (\lambda _A,y_A) on AA is a coupling of HH with Hδ:=Law(λA,yA)H^\delta :=\Law (\lambda _A,y_A); the tower property gives the exact barycentre and the conditional Jensen estimates of Lemma 4.1 give the exponential bound.

For the cost, ulogE[eu]u\mapsto \log \E [e^u] is monotone, so λA\lambda _A lies between the essential infimum and supremum of LL on AA; hence |LλA|δ|L-\lambda _A|\le \delta on every core cell, and |YyA|δ|Y-y_A|\le \delta on every cell, contributing at most (2δ)p(2\delta )^p. On a tail cell, write P:=H(A)CeqRP:=H(A)\le Ce^{-qR}. From xp(2p/(eq))peqx/2x^p\le (2p/(eq))^pe^{qx/2} for x0x\ge 0,

(34)E[|L|p1A](2p/(eq))pE[eq|L|/21A](2p/(eq))pCeqR/2.\begin{equation}\label {eq:tail-moment} \E \bigl [|L|^p\mathbf 1_A\bigr ] \le (2p/(eq))^p\,\E \bigl [e^{q|L|/2}\mathbf 1_A\bigr ] \le (2p/(eq))^p\,C\,e^{-qR/2}. \end{equation}
On the lower tail, Jensen gives λAE[LA]\lambda _A\ge \E [L\mid A] while λAR<0\lambda _A\le -R<0, so |λA|pPE[|L|p1A]|\lambda _A|^pP\le \E [|L|^p\mathbf 1_A]. On the upper tail, Hölder gives E[eL1A]C1/qP11/q\E [e^L\mathbf 1_A]\le C^{1/q}P^{1-1/q}, hence 0<λAq1log(C/P)0<\lambda _A\le q^{-1}\log (C/P); the map xxlog(C/x)px\mapsto x\log (C/x)^p is increasing on (0,Cep)(0,Ce^{-p}) and PCeqRCepP\le Ce^{-qR}\le Ce^{-p}, so |λA|pPqpCeqR(qR)p=CRpeqR|\lambda _A|^pP\le q^{-p}Ce^{-qR}(qR)^p=CR^pe^{-qR}. Since eqR/2=δpe^{-qR/2}=\delta ^p, both tail terms are O(δp(1+log(1/δ))p)O\bigl (\delta ^p(1+\log (1/\delta ))^p\bigr ). Taking pp-th roots gives (32). Finally, choosing δ(logM/M)1/(d+1)\delta \asymp (\log M/M)^{1/(d+1)} in (31) and substituting into (32) gives (33).

The second ingredient replaces the splitting at a scale δ\delta in the proof of Lemma 6.4, which would cost a square root if optimized: with the factor present, the ATM tail probability is the expectation of a Lipschitz function of the residual.

Lemma 8.2 (Linear readout response). Fix α>0\alpha >0 and let R=Z+GR=Z+G and R~=Z~+G\widetilde R=\widetilde Z+G with GγαG\sim \gamma _\alpha independent of ZZ and Z~\widetilde Z. Write C(k)=E[(eRek)+]C(k)=\E [(e^R-e^k)^+] and C~(k)=E[(eR~ek)+]\widetilde C(k)=\E [(e^{\widetilde R}-e^k)^+], and let ε\varepsilon be the LpL^p cost of a coupling of ZZ and Z~\widetilde Z, p>2p>2. If E[epZ]E[epZ~]K\E [e^{p'Z}]\vee \E [e^{p'\widetilde Z}]\le K with p=p/(p1)p'=p/(p-1), then

|C~(0)C(0)|2K1/pε,|C~(i)(0)C(i)(0)|ci(α1/2αi/2)ε(i1),\bigl |\widetilde C(0)-C(0)\bigr |\le 2K^{1/p'}\varepsilon , \qquad \bigl |\widetilde C^{(i)}(0)-C^{(i)}(0)\bigr |\le c_i\,\bigl (\alpha ^{-1/2}\vee \alpha ^{-i/2}\bigr )\,\varepsilon \quad (i\ge 1),
with cic_i absolute.

Proof.The level bound is |euev||uv|(eu+ev)|e^u-e^v|\le |u-v|(e^u+e^v) and Hölder, as in Lemma 6.4. For i=1i=1, C(k)=ekP(R>k)C'(k)=-e^k\mathbb P(R>k), so

(35)C(0)=P(Z+G>0)=E[Φ(Zα/2α)],\begin{equation}\label {eq:atm-lipschitz} C'(0)=-\mathbb P(Z+G>0)=-\E \left [\Phi \left (\frac {Z-\alpha /2}{\sqrt \alpha }\right )\right ], \end{equation}
because GN(α/2,α)G\sim \N (-\alpha /2,\alpha ). The map zΦ((zα/2)/α)z\mapsto \Phi ((z-\alpha /2)/\sqrt \alpha ) is Lipschitz with constant (2πα)1/2(2\pi \alpha )^{-1/2}, so the difference of the two expectations is at most (2πα)1/2W1(Z~,Z)(2πα)1/2ε(2\pi \alpha )^{-1/2}W_1(\widetilde Z,Z)\le (2\pi \alpha )^{-1/2}\varepsilon. For i2i\ge 2, the identity C(i)=C(i1)+ki2(ekf(k))C^{(i)}=C^{(i-1)}+\partial _k^{i-2}(e^kf(k)) expresses C(i)(0)C^{(i)}(0) through C(0)C'(0) and the values f(j)(0)f^{(j)}(0), ji2j\le i-2, where ff is the density of RR; the same holds for C~\widetilde C. Each return density is a residual law convolved with φα\varphi _\alpha, so Lemma 6.2 gives f~(j)f(j)φα(j+1)ε=cjα(j+2)/2ε\|\widetilde f^{(j)}-f^{(j)}\|_\infty \le \|\varphi _\alpha ^{(j+1)}\|_\infty \varepsilon =c_j\alpha ^{-(j+2)/2}\varepsilon; with the α1/2\alpha ^{-1/2} of C(0)C'(0) the largest power is α1/2αi/2\alpha ^{-1/2}\vee \alpha ^{-i/2}, attained at j=i2j=i-2 when α1\alpha \le 1.

Theorem 8.3 (Rate for the homogeneous class). Let PTSVP\in \mathfrak T_{\rm SV}, let NN be finite and L{2,,J}\mathcal L\subseteq \{2,\ldots ,J\} finite. For M2M\ge 2 let PMP^M be the finite marked GM chain of Theorem 5.2 built from Lemma 8.1 with at most MM atoms per node. Then every PMP^M is an exact asset martingale, Law(XjM)\Law (X_j^M) is a Gaussian mixture of at most MjM^j components, and

(36)RN,Lraw(PM)RN,Lraw(P)CM1/(d+1)(logM)(d+2)/(d+1),\begin{equation}\label {eq:rate} \bigl \|\Rcal ^{\rm raw}_{N,\mathcal L}(P^M)-\Rcal ^{\rm raw}_{N,\mathcal L}(P)\bigr \|_\infty \le C_*\,M^{-1/(d+1)}(\log M)^{(d+2)/(d+1)}, \end{equation}
where
C=c(N,p,q,C,D,d,w,w)(1+L)J(1+α(N+1)/2)(1+α01).C_*=c\bigl (N,p,q,C,D,d,\underline w,\overline w\bigr )\, (1+L)^J\bigl (1+\alpha ^{-(N+1)/2}\bigr )\bigl (1+\alpha _0^{-1}\bigr ).
On a chart with a10a_1^\ell \ne 0 the velocity jets satisfy the same bound, with a constant depending in addition on |a1|1|a_1^\ell |^{-1}.

Proof.Apply Lemma 8.1 at every selected node with the common budget MM, giving a per-edge error ε=ε(M)\varepsilon =\varepsilon (M) as in (33). The coupling recursion (25) gives maxjJej2ε((1+L)J1)/L\max _{j\le J}e_j\le 2\varepsilon \bigl ((1+L)^J-1\bigr )/L, and 2Jε2J\varepsilon when L=0L=0; the same recursion started from a common label bounds the continuation residuals of Lemma 6.1 by CtreeεC_{\rm tree}\varepsilon, uniformly over labels. The order-NN readout needs C(i)(0)C^{(i)}(0) only for iN+1i\le N+1, through aN+1a_{N+1}^\ell, and Lemma 8.2 converts this into maxiN+1|Cn(i)(0)C(i)(0)|c(1+α(N+1)/2)Ctreeε\max _{i\le N+1}|C_n^{(i)}(0)-C^{(i)}(0)|\le c\,(1+\alpha ^{-(N+1)/2})C_{\rm tree}\varepsilon, uniformly over reachable states, the exponential moments being uniform by the tower-property estimate in Proposition 6.5. By Assumption 6.3 the target’s ATM variances lie in a compact subset [w,w][\underline w,\overline w] of (0,)(0,\infty ). The approximants’ are at least α\alpha, by Jensen’s inequality in the independent factor, and bounded above through the exponential moments, so for every MM all of them lie in one compact window, on which the jet map of Lemma 2.1 is CC^\infty with denominator bounded away from zero; it is therefore Lipschitz in the call-derivative vector, which gives (36) for the static jets. For the rows, the quantization coupling is explicit, so am,n1,(Y1n)am1,(Y1)Lp\|a_{m,n}^{1,\ell }(Y_1^n)-a_m^{1,\ell }(Y_1)\|_{L^{p'}} is bounded by the uniform level error plus CtreeεC_{\rm tree}\varepsilon through the Lipschitz label dependence of Lemma 6.1; Hölder against L0nL0L_0^n-L_0 and Var(X1X0)α0\Var (X_1-X_0)\ge \alpha _0 then give the same bound for dm,ndmd_{m,n}^\ell -d_m^\ell through (29). The triangular inversion (8) is smooth near a target with a10a_1^\ell \ne 0, which transfers the rate to the velocity jets.

Proposition 8.4 (Lower bounds). Fix p1p\ge 1 and use on R×E\R \times E the metric |¯|+|yy¯||\ell -\bar \ell |+|y-\bar y|.

  1. If HP(R×E)H\in \Pp (\R \times E) has a density bounded below by ρ>0\rho >0 on a cube of side ss, then every law H~\widetilde H with at most MM atoms satisfies Wp(H~,H)cM1/(d+1)W_p(\widetilde H,H)\ge c\,M^{-1/(d+1)}, with c=c(ρ,s,d,p)>0c=c(\rho ,s,d,p)>0.
  2. If LawP(Y1)\Law _P(Y_1) has a density bounded below by ρ>0\rho >0 on a cube of side ss in EE, then every chain QQ whose date-11 mark takes at most MM values satisfies

    AWp(P,Q)Wp(LawP(Y1),LawQ(Y1))cM1/d.\AW _p(P,Q)\ge W_p\bigl (\Law _P(Y_1),\Law _Q(Y_1)\bigr )\ge c\,M^{-1/d}.

Proof.Both parts are one volume argument in dimension d=d+1d'=d+1, respectively d=dd'=d. Balls of radius rr about the MM atoms cover at most half of the cube when Mvdrd12sdMv_{d'}r^{d'}\le \frac 12s^{d'}, where vdv_{d'} is the volume of the unit ball; the uncovered half carries mass at least 12ρsd\frac 12\rho s^{d'} at distance at least rr from every atom, and every coupling must move that mass onto the atoms, so Wpp12ρsdrpW_p^p\ge \frac 12\rho s^{d'}r^p. Take r=s(2vdM)1/dr=s(2v_{d'}M)^{-1/d'}. For (2), the adapted distance dominates the Wasserstein distance of the date-11 marginals, and the mark marginal of QQ has at most MM atoms.

Remark 8.5 (What the rate does and does not say). By Proposition 8.4(1) the quantization exponent 1/(d+1)1/(d+1) is optimal up to the logarithm, so exact martingality costs nothing in it. When the date-11 mark law has a density bounded below on a cube, part (2), the recursion (25) and (33) give

cM1/dAWp(PM,P)CM1/(d+1)(logM)(d+2)/(d+1),c\,M^{-1/d}\le \AW _p(P^M,P)\le C\,M^{-1/(d+1)}(\log M)^{(d+2)/(d+1)},
the lower bound holding for every finite chain, martingale or not; the exponents differ by one because the return coordinate is smoothed rather than quantized, and whether a fixed factor closes the gap is open. Neither bound concerns the readout alone: Proposition 10.1 matches the order-NN readout of one edge exactly with at most 3N+83N+8 components, so a readout lower bound would have to come from the multi-step recursion, and we have none. The constant of Theorem 8.3 degrades like α(N+1)/2\alpha ^{-(N+1)/2}, so a high-order readout needs a factor that is not small relative to the residual scale, and the bound is per node, the date-jj marginal having up to MjM^j components.

9Local stochastic volatility

Local volatility rests on mimicking [2815]; the targets here are grid skeletons of local-stochastic-volatility models, whose return kernel depends on the log price and which supply no Gaussian factor. A finite proxy log price records where the local kernel is sampled, while the traded log price receives independent mean-one Gaussian innovations of vanishing variance; the class contains the homogeneous one (Remark 9.14).

9.1Local-SV kernels and readout regularity

A motivating diffusion is

(37)dXt=12σ(t,Xt,Yt)2dt+σ(t,Xt,Yt)(ρ(t,Yt)dWt+1ρ(t,Yt)2dBt),\begin{equation}\label {eq:lsv-X} dX_t=-\tfrac 12\sigma (t,X_t,Y_t)^2dt +\sigma (t,X_t,Y_t)\left ( \rho (t,Y_t)^\top dW_t+ \sqrt {1-\|\rho (t,Y_t)\|^2}\,dB_t \right ), \end{equation}

with YY as in (18), σ>0\sigma >0, ρ<1\|\rho \|<1, and BB independent of WW. Conditioning on the WW-path now leaves a state-dependent diffusion, so the argument of Proposition 3.3 does not apply, and we work directly with grid kernels.

Put Z:=R×E\mathsf Z:=\R \times E with dZ((x,y),(x¯,y¯)):=|xx¯|+|yy¯|d_{\mathsf Z}((x,y),(\bar x,\bar y)):=|x-\bar x|+|y-\bar y|, and for z=(x,y)z=(x,y) let Hj(z;du,dy):=Lawz(Xj+1Xj,Yj+1)H_j(z;du,dy'):=\Law _z(X_{j+1}-X_j,Y_{j+1}).

Assumption 9.1 (Local-SV Markov kernels). Fix p>2p>2 and q>pq>p. The state space ERdE\subset \R ^d is compact and convex, the root z0=(x0,y0)z_0=(x_0,y_0) is known, and (Xj,Yj)j=0J(X_j,Y_j)_{j=0}^J is a time-inhomogeneous Markov chain with kernels HjH_j. There are L,C<L,C<\infty such that, for every j<Jj<J and z,z¯Zz,\bar z\in \mathsf Z,

(38)euHj(z;du,dy)=1,(39)Wp(Hj(z),Hj(z¯))LdZ(z,z¯),(40)supzZeq|u|Hj(z;du,dy)C.\begin{align} \int e^uH_j(z;du,dy')&=1, \label {eq:lsv-barycentre}\\ W_p(H_j(z),H_j(\bar z))&\le Ld_{\mathsf Z}(z,\bar z), \label {eq:lsv-kernel-Lip}\\ \sup _{z\in \mathsf Z}\int e^{q|u|}H_j(z;du,dy')&\le C. \label {eq:lsv-kernel-exp} \end{align}

Assumption 9.2 (All-order terminal-kernel regularity). For every j<Jj<J, the return marginal has a density

(41)Hj(z;du,E)=qj(u;z)du.\begin{equation}\label {eq:lsv-one-edge-density} H_j(z;du,E)=q_j(u;z)\,du. \end{equation}
For each integer r0r\ge 0, (u,z)urqj(u;z)(u,z)\mapsto \partial _u^rq_j(u;z) is bounded and uniformly continuous on R×Z\R \times \mathsf Z. Equivalently, for suitable constants Mj,r<M_{j,r}<\infty and moduli ωj,r(s)0\omega _{j,r}(s)\downarrow 0 as s0s\downarrow 0,
(42)supu,z|urqj(u;z)|Mj,r,(43)|urqj(u;z)urqj(u¯;z¯)|ωj,r(|uu¯|+dZ(z,z¯)).\begin{align} \sup _{u,z}|\partial _u^rq_j(u;z)|&\le M_{j,r}, \label {eq:lsv-density-bound}\\ |\partial _u^rq_j(u;z)-\partial _u^rq_j(\bar u;\bar z)| &\le \omega _{j,r}(|u-\bar u|+d_{\mathsf Z}(z,\bar z)). \label {eq:lsv-density-modulus} \end{align}

Assumption 9.3 (Local-SV readout window). For every date pair used by the readout, the target ATM implied total variance satisfies

(44)0<wwj,(0;z)w<(zZ),\begin{equation}\label {eq:lsv-iv-window} 0<\underline w\le w_{j,\ell }(0;z)\le \overline w<\infty \qquad (z\in \mathsf Z), \end{equation}
and
(45)Varz0(X1X0)>0.\begin{equation}\label {eq:lsv-positive-variance} \Var _{z_0}(X_1-X_0)>0. \end{equation}

For (37), bounded volatility gives the exponential moments and Lipschitz coefficients give (39); Assumption 9.2 holds for the following class.

Proposition 9.4 (A class satisfying Assumptions 9.1–9.3). Fix ν>0\nu >0, let BB' be a Brownian motion independent of (W,B)(W,B), and write η(t):=tj\eta (t):=t_j for t[tj,tj+1)t\in [t_j,t_{j+1}). Consider the local-SV model with an independent variance floor and edge-frozen leverage,

(46)dXt=12(σt2+ν2)dt+σt(ρ(t,Yt)dWt+1ρ2dBt)+νdBt,σt=σ(t,Xη(t),Yt),\begin{equation}\label {eq:lsv-floor-model} dX_t=-\tfrac 12\bigl (\sigma _t^2+\nu ^2\bigr )dt +\sigma _t\bigl (\rho (t,Y_t)^\top dW_t+\sqrt {1-\|\rho \|^2}\,dB_t\bigr ) +\nu \,dB_t', \qquad \sigma _t=\sigma (t,X_{\eta (t)},Y_t), \end{equation}
with YY as in (18) and valued in a compact convex EE (for instance by the boundary conditions of Proposition 3.4), and σ,ρ,b,Γ\sigma ,\rho ,b,\Gamma bounded and globally Lipschitz with ρρ¯<1\|\rho \|\le \bar \rho <1. Then the grid chain (Xj,Yj)(X_j,Y_j) is Markov and satisfies Assumptions 9.1, 9.2 and 9.3, with Δj=tj+1tj\Delta _j=t_{j+1}-t_j and
Mj,r=φν2Δj(r),ωj,r(s)=φν2Δj(r+1)(1+L)s,M_{j,r}=\bigl \|\varphi _{\nu ^2\Delta _j}^{(r)}\bigr \|_\infty , \qquad \omega _{j,r}(s)=\bigl \|\varphi _{\nu ^2\Delta _j}^{(r+1)}\bigr \|_\infty (1+L)s,
where LL is any constant realizing the synchronous-coupling estimate (39).

Proof.On each edge the leverage reads the frozen spot XtjX_{t_j} and the WW-driven mark path, so from the current state z=(x,y)z=(x,y) the increment splits as Xj+1Xj=Rj+NjX_{j+1}-X_j=R_j+N_j with

Nj:=12ν2Δj+ν(Btj+1Btj)γν2Δj,Rj:=12tjtj+1σt2dt+tjtj+1σt(ρdWt+1ρ2dBt).\begin{align*} N_j&:=-\tfrac 12\nu ^2\Delta _j+\nu \,(B'_{t_{j+1}}-B'_{t_j}) \sim \gamma _{\nu ^2\Delta _j},\\ R_j&:=-\tfrac 12\int _{t_j}^{t_{j+1}}\sigma _t^2\,dt +\int _{t_j}^{t_{j+1}}\sigma _t\bigl (\rho ^\top dW_t+\sqrt {1-\|\rho \|^2}\,dB_t\bigr ). \end{align*}

With the leverage frozen at xx, RjR_j depends on zz and the (W,B)(W,B) increments on the edge, and NjN_j on the BB' increment only. Post-tjt_j increments of (W,B)(W,B) and of BB' are independent of each other and of Ftj\mathcal F_{t_j}, so RjNjR_j\perp N_j given the state and

Hj(z;du,E)=Lawz(Rj)γν2Δj,urqj(;z)=Lawz(Rj)φν2Δj(r),H_j(z;du,E)=\Law _z(R_j)*\gamma _{\nu ^2\Delta _j}, \qquad \partial _u^rq_j(\cdot \,;z)=\Law _z(R_j)*\varphi ^{(r)}_{\nu ^2\Delta _j},
which gives (42). For the modulus, the mean value theorem bounds the uu-increment by φν2Δj(r+1)|uu¯|\|\varphi _{\nu ^2\Delta _j}^{(r+1)}\|_\infty |u-\bar u|, and Lemma 6.2 gives
|urqj(u;z)urqj(u;z¯)|φν2Δj(r+1)W1(Lawz(Rj),Lawz¯(Rj)).|\partial _u^rq_j(u;z)-\partial _u^rq_j(u;\bar z)| \le \|\varphi _{\nu ^2\Delta _j}^{(r+1)}\|_\infty W_1\bigl (\Law _z(R_j),\Law _{\bar z}(R_j)\bigr ).
Driving copies from z,z¯z,\bar z by the same (W,B,B)(W,B,B') makes the NjN_j equal, and Gronwall, with 1ρ2\sqrt {1-\|\rho \|^2} Lipschitz because ρρ¯<1\|\rho \|\le \bar \rho <1, bounds the LpL^p distance of the return–mark pairs by a multiple of dZ(z,z¯)d_{\mathsf Z}(z,\bar z); this gives (39) and, adding the two moduli, (43). Since σ\sigma is bounded, Novikov’s condition gives Ez[eRj]=1\E _z[e^{R_j}]=1, so independence gives (38), and bounded coefficients give (40).

For Assumption 9.3, with σ¯:=supσ\bar \sigma :=\sup \sigma the instantaneous variance lies in [ν2,σ¯2+ν2][\nu ^2,\bar \sigma ^2+\nu ^2]; since a call is convex, [21] places the conditional call price between the Black–Scholes prices at these volatilities, and monotonicity of B(0,)B(0,\cdot ) places the ATM total variance in [ν2(ttj),(σ¯2+ν2)(ttj)][\nu ^2(t_\ell -t_j),(\bar \sigma ^2+\nu ^2)(t_\ell -t_j)], and Varz0(X1X0)Var(N0)=ν2Δ0>0\Var _{z_0}(X_1-X_0)\ge \Var (N_0)=\nu ^2\Delta _0>0.

Remark 9.5 (Why the leverage is frozen). Read continuously, σt=σ(t,Xt,Yt)\sigma _t=\sigma (t,X_t,Y_t) depends on the past of BB' within the edge, so RjR_j and NjN_j need not be independent and the convolution identity is lost; Lipschitz coefficients alone need not give bounded density derivatives of every order. The alternative is parabolic: smooth coefficients and a uniformly elliptic diffusion matrix for (X,Y)(X,Y) give Assumption 9.2 by [2232], but a chain invariant on a compact mark space degenerates at E\partial E, so we rely on Proposition 9.4 or on Assumption 9.2 directly.

For the LSV target the smiles, jets and rows are those of Section 2 with ζj=(Xj,Yj)\zeta _j=(X_j,Y_j): Cj,(k;x,y)=Ex,y[(eXXjek)+]C_{j,\ell }(k;x,y)=\E _{x,y}[(e^{X_\ell -X_j}-e^k)^+], am=am0,(x0,y0)a_m^\ell =a_m^{0,\ell }(x_0,y_0), Am=am1,(X1,Y1)A_m^\ell =a_m^{1,\ell }(X_1,Y_1), and dmd_m^\ell is the covariance ratio (6).

9.2Finite proxy-GM chains

Definition 9.6 (Finite proxy-GM chain). At date tjt_j let QjnZ\mathcal Q_j^n\subset \mathsf Z be finite and write a proxy label as ζ=(ξ,y)\zeta =(\xi ,y). At each ζQjn\zeta \in \mathcal Q_j^n choose positive weights and successor data

pj,ζ,rn,(λj,ζ,rn,yj,ζ,rn,+)R×E,rpj,ζ,rn=1,p_{j,\zeta ,r}^n, \qquad (\lambda _{j,\zeta ,r}^n,y_{j,\zeta ,r}^{n,+})\in \R \times E, \qquad \sum _rp_{j,\zeta ,r}^n=1,
such that
(47)rpj,ζ,rneλj,ζ,rn=1.\begin{equation}\label {eq:proxy-barycentre} \sum _rp_{j,\zeta ,r}^ne^{\lambda _{j,\zeta ,r}^n}=1. \end{equation}
The successor label is
ζj,ζ,rn,+:=(ξ+λj,ζ,rn,yj,ζ,rn,+)Qj+1n.\zeta _{j,\zeta ,r}^{n,+} :=(\xi +\lambda _{j,\zeta ,r}^n,y_{j,\zeta ,r}^{n,+})\in \mathcal Q_{j+1}^n.
For an edge variance βj,n>0\beta _{j,n}>0, the transition from the full approximating state (x^,ζ)(\widehat x,\zeta ) is
(48)Kj,npr((x^,ζ),dx^,dζ)=rpj,ζ,rnN(x^+λj,ζ,rnβj,n2,βj,n)(dx^)δζj,ζ,rn,+(dζ).\begin{align} &K_{j,n}^{\rm pr}((\widehat x,\zeta ),d\widehat x',d\zeta ') \label {eq:proxy-kernel}\\ &\quad =\sum _rp_{j,\zeta ,r}^n \N \left (\widehat x+\lambda _{j,\zeta ,r}^n-\frac {\beta _{j,n}}2, \beta _{j,n}\right )(d\widehat x') \delta _{\zeta _{j,\zeta ,r}^{n,+}}(d\zeta '). \nonumber \end{align}

The root is (X^0n,Ξ0n,Y^0n)=(x0,x0,y0)(\widehat X_0^n,\Xi _0^n,\widehat Y_0^n)=(x_0,x_0,y_0). The proxy coordinate Ξjn\Xi _j^n is part of the finite continuation label; the traded log price is X^jn\widehat X_j^n. Let F^j\widehat {\mathcal F}_j be the natural filtration of (X^hn,Ξhn,Y^hn)hj(\widehat X_h^n,\Xi _h^n,\widehat Y_h^n)_{h\le j}.

The normalized future return law depends only on the label ζ\zeta, so the approximating smiles and jets are written Cj,,n(k;ζ)C_{j,\ell ,n}(k;\zeta ) and am,nj,(ζ)a_{m,n}^{j,\ell }(\zeta ). Equivalently, after choosing branch rr at ζ=(ξ,y)\zeta =(\xi ,y), draw Gj,nγβj,nG_{j,n}\sim \gamma _{\beta _{j,n}}, independently of the complete proxy-label chain and of the other Gaussian innovations, and set

(49)Ξj+1n=Ξjn+λj,ζ,rn,Y^j+1n=yj,ζ,rn,+,X^j+1n=X^jn+λj,ζ,rn+Gj,n.\begin{equation}\label {eq:proxy-recursion} \Xi _{j+1}^n=\Xi _j^n+\lambda _{j,\zeta ,r}^n, \quad \widehat Y_{j+1}^n=y_{j,\zeta ,r}^{n,+}, \quad \widehat X_{j+1}^n=\widehat X_j^n+\lambda _{j,\zeta ,r}^n+G_{j,n}. \end{equation}

Lemma 9.7 (Exact martingality and finite marginals). Every finite proxy-GM chain is an exact asset martingale. Put Bj,n:=h<jβh,nB_{j,n}:=\sum _{h<j}\beta _{h,n}, with the convention γ0=δ0\gamma _0=\delta _0. Then

(50)X^jnΞjn=h<jGh,nγBj,n,\begin{equation}\label {eq:proxy-gap} \widehat X_j^n-\Xi _j^n=\sum _{h<j}G_{h,n}\sim \gamma _{B_{j,n}}, \end{equation}
independently of (Ξjn,Y^jn)(\Xi _j^n,\widehat Y_j^n). Consequently
(51)Law(X^jn)=(ξ,y)QjnP((Ξjn,Y^jn)=(ξ,y))N(ξBj,n2,Bj,n).\begin{equation}\label {eq:proxy-marginal} \Law (\widehat X_j^n) =\sum _{(\xi ,y)\in \mathcal Q_j^n} \mathbb P((\Xi _j^n,\widehat Y_j^n)=(\xi ,y)) \N \left (\xi -\frac {B_{j,n}}2,B_{j,n}\right ). \end{equation}
Thus every positive-date log-price marginal is a finite Gaussian mixture, with at most |Qjn||\mathcal Q_j^n| components.

Proof.Conditioning on the current state and using (47),

E[eX^j+1nF^j]=eX^jnrpj,ζ,rneλj,ζ,rnE[eGj,n]=eX^jn.\E [e^{\widehat X_{j+1}^n}\mid \widehat {\mathcal F}_j] =e^{\widehat X_j^n} \sum _rp_{j,\zeta ,r}^ne^{\lambda _{j,\zeta ,r}^n}\E [e^{G_{j,n}}] =e^{\widehat X_j^n}.
Subtracting the proxy recursion from the traded-price recursion in (49) gives (50). By the independent product construction in (49), the complete proxy-label chain is independent of (Gh,n)h<J(G_{h,n})_{h<J}. Conditioning on the finitely many proxy labels gives (51).

9.3Diagonal quantization and lifted adapted convergence

Let φβ\varphi _\beta denote the density of γβ\gamma _\beta.

Lemma 9.8 (Diagonal martingale smoothing). Let fz(u)duf_z(u)\,du, zZz\in \mathsf Z, be probability laws whose derivatives through order RR are bounded and uniformly continuous jointly in (u,z)(u,z), and suppose eufz(u)du=1\int e^uf_z(u)\,du=1. Let ΘnZ\Theta _n\subset \mathsf Z be finite and let

ηn,z=rpn,z,rδλn,z,r,zΘn,\eta _{n,z}=\sum _rp_{n,z,r}\delta _{\lambda _{n,z,r}}, \qquad z\in \Theta _n,
satisfy rpn,z,reλn,z,r=1\sum _rp_{n,z,r}e^{\lambda _{n,z,r}}=1 and
supzΘnW1(ηn,z,fz(u)du)εn.\sup _{z\in \Theta _n}W_1(\eta _{n,z},f_z(u)\,du)\le \varepsilon _n.
If βn0\beta _n\downarrow 0 and
(52)εnmax0iRφβn(i+1)0,\begin{equation}\label {eq:lsv-diagonal-rate} \varepsilon _n\max _{0\le i\le R}\|\varphi _{\beta _n}^{(i+1)}\|_\infty \longrightarrow 0, \end{equation}
then gn,z:=ηn,zφβng_{n,z}:=\eta _{n,z}*\varphi _{\beta _n} has unit exponential mean and
(53)max0iRsupzΘnuign,zuifz0.\begin{equation}\label {eq:lsv-diagonal-Cr} \max _{0\le i\le R}\sup _{z\in \Theta _n} \|\partial _u^ig_{n,z}-\partial _u^if_z\|_\infty \longrightarrow 0. \end{equation}

Proof.The exponential mean factorizes:

eugn,z(u)du=(rpn,z,reλn,z,r)E[eGβn]=1.\int e^ug_{n,z}(u)\,du =\left (\sum _rp_{n,z,r}e^{\lambda _{n,z,r}}\right )\E [e^{G_{\beta _n}}]=1.
Lemma 6.2 gives, for iRi\le R,
ui(ηn,zφβnfzφβn)φβn(i+1)εn.\|\partial _u^i(\eta _{n,z}*\varphi _{\beta _n} -f_z*\varphi _{\beta _n})\|_\infty \le \|\varphi _{\beta _n}^{(i+1)}\|_\infty \varepsilon _n.
Moreover (fzφβn)(i)=(uifz)φβn(f_z*\varphi _{\beta _n})^{(i)}=(\partial _u^if_z)*\varphi _{\beta _n}, and the common modulus of continuity gives
supz(uifz)φβnuifzE[ωi(|Gβn|)]0.\sup _z\|(\partial _u^if_z)*\varphi _{\beta _n}-\partial _u^if_z\|_\infty \le \E [\omega _i(|G_{\beta _n}|)]\longrightarrow 0.
The limit follows by bounded convergence, the modulus being bounded. Combining the two bounds proves (53). Since β\beta denotes variance, φβ(i+1)=ciβ(i+2)/2\|\varphi _\beta ^{(i+1)}\|_\infty =c_i\beta ^{-(i+2)/2}; thus εn=o(βn(R+2)/2)\varepsilon _n=o(\beta _n^{(R+2)/2}) is a sufficient explicit rate.

Equip the lifted state space R×(R×E)\R \times (\R \times E) with the metric d((x,(ξ,y)),(x¯,(ξ¯,y¯))):=|xx¯|+|ξξ¯|+|yy¯|d^\uparrow ((x,(\xi ,y)),(\bar x,(\bar \xi ,\bar y))):=|x-\bar x|+|\xi -\bar \xi |+|y-\bar y|, and use it in the corresponding adapted Wasserstein distance.

Theorem 9.9 (Lifted finite-tree approximation). Under Assumptions 9.1 and 9.2, there are integers rnr_n\uparrow \infty, variances βn0\beta _n\downarrow 0, and finite proxy-GM chains PnP^n with βj,n=βn\beta _{j,n}=\beta _n on every edge such that:

  1. every transition is an exact asset-martingale finite-GM transition and every Law(X^jn)\Law (\widehat X_j^n), j1j\ge 1, is a finite Gaussian mixture;
  2. after lifting the target and approximating states to the common space

    Zj=(Xj,(Xj,Yj)),Z^jn=(X^jn,(Ξjn,Y^jn)),Z_j^\uparrow =(X_j,(X_j,Y_j)), \qquad \widehat Z_j^n=(\widehat X_j^n,(\Xi _j^n,\widehat Y_j^n)),
    their path laws converge in AWp\AW _p;
  3. if gj,n(;ζ)g_{j,n}(\cdot \,;\zeta ) is the one-edge return density of (48), then

    (54)maxj<JmaxζQjnmax0irnuigj,n(;ζ)uiqj(;ζ)0.\begin{equation}\label {eq:lsv-node-Cr} \max _{j<J}\max _{\zeta \in \mathcal Q_j^n}\max _{0\le i\le r_n} \|\partial _u^ig_{j,n}(\cdot \,;\zeta ) -\partial _u^iq_j(\cdot \,;\zeta )\|_\infty \longrightarrow 0. \end{equation}

At each date,

(55)Wp(Law(X^jn,Y^jn),Law(Xj,Yj))0,\begin{equation}\label {eq:lsv-marginal-Wp} W_p(\Law (\widehat X_j^n,\widehat Y_j^n),\Law (X_j,Y_j))\longrightarrow 0, \end{equation}
and the spot marginals converge in W1W_1.

Proof.Choose any rnr_n\uparrow \infty. By Assumption 9.2 and the approximate-identity estimate in Lemma 9.8, one may choose βn0\beta _n\downarrow 0 with

(56)maxj<Jmax0irnsupzZ(uiqj(;z))φβnuiqj(;z)n1.\begin{equation}\label {eq:lsv-mollifier-diagonal} \max _{j<J}\max _{0\le i\le r_n}\sup _{z\in \mathsf Z} \|(\partial _u^iq_j(\cdot \,;z))*\varphi _{\beta _n} -\partial _u^iq_j(\cdot \,;z)\|_\infty \le n^{-1}. \end{equation}
Put Dn:=1+max0irnφβn(i+1)D_n:=1+\max _{0\le i\le r_n}\|\varphi _{\beta _n}^{(i+1)}\|_\infty and choose 0<εn(nDn)10<\varepsilon _n\le (nD_n)^{-1}.

Start with Q0n={(x0,y0)}\mathcal Q_0^n=\{(x_0,y_0)\}. Suppose Qjn\mathcal Q_j^n is finite. For every ζ=(ξ,y)Qjn\zeta =(\xi ,y)\in \mathcal Q_j^n, apply Lemma 4.1 to the full joint law Hj(ξ,y)H_j(\xi ,y). Choose

(57)Hj,nat(ζ)=rpj,ζ,rnδ(λj,ζ,rn,yj,ζ,rn,+)\begin{equation}\label {eq:lsv-atomic-kernel} H_{j,n}^{\rm at}(\zeta ) =\sum _rp_{j,\zeta ,r}^n \delta _{(\lambda _{j,\zeta ,r}^n,y_{j,\zeta ,r}^{n,+})} \end{equation}
with Wp(Hj,nat(ζ),Hj(ζ))εnW_p(H_{j,n}^{\rm at}(\zeta ),H_j(\zeta ))\le \varepsilon _n and exact exponential barycentre (47). The conditional-Jensen estimate in Lemma 4.1 gives the uniform qq-exponential bound for the selected return atoms. Define Qj+1n\mathcal Q_{j+1}^n as the finite set of successor labels (ξ+λj,ζ,rn,yj,ζ,rn,+)(\xi +\lambda _{j,\zeta ,r}^n,y_{j,\zeta ,r}^{n,+}). Recursion gives one coherent chain over the entire grid, including all option maturities.

Since W1WpW_1\le W_p and projection onto the return is 11-Lipschitz, the return marginal of (57) is within εn\varepsilon _n in W1W_1 of qj(u;ζ)duq_j(u;\zeta )\,du. Equations (56) and the definition of DnD_n therefore give (54) through the two estimates in the proof of Lemma 9.8, used directly because the order rnr_n grows: the quantization term is maxirnφβn(i+1)εnDnεnn1\max _{i\le r_n}\|\varphi _{\beta _n}^{(i+1)}\|_\infty \varepsilon _n\le D_n\varepsilon _n\le n^{-1} by the choice of εn\varepsilon _n, and the approximate-identity term is at most n1n^{-1} by (56), both already uniform in irni\le r_n, in jj, and in the node. Exact martingality and finite-GM marginals follow from Lemma 9.7.

It remains to prove the adapted convergence. Given a current target state (Xj,Yj)(X_j,Y_j) and proxy label (Ξjn,Y^jn)(\Xi _j^n,\widehat Y_j^n), the triangle inequality and (39) give

(58)Wp(Hj(Xj,Yj),Hj,nat(Ξjn,Y^jn))L(|XjΞjn|+|YjY^jn|)+εn.\begin{equation}\label {eq:lsv-edge-coupling} W_p\bigl (H_j(X_j,Y_j),H_{j,n}^{\rm at}(\Xi _j^n,\widehat Y_j^n)\bigr ) \le L\bigl (|X_j-\Xi _j^n|+|Y_j-\widehat Y_j^n|\bigr )+\varepsilon _n. \end{equation}
Choose a Borel measurable εn\varepsilon _n-optimal coupling kernel, which exists by measurable selection of optimal plans [36, Corollary 5.22], the map zHj(z)z\mapsto H_j(z) being continuous into Pp\Pp _p by (39). Denote the coupled target return–mark by (Uj,Yj+1)(U_j,Y_{j+1}) and the proxy atom by (Λjn,Y^j+1n)(\Lambda _j^n,\widehat Y_{j+1}^n). Set
Xj+1=Xj+Uj,Ξj+1n=Ξjn+Λjn,X_{j+1}=X_j+U_j, \qquad \Xi _{j+1}^n=\Xi _j^n+\Lambda _j^n,
and independently add Gj,nγβnG_{j,n}\sim \gamma _{\beta _n} to obtain X^j+1n\widehat X_{j+1}^n as in (49). The joint one-step kernel has first marginal Hj(Xj,Yj)H_j(X_j,Y_j) and second marginal Kj,npr((X^jn,(Ξjn,Y^jn)),)K_{j,n}^{\rm pr}((\widehat X_j^n,(\Xi _j^n,\widehat Y_j^n)),\cdot ), each depending only on its own current state. Recursion therefore gives a bicausal coupling.

With ej,n:=|XjΞjn|+|YjY^jn|Lpe_{j,n}:=\bigl \||X_j-\Xi _j^n|+|Y_j-\widehat Y_j^n|\bigr \|_{L^p}, the argument of Theorem 5.2 with (58) in place of (24) gives the recursion (25) from e0,n=0e_{0,n}=0, the roots coinciding, hence maxjJej,nCJ,Lεn\max _{j\le J}e_{j,n}\le C_{J,L}\varepsilon _n. By (50),

X^jnΞjnLpcpJβn+12Jβn.\|\widehat X_j^n-\Xi _j^n\|_{L^p} \le c_p\sqrt {J\beta _n}+\frac 12J\beta _n.
For the lifted-state distance,
|X^jnXj|+|ΞjnXj|+|Y^jnYj|2(|XjΞjn|+|YjY^jn|)+|X^jnΞjn|.|\widehat X_j^n-X_j|+|\Xi _j^n-X_j|+|\widehat Y_j^n-Y_j| \le 2\bigl (|X_j-\Xi _j^n|+|Y_j-\widehat Y_j^n|\bigr )+|\widehat X_j^n-\Xi _j^n|.
Summing over the finite grid gives the quantitative bound
(59)AWp(Law(Z^n),Law(Z))Cp,J,L(εn+Jβn+Jβn),\begin{equation}\label {eq:lsv-AW-bound} \AW _p(\Law (\widehat Z^n),\Law (Z^\uparrow )) \le C_{p,J,L}(\varepsilon _n+\sqrt {J\beta _n}+J\beta _n), \end{equation}
and hence adapted convergence. Projection by (x,(ξ,y))(x,y)(x,(\xi ,y))\mapsto (x,y) is 11-Lipschitz for dd^\uparrow, which proves (55). In particular X^jnXj\widehat X_j^n\to X_j in law, so eX^jneXje^{\widehat X_j^n}\Rightarrow e^{X_j}; both have first moment ex0e^{x_0}, by exact martingality, and for nonnegative laws weak convergence with converging first moments is W1W_1 convergence.

Since normalized proxy returns depend only on the label, the same recursion can be started from any ζ=(ξ,y)Qjn\zeta =(\xi ,y)\in \mathcal Q_j^n with both log prices normalized to ξ\xi; it gives for the coupled continuations (Xhζ,Yhζ)(X_h^\zeta ,Y_h^\zeta ) and (X^hn,ζ,Ξhn,ζ,Y^hn,ζ)(\widehat X_h^{n,\zeta },\Xi _h^{n,\zeta },\widehat Y_h^{n,\zeta }) the uniform bound

(60)supζQjn(Eζ[h=jJ(|XhζX^hn,ζ|+|XhζΞhn,ζ|+|YhζY^hn,ζ|)p])1/pCj,J(εn+βn+βn).\begin{align} &\sup _{\zeta \in \mathcal Q_j^n} \left (\E _\zeta \left [\sum _{h=j}^J \left ( |X_h^\zeta -\widehat X_h^{n,\zeta }| +|X_h^\zeta -\Xi _h^{n,\zeta }| +|Y_h^\zeta -\widehat Y_h^{n,\zeta }| \right )^p\right ]\right )^{1/p} \label {eq:lsv-uniform-continuation}\\ &\hspace {35mm}\le C_{j,J}(\varepsilon _n+\sqrt {\beta _n}+\beta _n). \nonumber \end{align}

9.4Continuation densities and smile jets

Fix j<j<\ell. Under the target continuation from zZz\in \mathsf Z, write V:=X1XjV:=X_{\ell -1}-X_j and Z1:=(X1,Y1)Z_{\ell -1}:=(X_{\ell -1},Y_{\ell -1}). Conditioning at t1t_{\ell -1} gives the density

(61)fj,(u;z)=Ez[q1(uV;Z1)].\begin{equation}\label {eq:lsv-continuation-density} f_{j,\ell }(u;z) =\E _z[q_{\ell -1}(u-V;Z_{\ell -1})]. \end{equation}

For the proxy continuation from ζ=(ξ,y)\zeta =(\xi ,y), normalize the traded starting log price to ξ\xi and write V^n=X^1nX^jn\widehat V^n=\widehat X_{\ell -1}^n-\widehat X_j^n. Its terminal density is

(62)fj,n(u;ζ)=Eζ[g1,n(uV^n;Ξ1n,Y^1n)].\begin{equation}\label {eq:lsv-proxy-continuation-density} f_{j,\ell }^n(u;\zeta ) =\E _\zeta [g_{\ell -1,n}(u-\widehat V^n; \Xi _{\ell -1}^n,\widehat Y_{\ell -1}^n)]. \end{equation}

Lemma 9.10 (Terminal-kernel transfer). For each fixed R<R<\infty and j<j<\ell,

(63)max0iRmaxζQjnuifj,n(;ζ)uifj,(;ζ)0.\begin{equation}\label {eq:lsv-continuation-Cr} \max _{0\le i\le R}\max _{\zeta \in \mathcal Q_j^n} \|\partial _u^if_{j,\ell }^n(\cdot \,;\zeta ) -\partial _u^if_{j,\ell }(\cdot \,;\zeta )\|_\infty \longrightarrow 0. \end{equation}
For each ii, (u,z)uifj,(u;z)(u,z)\mapsto \partial _u^if_{j,\ell }(u;z) is bounded and uniformly continuous.

Proof.For large nn, rnRr_n\ge R; couple the preterminal continuations by (60). Differentiation under the expectation is justified by dominated convergence, with the bounds M1,iM_{\ell -1,i} of (42) and, for large nn, M1,i+1M_{\ell -1,i}+1 for g1,ng_{\ell -1,n} by (54). Add and subtract uiq1(uV^n;Ξ1n,Y^1n)\partial _u^iq_{\ell -1}(u-\widehat V^n;\Xi _{\ell -1}^n,\widehat Y_{\ell -1}^n). The first difference is bounded uniformly by (54). With Dn,ζ:=|V^nV|+|Ξ1nX1|+|Y^1nY1|D_{n,\zeta }:=|\widehat V^n-V|+|\Xi _{\ell -1}^n-X_{\ell -1}|+|\widehat Y_{\ell -1}^n-Y_{\ell -1}|, the second is bounded, uniformly in uu, by Eζ[ω1,i(Dn,ζ)]\E _\zeta [\omega _{\ell -1,i}(D_{n,\zeta })]. Choose the modulus bounded by 2M1,i2M_{\ell -1,i}. For every δ>0\delta >0,

supζQjnEζ[ω1,i(Dn,ζ)]ω1,i(δ)+2M1,iδpsupζQjnEζ[Dn,ζp].\sup _{\zeta \in \mathcal Q_j^n}\E _\zeta [\omega _{\ell -1,i}(D_{n,\zeta })] \le \omega _{\ell -1,i}(\delta ) +2M_{\ell -1,i}\delta ^{-p} \sup _{\zeta \in \mathcal Q_j^n}\E _\zeta [D_{n,\zeta }^p].
The second term tends to zero by (60); letting δ0\delta \downarrow 0 proves (63). The same calculation for two target continuations, via (39), gives the uniform continuity.

Proposition 9.11 (Local-SV conditional static jets). Under Assumptions 9.19.3, for every j<j<\ell and every fixed mm,

(64)maxζQjn|am,nj,(ζ)amj,(ζ)|0.\begin{equation}\label {eq:lsv-static-limit} \max _{\zeta \in \mathcal Q_j^n} |a_{m,n}^{j,\ell }(\zeta )-a_m^{j,\ell }(\zeta )|\longrightarrow 0. \end{equation}
The target map zamj,(z)z\mapsto a_m^{j,\ell }(z) is bounded and uniformly continuous on Z\mathsf Z.

Proof.Apply Lemma 6.4 with Zn=Qjn\mathcal Z_n=\mathcal Q_j^n and the coupled normalized continuation returns. Hypothesis (1) is (60); (2) follows from (40) by the edgewise tower argument in the proof of Proposition 6.5, the Gaussian innovations contributing a factor bounded uniformly in nn and the atomic kernels inheriting the bound by Lemma 4.1; (3) is Lemma 9.10, the target bounds coming from (61) and (42); and (4) is Assumption 9.3. The same lemma applied to two target continuations coupled through (39), with the uniform continuity of Lemma 9.10, gives a common modulus for the jets, hence bounded and uniformly continuous zamj,(z)z\mapsto a_m^{j,\ell }(z).

Lemma 9.12 (Local-SV projected rows). For every fixed mm and 2\ell \ge 2,

(65)dm,ndm.\begin{equation}\label {eq:lsv-projected-limit} d_{m,n}^\ell \longrightarrow d_m^\ell . \end{equation}

Proof.Apply Lemma 7.1 under the first-edge coupling of Theorem 9.9, with Rn=X^1nx0R_n=\widehat X_1^n-x_0, R=X1x0R=X_1-x_0, ζn=(Ξ1n,Y^1n)\zeta _n=(\Xi _1^n,\widehat Y_1^n) and ζ=(X1,Y1)\zeta =(X_1,Y_1). The coupling gives ζnζ\zeta _n\to \zeta in LpL^p, and RnRR_n\to R in LpL^p because the proxy gap (50) vanishes with βn\beta _n. Proposition 9.11 supplies the jets, bounded and uniformly continuous, and their uniform convergence bounds the AnA_n uniformly in nn; (45) gives the variance.

9.5The second density theorem

Let TLSV\mathfrak T_{\rm LSV} be the rooted Markov chains on the grid, with state space R×E\R \times E and known root, whose kernels satisfy Assumptions 9.19.3, such as the chains of Proposition 9.4. Let Glift\mathfrak G_{\rm lift} be the finite proxy-GM chains of Definition 9.6 with the stated moments and well-defined raw readouts, and put

MLSV:=TLSVGlift,τmovLSV:=τmov(MLSV).\mathfrak M_{\rm LSV}:=\mathfrak T_{\rm LSV}\cup \mathfrak G_{\rm lift}, \qquad \tau _{\rm mov}^{\rm LSV}:=\tau _{\rm mov}(\mathfrak M_{\rm LSV}).

Theorem 9.13 (Local-SV density at every finite projected order). Let PTLSVP\in \mathfrak T_{\rm LSV} and let L{2,,J}\mathcal L\subseteq \{2,\ldots ,J\} be finite. The single diagonal sequence PnP^n of Theorem 9.9 satisfies items 1–4 of Theorem 7.3, with X^n\widehat X^n in place of XnX^n. Consequently

TLSVGliftτmovLSV,k(P;Glift)=.\mathfrak T_{\rm LSV} \subseteq \overline {\mathfrak G_{\rm lift}}^{\,\tau _{\rm mov}^{\rm LSV}}, \qquad k_*(P;\mathfrak G_{\rm lift})=\infty .

Proof.Items 1 and 2 are Theorem 9.9 and Lemma 9.7. Lemma 9.10 and Proposition 9.11 give every fixed finite collection of root and continuation static jets, the root jets being the case j=0j=0, whose only label is the root; the diagonal rnr_n\uparrow \infty makes one sequence sufficient for all finite orders. Lemma 9.12 gives the projected rows. On the regular chart, the triangular map (8) is continuous, which gives the velocity jets. Raw-coordinate and marginal convergence is convergence in every generator of (10), so PnPP^n\to P in τmovLSV\tau _{\rm mov}^{\rm LSV}; as PP was arbitrary, the class inclusion follows.

Remark 9.14 (Relation to the homogeneous theorem). The hypotheses of Theorem 7.3 imply those of Theorem 9.13: under Assumption 3.1 the return density is Law(Lj)φαj\Law (L_j)*\varphi _{\alpha _j}, so urqjφαj(r)\|\partial _u^rq_j\|_\infty \le \|\varphi _{\alpha _j}^{(r)}\|_\infty and, by Lemma 6.2 and (15),

|urqj(u;y)urqj(u;y¯)|Lφαj(r+1)|yy¯|,|\partial _u^rq_j(u;y)-\partial _u^rq_j(u;\bar y)|\le L\|\varphi _{\alpha _j}^{(r+1)}\|_\infty |y-\bar y|,
which is Assumption 9.2; the remaining conditions transfer by independence of GjG_j, and Var(X1X0)α0>0\Var (X_1-X_0)\ge \alpha _0>0. The inclusion is strict: in (46) with σ\sigma depending on xx, the residual law depends on XtjX_{t_j}. The homogeneous theorem is kept because its regularity follows from a checkable structural condition and its proof needs no diagonal: with a fixed factor one sequence serves every order, without the rate coupling of Lemma 9.8 or the component growth of Remark 9.16.

Remark 9.15 (Finite regularity). If Assumption 9.2 holds only through order R=max{N1,0}R=\max \{N-1,0\} on the terminal edges of a finite maturity set L\mathcal L, take rnRr_n\equiv R and impose (56) on those edges only, the only ones at which (54) enters (Lemma 9.10); the same proof gives closure in τmovN,L(TLSVN,LGlift)\tau _{\rm mov}^{N,\mathcal L}(\mathfrak T_{\rm LSV}^{N,\mathcal L}\cup \mathfrak G_{\rm lift}) for this finite-regularity class TLSVN,L\mathfrak T_{\rm LSV}^{N,\mathcal L}. The all-order assumption is what yields one sequence converging in the full topology.

Remark 9.16 (Component budget). With φβ(i+1)=ciβ(i+2)/2\|\varphi _\beta ^{(i+1)}\|_\infty =c_i\beta ^{-(i+2)/2} the proof needs εn(nDn)1\varepsilon _n\le (nD_n)^{-1}, Dn=1+maxirnciβn(i+2)/2D_n=1+\max _{i\le r_n}c_i\beta _n^{-(i+2)/2}. On the schedule rn=nr_n=n, βn=10n\beta _n=10^{-n} this is below 102510^{-25} at n=6n=6, and the component count in (51) grows accordingly — driven by the derivative order, unlike the homogeneous construction. At a fixed order (rnRr_n\equiv R, Remark 9.15) the admissible error is n1βn(R+2)/2n^{-1}\beta _n^{(R+2)/2} up to a constant.

Corollary 9.17 (Rate at a fixed readout order in the local-SV class). Run the lifted construction with per-node quantization error ε\varepsilon and a common innovation variance β\beta, and fix a readout order mNm\le N.

  1. For m=0m=0 the error is O(ε+β)O(\varepsilon +\beta ), so there is no balance to strike. For m1m\ge 1, if um2qj\partial _u^{m-2}q_j is Hölder-γ\gamma in uu uniformly in the label, γ(0,1]\gamma \in (0,1] — at m=1m=1 read u1qj\partial _u^{-1}q_j as the distribution function, Lipschitz because qjq_j is bounded, so γ=1\gamma =1 — the errors in ama_m^\ell and dmd_m^\ell are at most cm(βm/2ε+Mmβγ/2)c_m\bigl (\beta ^{-m/2}\varepsilon +M_m\beta ^{\gamma /2}\bigr ), optimized at βε2/(m+γ)\beta \asymp \varepsilon ^{2/(m+\gamma )}, giving O(εγ/(m+γ))O\bigl (\varepsilon ^{\gamma /(m+\gamma )}\bigr ).
  2. If in addition the target return densities are CmC^{m} with bounded derivatives — as in Proposition 9.4, where uiqjφν2Δj(i)\|\partial _u^iq_j\|_\infty \le \|\varphi _{\nu ^2\Delta _j}^{(i)}\|_\infty — the innovation bias is O(β)O(\beta ) and the balance improves to βε2/(m+2)\beta \asymp \varepsilon ^{2/(m+2)}, giving

    |dm,ndm||am,nam|=O(ε2/(m+2)),ε=ε(M) as in (33).\bigl |d_{m,n}^\ell -d_m^\ell \bigr |\vee \bigl |a_{m,n}^\ell -a_m^\ell \bigr | =O\bigl (\varepsilon ^{2/(m+2)}\bigr ), \qquad \varepsilon =\varepsilon (M)\ \text {as in }\eqref {eq:quant-rate-M}.

Proof.Compare three objects: the target; the smoothed target, whose return from any state over h1h\ge 1 edges is the target’s return plus an independent GγhβG\sim \gamma _{h\beta }; and the approximant. The error splits into a quantization part, approximant against smoothed target, and a bias, smoothed target against target.

Quantization part. From a label ζ\zeta, the approximant’s return over hh edges is the sum of its proxy increments plus the accumulated innovations, which by (50) have law γhβ\gamma _{h\beta } and are independent of the proxy-label chain. Coupling the proxy increments with the target’s returns from the point ζ\zeta, as in Theorem 9.9, costs at most CεC\varepsilon in LpL^p, uniformly in ζ\zeta; the innovations do not enter the proxy recursion, so this cost does not depend on β\beta. Lemma 8.2 with the common factor GG and α:=hββ\alpha :=h\beta \ge \beta then bounds the difference of C(i)(0)C^{(i)}(0) by cεc\,\varepsilon at i=0i=0 and by ciβi/2εc_i\beta ^{-i/2}\varepsilon for 1im1\le i\le m. Only density derivatives of order at most m2m-2 enter, which is where this differs from a bound through φβ(m+1)\|\varphi _\beta ^{(m+1)}\|_\infty. The exponential moments are uniform by the tower argument in the proof of Proposition 9.11, and on the window of Assumption 9.3, slightly enlarged, the jet map of Lemma 2.1 is Lipschitz, as in the proof of Theorem 8.3. The static jets therefore differ by cmβm/2εc_m\beta ^{-m/2}\varepsilon, uniformly in the label. The same two lemmas, applied to smoothed target continuations from two states coupled through (39), show that the smoothed jets are Lipschitz in the state with constant cmβm/2c_m\beta ^{-m/2}; no modulus of the target’s own jets in the state is needed.

Rows. The first-edge innovation is independent of the date-11 label ζ1=(Ξ1n,Y^1n)\zeta _1=(\Xi _1^n,\widehat Y_1^n), so the approximant’s row has numerator Cov(Am,n,Λ0)\Cov (A_{m,n}^\ell ,\Lambda _0), with Λ0\Lambda _0 the first proxy increment, and denominator Var(Λ0)+β\Var (\Lambda _0)+\beta. Split Am,nAmA_{m,n}^\ell -A_m^\ell into the approximant’s jets against the smoothed jets at ζ1\zeta _1, the smoothed jets at ζ1\zeta _1 against those at Z1=(X1,Y1)Z_1=(X_1,Y_1), and the smoothed jets at Z1Z_1 against the target’s. The first term is the quantization part, the second is at most cmβm/2E|ζ1Z1|cmβm/2Cεc_m\beta ^{-m/2}\E |\zeta _1-Z_1|\le c_m\beta ^{-m/2}C\varepsilon, and the third is the bias below, uniform in the state. Hölder’s inequality against the first-edge returns, which are CεC\varepsilon apart in LpL^p, and the denominators, which differ by at most Cε+βC\varepsilon +\beta and are bounded below by (45), give dmd_m^\ell the same bound as the jets.

Bias. For a target return UU from any state and an independent GβγβG_\beta \sim \gamma _\beta, the call function of U+GβU+G_\beta is Cβ(k)=E[eGβC(kGβ)]C_\beta (k)=\E [e^{G_\beta }C(k-G_\beta )], so Cβ(i)(k)=E[eGβC(i)(kGβ)]C_\beta ^{(i)}(k)=\E [e^{G_\beta }C^{(i)}(k-G_\beta )]. Taylor’s formula in GβG_\beta, whose first two moments are β/2-\beta /2 and β+β2/4\beta +\beta ^2/4, gives

(66)Cβ(i)(0)C(i)(0)=β2(C(i+2)(0)C(i+1)(0))+O(β2)\begin{equation}\label {eq:innovation-bias} C_\beta ^{(i)}(0)-C^{(i)}(0)=\tfrac \beta 2\bigl (C^{(i+2)}(0)-C^{(i+1)}(0)\bigr )+O(\beta ^2) \end{equation}
when the return density has i+2i+2 bounded derivatives, as in Proposition 9.4, and Cβ(i)(0)C(i)(0)=O(β)C_\beta ^{(i)}(0)-C^{(i)}(0)=O(\beta ) already when it has ii, the second-order remainder being O(E[Gβ2e|Gβ|])O(\E [G_\beta ^2e^{|G_\beta |}]); over hh edges, replace β\beta by hβh\beta. Under (2) the bias is therefore O(β)O(\beta ), uniformly in the state, with a constant controlled by umqj\|\partial _u^mq_j\|_\infty; at m=0m=0 it is O(β)O(\beta ) because the density is bounded. Under (1) only the modulus is available and the cruder MmE[eGβ|Gβ|γ]cMmβγ/2M_m\E [e^{G_\beta }|G_\beta |^\gamma ]\le cM_m\beta ^{\gamma /2} is used.

Equating the two parts gives the stated balances; at m=0m=0 the quantization part carries no power of β\beta, so no balance arises.

One might keep the target’s own floor instead: in Proposition 9.4 the increment splits as Rj+NjR_j+N_j with Njγν2ΔjN_j\sim \gamma _{\nu ^2\Delta _j} independent of RjR_j given the state, and quantizing RjR_j alone would give the fixed-factor situation of Section 8. But a finite chain samples its kernel at the finitely valued proxy label, the target at its own traded log price, and a fixed innovation keeps the two apart. The proxy enters through the approximant’s filtration, not through the cost: let Πbc(P,Q)\Pi _{\rm bc}(P,Q) be the couplings of PP and a finite proxy-GM chain QQ that are bicausal for the natural filtration of (X,Y)(X,Y) under PP and for F^\widehat {\mathcal F} under QQ, and put

AW1proj(P,Q):=infπΠbc(P,Q)Eπ[j=0J(|XjX^j|+|YjY^j|)].\AW _1^{\rm proj}(P,Q):=\inf _{\pi \in \Pi _{\rm bc}(P,Q)} \E _\pi \Bigl [\sum _{j=0}^J\bigl (|X_j-\widehat X_j|+|Y_j-\widehat Y_j|\bigr )\Bigr ].

It charges only the traded price and the mark and is at most a multiple of the lifted distance of Theorem 9.9, so a lower bound on it is the stronger statement.

Proposition 9.18 (A non-vanishing innovation obstructs adapted approximation). Let PTLSVP\in \mathfrak T_{\rm LSV}, and suppose some edge kernel depends on the log price at every mark: for some 1j<J1\le j<J and some bounded 11-Lipschitz ψ\psi, the map

xΨj(x,y):=ψ(u)Hj((x,y);du,E)x\longmapsto \Psi _j(x,y):=\int \psi (u)\,H_j((x,y);du,E)
is non-constant for every yEy\in E. If QnQ_n are finite proxy-GM chains whose total innovation variance before date jj, Bj,n:=h<jβh,nB_{j,n}:=\sum _{h<j}\beta _{h,n}, satisfies infnBj,n>0\inf _nB_{j,n}>0, then infnAW1proj(P,Qn)>0\inf _n\AW _1^{\rm proj}(P,Q_n)>0.

Proof.Suppose not; then there are indices, relabelled nn, and πnΠbc(P,Qn)\pi _n\in \Pi _{\rm bc}(P,Q_n) with cost tending to zero. Put ζj:=(Ξjn,Y^jn)\zeta _j:=(\Xi _j^n,\widehat Y_j^n) and G:=X^jnΞjnG:=\widehat X_j^n-\Xi _j^n, which by (50) has law γBj,n\gamma _{B_{j,n}} and is independent of ζj\zeta _j. The laws of X^jn=Ξjn+G\widehat X_j^n=\Xi _j^n+G converge, which is impossible if Bj,nB_{j,n} is unbounded, since then supcP(G[cR,c+R])0\sup _c\mathbb P(G\in [c-R,c+R])\to 0 for every RR; so along a further subsequence Bj,n[b,b¯]B_{j,n}\in [b,\bar b] with b>0b>0. Under πn\pi _n, conditionally on the joint past at date jj, bicausality and the Markov property make Xj+1XjX_{j+1}-X_j distributed as Hj((Xj,Yj);,E)H_j((X_j,Y_j);\cdot ,E) and X^j+1nX^jn\widehat X_{j+1}^n-\widehat X_j^n as a law κn(ζj)\kappa _n(\zeta _j) that depends on the proxy label alone, by (48). With kn:=ψdκnk_n:=\int \psi \,d\kappa _n and ψ\psi 11-Lipschitz,

Eπn|Ψj(Xj,Yj)kn(ζj)|Eπn|Xj+1X^j+1n|+Eπn|XjX^jn|0.\E _{\pi _n}\bigl |\Psi _j(X_j,Y_j)-k_n(\zeta _j)\bigr | \le \E _{\pi _n}|X_{j+1}-\widehat X_{j+1}^n|+\E _{\pi _n}|X_j-\widehat X_j^n|\longrightarrow 0.
By (39), Ψj\Psi _j is LL-Lipschitz, so (Xj,Yj)(X_j,Y_j) may be replaced by (X^jn,Y^jn)=(Ξjn+G,Y^jn)(\widehat X_j^n,\widehat Y_j^n)=(\Xi _j^n+G,\widehat Y_j^n) at a cost tending to zero, and E|Ψj(Ξjn+G,Y^jn)kn(ζj)|0\E |\Psi _j(\Xi _j^n+G,\widehat Y_j^n)-k_n(\zeta _j)|\to 0. For (ξ,y,B)(\xi ,y,B) put D(ξ,y,B):=infcRE|Ψj(ξ+GB,y)c|D(\xi ,y,B):=\inf _{c\in \R }\E |\Psi _j(\xi +G_B,y)-c| with GBγBG_B\sim \gamma _B. Since GG is independent of ζj\zeta _j, the last expectation is at least E[D(Ξjn,Y^jn,Bj,n)]\E [D(\Xi _j^n,\widehat Y_j^n,B_{j,n})]. Since Ψj\Psi _j is bounded and LL-Lipschitz, cc may be restricted to [ψ,ψ][-\|\psi \|_\infty ,\|\psi \|_\infty ], where the infimum is attained, and E|Ψj(ξ+GB,y)c|\E |\Psi _j(\xi +G_B,y)-c| is equicontinuous in (ξ,y,B)(\xi ,y,B) uniformly in cc, the GBG_B being coupled through one standard normal; so DD is continuous. It is positive because GBG_B has full support: D=0D=0 would make the continuous function Ψj(,y)\Psi _j(\cdot ,y) equal to the minimizing cc everywhere. The Ξjn=X^jnG\Xi _j^n=\widehat X_j^n-G are tight; a compact KK with P(ΞjnK)12\mathbb P(\Xi _j^n\in K)\ge \frac 12 for all nn and cK:=minK×E×[b,b¯]D>0c_K:=\min _{K\times E\times [b,\bar b]}D>0 give E[D]cK/2\E [D]\ge c_K/2, a contradiction.

Kernels that do not depend on the log price are homogeneous, and fall under Theorem 7.3 when they carry a factor (Assumptions 3.1 and 3.2). Keeping the floor of Proposition 9.4 gives Bj,n=ν2(tjt0)B_{j,n}=\nu ^2(t_j-t_0), so under the hypothesis it cannot converge in AW1proj\AW _1^{\rm proj}, and quantizing the floor instead removes the Gaussian component it was meant to supply (Proposition 2.5). The obstruction concerns adapted approximation: a fixed innovation can still match finitely many readout coordinates on one edge (Proposition 10.1), and the vanishing need not be tied to the derivative order (Proposition 9.20).

Proposition 9.19 (The innovation term is sharp). Let PTLSVP\in \mathfrak T_{\rm LSV}, and suppose there are 1j<J1\le j<J, a bounded 11-Lipschitz ψ\psi, constants κ>0\kappa >0 and s(0,1]s\in (0,1], and a Borel set SR×ES\subseteq \R \times E with ϖ:=P((Xj,Yj)S)>0\varpi :=\mathbb P((X_j,Y_j)\in S)>0, such that whenever (x,y)S(x,y)\in S and |yy|s|y'-y|\le s the map Ψj(,y)\Psi _j(\cdot ,y') is monotone on [x3s,x+3s][x-3s,x+3s] with |Ψj(u,y)Ψj(u,y)|κ|uu||\Psi _j(u',y')-\Psi _j(u,y')|\ge \kappa |u'-u| there. Then every finite proxy-GM chain QQ with Bj:=h<jβhb0:=s2/(8log(8/ϖ))B_j:=\sum _{h<j}\beta _h\le b_0:=s^2/(8\log (8/\varpi )) satisfies

AW1proj(P,Q)cBj,c:=ϖmin{14,p0κ8(1+L)},p0:=Φ(12)Φ(1).\AW _1^{\rm proj}(P,Q)\ge c\sqrt {B_j}, \qquad c:=\varpi \min \Bigl \{\frac 14,\frac {p_0\kappa }{8(1+L)}\Bigr \}, \quad p_0:=\Phi (-\tfrac 12)-\Phi (-1).
For the chains of Theorem 9.9, Bj,n=jβnB_{j,n}=j\beta _n, and with (59) this gives cjβnAW1proj(P,Pn)C(εn+βn)c\sqrt {j\beta _n}\le \AW _1^{\rm proj}(P,P^n)\le C(\varepsilon _n+\sqrt {\beta _n}) for large nn: the innovation term of (59) is optimal. The construction has εn(nDn)1\varepsilon _n\le (nD_n)^{-1} with Dnφβn=(2πe)1/2βn1D_n\ge \|\varphi _{\beta _n}'\|_\infty =(2\pi e)^{-1/2}\beta _n^{-1}, so εn=o(βn)\varepsilon _n=o(\beta _n) and its adapted error is of exact order βn\sqrt {\beta _n}.

Proof.Let πΠbc(P,Q)\pi \in \Pi _{\rm bc}(P,Q) have cost CπC_\pi, and keep the notation of the proof of Proposition 9.18. The two estimates at its start give Eπ|Ψj(Ξj+G,Y^j)k(ζj)|(1+L)Cπ\E _\pi |\Psi _j(\Xi _j+G,\widehat Y_j)-k(\zeta _j)|\le (1+L)C_\pi. If Cπϖs/4C_\pi \ge \varpi s/4, then Cπϖ4BjC_\pi \ge \frac \varpi 4\sqrt {B_j} because Bjs2B_j\le s^2. Otherwise Markov’s inequality gives P(|X^jXj|>s)+P(|Y^jYj|>s)<ϖ2\mathbb P(|\widehat X_j-X_j|>s)+\mathbb P(|\widehat Y_j-Y_j|>s)<\frac \varpi 2, and P(|G|>s)2es2/(8Bj)ϖ4\mathbb P(|G|>s)\le 2e^{-s^2/(8B_j)}\le \frac \varpi 4; on the intersection of {(Xj,Yj)S}\{(X_j,Y_j)\in S\} with the complements, |ΞjXj|2s|\Xi _j-X_j|\le 2s and |Y^jYj|s|\widehat Y_j-Y_j|\le s, so on this event, of probability at least ϖ/4\varpi /4, ζj\zeta _j lies in the set SS' of labels (ξ,y)(\xi ,y') within 2s2s and ss of a point of SS. For (ξ,y)S(\xi ,y')\in S' the map Ψj(,y)\Psi _j(\cdot ,y') is κ\kappa-monotone on [ξs,ξ+s][\xi -s,\xi +s], which contains ξ+G\xi +G on the events {G+Bj/2[Bj,12Bj]}\{G+B_j/2\in [-\sqrt {B_j},-\frac 12\sqrt {B_j}]\} and {G+Bj/2[12Bj,Bj]}\{G+B_j/2\in [\frac 12\sqrt {B_j},\sqrt {B_j}]\}, each of probability p0p_0; on them Ψj(ξ+G,y)\Psi _j(\xi +G,y') lies respectively below and above two values κBj\kappa \sqrt {B_j} apart, so every constant is at distance at least 12κBj\frac 12\kappa \sqrt {B_j} from one of them and D(ξ,y,Bj)12p0κBjD(\xi ,y',B_j)\ge \frac 12p_0\kappa \sqrt {B_j}. Independence of GG and ζj\zeta _j then gives (1+L)CπE[D(ζj,Bj)]ϖ8p0κBj(1+L)C_\pi \ge \E [D(\zeta _j,B_j)]\ge \frac \varpi 8p_0\kappa \sqrt {B_j}. Take the infimum over π\pi.

Proposition 9.20 (Linear rate when no evaluation follows the readout). Let the target be in the class of Proposition 9.4 with final-edge floor α:=ν2ΔJ1\alpha :=\nu ^2\Delta _{J-1}, and let L={J}\mathcal L=\{J\}. Run the construction of Theorem 9.9 with edge variances βj,n=εn2\beta _{j,n}=\varepsilon _n^2 for j<J1j<J-1 and βJ1,n=α\beta _{J-1,n}=\alpha, the final-edge atoms quantizing the residual RJ1R_{J-1} alone. Then for every finite NN,

RN,{J}raw(Pn)RN,{J}raw(P)Cεn,\bigl \|\Rcal ^{\rm raw}_{N,\{J\}}(P^n)-\Rcal ^{\rm raw}_{N,\{J\}}(P)\bigr \|_\infty \le C\,\varepsilon _n,
with C=c(N,p,q,C,D,d,w,w)(1+L)J(1+α(N+1)/2)(1+(ν2Δ0)1)C=c(N,p,q,C,D,d,\underline w,\overline w)\,(1+L)^J\bigl (1+\alpha ^{-(N+1)/2}\bigr )\bigl (1+(\nu ^2\Delta _0)^{-1}\bigr ), and with no diagonal and no order-dependent component budget.

Proof.The static jets at tJt_J, root and continuation, are functionals of return laws ending at JJ, each containing NJ1γαN_{J-1}\sim \gamma _\alpha common to target and approximant, so Lemma 8.2 gives maxiN+1|Cn(i)(0)C(i)(0)|c(1+α(N+1)/2)εn\max _{i\le N+1}|C_n^{(i)}(0)-C^{(i)}(0)|\le c\,(1+\alpha ^{-(N+1)/2})\varepsilon _n once the parts before that factor are O(εn)O(\varepsilon _n) apart in LpL^p. They are: the kernels are evaluated at the proxy, which (25) keeps within O(εn)O(\varepsilon _n) of the target’s log price (on the final edge for the residual, Lipschitz in the state with the same LL since the synchronous coupling of Proposition 9.4 makes the floors equal), and the innovations before the final edge, the only ones preceding a kernel evaluation, have total variance at most Jεn2J\varepsilon _n^2. For the rows, Lemma 8.2 on two synchronously coupled target continuations makes the continuation jets Lipschitz in the state, and the proof of Theorem 8.3 applies with denominator at least ν2Δ0\nu ^2\Delta _0. Definition 9.6 permits a different variance on each edge, and the jet map is Lipschitz on the window as in Theorem 8.3.

Remark 9.21 (Why the two classes separate quantitatively). Theorem 8.3 keeps α\alpha fixed, so the readout error is the quantization error times a constant; Corollary 9.17 must send β0\beta \downarrow 0 and pays ε2/(m+2)\varepsilon ^{2/(m+2)} even in case (2). With weights 2m2^{-m} on the order-mm coordinates, a metric compatible with the readout topology, the error is of order εθ\varepsilon ^{\theta }, θ=min{1,2log2/log(1/α)}\theta =\min \{1,2\log 2/\log (1/\alpha )\}, with a fixed factor and only exp(clog(1/ε))\exp (-c\sqrt {\log (1/\varepsilon )}) with vanishing innovations. Neither exponent is intrinsic — polynomially decaying weights turn both into logarithms — but the gap is, being present coordinate by coordinate. All of these are upper bounds; Section 10.1 reports where they are loose.

10Closed-form readout, cubature and numerics

The readout of a finite chain is available in closed form, which drives the discrete model of [29] and the computations below; one edge with fixed continuation admits exact finite-order matching by cubature.

If a normalized log return has the finite mixture law Ri=1MpiN(mi,si2)R\sim \sum _{i=1}^Mp_i\N (m_i,s_i^2), with si>0s_i>0 and ipiemi+si2/2=1\sum _ip_ie^{m_i+s_i^2/2}=1, its call price is, as in [13],

(67)C(k)=ipi[emi+si2/2Φ(mi+si2ksi)ekΦ(miksi)].\begin{equation}\label {eq:gm-call} C(k)=\sum _i p_i\left [ e^{m_i+s_i^2/2}\Phi \left (\frac {m_i+s_i^2-k}{s_i}\right ) -e^k\Phi \left (\frac {m_i-k}{s_i}\right ) \right ]. \end{equation}

All strike derivatives are finite sums, so Lemma 2.1 gives the ATM jets by one scalar inversion and implicit differentiation, and the projected rows follow from the mixture formulas of Section 2.3, which for branch-constant continuation jets give (29) with variance α0\alpha _0 (homogeneous) or β0,n\beta _{0,n} (proxy).

Proposition 10.1 (One-step finite-order attainment). Fix a current state, one maturity, a fixed continuation, and the standard pivot a10a_1\ne 0. Suppose the target one-step marked return has the common-variance representation

[N(λ(θ)α/2,α)δy(θ)]Λ(dθ),\int \bigl [\N (\lambda (\theta )-\alpha /2,\alpha )\otimes \delta _{y(\theta )}\bigr ]\,\Lambda (d\theta ),
where α>0\alpha >0 and eλ(θ)Λ(dθ)=1\int e^{\lambda (\theta )}\Lambda (d\theta )=1. Write m(θ)=λ(θ)α/2m(\theta )=\lambda (\theta )-\alpha /2, let Aj(θ)A_j(\theta ) be the fixed-continuation jet, and define
Cθ(k):=E[(eλ(θ)+G+Rθcontek)+],Gγα,GRθcont,C_\theta (k):=\E \!\left [ \left (e^{\lambda (\theta )+G+R_\theta ^{\rm cont}}-e^k\right )^+ \right ], \qquad G\sim \gamma _\alpha ,\quad G\perp R_\theta ^{\rm cont},
where RθcontR_\theta ^{\rm cont} is the continuation log return from y(θ)y(\theta ). Assume every coordinate of
(1,eλ,m,m2+α,Aj,Ajm:0jN,khCθ(0):0hN+1)\left (1,e^\lambda ,m,m^2+\alpha , A_j,A_jm:0\le j\le N, \partial _k^hC_\theta (0):0\le h\le N+1\right )
is in L1(Λ)L^1(\Lambda ). For every NN, there is a finite mixture with at most 3N+83N+8 components which is an exact asset martingale and matches exactly the raw static and projected inputs needed to recover v0,,vNv_0,\ldots ,v_N.

Proof.Use the displayed L1L^1 vector as the cubature test vector. The common positive Gaussian variance justifies differentiation under the Λ\Lambda-integral: Gaussian derivative bounds dominate the strike derivatives, while the exponential-moment coordinate controls the call value. Including the constant, the test vector has

1+1+2+2(N+1)+(N+2)=3N+81+1+2+2(N+1)+(N+2)=3N+8
coordinates. Tchakaloff’s theorem [6], applied to the law of the test vector with linear test functions, gives at most that many nodes in its support and positive weights pip_i matching every coordinate. The coordinates depend on θ\theta only through (λ(θ),y(θ))(\lambda (\theta ),y(\theta )) and are continuous there, the continuation jets by Propositions 6.5 and 9.11; as mm bounds λ\lambda along convergent sequences and EE is compact, the range on suppΛ\operatorname {supp}\Lambda is closed, so every node is the value at some θi\theta _i. A martingale version, with support inside that of the original law, is in [9]. Matching the martingale coordinate gives ipieλi=1\sum _i p_ie^{\lambda _i}=1. Matching the first two moments and the AjA_j cross-moments gives the exact regression numerators and denominator, including the within-component variances. Matching the call derivatives gives the exact static jets through the order needed by the triangular inversion. Hence the recovered projected velocity jets agree.

Proposition 10.1 concerns one edge with fixed continuation and variance, not the constructions, whose continuations are endogenous.

10.1Two checks of the rates

Take two edges, t0<t1<t2t_0<t_1<t_2, and a latent mark Y1Y_1 on E=[0.10,0.40]E=[0.10,0.40] with a smooth density. Conditionally on Y1=yY_1=y,

L0{Y1=y}N(μ(y)12s(y)2,s(y)2),s(y)=yΔ0,μ(y)=κ(yy¯)+c,L_0\mid \{Y_1=y\}\sim \N \!\left (\mu (y)-\tfrac 12 s(y)^2,\;s(y)^2\right ), \qquad s(y)=y\sqrt {\Delta _0}, \qquad \mu (y)=-\kappa (y-\bar y)+c,

with cc fixed by E[eL0]=1\E [e^{L_0}]=1, and from yy the continuation return is a two-component Gaussian mixture with skewed means and yy-proportional scales, renormalized to unit exponential mean. Both conditional laws are finite mixtures, so the maturity-t2t_2 jets Am=am1,2(Y1)A_m=a_m^{1,2}(Y_1) follow from (67) and Lemma 2.1, and the target readout dm:=dm2d_m:=d_m^2 of (6) is a one-dimensional quadrature in yy.

The approximant quantizes the joint law of (L0,Y1)(L_0,Y_1) into M=K2M=K^2 cells, KK mark bands times KK conditional return quantiles, so d=1d=1. Cell rr carries mass prp_r, mark yr=E[Y1r]y_r=\E [Y_1\mid r] and log-barycentre λr=logE[eL0r]\lambda _r=\log \E [e^{L_0}\mid r], so rpreλr=1\sum _rp_re^{\lambda _r}=1 at every resolution; the first-edge return is λr\lambda _r plus the independent factor, and, the target being homogeneous, the label is the mark alone.

Table 1. Theorem 8.3 at d=1d=1: readout error against the atom count M=K2M=K^2 with the factor variance held at α=103\alpha =10^{-3}. Last row: log-log slopes in KK over the four finest resolutions, against the predicted M1/(d+1)=K1M^{-1/(d+1)}=K^{-1} up to logarithms.
|amMam||a_{m}^{M}-a_m|
|dmMdm||d_{m}^{M}-d_m|
KK MM m=0m=0 m=1m=1 m=2m=2 m=0m=0 m=1m=1 m=2m=2
6 36 1.41031.4\cdot 10^{-3} 2.91032.9\cdot 10^{-3} 1.21021.2\cdot 10^{-2} 8.11048.1\cdot 10^{-4} 4.01034.0\cdot 10^{-3} 8.31038.3\cdot 10^{-3}
13 169 3.81043.8\cdot 10^{-4} 8.31048.3\cdot 10^{-4} 8.11048.1\cdot 10^{-4} 5.01045.0\cdot 10^{-4} 1.81031.8\cdot 10^{-3} 3.71033.7\cdot 10^{-3}
28 784 1.11041.1\cdot 10^{-4} 2.61042.6\cdot 10^{-4} 4.41044.4\cdot 10^{-4} 2.71042.7\cdot 10^{-4} 7.61047.6\cdot 10^{-4} 1.51031.5\cdot 10^{-3}
58 3364 3.41053.4\cdot 10^{-5} 9.31059.3\cdot 10^{-5} 2.51042.5\cdot 10^{-4} 1.31041.3\cdot 10^{-4} 3.21043.2\cdot 10^{-4} 6.31046.3\cdot 10^{-4}
121 14641 1.11051.1\cdot 10^{-5} 3.41053.4\cdot 10^{-5} 1.11041.1\cdot 10^{-4} 5.91055.9\cdot 10^{-5} 1.21041.2\cdot 10^{-4} 2.41042.4\cdot 10^{-4}
175 30625 6.51066.5\cdot 10^{-6} 2.11052.1\cdot 10^{-5} 7.31057.3\cdot 10^{-5} 3.71053.7\cdot 10^{-5} 7.41057.4\cdot 10^{-5} 1.41041.4\cdot 10^{-4}
slope in KK
1.49-1.49 1.35-1.35 1.13-1.13 1.15-1.15 1.32-1.32 1.34-1.34

Theorem 8.3 assumes a factor of fixed variance, so we split α=103\alpha =10^{-3} off both edges and hold it fixed, with s(y)2=y2Δ0α>0s(y)^2=y^2\Delta _0-\alpha >0 on EE, and vary only the resolution. With κ=2\kappa =2 and Δ0=Δ1=0.25\Delta _0=\Delta _1=0.25 its readout is

(a0,a1,a2)=(0.074497,0.133267,0.012041),(d0,d1,d2)=(0.117275,+0.216807,0.418134),\begin {aligned} (a_0,a_1,a_2)&=(0.074497,\,-0.133267,\,0.012041),\\ (d_0,d_1,d_2)&=(-0.117275,\,+0.216807,\,-0.418134), \end {aligned}

and Var(X1X0)=0.031458\Var (X_1-X_0)=0.031458. The continuation is kept exact at the cell’s mark barycentre, which isolates the quantization error of Lemma 8.1 and its propagation.

In Table 1, rpreλr\sum _rp_re^{\lambda _r} equals one to 710147\cdot 10^{-14} at every resolution, so martingality is exact rather than asymptotic, and every measured slope is at least 11 in modulus against a predicted M1/2M^{-1/2} up to logarithms, the ratio to M1/2(logM)3/2M^{-1/2}(\log M)^{3/2} falling monotonically over the last four rows. The bound therefore holds with room; the projected rows, which carry the covariance against the quantized edge, track it most tightly.

The same setup shows where Corollary 9.17 is loose. For a floored return U=S+NU=S+N, NγAN\sim \gamma _A with A=102A=10^{-2} and SS a two-component mixture, quantized into KK equal-probability log-barycentric cells and smoothed by γβ\gamma _\beta as in Section 9 (worst case over the phase of the cell grid relative to k=0k=0), the innovation bias is linear in β\beta — over β[2.8104,2.8103]\beta \in [2.8\cdot 10^{-4},2.8\cdot 10^{-3}] the error grows by factors 3.073.07 to 3.163.16 against a β\beta ratio of 3.123.12, as (66) predicts — and at m=2m=2 the measured error exponent in ε\varepsilon is 0.720.72 at the worst phase and 0.920.92 at the median, against 1/21/2 from Corollary 9.17(2). The exponent is right in kind but conservative: equal-probability cells make the distribution function accurate to O(ε)O(\varepsilon ) before any smoothing, which a bound through W1W_1 cannot see.

10.2A constrained band on an SPX surface

Figure 1: Forward-start and cliquet values of exact-martingale finite-GM chains that price the
same SPX vanillas inside bid–ask, 28 May 2026. Left: the three-month forward-start smile starting
in one month; the outer band is the range under the vanilla quotes alone, and the inner bands
are the ranges under chains that also match \(v_0\) , then \(v_1\) , then \(v_2\) , pinned at the three-month values
of [ 18 ] . Every band is exact: each endpoint is attained by an explicit chain and bounded by a
dual certificate, the two agreeing to within \(0.01\) volatility points. The two lines are the chains at the
cliquet endpoints under all three constraints. Right: the same for a two-period cliquet with \(\pm 5\%\) local
caps floored at zero.
Figure 1. Forward-start and cliquet values of exact-martingale finite-GM chains that price the same SPX vanillas inside bid–ask, 28 May 2026. Left: the three-month forward-start smile starting in one month; the outer band is the range under the vanilla quotes alone, and the inner bands are the ranges under chains that also match v0v_0, then v1v_1, then v2v_2, pinned at the three-month values of [18]. Every band is exact: each endpoint is attained by an explicit chain and bounded by a dual certificate, the two agreeing to within 0.010.01 volatility points. The two lines are the chains at the cliquet endpoints under all three constraints. Right: the same for a two-period cliquet with ±5%\pm 5\% local caps floored at zero.

Whether pinning a finite readout says anything about prices is an empirical question, and this section examines it on one surface in detail and then on sixty month-ends.

Data and fit. SPX end-of-day chains for 28 May 2026 at τ1=0.090\tau _1=0.090 and τ2=0.342\tau _2=0.342 years, with 227227 and 353353 two-sided quotes and forwards from put–call parity. A two-date chain with 3636 and 4747 atoms prices all 580580 quotes inside their bid–ask spreads, with implied-volatility errors of 0.0270.027 and 0.0780.078 points root mean square, and is a martingale to 610176\cdot 10^{-17}.

Experiment. Hold both fitted marginals fixed, so that every quote stays inside bid–ask, and search over the finite exact-martingale chains that join them, with the fitted atoms and component variances: a chain is a finite set of first-date branches, each at a fitted location yiy_i with its own transition law ρ\rho over the second-date components, kρkezkyi=1\sum _k\rho _ke^{z_k-y_i}=1, and total masses pip_i per location and rkr_k per component as fitted. Several branches may share a location — a latent state the price does not reveal — and the readout then averages their jets (Proposition 7.5). Maximize and minimize the forward-start implied volatility for [t1,t2][t_1,t_2] at strikes κ{0.85,0.90,,1.10}\kappa \in \{0.85,0.90,\ldots ,1.10\}, first under these constraints alone and then adding the velocity coordinates one at a time, each pinned to the three-month value of [18], (v0,v1,v2)=(1.336,0.901,37.26)(v_0,v_1,v_2)=(1.336,0.901,37.26), within (0.005,0.01,0.2)(0.005,0.01,0.2).

Computation. Without the readout these are linear programs in the joint weights, the martingale optimal transport problem of [726] on a finite grid. With it, each branch still enters linearly — through its mass, its law and its jets J(Cρ)J(C\rho ), with CρR3C\rho \in \R ^3 its continuation call price and first two strike derivatives at the money and JJ the map of Lemma 2.1 — and so does the velocity box, since with the date-22 marginal fixed the velocities are a fixed triangular image of the rows (8). Each program is thus a linear program over infinitely many candidate branches, solved by column generation; its optimum is an explicit chain, checked against every bid–ask interval and the exact readout, so its value is attained. Conversely, weak duality bounds the largest price by νr+σ(μ)+ipigi\nu \cdot r+\sigma (\mu )+\sum _ip_ig_i for any multipliers ν\nu on the second-date weights and μ\mu on the velocity box, with σ\sigma the support function of the box and gig_i the largest reduced value of one branch at location ii; each gig_i is nonconvex only through the three coordinates of CρC\rho and is bounded rigorously by branch-and-bound over them. Attained values and bounds agree to within 0.010.01 volatility points at every endpoint. Code is available from the authors; the ORATS option data are proprietary.

Result. On vanillas alone the forward-start smile is undetermined by 6.56.5 to 11.111.1 volatility points across the six strikes, 7.07.0 at the money (Figure 1). Matching the three readout coordinates as well narrows the range at every strike, to between 4.04.0 and 5.75.7 points and to 4.04.0 at the money, so that one half to three quarters of the vanilla range remains; at the money the skew response v1v_1 does most of the work and the curvature response v2v_2 the least. A two-period cliquet with ±5%\pm 5\% local caps floored at zero, resetting at the two quoted expiries so that its periods are 3333 and 9292 days and paid at t2t_2, ranges over 110110 basis points on a premium between 1.831.83 and 2.93%2.93\% of notional, and over 79.579.5 with the three coordinates matched. For a desk this is a measure of residual model risk, not a model.

Figure 2: The same experiment on the last trading day of each of the sixty months from July 2021
to June 2026, on the \(55\) whose fit places at least \(95\%\) of quotes inside bid–ask. Left: the at-the-money
forward-start range left open by the vanilla quotes, divided into what each coordinate removes as
it is added and what remains with all three pinned at the values of [ 18 ] , shaded as in Figure 1 ;
every endpoint is attained by an explicit chain and bounded by a dual certificate, the two agreeing
to \(0.01\) volatility points. Right: the fraction removed against the level of volatility, for both targets.
The forward-start fraction does not track that level; the cliquet fraction does, its \(\pm 5\%\) local caps
binding harder when volatility is high.
Figure 2. The same experiment on the last trading day of each of the sixty months from July 2021 to June 2026, on the 5555 whose fit places at least 95%95\% of quotes inside bid–ask. Left: the at-the-money forward-start range left open by the vanilla quotes, divided into what each coordinate removes as it is added and what remains with all three pinned at the values of [18], shaded as in Figure 1; every endpoint is attained by an explicit chain and bounded by a dual certificate, the two agreeing to 0.010.01 volatility points. Right: the fraction removed against the level of volatility, for both targets. The forward-start fraction does not track that level; the cliquet fraction does, its ±5%\pm 5\% local caps binding harder when volatility is high.

Across dates. Run on the last trading day of each of the sixty months from July 2021 to June 2026, with the two expiries chosen automatically near one and four months and the atom grids scaled to each surface, the same construction fits 5555 of them — on the rest it cannot place 95%95\% of the quotes inside bid–ask — and gives the same picture on every one (Figure 2): the velocity box is reachable, attained and certified endpoints agree at the money to 0.0020.002 volatility points, and the three coordinates remove a median 39%39\% of the at-the-money vanilla range, quartiles 3535 and 43%43\% and never below 29%29\%, a median 7.67.6 points narrowing to 4.44.4; for the cliquet the median removed is 25%25\%. The staged picture repeats too: at the money v0v_0, v1v_1 and v2v_2 remove a median 3232, 5353 and 14%14\% of the total narrowing, and v1v_1 is the largest single step on 5252 of the 5555 dates, its interval clear of the other two at every one of them. The pattern across strikes does not: on 5454 of the 5555 it is the 0.850.85 wing that is narrowed by the largest fraction, a median 55%55\% against 39%39\% at the money, where this surface has it the other way about. The forward-start figure does not track the level of volatility (correlation 0.12-0.12 with the one-month at-the-money volatility, which ranges over 9.99.9 to 28.3%28.3\%), whereas the cliquet figure does (0.56-0.56), its local caps binding harder when volatility is high. Holding the component volatility at the 5%5\% of this section instead of a quarter of the at-the-money volatility fits 5151 of the sixty and removes more, a median 50%50\%, with the same staged ordering.

What is and is not claimed. The ranges are exact over the class searched, and every value between two endpoints is attained, since mixing two chains branch by branch mixes prices and readouts. Chains with a single branch per location — continuation a function of the price location alone — can only have narrower constrained bands; sequential linear programming over them reaches 14.7814.78 to 18.2618.26 at the money, against the exact 14.3314.33 to 18.3518.35. Letting the marginals move within the spreads could only widen every band. The closest listed forward-starting claim, a VIX future, does not test the pinned coordinates: its square is a log contract over the whole smile rather than an at-the-money quantity, and on the VIX calendar of this surface chains matching the three coordinates still attain 0.7150.715 to 0.9940.994 of E[VIX2]\sqrt {\E [\mathrm {VIX}^2]}, against 0.6710.671 to 11 without them, while strike truncation in the variance-swap anchors pins the market ratio only to 0.970.971.031.03; the listed VIX prices therefore neither confirm nor contradict the readout. And the coefficients of [18] are physical-measure regressions of realized daily moves, whereas dmd_m^\ell is a risk-neutral finite-step projection: they fix the readout at a plausible magnitude to demonstrate capacity, not to calibrate. Nor does the narrowing depend on those values: sweeping v0v_0 from 11 to 22, or v1v_1 and v2v_2 each from zero to twice the values used, with the other two held, keeps the at-the-money range removed between 4141 and 47%47\% and the cliquet’s between 2727 and 31%31\%. It does depend on the readout being a response to the spot: replacing the regression on the spot move by one on any other contrast across first-date states, with the jets pinned as tightly relative to their spread, removes at most 16%16\% of the at-the-money range, whereas pinning spot-dependent moments of the two returns equally tightly removes about as much as the readout. What narrows the range is pinning how the continuation law responds to the spot, and the readout records that in quoted implied volatilities.

11Scope and extensions

Coverage. Imperfect correlation and Lipschitz dependence on the mark are mild; the implied-variance window and Var(X1X0)>0\Var (X_1-X_0)>0 need total volatility bounded above and away from zero; the variance floor of Assumption 3.1 and a finite-dimensional Markov mark are structural, whereas compactness of EE and uniform exponential moments are technical and should yield to localization. Finitely many known roots are handled by a disjoint union of trees.

target status reason
bounded SV with a variance floor Theorem 7.3 by construction
square-root variance, floored and capped Theorem 7.3 Proposition 3.4
square-root variance, unregularized outside as stated volatility degenerates
LSV with an independent variance floor Theorem 9.13 Proposition 9.4
LSV, kernel with RR derivatives orders up to R+1R+1 Remark 9.15
NN-factor forward-variance models Theorem 7.3 finite Markov state, truncated
curve-valued and rough volatility outside as stated EE assumed finite-dimensional

Square-root variance. An unregularized square-root variance fails both theorems through degeneracy, not only noncompactness: localizing to [v,v][\underline v,\overline v] gives a factor of variance θv(1ρ¯2)Δj0\theta \underline v(1-\bar \rho ^2)\Delta _j\to 0 as v0\underline v\downarrow 0, and one-step density derivatives of order V(r+1)/2V^{-(r+1)/2}. The regularizations that restore the hypotheses (Propositions 3.4 and 9.4) are of the kind implementations commonly apply.

Mark space. An NN-factor forward-variance model, truncated as in Proposition 3.4, is covered. A curve-valued or rough mark is not (ERdE\subset \R ^d), although the proofs use only that EE is compact and convex, as is any sup-norm-closed set of uniformly bounded curves with a common modulus of continuity.

Fixed marginals. For generic targets the moving-marginal formulation is the strongest one available, not a weakening (Section 2.6), and with the marginals held at prescribed finite-mixture laws the finite class is capped at the resolution of the first marginal (Proposition 7.5). A fixed-marginal density theorem can therefore hold only in a larger class, whose mixing weights depend on the continuous state. That route rests on stability of martingale couplings, which holds on the line [8374] and fails on R2\R ^2 [14]; whether it can be carried out with finitely many components per fibre is open.

Open directions. (i) Instantaneous limit: dmd_m is a regression over a fixed step; an instantaneous version needs a diagonal in the mesh. (ii) Noncompact states: Heston-type targets call for Lyapunov localization, as mimicking already handles degenerate covariance [15]. (iii) Budgets: the bound 3N+83N+8 of Proposition 10.1 is one-step; a fixed budget yields a capacity profile across orders. (iv) Other readouts: fixed-strike implied volatilities can be adjoined by (26); for targets with a finite moment index wing slopes cannot (Proposition 7.4).

12Conclusion

In the topology generated by the dynamics characteristics themselves, for the two target classes,

k(P;GSV)=(PTSV),k(P;Glift)=(PTLSV):k_*(P;\mathfrak G_{\rm SV})=\infty \quad (P\in \mathfrak T_{\rm SV}), \qquad k_*(P;\mathfrak G_{\rm lift})=\infty \quad (P\in \mathfrak T_{\rm LSV}):

every finite order of projected at-the-money smile dynamics is matched to arbitrary accuracy by exact-martingale finite Gaussian-mixture chains with converging marginals, using the target’s Gaussian factor (SV) or one manufactured at vanishing variance (LSV). The kernel that generates the dynamics is never observed; what these statements provide is a licence to search inside a closed-form class, in the coordinates that are observed, without excluding any finite-order behaviour these models produce.

Three qualifications travel with that licence. The rate is M1/(d+1)M^{-1/(d+1)} up to logarithms, with only the latent dimension in the exponent (Theorem 8.3), and our bound degrades when the Gaussian innovations must vanish, which adapted approximation forces for innovations preceding a price-dependent kernel (Proposition 9.18); we do not prove the degradation itself necessary. The licence is at the money: finite mixtures have zero implied-variance wing slopes, so it does not extend to a readout containing the wings (Proposition 7.4). And it is a capacity statement, not a calibration procedure — though on one SPX surface the vanillas leave the at-the-money forward-start volatility undetermined by 7.07.0 points, and fixing the first three characteristics narrows this to 4.04.0 (Section 10.2).

References

[1] Eduardo Abi Jaber and Shaun (Xiaoyuan) Li, Capturing smile dynamics with the quintic volatility model: SPX, skew-stickiness ratio and VIX, arXiv:2503.14158 (2025; revised 2026).

[2] Athanassia Bacharoglou, Approximation of probability distributions by convex mixtures of Gaussian measures, Proceedings of the American Mathematical Society 138 (2010), no. 7, 2619–2628, doi:10.1090/S0002-9939-10-10340-2.

[3] Julio Backhoff-Veraguas, Daniel Bartl, Mathias Beiglböck, and Manu Eder, Adapted Wasserstein distances and stability in mathematical finance, Finance and Stochastics 24 (2020), 601–632; arXiv:1901.07450, doi:10.1007/s00780-020-00426-3.

[4] Julio Backhoff-Veraguas and Gudmund Pammer, Stability of martingale optimal transport and weak optimal transport, Annals of Applied Probability 32 (2022), no. 1, 721–752, doi:10.1214/21-AAP1694.

[5] Daniel Bartl, Mathias Beiglböck, and Gudmund Pammer, The Wasserstein space of stochastic processes, Journal of the European Mathematical Society 28 (2026), no. 1, 393–454; arXiv:2104.14245.

[6] Christian Bayer and Josef Teichmann, The proof of Tchakaloff’s theorem, Proceedings of the American Mathematical Society 134 (2006), 3035–3040, doi:10.1090/S0002-9939-06-08249-9.

[7] Mathias Beiglböck, Pierre Henry-Labordère, and Friedrich Penkner, Model-independent bounds for option prices—a mass transport approach, Finance and Stochastics 17 (2013), 477–501, doi:10.1007/s00780-013-0205-8.

[8] Mathias Beiglböck, Benjamin Jourdain, William Margheriti, and Gudmund Pammer, Approximation of martingale couplings on the line in the adapted weak topology, Probability Theory and Related Fields 183 (2022), no. 1–2, 359–413, doi:10.1007/s00440-021-01103-y.

[9] Mathias Beiglböck and Marcel Nutz, Martingale inequalities and deterministic counterparts, Electronic Journal of Probability 19 (2014), no. 95, 1–15, doi:10.1214/EJP.v19-3270.

[10] Lorenzo Bergomi, Smile dynamics IV, Risk Magazine, December 2009, 94–100.

[11] Lorenzo Bergomi, Stochastic Volatility Modeling, Chapman & Hall/CRC, 2016.

[12] Lorenzo Bergomi and Julien Guyon, Stochastic volatility’s orderly smiles, Risk, May 2012, 60–66.

[13] Damiano Brigo and Fabio Mercurio, Lognormal-mixture dynamics and calibration to market volatility smiles, International Journal of Theoretical and Applied Finance 5 (2002), no. 4, 427–446, doi:10.1142/S0219024902001511.

[14] Martin Brückerhoff and Nicolas Juillet, Instability of martingale optimal transport in dimension d2d\ge 2, Electronic Communications in Probability 27 (2022), paper no. 24, doi:10.1214/22-ECP463.

[15] Gerard Brunick and Steven Shreve, Mimicking an Itô process by a solution of a stochastic differential equation, Annals of Applied Probability 23 (2013), no. 4, 1584–1628, doi:10.1214/12-AAP881.

[16] René Carmona and Sergey Nadtochiy, Local volatility dynamic models, Finance and Stochastics 13 (2009), no. 1, 1–48.

[17] René Carmona and Sergey Nadtochiy, Tangent models as a mathematical framework for dynamic calibration, International Journal of Theoretical and Applied Finance 14 (2011), no. 1, 107–135.

[18] Charlie Che and Pradeepta Das, Beyond the skew-stickiness ratio: transport geometry of spot-driven variance surface dynamics, arXiv:2608.12493 (2026).

[19] Rama Cont and Peter Tankov, Financial Modelling with Jump Processes, Chapman & Hall/CRC, Boca Raton, 2004.

[20] Valdo Durrleman, From implied to spot volatilities, Finance and Stochastics 14 (2010), no. 2, 157–177.

[21] Nicole El Karoui, Monique Jeanblanc-Picqué, and Steven E. Shreve, Robustness of the Black and Scholes formula, Mathematical Finance 8 (1998), no. 2, 93–126, doi:10.1111/1467-9965.00047.

[22] Avner Friedman, Partial Differential Equations of Parabolic Type, Prentice-Hall, 1964.

[23] Masaaki Fukasawa, Martingale expansion for stochastic volatility, SIAM Journal on Financial Mathematics 17 (2026), no. 2, SC1–SC12, doi:10.1137/26M1841318; arXiv:2601.09324.

[24] Masaaki Fukasawa, On the skew stickiness ratio, arXiv:2602.05241 (2026).

[25] Siegfried Graf and Harald Luschgy, Foundations of Quantization for Probability Distributions, Lecture Notes in Mathematics 1730, Springer, Berlin, 2000.

[26] Alfred Galichon, Pierre Henry-Labordère, and Nizar Touzi, A stochastic control approach to no-arbitrage bounds given marginals, with an application to lookback options, The Annals of Applied Probability 24 (2014), no. 1, 312–336, doi:10.1214/13-AAP925.

[27] Julien Guyon and Jordan Lekeufack, Volatility is (mostly) path-dependent, Quantitative Finance 23 (2023), no. 9, 1221–1258.

[28] István Gyöngy, Mimicking the one-dimensional marginal distributions of processes having an Itô differential, Probability Theory and Related Fields 71 (1986), 501–516, doi:10.1007/BF00699039.

[29] Shaosai Huang, SANOS-Evolve: a discrete stochastic-local-volatility model for European-option smile dynamics, SSRN working paper no. 7151258 (2026); ssrn.com/abstract=7151258.

[30] Benjamin Jourdain and Gilles Pagès, Quantization and martingale couplings, ALEA, Latin American Journal of Probability and Mathematical Statistics 19 (2022), 1–22; arXiv:2012.10370, doi:10.30757/ALEA.v19-01.

[31] Ioannis Karatzas and Steven E. Shreve, Brownian Motion and Stochastic Calculus, 2nd ed., Graduate Texts in Mathematics 113, Springer, New York, 1991.

[32] Olga A. Ladyzhenskaya, Vsevolod A. Solonnikov, and Nina N. Ural’tseva, Linear and Quasi-linear Equations of Parabolic Type, American Mathematical Society, Providence, RI, 1968.

[33] Roger W. Lee, The moment formula for implied volatility at extreme strikes, Mathematical Finance 14 (2004), no. 3, 469–480, doi:10.1111/j.0960-1627.2004.00200.x.

[34] Philipp J. Schönbucher, A market model for stochastic implied volatility, Philosophical Transactions of the Royal Society of London A 357 (1999), no. 1758, 2071–2092.

[35] Martin Schweizer and Johannes Wissel, Term structures of implied volatilities: absence of arbitrage and existence results, Mathematical Finance 18 (2008), no. 1, 77–114.

[36] Cédric Villani, Optimal Transport: Old and New, Grundlehren Math. Wiss. 338, Springer, 2009.

[37] Johannes Wiesel, Continuity of the martingale optimal transport problem on the real line, Annals of Applied Probability 33 (2023), no. 6A, 4645–4692, doi:10.1214/22-AAP1928.

How to cite

Zeyu Cao and Shaosai Huang (2026). Finite Gaussian-mixture martingale kernels: density for projected smile-jet readouts. Working paper, version of September 2026. Kspectra Research. SSRN 7444340 (doi:10.2139/ssrn.7444340). https://kspectra.ai/papers/finite-gaussian-mixture-kernels/

@misc{cao2026finite,
  author = {Cao, Zeyu and Huang, Shaosai},
  title  = {{Finite Gaussian-mixture martingale kernels: density for projected smile-jet readouts}},
  year   = {2026},
  month  = sep,
  note   = {Working paper, version of September 2026},
  doi    = {10.2139/ssrn.7444340},
  url    = {https://kspectra.ai/papers/finite-gaussian-mixture-kernels/}
}

For AI tools and text processing, the full paper is also available as Markdown with LaTeX formulas.