How the smile moves.
A no-arbitrage option surface pins down what the market believes about tomorrow’s range of outcomes. It says almost nothing about how that smile moves when the spot moves. SANOS-Evolve takes the split literally: fix the arbitrage-free smile once, then calibrate a two-timescale martingale kernel for the motion on top — closed-form readouts, arbitrage-free by construction. Here is the design, what nine SPX regimes say about it, and the limitations the evidence exposes.
1 — The gap European prices leave open
2 — Kernel-first: the transition law as the unknown
3 — Two timescales, two readouts
4 — Nine regimes, 2012–2024
5 — Limitations
The gap European prices leave open
Read an option smile at a single maturity and you recover the market’s whole risk-neutral view of where the underlying lands on that date. What you cannot read is what happens next — how the smile itself slides and tilts when the spot moves. Two surfaces that price every vanilla identically today can imply completely different tomorrows, and that difference is the whole game for anything path-dependent or hedged.
This is not a modelling preference. It is a gap in the data. European prices identify the risk-neutral marginals — one distribution per maturity — and they do not identify the transition law that connects them. Infinitely many martingale kernels carry the same sequence of marginals, and they disagree about exactly the thing a hedger cares about.
Each classical tool closes one half of the gap and opens the other. Local volatility reprices today’s smile exactly but gets the motion famously backwards. Stochastic volatility gets the motion but cannot reprice an arbitrary market surface. Stochastic-local volatility does both, at the cost of a fixed-point calibration in which the leverage function and the transition law are solved for together — so the statics and the dynamics are never cleanly separable, and neither is available in closed form.
SANOS-Evolve refuses to solve them together. The marginals are settled first, by a convex program, and then frozen. Only the motion is fitted. That single decision is what buys closed-form readouts and an arbitrage-free surface at the same time.
The approach
Kernel-first: the transition law as the unknown
The usual order of business is to pick a driving process — a diffusion, a jump process, a rough kernel — and then work out what smile dynamics it happens to imply. If the implied dynamics are wrong, you change the process and try again. The dynamics are an output.
Kernel-first inverts that. Conditional on a fixed arbitrage-free marginal surface, the transition kernel is the unknown, and observed dynamic behaviour is what selects it. Stated at that level the construction needs only three ingredients: a marginal engine that produces an arbitrage-free surface, a family of admissible martingale kernels wide enough to contain the dynamics worth arguing about, and a readout that maps a kernel to numbers you can actually measure. Fit the readout; what comes back is a kernel. None of that depends on the particular pieces used here.
Everything then rests on the choice of family, which has to clear two bars that pull against each other.
General enough to span the spot–volatility dynamics the readouts interrogate — the dispersion of the volatility state, how persistently it decays, and how strongly returns couple to it. Tractable enough that those readouts are formulas rather than Monte Carlo estimates, and that propagating the law forward does not leave the family.
A finite Gaussian mixture clears both, and it is worth being concrete about why. Conditional on the carried volatility state and on the within-step return branch, the log-return is Gaussian. So the transition law is a mixture of Gaussians whose weights and moments are explicit functions of the parameters — not a limit, not a discretisation of something else, but the object itself. Generality comes from the mixture: enough branches, and the conditional law can carry skew, excess kurtosis, and a return–state dependence that a single Gaussian cannot.
The tractability comes from a closure that is easy to miss. The marginals are themselves finite Gaussian mixtures — the SANOS program fits weights over a fixed grid of Gaussian bumps in log-moneyness, so the static layer arrives in the same class the kernel propagates. A mixture updated by a Gaussian-mixture kernel is again a mixture. That is the whole reason the readouts stay closed-form: propagation is an analytic mixture update, the local variance and its density are closed forms on the same object, and every quantity downstream is a finite sum rather than a simulation. This is why that marginal engine and not another — the architecture asks only for a convex-ordered engine with an explicit representation, but it is closure under propagation that keeps the arithmetic finite.
Specifying the transition law directly rather than a driving process has an architectural precedent in Markov-functional interest-rate models, which share a low-dimensional Markov driver, a deterministic calibration map and a cheap implementation. The analogy is structural only. What selects the kernel here is not a functional fit to prices but observed dynamic behaviour — and the honest consequence is that what a readout selects is a representative of an equivalence class, every member of which reproduces the same measured dynamics. That is not a defect to be engineered away; it is what the data supports, and Section 5 measures how wide the class is.
The instance
Two timescales, two readouts
The lower row of that diagram is the instance actually built and tested — the narrowest realisation of the architecture that still fits. The static layer is settled first, by the SANOS linear program, and is then fixed and never refit.
The dynamics are a martingale transition kernel laid on top — the rule that carries one maturity’s distribution into the next. It carries a latent volatility state with two mean-reversion timescales, a fast factor and a slow one, each a standardised Gaussian AR(1) process whose autocorrelation is its persistence parameter. That is the familiar fast/slow structure of multi-factor stochastic volatility, written in a form that stays finite and Gaussian at every step.
Two properties are engineered rather than hoped for. The return kernel is normalised by a log-sum-exp, which makes it a martingale by construction — no drift correction, no calibration step. And because the marginals are already fixed, a bare kernel would not reproduce them, so a deterministic leverage overlay aligns the kernel’s conditional variance to the target level.
σ2LV(z, Tj) = σ2Dupire(z, Tj) / 𝔼[ ν ∣ z ]
The discrete Gyöngy identity: Dupire’s local variance, read straight off the fixed SANOS surface, divided by the regime-conditional mean of the volatility factor ν. It sets the kernel’s level pointwise, leaving the kernel free to determine the shape of the motion.
Both guarantees come with a stated boundary, and the paper is careful about where each one stops. The martingale property is exact at the numerical state before the recompression that holds the mixture’s component budget flat; recompression is a law-level projection that preserves mass and the unconditional forward, so after it the property is measured rather than guaranteed. The leverage overlay aligns the variance level to leading order in the time step, and the residual term this leaves is not waved away — it is computed at the production nodes, where it runs to a weighted RMS of a few tenths of a percent to about three percent depending on the regime.
If the marginals are already fixed, prices cannot be the calibration target — they are matched by construction. So what does the kernel fit? Two observable readouts of leading-order at-the-money smile dynamics, chosen because between them they carry both halves of the spot–volatility response.
The first is the skew-stickiness ratio: a normalised measure of how much at-the-money implied volatility moves per unit of spot return, and how that coupling decays with maturity. Because the kernel is translation-invariant in log-moneyness, the SSR is not simulated — it is a closed-form function of the parameters. Five tenors are fitted, from one week to three months.
The second is a forward-variance dispersion readout, which carries the amplitude that the SSR’s normalisation divides out. Fit the SSR alone and the coupling’s shape is pinned while its scale floats; the dispersion block closes that. On SPX it is observed as VIX at-the-money implied volatility, across the six to twelve expiries that survive the filters.
This is not a joint SPX–VIX calibration, and the distinction is load-bearing. The VIX observation enters as a softly weighted reading — an instrument that helps identify the kernel — not as a coequal market to be fitted alongside SPX. The two targets live under different measures, and the paper states the identification restriction that links them rather than quietly asserting an equality. Joint SPX–VIX calibration is an adjacent problem with a harder consistency requirement; this is not a solution to it.
Seven free parameters are fitted against those two blocks under one uniform two-stage protocol — a cold stage, then a warm restart accepted only if the data cost falls — with the block weights scaled so neither dominates on length alone, and a ridge toward each stage’s own starting point.
The evidence
Nine regimes, 2012–2024
The test is whether one construction, run under one unchanged protocol, fits regimes that look nothing like each other. Nine June dates on real SPX end-of-day chains span the flattest year and the steepest, the COVID shock, and the low-volatility melt-ups on either side of it. Nothing is tuned per date: the same two stages, the same weights, the same ridge.
The skew-stickiness term structure comes out with the right shape everywhere — steepest at the short end, rolling off toward the sticky long end — and the closed-form curve sits on the realised estimate tenor by tenor. The level differs enormously between regimes, which is the point: 2024 runs between about 0.6 and 0.9 of a unit of SSR above 2022 at every tenor, and one protocol reaches both.
One caveat on that dashed line. The value 2 is the short-time limit of the skew-stickiness under a diffusion, not a bound on the statistic actually measured over a finite step. In the lowest-volatility years the realised one-week value runs above it, to roughly 2.1–2.5, and the kernel follows it there rather than being clipped.
Across the panel the skew-stickiness residual lands between 0.9% and 3.1% RMS, below the estimated standard error of the target at every single date — the fit is inside the noise of the thing being fitted, which is the most that can be claimed of it. The forward-variance block runs wider, 1.8% to 9.2% RMS, and it is worth being precise about why the two are not comparable: the VIX target carries no sampling band of its own, so its residual is measured against a chosen tolerance rather than a statistical one.
The limits
Limitations
Four limitations survive the study. None is a numerical defect, and each one points somewhere different.
The short end is the weak part. Withhold the one-week tenor from the objective and refit: the model then undershoots it by a median 11%, and misses low at seven of nine dates. Withhold the three-month tenor instead and it is recovered to 3%. So the long end is genuinely implied by the rest of the curve, and the short end is not — the model reaches the realised one-week value when that value is in the objective, and does not get there on its own. A two-timescale diffusion is too smooth at the short end; this is the point where a rough or multiscale kernel would earn its keep.
The decay floor. Porting the whole construction to NDX is where the two-factor family reaches its limit. The skew-stickiness fit survives the move essentially intact — eight of nine years still land near or below two percent — but the forward-variance block blows out, to as much as 24%.
The mechanism is specific rather than mysterious: off index, the realised variance-of-variance falls from thirty to ninety days at a ratio of about 0.4, and two mean-reverting timescales cannot decay faster than about 0.58. The model is being asked for a persistence structure it does not have. Part of that residual is also the physical measure the off-index target has to be quoted in, and with the data available there the two causes cannot be separated — which is itself reported rather than resolved.
What is identified is a class, not a parameter vector. The Jacobian of the readouts has a median effective rank of five against seven free parameters, so distinct parameter vectors reproduce the same readouts along the soft directions. Two accepted fits can differ by about a point of skew-stickiness RMS. What is stable is the readout — under perturbation of both the static date it is calibrated on and the sample it is used out of — and the readout is what gets carried downstream. Read as a contract, that says only quantities that are themselves functionals of the fitted readouts inherit the identification. Everything else is model-implied.
And there is no hedging edge. This one deserves stating plainly, because it is the claim a reader would most want and the evidence does not support it.
What is measured there is a smile-roll residual variance — the part of the at-the-money move that a one-parameter delta family leaves behind — and it is a proxy for hedging error, not a traded profit and loss: no vega weighting, no position, no transaction cost. On that proxy the family is worth a great deal against no adjustment: leaving the smile roll unhedged carries 2.8–3.4× the residual variance of the industry minimum-variance benchmark. Against a competent choice it is worth nothing. Out of sample on 2015–2023, under one train-only affine procedure fitted separately to each competitor and a twenty-row purge between train and test, the calibrated kernel’s own readout, a corrected constant, an analytic smile-implied estimate and a trailing regression all land within a few percent of the benchmark and of each other.
The reason is structural, not a tuning failure. The minimum-variance delta collapses to estimating one number, and one number is estimable many ways — so a model with genuine dynamic content has nowhere to show it. Where that content should show up instead is in quantities the marginals cannot price at all: forward-start and path-dependent structures, whose definition needs exactly the forward implied volatility the marginals leave open. The construction supplies deterministic finite-sum forward densities for that comparison. It has not been run yet.
What the split buys
Keep the smile exact
A convex program fixes strictly arbitrage-free marginals — the market surface, reproduced, then frozen and never refit.
Fit only the motion
A transition kernel with a fast and a slow timescale, a martingale by construction, carries the dynamics local volatility gets wrong.
Read it in closed form
Gaussian mixtures throughout — the skew-stickiness ratio is a formula, the forward-variance dispersion a readout, propagation a finite sum.
Know what you identified
A tolerance class of observationally equivalent kernels, with the soft directions located and the limits measured rather than assumed.