MaRK: Markov-adapted Recurrent Kernels for
Dynamic Operator Conditioning in State Space Models
Abstract
State Space Models (SSMs) offer an efficient alternative to Transformers for sequence modeling, yet conditioning pre-trained SSMs for iterative generation typically operates outside the recurrent operator, through input injection or activation modulation. While such mechanisms expose the model to conditioning information, they leave the underlying temporal dynamics fixed. We introduce MaRK (Markov-adapted Recurrent Kernels), a dynamic operator-conditioning framework that maps context vectors directly into bounded modulations of a frozen SSM's recurrence (A), read-in (B), read-out (C), skip (D), and discretization (Δ) parameters. Viewed through the lens of LPV-SSM systems, MaRK induces a context-indexed family of Markov parameter sequences, allowing each diffusion timestep to reshape the model's input-output memory kernel. We instantiate MaRK on a frozen 111M-parameter Hydra SSM backbone and study three adapter geometries: Hypernet, Chebyshev polynomial, and Discrete Cosine Transform kernels. Since these adapters modify the Markov parameter sequence through low-rank auxiliary maps on the frozen backbone, parameter-efficient fine-tuning arises as a structural consequence of the adaptation mechanism itself, requiring only 6.3–11M trainable auxiliary parameters to transition from a bidirectional objective to an iterative diffusion regime. The bounded recurrence parameterization further yields an analytic Affine Quadratic Stability certificate for the modulated recurrence. Through synthetic LPV recovery experiments and Markov-operator diagnostics, we show that MaRK recovers coordinate-invariant temporal operators under matched assumptions and produces distinct, stable timestep-conditioned memory profiles. Empirically, the Chebyshev variant yields the strongest performance, achieving an average validation loss of 2.55, followed by the DCT (2.59) and Hypernet (3.77) geometries.
Conditional generation asks a sequence model to change how it behaves as generation proceeds. For a State Space Model that behaviour is the recurrence: the input–output memory kernel that decides how far back the model looks. Today's conditioning mechanisms deliberately leave that kernel alone:
- Input-stream injection concatenates the condition into the sequence. It is cheap, but it asks the hidden state to carry that signal across positions and lags.
- Adaptive LayerNorm rescales activations. It works well inside Transformers, but it remains external to the continuous-time SSM operator.
MaRK moves conditioning inside the operator. A context vector (here, the diffusion timestep embedding) is mapped to bounded, low-rank modulations of the frozen backbone's recurrence (A), read-in (B), read-out (C), skip (D) and discretization (Δ) parameters. Because that map is bounded and low-rank, the result is a context-indexed family of stable input–output operators, and parameter-efficient adaptation falls out of the mechanism itself instead of being bolted on afterwards.
Operator identifiability: Hypernet | Chebyshev | DCT. Markov Operator Error (coordinate-invariant, the difference between the true and estimated Markov parameter sequences) against the number of synthetic trajectories, . All three kernels decay along the reference (red dashed); the fitted rates are for Hypernet, for Chebyshev and for DCT. Chebyshev reaches the lowest structured-basis error, and DCT's larger constant is consistent with its extra bandlimited restriction rather than a loss of recoverability.
Timestep-conditioned Markov operators. Head-resolved mean Markov parameter norm against lag, over lags up to 4096, for the Hypernet, Chebyshev and DCT adapters at . The dashed curve is the frozen Hydra backbone. Chebyshev forms an ordered family of smooth decay profiles, DCT redistributes energy and crosses the backbone at intermediate lags, and Hypernet gives the broadest upward shift at medium and long lags.
Stability certificate. Per-layer stability margin of the pre-trained 23-layer Hydra SSM evaluated at the extreme bounds of MaRK modulation. All 23 layers certify with using the constructive Lyapunov function, so the modulated recurrence stays stable for every admissible conditioning context.
If MaRK really conditions the operator, then each frozen diffusion timestep should come with its own memory kernel . We measure the head-averaged Markov norm over lags and compare it with the unmodulated backbone. The curves separate, and because the Markov sequence fixes the transfer function up to a change of state coordinates, that separation is a change of input–output operator, not merely a timestep-dependent activation gain.
Ablations
We ablate the three MaRK adapter geometries under the same frozen Hydra backbone, diffusion training objective, and CART-weighted validation protocol. The goal is to isolate the effect of the operator-modulation map: a direct Hypernet projection, a smooth Chebyshev polynomial basis, or a bandlimited DCT basis. Table 1 reports , where lower values indicate better reconstruction of masked tokens under the same context-adaptive weighting used for training. Trainable adapter size is listed alongside so that geometry effects can be separated from parameter count.
Table 2 isolates which state-space parameters carry the gain, and Table 3 measures the inference cost of the same adapters against two standard conditioning baselines. Every number is reported with a 95% confidence interval.
Both structured bases beat the direct Hypernet projection on every dataset. Chebyshev gives the lowest average loss (2.55) and leads DCT on four of five datasets, with DCT 0.04 nats behind (2.59). The gap is structural, not statistical. The timestep is smooth and adjacent steps do nearly the same denoising, so the modulation map should be smooth too, but the Hypernet MLP reads each timestep independently and has to learn that smoothness from data. Chebyshev and DCT build it into the basis instead.
Modulating the recurrence matrix A is necessary: freezing it costs 0.18 to 0.25 nats, and a Mamba-style B, C, Δ selection leaves 0.34 to 0.51 nats on the table, so MaRK cannot be reduced to input gating. The price is modest and well behaved. At sequence length 4096 the adapters add 14% to 16% latency and under 1% memory over the frozen backbone, and that overhead is flat in sequence length, so it amortizes as sequences grow.
These are internal ablations under a shared frozen backbone, objective and protocol, not a head-to-head comparison against separately trained diffusion language models. Code is in the project repository.
BibTeX
@inproceedings{ibrahim2026mark,
author={Syed Ibrahim Omer, Ginny Y. Wong, Xiangyu Zhao},
booktitle={Advances in Neural Information Processing Systems},
title={MaRK: Markov-adapted Recurrent Kernels for Dynamic Operator Conditioning in State Space Models},
url={https://neurips.cc},
volume={39, Main Conference},
year={2026},
}