MaRK: Markov-adapted Recurrent Kernels for
Dynamic Operator Conditioning in State Space Models

1 Department of Data Science, City University of Hong Kong
2 NVIDIA AI Technology Center (NVAITC-HK)
Overview of MaRK: a context vector (the diffusion timestep) is mapped to bounded, low-rank modulations of a frozen SSM's recurrence, read-in, read-out, skip and discretization parameters.

We propose an LPV-SSM framework that conditions SSMs via direct recurrent operator modulation by turning the SSM parameters into matrix-valued functions, with provable stability, operator identifiability, and empirical timestep-dependent memory adaptation.

Abstract

State Space Models (SSMs) offer an efficient alternative to Transformers for sequence modeling, yet conditioning pre-trained SSMs for iterative generation typically operates outside the recurrent operator, through input injection or activation modulation. While such mechanisms expose the model to conditioning information, they leave the underlying temporal dynamics fixed. We introduce MaRK (Markov-adapted Recurrent Kernels), a dynamic operator-conditioning framework that maps context vectors directly into bounded modulations of a frozen SSM's recurrence (A), read-in (B), read-out (C), skip (D), and discretization (Δ) parameters. Viewed through the lens of LPV-SSM systems, MaRK induces a context-indexed family of Markov parameter sequences, allowing each diffusion timestep to reshape the model's input-output memory kernel. We instantiate MaRK on a frozen 111M-parameter Hydra SSM backbone and study three adapter geometries: Hypernet, Chebyshev polynomial, and Discrete Cosine Transform kernels. Since these adapters modify the Markov parameter sequence through low-rank auxiliary maps on the frozen backbone, parameter-efficient fine-tuning arises as a structural consequence of the adaptation mechanism itself, requiring only 6.3–11M trainable auxiliary parameters to transition from a bidirectional objective to an iterative diffusion regime. The bounded recurrence parameterization further yields an analytic Affine Quadratic Stability certificate for the modulated recurrence. Through synthetic LPV recovery experiments and Markov-operator diagnostics, we show that MaRK recovers coordinate-invariant temporal operators under matched assumptions and produces distinct, stable timestep-conditioned memory profiles. Empirically, the Chebyshev variant yields the strongest performance, achieving an average validation loss of 2.55, followed by the DCT (2.59) and Hypernet (3.77) geometries.

Conditional generation asks a sequence model to change how it behaves as generation proceeds. For a State Space Model that behaviour is the recurrence: the input–output memory kernel that decides how far back the model looks. Today's conditioning mechanisms deliberately leave that kernel alone:

  1. Input-stream injection concatenates the condition into the sequence. It is cheap, but it asks the hidden state to carry that signal across positions and lags.
  2. Adaptive LayerNorm rescales activations. It works well inside Transformers, but it remains external to the continuous-time SSM operator.

MaRK moves conditioning inside the operator. A context vector (here, the diffusion timestep embedding) is mapped to bounded, low-rank modulations of the frozen backbone's recurrence (A), read-in (B), read-out (C), skip (D) and discretization (Δ) parameters. Because that map is bounded and low-rank, the result is a context-indexed family of stable input–output operators, and parameter-efficient adaptation falls out of the mechanism itself instead of being bolted on afterwards.

If MaRK really conditions the operator, then each frozen diffusion timestep should come with its own memory kernel H(t)={C(t)Ak(t)B(t)}k≥0\mathcal{H}(t) = \{C(t)A^k(t)B(t)\}_{k \geq 0}. We measure the head-averaged Markov norm ∥hk(t)∥2\|h_k(t)\|_2 over lags k≤4096k \le 4096 and compare it with the unmodulated backbone. The curves separate, and because the Markov sequence fixes the transfer function up to a change of state coordinates, that separation is a change of input–output operator, not merely a timestep-dependent activation gain.

2.55
Average CART loss for Chebyshev, the strongest of the three kernels
6.3–11M
Trainable adapter parameters on a frozen 111M-parameter Hydra SSM
O(1)
Conditioning overhead holds constant as sequence length grows

Ablations

We ablate the three MaRK adapter geometries under the same frozen Hydra backbone, diffusion training objective, and CART-weighted validation protocol. The goal is to isolate the effect of the operator-modulation map: a direct Hypernet projection, a smooth Chebyshev polynomial basis, or a bandlimited DCT basis. Table 1 reports LCART\mathcal{L}_{\text{CART}}, where lower values indicate better reconstruction of masked tokens under the same context-adaptive weighting used for training. Trainable adapter size is listed alongside so that geometry effects can be separated from parameter count.

Table 2 isolates which state-space parameters carry the gain, and Table 3 measures the inference cost of the same adapters against two standard conditioning baselines. Every number is reported with a 95% confidence interval.

Both structured bases beat the direct Hypernet projection on every dataset. Chebyshev gives the lowest average loss (2.55) and leads DCT on four of five datasets, with DCT 0.04 nats behind (2.59). The gap is structural, not statistical. The timestep is smooth and adjacent steps do nearly the same denoising, so the modulation map should be smooth too, but the Hypernet MLP reads each timestep independently and has to learn that smoothness from data. Chebyshev and DCT build it into the basis instead.

Modulating the recurrence matrix A is necessary: freezing it costs 0.18 to 0.25 nats, and a Mamba-style B, C, Δ selection leaves 0.34 to 0.51 nats on the table, so MaRK cannot be reduced to input gating. The price is modest and well behaved. At sequence length 4096 the adapters add 14% to 16% latency and under 1% memory over the frozen backbone, and that overhead is flat in sequence length, so it amortizes as sequences grow.

These are internal ablations under a shared frozen backbone, objective and protocol, not a head-to-head comparison against separately trained diffusion language models. Code is in the project repository.

BibTeX


    @inproceedings{ibrahim2026mark,
      author={Syed Ibrahim Omer, Ginny Y. Wong, Xiangyu Zhao},
      booktitle={Advances in Neural Information Processing Systems},
      title={MaRK: Markov-adapted Recurrent Kernels for Dynamic Operator Conditioning in State Space Models},
      url={https://neurips.cc},
      volume={39, Main Conference},
      year={2026},
    }