Reward at Test Time: What Happens When the Objective Reaches the State
Mamba, test-time training and Titans wrote down the same update rule. All three train it by likelihood. Nobody has trained it by reward.
Mamba, test-time training and Titans wrote down the same update rule. All three train it by likelihood. Nobody has trained it by reward.
Mechanistic interpretability was built for softmax transformers. SSMs, diffusion LMs, and sparse MoE each dissolved one of its assumptions. Here's what that actually cost.
HiPPO's polynomial-projection idea is the through-line from S4 to Mamba-3 to HiSS. And it argues that long-context LLMs and long-context prediction are not the same problem.
A practitioner's guide to Mamba and State Space Models: how selective state spaces achieve linear scaling, when to use SSMs vs Transformers vs hybrids, and production-ready models.