Reward at Test Time: What Happens When the Objective Reaches the State
Mamba, test-time training and Titans wrote down the same update rule. All three train it by likelihood. Nobody has trained it by reward.
Mamba, test-time training and Titans wrote down the same update rule. All three train it by likelihood. Nobody has trained it by reward.
Four generations of RL, each one deleting a piece of the pipeline. The last one deleted the text, and stopped publishing.
Mechanistic interpretability was built for softmax transformers. SSMs, diffusion LMs, and sparse MoE each dissolved one of its assumptions. Here's what that actually cost.
A city shipped a 'homegrown' AI; subtraction proved most of it was someone else's. The toolkit for catching model fraud, and its honest limits.
Spotify deleted the BPM API and underground techno has no genre tag anywhere, so I stopped looking the genre up and started computing it from the audio.
HiPPO's polynomial-projection idea is the through-line from S4 to Mamba-3 to HiSS. And it argues that long-context LLMs and long-context prediction are not the same problem.
An EBM finally crossed 800M parameters without collapsing. Nobody has independently reproduced the 35% scaling claim. Both halves matter.
A practitioner's guide to Mamba and State Space Models: how selective state spaces achieve linear scaling, when to use SSMs vs Transformers vs hybrids, and production-ready models.
VLA models give robots the ability to see, parse language, and execute physical actions through a single architecture. This post covers how they work, what shipped in 2025, and where the limitations are.
What if your AI worked offline, kept your secrets, and actually remembered you, without ever flinching at a spotty network? This post moves past the "API-everywhere" playbook. It lays out theory for a...