Ablating xLSTM blocks against LSTM and Transformers
Controlled block substitutions isolate whether xLSTM’s scalar-memory path, matrix-memory path, or their replacements explain the observed performance.
Compare the benchmark results.
A capacity-matched study of two-block xLSTM, LSTM, and Transformer models on associative recall and formal-language tasks.
MM5 seeds92.0%± 8.8
TT5 seeds28.0%± 5.2
Repetition counts differ by configuration and width, so each row shows its own. MM and TT at d = 8 are the thesis-reported five-seed figures.All configurations remain near chance, so the task does not distinguish the block choices.
Matched substitutions separate, supporting the component-level MQAR conclusion.
Configurations approach saturation, so the apparent differences largely disappear.
Which xLSTM component drives performance: the scalar-memory block, the matrix-memory block, or the LSTM and Transformer blocks used in their place?
- 01
Built a 12-configuration ablation grid across SS, SM, MS, and MM base stacks, replacing sLSTM with LSTM and mLSTM with Transformer blocks.
- 02
Used a shared wrapper with fixed embeddings, output heads, width, depth, sequence length, heads, normalization, and dropout policy for matched substitutions.
- 03
Evaluated each model-benchmark configuration with five independent seeds and reported the mean with a two-sided 95% confidence interval.
- Chose
- Swapped one block at a time inside a fixed wrapper
- Instead of
- Comparing published xLSTM, LSTM, and Transformer models against each other
- Because
- Only a matched substitution attributes a difference to the block itself rather than to width, depth, sequence length, or training setup.
The experiments cover small, two-block models on synthetic MQAR and formal-language tasks. The clearest component comparison comes from width 8 because width 4 was at the performance floor and width 16 approached saturation.
- PyTorch
- xLSTM
- LSTM
- Transformers
- MQAR
- 95% confidence intervals