PROJECT / 02Sequence models · Experimental research

Ablating xLSTM blocks against LSTM and Transformers

Controlled block substitutions isolate whether xLSTM’s scalar-memory path, matrix-memory path, or their replacements explain the observed performance.

Role
Bachelor's thesis research and implementation
Context
Bachelor's thesis · two-block sequence models
Period
2025
INTERACTIVE EVIDENCE

Compare the benchmark results.

A capacity-matched study of two-block xLSTM, LSTM, and Transformer models on associative recall and formal-language tasks.

Thesis experiment results
PRIMARY / MM
MmLSTM
MmLSTM
original mLSTM + mLSTM
Ablation[SS]
Ablation[SM]
Ablation[MS]
Ablation[MM]
MQAR · N=128 · D=32 · vocabulary 8192mean ± 95% CI · compare up to 4
MQAR validation accuracy across model widthMM (original mLSTM + mLSTM): 1.4 percent at width 4, 92.0 percent at width 8, 98.7 percent at width 16. TT (replace mLSTM with Transformer): 0.2 percent at width 4, 28.0 percent at width 8, 98.7 percent at width 160255075100d = 4d = 8d = 16model dimension
d = 8

MM5 seeds92.0%± 8.8

TT5 seeds28.0%± 5.2

Repetition counts differ by configuration and width, so each row shows its own. MM and TT at d = 8 are the thesis-reported five-seed figures.
d = 4Floor regime

All configurations remain near chance, so the task does not distinguish the block choices.

d = 8Informative regime

Matched substitutions separate, supporting the component-level MQAR conclusion.

d = 16Ceiling regime

Configurations approach saturation, so the apparent differences largely disappear.

THE QUESTION

Which xLSTM component drives performance: the scalar-memory block, the matrix-memory block, or the LSTM and Transformer blocks used in their place?

HOW I APPROACHED IT
  1. 01

    Built a 12-configuration ablation grid across SS, SM, MS, and MM base stacks, replacing sLSTM with LSTM and mLSTM with Transformer blocks.

  2. 02

    Used a shared wrapper with fixed embeddings, output heads, width, depth, sequence length, heads, normalization, and dropout policy for matched substitutions.

  3. 03

    Evaluated each model-benchmark configuration with five independent seeds and reported the mean with a two-sided 95% confidence interval.

The call I made
Chose
Swapped one block at a time inside a fixed wrapper
Instead of
Comparing published xLSTM, LSTM, and Transformer models against each other
Because
Only a matched substitution attributes a difference to the block itself rather than to width, depth, sequence length, or training setup.
Configurations12 two-block stacks
Repetitions5 independent seeds
UncertaintyMean ± 95% CI
RESEARCH SCOPE

The experiments cover small, two-block models on synthetic MQAR and formal-language tasks. The clearest component comparison comes from width 8 because width 4 was at the performance floor and width 16 approached saturation.

TOOLS & METHODS
  • PyTorch
  • xLSTM
  • LSTM
  • Transformers
  • MQAR
  • 95% confidence intervals
Inspect the source repository