VLDB 2026 Research / reviewers in the wild / expert
Markus Spanring
dblp:317/5457
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Deep learning architectures and training · 80% Language models and text generation · 20% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › neural network layer design
gating mechanism |
0.8 | 1 | 2024 | xLSTM: Extended Long Short-Term Memory · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model |
0.8 | 1 | 2024 | xLSTM: Extended Long Short-Term Memory · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM |
0.8 | 1 | 2024 | xLSTM: Extended Long Short-Term Memory · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.8 | 1 | 2024 | xLSTM: Extended Long Short-Term Memory · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › recurrent neural network › LSTM
xLSTM |
0.8 | 1 | 2024 | xLSTM: Extended Long Short-Term Memory · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
residual blocks · 0.8matrix memory · 0.8exponential gating · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | xLSTM: Extended Long Short-Term MemoryabstractIn the 1990s, the constant error carousel and gating were introduced as the central ideas of the Long Short-Term Memory (LSTM). Since then, LSTMs have stood the test of time and contributed to numerous deep learning success stories, in particular they constituted the first Large Language Models (LLMs). However, the advent of the Transformer technology with parallelizable self-attention at its core marked the dawn of a new era, outpacing LSTMs at scale. We now raise a simple question: How far do we get in language modeling when scaling LSTMs to billions of parameters, leveraging the latest techniques from modern LLMs, but mitigating known limitations of LSTMs? Firstly, we introduce exponential gating with appropriate normalization and stabilization techniques. Secondly, we modify the LSTM memory structure, obtaining: (i) sLSTM with a scalar memory, a scalar update, and new memory mixing, (ii) mLSTM that is fully parallelizable with a matrix memory and a covariance update rule. Integrating these LSTM extensions into residual block backbones yields xLSTM blocks that are then residually stacked into xLSTM architectures. Exponential gating and modified memory structures boost xLSTM capabilities to perform favorably when compared to state-of-the-art Transformers and State Space Models, both in performance and scaling. Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp 0001, Günter Klambauer, Johannes Brandstetter, Sepp Hochreiter |
NeurIPS | 3 |