EDBT 2026 Demo / reviewers in the wild / expert
Róbert Csordás
dblp:166/4773
· DBLP profile ↗
16ranked-venue papers
8as first author
15since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 8 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MrT5: Dynamic Token Merging for Efficient Byte-level Language ModelsabstractModels that rely on subword tokenization have significant drawbacks, such as sensitivity to character-level noise like spelling errors and inconsistent compression rates across different languages and scripts. While character- or byte-level models like ByT5 attempt to address these concerns, they have not gained widespread adoption—processing raw byte streams without tokenization results in significantly longer sequence lengths, making training and inference inefficient. This work introduces MrT5 (MergeT5), a more efficient variant of ByT5 that integrates a token deletion mechanism in its encoder to dynamically shorten the input sequence length. After processing through a fixed number of encoder layers, a learned delete gate determines which tokens are to be removed and which are to be retained for subsequent layers. MrT5 effectively "merges" critical information from deleted tokens into a more compact sequence, leveraging contextual information from the remaining tokens. In continued pre-training experiments, we find that MrT5 can achieve significant gains in inference runtime with minimal effect on performance, as measured by bits-per-byte. Additionally, with multilingual training, MrT5 adapts to the orthographic characteristics of each language, learning language-specific compression rates. Furthermore, MrT5 shows comparable accuracy to ByT5 on downstream evaluations such as XNLI, TyDi QA, and character-level tasks while reducing sequence lengths by up to 75%. Our approach presents a solution to the practical limitations of existing byte-level models. Julie Kallini, Shikhar Murty, Christopher D. Manning, Christopher Potts, Róbert Csordás |
ICLR | 5 |
| 2025 | Measuring In-Context Computation Complexity via Hidden State PredictionabstractDetecting when a neural sequence model does "interesting" computation is an open problem. The next token prediction loss is a poor indicator: Low loss can stem from trivially predictable sequences that are uninteresting, while high loss may reflect unpredictable but also irrelevant information that can be ignored by the model. We propose a better metric: measuring the model’s ability to predict its own future hidden states. We show empirically that this metric–in contrast to the next token prediction loss–correlates with the intuitive interestingness of the task. To measure predictability, we introduce the architecture-agnostic "prediction of hidden states" (PHi) layer that serves as an information bottleneck on the main pathway of the network (e.g., the residual stream in Transformers). We propose a novel learned predictive prior that enables us to measure the novel information gained in each computation step, which serves as our metric. We show empirically that our metric predicts the description length of formal languages learned in-context, the complexity of mathematical reasoning problems, and the correctness of self-generated reasoning chains. Vincent Herrmann, Róbert Csordás, Jürgen Schmidhuber |
ICML | 2 |
| 2025 | Do Language Models Use Their Depth Efficiently?abstractModern LLMs are increasingly deep, and depth correlates with performance, albeit with diminishing returns. However, do these models use their depth efficiently? Do they compose more features to create higher-order computations that are impossible in shallow models, or do they merely spread the same kinds of computation out over more layers? To address these questions, we analyze the residual stream of the Llama 3.1, Qwen 3, and OLMo 2 family of models. We find: First, comparing the output of the sublayers to the residual stream reveals that layers in the second half contribute much less than those in the first half, with a clear phase transition between the two halves. Second, skipping layers in the second half has a much smaller effect on future computations and output predictions. Third, for multihop tasks, we are unable to find evidence that models are using increased depth to compose subresults in examples involving many hops. Fourth, we seek to directly address whether deeper models are using their additional layers to perform new kinds of computation. To do this, we train linear maps from the residual stream of a shallow model to a deeper one. We find that layers with the same relative depth map best to each other, suggesting that the larger model simply spreads the same computations out over its many layers. All this evidence suggests that deeper models are not using their depth to learn new kinds of computation, but only using the greater depth to perform more fine-grained adjustments to the residual. This may help explain why increasing scale leads to diminishing returns for stacked Transformer architectures. Róbert Csordás, Christopher D. Manning, Christopher Potts |
NeurIPS | 1 |
| 2025 | Mindstorms in Natural Language-Based Societies of MindabstractInspired by Minsky's Society of Mind, Schmidhuber's Learning to Think, and other more recent works, this paper proposes and advocates for the concept of natural language-based societies of mind (NLSOMs). We imagine these societies as consisting of a collection of multimodal neural networks, including large language models, which engage in a “mindstorm” to solve problems using a shared natural language interface. Here, we work to identify and discuss key questions about the social structure, governance, and economic principles for NLSOMs, emphasizing their impact on the future of AI. Our demonstrations with NLSOMs-which feature up to 129 agents-show their effectiveness in various tasks, including visual question answering, image captioning, and prompt generation for text-to-image synthesis. Mingchen Zhuge, Francesco Faccio, Dylan R. Ashley, Róbert Csordás, Anand Gopalakrishnan, Abdullah Hamdi, Hasan Hammoud, Vincent Herrmann, Kazuki Irie, Louis Kirsch, Bing Li 0024, Guohao Li 0001, Shuming Liu 0001, Jinjie Mai, Piotr Piekos, Aditya A. Ramesh, Imanol Schlag, Aleksandar Stanic, Yuhui Wang 0004, Mengmeng Xu 0006, Deng-Ping Fan, Bernard Ghanem, Jürgen Schmidhuber |
Comput. Vis. Media | 5 |
| 2024 | Self-organising Neural Discrete Representation Learning à la Kohonen
Kazuki Irie, Róbert Csordás, Jürgen Schmidhuber |
ICANN (1) | 2 |
| 2024 | MoEUT: Mixture-of-Experts Universal TransformersabstractPrevious work on Universal Transformers (UTs) has demonstrated the importance of parameter sharing across layers. By allowing recurrence in depth, UTs have advantages over standard Transformers in learning compositional generalizations, but layer-sharing comes with a practical limitation of parameter-compute ratio: it drastically reduces the parameter count compared to the non-shared model with the same dimensionality. Naively scaling up the layer size to compensate for the loss of parameters makes its computational resource requirements prohibitive. In practice, no previous work has succeeded in proposing a shared-layer Transformer design that is competitive in parameter count-dominated tasks such as language modeling. Here we propose MoEUT (pronounced "moot"), an effective mixture-of-experts (MoE)-based shared-layer Transformer architecture, which combines several recent advances in MoEs for both feedforward and attention layers of standard Transformers together with novel layer-normalization and grouping schemes that are specific and crucial to UTs. The resulting UT model, for the first time, slightly outperforms standard Transformers on language modeling tasks such as BLiMP and PIQA, while using significantly less compute and memory. Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber, Christopher Potts, Christopher D. Manning |
NeurIPS | 1 |
| 2024 | SwitchHead: Accelerating Transformers with Mixture-of-Experts AttentionabstractDespite many recent works on Mixture of Experts (MoEs) for resource-efficient Transformer language models, existing methods mostly focus on MoEs for feedforward layers. Previous attempts at extending MoE to the self-attention layer fail to match the performance of the parameter-matched baseline. Our novel SwitchHead is an effective MoE method for the attention layer that successfully reduces both the compute and memory requirements, achieving wall-clock speedup, while matching the language modeling performance of the baseline Transformer. Our novel MoE mechanism allows SwitchHead to compute up to 8 times fewer attention matrices than the standard Transformer. SwitchHead can also be combined with MoE feedforward layers, resulting in fully-MoE "SwitchAll" Transformers. For our 262M parameter model trained on C4, SwitchHead matches the perplexity of standard models with only 44% compute and 27% memory usage. Zero-shot experiments on downstream tasks confirm the performance of SwitchHead, e.g., achieving more than 3.5% absolute improvements on BliMP compared to the baseline with an equal compute resource. Róbert Csordás, Piotr Piekos, Kazuki Irie, Jürgen Schmidhuber |
NeurIPS | 1 |
| 2023 | Practical Computational Power of Linear Transformers and Their Recurrent and Self-Referential ExtensionsabstractRecent studies of the computational power of recurrent neural networks (RNNs) reveal a hierarchy of RNN architectures, given real-time and finite-precision assumptions.Here we study auto-regressive Transformers with linearised attention, a.k.a.linear Transformers (LTs) or Fast Weight Programmers (FWPs).LTs are special in the sense that they are equivalent to RNN-like sequence processors with a fixed-size state, while they can also be expressed as the now-popular self-attention networks.We show that many well-known results for the standard Transformer directly transfer to LTs/FWPs.Our formal language recognition experiments demonstrate how recently proposed FWP extensions such as recurrent FWPs and self-referential weight matrices successfully overcome certain limitations of the LT, e.g., allowing for generalisation on the parity problem.Our code is public.1 Kazuki Irie, Róbert Csordás, Jürgen Schmidhuber |
EMNLP | 2 |
| 2022 | CTL++: Evaluating Generalization on Never-Seen Compositional Patterns of Known Functions, and Compatibility of Neural RepresentationsabstractWell-designed diagnostic tasks have played a key role in studying the failure of neural nets (NNs) to generalize systematically.Famous examples include SCAN and Compositional Table Lookup (CTL).Here we introduce CTL++, a new diagnostic dataset based on compositions of unary symbolic functions.While the original CTL is used to test length generalization or productivity, CTL++ is designed to test systematicity of NNs, that is, their capability to generalize to unseen compositions of known functions.CTL++ splits functions into groups and tests performance on group elements composed in a way not seen during training.We show that recent CTL-solving Transformer variants fail on CTL++.The simplicity of the task design allows for fine-grained control of task difficulty, as well as many insightful analyses.For example, we measure how much overlap between groups is needed by tested NNs for learning to compose.We also visualize how learned symbol representations in outputs of functions from different groups are compatible in case of success but not in case of failure.These results provide insights into failure cases reported on more complex compositions in the natural language domain.Our code is public.1 Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber |
EMNLP | 1 |
| 2022 | The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization
Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber |
ICLR | 1 |
| 2022 | The Dual Form of Neural Networks Revisited: Connecting Test Time Predictions to Training Patterns via Spotlights of AttentionabstractLinear layers in neural networks (NNs) trained by gradient descent can be expressed as a key-value memory system which stores all training datapoints and the initial weights, and produces outputs using unnormalised dot attention over the entire training experience. While this has been technically known since the 1960s, no prior work has effectively studied the operations of NNs in such a form, presumably due to prohibitive time and space complexities and impractical model sizes, all of them growing linearly with the number of training patterns which may get very large. However, this dual formulation offers a possibility of directly visualising how an NN makes use of training patterns at test time, by examining the corresponding attention weights. We conduct experiments on small scale supervised image classification tasks in single-task, multi-task, and continual learning settings, as well as language modelling, and discuss potentials and limits of this view for better understanding and interpreting how NNs exploit training patterns. Our code is public. Kazuki Irie, Róbert Csordás, Jürgen Schmidhuber |
ICML | 2 |
| 2022 | A Modern Self-Referential Weight Matrix That Learns to Modify ItselfabstractThe weight matrix (WM) of a neural network (NN) is its program. The programs of many traditional NNs are learned through gradient descent in some error function, then remain fixed. The WM of a self-referential NN, however, can keep rapidly modifying all of itself during runtime. In principle, such NNs can meta-learn to learn, and meta-meta-learn to meta-learn to learn, and so on, in the sense of recursive self-improvement. While NN architectures potentially capable of implementing such behaviour have been proposed since the ’90s, there have been few if any practical studies. Here we revisit such NNs, building upon recent successes of fast weight programmers and closely related linear Transformers. We propose a scalable self-referential WM (SRWM) that learns to use outer products and the delta update rule to modify itself. We evaluate our SRWM in supervised few-shot learning and in multi-task reinforcement learning with procedurally generated game environments. Our experiments demonstrate both practical applicability and competitive performance of the proposed SRWM. Our code is public. Kazuki Irie, Imanol Schlag, Róbert Csordás, Jürgen Schmidhuber |
ICML | 3 |
| 2021 | The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of TransformersabstractRecently, many datasets have been proposed to test the systematic generalization ability of neural networks.The companion baseline Transformers, typically trained with default hyper-parameters from standard tasks, are shown to fail dramatically.Here we demonstrate that by revisiting model configurations as basic as scaling of embeddings, early stopping, relative positional embedding, and Universal Transformer variants, we can drastically improve the performance of Transformers on systematic generalization.We report improvements on five popular datasets: SCAN, CFQ, PCFG, COGS, and Mathematics dataset.Our models improve accuracy from 50% to 85% on the PCFG productivity split, and from 35% to 81% on COGS.On SCAN, relative positional embedding largely mitigates the EOS decision problem (Newman et al., 2020), yielding 100% accuracy on the length split with a cutoff at 26. Importantly, performance differences between these models are typically invisible on the IID data split.This calls for proper generalization validation sets for developing neural networks that generalize systematically.We publicly release the code to reproduce our results 1 . Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber |
EMNLP (1) | 1 |
| 2021 | Are Neural Nets Modular? Inspecting Functional Modularity Through Differentiable Weight Masks
Róbert Csordás, Sjoerd van Steenkiste, Jürgen Schmidhuber |
ICLR | 1 |
| 2021 | Going Beyond Linear Transformers with Recurrent Fast Weight ProgrammersabstractTransformers with linearised attention (''linear Transformers'') have demonstrated the practical scalability and effectiveness of outer product-based Fast Weight Programmers (FWPs) from the '90s. However, the original FWP formulation is more general than the one of linear Transformers: a slow neural network (NN) continually reprograms the weights of a fast NN with arbitrary architecture. In existing linear Transformers, both NNs are feedforward and consist of a single layer. Here we explore new variations by adding recurrence to the slow and fast nets. We evaluate our novel recurrent FWPs (RFWPs) on two synthetic algorithmic tasks (code execution and sequential ListOps), Wikitext-103 language models, and on the Atari 2600 2D game environment. Our models exhibit properties of Transformers and RNNs. In the reinforcement learning setting, we report large improvements over LSTM in several Atari games. Our code is public. Kazuki Irie, Imanol Schlag, Róbert Csordás, Jürgen Schmidhuber |
NeurIPS | 3 |
| 2019 | Improving Differentiable Neural Computers Through Memory Masking, De-allocation, and Link Distribution Sharpness Control
Róbert Csordás, Jürgen Schmidhuber |
ICLR (Poster) | 1 |