Carmen Amo Alonso

dblp:234/2243 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0001-7593-5992ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Deep learning architectures and training · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
sequence modeling
1.622025
Lambda-Skip Connections: the architectural component that prevents Rank Collapse · ICLR 2025
Understanding the Differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks · NeurIPS 2024
Machine learning › Deep learning architectures and training
state space model
1.622025
Lambda-Skip Connections: the architectural component that prevents Rank Collapse · ICLR 2025
Understanding the Differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks · NeurIPS 2024
Machine learning › Deep learning architectures and training
attention mechanism
0.812024
Understanding the Differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks · NeurIPS 2024
Machine learning › Deep learning architectures and training › attention mechanism › efficient attention
linear attention
0.812024
Understanding the Differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks · NeurIPS 2024
Machine learning › Deep learning architectures and training
skip connections
0.312025
Lambda-Skip Connections: the architectural component that prevents Rank Collapse · ICLR 2025
Machine learning › Deep learning architectures and training
transformer
0.312025
Lambda-Skip Connections: the architectural component that prevents Rank Collapse · ICLR 2025
Machine learning › Deep learning architectures and training
recurrent neural network
0.212024
Understanding the Differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

analytical analysis · 0.9ablation study · 0.9dynamical systems framework · 0.8
YearPublicationVenuePosition
2025 Lambda-Skip Connections: the architectural component that prevents Rank Collapse
abstract
Rank collapse, a phenomenon where embedding vectors in sequence models rapidly converge to a uniform token or equilibrium state, has recently gained at- tention in the deep learning literature. This phenomenon leads to reduced expres- sivity and potential training instabilities due to vanishing gradients. Empirical ev- idence suggests that architectural components like skip connections, LayerNorm, and MultiLayer Perceptrons (MLPs) play critical roles in mitigating rank collapse. While this issue is well-documented for transformers, alternative sequence mod- els, such as State Space Models (SSMs), which have recently gained prominence, have not been thoroughly examined for similar vulnerabilities. This paper extends the theory of rank collapse from transformers to SSMs using a unifying frame- work that captures both architectures. We introduce a modification in the skip connection component, termed lambda-skip connections, that provides guaran- tees for rank collapse prevention. We present, via analytical results, a sufficient condition to achieve the guarantee for all of the aforementioned architectures. We also study the necessity of this condition via ablation studies and analytical exam- ples. To our knowledge, this is the first study that provides a general guarantee to prevent rank collapse, and that investigates rank collapse in the context of SSMs, offering valuable understanding for both theoreticians and practitioners. Finally, we validate our findings with experiments demonstrating the crucial role of archi- tectural components in preventing rank collapse.
Federico Arangath Joseph, Jerome Sieber, Melanie Nicole Zeilinger, Carmen Amo Alonso
ICLR4
2025 Bridging Expressivity and Scalability with Adaptive Unitary SSMs
abstract
Recent work has revealed that state space models (SSMs), while efficient for long-sequence processing, are fundamentally limited in their ability to represent formal languages—particularly due to time-invariant and real-valued recurrence structures. In this work, we draw inspiration from adaptive and structured dynamics observed in biological neural systems and introduce the Adaptive Unitary State Space Model (AUSSM): a novel class of SSMs that leverages skew-symmetric, input-dependent recurrence to achieve unitary evolution and high expressive power. Using algebraic automata theory, we prove that AUSSM can perform modulo counting and simulate solvable group automata at precision logarithmically bounded in the input length, enabling SSMs to model a broad class of regular languages out of reach for other SSM architectures. To overcome the practical inefficiencies of adaptive recurrence, we develop a separable convolution formulation and a CUDA implementation that enables scalable parallel training. Empirically, we show that AUSSM and its hybrid variant—interleaved with Mamba—outperform prior SSMs on formal algorithmic tasks such as parity and modular arithmetic, and achieve competent performance on real-world long time-series classification benchmarks. Our results demonstrate that adaptive unitary recurrence provides a powerful and efficient inductive bias for both symbolic and continuous sequence modeling. The code is available at https://github.com/arjunkaruvally/AUSSM
Arjun Karuvally, Franz Nowak, T. Anderson Keller 0001, Carmen Amo Alonso, Terrence J. Sejnowski, Hava T. Siegelmann
NeurIPS4
2024 NARRATE: Versatile Language Architecture for Optimal Control in Robotics
abstract
The impressive capabilities of Large Language Models (LLMs) have led to various efforts in enabling robots to be controlled through natural language instructions, opening exciting possibilities for human-robot interaction. The goal is for the motor-control task to be performed accurately, efficiently and safely while also enjoying the flexibility imparted by LLMs to specify and adjust the task through natural language. In this work, we demonstrate how a careful layering of an LLM in combination with a Model Predictive Control (MPC) formulation allows for accurate and flexible robotic control via natural language while taking into consideration safety constraints. In particular, we rely on the LLM to effectively frame constraints and objective functions as mathematical expressions, which are later used in the motor-control module via MPC. The transparency of the optimization formulation allows for interpretability of the task and enables adjustments through human feedback. We demonstrate the validity of our method through extensive experiments on long-horizon reasoning, contact-rich, and multi-object interaction tasks. Our evaluations show that NARRATE outperforms current existing methods on these benchmarks and effectively transfers to the real world on two different embodiments.Videos, Code and Prompts at narrate-mpc.github.io
Seif Ismail, Antonio Arbues, Ryan Cotterell, René Zurbrügg, Carmen Amo Alonso
IROS5
2024 Understanding the Differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks
abstract
Softmax attention is the principle backbone of foundation models for various artificial intelligence applications, yet its quadratic complexity in sequence length can limit its inference throughput in long-context settings. To address this challenge, alternative architectures such as linear attention, State Space Models (SSMs), and Recurrent Neural Networks (RNNs) have been considered as more efficient alternatives. While connections between these approaches exist, such models are commonly developed in isolation and there is a lack of theoretical understanding of the shared principles underpinning these architectures and their subtle differences, greatly influencing performance and scalability. In this paper, we introduce the Dynamical Systems Framework (DSF), which allows a principled investigation of all these architectures in a common representation. Our framework facilitates rigorous comparisons, providing new insights on the distinctive characteristics of each model class. For instance, we compare linear attention and selective SSMs, detailing their differences and conditions under which both are equivalent. We also provide principled comparisons between softmax attention and other model classes, discussing the theoretical conditions under which softmax attention can be approximated. Additionally, we substantiate these new insights with empirical validations and mathematical arguments. This shows the DSF's potential to guide the systematic development of future more efficient and scalable foundation models.
Jerome Sieber, Carmen Amo Alonso, Alexandre Didier, Melanie Nicole Zeilinger, Antonio Orvieto
NeurIPS2