EDBT 2026 Demo / reviewers in the wild / expert
Pedro A. M. Mediano
dblp:190/7253
· DBLP profile ↗
18ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0003-1789-5894ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Theory of computation · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Logarithmic Decomposition and a Signed Measure Space for EntropyabstractThe Shannon entropy behaves in many ways like a signed measure. Previous work has explored this connection by defining a signed measure on abstract sets, representing the information that different variables contain. This yields many measure-theoretic counterparts to information quantities such as the mutual information (set intersection), the joint entropy (set union), and the conditional entropy (set difference), but an explicit construction of the elements of these sets remains ambiguous. Here we provide a concrete characterisation of these abstract sets and a corresponding signed measure by extending Yeung’s I-measure to all possible outcomes in an outcome space Ω, and in doing so we demonstrate a much finer decomposition with intuitive properties that we call the logarithmic decomposition (LD). We show that this signed measure space has the useful property that its logarithmic atoms can be characterised as having negative or positive entropy (depending only on their structure), while being consistent with classical information quantities. Using this logarithmic decomposition, we re-examine the Gács-Körner common information, the functional common information, and minimally sufficient statistics from this new perspective and characterise them in terms of our logarithmic atoms—a property we call logarithmic decomposability. We present possible extensions of this construction to continuous probability distributions, before finally applying our new decomposition to the Dyadic and Triadic systems of James and Crutchfield as a motivating example. We show that, in contrast to the I-measure alone, our classically derived decomposition qualitatively distinguishes between them. Keenan J. A. Down, Pedro A. M. Mediano |
IEEE Trans. Inf. Theory | 2 |
| 2025 | From Lazy to Rich: Exact Learning Dynamics in Deep Linear NetworksabstractBiological and artificial neural networks develop internal representations that enable them to perform complex tasks. In artificial networks, the effectiveness of these models relies on their ability to build task specific representation, a process influenced by interactions among datasets, architectures, initialization strategies, and optimization algorithms. Prior studies highlight that different initializations can place networks in either a lazy regime, where representations remain static, or a rich/feature learning regime, where representations evolve dynamically. Here, we examine how initialization influences learning dynamics in deep linear neural networks, deriving exact solutions for lambda-balanced initializations-defined by the relative scale of weights across layers. These solutions capture the evolution of representations and the Neural Tangent Kernel across the spectrum from the rich to the lazy regimes. Our findings deepen the theoretical understanding of the impact of weight initialization on learning regimes, with implications for continual learning, reversal learning, and transfer learning, relevant to both neuroscience and practical applications. Clémentine C. J. Dominé, Nicolas Anguita, Alexandra Maria Proca, Lukas Braun, Daniel Kunin, Pedro A. M. Mediano, Andrew M. Saxe |
ICLR | 6 |
| 2025 | Grokking at the Edge of Numerical StabilityabstractGrokking, or sudden generalization that occurs after prolonged overfitting, is a surprising phenomenon that has challenged our understanding of deep learning. While a lot of progress has been made in understanding grokking, it is still not clear why generalization is delayed and why grokking often does not happen without regularization. In this work we argue that without regularization, grokking tasks push models to the edge of numerical stability, introducing floating point errors in the Softmax that we refer to as _Softmax Collapse_ (SC). We show that SC prevents grokking and that mitigating SC leads to grokking _without_ regularization. Investigating the root cause of SC, we find that beyond the point of overfitting, the gradients strongly align with what we call the _naïve loss minimization_ (NLM) direction. This component of the gradient does not change the predictions of the model but decreases the loss by scaling the logits, usually through the scaling of the weights along their current direction. We show that this scaling of the logits explains the delay in generalization characteristic of grokking, and eventually leads to SC, stopping learning altogether. To validate these hypotheses, we introduce two key contributions that mitigate the issues faced in grokking tasks: (i) $\mathrm{StableMax}$, a new activation function that prevents SC and enables grokking without regularization, and (ii) $\perp\mathrm{Grad}$, a training algorithm that leads to quick generalization in grokking tasks by preventing NLM altogether. These contributions provide new insights into grokking, shedding light on its delayed generalization, reliance on regularization, and the effectiveness of known grokking-inducing methods. Lucas Prieto, Melih Barsbey, Pedro A. M. Mediano, Tolga Birdal |
ICLR | 3 |
| 2025 | Learning dynamics in linear recurrent neural networksabstractRecurrent neural networks (RNNs) are powerful models used widely in both machine learning and neuroscience to learn tasks with temporal dependencies and to model neural dynamics. However, despite significant advancements in the theory of RNNs, there is still limited understanding of their learning process and the impact of the temporal structure of data. Here, we bridge this gap by analyzing the learning dynamics of linear RNNs (LRNNs) analytically, enabled by a novel framework that accounts for task dynamics. Our mathematical result reveals four key properties of LRNNs: (1) Learning of data singular values is ordered by both scale and temporal precedence, such that singular values that are larger and occur later are learned faster. (2) Task dynamics impact solution stability and extrapolation ability. (3) The loss function contains an effective regularization term that incentivizes small weights and mediates a tradeoff between recurrent and feedforward computation. (4) Recurrence encourages feature learning, as shown through a novel derivation of the neural tangent kernel for finite-width LRNNs. As a final proof-of-concept, we apply our theoretical framework to explain the behavior of LRNNs performing sensory integration tasks. Our work provides a first analytical treatment of the relationship between the temporal dependencies in tasks and learning dynamics in LRNNs, building a foundation for understanding how complex dynamic behavior emerges in cognitive models. Alexandra Maria Proca, Clémentine C. J. Dominé, Murray Shanahan, Pedro A. M. Mediano |
ICML | 4 |
| 2025 | Understanding the high-order network plasticity mechanisms of ultrasound neuromodulationabstractTranscranial ultrasound stimulation (TUS) is an emerging non-invasive neuromodulation technique, offering a potential alternative to pharmacological treatments for psychiatric and neurological disorders. While functional analysis has been instrumental in characterizing the TUS effects, understanding its indirect influence across the network remains challenging. Here, we developed a whole-brain model to represent functional changes as measured by fMRI, enabling us to investigate how TUS-induced effects propagate throughout the brain with increasing stimulus intensity. We implemented two mechanisms: one based on anatomical distance and another on broadcasting dynamics, to explore plasticity-driven changes in specific brain regions. Finally, we highlighted the role of higher-order functional interactions in localizing spatial effects of off-line TUS at two target areas-the right thalamus and inferior frontal cortex-revealing distinct patterns of functional reorganization. This work lays the foundation for mechanistic insights and predictive models of TUS, advancing its potential clinical applications. Marilyn Gatica, Cyril Atkinson-Clement, Carlos Coronel-Oliveros, Mohammad Alkhawashki, Pedro A. M. Mediano, Enzo Tagliazucchi, Fernando Rosas, Marcus Kaiser, Giovanni Petri |
PLoS Comput. Biol. | 5 |
| 2025 | Null models for comparing information decomposition across complex systemsabstractA key feature of information theory is its universality, as it can be applied to study a broad variety of complex systems. However, many information-theoretic measures can vary significantly even across systems with similar properties, making normalisation techniques essential for allowing meaningful comparisons across datasets. Inspired by the framework of Partial Information Decomposition (PID), here we introduce Null Models for Information Theory (NuMIT), a null model-based non-linear normalisation procedure which improves upon standard entropy-based normalisation approaches and overcomes their limitations. We provide practical implementations of the technique for systems with different statistics, and showcase the method on synthetic models and on human neuroimaging data. Our results demonstrate that NuMIT provides a robust and reliable tool to characterise complex systems of interest, allowing cross-dataset comparisons and providing a meaningful significance test for PID analyses. Alberto Liardi, Fernando Rosas, Robin L. Carhart-Harris, George Blackburne, Daniel Bor, Pedro A. M. Mediano |
PLoS Comput. Biol. | 6 |
| 2025 | The topology of synergy: Linking topological and information-theoretic approaches to higher-order interactions in complex systemsabstractThe study of irreducible higher-order interactions has become a core topic of study in complex systems, as they provide a formal scaffold around which to build a quantitative understanding of emergence and emergent properties. Two of the most well-developed frameworks, topological data analysis and multivariate information theory, aim to provide formal tools for identifying higher-order interactions in empirical data. Despite similar aims, however, these two approaches are built on markedly different mathematical foundations and have been developed largely in parallel - with limited interdisciplinary cross-talk between them. In this study, we present a head-to-head comparison of topological data analysis and information-theoretic approaches to describing higher-order interactions in multivariate data; with the goal of assessing the similarities, and differences, between how the frameworks define "higher-order structures." We begin with toy examples with known topologies (spheres, toroids, planes, and knots), before turning to more complex, naturalistic data: fMRI signals collected from the human brain. We find that intrinsic, higher-order synergistic information is associated with three-dimensional cavities in an embedded point cloud: shapes such as spheres and hollow toroids are synergy-dominated, regardless of how the data is rotated. In fMRI data, we find strong correlations between synergistic information and both the number and size of three-dimensional cavities. Furthermore, we find that dimensionality reduction techniques such as PCA preferentially represent higher-order redundancies, and largely fail to preserve both higher-order information and topological structure, suggesting that common manifold-based approaches to studying high-dimensional data are systematically failing to identify important features of the data. These results point towards the possibility of developing a rich theory of higher-order interactions that spans topological and information-theoretic approaches while simultaneously highlighting the profound limitations of more conventional methods. Thomas F. Varley, Pedro A. M. Mediano, Alice Patania, Josh C. Bongard |
PLoS Comput. Biol. | 2 |
| 2024 | Learning diverse causally emergent representations from time series dataabstractCognitive processes usually take place at a macroscopic scale in systems characterised by emergent properties, which make the whole ‘more than the sum of its parts.’ While recent proposals have provided quantitative, information-theoretic metrics to detect emergence in time series data, it is often highly non-trivial to identify the relevant macroscopic variables a priori. In this paper we leverage recent advances in representation learning and differentiable information estimators to put forward a data-driven method to find emergent variables. The proposed method successfully detects emergent variables and recovers the ground-truth emergence values in a synthetic dataset. Furthermore, we show the method can be extended to learn multiple independent features, extracting a diverse set of emergent quantities. We finally show that a modified method scales to real experimental data from primate brain activity, paving the ground for future analyses uncovering the emergent structure of cognitive representations in biological and artificial intelligence systems. David McSharry, Christos Kaplanis, Fernando Rosas, Pedro A. M. Mediano |
NeurIPS | 4 |
| 2024 | Synergistic information supports modality integration and flexible learning in neural networks solving multiple tasksabstractStriking progress has been made in understanding cognition by analyzing how the brain is engaged in different modes of information processing. For instance, so-called synergistic information (information encoded by a set of neurons but not by any subset) plays a key role in areas of the human brain linked with complex cognition. However, two questions remain unanswered: (a) how and why a cognitive system can become highly synergistic; and (b) how informational states map onto artificial neural networks in various learning modes. Here we employ an information-decomposition framework to investigate neural networks performing cognitive tasks. Our results show that synergy increases as networks learn multiple diverse tasks, and that in tasks requiring integration of multiple sources, performance critically relies on synergistic neurons. Overall, our results suggest that synergy is used to combine information from multiple modalities-and more generally for flexible and efficient learning. These findings reveal new ways of investigating how and why learning systems employ specific information-processing strategies, and support the principle that the capacity for general-purpose learning critically relies on the system's information dynamics. Alexandra Maria Proca, Fernando Rosas, Andrea I. Luppi, Daniel Bor, Matthew Crosby, Pedro A. M. Mediano |
PLoS Comput. Biol. | 6 |
| 2023 | A Logarithmic Decomposition for InformationabstractThe Shannon entropy of a random variable X has much behaviour analogous to a signed measure. Previous work has concretized this connection by defining a signed measure µ on an abstract information space $\tilde X$, which is taken to represent the information that X contains. This construction is sufficient to derive many measure-theoretical counterparts to information quantities such as the mutual information $I(X;Y) = \mu (\tilde X \cap \tilde Y)$, the joint entropy $H(X,Y) = \mu (\tilde X \cup \tilde Y)$, and the conditional entropy $H(X|Y) = \mu (\tilde X\backslash \tilde Y)$. We demonstrate that there exists a much finer decomposition with intuitive properties which we call the logarithmic decomposition (LD). We show that this signed measure space has the useful property that its logarithmic atoms are easily characterised with negative or positive entropy, while also being coherent with Yeung’s I-measure [14]. We present the usability of our approach by re-examining the Gács-Körner common information from this new geometric perspective and characterising it in terms of our logarithmic atoms. We then highlight that our geometric refinement can account for an entire class of information quantities, which we call logarithmically decomposable quantities. Keenan J. A. Down, Pedro A. M. Mediano |
ISIT | 2 |
| 2023 | Interaction Measures, Partition Lattices and Kernel Tests for High-Order InteractionsabstractModels that rely solely on pairwise relationships often fail to capture the complete statistical structure of the complex multivariate data found in diverse domains, such as socio-economic, ecological, or biomedical systems. Non-trivial dependencies between groups of more than two variables can play a significant role in the analysis and modelling of such systems, yet extracting such high-order interactions from data remains challenging. Here, we introduce a hierarchy of $d$-order ($d \geq 2$) interaction measures, increasingly inclusive of possible factorisations of the joint probability distribution, and define non-parametric, kernel-based tests to establish systematically the statistical significance of $d$-order interactions. We also establish mathematical links with lattice theory, which elucidate the derivation of the interaction measures and their composite permutation tests; clarify the connection of simplicial complexes with kernel matrix centring; and provide a means to enhance computational efficiency. We illustrate our results numerically with validations on synthetic data, and through an application to neuroimaging data. Zhaolu Liu, Robert L. Peach, Pedro A. M. Mediano, Mauricio Barahona |
NeurIPS | 3 |
| 2022 | High-order functional redundancy in ageing explained via alterations in the connectome in a whole-brain modelabstractThe human brain generates a rich repertoire of spatio-temporal activity patterns, which support a wide variety of motor and cognitive functions. These patterns of activity change with age in a multi-factorial manner. One of these factors is the variations in the brain's connectomics that occurs along the lifespan. However, the precise relationship between high-order functional interactions and connnectomics, as well as their variations with age are largely unknown, in part due to the absence of mechanistic models that can efficiently map brain connnectomics to functional connectivity in aging. To investigate this issue, we have built a neurobiologically-realistic whole-brain computational model using both anatomical and functional MRI data from 161 participants ranging from 10 to 80 years old. We show that the differences in high-order functional interactions between age groups can be largely explained by variations in the connectome. Based on this finding, we propose a simple neurodegeneration model that is representative of normal physiological aging. As such, when applied to connectomes of young participant it reproduces the age-variations that occur in the high-order structure of the functional data. Overall, these results begin to disentangle the mechanisms by which structural changes in the connectome lead to functional differences in the ageing brain. Our model can also serve as a starting point for modeling more complex forms of pathological ageing or cognitive deficits. Marilyn Gatica, Fernando Rosas, Pedro A. M. Mediano, Ibai Díez, Stephan P. Swinnen, Patricio Orio, Rodrigo Cofré, Jesús M. Cortés |
PLoS Comput. Biol. | 3 |
| 2020 | Learning, compression, and leakage: Minimising classification error via meta-universal compression principlesabstractLearning and compression are driven by the common aim of identifying and exploiting statistical regularities in data, which opens the door for fertile collaboration between these areas. A promising group of compression techniques for learning scenarios is normalised maximum likelihood (NML) coding, which provides strong guarantees for compression of small datasets — in contrast with more popular estimators whose guarantees hold only in the asymptotic limit. Here we consider a NMLbased decision strategy for supervised classification problems, and show that it attains heuristic PAC learning when applied to a wide variety of models. Furthermore, we show that the misclassification rate of our method is upper bounded by the maximal leakage, a recently proposed metric to quantify the potential of data leakage in privacy-sensitive scenarios. Fernando Rosas, Pedro A. M. Mediano, Michael Gastpar |
ITW | 2 |
| 2020 | Deep active inference agents using Monte-Carlo methodsabstractActive inference is a Bayesian framework for understanding biological intelligence. The underlying theory brings together perception and action under one single imperative: minimizing free energy. However, despite its theoretical utility in explaining intelligence, computational implementations have been restricted to low-dimensional and idealized situations. In this paper, we present a neural architecture for building deep active inference agents operating in complex, continuous state-spaces using multiple forms of Monte-Carlo (MC) sampling. For this, we introduce a number of techniques, novel to active inference. These include: i) selecting free-energy-optimal policies via MC tree search, ii) approximating this optimal policy distribution via a feed-forward `habitual' network, iii) predicting future parameter belief updates using MC dropouts and, finally, iv) optimizing state transition precision (a high-end form of attention). Our approach enables agents to learn environmental dynamics efficiently, while maintaining task performance, in relation to reward-based counterparts. We illustrate this in a new toy environment, based on the dSprites data-set, and demonstrate that active inference agents automatically create disentangled representations that are apt for modeling state transitions. In a more complex Animal-AI environment, our agents (using the same neural architecture) are able to simulate future state transitions and actions (i.e., plan), to evince reward-directed navigation - despite temporary suspension of visual input. These results show that deep active inference - equipped with MC methods - provides a flexible framework to develop biologically-inspired intelligent agents, with applications in both machine learning and cognitive science. Zafeirios Fountas, Noor Sajid, Pedro A. M. Mediano, Karl J. Friston |
NeurIPS | 3 |
| 2020 | Reconciling emergences: An information-theoretic approach to identify causal emergence in multivariate dataabstractThe broad concept of emergence is instrumental in various of the most challenging open scientific questions-yet, few quantitative theories of what constitutes emergent phenomena have been proposed. This article introduces a formal theory of causal emergence in multivariate systems, which studies the relationship between the dynamics of parts of a system and macroscopic features of interest. Our theory provides a quantitative definition of downward causation, and introduces a complementary modality of emergent behaviour-which we refer to as causal decoupling. Moreover, the theory allows practical criteria that can be efficiently calculated in large systems, making our framework applicable in a range of scenarios of practical interest. We illustrate our findings in a number of case studies, including Conway's Game of Life, Reynolds' flocking model, and neural activity as measured by electrocorticography. Fernando Rosas, Pedro A. M. Mediano, Henrik Jeldtoft Jensen, Anil K. Seth, Adam B. Barrett, Robin L. Carhart-Harris, Daniel Bor |
PLoS Comput. Biol. | 2 |
| 2019 | Relational Forward Models for Multi-Agent Learning
Andrea Tacchetti, H. Francis Song, Pedro A. M. Mediano, Vinícius Flores Zambaldi, János Kramár, Neil C. Rabinowitz, Thore Graepel, Matt M. Botvinick, Peter W. Battaglia |
ICLR (Poster) | 3 |
| 2018 | Adaptation Is Not Just Improvement over TimeabstractThe idea that an agent's actions can impact its actual long-term survival is a very appealing one, underlying influential treatments such as Di Paolo's (2005). However, this presents a tension with understanding the agent and environment as possessing specific objective physical microstates. More specifically, we show that such an approach leads to undesirable outcomes, for example, all organisms being maladaptive on average. We suggest that this problematic intuition of improvement over time may stem from Bayesian inference. We illustrate our arguments using a recent model of autopoietic agency in a model protocell, showing the limitations of previous approaches in this model and specific instantiations of Bayesian inference by ignorant observers in certain scenarios. Simon McGregor, Pedro A. M. Mediano |
Artif. Life | 2 |
| 2018 | Measuring Fitness Effects of Agent-Environment InteractionsabstractOne important sense of the term "adaptation" is the process by which an agent changes appropriately in response to new information provided by environmental stimuli. We propose a novel quantitative measure of this phenomenon, which extends a little-known definition of adaptation as "increased robustness to repeated perturbation" proposed by Klyubin (2002). Our proposed definition essentially corresponds to the average value (relative to some fitness function) of state changes that are caused by the environment (in some statistical ensemble of environments). We compute this value by comparing the agent's actual fitness with its fitness in a counterfactual world where the causal links between agent and environment are disrupted. The proposed measure is illustrated in a simple Markov chain model and also using a recent model of autopoietic agency in a simulated protocell. Simon McGregor, Pedro A. M. Mediano |
Artif. Life | 2 |