VLDB 2026 Research / reviewers in the wild / expert
Mattia Cerrato
dblp:249/6779
· DBLP profile ↗
14ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0001-7736-0547ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Science-Gym: a simple testbed for AI-driven scientific discoveryabstractAbstract Automating scientific discovery has been one of the motivating tasks in the development of AI methods. The task of Equation Discovery (also called Symbolic Regression) is to learn a free-form symbolic equation from experimental data. Equation Discovery benchmarks, however, assume the experimental data as given. Recent successes in protein folding and material optimization, powered by advancements, amongst others, in reinforcement learning and deep learning, have renewed the broader community’s interest in applications of AI in science. Nonetheless, these successful applications do not necessarily lead to an improved understanding of the underlying phenomena, just as super-human chess engines do not necessarily lead to improved understanding of chess theory and practice. In this paper, we propose Science-Gym: a new testbed for basic physics understanding. To the best of our knowledge, Science-Gym is the first scientific discovery benchmark that requires agents to autonomously perform data collection, experimental design, and discover the underlying equations of phenomena. Science-Gym is a Python software library with Gym-compatible bindings. It offers seven scientific simulations, which reproduce basic physics and epidemiology principles: the law of the lever, projectile motion, the inclined plane, Lagrangian points in space, brachistochrones, the SIRV model, and the friction force of a droplet. In these environments, agents may be evaluated not only on their ability in e.g. balancing objects on the two beams of a lever, but more importantly on finding equations that describe the overall behavior of the dynamical system at hand. Mattia Cerrato, Lennart Baur, Jannis Brugger, Sajjad Shumaly, Nicholas Schmitt, Edward Finkelstein, Selina Jukic, Lars Münzel, Felix Peter Paul, Pascal Pfannes, Benedikt Rohr, Julius Schellenberg, Stefan Kramer 0001 |
Mach. Learn. | 1 |
| 2026 | Automated Scientific Discovery: From Equation Discovery to Autonomous Discovery SystemsabstractAbstract The paper surveys automated scientific discovery, from equation discovery and symbolic regression to autonomous discovery systems and agents. It discusses the individual approaches from a "big picture" perspective and in context, but also discusses open issues and recent topics like the various roles of deep neural networks in this area, aiding in the discovery of human-interpretable knowledge. Further, we will present closed-loop scientific discovery systems, starting with the pioneering work on the Adam system up to current efforts in fields from material science to astronomy. Finally, we will elaborate on autonomy from a machine learning perspective, but also in analogy to the autonomy levels in autonomous driving. The maximal level, level five, is defined to require no human intervention at all in the production of scientific knowledge. Achieving this is one step towards solving the Nobel Turing Grand Challenge to develop AI Scientists: AI systems capable of making Nobel-quality scientific discoveries highly autonomously at a level comparable, and possibly superior, to the best human scientists by 2050. Stefan Kramer 0001, Mattia Cerrato, Jannis Brugger, Saso Dzeroski, Ross D. King |
Mach. Learn. | 2 |
| 2025 | Exploring the Design Space of Fair Tree Learning Algorithms
Kiara Stempel, Mattia Cerrato, Stefan Kramer 0001 |
DS | 2 |
| 2025 | A Neurosymbolic Approach to Counterfactual FairnessabstractIntegrating fairness into machine learning models has been an important consideration for the last decade. Here, neurosymbolic models offer a valuable opportunity, as they allow the specification of symbolic, logical constraints that are often guaranteed to be satisfied. However, research on neurosymbolic applications to algorithmic fairness is still in an early stage. With our work, we bridge this gap by integrating counterfactual fairness into the neurosymbolic framework of Logic Tensor Networks (LTN). We use LTN to express accuracy and counterfactual fairness constraints in first-order logic and employ them to achieve desirable levels of both performance and fairness at training time. Our approach is agnostic to the underlying causal model and data generation technique; as such, it may be easily integrated into existing pipelines that generate and extract counterfactual examples. We show, through concrete examples on three real-world datasets, that logical reasoning about counterfactual fairness has some important advantages, among which its intrinsic interpretability, and its flexibility in handling subgroup fairness. Compared to three recent methodologies in counterfactual fairness, our experiments show that a neurosymbolic, LTN-based approach attains better levels of counterfactual fairness. Xenia Heilmann, Chiara Manganini, Mattia Cerrato, Vaishak Belle |
NeSy | 3 |
| 2024 | Peer Learning: Learning Complex Policies in Groups from Scratch via Action RecommendationsabstractPeer learning is a novel high-level reinforcement learning framework for agents learning in groups. While standard reinforcement learning trains an individual agent in trial-and-error fashion, all on its own, peer learning addresses a related setting in which a group of agents, i.e., peers, learns to master a task simultaneously together from scratch. Peers are allowed to communicate only about their own states and actions recommended by others: "What would you do in my situation?". Our motivation is to study the learning behavior of these agents. We formalize the teacher selection process in the action advice setting as a multi-armed bandit problem and therefore highlight the need for exploration. Eventually, we analyze the learning behavior of the peers and observe their ability to rank the agents' performance within the study group and understand which agents give reliable advice. Further, we compare peer learning with single agent learning and a state-of-the-art action advice baseline. We show that peer learning is able to outperform single-agent learning and the baseline in several challenging discrete and continuous OpenAI Gym domains. Doing so, we also show that within such a framework complex policies from action recommendations beyond discrete action spaces can evolve. Cedric Derstroff, Mattia Cerrato, Jannis Brugger, Jan Peters 0001, Stefan Kramer 0001 |
AAAI | 2 |
| 2024 | Science-Gym: A Simple Testbed for AI-Driven Scientific Discovery
Mattia Cerrato, Nicholas Schmitt, Lennart Baur, Edward Finkelstein, Selina Jukic, Lars Münzel, Felix Peter Paul, Pascal Pfannes, Benedikt Rohr, Julius Schellenberg, Stefan Kramer 0001 |
DS (1) | 1 |
| 2024 | Differentially Private Sum-Product NetworksabstractDifferentially private ML approaches seek to learn models which may be publicly released while guaranteeing that the input data is kept private. One issue with this construction is that further model releases based on the same training data (e.g. for a new task) incur a further privacy budget cost. Privacy-preserving synthetic data generation is one possible solution to this conundrum. However, models trained on synthetic private data struggle to approach the performance of private, ad-hoc models. In this paper, we present a novel method based on sum-product networks that is able to perform both privacy-preserving classification and privacy-preserving data generation with a single model. To the best of our knowledge, ours is the first approach that provides both discriminative and generative capabilities to differentially private ML. We show that our approach outperforms the state of the art in terms of stability (i.e. number of training runs required for convergence) and utility of the generated data. Xenia Heilmann, Mattia Cerrato, Ernst Althaus |
ICML | 2 |
| 2024 | Partitioned least squaresabstractAbstract Linear least squares is one of the most widely used regression methods in many fields. The simplicity of the model allows this method to be used when data is scarce and allows practitioners to gather some insight into the problem by inspecting the values of the learnt parameters. In this paper we propose a variant of the linear least squares model allowing practitioners to partition the input features into groups of variables that they require to contribute similarly to the final result. We show that the new formulation is not convex and provide two alternative methods to deal with the problem: one non-exact method based on an alternating least squares approach; and one exact method based on a reformulation of the problem. We show the correctness of the exact method and compare the two solutions showing that the exact solution provides better results in a fraction of the time required by the alternating least squares solution (when the number of partitions is small). We also provide a branch and bound algorithm that can be used in place of the exact method when the number of partitions is too large as well as a proof of NP-completeness of the optimization problem. Roberto Esposito, Mattia Cerrato, Marco Locatelli 0001 |
Mach. Learn. | 2 |
| 2023 | Invariant Representations with Stochastically Quantized Neural NetworksabstractRepresentation learning algorithms offer the opportunity to learn invariant representations of the input data with regard to nuisance factors. Many authors have leveraged such strategies to learn fair representations, i.e., vectors where information about sensitive attributes is removed. These methods are attractive as they may be interpreted as minimizing the mutual information between a neural layer's activations and a sensitive attribute. However, the theoretical grounding of such methods relies either on the computation of infinitely accurate adversaries or on minimizing a variational upper bound of a mutual information estimate. In this paper, we propose a methodology for direct computation of the mutual information between neurons in a layer and a sensitive attribute. We employ stochastically-activated binary neural networks, which lets us treat neurons as random variables. Our method is therefore able to minimize an upper bound on the mutual information between the neural representations and a sensitive attribute. We show that this method compares favorably with the state of the art in fair representation learning and that the learned representations display a higher level of invariance compared to full-precision neural networks. Mattia Cerrato, Marius Köppel, Roberto Esposito, Stefan Kramer 0001 |
AAAI | 1 |
| 2023 | Privacy-Preserving Learning of Random Forests Without Revealing the Trees
Lukas-Malte Bammert, Stefan Kramer 0001, Mattia Cerrato, Ernst Althaus |
DS | 3 |
| 2023 | Hierarchical priors for Hyperspherical Prototypical NetworksabstractIn this paper, we explore the usage of hierarchical priors to improve learning in contexts where the number of available examples is extremely low.Specifically, we consider a Prototype Learning setting where deep neural networks are used to embed data in hyperspherical geometries.In this scenario, we propose an innovative way to learn the prototypes by combining class separation and hierarchical information.In addition, we introduce a contrastive loss function capable of balancing the exploitation of prototypes through a prototype pruning mechanism.We compare the proposed method with state-of-the-art approaches on two public datasets.This work has been partially supported by the Spoke 1 "FutureHPC & BigData" of ICSC -Centro Nazionale di Ricerca in High-Performance-Computing, Big Data and Quantum Computing, funded by European Union -NextGenerationEU. Samuele Fonio, Lorenzo Paletto, Mattia Cerrato, Dino Ienco, Roberto Esposito |
ESANN | 3 |
| 2023 | FairSwiRL: fair semi-supervised classification with representation learningabstractAbstract Semi-supervised learning has shown its potential in many real-world applications where only few labeled examples are available. However, when some fairness constraints need to be satisfied, semi-supervised classification models often struggle as they are required to cope with the lack of sufficient information for predicting the target variable while forgetting its relationships with any sensitive and potentially discriminatory attribute. To address this issue, we propose a fair semi-supervised representation learning architecture that leads to fair and accurate classification results even in very challenging scenarios with few labeled (but biased) instances. We show experimentally that our model can be easily adopted in very general settings, as the learned representations may be employed to train any supervised classifier. Moreover, when applied to several synthetic and real-world datasets, our method is competitive with state-of-the-art fair semi-supervised approaches. Mattia Cerrato, Dino Ienco, Ruggero G. Pensa, Roberto Esposito |
Mach. Learn. | 2 |
| 2020 | Fair pairwise learning to rankabstractRanking algorithms based on Neural Networks have been a topic of recent research. Ranking is employed in everyday applications like product recommendations, search results, or even in finding good candidates for hiring. However, Neural Networks are mostly opaque tools, and it is hard to evaluate why a specific candidate, for instance, was not considered. Therefore, for neural-based ranking methods to be trustworthy, it is crucial to guarantee that the outcome is fair and that the decisions are not discriminating people according to sensitive attributes such as gender, sexual orientation, or ethnicity.In this work we present a family of fair pairwise learning to rank approaches based on Neural Networks, which are able to produce balanced outcomes for underprivileged groups and, at the same time, build fair representations of data, i.e. new vectors having no correlation with regard to a sensitive attribute. We compare our approaches to recent work dealing with fair ranking and evaluate them using both relevance and fairness metrics. Our results show that the introduced fair pairwise ranking methods compare favorably to other methods when considering the fairness/relevance trade-off. Mattia Cerrato, Marius Köppel, Alexander Segner, Roberto Esposito, Stefan Kramer 0001 |
DSAA | 1 |
| 2019 | Taxonomic and Whole Object Constraints: A Deep Architecture
Mattia Cerrato, Edoardo Arnaudo, Valentina Gliozzi, Roberto Esposito |
CogSci | 1 |