Federico Adolfi

dblp:271/8491 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-4202-5252ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 91% Probabilistic and Bayesian machine learning · 9%
Theoretical computer science
1 paper
Computational complexity · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.622025
The Computational Complexity of Circuit Discovery for Inner Interpretability · ICLR 2025
Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience · ICML 2024
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
1.622025
The Computational Complexity of Circuit Discovery for Inner Interpretability · ICLR 2025
Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience · ICML 2024
Computational complexity
parameterized complexity
0.912025
The Computational Complexity of Circuit Discovery for Inner Interpretability · ICLR 2025
Machine learning › Trustworthy machine learning › interpretability › explainable AI
self-interpretable models
0.812024
Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.412020
Gibbs Sampling with People · NeurIPS 2020
Computational social science and digital humanities
cognitive science
0.412020
Gibbs Sampling with People · NeurIPS 2020
Machine learning › Trustworthy machine learning › language model interpretability
mechanistic understanding
0.212024
Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience · ICML 2024

Methods — techniques the papers use, named apart from their topics

parameterized complexity theory · 1.7markov chain monte carlo with people · 0.9gibbs sampling · 0.9cognitive neuroscience methodology · 0.8
YearPublicationVenuePosition
2025 Content-agnostic online segmentation as a core operation
Federico Adolfi, David Poeppel
CogSci1
2025 The Computational Complexity of Circuit Discovery for Inner Interpretability
abstract
Many proposed applications of neural networks in machine learning, cognitive/brain science, and society hinge on the feasibility of inner interpretability via circuit discovery. This calls for empirical and theoretical explorations of viable algorithmic options. Despite advances in the design and testing of heuristics, there are concerns about their scalability and faithfulness at a time when we lack understanding of the complexity properties of the problems they are deployed to solve. To address this, we study circuit discovery with classical and parameterized computational complexity theory: (1) we describe a conceptual scaffolding to reason about circuit finding queries in terms of affordances for description, explanation, prediction and control; (2) we formalize a comprehensive set of queries for mechanistic explanation, and propose a formal framework for their analysis; (3) we use it to settle the complexity of many query variants and relaxations of practical interest on multi-layer perceptrons. Our findings reveal a challenging complexity landscape. Many queries are intractable, remain fixed-parameter intractable relative to model/circuit features, and inapproximable under additive, multiplicative, and probabilistic approximation schemes. To navigate this landscape, we prove there exist transformations to tackle some of these hard problems with better-understood heuristics, and prove the tractability or fixed-parameter tractability of more modest queries which retain useful affordances. This framework allows us to understand the scope and limits of interpretability queries, explore viable options, and compare their resource demands on existing and future architectures.
Federico Adolfi, Martina G. Vilas, Todd Wareham
ICLR1
2024 Complexity-Theoretic Limits on the Promises of Artificial Neural Network Reverse-Engineering
Federico Adolfi, Martina G. Vilas, Todd Wareham
CogSci1
2024 Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience
abstract
Inner Interpretability is a promising emerging field tasked with uncovering the inner mechanisms of AI systems, though how to develop these mechanistic theories is still much debated. Moreover, recent critiques raise issues that question its usefulness to advance the broader goals of AI. However, it has been overlooked that these issues resemble those that have been grappled with in another field: Cognitive Neuroscience. Here we draw the relevant connections and highlight lessons that can be transferred productively between fields. Based on these, we propose a general conceptual framework and give concrete methodological strategies for building mechanistic explanations in AI inner interpretability research. With this conceptual framework, Inner Interpretability can fend off critiques and position itself on a productive path to explain AI systems.
Martina G. Vilas, Federico Adolfi, David Poeppel, Gemma Roig
ICML2
2023 Successes and critical failures of neural networks in capturing human-like speech recognition
abstract
Natural and artificial audition can in principle acquire different solutions to a given problem. The constraints of the task, however, can nudge the cognitive science and engineering of audition to qualitatively converge, suggesting that a closer mutual examination would potentially enrich artificial hearing systems and process models of the mind and brain. Speech recognition - an area ripe for such exploration - is inherently robust in humans to a number transformations at various spectrotemporal granularities. To what extent are these robustness profiles accounted for by high-performing neural network systems? We bring together experiments in speech recognition under a single synthesis framework to evaluate state-of-the-art neural networks as stimulus-computable, optimized observers. In a series of experiments, we (1) clarify how influential speech manipulations in the literature relate to each other and to natural speech, (2) show the granularities at which machines exhibit out-of-distribution robustness, reproducing classical perceptual phenomena in humans, (3) identify the specific conditions where model predictions of human performance differ, and (4) demonstrate a crucial failure of all artificial systems to perceptually recover where humans do, suggesting alternative directions for theory and model building. These findings encourage a tighter synergy between the cognitive science and engineering of audition.
Federico Adolfi, Jeffrey S. Bowers, David Poeppel
Neural Networks1
2022 Computational Complexity of Segmentation
Federico Adolfi, Todd Wareham, Iris van Rooij
CogSci1
2020 Gibbs Sampling with People
abstract
A core problem in cognitive science and machine learning is to understand how humans derive semantic representations from perceptual objects, such as color from an apple, pleasantness from a musical chord, or seriousness from a face. Markov Chain Monte Carlo with People (MCMCP) is a prominent method for studying such representations, in which participants are presented with binary choice trials constructed such that the decisions follow a Markov Chain Monte Carlo acceptance rule. However, while MCMCP has strong asymptotic properties, its binary choice paradigm generates relatively little information per trial, and its local proposal function makes it slow to explore the parameter space and find the modes of the distribution. Here we therefore generalize MCMCP to a continuous-sampling paradigm, where in each iteration the participant uses a slider to continuously manipulate a single stimulus dimension to optimize a given criterion such as ‘pleasantness’. We formulate both methods from a utility-theory perspective, and show that the new method can be interpreted as ‘Gibbs Sampling with People’ (GSP). Further, we introduce an aggregation parameter to the transition step, and show that this parameter can be manipulated to flexibly shift between Gibbs sampling and deterministic optimization. In an initial study, we show GSP clearly outperforming MCMCP; we then show that GSP provides novel and interpretable results in three other domains, namely musical chords, vocal emotions, and faces. We validate these results through large-scale perceptual rating experiments. The final experiments use GSP to navigate the latent space of a state-of-the-art image synthesis network (StyleGAN), a promising approach for applying GSP to high-dimensional perceptual spaces. We conclude by discussing future cognitive applications and ethical implications.
Peter M. C. Harrison, Raja Marjieh, Federico Adolfi, Pol van Rijn, Manuel Anglada-Tort, Ofer Tchernichovski, Pauline Larrouy-Maestri, Nori Jacoby
NeurIPS3