VLDB 2026 Research / reviewers in the wild / expert
Samuel Kessler
dblp:255/7238
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0007-4940-8575ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 46% Reinforcement learning · 26% Representation and self-supervised learning · 11% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning and data management
data valuation |
0.9 | 1 | 2025 | SAVA: Scalable Learning-Agnostic Data Valuation · ICLR 2025 |
Mathematical optimization
continuous optimization |
0.9 | 1 | 2025 | SAVA: Scalable Learning-Agnostic Data Valuation · ICLR 2025 |
Mathematical optimization
optimal transport |
0.9 | 1 | 2025 | SAVA: Scalable Learning-Agnostic Data Valuation · ICLR 2025 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Fisher Flow Matching for Generative Modeling over Discrete Data · NeurIPS 2024 |
Machine learning › Generative modeling › generative model › discrete generative model
discrete data generation |
0.8 | 1 | 2024 | Fisher Flow Matching for Generative Modeling over Discrete Data · NeurIPS 2024 |
Machine learning › Generative modeling › generative model
discrete generative model |
0.8 | 1 | 2024 | Fisher Flow Matching for Generative Modeling over Discrete Data · NeurIPS 2024 |
Machine learning › Generative modeling
flow matching |
0.8 | 1 | 2024 | Fisher Flow Matching for Generative Modeling over Discrete Data · NeurIPS 2024 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.6 | 1 | 2022 | Same State, Different Task: Continual Reinforcement Learning without Interference · AAAI 2022 |
Machine learning › Reinforcement learning › non-stationary reinforcement learning
continual reinforcement learning |
0.6 | 1 | 2022 | Same State, Different Task: Continual Reinforcement Learning without Interference · AAAI 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
interference |
0.6 | 1 | 2022 | Same State, Different Task: Continual Reinforcement Learning without Interference · AAAI 2022 |
Machine learning › Reinforcement learning
multi-armed bandit |
0.6 | 1 | 2022 | Same State, Different Task: Continual Reinforcement Learning without Interference · AAAI 2022 |
Machine learning › Reinforcement learning
policy selection |
0.6 | 1 | 2022 | Same State, Different Task: Continual Reinforcement Learning without Interference · AAAI 2022 |
Collaborative and social computing › team collaboration
leadership emergence |
0.6 | 1 | 2022 | Follow the Leader: Technical and Inspirational Leadership in Open Source Software · CHI 2022 |
Collaborative and social computing › team collaboration
virtual teams |
0.6 | 1 | 2022 | Follow the Leader: Technical and Inspirational Leadership in Open Source Software · CHI 2022 |
Empirical software engineering
mining software repositories |
0.6 | 1 | 2022 | Follow the Leader: Technical and Inspirational Leadership in Open Source Software · CHI 2022 |
Methods — techniques the papers use, named apart from their topics
stochastic optimization · 1.7optimal transport · 1.7entropic regularization · 1.7large-scale interaction analysis · 1.1riemannian optimal transport · 0.8geodesic flow · 0.8fisher-rao metric · 0.8replay buffer · 0.6multi-armed bandit · 0.6factorized policy · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAVA: Scalable Learning-Agnostic Data ValuationabstractSelecting data for training machine learning models is crucial since large, web-scraped, real datasets contain noisy artifacts that affect the quality and relevance of individual data points. These noisy artifacts will impact model performance. We formulate this problem as a data valuation task, assigning a value to data points in the training set according to how similar or dissimilar they are to a clean and curated validation set. Recently, *LAVA* (Just et al., 2023) demonstrated the use of optimal transport (OT) between a large noisy training dataset and a clean validation set, to value training data efficiently, without the dependency on model performance. However, the *LAVA* algorithm requires the entire dataset as an input, this limits its application to larger datasets. Inspired by the scalability of stochastic (gradient) approaches which carry out computations on *batches* of data points instead of the entire dataset, we analogously propose *SAVA*, a scalable variant of *LAVA* with its computation on batches of data points. Intuitively, *SAVA* follows the same scheme as *LAVA* which leverages the hierarchically defined OT for data valuation. However, while *LAVA* processes the whole dataset, *SAVA* divides the dataset into batches of data points, and carries out the OT problem computation on those batches. Moreover, our theoretical derivations on the trade-off of using entropic regularization for OT problems include refinements of prior work. We perform extensive experiments, to demonstrate that *SAVA* can scale to large datasets with millions of data points and does not trade off data valuation performance. Our Github repository is available at \url{https://github.com/skezle/sava}. Samuel Kessler, Tam Le |
ICLR | 1 |
| 2024 | Fisher Flow Matching for Generative Modeling over Discrete DataabstractGenerative modeling over discrete data has recently seen numerous success stories, with applications spanning language modeling, biological sequence design, and graph-structured molecular data. The predominant generative modeling paradigm for discrete data is still autoregressive, with more recent alternatives based on diffusion or flow-matching falling short of their impressive performance in continuous data settings, such as image or video generation. In this work, we introduce Fisher-Flow, a novel flow-matching model for discrete data. Fisher-Flow takes a manifestly geometric perspective
by considering categorical distributions over discrete data as points residing on a statistical manifold equipped with its natural Riemannian metric: the \emph{Fisher-Rao metric}. As a result, we demonstrate discrete data itself can be continuously reparameterised to points on the positive orthant of the $d$-hypersphere $\mathbb{S}^d_+$,
which allows us to define flows that map any source distribution to target in a principled manner by transporting mass along (closed-form) geodesics of $\mathbb{S}^d_+$. Furthermore, the learned flows in Fisher-Flow can be further bootstrapped by leveraging Riemannian optimal transport leading to improved training dynamics. We prove that the gradient flow induced by Fisher-FLow is optimal in reducing the forward KL divergence. We evaluate Fisher-Flow on an array of synthetic and diverse real-world benchmarks, including designing DNA Promoter, and DNA Enhancer sequences. Empirically, we find that Fisher-Flow improves over prior diffusion and flow-matching models on these benchmarks. Oscar Davis, Samuel Kessler, Mircea Petrache, Ismail Ilkan Ceylan, Michael M. Bronstein, Joey Bose |
NeurIPS | 2 |
| 2022 | Same State, Different Task: Continual Reinforcement Learning without InterferenceabstractContinual Learning (CL) considers the problem of training an agent sequentially on a set of tasks while seeking to retain performance on all previous tasks. A key challenge in CL is catastrophic forgetting, which arises when performance on a previously mastered task is reduced when learning a new task. While a variety of methods exist to combat forgetting, in some cases tasks are fundamentally incompatible with each other and thus cannot be learnt by a single policy. This can occur, in reinforcement learning (RL) when an agent may be rewarded for achieving different goals from the same observation. In this paper we formalize this "interference" as distinct from the problem of forgetting. We show that existing CL methods based on single neural network predictors with shared replay buffers fail in the presence of interference. Instead, we propose a simple method, OWL, to address this challenge. OWL learns a factorized policy, using shared feature extraction layers, but separate heads, each specializing on a new task. The separate heads in OWL are used to prevent interference. At test time, we formulate policy selection as a multi-armed bandit problem, and show it is possible to select the best policy for an unknown task using feedback from the environment. The use of bandit algorithms allows the OWL agent to constructively re-use different continually learnt policies at different times during an episode. We show in multiple RL environments that existing replay based CL methods fail, while OWL is able to achieve close to optimal performance when training sequentially. Samuel Kessler, Jack Parker-Holder, Philip J. Ball, Stefan Zohren, Stephen J. Roberts |
AAAI | 1 |
| 2022 | Follow the Leader: Technical and Inspirational Leadership in Open Source SoftwareabstractWe conduct the first comprehensive study of the behavioral factors which predict leader emergence within open source software (OSS) virtual teams. We leverage the full history of developers’ interactions with their teammates and projects at github.com between January 2010 and April 2017 (representing about 133 million interactions) to establish that – contrary to a common narrative describing open source as a pure “technical meritocracy” – developers’ communication abilities and community building skills are significant predictors of whether they emerge as team leaders. Inspirational communication therefore appears as central to the process of leader emergence in virtual teams, even in a setting like OSS, where technical contributions have often been conceptualized as the sole pathway to gaining community recognition. Those results should be of interest to researchers and practitioners theorizing about OSS in particular and, more generally, leadership in geographically dispersed virtual teams, as well as to online community managers. Jérôme Hergueux, Samuel Kessler |
CHI | 2 |
| 2022 | An Adapter Based Pre-Training for Efficient and Scalable Self-Supervised Speech Representation LearningabstractWe present a method for transferring pre-trained self-supervised (SSL) speech representations to multiple languages. There is an abundance of unannotated speech, so creating self-supervised representations from raw audio and fine-tuning on small annotated datasets is a promising direction to build speech recognition systems. SSL models generally perform SSL on raw audio in a pre-training phase and then fine-tune on a small fraction of annotated data. Such models have produced state of the art results for ASR. However, these models are very expensive to pre-train. We use an existing wav2vec 2.0 model and tackle the problem of learning new language representations while utilizing existing model knowledge. Crucially we do so without catastrophic forgetting of the existing language representation. We use adapter modules to speed up pre-training a new language task. Our model can decrease pre-training times by 32% when learning a new language task, and learn this new audio-language representation without forgetting previous language representation. We evaluate by applying these language representations to automatic speech recognition. Samuel Kessler, Bethan Thomas, Salah Karout |
ICASSP | 1 |
| 2022 | Efficient Adapter Transfer of Self-Supervised Speech Models for Automatic Speech RecognitionabstractSelf-supervised learning (SSL) is a powerful tool that allows learning of underlying representations from unlabeled data. Transformer based models such as wav2vec 2.0 and HuBERT are leading the field in the speech domain. Generally these models are fine-tuned on a small amount of labeled data for a downstream task such as Automatic Speech Recognition (ASR). This involves re-training the majority of the model for each task. Adapters are small lightweight modules which are commonly used in Natural Language Processing (NLP) to adapt pre-trained models to new tasks. In this paper we propose applying adapters to wav2vec 2.0 to reduce the number of parameters required for downstream ASR tasks, and increase scalability of the model to multiple tasks or languages. Using adapters we can perform ASR while training fewer than 10% of parameters per task compared to full fine-tuning with little degradation of performance. Ablations show that applying adapters into just the top few layers of the pre-trained network gives similar performance to full transfer, supporting the theory that higher pre-trained layers encode more phonemic information, and further optimizing efficiency. Bethan Thomas, Samuel Kessler, Salah Karout |
ICASSP | 2 |
| 2021 | Hierarchical Indian buffet neural networks for Bayesian continual learningabstractWe place an Indian Buffet process (IBP) prior over the structure of a Bayesian Neural Network (BNN), thus allowing the complexity of the BNN to increase and decrease automatically. We further extend this model such that the prior on the structure of each hidden layer is shared globally across all layers, using a Hierarchical-IBP (H-IBP). We apply this model to the problem of resource allocation in Continual Learning (CL) where new tasks occur and the network requires extra resources. Our model uses online variational inference with reparameterisation of the Bernoulli and Beta distributions, which constitute the IBP and H-IBP priors. As we automatically learn the number of weights in each layer of the BNN, overfitting and underfitting problems are largely overcome. We show empirically that our approach offers a competitive edge over existing methods in CL. Samuel Kessler, Stefan Zohren, Stephen J. Roberts |
UAI | 1 |