EDBT 2026 Demo / reviewers in the wild / expert
Samuel Robertson
dblp:265/4898
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 58% Learning theory · 26% Motion planning and robot control · 10% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 75% Information retrieval · 25% |
Topics — the 17 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Motion planning and robot control › stochastic optimal control
hamilton-jacobi-bellman equation |
1.0 | 1 | 2026 | Continuous time policy evaluation is easier with noisy dynamics · COLT 2026 |
Machine learning › Reinforcement learning › value function approximation
kernel-based value function approximation |
1.0 | 1 | 2026 | Continuous time policy evaluation is easier with noisy dynamics · COLT 2026 |
Machine learning › Reinforcement learning
policy evaluation |
1.0 | 1 | 2026 | Continuous time policy evaluation is easier with noisy dynamics · COLT 2026 |
Machine learning › Reinforcement learning
stochastic control |
1.0 | 1 | 2026 | Continuous time policy evaluation is easier with noisy dynamics · COLT 2026 |
Machine learning › Learning theory › online learning
eluder dimension |
0.9 | 1 | 2025 | Eluder dimension: localise it! · NeurIPS 2025 |
Machine learning › Reinforcement learning › regret minimization
first-order regret bounds |
0.9 | 1 | 2025 | Eluder dimension: localise it! · NeurIPS 2025 |
Machine learning › Learning theory › online learning
regret bounds |
0.9 | 1 | 2025 | Eluder dimension: localise it! · NeurIPS 2025 |
Recommender systems
collaborative filtering |
0.9 | 1 | 2025 | Does Weighting Improve Matrix Factorization for Recommender Systems? · WWW 2025 |
Recommender systems › collaborative filtering
implicit feedback |
0.9 | 1 | 2025 | Does Weighting Improve Matrix Factorization for Recommender Systems? · WWW 2025 |
Recommender systems › collaborative filtering
matrix factorization |
0.9 | 1 | 2025 | Does Weighting Improve Matrix Factorization for Recommender Systems? · WWW 2025 |
Information retrieval
weighting schemes |
0.9 | 1 | 2025 | Does Weighting Improve Matrix Factorization for Recommender Systems? · WWW 2025 |
Machine learning › Reinforcement learning › dynamic programming › value iteration
fitted q-iteration |
0.8 | 1 | 2024 | Switching the Loss Reduces the Cost in Batch Reinforcement Learning · ICML 2024 |
Machine learning › Learning theory › online learning
log-loss |
0.8 | 1 | 2024 | Switching the Loss Reduces the Cost in Batch Reinforcement Learning · ICML 2024 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.8 | 1 | 2024 | Switching the Loss Reduces the Cost in Batch Reinforcement Learning · ICML 2024 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel ridge regression |
0.3 | 1 | 2026 | Continuous time policy evaluation is easier with noisy dynamics · COLT 2026 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
reproducing kernel hilbert space |
0.3 | 1 | 2026 | Continuous time policy evaluation is easier with noisy dynamics · COLT 2026 |
Machine learning › Reinforcement learning
bandit |
0.3 | 1 | 2025 | Eluder dimension: localise it! · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
kernel ridge regression · 1.0elliptic partial differential equation theory · 1.0Matérn RKHS · 1.0regularization · 0.9optimization · 0.9matrix factorization · 0.9generalized linear model · 0.9eluder dimension localization · 0.9squared loss · 0.8log loss · 0.8fitted q-iteration · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Continuous time policy evaluation is easier with noisy dynamicsabstractIn this work, we study continuous-time stochastic control problems governed by controlled stochastic differential equations with unknown dynamics. We focus on the discounted infinite-horizon setting and restrict attention to feedback controllers. In general, the continuous time value function is the solution to the nonlinear Hamilton-Jacobi-Bellman (HJB) equation, which typical only admits viscosity solutions with no regularity. Our first contribution is to establish sharp regularity results for value functions using elliptic partial differential equation theory. Under mild growth and regularity assumptions on the controlled dynamics and a uniform ellipticity condition on the diffusion, we show that the value function belongs to a Matérn reproducing kernel Hilbert space (RKHS) that is strictly smoother than the running reward. Building on this analysis, we develop a kernel-based policy evaluation method that estimates value functions directly from online trajectory rollouts of a fixed policy. The resulting algorithm exploits the RKHS structure with a kernel ridge regression technique, reducing the infinite-dimensional learning problem to a finite-dimensional one. Our results establish a direct connection between stochastic control, elliptic regularity theory, and kernel methods, and provide a foundation for online policy evaluation and policy improvement in continuous time. Samuel Robertson, Thomas Newton, Csaba Szepesvári |
COLT | 1 |
| 2025 | Eluder dimension: localise it!abstractWe establish a lower bound on the eluder dimension in generalised linear model classes, showing that standard eluder dimension-based analysis cannot lead to first-order regret bounds. To address this, we introduce a localisation method for the eluder dimension; our analysis immediately recovers and improves on classic results for Bernoulli bandits, and allows for the first genuine first-order bounds for finite-horizon reinforcement learning tasks with bounded cumulative returns. Alireza Bakhtiari, Alex Ayoub, Samuel Robertson, David Janz, Csaba Szepesvári |
NeurIPS | 3 |
| 2025 | REINFORCE Converges to Optimal Policies with Any Learning RateabstractWe prove that the classic REINFORCE stochastic policy gradient (SPG) method converges to globally optimal policies in finite-horizon Markov Decision Processes (MDPs) with $\textit{any}$ constant learning rate. To avoid the need for small or decaying learning rates, we introduce two key innovations in the stochastic bandit setting, which we then extend to MDPs. $\textbf{First}$, we identify a new exploration property of SPG: the online SPG method samples every action infinitely often (i.o.), improving on previous results that only guaranteed at least two actions would be sampled i.o. This means SPG inherently achieves asymptotic exploration without modification. $\textbf{Second}$, we eliminate the assumption of unique mean reward values, a condition that previous convergence analyses in the bandit setting relied on, but that does not translate to MDPs. Our results deepen the theoretical understanding of SPG in both bandit problems and MDPs, with a focus on how it handles the exploration-exploitation trade-off when standard optimization and stochastic approximation methods cannot be applied, as is the case with large constant learning rates. Samuel Robertson, Thang Chu, Bo Dai 0001, Dale Schuurmans, Csaba Szepesvári, Jincheng Mei |
NeurIPS | 1 |
| 2025 | Does Weighting Improve Matrix Factorization for Recommender Systems?abstractMatrix factorization is a widely used approach for top-N recommendation and collaborative filtering. When implemented on implicit feedback data (such as clicks), a common heuristic is to upweight the observed interactions. This strategy has been shown to improve performance for certain algorithms. In this paper, we conduct a systematic study of various weighting schemes and matrix factorization algorithms. Somewhat surprisingly, we find that training with unweighted data can perform comparably to-and sometimes outperform-training with weighted data, especially for large models. This observation challenges the conventional wisdom. Nevertheless, we identify cases where weighting can be beneficial, particularly for models with lower capacity and specific regularization schemes. We also derive efficient algorithms for exactly minimizing several weighted objectives that were previously considered computationally intractable. Our work provides a comprehensive analysis of the interplay between weighting, regularization, and model capacity in matrix factorization for recommender systems. Alex Ayoub, Samuel Robertson, Dawen Liang, Harald Steck, Nathan Kallus |
WWW | 2 |
| 2025 | Structural network measures reveal the emergence of heavy-tailed degree distributions in lottery ticket multilayer perceptronsabstractArtificial neural networks (ANNs) were originally modeled after their biological counterparts, but have since conceptually diverged in many ways. The resulting network architectures are not well understood, and furthermore, we lack the quantitative tools to characterize their structures. Network science provides an ideal mathematical framework with which to characterize systems of interacting components, and has transformed our understanding across many domains, including the mammalian brain. Yet, little has been done to bring network science to ANNs. In this work, we propose tools that leverage and adapt network science methods to measure both global- and local-level characteristics of ANNs. Specifically, we focus on the structures of efficient multilayer perceptrons as a case study, which are sparse and systematically pruned such that they share many characteristics with real-world networks. We use adapted network science metrics to show that the pruning process leads to the emergence of a spanning subnetwork (lottery ticket multilayer perceptrons) with complex architecture. This complex network exhibits global and local characteristics, including heavy-tailed nodal degree distributions and dominant weighted pathways, that mirror patterns observed in human neuronal connectivity. Furthermore, alterations in network metrics precede catastrophic decay in performance as the network is heavily pruned. This network science-driven approach to the analysis of artificial neural networks serves as a valuable tool to establish and improve biological fidelity, increase the interpretability, and assess the performance of artificial neural networks. Significance Statement Artificial neural network architectures have become increasingly complex, often diverging from their biological counterparts in many ways. To design plausible "brain-like" architectures, whether to advance neuroscience research or to improve explainability, it is essential that these networks optimally resemble their biological counterparts. Network science tools offer valuable information about interconnected systems, including the brain, but have not attracted much attention for analyzing artificial neural networks. Here, we present the significance of our work: •We adapt network science tools to analyze the structural characteristics of artificial neural networks. •We demonstrate that organizational patterns similar to those observed in the mammalian brain emerge through the pruning process alone. The convergence on these complex network features in both artificial neural networks and biological brain networks is compelling evidence for their optimality in information processing capabilities. •Our approach is a significant first step towards a network science-based understanding of artificial neural networks, and has the potential to shed light on the biological fidelity of artificial neural networks. Chris Kang, Jasmine A. Moore, Samuel Robertson, Matthias Wilms, Emma K. Towlson, Nils Daniel Forkert |
Neural Networks | 3 |
| 2024 | Switching the Loss Reduces the Cost in Batch Reinforcement LearningabstractWe propose training fitted Q-iteration with log-loss (FQI-LOG) for batch reinforcement learning (RL). We show that the number of samples needed to learn a near-optimal policy with FQI-LOG scales with the accumulated cost of the optimal policy, which is zero in problems where acting optimally achieves the goal and incurs no cost. In doing so, we provide a general framework for proving small-cost bounds, i.e. bounds that scale with the optimal achievable cost, in batch RL. Moreover, we empirically verify that FQI-LOG uses fewer samples than FQI trained with squared loss on problems where the optimal policy reliably achieves the goal. Alex Ayoub, Samuel Robertson, James McInerney, Dawen Liang, Nathan Kallus, Csaba Szepesvári |
ICML | 4 |