EDBT 2026 Demo / reviewers in the wild / expert
Thomas S. Stepleton
dblp:82/6271 · also Tom Stepleton
· DBLP profile ↗
10ranked-venue papers
1as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 79% Trustworthy machine learning · 14% Segmentation and scene understanding · 6% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment |
0.5 | 1 | 2021 | Counterfactual Credit Assignment in Model-Free Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning › value function estimation
future-dependent value function |
0.5 | 1 | 2021 | Counterfactual Credit Assignment in Model-Free Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.5 | 1 | 2021 | Counterfactual Credit Assignment in Model-Free Reinforcement Learning · ICML 2021 |
Machine learning › Trustworthy machine learning
fairness |
0.4 | 1 | 2020 | A General Approach to Fairness with Optimal Transport · AAAI 2020 |
Mathematical optimization
optimal transport |
0.4 | 1 | 2020 | A General Approach to Fairness with Optimal Transport · AAAI 2020 |
Machine learning › Reinforcement learning
off-policy evaluation |
0.2 | 1 | 2016 | Safe and Efficient Off-Policy Reinforcement Learning · NIPS 2016 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
0.2 | 1 | 2016 | Safe and Efficient Off-Policy Reinforcement Learning · NIPS 2016 |
Machine learning › Reinforcement learning
policy evaluation |
0.2 | 1 | 2016 | Safe and Efficient Off-Policy Reinforcement Learning · NIPS 2016 |
Machine learning › Reinforcement learning
model-free reinforcement learning |
0.1 | 1 | 2021 | Counterfactual Credit Assignment in Model-Free Reinforcement Learning · ICML 2021 |
Computer vision › Segmentation and scene understanding
boundary detection |
0.1 | 1 | 2008 | Towards unsupervised whole-object segmentation: Combining automated matting with boundary detection · CVPR 2008 |
Computer vision › Segmentation and scene understanding › object segmentation
unsupervised object segmentation |
0.1 | 1 | 2008 | Towards unsupervised whole-object segmentation: Combining automated matting with boundary detection · CVPR 2008 |
Machine learning › Graph learning › graph clustering
spectral clustering |
0.0 | 1 | 2008 | Towards unsupervised whole-object segmentation: Combining automated matting with boundary detection · CVPR 2008 |
Methods — techniques the papers use, named apart from their topics
optimal transport · 0.9distribution transport · 0.9hindsight information constraints · 0.5future-conditional value functions · 0.5counterfactual reasoning · 0.5retrace(lambda) · 0.2q(lambda) · 0.2spectral clustering · 0.1image matting · 0.1boundary detection · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Gaps in the Safety Evaluation of Generative AIabstractGenerative AI systems produce a range of ethical and social risks. Evaluation of these risks is a critical step on the path to ensuring the safety of these systems. However, evaluation requires the availability of validated and established measurement approaches and tools. In this paper, we provide an empirical review of the methods and tools that are available for evaluating known safety of generative AI systems to date. To this end, we review more than 200 safety-related evaluations that have been applied to generative AI systems. We categorise each evaluation along multiple axes to create a detailed snapshot of the safety evaluation landscape to date. We release this data for researchers and AI safety practitioners (https://bitly.ws/3hUzu). Analysing the current safety evaluation landscape reveals three systemic ”evaluation gaps”. First, a ”modality gap” emerges as few safety evaluations exist for non-text modalities. Second, a ”risk coverage gap” arises as evaluations for several ethical and social risks are simply lacking. Third, a ”context gap” arises as most safety evaluations are model-centric and fail to take into account the broader context in which AI systems operate. Devising next steps for safety practitioners based on these findings, we present tactical ”low-hanging fruit” steps towards closing the identified evaluation gaps and their limitations. We close by discussing the role and limitations of safety evaluation to ensure the safety of generative AI systems. Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Ramona Comanescu, Canfer Akbulut, Thomas S. Stepleton, Juan Mateos-Garcia, A. Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, William Isaac 0001, Laura Weidinger |
AIES (1) | 7 |
| 2021 | Counterfactual Credit Assignment in Model-Free Reinforcement LearningabstractCredit assignment in reinforcement learning is the problem of measuring an action’s influence on future rewards. In particular, this requires separating skill from luck, i.e. disentangling the effect of an action on rewards from that of external factors and subsequent actions. To achieve this, we adapt the notion of counterfactuals from causality theory to a model-free RL setup. The key idea is to condition value functions on future events, by learning to extract relevant information from a trajectory. We formulate a family of policy gradient algorithms that use these future-conditional value functions as baselines or critics, and show that they are provably low variance. To avoid the potential bias from conditioning on future information, we constrain the hindsight information to not contain information about the agent’s actions. We demonstrate the efficacy and validity of our algorithm on a number of illustrative and challenging problems. Thomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor, Alaa Saade, Anna Harutyunyan, Will Dabney, Thomas S. Stepleton, Nicolas Heess, Arthur Guez, Eric Moulines, Marcus Hutter, Lars Buesing, Rémi Munos |
ICML | 8 |
| 2020 | A General Approach to Fairness with Optimal TransportabstractWe propose a general approach to fairness based on transporting distributions corresponding to different sensitive attributes to a common distribution. We use optimal transport theory to derive target distributions and methods that allow us to achieve fairness with minimal changes to the unfair model. Our approach is applicable to both classification and regression problems, can enforce different notions of fairness, and enable us to achieve a Pareto-optimal trade-off between accuracy and fairness. We demonstrate that it outperforms previous approaches in several benchmark fairness datasets. Silvia Chiappa, Ray Jiang, Thomas S. Stepleton, Aldo Pacchiano, Heinrich Jiang, John Aslanides |
AAAI | 3 |
| 2019 | Wasserstein Fair Classification
Ray Jiang, Aldo Pacchiano, Thomas S. Stepleton, Heinrich Jiang, Silvia Chiappa |
UAI | 3 |
| 2016 | Q(λ) with Off-Policy Corrections
Anna Harutyunyan, Marc G. Bellemare, Thomas S. Stepleton, Rémi Munos |
ALT | 3 |
| 2016 | Safe and Efficient Off-Policy Reinforcement LearningabstractIn this work, we take a fresh look at some old and new algorithms for off-policy, return-based reinforcement learning. Expressing these in a common form, we derive a novel algorithm, Retrace(lambda), with three desired properties: (1) it has low variance; (2) it safely uses samples collected from any behaviour policy, whatever its degree of "off-policyness"; and (3) it is efficient as it makes the best use of samples collected from near on-policy behaviour policies. We analyse the contractive nature of the related operator under both off-policy policy evaluation and control settings and derive online sample-based algorithms. We believe this is the first return-based off-policy control algorithm converging a.s. to Q* without the GLIE assumption (Greedy in the Limit with Infinite Exploration). As a corollary, we prove the convergence of Watkins' Q(lambda), which was an open problem since 1989. We illustrate the benefits of Retrace(lambda) on a standard suite of Atari 2600 games. Rémi Munos, Thomas S. Stepleton, Anna Harutyunyan, Marc G. Bellemare |
NIPS | 2 |
| 2008 | Towards unsupervised whole-object segmentation: Combining automated matting with boundary detectionabstractWe propose a novel step toward the unsupervised segmentation of whole objects by combining ldquohintsrdquo of partial scene segmentation offered by multiple soft, binary mattes. These mattes are implied by a set of hypothesized object boundary fragments in the scene. Rather than trying to find or define a single ldquobestrdquo segmentation, we generate multiple segmentations of an image. This reflects contemporary methods for unsupervised object discovery from groups of images, and it allows us to define intuitive evaluation metrics for our sets of segmentations based on the accurate and parsimonious delineation of scene objects. Our proposed approach builds on recent advances in spectral clustering, image matting, and boundary detection. It is demonstrated qualitatively and quantitatively on a dataset of scenes and is suitable for current work in unsupervised object discovery without top-down knowledge. Andrew N. Stein, Thomas S. Stepleton, Martial Hebert |
CVPR | 2 |
| 2007 | Iterative design of a Braille writing tutor to combat illiteracyabstractLess than 3% of the 145 million blind people living in developing countries are literate. This low literacy rate is partly due to the lack of trained teachers and the challenges associated with learning to write Braille on a traditional slate and stylus. These challenges include writing from right to left, writing mirrored images of letters, and receiving significantly delayed feedback. Extensive conversations with the Mathru School for the Blind near Bangalore, India, revealed the need for a robust, low-power, low-cost Braille writing tutor. We present an iterative and participatory design process resulting in the creation and refinement of a prototype Braille writing tutor system. This system uses a novel input device to capture a student's activity on a slate using a stylus and uses a range of techniques to teach Braille writing skills to both beginner and advanced students. We report on lessons learned from the implementation of this project and from a six-week pilot study at the Mathru school, and outline future directions for improvement. Nidhi Kalra, Tom Lauwers, Daniel Dewey, Thomas S. Stepleton, M. Bernardine Dias |
ICTD | 4 |
| 2003 | A real-time vision module for interactive perceptual agents
Bruce A. Maxwell, Nathaniel Fairfield, Nikolas Johnson, Pukar Malla, Paul E. Dickson, Suor Kim, Stephanie Wojtkowski, Thomas S. Stepleton |
Mach. Vis. Appl. | 8 |
| 2002 | Work-Augmented Laziness with the Los Task Request System
Thomas S. Stepleton |
LISA | 1 |