Thomas S. Stepleton

dblp:82/6271 · also Tom Stepleton · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 79% Trustworthy machine learning · 14% Segmentation and scene understanding · 6%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment
0.512021
Counterfactual Credit Assignment in Model-Free Reinforcement Learning · ICML 2021
Machine learning › Reinforcement learning › value function estimation
future-dependent value function
0.512021
Counterfactual Credit Assignment in Model-Free Reinforcement Learning · ICML 2021
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.512021
Counterfactual Credit Assignment in Model-Free Reinforcement Learning · ICML 2021
Machine learning › Trustworthy machine learning
fairness
0.412020
A General Approach to Fairness with Optimal Transport · AAAI 2020
Mathematical optimization
optimal transport
0.412020
A General Approach to Fairness with Optimal Transport · AAAI 2020
Machine learning › Reinforcement learning
off-policy evaluation
0.212016
Safe and Efficient Off-Policy Reinforcement Learning · NIPS 2016
Machine learning › Reinforcement learning
off-policy reinforcement learning
0.212016
Safe and Efficient Off-Policy Reinforcement Learning · NIPS 2016
Machine learning › Reinforcement learning
policy evaluation
0.212016
Safe and Efficient Off-Policy Reinforcement Learning · NIPS 2016
Machine learning › Reinforcement learning
model-free reinforcement learning
0.112021
Counterfactual Credit Assignment in Model-Free Reinforcement Learning · ICML 2021
Computer vision › Segmentation and scene understanding
boundary detection
0.112008
Towards unsupervised whole-object segmentation: Combining automated matting with boundary detection · CVPR 2008
Computer vision › Segmentation and scene understanding › object segmentation
unsupervised object segmentation
0.112008
Towards unsupervised whole-object segmentation: Combining automated matting with boundary detection · CVPR 2008
Machine learning › Graph learning › graph clustering
spectral clustering
0.012008
Towards unsupervised whole-object segmentation: Combining automated matting with boundary detection · CVPR 2008

Methods — techniques the papers use, named apart from their topics

optimal transport · 0.9distribution transport · 0.9hindsight information constraints · 0.5future-conditional value functions · 0.5counterfactual reasoning · 0.5retrace(lambda) · 0.2q(lambda) · 0.2spectral clustering · 0.1image matting · 0.1boundary detection · 0.1
YearPublicationVenuePosition
2024 Gaps in the Safety Evaluation of Generative AI
abstract
Generative AI systems produce a range of ethical and social risks. Evaluation of these risks is a critical step on the path to ensuring the safety of these systems. However, evaluation requires the availability of validated and established measurement approaches and tools. In this paper, we provide an empirical review of the methods and tools that are available for evaluating known safety of generative AI systems to date. To this end, we review more than 200 safety-related evaluations that have been applied to generative AI systems. We categorise each evaluation along multiple axes to create a detailed snapshot of the safety evaluation landscape to date. We release this data for researchers and AI safety practitioners (https://bitly.ws/3hUzu). Analysing the current safety evaluation landscape reveals three systemic ”evaluation gaps”. First, a ”modality gap” emerges as few safety evaluations exist for non-text modalities. Second, a ”risk coverage gap” arises as evaluations for several ethical and social risks are simply lacking. Third, a ”context gap” arises as most safety evaluations are model-centric and fail to take into account the broader context in which AI systems operate. Devising next steps for safety practitioners based on these findings, we present tactical ”low-hanging fruit” steps towards closing the identified evaluation gaps and their limitations. We close by discussing the role and limitations of safety evaluation to ensure the safety of generative AI systems.
Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Ramona Comanescu, Canfer Akbulut, Thomas S. Stepleton, Juan Mateos-Garcia, A. Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, William Isaac 0001, Laura Weidinger
AIES (1)7
2021 Counterfactual Credit Assignment in Model-Free Reinforcement Learning
abstract
Credit assignment in reinforcement learning is the problem of measuring an action’s influence on future rewards. In particular, this requires separating skill from luck, i.e. disentangling the effect of an action on rewards from that of external factors and subsequent actions. To achieve this, we adapt the notion of counterfactuals from causality theory to a model-free RL setup. The key idea is to condition value functions on future events, by learning to extract relevant information from a trajectory. We formulate a family of policy gradient algorithms that use these future-conditional value functions as baselines or critics, and show that they are provably low variance. To avoid the potential bias from conditioning on future information, we constrain the hindsight information to not contain information about the agent’s actions. We demonstrate the efficacy and validity of our algorithm on a number of illustrative and challenging problems.
Thomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor, Alaa Saade, Anna Harutyunyan, Will Dabney, Thomas S. Stepleton, Nicolas Heess, Arthur Guez, Eric Moulines, Marcus Hutter, Lars Buesing, Rémi Munos
ICML8
2020 A General Approach to Fairness with Optimal Transport
abstract
We propose a general approach to fairness based on transporting distributions corresponding to different sensitive attributes to a common distribution. We use optimal transport theory to derive target distributions and methods that allow us to achieve fairness with minimal changes to the unfair model. Our approach is applicable to both classification and regression problems, can enforce different notions of fairness, and enable us to achieve a Pareto-optimal trade-off between accuracy and fairness. We demonstrate that it outperforms previous approaches in several benchmark fairness datasets.
Silvia Chiappa, Ray Jiang, Thomas S. Stepleton, Aldo Pacchiano, Heinrich Jiang, John Aslanides
AAAI3
2019 Wasserstein Fair Classification
Ray Jiang, Aldo Pacchiano, Thomas S. Stepleton, Heinrich Jiang, Silvia Chiappa
UAI3
2016 Q(λ) with Off-Policy Corrections
Anna Harutyunyan, Marc G. Bellemare, Thomas S. Stepleton, Rémi Munos
ALT3
2016 Safe and Efficient Off-Policy Reinforcement Learning
abstract
In this work, we take a fresh look at some old and new algorithms for off-policy, return-based reinforcement learning. Expressing these in a common form, we derive a novel algorithm, Retrace(lambda), with three desired properties: (1) it has low variance; (2) it safely uses samples collected from any behaviour policy, whatever its degree of "off-policyness"; and (3) it is efficient as it makes the best use of samples collected from near on-policy behaviour policies. We analyse the contractive nature of the related operator under both off-policy policy evaluation and control settings and derive online sample-based algorithms. We believe this is the first return-based off-policy control algorithm converging a.s. to Q* without the GLIE assumption (Greedy in the Limit with Infinite Exploration). As a corollary, we prove the convergence of Watkins' Q(lambda), which was an open problem since 1989. We illustrate the benefits of Retrace(lambda) on a standard suite of Atari 2600 games.
Rémi Munos, Thomas S. Stepleton, Anna Harutyunyan, Marc G. Bellemare
NIPS2
2008 Towards unsupervised whole-object segmentation: Combining automated matting with boundary detection
abstract
We propose a novel step toward the unsupervised segmentation of whole objects by combining ldquohintsrdquo of partial scene segmentation offered by multiple soft, binary mattes. These mattes are implied by a set of hypothesized object boundary fragments in the scene. Rather than trying to find or define a single ldquobestrdquo segmentation, we generate multiple segmentations of an image. This reflects contemporary methods for unsupervised object discovery from groups of images, and it allows us to define intuitive evaluation metrics for our sets of segmentations based on the accurate and parsimonious delineation of scene objects. Our proposed approach builds on recent advances in spectral clustering, image matting, and boundary detection. It is demonstrated qualitatively and quantitatively on a dataset of scenes and is suitable for current work in unsupervised object discovery without top-down knowledge.
Andrew N. Stein, Thomas S. Stepleton, Martial Hebert
CVPR2
2007 Iterative design of a Braille writing tutor to combat illiteracy
abstract
Less than 3% of the 145 million blind people living in developing countries are literate. This low literacy rate is partly due to the lack of trained teachers and the challenges associated with learning to write Braille on a traditional slate and stylus. These challenges include writing from right to left, writing mirrored images of letters, and receiving significantly delayed feedback. Extensive conversations with the Mathru School for the Blind near Bangalore, India, revealed the need for a robust, low-power, low-cost Braille writing tutor. We present an iterative and participatory design process resulting in the creation and refinement of a prototype Braille writing tutor system. This system uses a novel input device to capture a student's activity on a slate using a stylus and uses a range of techniques to teach Braille writing skills to both beginner and advanced students. We report on lessons learned from the implementation of this project and from a six-week pilot study at the Mathru school, and outline future directions for improvement.
Nidhi Kalra, Tom Lauwers, Daniel Dewey, Thomas S. Stepleton, M. Bernardine Dias
ICTD4
2003 A real-time vision module for interactive perceptual agents
Bruce A. Maxwell, Nathaniel Fairfield, Nikolas Johnson, Pukar Malla, Paul E. Dickson, Suor Kim, Stephanie Wojtkowski, Thomas S. Stepleton
Mach. Vis. Appl.8
2002 Work-Augmented Laziness with the Los Task Request System
Thomas S. Stepleton
LISA1