EDBT 2026 Demo / reviewers in the wild / expert
Stephen Giguere 0001
dblp:14/8174 · also Stephen J. Giguere
· DBLP profile ↗
6ranked-venue papers
1as first author
2since 2021 · last 2022
0000-0001-6974-1139ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 48% Trustworthy machine learning · 36% Probabilistic and Bayesian machine learning · 9% | |
| Human-computer interaction and pervasive computing
1 paper |
User interface design and tools · 100% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
1.0 | 2 | 2022 | Fairness Guarantees under Demographic Shift · ICLR 2022 Offline Contextual Bandits with High Probability Fairness Guarantees · NeurIPS 2019 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.6 | 1 | 2022 | Fairness Guarantees under Demographic Shift · ICLR 2022 |
Machine learning › Reinforcement learning › off-policy evaluation
doubly robust estimation |
0.5 | 1 | 2021 | SOPE: Spectrum of Off-Policy Estimators · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
importance sampling |
0.5 | 1 | 2021 | SOPE: Spectrum of Off-Policy Estimators · NeurIPS 2021 |
Machine learning › Reinforcement learning
off-policy evaluation |
0.5 | 1 | 2021 | SOPE: Spectrum of Off-Policy Estimators · NeurIPS 2021 |
Machine learning › Reinforcement learning › bandit
contextual bandit |
0.4 | 1 | 2019 | Offline Contextual Bandits with High Probability Fairness Guarantees · NeurIPS 2019 |
Machine learning › Trustworthy machine learning › fairness
fairness guarantees |
0.4 | 1 | 2019 | Offline Contextual Bandits with High Probability Fairness Guarantees · NeurIPS 2019 |
Machine learning › Reinforcement learning › bandit › contextual bandit
offline contextual bandit |
0.4 | 1 | 2019 | Offline Contextual Bandits with High Probability Fairness Guarantees · NeurIPS 2019 |
Machine learning › Reinforcement learning
actor-critic methods |
0.2 | 1 | 2013 | Projected Natural Actor-Critic · NIPS 2013 |
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy optimization |
0.2 | 1 | 2013 | Projected Natural Actor-Critic · NIPS 2013 |
Machine learning › Reinforcement learning › actor-critic methods
natural actor-critic |
0.2 | 1 | 2013 | Projected Natural Actor-Critic · NIPS 2013 |
Machine learning › Optimization for machine learning › gradient-based optimization
proximal gradient method |
0.2 | 1 | 2013 | Basis Adaptation for Sparse Nonlinear Reinforcement Learning · AAAI 2013 |
Machine learning › Reinforcement learning
safe reinforcement learning |
0.2 | 1 | 2013 | Projected Natural Actor-Critic · NIPS 2013 |
Machine learning › Reinforcement learning
value function approximation |
0.2 | 1 | 2013 | Basis Adaptation for Sparse Nonlinear Reinforcement Learning · AAAI 2013 |
Machine learning › Learning theory › statistical learning theory
bias-variance tradeoff |
0.1 | 1 | 2021 | SOPE: Spectrum of Off-Policy Estimators · NeurIPS 2021 |
Visual content generation and editing
3d content creation |
0.0 | 1 | 2013 | Attribit: content creation with semantic attributes · UIST 2013 |
Methods — techniques the papers use, named apart from their topics
trajectory importance sampling · 0.5state-action visitation distribution · 0.5high-probability bounds · 0.4semantic attribute ranking · 0.3relative attribute learning · 0.3mirror descent · 0.3temporal difference · 0.2natural gradient descent · 0.2l1 regularization · 0.2bregman divergence · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Fairness Guarantees under Demographic Shift
Stephen Giguere 0001, Blossom Metevier, Bruno C. da Silva 0001, Yuriy Brun, Philip S. Thomas, Scott Niekum |
ICLR | 1 |
| 2021 | SOPE: Spectrum of Off-Policy EstimatorsabstractMany sequential decision making problems are high-stakes and require off-policy evaluation (OPE) of a new policy using historical data collected using some other policy. One of the most common OPE techniques that provides unbiased estimates is trajectory based importance sampling (IS). However, due to the high variance of trajectory IS estimates, importance sampling methods based on state-action visitation distributions (SIS) have recently been adopted. Unfortunately, while SIS often provides lower variance estimates for long horizons, estimating the state-action distribution ratios can be challenging and lead to biased estimates. In this paper, we present a new perspective on this bias-variance trade-off and show the existence of a spectrum of estimators whose endpoints are SIS and IS. Additionally, we also establish a spectrum for doubly-robust and weighted version of these estimators. We provide empirical evidence that estimators in this spectrum can be used to trade-off between the bias and variance of IS and SIS and can achieve lower mean-squared error than both IS and SIS. Christina J. Yuan, Yash Chandak, Stephen Giguere 0001, Philip S. Thomas, Scott Niekum |
NeurIPS | 3 |
| 2019 | Offline Contextual Bandits with High Probability Fairness GuaranteesabstractWe present RobinHood, an offline contextual bandit algorithm designed to satisfy a broad family of fairness constraints. Our algorithm accepts multiple fairness definitions and allows users to construct their own unique fairness definitions for the problem at hand. We provide a theoretical analysis of RobinHood, which includes a proof that it will not return an unfair solution with probability greater than a user-specified threshold. We validate our algorithm on three applications: a tutoring system in which we conduct a user study and consider multiple unique fairness definitions; a loan approval setting (using the Statlog German credit data set) in which well-known fairness definitions are applied; and criminal recidivism (using data released by ProPublica). In each setting, our algorithm is able to produce fair policies that achieve performance competitive with other offline and online contextual bandit algorithms. Blossom Metevier, Stephen Giguere 0001, Sarah Brockman, Ari Kobren, Yuriy Brun, Emma Brunskill, Philip S. Thomas |
NeurIPS | 2 |
| 2013 | Basis Adaptation for Sparse Nonlinear Reinforcement LearningabstractThis paper presents a new approach to representation discovery in reinforcement learning (RL) using basis adaptation. We introduce a general framework for basis adaptation as {\em nonlinear separable least-squares value function approximation} based on finding Frechet gradients of an error function using variable projection functionals. We then present a scalable proximal gradient-based approach for basis adaptation using the recently proposed mirror-descent framework for RL. Unlike traditional temporal-difference (TD) methods for RL, mirror descent based RL methods undertake proximal gradient updates of weights in a dual space, which is linked together with the primal space using a Legendre transform involving the gradient of a strongly convex function. Mirror descent RL can be viewed as a proximal TD algorithm using Bregman divergence as the distance generating function. We present a new class of regularized proximal-gradient based TD methods, which combine feature selection through sparse L1 regularization and basis adaptation. Experimental results are provided to illustrate and validate the approach. Sridhar Mahadevan, Stephen Giguere 0001, Nicholas Jacek |
AAAI | 2 |
| 2013 | Projected Natural Actor-CriticabstractNatural actor-critics are a popular class of policy search algorithms for finding locally optimal policies for Markov decision processes. In this paper we address a drawback of natural actor-critics that limits their real-world applicability - their lack of safety guarantees. We present a principled algorithm for performing natural gradient descent over a constrained domain. In the context of reinforcement learning, this allows for natural actor-critic algorithms that are guaranteed to remain within a known safe region of policy space. While deriving our class of constrained natural actor-critic algorithms, which we call Projected Natural Actor-Critics (PNACs), we also elucidate the relationship between natural gradient descent and mirror descent. Philip S. Thomas, William Dabney, Stephen Giguere 0001, Sridhar Mahadevan |
NIPS | 3 |
| 2013 | Attribit: content creation with semantic attributesabstractWe present AttribIt, an approach for people to create visual content using relative semantic attributes expressed in linguistic terms. During an off-line processing step, AttribIt learns semantic attributes for design components that reflect the high-level intent people may have for creating content in a domain (e.g. adjectives such as "dangerous", "scary" or "strong") and ranks them according to the strength of each learned attribute. Then, during an interactive design session, a person can explore different combinations of visual components using commands based on relative attributes (e.g. "make this part more dangerous"). Novel designs are assembled in real-time as the strengths of selected attributes are varied, enabling rapid, in-situ exploration of candidate designs. We applied this approach to 3D modeling and web design. Experiments suggest this interface is an effective alternative for novices performing tasks with high-level design goals. Siddhartha Chaudhuri, Evangelos Kalogerakis, Stephen Giguere 0001, Thomas A. Funkhouser |
UIST | 3 |