Nicholay Topin

dblp:165/3324 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
3since 2021 · last 2022
0009-0006-5376-2801ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 57% Reinforcement learning · 35% Robot manipulation · 8%
Human-computer interaction and pervasive computing
2 papers
Usability and user experience research · 46% Human-AI interaction · 46% Games and playful interaction · 9%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › interpretability
explainable reinforcement learning
0.922021
Iterative Bounding MDPs: Learning Interpretable Policies via Non-Interpretable Methods · AAAI 2021
Generation of Policy-Level Explanations for Reinforcement Learning · AAAI 2019
Machine learning › Trustworthy machine learning
interpretability
0.832022
Use-Case-Grounded Simulations for Explanation Evaluation · NeurIPS 2022
Iterative Bounding MDPs: Learning Interpretable Policies via Non-Interpretable Methods · AAAI 2021
Generation of Policy-Level Explanations for Reinforcement Learning · AAAI 2019
Machine learning › Trustworthy machine learning › interpretability
explanation evaluation
0.612022
Use-Case-Grounded Simulations for Explanation Evaluation · NeurIPS 2022
Human-AI interaction
simulation-based evaluation
0.612022
Use-Case-Grounded Simulations for Explanation Evaluation · NeurIPS 2022
Usability and user experience research
user study
0.612022
Use-Case-Grounded Simulations for Explanation Evaluation · NeurIPS 2022
Machine learning › Trustworthy machine learning › interpretability › explainable reinforcement learning
decision tree policy
0.512021
Iterative Bounding MDPs: Learning Interpretable Policies via Non-Interpretable Methods · AAAI 2021
Machine learning › Reinforcement learning
demonstration dataset
0.412019
MineRL: A Large-Scale Dataset of Minecraft Demonstrations · IJCAI 2019
Machine learning › Reinforcement learning
imitation learning
0.412019
MineRL: A Large-Scale Dataset of Minecraft Demonstrations · IJCAI 2019
Robotics › Robot manipulation
learning from demonstration
0.412019
MineRL: A Large-Scale Dataset of Minecraft Demonstrations · IJCAI 2019
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.212015
Portable Option Discovery for Automated Learning Transfer in Object-Oriented Markov Decision Processes · IJCAI 2015
Machine learning › Reinforcement learning › hierarchical reinforcement learning › temporal abstraction
options
0.212015
Portable Option Discovery for Automated Learning Transfer in Object-Oriented Markov Decision Processes · IJCAI 2015
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery
0.212015
Portable Option Discovery for Automated Learning Transfer in Object-Oriented Markov Decision Processes · IJCAI 2015
Machine learning › Reinforcement learning
transfer learning in reinforcement learning
0.212015
Portable Option Discovery for Automated Learning Transfer in Object-Oriented Markov Decision Processes · IJCAI 2015
Machine learning › Reinforcement learning
markov decision process
0.112015
Portable Option Discovery for Automated Learning Transfer in Object-Oriented Markov Decision Processes · IJCAI 2015

Methods — techniques the papers use, named apart from their topics

simulated evaluations · 1.1algorithmic agents · 1.1deep reinforcement learning · 0.8value update · 0.5masking procedure · 0.5markov decision process · 0.5markov chain · 0.4abstracted policy graph · 0.4
YearPublicationVenuePosition
2022 Use-Case-Grounded Simulations for Explanation Evaluation
abstract
A growing body of research runs human subject evaluations to study whether providing users with explanations of machine learning models can help them with practical real-world use cases. However, running user studies is challenging and costly, and consequently each study typically only evaluates a limited number of different settings, e.g., studies often only evaluate a few arbitrarily selected model explanation methods. To address these challenges and aid user study design, we introduce Simulated Evaluations (SimEvals). SimEvals involve training algorithmic agents that take as input the information content (such as model explanations) that would be presented to the user, to predict answers to the use case of interest. The algorithmic agent's test set accuracy provides a measure of the predictiveness of the information content for the downstream use case. We run a comprehensive evaluation on three real-world use cases (forward simulation, model debugging, and counterfactual reasoning) to demonstrate that SimEvals can effectively identify which explanation methods will help humans for each use case. These results provide evidence that \simevals{} can be used to efficiently screen an important set of user study design decisions, e.g., selecting which explanations should be presented to the user, before running a potentially costly user study.
Valerie Chen, Nari Johnson, Nicholay Topin, Gregory Plumb, Ameet Talwalkar
NeurIPS3
2022 MAVIPER: Learning Decision Tree Policies for Interpretable Multi-agent Reinforcement Learning
Stephanie Milani, Zhicheng Zhang 0003, Nicholay Topin, Zheyuan Shi, Charles A. Kamhoua, Evangelos E. Papalexakis, Fei Fang 0001
ECML/PKDD (4)3
2021 Iterative Bounding MDPs: Learning Interpretable Policies via Non-Interpretable Methods
abstract
Current work in explainable reinforcement learning generally produces policies in the form of a decision tree over the state space. Such policies can be used for formal safety verification, agent behavior prediction, and manual inspection of important features. However, existing approaches fit a decision tree after training or use a custom learning procedure which is not compatible with new learning techniques, such as those which use neural networks. To address this limitation, we propose a novel Markov Decision Process (MDP) type for learning decision tree policies: Iterative Bounding MDPs (IBMDPs). An IBMDP is constructed around a base MDP so each IBMDP policy is guaranteed to correspond to a decision tree policy for the base MDP when using a method-agnostic masking procedure. Because of this decision tree equivalence, any function approximator can be used during training, including a neural network, while yielding a decision tree policy for the base MDP. We present the required masking procedure as well as a modified value update step which allows IBMDPs to be solved using existing algorithms. We apply this procedure to produce IBMDP variants of recent reinforcement learning methods. We empirically show the benefits of our approach by solving IBMDPs to produce decision tree policies for the base MDPs.
Nicholay Topin, Stephanie Milani, Fei Fang 0001, Manuela M. Veloso
AAAI1
2019 Generation of Policy-Level Explanations for Reinforcement Learning
abstract
Though reinforcement learning has greatly benefited from the incorporation of neural networks, the inability to verify the correctness of such systems limits their use. Current work in explainable deep learning focuses on explaining only a single decision in terms of input features, making it unsuitable for explaining a sequence of decisions. To address this need, we introduce Abstracted Policy Graphs, which are Markov chains of abstract states. This representation concisely summarizes a policy so that individual decisions can be explained in the context of expected future transitions. Additionally, we propose a method to generate these Abstracted Policy Graphs for deterministic policies given a learned value function and a set of observed transitions, potentially off-policy transitions used during training. Since no restrictions are placed on how the value function is generated, our method is compatible with many existing reinforcement learning methods. We prove that the worst-case time complexity of our method is quadratic in the number of features and linear in the number of provided transitions, O(|F|2|tr samples|). By applying our method to a family of domains, we show that our method scales well in practice and produces Abstracted Policy Graphs which reliably capture relationships within these domains.
Nicholay Topin, Manuela M. Veloso
AAAI1
2019 MineRL: A Large-Scale Dataset of Minecraft Demonstrations
abstract
The sample inefficiency of standard deep reinforcement learning methods precludes their application to many real-world problems. Methods which leverage human demonstrations require fewer samples but have been researched less. As demonstrated in the computer vision and natural language processing communities, large-scale datasets have the capacity to facilitate research by serving as an experimental and benchmarking platform for new methods. However, existing datasets compatible with reinforcement learning simulators do not have sufficient scale, structure, and quality to enable the further development and evaluation of methods focused on using human examples. Therefore, we introduce a comprehensive, large-scale, simulator-paired dataset of human demonstrations: MineRL. The dataset consists of over 60 million automatically annotated state-action pairs across a variety of related tasks in Minecraft, a dynamic, 3D, open-world environment. We present a novel data collection scheme which allows for the ongoing introduction of new tasks and the gathering of complete state information suitable for a variety of methods. We demonstrate the hierarchality, diversity, and scale of the MineRL dataset. Further, we show the difficulty of the Minecraft domain along with the potential of MineRL in developing techniques to solve key research challenges within it.
William H. Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden R. Codel, Manuela M. Veloso, Ruslan Salakhutdinov
IJCAI3
2019 Online Planning for Autonomous Underwater Vehicles Performing Information Gathering Tasks in Large Subsea Environments
abstract
We present an anytime Monte Carlo tree search (MCTS) algorithm to generate real-time, near-optimal search paths in large subsea environments. The MCTS planner continuously builds a tree of the search space until either the allowed time per move is reached or the budget constraint for the search mission is met. In order to improve the performance of the MCTS planner, we propose a novel heuristic action selection policy to determine the value of a leaf node. The proposed heuristic is tailored to problems where making a turn incurs a higher cost than moving straight, such as the case on autonomous underwater vehicles. Through extensive simulations, we show that our heuristic yields a significant performance improvement over a lawnmover path planner - a commonly employed approach in subsea search applications - and over a simple MCTS planner where actions are selected uniformly at random. In our numerical illustrations, we use a real data set abstracted from sonar measurements acquired from the Boston Harbor.
Harun Yetkin, James McMahon, Nicholay Topin, Artur Wolek, Zachary Waters, Daniel J. Stilwell
IROS3
2015 Portable Option Discovery for Automated Learning Transfer in Object-Oriented Markov Decision Processes
Nicholay Topin, Nicholas Haltmeyer, Shawn Squire, John Winder, Marie desJardins, James MacGlashan
IJCAI1