Anil Yildiz

dblp:224/8114 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0001-8194-4895ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Planning, search and constraint satisfaction · 87% Learning theory · 4% Optimization for machine learning · 4%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
1.022021
Bayesian Optimized Monte Carlo Planning · AAAI 2021
Improved POMDP Tree Search Planning with Prioritized Action Branching · AAAI 2021
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
online planning
1.022021
Bayesian Optimized Monte Carlo Planning · AAAI 2021
Improved POMDP Tree Search Planning with Prioritized Action Branching · AAAI 2021
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
partially observable markov decision process
1.022021
Bayesian Optimized Monte Carlo Planning · AAAI 2021
Improved POMDP Tree Search Planning with Prioritized Action Branching · AAAI 2021
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.112021
Bayesian Optimized Monte Carlo Planning · AAAI 2021
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.112021
Bayesian Optimized Monte Carlo Planning · AAAI 2021
Machine learning › Learning theory › information-theoretic learning
information gain
0.112021
Improved POMDP Tree Search Planning with Prioritized Action Branching · AAAI 2021

Methods — techniques the papers use, named apart from their topics

score function · 0.5prioritized action branching · 0.5gaussian process · 0.5bayesian optimization · 0.5
YearPublicationVenuePosition
2026 Backward Monte Carlo Tree Search: Charting Unsafe Regions in the Belief-Space
abstract
Safety-critical systems often operate in partially observable environments, where assessing the safety of the underlying policy remains a fundamental challenge. This study focuses on evaluating policies by identifying regions of the belief-space that can lead the system’s policy to an undesirable state with a non-negligible probability. In this paper, we introduce Backward Monte Carlo Tree Search, the first Monte Carlo tree search framework that expands backward in time within the belief-space. The tree search begins from an undesired terminal belief and recursively explores its possible predecessors, constructing a tree of belief transitions that could lead to an unsafe outcome within a given horizon. Evaluations in gridworld and autonomous driving domains show that identifying beliefs from which failures may occur enables runtime risk forecasting and targeted policy retraining, marking a conceptual shift in how safety is validated under uncertainty.
Anil Yildiz, Esen Yel, Marcell Vazquez-Chanlatte, Kyle Hollins Wray, Mykel J. Kochenderfer, Stefan J. Witwicki
J. Artif. Intell. Res.1
2023 Experience Filter: Using Past Experiences on Unseen Tasks or Environments
abstract
One of the bottlenecks of training autonomous vehicle (AV) agents is the variability of training environments. Since learning optimal policies for unseen environments is often very costly and requires substantial data collection, it becomes computationally intractable to train the agent on every possible environment or task the AV may encounter.This paper introduces a zero-shot filtering approach to interpolate learned policies of past experiences to generalize to unseen ones. We use an experience kernel to correlate environments. These correlations are then exploited to produce policies for new tasks or environments from learned policies. We demonstrate our methods on an autonomous vehicle driving through T-intersections with different characteristics, where its behavior is modeled as a partially observable Markov decision process (POMDP). We first construct compact representations of learned policies for POMDPs with unknown transition functions given a dataset of sequential actions and observations. Then, we filter parameterized policies of previously visited environments to generate policies to new, unseen environments. We demonstrate our approaches on both an actual AV and a high-fidelity simulator. Results indicate that our experience filter offers a fast, low-effort, and near-optimal solution to create policies for tasks or environments never seen before. Furthermore, the generated new policies outperform the policy learned using the entire data collected from past environments, suggesting that the correlation among different environments can be exploited and irrelevant ones can be filtered out.
Anil Yildiz, Esen Yel, Anthony Corso 0001, Kyle Hollins Wray, Stefan J. Witwicki, Mykel J. Kochenderfer
IV1
2021 Improved POMDP Tree Search Planning with Prioritized Action Branching
abstract
Online solvers for partially observable Markov decision processes have difficulty scaling to problems with large action spaces. This paper proposes a method called PA-POMCPOW to sample a subset of the action space that provides varying mixtures of exploitation and exploration for inclusion in a search tree. The proposed method first evaluates the action space according to a score function that is a linear combination of expected reward and expected information gain. The actions with the highest score are then added to the search tree during tree expansion. Experiments show that PA-POMCPOW is able to outperform existing state-of-the-art solvers on problems with large discrete action spaces.
John Mern, Anil Yildiz, Lawrence Bush, Tapan Mukerji, Mykel J. Kochenderfer
AAAI2
2021 Bayesian Optimized Monte Carlo Planning
abstract
Online solvers for partially observable Markov decision processes have difficulty scaling to problems with large action spaces. Monte Carlo tree search with progressive widening attempts to improve scaling by sampling from the action space to construct a policy search tree. The performance of progressive widening search is dependent upon the action sampling policy, often requiring problem-specific samplers. In this work, we present a general method for efficient action sampling based on Bayesian optimization. The proposed method uses a Gaussian process to model a belief over the action-value function and selects the action that will maximize the expected improvement in the optimal action value. We implement the proposed approach in a new online tree search algorithm called Bayesian Optimized Monte Carlo Planning (BOMCP). Several experiments show that BOMCP is better able to scale to large action space POMDPs than existing state-of-the-art tree search solvers.
John Mern, Anil Yildiz, Zachary Sunberg, Tapan Mukerji, Mykel J. Kochenderfer
AAAI2