Xuyuan Xiong

dblp:388/1508 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 50% Trustworthy machine learning · 50%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › interpretability › explainable reinforcement learning
decision tree policy
0.912025
SPOT: Scalable Policy Optimization with Trees for Markov Decision Processes · NeurIPS 2025
Machine learning › Reinforcement learning
markov decision process
0.912025
SPOT: Scalable Policy Optimization with Trees for Markov Decision Processes · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability › explainable reinforcement learning
policy explanation
0.912025
SPOT: Scalable Policy Optimization with Trees for Markov Decision Processes · NeurIPS 2025
Machine learning › Reinforcement learning
policy optimization
0.912025
SPOT: Scalable Policy Optimization with Trees for Markov Decision Processes · NeurIPS 2025
Mathematical optimization › discrete optimization
mixed integer linear programming
0.312025
SPOT: Scalable Policy Optimization with Trees for Markov Decision Processes · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

parallel search · 1.7branch-and-bound · 1.7mixed-integer linear programming · 0.9mixed integer linear programming · 0.9
YearPublicationVenuePosition
2025 SPOT: Scalable Policy Optimization with Trees for Markov Decision Processes
abstract
Interpretable reinforcement learning policies are essential for high-stakes decision-making, yet optimizing decision tree policies in Markov Decision Processes (MDPs) remains challenging. We propose SPOT, a novel method for computing decision tree policies, which formulates the optimization problem as a mixed-integer linear program (MILP). To enhance efficiency, we employ a reduced-space branch-and-bound approach that decouples the MDP dynamics from tree-structure constraints, enabling efficient parallel search. This significantly improves runtime and scalability compared to previous methods. Our approach ensures that each iteration yields the optimal decision tree. Experimental results on standard benchmarks demonstrate that SPOT achieves substantial speedup and scales to larger MDPs with a significantly higher number of states. The resulting decision tree policies are interpretable and compact, maintaining transparency without compromising performance. These results demonstrate that our approach simultaneously achieves interpretability and scalability, delivering high-quality policies an order of magnitude faster than existing approaches.
Xuyuan Xiong, Pedro Chumpitaz-Flores, Kaixun Hua, Cheng Hua
NeurIPS1