EDBT 2026 Demo / reviewers in the wild / expert
Tong Wang 0011
dblp:51/6856-11
· DBLP profile ↗
15ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0001-8687-4208ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Theory of computation · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PIE - Partially Interpretable Estimators with Refinement
Tong Wang 0011, Yunyi Li |
INFORMS J. Comput. | 1 |
| 2026 | Fast Rashomon Sets of Sparse Rule SetsabstractAbstract A sparse rule set (SRS) is a small predictive model that is a disjunctive normal form – an “OR of ANDs”. SRS models are understandable to human experts and robust to outliers. Constructing an SRS efficiently has always been one of the fundamental problems of interpretable machine learning. However, the practical challenge goes beyond simply optimizing for a single sparse SRS; the first interpretable model an algorithm finds often has flaws that need to be fixed. Thus, constructing a model for practical use involves human interaction with an algorithm. If we work within the Rashomon set paradigm, we would generate many models that a human can explore; this simplifies the interaction, but requires generating a large pool of models, which is computationally demanding. This work presents FastSRS, an efficient algorithm for generating a large quantity of high quality SRS models. Its key advantage is a set of theoretical bounds that efficiently reduces the size of the search space to enable fast computation. FastSRS produces models on the accuracy vs. sparsity frontier more consistently and efficiently than previous approaches, scales better to larger datasets, and sets a new state of the art for generating Rashomon sets for SRSs. Cristina Molero-Río, Boxuan Li, Tong Wang 0011, Cynthia Rudin |
Mach. Learn. | 3 |
| 2025 | ConPro-GAIL: Interpretable Policy Learning via Conceptual Prototyping for Human Spatiotemporal Decision UnderstandingabstractThe problem of human spatiotemporal (ST) decision understanding, which consists of extracting faithful and interpretable decision strategies from human agents' behavioral records in space and time, is important for many applications, such as improving taxi drivers' route planning and efficiency. It is challenging because ST data are not as readily interpretable as images or text data, which leads to difficulties in constructing data-driven explanations. Existing research on this topic defines the problem as a Markov Decision Process (MDP) and uses imitation learning to extract a policy approximating the underlying human policy for post-hoc interpretation. However, such methods cannot provide direct interpretation through model training and may result in incomprehensible interpretations when using ST data. We address these limitations by designing ConPro-GAIL, a prototype-based interpretable GAIL model for intrinsically interpretable ST policy extraction. ConPro-GAIL learns and represents the optimal policy in terms of prototypical sets of concepts that correspond to general scenarios in the MDP. It explains a decision associated with an input state via inductive generalization from what occurred in the state's most similar prototypes to the input state itself. Experiments and case studies on two taxi trajectory datasets show that ConPro-GAIL achieves better policy faithfulness than its black-box competitors and better interpretability than post-hoc explainers. Ronilo J. Ragodos, Xun Zhou 0001, Tong Wang 0011, Yajun Pan 0002, Jun Luo 0007 |
SIGSPATIAL/GIS | 3 |
| 2025 | ProtoPairNet: Interpretable Regression through Prototypical Pair ReasoningabstractWe present Prototypical Pair Network (ProtoPairNet), a novel interpretable architecture that combines deep learning with case-based reasoning to predict continuous targets. While prototype-based models have primarily addressed image classification with discrete outputs, extending these methods to continuous targets, such as regression, poses significant challenges. Existing architectures which rely heavily on one-to-one comparison with prototypes lack the directional information necessary for continuous predictions. Our method redefines the role of prototypes in such tasks by incorporating prototypical pairs into the reasoning process. Predictions are derived based on the input's relative dissimilarities to these pairs, leveraging an intuitive geometric interpretation. Our method further reduces the complexity of the reasoning process by relying on the single most relevant pair of prototypes, rather than all prototypes in the model as was done in prior works. Our model is versatile enough to be used in both vision-based regression and continuous control in reinforcement learning. Our experiments demonstrate that ProtoPairNet achieves performance on par with its black-box counterparts across these tasks. Comprehensive analyses confirm the meaningfulness of prototypical pairs and the faithfulness of our model’s interpretations, and extensive user studies highlight our model's improved interpretability over existing methods. Rose Gurung, Ronilo J. Ragodos, Chiyu Ma, Tong Wang 0011 |
NeurIPS | 4 |
| 2024 | Sparse and Faithful Explanations Without Sparse ModelsabstractEven if a model is not globally sparse, it is possible for decisions made from that model to be accurately and faithfully described by a small number of features. For instance, an application for a large loan might be denied to someone because they have no credit history, which overwhelms any evidence towards their creditworthiness. In this work, we introduce the Sparse Explanation Value (SEV), a new way of measuring sparsity in machine learning models. In the loan denial example above, the SEV is 1 because only one factor is needed to explain why the loan was denied. SEV is a measure of decision sparsity rather than overall model sparsity, and we are able to show that many machine learning models – even if they are not sparse – actually have low decision sparsity, as measured by SEV. SEV is defined using movements over a hypercube, allowing SEV to be defined consistently over various model classes, with movement restrictions reflecting real-world constraints. Our algorithms reduce SEV without sacrificing accuracy, providing sparse and completely faithful explanations, even without globally sparse models. Yiyang Sun 0001, Zhi Chen 0009, Vittorio Orlandi, Tong Wang 0011, Cynthia Rudin |
AISTATS | 4 |
| 2024 | Improving Decision SparsityabstractSparsity is a central aspect of interpretability in machine learning. Typically, sparsity is measured in terms of the size of a model globally, such as the number of variables it uses. However, this notion of sparsity is not particularly relevant for decision making; someone subjected to a decision does not care about variables that do not contribute to the decision. In this work, we dramatically expand a notion of *decision sparsity* called the *Sparse Explanation Value* (SEV) so that its explanations are more meaningful. SEV considers movement along a hypercube towards a reference point. By allowing flexibility in that reference and by considering how distances along the hypercube translate to distances in feature space, we can derive sparser and more meaningful explanations for various types of function classes. We present cluster-based SEV and its variant tree-based SEV, introduce a method that improves credibility of explanations, and propose algorithms that optimize decision sparsity in machine learning models. Yiyang Sun 0001, Tong Wang 0011, Cynthia Rudin |
NeurIPS | 2 |
| 2022 | ProtoX: Explaining a Reinforcement Learning Agent via PrototypingabstractWhile deep reinforcement learning has proven to be successful in solving control tasks, the ``black-box'' nature of an agent has received increasing concerns. We propose a prototype-based post-hoc \emph{policy explainer}, ProtoX, that explains a black-box agent by prototyping the agent's behaviors into scenarios, each represented by a prototypical state. When learning prototypes, ProtoX considers both visual similarity and scenario similarity. The latter is unique to the reinforcement learning context since it explains why the same action is taken in visually different states. To teach ProtoX about visual similarity, we pre-train an encoder using contrastive learning via self-supervised learning to recognize states as similar if they occur close together in time and receive the same action from the black-box agent. We then add an isometry layer to allow ProtoX to adapt scenario similarity to the downstream task. ProtoX is trained via imitation learning using behavior cloning, and thus requires no access to the environment or agent. In addition to explanation fidelity, we design different prototype shaping terms in the objective function to encourage better interpretability. We conduct various experiments to test ProtoX. Results show that ProtoX achieved high fidelity to the original black-box agent while providing meaningful and understandable explanations. Ronilo J. Ragodos, Tong Wang 0011, Qihang Lin, Xun Zhou 0001 |
NeurIPS | 2 |
| 2022 | A holistic approach to interpretability in financial lending: Models, visualizations, and summary-explanations
Kangcheng Lin, Cynthia Rudin, Yaron Shaposhnik, Tong Wang 0011 |
Decis. Support Syst. | 6 |
| 2022 | Disjunctive Rule ListsabstractIn this study, we present an interpretable model, disjunctive rule list (DisRL) for regression. This research is motivated by the increasing need for model interpretability, especially in high-stakes decisions such as medicine, where decisions are made on or related to humans. DisRL is a generalized form of rule lists. A DisRL model consists of a list of disjunctive rules embedded in an if-else logic structure that stratifies the data space. Compared with traditional decision trees and other rule list models in the literature that stratify the feature space with single itemsets (an itemset is a conjunction of conditions), each disjunctive rule in DisRL uses a set of itemsets to collectively cover a subregion in the feature space. In addition, a DisRL model is constructed under a global objective that balances the predictive performance and model complexity. To train a DisRL model, we devise a hierarchical stochastic local search algorithm that exploits the properties of DisRL’s unique structure to improve search efficiency. The algorithm adopts the main structure of simulated annealing and customizes the proposing strategy for faster convergence. Meanwhile, the algorithm uses a prefix bound to locate a subset of the search area, effectively pruning the search space at each iteration. An ablation study shows the effectiveness of this strategy in pruning the search space. Experiments on public benchmark datasets demonstrate that DisRL outperforms baseline interpretable models, including decision trees and other rule-based regressors. History: Accepted by J. Paul Brooks, Area Editor for Applications in Biology, Medicine, & Healthcare. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplementary Information [ https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.1242 ] or is available from the IJOC GitHub software repository ( https://github.com/INFORMSJoC ) at [ http://dx.doi.org/10.5281/zenodo.6954927 ]. Ronilo J. Ragodos, Tong Wang 0011 |
INFORMS J. Comput. | 2 |
| 2022 | Causal Rule Sets for Identifying Subgroups with Enhanced Treatment EffectsabstractA key question in causal inference analyses is how to find subgroups with elevated treatment effects. This paper takes a machine learning approach and introduces a generative model, causal rule sets (CRS), for interpretable subgroup discovery. A CRS model uses a small set of short decision rules to capture a subgroup in which the average treatment effect is elevated. We present a Bayesian framework for learning a causal rule set. The Bayesian model consists of a prior that favors simple models for better interpretability as well as avoiding overfitting and a Bayesian logistic regression that captures the likelihood of data, characterizing the relation between outcomes, attributes, and subgroup membership. The Bayesian model has tunable parameters that can characterize subgroups with various sizes, providing users with more flexible choices of models from the treatment-efficient frontier. We find maximum a posteriori models using iterative discrete Monte Carlo steps in the joint solution space of rules sets and parameters. To improve search efficiency, we provide theoretically grounded heuristics and bounding strategies to prune and confine the search space. Experiments show that the search algorithm can efficiently recover true underlying subgroups. We apply CRS on public and real-world data sets from domains in which interpretability is indispensable. We compare CRS with state-of-the-art rule-based subgroup discovery models. Results show that CRS achieves consistently competitive performance on data sets from various domains, represented by high treatment-efficient frontiers. Summary of Contribution: This paper is motivated by the large heterogeneity of treatment effect in many applications and the need to accurately locate subgroups for enhanced treatment effect. Existing methods either rely on prior hypotheses to discover subgroups or greedy methods, such as tree-based recursive partitioning. Our method adopts a machine learning approach to find an optimal subgroup learned with a carefully global objective. Our model is more flexible in capturing subgroups by using a set of short decision rules compared with tree-based baselines. We evaluate our model using a novel metric, treatment-efficient frontier, that characterizes the trade-off between the subgroup size and achievable treatment effect, and our model demonstrates better performance than baseline models. Tong Wang 0011, Cynthia Rudin |
INFORMS J. Comput. | 1 |
| 2021 | Hybrid Predictive Models: When an Interpretable Model Collaborates with a Black-box ModelabstractInterpretable machine learning has become a strong competitor for black-box models. However, the possible loss of the predictive performance for gaining understandability is often inevitable, especially when it needs to satisfy users with diverse backgrounds or high standards for what is considered interpretable. This tension puts practitioners in a dilemma of choosing between high accuracy (black-box models) and interpretability (interpretable models). In this work, we propose a novel framework for building a Hybrid Predictive Model that integrates an interpretable model with any pre-trained black-box model to combine their strengths. The interpretable model substitutes the black-box model on a subset of data where the interpretable model is most competent, gaining transparency at a low cost of the predictive accuracy. We design a principled objective function that considers predictive accuracy, model interpretability, and model transparency (defined as the percentage of data processed by the interpretable substitute.) Under this framework, we propose two hybrid models, one substituting with association rules and the other with linear models, and design customized training algorithms for both models. We test the hybrid models on structured data and text data where interpretable models collaborate with various state-of-the-art black-box models. Results show that hybrid models obtain an efficient trade-off between transparency and predictive performance, characterized by pareto frontiers. Finally, we apply the proposed model on a real-world patients dataset for predicting cardiovascular disease and propose multi-model Pareto frontiers to assist model selection in real applications. Tong Wang 0011, Qihang Lin |
J. Mach. Learn. Res. | 1 |
| 2020 | Transparency Promotion with Model-Agnostic Linear CompetitorsabstractWe propose a novel type of hybrid model for multi-class classification, which utilizes competing linear models to collaborate with an existing black-box model, promoting transparency in the decision-making process. Our proposed hybrid model, Model-Agnostic Linear Competitors (MALC), brings together the interpretable power of linear models and the good predictive performance of the state-of-the-art black-box models. We formulate the training of a MALC model as a convex optimization problem, optimizing the predictive accuracy and transparency (defined as the percentage of data captured by the linear models) in the objective function. Experiments show that MALC offers more model flexibility for users to balance transparency and accuracy, in contrast to the currently available choice of either a pure black-box model or a pure interpretable model. The human evaluation also shows that more users are likely to choose MALC for this model flexibility compared with interpretable models and black-box models. Hassan Rafique, Tong Wang 0011, Qihang Lin, Arshia Singhani |
ICML | 2 |
| 2017 | A Bayesian Framework for Learning Rule Sets for Interpretable ClassificationabstractWe present a machine learning algorithm for building classifiers that are comprised of a small number of short rules. These are restricted disjunctive normal form models. An example of a classifier of this form is as follows: If $X$ satisfies (condition $A$ AND condition $B$) OR (condition $C$) OR $\cdots$, then $Y=1$. Models of this form have the advantage of being interpretable to human experts since they produce a set of rules that concisely describe a specific class. We present two probabilistic models with prior parameters that the user can set to encourage the model to have a desired size and shape, to conform with a domain-specific definition of interpretability. We provide a scalable MAP inference approach and develop theoretical bounds to reduce computation by iteratively pruning the search space. We apply our method (Bayesian Rule Sets -- BRS) to characterize and predict user behavior with respect to in-vehicle context-aware personalized recommender systems. Our method has a major advantage over classical associative classification methods and decision trees in that it does not greedily grow the model. Tong Wang 0011, Cynthia Rudin, Finale Doshi-Velez, Erica Klampfl, Perry MacNeille |
J. Mach. Learn. Res. | 1 |
| 2016 | Bayesian Rule Sets for Interpretable ClassificationabstractA Rule Set model consists of a small number of short rules for interpretable classification, where an instance is classified as positive if it satisfies at least one of the rules. The rule set provides reasons for predictions, and also descriptions of a particular class. We present a Bayesian framework for learning Rule Set models, with prior parameters that the user can set to encourage the model to have a desired size and shape in order to conform with a domain-specific definition of interpretability. We use an efficient inference approach for searching for the MAP solution and provide theoretical bounds to reduce computation. We apply Rule Set models to ten UCI data sets and compare the performance with other interpretable and non-interpretable models. Tong Wang 0011, Cynthia Rudin, Finale Doshi-Velez, Erica Klampfl, Perry MacNeille |
ICDM | 1 |
| 2013 | Learning to Detect Patterns of Crime
Tong Wang 0011, Cynthia Rudin, Rich Sevieri |
ECML/PKDD (3) | 1 |