Wei Sun 0031

dblp:09/5042-31 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-5352-7629ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Trustworthy machine learning · 23% Reinforcement learning · 22% Probabilistic and Bayesian machine learning · 14%
Theoretical computer science
3 papers
Mathematical optimization · 61% Algorithmic game theory and mechanism design · 25% Algorithms and data structures · 14%
Databases, data mining, and information retrieval
5 papers
Data mining · 66% Recommender systems · 26% Information retrieval · 7%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational finance and economics · 100%

Topics — the 25 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization › discrete optimization
mixed integer linear programming
1.222023
Scalable Optimal Multiway-Split Decision Trees with Constraints · AAAI 2023
Constrained Prescriptive Trees via Column Generation · AAAI 2022
Machine learning › Trustworthy machine learning
interpretability
1.222023
Learning Prescriptive ReLU Networks · ICML 2023
Model Distillation for Revenue Optimization: Interpretable Personalized Pricing · ICML 2021
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
decision making under uncertainty
0.912025
Causal LLM Routing: End-to-End Regret Minimization from Observational Data · NeurIPS 2025
Machine learning › Efficient and distributed learning › inference efficiency
LLM routing
0.912025
Causal LLM Routing: End-to-End Regret Minimization from Observational Data · NeurIPS 2025
Machine learning › Reinforcement learning
regret minimization
0.912025
Causal LLM Routing: End-to-End Regret Minimization from Observational Data · NeurIPS 2025
Machine learning › Reinforcement learning
bandit
0.822020
Fatigue-Aware Bandits for Dependent Click Models · AAAI 2020
Dynamic Learning of Sequential Choice Bandit Problem under Marketing Fatigue · AAAI 2019
Machine learning › Time series and sequential data › time series analysis
time series forecasting
0.812024
Learning Optimal Projection for Forecast Reconciliation of Hierarchical Time Series · ICML 2024
Computational finance and economics
pricing
0.722021
Model Distillation for Revenue Optimization: Interpretable Personalized Pricing · ICML 2021
Latent Variable Copula Inference for Bundle Pricing from Retail Transaction Data · ICML 2014
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.712023
Learning Prescriptive ReLU Networks · ICML 2023
Data mining
interpretable machine learning
0.712023
Scalable Optimal Multiway-Split Decision Trees with Constraints · AAAI 2023
Algorithms and data structures › decision tree
decision tree learning
0.712023
Scalable Optimal Multiway-Split Decision Trees with Constraints · AAAI 2023
Machine learning › Trustworthy machine learning › causal machine learning
counterfactual learning
0.612022
Enhancing Counterfactual Classification Performance via Self-Training · AAAI 2022
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training
0.612022
Enhancing Counterfactual Classification Performance via Self-Training · AAAI 2022
Data mining › business intelligence
prescriptive analytics
0.612022
Constrained Prescriptive Trees via Column Generation · AAAI 2022
Mathematical optimization › large-scale optimization › decomposition methods
column generation
0.612022
Constrained Prescriptive Trees via Column Generation · AAAI 2022
Mathematical optimization
discrete optimization
0.612022
Constrained Prescriptive Trees via Column Generation · AAAI 2022
Computational finance and economics › mechanism design
revenue maximization
0.512021
Model Distillation for Revenue Optimization: Interpretable Personalized Pricing · ICML 2021
Recommender systems
sequential recommendation
0.412019
Dynamic Learning with Frequent New Product Launches: A Sequential Multinomial Logit Bandit Problem · ICML 2019
Algorithmic game theory and mechanism design › multi-armed bandit
multinomial logit bandit
0.412019
Dynamic Learning with Frequent New Product Launches: A Sequential Multinomial Logit Bandit Problem · ICML 2019
Algorithmic game theory and mechanism design › learning in games
online learning in games
0.412019
Dynamic Learning with Frequent New Product Launches: A Sequential Multinomial Logit Bandit Problem · ICML 2019
Mathematical optimization
online optimization
0.412019
Dynamic Learning with Frequent New Product Launches: A Sequential Multinomial Logit Bandit Problem · ICML 2019
Algorithmic game theory and mechanism design
regret minimization
0.412019
Dynamic Learning with Frequent New Product Launches: A Sequential Multinomial Logit Bandit Problem · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning
copula models
0.212014
Latent Variable Copula Inference for Bundle Pricing from Retail Transaction Data · ICML 2014
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent variable inference
0.212014
Latent Variable Copula Inference for Bundle Pricing from Retail Transaction Data · ICML 2014
Information retrieval › user behavior › search behavior
click model
0.112020
Fatigue-Aware Bandits for Dependent Click Models · AAAI 2020

Methods — techniques the papers use, named apart from their topics

column generation · 2.5prescriptive tree · 1.7mixed-integer programming · 1.2mixed integer programming · 1.2bandit algorithms · 1.2surrogate objective · 0.9observational data · 0.9causal inference · 0.9oblique projection · 0.8neural forecaster · 0.8multinomial logit model · 0.8end-to-end learning · 0.8piecewise-linear neural network · 0.7self-training · 0.6input consistency loss · 0.6model distillation · 0.5decision tree · 0.5regret analysis · 0.4
YearPublicationVenuePosition
2025 Causal LLM Routing: End-to-End Regret Minimization from Observational Data
abstract
LLM routing aims to select the most appropriate model for each query, balancing competing performance metrics such as accuracy and cost across a pool of language models. Prior approaches typically adopt a decoupled strategy, where the metrics are first predicted and the model is then selected based on these estimates. This setup is prone to compounding errors and often relies on full-feedback data, where each query is evaluated by all candidate models, which is costly to obtain and maintain in practice. In contrast, we learn from observational data, which records only the outcome of the model actually deployed. We propose a causal end-to-end framework that learns routing policies by minimizing decision-making regret from observational data. To enable efficient optimization, we introduce two theoretically grounded surrogate objectives: a classification-based upper bound, and a softmax-weighted regret approximation shown to recover the optimal policy at convergence. We further extend our framework to handle heterogeneous cost preferences via an interval-conditioned architecture. Experiments on public benchmarks show that our method outperforms existing baselines, achieving state-of-the-art performance across different embedding models.
Asterios Tsiourvas, Wei Sun 0031, Georgia Perakis
NeurIPS2
2024 Manifold-Aligned Counterfactual Explanations for Neural Networks
abstract
We study the problem of finding optimal manifold-aligned counterfactual explanations for neural networks. Existing approaches that involve solving a complex mixed-integer optimization (MIP) problem frequently suffer from scalability issues, limiting their practical usefulness. Furthermore, the solutions are not guaranteed to follow the data manifold, resulting in unrealistic counterfactual explanations. To address these challenges, we first present a MIP formulation where we explicitly enforce manifold alignment by reformulating the highly nonlinear Local Outlier Factor (LOF) metric as mixed-integer constraints. To address the computational challenge, we leverage the geometry of a trained neural network and propose an efficient decomposition scheme that reduces the initial large, hard-to-solve optimization problem into a series of significantly smaller, easier-to-solve problems by constraining the search space to “live” polytopes, i.e., regions that contain at least one actual data point. Experiments on real-world datasets demonstrate the efficacy of our approach in producing both optimal and realistic counterfactual explanations, and computational traceability.
Asterios Tsiourvas, Wei Sun 0031, Georgia Perakis
AISTATS2
2024 Learning Optimal Projection for Forecast Reconciliation of Hierarchical Time Series
abstract
Hierarchical time series forecasting requires not only prediction accuracy but also coherency, i.e., forecasts add up appropriately across the hierarchy. Recent literature has shown that reconciliation via projection outperforms prior methods such as top-down or bottom-up approaches. Unlike existing work that pre-specifies a projection matrix (e.g., orthogonal), we study the problem of learning the optimal oblique projection from data for coherent forecasting of hierarchical time series. In addition to the unbiasedness-preserving property, oblique projection implicitly accounts for the hierarchy structure and assigns different weights to individual time series, providing significant adaptability over orthogonal projection which treats base forecast errors equally. We examine two broad classes of projections, namely Euclidean projection and general oblique projections. We propose to model the reconciliation step as a learnable, structured, projection layer in the neural forecaster architecture. The proposed approach allows for the efficient learning of the optimal projection in an end-to-end framework where both the neural forecaster and the projection layer are learned simultaneously. An empirical evaluation of real-world hierarchical time series datasets demonstrates the superior performance of the proposed method over existing state-of-the-art approaches.
Asterios Tsiourvas, Wei Sun 0031, Georgia Perakis, Yada Zhu
ICML2
2023 Scalable Optimal Multiway-Split Decision Trees with Constraints
abstract
There has been a surge of interest in learning optimal decision trees using mixed-integer programs (MIP) in recent years, as heuristic-based methods do not guarantee optimality and find it challenging to incorporate constraints that are critical for many practical applications. However, existing MIP methods that build on an arc-based formulation do not scale well as the number of binary variables is in the order of 2 to the power of the depth of the tree and the size of the dataset. Moreover, they can only handle sample-level constraints and linear metrics. In this paper, we propose a novel path-based MIP formulation where the number of decision variables is independent of dataset size. We present a scalable column generation framework to solve the MIP. Our framework produces a multiway-split tree which is more interpretable than the typical binary-split trees due to its shorter rules. Our framework is more general as it can handle nonlinear metrics such as F1 score, and incorporate a broader class of constraints. We demonstrate its efficacy with extensive experiments. We present results on datasets containing up to 1,008,372 samples while existing MIP-based decision tree models do not scale well on data beyond a few thousand points. We report superior or competitive results compared to the state-of-art MIP-based methods with up to a 24X reduction in runtime.
Shivaram Subramanian, Wei Sun 0031
AAAI2
2023 Learning Prescriptive ReLU Networks
abstract
We study the problem of learning optimal policy from a set of discrete treatment options using observational data. We propose a piecewise linear neural network model that can balance strong prescriptive performance and interpretability, which we refer to as the prescriptive ReLU network, or P-ReLU. We show analytically that this model (i) partitions the input space into disjoint polyhedra, where all instances that belong to the same partition receive the same treatment, and (ii) can be converted into an equivalent prescriptive tree with hyperplane splits for interpretability. We demonstrate the flexibility of the P-ReLU network as constraints can be easily incorporated with minor modifications to the architecture. Through experiments, we validate the superior prescriptive accuracy of P-ReLU against competing benchmarks. Lastly, we present examples of prescriptive trees extracted from trained P-ReLUs using a real-world dataset, for both the unconstrained and constrained scenarios.
Wei Sun 0031, Asterios Tsiourvas
ICML1
2022 Enhancing Counterfactual Classification Performance via Self-Training
abstract
Unlike traditional supervised learning, in many settings only partial feedback is available. We may only observe outcomes for the chosen actions, but not the counterfactual outcomes associated with other alternatives. Such settings encompass a wide variety of applications including pricing, online marketing and precision medicine. A key challenge is that observational data are influenced by historical policies deployed in the system, yielding a biased data distribution. We approach this task as a domain adaptation problem and propose a self-training algorithm which imputes outcomes with categorical values for finite unseen actions in the observational data to simulate a randomized trial through pseudolabelling, which we refer to as Counterfactual Self-Training (CST). CST iteratively imputes pseudolabels and retrains the model. In addition, we show input consistency loss can further improve CST performance which is shown in recent theoretical analysis of pseudolabelling. We demonstrate the effectiveness of the proposed algorithms on both synthetic and real datasets.
Ruijiang Gao, Max Biggs, Wei Sun 0031, Ligong Han
AAAI3
2022 Constrained Prescriptive Trees via Column Generation
abstract
With the abundance of available data, many enterprises seek to implement data-driven prescriptive analytics to help them make informed decisions. These prescriptive policies need to satisfy operational constraints, and proactively eliminate rule conflicts, both of which are ubiquitous in practice. It is also desirable for them to be simple and interpretable, so they can be easily verified and implemented. Existing approaches from the literature center around constructing variants of prescriptive decision trees to generate interpretable policies. However, none of the existing methods is able to handle constraints. In this paper, we propose a scalable method that solves the constrained prescriptive policy generation problem. We introduce a novel path-based mixed-integer program (MIP) formulation which identifies a (near) optimal policy efficiently via column generation. The policy generated can be represented as a multiway-split tree which is more interpretable and informative than binary-split trees due to its shorter rules. We demonstrate the efficacy of our method with extensive computational experiments on both synthetic and real datasets.
Shivaram Subramanian, Wei Sun 0031, Youssef Drissi, Markus Ettl
AAAI2
2021 Model Distillation for Revenue Optimization: Interpretable Personalized Pricing
abstract
Data-driven pricing strategies are becoming increasingly common, where customers are offered a personalized price based on features that are predictive of their valuation of a product. It is desirable for this pricing policy to be simple and interpretable, so it can be verified, checked for fairness, and easily implemented. However, efforts to incorporate machine learning into a pricing framework often lead to complex pricing policies that are not interpretable, resulting in slow adoption in practice. We present a novel, customized, prescriptive tree-based algorithm that distills knowledge from a complex black-box machine learning algorithm, segments customers with similar valuations and prescribes prices in such a way that maximizes revenue while maintaining interpretability. We quantify the regret of a resulting policy and demonstrate its efficacy in applications with both synthetic and real-world datasets.
Max Biggs, Wei Sun 0031, Markus Ettl
ICML2
2020 Fatigue-Aware Bandits for Dependent Click Models
Junyu Cao, Wei Sun 0031, Zuo-Jun Max Shen, Markus Ettl
AAAI2
2019 Dynamic Learning of Sequential Choice Bandit Problem under Marketing Fatigue
abstract
Motivated by the observation that overexposure to unwanted marketing activities leads to customer dissatisfaction, we consider a setting where a platform offers a sequence of messages to its users and is penalized when users abandon the platform due to marketing fatigue. We propose a novel sequential choice model to capture multiple interactions taking place between the platform and its user: Upon receiving a message, a user decides on one of the three actions: accept the message, skip and receive the next message, or abandon the platform. Based on user feedback, the platform dynamically learns users’ abandonment distribution and their valuations of messages to determine the length of the sequence and the order of the messages, while maximizing the cumulative payoff over a horizon of length T. We refer to this online learning task as the sequential choice bandit problem. For the offline combinatorial optimization problem, we show a polynomialtime algorithm. For the online problem, we propose an algorithm that balances exploration and exploitation, and characterize its regret bound. Lastly, we demonstrate how to extend the model with user contexts to incorporate personalization.
Junyu Cao, Wei Sun 0031
AAAI2
2019 Dynamic Learning with Frequent New Product Launches: A Sequential Multinomial Logit Bandit Problem
abstract
Motivated by the phenomenon that companies introduce new products to keep abreast with customers’ rapidly changing tastes, we consider a novel online learning setting where a profit-maximizing seller needs to learn customers’ preferences through offering recommendations, which may contain existing products and new products that are launched in the middle of a selling period. We propose a sequential multinomial logit (SMNL) model to characterize customers’ behavior when product recommendations are presented in tiers. For the offline version with known customers’ preferences, we propose a polynomial-time algorithm and characterize the properties of the optimal tiered product recommendation. For the online problem, we propose a learning algorithm and quantify its regret bound. Moreover, we extend the setting to incorporate a constraint which ensures every new product is learned to a given accuracy. Our results demonstrate the tier structure can be used to mitigate the risks associated with learning new products.
Junyu Cao, Wei Sun 0031
ICML2
2014 Latent Variable Copula Inference for Bundle Pricing from Retail Transaction Data
abstract
Bundle discounts are used by retailers in many industries. Optimal bundle pricing requires learning the joint distribution of consumer valuations for the items in the bundle, that is, how much they are willing to pay for each of the items. We suppose that a retailer has sales transaction data, and the corresponding consumer valuations are latent variables. We develop a statistically consistent and computationally tractable inference procedure for fitting a copula model over correlated valuations, using only sales transaction data for the individual items. Simulations and data experiments demonstrate consistency, scalability, and the importance of incorporating correlations in the joint distribution.
Benjamin Letham, Wei Sun 0031, Anshul Sheopuri
ICML2