VLDB 2026 Research / reviewers in the wild / expert
Xiaocheng Li
dblp:171/2155
· DBLP profile ↗
19ranked-venue papers
3as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reward Modeling with Ordinal Feedback: Wisdom of the CrowdabstractThe canonical setup of learning a reward model (RM) from human preferences with binary feedback discards potentially useful samples (such as "tied" between the two responses) and loses fine-grained information (such as "slightly better’"). This paper proposes a framework for learning RMs under ordinal feedback, generalizing the binary feedback to arbitrary granularity. We first identify a marginal unbiasedness condition, which generalizes the existing assumption of the binary feedback. The condition is validated via the sociological concept called "wisdom of the crowd". Under this condition, we develop a natural probability model and prove the benefits of fine-grained feedback in terms of reducing the Rademacher complexity, which may be of independent interest to another problem: the bias-variance trade-off in knowledge distillation. The framework also sheds light on designing guidelines for human annotators. Our numerical experiments validate that: (1) fine-grained feedback leads to better RM learning for both in- and out-of-distribution settings; (2) incorporating a certain proportion of tied samples boosts RM learning. Guanting Chen 0001, Xiaocheng Li |
ICML | 4 |
| 2025 | Out-of-distribution Robust OptimizationabstractIn this paper, we consider the contextual robust optimization problem under an out-of-distribution setting. The contextual robust optimization problem considers a risk-sensitive objective function for an optimization problem with the presence of a context vector (also known as covariates or side information) capturing related information. While the existing works mainly consider the in-distribution setting, and the resultant robustness achieved is in an out-of-sample sense, our paper studies an out-of-distribution setting where there can be a difference between the test environment and the training environment where the data are collected. We propose methods that handle this out-of-distribution setting, and the key relies on a density ratio estimation for the distribution shift. We show that additional structures such as covariate shift and label shift are not only helpful in defending distribution shift but also necessary in avoiding non-trivial solutions compared to other principled methods such as distributionally robust optimization. We also illustrate how the covariates can be useful in this procedure. Numerical experiments generate more intuitions and demonstrate that the proposed methods can help avoid over-conservative solutions. Zhongze Cai, Hansheng Jiang, Xiaocheng Li |
UAI | 3 |
| 2025 | Collaborative Prediction: To Join or To Disjoin DatasetsabstractWith the recent rise of generative Artificial Intelligence (AI), the need of selecting high-quality dataset to improve machine learning models has garnered increasing attention. However, some part of this topic remains underexplored, even for simple prediction models. In this work, we study the problem of developing practical algorithms that select appropriate dataset to minimize population loss of our prediction model with high probability. Broadly speaking, we investigate when datasets from different sources can be effectively merged to enhance the predictive model’s performance, and propose a practical algorithm with theoretical guarantees. By leveraging an oracle inequality and data-driven estimators, the algorithm reduces population loss with high probability. Numerical experiments demonstrate its effectiveness in both standard linear regression and broader deep learning applications. Kyung Rok Kim, Xiaocheng Li |
UAI | 3 |
| 2024 | When No-Rejection Learning is Consistent for Regression with RejectionabstractLearning with rejection has been a prototypical model for studying the human-AI interaction on prediction tasks. Upon the arrival of a sample instance, the model first uses a rejector to decide whether to accept and use the AI predictor to make a prediction or reject and defer the sample to humans. Learning such a model changes the structure of the original loss function and often results in undesirable non-convexity and inconsistency issues. For the classification with rejection problem, several works develop consistent surrogate losses for the joint learning of the predictor and the rejector, while there have been fewer works for the regression counterpart. This paper studies the regression with rejection (RwR) problem and investigates a no-rejection learning strategy that uses all the data to learn the predictor. We first establish the consistency for such a strategy under the weak realizability condition. Then for the case without the weak realizability, we show that the excessive risk can also be upper bounded with the sum of two parts: prediction error and calibration error. Lastly, we demonstrate the advantage of such a proposed learning strategy with empirical evidence. Xiaocheng Li, Chunlin Sun, Hanzhao Wang |
AISTATS | 1 |
| 2024 | Learning to Make Adherence-aware AdviceabstractAs artificial intelligence (AI) systems play an increasingly prominent role in human decision-making, challenges surface in the realm of human-AI interactions. One challenge arises from the suboptimal AI policies due to the inadequate consideration of humans disregarding AI recommendations, as well as the need for AI to provide advice selectively when it is most pertinent. This paper presents a sequential decision-making model that (i) takes into account the human's adherence level (the probability that the human follows/rejects machine advice) and (ii) incorporates a defer option so that the machine can temporarily refrain from making advice. We provide learning algorithms that learn the optimal advice policy and make advice only at critical time stamps. Compared to problem-agnostic reinforcement learning algorithms, our specialized learning algorithms not only enjoy better theoretical convergence properties but also show strong empirical performance. Guanting Chen 0001, Xiaocheng Li, Chunlin Sun, Hanzhao Wang |
ICLR | 2 |
| 2023 | Maximum Optimality Margin: A Unified Approach for Contextual Linear Programming and Inverse Linear ProgrammingabstractIn this paper, we study the predict-then-optimize problem where the output of a machine learning prediction task is used as the input of some downstream optimization problem, say, the objective coefficient vector of a linear program. The problem is also known as predictive analytics or contextual linear programming. The existing approaches largely suffer from either (i) optimization intractability (a non-convex objective function)/statistical inefficiency (a suboptimal generalization bound) or (ii) requiring strong condition(s) such as no constraint or loss calibration. We develop a new approach to the problem called maximum optimality margin which designs the machine learning loss function by the optimality condition of the downstream optimization. The max-margin formulation enjoys both computational efficiency and good theoretical properties for the learning procedure. More importantly, our new approach only needs the observations of the optimal solution in the training data rather than the objective function, which makes it a new and natural approach to the inverse linear programming problem under both contextual and context-free settings; we also analyze the proposed method under both offline and online settings, and demonstrate its performance using numerical experiments. Chunlin Sun, Xiaocheng Li |
ICML | 3 |
| 2023 | Distribution-Free Model-Agnostic Regression Calibration via Nonparametric MethodsabstractIn this paper, we consider the uncertainty quantification problem for regression models. Specifically, we consider an individual calibration objective for characterizing the quantiles of the prediction model. While such an objective is well-motivated from downstream tasks such as newsvendor cost, the existing methods have been largely heuristic and lack of statistical guarantee in terms of individual calibration. We show via simple examples that the existing methods focusing on population-level calibration guarantees such as average calibration or sharpness can lead to harmful and unexpected results. We propose simple nonparametric calibration methods that are agnostic of the underlying prediction model and enjoy both computational efficiency and statistical consistency. Our approach enables a better understanding of the possibility of individual calibration, and we establish matching upper and lower bounds for the calibration error of our proposed methods. Technically, our analysis combines the nonparametric analysis with a covering number argument for parametric analysis, which advances the existing theoretical analyses in the literature of nonparametric density estimation and quantile bandit problems. Importantly, the nonparametric perspective sheds new theoretical insights into regression calibration in terms of the curse of dimensionality and reconciles the existing results on the impossibility of individual calibration. To our knowledge, we make the first effort to reach both individual calibration and finite-sample guarantee with minimal assumptions in terms of conformal prediction. Numerical experiments show the advantage of such a simple approach under various metrics, and also under covariates shift. We hope our work provides a simple benchmark and a starting point of theoretical ground for future research on regression calibration. Zhongze Cai, Xiaocheng Li |
NeurIPS | 3 |
| 2023 | Predict-then-Calibrate: A New Perspective of Robust Contextual LPabstractContextual optimization, also known as predict-then-optimize or prescriptive analytics, considers an optimization problem with the presence of covariates (context or side information). The goal is to learn a prediction model (from the training data) that predicts the objective function from the covariates, and then in the test phase, solve the optimization problem with the covariates but without the observation of the objective function. In this paper, we consider a risk-sensitive version of the problem and propose a generic algorithm design paradigm called predict-then-calibrate. The idea is to first develop a prediction model without concern for the downstream risk profile or robustness guarantee, and then utilize calibration (or recalibration) methods to quantify the uncertainty of the prediction. While the existing methods suffer from either a restricted choice of the prediction model or strong assumptions on the underlying data, we show the disentangling of the prediction model and the calibration/uncertainty quantification has several advantages. First, it imposes no restriction on the prediction model and thus fully unleashes the potential of off-the-shelf machine learning methods. Second, the derivation of the risk and robustness guarantee can be made independent of the choice of the prediction model through a data-splitting idea. Third, our paradigm of predict-then-calibrate applies to both (risk-sensitive) robust and (risk-neutral) distributionally robust optimization (DRO) formulations. Theoretically, it gives new generalization bounds for the contextual LP problem and sheds light on the existing results of DRO for contextual LP. Numerical experiments further reinforce the advantage of the predict-then-calibrate paradigm in that an improvement on either the prediction model or the calibration model will lead to a better final performance. Chunlin Sun, Linyu Liu, Xiaocheng Li |
NeurIPS | 3 |
| 2022 | Learning from Stochastically Revealed PreferenceabstractWe study the learning problem of revealed preference in a stochastic setting: a learner observes the utility-maximizing actions of a set of agents whose utility follows some unknown distribution, and the learner aims to infer the distribution through the observations of actions. The problem can be viewed as a single-constraint special case of the inverse linear optimization problem. Existing works all assume that all the agents share one common utility which can easily be violated under practical contexts. In this paper, we consider two settings for the underlying utility distribution: a Gaussian setting where the customer utility follows the von Mises-Fisher distribution, and a $\delta$-corruption setting where the customer utility distribution concentrates on one fixed vector with high probability and is arbitrarily corrupted otherwise. We devise Bayesian approaches for parameter estimation and develop theoretical guarantees for the recovery of the true parameter. We illustrate the algorithm performance through numerical experiments. John R. Birge, Xiaocheng Li, Chunlin Sun |
NeurIPS | 2 |
| 2022 | Non-stationary Bandits with KnapsacksabstractIn this paper, we study the problem of bandits with knapsacks (BwK) in a non-stationary environment. The BwK problem generalizes the multi-arm bandit (MAB) problem to model the resource consumption associated with playing each arm. At each time, the decision maker/player chooses to play an arm, and s/he will receive a reward and consume certain amount of resource from each of the multiple resource types. The objective is to maximize the cumulative reward over a finite horizon subject to some knapsack constraints on the resources. Existing works study the BwK problem under either a stochastic or adversarial environment. Our paper considers a non-stationary environment which continuously interpolates between these two extremes. We first show that the traditional notion of variation budget is insufficient to characterize the non-stationarity of the BwK problem for a sublinear regret due to the presence of the constraints, and then we propose a new notion of global non-stationarity measure. We employ both non-stationarity measures to derive upper and lower bounds for the problem. Our results are based on a primal-dual analysis of the underlying linear programs and highlight the interplay between the constraints and the non-stationarity. Finally, we also extend the non-stationarity measure to the problem of online convex optimization with constraints and obtain new regret bounds accordingly. Jiashuo Jiang, Xiaocheng Li |
NeurIPS | 3 |
| 2021 | When Do You Need Billions of Words of Pretraining Data?abstractYian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. Bowman |
ACL/IJCNLP (1) | 3 |
| 2021 | The Symmetry between Arms and Knapsacks: A Primal-Dual Approach for Bandits with KnapsacksabstractIn this paper, we study the bandits with knapsacks (BwK) problem and develop a primal-dual based algorithm that achieves a problem-dependent logarithmic regret bound. The BwK problem extends the multi-arm bandit (MAB) problem to model the resource consumption, and the existing BwK literature has been mainly focused on deriving asymptotically optimal distribution-free regret bounds. We first study the primal and dual linear programs underlying the BwK problem. From this primal-dual perspective, we discover symmetry between arms and knapsacks, and then propose a new notion of suboptimality measure for the BwK problem. The suboptimality measure highlights the important role of knapsacks in determining algorithm regret and inspires the design of our two-phase algorithm. In the first phase, the algorithm identifies the optimal arms and the binding knapsacks, and in the second phase, it exhausts the binding knapsacks via playing the optimal arms through an adaptive procedure. Our regret upper bound involves the proposed suboptimality measure and it has a logarithmic dependence on length of horizon $T$ and a polynomial dependence on $m$ (the numbers of arms) and $d$ (the number of knapsacks). To the best of our knowledge, this is the first problem-dependent logarithmic regret bound for solving the general BwK problem. Xiaocheng Li, Chunlin Sun, Yinyu Ye 0001 |
ICML | 1 |
| 2020 | Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)abstractOne reason pretraining on self-supervised linguistic tasks is effective is that it teaches models features that are helpful for language understanding.However, we want pretrained models to learn not only to represent linguistic features, but also to use those features preferentially during fine-turning.With this goal in mind, we introduce a new English-language diagnostic set called MSGS (the Mixed Signals Generalization Set), which consists of 20 ambiguous binary classification tasks that we use to test whether a pretrained model prefers linguistic or surface generalizations during finetuning.We pretrain RoBERTa models from scratch on quantities of data ranging from 1M to 1B words and compare their performance on MSGS to the publicly available RoBERTa BASE .We find that models can learn to represent linguistic features with little pretraining data, but require far more data to learn to prefer linguistic generalizations over surface ones.Eventually, with about 30B words of pretraining data, RoBERTa BASE does demonstrate a linguistic bias with some regularity.We conclude that while self-supervised pretraining is an effective way to learn helpful inductive biases, there is likely room to improve the rate at which models learn which features matter. Feature type Feature description Positive example Negative example SurfaceAbsolute position Is the first token of S "the"?The cat chased a mouse.A cat chased a mouse.Length Is S longer than n (e.g., 3) words?The cat chased a mouse.The cat meowed.Lexical content Does S contain "the"?That cat chased the mouse.That cat chased a mouse.Relative position Does "the" precede "a"?The cat chased a mouse.A cat chased the mouse.Orthography Does S appear in title case?The Cat Chased a Mouse.The cat chased a mouse. LinguisticMorphology Does S have an irregular past verb?The cats slept.The cats meow.Syn.category Does S have an adjective?Lincoln was tall.Lincoln was president.Syn.construction Is S the control construction?Sue is eager to sleep.Sue is likely to sleep.Syn.position Is the main verb in "ing" form?Cats who eat mice are purring.Cats who are eating mice purr. Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu, Samuel R. Bowman |
EMNLP (1) | 3 |
| 2020 | Simple and Fast Algorithm for Binary Integer and Online Linear ProgrammingabstractIn this paper, we develop a simple and fast online algorithm for solving a class of binary integer linear programs (LPs) arisen in the general resource allocation problem. The algorithm requires only one single pass through the input data and is free of doing any matrix inversion. It can be viewed as both an approximate algorithm for solving binary integer LPs and a fast algorithm for solving online LP problems. The algorithm is inspired by an equivalent form of the dual problem of the relaxed LP and it essentially performs (one-pass) projected stochastic subgradient descent in the dual space. We analyze the algorithm under two different models, stochastic input and random permutation, with minimal technical assumptions on the input data. The algorithm achieves $O\left(m \sqrt{n}\right)$ expected regret under the stochastic input model and $O\left((m+\log n)\sqrt{n}\right)$ expected regret under the random permutation model, and it achieves $O(m \sqrt{n})$ expected constraint violation under both models, where $n$ is the number of decision variables and $m$ is the number of constraints. In addition, we employ the notion of permutational Rademacher complexity and derive regret bounds for two earlier online LP algorithms for comparison. Both algorithms improve the regret bound with a factor of $\sqrt{m}$ by paying more computational cost. Furthermore, we demonstrate how to convert the possibly infeasible solution to a feasible one through a randomized procedure. Numerical experiments illustrate the general applicability and effectiveness of the algorithms. Xiaocheng Li, Chunlin Sun, Yinyu Ye 0001 |
NeurIPS | 1 |
| 2018 | Recurrent Autoregressive Networks for Online Multi-object TrackingabstractThe main challenge of online multi-object tracking is to reliably associate object trajectories with detections in each video frame based on their tracking history. In this work, we propose the Recurrent Autoregressive Network (RAN), a temporal generative modeling framework to characterize the appearance and motion dynamics of multiple objects over time. The RAN couples an external memory and an internal memory. The external memory explicitly stores previous inputs of each trajectory in a time window, while the internal memory learns to summarize long-term tracking history and associate detections by processing the external memory. We conduct experiments on the MOT 2015 and 2016 datasets to demonstrate the robustness of our tracking method in highly crowded and occluded scenes. Our method achieves top-ranked results on the two benchmarks. Kuan Fang, Xiaocheng Li, Silvio Savarese |
WACV | 3 |
| 2017 | Deep Gaussian Process for Crop Yield Prediction Based on Remote Sensing DataabstractAgricultural monitoring, especially in developing countries, can help prevent famine and support humanitarian efforts. A central challenge is yield estimation, i.e., predicting crop yields before harvest. We introduce a scalable, accurate, and inexpensive method to predict crop yields using publicly available remote sensing data. Our approach improves existing techniques in three ways. First, we forego hand-crafted features traditionally used in the remote sensing community and propose an approach based on modern representation learning ideas. We also introduce a novel dimensionality reduction technique that allows us to train a Convolutional Neural Network or Long-short Term Memory network and automatically learn useful features even when labeled training data are scarce. Finally, we incorporate a Gaussian Process component to explicitly model the spatio-temporal structure of the data and further improve accuracy. We evaluate our approach on county-level soybean yield prediction in the U.S. and show that it outperforms competing techniques. Jiaxuan You, Xiaocheng Li, Melvin Low, David B. Lobell, Stefano Ermon |
AAAI | 2 |
| 2017 | Cooperative Control of Heterogeneous Uncertain Dynamical Networks: An Adaptive Explicit Synchronization FrameworkabstractThis paper proposes an adaptive explicit synchronization framework to address the cooperative control for heterogeneous uncertain dynamical networks under switching communication topologies. The main contribution is to develop an adaptive explicit synchronization algorithm, in which the synchronization state can be completely tracked by each agent in real time rather than only be measured after the synchronization process of all agents is over. By introducing appropriate assumptions, a class of adaptive explicit synchronization protocols is designed by using a combination of the virtual leader's states, the neighboring agents' relative information, distributed feedback gain, and distributed average weighted parameters. It is proved in the sense of Lyapunov that, if the dwell time is larger than a positive threshold, the cooperative control problem for the closed-loop heterogeneous uncertain dynamical networks under switching of strongly-connected communication topologies can be solved by the proposed adaptive explicit synchronization algorithm. Furthermore, by assuming that the topology is frequently strongly-connected, it shows that intermittent adaptive explicit synchronization can be achieved with well-designed control parameters. Two examples are presented to demonstrate the effectiveness of the proposed theory. Bohui Wang, Langwen Zhang, Bin Zhang 0008, Xiaocheng Li |
IEEE Trans. Cybern. | 5 |
| 2017 | Global Cooperative Control Framework for Multiagent Systems Subject to Actuator Saturation With Industrial ApplicationsabstractThis paper proposes a global cooperative control framework to address leader-follower consensus of constraints subsystems of industrial plants, in which each subsystem is modeled as an agent and all the subsystems and networks of information flow construct a multiagent system. The focus of this paper is to solve the global leader-follower consensus for multiagent systems with input saturation via low-high gain feedback approach and parametric algebraic Riccati equation approach, in which the feedback gain design is distributed and decoupled from network topologies. By introducing appropriate assumptions, a class of low-high gain feedback protocol is designed based on the states of local neighbors to reach the global stability. It is proved in the sense of Lyapunov that, if the dwell time is larger than a positive threshold, the global leader-follower consensus for the closed-loop linear multiagent systems with input saturation under the derived topology containing a directed spanning tree can be achieved. The results are further extended to leader-follower consensus for nonlinear multiagent systems with the design of nonlinear low-high gain feedback protocol. As industrial applications of the proposed low-high gain scheduling approaches, the controller design of vibration in mechanical systems and satellite formation systems are revisited. Numerical simulations with cooperative control of industries subsystems show the effectiveness of the proposed approach. Bohui Wang, Bin Zhang 0008, Xiaocheng Li |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2016 | Leader-follower consensus for multi-agent systems with three-layer network framework and dynamic interaction jointly connected topology
Bohui Wang, Bin Zhang 0008, Xiaocheng Li |
Neurocomputing | 5 |