VLDB 2026 Research / reviewers in the wild / expert
Xingye Qiao
dblp:21/10859
· DBLP profile ↗
13ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0003-0937-9822ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Trustworthy machine learning · 31% Learning theory · 22% Reinforcement learning · 19% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Medical and health informatics · 66% Computational social science and digital humanities · 34% |
Topics — the 22 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction |
1.6 | 2 | 2025 | Conformal Inference of Individual Treatment Effects Using Conditional Density Estimates · AAAI 2025 Efficient Online Set-valued Classification with Bandit Feedback · ICML 2024 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
1.4 | 3 | 2025 | Conformal Inference of Individual Treatment Effects Using Conditional Density Estimates · AAAI 2025 Learning Confidence Sets using Support Vector Machines · NeurIPS 2018 Efficient Online Set-valued Classification with Bandit Feedback · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.9 | 1 | 2025 | Conformal Inference of Individual Treatment Effects Using Conditional Density Estimates · AAAI 2025 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation › treatment effect estimation
individual treatment effect estimation |
0.9 | 1 | 2025 | Conformal Inference of Individual Treatment Effects Using Conditional Density Estimates · AAAI 2025 |
Machine learning › Learning theory › classification › nonparametric classification
nearest neighbor classification |
0.8 | 2 | 2020 | Statistical Guarantees of Distributed Nearest Neighbor Classification · NeurIPS 2020 Rates of Convergence for Large-scale Nearest Neighbor Classification · NeurIPS 2019 |
Machine learning › Reinforcement learning › bandit
bandit learning |
0.8 | 1 | 2024 | Efficient Online Set-valued Classification with Bandit Feedback · ICML 2024 |
Mathematical optimization › causal inference
instrumental variable estimation |
0.8 | 1 | 2024 | A Non-parametric Direct Learning Approach to Heterogeneous Treatment Effect Estimation under Unmeasured Confounding · NeurIPS 2024 |
Mathematical optimization
stochastic optimization |
0.8 | 1 | 2024 | A Non-parametric Direct Learning Approach to Heterogeneous Treatment Effect Estimation under Unmeasured Confounding · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection |
0.7 | 1 | 2023 | Set-valued Classification with Out-of-distribution Detection for Many Classes · J. Mach. Learn. Res. 2023 |
Machine learning › Transfer learning and domain adaptation
multi-source transfer learning |
0.4 | 1 | 2020 | Mutual Transfer Learning for Massive Data · ICML 2020 |
Machine learning › Transfer learning and domain adaptation › knowledge transfer
mutual transfer learning |
0.4 | 1 | 2020 | Mutual Transfer Learning for Massive Data · ICML 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
open-world learning |
0.4 | 1 | 2020 | Near-optimal Individualized Treatment Recommendations · J. Mach. Learn. Res. 2020 |
Medical and health informatics
precision medicine |
0.4 | 1 | 2020 | Near-optimal Individualized Treatment Recommendations · J. Mach. Learn. Res. 2020 |
Machine learning › Reinforcement learning
exploration |
0.4 | 1 | 2019 | Diverse Exploration via Conjugate Policies for Policy Gradient Methods · AAAI 2019 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.4 | 1 | 2019 | Diverse Exploration via Conjugate Policies for Policy Gradient Methods · AAAI 2019 |
Machine learning › Learning theory
classification |
0.3 | 1 | 2018 | Learning Confidence Sets using Support Vector Machines · NeurIPS 2018 |
Machine learning › Reinforcement learning › exploration
confidence sets |
0.3 | 1 | 2018 | Learning Confidence Sets using Support Vector Machines · NeurIPS 2018 |
Computational social science and digital humanities
causal inference |
0.2 | 1 | 2024 | A Non-parametric Direct Learning Approach to Heterogeneous Treatment Effect Estimation under Unmeasured Confounding · NeurIPS 2024 |
Machine learning › Reinforcement learning › dynamic programming › value iteration
approximate value iteration |
0.2 | 1 | 2015 | Improving Approximate Value Iteration with Complex Returns by Bounding · AAAI 2015 |
Machine learning › Learning theory › statistical learning theory
asymptotic analysis |
0.2 | 1 | 2015 | Flexible high-dimensional classification machines and their asymptotic properties · J. Mach. Learn. Res. 2015 |
Machine learning › Learning theory › classification
high-dimensional classification |
0.2 | 1 | 2015 | Flexible high-dimensional classification machines and their asymptotic properties · J. Mach. Learn. Res. 2015 |
Machine learning › Reinforcement learning › dynamic programming
value iteration |
0.2 | 1 | 2015 | Improving Approximate Value Iteration with Complex Returns by Bounding · AAAI 2015 |
Methods — techniques the papers use, named apart from their topics
non-parametric direct learning · 1.5instrumental variables · 1.5empirical risk minimization · 1.0conformal quantile regression · 0.9conditional density estimation · 0.9majority voting · 0.8unbiased estimation · 0.8stochastic gradient descent · 0.8kernel learning · 0.7kernel feature selection · 0.7parallel computation · 0.4generalization bound · 0.4consistency analysis · 0.4confidence distribution fusion · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Conformal Inference of Individual Treatment Effects Using Conditional Density EstimatesabstractIn an era where diverse and complex data are increasingly accessible, the precise prediction of individual treatment effects (ITE) becomes crucial across fields such as healthcare, economics, and social policy. Current state-of-the-art approaches, while providing valid prediction intervals through Conformal Quantile Regression (CQR) and related techniques, often yield overly conservative prediction intervals. In this work, we introduce a conformal inference approach to ITE using the conditional density of the outcome given the covariates. We leverage the reference distribution technique to efficiently estimate the conditional densities as the score functions under a two-stage conformal ITE framework. We show that our prediction intervals are not only marginally valid but are narrower than existing methods. Experimental results further validate the usefulness of our method. Baozhen Wang, Xingye Qiao |
AAAI | 2 |
| 2025 | Conformal Prediction Under Generalized Covariate Shift with Posterior DriftabstractIn many real applications of statistical learning, collecting sufficiently many training data is often expensive, time-consuming, or even unrealistic. In this case, a transfer learning approach, which aims to leverage knowledge from a related source domain to improve the learning performance in the target domain, is more beneficial. There have been many transfer learning methods developed under various distributional assumptions. In this article, we study a particular type of classification problem, called conformal prediction, under a new distributional assumption for transfer learning. Classifiers under the conformal prediction framework predict a set of plausible labels instead of one single label for each data instance, affording a more cautious and safer decision. We consider a generalization of the covariate shift with posterior drift setting for transfer learning. Under this setting, we propose a weighted conformal classifier that leverages both the source and target samples, with a coverage guarantee in the target domain. Theoretical studies demonstrate favorable asymptotic properties. Numerical studies further illustrate the usefulness of the proposed method. Baozhen Wang, Xingye Qiao |
AISTATS | 2 |
| 2024 | Efficient Online Set-valued Classification with Bandit FeedbackabstractConformal prediction is a distribution-free method that wraps a given machine learning model and returns a set of plausible labels that contain the true label with a prescribed coverage rate. In practice, the empirical coverage achieved highly relies on fully observed label information from data both in the training phase for model fitting and the calibration phase for quantile estimation. This dependency poses a challenge in the context of online learning with bandit feedback, where a learner only has access to the correctness of actions (i.e., pulled an arm) but not the full information of the true label. In particular, when the pulled arm is incorrect, the learner only knows that the pulled one is not the true class label, but does not know which label is true. Additionally, bandit feedback further results in a smaller labeled dataset for calibration, limited to instances with correct actions, thereby affecting the accuracy of quantile estimation. To address these limitations, we propose Bandit Class-specific Conformal Prediction (BCCP), offering coverage guarantees on a class-specific granularity. Using an unbiased estimation of an estimand involving the true label, BCCP trains the model and makes set-valued inferences through stochastic gradient descent. Our approach overcomes the challenges of sparsely labeled data in each iteration and generalizes the reliability and applicability of conformal prediction to online decision-making environments. Xingye Qiao |
ICML | 2 |
| 2024 | A Non-parametric Direct Learning Approach to Heterogeneous Treatment Effect Estimation under Unmeasured ConfoundingabstractIn many social, behavioral, and biomedical sciences, treatment effect estimation is a crucial step in understanding the impact of an intervention, policy, or treatment. In recent years, an increasing emphasis has been placed on heterogeneity in treatment effects, leading to the development of various methods for estimating Conditional Average Treatment Effects (CATE). These approaches hinge on a crucial identifying condition of no unmeasured confounding, an assumption that is not always guaranteed in observational studies or randomized control trials with non-compliance. In this paper, we proposed a general framework for estimating CATE with a possible unmeasured confounder using Instrumental Variables. We also construct estimators that exhibit greater efficiency and robustness against various scenarios of model misspecification. The efficacy of the proposed framework is demonstrated through simulation studies and a real data example. Xinhai Zhang, Xingye Qiao |
NeurIPS | 2 |
| 2023 | Set-valued Classification with Out-of-distribution Detection for Many ClassesabstractSet-valued classification, a new classification paradigm that aims to identify all the plausible classes that an observation belongs to, improves over the traditional classification paradigms in multiple aspects. Existing set-valued classification methods do not consider the possibility that the test set may contain out-of-distribution data, that is, the emergence of a new class that never appeared in the training data. Moreover, they are computationally expensive when the number of classes is large. We propose a Generalized Prediction Set (GPS) approach to set-valued classification while considering the possibility of a new class in the test data. The proposed classifier uses kernel learning and empirical risk minimization to encourage a small expected size of the prediction set while guaranteeing that the class-specific accuracy is at least some value specified by the user. For high-dimensional data, further improvement is obtained through kernel feature selection. Unlike previous methods, the proposed method achieves a good balance between accuracy, efficiency, and out-of-distribution detection rate. Moreover, our method can be applied in parallel to all the classes to alleviate the computational burden. Both theoretical analysis and numerical experiments are conducted to illustrate the effectiveness of the proposed method. Xingye Qiao |
J. Mach. Learn. Res. | 2 |
| 2020 | Mutual Transfer Learning for Massive DataabstractIn the transfer learning problem, the target and the source data domains are typically known. In this article, we study a new paradigm called mutual transfer learning where among many heterogeneous data domains, every data domain could potentially be the target of interest, and it could also be a useful source to help the learning in other data domains. However, it is important to note that given a target not every data domain can be a successful source; only data sets that are similar enough to be thought as from the same population can be useful sources for each other. Under this mutual learnability assumption, a confidence distribution fusion approach is proposed to recover the mutual learnability relation in the transfer learning regime. Our proposed method achieves the same oracle statistical inferential accuracy as if the true learnability structure were known. It can be implemented in an efficient parallel fashion to deal with large-scale data. Simulated and real examples are analyzed to illustrate the usefulness of the proposed method. Ching-Wei Cheng, Xingye Qiao, Guang Cheng 0003 |
ICML | 2 |
| 2020 | Statistical Guarantees of Distributed Nearest Neighbor ClassificationabstractNearest neighbor is a popular nonparametric method for classification and regression with many appealing properties. In the big data era, the sheer volume and spatial/temporal disparity of big data may prohibit centrally processing and storing the data. This has imposed considerable hurdle for nearest neighbor predictions since the entire training data must be memorized. One effective way to overcome this issue is the distributed learning framework. Through majority voting, the distributed nearest neighbor classifier achieves the same rate of convergence as its oracle version in terms of the regret, up to a multiplicative constant that depends solely on the data dimension. The multiplicative difference can be eliminated by replacing majority voting with the weighted voting scheme. In addition, we provide sharp theoretical upper bounds of the number of subsamples in order for the distributed nearest neighbor classifier to reach the optimal convergence rate. It is interesting to note that the weighted voting scheme allows a larger number of subsamples than the majority voting one. Our findings are supported by numerical studies. Jiexin Duan, Xingye Qiao, Guang Cheng 0003 |
NeurIPS | 2 |
| 2020 | Near-optimal Individualized Treatment RecommendationsabstractThe individualized treatment recommendation (ITR) is an important analytic framework for precision medicine. The goal of ITR is to assign the best treatments to patients based on their individual characteristics. From the machine learning perspective, the solution to the ITR problem can be formulated as a weighted classification problem to maximize the mean benefit from the recommended treatments given patients' characteristics. Several ITR methods have been proposed in both the binary setting and the multicategory setting. In practice, one may prefer a more flexible recommendation that includes multiple treatment options. This motivates us to develop methods to obtain a set of near-optimal individualized treatment recommendations alternative to each other, called alternative individualized treatment recommendations (A-ITR). We propose two methods to estimate the optimal A-ITR within the outcome weighted learning (OWL) framework. Simulation studies and a real data analysis for Type 2 diabetic patients with injectable antidiabetic treatments are conducted to show the usefulness of the proposed A-ITR framework. We also show the consistency of these methods and obtain an upper bound for the risk between the theoretically optimal recommendation and the estimated one. An R package "aitr" has been developed, found at https://github.com/menghaomiao/aitr. Haomiao Meng, Ying-Qi Zhao, Haoda Fu, Xingye Qiao |
J. Mach. Learn. Res. | 4 |
| 2019 | Diverse Exploration via Conjugate Policies for Policy Gradient MethodsabstractWe address the challenge of effective exploration while maintaining good performance in policy gradient methods. As a solution, we propose diverse exploration (DE) via conjugate policies. DE learns and deploys a set of conjugate policies which can be conveniently generated as a byproduct of conjugate gradient descent. We provide both theoretical and empirical results showing the effectiveness of DE at achieving exploration, improving policy performance, and the advantage of DE over exploration by random policy perturbations. Andrew Cohen, Xingye Qiao, Lei Yu 0001, Elliot Way, Xiangrong Tong |
AAAI | 2 |
| 2019 | Rates of Convergence for Large-scale Nearest Neighbor ClassificationabstractNearest neighbor is a popular class of classification methods with many desirable properties. For a large data set which cannot be loaded into the memory of a single machine due to computation, communication, privacy, or ownership limitations, we consider the divide and conquer scheme: the entire data set is divided into small subsamples, on which nearest neighbor predictions are made, and then a final decision is reached by aggregating the predictions on subsamples by majority voting. We name this method the big Nearest Neighbor (bigNN) classifier, and provide its rates of convergence under minimal assumptions, in terms of both the excess risk and the classification instability, which are proven to be the same rates as the oracle nearest neighbor classifier and cannot be improved. To significantly reduce the prediction time that is required for achieving the optimal rate, we also consider the pre-training acceleration technique applied to the bigNN method, with proven convergence rate. We find that in the distributed setting, the optimal choice of the neighbor k should scale with both the total sample size and the number of partitions, and there is a theoretical upper limit for the latter. Numerical studies have verified the theoretical findings. Xingye Qiao, Jiexin Duan, Guang Cheng 0003 |
NeurIPS | 1 |
| 2018 | Learning Confidence Sets using Support Vector MachinesabstractThe goal of confidence-set learning in the binary classification setting is to construct two sets, each with a specific probability guarantee to cover a class. An observation outside the overlap of the two sets is deemed to be from one of the two classes, while the overlap is an ambiguity region which could belong to either class. Instead of plug-in approaches, we propose a support vector classifier to construct confidence sets in a flexible manner. Theoretically, we show that the proposed learner can control the non-coverage rates and minimize the ambiguity with high probability. Efficient algorithms are developed and numerical studies illustrate the effectiveness of the proposed method. Xingye Qiao |
NeurIPS | 2 |
| 2015 | Improving Approximate Value Iteration with Complex Returns by Bounding
Robert William Wright, Xingye Qiao, Steven Loscalzo, Lei Yu 0001 |
AAAI | 2 |
| 2015 | Flexible high-dimensional classification machines and their asymptotic properties
Xingye Qiao, Lingsong Zhang |
J. Mach. Learn. Res. | 1 |