EDBT 2026 Demo / reviewers in the wild / expert
Xiao-Hua Zhou
dblp:60/2331
· DBLP profile ↗
16ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0001-7935-1222ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 9 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Denoising Entire Space Multi-Task Model for Post-Click Conversion Rate Prediction with Noisy Labels
Haoxuan Li 0001, Xiang Li 0067, Chunyuan Zheng 0001, Xiao-Hua Zhou |
SIGIR | 7 |
| 2025 | Temporal visiting-monitoring feature interaction learning for modelling structured electronic health records
Xiang Li 0112, Xiao-Hua Zhou |
Knowl. Based Syst. | 2 |
| 2024 | Relaxing the Accurate Imputation Assumption in Doubly Robust Learning for Debiased Collaborative FilteringabstractRecommender system aims to recommend items or information that may interest users based on their behaviors and preferences. However, there may be sampling selection bias in the data collection process, i.e., the collected data is not a representative of the target population. Many debiasing methods are developed based on pseudo-labelings. Nevertheless, the validity of these methods relies heavily on accurate pseudo-labelings (i.e., the imputed labels), which is difficult to satisfy in practice. In this paper, we theoretically propose several novel doubly robust estimators that are unbiased when either (a) the pseudo-labelings deviate from the true labels with an arbitrary user-specific inductive bias, item-specific inductive bias, or a combination of both, or (b) the learned propensities are accurate. We further propose a propensity reconstruction learning approach that adaptively updates the constraint weights using an attention mechanism and effectively controls the variance. Extensive experiments show that our approach outperforms the state-of-the-art on one semi-synthetic and three real-world datasets. Haoxuan Li 0001, Chunyuan Zheng 0001, Kunhan Wu, Hao Wang 0049, Peng Wu 0012, Zhi Geng, Xu Chen 0017, Xiao-Hua Zhou |
ICML | 9 |
| 2024 | Debiased Recommendation with Noisy FeedbackabstractRatings of a user to most items in recommender systems are usually missing not at random (MNAR), largely because users are free to choose which items to rate. To achieve unbiased learning of the prediction model under MNAR data, three typical solutions have been proposed, including error-imputation-based (EIB), inverse-propensity-scoring (IPS), and doubly robust (DR) methods. However, these methods ignore an alternative form of bias caused by the inconsistency between the observed ratings and the users' true preferences, also known as noisy feedback or outcome measurement errors (OME), e.g., due to public opinion or low-quality data collection process. In this work, we study intersectional threats to the unbiased learning of the prediction model from data MNAR and OME in the collected data. First, we design OME-EIB, OME-IPS, and OME-DR estimators, which largely extend the existing estimators to combat OME in real-world recommendation scenarios. Next, we theoretically prove the unbiasedness and generalization bound of the proposed estimators. We further propose an alternate denoising training approach to achieve unbiased learning of the prediction model under MNAR data with OME. Extensive experiments are conducted on three real-world datasets and one semi-synthetic dataset to show the effectiveness of our proposed approaches. The code is available at https://github.com/haoxuanli-pku/KDD24-OME-DR. Haoxuan Li 0001, Chunyuan Zheng 0001, Wenjie Wang 0007, Hao Wang 0049, Fuli Feng, Xiao-Hua Zhou |
KDD | 6 |
| 2024 | Integrating human learning and reinforcement learning: A novel approach to agent training
Yao-Hui Li, Qiang Hua, Xiao-Hua Zhou |
Knowl. Based Syst. | 4 |
| 2023 | Multiple Robust Learning for RecommendationabstractIn recommender systems, a common problem is the presence of various biases in the collected data, which deteriorates the generalization ability of the recommendation models and leads to inaccurate predictions. Doubly robust (DR) learning has been studied in many tasks in RS, with the advantage that unbiased learning can be achieved when either a single imputation or a single propensity model is accurate. In this paper, we propose a multiple robust (MR) estimator that can take the advantage of multiple candidate imputation and propensity models to achieve unbiasedness. Specifically, the MR estimator is unbiased when any of the imputation or propensity models, or a linear combination of these models is accurate. Theoretical analysis shows that the proposed MR is an enhanced version of DR when only having a single imputation and propensity model, and has a smaller bias. Inspired by the generalization error bound of MR, we further propose a novel multiple robust learning approach with stabilization. We conduct extensive experiments on real-world and semi-synthetic datasets, which demonstrates the superiority of the proposed approach over state-of-the-art methods. Haoxuan Li 0001, Quanyu Dai, Yuru Li, Zhenhua Dong, Xiao-Hua Zhou, Peng Wu 0012 |
AAAI | 6 |
| 2023 | ADRNet: A Generalized Collaborative Filtering Framework Combining Clinical and Non-Clinical Data for Adverse Drug Reaction PredictionabstractAdverse drug reaction (ADR) prediction plays a crucial role in both health care and drug discovery for reducing patient mortality and enhancing drug safety. Recently, many studies have been devoted to effectively predict the drug-ADRs incidence rates. However, these methods either did not effectively utilize non-clinical data, i.e., physical, chemical, and biological information about the drug, or did little to establish a link between content-based and pure collaborative filtering during the training phase. In this paper, we first formulate the prediction of multi-label ADRs as a drug-ADR collaborative filtering problem, and to the best of our knowledge, this is the first work to provide extensive benchmark results of previous collaborative filtering methods on two large publicly available clinical datasets. Then, by exploiting the easy accessible drug characteristics from non-clinical data, we propose ADRNet, a generalized collaborative filtering framework combining clinical and non-clinical data for drug-ADR prediction. Specifically, ADRNet has a shallow collaborative filtering module and a deep drug representation module, which can exploit the high-dimensional drug descriptors to further guide the learning of low-dimensional ADR latent embeddings, which incorporates both the benefits of collaborative filtering and representation learning. Extensive experiments are conducted on two publicly available real-world drug-ADR clinical datasets and two non-clinical datasets to demonstrate the accuracy and efficiency of the proposed ADRNet. The code is available at https://github.com/haoxuanli-pku/ADRnet. Haoxuan Li 0001, Taojun Hu, Zetong Xiong, Chunyuan Zheng 0001, Fuli Feng, Xiangnan He 0001, Xiao-Hua Zhou |
RecSys | 7 |
| 2022 | Convolutional Transformer Networks for Epileptic Seizure DetectionabstractEpilepsy is a chronic neurological disease that affects many people in the world. Automatic epileptic seizure detection based on electroencephalogram (EEG) signals is of great significance and has been widely studied. The current deep learning epilepsy detection algorithms are often designed to be relatively simple and seldom consider the characteristics of EEG signals. In this paper, we propose a promising epilepsy detection model based on convolutional transformer networks. We demonstrate that integrating convolution and transformer modules can achieve higher detection performance. Our convolutional transformer model is composed of two branches: one extracts time-domain features from multiple inputs of channel-exchanged EEG signals, and the other handle frequency-domain representations. Experiments on two EEG datasets show that our model offers state-of-the-art performance. Particularly on the CHB-MIT dataset, our model achieves 96.02% in average sensitivity and 97.94% in average specificity, outperforming other existing methods with clear margins. Nan Ke, Tong Lin 0002, Zhouchen Lin, Xiao-Hua Zhou, Taoyun Ji |
CIKM | 4 |
| 2022 | On the Opportunity of Causal Learning in Recommendation Systems: Foundation, Estimation, Prediction and ChallengesabstractRecently, recommender system (RS) based on causal inference has gained much attention in the industrial community, as well as the states of the art performance in many prediction and debiasing tasks. Nevertheless, a unified causal analysis framework has not been established yet. Many causal-based prediction and debiasing studies rarely discuss the causal interpretation of various biases and the rationality of the corresponding causal assumptions. In this paper, we first provide a formal causal analysis framework to survey and unify the existing causal-inspired recommendation methods, which can accommodate different scenarios in RS. Then we propose a new taxonomy and give formal causal definitions of various biases in RS from the perspective of violating the assumptions adopted in causal analysis. Finally, we formalize many debiasing and prediction tasks in RS, and summarize the statistical and machine learning-based causal estimation methods, expecting to provide new research opportunities and perspectives to the causal RS community. Peng Wu 0012, Haoxuan Li 0001, Quanyu Dai, Zhenhua Dong, Jie Sun 0007, Xiao-Hua Zhou |
IJCAI | 9 |
| 2022 | A Generalized Doubly Robust Learning Framework for Debiasing Post-Click Conversion Rate PredictionabstractPost-click conversion rate (CVR) prediction is an essential task for discovering user interests and increasing platform revenues in a range of industrial applications. One of the most challenging problems of this task is the existence of severe selection bias caused by the inherent self-selection behavior of users and the item selection process of systems. Currently, doubly robust (DR) learning approaches achieve the state-of-the-art performance for debiasing CVR prediction. However, in this paper, by theoretically analyzing the bias, variance and generalization bounds of DR methods, we find that existing DR approaches may have poor generalization caused by inaccurate estimation of propensity scores and imputation errors, which often occur in practice. Motivated by such analysis, we propose a generalized learning framework that not only unifies existing DR methods, but also provides a valuable opportunity to develop a series of new debiasing techniques to accommodate different application scenarios. Based on the framework, we propose two new DR methods, namely DR-BIAS and DR-MSE. DR-BIAS directly controls the bias of DR loss, while DR-MSE balances the bias and variance flexibly, which achieves better generalization performance. In addition, we propose a novel tri-level joint learning optimization method for DR-MSE in CVR prediction, and an efficient training algorithm correspondingly. We conduct extensive experiments on both real-world and semi-synthetic datasets, which validate the effectiveness of our proposed methods. Quanyu Dai, Haoxuan Li 0001, Peng Wu 0012, Zhenhua Dong, Xiao-Hua Zhou, Rui Zhang 0079, Rui Zhang 0003, Jie Sun 0007 |
KDD | 5 |
| 2021 | Prevent the Language Model from being Overconfident in Neural Machine TranslationabstractMengqi Miao, Fandong Meng, Yijin Liu, Xiao-Hua Zhou, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Mengqi Miao, Fandong Meng, Yijin Liu, Xiao-Hua Zhou, Jie Zhou 0016 |
ACL/IJCNLP (1) | 4 |
| 2020 | Chinese clinical named entity recognition with variant neural structures based on BERT methods
Xiao-Hua Zhou |
J. Biomed. Informatics | 3 |
| 2016 | Semiparametric Inference of the Complier Average Causal Effect with Nonignorable Missing OutcomesabstractNoncompliance and missing data often occur in randomized trials, which complicate the inference of causal effects. When both noncompliance and missing data are present, previous papers proposed moment and maximum likelihood estimators for binary and normally distributed continuous outcomes under the latent ignorable missing data mechanism. However, the latent ignorable missing data mechanism may be violated in practice, because the missing data mechanism may depend directly on the missing outcome itself. Under noncompliance and an outcome-dependent nonignorable missing data mechanism, previous studies showed the identifiability of complier average causal effect for discrete outcomes. In this article, we study the semiparametric identifiability and estimation of complier average causal effect in randomized clinical trials with both all-or-none noncompliance and outcome-dependent nonignorable missing continuous outcomes, and propose a two-step maximum likelihood estimator in order to eliminate the infinite dimensional nuisance parameter. Our method does not need to specify a parametric form for the missing data mechanism. We also evaluate the finite sample property of our method via extensive simulation studies and sensitivity analysis, with an application to a double-blinded psychiatric clinical trial. Zhi Geng, Xiao-Hua Zhou |
ACM Trans. Intell. Syst. Technol. | 4 |
| 1998 | Research Paper: Effects of Computer-based Prescribing on Pharmacist Work PatternsabstractOBJECTIVE: To measure the effect of computer-based outpatient prescription writing by internal medicine physicians on pharmacist work patterns. DESIGN: Work sampling at a hospital-based outpatient pharmacy. Data were collected from pharmacists wearing silent, random-signal generators before and after the implementation of computer-based prescribing. MEASUREMENTS: The type of work performed by pharmacists (activity), the reason for their work (function), and the people they contacted (contact) were measured. RESULTS: Total staff hours and prescriptions handled were similar before and after computer-based prescribing. Pharmacists recorded 4,687 observations before and 4,735 observations after implementation of computer-based outpatient prescription writing. After implementation, pharmacists spent 12.9 percent more time correcting prescription problems, had 3.9 percent less idle time, and spent 2.2 percent less time in discussions with others. Pharmacists also spent 34.0 percent less time filling prescriptions, 45.8 percent more time in problem-solving activities involving prescriptions, and 3.4 percent less time providing advice. Over 80 percent of pharmacist time was spent working alone both before and after computer-based outpatient prescription writing. CONCLUSION: Computer-based prescribing results in major changes in the type of work done by hospital-based outpatient pharmacists and in the reason for their work and small changes in the people contacted during their work. Michael D. Murray, Bonnie Loos, Wanzhu Tu, George J. Eckert, Xiao-Hua Zhou, William M. Tierney |
J. Am. Medical Informatics Assoc. | 5 |
| 1997 | Research Paper: A Randomized Trial of "Corollary Orders" to Prevent Errors of OmissionabstractOBJECTIVE: Errors of omission are a common cause of systems failures. Physicians often fail to order tests or treatments needed to monitor/ameliorate the effects of other tests or treatments. The authors hypothesized that automated, guideline-based reminders to physicians, provided as they wrote orders, could reduce these omissions. DESIGN: The study was performed on the inpatient general medicine ward of a public teaching hospital. Faculty and housestaff from the Indiana University School of Medicine, who used computer workstations to write orders, were randomized to intervention and control groups. As intervention physicians wrote orders for 1 of 87 selected tests or treatments, the computer suggested corollary orders needed to detect or ameliorate adverse reactions to the trigger orders. The physicians could accept or reject these suggestions. RESULTS: During the 6-month trial, reminders about corollary orders were presented to 48 intervention physicians and withheld from 41 control physicians. Intervention physicians ordered the suggested corollary orders in 46.3% of instances when they received a reminder, compared with 21.9% compliance by control physicians (p < 0.0001). Physicians discriminated in their acceptance of suggested orders, readily accepting some while rejecting others. There were one third fewer interventions initiated by pharmacists with physicians in the intervention than control groups. CONCLUSION: This study demonstrates that physician workstations, linked to a comprehensive electronic medical record, can be an efficient means for decreasing errors of omissions and improving adherence to practice guidelines. J. Marc Overhage, William M. Tierney, Xiao-Hua Zhou, Clement J. McDonald |
J. Am. Medical Informatics Assoc. | 3 |
| 1997 | Research Paper: Using Computer-based Medical Records to Predict Mortality Risk for Inner-city Patients with Reactive Airways DiseaseabstractObjective: To use routine data from a comprehensive electronic medical record system to predict death among patients with reactive airways disease. Design: Retrospective cohort study conducted in an academic primary care internal medicine practice. Subjects were 1,536 adults with reactive airways disease: 542 with asthma and 994 with chronic obstructive pulmonary disease (COPD). Measurements: The dependent variable was death from any cause within 3 years following patients' first primary care appointment in 1992. Multivariable logistic regression was used to identify independent predictors of 3-year mortality, with half of the patients used to derive the predictive model and the other half used to assess its predictability. Results: Of the 1,536 study patients, 191 (12%) died in the 3-year follow-up period. From information available on or before patients' first primary care visit in 1992, multivariable predictors of 3-year mortality were coincidental heart failure, male sex, presence of COPD, lower weight, low serum albumin concentration level, and a prior arterial PO2 of less than 60 mmHg; use of an inhaled corticosteroid was protective. The c-statistic (ROC curve area) in the validation cohort was 0.76, indicating good discrimination, and goodness of fit was excellent by Hosmer-Lemeshow chi-square (P > 0.5). Only 24% of the patients in the validation cohort were designated at high risk (estimated ≥15% 3-year mortality), but this group contained more than half of the deaths within 3 years for the entire cohort. Conclusions: Data generated during routine care and stored in a comprehensive electronic medical record can accurately predict mortality among patients with reactive airways disease. Such technology can be used by practices to control for severity of illness when assessing clinical practice and to identify high-risk patients for interventions to improve prognosis. William M. Tierney, Michael D. Murray, Denise L. Gaskins, Xiao-Hua Zhou |
J. Am. Medical Informatics Assoc. | 4 |