EDBT 2026 Demo / reviewers in the wild / expert
Doudou Zhou
dblp:205/8849
· DBLP profile ↗
12ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-0830-2287ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data mining · 66% Data integration and cleaning · 17% Machine learning and data management · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 57% Medical and health informatics · 38% Smart cities and intelligent transportation · 6% | |
| Artificial intelligence
2 papers |
Transfer learning and domain adaptation · 33% Learning theory · 33% Representation and self-supervised learning · 25% | |
| Theoretical computer science
1 paper |
Information theory · 50% Graph algorithms and graph theory · 50% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 100% |
Topics — the 15 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
anomaly detection |
0.9 | 1 | 2025 | Asymptotic Distribution-Free Change-Point Detection for Modern Data Based on a New Ranking Scheme · IEEE Trans. Inf. Theory 2025 |
Data mining › time series analysis
change point detection |
0.9 | 1 | 2025 | Asymptotic Distribution-Free Change-Point Detection for Modern Data Based on a New Ranking Scheme · IEEE Trans. Inf. Theory 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.7 | 1 | 2023 | Refining the Unseen: Self-supervised Two-stream Feature Extraction for Image Quality Assessment · ICDM 2023 |
Bioinformatics and computational biology › drug discovery › drug repositioning
drug-disease association prediction |
0.7 | 1 | 2023 | Multimodal representation learning for predicting molecule-disease relations · Bioinform. 2023 |
Bioinformatics and computational biology › drug discovery
drug repurposing |
0.7 | 1 | 2023 | Multimodal representation learning for predicting molecule-disease relations · Bioinform. 2023 |
Bioinformatics and computational biology › drug discovery
drug side effect prediction |
0.7 | 1 | 2023 | Multimodal representation learning for predicting molecule-disease relations · Bioinform. 2023 |
Medical and health informatics
electronic health records |
0.7 | 1 | 2023 | Multi-source Learning via Completion of Block-wise Overlapping Noisy Matrices · J. Mach. Learn. Res. 2023 |
Medical and health informatics
pharmacovigilance |
0.7 | 1 | 2023 | Multimodal representation learning for predicting molecule-disease relations · Bioinform. 2023 |
Machine learning and data management
matrix completion |
0.7 | 1 | 2023 | Multi-source Learning via Completion of Block-wise Overlapping Noisy Matrices · J. Mach. Learn. Res. 2023 |
Data integration and cleaning
multi-source data integration |
0.7 | 1 | 2023 | Multi-source Learning via Completion of Block-wise Overlapping Noisy Matrices · J. Mach. Learn. Res. 2023 |
Image and video coding
image quality assessment |
0.7 | 1 | 2023 | Refining the Unseen: Self-supervised Two-stream Feature Extraction for Image Quality Assessment · ICDM 2023 |
Graph algorithms and graph theory › network analysis
graph ranking |
0.7 | 1 | 2023 | A new ranking scheme for modern data and its application to two-sample hypothesis testing · COLT 2023 |
Information theory › statistical inference
nonparametric statistics |
0.7 | 1 | 2023 | A new ranking scheme for modern data and its application to two-sample hypothesis testing · COLT 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
distribution regression |
0.3 | 1 | 2025 | Wasserstein Transfer Learning · NeurIPS 2025 |
Smart cities and intelligent transportation › urban computing
urban analytics |
0.2 | 1 | 2023 | A new ranking scheme for modern data and its application to two-sample hypothesis testing · COLT 2023 |
Methods — techniques the papers use, named apart from their topics
word embeddings · 1.3two-stream network · 1.3similarity graph · 1.3pointwise mutual information · 1.3permutation test · 1.3matrix factorization · 1.3contrastive learning · 1.3attention mechanism · 1.3wasserstein space · 0.9graph-induced ranking · 0.9asymptotic distribution-free testing · 0.9asymptotic analysis · 0.9multimodal representation learning · 0.7electronic health records · 0.7deep neural network · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Wasserstein Transfer LearningabstractTransfer learning is a powerful paradigm for leveraging knowledge from source domains to enhance learning in a target domain. However, traditional transfer learning approaches often focus on scalar or multivariate data within Euclidean spaces, limiting their applicability to complex data structures such as probability distributions. To address this, we introduce a novel framework for transfer learning in regression models, where outputs are probability distributions residing in the Wasserstein space. When the informative subset of transferable source domains is known, we propose an estimator with provable asymptotic convergence rates, quantifying the impact of domain similarity on transfer efficiency. For cases where the informative subset is unknown, we develop a data-driven transfer learning procedure designed to mitigate negative transfer. The proposed methods are supported by rigorous theoretical analysis and are validated through extensive simulations and real-world applications. Sinian Zhang, Doudou Zhou, Yidong Zhou 0001 |
NeurIPS | 3 |
| 2025 | ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis
Ziming Gan, Doudou Zhou, Everett Neil Rush, Vidul Ayakulangara Panickan, Yuk-Lam Ho, George Ostrouchov, Shuting Shen, Xin Xiong 0006, Kimberly F. Greco, Chuan Hong, Clara-Lea Bonzel, Jun Wen 0001, Lauren Costa, Tianrun A. Cai, Edmon Begoli, Zongqi Xia, John Michael Gaziano, Katherine P. Liao, Kelly Cho, Tianxi Cai |
J. Biomed. Informatics | 2 |
| 2025 | DOME: Directional medical embedding vectors from Electronic Health RecordsabstractMOTIVATION: The increasing availability of Electronic Health Record (EHR) systems has created enormous potential for translational research. Recent developments in representation learning techniques have led to effective large-scale representations of EHR concepts along with knowledge graphs that empower downstream EHR studies. However, most existing methods require training with patient-level data, limiting their abilities to expand the training with multi-institutional EHR data. On the other hand, scalable approaches that only require summary-level data do not incorporate temporal dependencies between concepts. METHODS: We introduce a DirectiOnal Medical Embedding (DOME) algorithm to encode temporally directional relationships between medical concepts, using summary-level EHR data. Specifically, DOME first aggregates patient-level EHR data into an asymmetric co-occurrence matrix. Then it computes two Positive Pointwise Mutual Information (PPMI) matrices to correspondingly encode the pairwise prior and posterior dependencies between medical concepts. Following that, a joint matrix factorization is performed on the two PPMI matrices, which results in three vectors for each concept: a semantic embedding and two directional context embeddings. They collectively provide a comprehensive depiction of the temporal relationship between EHR concepts. RESULTS: We highlight the advantages and translational potential of DOME through three sets of validation studies. First, DOME consistently improves existing direction-agnostic embedding vectors for disease risk prediction in several diseases, for example achieving a relative gain of 5.5% in the area under the receiver operating characteristic (AUROC) for lung cancer. Second, DOME excels in directional drug-disease relationship inference by successfully differentiating between drug side effects and indications, correspondingly achieving relative AUROC gain over the state-of-the-art methods by 10.8% and 6.6%. Finally, DOME effectively constructs directional knowledge graphs, which distinguish disease risk factors from comorbidities, thereby revealing disease progression trajectories. The source codes are provided at https://github.com/celehs/Directional-EHR-embedding. Jun Wen 0001, Hao Xue 0005, Everett Neil Rush, Vidul Ayakulangara Panickan, Tianrun A. Cai, Doudou Zhou, Yuk-Lam Ho, Lauren Costa, Edmon Begoli, Chuan Hong, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Tianxi Cai |
J. Biomed. Informatics | 6 |
| 2025 | Asymptotic Distribution-Free Change-Point Detection for Modern Data Based on a New Ranking SchemeabstractChange-point detection (CPD) involves identifying distributional changes in a sequence of independent observations. Among nonparametric methods, rank-based methods are attractive due to their robustness and effectiveness, and have been extensively studied for univariate data. However, they are not well explored for high-dimensional or non-Euclidean data. This paper proposes a new method, Rank INduced by Graph Change-Point Detection (RING-CPD), which utilizes graph-induced ranks to handle high-dimensional and non-Euclidean data. The new method is asymptotically distribution-free under the null hypothesis, and an analytic p-value approximation is provided for easy type-I error control. Simulation studies show that RING-CPD effectively detects change points across a wide range of alternatives and is also robust to heavy-tailed distribution and outliers. The new method is illustrated by the detection of seizures in a functional connectivity network dataset, changes in digit images, and travel pattern changes in the New York City Taxi dataset. Doudou Zhou, Hao Chen 0064 |
IEEE Trans. Inf. Theory | 1 |
| 2023 | A new ranking scheme for modern data and its application to two-sample hypothesis testingabstractRank-based approaches are among the most popular nonparametric methods for univariate data in tackling statistical problems such as hypothesis testing due to their robustness and effectiveness. However, they are unsatisfactory for more complex data. In the era of big data, high-dimensional and non-Euclidean data, such as networks and images, are ubiquitous and pose challenges for statistical analysis. Existing multivariate ranks such as component-wise, spatial, and depth-based ranks do not apply to non-Euclidean data and have limited performance for high-dimensional data. Instead of dealing with the ranks of observations, we propose two types of ranks applicable to complex data based on a similarity graph constructed on observations: a graph-induced rank defined by the inductive nature of the graph and an overall rank defined by the weight of edges in the graph. To illustrate their utilization, both the new ranks are used to construct test statistics for the two-sample hypothesis testing, which converge to the $\chi_2^2$ distribution under the permutation null distribution and some mild conditions of the ranks, enabling an easy type-I error control. Simulation studies show that the new method exhibits good power under a wide range of alternatives compared to existing methods. The new test is illustrated on the New York City taxi data for comparing travel patterns in consecutive months and a brain network dataset comparing male and female subjects. Doudou Zhou |
COLT | 1 |
| 2023 | Refining the Unseen: Self-supervised Two-stream Feature Extraction for Image Quality AssessmentabstractThe inadequacy of labeled datasets for image quality assessment has led to the development and popularity of self-supervised approaches. However, most existing self-supervised methods primarily focus on content and fidelity features extracted with convolutional neural networks, overlooking the crucial importance of structural features in quality assessment. To address this problem, we present a novel self-supervised two-stream feature extraction and representation approach. In our approach, the first stream leverages a contrastive learning framework to extract image fidelity features, while the second stream emphasizes structural features by incorporating an attention mechanism. This innovative combination results in a comprehensive feature representation for quality assessment. Moreover, our proposed method facilitates transfer learning, allowing the pre-trained two-stream model in the source domain to be seamlessly applied to target domains for quality regression. This compatibility with transfer learning enhances the adaptability and generalization of the model. Extensive experiments are carried out on three synthetic distortion datasets to validate the effectiveness of our approach. The results demonstrate that our work not only competes with state-of-the-art self-supervised methods but also outperforms some supervised approaches. Yiwei Lou, Yanyuan Chen, Dexuan Xu, Doudou Zhou, Yongzhi Cao, Hanpin Wang, Yu Huang 0004 |
ICDM | 4 |
| 2023 | Multimodal representation learning for predicting molecule-disease relationsabstractMOTIVATION: Predicting molecule-disease indications and side effects is important for drug development and pharmacovigilance. Comprehensively mining molecule-molecule, molecule-disease and disease-disease semantic dependencies can potentially improve prediction performance. METHODS: We introduce a Multi-Modal REpresentation Mapping Approach to Predicting molecular-disease relations (M2REMAP) by incorporating clinical semantics learned from electronic health records (EHR) of 12.6 million patients. Specifically, M2REMAP first learns a multimodal molecule representation that synthesizes chemical property and clinical semantic information by mapping molecule chemicals via a deep neural network onto the clinical semantic embedding space shared by drugs, diseases and other common clinical concepts. To infer molecule-disease relations, M2REMAP combines multimodal molecule representation and disease semantic embedding to jointly infer indications and side effects. RESULTS: We extensively evaluate M2REMAP on molecule indications, side effects and interactions. Results show that incorporating EHR embeddings improves performance significantly, for example, attaining an improvement over the baseline models by 23.6% in PRC-AUC on indications and 23.9% on side effects. Further, M2REMAP overcomes the limitation of existing methods and effectively predicts drugs for novel diseases and emerging pathogens. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/celehs/M2REMAP, and prediction results are provided at https://shiny.parse-health.org/drugs-diseases-dev/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jun Wen 0001, Xiang Zhang 0012, Everett Neil Rush, Vidul Ayakulangara Panickan, Tianrun A. Cai, Doudou Zhou, Yuk-Lam Ho, Lauren Costa, Edmon Begoli, Chuan Hong, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Marinka Zitnik, Tianxi Cai |
Bioinform. | 7 |
| 2023 | Multi-source Learning via Completion of Block-wise Overlapping Noisy MatricesabstractElectronic healthcare records (EHR) provide a rich resource for healthcare research. An important problem for the efficient utilization of the EHR data is the representation of the EHR features, which include the unstructured clinical narratives and the structured codified data. Matrix factorization-based embeddings trained using the summary-level co-occurrence statistics of EHR data have provided a promising solution for feature representation while preserving patients' privacy. However, such methods do not work well with multi-source data when these sources have overlapping but non-identical features. To accommodate multi-sources learning, we propose a novel word embedding generative model. To obtain multi-source embeddings, we design an efficient Block-wise Overlapping Noisy Matrix Integration (BONMI) algorithm to aggregate the multi-source pointwise mutual information matrices optimally with a theoretical guarantee. Our algorithm can also be applied to other multi-source data integration problems with a similar data structure. A by-product of BONMI is the contribution to the field of matrix completion by considering the missing mechanism other than the entry-wise independent missing. We show that the entry-wise missing assumption, despite its prevalence in the works of matrix completion, is not necessary to guarantee recovery. We prove the statistical rate of our estimator, which is comparable to the rate under independent missingness. Simulation studies show that BONMI performs well under a variety of configurations. We further illustrate the utility of BONMI by integrating multi-lingual multi-source medical text and EHR data to perform two tasks: (i) co-training semantic embeddings for medical concepts in both English and Chinese and (ii) the translation between English and Chinese medical concepts. Our method shows an advantage over existing methods. Doudou Zhou, Tianxi Cai |
J. Mach. Learn. Res. | 1 |
| 2022 | Multiview Incomplete Knowledge Graph Integration with application to cross-institutional EHR data harmonizationabstractOBJECTIVE: The growing availability of electronic health records (EHR) data opens opportunities for integrative analysis of multi-institutional EHR to produce generalizable knowledge. A key barrier to such integrative analyses is the lack of semantic interoperability across different institutions due to coding differences. We propose a Multiview Incomplete Knowledge Graph Integration (MIKGI) algorithm to integrate information from multiple sources with partially overlapping EHR concept codes to enable translations between healthcare systems. METHODS: The MIKGI algorithm combines knowledge graph information from (i) embeddings trained from the co-occurrence patterns of medical codes within each EHR system and (ii) semantic embeddings of the textual strings of all medical codes obtained from the Self-Aligning Pretrained BERT (SAPBERT) algorithm. Due to the heterogeneity in the coding across healthcare systems, each EHR source provides partial coverage of the available codes. MIKGI synthesizes the incomplete knowledge graphs derived from these multi-source embeddings by minimizing a spherical loss function that combines the pairwise directional similarities of embeddings computed from all available sources. MIKGI outputs harmonized semantic embedding vectors for all EHR codes, which improves the quality of the embeddings and enables direct assessment of both similarity and relatedness between any pair of codes from multiple healthcare systems. RESULTS: With EHR co-occurrence data from Veteran Affairs (VA) healthcare and Mass General Brigham (MGB), MIKGI algorithm produces high quality embeddings for a variety of downstream tasks including detecting known similar or related entity pairs and mapping VA local codes to the relevant EHR codes used at MGB. Based on the cosine similarity of the MIKGI trained embeddings, the AUC was 0.918 for detecting similar entity pairs and 0.809 for detecting related pairs. For cross-institutional medical code mapping, the top 1 and top 5 accuracy were 91.0% and 97.5% when mapping medication codes at VA to RxNorm medication codes at MGB; 59.1% and 75.8% when mapping VA local laboratory codes to LOINC hierarchy. When trained with 500 labels, the lab code mapping attained top 1 and 5 accuracy at 77.7% and 87.9%. MIKGI also attained best performance in selecting VA local lab codes for desired laboratory tests and COVID-19 related features for COVID EHR studies. Compared to existing methods, MIKGI attained the most robust performance with accuracy the highest or near the highest across all tasks. CONCLUSIONS: The proposed MIKGI algorithm can effectively integrate incomplete summary data from biomedical text and EHR data to generate harmonized embeddings for EHR codes for knowledge graph modeling and cross-institutional translation of EHR codes. Doudou Zhou, Ziming Gan, Alina Patwari, Everett Neil Rush, Clara-Lea Bonzel, Vidul Ayakulangara Panickan, Chuan Hong, Yuk-Lam Ho, Tianrun A. Cai, Lauren Costa, Victor M. Castro, Shawn N. Murphy, Gabriel A. Brat, Griffin M. Weber, Paul Avillach, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Tianxi Cai |
J. Biomed. Informatics | 1 |
| 2020 | sureLDA: A multidisease automated phenotyping method for the electronic health recordabstractOBJECTIVE: A major bottleneck hindering utilization of electronic health record data for translational research is the lack of precise phenotype labels. Chart review as well as rule-based and supervised phenotyping approaches require laborious expert input, hampering applicability to studies that require many phenotypes to be defined and labeled de novo. Though International Classification of Diseases codes are often used as surrogates for true labels in this setting, these sometimes suffer from poor specificity. We propose a fully automated topic modeling algorithm to simultaneously annotate multiple phenotypes. MATERIALS AND METHODS: Surrogate-guided ensemble latent Dirichlet allocation (sureLDA) is a label-free multidimensional phenotyping method. It first uses the PheNorm algorithm to initialize probabilities based on 2 surrogate features for each target phenotype, and then leverages these probabilities to constrain the LDA topic model to generate phenotype-specific topics. Finally, it combines phenotype-feature counts with surrogates via clustering ensemble to yield final phenotype probabilities. RESULTS: sureLDA achieves reliably high accuracy and precision across a range of simulated and real-world phenotypes. Its performance is robust to phenotype prevalence and relative informativeness of surogate vs nonsurrogate features. It also exhibits powerful feature selection properties. DISCUSSION: sureLDA combines attractive properties of PheNorm and LDA to achieve high accuracy and precision robust to diverse phenotype characteristics. It offers particular improvement for phenotypes insufficiently captured by a few surrogate features. Moreover, sureLDA's feature selection ability enables it to handle high feature dimensions and produce interpretable computational phenotypes. CONCLUSIONS: sureLDA is well suited toward large-scale electronic health record phenotyping for highly multiphenotype applications such as phenome-wide association studies . Yuri Ahuja, Doudou Zhou, Zeling He, Jiehuan Sun, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, Chuan Hong, Tianxi Cai |
J. Am. Medical Informatics Assoc. | 2 |
| 2019 | sureLDA: A Novel Multi-Disease Automated Phenotyping Method for the Electronic Health Record
Yuri Ahuja, Doudou Zhou, Zeling He, Jiehuan Sun, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, Chuan Hong, Tianxi Cai |
AMIA | 2 |
| 2017 | Identify Biological Modules and Hub MiRNAs for Oral Squamous Cell CarcinomasabstractOral squamous cell carcinomas (OSCC) is the most common head and neck cancer worldwide, with more than 300,000 new cases being diagnosed annually. Studies have shown that miRNAs are involved in the process of growth, differentiation, apoptosis, invasion and metastasis of OSCC tumor cells. How miRNAs work together to contribute to this process is still largely unknown. The goal of our study was to characterize the coexpression network of miRNAs and to identify the miRNA subnetworks (modules) that were significantly associated with the OSCC cancer status. We also searched hub miRNAs that might play a vital role in the development of OSCC. We applied the weighted gene co-expression network analysis (WGCNA) to the miRNA expression profile data from a paired design study contributed by Shiah et al. To account for the within-pair correlation, a linear mixed model (LMM) was constructed to test the associations of miRNA modules to cancer status. Two significant modules (turquoise module with 254 miRNAs and grey module with 309 miRNAs) were identified. The miRNA miR-let-7c was the hub miRNA in the turquoise module in terms of node degree. Finally, we used miRsystem to perform the target gene prediction and KEGG pathway enrichment analysis of miRNAs within the two modules. Interestingly, the two modules have similar sets of target genes so that the top 6 enriched KEGG pathways for the 2 modules were the same. Compared with the probe-wise test used by Shiah et al., we took the network approach and identified significant OSCC-associated miRNA modules, which could help uncover the mechanism that miRNAs interplay each other to contribute to OSCC. Doudou Zhou, Jianqiang Li 0002, Qing Wang 0003, Weiliang Qiu, Shi Chen 0002, Minhua Lu |
COMPSAC (2) | 1 |