VLDB 2026 Research / reviewers in the wild / expert
Dong Qian
dblp:82/7696
· DBLP profile ↗
5ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 69% Probabilistic and Bayesian machine learning · 13% Representation and self-supervised learning · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Computational science and engineering · 59% Medical and health informatics · 41% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
variational autoencoder |
1.0 | 2 | 2023 | Learning Hierarchical Variational Autoencoders With Mutual Information Maximization for Autoregressive Sequence Modeling · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Enhancing Variational Autoencoders with Mutual Information Neural Estimation for Text Generation · EMNLP/IJCNLP (1) 2019 |
Machine learning › Generative modeling › diffusion model
crystal structure generation |
0.9 | 1 | 2025 | WyckoffDiff - A Generative Diffusion Model for Crystal Symmetry · ICML 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | WyckoffDiff - A Generative Diffusion Model for Crystal Symmetry · ICML 2025 |
Computational science and engineering
materials science |
0.9 | 1 | 2025 | WyckoffDiff - A Generative Diffusion Model for Crystal Symmetry · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization |
0.8 | 1 | 2024 | PATNet: Propensity-Adjusted Temporal Network for Joint Imputation and Prediction Using Binary EHRs With Observation Bias · IEEE Trans. Knowl. Data Eng. 2024 |
Computational science and engineering › statistical computing
data imputation |
0.8 | 1 | 2024 | PATNet: Propensity-Adjusted Temporal Network for Joint Imputation and Prediction Using Binary EHRs With Observation Bias · IEEE Trans. Knowl. Data Eng. 2024 |
Medical and health informatics
electronic health records |
0.8 | 1 | 2024 | PATNet: Propensity-Adjusted Temporal Network for Joint Imputation and Prediction Using Binary EHRs With Observation Bias · IEEE Trans. Knowl. Data Eng. 2024 |
Machine learning › Generative modeling › autoregressive model
autoregressive sequence modeling |
0.7 | 1 | 2023 | Learning Hierarchical Variational Autoencoders With Mutual Information Maximization for Autoregressive Sequence Modeling · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Generative modeling › variational autoencoder
hierarchical VAE |
0.7 | 1 | 2023 | Learning Hierarchical Variational Autoencoders With Mutual Information Maximization for Autoregressive Sequence Modeling · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Representation and self-supervised learning › mutual information
mutual information estimation |
0.4 | 1 | 2019 | Enhancing Variational Autoencoders with Mutual Information Neural Estimation for Text Generation · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Language models and text generation
text generation |
0.4 | 1 | 2019 | Enhancing Variational Autoencoders with Mutual Information Neural Estimation for Text Generation · EMNLP/IJCNLP (1) 2019 |
Medical and health informatics › clinical data analysis › phenotyping
computational phenotyping |
0.4 | 1 | 2019 | Learning Phenotypes and Dynamic Patient Representations via RNN Regularized Collective Non-Negative Tensor Factorization · AAAI 2019 |
Machine learning › Representation and self-supervised learning › representation learning
latent representation learning |
0.2 | 1 | 2023 | Learning Hierarchical Variational Autoencoders With Mutual Information Maximization for Autoregressive Sequence Modeling · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Representation and self-supervised learning
mutual information |
0.1 | 1 | 2019 | Enhancing Variational Autoencoders with Mutual Information Neural Estimation for Text Generation · EMNLP/IJCNLP (1) 2019 |
Methods — techniques the papers use, named apart from their topics
wyckoff representation · 1.7neural network architecture · 1.7discrete diffusion · 1.7temporal network · 1.5propensity score · 1.5PARAFAC2 · 1.5neural network estimation · 0.7mutual information maximization · 0.7autoregressive decoder · 0.7non-negative tensor factorization · 0.4mutual information neural estimation · 0.4RNN regularization · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | WyckoffDiff - A Generative Diffusion Model for Crystal SymmetryabstractCrystalline materials often exhibit a high level of symmetry. However, most generative models do not account for symmetry, but rather model each atom without any constraints on its position or element. We propose a generative model, Wyckoff Diffusion (WyckoffDiff), which generates symmetry-based descriptions of crystals. This is enabled by considering a crystal structure representation that encodes all symmetry, and we design a novel neural network architecture which enables using this representation inside a discrete generative model framework. In addition to respecting symmetry by construction, the discrete nature of our model enables fast generation. We additionally present a new metric, Fréchet Wrenformer Distance, which captures the symmetry aspects of the materials generated, and we benchmark WyckoffDiff against recently proposed generative models for crystal generation. As a proof-of-concept study, we use WyckoffDiff to find new materials below the convex hull of thermodynamical stability. Filip Ekström Kelvinius, Oskar B. Andersson, Abhijith S. Parackal, Dong Qian, Rickard Armiento, Fredrik Lindsten |
ICML | 4 |
| 2024 | PATNet: Propensity-Adjusted Temporal Network for Joint Imputation and Prediction Using Binary EHRs With Observation BiasabstractPredictive analysis of electronic health records (EHR) is a fundamental task that could provide actionable insights to help clinicians improve the efficiency and quality of care. EHR are commonly recorded in binary format and contain inevitable missing data. The nature of missingness may vary by patients, clinical features, and time, which incurs observation bias. It is essential to account for the binary missingness and observation bias or the predictive performance could be substantially compromised. In this paper, we develop a propensity-adjusted temporal network (PATNet) to conduct data imputation and predictive analysis simultaneously. PATNet contains three subnetworks: 1) an imputation subnetwork that generates the initial imputation based on historical observations, 2) a propensity subnetwork that infers the patient-, feature-, and time-dependent propensity scores, and 3) a prediction subnetwork that produces the missing-informative prediction using the propensity-adjusted imputations and the missing probabilities. To allow the propensity scores to be inferred from data, we use the expectation-maximization (EM) algorithm to learn the imputation and propensity subnetworks and incorporate a low-rank constraint via PARAFAC2 approximation. Extensive evaluation using the MIMIC-III and eICU datasets demonstrates that PATNet outperforms the state-of-the-art methods in terms of binary data imputation, disease progression modeling, and mortality prediction tasks. Kejing Yin, Dong Qian, William Kwok-Wai Cheung |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Learning Hierarchical Variational Autoencoders With Mutual Information Maximization for Autoregressive Sequence ModelingabstractVariational autoencoders (VAEs) are a class of effective deep generative models, with the objective to approximate the true, but unknown data distribution. VAEs make use of latent variables to capture high-level semantics so as to reconstruct the data well with the help of informative latent variables. Yet, training VAEs tends to suffer from posterior collapse, when the decoder is parameterized by an autoregressive model for sequence generation. VAEs can be further enhanced by introducing multiple layers of latent variables, but the posterior collapse issue hinders the adoption of such hierarchical VAEs in real-world applications. In this paper, we introduce InfoMaxHVAE, which integrates mutual information estimated via neural networks into hierarchical VAEs to alleviate posterior collapse, when powerful autoregressive models are used for modeling sequences. Experimental results on a number of text and image datasets show that InfoMaxHVAE can outperform the state-of-the-art baselines and exhibits less posterior collapse. We further show that InfoMaxHVAE can shape a coarse-to-fine hierarchical organization of the latent space. Dong Qian, William Kwok-Wai Cheung |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Learning Phenotypes and Dynamic Patient Representations via RNN Regularized Collective Non-Negative Tensor FactorizationabstractNon-negative Tensor Factorization (NTF) has been shown effective to discover clinically relevant and interpretable phenotypes from Electronic Health Records (EHR). Existing NTF based computational phenotyping models aggregate data over the observation window, resulting in the learned phenotypes being mixtures of disease states appearing at different times. We argue that by separating the clinical events happening at different times in the input tensor, the temporal dynamics and the disease progression within the observation window could be modeled and the learned phenotypes will correspond to more specific disease states. Yet how to construct the tensor for data samples with different temporal lengths and properly capture the temporal relationship specific to each individual data sample remains an open challenge. In this paper, we propose a novel Collective Non-negative Tensor Factorization (CNTF) model where each patient is represented by a temporal tensor, and all of the temporal tensors are factorized collectively with the phenotype definitions being shared across all patients. The proposed CNTF model is also flexible to incorporate non-temporal data modality and RNN-based temporal regularization. We validate the proposed model using MIMIC-III dataset, and the empirical results show that the learned phenotypes are clinically interpretable. Moreover, the proposed CNTF model outperforms the state-of-the-art computational phenotyping models for the mortality prediction task. Kejing Yin, Dong Qian, William Kwok-Wai Cheung, Benjamin C. M. Fung, Jonathan Poon |
AAAI | 2 |
| 2019 | Enhancing Variational Autoencoders with Mutual Information Neural Estimation for Text GenerationabstractDong Qian, William K. Cheung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Dong Qian, William Kwok-Wai Cheung |
EMNLP/IJCNLP (1) | 1 |