Dong Qian

dblp:82/7696 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 69% Probabilistic and Bayesian machine learning · 13% Representation and self-supervised learning · 12%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Computational science and engineering · 59% Medical and health informatics · 41%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
variational autoencoder
1.022023
Learning Hierarchical Variational Autoencoders With Mutual Information Maximization for Autoregressive Sequence Modeling · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Enhancing Variational Autoencoders with Mutual Information Neural Estimation for Text Generation · EMNLP/IJCNLP (1) 2019
Machine learning › Generative modeling › diffusion model
crystal structure generation
0.912025
WyckoffDiff - A Generative Diffusion Model for Crystal Symmetry · ICML 2025
Machine learning › Generative modeling
diffusion model
0.912025
WyckoffDiff - A Generative Diffusion Model for Crystal Symmetry · ICML 2025
Computational science and engineering
materials science
0.912025
WyckoffDiff - A Generative Diffusion Model for Crystal Symmetry · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization
0.812024
PATNet: Propensity-Adjusted Temporal Network for Joint Imputation and Prediction Using Binary EHRs With Observation Bias · IEEE Trans. Knowl. Data Eng. 2024
Computational science and engineering › statistical computing
data imputation
0.812024
PATNet: Propensity-Adjusted Temporal Network for Joint Imputation and Prediction Using Binary EHRs With Observation Bias · IEEE Trans. Knowl. Data Eng. 2024
Medical and health informatics
electronic health records
0.812024
PATNet: Propensity-Adjusted Temporal Network for Joint Imputation and Prediction Using Binary EHRs With Observation Bias · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Generative modeling › autoregressive model
autoregressive sequence modeling
0.712023
Learning Hierarchical Variational Autoencoders With Mutual Information Maximization for Autoregressive Sequence Modeling · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Generative modeling › variational autoencoder
hierarchical VAE
0.712023
Learning Hierarchical Variational Autoencoders With Mutual Information Maximization for Autoregressive Sequence Modeling · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Representation and self-supervised learning › mutual information
mutual information estimation
0.412019
Enhancing Variational Autoencoders with Mutual Information Neural Estimation for Text Generation · EMNLP/IJCNLP (1) 2019
Natural language and speech › Language models and text generation
text generation
0.412019
Enhancing Variational Autoencoders with Mutual Information Neural Estimation for Text Generation · EMNLP/IJCNLP (1) 2019
Medical and health informatics › clinical data analysis › phenotyping
computational phenotyping
0.412019
Learning Phenotypes and Dynamic Patient Representations via RNN Regularized Collective Non-Negative Tensor Factorization · AAAI 2019
Machine learning › Representation and self-supervised learning › representation learning
latent representation learning
0.212023
Learning Hierarchical Variational Autoencoders With Mutual Information Maximization for Autoregressive Sequence Modeling · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Representation and self-supervised learning
mutual information
0.112019
Enhancing Variational Autoencoders with Mutual Information Neural Estimation for Text Generation · EMNLP/IJCNLP (1) 2019

Methods — techniques the papers use, named apart from their topics

wyckoff representation · 1.7neural network architecture · 1.7discrete diffusion · 1.7temporal network · 1.5propensity score · 1.5PARAFAC2 · 1.5neural network estimation · 0.7mutual information maximization · 0.7autoregressive decoder · 0.7non-negative tensor factorization · 0.4mutual information neural estimation · 0.4RNN regularization · 0.4
YearPublicationVenuePosition
2025 WyckoffDiff - A Generative Diffusion Model for Crystal Symmetry
abstract
Crystalline materials often exhibit a high level of symmetry. However, most generative models do not account for symmetry, but rather model each atom without any constraints on its position or element. We propose a generative model, Wyckoff Diffusion (WyckoffDiff), which generates symmetry-based descriptions of crystals. This is enabled by considering a crystal structure representation that encodes all symmetry, and we design a novel neural network architecture which enables using this representation inside a discrete generative model framework. In addition to respecting symmetry by construction, the discrete nature of our model enables fast generation. We additionally present a new metric, Fréchet Wrenformer Distance, which captures the symmetry aspects of the materials generated, and we benchmark WyckoffDiff against recently proposed generative models for crystal generation. As a proof-of-concept study, we use WyckoffDiff to find new materials below the convex hull of thermodynamical stability.
Filip Ekström Kelvinius, Oskar B. Andersson, Abhijith S. Parackal, Dong Qian, Rickard Armiento, Fredrik Lindsten
ICML4
2024 PATNet: Propensity-Adjusted Temporal Network for Joint Imputation and Prediction Using Binary EHRs With Observation Bias
abstract
Predictive analysis of electronic health records (EHR) is a fundamental task that could provide actionable insights to help clinicians improve the efficiency and quality of care. EHR are commonly recorded in binary format and contain inevitable missing data. The nature of missingness may vary by patients, clinical features, and time, which incurs observation bias. It is essential to account for the binary missingness and observation bias or the predictive performance could be substantially compromised. In this paper, we develop a propensity-adjusted temporal network (PATNet) to conduct data imputation and predictive analysis simultaneously. PATNet contains three subnetworks: 1) an imputation subnetwork that generates the initial imputation based on historical observations, 2) a propensity subnetwork that infers the patient-, feature-, and time-dependent propensity scores, and 3) a prediction subnetwork that produces the missing-informative prediction using the propensity-adjusted imputations and the missing probabilities. To allow the propensity scores to be inferred from data, we use the expectation-maximization (EM) algorithm to learn the imputation and propensity subnetworks and incorporate a low-rank constraint via PARAFAC2 approximation. Extensive evaluation using the MIMIC-III and eICU datasets demonstrates that PATNet outperforms the state-of-the-art methods in terms of binary data imputation, disease progression modeling, and mortality prediction tasks.
Kejing Yin, Dong Qian, William Kwok-Wai Cheung
IEEE Trans. Knowl. Data Eng.2
2023 Learning Hierarchical Variational Autoencoders With Mutual Information Maximization for Autoregressive Sequence Modeling
abstract
Variational autoencoders (VAEs) are a class of effective deep generative models, with the objective to approximate the true, but unknown data distribution. VAEs make use of latent variables to capture high-level semantics so as to reconstruct the data well with the help of informative latent variables. Yet, training VAEs tends to suffer from posterior collapse, when the decoder is parameterized by an autoregressive model for sequence generation. VAEs can be further enhanced by introducing multiple layers of latent variables, but the posterior collapse issue hinders the adoption of such hierarchical VAEs in real-world applications. In this paper, we introduce InfoMaxHVAE, which integrates mutual information estimated via neural networks into hierarchical VAEs to alleviate posterior collapse, when powerful autoregressive models are used for modeling sequences. Experimental results on a number of text and image datasets show that InfoMaxHVAE can outperform the state-of-the-art baselines and exhibits less posterior collapse. We further show that InfoMaxHVAE can shape a coarse-to-fine hierarchical organization of the latent space.
Dong Qian, William Kwok-Wai Cheung
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Learning Phenotypes and Dynamic Patient Representations via RNN Regularized Collective Non-Negative Tensor Factorization
abstract
Non-negative Tensor Factorization (NTF) has been shown effective to discover clinically relevant and interpretable phenotypes from Electronic Health Records (EHR). Existing NTF based computational phenotyping models aggregate data over the observation window, resulting in the learned phenotypes being mixtures of disease states appearing at different times. We argue that by separating the clinical events happening at different times in the input tensor, the temporal dynamics and the disease progression within the observation window could be modeled and the learned phenotypes will correspond to more specific disease states. Yet how to construct the tensor for data samples with different temporal lengths and properly capture the temporal relationship specific to each individual data sample remains an open challenge. In this paper, we propose a novel Collective Non-negative Tensor Factorization (CNTF) model where each patient is represented by a temporal tensor, and all of the temporal tensors are factorized collectively with the phenotype definitions being shared across all patients. The proposed CNTF model is also flexible to incorporate non-temporal data modality and RNN-based temporal regularization. We validate the proposed model using MIMIC-III dataset, and the empirical results show that the learned phenotypes are clinically interpretable. Moreover, the proposed CNTF model outperforms the state-of-the-art computational phenotyping models for the mortality prediction task.
Kejing Yin, Dong Qian, William Kwok-Wai Cheung, Benjamin C. M. Fung, Jonathan Poon
AAAI2
2019 Enhancing Variational Autoencoders with Mutual Information Neural Estimation for Text Generation
abstract
Dong Qian, William K. Cheung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Dong Qian, William Kwok-Wai Cheung
EMNLP/IJCNLP (1)1