Hongmei Jiang

dblp:36/7453 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Sparse CCA-based mediation analysis with high-dimensional exposures and mediators
abstract
MOTIVATION: Mediation analysis plays a crucial role in understanding how exposure variables influence health outcomes via intermediate variables, or mediators, in environmental studies. When analysing a large number of environmental exposures, such as chemical mixtures or pollutants, together with multiple potential mediators such as metabolites, advanced methodologies are necessary to accurately separate direct and indirect effects. This paper proposes a novel mediation analysis method based on Sparse Canonical Correlation Analysis (SCCA), designed specifically for settings where both exposures and mediators are high-dimensional. The effectiveness of the proposed method is evaluated through simulation studies and an application to real-world data. RESULTS: The proposed SCCA-based mediation framework improved identification of relevant mediators and pathways in simulation studies, particularly in high-dimensional and noisy settings. The two-step screening extension further enhanced feature selection while maintaining stable estimation. In the real-data application, the method identified interpretable exposure-metabolite pathways associated with MELD score, with several pathways showing moderate selection stability and robustness to potential unmeasured confounding. AVAILABILITY: The R code for implementing the proposed method and the simulation studies is available at https://github.com/MaggieLi2001/HDM-SCCA2.
Maiying Kong, Matthew Ryan Smith, Yongliang Liang, Sami Teeny, Vilinh T. Ly, Young-Mi Go, Niharika Samala, Dean P. Jones, Jianzhu Luo, Walter H. Watson, Craig McClain, Vatsalya Vatsalya, Gyongyi Szabo, Srinivasan Dasarathy, Mack Mitchell, Laura E. Nagy, Bruce Barton, Matthew C. Cave, Hongmei Jiang
Bioinform.20
2025 PhyImpute and UniFracImpute: two imputation approaches incorporating phylogeny information for microbial count data
abstract
Sequencing-based microbial count data analysis is a challenging task due to the presence of numerous non-biological zeros, which can impede downstream analysis. To tackle this issue, we introduce two novel approaches, PhyImpute and UniFracImpute, which leverage similar microbial samples to identify and impute non-biological zeros in microbial count data. Our proposed methods utilize the probability of non-biological zeros and phylogenetic trees to estimate sample-to-sample similarity, thus addressing this challenge. To evaluate the performance of our proposed methods, we conduct experiments using both simulated and real microbial data. The results demonstrate that PhyImpute and UniFracImpute outperform existing methods in recovering the zeros and empowering downstream analyses such as differential abundance analysis, and disease status classification.
Qianwen Luo, Hamza Butt, Hongmei Jiang, Lingling An
Briefings Bioinform.5
2024 Deep learning with noisy labels in medical prediction problems: a scoping review
abstract
OBJECTIVES: Medical research faces substantial challenges from noisy labels attributed to factors like inter-expert variability and machine-extracted labels. Despite this, the adoption of label noise management remains limited, and label noise is largely ignored. To this end, there is a critical need to conduct a scoping review focusing on the problem space. This scoping review aims to comprehensively review label noise management in deep learning-based medical prediction problems, which includes label noise detection, label noise handling, and evaluation. Research involving label uncertainty is also included. METHODS: Our scoping review follows the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. We searched 4 databases, including PubMed, IEEE Xplore, Google Scholar, and Semantic Scholar. Our search terms include "noisy label AND medical/healthcare/clinical," "uncertainty AND medical/healthcare/clinical," and "noise AND medical/healthcare/clinical." RESULTS: A total of 60 papers met inclusion criteria between 2016 and 2023. A series of practical questions in medical research are investigated. These include the sources of label noise, the impact of label noise, the detection of label noise, label noise handling techniques, and their evaluation. Categorization of both label noise detection methods and handling techniques are provided. DISCUSSION: From a methodological perspective, we observe that the medical community has been up to date with the broader deep-learning community, given that most techniques have been evaluated on medical data. We recommend considering label noise as a standard element in medical research, even if it is not dedicated to handling noisy labels. Initial experiments can start with easy-to-implement methods, such as noise-robust loss functions, weighting, and curriculum learning.
Yishu Wei, Cong Sun 0004, Mingquan Lin, Hongmei Jiang, Yifan Peng 0002
J. Am. Medical Informatics Assoc.5
2024 Parallel Alternating Iterative Optimization for Cardiac Magnetic Resonance Image Blind Super-Resolution
abstract
Cardiac magnetic resonance imaging (CMRI) super-resolution (SR) reconstruction technology can enhance the resolution and quality of CMRI, providing experts with clearer and more accurate information about cardiac structure and function. This technology aids in the rapid and accurate diagnosis of cardiac abnormalities and the development of personalized treatment plans. In the processing of CMRI, existing bicubic degradation-based SR methods often suffer from performance degradation, resulting in blurred SR images. To address the aforementioned problem, we present a parallel alternating iterative optimization for CMRI image blind SR method (PAIBSR). Specifically, we propose a parallel alternating iterative optimization strategy, which employs dynamically corrected blur kernels and dynamically extracted intermediate low-resolution features as prior knowledge for both the blind SR process and the blur kernel correction process. Meanwhile, we propose a blur kernel update module composed of a blur kernel extractor and a low-resolution kernel extractor to correct the blur kernel. Furthermore, we propose an enhanced spatial feature transformation residual block, leveraging the corrected blur kernel as prior knowledge for the blind SR process. Through extensive experiments conducted on synthetic datasets, we have validated the superiority of PAIBSR method. It outperforms state-of-the-art SR methods in terms of performance and produces visually pleasing results.
Zhaoyang Song, Defu Qiu, Ruijuan Liu, Yongyong Hui, Hongmei Jiang
IEEE J. Biomed. Health Informatics6
2023 Attention hierarchical network for super-resolution
Zhaoyang Song, Yongyong Hui, Hongmei Jiang
Multim. Tools Appl.4
2021 NinimHMDA: neural integration of neighborhood information on a multiplex heterogeneous network for multiple types of human Microbe-Disease association
abstract
MOTIVATION: Many computational methods have been recently proposed to identify differentially abundant microbes related to a single disease; however, few studies have focused on large-scale microbe-disease association prediction using existing experimentally verified associations. This area has critical meanings. For example, it can help to rank and select potential candidate microbes for different diseases at-scale for downstream lab validation experiments and it utilizes existing evidence instead of the microbiome abundance data which usually costs money and time to generate. RESULTS: We construct a multiplex heterogeneous network (MHEN) using human microbe-disease association database, Disbiome and other prior biological databases, and define the large-scale human microbe-disease association prediction as link prediction problems on MHEN. We develop an end-to-end graph convolutional neural network-based mining model NinimHMDA which can not only integrate different prior biological knowledge but also predict different types of microbe-disease associations (e.g. a microbe may be reduced or elevated under the impact of a disease) using one-time model training. To the best of our knowledge, this is the first method that targets on predicting different association types between microbes and diseases. Results from large-scale cross validation and case studies show that our model is highly competitive compared to other commonly used approaches. AVAILABILITYAND IMPLEMENTATION: The codes are available at Github https://github.com/yuanjing-ma/NinimHMDA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yuanjing Ma, Hongmei Jiang
Bioinform.2
2021 Gradual deep residual network for super-resolution
Zhaoyang Song, Hongmei Jiang
Multim. Tools Appl.3
2021 Extended graphical lasso for multiple interaction networks for high dimensional omics data
abstract
There has been a spate of interest in association networks in biological and medical research, for example, genetic interaction networks. In this paper, we propose a novel method, the extended joint hub graphical lasso (EDOHA), to estimate multiple related interaction networks for high dimensional omics data across multiple distinct classes. To be specific, we construct a convex penalized log likelihood optimization problem and solve it with an alternating direction method of multipliers (ADMM) algorithm. The proposed method can also be adapted to estimate interaction networks for high dimensional compositional data such as microbial interaction networks. The performance of the proposed method in the simulated studies shows that EDOHA has remarkable advantages in recognizing class-specific hubs than the existing comparable methods. We also present three applications of real datasets. Biological interpretations of our results confirm those of previous studies and offer a more comprehensive understanding of the underlying mechanism in disease.
Yang Xu 0069, Hongmei Jiang, Wenxin Jiang 0003
PLoS Comput. Biol.2
2020 A novel normalization and differential abundance test framework for microbiome data
abstract
MOTIVATION: Microbial communities have been proved to have close relationship with many diseases. The identification of differentially abundant microbial species is clinically meaningful for finding disease-related pathogenic or probiotic bacteria. However, certain characteristics of microbiome data have hurdled the accuracy and effectiveness of differential abundance analysis. The abundances or counts of microbiome species are usually on different scales and exhibit zero-inflation and over-dispersion. Normalization is a crucial step before the differential abundance test. However, existing normalization methods typically try to adjust counts on different scales to a common scale by constructing size factors with the assumption that count distributions across samples are equivalent up to a certain percentile. These methods often yield undesirable results when differentially abundant species are of low to medium abundance level. For differential abundance analysis, existing methods often use a single distribution to model the dispersion of species which lacks flexibility to catch a single species' distinctiveness. These methods tend to detect a lot of false positives and often lack of power when the effect size is small. RESULTS: We develop a novel framework for differential abundance analysis on sparse high-dimensional marker gene microbiome data. Our methodology relies on a novel network-based normalization technique and a two-stage zero-inflated mixture count regression model (RioNorm2). Our normalization method aims to find a group of relatively invariant microbiome species across samples and conditions in order to construct the size factor. Another contribution of the paper is that our testing approach can take under-sampling and over-dispersion into consideration by separating microbiome species into two groups and model them separately. Through comprehensive simulation studies, the performance of our method is consistently powerful and robust across different settings with different sample size, library size and effect size. We also demonstrate the effectiveness of our novel framework using a published dataset of metastatic melanoma and find biological insights from the results. AVAILABILITY AND IMPLEMENTATION: The R package 'RioNorm2' can be installed from Github athttps://github.com/yuanjing-ma/RioNorm2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yuanjing Ma, Yuan Luo 0001, Hongmei Jiang
Bioinform.3
2018 A marginalized two-part Beta regression model for microbiome compositional data
abstract
In microbiome studies, an important goal is to detect differential abundance of microbes across clinical conditions and treatment options. However, the microbiome compositional data (quantified by relative abundance) are highly skewed, bounded in [0, 1), and often have many zeros. A two-part model is commonly used to separate zeros and positive values explicitly by two submodels: a logistic model for the probability of a specie being present in Part I, and a Beta regression model for the relative abundance conditional on the presence of the specie in Part II. However, the regression coefficients in Part II cannot provide a marginal (unconditional) interpretation of covariate effects on the microbial abundance, which is of great interest in many applications. In this paper, we propose a marginalized two-part Beta regression model which captures the zero-inflation and skewness of microbiome data and also allows investigators to examine covariate effects on the marginal (unconditional) mean. We demonstrate its practical performance using simulation studies and apply the model to a real metagenomic dataset on mouse skin microbiota. We find that under the proposed marginalized model, without loss in power, the likelihood ratio test performs better in controlling the type I error than those under conventional methods.
Haitao Chai, Hongmei Jiang, Lei Liu 0004
PLoS Comput. Biol.2
2015 Investigating microbial co-occurrence patterns based on metagenomic compositional data
abstract
MOTIVATION: The high-throughput sequencing technologies have provided a powerful tool to study the microbial organisms living in various environments. Characterizing microbial interactions can give us insights into how they live and work together as a community. Metagonomic data are usually summarized in a compositional fashion due to varying sampling/sequencing depths from one sample to another. We study the co-occurrence patterns of microbial organisms using their relative abundance information. Analyzing compositional data using conventional correlation methods has been shown prone to bias that leads to artifactual correlations. RESULTS: We propose a novel method, regularized estimation of the basis covariance based on compositional data (REBACCA), to identify significant co-occurrence patterns by finding sparse solutions to a system with a deficient rank. To be specific, we construct the system using log ratios of count or proportion data and solve the system using the l1-norm shrinkage method. Our comprehensive simulation studies show that REBACCA (i) achieves higher accuracy in general than the existing methods when a sparse condition is satisfied; (ii) controls the false positives at a pre-specified level, while other methods fail in various cases and (iii) runs considerably faster than the existing comparable method. REBACCA is also applied to several real metagenomic datasets. AVAILABILITY AND IMPLEMENTATION: The R codes for the proposed method are available at http://faculty.wcas.northwestern.edu/∼hji403/REBACCA.htm CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yuguang Ban, Lingling An, Hongmei Jiang
Bioinform.3
2015 A two-stage statistical procedure for feature selection and comparison in functional analysis of metagenomes
abstract
MOTIVATION: With the advance of new sequencing technologies producing massive short reads data, metagenomics is rapidly growing, especially in the fields of environmental biology and medical science. The metagenomic data are not only high dimensional with large number of features and limited number of samples but also complex with a large number of zeros and skewed distribution. Efficient computational and statistical tools are needed to deal with these unique characteristics of metagenomic sequencing data. In metagenomic studies, one main objective is to assess whether and how multiple microbial communities differ under various environmental conditions. RESULTS: We propose a two-stage statistical procedure for selecting informative features and identifying differentially abundant features between two or more groups of microbial communities. In the functional analysis of metagenomes, the features may refer to the pathways, subsystems, functional roles and so on. In the first stage of the proposed procedure, the informative features are selected using elastic net as reducing the dimension of metagenomic data. In the second stage, the differentially abundant features are detected using generalized linear models with a negative binomial distribution. Compared with other available methods, the proposed approach demonstrates better performance for most of the comprehensive simulation studies. The new method is also applied to two real metagenomic datasets related to human health. Our findings are consistent with those in previous reports. AVAILABILITY: R code and two example datasets are available at http://cals.arizona.edu/∼anling/software.htm. SUPPLEMENTARY INFORMATION: Supplementary file is available at Bioinformatics online.
Naruekamol Pookhao, Michael B. Sohn, Qike Li, Isaac Jenkins, Ruofei Du, Hongmei Jiang, Lingling An
Bioinform.6