VLDB 2026 Research / reviewers in the wild / expert
Huaying Fang
dblp:168/5757
· DBLP profile ↗
6ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0002-7693-8260ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 70% Computational science and engineering · 30% | |
| Artificial intelligence
1 paper |
Probabilistic and Bayesian machine learning · 100% |
Topics — the 9 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › computational microbiology
microbiome analysis |
1.4 | 2 | 2024 | fastCCLasso: a fast and efficient algorithm for estimating correlation matrix from compositional data · Bioinform. 2024 gmcoda: Graphical model for multiple compositional vectors in microbiome studies · Bioinform. 2023 |
Computational science and engineering
compositional data analysis |
1.0 | 2 | 2024 | fastCCLasso: a fast and efficient algorithm for estimating correlation matrix from compositional data · Bioinform. 2024 CCLasso: correlation inference for compositional data through Lasso · Bioinform. 2015 |
Bioinformatics and computational biology › metagenomics
microbial interaction network inference |
1.0 | 2 | 2024 | fastCCLasso: a fast and efficient algorithm for estimating correlation matrix from compositional data · Bioinform. 2024 CCLasso: correlation inference for compositional data through Lasso · Bioinform. 2015 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.6 | 1 | 2022 | Estimating graphical models for count data with applications to single-cell gene network · NeurIPS 2022 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › gaussian graphical model
precision matrix estimation |
0.6 | 1 | 2022 | Estimating graphical models for count data with applications to single-cell gene network · NeurIPS 2022 |
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference |
0.6 | 1 | 2022 | Estimating graphical models for count data with applications to single-cell gene network · NeurIPS 2022 |
Computational science and engineering
graphical models |
0.6 | 1 | 2022 | Estimating graphical models for count data with applications to single-cell gene network · NeurIPS 2022 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis › gene co-expression network analysis
gene coexpression network construction |
0.3 | 1 | 2017 | VCNet: vector-based gene co-expression network construction and its application to RNA-seq data · Bioinform. 2017 |
Bioinformatics and computational biology
metagenomics |
0.2 | 1 | 2015 | CCLasso: correlation inference for compositional data through Lasso · Bioinform. 2015 |
Methods — techniques the papers use, named apart from their topics
poisson log-normal model · 1.1maximum marginal likelihood · 1.1d-trace loss · 1.1penalized weighted least squares · 0.8majorization-minimization · 0.7additive logistic normal distribution · 0.7frobenius norm hypothesis test · 0.3correlation matrix · 0.3augmented lagrangian · 0.2alternating direction method · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | fastCCLasso: a fast and efficient algorithm for estimating correlation matrix from compositional dataabstractMOTIVATION: The composition and structure of microbial communities on the body surface are closely related to human health. The interaction relationship among microbes can help us understand the formation of the microecological environment and the biological mechanism by which microorganisms influence host health. With the help of high-throughput sequencing technologies, microbial abundances in a natural environment can be directly measured without the isolation of microorganisms in culture. Sequencing experiments in microbiome studies can measure the relative abundance of microbes, which is called compositional data. Although there are already many methods for correlation analysis for compositional data, the computation time or accuracy still needs to be improved for current microbiome studies. RESULTS: We develop a fast and efficient algorithm, called fastCCLasso, based on a penalized weighted least squares for inferring the correlation structure of microbes from compositional data in microbiome studies. We perform a large number of numerical experiments and the simulation results show that fastCCLasso outperforms its competitors in edge detection for inferring the correlation network. We also apply fastCCLasso for estimating microbial networks in microbiome studies and fastCCLasso provides a conservative network with comparable false discovery counts that are derived from shuffled data. AVAILABILITY AND IMPLEMENTATION: FastCCLasso is open source and freely available from https://github.com/ShenZhang-Statistics/fastCCLasso under GNU LGPL v3. Huaying Fang, Tao Hu 0001 |
Bioinform. | 2 |
| 2023 | gmcoda: Graphical model for multiple compositional vectors in microbiome studiesabstractMOTIVATION: Microbes are essential components in the ecosystem and participate in most biological procedures in environments. The high-throughput sequencing technologies help researchers directly quantify the abundance of microbes in a natural environment. Microbiome studies explore the construction, stability, and function of microbial communities with the aid of sequencing technology. However, sequencing technologies only provide relative abundances of microbes, and this kind of data is called compositional data in statistics. The constraint of the constant-sum requires flexible statistical methods for analyzing microbiome data. Current statistical analysis of compositional data mainly focuses on one compositional vector such as bacterial communities. The fungi are also an important component in microbial communities and are always measured by sequencing internal transcribed spacer instead of 16S rRNA genes for bacteria. The different sequencing methods between fungi and bacteria bring two compositional vectors in microbiome studies. RESULTS: We propose a novel statistical method, called gmcoda, based on an additive logistic normal distribution for estimating the partial correlation matrix for cross-domain interactions. A majorization-minimization algorithm is proposed to solve the optimization problem involved in gmcoda. Through simulation studies, gmcoda is demonstrated to work well in estimating partial correlations between two compositional vectors. Gmcoda is also applied to infer cross-domain interactions in a real microbiome dataset and finds potential interactions between bacteria and fungi. AVAILABILITY AND IMPLEMENTATION: Gmcoda is open source and freely available from https://github.com/huayingfang/gmcoda under GNU LGPL v3. Huaying Fang |
Bioinform. | 1 |
| 2022 | Estimating graphical models for count data with applications to single-cell gene networkabstractGraphical models such as Gaussian graphical models have been widely applied for direct interaction inference in many different areas. In many modern applications, such as single-cell RNA sequencing (scRNA-seq) studies, the observed data are counts and often contain many small counts. Traditional graphical models for continuous data are inappropriate for network inference of count data. We consider the Poisson log-normal (PLN) graphical model for count data and the precision matrix of the latent normal distribution represents the network. We propose a two-step method PLNet to estimate the precision matrix. PLNet first estimates the latent covariance matrix using the maximum marginal likelihood estimator (MMLE) and then estimates the precision matrix by minimizing the lasso-penalized D-trace loss function. We establish the convergence rate of the MMLE of the covariance matrix and further establish the convergence rate and the sign consistency of the proposed PLNet estimator of the precision matrix in the high dimensional setting. Importantly, although the PLN model is not sub-Gaussian, we show that the PLNet estimator is consistent even if the model dimension goes to infinity exponentially as the sample size increases. The performance of PLNet is evaluated and compared with available methods using simulation and gene regulatory network analysis of real scRNA-seq data. Feiyi Xiao, Huaying Fang, Ruibin Xi |
NeurIPS | 3 |
| 2019 | Network Clustering Analysis Using Mixture Exponential-Family Random Graph Models and Its Application in Genetic Interaction DataabstractMOTIVATION: Epistatic miniarrary profile (EMAP) studies have enabled the mapping of large-scale genetic interaction networks and generated large amounts of data in model organisms. It provides an incredible set of molecular tools and advanced technologies that should be efficiently understanding the relationship between the genotypes and phenotypes of individuals. However, the network information gained from EMAP cannot be fully exploited using the traditional statistical network models. Because the genetic network is always heterogeneous, for example, the network structure features for one subset of nodes are different from those of the left nodes. Exponential-family random graph models (ERGMs) are a family of statistical models, which provide a principled and flexible way to describe the structural features (e.g., the density, centrality, and assortativity) of an observed network. However, the single ERGM is not enough to capture this heterogeneity of networks. In this paper, we consider a mixture ERGM (MixtureEGRM) networks, which model a network with several communities, where each community is described by a single EGRM. RESULTS: EM algorithm is a classical method to solve the mixture problem, however, it will be very slow when the data size is huge in the numerous applications. We adopt an efficient novel online graph clustering algorithm to classify the graph nodes and estimate the ERGM parameters for the MixtureERGM. In comparison studies, the MixtureERGM outperforms the role analysis for the network cluster in which the mixture of exponential-family random graph model is developed for many ego-network according to their roles. One genetic interaction network of yeast and two real social networks (provided as supplemental materials, which can be found on the Computer Society Digital Library at http://doi.ieeecomputersociety.org/10.1109/TCBB.2017.2743711) show the wide potential application of the MixtureERGM. Yishu Wang 0003, Huaying Fang, Dejie Yang, Hongyu Zhao 0003, Minghua Deng |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2017 | VCNet: vector-based gene co-expression network construction and its application to RNA-seq dataabstractMOTIVATION: Building gene co-expression network (GCN) from gene expression data is an important field of bioinformatic research. Nowadays, RNA-seq data provides high dimensional information to quantify gene expressions in term of read counts for individual exons of genes. Such an increase in the dimension of expression data during the transition from microarray to RNA-seq era made many previous co-expression analysis algorithms based on simple univariate correlation no longer applicable. Recently, two vector-based methods, SpliceNet and RNASeqNet, have been proposed to build GCN. However, they failed to work when sample size is less than the number of exons. RESULTS: We develop an algorithm called VCNet to construct GCN from RNA-seq data to overcome this dimensional problem. VCNet performs a new statistical hypothesis test based on the correlation matrix of a gene-gene pair using the Frobenius norm. The asymptotic distribution of the new test is obtained under the null model. Simulation studies demonstrate that VCNet outperforms SpliceNet and RNASeqNet for detecting edges of GCN. We also apply VCNet to two expression datasets from TCGA database: the normal breast tissue and kidney tumour tissue, and the results show that the GCNs constructed by VCNet contain more biologically meaningful interactions than existing methods. CONCLUSION: VCNet is a useful tool to construct co-expression network. AVAILABILITY AND IMPLEMENTATION: VCNet is open source and freely available from https://github.com/wangzengmiao/VCNet under GNU LGPL v3. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zengmiao Wang, Huaying Fang, Nelson L. S. Tang, Minghua Deng |
Bioinform. | 2 |
| 2015 | CCLasso: correlation inference for compositional data through LassoabstractMOTIVATION: Direct analysis of microbial communities in the environment and human body has become more convenient and reliable owing to the advancements of high-throughput sequencing techniques for 16S rRNA gene profiling. Inferring the correlation relationship among members of microbial communities is of fundamental importance for genomic survey study. Traditional Pearson correlation analysis treating the observed data as absolute abundances of the microbes may lead to spurious results because the data only represent relative abundances. Special care and appropriate methods are required prior to correlation analysis for these compositional data. RESULTS: In this article, we first discuss the correlation definition of latent variables for compositional data. We then propose a novel method called CCLasso based on least squares with [Formula: see text] penalty to infer the correlation network for latent variables of compositional data from metagenomic data. An effective alternating direction algorithm from augmented Lagrangian method is used to solve the optimization problem. The simulation results show that CCLasso outperforms existing methods, e.g. SparCC, in edge recovery for compositional data. It also compares well with SparCC in estimating correlation network of microbe species from the Human Microbiome Project. AVAILABILITY AND IMPLEMENTATION: CCLasso is open source and freely available from https://github.com/huayingfang/CCLasso under GNU LGPL v3. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Huaying Fang, Chengcheng Huang, Hongyu Zhao 0003, Minghua Deng |
Bioinform. | 1 |