VLDB 2026 Research / reviewers in the wild / expert
Zhong Wang 0001
dblp:02/2025-1
· DBLP profile ↗
31ranked-venue papers
5as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 29 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MambaHM: High-Resolution Histone Modification Prediction from ATAC-seq and DNA SequenceabstractHistone modifications play a crucial role in transcriptional regulation and are essential targets for genome annotation and gene expression modeling. However, experimentally profiling histone modifications across species, tissues, and cellular states is costly and often impractical. While numerous deep learning models have been proposed to predict histone modifications, most operate at relatively low resolution (128 bp), limiting their utility in fine-scale genomic analysis. In this study, we present MambaHM (Mamba for predicting Histone Modifications), a novel deep learning framework built upon Mamba architecture. MambaHM integrates DNA sequence features with chromatin accessibility data (ATAC-seq) to predict ten histone modifications. Leveraging the linear computational complexity of Mamba's state-space modeling, our model avoids excessive compression of feature length during embedding. This enables the model to retain sufficient contextual information while simultaneously achieving 16-bp resolution, thereby enhancing the granularity and accuracy of prediction. Experiments demonstrate that MambaHM achieves a mean Pearson correlation of 0.836 (± 0.063) on the K562 cell line test set. Furthermore, the model generalizes well across cell types, tissues, and species, with a mean cross-context correlation of$0.557 (\pm 0.105)$, approaching the reliability of experimental assays. Compared to state-of-the-art models, MambaHM achieves a$\mathbf{7. 0 3 3 \%}(\mathbf{\pm 0. 0 4 1})$improvement in cross-cell-type performance and over a 58.978 % (877 bp) improvement in prediction deviation evaluation for peak calling. Overall, MambaHM provides a powerful, cost-effective, and fineresolution tool for histone modification prediction, offering precise epigenomic references for downstream analysis and potential applications in drug discovery. The source code is available at https://github.com/zhichunlizzx/MambaHM. Lijuan Jia, Zhixiang Xu, Zengyou He, Zhong Wang 0001, Xiaoya Fan |
BIBM | 5 |
| 2025 | Visual Feature Learning from Randomized EEG Trials for Object RecognitionabstractObject recognition from electroencephalography (EEG) responses to visual stimuli has received growing attention but has remained challenging, with prior methods achieving only marginally above-chance accuracy for randomized EEG trials. This paper introduces GVFL-EEG (Guided Visual Feature Learning from EEG), a novel framework that leverages well-trained computer vision models to enhance EEG object decoding. Specifically, a pre-trained image encoder is used to guide EEG feature learning via contrastive learning, aligning EEG embeddings with image representations. We present EEGMambaformer, a custom EEG encoder incorporating a residual Mamba block to capture temporal dynamics, an inverted transformer encoder to extract spatial dependencies, a temporal-spatial convolution block for feature fusion, and a projection layer for dimension alignment with the image encoder. Evaluation on the EEG40000 dataset show that our framework achieves 60.94% accuracy in 40-way classification, outperforming existing state-of-the-art methods. This work advances neural decoding and provides insights into brain-computer interface development. The code is available at https://github.com/xuehaixiao/GVFL_EEG/tree/main/GVFL. Xiaoya Fan, Haixiao Xue, Yufan Feng, Qi Zhao 0008, Zhong Wang 0001 |
ICME | 6 |
| 2024 | Partition, Predict and Assemble: Targeting Long RNA Secondary Structure PredictionabstractAccurately predicting the secondary structure of Non-coding RNAs is crucial for understanding their biological roles. However, current methods face challenges in accuracy and computational efficiency, particularly for long RNA sequences. Here, we present a novel three-step framework for RNA secondary structure prediction, PPA (Partition, Predict, and Assemble). It first partitions the RNA sequence into independent fragments based on exterior loops, then independently predicts the secondary structure of each fragment, and finally assembles them to construct the complete RNA secondary structure. Following the PPA framework, we introduce ELPBert, a model for RNA exterior loop prediction. We partition the RNA sequence at the central nucleotide of the predicted exterior-loop bases. The performance of six state-of-the-art RNA secondary structure prediction methods with and without RNA partition using our ELPBert were compared. Results demonstrate a great improvement of these methods with RNA partition, especially for long sequences. We believe our PPA framework could serve as an universal framework for RNA secondary structure prediction, particularly for long RNA sequences. Data and code are provided at https://github.com/Kali-ym/PPAwithELPBert. Xiaoya Fan, Yuming Cui, Zengyou He, Qi Zhao 0008, Zhong Wang 0001 |
BIBM | 6 |
| 2024 | EEG-based Seizure Type Classification with Temporal-Spatial-Spectral AttentionabstractSeizure detection and type classification from electroencephalogram (EEG) has the potential to improve the diagnosis and treatment of epilepsy. Although the task of seizure detection has been well-investigated, seizure type classification remains largely unexplored. The intricate nature of seizure dynamics presents a significant challenge in effectively extracting distinguishing features from noisy and high-dimensional EEG signals. Previous studies mainly focus on extracting features from temporal, spectral or special domain. Extracting temporal-spatial-spectral features simultaneously from EEG remains challenging. In this paper, we introduce an attention-based neural network to effectively extract temporal-spatial-spectral EEG features, for seizure type classification. It utilizes an attention module with one-shot aggregation to extract multi-level temporal-spatial-spectral EEG features, aiming to differentiate the complex patterns of various seizure types. Specifically, we construct the 3-dimensional representation of EEG by stacking the time-frequency matrices obtained from short time Fourier transform. The attention module consists of paralleled temporal and spatial-spectral attention blocks, allowing the model to focus on the most distinctive time stamps, sensor locations, and frequency bands. The proposed approach is validated on the largest public seizure EEG database, TUSZ v1.5.2. Five-fold cross validation demonstrates that our framework achieved 0.951 weighted F1 score on seizure type classification, achieving state-of-the-art performance. Ablation study confirmed the effectiveness of the temporal and spatial-spectral attention blocks. Xiaoya Fan, Pengzhi Xu, Wenkui Sun, Qi Zhao 0008, Chenru Hao, Zengyou He, Zhong Wang 0001 |
BIBM | 9 |
| 2024 | KAS-former: a transformer-based model for predicting histone modifications using KAS-seqabstractHistone modifications (HMs) play a critical role in various biological processes, but annotating histone modifications across different cell types using experimental methods alone is extremely challenging. Although many deep learning methods have been developed to predict histone modifications, most rely solely on DNA sequences and do not incorporate novel cell-specific features. In this study, we propose KAS-former, a transformer-based model that integrates DNA sequences with cell-specific features derived from KAS-seq data, enabling effective prediction of histone modifications. Leveraging this transformer architecture coupled with dilated convolution, KAS-former achieves a broad receptive field, effectively capturing cell type-specific specificity from KAS-seq data. Our results demonstrate that KAS-former achieves high accuracy in predicting histone modifications across multiple cell types and shows strong potential for transcription factor prediction. By capturing cell-specific features, this approach not only improves the accuracy of histone modification predictions but also offers valuable insights into the interplay between histone modifications and transcription regulation. The code for KAS-former is available on GitHub at https://github.com/wzhy2000/KAS-former. Lijuan Jia, Wen Wen 0004, Xiaoya Fan, Zengyou He, Ruitu Lyu, Zhong Wang 0001 |
BIBM | 8 |
| 2024 | Integrating topology and biological information to predict essential proteins via Shannon entropyabstractIdentifying essential proteins is vital for deciphering the intricacies of disease mechanisms and devising efficacious therapeutic strategies. Over the past several decades, a plethora of algorithms have been proposed, aimed at synthesizing topological and biological information to address the complex challenge of essential protein identification. Nevertheless, a critical examination of the current methodologies reveals certain limitations: (1) the aggregation of diverse features in various methods often requires parameter tuning to maintain balance, potentially introducing instability and increasing complexity in practical scenarios; (2) traditional methods commonly combine various features without in-depth mathematical or physical interpretation, possibly falling short of fully revealing the principles behind the observed phenomena. Hence, we propose a new algorithm for essential protein detection, which is named ITBSE. The basic idea behind this method is to reconstruct the PPI network by removing false positive edges and the subsequent allocation of a protein score. This scoring process integrates both the topological attributes of the reconstructed PPI network and biological data, utilizing the computational framework of Shannon entropy for a comprehensive assessment. To evaluate the effectiveness of our method, we conduct the experiments on real PPI networks and compare with 10 popular methods including DC, BC, CC, LID, PR, DMNC, LAC, NC, PeC and esPOS. The comparison results demonstrate that ITBSE is able to achieve better performance than those competing algorithms. Yan Liu 0085, Zhong Wang 0001, Zengyou He, Hexin Zhang, Jing Qin 0007 |
BIBM | 2 |
| 2024 | Essential protein discovery on weighted PPI networks via statistical information fusionabstractIdentifying essential proteins is crucial for understanding disease mechanisms and developing therapeutic strategies. During the past several decades, numerous algorithms have been introduced to integrate topological and biological information to tackle the challenge of identifying essential proteins. However, existing methods still have some drawbacks: (1) the lack of rigorous mathematical interpretation for the parameters that determine their respective weights when integrating topological and biological information; (2) the lack of flexibility for adding or removing topological and biological information from integration process in real applications. To overcome these limitations, we propose a novel essential protein discovery method, called EPSIF, which assigns weights to PPI network interactions and integrates diverse topological and biological information via a statistical ensemble model. To assess the performance of EPSIF, we conduct experiments on real PPI networks and compare with eight state-of-the-art algorithms, which confirm the effectiveness and flexibility of EPSIF. Yan Liu 0085, Zhong Wang 0001, Zengyou He, Hexin Zhang, Hongwei Wu, Ya-Dong Wang |
BIBM | 3 |
| 2024 | Electroencephalogram Helps Few-Shot LearningabstractLearning to categorize images with limited samples is a challenge for machines. However, humans can easily generalize from just a few examples. In this study, we propose that the remarkable ability of the human brain to generalize can be reflected in electroencephalogram (EEG) signals. These EEG-related features have the potential to enhance few-shot image classification. Our novel two-stage approach involves the following: first, we learn transferable knowledge from large labeled auxiliary sets by multimodal learning of images and EEG signals using contrastive learning. Then, we finetune the image encoder with novel classes that have only a few samples. We integrate this approach with the Triplet and ProxyNCA framework. Experimental results demonstrate an average improvement of 6.1% and 8.5% in terms of top-1 recall compared to the original Triplet and ProxyNCA methods, respectively. This work showcases the feasibility of leveraging brain signals to enhance few-shot learning. Xiaoya Fan, Zhong Wang 0001 |
ICASSP | 3 |
| 2024 | A Domain Adaption Approach for EEG-Based Automated Seizure Classification with Temporal-Spatial-Spectral Attention
Xiaoya Fan, Pengzhi Xu, Qi Zhao 0008, Chenru Hao, Zhong Wang 0001 |
MICCAI (5) | 6 |
| 2024 | dHICA: a deep transformer-based model enables accurate histone imputation from chromatin accessibilityabstractHistone modifications (HMs) are pivotal in various biological processes, including transcription, replication, and DNA repair, significantly impacting chromatin structure. These modifications underpin the molecular mechanisms of cell-type-specific gene expression and complex diseases. However, annotating HMs across different cell types solely using experimental approaches is impractical due to cost and time constraints. Herein, we present dHICA (deep histone imputation using chromatin accessibility), a novel deep learning framework that integrates DNA sequences and chromatin accessibility data to predict multiple HM tracks. Employing the transformer architecture alongside dilated convolutions, dHICA boasts an extensive receptive field and captures more cell-type-specific information. dHICA outperforms state-of-the-art baselines and achieves superior performance in cell-type-specific loci and gene elements, aligning with biological expectations. Furthermore, dHICA's imputations hold significant potential for downstream applications, including chromatin state segmentation and elucidating the functional implications of SNPs (Single Nucleotide Polymorphisms). In conclusion, dHICA serves as a valuable tool for advancing the understanding of chromatin dynamics, offering enhanced predictive capabilities and interpretability. Wen Wen 0004, Lijuan Jia, Tinyi Chu, Nating Wang, Charles G. Danko, Zhong Wang 0001 |
Briefings Bioinform. | 8 |
| 2020 | HiGwas: how to compute longitudinal GWAS data in population designsabstractSUMMARY: Genome-wide association studies (GWAS), particularly designed with thousands and thousands of single-nucleotide polymorphisms (SNPs) (big p) genotyped on tens of thousands of subjects (small n), are encountered by a major challenge of p ≪ n. Although the integration of longitudinal information can significantly enhance a GWAS's power to comprehend the genetic architecture of complex traits and diseases, an additional challenge is generated by an autocorrelative process. We have developed several statistical models for addressing these two challenges by implementing dimension reduction methods and longitudinal data analysis. To make these models computationally accessible to applied geneticists, we wrote an R package of computer software, HiGwas, designed to analyze longitudinal GWAS datasets. Functions in the package encompass single SNP analyses, significance-level adjustment, preconditioning and model selection for a high-dimensional set of SNPs. HiGwas provides the estimates of genetic parameters and the confidence intervals of these estimates. We demonstrate the features of HiGwas through real data analysis and vignette document in the package. AVAILABILITY AND IMPLEMENTATION: https://github.com/wzhy2000/higwas. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhong Wang 0001, Nating Wang, Libo Jiang, Yaqun Wang, Jiahan Li, Rongling Wu, Janet Kelso |
Bioinform. | 1 |
| 2016 | RTFBSDB: an integrated framework for transcription factor binding site analysisabstractUNLABELLED: Transcription factors (TFs) regulate complex programs of gene transcription by binding to short DNA sequence motifs. Here, we introduce rtfbsdb, a unified framework that integrates a database of more than 65 000 TF binding motifs with tools to easily and efficiently scan target genome sequences. Rtfbsdb clusters motifs with similar DNA sequence specificities and integrates RNA-seq or PRO-seq data to restrict analyses to motifs recognized by TFs expressed in the cell type of interest. Our package allows common analyses to be performed rapidly in an integrated environment. AVAILABILITY AND IMPLEMENTATION: rtfbsdb available at (https://github.com/Danko-Lab/rtfbs_db). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhong Wang 0001, André L. Martins, Charles G. Danko |
Bioinform. | 1 |
| 2015 | Functional mapping of seasonal transition in perennial plantsabstractUnlike annuals, all perennial plants undergo seasonal transitions during ontogeny. As an adaptive response to seasonal changes in climate, the seasonal pattern of growth is likely to be under genetic control, although its underlying genetic basis remains unknown. Here, we develop a computational model that can map specific quantitative trait loci (QTLs) responsible for seasonal transitions of growth in perennials. The model is founded on functional mapping, a statistical framework to map developmental dynamics, which is reformed to integrate a seasonally adjusted growth function. The new model is equipped with a capacity to characterize the genetic effects of QTLs on seasonal alternation at different ages and then to better elucidate the genetic architecture of development. The model is implemented with a series of testing procedures, including (i) how a QTL controls an overall ontogenetic growth curve, (ii) how the QTL determines seasonal trajectories of growth within years and (iii) how it determines the dynamic nature of age-specific season response. The model was validated through computer simulation. The extension of season adjustment to other types of biological curves is statistically straightforward, facilitating a wider variety of genetic studies into ontogenetic growth and development in perennial plants. Meixia Ye, Libo Jiang, Ke Mao, Yaqun Wang, Zhong Wang 0001, Rongling Wu |
Briefings Bioinform. | 5 |
| 2015 | A multi-Poisson dynamic mixture model to cluster developmental patterns of gene expression by RNA-seqabstractDynamic changes of gene expression reflect an intrinsic mechanism of how an organism responds to developmental and environmental signals. With the increasing availability of expression data across a time-space scale by RNA-seq, the classification of genes as per their biological function using RNA-seq data has become one of the most significant challenges in contemporary biology. Here we develop a clustering mixture model to discover distinct groups of genes expressed during a period of organ development. By integrating the density function of multivariate Poisson distribution, the model accommodates the discrete property of read counts characteristic of RNA-seq data. The temporal dependence of gene expression is modeled by the first-order autoregressive process. The model is implemented with the Expectation-Maximization algorithm and model selection to determine the optimal number of gene clusters and obtain the estimates of Poisson parameters that describe the pattern of time-dependent expression of genes from each cluster. The model has been demonstrated by analyzing a real data from an experiment aimed to link the pattern of gene expression to catkin development in white poplar. The usefulness of the model has been validated through computer simulation. The model provides a valuable tool for clustering RNA-seq data, facilitating our global view of expression dynamics and understanding of gene regulation mechanisms. Meixia Ye, Zhong Wang 0001, Yaqun Wang, Rongling Wu |
Briefings Bioinform. | 2 |
| 2014 | Systems mapping: how to map genes for biomass allocation toward an ideotypeabstractThe recent availability of high-throughput genetic and genomic data allows the genetic architecture of complex traits to be systematically mapped. The application of these genetic results to design and breed new crop types can be made possible through systems mapping. Systems mapping is a computational model that dissects a complex phenotype into its underlying components, coordinates different components in terms of biological laws through mathematical equations and maps specific genes that mediate each component and its connection with other components. Here, we present a new direction of systems mapping by integrating this tool with carbon economy. With an optimal spatial distribution of carbon fluxes between sources and sinks, plants tend to maximize whole-plant growth and competitive ability under limited availability of resources. We argue that such an economical strategy for plant growth and development, once integrated with systems mapping, will not only provide mechanistic insights into plant biology, but also help to spark a renaissance of interest in ideotype breeding in crops and trees. Wenhao Bo, Guifang Fu, Zhong Wang 0001, Jichen Xu, Zhongwen Huang, Junyi Gai, C. Eduardo Vallejos, Rongling Wu |
Briefings Bioinform. | 3 |
| 2014 | Shape mapping: genetic mapping meets geometric morphometricsabstractKnowledge about biological shape has important implications in biology and biomedicine, but the underlying genetic mechanisms for shape variation have not been well studied. Statistical models play a pivotal role in mapping specific quantitative trait loci (QTLs) that contribute to biological shape and its developmental trajectories. We describe and assess a statistical framework for shape gene identification that incorporates shape and image analysis into a mixture-model framework for QTL mapping. Statistical parameters that define genotype-specific differences in biological shape are estimated by implementing statistical and computational algorithms. A state-of-the-art procedure is described to examine the control patterns of specific QTLs on the origin, properties and functions of biological shape. The statistical framework described will help to address many integrative biological and genetic questions and challenges in shape variation faced by the life sciences community. Wenhao Bo, Zhong Wang 0001, Guifang Fu, Yihan Sui, Weimiao Wu, Xuli Zhu, Danni Yin, Qin Yan, Rongling Wu |
Briefings Bioinform. | 2 |
| 2014 | An allometric model for mapping seed development in plantsabstractDespite a tremendous effort to map quantitative trait loci (QTLs) responsible for agriculturally and biologically important traits in plants, our understanding of how a QTL governs the developmental process of plant seeds remains elusive. In this article, we address this issue by describing a model for functional mapping of seed development through the incorporation of the relationship between vegetative and reproductive growth. The time difference of reproductive from vegetative growth is described by Reeve and Huxley’s allometric equation. Thus, the implementation of this equation into the framework of functional mapping allows dynamic QTLs for seed development to be identified more precisely. By estimating and testing mathematical parameters that define Reeve and Huxley’s allometric equations of seed growth, the dynamic pattern of the genetic effects of the QTLs identified can be analyzed. We used the model to analyze a soybean data, leading to the detection of QTLs that control the growth of seed dry weight. Three dynamic QTLs, located in two different linkage groups, were detected to affect growth curves of seed dry weight. The QTLs detected may be used to improve seed yield with marker-assisted selection by altering the pattern of seed development in a hope to achieve a maximum size of seeds at a harvest time. Zhongwen Huang, Chunfa Tong, Wenhao Bo, Xiaoming Pang, Zhong Wang 0001, Jichen Xu, Junyi Gai, Rongling Wu |
Briefings Bioinform. | 5 |
| 2014 | A case-control design for testing and estimating epigenetic effects on complex diseasesabstractEpigenetic modifications may play an important role in the formation and progression of complex diseases through the regulation of gene expression. The systematic identification of epigenetic variants that contribute to human diseases can be made possible using genome-wide association studies (GWAS), although epigenetic effects are currently not included in commonly used case-control designs for GWAS. Here, we show that epigenetic modifications can be integrated into a case-control setting by dissolving the overall genetic effect into its different components, additive, dominant and epigenetic. We describe a general procedure for testing and estimating the significance of each component based on a conventional chi-squared test approach. Simulation studies were performed to investigate the power and false-positive rate of this procedure, providing recommendations for its practical use. The integration of epigenetic variants into GWAS can potentially improve our understanding of how genetic, environmental and stochastic factors interact with epialleles to construct the genetic architecture of complex diseases. Yihan Sui, Weimiao Wu, Zhong Wang 0001, Jianxin Wang 0004, Zuoheng Wang, Rongling Wu |
Briefings Bioinform. | 3 |
| 2014 | Structural mapping: how to study the genetic architecture of a phenotypic trait through its formation mechanismabstractTraditional approaches for genetic mapping are to simply associate the genotypes of a quantitative trait locus (QTL) with the phenotypic variation of a complex trait. A more mechanistic strategy has emerged to dissect the trait phenotype into its structural components and map specific QTLs that control the mechanistic and structural formation of a complex trait. We describe and assess such a strategy, called structural mapping, by integrating the internal structural basis of trait formation into a QTL mapping framework. Electrical impedance spectroscopy (EIS) has been instrumental for describing the structural components of a phenotypic trait and their interactions. By building robust mathematical models on circuit EIS data and embedding these models within a mixture model-based likelihood for QTL mapping, structural mapping implements the EM algorithm to obtain maximum likelihood estimates of QTL genotype-specific EIS parameters. The uniqueness of structural mapping is to make it possible to test a number of hypotheses about the pattern of the genetic control of structural components. We validated structural mapping by analyzing an EIS data collected for QTL mapping of frost hardiness in a controlled cross of jujube trees. The statistical properties of parameter estimates were examined by simulation studies. Structural mapping can be a powerful alternative for genetic mapping of complex traits by taking account into the biological and physical mechanisms underlying their formation. Chunfa Tong, Lianying Shen, Yafei Lv, Zhong Wang 0001, Sisi Feng, Xin Li 0025, Yihan Sui, Xiaoming Pang, Rongling Wu |
Briefings Bioinform. | 4 |
| 2014 | A bi-Poisson model for clustering gene expression profiles by RNA-seqabstractWith the availability of gene expression data by RNA-seq, powerful statistical approaches for grouping similar gene expression profiles across different environments have become increasingly important. We describe and assess a computational model for clustering genes into distinct groups based on the pattern of gene expression in response to changing environment. The model capitalizes on the Poisson distribution to capture the count property of RNA-seq data. A two-stage hierarchical expectation–maximization (EM) algorithm is implemented to estimate an optimal number of groups and mean expression amounts of each group across two environments. A procedure is formulated to test whether and how a given group shows a plastic response to environmental changes. The impact of gene–environment interactions on the phenotypic plasticity of the organism can also be visualized and characterized. The model was used to analyse an RNA-seq dataset measured from two cell lines of breast cancer that respond differently to an anti-cancer drug, from which genes associated with the resistance and sensitivity of the cell lines are identified. We performed simulation studies to validate the statistical behaviour of the model. The model provides a useful tool for clustering gene expression data by RNA-seq, facilitating our understanding of gene functions and networks. Ningtao Wang, Yaqun Wang, Luojun Wang, Zhong Wang 0001, Jianxin Wang 0004, Rongling Wu |
Briefings Bioinform. | 5 |
| 2014 | Towards a comprehensive picture of the genetic landscape of complex traitsabstractThe formation of phenotypic traits, such as biomass production, tumor volume and viral abundance, undergoes a complex process in which interactions between genes and developmental stimuli take place at each level of biological organization from cells to organisms. Traditional studies emphasize the impact of genes by directly linking DNA-based markers with static phenotypic values. Functional mapping, derived to detect genes that control developmental processes using growth equations, has proven powerful for addressing questions about the roles of genes in development. By treating phenotypic formation as a cohesive system using differential equations, a different approach-systems mapping-dissects the system into interconnected elements and then map genes that determine a web of interactions among these elements, facilitating our understanding of the genetic machineries for phenotypic development. Here, we argue that genetic mapping can play a more important role in studying the genotype-phenotype relationship by filling the gaps in the biochemical and regulatory process from DNA to end-point phenotype. We describe a new framework, named network mapping, to study the genetic architecture of complex traits by integrating the regulatory networks that cause a high-order phenotype. Network mapping makes use of a system of differential equations to quantify the rule by which transcriptional, proteomic and metabolomic components interact with each other to organize into a functional whole. The synthesis of functional mapping, systems mapping and network mapping provides a novel avenue to decipher a comprehensive picture of the genetic landscape of complex phenotypes that underlie economically and biomedically important traits. Zhong Wang 0001, Yaqun Wang, Ningtao Wang, Jianxin Wang 0004, Zuoheng Wang, C. Eduardo Vallejos, Rongling Wu |
Briefings Bioinform. | 1 |
| 2013 | A multivalent three-point linkage analysis model of autotetraploidsabstractBecause of its widespread occurrence and role in shaping evolutionary processes in the biological kingdom, especially in plants, polyploidy has been increasingly studied from cytological to molecular levels. By inferring gene order, gene distances and gene homology, linkage mapping with molecular markers has proven powerful for investigating genome structure and organization. Here we review and assess a general statistical model for three-point linkage analysis in autotetraploids by integrating double reduction, a phenomenon that commonly occurs in autopolyploids whose chromosomes are derived from a single ancestral species. This model does not require any assumption on the distribution of the occurrence of double reduction and can handle the complexity of multilocus linkage in terms of crossover interference. Implemented with the expectation-maximization (EM) algorithms, the model can estimate and test the recombination fractions between less informative dominant markers, thus facilitating its practical implications for any autopolyploids in most of which inexpensive dominant markers are still used for their genetic and evolutionary studies. The model was applied to reanalyze a published data in tetraploid switchgrass, validating its practical usefulness and utilization. Yafei Lv, Chunfa Tong, Xin Li 0025, Sisi Feng, Zhong Wang 0001, Xiaoming Pang, Yaqun Wang, Ningtao Wang, Christian M. Tobias, Rongling Wu |
Briefings Bioinform. | 6 |
| 2013 | A statistical procedure to map high-order epistasis for complex traitsabstractGenetic interactions or epistasis have been thought to play a pivotal role in shaping the formation, development and evolution of life. Previous work focused on lower-order interactions between a pair of genes, but it is obviously inadequate to explain a complex network of genetic interactions and pathways. We review and assess a statistical model for characterizing high-order epistasis among more than two genes or quantitative trait loci (QTLs) that control a complex trait. The model includes a series of start-of-the-art standard procedures for estimating and testing the nature and magnitude of QTL interactions. Results from simulation studies and real data analysis warrant the statistical properties of the model and its usefulness in practice. High-order epistatic mapping will provide a routine procedure for charting a detailed picture of the genetic regulation mechanisms underlying the phenotypic variation of complex traits. Xiaoming Pang, Zhong Wang 0001, John S. Yap, Jianxin Wang 0004, Junjia Zhu, Wenhao Bo, Yafei Lv, Shaofeng Peng, Dengfeng Shen, Rongling Wu |
Briefings Bioinform. | 2 |
| 2013 | A dynamic framework for quantifying the genetic architecture of phenotypic plasticityabstractDespite its central role in the adaptation and microevolution of traits, the genetic architecture of phenotypic plasticity, i.e. multiple phenotypes produced by a single genotype in changing environments, remains elusive. We know little about the genes that underlie the plastic response of traits to the environment, their number, chromosomal locations and genetic interactions as well as environment impact on their effects. Here we review key statistical approaches for analyzing the genetic variation of phenotypic plasticity due to genotype-environment interactions and describe the implementation of a dynamic model to map specific quantitative trait loci (QTLs) that affect the gradient expression of a quantitative trait across a range of environments. This dynamic model is distinct by incorporating mathematical aspects of phenotypic plasticity into a QTL mapping framework, thereby better unraveling the quantitative attribute of trait response to the environment. By testing the curve parameters that specify environment-dependent trajectories of the trait, the model allows a series of fundamental hypotheses to be tested in a quantitative way about the interplay between gene action/interaction and environmental sensitivity. The model can also make the dynamic prediction of genetic control over phenotypic plasticity within the context of changing environments. We demonstrate the usefulness of the model by reanalyzing a QTL data set for rice, gleaning new insights into the genetic basis for phenotypic plasticity in plant height growth. Zhong Wang 0001, Xiaoming Pang, Yafei Lv, Xin Li 0025, Sisi Feng, Jiahan Li, Rongling Wu |
Briefings Bioinform. | 1 |
| 2013 | A unifying framework for bivalent multilocus linkage analysis of allotetraploidsabstractAn allotetraploid has four paired sets of chromosomes derived from different diploid species, whose meiotic behavior is qualitatively different from the underlying diploids. According to a traditional view, meiotic pairing occurs only between homologous chromosomes, but new evidence indicates that homoeologous chromosomes may also pair to a lesser extent compared with homolog pairing. Here, we describe and assess a unifying analytical framework that incorporates differential chromosomal pairing into a multilocus linkage model. The preferential pairing factor is used to quantify the probability difference of pairing occurring between homologous chromosomes and homoeologous chromosomes. The unifying framework allows simultaneous estimation of the linkage, genetic interference and preferential pairing factor using commonly existing multiplex markers. We compared the unifying approach and traditional approaches assuming random chromosomal pairing by analyzing marker data collected in a full-sib family of tetraploid switchgrass, a bioenergy species whose diploid origins are undefined, but with subgenomes that are genetically well differentiated. The unifying framework provides a better tool for estimating the meiotic linkage and constructing a genetic map for allotetraploids. Yafei Lv, Xiaoming Pang, Chunfa Tong, Zhong Wang 0001, Xin Li 0025, Sisi Feng, Christian M. Tobias, Rongling Wu |
Briefings Bioinform. | 5 |
| 2013 | A quantitative model of transcriptional differentiation driving host-pathogen interactionsabstractDespite our expanding knowledge about the biochemistry of gene regulation involved in host-pathogen interactions, a quantitative understanding of this process at a transcriptional level is still limited. We devise and assess a computational framework that can address this question. This framework is founded on a mixture model-based likelihood, equipped with functionality to cluster genes per dynamic and functional changes of gene expression within an interconnected system composed of the host and pathogen. If genes from the host and pathogen are clustered in the same group due to a similar pattern of dynamic profiles, they are likely to be reciprocally co-evolving. If genes from the two organisms are clustered in different groups, this means that they experience strong host-pathogen interactions. The framework can test the rates of change for individual gene clusters during pathogenic infection and quantify their impacts on host-pathogen interactions. The framework was validated by a pathological study of poplar leaves infected by fungal Marssonina brunnea in which co-evolving and interactive genes that determine poplar-fungus interactions are identified. The new framework should find its wide application to studying host-pathogen interactions for any other interconnected systems. Zhong Wang 0001, Jianxin Wang 0004, Yaqun Wang, Ningtao Wang, Zuoheng Wang, Xiaohua Su, Mingxiu Wang, Shougong Zhang, Minren Huang, Rongling Wu |
Briefings Bioinform. | 2 |
| 2012 | A computational framework for the inheritance pattern of genomic imprinting for complex traitsabstractGenetic imprinting, by which the expression of a gene depends on the parental origin of its alleles, may be subjected to reprogramming through each generation. Currently, such reprogramming is limited to qualitative description only, lacking more precise quantitative estimation for its extent, pattern and mechanism. Here, we present a computational framework for analyzing the magnitude of genetic imprinting and its transgenerational inheritance mode. This quantitative model is based on the breeding scheme of reciprocal backcrosses between reciprocal F(1) hybrids and original inbred parents, in which the transmission of genetic imprinting across generations can be tracked. We define a series of quantitative genetic parameters that describe the extent and transmission mode of genetic imprinting and further estimate and test these parameters within a genetic mapping framework using a new powerful computational algorithm. The model and algorithm described will enable geneticists to identify and map imprinted quantitative trait loci and dictate a comprehensive atlas of developmental and epigenetic mechanisms related to genetic imprinting. We illustrate the new discovery of the role of genetic imprinting in regulating hyperoxic acute lung injury survival time using a mouse reciprocal backcross design. Zhong Wang 0001, Daniel R. Prows, Rongling Wu |
Briefings Bioinform. | 2 |
| 2012 | How to cluster gene expression dynamics in response to environmental signalsabstractOrganisms usually cope with change in the environment by altering the dynamic trajectory of gene expression to adjust the complement of active proteins. The identification of particular sets of genes whose expression is adaptive in response to environmental changes helps to understand the mechanistic base of gene-environment interactions essential for organismic development. We describe a computational framework for clustering the dynamics of gene expression in distinct environments through Gaussian mixture fitting to the expression data measured at a set of discrete time points. We outline a number of quantitative testable hypotheses about the patterns of dynamic gene expression in changing environments and gene-environment interactions causing developmental differentiation. The future directions of gene clustering in terms of incorporations of the latest biological discoveries and statistical innovations are discussed. We provide a set of computational tools that are applicable to modeling and analysis of dynamic gene expression data measured in multiple environments. Zhong Wang 0001, Junjia Zhu, Li Wang 0021, Runze Li 0001, Scott A. Berceli, Rongling Wu |
Briefings Bioinform. | 3 |
| 2012 | Functional mapping of ontogeny in flowering plantsabstractAll organisms face the problem of how to perform a sequence of developmental changes and transitions during ontogeny. We revise functional mapping, a statistical model originally derived to map genes that determine developmental dynamics, to take into account the entire process of ontogenetic growth from embryo to adult and from the vegetative to reproductive phase. The revised model provides a framework that reconciles the genetic architecture of development at different stages and elucidates a comprehensive picture of the genetic control mechanisms of growth that change gradually from a simple to a more complex level. We use an annual flowering plant, as an example, to demonstrate our model by which to map genes and their interactions involved in embryo and postembryonic growth. The model provides a useful tool to study the genetic control of ontogenetic growth in flowering plants and any other organisms through proper modifications based on their biological characteristics. Xiyang Zhao, Chunfa Tong, Xiaoming Pang, Zhong Wang 0001, Yunqian Guo, Fang Du, Rongling Wu |
Briefings Bioinform. | 4 |
| 2012 | A quantitative genetic and epigenetic model of complex traitsabstractBACKGROUND: Despite our increasing recognition of the mechanisms that specify and propagate epigenetic states of gene expression, the pattern of how epigenetic modifications contribute to the overall genetic variation of a phenotypic trait remains largely elusive. RESULTS: We construct a quantitative model to explore the effect of epigenetic modifications that occur at specific rates on the genome. This model, derived from, but beyond, the traditional quantitative genetic theory that is founded on Mendel's laws, allows questions concerning the prevalence and importance of epigenetic variation to be incorporated and addressed. CONCLUSIONS: It provides a new avenue for bringing chromatin inheritance into the realm of complex traits, facilitating our understanding of the means by which phenotypic variation is generated. Zhong Wang 0001, Zuoheng Wang, Jianxin Wang 0004, Yihan Sui, Jian Zhang 0106, Duanping Liao, Rongling Wu |
BMC Bioinform. | 1 |
| 2011 | 3FunMap: full-sib family functional mapping of dynamic traitsabstractMOTIVATION: Functional mapping that embeds the developmental mechanisms of complex traits shows great power to study the dynamic pattern of genetic effects triggered by individual quantitative trait loci (QTLs). A full-sib family, produced by crossing two heterozygous parents, is characteristic of uncertainties about cross-type at a locus and linkage phase between different loci. Integrating functional mapping into a full-sib family requires a model selection procedure capable of addressing these uncertainties. 3FunMap, written in VC++ 6.0, provides a flexible and extensible platform to perform full-sib functional mapping of dynamic traits. Functions in the package encompass linkage phase determination, marker map construction and the pattern identification of QTL segregation, dynamic tests of QTL effects, permutation tests and numerical simulation. We demonstrate the features of 3FunMap through real data analysis and computer simulation. AVAILABILITY: http://statgen.psu.edu/software. Chunfa Tong, Zhong Wang 0001, Jisen Shi, Rongling Wu |
Bioinform. | 2 |