Wei Zhang 0076

dblp:10/4661-76 · DBLP profile ↗
← Back
30ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0003-3605-9373ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 2 first-author · 13 since 2021Systems, architecture and hardware · 6 · 1 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 2Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EpGAT: integrating epigenetics and 3D genome structure to predict alternative splicing and polyadenylation
abstract
Understanding how the 3D structure of the genome influences gene regulation is a growing area of interest, particularly in the context of alternative post-transcriptional regulatory events such as alternative splicing (AS) and alternative polyadenylation (APA). These processes are essential for generating transcript and protein diversity, and they are tightly coordinated with transcription. However, despite their biological importance, the relationship between chromatin interactions and alternative pre-messenger RNA regulation remains poorly understood. This gap largely stems from a lack of computational tools capable of integrating structural genomic data with RNA processing dynamics. Exploring how chromatin interactions and epigenetic landscapes shape these events is essential for uncovering the multilayered regulation of gene expression. To bridge this gap, we present EpGAT, a graph attention network-based model that integrates epigenetic read coverage and chromatin interaction data to predict and quantify AS and APA events. By explicitly modeling the spatial organization of the genome, EpGAT captures the regulatory influence of chromatin looping and long-range genomic interactions on RNA processing. The model's predictions are validated through rigorous cross-cell line and cross-chromosome evaluations, affirming its generalizability and reliability. Beyond prediction, EpGAT offers interpretability by tracing learned parameters back to genomic features, enabling the identification of active enhancers, mapping promoter-enhancer connectivity, and pinpointing the epigenetic factors most critical to specific RNA processing events. These capabilities make EpGAT a powerful tool for dissecting the complex interplay between genome architecture and transcriptomic regulation. More broadly, it provides a generalizable framework for multiple tasks to study the link between 3D genome organization, epigenetic signals, and RNA processing.
Sudipto Baul, Naima Ahmed Fahmi, Ahmed Louri, Jeongsik Yong, Wei Zhang 0076
Briefings Bioinform.7
2026 X-intNMF: a cross- and intra-omics regularized NMF framework for multi-omics integration
abstract
MOTIVATION: The rapid accumulation of multi-omics data presents a valuable opportunity to advance our understanding of complex diseases and biological systems, driving the development of integrative computational methods. However, the complexity of biological processes, spanning multiple molecular layers and involving intricate regulatory interactions, requires models that can capture both intra- and cross-omics relationships. Most existing integration methods primarily focus on sample-level similarities or intra-omics feature interactions, often neglecting the interactions across different omics layers. This limitation can result in the loss of critical biological information and suboptimal performance. To address this gap, we propose X-intNMF, a network-regularized non-negative matrix factorization (NMF) framework that simultaneously integrates intra- and cross-omics feature interactions into a shared low-dimensional representation (see Fig. 1). By modeling these multi-layered relationships, X-intNMF enhances the representation of biological interactions and improves integration quality and prediction accuracy. RESULTS: For evaluation, we applied X-intNMF to predict breast cancer phenotypes and classify clinical outcomes in lung and ovarian cancers using mRNA expression, microRNA expression, and DNA methylation data from TCGA. The results show that X-intNMF consistently outperforms state-of-the-art methods. Ablation studies confirm that incorporating both cross-omics and intra-omics interactions contributes significantly to the model's improved performance. Additionally, survival analysis on 25 TCGA cancer datasets demonstrates that the integrated multi-omics representation offers strong prognostic value for both overall survival and disease-free status. These findings highlight X-intNMF's ability to effectively model multi-layered molecular interactions while maintaining interpretability, robustness, and scalability within the NMF framework. AVAILABILITY AND IMPLEMENTATION: The source code and datasets supporting this study are publicly available at GitHub (https://github.com/compbiolabucf/X-intNMF) and archived on Zenodo (https://doi.org/10.5281/zenodo.18238385).
Tien-Thanh Bui, Rui Xie 0002, Wei Zhang 0076
Bioinform.3
2026 SoK: Can Fully Homomorphic Encryption Support General AI Computation? A Functional and Cost Analysis
abstract
Artificial intelligence (AI) increasingly powers sensitive applications in domains such as healthcare and finance, relying on both extit{linear operations} (e.g., matrix multiplications in large language models) and extit{non-linear operations} (e.g., sorting in retrieval-augmented generation). Fully homomorphic encryption (FHE) has emerged as a promising tool for privacy-preserving computation, but it remains unclear whether existing methods can support the full spectrum of AI workloads that combine these operations. In this SoK, we ask: extit{Can FHE support general AI computation?} We provide both a functional analysis and a cost analysis. First, we categorize ten distinct FHE approaches and evaluate their ability to support general computation. We then identify three promising candidates and benchmark workloads that mix linear and non-linear operations across different bit lengths and SIMD parallelization settings. Finally, we evaluate five real-world, privacy-sensitive AI applications that instantiate these workloads. Our results quantify the costs of achieving general computation in FHE and offer practical guidance on selecting FHE methods that best fit specific AI application requirements. Our codes are available at https://github.com/UCF-ML-Research/FHE-AI-Generality.
Wei Zhang 0076, Mengxin Zheng, Minxuan Zhou, Yushun Dong, Dongjie Wang 0001, Jiafeng Xie, David Mohaisen, Hongyi Wu, Qian Lou
Proc. Priv. Enhancing Technol.3
2025 Biological Pathway Guided Gene Selection Through Collaborative Reinforcement Learning
abstract
Gene selection in high-dimensional genomic data is essential for understanding disease mechanisms and improving therapeutic outcomes. Traditional feature selection methods effectively identify predictive genes but often ignore complex biological pathways and regulatory networks, leading to unstable and biologically irrelevant signatures. Prior approaches, such as Lasso-based methods and statistical filtering, either focus solely on individual gene-outcome associations or fail to capture pathway-level interactions, presenting a key challenge: how to integrate biological pathway knowledge while maintaining statistical rigor in gene selection? To address this gap, we propose a novel two-stage framework that integrates statistical selection with biological pathway knowledge using multi-agent reinforcement learning (MARL). First, we introduce a pathway-guided pre-filtering strategy that leverages multiple statistical methods alongside KEGG pathway information for initial dimensionality reduction. Next, for refined selection, we model genes as collaborative agents in a MARL framework, where each agent optimizes both predictive power and biological relevance. Our framework incorporates pathway knowledge through Graph Neural Network-based state representations, a reward mechanism combining prediction performance with gene centrality and pathway coverage, and collaborative learning strategies using shared memory and a centralized critic component. Extensive experiments on multiple gene expression datasets demonstrate that our approach significantly improves both prediction accuracy and biological interpretability compared to traditional methods.
Ehtesamul Azim, Dongjie Wang 0001, Taehyun Hwang, Yanjie Fu, Wei Zhang 0076
KDD (2)5
2025 SynOmics: integrating multi-omics data through feature interaction networks
abstract
The integration of multi-omics data is essential for achieving a comprehensive understanding of molecular systems and enhancing the performance of predictive models in biomedical research. However, many existing models have limited capacity to capture cross-omics feature interactions, which hinders the depth of integration. In this study, we introduce SynOmics, a graph convolutional network framework designed to improve multi-omics integration by constructing omics networks in the feature space and modeling both within- and cross-omics dependencies. By incorporating both omics-specific networks and cross-omics bipartite networks, SynOmics enables simultaneous learning of intra-omics and inter-omics relationships. Unlike traditional approaches that rely on early or late integration strategies, SynOmics adopts a parallel learning strategy to process feature-level interactions at each layer of the model. Experimental results demonstrate that SynOmics consistently outperforms state-of-the-art multi-omics integration methods across a range of biomedical classification tasks, highlighting its potential for biomarker discovery and clinical applications.
Muhtasim Noor Alif, Wei Zhang 0076
Briefings Bioinform.2
2025 IPScan: Detecting novel intronic PolyAdenylation events with RNA-seq data
abstract
Intronic PolyAdenylation (IPA) is an important post-transcriptional mechanism that can alter transcript coding potential by truncating translation regions, thereby increasing transcriptome and proteome diversity. This process generates novel protein isoforms with altered peptide sequences, some of which are implicated in disease progression, including cancer. Truncated proteins may lose tumor-suppressive functions, contributing to oncogenesis. Despite advancements in Alternative PolyAdenylation (APA) analysis using RNA-seq, detecting and quantifying novel IPA events remains challenging. To address this, we developed IPScan, a computational pipeline for precise IPA event identification, quantification, and visualization. IPScan has been benchmarked against existing methods using simulated data, different human and mouse cell lines, and TCGA (The Cancer Genome Atlas) breast cancer datasets. Differential IPA events under different biological conditions were quantified and validated via qPCR.
Naima Ahmed Fahmi, Sze Cheng, Jeovani Overstreet, Jeongsik Yong, Wei Zhang 0076
PLoS Comput. Biol.6
2024 Improving Automated Feature Selection in Cancer Genomics Using Reinforcement Learning
abstract
Genomic datasets are often characterized by high dimensionality and a limited number of samples, making it difficult to identify biologically meaningful features. Effective feature selection is crucial in biomedical research, where both genomic and clinical variables are abundant. In this study, we evaluate multiple feature selection techniques on large-scale gene expression datasets from breast cancer, lung adenocarcinoma, and ovarian cancer. We specifically adapt and optimize a novel reinforcement learning (RL) approach to address the challenges of high-dimensional genomic data. The performance of the RL-based method is compared against widely-used techniques, including Least Absolute Shrinkage and Selection Operator (LASSO), Recursive Feature Elimination (RFE), Random Forest, and Principal Component Analysis (PCA). Our experimental results demonstrate that the RL approach outperforms traditional methods in predicting cancer outcomes in most cases, highlighting its potential for identifying biologically meaningful features in genomic analysis.
Hannah Z. Wang, Wei Zhang 0076
IEEE Big Data2
2024 Feature Interaction Aware Automated Data Representation Transformation
abstract
Creating an effective representation space is crucial for mitigating the curse of dimensionality, enhancing model generalization, addressing data sparsity, and leveraging classical models more effectively. Recent advancements in automated feature engineering (AutoFE) have made significant progress in addressing various challenges associated with representation learning, issues such as heavy reliance on intensive labor and empirical experiences, lack of explainable explic-itness, and inflexible feature space reconstruction embedded into downstream tasks. However, these approaches are constrained by: 1) generation of potentially unintelligible and illogical reconstructed feature spaces, stemming from the neglect of expert-level cognitive processes; 2) lack of systematic exploration, which subsequently results in slower model convergence for identification of optimal feature space. To address these, we introduce an interaction-aware reinforced generation perspective. We redefine feature space reconstruction as a nested process of creating meaningful features and controlling feature set size through selection. We develop a hierarchical reinforcement learning structure with cascading Markov Decision Processes to automate feature and operation selection, as well as feature crossing. By incorporating statistical measures, we reward agents based on the interaction strength between selected features, resulting in intelligent and efficient exploration of the feature space that emulates human decision-making. Extensive experiments are conducted to validate our proposed approach.
Ehtesamul Azim, Dongjie Wang 0001, Kunpeng Liu 0001, Wei Zhang 0076, Yanjie Fu
SDM4
2024 Integrating spatial transcriptomics and bulk RNA-seq: predicting gene expression with enhanced resolution through graph attention networks
abstract
Spatial transcriptomics data play a crucial role in cancer research, providing a nuanced understanding of the spatial organization of gene expression within tumor tissues. Unraveling the spatial dynamics of gene expression can unveil key insights into tumor heterogeneity and aid in identifying potential therapeutic targets. However, in many large-scale cancer studies, spatial transcriptomics data are limited, with bulk RNA-seq and corresponding Whole Slide Image (WSI) data being more common (e.g. TCGA project). To address this gap, there is a critical need to develop methodologies that can estimate gene expression at near-cell (spot) level resolution from existing WSI and bulk RNA-seq data. This approach is essential for reanalyzing expansive cohort studies and uncovering novel biomarkers that have been overlooked in the initial assessments. In this study, we present STGAT (Spatial Transcriptomics Graph Attention Network), a novel approach leveraging Graph Attention Networks (GAT) to discern spatial dependencies among spots. Trained on spatial transcriptomics data, STGAT is designed to estimate gene expression profiles at spot-level resolution and predict whether each spot represents tumor or non-tumor tissue, especially in patient samples where only WSI and bulk RNA-seq data are available. Comprehensive tests on two breast cancer spatial transcriptomics datasets demonstrated that STGAT outperformed existing methods in accurately predicting gene expression. Further analyses using the TCGA breast cancer dataset revealed that gene expression estimated from tumor-only spots (predicted by STGAT) provides more accurate molecular signatures for breast cancer sub-type and tumor stage prediction, and also leading to improved patient survival and disease-free analysis. Availability: Code is available at https://github.com/compbiolabucf/STGAT.
Sudipto Baul, Khandakar Tanvir Ahmed, Qibing Jiang, Qian Li 0023, Jeongsik Yong, Wei Zhang 0076
Briefings Bioinform.7
2024 DTI-LM: language model powered drug-target interaction prediction
abstract
MOTIVATION: The identification and understanding of drug-target interactions (DTIs) play a pivotal role in the drug discovery and development process. Sequence representations of drugs and proteins in computational model offer advantages such as their widespread availability, easier input quality control, and reduced computational resource requirements. These make them an efficient and accessible tools for various computational biology and drug discovery applications. Many sequence-based DTI prediction methods have been developed over the years. Despite the advancement in methodology, cold start DTI prediction involving unknown drug or protein remains a challenging task, particularly for sequence-based models. Introducing DTI-LM, a novel framework leveraging advanced pretrained language models, we harness their exceptional context-capturing abilities along with neighborhood information to predict DTIs. DTI-LM is specifically designed to rely solely on sequence representations for drugs and proteins, aiming to bridge the gap between warm start and cold start predictions. RESULTS: Large-scale experiments on four datasets show that DTI-LM can achieve state-of-the-art performance on DTI predictions. Notably, it excels in overcoming the common challenges faced by sequence-based models in cold start predictions for proteins, yielding impressive results. The incorporation of neighborhood information through a graph attention network further enhances prediction accuracy. Nevertheless, a disparity persists between cold start predictions for proteins and drugs. A detailed examination of DTI-LM reveals that language models exhibit contrasting capabilities in capturing similarities between drugs and proteins. AVAILABILITY AND IMPLEMENTATION: Source code is available at: https://github.com/compbiolabucf/DTI-LM.
Khandakar Tanvir Ahmed, Md. Istiaq Ansari, Wei Zhang 0076
Bioinform.3
2024 Optimizing multi-omics data imputation with NMF and GAN synergy
abstract
MOTIVATION: Integrating multiple omics datasets can significantly advance our understanding of disease mechanisms, physiology, and treatment responses. However, a major challenge in multi-omics studies is the disparity in sample sizes across different datasets, which can introduce bias and reduce statistical power. To address this issue, we propose a novel framework, OmicsNMF, designed to impute missing omics data and enhance disease phenotype prediction. OmicsNMF integrates Generative Adversarial Networks (GANs) with Non-Negative Matrix Factorization (NMF). NMF is a well-established method for uncovering underlying patterns in omics data, while GANs enhance the imputation process by generating realistic data samples. This synergy aims to more effectively address sample size disparity, thereby improving data integration and prediction accuracy. RESULTS: For evaluation, we focused on predicting breast cancer subtypes using the imputed data generated by our proposed framework, OmicsNMF. Our results indicate that OmicsNMF consistently outperforms baseline methods. We further assessed the quality of the imputed data through survival analysis, revealing that the imputed omics profiles provide significant prognostic power for both overall survival and disease-free status. Overall, OmicsNMF effectively leverages GANs and NMF to impute missing samples while preserving key biological features. This approach shows potential for advancing precision oncology by improving data integration and analysis. AVAILABILITY AND IMPLEMENTATION: Source code is available at: https://github.com/compbiolabucf/OmicsNMF.
Md. Istiaq Ansari, Khandakar Tanvir Ahmed, Wei Zhang 0076
Bioinform.3
2023 Attention-Based Multi-modal Missing Value Imputation for Time Series Data with High Missing Rate
abstract
Multivariate time series data is prone to a high missing rate which presents an obstacle to statistical analysis of the data. Imputation has become the standard measure to handle this challenge. However, existing time series missing value imputation methods are mostly uni-modal that relies on self- imputation. With an unprecedented rate of data collection, the availability of multi-modal data is increasing, allowing us the opportunity to impute the time series missing values using other datasets generated from the same cohort. In this paper, we propose a multi-modal time series missing value imputation framework, TSEst, that can utilize multiple data modalities to overcome the limitations of self-imputation. The framework uses additional cross-sectional or time series data for the imputation and therefore, is less affected by a high missing rate in the time series data. A comprehensive set of experiments on two datasets shows an improvement in imputation accuracy over the baselines. Experimental results also demonstrate that the improvement is caused by the effective integration of the additional data modality. The proposed framework can impute missing values in the samples with no time series data available, reducing the reliance on long-term data collection. Availability: Code is available at https://github.com/compbiolabucf/TSEst
Khandakar Tanvir Ahmed, Sudipto Baul, Yanjie Fu, Wei Zhang 0076
SDM4
2023 Incomplete time-series gene expression in integrative study for islet autoimmunity prediction
abstract
Type 1 diabetes (T1D) outcome prediction plays a vital role in identifying novel risk factors, ensuring early patient care and designing cohort studies. TEDDY is a longitudinal cohort study that collects a vast amount of multi-omics and clinical data from its participants to explore the progression and markers of T1D. However, missing data in the omics profiles make the outcome prediction a difficult task. TEDDY collected time series gene expression for less than 6% of enrolled participants. Additionally, for the participants whose gene expressions are collected, 79% time steps are missing. This study introduces an advanced bioinformatics framework for gene expression imputation and islet autoimmunity (IA) prediction. The imputation model generates synthetic data for participants with partially or entirely missing gene expression. The prediction model integrates the synthetic gene expression with other risk factors to achieve better predictive performance. Comprehensive experiments on TEDDY datasets show that: (1) Our pipeline can effectively integrate synthetic gene expression with family history, HLA genotype and SNPs to better predict IA status at 2 years (sensitivity 0.622, AUC 0.715) compared with the individual datasets and state-of-the-art results in the literature (AUC 0.682). (2) The synthetic gene expression contains predictive signals as strong as the true gene expression, reducing reliance on expensive and long-term longitudinal data collection. (3) Time series gene expression is crucial to the proposed improvement and shows significantly better predictive ability than cross-sectional gene expression. (4) Our pipeline is robust to limited data availability. Availability: Code is available at https://github.com/compbiolabucf/TEDDY.
Khandakar Tanvir Ahmed, Sze Cheng, Qian Li 0023, Jeongsik Yong, Wei Zhang 0076
Briefings Bioinform.5
2022 CancerCellTracker: a brightfield time-lapse microscopy framework for cancer drug sensitivity estimation
abstract
MOTIVATION: Time-lapse microscopy is a powerful technique that relies on images of live cells cultured ex vivo that are captured at regular intervals of time to describe and quantify their behavior under certain experimental conditions. This imaging method has great potential in advancing the field of precision oncology by quantifying the response of cancer cells to various therapies and identifying the most efficacious treatment for a given patient. Digital image processing algorithms developed so far require high-resolution images involving very few cells originating from homogeneous cell line populations. We propose a novel framework that tracks cancer cells to capture their behavior and quantify cell viability to inform clinical decisions in a high-throughput manner. RESULTS: The brightfield microscopy images a large number of patient-derived cells in an ex vivo reconstruction of the tumor microenvironment treated with 31 drugs for up to 6 days. We developed a robust and user-friendly pipeline CancerCellTracker that detects cells in co-culture, tracks these cells across time and identifies cell death events using changes in cell attributes. We validated our computational pipeline by comparing the timing of cell death estimates by CancerCellTracker from brightfield images and a fluorescent channel featuring ethidium homodimer. We benchmarked our results using a state-of-the-art algorithm implemented in ImageJ and previously published in the literature. We highlighted CancerCellTracker's efficiency in estimating the percentage of live cells in the presence of bone marrow stromal cells. AVAILABILITY AND IMPLEMENTATION: https://github.com/compbiolabucf/CancerCellTracker. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qibing Jiang, Praneeth Sudalagunta, Maria C. Silva, Rafael R. Canevarolo, Xiaohong Zhao, Khandakar Tanvir Ahmed, Raghunandan Reddy Alugubelli, Gabe DeAvila, Alexandre Tungesvik, Lia Perez, Robert A. Gatenby, Robert J. Gillies, Rachid Baz, Mark B. Meads, Kenneth H. Shain, Ariosto Siqueira Silva, Wei Zhang 0076
Bioinform.17
2022 APA-Scan: detection and visualization of 3′-UTR alternative polyadenylation with RNA-seq and 3′-end-seq data
abstract
BACKGROUND: The eukaryotic genome is capable of producing multiple isoforms from a gene by alternative polyadenylation (APA) during pre-mRNA processing. APA in the 3'-untranslated region (3'-UTR) of mRNA produces transcripts with shorter or longer 3'-UTR. Often, 3'-UTR serves as a binding platform for microRNAs and RNA-binding proteins, which affect the fate of the mRNA transcript. Thus, 3'-UTR APA is known to modulate translation and provides a mean to regulate gene expression at the post-transcriptional level. Current bioinformatics pipelines have limited capability in profiling 3'-UTR APA events due to incomplete annotations and a low-resolution analyzing power: widely available bioinformatics pipelines do not reference actionable polyadenylation (cleavage) sites but simulate 3'-UTR APA only using RNA-seq read coverage, causing false positive identifications. To overcome these limitations, we developed APA-Scan, a robust program that identifies 3'-UTR APA events and visualizes the RNA-seq short-read coverage with gene annotations. METHODS: APA-Scan utilizes either predicted or experimentally validated actionable polyadenylation signals as a reference for polyadenylation sites and calculates the quantity of long and short 3'-UTR transcripts in the RNA-seq data. APA-Scan works in three major steps: (i) calculate the read coverage of the 3'-UTR regions of genes; (ii) identify the potential APA sites and evaluate the significance of the events among two biological conditions; (iii) graphical representation of user specific event with 3'-UTR annotation and read coverage on the 3'-UTR regions. APA-Scan is implemented in Python3. Source code and a comprehensive user's manual are freely available at https://github.com/compbiolabucf/APA-Scan . RESULT: APA-Scan was applied to both simulated and real RNA-seq datasets and compared with two widely used baselines DaPars and APAtrap. In simulation APA-Scan significantly improved the accuracy of 3'-UTR APA identification compared to the other baselines. The performance of APA-Scan was also validated by 3'-end-seq data and qPCR on mouse embryonic fibroblast cells. The experiments confirm that APA-Scan can detect unannotated 3'-UTR APA events and improve genome annotation. CONCLUSION: APA-Scan is a comprehensive computational pipeline to detect transcriptome-wide 3'-UTR APA events. The pipeline integrates both RNA-seq and 3'-end-seq data information and can efficiently identify the significant events with a high-resolution short reads coverage plots.
Naima Ahmed Fahmi, Khandakar Tanvir Ahmed, Jae-Woong Chang, Heba Nassereddeen, Deliang Fan, Jeongsik Yong, Wei Zhang 0076
BMC Bioinform.7
2021 PIM-Quantifier: A Processing-in-Memory Platform for mRNA Quantification
abstract
Processing-in-memory (PIM) architecture has been considered as a promising solution for the “memory-wall” issue in many data-intensive applications, especially in bioinformatics. Recent works of developing PIM for genome alignment and assembling have achieved tremendous improvement, while another important genome analysis - mRNA quantification has not been explored. Efficient and accurate mRNA quantification is a crucial step for molecular signature identification, disease outcome prediction and drug development. In this paper, for the first time, we propose a SOT-MRAM based PIM platform, named PIM-Quantifier, for efficient mRNA quantification. A PIM-friendly alignment-free quantification algorithm is first proposed. Then, we present the optimized PIM architecture/circuit designs and mapping method to efficiently accelerate mRNA quantification. Extensive experiments show that PIM-Quantifier significantly improves mRNA quantification performance than CPU and recent other PIM platforms in efficiency defined as throughput/power.
Fan Zhang 0069, Shaahin Angizi, Naima Ahmed Fahmi, Wei Zhang 0076, Deliang Fan
DAC4
2021 Multi-Armed Bandit Based Feature Selection
abstract
Effective feature selection can help reduce dimensionality, improve prediction accuracy, and increase result comprehensibility. Classic feature selection methods typically select and test feature subset in multiple iterations, and thus can be regarded as an exploratory process. In recent literature, the multi-armed bandit has become an emerging method to automate exploration for searching optimal solutions in large spaces. In this paper, our research question is: Can the multi-armed bandit formulation help us to automate feature selection? Along this line, we reformulate the feature selection problem with the combinatorial multi-armed bandit (CMAB) framework by regarding each feature as an arm. We propose two novel oracles and investigate how the super arm is formed under different oracles, and how the coordination between various features can be improved by a novel reward scheme. We present extensive experimental results to demonstrate the improved performance of the proposed methods over conventional feature selection approaches.
Kunpeng Liu 0001, Wei Zhang 0076, Ahmad Hariri, Yanjie Fu, Kien A. Hua
SDM3
2021 In silico model for miRNA-mediated regulatory network in cancer
abstract
Deregulation of gene expression is associated with the pathogenesis of numerous human diseases including cancer. Current data analyses on gene expression are mostly focused on differential gene/transcript expression in big data-driven studies. However, a poor connection to the proteome changes is a widespread problem in current data analyses. This is partly due to the complexity of gene regulatory pathways at the post-transcriptional level. In this study, we overcome these limitations and introduce a graph-based learning model, PTNet, which simulates the microRNAs (miRNAs) that regulate gene expression post-transcriptionally in silico. Our model does not require large-scale proteomics studies to measure the protein expression and can successfully predict the protein levels by considering the miRNA-mRNA interaction network, the mRNA expression, and the miRNA expression. Large-scale experiments on simulations and real cancer high-throughput datasets using PTNet validated that (i) the miRNA-mediated interaction network affects the abundance of corresponding proteins and (ii) the predicted protein expression has a higher correlation with the proteomics data (ground-truth) than the mRNA expression data. The classification performance also shows that the predicted protein expression has an improved prediction power on cancer outcomes compared to the prediction done by the mRNA expression data only or considering both mRNA and miRNA. Availability: PTNet toolbox is available at http://github.com/CompbioLabUCF/PTNet.
Khandakar Tanvir Ahmed, Jiao Sun, Irene Martinez, Sze Cheng, Wencai Zhang, Jeongsik Yong, Wei Zhang 0076
Briefings Bioinform.8
2021 Multi-omics data integration by generative adversarial network
abstract
MOTIVATION: Accurate disease phenotype prediction plays an important role in the treatment of heterogeneous diseases like cancer in the era of precision medicine. With the advent of high throughput technologies, more comprehensive multi-omics data is now available that can effectively link the genotype to phenotype. However, the interactive relation of multi-omics datasets makes it particularly challenging to incorporate different biological layers to discover the coherent biological signatures and predict phenotypic outcomes. In this study, we introduce omicsGAN, a generative adversarial network model to integrate two omics data and their interaction network. The model captures information from the interaction network as well as the two omics datasets and fuse them to generate synthetic data with better predictive signals. RESULTS: Large-scale experiments on The Cancer Genome Atlas breast cancer, lung cancer and ovarian cancer datasets validate that (i) the model can effectively integrate two omics data (e.g. mRNA and microRNA expression data) and their interaction network (e.g. microRNA-mRNA interaction network). The synthetic omics data generated by the proposed model has a better performance on cancer outcome classification and patients survival prediction compared to original omics datasets. (ii) The integrity of the interaction network plays a vital role in the generation of synthetic data with higher predictive quality. Using a random interaction network does not allow the framework to learn meaningful information from the omics datasets; therefore, results in synthetic data with weaker predictive signals. AVAILABILITY AND IMPLEMENTATION: Source code is available at: https://github.com/CompbioLabUCF/omicsGAN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Khandakar Tanvir Ahmed, Jiao Sun, Sze Cheng, Jeongsik Yong, Wei Zhang 0076
Bioinform.5
2020 PIM-Assembler: A Processing-in-Memory Platform for Genome Assembly
abstract
In this paper, for the first time, we propose a high-throughput and energy-efficient Processing-in-DRAM-accelerated genome assembler called PIM-Assembler based on an optimized and hardware-friendly genome assembly algorithm. PIM-Assembler can assemble large-scale DNA sequence dataset from all-pair overlaps. We first develop PIM-Assembler platform that harnesses DRAM as computational memory and transforms it to a fundamental processing unit for genome assembly. PIM-Assembler can perform efficient X(N)OR-based operations inside DRAM incurring low cost on top of commodity DRAM designs (~5% of chip area). PIM-Assembler is then optimized through a correlated data partitioning and mapping methodology that allows local storage and processing of DNA short reads to fully exploit the genome assembly algorithm-level's parallelism. The simulation results show that PIM-Assembler achieves on average 8.4× and 2.3 wise× higher throughput for performing bulk bit-XNOR-based comparison operations compared with CPU and recent processing-in-DRAM platforms, respectively. As for comparison/addition-extensive genome assembly application, it reduces the execution time and power by ~5× and ~ 7.5× compared to GPU.
Shaahin Angizi, Naima Ahmed Fahmi, Wei Zhang 0076, Deliang Fan
DAC3
2020 PIM-Aligner: A Processing-in-MRAM Platform for Biological Sequence Alignment
abstract
In this paper, we propose a high-throughput and energy-efficient Processing-in-Memory accelerator (PIM-Aligner) to execute DNA short read alignment based on an optimized and hardware-friendly alignment algorithm. We first reconstruct the existing sequence alignment algorithm based on BWT and FM-index such that it can be fully implemented in PIM platforms. It supports exact alignment and also handles mismatches to reduce excessive backtracking. We then develop PIM-Aligner platform that transforms SOT-MRAM array to a potential computational memory to accelerate the reconstructed alignment-in-memory algorithm incurring a low cost on top of original SOT-MRAM chips (less than 10% of chip area). Accordingly, we present a local data partitioning, mapping, and pipeline technique to maximize the parallelism in multiple computational sub-array while doing the alignment task. The simulation results show that PIM-Aligner outperforms recent platforms based on dynamic programming with ~ 3.1× higher throughput per Watt. Besides, PIM-Aligner improves the short read alignment throughput per Watt per mm2by ~ 9× and 1.9× compared to FM-index-based ASIC and processing-in-ReRAM designs, respectively.
Shaahin Angizi, Jiao Sun, Wei Zhang 0076, Deliang Fan
DATE3
2020 Exploring DNA Alignment-in-Memory Leveraging Emerging SOT-MRAM
abstract
In this work, we review two alternative Processing-in-Memory (PIM) accelerators based on Spin-Orbit-Torque Magnetic Random Access Memory (SOT-MRAM) to execute DNA short read alignment based on an optimized and hardware-friendly alignment algorithm. We first discuss the reconstruction of the existing sequence alignment algorithm based on BWT and FM-index such that it can be fully implemented leveraging PIM functions. We then transform SOT-MRAM array to a potential computational memory by presenting two different reconfigurable sense amplifiers to accelerate the reconstructed alignment-in-memory algorithm. The cross-layer simulation results show that such PIM platforms are able to achieve a nearly ten-fold and two-fold increases in throughput/power/area measure compared with recent ASIC and processing-in-ReRAM designs, respectively.
Shaahin Angizi, Wei Zhang 0076, Deliang Fan
ACM Great Lakes Symposium on VLSI2
2020 Platform-integrated mRNA isoform quantification
abstract
MOTIVATION: Accurate estimation of transcript isoform abundance is critical for downstream transcriptome analyses and can lead to precise molecular mechanisms for understanding complex human diseases, like cancer. Simplex mRNA Sequencing (RNA-Seq) based isoform quantification approaches are facing the challenges of inherent sampling bias and unidentifiable read origins. A large-scale experiment shows that the consistency between RNA-Seq and other mRNA quantification platforms is relatively low at the isoform level compared to the gene level. In this project, we developed a platform-integrated model for transcript quantification (IntMTQ) to improve the performance of RNA-Seq on isoform expression estimation. IntMTQ, which benefits from the mRNA expressions reported by the other platforms, provides more precise RNA-Seq-based isoform quantification and leads to more accurate molecular signatures for disease phenotype prediction. RESULTS: In the experiments to assess the quality of isoform expression estimated by IntMTQ, we designed three tasks for clustering and classification of 46 cancer cell lines with four different mRNA quantification platforms, including newly developed NanoString's nCounter technology. The results demonstrate that the isoform expressions learned by IntMTQ consistently provide more and better molecular features for downstream analyses compared with five baseline algorithms which consider RNA-Seq data only. An independent RT-qPCR experiment on seven genes in twelve cancer cell lines showed that the IntMTQ improved overall transcript quantification. The platform-integrated algorithms could be applied to large-scale cancer studies, such as The Cancer Genome Atlas (TCGA), with both RNA-Seq and array-based platforms available. AVAILABILITY AND IMPLEMENTATION: Source code is available at: https://github.com/CompbioLabUcf/IntMTQ. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiao Sun, Jae-Woong Chang, Teng Zhang 0002, Jeongsik Yong, Rui Kuang, Wei Zhang 0076
Bioinform.6
2020 Network-based multi-task learning models for biomarker selection and cancer outcome prediction
abstract
MOTIVATION: Detecting cancer gene expression and transcriptome changes with mRNA-sequencing or array-based data are important for understanding the molecular mechanisms underlying carcinogenesis and cellular events during cancer progression. In previous studies, the differentially expressed genes were detected across patients in one cancer type. These studies ignored the role of mRNA expression changes in driving tumorigenic mechanisms that are either universal or specific in different tumor types. To address the problem, we introduce two network-based multi-task learning frameworks, NetML and NetSML, to discover common differentially expressed genes shared across different cancer types as well as differentially expressed genes specific to each cancer type. The proposed frameworks consider the common latent gene co-expression modules and gene-sample biclusters underlying the multiple cancer datasets to learn the knowledge crossing different tumor types. RESULTS: Large-scale experiments on simulations and real cancer high-throughput datasets validate that the proposed network-based multi-task learning frameworks perform better sample classification compared with the models without the knowledge sharing across different cancer types. The common and cancer-specific molecular signatures detected by multi-task learning frameworks on The Cancer Genome Atlas ovarian, breast and prostate cancer datasets are correlated with the known marker genes and enriched in cancer-relevant Kyoto Encyclopedia of Genes and Genome pathways and gene ontology terms. AVAILABILITY AND IMPLEMENTATION: Source code is available at: https://github.com/compbiolabucf/NetML. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhezhi He, Milan Shah, Teng Zhang 0002, Deliang Fan, Wei Zhang 0076
Bioinform.6
2019 AlignS: A Processing-In-Memory Accelerator for DNA Short Read Alignment Leveraging SOT-MRAM
abstract
Classified as a complex big data analytics problem, DNA short read alignment serves as a major sequential bottleneck to massive amounts of data generated by next-generation sequencing platforms. With Von-Neumann computing architectures struggling to address such computationally-expensive and memory-intensive task today, Processing-in-Memory (PIM) platforms are gaining growing interests. In this paper, an energy-efficient and parallel PIM accelerator (AlignS) is proposed to execute DNA short read alignment based on an optimized and hardware-friendly alignment algorithm. We first develop AlignS platform that harnesses SOT-MRAM as computational memory and transforms it to a fundamental processing unit for short read alignment. Accordingly, we present a novel, customized, highly parallel read alignment algorithm that only seeks the proposed simple and parallel in-memory operations (i.e. comparisons and additions). AlignS is then optimized through a new correlated data partitioning and mapping methodology that allows local storage and processing of DNA sequence to fully exploit the algorithm-level's parallelism, and to accelerate both exact and inexact matches. The device-to-architecture co-simulation results show that AlignS improves the short read alignment throughput per Watt per mm2 by ~12× compared to the ASIC accelerator. Compared to recent FM-index-based ReRAM platform, AlignS achieves 1.6× higher throughput per Watt.
Shaahin Angizi, Jiao Sun, Wei Zhang 0076, Deliang Fan
DAC3
2019 GraphS: A Graph Processing Accelerator Leveraging SOT-MRAM
abstract
In this work, we present GraphS architecture, which transforms current Spin Orbit Torque Magnetic Random Access Memory (SOT-MRAM) to massively parallel computational units capable of accelerating graph processing applications. GraphS can be leveraged to greatly reduce energy consumption dealing with underlying adjacency matrix computations, eliminating unnecessary off-chip accesses and providing ultra-high internal bandwidth. The device-to-architecture co-simulation for three social network data-sets indicate roughly 3.6× higher energy-efficiency and 5.3× speed-up over recent ReRAM crossbar. It achieves ~4× higher energy-efficiency and 5.1× speed-up over recent processing-in-DRAM acceleration methods.
Shaahin Angizi, Jiao Sun, Wei Zhang 0076, Deliang Fan
DATE3
2019 Learning a Low-Rank Tensor of Pharmacogenomic Multi-relations from Biomedical Networks
abstract
Learning pharmacogenomic multi-relations among diseases, genes and chemicals from content-rich biomedical and biological networks can provide important guidance for drug discovery, drug repositioning and disease treatment. Most of the existing methods focus on imputing missing values in the disease-gene, disease chemical and gene-chemical pairwise relations from the observed relations instead of being designed for learning high-order disease-gene-chemical multi-relations. To achieve the goal, we propose a general tensor-based optimization framework and a scalable Graph-Regularized Tensor Completion from Observed Pairwise Relations (GT-COPR) algorithm to infer the multi-relations among the entities across multiple networks in a low-rank tensor, based on manifold regularization with the graph Laplacian of a Cartesian, tensor or strong product of the networks, and consistencies between the collapsed tensors and the observed bipartite relations. Our theoretical analyses also prove the convergence and efficiency of GT-COPR. In the experiments, the tensor fiber-wise and slice-wise evaluations demonstrate the accuracy of GT-COPR for predicting the diseasegene-chemical associations across the large-scale protein-protein interactions network, chemical structural similarity network and phenotype-based human disease network; and the validation on Genomics of Drug Sensitivity in Cancer cell line dataset shows a potential clinical application of GT-COPR for learning diseasespecific chemical-gene interactions. Statistical enrichment analysis demonstrates that GT-COPR is also capable of producing both topologically and biologically relevant disease, gene and chemical components with high significance.
Zhuliu Li, Wei Zhang 0076, Rong Stephanie Huang, Rui Kuang
ICDM2
2015 Network-Based Isoform Quantification with RNA-Seq Data for Cancer Transcriptome Analysis
abstract
High-throughput mRNA sequencing (RNA-Seq) is widely used for transcript quantification of gene isoforms. Since RNA-Seq data alone is often not sufficient to accurately identify the read origins from the isoforms for quantification, we propose to explore protein domain-domain interactions as prior knowledge for integrative analysis with RNA-Seq data. We introduce a Network-based method for RNA-Seq-based Transcript Quantification (Net-RSTQ) to integrate protein domain-domain interaction network with short read alignments for transcript abundance estimation. Based on our observation that the abundances of the neighboring isoforms by domain-domain interactions in the network are positively correlated, Net-RSTQ models the expression of the neighboring transcripts as Dirichlet priors on the likelihood of the observed read alignments against the transcripts in one gene. The transcript abundances of all the genes are then jointly estimated with alternating optimization of multiple EM problems. In simulation Net-RSTQ effectively improved isoform transcript quantifications when isoform co-expressions correlate with their interactions. qRT-PCR results on 25 multi-isoform genes in a stem cell line, an ovarian cancer cell line, and a breast cancer cell line also showed that Net-RSTQ estimated more consistent isoform proportions with RNA-Seq data. In the experiments on the RNA-Seq data in The Cancer Genome Atlas (TCGA), the transcript abundances estimated by Net-RSTQ are more informative for patient sample classification of ovarian cancer, breast cancer and lung cancer. All experimental results collectively support that Net-RSTQ is a promising approach for isoform quantification. Net-RSTQ toolbox is available at http://compbio.cs.umn.edu/Net-RSTQ/.
Wei Zhang 0076, Jae-Woong Chang, Lilong Lin, Kay Minn, Baolin Wu, Jeremy Chien, Jeongsik Yong, Rui Kuang
PLoS Comput. Biol.1
2013 Network-based Survival Analysis Reveals Subnetwork Signatures for Predicting Outcomes of Ovarian Cancer Treatment
abstract
Cox regression is commonly used to predict the outcome by the time to an event of interest and in addition, identify relevant features for survival analysis in cancer genomics. Due to the high-dimensionality of high-throughput genomic data, existing Cox models trained on any particular dataset usually generalize poorly to other independent datasets. In this paper, we propose a network-based Cox regression model called Net-Cox and applied Net-Cox for a large-scale survival analysis across multiple ovarian cancer datasets. Net-Cox integrates gene network information into the Cox's proportional hazard model to explore the co-expression or functional relation among high-dimensional gene expression features in the gene network. Net-Cox was applied to analyze three independent gene expression datasets including the TCGA ovarian cancer dataset and two other public ovarian cancer datasets. Net-Cox with the network information from gene co-expression or functional relations identified highly consistent signature genes across the three datasets, and because of the better generalization across the datasets, Net-Cox also consistently improved the accuracy of survival prediction over the Cox models regularized by L(2) or L(1). This study focused on analyzing the death and recurrence outcomes in the treatment of ovarian carcinoma to identify signature genes that can more reliably predict the events. The signature genes comprise dense protein-protein interaction subnetworks, enriched by extracellular matrix receptors and modulators or by nuclear signaling components downstream of extracellular signal-regulated kinases. In the laboratory validation of the signature genes, a tumor array experiment by protein staining on an independent patient cohort from Mayo Clinic showed that the protein expression of the signature gene FBN1 is a biomarker significantly associated with the early recurrence after 12 months of the treatment in the ovarian cancer patients who are initially sensitive to chemotherapy. Net-Cox toolbox is available at http://compbio.cs.umn.edu/Net-Cox/.
Wei Zhang 0076, Takayo Ota, Viji Shridhar, Jeremy Chien, Baolin Wu, Rui Kuang
PLoS Comput. Biol.1
2011 Inferring disease and gene set associations with rank coherence in networks
abstract
MOTIVATION: To validate the candidate disease genes identified from high-throughput genomic studies, a necessary step is to elucidate the associations between the set of candidate genes and disease phenotypes. The conventional gene set enrichment analysis often fails to reveal associations between disease phenotypes and the gene sets with a short list of poorly annotated genes, because the existing annotations of disease-causative genes are incomplete. This article introduces a network-based computational approach called rcNet to discover the associations between gene sets and disease phenotypes. A learning framework is proposed to maximize the coherence between the predicted phenotype-gene set relations and the known disease phenotype-gene associations. An efficient algorithm coupling ridge regression with label propagation and two variants are designed to find the optimal solution to the objective functions of the learning framework. RESULTS: We evaluated the rcNet algorithms with leave-one-out cross-validation on Online Mendelian Inheritance in Man (OMIM) data and an independent test set of recently discovered disease-gene associations. In the experiments, the rcNet algorithms achieved best overall rankings compared with the baselines. To further validate the reproducibility of the performance, we applied the algorithms to identify the target diseases of novel candidate disease genes obtained from recent studies of Genome-Wide Association Study (GWAS), DNA copy number variation analysis and gene expression profiling. The algorithms ranked the target disease of the candidate genes at the top of the rank list in many cases across all the three case studies. AVAILABILITY: http://compbio.cs.umn.edu/dgsa_rcNet CONTACT: [email protected].
Taehyun Hwang, Wei Zhang 0076, Maoqiang Xie, Jinfeng Liu 0003, Rui Kuang
Bioinform.2