EDBT 2026 Demo / reviewers in the wild / expert
Yufei Huang 0001
dblp:68/1946-1
· DBLP profile ↗
64ranked-venue papers
11as first author
10since 2021 · last 2026
0000-0001-6268-5357ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 38 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 8Computer networks · 4 · 3 first-authorArtificial intelligence and machine learning · 3 · 1 since 2021Security and privacy · 3 · 3 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ST2HE: enhancing spatial transcriptomics interpretability via virtual staining for histological annotationabstractHigh-resolution spatial transcriptomics (HR-ST) technologies offer unprecedented insights into tissue architecture but lack standardized frameworks for histological annotation. We present ST2HE, a cross-platform generative framework that synthesizes virtual hematoxylin and eosin images directly from HR-ST data. ST2HE integrates nuclei morphology and spatial transcript coordinates using a one-step diffusion model, enabling histologically informative image generation across diverse tissue types and HR-ST platforms. Conditional and tissue-independent variants support both known and novel tissue contexts. Evaluations on breast cancer, non-small cell lung cancer, and Kaposi's sarcoma demonstrate ST2HE's ability to preserve morphological features and support downstream annotations of tissue histology and phenotype classification. Ablation studies reveal that larger context windows, balanced loss functions, and multi-colored transcript visualization enhance image fidelity. ST2HE bridges molecular and histological domains, enabling interpretable, scalable annotation of HR-ST data and advancing computational pathology. Arun Das 0001, Wen Meng, Yu-Chiao Chiu, Shou-Jiang Gao, Yufei Huang 0001 |
Briefings Bioinform. | 6 |
| 2025 | A comparative evaluation of computational models for RNA modification detection using nanopore sequencing with RNA004 chemistryabstractDirect RNA sequencing from Oxford Nanopore Technologies has become a valuable method for studying RNA modifications such as N6-methyladenosine (m6A) and pseudouridine (pseU). Recent advancements in the RNA004 chemistry substantially reduce sequencing errors compared to previous chemistries, promising enhanced accuracy for epitranscriptomic analysis. Here we benchmark the performance of two RNA modification detection models for RNA004 data, Dorado and m6Anet, using two wild-type (WT) cell lines (HEK293T and HeLa), with respective ground truths from GLORI and eTAM-seq, and in vitro transcribed (IVT) RNA as negative controls. We found that for m6A sites with ≥10% modification ratio and ≥ 10X coverage, Dorado has higher recall (~0.92) than m6Anet (~0.51). Among true positive predictions, there are high correlations of m6A modification stoichiometry (correlation coefficient of ~0.89 for Dorado-truth and ~ 0.72 for m6Anet-truth). However, combined assessment of WT and IVT datasets show that while the per-site false positive rate can be lower (~8% for Dorado and ~ 33% for m6Anet), both tools can have high per-site false discovery rate of m6A (~40% for Dorado and ~ 80% for m6Anet), or for pseU (~95% for Dorado). Motif analysis reveals that both tools exhibit high heterogeneity of false positive calls across sequence contexts. There is also a substantial overlap of false positive calls between the two IVT samples, suggesting a filtering strategy by compiling a set of low-confidence sites from diverse IVT samples. Our analysis highlights key strengths and limitations of the current generation of m6A detection algorithms and offers insights into optimizing thresholds and interpretability. Yongji Zou, Mian Umair Ahsan, Joe Chan, Wen Meng, Shou-Jiang Gao, Yufei Huang 0001, Kai Wang 0049 |
Briefings Bioinform. | 6 |
| 2024 | Understanding YTHDF2-mediated mRNA degradation by m6A-BERT-DegabstractN6-methyladenosine (m6A) is the most abundant mRNA modification within mammalian cells, holding pivotal significance in the regulation of mRNA stability, translation and splicing. Furthermore, it plays a critical role in the regulation of RNA degradation by primarily recruiting the YTHDF2 reader protein. However, the selective regulation of mRNA decay of the m6A-methylated mRNA through YTHDF2 binding is poorly understood. To improve our understanding, we developed m6A-BERT-Deg, a BERT model adapted for predicting YTHDF2-mediated degradation of m6A-methylated mRNAs. We meticulously assembled a high-quality training dataset by integrating multiple data sources for the HeLa cell line. To overcome the limitation of small training samples, we employed a pre-training-fine-tuning strategy by first performing a self-supervised pre-training of the model on 427 760 unlabeled m6A site sequences. The test results demonstrated the importance of this pre-training strategy in enabling m6A-BERT-Deg to outperform other benchmark models. We further conducted a comprehensive model interpretation and revealed a surprising finding that the presence of co-factors in proximity to m6A sites may disrupt YTHDF2-mediated mRNA degradation, subsequently enhancing mRNA stability. We also extended our analyses to the HEK293 cell line, shedding light on the context-dependent YTHDF2-mediated mRNA degradation. Tinghe Zhang, Sumin Jo, Michelle Zhang, Kai Wang 0049, Shou-Jiang Gao, Yufei Huang 0001 |
Briefings Bioinform. | 6 |
| 2023 | Fair patient model: Mitigating bias in the patient representation learned from the electronic health records
Sonish Sivarajkumar, Yufei Huang 0001, Yanshan Wang |
J. Biomed. Informatics | 2 |
| 2022 | Toward Deep Learning Based Access ControlabstractA common trait of current access control approaches is the challenging need to engineer abstract and intuitive access control models. This entails designing access control information in the form of roles (RBAC), attributes (ABAC), or relationships (ReBAC) as the case may be, and subsequently, designing access control rules. This framework has its benefits but has significant limitations in the context of modern systems that are dynamic, complex, and large-scale, due to which it is difficult to maintain an accurate access control state in the system for a human administrator. This paper proposes Deep Learning Based Access Control (DLBAC) by leveraging significant advances in deep learning technology as a potential solution to this problem. We envision that DLBAC could complement and, in the long-term, has the potential to even replace, classical access control models with a neural network that reduces the burden of access control model engineering and updates. Without loss of generality, we conduct a thorough investigation of a candidate DLBAC model, called DLBAC_alpha, using both real-world and synthetic datasets. We demonstrate the feasibility of the proposed approach by addressing issues related to accuracy, generalization, and explainability. We also discuss challenges and future research directions. Mohammad Nur Nobi, Ram Krishnan, Yufei Huang 0001, Mehrnoosh Shakarami, Ravi S. Sandhu |
CODASPY | 3 |
| 2022 | Administration of Machine Learning Based Access Control
Mohammad Nur Nobi, Ram Krishnan, Yufei Huang 0001, Ravi S. Sandhu |
ESORICS (2) | 3 |
| 2022 | Deep learning tackles single-cell analysis - a survey of deep learning for scRNA-seq analysisabstractSince its selection as the method of the year in 2013, single-cell technologies have become mature enough to provide answers to complex research questions. With the growth of single-cell profiling technologies, there has also been a significant increase in data collected from single-cell profilings, resulting in computational challenges to process these massive and complicated datasets. To address these challenges, deep learning (DL) is positioned as a competitive alternative for single-cell analyses besides the traditional machine learning approaches. Here, we survey a total of 25 DL algorithms and their applicability for a specific step in the single cell RNA-seq processing pipeline. Specifically, we establish a unified mathematical representation of variational autoencoder, autoencoder, generative adversarial network and supervised DL models, compare the training strategies and loss functions for these models, and relate the loss functions of these models to specific objectives of the data processing step. Such a presentation will allow readers to choose suitable algorithms for their particular objective at each step in the pipeline. We envision that this survey will serve as an important information portal for learning the application of DL for scRNA-seq analysis and inspire innovative uses of DL to address a broader range of new challenges in emerging multi-omics and spatial single-cell sequencing. Mario Flores, Tinghe Zhang, Md Musaddaqui Hasib, Yu-Chiao Chiu, Zhenqing Ye, Karla Paniagua, Sumin Jo, Jianqiu Zhang 0002, Shou-Jiang Gao, Yu-Fang Jin, Yidong Chen 0002, Yufei Huang 0001 |
Briefings Bioinform. | 13 |
| 2021 | Interpretable Self-Supervised Facial Micro-Expression Learning to Predict Cognitive State and Neurological DisordersabstractHuman behavior is the confluence of output from voluntary and involuntary motor systems. The neural activities that mediate behavior, from individual cells to distributed networks, are in a state of constant flux. Artificial intelligence (AI) research over the past decade shows that behavior, in the form of facial muscle activity, can reveal information about fleeting voluntary and involuntary motor system activity related to emotion, pain, and deception. However, the AI algorithms often lack an explanation for their decisions, and learning meaningful representations requires large datasets labeled by a subject-matter expert. Motivated by the success of using facial muscle movements to classify brain states and the importance of learning from small amounts of data, we propose an explainable self-supervised representation-learning paradigm that learns meaningful temporal facial muscle movement patterns from limited samples. We validate our methodology by carrying out comprehensive empirical study to predict future speech behavior in a real-world dataset of adults who stutter (AWS). Our explainability study found facial muscle movements around the eyes (p Arun Das 0001, Jeffrey Mock, Yufei Huang 0001, Edward J. Golob, Peyman Najafirad |
AAAI | 3 |
| 2021 | Access Control Policy Generation from User Stories Using Machine Learning
John Heaps, Ram Krishnan, Yufei Huang 0001, Jianwei Niu 0001, Ravi S. Sandhu |
DBSec | 3 |
| 2021 | CancerSiamese: one-shot learning for predicting primary and metastatic tumor types unseen during model trainingabstractBACKGROUND: The state-of-the-art deep learning based cancer type prediction can only predict cancer types whose samples are available during the training where the sample size is commonly large. In this paper, we consider how to utilize the existing training samples to predict cancer types unseen during the training. We hypothesize the existence of a set of type-agnostic expression representations that define the similarity/dissimilarity between samples of the same/different types and propose a novel one-shot learning model called CancerSiamese to learn this common representation. CancerSiamese accepts a pair of query and support samples (gene expression profiles) and learns the representation of similar or dissimilar cancer types through two parallel convolutional neural networks joined by a similarity function. RESULTS: We trained CancerSiamese for cancer type prediction for primary and metastatic tumors using samples from the Cancer Genome Atlas (TCGA) and MET500. Network transfer learning was utilized to facilitate the training of the CancerSiamese models. CancerSiamese was tested for different N-way predictions and yielded an average accuracy improvement of 8% and 4% over the benchmark 1-Nearest Neighbor (1-NN) classifier for primary and metastatic tumors, respectively. Moreover, we applied the guided gradient saliency map and feature selection to CancerSiamese to examine 100 and 200 top marker-gene candidates for the prediction of primary and metastatic cancers, respectively. Functional analysis of these marker genes revealed several cancer related functions between primary and metastatic tumors. CONCLUSION: This work demonstrated, for the first time, the feasibility of predicting unseen cancer types whose samples are limited. Thus, it could inspire new and ingenious applications of one-shot and few-shot learning solutions for improving cancer diagnosis, prognostic, and our understanding of cancer. Milad Mostavi, Yu-Chiao Chiu, Yidong Chen 0002, Yufei Huang 0001 |
BMC Bioinform. | 4 |
| 2020 | Automatic Detection and Prediction of Cybersickness Severity using Deep Neural Networks from user's Physiological SignalsabstractCybersickness is one of the primary challenges to the usability and acceptability of virtual reality (VR). Cybersickness can cause motion sickness-like discomforts, including disorientation, headache, nausea, and fatigue, both during and after the VR immersion. Prior research suggested a significant correlation between physiological signals and cybersickness severity, as measured by the simulator sickness questionnaire (SSQ). However, SSQ may not be suitable for automatic detection of cybersickness severity during immersion, as it is usually reported before and after the immersion. In this study, we introduced an automated approach for the detection and prediction of cybersickness severity from the user's physiological signals. We collected heart rate, breathing rate, heart rate variability, and galvanic skin response data from 31 healthy participants while immersed in a VR roller coaster simulation. We found a significant difference in the participants' physiological signals during their cybersickness state compared to their resting baseline. We compared a support vector machine classifier and three deep neural classifiers for cybersickness severity detection and prediction in two minutes' future, given the previous two minutes of physiological signals. Our proposed simplified convolutional long short-term memory classifier achieved an accuracy of 97.44% for detecting current cybersickness severity and 87.38% for predicting future cybersickness severity from the physiological signals. Rifatul Islam, Yonggun Lee, Mehrad Jaloli, Imtiaz Muhammad, Dakai Zhu 0001, Peyman Najafirad, Yufei Huang 0001, John Quarles |
ISMAR | 7 |
| 2020 | Deep learning of pharmacogenomics resources: moving towards precision oncologyabstractThe recent accumulation of cancer genomic data provides an opportunity to understand how a tumor's genomic characteristics can affect its responses to drugs. This field, called pharmacogenomics, is a key area in the development of precision oncology. Deep learning (DL) methodology has emerged as a powerful technique to characterize and learn from rapidly accumulating pharmacogenomics data. We introduce the fundamentals and typical model architectures of DL. We review the use of DL in classification of cancers and cancer subtypes (diagnosis and treatment stratification of patients), prediction of drug response and drug synergy for individual tumors (treatment prioritization for a patient), drug repositioning and discovery and the study of mechanism/mode of action of treatments. For each topic, we summarize current genomics and pharmacogenomics data resources such as pan-cancer genomics data for cancer cell lines (CCLs) and tumors, and systematic pharmacologic screens of CCLs. By revisiting the published literature, including our in-house analyses, we demonstrate the unprecedented capability of DL enabled by rapid accumulation of data resources to decipher complex drug response patterns, thus potentially improving cancer medicine. Overall, this review provides an in-depth summary of state-of-the-art DL methods and up-to-date pharmacogenomics resources and future opportunities and challenges to realize the goal of precision oncology. Yu-Chiao Chiu, Hung-I Harry Chen, Aparna Gorthi, Milad Mostavi, Siyuan Zheng, Yufei Huang 0001, Yidong Chen 0002 |
Briefings Bioinform. | 6 |
| 2019 | Predicting Auditory Spatial Attention from EEG using Single- and Multi-task Convolutional Neural NetworksabstractRecent behavioral and electroencephalography (EEG) studies have defined ways that auditory spatial attention can be allocated over large regions of space. As with most experimental studies, behavior and EEG were averaged over 10s of minutes because identifying abstract feature spatial codes from raw EEG data is extremely challenging. The goal of this study is to design a deep learning model that can learn from raw EEG data and predict auditory spatial information on a trial-by-trial basis. We designed a convolutional neural network (CNN) model to predict the attended location or other stimulus locations relative to the attended location. A multi-task model was also used to predict the attended and stimulus locations at the same time. Based on the visualization of our models, we investigated features of individual classification tasks and joint feature of the multi-task model. Our model achieved an average 72.4% in relative location prediction and 90.0% in attended location prediction individually (AUROC's). The multi-task model improved the performance of attended location prediction by 3%. Our results show that deep learning methods are able to define abstract neural codes in EEG thought to neural mechanisms of human spatial cognition and attention. Jeffrey Mock, Yufei Huang 0001, Edward J. Golob |
SMC | 3 |
| 2019 | Target Classification in a Novel SSVEP-RSVP Based BCI Gaming SystemabstractRecently game-based brain-computer interface (BCI) systems using electroencephalography (EEG) has been gaining popularity, providing a sophisticated experience to its users. Here we present such a novel hybrid system based on rapid serial visual presentation (RSVP) in conjunction with steady-state visual evoked potentials (SSVEP). Based on a matching computer game Jewel Quest a game is designed wherein a sequence of jewel images containing rare targets (<; 3%) in an RSVP paradigm is presented on a display at four distinct locations each flickering at different rates (4, 5, 6 and 7 Hz). A score is awarded upon successful detection of target image from neural signals. During real-time implementation to achieve higher classification speeds, EEG signals were epoched at the onset of each image, creating a high degree of class overlap and imbalance. Given these challenges in our EEG datasets, we present classifiers that can classify single-trial EEG epochs at the onset of target image presentation accurately. Initial results from 14 subjects indicate Hidden Markov Model (HMM) with Dirichlet emission probabilities provide ~1% higher, on average, the area under the precision-recall curve (AUC-PR) compared to the ensemble technique Bagging, commonly used to handle class imbalance. Tapsya Nayak, Li-Wei Ko, Tzyy-Ping Jung, Yufei Huang 0001 |
SMC | 4 |
| 2019 | A Semi-Supervised Wasserstein Generative Adversarial Network for Classifying Driving Fatigue from EEG signalsabstractPredicting driver's cognitive states using deep learning from electroencephalography (EEG) signals is considered this paper. To address the challenge posed by limited labeled training samples, a semi-supervised Wasserstein Generative Adversarial Network with gradient penalty (sWGAN-GP) is proposed. The proposed sWGAN-GP includes a classifier with the shared architecture with the discriminator in GAN and its loss function enables the augmentation of limited training samples with generated EEG samples during training, thus resulting in improved classification performance. The several modeling challenges including frequency artifacts and training instability, are also considered. The test results on predicting the alert and drowsy states from a simulated driving experiment demonstrate improved prediction performance and training stability over the baseline semi-supervised GAN and a convolutional neural network model. Sharaj Panwar, Peyman Najafirad, John Quarles, Edward J. Golob, Yufei Huang 0001 |
SMC | 5 |
| 2019 | Generating EEG signals of an RSVP Experiment by a Class Conditioned Wasserstein Generative Adversarial NetworkabstractElectroencephalography (EEG) data is difficult to obtain due to complex experimental setups and reduced comfort due to prolonged wearing. This poses challenges to train powerful deep learning model due to the limited EEG data. Hence, being able to generate EEG data computationally is highly desirable. We propose a novel Conditional Wasserstein Generative Adversarial Network with gradient penalty (cWGAN-GP) that can be trained to synthesize EEG data for different cognitive events. This network addresses several modeling challenges, including frequency artifacts and training instability. The proposed GAN model is tested to generate one channel EEG data for the rapid serial visual presentation. We demonstrated the validity of the generated samples using several evaluation metrics and show that the synthesized EEG data can augment the real EEG data to achieve improved event classification performance. Sharaj Panwar, Peyman Najafirad, John Quarles, Yufei Huang 0001 |
SMC | 4 |
| 2019 | FunDMDeep-m6A: identification and prioritization of functional differential m6A methylation genesabstractMOTIVATION: As the most abundant mammalian mRNA methylation, N6-methyladenosine (m6A) exists in >25% of human mRNAs and is involved in regulating many different aspects of mRNA metabolism, stem cell differentiation and diseases like cancer. However, our current knowledge about dynamic changes of m6A levels and how the change of m6A levels for a specific gene can play a role in certain biological processes like stem cell differentiation and diseases like cancer is largely elusive. RESULTS: To address this, we propose in this paper FunDMDeep-m6A a novel pipeline for identifying context-specific (e.g. disease versus normal, differentiated cells versus stem cells or gene knockdown cells versus wild-type cells) m6A-mediated functional genes. FunDMDeep-m6A includes, at the first step, DMDeep-m6A a novel method based on a deep learning model and a statistical test for identifying differential m6A methylation (DmM) sites from MeRIP-Seq data at a single-base resolution. FunDMDeep-m6A then identifies and prioritizes functional DmM genes (FDmMGenes) by combing the DmM genes (DmMGenes) with differential expression analysis using a network-based method. This proposed network method includes a novel m6A-signaling bridge (MSB) score to quantify the functional significance of DmMGenes by assessing functional interaction of DmMGenes with their signaling pathways using a heat diffusion process in protein-protein interaction (PPI) networks. The test results on 4 context-specific MeRIP-Seq datasets showed that FunDMDeep-m6A can identify more context-specific and functionally significant FDmMGenes than m6A-Driver. The functional enrichment analysis of these genes revealed that m6A targets key genes of many important context-related biological processes including embryonic development, stem cell differentiation, transcription, translation, cell death, cell proliferation and cancer-related pathways. These results demonstrate the power of FunDMDeep-m6A for elucidating m6A regulatory functions and its roles in biological processes and diseases. AVAILABILITY AND IMPLEMENTATION: The R-package for DMDeep-m6A is freely available from https://github.com/NWPU-903PR/DMDeepm6A1.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Songyao Zhang, Shaowu Zhang 0001, Xiaonan Fan 0001, Jia Meng 0001, Yufei Huang 0001 |
Bioinform. | 6 |
| 2019 | Global analysis of N6-methyladenosine functions and its disease association using deep learning and network-based methodsabstractN6-methyladenosine (m6A) is the most abundant methylation, existing in >25% of human mRNAs. Exciting recent discoveries indicate the close involvement of m6A in regulating many different aspects of mRNA metabolism and diseases like cancer. However, our current knowledge about how m6A levels are controlled and whether and how regulation of m6A levels of a specific gene can play a role in cancer and other diseases is mostly elusive. We propose in this paper a computational scheme for predicting m6A-regulated genes and m6A-associated disease, which includes Deep-m6A, the first model for detecting condition-specific m6A sites from MeRIP-Seq data with a single base resolution using deep learning and Hot-m6A, a new network-based pipeline that prioritizes functional significant m6A genes and its associated diseases using the Protein-Protein Interaction (PPI) and gene-disease heterogeneous networks. We applied Deep-m6A and this pipeline to 75 MeRIP-seq human samples, which produced a compact set of 709 functionally significant m6A-regulated genes and nine functionally enriched subnetworks. The functional enrichment analysis of these genes and networks reveal that m6A targets key genes of many critical biological processes including transcription, cell organization and transport, and cell proliferation and cancer-related pathways such as Wnt pathway. The m6A-associated disease analysis prioritized five significantly associated diseases including leukemia and renal cell carcinoma. These results demonstrate the power of our proposed computational scheme and provide new leads for understanding m6A regulatory functions and its roles in diseases. Songyao Zhang, Shaowu Zhang 0001, Xiaonan Fan 0001, Jia Meng 0001, Yidong Chen 0002, Shou-Jiang Gao, Yufei Huang 0001 |
PLoS Comput. Biol. | 7 |
| 2018 | Malware Detection in Cloud Infrastructures Using Convolutional Neural NetworksabstractA major challenge in Infrastructure as a Service (IaaS) clouds is its exposure to malware. Malware can spread rapidly within a datacenter and can cause major disruption to a cloud service provider and its clients. This paper introduces and discusses an effective malware detection approach in cloud infrastructure using Convolutional Neural Network (CNN), a deep learning approach. We initially employ a standard 2d CNN by training on metadata available for each of the processes in a virtual machine (VM) obtained by means of the hypervisor. We enhance the CNN classifier accuracy by using a novel 3d CNN (where an input is a collection of samples over a time interval), which greatly helps reduce mislabelled samples during data collection and training. Our experiments are performed on data collected by running various malware (mostly Trojans and Rootkits) on VMs. The malware used in our experiments are randomly selected. This reduces the selection bias of known-to-be highly active malware for easy detection. We demonstrate that our 2d CNN model reaches an accuracy of ≃ 79%, and our 3d CNN model significantly improves the accuracy to ≃ 90%. Mahmoud Abdelsalam, Ram Krishnan, Yufei Huang 0001, Ravi S. Sandhu |
IEEE CLOUD | 3 |
| 2018 | CLIPSeed: Achieving High Precision miRNA Binding Sites Prediction using PAR-CLIP Data
Mingzhu Lu, Yufei Huang 0001 |
BIBM | 2 |
| 2018 | Base-pair resolution detection of transcription factor binding site by deep deconvolutional networkabstractMotivation: Transcription factor (TF) binds to the promoter region of a gene to control gene expression. Identifying precise TF binding sites (TFBSs) is essential for understanding the detailed mechanisms of TF-mediated gene regulation. However, there is a shortage of computational approach that can deliver single base pair resolution prediction of TFBS. Results: In this paper, we propose DeepSNR, a Deep Learning algorithm for predicting TF binding location at Single Nucleotide Resolution de novo from DNA sequence. DeepSNR adopts a novel deconvolutional network (deconvNet) model and is inspired by the similarity to image segmentation by deconvNet. The proposed deconvNet architecture is constructed on top of 'DeepBind' and we trained the entire model using TF-specific data from ChIP-exonuclease (ChIP-exo) experiments. DeepSNR has been shown to outperform motif search-based methods for several evaluation metrics. We have also demonstrated the usefulness of DeepSNR in the regulatory analysis of TFBS as well as in improving the TFBS prediction specificity using ChIP-seq data. Availability and implementation: DeepSNR is available open source in the GitHub repository (https://github.com/sirajulsalekin/DeepSNR). Supplementary information: Supplementary data are available at Bioinformatics online. Sirajul Salekin, Jianqiu Zhang 0002, Yufei Huang 0001 |
Bioinform. | 3 |
| 2018 | A Bayesian framework for the inference of gene regulatory networks from time and pseudo-time series dataabstractMotivation: Molecular profiling techniques have evolved to single-cell assays, where dense molecular profiles are screened simultaneously for each cell in a population. High-throughput single-cell experiments from a heterogeneous population of cells can be experimentally and computationally sorted as a sequence of samples pseudo-temporally ordered samples. The analysis of these datasets, comprising a large number of samples, has the potential to uncover the dynamics of the underlying regulatory programmes. Results: We present a novel approach for modelling and inferring gene regulatory networks from high-throughput time series and pseudo-temporally sorted single-cell data. Our method is based on a first-order autoregressive moving-average model and it infers the gene regulatory network within a variational Bayesian framework. We validate our method with synthetic data and we apply it to single cell qPCR and RNA-Seq data for mouse embryonic cells and hematopoietic cells in zebra fish. Availability and implementation: The method presented in this article is available at https://github.com/mscastillo/GRNVBEM. Contact: [email protected]. Manuel Sanchez-Castillo, Isabel M. Tienda-Luna, Maria Carmen Carrion Perez, Yufei Huang 0001 |
Bioinform. | 5 |
| 2018 | MeTDiff: A Novel Differential RNA Methylation Analysis for MeRIP-Seq DataabstractN6-Methyladenosine (m6A) transcriptome methylation is an exciting new research area that just captures the attention of research community. We present in this paper, MeTDiff, a novel computational tool for predicting differential m6A methylation sites from Methylated RNA immunoprecipitation sequencing (MeRIP-Seq) data. Compared with the existing algorithm exomePeak, the advantages of MeTDiff are that it explicitly models the reads variation in data and also devices a more power likelihood ratio test for differential methylation site prediction. Comprehensive evaluation of MeTDiff's performance using both simulated and real datasets showed that MeTDiff is much more robust and achieved much higher sensitivity and specificity over exomePeak. Lin Zhang 0015, Jia Meng 0001, Manjeet K. Rao, Yidong Chen 0002, Yufei Huang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2017 | Driver's fatigue prediction by deep covariance learning from EEGabstractWe present here deep covariance learning models for predicting drivers' drowsy and alert states from Electroencephalography (EEG). Three types of deep covariance learning models are proposed: SPDNet, CNN, and DNN on covariance matrices. Our test results show that all the deep covariance learning methods reported better performance than shallow learning methods including Riemannian methods and STCNN, a previously proposed CNN model for EEG classification. Among the deep covariance learning methods, the best classification performance is obtained by a CNN model applied on sample spatial EEG covariance matrices and it improved the AUC of the best shallow algorithm (logistic regression + Log-Euclidean Metric) by 12.32% from 70.96% to 86.14%. Our study showed that deep covariance learning is a very promising approach for drivers' fatigue prediction. Mehdi Hajinoroozi, Jianqiu Zhang 0002, Yufei Huang 0001 |
SMC | 3 |
| 2017 | QNB: differential RNA methylation analysis for count-based small-sample sequencing data with a quad-negative binomial modelabstractBACKGROUND: As a newly emerged research area, RNA epigenetics has drawn increasing attention recently for the participation of RNA methylation and other modifications in a number of crucial biological processes. Thanks to high throughput sequencing techniques, such as, MeRIP-Seq, transcriptome-wide RNA methylation profile is now available in the form of count-based data, with which it is often of interests to study the dynamics at epitranscriptomic layer. However, the sample size of RNA methylation experiment is usually very small due to its costs; and additionally, there usually exist a large number of genes whose methylation level cannot be accurately estimated due to their low expression level, making differential RNA methylation analysis a difficult task. RESULTS: We present QNB, a statistical approach for differential RNA methylation analysis with count-based small-sample sequencing data. Compared with previous approaches such as DRME model based on a statistical test covering the IP samples only with 2 negative binomial distributions, QNB is based on 4 independent negative binomial distributions with their variances and means linked by local regressions, and in the way, the input control samples are also properly taken care of. In addition, different from DRME approach, which relies only the input control sample only for estimating the background, QNB uses a more robust estimator for gene expression by combining information from both input and IP samples, which could largely improve the testing performance for very lowly expressed genes. CONCLUSION: A-Seq, Par-CLIP, RIP-Seq, etc. Shaowu Zhang 0001, Yufei Huang 0001, Jia Meng 0001 |
BMC Bioinform. | 3 |
| 2017 | Cancer Progression Prediction Using Gene Interaction Regularized Elastic NetabstractDifferent types of genomic aberration may simultaneously contribute to tumorigenesis. To obtain a more accurate prognostic assessment to guide therapeutic regimen choice for cancer patients, the heterogeneous multi-omics data should be integrated harmoniously, which can often be difficult. For this purpose, we propose a Gene Interaction Regularized Elastic Net (GIREN) model that predicts clinical outcome by integrating multiple data types. GIREN conveniently embraces both gene measurements and gene-gene interaction information under an elastic net formulation, enforcing structure sparsity, and the "grouping effect" in solution to select the discriminate features with prognostic value. An iterative gradient descent algorithm is also developed to solve the model with regularized optimization. GIREN was applied to human ovarian cancer and breast cancer datasets obtained from The Cancer Genome Atlas, respectively. Result shows that, the proposed GIREN algorithm obtained more accurate and robust performance over competing algorithms (LASSO, Elastic Net, and Semi-supervised PCA, with or without average pathway expression features) in predicting cancer progression on both two datasets in terms of median area under curve (AUC) and interquartile range (IQR), suggesting a promising direction for more effective integration of gene measurement and gene interaction information. Lin Zhang 0015, Hui Liu 0024, Yufei Huang 0001, Xuesong Wang 0001, Yidong Chen 0002, Jia Meng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2016 | A novel algorithm for calling mRNA m6A peaks by modeling biological variances in MeRIP-seq dataabstractMOTIVATION: N(6)-methyl-adenosine (m(6)A) is the most prevalent mRNA methylation but precise prediction of its mRNA location is important for understanding its function. A recent sequencing technology, known as Methylated RNA Immunoprecipitation Sequencing technology (MeRIP-seq), has been developed for transcriptome-wide profiling of m(6)A. We previously developed a peak calling algorithm called exomePeak. However, exomePeak over-simplifies data characteristics and ignores the reads' variances among replicates or reads dependency across a site region. To further improve the performance, new model is needed to address these important issues of MeRIP-seq data. RESULTS: We propose a novel, graphical model-based peak calling method, MeTPeak, for transcriptome-wide detection of m(6)A sites from MeRIP-seq data. MeTPeak explicitly models read count of an m(6)A site and introduces a hierarchical layer of Beta variables to capture the variances and a Hidden Markov model to characterize the reads dependency across a site. In addition, we developed a constrained Newton's method and designed a log-barrier function to compute analytically intractable, positively constrained Beta parameters. We applied our algorithm to simulated and real biological datasets and demonstrated significant improvement in detection performance and robustness over exomePeak. Prediction results on publicly available MeRIP-seq datasets are also validated and shown to be able to recapitulate the known patterns of m(6)A, further validating the improved performance of MeTPeak. AVAILABILITY AND IMPLEMENTATION: The package 'MeTPeak' is implemented in R and C ++, and additional details are available at https://github.com/compgenomics/MeTPeak CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jia Meng 0001, Shaowu Zhang 0001, Yidong Chen 0002, Yufei Huang 0001 |
Bioinform. | 5 |
| 2016 | m6A-Driver: Identifying Context-Specific mRNA m6A Methylation-Driven Gene Interaction NetworksabstractAs the most prevalent mammalian mRNA epigenetic modification, N6-methyladenosine (m6A) has been shown to possess important post-transcriptional regulatory functions. However, the regulatory mechanisms and functional circuits of m6A are still largely elusive. To help unveil the regulatory circuitry mediated by mRNA m6A methylation, we develop here m6A-Driver, an algorithm for predicting m6A-driven genes and associated networks, whose functional interactions are likely to be actively modulated by m6A methylation under a specific condition. Specifically, m6A-Driver integrates the PPI network and the predicted differential m6A methylation sites from methylated RNA immunoprecipitation sequencing (MeRIP-Seq) data using a Random Walk with Restart (RWR) algorithm and then builds a consensus m6A-driven network of m6A-driven genes. To evaluate the performance, we applied m6A-Driver to build the context-specific m6A-driven networks for 4 known m6A (de)methylases, i.e., FTO, METTL3, METTL14 and WTAP. Our results suggest that m6A-Driver can robustly and efficiently identify m6A-driven genes that are functionally more enriched and associated with higher degree of differential expression than differential m6A methylated genes. Pathway analysis of the constructed context-specific m6A-driven gene networks further revealed the regulatory circuitry underlying the dynamic interplays between the methyltransferases and demethylase at the epitranscriptomic layer of gene regulation. Songyao Zhang, Shaowu Zhang 0001, Jia Meng 0001, Yufei Huang 0001 |
PLoS Comput. Biol. | 5 |
| 2016 | EEG-based prediction of driver's cognitive performance by deep convolutional neural network
Mehdi Hajinoroozi, Zijing Mao, Tzyy-Ping Jung, Chin-Teng Lin, Yufei Huang 0001 |
Signal Process. Image Commun. | 5 |
| 2015 | Sketching the distribution of transcriptomic features on RNA transcripts with Travis coordinatesabstractBiological features, such as, genes, transcription factor binding sites, SNPs, etc., are usually denoted with genome-based coordinates as the genomic features. While genome-based representation is usually very effective, it can be tedious to examine the distribution of RNA-related genomic features on RNA transcripts with existing tools due to the conversion and comparison between genome-based coordinates to RNA-based coordinates. We developed here an open source R package Travis for sketching the transcriptomic view of genomic features so as to facilitate the analysis of RNA-related but genome-based coordinates. Internally, Travis package extracts the coordinates relative to the landmarks of transcripts, with which the distribution of RNA-related genomic features can then be conveniently analyzed. We demonstrated the usage of Travis package in analyzing post-transcriptional RNA modifications (5-MethylCytosine and N6-MethylAdenosine) derived from high-throughput sequencing approaches (MeRIP-Seq and RNA BS-Seq). The Travis R package is now publicly available from GitHub: https://github.com/lzcyzm/Travis. Lin Zhang 0015, Hui Liu 0024, Shaowu Zhang 0001, Yufei Huang 0001, Jia Meng 0001 |
BIBM | 7 |
| 2014 | BIMMER: a novel algorithm for detecting differential DNA methylation regions from MBDCap-seq dataabstractDNA methylation is a common epigenetic marker that regulates gene expression. A robust and cost-effective way for measuring whole genome methylation is Methyl-CpG binding domain-based capture followed by sequencing (MBDCap-seq). In this study, we proposed BIMMER, a Hidden Markov Model (HMM) for differential Methylation Regions (DMRs) identification, where HMMs were proposed to model the methylation status in normal and cancer samples in the first layer and another HMM was introduced to model the relationship between differential methylation and methylation statuses in normal and cancer samples. To carry out the prediction for BIMMER, an Expectation-Maximization algorithm was derived. BIMMER was validated on the simulated data and applied to real MBDCap-seq data of normal and cancer samples. BIMMER revealed that 8.83% of the breast cancer genome are differentially methylated and the majority are hypo-methylated in breast cancer. Zijing Mao, Chifeng Ma, Tim Hui-Ming Huang, Yidong Chen 0002, Yufei Huang 0001 |
BMC Bioinform. | 5 |
| 2014 | Selected Articles from the 2012 IEEE International Workshop on Genomic Signal Processing and Statistics (GENSIPS 2012)abstractThe articles in this special section were presented at the 2012 IEEE International Workshop on Genomic Signal Processing and Statistics (GENSIPS 2012) that was held in Washington DC from December 2nd to 4th. Yufei Huang 0001, Yidong Chen 0002, Xiaoning Qian |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2013 | BIMMER: A Bi-layer hidden Markov model for differential methylation analysisabstractMethyl-CpG binding domain-based capture followed by sequencing (MBDCap-seq) is a cost-effective method for genome-wide methylation analyses especially in CpG-rich regions. In this study, we developed BIMMER, a BI-layer hidden Markov model for differential Methylation Regions (DMRs) identification BIMMER using MBDCap-seq samples derived from two different phenotypes. BIMMER models and generates a posterior probability for a 100bp bin to be a methylation site in either normal or disease samples by its first hidden layer, and then integrate these posterior probabilities in the second hidden layer to obtain the posterior probability of bin-specific differential methylation between the normal and disease samples. Based on these posterior probabilities, the decisions on the methylation and differential statuses for each bin can be calculated. Simulated results showed 94.3% area under precision-recall curve for BIMMER (BIMMER is programmed in Java and available by request). Zijing Mao, Tim Hui-Ming Huang, Yidong Chen 0002, Yufei Huang 0001 |
BIBM | 4 |
| 2013 | Unveiling the dynamics in RNA epigenetic regulationsabstractDespite the prevalent studies of DNA/Chromatin related epigenetics, such as, histone modifications and DNA methylation, RNA epigenetics did not receive deserved attention due to the lack of high throughput approach for profiling epitranscriptome. Recently, a new affinity-based sequencing approach MeRIPseq was developed and applied to survey the global mRNA N6-methyladenosine (m6A) in mammalian cells. As a marriage of ChIPseq and RNAseq, MeRIPseq has the potential to study, for the first time, the transcriptome-wide distribution of different types of post-transcriptional RNA modifications. Yet, this technology introduced new computational challenges that have not been adequately addressed. We have previously developed a MATLAB-based package ‘exomePeak’ for detection of RNA methylation sites from MeRIPseq data. Here, we extend the features of exomePeak by including a novel computational framework that enables differential analysis to unveil the dynamics in RNA epigenetic regulations. The novel differential analysis monitors the percentage of modified RNA molecules among the total transcribed RNAs, which directly reflects the impact of RNA epigenetic regulations. In contrast, current available software packages developed for sequencing-based differential analysis such as DESeq or edgeR monitors the changes in the absolute amount of molecules, and, if applied to MeRIPseq data, might be dominated by transcriptional gene differential expression. The algorithm is implemented as an R-package ‘exomePeak’ and freely available. It takes directly the aligned BAM files as input, statistically supports biological replicates, corrects PCR artifacts, and outputs exome-based results in BED format, which is compatible with all major genome browsers for convenient visualization and manipulation. Examples are also provided to depict how exomePeak R-package is integrated with exiting tools for MeRIPseq based peak calling and differential analysis. Particularly, the rationales behind each processing step as well as the specific method used, the best practice, and possible alternative strategies are briefly discussed. The algorithm was applied to the human HepG2 cell MeRIPseq data sets and detects more than 16000 RNA m6A sites, many of which are differentially methylated under ultraviolet radiation. The challenges and potentials of MeRIPseq in epitranscriptome studies are discussed in the end. Jia Meng 0001, Hui Liu 0024, Lin Zhang 0015, Shaowu Zhang 0001, Manjeet K. Rao, Yidong Chen 0002, Yufei Huang 0001 |
BIBM | 8 |
| 2013 | Integration of gene expression, genome wide DNA methylation, and gene networks for clinical outcome prediction in ovarian cancerabstractIntegrative clinical outcome prediction model called gene interaction regularized elastic net (GIREN) method is proposed in this paper. GIREN combines gene expression, methylation profiles, and gene interaction networks in order to reveal genomic and epigenomic features that bear important prognostic value. With GIREN, gene expression and DNA methylation profiles are first jointly analyzed in a linear regression model, and additional gene interaction network is simultaneously integrated as a regularizing penalty that follow an elastic net formulation. Such regularization also enforce sparsity in the solution so that features with prognostic values are automatically selected. To solve the regularized optimization, an iterative gradient descent algorithm is also developed. We applied GIREN to a set of 87 human ovarian cancer samples, which underwent a rigorous sample selection. The predicted outcome was used to group patients into high-risk vs. low-risk. Validation showed that GIREN outperformed other competing algorithms including SuperPCA. Lin Zhang 0015, Hui Liu 0024, Jia Meng 0001, Xuesong Wang 0001, Yidong Chen 0002, Yufei Huang 0001 |
BIBM | 6 |
| 2013 | A bag-of-words model for task-load prediction from EEG in complex environmentsabstractNeurotechnologies based on electroencephalography (EEG) and other physiological measures to improve task performance in complex environments will require tools and analysis methods that can account for increased environmental noise and task complexity compared to traditional neuroscience laboratory experiments. We propose a bag-of-words (BoW) model to address the difficulties associated with realistic applications in complex environments. In this paper, our proof-of-concept results show that a BoW classifier can discriminate two task-relevant states (high versus low task-load) while an individual performs a simulated security patrol mission with complex, concurrent tasking. Classifier performance is largely consistent across six simulation missions for a given participant, but performance decreases when trying to predict between two individuals. Overall, these initial results suggest that this BoW approach holds promise for detecting task-relevant states in real-world settings. Lenis Mauricio Merino, Jia Meng 0001, Stephen M. Gordon, Brent Lance, Tony Johnson, Victor Paul, Kay A. Robbins, Jean M. Vettel, Yufei Huang 0001 |
ICASSP | 9 |
| 2013 | Exome-based analysis for RNA epigenome sequencing dataabstractMOTIVATION: Fragmented RNA immunoprecipitation combined with RNA sequencing enabled the unbiased study of RNA epigenome at a near single-base resolution; however, unique features of this new type of data call for novel computational techniques. RESULT: Through examining the connections of RNA epigenome sequencing data with two well-studied data types, ChIP-Seq and RNA-Seq, we unveiled the salient characteristics of this new data type. The computational strategies were discussed accordingly, and a novel data processing pipeline was proposed that combines several existing tools with a newly developed exome-based approach 'exomePeak' for detecting, representing and visualizing the post-transcriptional RNA modification sites on the transcriptome. AVAILABILITY: The MATLAB package 'exomePeak' and additional details are available at http://compgenomics.utsa.edu/exomePeak/. Jia Meng 0001, Manjeet K. Rao, Yidong Chen 0002, Yufei Huang 0001 |
Bioinform. | 5 |
| 2011 | Uncover cooperative gene regulations by microRNAs and transcription factors in glioblastoma using a nonnegative hybrid factor modelabstractTranscriptional regulation by transcription factors (TFs) and microRNAs controls when and how much RNA is created. Due to technical limitations, the protein level expressions of TFs are usually unknown, making computational reconstruction of transcriptional network a difficult task. We proposed here a novel Bayesian non negative hybrid factor model for transcriptional network modeling, which is capable to estimate both the non-negative abundances of the transcription factors, the regulatory effects of TFs and microRNAs, and the sample clustering information by integrating microarray data and existing knowledge regarding TFs and microRNAs regulated target genes. The results demonstrated its validity and effectiveness to reconstructing transcriptional networks through simulated systems and real data. Jia Meng 0001, Hung-I Harry Chen, Jianqiu Zhang 0002, Yidong Chen 0002, Yufei Huang 0001 |
ICASSP | 5 |
| 2011 | PlantMiRNAPred: efficient classification of real and pseudo plant pre-miRNAsabstractMOTIVATION: MicroRNAs (miRNAs) are a set of short (21-24 nt) non-coding RNAs that play significant roles as post-transcriptional regulators in animals and plants. While some existing methods use comparative genomic approaches to identify plant precursor miRNAs (pre-miRNAs), others are based on the complementarity characteristics between miRNAs and their target mRNAs sequences. However, they can only identify the homologous miRNAs or the limited complementary miRNAs. Furthermore, since the plant pre-miRNAs are quite different from the animal pre-miRNAs, all the ab initio methods for animals cannot be applied to plants. Therefore, it is essential to develop a method based on machine learning to classify real plant pre-miRNAs and pseudo genome hairpins. RESULTS: A novel classification method based on support vector machine (SVM) is proposed specifically for predicting plant pre-miRNAs. To make efficient prediction, we extract the pseudo hairpin sequences from the protein coding sequences of Arabidopsis thaliana and Glycine max, respectively. These pseudo pre-miRNAs are extracted in this study for the first time. A set of informative features are selected to improve the classification accuracy. The training samples are selected according to their distributions in the high-dimensional sample space. Our classifier PlantMiRNAPred achieves >90% accuracy on the plant datasets from eight plant species, including A.thaliana, Oryza sativa, Populus trichocarpa, Physcomitrella patens, Medicago truncatula, Sorghum bicolor, Zea mays and G.max. The superior performance of the proposed classifier can be attributed to the extracted plant pseudo pre-miRNAs, the selected training dataset and the carefully selected features. The ability of PlantMiRNAPred to discern real and pseudo pre-miRNAs provides a viable method for discovering new non-homologous plant pre-miRNAs. Ping Xuan, Maozu Guo 0001, Yangchao Huang, Yufei Huang 0001 |
Bioinform. | 6 |
| 2010 | An Iterated Conditional Modes solution for sparse Bayesian factor modeling of transcriptional regulatory networksabstractThe problem of uncovering transcriptional regulation by transcription factors (TFs) based on microarray data is considered. A novel Bayesian sparse correlated rectified factor model (BSCRFM) coupled with its ICM solution is proposed. BSCRFM models the unknown TF protein level activity, the correlated regulations between TFs, and the sparse nature of TF regulated genes and it admits prior knowledge from existing database regarding TF regulated target genes. An efficient Iterated Conditional Modes (ICM) algorithm is developed, and a maximum a posterior (MAP) solution is calculated from multiple ICM results to avoid the local maximum problem, a context-specific transcriptional regulatory network specific to the experimental condition of the microarray data can then be obtained. The proposed model's ICM algorithm and MAP solution are evaluated on the simulated systems and results demonstrated the validity and effectiveness of the proposed approach. The proposed model is also applied to the breast cancer microarray data and a TF regulated network is obtained. Jia Meng 0001, Jianqiu Zhang 0002, Yidong Chen 0002, Yufei Huang 0001 |
BIBM | 4 |
| 2010 | SysMicrO: A Novel Systems Approach for miRNA Target Prediction
Hui Liu 0024, Lin Zhang 0015, Qilong Sun, Yidong Chen 0002, Yufei Huang 0001 |
ICIC (2) | 5 |
| 2010 | miRNA Target Prediction Method Based on the Combination of Multiple Algorithms
Lin Zhang 0015, Hui Liu 0024, Dong Yue 0002, Yufei Huang 0001 |
ICIC (1) | 5 |
| 2010 | Improving performance of mammalian microRNA target predictionabstractBACKGROUND: MicroRNAs (miRNAs) are single-stranded non-coding RNAs known to regulate a wide range of cellular processes by silencing the gene expression at the protein and/or mRNA levels. Computational prediction of miRNA targets is essential for elucidating the detailed functions of miRNA. However, the prediction specificity and sensitivity of the existing algorithms are still poor to generate meaningful, workable hypotheses for subsequent experimental testing. Constructing a richer and more reliable training data set and developing an algorithm that properly exploits this data set would be the key to improve the performance current prediction algorithms. RESULTS: A comprehensive training data set is constructed for mammalian miRNAs with its positive targets obtained from the most up-to-date miRNA target depository called miRecords and its negative targets derived from 20 microarray data. A new algorithm SVMicrO is developed, which assumes a 2-stage structure including a site support vector machine (SVM) followed by a UTR-SVM. SVMicrO makes prediction based on 21 optimal site features and 18 optimal UTR features, selected by training from a comprehensive collection of 113 site and 30 UTR features. Comprehensive evaluation of SVMicrO performance has been carried out on the training data, proteomics data, and immunoprecipitation (IP) pull-down data. Comparisons with some popular algorithms demonstrate consistent improvements in prediction specificity, sensitivity and precision in all tested cases. All the related materials including source code and genome-wide prediction of human targets are available at http://compgenomics.utsa.edu/svmicro.html. CONCLUSIONS: A 2-stage SVM based new miRNA target prediction algorithm called SVMicrO is developed. SVMicrO is shown to be able to achieve robust performance. It holds the promise to achieve continuing improvement whenever better training data that contain additional verified or high confidence positive targets and properly selected negative targets are available. Hui Liu 0024, Dong Yue 0002, Yidong Chen 0002, Shou-Jiang Gao, Yufei Huang 0001 |
BMC Bioinform. | 5 |
| 2009 | Enrichment constrained time-dependent clustering analysis for finding meaningful temporal transcription modulesabstractMOTIVATION: Clustering is a popular data exploration technique widely used in microarray data analysis. When dealing with time-series data, most conventional clustering algorithms, however, either use one-way clustering methods, which fail to consider the heterogeneity of temporary domain, or use two-way clustering methods that do not take into account the time dependency between samples, thus producing less informative results. Furthermore, enrichment analysis is often performed independent of and after clustering and such practice, though capable of revealing biological significant clusters, cannot guide the clustering to produce biologically significant result. RESULT: We present a new enrichment constrained framework (ECF) coupled with a time-dependent iterative signature algorithm (TDISA), which, by applying a sliding time window to incorporate the time dependency of samples and imposing an enrichment constraint to parameters of clustering, allows supervised identification of temporal transcription modules (TTMs) that are biologically meaningful. Rigorous mathematical definitions of TTM as well as the enrichment constraint framework are also provided that serve as objective functions for retrieving biologically significant modules. We applied the enrichment constrained time-dependent iterative signature algorithm (ECTDISA) to human gene expression time-series data of Kaposi's sarcoma-associated herpesvirus (KSHV) infection of human primary endothelial cells; the result not only confirms known biological facts, but also reveals new insight into the molecular mechanism of KSHV infection. AVAILABILITY: Data and Matlab code are available at http://engineering.utsa.edu/ approximately yfhuang/ECTDISA.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jia Meng 0001, Shou-Jiang Gao, Yufei Huang 0001 |
Bioinform. | 3 |
| 2009 | Adaptive Sensor Fault Detection and Identification Using Particle Filter AlgorithmsabstractSensor fault detection and identification (FDI) is a process of detecting and validating sensor's fault status. Because FDI guarantees system reliable performance, it has received much attention recently. In this paper, we address the problem of online sensor fault identification and validation. For a physical sensor validation system, it contains transitions between sensor normal and faulty states, change of system parameters, and a fusion of noisy readings. A common dynamic state-space model with continuous state variables and observations cannot handle this problem. To circumvent this limitation, we adopt a Markov switch dynamic state-space model to simulate the system: we use discrete-state variables to model sensor states and continuous variables to track the change of the system parameters. Problems in Markov switch dynamic state-space model can be well solved by particle filters, which are popularly used in solving problems in digital communications. Among them, mixture Kalman filter (MKF) and stochasticM-algorithm (SMA) have very good performance, both in accuracy and efficiency. In this paper, we plan to incorporate these two algorithms into the sensor validation problem, and compare the effectiveness and complexity of MKF and SMA methods under different situations in the simulation with an existing algorithm - interactive multiple models. Yufei Huang 0001, C. L. Philip Chen |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2008 | Bayesian peak detection for Pro-TOF MS MALDI dataabstractIn this paper, a novel Bayesian peak detection algorithm is proposed for peptide peak detection in high resolution prOTOFtrade MALDI Mass Spectrometry(MS) data. A nonlinear parametric model is proposed for modeling the peptide signals, chemical noise, and thermal noise. A metropolized Gibbs sampling algorithm is derived for Bayesian peak detection. The proposed algorithm is compared with a popular wavelet-based algorithm and the results show a significant improvement in performance on simulated data. The algorithm is finally tested on real MS MALDI data and the results agree with visual inspection very well. Jianqiu Zhang 0002, Anthony Suffredini, Denise Gonzales, Elias Gonzalez, Yufei Huang 0001, Xiaobo Zhou 0001 |
ICASSP | 6 |
| 2008 | Constructing Gene Networks Using Variational Bayesian Variable SelectionabstractWe propose a Bayesian approach for constructing gene networks based on microarray data. Especially, we focus on Bayesian methods that can provide soft (probabilistic) information. This soft information is attractive not only for its ability to measure the level of confidence of the solution, but also because it can be used to realize Bayesian data integration, an extremely important task in gene network research. We propose a variable selection formulation of gene regulation and develop an inference solution based on a variational Bayesian expectation maximization (VBEM) learning rule. This solution has better performance and lower complexity than the popular Monte Carlo sampling techniques. In addition, we develop a method to incorporate the often needed constraints into the VBEM algorithm, making it much more suitable for common cases of small data size. To further illustrate the advantage of the VBEM algorithm, we demonstrate a Bayesian data integration scheme using the soft information obtained from the VBEM algorithm. The efficacy of the proposed VBEM algorithm and the corresponding Bayesian data integration scheme is evaluated on both simulated data and the yeast cell cycle microarray data sets. Isabel M. Tienda-Luna, Yufang Yin, Yufei Huang 0001, Diego P. Ruiz 0001, Maria Carmen Carrion Perez, Yufeng Wang 0002 |
Artif. Life | 3 |
| 2006 | Reverse Engineering Yeast Gene Regulatory Networks using Graphical ModelsabstractWe investigate in this paper reverse engineering of gene regulatory networks from time series microarray data. We propose a dynamic Bayesian networks (DBNs) modeling and a full Bayesian learning scheme. The proposed DBN models directly the continuous expression levels and also is associated with parameters that indicate the degree as well as the types of regulations. To learn the network from data, we proposed a reversible jump Markov chain Monte Carlo (RJMCMC) algorithm. The RJMCMC algorithm can provide not only more accurate inference results than the deterministic alternative algorithms but also an estimate on the a posteriori probabilities (APPs) of the network topology. The estimated APPs provide useful information on the confidence of the inferred results and can also be used for efficient Bayesian data integration. The proposed approach was tested on yeast cell cycle microarray data and the results were compared with the KEGG pathway map Yufei Huang 0001, Maribel Sanchez, Yufeng Wang 0002, Jianqiu Zhang 0002 |
ICASSP (2) | 2 |
| 2006 | Particle Filtering for Adaptive Sensor Fault Detection and IdentificationabstractIn this paper, we address the problem of adaptive sensor fault identification and validation by particle filtering. The model-based approaches are developed, where the sensor system is modeled by a Markov switch dynamic state-space model. To handle the nonlinearity of the problem, two different particle filters: mixture Kalman filter (MKF) and stochastic M-algorithm (SMA) are proposed. Simulation results are presented to compare the effectiveness and complexity of MKF and SMA methods Yufei Huang 0001, C. L. Philip Chen |
ICRA | 2 |
| 2005 | Belief-directed sequential probabilistic data association multiuser detectorabstractWe propose in this paper a novel soft-input-soft-output (SISO) multiuser detector for synchronous CDMA systems. The detector is called the belief-directed sequential probabilistic data association detector (BD-SPDAD). The BD-SPDAD is developed based on a general framework of the probabilistic data association detector (PDAD) proposed in (Y Huang et al, IEEE Int. Symp. on Inf. Theo., 2004). However, novel extensions to the general framework are proposed for the BD-SPDAD, which result in a low complexity sequential implementation. Specifically, the complexity of the BD-SPDAD is reduced from O(K/sup 3/) of the original PDAD to O(K/sup 2/), where K is the number of users in the system. Moreover, we show through simulation that the performance of the BD-SPDAD is comparable and even better than the original PDAD especially in high signal-to-noise regions. Yufei Huang 0001, Yufang Yin, Jianqiu Zhang 0002 |
ICASSP (3) | 1 |
| 2005 | Symbol detection with time-varying unknown phase by expectation propagationabstractIn digital communications, symbol detection in phase noise is an important topic that has been discussed in many papers under different conditions. In this paper, we consider symbol detection with time-varying unknown phase. We propose a solution based on expectation propagation (EP). EP is an extension to belief propagation and developed in machine learning. We point out that the developed EP solution can be considered as an iterated extended Kalman smoother (EKS). However, a crucial step of recycling the likelihoods in EP makes possible the further improvement over EKS. We show in the simulation that EP can produce very good performance with relatively low complexity. Since it produce soft information, the EP solution can be readily applied to iterative detection of coded systems. Yufei Huang 0001, Yuan Qi 0001 |
ICASSP (3) | 2 |
| 2005 | Remaining engine life estimation for a sensor-based aircraft engineabstractIt is generally known that an engine component will accumulate damage (life usage) during its lifetime of use in a harsh operating environment. The commonly used cycle count for engine component usage monitoring has an inherent range of uncertainty that can be overly costly or potentially less safe from an operational standpoint. This paper describes an approach to quantify the effects of engine operating parameter uncertainties on the thermomechanical fatigue (TMF) life of a selected engine part. A closed-loop engine simulation with a TMF life model is used to calculate the life consumption of different mission cycles. A Monte Carlo simulation approach is used to generate the statistical life usage profile for different operating assumptions. The probabilities of failure of different operating conditions are compared to illustrate the importance of the engine component life calculation using sensor information. The results of this study clearly show that a sensor-based life cycle calculation can greatly reduce the risk of component failure as well as extend on-wing component life by avoiding unnecessary maintenance actions. Ten-Huei Guo, C. L. Philip Chen, Yufei Huang 0001 |
SMC | 3 |
| 2004 | Adaptive blind multiuser detection over flat fast fading channels using particle filteringabstractIn this paper, we propose a method for blind multiuser detection (MUD) in synchronous systems over flat and fast Rayleigh fading channels. We adopt an autoregressive-moving-average (ARMA) process to model the temporal correlation of the channels. Based on the ARMA process, we propose a novel time-observation state space model (TOSSM) that describes the dynamics of the addressed multiuser system. The TOSSM allows an MUD with natural blending of low complexity particle filtering (PF) and mixture Kalman filtering (for channel estimation). We further propose to use a more efficient PF algorithm known as the stochastic M-algorithm (SMA), which, although having lower complexity than the generic PF implementation, maintains comparable performance. Yufei Huang 0001, Jianqiu Zhang 0002, Isabel M. Tienda-Luna, Petar M. Djuric, Diego P. Ruiz 0001 |
GLOBECOM | 1 |
| 2004 | Turbo equalization using probabilistic data associationabstractWe investigate turbo equalization using an algorithm called probabilistic data association (PDA). We first propose a general structure for PDA, which consists of a linear interference cancellation step followed by a probabilistic data association step in every iteration. Based on the general structure, we show that the original PDA belongs to one variation and it is computationally inefficient. We then unveil that the popular soft linear MMSE (SLMMSE) equalizer can be considered as one sweep within a generalized PDA. Such a connection implies that further performance improvement over the SLMMSE equalizer is possible if the PDA is applied instead in turbo equalization. We also provide a way for the PDA equalizer to incorporate the a priori probability, which makes the PDA readily applicable to turbo equalization. Yufang Yin, Yufei Huang 0001, Jianqiu Zhang 0002 |
GLOBECOM | 2 |
| 2004 | Joint symbol detection and timing estimation with stochastic M-algorithmabstractIn digital communications, the symbol timing estimation is an very important element for high quality data detection. This paper considers the problem of joint symbol detection and timing estimation, whose optimal solution is analytically intractable. In this paper, a stochastic M-algorithm is proposed for the solution. The stochastic M-algorithm is a novel efficient particle filtering algorithm designed for discrete unknowns. To accommodate, in the stochastic M-algorithm, the continuous unknown symbol timing, the unscented Kalman filter is introduced, which leads to a very efficient implementation. The simulation results illustrated that the stochastic M-algorithm achieves similar performance to particle filtering with less than 1/25 of the complexity. Yufang Yin, Yufei Huang 0001, Jianqiu Zhang 0002 |
ICASSP (4) | 2 |
| 2004 | A generalized probabilistic data association multiuser detectorabstractWe present in this paper a general framework for the probabilistic data association multiuser detector (PDAD). This generalization allows for different algorithm variations though the design of interference cancellation (IC) filters. We examine two possible variations. In the first variation, we demonstrate a computational more efficient algorithm than the original PDAD. In the second variation, we draw the connection between the PDAD and some popular soft IC detectors and this connection implies that further performance improvement over the soft IC detectors is possible in a turbo multiuser detection (MUD) by the PDAD. Moreover, we show the equivalence of the two variations, which suggests that the IC step is computationally redundant for both the second variation and the soft IC detectors. Yufei Huang 0001, Jianqiu Zhang 0002 |
ISIT | 1 |
| 2004 | A hybrid importance function for particle filteringabstractParticle filtering has drawn much attention in recent years due to its capacity to handle nonlinear and non-Gaussian dynamic problems. One crucial issue in particle filtering is the selection of the importance function that generates the particles. In this letter, we propose a new type of importance function that possesses the advantages of the posterior and the prior importance functions. We demonstrate its use on the problem of blind detection in flat fading channels and provide simulation results that show its efficiency and performance. Yufei Huang 0001, Petar M. Djuric |
IEEE Signal Process. Lett. | 1 |
| 2003 | Joint velocity estimation and symbol detection in non-stationary fading channels by particle filteringabstractThe paper addresses the problem of joint velocity estimation and data detection in a realistic scenario where mobile velocity changes continuously, resulting in non-stationary fast fading channels. A time-varying AR model and a Gauss-Markov model are used to describe the respective fading channel and variation of velocity. A connection is shown between the coefficients of the TVAR model and mobile velocity which makes the joint estimation and detection possible. A hierarchical dynamic state space model is formed for the problem, and a particle filtering algorithm is proposed. In particular, a hybrid importance function and the mixture Kalman filter are used to achieve efficient implementation of particle filtering. Simulation results are provided that show the performance of the particle filtering algorithm. Yufei Huang 0001, Petar M. Djuric, Jianqiu Zhang 0002 |
GLOBECOM | 1 |
| 2003 | Lower bounds on the variance of deterministic signal parameter estimators using Bayesian inferenceabstractBayesian estimators have been applied for both random and deterministic parameters. In the Bayesian estimation of deterministic parameters, the randomness is introduced only through the observations, and the prior distributions are adopted to impose certain constraints. In such cases, neither the well known Cramer-Rao lower bound (CRLB) or the posterior CRLB can be used reasonably as the performance lower bounds. We extend the theory of CRLB under the Bayesian framework to provide the lower bounds for both unbiased and biased Bayesian estimators of deterministic parameters. An example is provided to show the effectiveness of the proposed lower bound over other popular lower bounds. Yufei Huang 0001, Jianqiu Zhang 0002 |
ICASSP (6) | 1 |
| 2003 | Detection with particle filtering in BLAST systemsabstractThis work demonstrates the use of particle filtering for detection in BLAST systems. A novel dynamic state-space model (DSSM) is constructed for BLAST systems that are crucial for development of particle filtering algorithms. The proposed DSSM is based on QR decomposition and the output of the feedforward filter, and it evolves in space. The particle filtering solution does not suffer from error propagation, and our simulation show that it greatly outperforms the V-BLAST and provides near optimum performance. Yufei Huang 0001, Jianqiu Zhang 0002, Petar M. Djuric |
ICC | 1 |
| 2002 | A new importance function for particle filtering and its application to blind detection in flat fading channelsabstractParticle filtering has drawn much attention in recent years due to its capacity to handle nonlinear and non-Gaussian problems. One crucial issue in particle filtering is-the selection of importance function. In this paper, we propose a new type of importance function, which possess advantages over both the posterior and the prior importance functions. In addition, we demonstrate the use of the proposed importance function in blind detection in flat fading channels. Simulation results show its efficiency and performance. Yufei Huang 0001, Petar M. Djuric |
ICASSP | 1 |
| 2001 | Multiuser detection of synchronous CDMA signals by the Gibbs couplerabstractCode-division multiple-access (CDMA) is a multiplexing technique which has become a driving force behind the rapidly advancing communications industry. In order to recover transmitted signals at the receiver when CDMA is used, multiuser detection techniques are engaged. In the past, various techniques have been developed to approach the performance of the optimum multiuser detector, and we have the same objective. We propose a new multiuser detector which is developed under the Bayesian framework and implemented by a novel efficient perfect sampling algorithm called the Gibbs coupler. The simulation results demonstrate the excellent performance of the proposed detector. Yufei Huang 0001, Petar M. Djuric |
ICASSP | 1 |
| 2000 | Bayesian detection of transient signals in colored noiseabstractThe problem of detecting transient signals in colored noise is addressed. The generalized likelihood ratio test fails to provide good performance, and as an alternative, a Bayesian approach is proposed. Its implementation is simplified by adopting an invariant transformation, and a sequential Monte Carlo sampling procedure is provided to compute multidimensional integrations. The simulation results demonstrate the ability of this approach to detect weak and fast decaying signals in colored noise. Yufei Huang 0001, Petar M. Djuric |
ICASSP | 1 |
| 2000 | Estimation of a Bernoulli parameter p from imperfect trialsabstractImperfect Bernoulli trials arise when the outcome of a Bernoulli experiment is not known with certainty. In signal processing, we often need to estimate a probability of occurrence p of an event from imperfect Bernoulli trials. A typical example is the estimation of the probability of a signal being present in noisy data. In his famous essay, Bayes solved the same problem but for perfect trials. In this letter, a solution is provided for imperfect trials. It is shown that it includes Bayes' solution as a special case. Petar M. Djuric, Yufei Huang 0001 |
IEEE Signal Process. Lett. | 2 |