Jean Gao

dblp:18/4336 · also Jean X. Gao · DBLP profile ↗
← Back
89ranked-venue papers
10as first author
8since 2021 · last 2026
0000-0002-9077-7803ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 55 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Learning from Guidelines: Structured Prompt Optimization for Expert Annotation Tasks
abstract
Deep learning has significantly advanced numerous fields by training on extensive annotated datasets. However, this data-driven paradigm faces limitations such as limited adaptability and high annotation costs, particularly when precise adherence to detailed, domain-specific guidelines is required in annotation. This challenge raises a critical question: Can models effectively shift from data-driven learning to autonomously leveraging guidelines with minimal annotated examples? To address this, we propose the Guideline-Driven Prompt (GDP) optimization framework, which shifts the learning paradigm from data-driven training to guideline-driven reasoning. GDP leverages Retrieval Augmented Generation (RAG) to retrieve essential fragments from complex guidelines and synthesize them into structured, executable prompts. A tree-based optimization algorithm systematically constructs and refines these prompts, explicitly capturing the intricate logic embedded in professional guidelines through a latent pipeline structure. Empirical evaluations on four datasets ranging from diverse domains and different tasks demonstrate that GDP effectively transitions the learning process from data-intensive methods to a guideline-driven approach in tasks requiring detailed and complex guideline adherence, reducing dependence on extensive annotated datasets.
Leon Wenliang Zhong, Thao M. Dang, Feng Jiang 0012, Hehuan Ma, Yuzhi Guo, Jean Gao, Junzhou Huang
AAAI7
2026 Transcriptome graph transformer: a graph transformer-based unsupervised model for transcriptome data analysis
abstract
BACKGROUND: Rapidly growing transcriptomic datasets pose challenges for traditional analytical methods, which struggle with high dimensionality, heterogeneity, and nonlinear gene relationships. Existing deep learning models often require fixed-length inputs and fail to integrate biological network information. METHODS: We introduce Transcriptome Graph Transformer (TGT), an unsupervised graph Transformer framework that constructs a heterogeneous gene-pathway graph using expression data, STRING interactions, and GO/KEGG/Reactome pathway annotations. TGT is pretrained with a masked-node prediction task and fine-tuned for disease classification, biomarker discovery, and zero-shot clustering of single-cell and spatial transcriptomics. RESULTS: TGT achieves superior performance across Alzheimer's disease, cancer, acute kidney injury, and COVID-19 datasets, outperforming state-of-the-art baselines. The model generalizes well across platforms and yields biologically meaningful gene and pathway importance scores consistent with known disease mechanisms. CONCLUSIONS: TGT provides an effective and generalizable approach for transcriptomic representation learning by integrating biological network knowledge with graph Transformer architectures. Its strong performance highlights its utility for broad transcriptomic applications and precision medicine.
Sachit Satyal, Jean Gao
BMC Bioinform.3
2025 HAGE: Hierarchical Alignment Gene-Enhanced Pathology Representation Learning with Spatial Transcriptomics
Thao M. Dang, Yuzhi Guo, Hehuan Ma, Feng Jiang 0012, Yuwei Miao, Qifeng Zhou, Jean Gao, Junzhou Huang
MICCAI (1)8
2025 Text-Guided Multi-instance Learning for Scoliosis Screening via Gait Video Analysis
Yuzhi Guo, Feng Jiang 0012, Thao M. Dang, Hehuan Ma, Qifeng Zhou, Jean Gao, Junzhou Huang
MICCAI (6)7
2025 Segment Any Cell: A SAM-Based Auto-Prompting Fine-Tuning Framework for Nuclei Segmentation
abstract
In the rapidly evolving field of AI research, foundational models like BERT and GPT have significantly advanced language and vision tasks. The advent of pretrain-prompting models, such as ChatGPT and segment anything model (SAM), has further revolutionized image segmentation. However, their applications in specialized areas, particularly in nuclei segmentation within medical imaging, reveal a key challenge: the generation of high-quality, informative prompts is as crucial as applying state-of-the-art (SOTA) fine-tuning techniques on foundation models. To address this, we introduce segment any cell (SAC), an innovative framework that enhances SAM specifically for nuclei segmentation. SAC integrates a low-rank adaptation (LoRA) within the attention layer of the Transformer to improve the fine-tuning process, outperforming existing SOTA methods. It also introduces an innovative auto-prompt generator that produces effective prompts to guide segmentation, a critical factor in handling the complexities of nuclei segmentation in biomedical imaging. Our extensive experiments demonstrate the superiority of SAC in nuclei segmentation tasks, proving its effectiveness as a tool for pathologists and researchers. Our contributions include a novel prompt generation strategy, automated adaptability for diverse segmentation tasks, the innovative application of low-rank attention adaptation in SAM, and a versatile framework for semantic and instance segmentation challenges.
Saiyang Na, Yuzhi Guo, Feng Jiang 0012, Hehuan Ma, Jean Gao, Junzhou Huang
IEEE Trans. Neural Networks Learn. Syst.5
2023 Integrating Heterogeneous Biological Networks and Ontologies for Improved Protein Function Prediction with Graph Neural Networks
abstract
Elucidating protein functions is critical to advance our understanding of biological systems. However, the majority of proteins lack functional annotations due to the pace of manual knowledge-driven curation of the Gene Ontology (GO) terms. Automatic function prediction (AFP) aims to computationally infer these missing annotations by integrating sequence, interaction, and ontology data. Current AFP methods have yet to fully utilize heterogeneous relationships across multi-omics networks between non-coding RNAs, proteins, and the GO hierarchies. To address this, we introduce LATTE2GO, a novel graph neural network that integrates protein, RNA, and GO function entities and their interactions into a unified representation. By extracting higher-order associations with attention mechanism, LATTE2GO achieved significant gains over previous graph-based AFP techniques on CAFA4 benchmarks. Our analyses revealed modeling specific protein-protein interactions (PPI) and GO relationships increases accuracy in predicting molecular functions and biological processes. Overall, LATTE2GO demonstrates that heterogeneous graph neural networks can integrate diverse omics knowledge to advance systems-level understanding of protein roles within the complex milieu of functional and physical interaction networks.
Nhat C. Tran, Jean Gao
BIBM2
2022 Causal Discovery in Biological Data Using Directed Topological Overlap Matrix
abstract
Most causal discovery tools assume the local causal Markov condition. However, the theoretical assumptions that underline the local causal Markov condition are often not met in practice. This is especially marked in genomics, where the unwanted presence of measurement errors, averaging effects, and feedback loops significantly undermine the legitimacy of the local causal Markov condition. Furthermore, even the most sample efficient causal discovery algorithms require very large samples, orders above what is available even for the largest genomics databases out there. In this paper, by replacing the local causal Markov condition with Reichenbach’s common cause principle, we present directed topological overlap matrix (DTOM): a more flexible approach to causal discovery in genomics settings, robust against the presence of measurement errors, averaging effects, and feedback loops. We study the utility of DTOM for discovering causal relations in biological data using two real gene expression data-sets. We first consider a large-scale gene deletion study in yeast. We show that DTOM allows us to distinguish the deleted gene in a sample knowing only the set of differentially expressed genes in that sample. We then examine the progression of Alzheimer’s disease (AD) under the lens of DTOM. The genes implicated as having a causal role in the progression of AD by our DTOM analysis were significantly enriched in cellular components that had been repeatedly implicated in the progression of AD.
Borzou Alipourfard, Jean Gao
BIBM2
2021 Deep Fusion of Brain Structure-Function in Mild Cognitive Impairment
Lu Zhang 0050, Li Wang 0033, Jean Gao, Shannon L. Risacher, Gang Li 0001, Tianming Liu 0001, Dajiang Zhu
Medical Image Anal.3
2020 Causal Gene Selection Using Marginal Independencies
abstract
Most causal learning algorithms, especially the ones using conditional independence tests, suffer in the low sample regime. In this paper, we examine the possibility of causal discovery using only the pairwise dependency structure of a domain and investigate if this approach can lead to more robust and reliable solutions to causal discovery when the sample size is limited. We show that marginal independence relations can be used to construct a novel causal discovery tool capable of distinguishing between direct and indirect causal effects and detecting a lower bound on direct causal effects. In our simulations we found that causal discovery using marginal independency information can be significantly more accurate than causal knowledge obtained through conditional independence tests when the sample size is limited: our experiments suggests avoiding conditional independence tests can reduce error rate by up to 30 percent. Finally, using this marginal independency based causal discovery tool, we propose four candidate genes that possibly contain the causal mutation in Major Depressive Disorder utilizing a database of only 59 samples.
Borzou Alipourfard, Jean Gao
BIBM2
2019 Semi-Supervised Discriminative Transfer Learning in Cross-Language Text Classification
abstract
Cross-Language Text Classification (CLTC) has been increasing its attention to multilingual data due to its exponentially growing. CLTC aims to classify text documents in a label-scarce language, by leveraging classification information in a label-rich language. We propose a novel semi-supervised Discriminative Transfer Learning method (DTL) for the CLTC problem of a semi-supervised setting. A small number of paired labeled data in bilingual documents constructs a discriminative transfer model that maximizes the correlations of the documents in both languages, while a large number of unlabeled data are used for accurate data reconstruction. The discriminative transfer model minimizes the discrepancy between bilingual subspaces prioritizing discriminative features to improve the text classification performance without an automatic machine translation that most state-of-the-art methods require. The performance of DTL is empirically and statistically assessed by intensive experiments with the publicly available data, Reuters RCV1/RCV2 collections. The experimental results demonstrate that DTL outperforms several representative state-of-the-art methods in CLTC in terms of accuracy and efficiency.
Mingon Kang, Ashis Kumer Biswas, Dong-Chul Kim, Jean Gao
ICMLA4
2019 Robust Inductive Matrix Completion Strategy to Explore Associations Between LincRNAs and Human Disease Phenotypes
abstract
Over the past few years, it has been established that a number of long intergenic non-coding RNAs (lincRNAs) are linked to a wide variety of human diseases. The relationship among many other lincRNAs still remains as puzzle. Validation of such link between the two entities through biological experiments is expensive. However, piles of information about the two are becoming available, thanks to the High Throughput Sequencing (HTS) platforms, Genome Wide Association Studies (GWAS), etc., thereby opening opportunity for cutting-edge machine learning and data mining approaches. However, there are only a few in silico lincRNA-disease association inference tools available to date, and none of these utilizes side information of both the entities. The recently developed Inductive Matrix Completion (IMC) technique provides a recommendation platform among two entities considering respective side information. But, the formulation of IMC is incapable of handling noise and outliers that may present in the dataset, while data sparsity consideration is another issue with the standard IMC method. Thus, a robust version of IMC is needed that can solve these two issues. As a remedy, in this paper, we propose Robust Inductive Matrix Completion (RIMC) using 12;1 norm loss function aswell as 12;1 norm based regularization. We applied RIMC to the available association data between human lincRNAs and OMIM disease phenotypes as well as a diverse set of side information about the lincRNAs and the diseases. Our method performs better than the state-of-the-art methods in terms of precision@k and recall@k at the top-k disease prioritization to the subject lincRNAs. We also demonstrate that RIMC is equally effective forquerying about novel lincRNAs, as well as predicting rankof a newly known disease for a set of well-characterized lincRNAs. Availability: All the supporting datasets are available at the publicly accessible URL located at http://biomecis.uta.edu/~ashis/res/RIMC/.
Ashis Kumer Biswas, Dong-Chul Kim, Mingon Kang, Jean Gao
IEEE ACM Trans. Comput. Biol. Bioinform.4
2018 MicroRNA dysregulational synergistic network: discovering microRNA dysregulatory modules across subtypes in non-small cell lung cancers
abstract
BACKGROUND: The majority of cancer-related deaths are due to lung cancer, and there is a need for reliable diagnostic biomarkers to predict stages in non-small cell lung cancer cases. Recently, microRNAs were found to have potential as both biomarkers and therapeutic targets for lung cancer. However, some of the microRNA's functions are unknown, and their roles in cancer stage progression have been mostly undiscovered in this clinically and genetically heterogeneous disease. As evidence suggests that microRNA dysregulations are implicated in many diseases, it is essential to consider the changes in microRNA-target regulation across different lung cancer subtypes. RESULTS: We proposed a pipeline to identify microRNA synergistic modules with similar dysregulation patterns across multiple subtypes by constructing the MicroRNA Dysregulational Synergistic Network. From the network, we extracted microRNA modules and incorporated them as prior knowledge to the Sparse Group Lasso classifier. This leads to a more relevant selection of microRNA biomarkers, thereby improving the cancer stage classification accuracy. We applied our method to the TCGA Lung Adenocarcinoma and the Lung Squamous Cell Carcinoma datasets. In cross-validation tests, the area under ROC curve rate for the cancer stages prediction has increased considerably when incorporating the learned microRNA dysregulation modules. The extracted modules from multiple independent subtypes differential analyses were found to have high agreement with microRNA family annotations, and they can also be used to identify mutual biomarkers between different subtypes. Among the top-ranked candidate microRNAs selected by the model, 87% were reported to be related to Lung Adenocarcinoma. The overall result demonstrates that clustering microRNAs from the dysregulation pattern between microRNAs and their targets leads to biomarkers with high precision and recall rate to known differentially expressed disease-associated microRNAs. CONCLUSIONS: The results indicated that our method improves microRNA biomarker selection by detecting similar microRNA dysregulational synergistic patterns across the multiple subtypes. Since microRNA-target dysregulations are implicated in many cancers, we believe this tool can have broad applications for discovery of novel microRNA biomarkers in heterogeneous cancer diseases.
Nhat Tran, Vinay Abhyankar, Kytai Nguyen, Jon Weidanz, Jean Gao
BMC Bioinform.5
2017 LiDiAimc: LincRNA-disease associations through inductive matrix completion
abstract
The dysregulations of long intergenic non-coding RNAs (lincRNAs) have shown to be linked with a wide variety of human diseases over the past few years. However, there are only a few lincRNA-disease association inference tools available with most of them relying on very specific type of prior knowledge about the lincRNAs and the diseases. They fall short in generalized association predictions when new type of knowledge becomes available, or to rank disease implications by a novel lincRNA. In this article, we proposed LiDiAimc, a method based on Inductive Matrix Completion strategy, offering generalized integration platform of any type of prior knowledge about the lincRNAs and the diseases. A benefit of the approach is that being an inductive learner, it can be applied to lincRNAs for disease association predictions that are not listed during model build-up phase, and vice versa, an approach unlike traditional matrix factorization frameworks. The proposed LiDiAimc method was applied to association data between human lincRNAs and OMIM disease phenotypes as well as a diverse set of knowledgebase of lincRNAs and diseases. Our method performs better than the state-of-the-art methods in terms of precision@k and recall@k at the top-k disease prioritization to the subject lincRNAs.
Ashis Kumer Biswas, Jean Gao
BIBM2
2017 MicroRNA dysregulational synergistic network: Learning context-specific MicroRNA dysregulations in lung cancer subtypes
abstract
Recently, microRNAs were found to have potential as both diagnostic biomarkers and therapeutic targets for lung cancer, especially in identifying early-stage cancer. However, a miRNA biomarker derived from the standard normal v.s tumor differential expression analysis is not robust, as its functional interactions with messenger-RNA targets may change between different lung cancer subtypes. Furthermore, as evidence suggests miRNA dysregulations and synergistic regulations are important to understanding many cancer diseases, it is important to consider the changes in miRNA-target associations among different lung cancer subtypes. We proposed a pipeline to identify miRNA synergistic modules with high context-specific functional similarity by constructing the MicroRNA Dysregulational Synergistic Network (MDSN). Then, we incorporated the extracted miRNA modules as prior knowledge to a Sparse Group Lasso classifier for more relevant selection of microRNA biomarkers, thereby improving classification results. We applied the method to the TCGA Lung Adenocarcinoma dataset and found extracted miRNA modules from independent subtypes differential analyses to have high agreement with known miRNA family assignments. Cross-validation test also demonstrates clustering miRNAs by considering the dysregulation between miRNAs and its targets results in more accurate prediction of cancer stage.
Nhat Tran, Vinay Abhyankar, Kytai Nguyen, Ishfaq Ahmad 0001, Jon Weidanz, Jean Gao
BIBM6
2017 Multi-Block Bipartite Graph for Integrative Genomic Analysis
abstract
Human diseases involve a sequence of complex interactions between multiple biological processes. In particular, multiple genomic data such as Single Nucleotide Polymorphism (SNP), Copy Number Variation (CNV), DNA Methylation (DM), and their interactions simultaneously play an important role in human diseases. However, despite the widely known complex multi-layer biological processes and increased availability of the heterogeneous genomic data, most research has considered only a single type of genomic data. Furthermore, recent integrative genomic studies for the multiple genomic data have also been facing difficulties due to the high-dimensionality and complexity, especially when considering their intra- and inter-block interactions. In this paper, we introduce a novel multi-block bipartite graph and its inference methods, MB2I and sMB2I, for the integrative genomic study. The proposed methods not only integrate multiple genomic data but also incorporate intra/inter-block interactions by using a multi-block bipartite graph. In addition, the methods can be used to predict quantitative traits (e.g., gene expression, survival time) from the multi-block genomic data. The performance was assessed by simulation experiments that implement practical situations. We also applied the method to the human brain data of psychiatric disorders. The experimental results were analyzed by maximum edge biclique and biclustering, and biological findings were discussed.
Mingon Kang, Juyoung Park, Dong-Chul Kim, Ashis Kumer Biswas, Chunyu Liu 0001, Jean Gao
IEEE ACM Trans. Comput. Biol. Bioinform.6
2016 Is EEG causal to fNIRs?
abstract
Causality analysis of simultaneous measurements of the brain's electrical activity and its hemodynamic activity provides the opportunity to study the neural underpinning of hemodynamic fluctuations. This multimodal analysis can also be used to extract valuable information regarding the location of the generators of various electrical events such as Alpha rhythms or epileptiform activity. To best of our knowledge, we are the first propose a method to assess causality from EEG to the hemodynamic activity measured using functional near-infrared spectroscopy (fNIRs). The main challenge in studying causality within this setting arises from the low sampling rate of the fNIRs and the mixed frequency nature of the data. Our method of analysis consists of two parts. Through a simple modification of Geweke's formulation of contamination, we first show that the low sampling frequency of the fNIRs does not cause contamination in estimating causality from EEG to fNIRs. We then apply a novel causality test to avoid the down-sampling of the EEG when measuring for causality. The method of analysis proposed here can be generalized to study causality in other biomedical signal analysis applications and mixed frequency settings.
Borzou Alipourfard, Jean Gao, Olajide Babawale, Hanli Liu
BIBM2
2016 Robust Inductive Matrix Completion strategy to explore associations between lincRNAs and human disease phenotypes
abstract
Long intergenic non-coding RNAs (lincRNAs) are associated with a wide variety of human diseases. Piles of data about the lincRNAs are becoming available, thanks to the High Throughput Sequencing (HTS) platforms, which open opportunity for cutting-edge machine learning and data mining approaches to analyze the disease association better. However, there are only a few in silico association inference tools available to date, and none of them utilizes the heterogeneous data about the lincRNAs and diseases. The standard Inductive Matrix Completion (IMC) technique provides with a platform among the two entities considering respective side information. But, it has two major issues pertaining to the noise and sparsity in the dataset. Thus, a robust version of IMC is needed to adequately address the issues. In this paper, we propose Robust Inductive Matrix Completion (RIMC) to address these challenges. Then, we applied RIMC to the available association dataset between the lincRNAs and OMIM disease phenotypes with a diverse set of side information of the both. The proposed method performs better than the state-of-the-art methods in terms of precision@k and recall@k at the top-k disease prioritization to the subject lincRNAs. Moreover, with an induction experiment we showed that RIMC performs superior than the standard IMC for ranking unexplored disease phenotypes to a set of known lincRNAs.
Ashis Kumer Biswas, Dong-Chul Kim, Mingon Kang, Jean Gao
BIBM4
2016 Integrative Gene Regulatory Network inference using multi-omics data
abstract
Biological network inference is of importance to understand underlying biological mechanisms. Gene regulatory networks describe molecular interactions of complex biological processes. Graph models are mainly used for gene regulatory networks, where nodes and edges represent genes and their regulations respectively. In the most research, the molecular interactions (edges) of gene regulatory networks are inferred from a single type of genomic data, e.g., gene expression data. However, gene expression is a product of sequential interactions of DNA sequence variations, single nucleotide polymorphism, copy number variation, histone modifications, transcription factor, DNA methylation, and many other factors. There are high-throughput genomic data that measure the various biological processes. We call the multiple types of genomics data as ‘multi-omics data’. In this paper, we propose an Integrative Gene Regulatory Network inference method (iGRN) that can incorporate multi-omics data and their interactions in the graph model of gene regulatory network. Copy number variation and DNA methylation were considered for multi-omics data in this paper. The proposed method, iGRN, was applied to the human brain data of psychiatric disorder. Through the experiments, iGRN showed its better performance on model representation and interpretation than other integrative methods in gene regulatory network inference.
Neda Zarayeneh, Jung Hun Oh, Donghyun Kim 0001, Chunyu Liu 0001, Jean Gao, Sang C. Suh, Mingon Kang
BIBM5
2016 Computational modeling of phagocyte transmigration for foreign body responses to subcutaneous biomaterial implants in mice
abstract
BACKGROUND: Computational modeling and simulation play an important role in analyzing the behavior of complex biological systems in response to the implantation of biomedical devices. Quantitative computational modeling discloses the nature of foreign body responses. Such understanding will shed insight on the cause of foreign body responses, which will lead to improved biomaterial design and will reduce foreign body reactions. One of the major obstacles in computational modeling is to build a mathematical model that represents the biological system and to quantitatively define the model parameters. RESULTS: In this paper, we considered quantitative inter connections and logical relationships among diverse proteins and cells, which have been reported in biological experiments and literature. Based on the established biological discovery, we have built a mathematical model while unveiling the key components that contribute to biomaterial-mediated inflammatory responses. For the parameter estimation of the mathematical model, we proposed a global optimization algorithm, called Discrete Selection Levenberg-Marquardt (DSLM). This is an extension of Levenberg-Marquardt (LM) algorithm which is a gradient-based local optimization algorithm. The proposed DSLM suggests a new approach for the selection of optimal parameters in the discrete space with fast computational convergence. CONCLUSIONS: The computational modeling not only provides critical clues to recognize current knowledge of fibrosis development but also enables the prediction of yet-to-be observed biological phenomena.
Mingon Kang, Jean Gao
BMC Bioinform.3
2015 An integrative genomic study for multimodal genomic data using multi-block bipartite graph
abstract
Human diseases involve a sequence of complex interactions in multiple biological processes. In particular, multiple genomic data such as Single Nucleotide Polymorphism (SNP), Copy Number Variation (CNV), and DNA Methylation (DM) and their interactions simultaneously play an important role in the variation of mRNA transcription in human diseases. However, despite of the widely known complex multi-layer biological processes and increased availability of the heterogeneous genomic data, most research has considered only a single type of the genomic data. Furthermore, recent integrative genomic studies for the multiple genomic data have also been facing difficulties due to the high-dimensionality and complexity, especially when considering their intra- and inter-block interactions. In this paper, we introduce a novel multi-block bipartite graph and its inference methods, MB2I and sMB2I, for the integrative genomic study. The proposed methods not only integrate the multiple genomic data but also incorporate their intra/inter-block interactions by using a multi-block bipartite graph. In addition, the methods can be used to predict quantitative traits (e.g. gene expression, survival time) from the multi-block genomic data. The outstanding performance was assessed by simulation experiments that implement practical situations.
Mingon Kang, Juyoung Park, Dong-Chul Kim, Ashis Kumer Biswas, Chunyu Liu 0001, Jean Gao
BIBM6
2015 Integrative approach for inference of gene regulatory networks using lasso-based random featuring and application to Psychiatric disorders
abstract
Inferring gene regulatory networks is one of the most interesting research areas in the systems biology. Many inference methods have been developed by using a variety of computational models and approaches. In this paper, we propose two network inference methods based on a lasso-based random feature selection algorithm (LARF). There are three main contributions. First, our z score-based method to measure gene expression variations from knockout data is more effective than similar criteria of related works. Second, we confirmed that the true regulator selection can be effectively improved by LARF. Lastly, we verified that an integrative approach can clearly outperform a single method when two different methods are effectively jointed. In the experiments, our method outperformed state of the art methods on simulated data, and LARF also was applied to the inference of gene regulatory networks associated with Psychiatric disorders.
Dong-Chul Kim, Mingon Kang, Ashis Kumer Biswas, Chunyu Liu 0001, Jean Gao
BIBM5
2015 eQTL epistasis: detecting epistatic effects and inferring hierarchical relationships of genes in biological pathways
abstract
MOTIVATION: Epistasis is the interactions among multiple genetic variants. It has emerged to explain the 'missing heritability' that a marginal genetic effect does not account for by genome-wide association studies, and also to understand the hierarchical relationships between genes in the genetic pathways. The Fisher's geometric model is common in detecting the epistatic effects. However, despite the substantial successes of many studies with the model, it often fails to discover the functional dependence between genes in an epistasis study, which is an important role in inferring hierarchical relationships of genes in the biological pathway. RESULTS: We justify the imperfectness of Fisher's model in the simulation study and its application to the biological data. Then, we propose a novel generic epistasis model that provides a flexible solution for various biological putative epistatic models in practice. The proposed method enables one to efficiently characterize the functional dependence between genes. Moreover, we suggest a statistical strategy for determining a recessive or dominant link among epistatic expression quantitative trait locus to enable the ability to infer the hierarchical relationships. The proposed method is assessed by simulation experiments of various settings and is applied to human brain data regarding schizophrenia. AVAILABILITY AND IMPLEMENTATION: The MATLAB source codes are publicly available at: http://biomecis.uta.edu/epistasis.
Mingon Kang, Chunling Zhang, Hyung-Wook Chun, Chris Ding, Chunyu Liu 0001, Jean Gao
Bioinform.6
2014 NMF-Based LncRNA-Disease Association Inference and Bi-Clustering
abstract
Long non-coding RNAs (lncRNAs) have been implicated in various biological processes, and are linked in many dysregulations. Researchers have reported large number of lncRNA associated human diseases over the past decade. In this article we employed the Non-negative Matrix Factorization method to develop a low-dimensional computational model that can describe the existing knowledge about lncRNA-disease associations represented in a two dimensional association matrix. The non-negativity constraints of the matrix and its corresponding factors ensure that each lncRNA's disease profile can be represented as an additive linear combination of the latent coordinates. To learn such a constrained model from an incomplete association matrix, several NMF formulations were developed. Based on our experiments, we found that the Sparse NMF obtained the best model among all the other models. Moreover, by exploiting the inherent bi-clustering ability of the NMF models, we extracted several lncRNA groups and disease groups that possess biological significance.
Ashis Kumer Biswas, Jean Gao, Baoju Zhang
BIBE2
2014 Multi-block and Multi-task Learning for Integrative Genomic Study
abstract
The importance of an integrative genomic study is steadily increasing in an emerging era of various high-throughput genomic data. Mechanisms of human diseases consist of complex interactions of multiple biological processes such as genetic, epigenetic, and transcriptional regulation. The collection of the multiple genomic data that represents the multiple processes is called 'multi-block data'. The multi-block data profiled from human disease samples provide comprehensive global snapshots of the diseases. Due to the rapid development of high-throughput technologies, the integrative genomic study using the multi-block data has been more highlighted than ever. However, in spite of its importance, there are only a few methodologies that can analyze such data. In this paper, we propose a novel Multi-Block and Multi-Task Learning (MBMTL) method for the integrative genomic study. We consider Single Nucleotide Polymorphism (SNP), Copy Number Variation (CNV), DNA methylation, and gene expression data as the multi-block data from four group samples of three major psychiatric disorders as well as data from a normal control. MBMTL identifies biomarkers that play important roles in explaining mechanisms of the human diseases from the multi-block data. We also take a multi-task problem into account so that we can identify different functions of the mechanisms. The performance of the proposed MBMTL was assessed by comparing it to a number of existing multi-block methods through simulation studies. We applied MBMTL to the multi-block data of the major psychiatric disorder samples.
Mingon Kang, Dong-Chul Kim, Chunyu Liu 0001, Baoju Zhang, Jean Gao
BIBE6
2014 Integration of DNA Methylation, Copy Number Variation, and Gene Expression for Gene Regulatory Network Inference and Application to Psychiatric Disorders
abstract
Biological network inference is a crucial problem to solve in Bioinformatics as most of biological process are based on bio molecular interactions. Many researchers have worked on especially the inference of gene regulatory networks where a node and edge represent a gene and regulation relationship respectively assuming that a gene can regulate another gene indirectly. However, a gene expression level can be influenced by not only genes and proteins but also other biological factors. Therefore, the inference could be more effective if those factors are considered in gene regulatory network inferences. In this paper, we propose an integrative approach to infer gene regulatory networks where a gene can be regulated by not only gene and but also DNA Methylation and copy number variation. It is assumed that a gene can be directly regulated by a single DNA Methylation and copy number variation at most. The simulation results show that our method outperforms popular and state-of-the-art methods of biological network inference. In addition, we applied the proposed method to psychiatric disorder data. The inferred networks provide the relationships within a set of genes that are more likely to be regulated by DNA Methylation and copy number variation of the genes.
Dong-Chul Kim, Mingon Kang, Baoju Zhang, Chunyu Liu 0001, Jean Gao
BIBE6
2014 Guest Editorial: Data Mining in Bioinformatics, Biomedicine, and Healthcare Informatics
abstract
The eight papers in this special issue focus on data mining in bioinformatics, biomedicine, and healthcare informatics. Four of the papers in this special issue were selected from papers presented at the 2012 IEEE Conference on Bioinformatics and Biomedicine and the other four came from open solicitation with a wide range of authors.
Jean Gao, Werner Dubitzky
IEEE J. Biomed. Health Informatics1
2014 Models and Methods for Quantitative Analysis of Surface-Enhanced Raman Spectra
abstract
The quantitative analysis of surface-enhanced Raman spectra using scattering nanoparticles has shown the potential and promising applications in in vivo molecular imaging. The diverse approaches have been used for quantitative analysis of Raman pectra information, which can be categorized as direct classical least squares models, full spectrum multivariate calibration models, selected multivariate calibration models, and latent variable regression (LVR) models. However, the working principle of these methods in the Raman spectra application remains poorly understood and a clear picture of the overall performance of each model is missing. Based on the characteristics of the Raman spectra, in this paper, we first provide the theoretical foundation of the aforementioned commonly used models and show why the LVR models are more suitable for quantitative analysis of the Raman spectra. Then, we demonstrate the fundamental connections and differences between different LVR methods, such as principal component regression, reduced-rank regression, partial least square regression (PLSR), canonical correlation regression, and robust canonical analysis, by comparing their objective functions and constraints.We further prove that PLSR is literally a blend of multivariate calibration and feature extraction model that relates concentrations of nanotags to spectrum intensity. These features (a.k.a. latent variables) satisfy two purposes: the best representation of the predictor matrix and correlation with the response matrix. These illustrations give a new understanding of the traditional PLSR and explain why PLSR exceeds other methods in quantitative analysis of the Raman spectra problem. In the end, all the methods are tested on the Raman spectra datasets with different evaluation criteria to evaluate their performance.
Shuo Li 0009, James O. Nyagilo, Digant P. Dave, Jean Gao
IEEE J. Biomed. Health Informatics4
2013 QLZCClust: Quaternary lempel-Ziv complexity based clustering of the RNA-seq read block segments
abstract
The Next Generation Sequencing platform, RNA-seq provides quantitative expression data that exhibit distinctive sequence patterns in the segments of the short-reads level and are found useful in clustering of those segments. However, the result does not reflect the functional chemistry of the non-coding RNAs (ncRNAs). The functions of the ncRNAs are deeply related to their secondary structures. Thus by exploring the clustering in terms of structural profiles of the read block segments rather than their sequence patterns would be essential and useful. We proposed the QLZCClust (Quaternary Lempel-Ziv complexity based Clustering) method which is an extension to the popular Lempel-Ziv algorithm to compute pairwise secondary structure distance. We applied QLZCClust on the short-read segments obtained from the RNA-seq experient and found that it can separate most miRNAs and the tRNAs. Moreover, it can be used to detect structural similarities among different classes of ncRNAs. We compared our algorithm with the clustering of two other structural distance measures - SimTree edit distance and RNAz based distance, and found that our method performs superior.
Ashis Kumer Biswas, Baoju Zhang, Jean Gao
BIBE4
2013 eQTL epistasis: Detecting complex interaction effects between multiple loci from eQTL data
abstract
The identification of expression quantitative trait loci (eQTL) epistasis, which is a non-linear interaction effect between two or more genetic loci that control quantitative traits, play an essential role in understanding the mechanisms of gene interactions in the complex diseases. However, many studies have ignored the possibility of the epistasis in spite of the countless empirical evidence of epistasis in Genetics. Thus, the epistasis research is still in its infancy. Furthermore, most eQTL epistasis studies have involved only the interaction model with additive effects, which lacks the power to represent either 1) other biologically putative interaction models or 2) the genetic dominance that is a superior relationship of an allele against another one at the same locus. To tackle the problems, we propose a general eQTL epistasis model that provides the capability to incorporate diverse interaction models so that it can be widely utilized as a generic equation of epistasis in the multiple testing of eQTL studies. The extensibility of the eQTL epistasis model into multivariate methods is also proposed to reduce the computational burden of the multiple testing and to consider group effects. As the example of the extensibility, we provide the methodology that embeds the eQTL epistasis model into the sparse canonical correlation analysis (SCCA) method. A globally optimal solution is provided and the performance was assessed by realistic simulation experiments. A study of psychiatric disorder diseases with the method was conducted as a target application.
Mingon Kang, Shuo Li 0009, Chunyu Liu 0001, Jean Gao
BIBM4
2013 Inference of gene regulatory networks by integrating gene expressions and genetic perturbations
abstract
To elucidate the overall relationships between gene expressions and genetic perturbations, many approaches for SNPs to be involved in gene regulatory network (GRNs) inference have been suggested. In the most of the inferences of networks named as SNP-Gene Regulatory Networks (SGRNs) inference, pairs of SNP-gene are still separately given by performing eQTL mappings but eQTLs are not identified during network constructions. In order to build a SGRN for a given set of genes and SNPs without pre-defined eQTL information, we propose a method that is based on a structural equation model. The method consists of three steps: (i) ridge regression, (ii) elastic net regression, and (iii) iterative adaptive lasso regression. In the inference, it is assumed that each gene has a single unknown eQTL. The first two steps are to remove false positive edges keeping as many true positive edges as possible, and then lastly, final edges are selected by iterative adaptive lasso regression that iteratively gives more weight to the edge whose coefficient value is relatively high. To evaluate the performance, the method is applied to data that is randomly generated from the simulated networks and parameters. The method is also applied to psychiatric disorder data. There are three main contributions in this work. First, the proposed method provides both the gene regulatory inference and the identification of eQTL that affect genes in the network. Second, the simulation result proves that an integrative approach of multiple regression methods can effectively detect true edges as well as filter false positive edges. Lastly, it is demonstrated by applying it to psychiatric disorder data that our inference without eQTL information can discover the SGRNs that partially confirms eQTLs identified by our previous work.
Dong-Chul Kim, Chunyu Liu 0001, Baoju Zhang, Jean Gao
BIBM5
2013 eQTL Mapping Study via Regularized Sparse Canonical Correlation Analysis
abstract
While genome-wide association studies (GWAS) have focused on discovering genetic loci mapped to a disease, expression quantitative trait loci (eQTL) studies combine micro array data and provide a powerful approach. Micro arrays allow one to measure thousands of gene expressions simultaneously and the advances in eQTL studies enable one to capture the insight of the genetic architecture of gene expression. A number of multivariate methods have been recently proposed to identify genetic loci which are linked to gene expression taking into account joint effects and relationships between the units rather than the single locus alone independently. However, the previous research has limitations, such as the lack of supporting the cis/tran-eQTL model into being accepted as a general genetics model. We propose a novel regularized eQTL association mapping detection (Reg-AMADE) method. We have focused on the following three problems. First, we need to take into account co-expressed genes without using clustering or partitioning techniques, as well as detecting linkage disequilibrium and the joint effect of multiple genetic markers. Secondly, we need to build a regularized model to support the cis- and trans-eQTL model observed in most association studies. Lastly, we need to discover the significant genes underlying within diseases rather than a common component. We also propose a new simulation experiment method that implements practical situations so that the results can be evaluated in the true sense instead of the assessment with random samples generated from multivariate normal distributions that most research has mainly used. The power to detect both the joint effect and grouping effect of SNPs and gene expressions is assessed in the simulation study.
Mingon Kang, Shuo Li 0009, Dong-Chul Kim, Chunyu Liu 0001, Baoju Zhang, Jean Gao
ICMLA (1)7
2013 A Fast Multi-component Latent Variable Regression Framework for Quantitative Analysis of Surface-Enhanced Raman Spectra
abstract
Surface-enhanced Raman spectroscopy (SERS) has been a routine method for the quantitative analysis of Nano-tags or biomarkers. The multivariate calibration (MC) model is normally used to reduce the bias from the inherent instability of Raman signals. To solve the more variables than observations, ill-conditioned problem within the MC model, latent variable regression (LVR) methods are usually used. In order to decide the optimized number of latent variables (LVs) used in the model, cross-validation methods are commonly used to test every possible number, and the one gives the minimum estimated error is returned as the optimized number. In this paper we present a new multi-component LVR together with a cross-validation framework to accelerate the time-consuming processes of optimizing number of LVs. It reduces the growth rate of the algorithms from O(k^2) to O(k), where k is the possible numbers of LVs. Experimental results show the estimated results of the two frameworks are equivalent and the running time of our new framework is evidently reduced.
Shuo Li 0009, James O. Nyagilo, Digant P. Dave, Baoju Zhang, Jean Gao
ICMLA (1)6
2012 CWT-PLSR for quantitative analysis of Raman spectrum
abstract
Quantitative analysis of Raman spectra using Surface Enhanced Raman scattering (SERS) nanoparticles has shown the potential and promising trend of development in vivo molecular imaging. Partial Least Square Regression (PLSR) methods are the state-of-the-art methods. But they rely on the whole intensities of Raman spectra and can not avoid the instable background. In this paper we design a new CWT-PLSR algorithm that uses mixing concentrations and the average continuous wavelet transform (CWT) coefficients of Raman spectra to do PLSR. The average CWT coefficients with a Mexican hat mother wavelet are robust representations of the Raman peaks, and the method can reduce the influences of the instable baseline and random noises during the prediction process. In the end, the algorithm is tested on three Raman spectrum data sets with three cross-validation methods, and the results show its robustness and effectiveness.
Shuo Li 0009, James O. Nyagilo, Digant P. Dave, Jean Gao
BIBM4
2012 A New Continuum Regression Method for Quantitative Analysis of Raman Spectrum
abstract
Quantitative analysis of Raman spectrum using Surface Enhanced Raman scattering (SERS) nanoparticles has shown the potential and promising trend of development in vivo molecular imaging. Because of the high dimension of Raman spectra and limited number of samples, latent variable regression methods, e.g. principal component regression (PCR), reduced-rank regression (RRR) and partial least squares (PLS), are commonly used. According to different criteria, these methods tend to seek different latent variables of the spectra data. For PCR and RRR, the latent variables tend to best represent the Raman spectra and best predict the concentrations. PLS balances the two criteria with an equal weight. We design a new continuum regression (NCR) method that uses a weight parameter α to control the portion of each criterion in the objective function, and embraces RRR (α = 0), PLS2 (α = 1) and PCR (α = ∞) as its special cases. The experimental results show that its performance is better than the other two continuum regression methods.
Shuo Li 0009, Jean Gao, James O. Nyagilo, Digant P. Dave
ICMLA (1)2
2011 A Framework for Personalized Medicine with Reverse Phase Protein Array and Drug Sensitivity
abstract
In this paper, we propose a framework for personalized medicine with Reverse-Phase Protein Array (RPPA) and drug sensitivity. The goal of personalized medicine is to provide an optimal drug to a patient by predicting the drug sensitivity. For the prediction, our method is based on naive Bayes classifier assuming that all features (proteins) are independent. Once the classifier is trained by RPPA data for the cancer, the sensitivity of target drug is predicted for a patient's sample. As a result, a set of drugs that has low sensitivity can be provided to a patient. Through this individualized therapy, the patients can be treated more effectively without the risk of wasting time and cost. In addition, we explore if naive Bayes classifier can be improved by using the dependency between features. To this end, learning Bayesian network is performed to infer the dependency map, and then the selected edges from the estimated network model is combined with the network structure of naive Bayes classifier. As our contribution, the experiments with lung cancer data prove that RPPA data can be used to profile patient for drug sensitivity prediction, and also our proposed personalize medicine system achieved approximately 94% prediction accuracy.
Dong-Chul Kim, Jean Gao, Chin-Rang Yang
BIBM2
2011 Probabilistic Partial Least Square Regression: A Robust Model for Quantitative Analysis of Raman Spectroscopy Data
abstract
Raman spectroscopy has been one of the most sensitive techniques widely used in chemical and pharmaceutical material identification research ever since it is invented based on Raman scattering theory, because of the fingerprints property of Raman signals to different materials. With the latest development of surface enhanced Raman scattering (SERS) nanoparticles, Raman spectroscopy is now used in more and more quantitative analysis applications. But due to the unavoidable instable problem of Raman spectroscopy signal, as well as the high signal dimension and small sample number problem, it is badly in need of a robust and accurate signal quantitative analysis method. Based on Partial Least Square Regression (PLSR) method, Probabilistic PCA and Probabilistic curve-fitting idea, we propose a new Probabilistic-PLSR (PPLSR) model. It explains PLSR from a probabilistic viewpoint and deeply describes the physical meaning of PLSR model. It is a solid foundation to develop more robust and accurate probabilistic PLSR models with Bayesian model in order to solve the over-fitting problem. And since this model adds a regularization term in the matrix of regression coefficients, the estimated result is more robust than PLSR model. We also provide an EM Algorithm to estimate the parameters of the model from sample data. To take fully use of the valuable data, we design two experiments, leave-one-out and cross-validation-on-average-signal, on one real Raman spectroscopy signal data set. By comparing with results from traditional Least Square (LS) method and traditional PLSR, we demonstrate PPLSR is more robust and accurate.
Shuo Li 0009, Jean Gao, James O. Nyagilo, Digant P. Dave
BIBM2
2011 Fast Kernel Discriminant Analysis for Classification of Liver Cancer Mass Spectra
abstract
The classification of serum samples based on mass spectrometry (MS) has been increasingly used for monitoring disease progression and for diagnosing early disease. However, the classification task in mass spectrometry data is extremely challenging due to the very huge size of peaks (features) on mass spectra. Linear discriminant analysis (LDA) has been widely used for dimension reduction and feature extraction in many applications. However, the conversional LDA suffers from the singularity problem when dealing with high-dimensional features. Another critical limitation is its linearity property which results in failing in classification problems over nonlinearly clustered data sets. To overcome such problems, we develop a new fast kernel discriminant analysis (FKDA) that is pretty fast in the calculation of optimal discriminant vectors. FKDA is applied to the classification of liver cancer mass spectrometry data that consist of three categories: hepatocellular carcinoma, cirrhosis, and healthy that was originally analyzed by Ressom et al. We demonstrate the superiority and effectiveness of FKDA when compared to other classification techniques.
Jung Hun Oh, Jean Gao
IEEE ACM Trans. Comput. Biol. Bioinform.2
2011 Branch-and-Bound for Model Selection and Its Computational Complexity
abstract
Branch-and-bound methods are used in various data analysis problems, such as clustering, seriation and feature selection. Classical approaches of branch-and-bound based clustering search through combinations of various partitioning possibilities to optimize a clustering cost. However, these approaches are not practically useful for clustering of image data where the size of data is large. Additionally, the number of clusters is unknown in most of the image data analysis problems. By taking advantage of the spatial coherency of clusters, we formulate an innovative branch-and-bound approach, which solves clustering problem as a model-selection problem. In this generalized approach, cluster parameter candidates are first generated by spatially coherent sampling. A branch-and-bound search is carried out through the candidates to select an optimal subset. This paper formulates this approach and investigates its average computational complexity. Improved clustering quality and robustness to outliers compared to conventional iterative approach are demonstrated with experiments.
Ninad Thakoor, Jean Gao
IEEE Trans. Knowl. Data Eng.2
2010 Learning Proteomic Network Structure by a New Hill Climbing Algorithm
abstract
As a progressive, degenerative disease, ataxia telangiectasia (A-T) is caused by a gene mutation (ATM) and is a predisposition to cancer. Understanding the impaired signaling networks caused by ATM will help minimizing the damage and finding effective therapies. The goal of this work is to investigate the dynamic change of ATM-dependent signaling pathways under the treatment of different radiation dosages. A reverse-phase protein microarray (RPPM) in conjunction with quantum dots nano-crystal technology is used for the quantitative measurement. To discover the proteomic pathways affected in ATM cells, a new hill climbing algorithm is developed based on mutual information, the classical hill-climbing method, and the optimization of the local structure. More trusted biology networks are thus defined by the new approach. The study was carried out at different time points under different dosages for cell lines with and without ATM mutation. To validate the performance of the proposed algorithm, comparison experiments were also implemented using public networks.
Dong-Chul Kim, Jean Gao, Chin-Rang Yang
BIBE2
2010 Computational modeling of phagocyte transmigration during biomaterial-mediated foreign body responses
abstract
One of the major obstacles in computational modeling of a biological system is to determine a large number of parameters in the mathematical equations representing biological properties of the system. To tackle this problem, we have developed a global optimization method, called Discrete Selection Levenberg-Marquardt (DSLM), for parameter estimation. For fast computational convergence, DSLM suggests a new approach for the selection of optimal parameters in the discrete spaces, while other global optimization methods such as genetic algorithm and simulated annealing use heuristic approaches that do not guarantee the convergence. As a specific application example, we have targeted understanding phagocyte transmigration which is involved in the fibrosis process for biomedical device implantation. The goal of computational modeling is to construct an analyzer to understand the nature of the system. Also, the simulation by computational modeling for phagocyte transmigration provides critical clues to recognize current knowledge of the system and to predict yet-to-be observed biological phenomenon.
Mingon Kang, Jean Gao
BIBM2
2010 Eigenspectra, a robust regression method for multiplexed Raman spectra analysis
abstract
Raman spectroscopy has been one of the most sensitive techniques widely used in chemical and pharmaceutical research. With the latest development of surface enhanced Raman scattering (SERS) nanoparticles, the application now can be extended to bioimaging and biosensing. In this study, we demonstrate the ability of Raman spectroscopy to separate multiple spectral fingerprints using Raman nanotags after injection. The competence will further be used as functional agents for diagnostic molecular imaging applications. In this paper, a machine learning method is proposed to estimate the mixing ratios of each source signal from a mixture signal. The method first decomposes the training mixture signal matrix into a number of components and meanwhile keeps the maximum linear relationship between the new coordinate and ground truth ratio matrix. Then a regression coefficient matrix is formed by the component matrix. Traditional regression methods provide poor decomposition results due to various factors in sample preparation and machine operation that lead to the stochastic nature of Raman spectrum. The robustness of the proposed method was compared with least square and weighted least square methods.
Shuo Li 0009, Jean Gao, James O. Nyagilo, Digant P. Dave
BIBM2
2010 Functional proteomic pattern identification under low dose ionizing radiation
Young Bun Kim, Chin-Rang Yang, Jean Gao
Artif. Intell. Medicine3
2010 Learning biological network using mutual information and conditional independence
abstract
BACKGROUND: Biological networks offer us a new way to investigate the interactions among different components and address the biological system as a whole. In this paper, a reverse-phase protein microarray (RPPM) is used for the quantitative measurement of proteomic responses. RESULTS: To discover the signaling pathway responsive to RPPM, a new structure learning algorithm of Bayesian networks is developed based on mutual Information, conditional independence, and graph immorality. Trusted biology networks are thus predicted by the new approach. As an application example, we investigate signaling networks of ataxia telangiectasis mutation (ATM). The study was carried out at different time points under different dosages for cell lines with and without gene transfection. To validate the performance of the proposed algorithm, comparison experiments were also implemented using three well-known networks. From the experiment results, our approach produces more reliable networks with a relatively small number of wrong connection especially in mid-size networks. By using the proposed method, we predicted different networks for ATM under different doses of radiation treatment, and those networks were compared with results from eight different protein protein interaction (PPI) databases. CONCLUSIONS: By using a new protein microarray technology in combination with a new computational framework, we demonstrate an application of the methodology to the study of biological networks of ATM cell lines under low dose ionization radiation.
Dong-Chul Kim, Chin-Rang Yang, Jean Gao
BMC Bioinform.4
2010 Embedded planar surface segmentation system for stereo images
Ninad Thakoor, Jean Gao, Sungyong Jung
Mach. Vis. Appl.2
2010 Multibody Structure-and-Motion Segmentation by Branch-and-Bound Model Selection
abstract
An efficient and robust framework is proposed for two-view multiple structure-and-motion segmentation of unknown number of rigid objects. The segmentation problem has three unknowns, namely the object memberships, the corresponding fundamental matrices, and the number of objects. To handle this otherwise recursive problem, hypotheses for fundamental matrices are generated through local sampling. Once the hypotheses are available, a combinatorial selection problem is formulated to optimize a model selection cost which takes into account the hypotheses likelihoods and the model complexity. An explicit model for outliers is also added for robust segmentation. The model selection cost is minimized through the branch-and-bound technique of combinatorial optimization. The proposed branch-and-bound approach efficiently searches the solution space and guarantees optimality over the current set of hypotheses. The efficiency and the guarantee of optimality of the method is due to its ability to reject solutions without explicitly evaluating them. The proposed approach was validated with synthetic data, and segmentation results are presented for real images.
Ninad Thakoor, Jean Gao, Venkat Devarajan
IEEE Trans. Image Process.2
2009 Computation complexity of branch-and-bound model selection
abstract
Segmentation problems are one of the most important areas of research in computer vision. While segmentation problems are generally solved with clustering paradigms, they formulate the problem as recursive. Additionally, most approaches need the number of clusters to be known beforehand. This requirement is unreasonable for majority of the computer vision problems. This paper analyzes the model selection perspective which can overcome these limitations. Under this framework multiple hypotheses for cluster centers are generated using spatially coherent sampling. An optimal subset of these hypotheses is selected according to a model selection criterion. The selection can be carried out with a branch-and-bound procedure. The worst case complexity of any branch-and-bound algorithm is exponential. However, the average complexity of the algorithm is significantly lower. In this paper, we develop a framework for analysis of average complexity of the algorithm from the statistics of model selection costs.
Ninad Thakoor, Venkat Devarajan, Jean Gao
ICCV3
2009 A kernel-based approach for detecting outliers of high-dimensional biological data
abstract
BACKGROUND: In many cases biomedical data sets contain outliers that make it difficult to achieve reliable knowledge discovery. Data analysis without removing outliers could lead to wrong results and provide misleading information. RESULTS: We propose a new outlier detection method based on Kullback-Leibler (KL) divergence. The original concept of KL divergence was designed as a measure of distance between two distributions. Stemming from that, we extend it to biological sample outlier detection by forming sample sets composed of nearest neighbors. KL divergence is defined between two sample sets with and without the test sample. To handle the non-linearity of sample distribution, original data is mapped into a higher feature space. We address the singularity problem due to small sample size during KL divergence calculation. Kernel functions are applied to avoid direct use of mapping functions. The performance of the proposed method is demonstrated on a synthetic data set, two public microarray data sets, and a mass spectrometry data set for liver cancer study. Comparative studies with Mahalanobis distance based method and one-class support vector machine (SVM) are reported showing that the proposed method performs better in finding outliers. CONCLUSION: Our idea was derived from Markov blanket algorithm that is a feature selection method based on KL divergence. That is, while Markov blanket algorithm removes redundant and irrelevant features, our proposed method detects outliers. Compared to other algorithms, our proposed method shows better or comparable performance for small sample and high-dimensional biological data. This indicates that the proposed method can be used to detect outliers in biological data sets.
Jung Hun Oh, Jean Gao
BMC Bioinform.2
2009 An Extended Markov Blanket Approach to Proteomic Biomarker Detection From High-Resolution Mass Spectrometry Data
abstract
High-resolution matrix-assisted laser desorption/ionization time-of-flight mass spectrometry has recently shown promise as a screening tool for detecting discriminatory peptide/protein patterns. The major computational obstacle in finding such patterns is the large number of mass/charge peaks (features, biomarkers, data points) in a spectrum. To tackle this problem, we have developed methods for data preprocessing and biomarker selection. The preprocessing consists of binning, baseline correction, and normalization. An algorithm, extended Markov blanket, is developed for biomarker detection, which combines redundant feature removal and discriminant feature selection. The biomarker selection couples with support vector machine to achieve sample prediction from high-resolution proteomic profiles. Our algorithm is applied to recurrent ovarian cancer study that contains platinum-sensitive and platinum-resistant samples after treatment. Experiments show that the proposed method performs better than other feature selection algorithms. In particular, our algorithm yields good performance in terms of both sensitivity and specificity as compared to other methods.
Jung Hun Oh, Prem Gurnani, John Schorge, Kevin P. Rosenblatt, Jean Gao
IEEE Trans. Inf. Technol. Biomed.5
2008 Signaling biomarker pattern discovery using reverse phase protein microarray
abstract
Understanding the molecular mechanism of aging will allow human beings to develop rationale strategies for therapeutic interventions in aging related diseases. The goal of this study is to investigate the effect of a new protein, Klotho protein, on FGF (fibroblast growth factor) signaling. To identify the signaling molecules, two emerging technologies, high-throughput siRNA (small inference RNA) and reverse phase protein microarray (RPPM) are utilized. To quantitatively analyze the patterns of siRNA knockdowns, we present a Discriminative Feature Pattern Identification System (DFPIS) to identify contributing nuclear hormone receptors. Computational analysis results using HEK293 (human embryonic kidney) cells knocked down from siRNAs and screened by protein microarray were presented.
Young Bun Kim, Jean Gao, Johanne V. Pastor, Kevin P. Rosenblatt
BIBE2
2008 A hybrid computational model for phagocyte transmigration
abstract
Phagocyte transmigration is the initiation of a series of phagocyte responses that are believed important in the formation of fibrotic capsules surrounding implanted medical devices. Understanding the molecular mechanisms governing phagocyte transmigration is highly desired in order to improve the stability and functionality of the implanted devices. A hybrid computational model that combines control theory and kinetics Monte Carlo (KMC) algorithm is proposed to simulate and predict phagocytes responses at molecular level. In order to mimic various biological knockout experiments, a general external control scenario is designed. The stochastic nature inherent to phagocyte transmigration is captured by KMC. A new formula is derived to calculate the transition rates as inputs to KMC. This formulation might quantify biological interactions in a general manner which is beyond the scope of the traditional chemical reaction kinetics.
Jiaxing Xue, Jean Gao
BIBE2
2008 Functional Proteomic Pattern Identification under Low Dose Ionizing Radiation
abstract
The goal of this study is to explore and to understand the dynamic responses of signaling pathways to low dose ionizing radiation (IR). Low dose radiation (10 cGy or lower) affects several signaling pathways including DNA repair, survival, cell cycle, cell growth, and cell death. To detect the possibly regulatory protein/kinase functions, an emerging reverse-phase protein microarray (RPPM) in conjunction with quantum dots nano-crystal technology is used as a quantitative detection system. The dynamic responses are observed under different time points and radiation doses. To quantitatively determine the responsive protein/kinases and to discover the network motifs, we present a Discriminative Network Pattern Identification System (DiNPIS). Instead of simply identifying proteins contributing to the pathways, this methodology takes into consideration of protein dependencies which are represented as Strong Jumping Emerging Patterns (SJEP). Furthermore, infrequent patterns though occurred will be considered irrelevant. The whole framework consists of three steps: protein selection, protein pattern identification, and pattern annotation. Computational results of analyzing ATM (ataxia-telangiectasia mutated) cells treated with six different IR doses up to 72 hours are presented.
Young Bun Kim, Jean Gao, Chin-Rang Yang
BIBM2
2008 Biological Data Outlier Detection Based on Kullback-Leibler Divergence
abstract
Outlier detection is imperative in biomedical data analysis to achieve reliable knowledge discovery. In this paper, a new outlier detection method based on Kullback-Leibler (KL) divergence is presented. The original concept of KL divergence was designed as a measure of distance between two distributions. Stemming from that, we extend it to biological sample outlier detection by forming sample sets composed of nearest neighbors. To handle the non-linearity during the KL divergence calculation and to tackle with the singularity problem due to small sample size, we map the original data into a higher feature space and apply kernel functions without resorting to a mapping function. A sample possessing the largest KL divergence is detected as an outlier. The proposed method is tested with one synthetic data, two public gene expression data sets, and our own mass spectrometry data generated for prostate cancer study.
Jung Hun Oh, Jean Gao, Kevin P. Rosenblatt
BIBM2
2008 Branch-and-bound hypothesis selection for two-view multiple structure and motion segmentation
abstract
An efficient and robust framework for two-view multiple structure and motion segmentation is proposed. To handle this otherwise recursive problem, hypotheses for the models are generated by local sampling. Once these hypotheses are available, a model selection problem is formulated which takes into account the hypotheses likelihoods and model complexity. An explicit model for outliers is also added for robust model selection. The model selection criterion is optimized through branch-and-bound technique of combinatorial optimization which guaranties optimality over current set of hypotheses by efficient search of solution space.
Ninad Thakoor, Jean Gao
CVPR2
2008 Multi-stage branch-and-bound for maximum variance disparity clustering
abstract
A split-and-merge framework based on a maximum variance criterion is proposed for disparity clustering. The proposed algorithm transforms low-level stereo disparity information to mid-level planar surface information which can be used further to carry out high-level computer vision tasks such as shape classification. Unlike conventional clustering, the proposed algorithm assumes that the number of clusters is unknown. Instead, a maximum variance criterion is applied to extract planar surfaces from the disparity image. The split phase of the algorithm creates clusters based on spatial continuity and the merge phase combines these clusters such that variance per cluster does not exceeded an allowable value. For efficient maximum variance clustering, a greedy branch-and-bound procedure is introduced. Efficiency of the approach is verified through experiments.
Ninad Thakoor, Venkat Devarajan, Jean Gao
ICPR3
2008 Biomarker selection and sample prediction for multi-category disease on MALDI-TOF data
abstract
MOTIVATION: Diseases normally progress through several stages. Therefore, biomarkers corresponding to each stage may exist. To deal with such a multi-category problem, including sample stage prediction and biomarker selection, we propose methods for classification and feature selection. The proposed classification method is based on two schemes: error-correcting output coding (ECOC) and pairwise coupling (PWC). The final decision for a test sample prediction is an integration of these two schemes. The biomarker pattern for distinguishing each disease category from another one is achieved by the development of an extended Markov blanket (EMB) feature selection method. RESULTS: In this study, a liver cancer matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) mass spectrometry (MS) dataset was used, which comprises hepatocellular carcinoma (HCC), cirrhosis, and healthy spectra. Peak patterns were discovered for distinguishing pairwise categories among the three classes. Importance and reliability of individual peaks were presented by the measurements of certain weight values and frequencies. The classification capability of the proposed approach was compared with classical ECOC, random forest, Naive Bayes, and J48 methods. AVAILABILITY: Supplementary materials are available at http://visionlab.uta.edu/biomarker/bioinfo.htm.
Jung Hun Oh, Young Bun Kim, Prem Gurnani, Kevin P. Rosenblatt, Jean Gao
Bioinform.5
2008 Multihypothesis Prior for Segmentation of Stereo Disparity
abstract
An iterative segmentation-estimation framework for segmentation of planar surfaces in the disparity space is proposed. Disparity of a scene is modeled by approximating various surfaces in the scene to be planar. Uncertainty in the stereo disparity of these surfaces is modeled with a hidden Markov random field. Planar surface labels are the hidden variables of the model. Surface labels are estimated during the segmentation phase of the framework with help of underlying plane parameters and the spatial coherency prior. After segmentation, the planar surfaces are separated into spatially continuous regions. Each spatially continuous region is treated as a hypothesis for the underlying plane. Hypotheses for each plane are then combined to form a multihypothesis Gaussian mixture prior. The underlying planar surface parameters are estimated with maximum a posteriori estimation from the multihypothesis prior. The iterative process is continued until the labels converge or satisfactory segmentation is achieved. Experimental results are presented with synthetic as well as real-life scenes to demonstrate the success of the proposed method.
Ninad Thakoor, Jean Gao, Venkat Devarajan
IEEE Signal Process. Lett.2
2008 Wireless Sensor Network Lifetime Analysis Using Interval Type-2 Fuzzy Logic Systems
abstract
Extending the lifetime of the energy constrained wireless sensor networks is a crucial challenge in sensor network research. In this paper, we present a novel approach based on fuzzy logic systems to analyze the lifetime of a wireless sensor network. We demonstrate that a type-2 fuzzy membership function (MF), i.e., a Gaussian MF with uncertain standard deviation (std) is most appropriate to model a single node lifetime in wireless sensor networks. In our research, we study two basic sensor placement schemes: square-grid and hex-grid. Two fuzzy logic systems (FLSs): a singleton type-1 FLS and an interval type-2 FLS are designed to perform lifetime estimation of the sensor network. We compare our fuzzy approach with other nonfuzzy schemes in previous papers. Simulation results show that FLS offers a feasible method to analyze and estimate the sensor network lifetime and the interval type-2 FLS in which the antecedent and the consequent membership functions are modeled as Gaussian with uncertain std outperforms the singleton type-1 FLS and the nonfuzzy schemes.
Haining Shu, Qilian Liang, Jean Gao
IEEE Trans. Fuzzy Syst.3
2008 Multistage Branch-and-Bound Merging for Planar Surface Segmentation in Disparity Space
abstract
An iterative split-and-merge framework for the segmentation of planar surfaces in the disparity space is presented. Disparity of a scene is modeled by approximating various surfaces in the scene to be planar. In the split phase, the number of planar surfaces along with the underlying plane parameters is assumed to be known from the initialization or from the previous merge phase. Based on these parameters, planar surfaces in the disparity image are labeled to minimize the residuals between the actual disparity and the modeled disparity. The labeled planar surfaces are separated into spatially continuous regions which are treated as candidates for the merging that follows. The regions are merged together under a maximum variance constraint while maximizing the merged area. A multistage branch-and-bound algorithm is proposed to carry out this optimization efficiently. Each stage of the branch-and-bound algorithm separates a planar surface from the set of spatially continuous regions. The multistage merging estimates the number of planar surfaces and their labeling. The splitting and the multistage merging is repeated till convergence is reached or satisfactory results are achieved. Experimental results are presented for variety of stereo image data.
Ninad Thakoor, Jean Gao, Venkat Devarajan
IEEE Trans. Image Process.2
2007 Biomarker Selection for Predicting Alzheimer Disease Using High-Resolution MALDI-TOF Data
abstract
High-resolution MALDI-TOF (matrix-assisted laser desorption/ionization time-of-flight) mass spectrometry has shown promise as a screening tool for detecting discriminatory peptide/protein patterns. The major computational obstacle in analyzing MALDI-TOF data is the large number of mass/charge peaks (a.k.a. features, data points). With such a huge number of data points for a single sample, efficient feature selection is critical for unequivocal protein pattern discovery. In this paper, we propose a feature selection method and a new biclassification algorithm based on error-correcting output coding (ECOC) in multiclass problems. Our scheme is applied to the analysis of alzheimer's disease (AD) data. To validate the performance of the proposed algorithm, experiments are performed in comparison with other methods. We show that our proposed framework outperforms not only the standard ECOC framework but also other algorithms.
Jung Hun Oh, Young Bun Kim, Prem Gurnani, Kevin P. Rosenblatt, Jean Gao
BIBE5
2007 Phagocyte Transmigration Modeling Using System Dynamic Controls
abstract
Phagocyte transmigration is one of major phagocyte responses which are important in pathogenesis of flbrotic responses to implanted medical devices. Excessive flbrotic responses may lead to the failure of various medical implants. Therefore, better understanding of molecular mechanisms governing phagocyte transmigration is important in the success of development of implanted medical devices. This paper hypothesizes the biological components involved in phagocyte transmigration and the relationship between these components. A mathematical model based on system theory and control theory is built to support the hypothesis, so as to provide the prediction of some un-measured dynamic features in the phagocyte transmigration process. External controls are also supplied by the mathematical model to analogize diverse knockout experiments which contribute the construction of the biological hypothesis. Simulation results demonstrate the effectiveness of the mathematical model.
Jiaxing Xue, Jean Gao
BIBE2
2007 A Novel Classification Method for Analysis of Multi-stage Diseases via Mass Spectrometric Data
abstract
Multi-category classification is one of the challenging issues in medical data analysis. We propose a new bi- classification algorithm for the multi-class classification, which is comprised of two schemes: error-correcting output coding (ECOC) and pairwise coupling (PWC). After fea- ture reduction in both schemes, each corresponding classi- fication strategy is performed. For a test sample, two class labels that are predicted in both schemes are compared. If two class labels are the same, we assign the test sample to an identical label; otherwise, only for samples belonging to different classes predicted from two schemes, a retraining method is employed. Our scheme is applied to the analysis of a MALDI-TOF data set which consists of hepatocellular carcinoma (HCC) patients, cirrhosis patients and healthy individuals. To validate the performance of our proposed algorithm, experiments were performed in comparison with other classification methods.
Jung Hun Oh, Young Bun Kim, Jean Gao
BIBM3
2007 Multiple Interacting Subcellular Structure Tracking by Sequential Monte Carlo Method
abstract
With the wide application of green fluorescent protein (GFP) in the study of live cells, there is a surging need for the computer-aided analysis on the huge amount of im- age sequence data acquired by the advanced microscopy devices. One of such tasks is the motility analysis of the multiple subcellular structures. In this paper, an algorithm using sequential Monte Carlo (SMC) method for multiple interacting object tracking is proposed. First, marker resid- ual image is applied to detect individual subcellular struc- ture automatically, and to represent all the objects together using the joint state. Then the interaction between ob- jects in the 2D plane is modeled by augmenting an extra dimension and evaluating the overlapping relationship in the 3D space. Finally, the distribution of the dimension varying joint state is sampled efficiently by Reversible jump Markov chain Monte Carlo (RJMCMC) algorithm with a novel height swap move. The experimental results show that our method is promising.
Quan Wen 0001, Jean Gao, Katherine Luby-Phelps
BIBM2
2007 Mathematical Modeling of Phagocyte Chemotaxis toward and Adherence to Biomaterial Implants
abstract
Medical implanted devices usually prompt foreign body reactions that lead to the formation of fibrotic capsule surrounding the implanted devices. Cascade phagocyte responses, such as phagocyte chemotaxis and phagocyte adherence, are believed important in pathogenesis of fibrotic responses. Therefore, in-depth understanding of phagocyte chemotaxis and adherence processes is important to successfully reduce the fibrotic responses to improve the stability and functionality of implanted medical devices. This paper hypothesizes the molecular mechanisms governing phagocyte chemotaxis and adherence. To support our hypotheses and to provide the prediction of un-measured dynamic features in phagocyte chemotaxis and adherence processes, mathematical models based on system and control theory are built. To maximize the prediction ability and to guide future experimental design, external controls are provided by the mathematical model. The effectiveness of the mathematical model is demonstrated by computer simulation.
Jiaxing Xue, Jean Gao
BIBM2
2007 Real-time Planar Surface Segmentation in Disparity Space
abstract
An iterative Segmentation-Estimation framework for segmentation of planar surfaces in the disparity space is implemented on a Digital Signal Processor (DSP). Disparity of a scene is modeled by approximating various surfaces in the scene to be planar. The surface labels are estimated during the segmentation phase of the framework with help of the underlying plane parameters. After segmentation, planar surfaces are separated into spatially continuous regions. The largest of these regions is used to compute the estimates for the plane parameters. The iterative process is continued till convergence. The algorithm was optimized and implemented on TMS320DM642 based embedded system that operates at 3 to 5 frames per second on images of size 320 x 240.
Ninad Thakoor, Sungyong Jung, Jean Gao
CVPR3
2007 A Statistical Approach for Intensity Loss Compensation of Confocal Microscopy Images
abstract
In this paper a probabilistic technique for compensation of intensity loss in the confocal microscopy images is presented. Confocal microscopy images are modeled as a mixture of two Gaussians, one representing the background and another corresponding to the foreground. Images are segmented into foreground and background by applying expectation maximization (EM) algorithm to the mixture. Final intensity compensation is carried out by scaling and shifting the original intensities with help of parameters estimated for the foreground. Since foreground is separated to calculate the compensation parameters, the method is effective even when image structure changes from frame to frame. As intensity decay function (IDF) is not used, complexity associated with estimation of IDF parameters is eliminated. Also, images can be compensated out of order as only information from the reference image is required for compensation of any image. These properties make our method an ideal tool for intensity compensation of confocal microscopy images which can suffer intensity loss due to absorption/scattering of light as well as photobleaching and can change structure from optical/temporal section to section due to change in the depth of specimen or due to a living specimen. The proposed method was tested with number of image stacks and results for one of the stacks are presented here to demonstrate the effectiveness of the method.
Sowmya Gopinath, Ninad Thakoor, Jean Gao, Katherine Luby-Phelps
ICIP (6)3
2007 Hidden Markov Model-Based Weighted Likelihood Discriminant for 2-D Shape Classification
abstract
The goal of this paper is to present a weighted likelihood discriminant for minimum error shape classification. Different from traditional maximum likelihood (ML) methods, in which classification is based on probabilities from independent individual class models as is the case for general hidden Markov model (HMM) methods, proposed method utilizes information from all classes to minimize classification error. The proposed approach uses a HMM for shape curvature as its 2-D shape descriptor. We introduce a weighted likelihood discriminant function and present a minimum classification error strategy based on generalized probabilistic descent method. We show comparative results obtained with our approach and classic ML classification with various HMM topologies alongside Fourier descriptor and Zernike moments-based support vector machine classification for a variety of shapes.
Ninad Thakoor, Jean Gao, Sungyong Jung
IEEE Trans. Image Process.2
2006 Unsupervised Gene Selection For High Dimensional Data
abstract
In this paper, we present a new hybrid approach for unsupervised gene selection. Hybrid approaches try to utilize different evaluation criteria of the filter approaches and wrapper approaches in different search stages. Our method thus uses a two-step approach to identify informative genes. The first step retrieves gene subsets with original physical meaning based on their capacities to reproduce sample projections on principle components by applying the least-square-estimation based evaluation. The second step then searches for the best gene subsets that maximize clustering performance. When applied to a gene expression dataset of leukemia, the method identified a small set of genes whose expression is highly predictive
Young Bun Kim, Jean Gao
BIBE2
2006 Prediction of labor for pregnant women using high-resolution mass spectrometry data
abstract
High-resolution MALDI-TOF (matrix-assisted laser desorption/ionization time-of-flight) mass spectrometry has shown promise as a screening tool for detecting discriminatory protein patterns. The major computational obstacle in analyzing MALDI-TOF data is a large number of mass/charge peaks (a.k.a. features, data points). With the number of data points easily going beyond one million for a single sample, efficient feature selection is critical for unequivocal protein pattern discovery. To tackle this problem, we have developed a multi-step strategy for data preprocessing and afterwards feature selection. The preprocessing is composed of binning, baseline correction, and normalization. For the preprocessed data, we propose a new feature subset selection method that is a hybrid filter/wrapper approach. Based on the two feature subsets for each feature, high and low correlated subsets, a feature is assigned a weight which indicates the extent of feature importance. Our scheme is applied to the analysis of labor dataset to predict delivery time of pregnant women. To validate the performance of the proposed algorithm, experiments are performed in comparison with other feature selection and classification methods. We show that our proposed approach outperforms other algorithms
Jung Hun Oh, Animesh Nandi, Prem Gurnani, Peter Bryant-Greenwood, Kevin P. Rosenblatt, Jean Gao
BIBE6
2006 A New Hybrid Approach for Unsupervised Gene Selection
abstract
In recent years, unsupervised gene (feature) selection has become an integral part of microarray analysis because of the large number of genes and complexity in biological systems. Principal components analysis (PCA) is one of the approaches which has been applied, even though principal components (PCs) have no clear physical meanings. In this paper, we present a PCA based feature selection within a wrapper framework called PFSBEM (hybrid PCA based feature selection and boost-expectation-maximization clustering). PFSBEM uses a two-step approach to select features. The first step retrieves feature subsets with original physical meaning based on their capacities to reproduce sample projections on PCs. The second step then searches for the best feature subsets that maximize clustering performance. Experiment results clearly show that our feature sets improve the class prediction with respect to the chosen performance criteria
Young Bun Kim, Jean Gao
CIBCB2
2006 Classification of Relapse Ovarian Cancer on MALDI-TOF Mass Spectrometry Data
abstract
Ovarian cancer recurs at the rate of 75% within a few months or several years later after therapy. Early recurrence, though responding better to treatment, is difficult to detect. Recently, high-resolution MALDI-TOF (matrix-assisted laser desorption/ionization time-of-flight) mass spectrometry has shown promise as a screening tool for detecting discriminatory protein patterns. The major computational obstacle in analyzing MALDI-TOF data is a large number of mass/charge peaks (a.k.a. features, data points). To tackle this problem, we have developed a multi-step strategy for data preprocessing and afterwards feature selection. The preprocessing is composed of binning, baseline correction, and normalization. For the preprocessed data, we propose a new feature subset selection method. Our scheme is applied to the analysis of ovarian cancer dataset to predict early relapse in ovarian cancer. To validate the performance of the proposed algorithm, experiments are performed in comparison with other feature selection and classification methods. We show that our proposed approach outperforms other algorithms
Jung Hun Oh, Animesh Nandi, Prem Gurnani, Lynne Knowles, John Schorge, Kevin P. Rosenblatt, Jean Gao
CIBCB7
2006 Detecting Occlusion for Hidden Markov Modeled Shapes
abstract
In this paper, we present a novel occlusion detection scheme for hidden Markov modeled shapes. First, hidden Markov model (HMM) is built using multiple examples of the shape. A reference path for the shape is built from the HMM, which is nothing but optimal path followed by the most likely example. The reference path stores temporal information about the entire shape, while the HMM only retains relationship between temporal information. For the shape of interest, its optimal path through HMM is calculated and warped to match the reference path using dynamic time warping (DTW). Occluded part of the shape is detected by identifying imbalance among various components of the matching cost. Detection results obtained for two shape data sets are presented for varying degrees of occlusion.
Ninad Thakoor, Sungyong Jung, Jean Gao
ICIP3
2006 A multi-Kalman filtering approach for video tracking of human-delineated objects in cluttered environments
Jean Gao, Akio Kosaka, Avinash C. Kak
Comput. Vis. Image Underst.1
2005 Hidden Markov Model Based 2D Shape Classification
Ninad Thakoor, Jean Gao
ACIVS2
2005 Peptide Identification by Tandem Mass Spectra: An Efficient Parallel Searching
abstract
De novo peptide sequencing that determines the amino acid sequence of a peptide via tandem mass spectrometry (MS/MS) has been increasingly used nowadays in proteomics for protein identification. Current de novo methods generally employ a graph theory which usually produces a large number of candidate sequences and causes heavy computational cost while trying to determine a sequence with less ambiguity. We present a novel de novo sequencing algorithm which greatly reduces the number of candidate sequences. By utilizing certain properties of b- and y-ion series in MS/MS spectrum, we propose a reliable two-way parallel searching algorithm to filter out the peptide candidates which are further pruned by an intensity evidence based screening criterion. And we find an adjusted value required to determine the position of end node of b- and y-ion series for the charged +2 precursor in our graph. Results of our algorithm are compared with those of PEAKS, a well-known de novo sequencing software. Experimental results demonstrate the six sequences are identical with the correct sequences. And for the further pruning, rankings of our result remain unchanged even though the screening criterion changes. Therefore we can reduce the number of candidate sequences by adopting a proper screening criterion.
Jung Hun Oh, Jean Gao
BIBE2
2005 A New Semi-Supervised Subspace Clustering Algorithm on Fitting Mixture Models
Young Bun Kim, Jean Gao
CIBCB2
2005 Multicategory Classification using Extended SVM-RFE and Markov Blanket on SELDITOF Mass Spectrometry Data
Jung Hun Oh, Jean Gao, Animesh Nandi, Prem Gurnani, Lynne Knowles, John Schorge, Kevin P. Rosenblatt
CIBCB2
2005 Shape Classifer Based on Generalized Probabilistic Descent Method with Hidden Markov Descriptor
abstract
The goal of this paper is to present a weighted likelihood discriminant for minimum error shape classification. Different from traditional maximum likelihood (ML) methods, in which classification is based on probabilities from independent individual class models as is the case for general hidden Markov model (HMM) methods, proposed method utilizes information from all classes to minimize classification error. The proposed approach uses a HMM for shape curvature as its 2D shape descriptor. In this contribution, we introduce a weighted likelihood discriminant function and present a minimum error classification strategy based on generalized probabilistic descent (GPD) method. We believe our sound theory based implementation reduces classification error by combining HMM with GPD theory. We show comparative results obtained with our approach and classic ML classification along with Fourier descriptor and Zernike moments based classification for fighter planes and vehicle shapes.
Ninad Thakoor, Jean Gao
ICCV2
2005 Automatic video object shape extraction and its classification with camera in motion
abstract
In this paper, we present an automatic moving object extraction and classification system. For automatic extraction of object taken by a moving camera, a novel technique is proposed, in which optical flow handles the background modeling and camera motion estimation, and frame difference information yields the exact object shape. We also use forward region boundary based change detection approach for frame difference. This approach assures change detection for uniform intensity regions. For classification, weighted likelihood discriminant based shape classifier is designed. Unlike maximum likelihood (ML) methods, our proposed method utilizes information from all classes to design the classifier. In the description phase of the classifier, curvature features are extracted from the shape and are utilized to build a hidden Markov model (HMM). The HMM provides a robust ML description of the shape. In the discrimination phase, a weighted likelihood discriminant function is introduced, which weights the likelihoods of curvature at individual points of the shape to minimize the classification error. The weights are estimated by generalized probabilistic descent (GPD) method. To demonstrate the performance of the proposed method, we present results achieved for car shapes extraction and classification.
Ninad Thakoor, Jean Gao
ICIP (3)2
2005 A particle filter framework using optimal importance function for protein molecules tracking
abstract
Tagging and tracking protein molecules are a key to a better understanding of proteomics in diverse aspects. In this paper, a common framework of particle filter using optimal importance function is proposed for confocal protein molecules tracking. To deal with the challenges stemming from small size, deformable shape, noisy environment, and multi-modality motion, a stochastic process based particle filter is used. Partial Gaussian state space (PGSS) model is developed as the importance function to incorporate the latest measurement in the state estimation. Experimental results have demonstrated the performance of the proposed algorithm for both Brownian and translational motion.
Quan Wen 0001, Jean Gao, Akio Kosaka, Hidekazu Iwaki, Katherine Luby-Phelps, Dorothy Mundy
ICIP (1)2
2005 Hidden Markov Model Based Weighted Likelihood Discriminant for Minimum Error Shape Classification
abstract
The goal of this communication is to present a weighted likelihood discriminant for minimum error shape classification. Different from traditional Maximum Likelihood (ML) methods in which classification is carried out based on probabilities from independent individual class models as is the case for general hidden Markov model (HMM) methods, our proposed method utilizes information from all classes to minimize classification error. Proposed approach uses a Hidden Markov Model as a curvature feature based 2D shape descriptor. In this contribution we present a Generalized Probabilistic Descent (GPD) method to weight the curvature likelihoods to achieve a discriminant function with minimum classification error. In contrast with other approaches, a weighted likelihood discriminant function is introduced. We believe that our sound theory based implementation reduces classification error by combining hidden Markov model with generalized probabilistic descent theory. We show comparative results obtained with our approach and classic maximum-likelihood calculation for fighter planes in terms of classification accuracies.
Ninad Thakoor, Sungyong Jung, Jean Gao
ICME3
2005 Automatic Extraction and Localization of Multiple Moving Objects with Stereo Camera in Motion
abstract
In this paper, we propose a technique for automatic extraction and localization of multiple moving objects captured by a moving stereo camera. Depth information obtained from the stereo matching is applied to initialize a background mask which is refined further by clustering optical flow. Simultaneous estimate for camera motion is also obtained and video frames are compensated accordingly. Refined background mask is used to model displaced frame difference (DFD) of the background. The objects are detected as outliers to this model. We also present a region boundary based change detection approach for frame difference. This approach assures change detection for uniform intensity regions. Finally we present experimental results for indoor and outdoor stereo video sequences captured with moving camera
Ninad Thakoor, Jean Gao
SMC2
2005 A multi-Kalman filtering approach for video tracking of human-delineated objects in cluttered environments
Jean Gao, Akio Kosaka, Avinash C. Kak
Comput. Vis. Image Underst.1
2004 A motion field reconstruction scheme for smooth boundary video object segmentation
abstract
Motion segmentation is a classic and on-going research topic which is an important pre-stage for many video processes. The reliability of the motion field calculation directly determines either success or failure of such segmentation. Due to the aperture property of optical flow calculation, motion estimation at a moving object boundary remains a challenging ill-posed problem. We propose a reliable optical-flow estimation method from the view of missing data reconstruction. Furthermore, to overcome the occlusion problem at motion boundaries, we apply a motion segmentation scheme which integrates with spatial segmentation. Experimental results of the proposed and other methods are presented.
Jean Gao, Ninad Thakoor, Sungyong Jung
ICIP1
2003 Self-Occlusion Immune Video Tracking of Objects in Cluttered Environments
abstract
We propose a new approach that uses a motion-estimation based framework for video tracking of objects in the presence of self-occlusion in cluttered environments. What makes our work different from others is that, instead of carrying out the motion estimation between two adjacent frames, we tackle the self-occlusion problem from the view of multiple frames. The heart of our approach lies in extracting features appearing in different time frames, genesis frames, and setting up a motion estimation scheme through multiple applications of Kalman filtering based on the different genesis indices. To make the tracked object look visually familiar to the human observer, the system also makes its best attempt at extracting the boundary contour of the object - a difficult problem in its own right, since self-occlusion created by any rotational motion of the tracked object would cause large sections of the boundary contour in the previous frame to disappear in the current frame. Our approach has been tested on a wide variety of video sequences, some of which are presented.
Jean Gao
AVSS1
2003 Object shape delineation during tracking process
abstract
Object shape delineation during the tracking process plays important roles in correctly interpreting tracked results, providing visually meaningful outcomes, and furthermore assisting better motion estimation. For the majority of object tracking scenario, the emphasis has been put on achieving robust motion estimation in different situations; and object shape delineation, though critical, has not been paid enough attention due to its ill-posed nature. Approaches have been proposed by assuming the similarity of object pixels in the vicinities of the boundaries between the current frame and the previous one. Such an assumption is usually broken down when occlusion occurs; instead, our implementation is based on a stronger assumption: the local properties of object silhouette should be similar to those of the nearby object pixels. In this paper, we are going to address how to depict object boundary by a novel double-region growing and statistical pattern classification approach. Different from using a single point as a seed as which is a typical way for region growing, our seeds are segmented contours; also instead of growing outward in a single direction from the seed, we propose a two-directional region growing approach. Finally the best object boundary candidates are arbitrated from the dual-region growing results by a statistical classification approach.
Jean Gao
ICIP (3)1
2002 A multi-frame based motion estimation for semantic object tracking in the presence of occlusion
abstract
Motion estimation in the presence of occlusion is a wide open research field. As such, a critical component of the research is formulating a computational framework. A distinguishing aspect of our motion estimation scheme for solving the occlusion problem is that it cleanly handles the situation where the tracked features appear in different frames and, as features become unobservable, new features may need to be incorporated. As opposed to employing a 2D parametric motion model which restricts the object to a planar surface, we use a 3D motion model to capture the object motion and shape vectors. The precise 3D contour of the tracked object is fine-tuned within a predicted potential field.
Jean Gao, Avinash C. Kak
ICIP (3)1
2001 Multi-Visit of Kalman Filtering for Semantic Object Tracking
abstract
In this paper, we propose a new approach to simultaneously estimate the motion parameters and shape parameters of a rigid object over the image sequence while the object is in motion. Our approach applies the Kalman filter multiple times to sequentially reduce the error of feature point correspondences and to generate an accurate motion parameter by propagating the object motion uncertainty over the image frame as well as optimally establishing the feature correspondences on the basis of template-based optical flows. Some preliminarily experimental results are shown to verify our proposed approach.
Jean Gao, Akio Kosaka
ICME1
1999 Interactive Color Image Segmentation Editor Driven by Active Contour Model
abstract
A new general-purpose color image segmentation editor (CISE) for the purpose of extracting a semantic object is designed, implemented and tested on a number of various natural scene images. Our editor integrates a deformable model and image statistics including intensity, color, gradient and texture. The editor starts with a coarse region segmentation which applies the Canny's operator followed by a low-complexity edge linking algorithm. This segmentation basically builds regions of smooth intensity by closing all dangling edges. Next, a topologically-based region labeling method makes full use of relationship among pixels and produces useful labeled image. Finally, to refine the extracted object of interest (O/sup 2/I), a deformable model based on energy minimization is applied by incorporating both the gradient and region criteria to the external constraint force. These processes are demonstrated through examples on natural scene color images. Experimental results suggest the efficiency and accuracy of the algorithm in its segmentation operations.
Jean Gao, Akio Kosaka, Avinash C. Kak
ICIP (3)1
1998 A Deformable Model for Human Organ Extraction
Jean Gao, Akio Kosaka, Avinash C. Kak
ICIP (3)1