VLDB 2026 Research / reviewers in the wild / expert
Jung Hun Oh
dblp:30/1854
· DBLP profile ↗
27ranked-venue papers
17as first author
5since 2021 · last 2025
0000-0001-8791-2755ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 14 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ORCO: Ollivier-Ricci Curvature-Omics - an unsupervised method for analyzing robustness in biological systemsabstractMOTIVATION: Although recent advanced sequencing technologies have improved the resolution of genomic and proteomic data to better characterize molecular phenotypes, efficient computational tools to analyze and interpret large-scale omic data are still needed. RESULTS: To address this, we have developed a network-based bioinformatic tool called Ollivier-Ricci curvature for omics (ORCO). ORCO incorporates omics data and a network describing biological relationships between the genes or proteins and computes Ollivier-Ricci curvature (ORC) values for individual interactions. ORC is an edge-based measure that assesses network robustness. It captures functional cooperation in gene signaling using a consistent information-passing measure, which can help investigators identify therapeutic targets and key regulatory modules in biological systems. ORC has identified novel insights in multiple cancer types using genomic data and in neurodevelopmental disorders using brain imaging data. This tool is applicable to any data that can be represented as a network. AVAILABILITY AND IMPLEMENTATION: ORCO is an open-source Python package and is publicly available on GitHub at https://github.com/aksimhal/ORC-Omics. Anish K. Simhal, Corey Weistuch, Kevin A. Murgas, Daniel Grange, Jiening Zhu, Jung Hun Oh, Rena Elkin, Joseph O. Deasy |
Bioinform. | 6 |
| 2024 | Wasserstein HOG: Local Directionality Extraction via Optimal TransportabstractDirectionally sensitive radiomic features including the histogram of oriented gradient (HOG) have been shown to provide objective and quantitative measures for predicting disease outcomes in multiple cancers. However, radiomic features are sensitive to imaging variabilities including acquisition differences, imaging artifacts and noise, making them impractical for using in the clinic to inform patient care. We treat the problem of extracting robust local directionality features by mapping via optimal transport a given local image patch to an iso-intense patch of its mean. We decompose the transport map into sub-work costs each transporting in different directions. To test our approach, we evaluated the ability of the proposed approach to quantify tumor heterogeneity from magnetic resonance imaging (MRI) scans of brain glioblastoma multiforme, computed tomography (CT) scans of head and neck squamous cell carcinoma as well as longitudinal CT scans in lung cancer patients treated with immunotherapy. By considering the entropy difference of the extracted local directionality within tumor regions, we found that patients with higher entropy in their images, had significantly worse overall survival for all three datasets, which indicates that tumors that have images exhibiting flows in many directions may be more malignant. This may seem to reflect high tumor histologic grade or disorganization. Furthermore, by comparing the changes in entropy longitudinally using two imaging time points, we found patients with reduction in entropy from baseline CT are associated with longer overall survival (hazard ratio = 1.95, 95% confidence interval of 1.4-2.8, p = 1.65e-5). The proposed method provides a robust, training free approach to quantify the local directionality contained in images. Jiening Zhu, Harini Veeraraghavan, Jue Jiang, Jung Hun Oh, Larry Norton, Joseph O. Deasy, Allen R. Tannenbaum |
IEEE Trans. Medical Imaging | 4 |
| 2022 | aWCluster: A Novel Integrative Network-Based Clustering of Multiomics for Subtype Analysis of Cancer DataabstractThe remarkable growth of multi-platform genomic profiles has led to the challenge of multiomics data integration. In this study, we present a novel network-based multiomics clustering founded on the Wasserstein distance from optimal mass transport. This distance has many important geometric properties making it a suitable choice for application in machine learning and clustering. Our proposed method of aggregating multiomics and Wasserstein distance clustering (aWCluster) is applied to breast carcinoma as well as bladder carcinoma, colorectal adenocarcinoma, renal carcinoma, lung non-small cell adenocarcinoma, and endometrial carcinoma from The Cancer Genome Atlas project. Subtypes were characterized by the concordant effect of mRNA expression, DNA copy number alteration, and DNA methylation of genes and their neighbors in the interaction network. aWCluster successfully clusters all cancer types into classes with significantly different survival rates. Also, a gene ontology enrichment analysis of significant genes in the low survival subgroup of breast cancer leads to the well-known phenomenon of tumor hypoxia and the transcription factor ETS1 whose expression is induced by hypoxia. We believe aWCluster has the potential to discover novel subtypes and biomarkers by accentuating the genes that have concordant multiomics measurements in their interaction network, which are challenging to find without the network inference or with single omics analysis. Maryam Pouryahya, Jung Hun Oh, Pedram Javanmard, James C. Mathews, Zehor Belkhatir, Joseph O. Deasy, Allen R. Tannenbaum |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | Highly accurate diagnosis of papillary thyroid carcinomas based on personalized pathways coupled with machine learningabstractThyroid nodules are neoplasms commonly found among adults, with papillary thyroid carcinoma (PTC) being the most prevalent malignancy. However, current diagnostic methods often subject patients to unnecessary surgical burden. In this study, we developed and validated an automated, highly accurate multi-study-derived diagnostic model for PTCs using personalized biological pathways coupled with a sophisticated machine learning algorithm. Surprisingly, the algorithm achieved near-perfect performance in discriminating PTCs from non-tumoral thyroid samples with an overall cross-study-validated area under the receiver operating characteristic curve (AUROC) of 0.999 (95% confidence interval [CI]: 0.995-1) and a Brier score of 0.013 on three independent development cohorts. In addition, the algorithm showed excellent generalizability and transferability on two large-scale external blind PTC cohorts consisting of The Cancer Genome Atlas (TCGA), which is the largest genomic PTC cohort studied to date, and the post-Chernobyl cohort, which includes PTCs reported after exposure to radiation from the Chernobyl accident. When applied to the TCGA cohort, the model yielded an AUROC of 0.969 (95% CI: 0.950-0.987) and a Brier score of 0.109. On the post-Chernobyl cohort, it yielded an AUROC of 0.962 (95% CI: 0.918-1) and a Brier score of 0.073. This algorithm also is robust against other various types of clinical scenarios, discriminating malignant from benign lesions as well as clinically aggressive thyroid cancer with poor prognosis from indolent ones. Furthermore, we discovered novel pathway alterations and prognostic signatures for PTC, which can provide directions for follow-up studies. Kyoung Sik Park, Seong Hoon Kim, Jung Hun Oh, Sung Young Kim |
Briefings Bioinform. | 3 |
| 2021 | PathCNN: interpretable convolutional neural networks for survival prediction and pathway analysis applied to glioblastomaabstractMOTIVATION: Convolutional neural networks (CNNs) have achieved great success in the areas of image processing and computer vision, handling grid-structured inputs and efficiently capturing local dependencies through multiple levels of abstraction. However, a lack of interpretability remains a key barrier to the adoption of deep neural networks, particularly in predictive modeling of disease outcomes. Moreover, because biological array data are generally represented in a non-grid structured format, CNNs cannot be applied directly. RESULTS: To address these issues, we propose a novel method, called PathCNN, that constructs an interpretable CNN model on integrated multi-omics data using a newly defined pathway image. PathCNN showed promising predictive performance in differentiating between long-term survival (LTS) and non-LTS when applied to glioblastoma multiforme (GBM). The adoption of a visualization tool coupled with statistical analysis enabled the identification of plausible pathways associated with survival in GBM. In summary, PathCNN demonstrates that CNNs can be effectively applied to multi-omics data in an interpretable manner, resulting in promising predictive power while identifying key biological correlates of disease. AVAILABILITY AND IMPLEMENTATION: The source code is freely available at: https://github.com/mskspi/PathCNN. Jung Hun Oh, Euiseong Ko, Mingon Kang, Allen R. Tannenbaum, Joseph O. Deasy |
Bioinform. | 1 |
| 2019 | Gene- and Pathway-Based Deep Neural Network for Multi-omics Data Integration to Predict Cancer Survival Outcomes
Mohammad Masum, Jung Hun Oh, Mingon Kang |
ISBRA | 3 |
| 2018 | Cox-PASNet: Pathway-based Sparse Deep Neural Network for Survival Analysis
Youngsoon Kim, Tejaswini Mallavarapu, Jung Hun Oh, Mingon Kang |
BIBM | 4 |
| 2018 | PASCL: Pathway-based Sparse Deep Clustering for Identifying Unknown Cancer Subtypes
Tejaswini Mallavarapu, Youngsoon Kim, Jung Hun Oh, Mingon Kang |
BIBM | 4 |
| 2017 | R-PathCluster: Identifying cancer subtype of glioblastoma multiforme using pathway-based restricted boltzmann machineabstractGlioblastoma multiforme (GBM) is the most fatal malignant type of brain tumor with a very poor prognosis with a median survival of around one year. Numerous studies have reported tumor subtypes that consider different characteristics on individual patients, which may play important roles in determining the survival rates in GBM. In this study, we present a pathway-based clustering method using Restricted Boltzmann Machine (RBM), called R-PathCluster, for identifying unknown subtypes with pathway markers of gene expressions. In order to assess the performance of R-PathCluster, we conducted experiments with several clustering methods such as k-means, hierarchical clustering, and RBM models with different input data. R-PathCluster showed the best performance in clustering long-term and short-term survivals, although its clustering score was not the highest among them in experiments. R-PathCluster provides a solution to interpret the model in biological sense, since it takes pathway markers that represent biological process of pathways. We discussed that our findings from R-PathCluster are supported by many biological literatures. Tejaswini Mallavarapu, Youngsoon Kim, Jung Hun Oh, Mingon Kang |
BIBM | 3 |
| 2016 | Transcriptional responses to ultraviolet and ionizing radiation: An approach based on graph curvatureabstractMore than half of all cancer patients receive radiotherapy in their treatment process. However, our understanding of abnormal transcriptional responses to radiation remains poor. In this study, we employ an extended definition of Ollivier-Ricci curvature based on LI-Wasserstein distance to investigate genes and biological processes associated with ionizing radiation (IR) and ultraviolet radiation (UV) exposure using a microarray dataset. Gene expression levels were modeled on a gene interaction topology downloaded from the Human Protein Reference Database (HPRD). This was performed for IR, UV, and mock datasets, separately. The difference curvature value between IR and mock graphs (also between UV and mock) for each gene was used as a metric to estimate the extent to which the gene responds to radiation. We found that in comparison of the top 200 genes identified from IR and UV graphs, about 20~30% genes were overlapping. Through gene ontology enrichment analysis, we found that the metabolic-related biological process was highly associated with both IR and UV radiation exposure. Yongxin Chen 0002, Jung Hun Oh, Romeil Sandhu, Sangkyu Lee, Joseph O. Deasy, Allen R. Tannenbaum |
BIBM | 2 |
| 2016 | Integrative Gene Regulatory Network inference using multi-omics dataabstractBiological network inference is of importance to understand underlying biological mechanisms. Gene regulatory networks describe molecular interactions of complex biological processes. Graph models are mainly used for gene regulatory networks, where nodes and edges represent genes and their regulations respectively. In the most research, the molecular interactions (edges) of gene regulatory networks are inferred from a single type of genomic data, e.g., gene expression data. However, gene expression is a product of sequential interactions of DNA sequence variations, single nucleotide polymorphism, copy number variation, histone modifications, transcription factor, DNA methylation, and many other factors. There are high-throughput genomic data that measure the various biological processes. We call the multiple types of genomics data as ‘multi-omics data’. In this paper, we propose an Integrative Gene Regulatory Network inference method (iGRN) that can incorporate multi-omics data and their interactions in the graph model of gene regulatory network. Copy number variation and DNA methylation were considered for multi-omics data in this paper. The proposed method, iGRN, was applied to the human brain data of psychiatric disorder. Through the experiments, iGRN showed its better performance on model representation and interpretation than other integrative methods in gene regulatory network inference. Neda Zarayeneh, Jung Hun Oh, Donghyun Kim 0001, Chunyu Liu 0001, Jean Gao, Sang C. Suh, Mingon Kang |
BIBM | 2 |
| 2016 | A literature mining-based approach for identification of cellular pathways associated with chemoresistance in cancerabstractChemoresistance is a major obstacle to the successful treatment of many human cancer types. Increasing evidence has revealed that chemoresistance involves many genes and multiple complex biological mechanisms including cancer stem cells, drug efflux mechanism, autophagy and epithelial-mesenchymal transition. Many studies have been conducted to investigate the possible molecular mechanisms of chemoresistance. However, understanding of the biological mechanisms in chemoresistance still remains limited. We surveyed the literature on chemoresistance-related genes and pathways of multiple cancer types. We then used a curated pathway database to investigate significant chemoresistance-related biological pathways. In addition, to investigate the importance of chemoresistance-related markers in protein-protein interaction networks identified using the curated database, we used a gene-ranking algorithm designed based on a graph-based scoring function in our previous study. Our comprehensive survey and analysis provide a systems biology-based overview of the underlying mechanisms of chemoresistance. Jung Hun Oh, Joseph O. Deasy |
Briefings Bioinform. | 1 |
| 2014 | Inference of radio-responsive gene regulatory networks using the graphical lasso algorithmabstractBACKGROUND: Inference of gene regulatory networks (GRNs) from gene microarray expression data is of great interest and remains a challenging task in systems biology. Despite many efforts to develop efficient computational methods, the successful modeling of GRNs thus far has been quite limited. To tackle this problem, we propose a novel framework to reconstruct radio-responsive GRNs based on the graphical lasso algorithm. In our attempt to study radiosensitivity, we reviewed the literature and analyzed two publicly available gene microarray datasets. The graphical lasso algorithm was applied to expression measurements for genes commonly found to be significant in these different analyses. RESULTS: Assuming that a protein-protein interaction network obtained from a reliable pathway database is a gold-standard network, a comparison between the networks estimated by the graphical lasso algorithm and the gold-standard network was performed. Statistically significant p-values were achieved when comparing the gold-standard network with networks estimated from one microarray dataset and when comparing the networks estimated from two microarray datasets. CONCLUSION: Our results show the potential to identify new interactions between genes that are not present in a curated database and GRNs using microarray datasets via the graphical lasso algorithm. Jung Hun Oh, Joseph O. Deasy |
BMC Bioinform. | 1 |
| 2011 | Fast Kernel Discriminant Analysis for Classification of Liver Cancer Mass SpectraabstractThe classification of serum samples based on mass spectrometry (MS) has been increasingly used for monitoring disease progression and for diagnosing early disease. However, the classification task in mass spectrometry data is extremely challenging due to the very huge size of peaks (features) on mass spectra. Linear discriminant analysis (LDA) has been widely used for dimension reduction and feature extraction in many applications. However, the conversional LDA suffers from the singularity problem when dealing with high-dimensional features. Another critical limitation is its linearity property which results in failing in classification problems over nonlinearly clustered data sets. To overcome such problems, we develop a new fast kernel discriminant analysis (FKDA) that is pretty fast in the calculation of optimal discriminant vectors. FKDA is applied to the classification of liver cancer mass spectrometry data that consist of three categories: hepatocellular carcinoma, cirrhosis, and healthy that was originally analyzed by Ressom et al. We demonstrate the superiority and effectiveness of FKDA when compared to other classification techniques. Jung Hun Oh, Jean Gao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2010 | Predicting Local Failure in Lung Cancer Using Bayesian NetworksabstractDespite various efforts to develop new predictive models for early detection of tumor local failure in locally advanced non-small cell lung cancer (NSCLC), many patients still suffer from a high local failure rate after radiotherapy. Based on recent studies of biomarker proteins' role in predicting tumor response following radiotherapy, we hypothesize that incorporation of physical and biological factors with a suitable framework could improve the overall prediction. To this end, we propose a graphical Bayesian network framework for predicting local failure in lung cancer. The proposed approach was tested using a dataset of locally advanced NSCLC patients treated with radiotherapy. This dataset was collected prospectively, which consisted of physical variables and blood-based biomarkers. Our experimental results demonstrate that the proposed method can be used as an efficient method to develop predictive models of local failure in these patients and to interpret relationships among the different variables. The combined model of physical and biological factors outperformed individual physical and biological models, achieving an accuracy (acc) of 87.78%, Matthew's correlation coefficient (r) of 0.74, and Spearman's rank correlation coefficient (rs) of 0.75 on leave-one-out cross-validation analysis. Jung Hun Oh, Jeffrey Craft, Rawan Al-Lozi, Manushka Vaidya, Yifan Meng, Joseph O. Deasy, Jeffrey D. Bradley, Issam El-Naqa |
ICMLA | 1 |
| 2009 | Application of Machine Learning Techniques for Prediction of Radiation Pneumonitis in Lung Cancer PatientsabstractLung cancer patients who receive radiotherapy as part of their treatment are at risk radiation-induced lung injury known as radiation pneumonitis (RP). RP is a potentially fatal side effect to treatment. Hence, new methods are needed to guide physicians to prescribe targeted therapy dosage to patients at high risk of RP. Several predictive models based on traditional statistical methods and machine learning techniques have been reported, however, no guidance to variation in performance has not been provided to date. Therefore, in this study, we compare several widely used classification algorithms in the machine learning field are used to distinguish between different risk groups of RP. The performance of these classification algorithms is evaluated in conjunction with several feature selection strategy and the impact of the feature selection on performance is further evaluated. Jung Hun Oh, Rawan Al-Lozi, Issam El-Naqa |
ICMLA | 1 |
| 2009 | Bayesian Network Learning for Detecting Reliable Interactions of Dose-Volume Related Parameters in Radiation PneumonitisabstractDue to the high fatality rate of patients with radiation pneumonitis (RP), a complication of the radiation therapy (radiotherapy), great attention has been paid to the treatment plan of individual RP patients. Therefore, not only technological advances in the development of treatment planning systems but also new prognostic models are urgently required to lessen the complication and to predict the state of patients more accurately. The Bayesian network is a useful tool for finding interactions among features and for developing prognostic models that enable physician to predict the outcome of radiotherapy. In this paper, we show the interactions among dosimetric features through Bayesian network structures and the performance of Bayesian classifiers with different search algorithms on a non-small cell lung cancer (NSCLC) dataset. Jung Hun Oh, Issam El-Naqa |
ICMLA | 1 |
| 2009 | A kernel-based approach for detecting outliers of high-dimensional biological dataabstractBACKGROUND: In many cases biomedical data sets contain outliers that make it difficult to achieve reliable knowledge discovery. Data analysis without removing outliers could lead to wrong results and provide misleading information. RESULTS: We propose a new outlier detection method based on Kullback-Leibler (KL) divergence. The original concept of KL divergence was designed as a measure of distance between two distributions. Stemming from that, we extend it to biological sample outlier detection by forming sample sets composed of nearest neighbors. KL divergence is defined between two sample sets with and without the test sample. To handle the non-linearity of sample distribution, original data is mapped into a higher feature space. We address the singularity problem due to small sample size during KL divergence calculation. Kernel functions are applied to avoid direct use of mapping functions. The performance of the proposed method is demonstrated on a synthetic data set, two public microarray data sets, and a mass spectrometry data set for liver cancer study. Comparative studies with Mahalanobis distance based method and one-class support vector machine (SVM) are reported showing that the proposed method performs better in finding outliers. CONCLUSION: Our idea was derived from Markov blanket algorithm that is a feature selection method based on KL divergence. That is, while Markov blanket algorithm removes redundant and irrelevant features, our proposed method detects outliers. Compared to other algorithms, our proposed method shows better or comparable performance for small sample and high-dimensional biological data. This indicates that the proposed method can be used to detect outliers in biological data sets. Jung Hun Oh, Jean Gao |
BMC Bioinform. | 1 |
| 2009 | An Extended Markov Blanket Approach to Proteomic Biomarker Detection From High-Resolution Mass Spectrometry DataabstractHigh-resolution matrix-assisted laser desorption/ionization time-of-flight mass spectrometry has recently shown promise as a screening tool for detecting discriminatory peptide/protein patterns. The major computational obstacle in finding such patterns is the large number of mass/charge peaks (features, biomarkers, data points) in a spectrum. To tackle this problem, we have developed methods for data preprocessing and biomarker selection. The preprocessing consists of binning, baseline correction, and normalization. An algorithm, extended Markov blanket, is developed for biomarker detection, which combines redundant feature removal and discriminant feature selection. The biomarker selection couples with support vector machine to achieve sample prediction from high-resolution proteomic profiles. Our algorithm is applied to recurrent ovarian cancer study that contains platinum-sensitive and platinum-resistant samples after treatment. Experiments show that the proposed method performs better than other feature selection algorithms. In particular, our algorithm yields good performance in terms of both sensitivity and specificity as compared to other methods. Jung Hun Oh, Prem Gurnani, John Schorge, Kevin P. Rosenblatt, Jean Gao |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2008 | Biological Data Outlier Detection Based on Kullback-Leibler DivergenceabstractOutlier detection is imperative in biomedical data analysis to achieve reliable knowledge discovery. In this paper, a new outlier detection method based on Kullback-Leibler (KL) divergence is presented. The original concept of KL divergence was designed as a measure of distance between two distributions. Stemming from that, we extend it to biological sample outlier detection by forming sample sets composed of nearest neighbors. To handle the non-linearity during the KL divergence calculation and to tackle with the singularity problem due to small sample size, we map the original data into a higher feature space and apply kernel functions without resorting to a mapping function. A sample possessing the largest KL divergence is detected as an outlier. The proposed method is tested with one synthetic data, two public gene expression data sets, and our own mass spectrometry data generated for prostate cancer study. Jung Hun Oh, Jean Gao, Kevin P. Rosenblatt |
BIBM | 1 |
| 2008 | Biomarker selection and sample prediction for multi-category disease on MALDI-TOF dataabstractMOTIVATION: Diseases normally progress through several stages. Therefore, biomarkers corresponding to each stage may exist. To deal with such a multi-category problem, including sample stage prediction and biomarker selection, we propose methods for classification and feature selection. The proposed classification method is based on two schemes: error-correcting output coding (ECOC) and pairwise coupling (PWC). The final decision for a test sample prediction is an integration of these two schemes. The biomarker pattern for distinguishing each disease category from another one is achieved by the development of an extended Markov blanket (EMB) feature selection method. RESULTS: In this study, a liver cancer matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) mass spectrometry (MS) dataset was used, which comprises hepatocellular carcinoma (HCC), cirrhosis, and healthy spectra. Peak patterns were discovered for distinguishing pairwise categories among the three classes. Importance and reliability of individual peaks were presented by the measurements of certain weight values and frequencies. The classification capability of the proposed approach was compared with classical ECOC, random forest, Naive Bayes, and J48 methods. AVAILABILITY: Supplementary materials are available at http://visionlab.uta.edu/biomarker/bioinfo.htm. Jung Hun Oh, Young Bun Kim, Prem Gurnani, Kevin P. Rosenblatt, Jean Gao |
Bioinform. | 1 |
| 2007 | Biomarker Selection for Predicting Alzheimer Disease Using High-Resolution MALDI-TOF DataabstractHigh-resolution MALDI-TOF (matrix-assisted laser desorption/ionization time-of-flight) mass spectrometry has shown promise as a screening tool for detecting discriminatory peptide/protein patterns. The major computational obstacle in analyzing MALDI-TOF data is the large number of mass/charge peaks (a.k.a. features, data points). With such a huge number of data points for a single sample, efficient feature selection is critical for unequivocal protein pattern discovery. In this paper, we propose a feature selection method and a new biclassification algorithm based on error-correcting output coding (ECOC) in multiclass problems. Our scheme is applied to the analysis of alzheimer's disease (AD) data. To validate the performance of the proposed algorithm, experiments are performed in comparison with other methods. We show that our proposed framework outperforms not only the standard ECOC framework but also other algorithms. Jung Hun Oh, Young Bun Kim, Prem Gurnani, Kevin P. Rosenblatt, Jean Gao |
BIBE | 1 |
| 2007 | A Novel Classification Method for Analysis of Multi-stage Diseases via Mass Spectrometric DataabstractMulti-category classification is one of the challenging issues in medical data analysis. We propose a new bi- classification algorithm for the multi-class classification, which is comprised of two schemes: error-correcting output coding (ECOC) and pairwise coupling (PWC). After fea- ture reduction in both schemes, each corresponding classi- fication strategy is performed. For a test sample, two class labels that are predicted in both schemes are compared. If two class labels are the same, we assign the test sample to an identical label; otherwise, only for samples belonging to different classes predicted from two schemes, a retraining method is employed. Our scheme is applied to the analysis of a MALDI-TOF data set which consists of hepatocellular carcinoma (HCC) patients, cirrhosis patients and healthy individuals. To validate the performance of our proposed algorithm, experiments were performed in comparison with other classification methods. Jung Hun Oh, Young Bun Kim, Jean Gao |
BIBM | 1 |
| 2006 | Prediction of labor for pregnant women using high-resolution mass spectrometry dataabstractHigh-resolution MALDI-TOF (matrix-assisted laser desorption/ionization time-of-flight) mass spectrometry has shown promise as a screening tool for detecting discriminatory protein patterns. The major computational obstacle in analyzing MALDI-TOF data is a large number of mass/charge peaks (a.k.a. features, data points). With the number of data points easily going beyond one million for a single sample, efficient feature selection is critical for unequivocal protein pattern discovery. To tackle this problem, we have developed a multi-step strategy for data preprocessing and afterwards feature selection. The preprocessing is composed of binning, baseline correction, and normalization. For the preprocessed data, we propose a new feature subset selection method that is a hybrid filter/wrapper approach. Based on the two feature subsets for each feature, high and low correlated subsets, a feature is assigned a weight which indicates the extent of feature importance. Our scheme is applied to the analysis of labor dataset to predict delivery time of pregnant women. To validate the performance of the proposed algorithm, experiments are performed in comparison with other feature selection and classification methods. We show that our proposed approach outperforms other algorithms Jung Hun Oh, Animesh Nandi, Prem Gurnani, Peter Bryant-Greenwood, Kevin P. Rosenblatt, Jean Gao |
BIBE | 1 |
| 2006 | Classification of Relapse Ovarian Cancer on MALDI-TOF Mass Spectrometry DataabstractOvarian cancer recurs at the rate of 75% within a few months or several years later after therapy. Early recurrence, though responding better to treatment, is difficult to detect. Recently, high-resolution MALDI-TOF (matrix-assisted laser desorption/ionization time-of-flight) mass spectrometry has shown promise as a screening tool for detecting discriminatory protein patterns. The major computational obstacle in analyzing MALDI-TOF data is a large number of mass/charge peaks (a.k.a. features, data points). To tackle this problem, we have developed a multi-step strategy for data preprocessing and afterwards feature selection. The preprocessing is composed of binning, baseline correction, and normalization. For the preprocessed data, we propose a new feature subset selection method. Our scheme is applied to the analysis of ovarian cancer dataset to predict early relapse in ovarian cancer. To validate the performance of the proposed algorithm, experiments are performed in comparison with other feature selection and classification methods. We show that our proposed approach outperforms other algorithms Jung Hun Oh, Animesh Nandi, Prem Gurnani, Lynne Knowles, John Schorge, Kevin P. Rosenblatt, Jean Gao |
CIBCB | 1 |
| 2005 | Peptide Identification by Tandem Mass Spectra: An Efficient Parallel SearchingabstractDe novo peptide sequencing that determines the amino acid sequence of a peptide via tandem mass spectrometry (MS/MS) has been increasingly used nowadays in proteomics for protein identification. Current de novo methods generally employ a graph theory which usually produces a large number of candidate sequences and causes heavy computational cost while trying to determine a sequence with less ambiguity. We present a novel de novo sequencing algorithm which greatly reduces the number of candidate sequences. By utilizing certain properties of b- and y-ion series in MS/MS spectrum, we propose a reliable two-way parallel searching algorithm to filter out the peptide candidates which are further pruned by an intensity evidence based screening criterion. And we find an adjusted value required to determine the position of end node of b- and y-ion series for the charged +2 precursor in our graph. Results of our algorithm are compared with those of PEAKS, a well-known de novo sequencing software. Experimental results demonstrate the six sequences are identical with the correct sequences. And for the further pruning, rankings of our result remain unchanged even though the screening criterion changes. Therefore we can reduce the number of candidate sequences by adopting a proper screening criterion. Jung Hun Oh, Jean Gao |
BIBE | 1 |
| 2005 | Multicategory Classification using Extended SVM-RFE and Markov Blanket on SELDITOF Mass Spectrometry Data
Jung Hun Oh, Jean Gao, Animesh Nandi, Prem Gurnani, Lynne Knowles, John Schorge, Kevin P. Rosenblatt |
CIBCB | 1 |