Byung-Jun Yoon

dblp:14/1887 · DBLP profile ↗
← Back
62ranked-venue papers
11as first author
20since 2021 · last 2026
0000-0001-9328-1101ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 28 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 15 · 13 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Iterative Decoding of Stabilizer Codes under Radiation-Induced Correlated Noise
abstract
Fault-tolerant quantum computation demands extremely low logical error rates, yet superconducting qubit arrays are subject to radiation-induced correlated noise arising from cosmic-ray muon-generated quasiparticles. The quasiparticle density is unknown and time-varying, resulting in a mismatch between the true noise statistics and the priors assumed by standard decoders, and consequently, degraded logical performance. We formalize joint noise sensing and decoding using syndrome measurements by modeling the QP density as a latent variable, which governs correlation in physical errors and syndrome measurements. Starting from a variational expectation--maximization approach, we derive an iterative algorithm that alternates between QP density estimation and syndrome-based decoding under the updated noise model. Simulations of surface-code and bivariate bicycle quantum memory under radiation-induced correlated noise demonstrate a measurable reduction in logical error probability relative to baseline decoding with a uniform prior. Beyond improved decoding performance, the inferred QP density provides diagnostic information relevant to device characterization, shielding, and chip design. These results indicate that integrating physical noise estimation into decoding can mitigate correlated noise effects and relax effective error-rate requirements for fault-tolerant quantum computation.
Anuj K. Nayak, Paul Baity, Peter J. Love, Nicholas Jeon, Byung-Jun Yoon, Adolfy Hoisie, Lav R. Varshney
ISIT5
2025 Pareto Prompt Optimization
abstract
Natural language prompt optimization, or prompt engineering, has emerged as a powerful technique to unlock the potential of Large Language Models (LLMs) for various tasks. While existing methods primarily focus on maximizing a single task-specific performance metric for LLM outputs, real-world applications often require considering trade-offs between multiple objectives. In this work, we address this limitation by proposing an effective technique for multi-objective prompt optimization for LLMs. Specifically, we propose **ParetoPrompt**, a reinforcement learning~(RL) method that leverages dominance relationships between prompts to derive a policy model for prompts optimization using preference-based loss functions. By leveraging multi-objective dominance relationships, ParetoPrompt enables efficient exploration of the entire Pareto front without the need for a predefined scalarization of multiple objectives. Our experimental results show that ParetoPrompt consistently outperforms existing algorithms that use specific objective values. ParetoPrompt also yields robust performances when the objective metrics differ between training and testing.
Byung-Jun Yoon, Gilchan Park, Shantenu Jha, Shinjae Yoo, Xiaoning Qian
ICLR2
2025 C-LoRA: Contextual Low-Rank Adaptation for Uncertainty Estimation in Large Language Models
abstract
Low-Rank Adaptation (LoRA) offers a cost-effective solution for fine-tuning large language models (LLMs), but it often produces overconfident predictions in data-scarce few-shot settings. To address this issue, several classical statistical learning approaches have been repurposed for scalable uncertainty-aware LoRA fine-tuning. However, these approaches neglect how input characteristics affect the predictive uncertainty estimates. To address this limitation, we propose Contextual Low-Rank Adaptation (**C-LoRA**) as a novel uncertainty-aware and parameter efficient fine-tuning approach, by developing new lightweight LoRA modules contextualized to each input data sample to dynamically adapt uncertainty estimates. Incorporating data-driven contexts into the parameter posteriors, C-LoRA mitigates overfitting, achieves well-calibrated uncertainties, and yields robust predictions. Extensive experiments on LLaMA2-7B models demonstrate that C-LoRA consistently outperforms the state-of-the-art uncertainty-aware LoRA methods in both uncertainty quantification and model generalization. Ablation studies further confirm the critical role of our contextual modules in capturing sample-specific uncertainties. C-LoRA sets a new standard for robust, uncertainty-aware LLM fine-tuning in few-shot regimes. Although our experiments are limited to 7B models, our method is architecture-agnostic and, in principle, applies beyond this scale; studying its scaling to larger models remains an open problem. Our code is available at https://github.com/ahra99/c_lora.
Amir Hossein Rahmati, Sanket R. Jantre, Byung-Jun Yoon, Nathan M. Urban, Xiaoning Qian
NeurIPS5
2025 A Plug-and-Play Query Synthesis Active Learning Framework for Neural PDE Solvers
abstract
In recent developments in scientific machine learning (SciML), neural surrogate solvers for partial differential equations (PDEs) have become powerful tools for accelerating scientific computation for various science and engineering applications. However, training neural PDE solvers often demands a large amount of high-fidelity PDE simulation data, which are expensive to generate. Active learning (AL) offers a promising solution by adaptively selecting training data from the PDE settings--including parameters, initial and boundary conditions--that are expected to be most informative to help reduce this data burden. In this work, we introduce PaPQS, a Plug-and-Play Query Synthesis AL framework that synthesizes informative PDE settings directly in the continuous design space. PaPQS optimizes the Expected Information Gain (EIG) while encouraging batch diversity, enabling model-aware exploration of the design space via backpropagation through the neural PDE solution trajectories. The framework is applicable to general PDE systems and surrogate architectures, and can be seamlessly integrated with existing AL strategies. Extensive experiments across different PDE systems demonstrate that our AL framework, PaPQS, consistently improves sample efficiency over existing AL baselines.
Jinwoo Go, Byung-Jun Yoon, Nathan M. Urban, Xiaoning Qian
NeurIPS3
2024 Uncertainty-aware Continuous Implicit Neural Representations for Remote Sensing Object Counting
abstract
Many existing object counting methods rely on density map estimation (DME) of the discrete grid representation by decoding extracted image semantic features from designed convolutional neural networks (CNNs). Relying on discrete density maps not only leads to information loss dependent on the original image resolution, but also has a scalability issue when analyzing high-resolution images with cubically increasing memory complexity. Furthermore, none of the existing methods can offer reliable uncertainty quantification (UQ) for the derived count estimates. To overcome these limitations, we design UNcertainty-aware, hypernetwork-based Implicit neural representations for Counting (UNIC) to assign probabilities and the corresponding counting confidence over continuous spatial coordinates. We derive a sampling-based Bayesian counting loss function and develop the corresponding model training algorithm. UNIC outperforms existing methods on the Remote Sensing Object Counting (RSOC) dataset with reliable UQ and improved interpretability of the derived count estimates. Our code is available at https://github.com/SiyuanXu-tamu/UNIC.
Mingzhou Fan, Byung-Jun Yoon, Xiaoning Qian
AISTATS4
2024 Privacy-Preserving Federated Learning for Science: Challenges and Research Directions
abstract
This paper discusses the key challenges and future research directions for privacy-preserving federated learning (PPFL), with a focus on its application to large-scale scientific artificial intelligence models, in particular, foundation models (FMs). PPFL enables collaborative model training across distributed datasets while preserving privacy—an important collaborative approach for science. We discuss the need for efficient and scalable algorithms to address the increasing complexity of FMs, particularly when dealing with heterogeneous clients. In addition, we underscore the need for developing advance privacy-preserving techniques, such as differential privacy, to balance privacy and utility in large FMs emphasizing fairness and incentive mechanisms to ensure equitable participation among heterogeneous clients. Finally, we emphasize the need for a robust software stack supporting scalable and secure PPFL deployments across multiple high-performance computing facilities. We envision that PPFL would play a crucial role to advance scientific discovery and enable large-scale, privacy-aware collaborations across science domains.
Kibaek Kim, Raghavan Krishnan, Olivera Kotevska, Matthieu Dorier, Ravi K. Madduri, Minseok Ryu, Todd S. Munson, Robert B. Ross, Thomas Flynn 0001, Ai Kagawa, Byung-Jun Yoon, Christian Engelmann, Farzad Yousefian
IEEE Big Data11
2024 Learning Active Subspaces for Effective and Scalable Uncertainty Quantification in Deep Neural Networks
abstract
Bayesian inference for neural networks, or Bayesian deep learning, has the potential to provide well-calibrated predictions with quantified uncertainty and robustness. However, the main hurdle for Bayesian deep learning is its computational complexity due to the high dimensionality of the parameter space. In this work, we propose a novel scheme that addresses this limitation by constructing a low-dimensional subspace of the neural network parameters–referred to as an active subspace–by identifying the parameter directions that have the most significant influence on the output of the neural network. We demonstrate that the significantly reduced active subspace enables effective and scalable Bayesian inference via either Monte Carlo (MC) sampling methods, otherwise computationally intractable, or variational inference. Empirically, our approach provides reliable predictions with robust uncertainty estimates for various regression tasks.
Sanket R. Jantre, Nathan M. Urban, Xiaoning Qian, Byung-Jun Yoon
ICASSP4
2024 Hierarchical Neural Operator Transformer with Learnable Frequency-aware Loss Prior for Arbitrary-scale Super-resolution
abstract
In this work, we present an arbitrary-scale super-resolution (SR) method to enhance the resolution of scientific data, which often involves complex challenges such as continuity, multi-scale physics, and the intricacies of high-frequency signals. Grounded in operator learning, the proposed method is resolution-invariant. The core of our model is a hierarchical neural operator that leverages a Galerkin-type self-attention mechanism, enabling efficient learning of mappings between function spaces. Sinc filters are used to facilitate the information transfer across different levels in the hierarchy, thereby ensuring representation equivalence in the proposed neural operator. Additionally, we introduce a learnable prior structure that is derived from the spectral resizing of the input data. This loss prior is model-agnostic and is designed to dynamically adjust the weighting of pixel contributions, thereby balancing gradients effectively across the model. We conduct extensive experiments on diverse datasets from different domains and demonstrate consistent improvements compared to strong baselines, which consist of various state-of-the-art SR methods.
Xihaier Luo, Xiaoning Qian, Byung-Jun Yoon
ICML3
2024 When Uncertainty-Based Active Learning May Fail?
Amir Hossein Rahmati, Mingzhou Fan, Ruida Zhou, Nathan M. Urban, Byung-Jun Yoon, Xiaoning Qian
ICPR (1)5
2024 Multi-fidelity Bayesian Optimization with Multiple Information Sources of Input-dependent Fidelity
abstract
By querying approximate surrogate models of different fidelity as available information sources, Multi-Fidelity Bayesian Optimization (MFBO) aims at optimizing unknown functions that are costly if not infeasible to evaluate. Existing MFBO methods often assume that approximate surrogates have consistently high/low fidelity across the input domain. However, approximate evaluations from the same surrogate can have different fidelity at different input regions due to data availability and model constraints, especially when considering machine learning surrogates. In this work, we investigate MFBO when multi-fidelity approximations have input-dependent fidelity. By explicitly capturing input dependency for multi-fidelity queries in Gaussian Process (GP), our new input-dependent MFBO (iMFBO) with learnable noise models better captures the fidelity of each information source in an intuitive way. We further design a new acquisition function for iMFBO and prove that the queries selected by iMFBO have higher quality than those by naive MFBO methods, with the derived sub-linear regret bound. Experiments on both synthetic and real-world data demonstrate its superior empirical performance.
Mingzhou Fan, Byung-Jun Yoon, Edward R. Dougherty, Nathan M. Urban, Francis J. Alexander, Raymundo Arróyave, Xiaoning Qian
UAI2
2023 Neural message-passing for objective-based uncertainty quantification and optimal experimental design
Qihua Chen, Xuejin Chen, Hyun-Myung Woo, Byung-Jun Yoon
Eng. Appl. Artif. Intell.4
2023 Geometric Affinity Propagation for Clustering With Network Knowledge
abstract
Clustering data into meaningful subsets is a major task in scientific data analysis. To date, various strategies ranging from model-based approaches to data-driven schemes, have been devised for efficient and accurate clustering. One important class of clustering methods that is of a particular interest is the class of exemplar-based approaches. This interest primarily stems from the amount of compressed information encoded in these exemplars that effectively reflect the major characteristics of the corresponding clusters. Affinity propagation (AP) has proven to be a powerful exemplar-based approach that refines the set of optimal exemplars by iterative pairwise message updates. However, a critical limitation is its inability to capitalize on known networked relations between data points often available for various scientific datasets. To address this shortcoming, we propose Geometric-AP, a novel clustering algorithm that effectively extends the original AP to take advantage of the network topology. Geometric-AP obeys network constraints and uses max-sum belief propagation to leverage the available network topology for generating smooth clusters over the network. Extensive performance assessment shows that Geometric-AP leads to a significant quality enhancement of the clustering results when compared to existing schemes. Especially, we demonstrate that Geometric-AP performs extremely well even in cases where the original AP fails drastically.
Omar Maddouri, Xiaoning Qian, Byung-Jun Yoon
IEEE Trans. Knowl. Data Eng.3
2022 Comprehensive analysis of gene expression profiles to radiation exposure reveals molecular signatures of low-dose radiation response
abstract
There are various sources of ionizing radiation exposure, where medical exposure for radiation therapy or diagnosis is the most common human-made source. Understanding how gene expression is modulated after ionizing radiation exposure and investigating the presence of any dose-dependent gene expression patterns have broad implications for health risks from radiotherapy, medical radiation diagnostic procedures, as well as other environmental exposure. In this paper, we perform a comprehensive pathway-based analysis of gene expression profiles in response to low-dose radiation exposure, in order to examine the potential mechanism of gene regulation underlying such responses. To accomplish this goal, we employ a statistical framework to determine whether a specific group of genes belonging to a known pathway display coordinated expression patterns that are modulated in a manner consistent with the radiation level. Findings in our study suggest that there exist complex yet consistent signatures that reflect the molecular response to radiation exposure, which differ between low-dose and high-dose radiation.
Xihaier Luo, Sean McCorkle, Gilchan Park, Vanessa López-Marrero, Shinjae Yoo, Edward R. Dougherty, Xiaoning Qian, Francis J. Alexander, Byung-Jun Yoon
BIBM9
2022 Adaptive Group Testing with Mismatched Models
abstract
Accurate detection of infected individuals is one of the critical steps in stopping any pandemic. When the underlying infection rate of the disease is low, testing people in groups, instead of testing each individual in the population, can be more efficient. In this work, we consider noisy adaptive group testing design with specific test sensitivity and specificity that select the optimal group given previous test results based on pre-selected utility function. As in prior studies on group testing, we model this problem as a sequential Bayesian Optimal Experimental Design (BOED) to adaptively design the groups for each test. We analyze the required number of group tests when using the updated posterior on the infection status and the corresponding Mutual Information (MI) as our utility function for selecting new groups. More importantly, we study how the potential bias on the ground-truth noise of group tests may affect the group testing sample complexity.
Mingzhou Fan, Byung-Jun Yoon, Francis J. Alexander, Edward R. Dougherty, Xiaoning Qian
ICASSP2
2022 Deep graph representations embed network information for robust disease marker identification
abstract
MOTIVATION: Accurate disease diagnosis and prognosis based on omics data rely on the effective identification of robust prognostic and diagnostic markers that reflect the states of the biological processes underlying the disease pathogenesis and progression. In this article, we present GCNCC, a Graph Convolutional Network-based approach for Clustering and Classification, that can identify highly effective and robust network-based disease markers. Based on a geometric deep learning framework, GCNCC learns deep network representations by integrating gene expression data with protein interaction data to identify highly reproducible markers with consistently accurate prediction performance across independent datasets possibly from different platforms. GCNCC identifies these markers by clustering the nodes in the protein interaction network based on latent similarity measures learned by the deep architecture of a graph convolutional network, followed by a supervised feature selection procedure that extracts clusters that are highly predictive of the disease state. RESULTS: By benchmarking GCNCC based on independent datasets from different diseases (psychiatric disorder and cancer) and different platforms (microarray and RNA-seq), we show that GCNCC outperforms other state-of-the-art methods in terms of accuracy and reproducibility. AVAILABILITY AND IMPLEMENTATION: https://github.com/omarmaddouri/GCNCC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Omar Maddouri, Xiaoning Qian, Byung-Jun Yoon
Bioinform.3
2021 Physics-constrained Automatic Feature Engineering for Predictive Modeling in Materials Science
abstract
Automatic Feature Engineering (AFE) aims to extract useful knowledge for interpretable predictions given data for the machine learning tasks. Here, we develop AFE to extract dependency relationships that can be interpreted with functional formulas to discover physics meaning or new hypotheses for the problems of interest. We focus on materials science applications, where interpretable predictive modeling may provide principled understanding of materials systems and guide new materials discovery. It is often computationally prohibitive to exhaust all the potential relationships to construct and search the whole feature space to identify interpretable and predictive features. We develop and evaluate new AFE strategies by exploring a feature generation tree (FGT) with deep Q-network (DQN) for scalable and efficient exploration policies. The developed DQN-based AFE strategies are benchmarked with the existing AFE methods on several materials science datasets.
Ziyu Xiang 0003, Mingzhou Fan, Guillermo Vázquez Tovar, William Trehern, Byung-Jun Yoon, Xiaofeng Qian, Raymundo Arróyave, Xiaoning Qian
AAAI5
2021 Bayesian Active Learning by Soft Mean Objective Cost of Uncertainty
abstract
To achieve label efficiency for training supervised learning models, pool-based active learning sequentially selects samples from a set of candidates as queries to label by optimizing an acquisition function. One category of existing methods adopts one-step-look-ahead strategies based on acquisition functions tailored with the learning objectives, for example based on the expected loss reduction (ELR) or the mean objective cost of uncertainty (MOCU) proposed recently. These active learning methods are optimal with the maximum classification error reduction when one considers a single query. However, it is well-known that there is no performance guarantee in the long run for these myopic methods. In this paper, we show that these methods are not guaranteed to converge to the optimal classifier of the true model because MOCU is not strictly concave. Moreover, we suggest a strictly concave approximation of MOCU—Soft MOCU—that can be used to define an acquisition function to guide Bayesian active learning with theoretical convergence guarantee. For training Bayesian classifiers with both synthetic and real-world data, our experiments demonstrate the superior performance of active learning by Soft MOCU compared to other existing methods.
Edward R. Dougherty, Byung-Jun Yoon, Francis J. Alexander, Xiaoning Qian
AISTATS3
2021 Uncertainty-aware Active Learning for Optimal Bayesian Classifier
Edward R. Dougherty, Byung-Jun Yoon, Francis J. Alexander, Xiaoning Qian
ICLR3
2021 Efficient Active Learning for Gaussian Process Classification by Error Reduction
abstract
Active learning sequentially selects the best instance for labeling by optimizing an acquisition function to enhance data/label efficiency. The selection can be either from a discrete instance set (pool-based scenario) or a continuous instance space (query synthesis scenario). In this work, we study both active learning scenarios for Gaussian Process Classification (GPC). The existing active learning strategies that maximize the Estimated Error Reduction (EER) aim at reducing the classification error after training with the new acquired instance in a one-step-look-ahead manner. The computation of EER-based acquisition functions is typically prohibitive as it requires retraining the GPC with every new query. Moreover, as the EER is not smooth, it can not be combined with gradient-based optimization techniques to efficiently explore the continuous instance space for query synthesis. To overcome these critical limitations, we develop computationally efficient algorithms for EER-based active learning with GPC. We derive the joint predictive distribution of label pairs as a one-dimensional integral, as a result of which the computation of the acquisition function avoids retraining the GPC for each query, remarkably reducing the computational overhead. We also derive the gradient chain rule to efficiently calculate the gradient of the acquisition function, which leads to the first query synthesis active learning algorithm implementing EER-based strategies. Our experiments clearly demonstrate the computational efficiency of the proposed algorithms. We also benchmark our algorithms on both synthetic and real-world datasets, which show superior performance in terms of sampling efficiency compared to the existing state-of-the-art algorithms.
Edward R. Dougherty, Byung-Jun Yoon, Francis J. Alexander, Xiaoning Qian
NeurIPS3
2021 MONACO: accurate biological network alignment through optimal neighborhood matching between focal nodes
abstract
MOTIVATION: Alignment of protein-protein interaction networks can be used for the unsupervised prediction of functional modules, such as protein complexes and signaling pathways, that are conserved across different species. To date, various algorithms have been proposed for biological network alignment, many of which attempt to incorporate topological similarity between the networks into the alignment process with the goal of constructing accurate and biologically meaningful alignments. Especially, random walk models have been shown to be effective for quantifying the global topological relatedness between nodes that belong to different networks by diffusing node-level similarity along the interaction edges. However, these schemes are not ideal for capturing the local topological similarity between nodes. RESULTS: In this article, we propose MONACO, a novel and versatile network alignment algorithm that finds highly accurate pairwise and multiple network alignments through the iterative optimal matching of 'local' neighborhoods around focal nodes. Extensive performance assessment based on real networks as well as synthetic networks, for which the ground truth is known, demonstrates that MONACO clearly and consistently outperforms all other state-of-the-art network alignment algorithms that we have tested, in terms of accuracy, coherence and topological quality of the aligned network regions. Furthermore, despite the sharply enhanced alignment accuracy, MONACO remains computationally efficient and it scales well with increasing size and number of networks. AVAILABILITY AND IMPLEMENTATION: Matlab implementation is freely available at https://github.com/bjyoontamu/MONACO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hyun-Myung Woo, Byung-Jun Yoon
Bioinform.2
2019 TOPAS: network-based structural alignment of RNA sequences
abstract
MOTIVATION: For many RNA families, the secondary structure is known to be better conserved among the member RNAs compared to the primary sequence. For this reason, it is important to consider the underlying folding structures when aligning RNA sequences, especially for those with relatively low sequence identity. Given a set of RNAs with unknown structures, simultaneous RNA alignment and folding algorithms aim to accurately align the RNAs by jointly predicting their consensus secondary structure and the optimal sequence alignment. Despite the improved accuracy of the resulting alignment, the computational complexity of simultaneous alignment and folding for a pair of RNAs is O(N6), which is too costly to be used for large-scale analysis. RESULTS: In order to address this shortcoming, in this work, we propose a novel network-based scheme for pairwise structural alignment of RNAs. The proposed algorithm, TOPAS, builds on the concept of topological networks that provide structural maps of the RNAs to be aligned. For each RNA sequence, TOPAS first constructs a topological network based on the predicted folding structure, which consists of sequential edges and structural edges weighted by the base-pairing probabilities. The obtained networks can then be efficiently aligned by using probabilistic network alignment techniques, thereby yielding the structural alignment of the RNAs. The computational complexity of our proposed method is significantly lower than that of the Sankoff-style dynamic programming approach, while yielding favorable alignment results. Furthermore, another important advantage of the proposed algorithm is its capability of handling RNAs with pseudoknots while predicting the RNA structural alignment. We demonstrate that TOPAS generally outperforms previous RNA structural alignment methods on RNA benchmarks in terms of both speed and accuracy. AVAILABILITY AND IMPLEMENTATION: Source code of TOPAS and the benchmark data used in this paper are available at https://github.com/bjyoontamu/TOPAS.
Chun-Chi Chen, Hyundoo Jeong, Xiaoning Qian, Byung-Jun Yoon
Bioinform.4
2019 RNAdetect: efficient computational detection of novel non-coding RNAs
abstract
MOTIVATION: Non-coding RNAs (ncRNAs) are known to play crucial roles in various biological processes, and there is a pressing need for accurate computational detection methods that could be used to efficiently scan genomes to detect novel ncRNAs. However, unlike coding genes, ncRNAs often lack distinctive sequence features that could be used for recognizing them. Although many ncRNAs are known to have a well conserved secondary structure, which provides useful cues for computational prediction, it has been also shown that a structure-based approach alone may not be sufficient for detecting ncRNAs in a single sequence. Currently, the most effective ncRNA detection methods combine structure-based techniques with a comparative genome analysis approach to improve the prediction performance. RESULTS: In this paper, we propose RNAdetect, a computational method incorporating novel features for accurate detection of ncRNAs in combination with comparative genome analysis. Given a sequence alignment, RNAdetect can accurately detect the presence of functional ncRNAs by incorporating novel predictive features based on the concept of generalized ensemble defect (GED), which assesses the degree of structure conservation across multiple related sequences and the conformation of the individual folding structures to a common consensus structure. Furthermore, n-gram models (NGMs) are used to extract features that can effectively capture sequence homology to known ncRNA families. Utilization of NGMs can enhance the detection of ncRNAs that have sparse folding structures with many unpaired bases. Extensive performance evaluation based on the Rfam database and bacterial genomes demonstrate that RNAdetect can accurately and reliably detect novel ncRNAs, outperforming the current state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The source code for RNAdetect and the benchmark data used in this paper can be downloaded at https://github.com/bjyoontamu/RNAdetect.
Chun-Chi Chen, Xiaoning Qian, Byung-Jun Yoon
Bioinform.3
2019 Selected research articles from the 2018 International Workshop on Computational Network Biology: Modeling, Analysis, and Control (CNB-MAC)
abstract
The Fifth International Workshop on Computational Network Biology: Modeling, Analysis, and Control (CNB-MAC 2018) was held in Washington, D.C. on August 29, 2018. The workshop was organized in conjunction with the ACM Conference on Bioinformatics, Computational Biology, and Health Informatics (ACM-BCB), the flagship conference of the ACM SIGBio. The CNB-MAC workshop aims to provide an international scientific forum for presenting recent advances in computational network biology that involve modeling, analysis, and control of biological systems and system-oriented analysis of large-scale OMICS data.
Byung-Jun Yoon, Xiaoning Qian, Tamer Kahveci, Ranadip Pal
BMC Bioinform.1
2018 Selected research articles from the 2017 International Workshop on Computational Network Biology: Modeling, Analysis, and Control (CNB-MAC)
abstract
The Fourth International Workshop on Computational Network Biology: Modeling, Analysis, and Control (CNB-MAC 2017) was held in Boston, Massachusetts on August 20, 2017. The workshop was organized in conjunction with the ACM Conference on Bioinformatics, Computational Biology, and Health Informatics (ACM-BCB), the flagship conference of the ACM SIGBio, as in previous years. The CNB-MAC workshop aims to provide an international scientific forum for presenting recent advances in computational network biology that involve modeling, analysis, and control of biological systems and system-oriented analysis of large-scale OMICS data.
Byung-Jun Yoon, Xiaoning Qian, Tamer Kahveci, Ranadip Pal
BMC Bioinform.1
2018 Examining De Novo Transcriptome Assemblies via a Quality Assessment Pipeline
abstract
New de novo transcriptome assembly and annotation methods provide an incredible opportunity to study the transcriptome of organisms that lack an assembled and annotated genome. There are currently a number of de novo transcriptome assembly methods, but it has been difficult to evaluate the quality of these assemblies. In order to assess the quality of the transcriptome assemblies, we composed a workflow of multiple quality check measurements that in combination provide a clear evaluation of the assembly performance. We presented novel transcriptome assemblies and functional annotations for Pacific Whiteleg Shrimp (Litopenaeus vannamei ), a mariculture species with great national and international interest, and no solid transcriptome/genome reference. We examined Pacific Whiteleg transcriptome assemblies via multiple metrics, and provide an improved gene annotation. Our investigations show that assessing the quality of an assembly purely based on the assembler's statistical measurements can be misleading; we propose a hybrid approach that consists of statistical quality checks and further biological-based evaluations.
Noushin Ghaffari, Osama A. Arshad, Hyundoo Jeong, John Thiltges, Michael F. Criscitiello, Byung-Jun Yoon, Aniruddha Datta, Charles D. Johnson
IEEE ACM Trans. Comput. Biol. Bioinform.6
2018 Computational Prediction of Pathogenic Network Modules in Fusarium verticillioides
abstract
Fusarium verticillioides is a fungal pathogen that triggers stalk rots and ear rots in maize. In this study, we performed a comparative analysis of wild type and loss-of-virulence mutant F. verticillioides co-expression networks to identify subnetwork modules that are associated with its pathogenicity. We constructed the F. verticillioides co-expression networks from RNA-Seq data and searched through these networks to identify subnetwork modules that are differentially activated between the wild type and mutant F. verticillioides, which considerably differ in terms of pathogenic potentials. A greedy seed-and-extend approach was utilized in our search, where we also used an efficient branch-out technique for reliable prediction of functional subnetwork modules in the fungus. Through our analysis, we identified four potential pathogenicity-associated subnetwork modules, each of which consists of interacting genes with coordinated expression patterns, but whose activation level is significantly different in the wild type and the mutant. The predicted modules were comprised of functionally coherent genes and topologically cohesive. Furthermore, they contained several orthologs of known pathogenic genes in other fungi, which may play important roles in the fungal pathogenesis.
Mansuck Kim, Charles Woloshuk, Won-Bo Shim, Byung-Jun Yoon
IEEE ACM Trans. Comput. Biol. Bioinform.5
2017 Effective computational detection of piRNAs using n-gram models and support vector machine
abstract
BACKGROUND: Piwi-interacting RNAs (piRNAs) are a new class of small non-coding RNAs that are known to be associated with RNA silencing. The piRNAs play an important role in protecting the genome from invasive transposons in the germline. Recent studies have shown that piRNAs are linked to the genome stability and a variety of human cancers. Due to their clinical importance, there is a pressing need for effective computational methods that can be used for computational identification of piRNAs. However, piRNAs lack conserved structural motifs and show relatively low sequence similarity across different species, which makes accurate computational prediction of piRNAs challenging. RESULTS: In this paper, we propose a novel method, piRNAdetect, for reliable computational prediction of piRNAs in genome sequences. In the proposed method, we first classify piRNA sequences in the training dataset that share similar sequence motifs and extract effective predictive features through the use of n-gram models (NGMs). The extracted NGM-based features are then used to construct a support vector machine that can be used for accurate prediction of novel piRNAs. CONCLUSIONS: We demonstrate the effectiveness of the proposed piRNAdetect algorithm through extensive performance evaluation based on piRNAs in three different species - H. sapiens, R. norvegicus, and M. musculus - obtained from the piRBase and show that piRNAdetect outperforms the current state-of-the-art methods in terms of efficiency and accuracy.
Chun-Chi Chen, Xiaoning Qian, Byung-Jun Yoon
BMC Bioinform.3
2017 CUFID-query: accurate network querying through random walk based network flow estimation
abstract
BACKGROUND: Functional modules in biological networks consist of numerous biomolecules and their complicated interactions. Recent studies have shown that biomolecules in a functional module tend to have similar interaction patterns and that such modules are often conserved across biological networks of different species. As a result, such conserved functional modules can be identified through comparative analysis of biological networks. RESULTS: In this work, we propose a novel network querying algorithm based on the CUFID (Comparative network analysis Using the steady-state network Flow to IDentify orthologous proteins) framework combined with an efficient seed-and-extension approach. The proposed algorithm, CUFID-query, can accurately detect conserved functional modules as small subnetworks in the target network that are expected to perform similar functions to the given query functional module. The CUFID framework was recently developed for probabilistic pairwise global comparison of biological networks, and it has been applied to pairwise global network alignment, where the framework was shown to yield accurate network alignment results. In the proposed CUFID-query algorithm, we adopt the CUFID framework and extend it for local network alignment, specifically to solve network querying problems. First, in the seed selection phase, the proposed method utilizes the CUFID framework to compare the query and the target networks and to predict the probabilistic node-to-node correspondence between the networks. Next, the algorithm selects and greedily extends the seed in the target network by iteratively adding nodes that have frequent interactions with other nodes in the seed network, in a way that the conductance of the extended network is maximally reduced. Finally, CUFID-query removes irrelevant nodes from the querying results based on the personalized PageRank vector for the induced network that includes the fully extended network and its neighboring nodes. CONCLUSIONS: Through extensive performance evaluation based on biological networks with known functional modules, we show that CUFID-query outperforms the existing state-of-the-art algorithms in terms of prediction accuracy and biological significance of the predictions.
Hyundoo Jeong, Xiaoning Qian, Byung-Jun Yoon
BMC Bioinform.3
2017 Selected research articles from the 2016 International Workshop on Computational Network Biology: Modeling, Analysis, and Control (CNB-MAC)
abstract
The Third International Workshop on Computational Network Biology: Modeling, Analysis, and Control (CNB-MAC 2016) was held in Seattle, Washington on October 2, 2016. As in previous years, the workshop was organized in conjunction with the ACM Conference on Bioinformatics, Computational Biology, and Health Informatics (ACM-BCB), the flagship conference of the ACM SIGBio. This workshop aims to provide an international scientific forum for presenting recent advances in computational network biology that involve modeling, analysis, and control of biological systems and system-oriented analysis of large-scale OMICS data.
Byung-Jun Yoon, Xiaoning Qian, Tamer Kahveci
BMC Bioinform.1
2016 Co-segmentation of multiple images through random walk on graphs
abstract
We present a new image co-segmentation framework to simultaneously segment multiple images by formulating the co-segmentation problem as a multiple graph clustering problem. For each image, we first construct a corresponding segment graph by extracting superpixels as vertices and assigning edge weights between superpixels according to their feature and spatial proximity. To integrate the related information across images, we further compose a similarity graph across all constructed segment graphs, in which edges capture the similarity among superpixels across images. We propose to solve the co-segmentation problem by applying an alternating random walk strategy on both the segment graphs and the similarity graph to borrow strengths across images for better segmentation. The common objects shared in images can be identified by finding low conductance sets based on the transition probability matrix of the alternating random walk on these graphs. Experiments on iCoseg and a sequence of echocardiac images demonstrate that our novel formulation yields promising results and performs better than image segmentation on individual images separately.
Yijie Wang 0003, Byung-Jun Yoon, Xiaoning Qian
ICASSP2
2016 Effective comparative analysis of protein-protein interaction networks by measuring the steady-state network flow using a Markov model
abstract
BACKGROUND: Comparative analysis of protein-protein interaction (PPI) networks provides an effective means of detecting conserved functional network modules across different species. Such modules typically consist of orthologous proteins with conserved interactions, which can be exploited to computationally predict the modules through network comparison. RESULTS: In this work, we propose a novel probabilistic framework for comparing PPI networks and effectively predicting the correspondence between proteins, represented as network nodes, that belong to conserved functional modules across the given PPI networks. The basic idea is to estimate the steady-state network flow between nodes that belong to different PPI networks based on a Markov random walk model. The random walker is designed to make random moves to adjacent nodes within a PPI network as well as cross-network moves between potential orthologous nodes with high sequence similarity. Based on this Markov random walk model, we estimate the steady-state network flow - or the long-term relative frequency of the transitions that the random walker makes - between nodes in different PPI networks, which can be used as a probabilistic score measuring their potential correspondence. Subsequently, the estimated scores can be used for detecting orthologous proteins in conserved functional modules through network alignment. CONCLUSIONS: Through evaluations based on multiple real PPI networks, we demonstrate that the proposed scheme leads to improved alignment results that are biologically more meaningful at reduced computational cost, outperforming the current state-of-the-art algorithms. The source code and datasets can be downloaded from http://www.ece.tamu.edu/~bjyoon/CUFID .
Hyundoo Jeong, Xiaoning Qian, Byung-Jun Yoon
BMC Bioinform.3
2016 Incorporating topological information for predicting robust cancer subnetwork markers in human protein-protein interaction network
abstract
BACKGROUND: Discovering robust markers for cancer prognosis based on gene expression data is an important yet challenging problem in translational bioinformatics. By integrating additional information in biological pathways or a protein-protein interaction (PPI) network, we can find better biomarkers that lead to more accurate and reproducible prognostic predictions. In fact, recent studies have shown that, "modular markers," that integrate multiple genes with potential interactions can improve disease classification and also provide better understanding of the disease mechanisms. RESULTS: In this work, we propose a novel algorithm for finding robust and effective subnetwork markers that can accurately predict cancer prognosis. To simultaneously discover multiple synergistic subnetwork markers in a human PPI network, we build on our previous work that uses affinity propagation, an efficient clustering algorithm based on a message-passing scheme. Using affinity propagation, we identify potential subnetwork markers that consist of discriminative genes that display coherent expression patterns and whose protein products are closely located on the PPI network. Furthermore, we incorporate the topological information from the PPI network to evaluate the potential of a given set of proteins to be involved in a functional module. Primarily, we adopt widely made assumptions that densely connected subnetworks may likely be potential functional modules and that proteins that are not directly connected but interact with similar sets of other proteins may share similar functionalities. CONCLUSIONS: Incorporating topological attributes based on these assumptions can enhance the prediction of potential subnetwork markers. We evaluate the performance of the proposed subnetwork marker identification method by performing classification experiments using multiple independent breast cancer gene expression datasets and PPI networks. We show that our method leads to the discovery of robust subnetwork markers that can improve cancer classification.
Navadon Khunlertgit, Byung-Jun Yoon
BMC Bioinform.2
2015 Erratum to: Efficient experimental design for uncertainty reduction in gene regulatory networks
abstract
Erratum to: Efficient experimental design for uncertainty reduction in gene regulatory networks https://dx.doi.org/10.1186/1471-2105-16-s13-s2, published online 09 September 2018. During the production of this article [1], errors occurred in equations and algorithms. The Editorial Department of BMC Bioinformatics would like to apologise and inform its readers that an updated version is now available on the BMC Bioinformatics website. Other Information Published in: BMC Bioinformatics License: https://creativecommons.org/licenses/by/4.0/ See article on publisher's website: https://dx.doi.org/10.1186/s12859-015-0839-y
Roozbeh Dehghannasiri, Byung-Jun Yoon, Edward R. Dougherty
BMC Bioinform.2
2015 Efficient experimental design for uncertainty reduction in gene regulatory networks
abstract
BACKGROUND: An accurate understanding of interactions among genes plays a major role in developing therapeutic intervention methods. Gene regulatory networks often contain a significant amount of uncertainty. The process of prioritizing biological experiments to reduce the uncertainty of gene regulatory networks is called experimental design. Under such a strategy, the experiments with high priority are suggested to be conducted first. RESULTS: The authors have already proposed an optimal experimental design method based upon the objective for modeling gene regulatory networks, such as deriving therapeutic interventions. The experimental design method utilizes the concept of mean objective cost of uncertainty (MOCU). MOCU quantifies the expected increase of cost resulting from uncertainty. The optimal experiment to be conducted first is the one which leads to the minimum expected remaining MOCU subsequent to the experiment. In the process, one must find the optimal intervention for every gene regulatory network compatible with the prior knowledge, which can be prohibitively expensive when the size of the network is large. In this paper, we propose a computationally efficient experimental design method. This method incorporates a network reduction scheme by introducing a novel cost function that takes into account the disruption in the ranking of potential experiments. We then estimate the approximate expected remaining MOCU at a lower computational cost using the reduced networks. CONCLUSIONS: Simulation results based on synthetic and real gene regulatory networks show that the proposed approximate method has close performance to that of the optimal method but at lower computational cost. The proposed approximate method also outperforms the random selection policy significantly. A MATLAB software implementing the proposed experimental design method is available at http://gsp.tamu.edu/Publications/supplementary/roozbeh15a/.
Roozbeh Dehghannasiri, Byung-Jun Yoon, Edward R. Dougherty
BMC Bioinform.2
2015 Computational identification of genetic subnetwork modules associated with maize defense response to Fusarium verticillioides
abstract
BACKGROUND: Maize, a crop of global significance, is vulnerable to a variety of biotic stresses resulting in economic losses. Fusarium verticillioides (teleomorph Gibberella moniliformis) is one of the key fungal pathogens of maize, causing ear rots and stalk rots. To better understand the genetic mechanisms involved in maize defense as well as F. verticillioides virulence, a systematic investigation of the host-pathogen interaction is needed. The aim of this study was to computationally identify potential maize subnetwork modules associated with its defense response against F. verticillioides. RESULTS: We obtained time-course RNA-seq data from B73 maize inoculated with wild type F. verticillioides and a loss-of-virulence mutant, and subsequently established a computational pipeline for network-based comparative analysis. Specifically, we first analyzed the RNA-seq data by a cointegration-correlation-expression approach, where maize genes were jointly analyzed with known F. verticillioides virulence genes to find candidate maize genes likely associated with the defense mechanism. We predicted maize co-expression networks around the selected maize candidate genes based on partial correlation, and subsequently searched for subnetwork modules that were differentially activated when inoculated with two different fungal strains. Based on our analysis pipeline, we identified four potential maize defense subnetwork modules. Two were directly associated with maize defense response and were associated with significant GO terms such as GO:0009817 (defense response to fungus) and GO:0009620 (response to fungus). The other two predicted modules were indirectly involved in the defense response, where the most significant GO terms associated with these modules were GO:0046914 (transition metal ion binding) and GO:0046686 (response to cadmium ion). CONCLUSION: Through our RNA-seq data analysis, we have shown that a network-based approach can enhance our understanding of the complicated host-pathogen interactions between maize and F. verticillioides by interpreting the transcriptome data in a system-oriented manner. We expect that the proposed analytic pipeline can also be adapted for investigating potential functional modules associated with host defense response in diverse plant-pathogen interactions.
Mansuck Kim, Charles Woloshuk, Won-Bo Shim, Byung-Jun Yoon
BMC Bioinform.5
2015 Discrete optimal Bayesian classification with error-conditioned sequential sampling
Ariana Broumand, Mohammad Shahrokh Esfahani, Byung-Jun Yoon, Edward R. Dougherty
Pattern Recognit.3
2015 Effective Estimation of Node-to-Node Correspondence Between Different Graphs
abstract
In this work, we propose a novel method for accurately estimating the node-to-node correspondence between two graphs. Given two graphs and their pairwise node similarity scores, our goal is to quantitatively measure the overall similarity-or the correspondence-between nodes that belong to different graphs. The proposed method is based on a Markov random walk model that performs a simultaneous random walk on two graphs. Unlike previous random walk models, the proposed random walker examines the neighboring nodes at each step and adjusts its mode of random walk, where it can switch between a simultaneous walk on both graphs and an individual walk on one of the two graphs. Based on extensive simulation results, we demonstrate that our random walk model yields better node correspondence scores that can more accurately identify nodes and edges that are conserved across graphs.
Hyundoo Jeong, Byung-Jun Yoon
IEEE Signal Process. Lett.2
2015 Optimal Experimental Design for Gene Regulatory Networks in the Presence of Uncertainty
abstract
Of major interest to translational genomics is the intervention in gene regulatory networks (GRNs) to affect cell behavior; in particular, to alter pathological phenotypes. Owing to the complexity of GRNs, accurate network inference is practically challenging and GRN models often contain considerable amounts of uncertainty. Considering the cost and time required for conducting biological experiments, it is desirable to have a systematic method for prioritizing potential experiments so that an experiment can be chosen to optimally reduce network uncertainty. Moreover, from a translational perspective it is crucial that GRN uncertainty be quantified and reduced in a manner that pertains to the operational cost that it induces, such as the cost of network intervention. In this work, we utilize the concept of mean objective cost of uncertainty (MOCU) to propose a novel framework for optimal experimental design. In the proposed framework, potential experiments are prioritized based on the MOCU expected to remain after conducting the experiment. Based on this prioritization, one can select an optimal experiment with the largest potential to reduce the pertinent uncertainty present in the current network model. We demonstrate the effectiveness of the proposed method via extensive simulations based on synthetic and real gene regulatory networks.
Roozbeh Dehghannasiri, Byung-Jun Yoon, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.2
2013 Classifier design given an uncertainty class of feature distributions via regularized maximum likelihood and the incorporation of biological pathway knowledge in steady-state phenotype classification
Mohammad Shahrokh Esfahani, Jason M. Knight, Amin Zollanvari, Byung-Jun Yoon, Edward R. Dougherty
Pattern Recognit.4
2012 Structural intervention of gene regulatory networks by general rank-k matrix perturbation
abstract
One of the ultimate objectives of studying gene regulatory networks is to derive potential intervention strategies to avoid aberrant cellular behavior. Boolean networks (BNs) and their stochastic extension, probabilistic Boolean networks (PBNs), provide a convenient framework to design different types of intervention strategies. In this paper, we focus on studying structural intervention, in which we perturb regulatory Boolean functions to alter the long-term network dynamics to obtain desirable behavior. Specifically, we extend our previous work that derives optimal structural intervention for rank-1 function perturbations to more general solutions for arbitrary rank-k function perturbations. The analytic solution is derived using the Sherman-Morrison-Woodbury (SMW) formula. We apply the derived structural intervention to a mutated mammalian cell cycle network. Our results show that our intervention strategy correctly identifies the main targets to stop uncontrolled cell growth in the mutated cell cycle network.
Xiaoning Qian, Byung-Jun Yoon, Edward R. Dougherty
ICASSP2
2012 RESQUE: Network Reduction Using Semi-Markov Random Walk Scores for Efficient Querying of Biological Networks (Extended Abstract)
Sayed Mohammad Ebrahim Sahraeian, Byung-Jun Yoon
RECOMB2
2012 RESQUE: Network reduction using semi-Markov random walk scores for efficient querying of biological networks
abstract
MOTIVATION: Recent technological advances in measuring molecular interactions have resulted in an increasing number of large-scale biological networks. Translation of these enormous network data into meaningful biological insights requires efficient computational techniques that can unearth the biological information that is encoded in the networks. One such example is network querying, which aims to identify similar subnetwork regions in a large target network that are similar to a given query network. Network querying tools can be used to identify novel biological pathways that are homologous to known pathways, thereby enabling knowledge transfer across different organisms. RESULTS: In this article, we introduce an efficient algorithm for querying large-scale biological networks, called RESQUE. The proposed algorithm adopts a semi-Markov random walk (SMRW) model to probabilistically estimate the correspondence scores between nodes that belong to different networks. The target network is iteratively reduced based on the estimated correspondence scores, which are also iteratively re-estimated to improve accuracy until the best matching subnetwork emerges. We demonstrate that the proposed network querying scheme is computationally efficient, can handle any network query with an arbitrary topology and yields accurate querying results. AVAILABILITY: The source code of RESQUE is freely available at http://www.ece.tamu.edu/~bjyoon/RESQUE/
Sayed Mohammad Ebrahim Sahraeian, Byung-Jun Yoon
Bioinform.2
2011 Contour-based hidden Markov model to segment 2D ultrasound images
abstract
The segmentation of ultrasound images is challenging due to the difficulty of appropriate modeling of their appearance variations including speckle as well as signal dropout. We propose a novel automatic segmentation method for 2D cardiac ultrasound images based on hidden Markov models (HMMs). By directly exploiting the local image characteristics around contour points in images and integrating them into contour-based HMMs, we solve the segmentation problem by graph matching using an efficient dynamic programming algorithm. Due to the direct integration of local properties in our HMMs, our segmentation method automatically deals with inhomogeneity but avoids the complexities of explicit appearance modeling in classical Maximum A Posteriori (MAP) approaches. The optimization for contour extraction is straightforward and guarantees the global optimal results. We implemented our method to segment the endocardium in short-axis cardiac ultrasound images successfully. The method can also be used for other image modalities with the presence of image inhomogeneity.
Xiaoning Qian, Byung-Jun Yoon
ICASSP2
2011 Fast network querying algorithm for searching large-scale biological networks
abstract
Network querying aims to search a large network for subnetwork regions that are similar to a given query network. In this paper, we propose a novel algorithm for querying large scale protein interaction networks. In this algorithm, we iteratively compute the correspondence scores between nodes in the query and the target networks using semi-Markov random walk. Based on these scores, we reduce the search space in the target network by discarding irrelevant nodes. The scores are re-estimated in each iteration after removing such nodes, which ultimately leads to more accurate querying result. Numerical experiments based on both synthetic and real networks show that the algorithm can efficiently find accurate querying results.
Sayed Mohammad Ebrahim Sahraeian, Byung-Jun Yoon
ICASSP2
2011 Probabilistic reconstruction of the tumor progression process in gene regulatory networks in the presence of uncertainty
abstract
BACKGROUND: Accumulation of gene mutations in cells is known to be responsible for tumor progression, driving it from benign states to malignant states. However, previous studies have shown that the detailed sequence of gene mutations, or the steps in tumor progression, may vary from tumor to tumor, making it difficult to infer the exact path that a given type of tumor may have taken. RESULTS: In this paper, we propose an effective probabilistic algorithm for reconstructing the tumor progression process based on partial knowledge of the underlying gene regulatory network and the steady state distribution of the gene expression values in a given tumor. We take the BNp (Boolean networks with pertubation) framework to model the gene regulatory networks. We assume that the true network is not exactly known but we are given an uncertainty class of networks that contains the true network. This network uncertainty class arises from our partial knowledge of the true network, typically represented as a set of local pathways that are embedded in the global network. Given the SSD of the cancerous network, we aim to simultaneously identify the true normal (healthy) network and the set of gene mutations that drove the network into the cancerous state. This is achieved by analyzing the effect of gene mutation on the SSD of a gene regulatory network. At each step, the proposed algorithm reduces the uncertainty class by keeping only those networks whose SSDs get close enough to the cancerous SSD as a result of additional gene mutation. These steps are repeated until we can find the best candidate for the true network and the most probable path of tumor progression. CONCLUSIONS: Simulation results based on both synthetic networks and networks constructed from actual pathway knowledge show that the proposed algorithm can identify the normal network and the actual path of tumor progression with high probability. The algorithm is also robust to model mismatch and allows us to control the trade-off between efficiency and accuracy.
Mohammad Shahrokh Esfahani, Byung-Jun Yoon, Edward R. Dougherty
BMC Bioinform.2
2011 Enhancing the accuracy of HMM-based conserved pathway prediction using global correspondence scores
abstract
BACKGROUND: Comparative network analysis aims to identify common subnetworks in biological networks. It can facilitate the prediction of conserved functional modules across different species and provide deep insights into their underlying regulatory mechanisms. Recently, it has been shown that hidden Markov models (HMMs) can provide a flexible and computationally efficient framework for modeling and comparing biological networks. RESULTS: In this work, we show that using global correspondence scores between molecules can improve the accuracy of the HMM-based network alignment results. The global correspondence scores are computed by performing a semi-Markov random walk on the networks to be compared. The resulting score naturally integrates the sequence similarity between molecules and the topological similarity between their molecular interactions, thereby providing a more effective measure for estimating the functional similarity between molecules. By incorporating the global correspondence scores, instead of relying on sequence similarity or functional annotation scores used by previous approaches, our HMM-based network alignment method can identify conserved subnetworks that are functionally more coherent. CONCLUSIONS: Performance analysis based on synthetic and microbial networks demonstrates that the proposed network alignment strategy significantly improves the robustness and specificity of the predicted alignment results, in terms of conserved functional similarity measured based on KEGG ortholog (KO) groups. These results clearly show that the HMM-based network alignment framework using global correspondence scores can effectively find conserved biological pathways and has the potential to be used for automatic functional annotation of biomolecules.
Xiaoning Qian, Sayed Mohammad Ebrahim Sahraeian, Byung-Jun Yoon
BMC Bioinform.3
2011 Comparative analysis of protein interaction networks reveals that conserved pathways are susceptible to HIV-1 interception
abstract
BACKGROUND: Human immunodeficiency virus type one (HIV-1) is the major pathogen that causes the acquired immune deficiency syndrome (AIDS). With the availability of large-scale protein-protein interaction (PPI) measurements, comparative network analysis can provide a promising way to study the host-virus interactions and their functional significance in the pathogenesis of AIDS. Until now, there have been a large number of HIV studies based on various animal models. In this paper, we present a novel framework for studying the host-HIV interactions through comparative network analysis across different species. RESULTS: Based on the proposed framework, we test our hypothesis that HIV-1 attacks essential biological pathways that are conserved across species. We selected the Homo sapiens and Mus musculus PPI networks with the largest coverage among the PPI networks that are available from public databases. By using a local network alignment algorithm based on hidden Markov models (HMMs), we first identified the pathways that are conserved in both networks. Next, we analyzed the HIV-1 susceptibility of these pathways, in comparison with random pathways in the human PPI network. Our analysis shows that the conserved pathways have a significantly higher probability of being intercepted by HIV-1. Furthermore, Gene Ontology (GO) enrichment analysis shows that most of the enriched GO terms are related to signal transduction, which has been conjectured to be one of the major mechanisms targeted by HIV-1 for the takeover of the host cell. CONCLUSIONS: This proof-of-concept study clearly shows that the comparative analysis of PPI networks across different species can provide important insights into the host-HIV interactions and the detailed mechanisms of HIV-1. We expect that comparative multiple network analysis of various species that have different levels of susceptibility to similar lentiviruses may provide a very effective framework for generating novel, and experimentally verifiable hypotheses on the mechanisms of HIV-1. We believe that the proposed framework has the potential to expedite the elucidation of the important mechanisms of HIV-1, and ultimately, the discovery of novel anti-HIV drugs.
Xiaoning Qian, Byung-Jun Yoon
BMC Bioinform.2
2011 PicXAA-R: Efficient structural alignment of multiple RNA sequences using a greedy approach
abstract
BACKGROUND: Accurate and efficient structural alignment of non-coding RNAs (ncRNAs) has grasped more and more attentions as recent studies unveiled the significance of ncRNAs in living organisms. While the Sankoff style structural alignment algorithms cannot efficiently serve for multiple sequences, mostly progressive schemes are used to reduce the complexity. However, this idea tends to propagate the early stage errors throughout the entire process, thereby degrading the quality of the final alignment. For multiple protein sequence alignment, we have recently proposed PicXAA which constructs an accurate alignment in a non-progressive fashion. RESULTS: Here, we propose PicXAA-R as an extension to PicXAA for greedy structural alignment of ncRNAs. PicXAA-R efficiently grasps both folding information within each sequence and local similarities between sequences. It uses a set of probabilistic consistency transformations to improve the posterior base-pairing and base alignment probabilities using the information of all sequences in the alignment. Using a graph-based scheme, we greedily build up the structural alignment from sequence regions with high base-pairing and base alignment probabilities. CONCLUSIONS: Several experiments on datasets with different characteristics confirm that PicXAA-R is one of the fastest algorithms for structural alignment of multiple RNAs and it consistently yields accurate alignment results, especially for datasets with locally similar sequences. PicXAA-R source code is freely available at: http://www.ece.tamu.edu/~bjyoon/picxaa/.
Sayed Mohammad Ebrahim Sahraeian, Byung-Jun Yoon
BMC Bioinform.2
2011 Enhanced stochastic optimization algorithm for finding effective multi-target therapeutics
abstract
BACKGROUND: For treating a complex disease such as cancer, we need effective means to control the biological network that underlies the disease. However, biological networks are typically robust to external perturbations, making it difficult to beneficially alter the network dynamics by controlling a single target. In fact, multi-target therapeutics is often more effective compared to monotherapies, and combinatory drugs are commonly used these days for treating various diseases. A practical challenge in combination therapy is that the number of possible drug combinations increases exponentially, which makes the prediction of the optimal drug combination a difficult combinatorial optimization problem. Recently, a stochastic optimization algorithm called the Gur Game algorithm was proposed for drug optimization, which was shown to be very efficient in finding potent drug combinations. RESULTS: In this paper, we propose a novel stochastic optimization algorithm that can be used for effective optimization of combinatory drugs. The proposed algorithm analyzes how the concentration change of a specific drug affects the overall drug response, thereby making an informed guess on how the concentration should be updated to improve the drug response. We evaluated the performance of the proposed algorithm based on various drug response functions, and compared it with the Gur Game algorithm. CONCLUSIONS: Numerical experiments clearly show that the proposed algorithm significantly outperforms the original Gur Game algorithm, in terms of reliability and efficiency. This enhanced optimization algorithm can provide an effective framework for identifying potent drug combinations that lead to optimal drug response.
Byung-Jun Yoon
BMC Bioinform.1
2011 A Novel Low-Complexity HMM Similarity Measure
abstract
In this letter, we propose a novel similarity measure for comparing Hidden Markov models (HMMs) and an efficient scheme for its computation. In the proposed approach, we probabilistically evaluate the correspondence, or goodness of match, between every pair of states in the respective HMMs, based on the concept of semi-Markov random walk. We show that this correspondence score reflects the contribution of a given state pair to the overall similarity between the two HMMs. For similar HMMs, each state in one HMM is expected to have only a few matching states in the other HMM, resulting in a sparse state correspondence score matrix. This allows us to measure the similarity between HMMs by evaluating the sparsity of the state correspondence matrix. Estimation of the proposed similarity score does not require time-consuming Monte-Carlo simulations, hence it can be computed much more efficiently compared to the Kullback–Leibler divergence (KLD) thas has been widely used. We demonstrate the effectiveness of the proposed measure through several examples.
Sayed Mohammad Ebrahim Sahraeian, Byung-Jun Yoon
IEEE Signal Process. Lett.2
2010 Shape matching based on graph alignment using hidden Markov models
abstract
We present a novel framework based on hidden Markov models (HMMs) for matching feature point sets, which capture the shapes of object contours of interest. Point matching algorithms provide effective tools for shape analysis, an important problem in computer vision and image processing applications. Typically, it is computationally expensive to find the optimal correspondence between feature points in different sets, hence existing algorithms often resort to various heuristics that find suboptimal solutions. Unlike most of the previous algorithms, the proposed HMM-based framework allows us to find the optimal correspondence using an efficient dynamic programming algorithm, where the computational complexity of the resulting shape matching algorithm grows only linearly with the size of the respective point sets. We demonstrate the promising potential of the proposed algorithm based on several benchmark data sets.
Xiaoning Qian, Byung-Jun Yoon
ICASSP2
2010 Identifying reliable subnetwork markers in protein-protein interaction network for classification of breast cancer metastasis
abstract
Due to the inherent measurement noise in microarray experiments, heterogeneity across samples, and limited sample size, it is often hard to find reliable gene markers for classification. For this reason, several studies proposed to analyze the expression data at the level of groups of functionally related genes such as pathways. One practical problem of these pathway-based approaches is the limited coverage of genes by known pathways. To overcome this problem, we propose a new method for identifying effective subnetwork markers by overlaying the gene expression data with a genome-scale protein-protein interaction network. Experimental results on two independent breast cancer datasets show that the subnetwork markers lead to more accurate classification of breast cancer metastasis and are more reproducible than both gene and pathway markers.
Junjie Su, Byung-Jun Yoon
ICASSP2
2010 Identification of diagnostic subnetwork markers for cancer in human protein-protein interaction network
abstract
BACKGROUND: Finding reliable gene markers for accurate disease classification is very challenging due to a number of reasons, including the small sample size of typical clinical data, high noise in gene expression measurements, and the heterogeneity across patients. In fact, gene markers identified in independent studies often do not coincide with each other, suggesting that many of the predicted markers may have no biological significance and may be simply artifacts of the analyzed dataset. To find more reliable and reproducible diagnostic markers, several studies proposed to analyze the gene expression data at the level of groups of functionally related genes, such as pathways. Studies have shown that pathway markers tend to be more robust and yield more accurate classification results. One practical problem of the pathway-based approach is the limited coverage of genes by currently known pathways. As a result, potentially important genes that play critical roles in cancer development may be excluded. To overcome this problem, we propose a novel method for identifying reliable subnetwork markers in a human protein-protein interaction (PPI) network. RESULTS: In this method, we overlay the gene expression data with the PPI network and look for the most discriminative linear paths that consist of discriminative genes that are highly correlated to each other. The overlapping linear paths are then optimally combined into subnetworks that can potentially serve as effective diagnostic markers. We tested our method on two independent large-scale breast cancer datasets and compared the effectiveness and reproducibility of the identified subnetwork markers with gene-based and pathway-based markers. We also compared the proposed method with an existing subnetwork-based method. CONCLUSIONS: The proposed method can efficiently find reliable subnetwork markers that outperform the gene-based and pathway-based markers in terms of discriminative power, reproducibility and classification performance. Subnetwork markers found by our method are highly enriched in common GO terms, and they can more accurately classify breast cancer metastasis compared to markers found by a previous method.
Junjie Su, Byung-Jun Yoon, Edward R. Dougherty
BMC Bioinform.2
2008 Coding Overcomplete Representations of Audio Using the MCLT
abstract
We propose a system for audio coding using the modulated complex lapped transform (MCLT). In general, it is difficult to encode signals using overcomplete representations without avoiding a penalty in rate-distortion performance. We show that the penalty can be significantly reduced for MCLT-based representations, without the need for iterative methods of sparsity reduction. We achieve that via a magnitude-phase polar quantization and the use of magnitude and phase prediction. Compared to systems based on quantization of orthogonal representations such as the modulated lapped transform (MLT), the new system allows for reduced warbling artifacts and more precise computation of frequency-domain auditory masking functions.
Byung-Jun Yoon, Henrique S. Malvar
DCC1
2007 Robust Adaptive Beamforming Algorithm using Instantaneous Direction of Arrival with Enhanced Noise Suppression Capability
abstract
In this paper, we propose a novel adaptive beamforming algorithm with enhanced noise suppression capability. The proposed algorithm incorporates the sound-source presence probability into the adaptive blocking matrix, which is estimated based on the instantaneous direction of arrival of the input signals and voice activity detection. The proposed algorithm guarantees robustness to steering vector errors without imposing ad hoc constraints on the adaptive filter coefficients. It can provide good suppression performance for both directional interference signals as well as isotropic ambient noise. For in-car environment the proposed beamformer shows SNR improvement up to 12 dB without using an additional noise suppressor.
Byung-Jun Yoon, Ivan Tashev, Alex Acero
ICASSP (1)1
2007 Fast Search of Sequences with Complex Symbol Correlations using Profile Context-Sensitive HMMS and Pre-Screening Filters
abstract
Recently, profile context-sensitive HMMs (profile-csHMMs) have been proposed which are very effective in modeling the common patterns and motifs in related symbol sequences. Profile-csHMMs are capable of representing long-range correlations between distant symbols, even when these correlations are entangled in a complicated manner. This makes profile-csHMMs an useful tool in computational biology, especially in modeling noncoding RNAs (ncRNAs) and finding new ncRNA genes. However, a profile-csHMM based search is quite slow, hence not practical for searching a large database. In this paper, we propose a practical scheme for making the search speed significantly faster without any degradation in the prediction accuracy. The proposed method utilizes a pre-screening filter based on a profile-HMM, which filters out most sequences that will not be predicted as a match by the original profile-csHMM. Experimental results show that the proposed approach can make the search speed eighty times faster.
Byung-Jun Yoon, P. P. Vaidyanathan
ICASSP (1)1
2006 Profile Context-Sensitive HMMs for Probabilistic Modeling of Sequences With Complex Correlations
abstract
The profile hidden Markov model is a specific type of HMM that is well suited for describing the common features of a set of related sequences. It has been extensively used in computational biology, where it is still one of the most popular tools. In this paper, we propose a new model called the profile context-sensitive HMM. Unlike traditional profile-HMMs, the proposed model is capable of describing complex long-range correlations between distant symbols in a consensus sequence. We also introduce a general algorithm that can be used for finding the optimal state-sequence of an observed symbol sequence based on the given profile-csHMM. The proposed model has an important application in RNA sequence analysis, especially in modeling and analyzing RNA pseudoknots.
Byung-Jun Yoon, P. P. Vaidyanathan
ICASSP (3)1
2006 A practical approach for the design of nonuniform lapped transforms
abstract
We propose a simple method for the design of lapped transforms with nonuniform frequency resolution and good time localization. The method is a generalization of an approach previously proposed by Princen, where the nonuniform filter bank is obtained by joining uniform cosine-modulated filter banks (CMFBs) using a transition filter. We use several transition filters to obtain a near perfect-reconstruction (PR) nonuniform lapped transform with significantly reduced overall distortion. The main advantage of the proposed method is in reducing the length of the transition filters, which leads to a reduction in processing delay that can be useful for applications such as real-time audio coding.
Byung-Jun Yoon, Henrique S. Malvar
IEEE Signal Process. Lett.1
2005 Optimal alignment algorithm for context-sensitive hidden Markov models
abstract
The hidden Markov model is well-known for its efficiency in modeling short-term dependencies between adjacent samples. However, it cannot be used for modeling longer-range interactions between symbols that are distant from each other. In this paper, we introduce the concept of context-sensitive HMM that is capable of modeling strong pairwise correlations between distant symbols. Based on this model, we propose a polynomial-time algorithm that can be used for finding the optimal state sequence of an observed symbol string. The proposed model is especially useful in modeling palindromes, which has an important application in RNA secondary structure analysis.
Byung-Jun Yoon, P. P. Vaidyanathan
ICASSP (4)1
2004 Wavelet-based denoising by customized thresholding
abstract
The problem of estimating a signal that is corrupted by additive noise has been of interest to many researchers for practical, as well as theoretical, reasons. Many of the traditional denoising methods use linear methods such as Wiener filtering. Recently, nonlinear methods, especially those based on wavelets, have become increasingly popular, due to a number of advantages over the linear methods. It has been shown that wavelet-thresholding has near-optimal properties in the minimax sense, and guarantees a better rate of convergence, despite its simplicity. Even though much work has been done in the field of wavelet-thresholding, most of it was focused on statistical modeling of the wavelet coefficients and the optimal choice of the thresholds. We propose a custom thresholding function which can improve the denoised results significantly. Simulation results are given to demonstrate the advantage of the new thresholding function.
Byung-Jun Yoon, P. P. Vaidyanathan
ICASSP (2)1
2003 Discrete probability density estimation using multirate DSP models
abstract
We propose a model based approach for estimation of probability mass functions for discrete random variables. The model is based on tools from multirate signal processing. Similar in principle to the kernel based methods, the approach takes advantage of well-known results from multirate signal processing theory. Similarities to and differences from wavelet based approaches are also indicated where appropriate. In the final form, the probability estimates are obtained by filtering the square root of the histogram through a multirate system whose components are biorthogonal partners of each other.
P. P. Vaidyanathan, Byung-Jun Yoon
ICASSP (6)2
2003 Discrete probability density estimation using multirate DSP models
abstract
We propose a model based approach for estimation of probability mass functions for discrete random variables. The model is based on tools from multirate signal processing. Similar in principle to the kernel based methods, the approach takes advantage of well-known results from multirate signal processing theory. Similarities to and differences from wavelet based approaches is also indicated where appropriate. In the final form, the probability estimates are obtained by filtering the square root of the histogram through a multirate system whose components are biorthogonal partners of each other.
P. P. Vaidyanathan, Byung-Jun Yoon
ICME2