Tae-Hyuk Ahn

dblp:24/1439 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-7281-9459ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021
YearPublicationVenuePosition
2025 Intelligent Code Completion by a Unified Multi-task Learning with a Large Language Model
abstract
Code completion has become an essential tool in modern software development. It helps developers by predicting the next token (e.g., an API function call) based on the current coding context. Its widespread use highlights the need for efficient and context-aware solutions that streamline the development process. Despite ongoing efforts to enhance code completion performance, many existing approaches remain limited, offering ranked lists primarily based on alphabetical order or usage frequency from partially typed code fragments. While these studies have seen incremental improvements, the level of meaningful assistance provided to developers has not advanced in parallel. To address this limitation, we propose CODECOM, a deep learning-based code completion technique that leverages a large language model (LLM) with multi-task learning. Our approach processes sequences of source code tokens, their corresponding abstract syntax trees (ASTs), and program dependencies, enabling more context-aware and accurate code predictions. In our case study, CODECOM demonstrates state-of-the-art performance in the code completion downstream task. The evaluation results highlight significant improvements over the baseline, achieving 29.49% in BLEU, 71.16% in Acc@1, and 67.19% in Acc@5. These advancements are further validated by the Wilcoxon signed-rank test, which confirms strong statistical significance across all metrics. These findings indicate that Code Com can accelerate software development and assist developers in reducing potential errors effectively. Index Terms-software maintenance, automated code completion, deep learning, large language model.
Shradha Maharjan, Tae-Hyuk Ahn, Myoungkyu Song
SERA3
2025 Bioinformatic approaches to blood and tissue microbiome analyses: challenges and perspectives
abstract
Advances in next-generation sequencing have resulted in a growing understanding of the microbiome and its role in human health. Unlike traditional microbiome analysis, blood and tissue microbiome analyses focus on the detection and characterization of microbial DNA in blood and tissue, previously considered a sterile environment. In this review, we discuss the challenges and methodologies associated with analyzing these samples, particularly emphasizing blood and tissue microbiome research. Key preprocessing steps-including the removal of ribosomal RNA, host DNA, and other contaminants-are critical to reducing noise and accurately capturing microbial evidence. We also explore how taxonomic profiling tools, machine learning, and advanced normalization techniques address contamination and low microbial biomass, thereby improving reliability. While it offers the potential for identifying microbial involvement in systemic diseases previously undetectable by traditional methods, this methodology also carries risks and lacks universal acceptance due to concerns over reliability and interpretation errors. This paper critically reviews these factors, highlighting both the promise and pitfalls of using blood and tissue microbiome analyses as a tool for biomarker discovery.
Jammi Prasanthi Sirasani, Cory Gardner, Gihwan Jung, Tae-Hyuk Ahn
Briefings Bioinform.5
2022 An IDE Support for Validating Machine Learning Applications in Bioengineering Text Corpora
abstract
Modeling in machine learning (ML) is critical for software systems in practice. ML applications are required to validate their models and implementations but quality validation is a challenging and time-consuming process for developers. To address this limitation, we present a novel validation technique for ML applications to help developers or researchers (e.g., bioengineering domain) inspect (1) software code (ML API usages) and (2) ML model (extracted features).
Piyush Basia, Tae-Hyuk Ahn, Myoungkyu Song
BIBM2
2022 Debugging Support for Machine Learning Applications in Bioengineering Text Corpora
abstract
Modeling in machine learning (ML) is becoming an essential part of software systems in practice. Validating ML applications is a challenging and time-consuming process for developers since the accuracy of prediction heavily relies on generated models. ML applications are written by relatively more data-driven programming based on the blackbox of ML frameworks. If all of the datasets and the ML application need to be individually investigated, the ML debugging tasks would take a lot of time and effort. To address this limitation, we present a novel debugging technique for machine learning applications, called MLDBUG that helps ML application developers inspect the training data and the generated features for the ML model. Inspired by software debugging for reproducing the potential reported bugs, MLDBUG takes as input an ML application and its training datasets to build the ML models, helping ML application developers easily reproduce and understand anomalies on the ML application. We have implemented an Eclipse plugin for MLDBUG which allows developers to validate the prediction behavior of their ML applications, the ML model, and the training data on the Eclipse IDE. In our evaluation, we used 23,500 documents in the bioengineering research domain. We assessed the MLDBUG's capability of how effectively our debugging technique can help ML application developers investi-gate the connection between the produced features and the labels in the training model and the relationship between the training instances and the instances the model predicts.
Kwok Sun Cheng, Tae-Hyuk Ahn, Myoungkyu Song
COMPSAC2
2021 MegaR: an interactive R package for rapid sample classification and phenotype prediction using metagenome profiles and machine learning
abstract
BACKGROUND: Diverse microbiome communities drive biogeochemical processes and evolution of animals in their ecosystems. Many microbiome projects have demonstrated the power of using metagenomics to understand the structures and factors influencing the function of the microbiomes in their environments. In order to characterize the effects from microbiome composition for human health, diseases, and even ecosystems, one must first understand the relationship of microbes and their environment in different samples. Running machine learning model with metagenomic sequencing data is encouraged for this purpose, but it is not an easy task to make an appropriate machine learning model for all diverse metagenomic datasets. RESULTS: We introduce MegaR, an R Shiny package and web application, to build an unbiased machine learning model effortlessly with interactive visual analysis. The MegaR employs taxonomic profiles from either whole metagenome sequencing or 16S rRNA sequencing data to develop machine learning models and classify the samples into two or more categories. It provides various options for model fine tuning throughout the analysis pipeline such as data processing, multiple machine learning techniques, model validation, and unknown sample prediction that can be used to achieve the highest prediction accuracy possible for any given dataset while still maintaining a user-friendly experience. CONCLUSIONS: Metagenomic sample classification and phenotype prediction is important particularly when it applies to a diagnostic method for identifying and predicting microbe-related human diseases. MegaR provides various interactive visualizations for user to build an accurate machine-learning model without difficulty. Unknown sample prediction with a properly trained model using MegaR will enhance researchers to identify the sample property in a fast turnaround time.
Eliza Dhungel, Yassin Mreyoud, Ho-Jin Gwak, Ahmad Rajeh, Mina Rho, Tae-Hyuk Ahn
BMC Bioinform.6
2019 Tool support for managing repetitive program changes in evolving software
abstract
Software modification often requires consistent program changes , a group of similar, related changes, at multiple locations in a program. Developers are typically uneasy to (i) detect potential change anomalies such as omission errors or incorrect edits and (ii) determine related locations to apply consistent changes, which is a tedious and error‐prone process. To address this problem, this study presents a technique for managing consistent program changes, checking and applying repetitive program transformation (CARP). Given program differencing results between original and edited program versions, CARP (i) infers change patterns to detect change anomalies , (ii) identifies required edit locations , and (iii) automatically applies adequate edits . It has been implemented in the context of the integrated development environment as an Eclipse plug‐in. The authors evaluated CARP on three open‐source projects and found that CARP detects seeded anomalies with 99.1% accuracy on average. Furthermore, it identifies change locations and transforms them with 98% accuracy. Their results show that CARP should help developers detect potential change anomalies in repetitive program changes and perform consistent changes automatically.
Vamshi Krishna Epuri, Sushma Sakala, Tae-Hyuk Ahn, Myoungkyu Song
IET Softw.3
2018 SORA: Scalable Overlap-graph Reduction Algorithms for Genome Assembly using Apache Spark in the Cloud
Alexander J. Paul, Dylan Lawrence, Myoungkyu Song, Seung-Hwan Lim, Chongle Pan, Tae-Hyuk Ahn
BIBM6
2018 LONGO: an R package for interactive gene length dependent analysis for neuronal identity
abstract
Motivation: Reprogramming somatic cells into neurons holds great promise to model neuronal development and disease. The efficiency and success rate of neuronal reprogramming, however, may vary between different conversion platforms and cell types, thereby necessitating an unbiased, systematic approach to estimate neuronal identity of converted cells. Recent studies have demonstrated that long genes (>100 kb from transcription start to end) are highly enriched in neurons, which provides an opportunity to identify neurons based on the expression of these long genes. Results: We have developed a versatile R package, LONGO, to analyze gene expression based on gene length. We propose a systematic analysis of long gene expression (LGE) with a metric termed the long gene quotient (LQ) that quantifies LGE in RNA-seq or microarray data to validate neuronal identity at the single-cell and population levels. This unique feature of neurons provides an opportunity to utilize measurements of LGE in transcriptome data to quickly and easily distinguish neurons from non-neuronal cells. By combining this conceptual advancement and statistical tool in a user-friendly and interactive software package, we intend to encourage and simplify further investigation into LGE, particularly as it applies to validating and improving neuronal differentiation and reprogramming methodologies. Availability and implementation: LONGO is freely available for download at https://github.com/biohpc/longo. Supplementary information: Supplementary data are available at Bioinformatics online.
Matthew J. McCoy, Alexander J. Paul, Matheus B. Victor, Michelle Richner, Harrison W. Gabel, Haijun Gong, Andrew S. Yoo, Tae-Hyuk Ahn
Bioinform.8
2015 Sigma: Strain-level inference of genomes from metagenomic analysis for biosurveillance
abstract
MOTIVATION: Metagenomic sequencing of clinical samples provides a promising technique for direct pathogen detection and characterization in biosurveillance. Taxonomic analysis at the strain level can be used to resolve serotypes of a pathogen in biosurveillance. Sigma was developed for strain-level identification and quantification of pathogens using their reference genomes based on metagenomic analysis. RESULTS: Sigma provides not only accurate strain-level inferences, but also three unique capabilities: (i) Sigma quantifies the statistical uncertainty of its inferences, which includes hypothesis testing of identified genomes and confidence interval estimation of their relative abundances; (ii) Sigma enables strain variant calling by assigning metagenomic reads to their most likely reference genomes; and (iii) Sigma supports parallel computing for fast analysis of large datasets. The algorithm performance was evaluated using simulated mock communities and fecal samples with spike-in pathogen strains. AVAILABILITY AND IMPLEMENTATION: Sigma was implemented in C++ with source codes and binaries freely available at http://sigma.omicsbio.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tae-Hyuk Ahn, Juanjuan Chai, Chongle Pan
Bioinform.1
2014 Omega: an Overlap-graph de novo Assembler for Metagenomics
abstract
MOTIVATION: Metagenomic sequencing allows reconstruction of microbial genomes directly from environmental samples. Omega (overlap-graph metagenome assembler) was developed for assembling and scaffolding Illumina sequencing data of microbial communities. RESULTS: Omega found overlaps between reads using a prefix/suffix hash table. The overlap graph of reads was simplified by removing transitive edges and trimming short branches. Unitigs were generated based on minimum cost flow analysis of the overlap graph and then merged to contigs and scaffolds using mate-pair information. In comparison with three de Bruijn graph assemblers (SOAPdenovo, IDBA-UD and MetaVelvet), Omega provided comparable overall performance on a HiSeq 100-bp dataset and superior performance on a MiSeq 300-bp dataset. In comparison with Celera on the MiSeq dataset, Omega provided more continuous assemblies overall using a fraction of the computing time of existing overlap-layout-consensus assemblers. This indicates Omega can more efficiently assemble longer Illumina reads, and at deeper coverage, for metagenomic datasets. AVAILABILITY AND IMPLEMENTATION: Implemented in C++ with source code and binaries freely available at http://omega.omicsbio.org.
Bahlul Haider, Tae-Hyuk Ahn, Brian Bushnell, Juanjuan Chai, Alex Copeland, Chongle Pan
Bioinform.2
2013 Sipros/ProRata: a versatile informatics system for quantitative community proteomics
abstract
SUMMARY: Sipros/ProRata is an open-source software package for end-to-end data analysis in a wide variety of community proteomics measurements. A database-searching program, Sipros 3.0, was developed for accurate general-purpose protein identification and broad-range post-translational modification searches. Hybrid Message Passing Interface/OpenMP parallelism of the new Sipros architecture allowed its computation to be scalable from desktops to supercomputers. The upgraded ProRata 3.0 performs label-free quantification and isobaric chemical labeling quantification in addition to metabolic labeling quantification. Sipros/ProRata is a versatile informatics system that enables identification and quantification of proteins and their variants in many types of community proteomics studies. AVAILABILITY: Both programs are freely available under the GNU GPL license at Sipros.omicsbio.org and ProRata.omicsbio.org.
Yingfeng Wang, Tae-Hyuk Ahn, Chongle Pan
Bioinform.2
2011 Evaluating Performance Optimizations of Large-scale Genomic Sequence Search Applications using SST/macro
Tae-Hyuk Ahn, Damian Dechev, Heshan Lin, Helgi Adalsteinsson, Curtis L. Janssen
SIMULTECH1