Giltae Song

dblp:38/3019 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0001-8796-4678ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Forecasting bitcoin prices based on hybrid LSTM and deep neural network architectures
Chae-Deug Yi, Giltae Song
Soft Comput.2
2025 Therapeutic gene target prediction using novel deep hypergraph representation learning
abstract
Identifying therapeutic genes is crucial for developing treatments targeting genetic causes of diseases, but experimental trials are costly and time-consuming. Although many deep learning approaches aim to identify biomarker genes, predicting therapeutic target genes remains challenging due to the limited number of known targets. To address this, we propose HIT (Hypergraph Interaction Transformer), a deep hypergraph representation learning model that identifies a gene's therapeutic potential, biomarker status, or lack of association with diseases. HIT uses hypergraph structures of genes, ontologies, diseases, and phenotypes, employing attention-based learning to capture complex relationships. Experiments demonstrate HIT's state-of-the-art performance, explainability, and ability to identify novel therapeutic targets.
Kibeom Kim, Juseong Kim, Minwook Kim, Giltae Song
Briefings Bioinform.5
2025 Gene spatial integration: enhancing spatial transcriptomics analysis via deep learning and batch effect mitigation
abstract
MOTIVATION: Spatial transcriptomics (ST) is a groundbreaking technique for studying the correlation between cellular organization within a tissue and its physiological and pathological properties. Every facet of spatial information, including cell/spot proximity, distribution, and dimensionality, is significant. Most methods lean heavily on proximity for ST analysis, each resulting in useful insights but still leaving other aspects untapped. In addition, samples procured at different times, by different donors, and by different technologies introduce a batch effects problem that hinders the statistical approach employed by most analysis tools. Addressing these challenges, we have developed a deep learning method for analyzing integrated multiple ST data, focusing on the distribution aspect. Furthermore, our method aims to leverage single-cell analysis tools. RESULTS: Our study introduces Gene Spatial Integration (GSI), a data integration pipeline utilizing a representation learning approach to extract the spatial distribution of genes into the same feature space as gene expression features. We employ an autoencoder network to extract spatial embedding, facilitating the projection of spatial features into gene expression feature space. Our approach allows for seamless integration of multiple samples with minimum detriment, increasing the performance of the ST data analysis tool. We show the application of our method on the human dorsolateral prefrontal cortex dataset. Our method consistently improves the performance of the clustering of Seurat tools, with the most significant increase observed in sample 151673, almost doubling the ARI score from 0.225 to 0.405. We also combine our pipeline with the clustering of GraphST, achieving a significantly higher ARI score in sample 151672 from 0.614 to 0.795. This result reveals the potential of gene distribution spatial aspect, also emphasizes the impact of integration and batch effect removal in developing a refined analysis in understanding tissue characteristics. AVAILABILITY AND IMPLEMENTATION: Implementation of GSI is accessible at https://github.com/Riandanis/Spatial_Integration_GSI.
Rian Pratama, Jason A. Hilton, J. Michael Cherry, Giltae Song
Bioinform.4
2024 Acute myocardial infarction prognosis prediction with reliable and interpretable artificial intelligence system
abstract
OBJECTIVE: Predicting mortality after acute myocardial infarction (AMI) is crucial for timely prescription and treatment of AMI patients, but there are no appropriate AI systems for clinicians. Our primary goal is to develop a reliable and interpretable AI system and provide some valuable insights regarding short, and long-term mortality. MATERIALS AND METHODS: We propose the RIAS framework, an end-to-end framework that is designed with reliability and interpretability at its core and automatically optimizes the given model. Using RIAS, clinicians get accurate and reliable predictions which can be used as likelihood, with global and local explanations, and "what if" scenarios to achieve desired outcomes as well. RESULTS: We apply RIAS to AMI prognosis prediction data which comes from the Korean Acute Myocardial Infarction Registry. We compared FT-Transformer with XGBoost and MLP and found that FT-Transformer has superiority in sensitivity and comparable performance in AUROC and F1 score to XGBoost. Furthermore, RIAS reveals the significance of statin-based medications, beta-blockers, and age on mortality regardless of time period. Lastly, we showcase reliable and interpretable results of RIAS with local explanations and counterfactual examples for several realistic scenarios. DISCUSSION: RIAS addresses the "black-box" issue in AI by providing both global and local explanations based on SHAP values and reliable predictions, interpretable as actual likelihoods. The system's "what if" counterfactual explanations enable clinicians to simulate patient-specific scenarios under various conditions, enhancing its practical utility. CONCLUSION: The proposed framework provides reliable and interpretable predictions along with counterfactual examples.
Minwook Kim, Donggil Kang, Min Sun Kim, Jeong Cheon Choe, Sun-Hack Lee, Jin-Hee Ahn, Jun-Hyok Oh, Jung Hyun Choi, Han Cheol Lee, Kwang Soo Cha, Kyungtae Jang, Woor I Bong, Giltae Song
J. Am. Medical Informatics Assoc.13
2023 DeepFold: enhancing protein structure prediction through optimized loss functions, improved template features, and re-optimized energy function
abstract
MOTIVATION: Predicting protein structures with high accuracy is a critical challenge for the broad community of life sciences and industry. Despite progress made by deep neural networks like AlphaFold2, there is a need for further improvements in the quality of detailed structures, such as side-chains, along with protein backbone structures. RESULTS: Building upon the successes of AlphaFold2, the modifications we made include changing the losses of side-chain torsion angles and frame aligned point error, adding loss functions for side chain confidence and secondary structure prediction, and replacing template feature generation with a new alignment method based on conditional random fields. We also performed re-optimization by conformational space annealing using a molecular mechanics energy function which integrates the potential energies obtained from distogram and side-chain prediction. In the CASP15 blind test for single protein and domain modeling (109 domains), DeepFold ranked fourth among 132 groups with improvements in the details of the structure in terms of backbone, side-chain, and Molprobity. In terms of protein backbone accuracy, DeepFold achieved a median GDT-TS score of 88.64 compared with 85.88 of AlphaFold2. For TBM-easy/hard targets, DeepFold ranked at the top based on Z-scores for GDT-TS. This shows its practical value to the structural biology community, which demands highly accurate structures. In addition, a thorough analysis of 55 domains from 39 targets with publicly available structures indicates that DeepFold shows superior side-chain accuracy and Molprobity scores among the top-performing groups. AVAILABILITY AND IMPLEMENTATION: DeepFold tools are open-source software available at https://github.com/newtonjoo/deepfold.
Jae-Won Lee, Jong-Hyun Won, Seonggwang Jeon, Yujin Choo, Yubin Yeon, Jin-Seon Oh, Seonhwa Kim, InSuk Joung, Cheongjae Jang, Sung Jong Lee, Kyong Hwan Jin, Giltae Song, Eun-Sol Kim, Jejoong Yoo, Eunok Paek, Yung-Kyun Noh, Keehyoung Joo
Bioinform.14
2023 AptaTrans: a deep neural network for predicting aptamer-protein interaction using pretrained encoders
abstract
BACKGROUND: Aptamers, which are biomaterials comprised of single-stranded DNA/RNA that form tertiary structures, have significant potential as next-generation materials, particularly for drug discovery. The systematic evolution of ligands by exponential enrichment (SELEX) method is a critical in vitro technique employed to identify aptamers that bind specifically to target proteins. While advanced SELEX-based methods such as Cell- and HT-SELEX are available, they often encounter issues such as extended time consumption and suboptimal accuracy. Several In silico aptamer discovery methods have been proposed to address these challenges. These methods are specifically designed to predict aptamer-protein interaction (API) using benchmark datasets. However, these methods often fail to consider the physicochemical interactions between aptamers and proteins within tertiary structures. RESULTS: In this study, we propose AptaTrans, a pipeline for predicting API using deep learning techniques. AptaTrans uses transformer-based encoders to handle aptamer and protein sequences at the monomer level. Furthermore, pretrained encoders are utilized for the structural representation. After validation with a benchmark dataset, AptaTrans has been integrated into a comprehensive toolset. This pipeline synergistically combines with Apta-MCTS, a generative algorithm for recommending aptamer candidates. CONCLUSION: The results show that AptaTrans outperforms existing models for predicting API, and the efficacy of the AptaTrans pipeline has been confirmed through various experimental tools. We expect AptaTrans will enhance the cost-effectiveness and efficiency of SELEX in drug discovery. The source code and benchmark dataset for AptaTrans are available at https://github.com/pnumlb/AptaTrans .
Incheol Shin, Keumseok Kang, Juseong Kim, Sanghun Sel, Jeonghoon Choi, Ho-Young Kang, Giltae Song
BMC Bioinform.8
2022 FastqCLS: a FASTQ compressor for long-read sequencing via read reordering using a novel scoring model
abstract
MOTIVATION: Over the past decades, vast amounts of genome sequencing data have been produced, requiring an enormous level of storage capacity. The time and resources needed to store and transfer such data cause bottlenecks in genome sequencing analysis. To resolve this issue, various compression techniques have been proposed to reduce the size of original FASTQ raw sequencing data, but these remain suboptimal. Long-read sequencing has become dominant in genomics, whereas most existing compression methods focus on short-read sequencing only. RESULTS: We designed a compression algorithm based on read reordering using a novel scoring model for reducing FASTQ file size with no information loss. We integrated all data processing steps into a software package called FastqCLS and provided it as a Docker image for ease of installation and execution to help users easily install and run. We compared our method with existing major FASTQ compression tools using benchmark datasets. We also included new long-read sequencing data in this validation. As a result, FastqCLS outperformed in terms of compression ratios for storing long-read sequencing data. AVAILABILITY AND IMPLEMENTATION: FastqCLS can be downloaded from https://github.com/krlucete/FastqCLS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dohyeon Lee, Giltae Song
Bioinform.2
2011 Evaluation of methods for detecting conversion events in gene clusters
abstract
BACKGROUND: Gene clusters are genetically important, but their analysis poses significant computational challenges. One of the major reasons for these difficulties is gene conversion among the duplicated regions of the cluster, which can obscure their true relationships. Many computational methods for detecting gene conversion events have been released, but their performance has not been assessed for wide deployment in evolutionary history studies due to a lack of accurate evaluation methods. RESULTS: We designed a new method that simulates gene cluster evolution, including large-scale events of duplication, deletion, and conversion as well as small mutations. We used this simulation data to evaluate several different programs for detecting gene conversion events. CONCLUSIONS: Our evaluation identifies strengths and weaknesses of several methods for detecting gene conversion, which can contribute to more accurate analysis of gene cluster evolution.
Giltae Song, Chih-Hao Hsu, Cathy Riemer, Webb Miller
BMC Bioinform.1
2008 Reconstructing the Evolutionary History of Complex Human Gene Clusters
Yu Zhang 0002, Giltae Song, Tomás Vinar, Eric D. Green, Adam C. Siepel, Webb Miller
RECOMB2