Byunghan Lee

dblp:124/7080 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-6727-0975ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Time series and sequential data · 39% Representation and self-supervised learning · 39% Trustworthy machine learning · 15%
Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
1 paper
Data mining · 87% Recommender systems · 13%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Time series and sequential data
anomaly detection
0.812024
Contrastive Time-Series Anomaly Detection · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
Contrastive Time-Series Anomaly Detection · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Time series and sequential data › time series analysis
time series anomaly detection
0.812024
Contrastive Time-Series Anomaly Detection · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Representation and self-supervised learning › contrastive learning › temporal contrastive learning
time series contrastive learning
0.812024
Contrastive Time-Series Anomaly Detection · IEEE Trans. Knowl. Data Eng. 2024
Bioinformatics and computational biology
metagenomics
0.622021
MUGAN: multi-GPU accelerated AmpliconNoise server for rapid microbial diversity assessment · Bioinform. 2021
Rapid and Robust Denoising of Pyrosequenced Amplicons for Metagenomics · ICDM 2012
Bioinformatics and computational biology
gene regulation
0.612022
TargetNet: functional microRNA target prediction with deep neural networks · Bioinform. 2022
Bioinformatics and computational biology › gene regulation
microRNA target prediction
0.612022
TargetNet: functional microRNA target prediction with deep neural networks · Bioinform. 2022
Data mining
anomaly detection
0.612022
Towards a Rigorous Evaluation of Time-Series Anomaly Detection · AAAI 2022
Data mining › anomaly detection
time series anomaly detection
0.612022
Towards a Rigorous Evaluation of Time-Series Anomaly Detection · AAAI 2022
Bioinformatics and computational biology › transcriptomics › non-coding RNA analysis
long non-coding RNA identification
0.312018
LncRNAnet: long non-coding RNA identification using deep learning · Bioinform. 2018
Bioinformatics and computational biology › sequence analysis
non-coding RNA identification
0.312018
LncRNAnet: long non-coding RNA identification using deep learning · Bioinform. 2018
Image and video processing › image restoration
denoising
0.212016
Neural Universal Discrete Denoiser · NIPS 2016
Recommender systems › recommender system evaluation
evaluation protocol
0.212022
Towards a Rigorous Evaluation of Time-Series Anomaly Detection · AAAI 2022
GPUs and heterogeneous computing
multi-GPU computing
0.112021
MUGAN: multi-GPU accelerated AmpliconNoise server for rapid microbial diversity assessment · Bioinform. 2021
Bioinformatics and computational biology › metagenomics
operational taxonomic unit clustering
0.112012
Rapid and Robust Denoising of Pyrosequenced Amplicons for Metagenomics · ICDM 2012
Bioinformatics and computational biology › sequence analysis
RNA sequence analysis
0.112018
LncRNAnet: long non-coding RNA identification using deep learning · Bioinform. 2018

Methods — techniques the papers use, named apart from their topics

point adjustment protocol · 1.1data-level parallelism · 1.0GPU co-processing · 1.0one-class learning · 0.8data augmentation · 0.8contrastive learning · 0.8LSTM · 0.8sequence alignment · 0.6k-mer encoding · 0.6deep residual network · 0.6pseudo-labeling · 0.5multicore CPU · 0.5multi-core CPU · 0.5deep neural network · 0.5recurrent neural network · 0.3deep learning · 0.3convolutional neural network · 0.3density-based clustering · 0.1
YearPublicationVenuePosition
2026 SPreT: Style-guided prefix tuning for efficient text style transfer
Minseo Kang, Seunggi Shin, Byunghan Lee
Neurocomputing3
2024 Contrastive Time-Series Anomaly Detection
abstract
In addition to its success in representation learning, contrastive learning is effective in image anomaly detection. Although contrastive learning depends significantly on data augmentation methods, time-series data augmentation for time-series anomaly detection is not investigated sufficiently. Additionally, although time-series data share a temporal context, the existing contrastive loss contrasts temporally related samples, in which deteriorated anomaly detection performance is observed on time-series data. Herein, we propose contrastive multivariate time-series anomaly detection (CTAD), a multivariate time-series anomaly detection framework that addresses these challenges by incorporating a one-class learning scheme into the contrastive loss based on meticulously designed time-series data augmentations. Specifically, we propose seven types of general time-series data augmentations to be applied variable- and point- wise, and provide guidance on data augmentation methods for contrastive time-series anomaly detection. The superiority of the one-class contrastive loss and the appropriate selection of time-series data augmentation allow CTAD to achieve outstanding performance in multiple datasets, even using a simple long short-term memory network. Furthermore, CTAD is robust to noise as it trains a noise-invariant network. This enables up to 47× faster and 20× more memory-efficient anomaly detection performance compared with existing methods while affording robustness, which are essential considerations in real-world applications.
HyunGi Kim, Siwon Kim, Seonwoo Min, Byunghan Lee
IEEE Trans. Knowl. Data Eng.4
2022 Towards a Rigorous Evaluation of Time-Series Anomaly Detection
abstract
In recent years, proposed studies on time-series anomaly detection (TAD) report high F1 scores on benchmark TAD datasets, giving the impression of clear improvements in TAD. However, most studies apply a peculiar evaluation protocol called point adjustment (PA) before scoring. In this paper, we theoretically and experimentally reveal that the PA protocol has a great possibility of overestimating the detection performance; even a random anomaly score can easily turn into a state-of-the-art TAD method. Therefore, the comparison of TAD methods after applying the PA protocol can lead to misguided rankings. Furthermore, we question the potential of existing TAD methods by showing that an untrained model obtains comparable detection performance to the existing methods even when PA is forbidden. Based on our findings, we propose a new baseline and an evaluation protocol. We expect that our study will help a rigorous evaluation of TAD and lead to further improvement in future researches.
Siwon Kim, Kukjin Choi, Hyun-Soo Choi, Byunghan Lee, Sungroh Yoon
AAAI4
2022 TargetNet: functional microRNA target prediction with deep neural networks
abstract
MOTIVATION: MicroRNAs (miRNAs) play pivotal roles in gene expression regulation by binding to target sites of messenger RNAs (mRNAs). While identifying functional targets of miRNAs is of utmost importance, their prediction remains a great challenge. Previous computational algorithms have major limitations. They use conservative candidate target site (CTS) selection criteria mainly focusing on canonical site types, rely on laborious and time-consuming manual feature extraction, and do not fully capitalize on the information underlying miRNA-CTS interactions. RESULTS: In this article, we introduce TargetNet, a novel deep learning-based algorithm for functional miRNA target prediction. To address the limitations of previous approaches, TargetNet has three key components: (i) relaxed CTS selection criteria accommodating irregularities in the seed region, (ii) a novel miRNA-CTS sequence encoding scheme incorporating extended seed region alignments and (iii) a deep residual network-based prediction model. The proposed model was trained with miRNA-CTS pair datasets and evaluated with miRNA-mRNA pair datasets. TargetNet advances the previous state-of-the-art algorithms used in functional miRNA target classification. Furthermore, it demonstrates great potential for distinguishing high-functional miRNA targets. AVAILABILITY AND IMPLEMENTATION: The codes and pre-trained models are available at https://github.com/mswzeus/TargetNet.
Seonwoo Min, Byunghan Lee, Sungroh Yoon
Bioinform.2
2021 MUGAN: multi-GPU accelerated AmpliconNoise server for rapid microbial diversity assessment
abstract
MOTIVATION: Metagenomic sequencing has become a crucial tool for obtaining a gene catalogue of operational taxonomic units (OTUs) in a microbial community. A typical metagenomic sequencing produces a large amount of data (often in the order of terabytes or more), and computational tools are indispensable for efficient processing. In particular, error correction in metagenomics is crucial for accurate and robust genetic cataloging of microbial communities. However, many existing error-correction tools take a prohibitively long time and often bottleneck the whole analysis pipeline. RESULTS: To overcome this computational hurdle, we analyzed and exploited the data-level parallelism that exists in the error-correction procedure and proposed a tool named MUGAN that exploits both multi-core central processing units and multiple graphics processing units for co-processing. According to the experimental results, our approach reduced not only the time demand for denoising amplicons from approximately 59 h to only 46 min, but also the overestimation of the number of OTUs, estimating 6.7 times less species-level OTUs than the baseline. In addition, our approach provides web-based intuitive visualization of results. Given its efficiency and convenience, we anticipate that our approach would greatly facilitate denoising efforts in metagenomics studies. AVAILABILITY AND IMPLEMENTATION: http://data.snu.ac.kr/pub/mugan. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Byunghan Lee, Hyeyoung Min, Sungroh Yoon
Bioinform.1
2018 LncRNAnet: long non-coding RNA identification using deep learning
abstract
Motivation: Long non-coding RNAs (lncRNAs) are important regulatory elements in biological processes. LncRNAs share similar sequence characteristics with messenger RNAs, but they play completely different roles, thus providing novel insights for biological studies. The development of next-generation sequencing has helped in the discovery of lncRNA transcripts. However, the experimental verification of numerous transcriptomes is time consuming and costly. To alleviate these issues, a computational approach is needed to distinguish lncRNAs from the transcriptomes. Results: We present a deep learning-based approach, lncRNAnet, to identify lncRNAs that incorporates recurrent neural networks for RNA sequence modeling and convolutional neural networks for detecting stop codons to obtain an open reading frame indicator. lncRNAnet performed clearly better than the other tools for sequences of short lengths, on which most lncRNAs are distributed. In addition, lncRNAnet successfully learned features and showed 7.83%, 5.76%, 5.30% and 3.78% improvements over the alternatives on a human test set in terms of specificity, accuracy, F1-score and area under the curve, respectively. Availability and implementation: Data and codes are available in http://data.snu.ac.kr/pub/lncRNAnet.
Junghwan Baek, Byunghan Lee, Sunyoung Kwon, Sungroh Yoon
Bioinform.2
2018 hc-OTU: A Fast and Accurate Method for Clustering Operational Taxonomic Units Based on Homopolymer Compaction
abstract
To assess the genetic diversity of an environmental sample in metagenomics studies, the amplicon sequences of 16s rRNA genes need to be clustered into operational taxonomic units (OTUs). Many existing tools for OTU clustering trade off between accuracy and computational efficiency. We propose a novel OTU clustering algorithm, hc-OTU, which achieves high accuracy and fast runtime by exploiting homopolymer compaction and k-mer profiling to significantly reduce the computing time for pairwise distances of amplicon sequences. We compare the proposed method with other widely used methods, including UCLUST, CD-HIT, MOTHUR, ESPRIT, ESPRIT-TREE, and CLUSTOM, comprehensively, using nine different experimental datasets and many evaluation metrics, such as normalized mutual information, adjusted Rand index, measure of concordance, and F-score. Our evaluation reveals that the proposed method achieves a level of accuracy comparable to the respective accuracy levels of MOTHUR and ESPRIT-TREE, two widely used OTU clustering methods, while delivering orders-of-magnitude speedups.
Seunghyun Park 0001, Hyun-Soo Choi, Byunghan Lee, Jongsik Chun, Joong-Ho Won, Sungroh Yoon
IEEE ACM Trans. Comput. Biol. Bioinform.3
2017 Deep learning in bioinformatics
abstract
In the era of big data, transformation of biomedical big data into valuable knowledge has been one of the most important challenges in bioinformatics. Deep learning has advanced rapidly since the early 2000s and now demonstrates state-of-the-art performance in various fields. Accordingly, application of deep learning in bioinformatics to gain insight from data has been emphasized in both academia and industry. Here, we review deep learning in bioinformatics, presenting examples of current research. To provide a useful and comprehensive perspective, we categorize research both by the bioinformatics domain (i.e. omics, biomedical imaging, biomedical signal processing) and deep learning architecture (i.e. deep neural networks, convolutional neural networks, recurrent neural networks, emergent architectures) and present brief descriptions of each study. Additionally, we discuss theoretical and practical issues of deep learning in bioinformatics and suggest future research directions. We believe that this review will provide valuable insights and serve as a starting point for researchers to apply deep learning approaches in their bioinformatics studies.
Seonwoo Min, Byunghan Lee, Sungroh Yoon
Briefings Bioinform.2
2016 Neural Universal Discrete Denoiser
abstract
We present a new framework of applying deep neural networks (DNN) to devise a universal discrete denoiser. Unlike other approaches that utilize supervised learning for denoising, we do not require any additional training data. In such setting, while the ground-truth label, i.e., the clean data, is not available, we devise ``pseudo-labels'' and a novel objective function such that DNN can be trained in a same way as supervised learning to become a discrete denoiser. We experimentally show that our resulting algorithm, dubbed as Neural DUDE, significantly outperforms the previous state-of-the-art in several applications with a systematic rule of choosing the hyperparameter, which is an attractive feature in practice.
Taesup Moon, Seonwoo Min, Byunghan Lee, Sungroh Yoon
NIPS3
2014 CASPER: context-aware scheme for paired-end reads from high-throughput amplicon sequencing
abstract
Merging the forward and reverse reads from paired-end sequencing is a critical task that can significantly improve the performance of downstream tasks, such as genome assembly and mapping, by providing them with virtually elongated reads. However, due to the inherent limitations of most paired-end sequencers, the chance of observing erroneous bases grows rapidly as the end of a read is approached, which becomes a critical hurdle for accurately merging paired-end reads. Although there exist several sophisticated approaches to this problem, their performance in terms of quality of merging often remains unsatisfactory. To address this issue, here we present a context-aware scheme for paired-end reads (CASPER): a computational method to rapidly and robustly merge overlapping paired-end reads. Being particularly well suited to amplicon sequencing applications, CASPER is thoroughly tested with both simulated and real high-throughput amplicon sequencing data. According to our experimental results, CASPER significantly outperforms existing state-of-the art paired-end merging tools in terms of accuracy and robustness. CASPER also exploits the parallelism in the task of paired-end merging and effectively speeds up by multithreading. CASPER is freely available for academic use at http://best.snu.ac.kr/casper.
Sunyoung Kwon, Byunghan Lee, Sungroh Yoon
BMC Bioinform.2
2012 Rapid and Robust Denoising of Pyrosequenced Amplicons for Metagenomics
abstract
Metagenomic sequencing has become a crucial tool for obtaining a gene catalogue of operational taxonomic units (OTUs) in a microbial community. High-throughput pyrosequencing is a next-generation sequencing technique very popular in microbial community analysis due to its longer read length compared to alternative methods. Computational tools are inevitable to process raw data from pyrosequencers, and in particular, noise removal is a critical data-mining step to obtain robust sequence reads. However, the slow rate of existing denoisers has bottlenecked the whole pyrosequencing process, let alone hindering efforts to improve robustness. To address these, we propose a new approach that can accelerate the denoising process substantially. By using our approach, it now takes only about 2 hours to denoise 62,873 pyrosequenced amplicons from a mixture of 91 full-length 16S rRNA clones. It would otherwise take nearly 2.5 days if existing software tools were used. Furthermore, our approach can effectively reduce overestimating the number of OTUs, producing 6.7 times fewer species-level OTUs on average than a state-of-the-art alternative under the same condition. Leveraged by our approach, we hope that metagenomic sequencing will become an even more appealing tool for microbial community analysis.
Byunghan Lee, Joonhong Park, Sungroh Yoon
ICDM1