Seunghak Lee

dblp:22/4961 · DBLP profile ↗
← Back
23ranked-venue papers
13as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 5 first-authorSystems, architecture and hardware · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 S-Tiering: A Unified HW/SW Solution for Memory Tiering Based on the Standard CXL Hotness Monitoring Unit
abstract
In this paper, we proposeS-Tiering, a unified hardware and software solution for memory tiering based on the standard CXL Hotness Monitoring Unit (CHMU).S-Tieringconsists of hardware components that comply with CHMU hardware specification defined in the CXL 3.2 Specification, and software components that control hardware components. Based on these components,S-Tieringminimizes the access to CXL memory by properly steering page migration.We evaluate various probabilistic data structure algorithms and adopt a Count-Min Sketch-based Hot Page Tracker that achieves 99% accuracy with only 0.3% tracking buffer overhead compared to assigning a dedicated counter for every 4KB page. We implement hardware components ofS-Tieringon an Field-Programmable Gate Array board and software components ofS-Tieringon Ubuntu 22.04 with Linux-v6.8 kernel. We evaluate the performance impact ofS-Tieringon benchmarks representative of real applications (e.g., High Performance Computing, Graph-processing, In-Memory Database).S-Tieringachieves a performance improvement of up to 193% compared to first-touch allocation and outperforms AutoNUMA memory tiering by 184%p.S-Tieringminimizes the memory access to CXL memory and increases the bandwidth utilization of DDR memory up to ×11.
Seunghak Lee, Wonjae Lee 0001, Hojin Nam, Jehoon Park, Youngshin Park, Junhyeok Im, Jinin So, Raghu Vamsi Krishna Talanki, Praful Ramesh O, Rajeev Verma, Taeksang Song, Wonhwa Shin, Sangjoon Hwang 0001
IEEE Trans. Computers1
2025 Beyond Page Migration: Enhancing Tiered Memory Performance via Integrated Last-Level Cache Management and Page Migration
Hwanjun Lee, Yeji Jung, Seonmu Oh, Ki-Dong Kang, Seunghak Lee, Daehoon Kim 0001
MICRO6
2025 TSB-AutoAD: Towards Automated Solutions for Time-Series Anomaly Detection [E, A & B]
abstract
Despite decades of research on time-series anomaly detection, the effectiveness of existing anomaly detectors remains constrained to specific domains - a model that performs well on one dataset may fail on another. Consequently, developing automated solutions for anomaly detection remains a pressing challenge. However, the AutoML community has predominantly focused on supervised learning solutions, which are impractical for anomaly detection due to the lack of labeled data and the absence of a well-defined objective function for model evaluation. While recent studies have evaluated standalone anomaly detectors, no study has ever evaluated automated solutions for selecting or generating scores in an automated manner. In this study, we (i) provide a systematic review and taxonomy of automated solutions for time-series anomaly detection, categorizing them into selection, ensembling, and generation methods; (ii) introduce TSB-AutoAD, a comprehensive benchmark encompassing 20 standalone methods and 70 variants; and (iii) conduct the most extensive evaluation in this area to date. Our benchmark includes state-of-the-art methods across all three categories, evaluated on TSB-AD, a recently curated heterogeneous testbed from nine domains. Our findings reveal a significant gap, where over half of the existing solutions do not statistically outperform a simple random choice. Foundation models that claim to offer generalized, one-size-fits-all solutions have yet to deliver on this promise. While naive ensembling achieves high accuracy, it comes at a substantial computational overhead. Conversely, methods leveraging historical datasets enable fast inference but suffer under out-of-distribution conditions. To address this trade-off, we propose a selective ensembling solution, which combines model selection with ensembling to offer a lightweight, practical balance between accuracy and efficiency. We open-source TSB-AutoAD and highlight the need for more robust and efficient solutions.
Seunghak Lee, John Paparrizos
Proc. VLDB Endow.2
2025 EasyAD: A Demonstration of Automated Solutions for Time-Series Anomaly Detection
abstract
Despite the recent focus on time-series anomaly detection, the effectiveness of the proposed anomaly detectors is restricted to specific domains. A model that performs well on one dataset may not perform well on another. Therefore, how to develop automated solutions for anomaly detection for a particular dataset has emerged as a pressing issue. However, there is a noticeable gap in the literature regarding providing a comprehensive review of the ongoing efforts toward automated solutions for selecting or generating scores in an automated manner. Conducting a meta-analysis of proposed methods is challenging due to: (i) their evaluation across limited datasets; (ii) different assumptions on application scenarios; and (iii) the absence of evaluations for out-of-distribution performance. Motivated by the limitations above, we introduce the EasyAD, a modular web engine designed to facilitate the exploration of the first comprehensive benchmark for automated time-series anomaly detection. The EasyAD engine enables rigorous statistical analysis of 20 automated methods and 70 of their variants across the TSB-AD benchmark, a recently curated, heterogeneous dataset spanning nine application domains. The engine supports a two-dimensional evaluation framework, incorporating both accuracy and runtime performance. Our engine allows users to assess the performance of various methods per dataset and per instance, which offers fine-grained analysis per time series. Furthermore, the engine accommodates the processing of user-uploaded data, enabling users to experiment with different model selection strategies on their own datasets.
Seunghak Lee, John Paparrizos
Proc. VLDB Endow.2
2021 GreenDIMM: OS-assisted DRAM Power Management for DRAM with a Sub-array Granularity Power-Down State
abstract
Power and energy consumed by DRAM comprising main memory of data-center servers have increased substantially as the capacity and bandwidth of memory increase. Especially, the fraction of DRAM background power in DRAM total power is already high, and it will continue to increase with the decelerating DRAM technology scaling as we will have to plug more DRAM modules in servers or stack more DRAM dies in a DRAM package to provide necessary DRAM capacity in the future. To reduce the background power, we may exploit low average utilization of the DRAM capacity in data-center servers (i.e., 40–60%) for DRAM power management. Nonetheless, the current DRAM power management supports low-power states only at the rank granularity, which becomes ineffective with memory interleaving techniques devised to disperse memory requests across ranks. That is, ranks need to be frequently woken up from low-power states with aggressive power management, which can significantly degrade system performance, or they do not get a chance to enter low-power states with conservative power management.
Seunghak Lee, Ki-Dong Kang, Hwanjun Lee, Hyungwon Park 0001, Young Hoon Son, Nam Sung Kim, Daehoon Kim 0001
MICRO1
2018 Ensembles of Lasso Screening Rules
abstract
In order to solve large-scale lasso problems, screening algorithms have been developed that discard features with zero coefficients based on a computationally efficient screening rule. Most existing screening rules were developed from a spherical constraint and half-space constraints on a dual optimal solution. However, existing rules admit at most two half-space constraints due to the computational cost incurred by the half-spaces, even though additional constraints may be useful to discard more features. In this paper, we present AdaScreen, an adaptive lasso screening rule ensemble, which allows to combine any one sphere with multiple half-space constraints on a dual optimal solution. Thanks to geometrical considerations that lead to a simple closed form solution for AdaScreen, we can incorporate multiple half-space constraints at small computational cost. In our experiments, we show that AdaScreen with multiple half-space constraints simultaneously improves screening performance and speeds up lasso solvers.
Seunghak Lee, Nico Görnitz, Eric P. Xing, David Heckerman, Christoph Lippert
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 STRADS: a distributed framework for scheduled model parallel machine learning
abstract
Machine learning (ML) algorithms are commonly applied to big data, using distributed systems that partition the data across machines and allow each machine to read and update all ML model parameters --- a strategy known as data parallelism. An alternative and complimentary strategy, model parallelism, partitions the model parameters for non-shared parallel access and updates, and may periodically repartition the parameters to facilitate communication. Model parallelism is motivated by two challenges that data-parallelism does not usually address: (1) parameters may be dependent, thus naive concurrent updates can introduce errors that slow convergence or even cause algorithm failure; (2) model parameters converge at different rates, thus a small subset of parameters can bottleneck ML algorithm completion. We propose scheduled model parallelism (SchMP), a programming approach that improves ML algorithm convergence speed by efficiently scheduling parameter updates, taking into account parameter dependencies and uneven convergence. To support SchMP at scale, we develop a distributed framework STRADS which optimizes the throughput of SchMP programs, and benchmark four common ML applications written as SchMP programs: LDA topic modeling, matrix factorization, sparse least-squares (Lasso) regression and sparse logistic regression. By improving ML progress per iteration through SchMP programming whilst improving iteration throughput through STRADS we show that SchMP programs running on STRADS outperform non-model-parallel ML implementations: for example, SchMP LDA and SchMP Lasso respectively achieve 10x and 5x faster convergence than recent, well-established baselines.
Jin Kyu Kim, Qirong Ho, Seunghak Lee, Xun Zheng, Wei Dai 0003, Garth A. Gibson, Eric P. Xing
EuroSys3
2016 A network-driven approach for genome-wide association mapping
abstract
MOTIVATION: It remains a challenge to detect associations between genotypes and phenotypes because of insufficient sample sizes and complex underlying mechanisms involved in associations. Fortunately, it is becoming more feasible to obtain gene expression data in addition to genotypes and phenotypes, giving us new opportunities to detect true genotype-phenotype associations while unveiling their association mechanisms. RESULTS: In this article, we propose a novel method, NETAM, that accurately detects associations between SNPs and phenotypes, as well as gene traits involved in such associations. We take a network-driven approach: NETAM first constructs an association network, where nodes represent SNPs, gene traits or phenotypes, and edges represent the strength of association between two nodes. NETAM assigns a score to each path from an SNP to a phenotype, and then identifies significant paths based on the scores. In our simulation study, we show that NETAM finds significantly more phenotype-associated SNPs than traditional genotype-phenotype association analysis under false positive control, taking advantage of gene expression data. Furthermore, we applied NETAM on late-onset Alzheimer's disease data and identified 477 significant path associations, among which we analyzed paths related to beta-amyloid, estrogen, and nicotine pathways. We also provide hypothetical biological pathways to explain our findings. AVAILABILITY AND IMPLEMENTATION: Software is available at http://www.sailing.cs.cmu.edu/ CONTACT: : [email protected].
Seunghak Lee, Soonho Kong, Eric P. Xing
Bioinform.1
2015 Petuum: A New Platform for Distributed Machine Learning on Big Data
abstract
How can one build a distributed framework that allows efficient deployment of a wide spectrum of modern advanced machine learning (ML) programs for industrial-scale problems using Big Models (100s of billions of parameters) on Big Data (terabytes or petabytes)- Contemporary parallelization strategies employ fine-grained operations and scheduling beyond the classic bulk-synchronous processing paradigm popularized by MapReduce, or even specialized operators relying on graphical representations of ML programs. The variety of approaches tends to pull systems and algorithms design in different directions, and it remains difficult to find a universal platform applicable to a wide range of different ML programs at scale. We propose a general-purpose framework that systematically addresses data- and model-parallel challenges in large-scale ML, by leveraging several fundamental properties underlying ML programs that make them different from conventional operation-centric programs: error tolerance, dynamic structure, and nonuniform convergence; all stem from the optimization-centric nature shared in ML programs' mathematical definitions, and the iterative-convergent behavior of their algorithmic solutions. These properties present unique opportunities for an integrative system design, built on bounded-latency network synchronization and dynamic load-balancing scheduling, which is efficient, programmable, and enjoys provable correctness guarantees. We demonstrate how such a design in light of ML-first principles leads to significant performance improvements versus well-known implementations of several ML programs, allowing them to run in much less time and at considerably larger model sizes, on modestly-sized computer clusters.
Eric P. Xing, Qirong Ho, Wei Dai 0003, Jin Kyu Kim, Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, Yaoliang Yu
KDD6
2015 An Efficient Nonlinear Regression Approach for Genome-Wide Detection of Marginal and Interacting Genetic Variations
Seunghak Lee, Aurélie C. Lozano, Prabhanjan Kambadur, Eric P. Xing
RECOMB1
2015 Petuum: A New Platform for Distributed Machine Learning on Big Data
abstract
What is a systematic way to efficiently apply a wide spectrum of advanced ML programs to industrial scale problems, using Big Models (up to 100 s of billions of parameters) on Big Data (up to terabytes or petabytes)? Modern parallelization strategies employ fine-grained operations and scheduling beyond the classic bulk-synchronous processing paradigm popularized by MapReduce, or even specialized graph-based execution that relies on graph representations of ML programs. The variety of approaches tends to pull systems and algorithms design in different directions, and it remains difficult to find a universal platform applicable to a wide range of ML programs at scale. We propose a general-purpose framework, Petuum, that systematically addresses data- and model-parallel challenges in large-scale ML, by observing that many ML programs are fundamentally optimization-centric and admit error-tolerant, iterative-convergent algorithmic solutions. This presents unique opportunities for an integrative system design, such as bounded-error network synchronization and dynamic scheduling based on ML program structure. We demonstrate the efficacy of these system designs versus well-known implementations of modern ML algorithms, showing that Petuum allows ML programs to run in much less time and at considerably larger model sizes, even on modestly-sized compute clusters.
Eric P. Xing, Qirong Ho, Wei Dai 0003, Jin Kyu Kim, Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, Yaoliang Yu
IEEE Trans. Big Data6
2014 On Model Parallelization and Scheduling Strategies for Distributed Machine Learning
Seunghak Lee, Jin Kyu Kim, Xun Zheng, Qirong Ho, Garth A. Gibson, Eric P. Xing
NIPS1
2014 Exploiting Bounded Staleness to Speed Up Big Data Analytics
Henggang Cui, James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Abhimanu Kumar, Jinliang Wei, Wei Dai 0003, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, Eric P. Xing
USENIX ATC5
2013 Solving the Straggler Problem with Bounded Staleness
James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Gregory R. Ganger, Garth A. Gibson, Kimberly Keeton, Eric P. Xing
HotOS4
2013 More Effective Distributed ML via a Stale Synchronous Parallel Parameter Server
abstract
We propose a parameter server system for distributed ML, which follows a Stale Synchronous Parallel (SSP) model of computation that maximizes the time computational workers spend doing useful work on ML algorithms, while still providing correctness guarantees. The parameter server provides an easy-to-use shared interface for read/write access to an ML model's values (parameters and variables), and the SSP model allows distributed workers to read older, stale versions of these values from a local cache, instead of waiting to get them from a central storage. This significantly increases the proportion of time workers spend computing, as opposed to waiting. Furthermore, the SSP model ensures ML algorithm correctness by limiting the maximum age of the stale values. We provide a proof of correctness under SSP, as well as empirical results demonstrating that the SSP model achieves faster algorithm convergence on several different ML problems, compared to fully-synchronous and asynchronous schemes.
Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B. Gibbons, Garth A. Gibson, Gregory R. Ganger, Eric P. Xing
NIPS4
2012 Leveraging input and output structures for joint mapping of epistatic and marginal eQTLs
abstract
MOTIVATION: As many complex disease and expression phenotypes are the outcome of intricate perturbation of molecular networks underlying gene regulation resulted from interdependent genome variations, association mapping of causal QTLs or expression quantitative trait loci must consider both additive and epistatic effects of multiple candidate genotypes. This problem poses a significant challenge to contemporary genome-wide-association (GWA) mapping technologies because of its computational complexity. Fortunately, a plethora of recent developments in biological network community, especially the availability of genetic interaction networks, make it possible to construct informative priors of complex interactions between genotypes, which can substantially reduce the complexity and increase the statistical power of GWA inference. RESULTS: In this article, we consider the problem of learning a multitask regression model while taking advantage of the prior information on structures on both the inputs (genetic variations) and outputs (expression levels). We propose a novel regularization scheme over multitask regression called jointly structured input-output lasso based on an ℓ(1)/ℓ(2) norm, which allows shared sparsity patterns for related inputs and outputs to be optimally estimated. Such patterns capture multiple related single nucleotide polymorphisms (SNPs) that jointly influence multiple-related expression traits. In addition, we generalize this new multitask regression to structurally regularized polynomial regression to detect epistatic interactions with manageable complexity by exploiting the prior knowledge on candidate SNPs for epistatic effects from biological experiments. We demonstrate our method on simulated and yeast eQTL datasets. AVAILABILITY: Software is available at http://www.sailing.cs.cmu.edu/.
Seunghak Lee, Eric P. Xing
Bioinform.1
2010 Adaptive Multi-Task Lasso: with Application to eQTL Detection
abstract
To understand the relationship between genomic variations among population and complex diseases, it is essential to detect eQTLs which are associated with phenotypic effects. However, detecting eQTLs remains a challenge due to complex underlying mechanisms and the very large number of genetic loci involved compared to the number of samples. Thus, to address the problem, it is desirable to take advantage of the structure of the data and prior information about genomic locations such as conservation scores and transcription factor binding sites. In this paper, we propose a novel regularized regression approach for detecting eQTLs which takes into account related traits simultaneously while incorporating many regulatory features. We first present a Bayesian network for a multi-task learning problem that includes priors on SNPs, making it possible to estimate the significance of each covariate adaptively. Then we find the maximum a posteriori (MAP) estimation of regression coefficients and estimate weights of covariates jointly. This optimization procedure is efficient since it can be achieved by using convex optimization and a coordinate descent procedure iteratively. Experimental results on simulated and real yeast datasets confirm that our model outperforms previous methods for finding eQTLs.
Seunghak Lee, Jun Zhu 0001, Eric P. Xing
NIPS1
2010 MoGUL: Detecting Common Insertions and Deletions in a Population
Seunghak Lee, Eric P. Xing, Michael Brudno
RECOMB1
2009 Ensembles of landmark multidimensional scalings
abstract
Landmark multidimensional scaling (LMDS) uses a subset of data (landmark points) to solve classical MDS, where the scalability is increased but the approximation is noise-sensitive. In this paper we present an ensemble of LMDSs, referred to as landmark MDS ensemble (LMDSE), where we use a portion of the input in a piecewise manner to solve classical MDS, combining individual LMDS solutions which operate on different partitions of the input. Ground control points (GCPs) that are shared by partitions considered in the ensemble, allow us to align individual LMDS solutions in a common coordinate system through affine transformations. LMDSE solution is determined by averaging aligned LMDS solutions. We show that LMDSE is less noise-sensitive while maintaining the scalability as well as the speed of LMDS. Experiments on synthetic data (noisy grid) and real-world data (similar image retrieval) confirm the high performance of the proposed LMDSE.
Seunghak Lee, Seungjin Choi 0001
ICASSP1
2009 Landmark MDS ensemble
Seunghak Lee, Seungjin Choi 0001
Pattern Recognit.1
2008 A robust framework for detecting structural variations in a genome
abstract
MOTIVATION: Recently, structural genomic variants have come to the forefront as a significant source of variation in the human population, but the identification of these variants in a large genome remains a challenge. The complete sequencing of a human individual is prohibitive at current costs, while current polymorphism detection technologies, such as SNP arrays, are not able to identify many of the large scale events. One of the most promising methods to detect such variants is the computational mapping of clone-end sequences to a reference genome. RESULTS: Here, we present a probabilistic framework for the identification of structural variants using clone-end sequencing. Unlike previous methods, our approach does not rely on an a priori determined mapping of all reads to the reference. Instead, we build a framework for finding the most probable assignment of sequenced clones to potential structural variants based on the other clones. We compare our predictions with the structural variants identified in three previous studies. While there is a statistically significant correlation between the predictions, we also find a significant number of previously uncharacterized structural variants. Furthermore, we identify a number of putative cross-chromosomal events, primarily located proximally to the centromeres of the chromosomes. AVAILABILITY: Our dataset, results and source code are available at http://compbio.cs.toronto.edu/structvar/.
Seunghak Lee, Elango Cheran, Michael Brudno
ISMB1
2008 Efficient service discovery mechanism for wireless sensor networks
Seunghak Lee, Namgi Kim, Hyunsoo Yoon
Comput. Commun.2
2007 Dynamically Weighted Hidden Markov Model for Spam Deobfuscation
Seunghak Lee, Iryoung Jeong, Seungjin Choi 0001
IJCAI1