EDBT 2026 Demo / reviewers in the wild / expert
Sangsoo Park
dblp:41/2400
· DBLP profile ↗
17ranked-venue papers
9as first author
5since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language ModelsabstractTransformer-based large language models (LLMs) such as Generative Pre-trained Transformer (GPT) have become popular due to their remarkable performance across diverse applications, including text generation and translation. For LLM training and inference, the GPU has been the predominant accelerator with its pervasive software development ecosystem and powerful computing capability. However, as the size of LLMs keeps increasing for higher performance and/or more complex applications, a single GPU cannot efficiently accelerate LLM training and inference due to its limited memory capacity, which demands frequent transfers of the model parameters needed by the GPU to compute the current layer(s) from the host CPU memory/storage. A GPU appliance may provide enough aggregated memory capacity with multiple GPUs, but it suffers from frequent transfers of intermediate values among GPU devices, each accelerating specific layers of a given LLM. As the frequent transfers of these model parameters and intermediate values are performed over relatively slow device-to-device interconnects such as PCIe or NVLink, they become the key bottleneck for efficient acceleration of LLMs. Focusing on accelerating LLM inference, which is essential for many commercial services, we develop CXL-PNM, a processing near memory (PNM) platform based on the emerging interconnect technology, Compute eXpress Link (CXL). Specifically, we first devise an LPDDR5X-based CXL memory architecture with 512GB of capacity and 1.1TB/s of bandwidth, which boasts 16× larger capacity and 10× higher bandwidth than GDDR6and DDR5-based CXL memory architectures, respectively, under a module form-factor constraint. Second, we design a CXLPNM controller architecture integrated with an LLM inference accelerator, exploiting the unique capabilities of such CXL memory to overcome the disadvantages of competing technologies such as HBM-PIM and AxDIMM. Lastly, we implement a CXLPNM software stack that supports seamless and transparent use of CXL-PNM for Python-based LLM programs. Our evaluation shows that a CXL-PNM appliance with 8 CXL-PNM devices offers 23% lower latency, 31% higher throughput, and 2.8× higher energy efficiency at 30% lower hardware cost than a GPU appliance with 8 GPU devices for an LLM inference service. Sangsoo Park, Kyungsoo Kim 0003, Jinin So, Jin Jung, Jonggeon Lee, Kyoungwan Woo, Nayeon Kim 0006, Younghyun Lee, Hyungyo Kim, Yongsuk Kwon, Jinhyun Kim, Yeongon Cho, Yongmin Tai, Jeonghyeon Cho, Hoyoung Song, Jung Ho Ahn, Nam Sung Kim |
HPCA | 1 |
| 2023 | Samsung PIM/PNM for Transfmer Based AI : Energy Efficiency on PIM/PNM Cluster
Jin Hyun Kim, Yuhwan Ro, Jinin So, Sukhan Lee 0002, Shinhaeng Kang, Yeongon Cho, Byeongho Kim, Kyungsoo Kim 0003, Sangsoo Park, Jin-Seong Kim, Sanghoon Cha, Won-Jo Lee, Jin Jung, Jonggeon Lee, Joon-Ho Song, Seungwon Lee 0006, Jeonghyeon Cho, Jaehoon Yu, Kyomin Sohn |
HCS | 10 |
| 2023 | High-Speed Counter With Novel LFSR State ExtensionabstractThis paper presents a high-speed counter architecture associated with novel LFSR state extension. By employing the proposed state extension, an$\emph {m}$-bit LFSR counter with$(\mathrm{2^{\mathit{m}}}-1)$states is modified to cover$\mathrm{2^{\mathit{m}}}$states without degrading the counting rate. Based on the property that only the low-order bits are frequently switched, the proposed counter consists of two sub-counters to achieve a high counting rate and reduce the hardware complexity needed to convert an LFSR state into a binary state. The low-order sub-counter is implemented with the proposed LFSR counter, and the high-order sub-counter is designed by employing the conventional synchronous binary counter. In addition, the implemented counter takes into account the speed degradation caused by the large fan-out of the high-order sub-counter. The proposed counter designed with standard cells operates at 2.08 GHz in a 65 nm CMOS technology, and its counting rate is almost independent of the counter size. Hyungjoon Bae 0001, Yujin Hyun, Suchang Kim, Sangsoo Park, Jaeyoung Lee 0004, Boseon Jang, Suyoung Choi, In-Cheol Park |
IEEE Trans. Computers | 4 |
| 2022 | Multi-Mode QC-LDPC Decoding Architecture With Novel Memory Access Scheduling for 5G New-Radio StandardabstractAs the low-density parity-check (LDPC) code has a powerful error-correcting performance and can achieve high throughput, it is being used in many application areas and recently adopted as a channel coding method in the 5G New-Radio communication standard. Unlike other LDPC codes, the 5G LDPC code has various irregular lifting sizes to support diverse message lengths. To meet the demanding requirements of the 5G standard, many solutions have been presented, but all of them are either impractical or fail to satisfy all the requirements. This paper, for the first time, proposes an area-efficient QC-LDPC decoder that satisfies the peak throughput requirements of the 5G standard and supports all the lifting sizes specified in the 5G standard. Instead of relying on full parallelism like in the previous works, this work tries partial parallelism to mitigate the hardware complexity, which leads to high efficiency in hardware complexity. In addition, a novel memory access scheduling method is proposed to solve the data access and alignment problems caused by the partially parallel structure, which is effective in supporting all the lifting sizes. A LDPC decoder realized in 65-nm CMOS technology demonstrates that its decoding throughput is greater than 20Gbps and its area is smaller than the existing decoders. Seongjin Lee, Sangsoo Park, Boseon Jang, In-Cheol Park |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | A Linear Time Algorithm for Constructing Hierarchical Overlap GraphsabstractThe hierarchical overlap graph (HOG) is a graph that encodes overlaps from a given set P of n strings, as the overlap graph does. A best known algorithm constructs HOG in O(||P|| log n) time and O(||P||) space, where ||P|| is the sum of lengths of the strings in P. In this paper we present a new algorithm to construct HOG in O(||P||) time and space. Hence, the construction time and space of HOG are better than those of the overlap graph, which are O(||P|| + n^2). Sangsoo Park, Sung Gwan Park, Bastien Cazaux, Kunsoo Park, Eric Rivals |
CPM | 1 |
| 2017 | Exploiting I/O Reordering and I/O Interleaving to Improve Application Launch PerformanceabstractApplication prefetchers improve application launch performance through either I/O reordering or I/O interleaving. However, there has been no proposal to combine the two techniques together, missing the opportunity for further optimization. We present a new application prefetching technique to take advantage of both the approaches. We evaluated our method with a set of applications to demonstrate that it reduces cold start application launch time by 50%, which is an improvement of 22% from the I/O reordering technique. Yongsoo Joo, Sangsoo Park, Hyokyung Bahn |
ACM Trans. Storage | 2 |
| 2014 | Task-I/O Co-scheduling for Pfair Real-Time Scheduler in Embedded Multi-core SystemsabstractReal-time embedded systems often require the ability of robust controls because interactions between the embedded systems and their physical environment that dynamically changes. Multi-core chips are regarded as ideal candidate hardware components for such environments, since each of them carries two or more cores on a single die, and has potential for providing execution parallelism as well as better performance at low cost. Parallelism, on the other hand, necessitates complex analysis of computation problems, such as task scheduling, while improving the realization of embedded controls. Pfair is an optimal scheduling algorithm that can fully utilize all cores in the system, but it incurs an excessive scheduling overhead which, in turn, diminishes its practicality in embedded systems. To mitigate this problem, hybrid partitioned-global Pfair (HPGP) scheduler was proposed, which significantly reduces the number of task migrations and global scheduling points by performing global scheduling only when absolutely necessary, while still achieving full processor utilization. This paper further extends the HPGP scheduler to support the robustness to interactions with the physical environment. Our evaluation results have shown that the extended HPGP can successfully limits the increase in response time caused by the hardware interrupts for physical interactions under a wide range of system utilization conditions., thus making it suitable for embedded real-time systems. Sangsoo Park |
EUC | 1 |
| 2014 | Rapid Prototyping and Evaluation of Intelligence Functions of Active Storage DevicesabstractActive storage devices further improve their performance by executing “intelligence functions,” such as prefetching and data deduplication, in addition to handling the usual I/O requests they receive. Significant research has been carried out to develop effective intelligence functions for the active storage devices. However, laborious and time-consuming efforts are usually required to set up a suitable experimental platform to evaluate each new intelligence function. Moreover, it is difficult to make such prototypes available to other researchers and users to gain valuable experience and feedback. To overcome these difficulties, we propose$\tt {IOLab}$, a virtual machine (VM)-based platform for evaluating intelligence functions of active storage devices. The VM-based structure of$\tt {IOLab}$enables the evaluation of new (and existing) intelligence functions for different types of OSes and active storage devices with little additional effort.$\tt {IOLab}$also supports real-time execution of intelligence functions, providing users opportunities to experience latest intelligence functions without waiting for their deployment in commercial products. Using a set of interesting case studies, we demonstrate the utility of$\tt {IOLab}$with negligible performance overhead except for the VM’s virtualization overhead. Yongsoo Joo, Junhee Ryu, Sangsoo Park, Heonshik Shin, Kang G. Shin |
IEEE Trans. Computers | 3 |
| 2013 | Robust Scheduling of Dynamic Real-Time Tasks with Low Overhead for Multi-Core Systems
Sangsoo Park |
ICA3PP (2) | 1 |
| 2012 | Improving Application Launch Performance on Solid State Drives
Yongsoo Joo, Junhee Ryu, Sangsoo Park, Kang G. Shin |
J. Comput. Sci. Technol. | 3 |
| 2011 | FAST: Quick Application Launch on Solid-State Drives
Yongsoo Joo, Junhee Ryu, Sangsoo Park, Kang G. Shin |
FAST | 3 |
| 2010 | Integration of Collaborative Analyses for Development of Embedded Control SoftwareabstractModel-based methodologies have been widely used to handle the increasing demand for rapid development of high-quality, real-time embedded control software. A key challenge in such model-based design is integration of various “collaborative” analysis methods to support the automation of the design process. Traditional analysis methods developed for analyzing specific system properties, however, are not designed for such integration, and thus cannot ensure that the information for the analysis will be provided at the design stage where the information is needed. Moreover, many traditional analysis methods depend heavily on complete and accurate design models which can only be applied to post-design verification and are unavailable for automation of the design process involving an early design stage, where implementation details are unknown. This challenge can be met by integrating analysis methods with the design process. We have developed such a framework combined with software modeling, execution platform configuration, and run-time monitoring mechanisms to enable accurate assessment of embedded software quality at early design stages. We have implemented and demonstrated the framework with a toolkit, called AIRES, that integrates software models, a virtual execution platform, and timing and schedulability analysis methods. Sangsoo Park, Kang G. Shin, Shige Wang |
Proc. IEEE | 1 |
| 2007 | Integrating Virtual Execution Platform for Accurate Analysis in Distributed Real-Time Control System DevelopmentabstractA distributed real-time control system is modeled by au- tomatically generating a virtual execution platform and in- tegrating it with an abstract run-time model. This allows us to capture the dynamic effects of non-deterministic behavior of the underlying hardware and real-time operating system (RTOS), which cannot be accurately evaluated by existing static approaches. Our framework has been implemented and integrated with the existing AIRES toolkit. The con- struction of a virtual execution platform for a control appli- cation is observed to take 422ms. The integrated platform and control application consists of 56 software components connected by 1280 links for data exchanges and precedence relations, and the constructed platform executes the appli- cation at a normalized speed--defined as the ratio of simu- lation time to real time -- 10.6. Our preliminary evaluation has demonstrated the virtual execution platform's capabil- ity of providing accurate run-time information at a reason- able time-cost. Sangsoo Park, Walter Olds, Kang G. Shin, Shige Wang |
RTSS | 1 |
| 2004 | An experimental analysis of the effect of the operating system on memory performance in embedded multimedia computingabstractAs embedded systems grow in size and complexity, an operating system has become essential to simplify the design of system software, for which more accurate analysis of its impact on memory performance is required. In this paper, we intend to investigate how the OS influences memory performance at run time by quantitatively evaluating the memory system behavior of an MPEG-4 application running on embedded Linux. Through the use of extensive simulations we have confirmed that the OS has poor memory performance with less memory locality than applications. The results of our experimental analysis are deemed useful for helping embedded system designers understand the memory performance of the OS and the application within a system, extending their capability to design a more power-aware and faster system. Sangsoo Park, Yonghee Lee, Heonshik Shin |
EMSOFT | 1 |
| 2004 | Experimental Performance Evaluation of Embedded Linux Using Alternative CPU Core Organizations
Sangsoo Park, Yonghee Lee, Heonshik Shin |
EUC | 1 |
| 2003 | Rigorous Modeling of Disk Performance for Real-Time Applications
Sangsoo Park, Heonshik Shin |
RTCSA | 1 |
| 2002 | Comparative performance evaluation of Java threads for embedded applications: Linux Thread vs. Green Thread
Minyoung Sung, Sangsoo Park, Naehyuck Chang, Heonshik Shin |
Inf. Process. Lett. | 3 |