Changsu Kim 0004

dblp:02/4380-4 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-6688-9322ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FAST: An Efficient Scheduler for All-to-All GPU Communication
Yiran Lei, Dongjoo Lee 0001, Liangyu Zhao, Daniar Kurniawan, Chanmyeong Kim, Heetaek Jeong, Changsu Kim 0004, Hyeonseong Choi, Liangcheng Yu, Arvind Krishnamurthy, Justine Sherry, Eriko Nurvitadhi
NSDI7
2023 F4T: A Fast and Flexible FPGA-based Full-stack TCP Acceleration Framework
abstract
As complex workloads that run on many servers are pursuing higher networking throughput, more CPU cycles are consumed to support the TCP stack. To mitigate the high CPU burden from executing the compute-intensive TCP, prior works have proposed to offload TCP processing to the embedded processors, ASICs, or FPGAs in network devices. However, none of the approaches satisfy all of the critical requirements of TCP simultaneously, which are high performance, many connections, and high flexibility. Embedded processors do not provide enough performance to fully offload the TCP stack, while ASICs fail to provide enough flexibility. Meanwhile, existing FPGA-based TCP accelerators either fail to provide high performance or give up some of the critical features and requirements to achieve high performance due to their inefficient processing architecture.
Junehyuk Boo, Yujin Chung, Eunjin Baek, Seongmin Na, Changsu Kim 0004, Jangwoo Kim
ISCA5
2022 HECATE: Performance-Aware Scale Optimization for Homomorphic Encryption Compiler
abstract
Despite the benefit of Fully Homomorphic Encryption (FHE) that supports encrypted computation, writing an efficient FHE application is challenging due to magnitude scale management. Each FHE operation increases scales of ciphertext and leaving the scales high harms performance of the following FHE operations. Thus, rescaling ciphertext is inevitable to optimize an FHE application, but since FHE requires programmers to match the rescaling levels of operands of each FHE operation, programmers should rescale ciphertext reflecting the entire FHE application. Although recently proposed FHE compilers reduce the programming burden by automatically manipulating ciphertext scales, they fail to fully optimize the FHE application because they greedily rescale the ciphertext without considering their performance impacts throughout the entire application. This work proposes HECATE, a new FHE compiler framework that optimizes scales of ciphertext reflecting their rescaling levels and performance impact. With a new type system that embeds the scale and rescaling level, and a new rescaling operation called downscale, HECATE makes various scale management plans, analyzes their expected performance, and finds the optimal rescaling points throughout the entire FHE application. This work implements HECATE on top of the MLIR framework with a Python frontend and shows that HECATE achieves 27% speedup over the state-of-the-art approach for various FHE applications.
Yongwoo Lee 0001, Seonyeong Heo, Seonyoung Cheon, Shinnung Jeong, Changsu Kim 0004, Eunkyung Kim 0002, Hanjun Kim 0001
CGO5
2021 Thread-Aware Area-Efficient High-Level Synthesis Compiler for Embedded Devices
abstract
In the embedded device market, custom hardware platforms such as an application specific integrated circuit (ASIC) and a field programmable gate array (FPGA) are attractive thanks to their high performance and power efficiency. However, its huge design costs make it challenging for manufacturers to timely launch new devices. High-level synthesis (HLS) helps significantly reduce the design costs by automating the translation of service algorithms into hardware logics; however, current HLS compilers do not fit well to embedded devices as they fail to produce area-efficient solutions while supporting concurrent events from diverse peripherals such as sensors, actuators and network modules. This paper proposes a new thread-aware HLS compiler named Duro that produces area-efficient embedded devices. Duro shares commonly-invoked functions and operators across different callers and threads with a new thread-aware area cost model, and thus effectively reduces the logic size. Moreover, Duro supports a variety of device peripherals by automatically integrating peripheral controllers and interfaces as peripheral drivers. The experiment results of six embedded devices with ten peripherals demonstrate that Duro reduces the area and energy dissipation of embedded devices by 28.5% and 25.3% compared with the designs generated by the state-of-the-art HLS compiler. This work also implements FPGA prototypes of the six devices using Duro, and the measurement results show 65.3% energy saving over Raspberry Pi Zero with slightly better computation performance.
Changsu Kim 0004, Shinnung Jeong, Sungjun Cho, Yongwoo Lee 0001, William Song, Youngsok Kim, Hanjun Kim 0001
CGO1
2021 Precise Correlation Extraction for IoT Fault Detection With Concurrent Activities
abstract
In the Internet of Things (IoT) environment, detecting a faulty device is crucial to guarantee the reliable execution of IoT services. To detect a faulty device, existing schemes trace a series of events among IoT devices within a certain time window, extract correlations among them, and find a faulty device that violates the correlations. However, if a few users share the same IoT environment, since their concurrent activities make non-correlated devices react together in the same time window, the existing schemes fail to detect a faulty device without differentiating the concurrent activities. To correctly detect a faulty device in the multiple concurrent activities, this work proposes a new precise correlation extraction scheme, called PCoExtractor. Instead of using a time window, PCoExtractor continuously traces the events, removes unrelated device statuses that inconsistently react for the same activity, and constructs fine-grained correlations. Moreover, to increase the detection precision, this work newly defines a fine-grained correlation representation that reflects not only sensor values and functionalities of actuators but also their transitions and program states such as contexts. Compared to existing schemes, PCoExtractor detects and identifies 40.06% more faults for 4 IoT services with concurrent activities of 12 users while reducing 80.3% of detection and identification times.
Gyeongmin Lee, Bongjun Kim, Seungbin Song, Changsu Kim 0004, Jong Kim 0001, Hanjun Kim 0001
ACM Trans. Embed. Comput. Syst.4
2020 Pipeline-aware Logic Deduplication in High-Level Synthesis for Post-Quantum Cryptography Algorithms
abstract
With the technical advance of quantum computers that can solve intractable problems for conventional computers, many of the currently used public-key cryptosystems become vulnerable. Recently proposed post-quantum cryptography (PQC) is secure against both classical and quantum computers, but existing embedded systems such as smart card can not easily support the PQC algorithms due to their much larger key sizes and more complex arithmetics. To accelerate the PQC algorithms, embedded systems have to embed the PQC hardware blocks, which can lead to huge hardware design costs. Although High-Level Synthesis (HLS) helps significantly reduce the design costs, current HLS frameworks produce inefficient hardware design for the PQC algorithms in terms of area and performance. This work analyzes common features of the PQC algorithms and proposes a new pipeline-aware logic deduplication method in HLS. The proposed method shares commonly invoked logic across hardware design while considering load balancing in pipeline and resolving dynamic memory accesses. This work implements FPGA hardware design of seven PQC algorithms in the round 2 candidates from the National Institute of Standards and Technology (NIST) PQC standardization process. Compared to commercial HLS framework, the proposed method achieves an area-delay-product reduction by 34.5%.
Changsu Kim 0004, Yongwoo Lee 0001, Shinnung Jeong, Wen Wang 0007, Jakub Szefer, Hanjun Kim 0001
FPGA1
2017 Context-Aware Memory Profiling for Speculative Parallelism
abstract
To expose hidden parallelism from programs with complex dependences, modern compilers employ memory profilers to augment imprecise static analyses. Since dynamic dependence patterns among instructions can vary widely depending on the context, such as function call site stack and loop nest level, context-aware memory profiling is of great value for precise memory profiling. However, recording memory dependences with full context information causes huge overheads in terms of CPU cycles and memory space. Existing profilers mitigate this problem by compromising precision, coverage, or both. This paper proposes a new precise Context-Aware Memory Profiling (CAMP) framework that efficiently traces all the memory dependences with full context information. CAMP statically analyzes a context tree of a program that illustrates all the possible dynamic contexts, and simplifies context management during profiling. For 14 programs from SPEC CINT2000 and CINT2006 benchmark suites, CAMP increases speculative parallelism opportunities by 12.6% on average and by up to 63.0% compared to the baseline context-oblivious, loop-aware memory profiler.
Changsu Kim 0004, Juhyun Kim, Juwon Kang, Jae W. Lee, Hanjun Kim 0001
HiPC1