VLDB 2026 Research / reviewers in the wild / expert
Seungkyun Kim
dblp:34/644
· DBLP profile ↗
6ranked-venue papers
1as first author
0since 2021 · last 2011
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 64% Embedded and real-time systems · 28% GPUs and heterogeneous computing · 8% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 50% Compilers and program optimization · 50% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
code layout optimization |
0.1 | 1 | 2010 | Scratchpad Memory Management Techniques for Code in Embedded Systems without an MMU · IEEE Trans. Computers 2010 |
Operating systems › resource management › memory management
memory allocation |
0.1 | 1 | 2010 | Scratchpad Memory Management Techniques for Code in Embedded Systems without an MMU · IEEE Trans. Computers 2010 |
Memory systems › shared memory
distributed shared memory |
0.1 | 1 | 2010 | COMIC++: A software SVM system for heterogeneous multicore accelerator clusters · HPCA 2010 |
Embedded and real-time systems › embedded software › embedded operating systems › embedded memory management
scratchpad memory management |
0.1 | 1 | 2010 | Scratchpad Memory Management Techniques for Code in Embedded Systems without an MMU · IEEE Trans. Computers 2010 |
Memory systems › virtual memory management
shared virtual memory |
0.1 | 1 | 2010 | COMIC++: A software SVM system for heterogeneous multicore accelerator clusters · HPCA 2010 |
Memory systems
cache |
0.0 | 1 | 2010 | Scratchpad Memory Management Techniques for Code in Embedded Systems without an MMU · IEEE Trans. Computers 2010 |
GPUs and heterogeneous computing › heterogeneous cluster computing
heterogeneous accelerator clusters |
0.0 | 1 | 2010 | COMIC++: A software SVM system for heterogeneous multicore accelerator clusters · HPCA 2010 |
Methods — techniques the papers use, named apart from their topics
profiling · 0.2mixed integer linear programming · 0.2demand paging · 0.2software-managed caches · 0.1hierarchical centralized release consistency · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | Demand Paging Techniques for Flash Memory Using Compiler Post-Pass OptimizationsabstractIn this article, we propose an application-specific demand paging mechanism for low-end embedded systems that have flash memory as secondary storage. These systems are not equipped with virtual memory. A small memory space called an execution buffer is used to page the code of an application. An application-specific page manager manages the buffer. The page manager is automatically generated by a compiler post-pass optimizer and combined with the application image. The post-pass optimizer analyzes the executable image and transforms function call/return instructions into calls to the page manager. As a result, each function in the code can be loaded into the memory on demand at runtime. To minimize the overhead incurred by the demand paging technique, code clustering algorithms are also presented. We evaluate our techniques with ten embedded applications, and our approach can reduce the code memory size by on average 39.5% with less than 10% performance degradation and on average 14% more energy consumption. Our demand paging technique provides embedded system designers with a trade-off control mechanism between the cost, performance, and energy efficiency in designing embedded systems. Embedded system designers can choose the code memory size depending on their cost, energy, and performance requirements. Seungkyun Kim, Ki-Won Kwon, Chihun Kim, Choonki Jang, Jaejin Lee, Sang Lyul Min |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2010 | An OpenCL framework for heterogeneous multicores with local memoryabstractIn this paper, we present the design and implementation of an Open Computing Language (OpenCL) framework that targets heterogeneous accelerator multicore architectures with local memory. The architecture consists of a general-purpose processor core and multiple accelerator cores that typically do not have any cache. Each accelerator core, instead, has a small internal local memory. Our OpenCL runtime is based on software-managed caches and coherence protocols that guarantee OpenCL memory consistency to overcome the limited size of the local memory. To boost performance, the runtime relies on three source-code transformation techniques, work-item coalescing, web-based variable expansion and preload-poststore buffering, performed by our OpenCL C source-to-source translator. Work-item coalescing is a procedure to serialize multiple SPMD-like tasks that execute concurrently in the presence of barriers and to sequentially run them on a single accelerator core. It requires the web-based variable expansion technique to allocate local memory for private variables. Preload-poststore buffering is a buffering technique that eliminates the overhead of software cache accesses. Together with work-item coalescing, it has a synergistic effect on boosting performance. We show the effectiveness of our OpenCL framework, evaluating its performance with a system that consists of two Cell BE processors. The experimental result shows that our approach is promising. Jaejin Lee, Seungkyun Kim, Jung-Ho Park, Honggyu Kim, Thanh Tuan Dao, Yongjin Cho, Sung Jong Seo, Seung Hak Lee, Seung Mo Cho, Hyo Jung Song, Sang-Bum Suh, Jong-Deok Choi |
PACT | 4 |
| 2010 | Parallelizing the H.264 decoder on the cell BE architectureabstractIn this paper, we propose parallelization and optimization techniques of the H.264 decoder for the Cell BE processor. We exploit both frame-level parallelism and macroblock pipelining. The major bottleneck in achieving the real-time performance is the entropy decoding stage, CABAC. Our decoder eliminates this bottleneck by exploiting the frame-level parallelism available in the entropy decoding stage. A macroblock software cache and a prefetching technique for the cache are used to facilitate macroblock pipelining. In addition, an asynchronous macroblock buffering technique is used to eliminate the effect of load imbalance between pipeline stages. We evaluate the effectiveness of our approach by implementing a parallel H.264 decoder on an IBM Cell blade server. The evaluation results indicate that our parallel H.264 decoder (with CABAC entropy decoding) on a single Cell BE processor meets the real-time requirement of the full HD standard at level 4.0. Moreover, our decoder also satisfies the real-time requirement at level 4.1 when an additional Cell BE processor is used. Yongjin Cho, Seungkyun Kim, Jaejin Lee, Heonshik Shin |
EMSOFT | 2 |
| 2010 | COMIC++: A software SVM system for heterogeneous multicore accelerator clustersabstractIn this paper, we propose a software shared virtual memory (SVM) system for heterogeneous multicore accelerator clusters with explicitly managed memory hierarchies. The target cluster consists of a single manager node and many compute nodes. The manager node contains a generalpurpose processor and larger main memory, and each compute node contains a heterogeneous multicore processor and smaller main memory. These nodes are connected with an interconnection network, such as Gigabit Ethernet. The heterogeneous multicore processor in each compute node consists of a general-purpose processor element (GPE) and multiple accelerator processor elements (APEs). The GPE runs an OS and the multiple APEs are dedicated to compute-intensive workloads. The GPE is typically backed by a deep on-chip cache hierarchy and hardware cache coherence. On the other hand, the APEs have small explicitly-addressed local memory instead of caches. This APE local memory is not coherent with the main memory. Different main and local memory units in the accelerator cluster can be viewed as an explicitly managed memory hierarchy: global memory, node local memory, and APE local memory. Since coherence protocols of previous software SVM proposals cannot effectively handle such a memory hierarchy, we propose a new coherence and consistency protocol, called hierarchical centralized release consistency (HCRC). Our software SVM system is built on top of HCRC and software-managed caches. We evaluate the effectiveness and analyze the performance of our software SVM system on a 32-node heterogeneous multicore cluster (a total of 192 APEs). Jaejin Lee, Seungkyun Kim, Zehra Sura |
HPCA | 5 |
| 2010 | Scratchpad Memory Management Techniques for Code in Embedded Systems without an MMUabstractWe propose a code scratchpad memory (SPM) management technique with demand paging for embedded systems that have no memory management unit. Based on profiling information, a postpass optimizer analyzes and optimizes application binaries in a fully automated process. It classifies the code of the application including libraries into three classes based on a mixed integer linear programming formulation: External code is executed directly from the external memory. Pinned code is loaded into the SPM when the application starts and stays in the SPM. Paged code is loaded into/unloaded from the SPM on demand. We evaluate the proposed technique by running 14 embedded applications both on a cycle-accurate ARM processor simulator and an ARM1136JF-S core. On the simulator, the reference case, a four-way set-associative cache, is compared to a direct-mapped cache and an SPM of comparable die area. On average, we observe an improvement of 12 percent in runtime performance and a 21 percent reduction in energy consumption. On the ARM11 board, the reference case run on the 16-KB four-way set-associative cache is compared to the demand paging solution on the 16-KB SPM, optionally supported by the cache. The measured results show both a runtime performance improvement and a reduction of the energy consumption by 23 percent, on average. Bernhard Egger 0002, Seungkyun Kim, Choonki Jang, Jaejin Lee, Sang Lyul Min, Heonshik Shin |
IEEE Trans. Computers | 2 |
| 2008 | FaCSim: a fast and cycle-accurate architecture simulator for embedded systemsabstractThere have been strong demands for a fast and cycle-accurate virtual platforms in the embedded systems area where developers can do meaningful software development including performance debugging in the context of the entire platform. In this paper, we describe the design and implementation of a fast and cycle-accurate architecture simulator called FaCSim as a first step towards such a virtual platform. FacSim accurately models the ARM9E-S processor core and ARM926EJ-S processor's memory subsystem. It accurately simulates exceptions and interrupts to enable whole-system simulation including the OS. Since it is implemented in a modular manner in C++, it can be easily extended with other system components by subclassing or adding new classes. FaCSim is based on an interpretive simulation technique to provide flexibility, yet achieving high speed. It enables fast cycle-accurate architecture simulation by means of three mechanisms. First, it computes elapsed cycles in each pipeline stage as a chunk and incrementally adds it up to advance the core clock instead of performing cycle-by-cycle simulation. Second, it uses a basic-block cache that caches decoded instructions at the basic-block level. Finally, it is parallelized to exploit multicore systems that are available everywhere these days. Using 21 applications from the EEMBC benchmark suite, FaCSim's accuracy is validated against the ARM926EJ-S development board from ARM, and is accurate in a ±7% error margin. Due to basic-block level caching and parallelization, FaCSim is, on average, more than three times faster than ARMulator and more than six times faster than SimpleScalar. Jaejin Lee, Choonki Jang, Seungkyun Kim, Bernhard Egger 0002, Kwangsub Kim, Sang-Yong Han |
LCTES | 4 |