VLDB 2026 Research / reviewers in the wild / expert
Youngsun Han
dblp:67/4306
· DBLP profile ↗
8ranked-venue papers
2as first author
0since 2021 · last 2019
0000-0001-7712-2514ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 64% GPUs and heterogeneous computing · 19% Reconfigurable computing and FPGAs · 16% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 94% Runtime systems and virtual machines · 6% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization › parallelization
automatic parallelization |
0.2 | 1 | 2016 | JavaScript Parallelizing Compiler for Exploiting Parallelism from Data-Parallel HTML5 Applications · ACM Trans. Archit. Code Optim. 2016 |
Parallel and multicore computing
data-parallel programming |
0.2 | 1 | 2016 | JavaScript Parallelizing Compiler for Exploiting Parallelism from Data-Parallel HTML5 Applications · ACM Trans. Archit. Code Optim. 2016 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2016 | JavaScript Parallelizing Compiler for Exploiting Parallelism from Data-Parallel HTML5 Applications · ACM Trans. Archit. Code Optim. 2016 |
Compilers and program optimization
hardware compilation |
0.1 | 1 | 2006 | Jaguar: a compiler infrastructure for Java reconfigurable computing · FPGA 2006 |
Reconfigurable computing and FPGAs
FPGA compilation |
0.1 | 1 | 2006 | Jaguar: a compiler infrastructure for Java reconfigurable computing · FPGA 2006 |
Methods — techniques the papers use, named apart from their topics
speculation · 0.5idempotence · 0.5affine loop analysis · 0.5program analysis · 0.1optimization · 0.1intermediate representation · 0.1bytecode parsing · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Fault Tolerance Technique Offlining Faulty Blocks by Heap Memory ManagementabstractAs dynamic random access memory (DRAM) cells continue to be scaled down for higher density and capacity, they have more faults. Thus, DRAM reliability becomes a major concern in computer systems. Previous studies have proposed many techniques preserving the reliability in various system components, such as DRAM internal, memory controller, caches, and operating systems. By reviewing the techniques, we identified the following two considerations: First, it is possible to recover faults with reasonable overhead at high fault rate only if the recovery unit is fine-grained. Second, since hardware modification requires additional cost in the employment of a technique, a pure software-based recovery technique is preferable. However, in the existing software-based recovery technique, the recovery unit is too coarse-grained to tolerate the high fault rate. In this article, we propose a pure software-based recovery technique with fine-granularity. Our key idea is based on heap segments being managed by the system library with variable-sized chunks to handle dynamic allocation in user applications. In our technique, faulty blocks in pages are offlined by marking them as allocated chunks. Thus, not only fault-free pages but also the remaining clean blocks in faulty pages are allowed to be usable space. Our technique is implemented by modifying the operating system and the system library. Since hardware assistance is unnecessary in the implementation, we evaluated our method on a real machine. Our evaluation results show that our technique has negligible performance overhead at high bit error rate (BER) 5.12e-5, which a hardware-based recovery technique could not tolerate without unacceptable area overhead. Also, at the same BER, our method provides 5.22× usable space, compared with page-offline, which is the state-of-the-art pure software-based technique. Jaeyung Jun, Yoonah Paik, Gyeong Il Min, Seon Wook Kim, Youngsun Han |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2018 | Recovering from Biased Distribution of Faulty Cells in Memory by Reorganizing Replacement Regions through Universal HashingabstractRecently, scaling down dynamic random access memory (DRAM) has become more of a challenge, with more faults than before and a significant degradation in yield. To improve the yield in DRAM, a redundancy repair technique with intra-subarray replacement has been extensively employed to replace faulty elements (i.e., rows or columns with defective cells) with spare elements in each subarray. Unfortunately, such technique cannot efficiently handle a biased distribution of faulty cells because each subarray has a fixed number of spare elements. In this article, we propose a novel redundancy repair technique that uses a hashing method to solve this problem. Our hashing technique reorganizes replacement regions by changing the way in which their replacement information is referred, thus making faulty cells become evenly distributed to the regions. We also propose a fast repair algorithm to find the best hash function among all possible candidates. Even if our approach requires little hardware overhead, it significantly improves the yield when compared with conventional redundancy techniques. In particular, the results of our experiment show that our technique saves spare elements by about 57% and 55% for a yield of 99% at BER 1e-6 and 5e-7, respectively. Jaeyung Jun, Kyu Hyun Choi, Hokwon Kim, Sang Ho Yu, Seon Wook Kim, Youngsun Han |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2017 | Content-Aware Bit Shuffling for Maximizing PCM EnduranceabstractRecently, phase change memory (PCM) has been emerging as a strong replacement for DRAM owing to its many advantages such as nonvolatility, high capacity, low leakage power, and so on. However, PCM is still restricted for use as main memory because of its limited write endurance. There have been many methods introduced to resolve the problem by either reducing or spreading out bit flips. Although many previous studies have significantly contributed to reducing bit flips, they still have the drawback that lower bits are flipped more often than higher bits because the lower bits frequently change their bit values. Also, interblock wear-leveling schemes are commonly employed for spreading out bit flips by shifting input data, but they increase the number of bit flips per write. In this article, we propose a noble content-aware bit shuffling (CABS) technique that minimizes bit flips and evenly distributes them to maximize the lifetime of PCM at the bit level. We also introduce two additional optimizations, namely, addition of an inversion bit and use of an XOR key, to further reduce bit flips. Moreover, CABS is capable of recovering from stuck-at faults by restricting the change in values of stuck-at cells. Experimental results showed that CABS outperformed the existing state-of-the-art methods in the aspect of PCM lifetime extension with minimal overhead. CABS achieved up to 48.5% enhanced lifetime compared to the data comparison write (DCW) method only with a few metadata bits. Moreover, CABS obtained approximately 9.7% of improved write throughput than DCW because it significantly reduced bit flips and evenly distributed them. Also, CABS reduced about 5.4% of write dynamic energy compared to DCW. Finally, we have also confirmed that CABS is fully applicable to BCH codes as it was able to reduce the maximum number of bit flips in metadata cells by 32.1%. Miseon Han, Youngsun Han, Seon Wook Kim, Hokyoon Lee, Il Park 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | JavaScript Parallelizing Compiler for Exploiting Parallelism from Data-Parallel HTML5 ApplicationsabstractWith the advent of the HTML5 standard, JavaScript is increasingly processing computationally intensive, data-parallel workloads. Thus, the enhancement of JavaScript performance has been emphasized because the performance gap between JavaScript and native applications is still substantial. Despite this urgency, conventional JavaScript compilers do not exploit much of parallelism even from data-parallel JavaScript applications, despite contemporary mobile devices being equipped with expensive parallel hardware platforms, such as multicore processors and GPGPUs. In this article, we propose an automatically parallelizing JavaScript compiler that targets emerging, data-parallel HTML5 applications by leveraging the mature affine loop analysis of conventional static compilers. We identify that the most critical issues when parallelizing JavaScript with a conventional static analysis are ensuring correct parallelization, minimizing compilation overhead, and conducting low-cost recovery when there is a speculation failure during parallel execution. We propose a mechanism for safely handling the failure at a low cost, based on compiler techniques and the property of idempotence. Our experiment shows that the proposed JavaScript parallelizing compiler detects most affine parallel loops. Also, we achieved a maximum speedup of 3.22 times on a quad-core system, while incurring negligible compilation and recovery overheads with various sets of data-parallel HTML5 applications. Yeoul Na, Seon Wook Kim, Youngsun Han |
ACM Trans. Archit. Code Optim. | 3 |
| 2015 | O2WebCL: an automatic OpenCL-to-WebCL translator for high performance web computing
Myeongjin Cho, Youngsun Han, Seon Wook Kim |
J. Supercomput. | 2 |
| 2012 | Resource Efficient Implementation of Low Power MB-OFDM PHY Baseband Modem With Highly Parallel ArchitectureabstractThe multi-band orthogonal frequency-division multiplexing modem needs to process large amount of computations in short time for support of high data rates, i.e., up to 480 Mbps. In order to satisfy the performance requirement while reducing power consumption, a multi-way parallel architecture has been proposed. But the use of the high degree parallel architecture would increase chip resource significantly, thus a resource efficient design is essential. In this paper, we introduce several novel optimization techniques for resource efficient implementation of the baseband modem which has highly, i.e., 8-way, parallel architecture, such as new processing structures for a (de)interleaver and a packet synchronizer and algorithm reconstruction for a carrier frequency offset compensator. Also, we describe how to efficiently design several other components. The detailed analysis shows that our optimization technique could reduce the gate count by 27.6% on average, while none of techniques degraded the overall system performance. With 0.18-μm CMOS process, the gate count and power consumption of the entire baseband modem were about 785 kgates and less than 381 mW at 66 MHz clock rate, respectively. Seokjoong Hwang, Youngsun Han, Seon Wook Kim, Jongsun Park 0001, Byung Gueon Min |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | A Novel Architecture for Block Interleaving Algorithm in MB-OFDM Using Mixed Radix SystemabstractIn this paper, we present a novel architecture of a block interleaver in MB-OFDM systems based on Mixed Radix System (MRS). We prove mathematically that the proposed architecture can support bit permutations in the interleaving process. The hierarchical property of our proposed MRS-based design methodology allows the proposed architecture to support all the required data rates in the MB-OFDM systems with simple modular design. Furthermore, the same design to be used for the interleaver can also be used for the operation of de-interleaving, which reduces the implementation complexity significantly. The latency of our architecture is as low as 6 MB-OFDM symbols. In addition, when comparing our proposed architecture with the conventional approach, we are able to reduce the implementation complexity by 85.5%, 69.4%, and 40.3% for 80, 200, and 480 Mb/s data rates, respectively, while improving our operating maximum clock frequency by more than 3.3 times over the conventional design. We also show that the power consumption is reduced by 87.4%, 73.6%, and 39.8% for 80, 200, and 480 Mb/s, respectively. Youngsun Han, Peter Harliman, Seon Wook Kim, Jong-Kook Kim, Chulwoo Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2006 | Jaguar: a compiler infrastructure for Java reconfigurable computingabstractIn this paper, we present our compiler infrastructure, called Jaguar for Java reconfigurable computing. The Jaguar compiler translates compiled Java methods, i.e. sequence of bytecodes into Verilog synthesizable code modules with exploiting the maximum operational parallelism in applications. Our compiler infrastructure consists of two major components: One is a compiler to generate synthesizable Verilog codes from Java applications, which performs full compilation passes, such as bytecode parsing, intermediate representation (IR) construction, program analysis, optimization, and code emission. The other component is the Java Virtual Machine (JVM), which provides Java execution environment to compiler-generated hardware. The JVM runs on a host processor and the generated hardware does on FPGA. Differently from previous work, our compiler infrastructure is a complete and solid solution for Java reconfigurable computing. We present how to design our compiler framework. Our infrastructure improves the performance by 66% on average and by up to 174% in measured benchmarks. Also we discuss the performance issues in detail, especially focusing on overhead of interactions between JVM and Jaguar hardware. Youngsun Han, Seokjoong Hwang, Seon Wook Kim |
FPGA | 1 |