Donghoon Yoo

dblp:123/7766 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
2 papers
Cryptographic primitives and cryptanalysis · 87% Cryptographic protocols and secure computation · 13%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 61% Distributed systems · 30% Processor architecture and microarchitecture · 9%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cryptographic primitives and cryptanalysis › homomorphic encryption
bootstrapping
1.422024
General Bootstrapping Approach for RLWE-Based Homomorphic Encryption · IEEE Trans. Computers 2024
Efficient FHEW Bootstrapping with Small Evaluation Keys, and Applications to Threshold Homomorphic Encryption · EUROCRYPT (3) 2023
Cryptographic primitives and cryptanalysis
homomorphic encryption
1.422024
General Bootstrapping Approach for RLWE-Based Homomorphic Encryption · IEEE Trans. Computers 2024
Efficient FHEW Bootstrapping with Small Evaluation Keys, and Applications to Threshold Homomorphic Encryption · EUROCRYPT (3) 2023
Cryptographic primitives and cryptanalysis › post-quantum cryptography
lattice-based cryptography
0.812024
General Bootstrapping Approach for RLWE-Based Homomorphic Encryption · IEEE Trans. Computers 2024
Cryptographic primitives and cryptanalysis › homomorphic encryption
RLWE-based homomorphic encryption
0.812024
General Bootstrapping Approach for RLWE-Based Homomorphic Encryption · IEEE Trans. Computers 2024
Cryptographic protocols and secure computation › threshold cryptography
threshold homomorphic encryption
0.712023
Efficient FHEW Bootstrapping with Small Evaluation Keys, and Applications to Threshold Homomorphic Encryption · EUROCRYPT (3) 2023
Memory systems
cache design
0.312018
Benzene: An Energy-Efficient Distributed Hybrid Cache Architecture for Manycore Systems · ACM Trans. Archit. Code Optim. 2018
Distributed systems
distributed caching
0.312018
Benzene: An Energy-Efficient Distributed Hybrid Cache Architecture for Manycore Systems · ACM Trans. Archit. Code Optim. 2018
Memory systems › cache › cache technology
hybrid cache
0.312018
Benzene: An Energy-Efficient Distributed Hybrid Cache Architecture for Manycore Systems · ACM Trans. Archit. Code Optim. 2018
Processor architecture and microarchitecture
many-core architecture
0.112018
Benzene: An Energy-Efficient Distributed Hybrid Cache Architecture for Manycore Systems · ACM Trans. Archit. Code Optim. 2018

Methods — techniques the papers use, named apart from their topics

noise refreshing · 0.8RLWE · 0.8intra-bank and inter-bank optimization · 0.3
YearPublicationVenuePosition
2024 General Bootstrapping Approach for RLWE-Based Homomorphic Encryption
abstract
Homomorphic Encryption (HE) makes it possible to compute on encrypted data without decryption. In lattice-based HE, a ciphertext contains noise, which accumulates along with homomorphic computations. Bootstrapping refreshes the noise and it is possible to perform arbitrary-depth computations on HE with bootstrapping, which we call Fully Homomorphic Encryption (FHE). In this article, we propose a new general bootstrapping technique for RLWE-based schemes and its practical instantiation for FHE. It can be applied to all three RLWE-based leveled FHE schemes: Brakerski-Gentry-Vaikuntanathan (BGV), Brakerski/Fan-Vercauteren (BFV), and Cheon-Kim-Kim-Song (CKKS) with minor deviations in the algorithms. Our new construction of bootstrapping extracts a noiseless ciphertext for a part of the input, scales it, and finally removes it. In contrast with previous bootstrapping algorithms, the proposed method consumes only 1–2 levels and uses smaller parameters. For BGV and BFV, our new bootstrapping does not have any restrictions on a plaintext modulus unlike typical cases of the previous methods. The error introduced by our approach for CKKS is comparable to a rescaling error, allowing us to preserve a large amount of precision after bootstrapping.
Andrey Kim, Maxim Anatolievich Deryabin, Jieun Eom, Rakyong Choi, Yongwoo Lee 0002, Whan Ghang, Donghoon Yoo
IEEE Trans. Computers7
2023 Efficient FHEW Bootstrapping with Small Evaluation Keys, and Applications to Threshold Homomorphic Encryption
Yongwoo Lee 0002, Daniele Micciancio, Andrey Kim, Rakyong Choi, Maxim Anatolievich Deryabin, Jieun Eom, Donghoon Yoo
EUROCRYPT (3)7
2023 Area-Efficient Number Theoretic Transform Architecture for Homomorphic Encryption
abstract
Homomorphic encryption (HE) has emerged as an ideal cryptographic technology for meaningful computations on encrypted data. Not only does HE secure private information even if the ciphertext is leaked, but it also maintains data integrity when inferring cloud-side services. However, homomorphic computations include expensive polynomial arithmetic, especially polynomial multiplication. Prior studies proposed number theoretic transform (NTT) hardware designs to accelerate polynomial multiplication. However, the trade-off between hardware complexity and throughput of NTT designs was not considered carefully. This paper proposes an area-efficient NTT architecture suitable for HE schemes. Center of the proposed NTT architecture is a high-throughput butterfly unit array, which communicates with a single data memory unit through a conflict-free memory access pattern. Additionally, we developed a twiddle factor generator to reduce memory consumption. The proposed NTT architecture was successfully accelerated on the Xilinx FPGA devices. Performing with a large number of moduli, the proposed NTT design achieves higher hardware efficiency than the prior arts. Especially, our NTT design consumes less on-chip memory with efficiency improvement of$8.8\times $over the most related work. The implementation results confirm that our design methodology has advantages to deploy many NTT accelerators on an FPGA device for practical HE-based applications.
Phap Duong-Ngoc, Sunmin Kwon, Donghoon Yoo, Hanho Lee
IEEE Trans. Circuits Syst. I Regul. Pap.3
2018 Benzene: An Energy-Efficient Distributed Hybrid Cache Architecture for Manycore Systems
abstract
This article proposes Benzene, an energy-efficient distributed SRAM/STT-RAM hybrid cache for manycore systems running multiple applications. It is based on the observation that a naïve application of hybrid cache techniques to distributed caches in a manycore architecture suffers from limited energy reduction due to uneven utilization of scarce SRAM. We propose two-level optimization techniques: intra-bank and inter-bank. Intra-bank optimization leverages highly associative cache design, achieving more uniform distribution of writes within a bank. Inter-bank optimization evenly balances the amount of write-intensive data across the banks. Our evaluation results show that Benzene significantly reduces energy consumption of distributed hybrid caches.
Namhyung Kim, Junwhan Ahn, Kiyoung Choi, Daniel Sánchez 0003, Donghoon Yoo, Soojung Ryu
ACM Trans. Archit. Code Optim.5
2016 TQSIM: A fast cycle-approximate processor simulator based on QEMU
Shin-Haeng Kang, Donghoon Yoo, Soonhoi Ha
J. Syst. Archit.2
2014 Retargetable automatic generation of compound instructions for CGRA based reconfigurable processor applications
abstract
Reconfigurable processors such as SRP (Samsung Reconfigurable Processors) have become increasingly important, which enables just enough flexibility of accepting software solutions and providing application specific hardware configurability for faster time-to-market, lower development cost and higher performance while maintaining lower energy consumption and area. The reconfigurable processor compilation framework supports wide range of architectures through architecture description template for different domains of applications such as image processing, multimedia, video, and graphics. These architectures support several domain specific compound instructions (also called as intrinsics), which are computationally efficient when compared to the set of general instructions in the processor. Application developers have to use these intrinsics in their programs according to the architecture, which can result very inefficient usage, tedious and more error-prone. Moreover, the intrinsics provided by the architecture need constant reference to the intrinsics file during development. In this paper, we propose a retargetable novel methodology for the automatic generation of compound instructions for a given architecture and application source code at compile time. Our approach is able to consider ~75% of total intrinsics in the architectures with the success rate of > 90% in identifying the intrinsics in the benchmarks such as AVC, OpenGL Full Engine and OpenGL Vector benchmarks.
Narasinga Rao Miniskar, Soma Kohli, Haewoo Park, Donghoon Yoo
CASES4
2013 An OpenCL optimizing compiler for reconfigurable processors
abstract
This paper presents simple and efficient optimization techniques for an OpenCL compiler that targets reconfigurable processors. The target architecture consists of a generalpurpose processor core and an embedded reconfigurable accelerator with vector units. The accelerator is able to switch its architecture between the VLIW mode and the Coarse Grained Reconfigurable Array (CGRA) mode to achieve high performance. One big problem of this architecture is programming difficulty and OpenCL can be a good solution. However, since OpenCL does not guarantee performance portability, hardware dependent optimization is still necessary. Hence, we develop an OpenCL compiler framework that exploits the mode switching capability and vector units. To measure the effectiveness of the techniques, we have implemented the OpenCL framework and evaluate their performance with fourteen OpenCL benchmark applications.
Jeongho Nah, Hongjune Kim, Seok Joong Hwang, Donghoon Yoo, Jaejin Lee
FPT6
2012 Function inlining and loop unrolling for loop acceleration in reconfigurable processors
abstract
The next generation SoCs for consumer electronics need software solutions for faster time-to-market, lower development cost and higher performance while maintaining lower energy consumption and area. As a result, reconfigurable processors (RPs) have become increasingly important, which enables just enough exibility of accepting software solutions and providing application-specific hardware reconfigurability. Samsung Electronics has developed a reconfigurable processor called Samsung Reconfigurable Processor (SRP), which is the basis of our work. Though, the SRP is a powerful processor, it requires a smart and intelligent compiler to compile the application software while exploring its reconfigurable architecture. The existing compiler for the SRP does not support functional inlining and loop unrolling, and no study has yet been done on these optimizations for the RPs. In this paper, we study the impact of these optimizations on the performance of applications for the SRP processor and we also show how these optimizations are supported in the SRP compiler. We analyze the performance improvement due to these optimizations on various benchmarks namely Sobel Edge filter, JPEG decoder, and Luma Deblocking filter of the H.264 standard. Our experimental results have shown about 83% gain on performance with the functional inlining optimization and the loop unrolling optimization when compared to the original code for Sobel filter and JPEG encoder, and 11% gain on performance for Luma Deblock filter.
Narasinga Rao Miniskar, Pankaj Shailendra Gode, Soma Kohli, Donghoon Yoo
CASES4
2012 SCC based modulo scheduling for coarse-grained reconfigurable processors
abstract
Coarse-grained reconfigurable arrays (CGRAs) architectures aim to offer high performance at low power consumption, especially for digital signal processing and streaming applications. To fully exploit the computing capability of CGRA, it is essential to develop a scheduling algorithm which maps operations over processing elements in CGRA. Modulo scheduling [1] is known as the state-of-art algorithm for CGRA scheduling, and there are many variants [2][3][4]. However, they suffer from dealing with inter-iteration dependences called as recurrences that form cyclic dependences. Hence we propose a new scheduling technique that efficiently handles the cyclic dependences. The key techniques are grouping all the mutually-dependent recurrence cycles into a strongly connected component (SCC), and scheduling the input data flow graph (DFG) based on SCCs. Since grouping removes all the recurrence cycles from DFG, the resulting SCC graph becomes a form of directed acyclic graph (DAG) in which the scheduler can track the total order of SCCs. While processing SCCs one by one, our intra-SCC scheduler analyzes the dependences between every pair of two different operations inside of SCC and produces the schedule of them. Thanks to the well-structured form of the SCC-based graph, we obtain more efficient schedule compared to the previous CGRA scheduling algorithm [2]. The experimental results show that the proposed technique enhances the performance of recurrence-dominant loops up to 3.5X and raises the success rate of modulo-scheduling compared to the previous CGRA scheduling algorithm [2].
Wonsub Kim, Donghoon Yoo, Haewoo Park, Minwook Ahn
FPT2
2011 An instruction-scheduling-aware data partitioning technique for coarse-grained reconfigurable architectures
abstract
In this paper, we propose a data partitioning technique for the memory subsystem that consists of a multi-ported scratchpad memory (SPM) unit and a single-ported data cache in coarse-grained reconfigurable arrays (CGRA) architecture. The embedded reconfigurable processor executes programs by switching between the Non-VLIW and VLIW modes depending on the type of the code region to achieve high performance. The VLIW mode exploits code regions with high ILP that require high memory bandwidth and the Non-VLIW mode exploits those with low ILP that require low memory latency. Our data partitioning technique between the SPM and the data cache is based on data interference graph reduction and profiling information. Given an SPM size, it finds the optimal data partitions by taking the VLIW instruction schedule into consideration. We evaluate our data partitioning technique for the CGRA architecture with three representative multimedia applications.
Choonki Jang, Jaejin Lee, Hee-Seok Kim, Donghoon Yoo, Sukjin Kim, Hongseok Kim, Soojung Ryu
LCTES5