VLDB 2026 Research / reviewers in the wild / expert
Hanho Lee
dblp:26/539
· DBLP profile ↗
44ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0001-8815-1927ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unified Architecture of Random Sampler for Fully Homomorphic Encryption
Muhammad Ogin Hasanuddin, Muhammad Fajri Sachruddin, Hanho Lee |
ISCAS | 3 |
| 2026 | Time-Area Efficient RNS Base-conversion Architecture for HPS-BFV Homomorphic Encryption using Generalized Solinas Primes
Rella Mareta, Ardianto Satriawan, Hanho Lee |
ISCAS | 3 |
| 2025 | Base Conversion RNS using Hybrid Barret-Karatsuba-Based Modulus Prime Cores for Homomorphic MultiplicationabstractResidue Number Systems (RNS) offer a promising approach to enhancing the computational efficiency of Fully Homomorphic Encryption (FHE). This paper proposes an efficient base conversion method for RNS, leveraging parallel Processing Unit (PU) cores, where each PU core operates under different moduli, along with a Karatsuba-based Barrett modular multiplier to accelerate modular arithmetic in homomorphic encryption schemes. Our approach addresses the computational bottleneck of base conversions in FHE by utilizing optimized parallelism and efficient multiplication methods. Experimental results show that, with nearly the same latency for the computing unit, the proposed design can achieve higher throughput for implementing base conversion (BConv). The performance is further evaluated with different parallel core configurations (2, 3, and 6 cores), and these scenarios are compared against a single-core implementation. The design with 6 parallel cores achieves more than 5Gbps, demonstrating the significant benefit of increased parallelism for BConv in FHE operations. Muhammad Ogin Hasanuddin, Rafael Aditya Cahyo W, Infall Syafalni, Nana Sutisna, Hanho Lee, Trio Adiono |
ISCAS | 5 |
| 2025 | An Efficient NTT-Based Polynomial Multiplication Architecture for BFV Homomorphic EncryptionabstractHomomorphic encryption (HE) ensures data security and privacy in cloud computing by permitting operations on encrypted information without revealing the original data. Reliable high-speed polynomial multiplication is an important operation to build HE and other lattice-based cryptography. There are two approaches in realizing polynomial multiplication:first school book multiplication and NTT-based polynomial multiplication. This paper proposed efficient NTT-based polynomial multiplication using improved folding Barrett modular multiplication (BMM) for Brakerski/Fan-Vercauteren (BFV) HE. We also implemented three types of BMM and integrated them into the polynomial multiplication to see and compare the performance. The result shows our architecture can reach 1.6× faster TPS compared to previous works. Rella Mareta, Ardianto Satriawan, Hanho Lee |
ISCAS | 3 |
| 2025 | Exploring Possibilities of BFV-based Homomorphic Encryption for Privacy-Preserving Image ProcessingabstractWe explore how homomorphic encryption (HE) can be used for secure medical image processing. While the CKKS scheme is the obvious choice because it supports floating-point numbers, it has drawbacks like larger ciphertexts and slower performance compared to the BFV and BGV schemes. BFV and BGV are faster and more efficient, but they can only handle integers, which limits their ability to perform more advanced image processing tasks. To overcome this, we modify the BFV scheme to work with fixed-point arithmetic, allowing it to handle basic real-number calculations. We tested this approach on simple image processing tasks like brightness enhancement, greyscaling, and sharpening filter, showing it can securely process images in the encrypted domain. Ardianto Satriawan, Rella Mareta, Hanho Lee |
ISCAS | 3 |
| 2025 | Hybrid Number Theoretic Transform Architecture for Homomorphic EncryptionabstractFully homomorphic encryption (FHE) is an innovative cryptographic technology that has the potential to protect the privacy and confidentiality of data in the untrusted environments, such as public clouds or external parties. However, due to the inclusion of time-consuming polynomial arithmetic, FHE remains a challenge for computationally heavy applications. The number theoretic transform (NTT) is widely used in HE to reduce the complexity of polynomial multiplication. Therefore, implementing NTT in hardware for FHE has been explored in prior studies. However, due to the high hardware resource requirements, especially with a large number of moduli, hardware architecture supporting both NTT and its inverse transform (INTT) is still missing. This brief presents a hardware architecture for$2^{17}$NTT and INTT suitable for high-circuit depth CKKS-based HE schemes, satisfying both criteria of high speed and affordability for various FPGA platforms. The implementation results highlight that this design is area-efficient compared to the most related work and hardware-friendly for practical HE-based applications on FPGA devices. Quang Dang Truong, Phap Duong-Ngoc, Hanho Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | Compact 217 NTT Architecture for Fully Homomorphic EncryptionabstractHomomorphic encryption (HE) ensures data security and privacy in cloud computing by permitting operations on encrypted information without revealing the original data. However, many of these advances only achieve Partial or Some-what Homomorphic Encryption. Fully Homomorphic Encryption (FHE), which overcomes the number of operation limitations, is being actively researched, with notable schemes like BGV, GSW, and CKKS. Particularly, the CKKS scheme focuses on real and complex number computations but has big parameters requiring many hardware resources. This paper presents an efficient hardware architecture for the number theoretical transform (NTT) 217and an optimized Barrett-based integer modular multiplier to address these challenges. The designs improve the efficiency of polynomial multiplication and modular multiplication processes. Rella Mareta, Hanho Lee |
ISCAS | 2 |
| 2023 | CKKS-Based Homomorphic Encryption Architecture using Parallel NTT MultiplierabstractThis paper presents a high-throughput CKKS-based encryption architecture for homomorphic encryption. By de-ploying a parallel number theoretic transform (NTT) multiplier architecture, the polynomial multiplication is significantly accel-erated. Additionally, the modular multiplier is also improved by efficiently implementing using digital signal processing resources. The proposed NTT multiplier and homomorphic encryption architecture are evaluated using Xilinx Vivado and Xilinx XCU250 FPGA board. The evaluation results demonstrate that the proposed NTT multiplier helps improve the throughput of polynomial multiplication by at least 1.5 x compared to the most recent works. The efficiency of the proposal NTT multiplier, calculated by throughput per L UT or Slice, is much better than that of existing studies. The proposed homomorphic encryption architecture using the proposed NTT multiplier offers a high throughput of 32.7 Gbps. Tuy Nguyen Tan, Hanho Lee |
ISCAS | 3 |
| 2023 | Area-Efficient Number Theoretic Transform Architecture for Homomorphic EncryptionabstractHomomorphic encryption (HE) has emerged as an ideal cryptographic technology for meaningful computations on encrypted data. Not only does HE secure private information even if the ciphertext is leaked, but it also maintains data integrity when inferring cloud-side services. However, homomorphic computations include expensive polynomial arithmetic, especially polynomial multiplication. Prior studies proposed number theoretic transform (NTT) hardware designs to accelerate polynomial multiplication. However, the trade-off between hardware complexity and throughput of NTT designs was not considered carefully. This paper proposes an area-efficient NTT architecture suitable for HE schemes. Center of the proposed NTT architecture is a high-throughput butterfly unit array, which communicates with a single data memory unit through a conflict-free memory access pattern. Additionally, we developed a twiddle factor generator to reduce memory consumption. The proposed NTT architecture was successfully accelerated on the Xilinx FPGA devices. Performing with a large number of moduli, the proposed NTT design achieves higher hardware efficiency than the prior arts. Especially, our NTT design consumes less on-chip memory with efficiency improvement of$8.8\times $over the most related work. The implementation results confirm that our design methodology has advantages to deploy many NTT accelerators on an FPGA device for practical HE-based applications. Phap Duong-Ngoc, Sunmin Kwon, Donghoon Yoo, Hanho Lee |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | An Efficient Unified Polynomial Arithmetic Unit for CRYSTALS-DilithiumabstractThe CRYSTALS-Dilithium protocol is considered as one of the most promising digital signature schemes in NIST’s post-quantum cryptography standardization process. While separating arithmetic computation units can be advantageous in some cases, it can lead to increased hardware resource consumption and performance degradation. To overcome this issue, this paper proposes a novel architecture called the Unified Polynomial Arithmetic Unit (UniPAU), specifically designed for the Dilithium signature scheme. The proposed UniPAU offers a unique hardware module that can execute all the polynomial operations required for the Dilithium signature scheme. To demonstrate the effectiveness of our design, we implemented it on the Xilinx Zynq UltraScale+ ZCU102 (xczu9eg-ffvb1156-2-e) FPGA platform and evaluated its hardware efficiency and performance. Our implementation results indicate that the proposed UniPAU can achieve comparable throughput while consuming fewer hardware resources compared to state-of-the-art studies. These findings suggest that our UniPAU can provide an optimized and efficient hardware solution for polynomial arithmetic operations in the Dilithium signature scheme. Thang Xuan Pham, Phap Duong-Ngoc, Hanho Lee |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | High-Efficient Nonbinary LDPC Decoder with Early Layer Decoding ScheduleabstractIncreasing nonbinary low density parity check (NB-LDPC) decoder throughput is challenging. This paper considers nonbinary quasi-cyclic LDPC code features to propose an early layered decoding schedule. The proposed method can eliminate idle time introduced by emptying pipeline stages after each layered decoding process, as well as increase decoder throughput. Layout results using TSMC 90-nm CMOS technology confirm that the proposed decoding schedule improved throughput with almost the same hardware complexity compared to the state-of- the-art NB-LDPC decoder. In particular, the proposed approach achieved considerably improved throughput and efficiency compared with predecessors when both early layer decoding schedule and early decoding termination were enabled. Thang Xuan Pham, Hanho Lee |
ISCAS | 2 |
| 2019 | Low-complexity multi-mode multi-way split-row layered LDPC decoder for gigabit wireless communications
Tram Thi Bao Nguyen, Hanho Lee |
Integr. | 2 |
| 2019 | Half-row modified two-extra-column trellis min-max decoder architecture for nonbinary LDPC codes
Huyen Pham Thi, Hanho Lee, Xuan Nghia Pham |
Integr. | 2 |
| 2018 | Reduced-Complexity Trellis Min-Max Decoder for Non-Binary Ldpc CodesabstractIn this paper, a novel algorithm and corresponding reduced-complexity decoder architecture are proposed for decoding the trellis min-max NB-LDPC code. This proposal reduces the number of messages exchanged between check node and variable node as well as the hardware complexity. Thus, the memory requirement and the wiring congestion is decreased, which increases the throughput of the decoder with a negligible error-correcting performance loss. A layered decoder architecture is implemented for the (2304, 2048) NB-LDPC code over GF(16) based on the proposed algorithm with a 90-nm CMOS technology. The results show an area reduction of 19.4% for the check node unit, 26.56% for the whole decoder and a throughput of 1396 Mbps with almost similar error-correcting performance, compared to previous works. Huyen Pham Thi, Hanho Lee |
ICASSP | 2 |
| 2018 | High-Secure Low-Latency Ring-LWE Cryptography Scheme for Biomedical Images Storing and TransmittingabstractIn this paper, a low-latency scheme for biomedical images storing and transmitting using ring-learning with errors (ring-LWE) cryptography is presented. The proposed scheme significantly reduces the encryption time and decryption time for biomedical images compared to existing works. Particularly, the encryption time and decryption time of proposed ring-LWE scheme for biomedical images can be reduced up to 70.1% and 52.7% compared to previous studies. In addition, by storing biomedical images under encrypted form at the central server, biomedical data are completely protected. The analysis in similarity and entropy of encrypted image proves the outperformance of the proposed scheme in terms of security level compared to conventional schemes. Tuy Nguyen Tan, Hanho Lee |
ISCAS | 2 |
| 2018 | Basic-Set Trellis Min-Max Decoder Architecture for Nonbinary LDPC Codes With High-Order Galois FieldsabstractNonbinary low-density parity-check (NB-LDPC) codes outperform their binary counterparts in terms of error-correction performance. However, the drawback of NB-LDPC decoders is high complexity, especially for the check node unit (CNU), and the complexity increases considerably when increasing the Galois-field (GF) order. In this paper, a novel basic-set trellis min-max algorithm is proposed to greatly reduce not only the CNU complexity but also the number of messages exchanged between the check node and the variable node compared with previous studies, which is highly efficient for higher order GFs. In addition, the proposed CNU is designed to compute the messages in a parallel way. Layered decoder architectures based on the proposed algorithm were implemented for the (837, 726) NB-LDPC code over GF(32) and the (1512, 1323) code over GF(64) using 90-nm CMOS technology, and obtained a reduction in the complexity by 30% and 37% for the CNU, and 40% and 37.4% for the whole decoder, respectively. Moreover, the proposed decoder achieves a higher throughput at 1.67 Gbit/s and 1.4 Gbit/s compared with the other state-of-the-art high-rate NB-LDPC decoders with high-order GFs. Huyen Thi Pham, Hanho Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | A delay-efficient ring-LWE cryptography architecture for biometric securityabstractThis paper presents an efficient scheme and architecture based on ring-LWE cryptography for biometric data security. The proposed architecture significantly reduces the number of polynomial multipliers compared to the traditional ECC cryptosystems. As a result, the encryption and decryption time for a single ring-LWE operation can be reduced up to 49.7% and 19.6% compared to the existing works, respectively. The simulation results show that the proposed architecture outperforms conventional ECC architectures in terms of computation complexity and processing time. Tuy Nguyen Tan, Hanho Lee |
ISCAS | 2 |
| 2017 | High-throughput partial-parallel block-layered decoding architecture for nonbinary LDPC codes
Huyen Thi Pham, Sabooh Ajaz, Hanho Lee |
Integr. | 3 |
| 2017 | Two-Extra-Column Trellis Min-Max Decoder Architecture for Nonbinary LDPC CodesabstractIn this brief, a novel two-extra-column trellis min-max algorithm and the decoder architecture based on only the first minimum values are proposed for nonbinary low-density parity-check (NB-LDPC) codes. The algorithm greatly reduces the hardware complexity and improves the latency as well as the throughput of the proposed decoder architecture compared with the previous works. A layered decoder architecture based on the proposed algorithm for (837, 726) NB-LDPC code over GF(32) is implemented with a 90-nm CMOS technology. The results show a decrease in the area of 24.6% for the check node unit and 75.6% for the whole decoder with a throughput of 1.27 Gb/s. The proposed decoder provides a lower area and a higher efficiency compared with the state of the art of high-rate NB-LDPC codes with high Galois-field order. Huyen Thi Pham, Hanho Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Efficient multi-Gb/s multi-mode LDPC decoder architecture for IEEE 802.11ad applications
Sabooh Ajaz, Hanho Lee |
Integr. | 2 |
| 2013 | High-performance iterative BCH decoder architecture for 100 Gb/s optical communicationsabstractThis paper presents a iterative Bose-Chaudhuri-Hocquenghem (i-BCH) code and its high-speed decoder architecture for 100 Gb/s optical communications. The proposed architecture features a very high data processing rate as well as excellent error correction capability. The proposed 6-iteration i-BCH code structure with interleaving method allows the decoder to achieve 9.34 dB net coding gain performance at 10-15decoder output bit error rate to compensate for serious transmission quality degradation. The proposed high-speed i-BCH decoder architecture is synthesized using a 90-nm CMOS technology. It can operate at a clock frequency of 430 MHz and achieve a data processing rate of 100Gb/s. Thus, it has potential applications in next generation forward error correction (FEC) schemes for 100 Gb/s optical communications. Jewong Yeon, Hanho Lee |
ISCAS | 2 |
| 2013 | A High-Speed Low-Complexity Modified Radix-25 FFT Processor for High Rate WPAN ApplicationsabstractThis paper presents a high-speed low-complexity modified radix-25512-point fast Fourier transform (FFT) processor using an eight data-path pipelined approach for high rate wireless personal area network applications. A novel modified radix-25FFT algorithm that reduces the hardware complexity is proposed. This method can reduce the number of complex multiplications and the size of the twiddle factor memory. It also uses a complex constant multiplier instead of a complex Booth multiplier. The proposed FFT processor achieves a signal-to-quantization noise ratio of 35 dB at 12 bit internal word length. The proposed processor has been designed and implemented using 90-nm CMOS technology with a supply voltage of 1.2 V. The results demonstrate that the total gate count of the proposed FFT processor is 290 K. Furthermore, the highest throughput rate is up to 2.5 GS/s at 310 MHz while requiring much less hardware complexity. Taesang Cho, Hanho Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Block-Circulant RS-LDPC Code: Code Construction and Efficient Decoder DesignabstractThis brief presents a method for constructing block-circulant (BC) Reed-Solomon-based low-density parity-check (RS-LDPC) codes and an efficient decoder design. The proposed construction method results in a BC form of a parity-check matrix from a random parity-check matrix for RS-LDPC codes. A decoder architecture and switch network for BC-RS-LDPC code are then developed based on the new BC parity-check matrix. Thus, an efficient decoder architecture dedicated to a promising class of high-performance BC-RS-LDPC codes is presented for the first time. Moreover, a (2048, 1723) BC-RS-LDPC decoder architecture is designed to demonstrate the efficiency of the presented techniques. Synthesis results show that the proposed decoder requires 1.3-M gates and can operate at 450 MHz to achieve a data throughput of 41 Gb/s with eight iterations. Seong-In Hwang, Hanho Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Concatenated non-binary LDPC and HD-FEC codes for 100Gb/s optical transport systemsabstractIn this paper, we propose a soft-decision-based FEC scheme which is the concatenation of a non-binary LDPC (NB-LDPC) code and hard-decision FEC code having a compatibility with existing OTU-4 frame structure. The proposed concatenated NB-LDPC + RS frame structure is also provided and the entire frame size is 18,368 bytes. The proposed NB-LDPC(2304,2048) code over GF(24) provides a superior performance at the higher BER region and can be concatenated with HD-FEC code. The simulation result shows that the proposed concatenated NB-LDPC(2304,2048) and RS(255,239) FEC can provide a superior NCG performance over 10.3 dB at a post-FEC BER 10−15and over 10.8 dB with enhanced HD-FEC in the outer code. As a result, the proposed NB-LDPC codes are expected to be strong FEC candidates of soft-decision FEC for 100Gb/s optical transmission system. Chang-Seok Choi, Hanho Lee, Noriaki Kaneda, Young-Kai Chen |
ISCAS | 2 |
| 2012 | A novel method of constructing Quasi-Cyclic RS-LDPC codes for 10GBASE-T EthernetabstractThis paper presents a novel method of constructing of Quasi-Cyclic Reed-Solomon-based LDPC (QC-RS-LDPC) codes for the application of 10GBASE-T Ethernet. The proposed code construction method makes RS-LDPC codes to be quasi-cyclic codes without bit error rate performance degradation. Therefore, QC-RS-LDPC decoder using the proposed method can use Banyan network or QSN that are efficient switch networks for QC codes. For performance comparison, the switch networks have been implemented. The results show that the switch network using the proposed method requires less memory size than existing switch network like Benes network. Since the switch network optimized for QC codes reduces critical path delay, the clock speed of QC-RS-LDPC decoder can be improved. The proposed method is able to construct QC-RS-LDPC codes, which reduces hardware size and improves clock speed. Seong-In Hwang, Hanho Lee, Shin-Il Lim |
ISCAS | 2 |
| 2011 | A high-throughput LDPC decoder architecture for high-rate WPAN systemsabstractThis paper presents a high-throughput memory- efficient decoder architecture for Quasi-Cyclic Low-Density Parity-Check (QC-LDPC) codes in the high-rate wireless personal area network applications. Two novel techniques which can apply to our selected QC-LDPC codes are proposed, including four-parallel block layered decoding architecture and simplification of the switch networks. The proposed architecture based on a block parallel decoding scheme replaces a crossbar- based interconnect network with a fixed wire network for a switch network. In addition, two-stage pipelining is used to improve the clock speed. A 672-bit, rate-1/2 LDPC decoder is implemented using 90 nm CMOS technology. The design achieves an information throughput of 1.45 Gbps at a clock speed of 285 MHz with a maximum of 16 iterations. Kyung-Il Baek, Hanho Lee, Chang-Seok Choi, Gerald E. Sobelman |
ISCAS | 2 |
| 2011 | A high-speed low-complexity modified radix-25 FFT processor for gigabit WPAN applicationsabstractIn this paper, we present a novel modified radix-25algorithm for 512-point fast Fourier transform (FFT) computation and high-speed eight-parallel data-path architecture for multi-gigabit wireless personal area network (WPAN) systems. The proposed FFT processor can provide a high data throughput and low hardware complexity by using eight-parallel data-path and multi-path delay-feedback (MDF) structure. The modified radix-25FFT algorithm is also realized in our processor to reduce the number of complex multiplications and twiddle factor look-up tables. The proposed FFT processor has been designed and implemented with 90nm CMOS technology in a supply voltage of 1.2V. The proposed 512-point modified radix-25FFT/IFFT processor has a throughput rate of up to 2.8 GS/s at 350 MHz while requiring much smaller hardware complexity and low power consumption. Taesang Cho, Hanho Lee, Jounsup Park, Chulgyun Park |
ISCAS | 2 |
| 2011 | An area-efficient truncated inversionless Berlekamp-Massey architecture for Reed-Solomon decodersabstractThis paper presents a novel area-efficient truncated inversionless Berlekamp-Massey (TiBM) architecture for Reed- Solomon decoders. Especially this paper proposes how to truncate processing elements (PE) in order to reduce the hardware complexity of key equation solver (KES) block. The RS decoder using proposed TiBM architecture has been designed and implemented by 90-nm CMOS standard cell technology with a supply voltage of 1.1 V. The RS decoder using proposed TiBM architecture operates at a clock frequency of 400 MHz and has a throughput of 3.2 Gb/s. The proposed architecture requires approximately 25.4% fewer gate counts than architecture based on the conventional RiBM algorithm. Jeong-In Park, Hanho Lee, Seongsoo Lee |
ISCAS | 2 |
| 2011 | A Reduced-Complexity Architecture for LDPC Layered Decoding SchemesabstractA reduced-complexity low density parity check (LDPC) layered decoding architecture is proposed using an offset permutation scheme in the switch networks. This method requires only one shuffle network, rather than the two shuffle networks which are used in conventional designs. In addition, we use a block parallel decoding scheme by suitably mapping between required memory banks and processing units in order to increase the decoding throughput. The proposed architecture is realized for a 672-bit, rate-1/2 irregular LDPC code on a Xilinx Virtex-4 FPGA device. The design achieves an information throughput of 822 Mb/s at a clock speed of 335 MHz with a maximum of 8 iterations. Gerald E. Sobelman, Hanho Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | Two bit-level pipelined viterbi decoder for high-performance UWB applicationsabstractThis paper presents a high-speed low-complexity two bit-level pipelined Viterbi decoder architecture for MB-OFDM UWB systems. As the add-compare-select unit (ACSU) is the main bottleneck of the Viterbi decoder, this paper proposes two bit-level pipelined MSB-first ACSU, which is based on 2-step look-ahead techniques, to reduce a critical path of the ACSU. The proposed ACSU architecture requires approximately 12% fewer gate counts and 9% faster speed than the conventional MSB-first ACSU. The proposed Viterbi decoder was implemented with 0.18-μm CMOS standard cell technology in a supply voltage of 1.8V. It operates at a clock frequency of 870 MHz and has a throughput of 1.74 Gb/s. Yong-Je Goo, Hanho Lee |
ISCAS | 2 |
| 2008 | Adaptive quantization in min-sum based irregular LDPC decoderabstractIn this paper, we present adaptive quantization schemes in the normalized min-sum decoding algorithm considering scaling effects to improve the performance of irregular low-density parity-check (LDPC) decoder for WirelessMAN (IEEE 802.16e) applications. We discuss the finite precision effects on the performance of irregular LDPC codes and develop optimal finite word lengths of variables over an SNR. For floating point simulation, it is known that in the normalized min-sum or offset min-sum algorithms the performance of a min-sum based decoder is not sensitive to scaling in the log-likelihood ratio (LLR) values. However, when considering the finite precision for hardware implementation, the scaling affects the dynamic range of the LLR values. The proposed adaptive quantization approach provides the optimal performance in selecting suitable input LLR values to the decoder as far as the tradeoffs between error performance and hardware complexity are concerned. Gerald E. Sobelman, Hanho Lee |
ISCAS | 3 |
| 2008 | A high-speed four-parallel radix-24 FFT/IFFT processor for UWB applicationsabstractIn this paper, we present a novel high-speed low-complexity four data-path 128-point radix-24FFT/IFFT processor for high-throughput MB-OFDM UWB systems. The high radix radix-24multi-path delay feed-back (MDF) FFT architecture provides a higher throughput rate and low hardware complexity by using a four-parallel data-path scheme. A method for compensating the truncation error of fixed-width Booth multipliers with a Dadda reduction network is also employed, which maintains the input and output at 10-bit width with 33dB SQNR. This method leads to reduction of truncation errors compared with direct-truncated Booth multipliers. The proposed FFT/IFFT processor has been designed and implemented with 0.18-μm CMOS technology and a supply voltage of 1.8V. The proposed four-parallel FFT/IFFT processor has a throughput rate of up to 1.8 Gsample/s at 450 MHz while requiring much smaller hardware complexity. Minhyeok Shin, Hanho Lee |
ISCAS | 2 |
| 2007 | A High-Speed Pipelined Degree-Computationless Modified Euclidean Algorithm Architecture for Reed-Solomon DecodersabstractThis paper presents a novel high-speed low-complexity pipelined degree-computationless modified Euclidean (pDCME) algorithm architecture for high-speed RS decoders. The pDCME algorithm allows elimination of the degree-computation so as to reduce hardware complexity and obtain high-speed processing. A high-speed RS decoder based on the pDCME algorithm has been designed and implemented with 0.13-μm CMOS standard cell technology in a supply voltage of 1.1 V. The proposed RS decoder operates at a clock frequency of 660 MHz and has a throughput of 5.3 Gb/s. The proposed architecture requires approximately 15% fewer gate counts and a simpler control logic than architectures based on the popular modified Euclidean algorithm. Seungbeom Lee, Hanho Lee, Jongyoon Shin, Je-Soo Ko |
ISCAS | 2 |
| 2006 | A high-speed, low-complexity radix-24 FFT processor for MB-OFDM UWB systemsabstractThis paper presents the architecture design of a high-speed, low-complexity 128-point radix-24FFT processor for ultra-wideband (UWB) systems. The proposed high-speed, low-complexity FFT architecture can provide a higher throughput rate and low hardware complexity by using 2-parallel data-path scheme and single-path delay-feedback (SDF) structure. This paper presents the key ideas applied to the design of high-speed, low-complexity FFT processor, especially that for achieving high throughput rate and reducing hardware complexity. The proposed FFT processor has been designed and implemented with the 0.18-mum CMOS technology in a supply voltage of 1.8 V. The throughput rate of proposed FFT processor is up to 1 Gsample/s while it requires much smaller hardware complexity Jeesung Lee, Hanho Lee, Sang-in Cho, Sangsung Choi |
ISCAS | 2 |
| 2006 | A reconfigurable FIR filter design using dynamic partial reconfigurationabstractThis paper presents a novel partially reconfigurable FIR filter design that employs dynamic partial reconfiguration. Our scope is to implement a low-power, area-efficient autonomously reconfigurable digital signal processing architecture that is tailored for the realization of arbitrary response FIR filters. The implementation of design addresses area efficiency and flexibility allowing dynamically inserting and/or removing the partial modules. This design method shows the configuration time improvement by small configured slice and good area efficiency as compared to the method of conventional FIR filters. Yeong-Jae Oh, Hanho Lee, Chong Ho Lee |
ISCAS | 2 |
| 2006 | Implementation of a FIR Filter on a Partial Reconfigurable Platform
Hanho Lee, Chang-Seok Choi |
KES (3) | 1 |
| 2005 | An Evolvable Hardware System Under Uneven Environment
In Ja Jeon, Phill-Kyu Rhee, Hanho Lee |
KES (2) | 3 |
| 2005 | Reconfigurable Power-Aware Scalable Booth Multiplier
Hanho Lee |
KES (1) | 1 |
| 2003 | High-speed VLSI architecture for parallel Reed-Solomon decoderabstractThis paper presents high-speed parallel Reed-Solomon (RS) (255,239) decoder architecture using modified Euclidean algorithm for the high-speed multigigabit-per-second fiber optic systems. Pipelining and parallelizing allow inputs to be received at very high fiber-optic rates and outputs to be delivered at correspondingly high rates with minimum delay. A parallel processing architecture results in speed-ups of as much as or more than 10 Gb, since the maximum achievable clock frequency is generally bounded by the critical path of the modified Euclidean algorithm block. The parallel RS decoders have been designed and implemented with the 0.13-/spl mu/m CMOS standard cell technology in a supply voltage of 1.1 V. It is suggested that a parallel RS decoder, which can keep up with optical transmission rates, i.e., 10 Gb/s and beyond, could be implemented. The proposed channel = 4 parallel RS decoder operates at a clock frequency of 770 MHz and has a data processing rate of 26.6 Gb/s. Hanho Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2000 | VLSI design of Reed-Solomon decoder architecturesabstractThis paper presents VLSI implementations of an 8-error correcting (255, 239) Reed-Solomon (RS) decoder architecture for the optical fibre systems. We present the RS decoders using Euclidean and modified Euclidean algorithms which are regular and simple, and naturally suitable for VLSI implementation. We investigate hardware complexity, clock frequency and data processing rate for those RS decoders. The RS decoder based on the modified Euclidean algorithm operates at a clock frequency of 75 MHz and has a data processing rate of 600 Mbits/s in 0.25-/spl mu/m CMOS technology with a supply voltage of 2.5 V. Hanho Lee, Meng-Lin Yu, Leilei Song |
ISCAS | 1 |
| 1999 | A Compact Fast Variable Key Size Elliptic Curve Cryptosystem CoprocessorabstractElliptic curve (EC) cryptosystems have become more attractive due to their small key sizes and varieties of choices of the curves available. However, it is not efficient to implement them with a general-purpose microprocessor because of word size mismatch, less parallel computation, no hardware supported wire permutation and algorithm/architecture mismatch. The solution to this problem is to build a coprocessor. This coprocessor can be optimized for the algorithm of a particular application to enhance performance. Thus, the total hardware utilization can be kept at a very high rate and the computation is speeded up. A compact fast elliptic curve crypto coprocessor with variable key size is introduced, which utilizes the internal SRAM/registers in an FPGA. The generic hardware architecture for the coprocessor is implemented with a parameterized (in term of key size) VHDL description and is synthesized/mapped to a Xilinx FPGA. The algorithms adopted and the architecture developed are suitable for massively parallel computation. The experimental results show that the design can achieve a high utilization of CLBs for the Xilinx 4000 series. Sarvesh Shrivastava, Hanho Lee, Gerald E. Sobelman |
FCCM | 3 |
| 1998 | Digit-Serial DSP Library for Optimized FPGA ConfigurationabstractThis paper gives the digit-serial DSP libraries used to implement the digit-serial DSP architecture for field programmable gate arrays (FPGAs) and compares schematic-based FPGA design with design based on logic synthesis for digit-serial DSP libraries. It describes the design of digit-serial addition/subtraction, multiplication and delay elements and indicates also how digit-serial FIR filter can be implemented. The FPGA device utilization and critical path delay of digit-serial DSP libraries are calculated and described. Hanho Lee, Gerald E. Sobelman |
FCCM | 1 |
| 1998 | FPGA Logic Block Architecture for Digit-Serial DSP Applications (Abstract)abstractNo abstract available. Hanho Lee, Sarvesh Shrivastava, Gerald E. Sobelman |
FPGA | 1 |
| 1997 | A New Low-Voltage Full Adder CircuitabstractA new circuit based on combining XOR gates and double pass-transistor logic has been developed for implementing a full adder. The main design objectives for these new circuits are low power consumption and full-voltage swing at a low supply voltage. The proposed full adder circuit is compared with previously known circuits and is shown to provide superior performance. The new and previous full adder circuits have been fully simulated using HSPICE with 0.4 /spl mu/m CMOS technology at a 2.0 V supply voltage. An extensive analysis of a 8-bit carry-select adder establishes the superiority of the proposed circuit in that application. Hanho Lee, Gerald E. Sobelman |
Great Lakes Symposium on VLSI | 1 |