Seok-Bum Ko

dblp:41/3994 · DBLP profile ↗
← Back
79ranked-venue papers
4as first author
31since 2021 · last 2026
0000-0002-9287-317XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 60 · 4 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Computer networks · 1Theory of computation · 1
YearPublicationVenuePosition
2026 When Posit Meets Microscaling: Energy Efficient Posit-Based Processing Element for Edge AI Computation
abstract
Low-precision computation is an effective method to improve energy efficiency when processing AI models at the edge. The design of numeric format is important to maintain good accuracy while reducing energy consumption. However, current fixed-point based formats or floating-point based formats have either limitations in representation range or precision, and thus efficiency or accuracy is compromised. Posit formats can achieve both large dynamic range and high precision, however, the computation overhead is too high. Inspired by the recent microscaling format, in this paper, a novel microscaling posit format and its corresponding dot-product based processing element are proposed. By designing a specific format, the dotproduct computation overhead of the original posit format is significantly reduced. Implementation results show that the proposed processing element can achieve up to $79 \%$ area reduction and $74 \%$ power reduction when compared with other designs available in the literature, which makes the proposed designs especially suitable for edge AI computation.
Seok-Bum Ko, Zhiqiang Wei 0002, Hao Zhang 0041
ASP-DAC2
2026 Power-Efficient and Reconfigurable Compute Unit for Multi-Precision AI Inference at the Edge
Muhammad Hamis Haider, Hao Zhang 0041, Seok-Bum Ko
ISCAS3
2026 LYRA: Low-Frequency Rank Adaptation via Factored DCT Coefficients for Parameter Efficient Fine Tuning of Transformers
Sayed Muhsin, Seok-Bum Ko
IEEE Signal Process. Lett.2
2026 FlexPWL: A Flexible, Scalable, and Multiplier-Free Approach for Activation Functions on FPGA
abstract
The hardware implementation of nonlinear activation functions (AFs), such as the Sigmoid and Hyperbolic Tangent (Tanh), presents a significant bottleneck for deploying Recurrent Neural Networks (RNNs) on resource-constrained edge devices. Field Programmable Gate Arrays (FPGAs), as a leading platform for edge AI, face challenges in efficiently executing these functions due to their complex mathematical nature and limited computational resources. This paper proposes a novel, hardware-friendly piecewise linear (PWL) approximation method for implementing Sigmoid and Tanh functions. Our approach introduces new mathematical formulations for computing the y-intercept that significantly reduce Mean Squared Error (MSE) and Maximum Absolute Error (MXE) at no additional hardware cost. By consolidating the features of recent works, the proposed architecture eliminates a costly segment address encoder, supports pipelining for low-latency and high-frequency operation, and avoids the complexity and low precision of other methods. Moreover, it offers high configurability in bit width, number of segments, and input ranges, enabling adaptable deployment across diverse hardware and precision targets. A unified architecture is presented for both AFs, maintaining identical FPGA resource usage across functions. Experimental results demonstrate that the proposed method outperforms prior state-of-the-art designs by up to 18.52× in accuracy, 437.29× in latency, and 3.71× in frequency, and achieves LUT and flip-flop savings of up to 11.63× and 12.42×, respectively.
Ebrahim Fard, Janier Arias-Garcia, Hao Zhang 0041, Seok-Bum Ko
IEEE Trans. Computers4
2026 LightIDS: a lightweight neural network-based intrusion detection system
Ebrahim Fard, Mahdi Soltani, Amir Hossein Jahangir, Seok-Bum Ko
J. Supercomput.4
2026 T3: Transformer Accelerator With Efficient Top-K Sorting and Dynamic Tanh Computation for Marine Edge Computing
abstract
The Transformer model has demonstrated superior performance across numerous marine applications. However, the complexity of self-attention and layer normalization in addition to the high precision requirement poses challenges for their efficient deployment in marine edge devices. To address this issue, in this article, an FPGA-based Transformer accelerator, T3, is proposed for marine edge computing. Two main techniques are proposed and utilized in the proposed T3accelerator. The first technique is the design of a floating-point (FP)-based top-$K$sorting method for self-attention pruning, which can be used to reduce the computational cost of self-attention modules. The other technique is an efficient implementation of the dynamic Tanh (DyT) module, which utilizes error-controlled piecewise linear (PWL) approximation and coefficient fusion method, which can be used to take the place of the costly layer normalization. The two proposed modules and the whole T3accelerator are implemented in Xilinx UltraScale+ FPGA devices. The proposed top-$K$sorting method can achieve up to 79.6% reduction in lookup tables (LUTs), 80.3% reduction in flip-flops (FFs), and 90.8% reduction in power consumption when compared with the previous top-$K$sorting method. The proposed DyT module can consume 38.3% fewer LUTs while achieving better accuracy compared to the state-of-the-art designs. Finally, due to the effectiveness of the proposed techniques, the proposed T3accelerator can achieve 15.2% higher energy efficiency while consuming fewer number of logic resources.
Dingyang Yu, Changlong Chen, Seok-Bum Ko, Hao Zhang 0041
IEEE Trans. Very Large Scale Integr. Syst.4
2025 Outdoor Check Stop HAZMAT Placard Detection Using Synthetic Images and YOLOv5-Small
abstract
Deep learning training is frequently supplemented with synthetic data. We propose a scheme to synthesize images of HAZMAT placards in outdoor environments (e.g., check stops) to train models for detection and classification. Our process has 3 levels of realism, with noise and other distortions, and can simulate day and nighttime images. We used our data to train YOLOv5-Small models and evaluated models over a test dataset of 4321 real outdoor check stop images containing 6738 placards across 16 classes. Respectively, the best models for similar and single-class placard groupings had 0.867 and 0.927 best half-precision test set [email protected]. Models run at 65 FPS on Nvidia’s AGX Xavier edge GPU single-board computer, and at 12 FPS on Nvidia’s Nano.
Riel Castro-Zunti, Juan Yepez, Seok-Bum Ko
ISCAS3
2025 Optimized COVID-19 detection using sparse deep learning models from multimodal imaging data
MohammadMahdi Moradi, Alireza Hassanzadeh, Arman Haghanifar, Seok-Bum Ko
Multim. Tools Appl.4
2025 A lightweight convolutional neural network based on U shape structure and attention mechanism for anterior mediastinum segmentation
Sina Soleimani Fard, Won Gi Jeong, Francis Ferri Ripalda, Hasti Sasani, Younhee Choi, S. Deiva, Gong Yong Jin, Seok-Bum Ko
Neural Comput. Appl.8
2024 Energy Efficient FPGA-Based Binary Transformer Accelerator for Edge Devices
abstract
Transformer-based large language models have gained much attention recently. Due to their superior performance, they are expected to take the place of conventional deep learning methods in many fields of applications, including edge computing. However, transformer models have even more amount of computations and parameters than convolutional neural networks which makes them challenging to be deployed at resource-constrained edge devices. To tackle this problem, in this paper, an efficient FPGA-based binary transformer accelerator is proposed. Within the proposed architecture, an energy efficient matrix multiplication decomposition method is proposed to reduce the amount of computation. Moreover, an efficient binarized Softmax computation method is also proposed to reduce the memory footprint during Softmax computation. The proposed architecture is implemented on Xilinx Zynq Untrascale+ device and implementation results show that the proposed matrix multiplication decomposition method can reduce up to 78% of computation at runtime. The proposed transformer accelerator can achieve improved throughput and energy efficiency compared to previous transformer accelerator designs.
Congpeng Du, Seok-Bum Ko, Hao Zhang 0041
ISCAS2
2024 Optimized Transformer Models: ℓ′ BERT with CNN-like Pruning and Quantization
abstract
Optimizing techniques for neural network architectures aimed at the edge are complex and intricate, which makes them non-universal. Edge computing and artificial intelligence overlap to enhance data security by enabling data processing at the source, mitigating any risk during data transfer. As data security concerns are growing among world governments, AI on edge has become a highly relevant field of modern research. There is a strict need to harness the power of Convolutional Neural Networks (CNNs) and Transformer networks on resource-constrained edge devices. Although many pruning and quantization techniques have been proposed for CNNs, they may not be directly applied to transformers due to the different computation patterns. This paper will explore the implications of two fundamental techniques: pruning and quantization. We will conduct a comparative analysis to explore the applicability of optimization techniques in Transformers, originally designed for CNNs, for real-world edge deployment. Experimental results show that significant improvement in compression ratio can be achieved while the accuracy of the transformer models is maintained.
Muhammad Hamis Haider, Stephany Valarezo-Plaza, Sayed Muhsin, Seok-Bum Ko
ISCAS5
2024 Anterior mediastinal nodular lesion segmentation from chest computed tomography imaging using UNet based neural network with attention mechanisms
Yi Wang 0064, Won Gi Jeong, Hao Zhang 0041, Younhee Choi, Gong Yong Jin, Seok-Bum Ko
Multim. Tools Appl.6
2024 Res-MGCA-SE: a lightweight convolutional neural network based on vision transformer for medical image classification
Sina Soleimani Fard, Seok-Bum Ko
Neural Comput. Appl.2
2024 Decoder Reduction Approximation Scheme for Booth Multipliers
abstract
Existing approximate Booth multipliers fail to keep up with modern approximate multipliers such as truncation-based approximate logarithmic multipliers. This paper introduces a new approximation scheme for Booth multipliers that can operate with negligible error rates using only$N/4$Booth decoders, instead of the traditional$N/2$Booth decoders. The proposed 16-bit BD16.4 approximate Booth multiplier reduces the Normalized Mean Error Deviation (NMED) by 96.5% and the Power-Area-Product (PAP) by 69.6%, when compared to a state-of-the-art approximate logarithmic multiplier. Additionally, the proposed BD16.4 approximate multiplier reduces the NMED by 94.4% and PAP by 74.8%, when compared to a state-of-the-art higher-radix approximate Booth multiplier. The proposed 8-bit approximate Booth multipliers reduce the NMED by up to 74% and PAP by up to 5% when compared to the existing state-of-the-art approximate logarithmic multipliers. We validated the results derived in this paper through a neural network inference experiment, where the proposed approximate multipliers showed a negligible drop in inference accuracy compared to the exact Booth multipliers and the state-of-the-art approximate logarithmic multipliers (ALM). The proposed approximate multipliers achieved a Power-Delay-Product reduction of 63% (vs. exact) and 21.22% (vs. ALM) in 16-bit experiments and a reduction of 67% (vs. exact) and 8.75% (vs. ALM) in 8-bit experiments.
Muhammad Hamis Haider, Hao Zhang 0041, Seok-Bum Ko
IEEE Trans. Computers3
2023 PaXNet: Tooth segmentation and dental caries detection in panoramic X-ray using ensemble transfer learning and capsule classifier
Arman Haghanifar, Mahdiyar Molahasani Majdabadi, Sina Haghanifar, Younhee Choi, Seok-Bum Ko
Multim. Tools Appl.5
2023 An automated multi-class skin lesion diagnosis by embedding local and global features of Dermoscopy images
Ravindranath Kadirappa, Deivalakshmi Subbian, Pandeeswari Ramasamy, Seok-Bum Ko
Multim. Tools Appl.4
2023 SaHNoC: an optimal energy efficient hybrid networks-on-chip architecture
Aravindhan Alagarsamy, Sundarakannan Mahilmaran, Gopalakrishnan Lakshminarayanan, Seok-Bum Ko
J. Supercomput.4
2022 Lightweight and CCA2-Secure Hardware Implementation of Binary Ring-LWE
abstract
Due to increasing the number of connected devices to IoT network, providing end-to-end security is essential. Lattice-based cryptography (LBC) is a promising method for IoT by providing the reasonable security against classic and quantum attacks. Binary Ring-LWE is a type of LBC that is suitable for IoT devices. However, a reliable cryptosystem should also be secure against different side-channel attacks, such as power analysis or fault injection ones. In this work, a fault resilient hardware implementation for an optimized hardware design of Ring Binary LWE for resource-constraint IoT devices is presented. The design was implemented on the FPGA platform. The maximum frequency and occupied slices on Virtex-7 are 210.5MHz and 423, respectively. Based on the result, the proposed design occupied only 1% of the total available slices.
Karim Shahbazi, Seok-Bum Ko
ISCAS2
2022 Optimizing a Multispectral-Images-Based DL Model, Through Feature Selection, Pruning and Quantization
abstract
The inclusion of technology in agriculture is highly relevant given the increasing global demand for food and our growing population. This paper focuses on an application for the analysis of multispectral images of wheat fields, and how they could be used to predict yield in a way that is resource-efficient and time convenient. Thus, our main goal is to optimize a deep learning model already proposed in the literature, through feature selection, pruning, and quantization, to be efficient enough that it could be deployed on a computer with limited resources. The main results of this work show that the size of the model was reduced by almost 94%, its inference time was almost 73% faster compared to the original model, while reducing its performance by 19% (still better than what was found in literature). This could be an important step towards the deployment of edge intelligence for plant phenotyping.
Julio Torres-Tello, Seok-Bum Ko
ISCAS2
2022 Factorized multi-scale multi-resolution residual network for single image deraining
Shivakanth Sujit, Deivalakshmi Subbian, Seok-Bum Ko
Appl. Intell.3
2022 Segmentation for document layout analysis: not dead yet
Logan Markewich, Hao Zhang 0041, Yubin Xing, Navid Lambert-Shirzad, Zhexin Jiang, Roy Ka-Wei Lee, Seok-Bum Ko
Int. J. Document Anal. Recognit.8
2022 FRDS: An efficient unique on-Chip interconnection network architecture
Aravindhan Alagarsamy, Sundarakannan Mahilmaran, Gopalakrishnan Lakshminarayanan, Seok-Bum Ko
Integr.4
2022 COVID-CXNet: Detecting COVID-19 in frontal chest X-ray images using deep learning
abstract
One of the primary clinical observations for screening the novel coronavirus is capturing a chest x-ray image. In most patients, a chest x-ray contains abnormalities, such as consolidation, resulting from COVID-19 viral pneumonia. In this study, research is conducted on efficiently detecting imaging features of this type of pneumonia using deep convolutional neural networks in a large dataset. It is demonstrated that simple models, alongside the majority of pretrained networks in the literature, focus on irrelevant features for decision-making. In this paper, numerous chest x-ray images from several sources are collected, and one of the largest publicly accessible datasets is prepared. Finally, using the transfer learning paradigm, the well-known CheXNet model is utilized to develop COVID-CXNet. This powerful model is capable of detecting the novel coronavirus pneumonia based on relevant and meaningful features with precise localization. COVID-CXNet is a step towards a fully automated and robust COVID-19 detection system.
Arman Haghanifar, Mahdiyar Molahasani Majdabadi, Younhee Choi, S. Deivalakshmi, Seok-Bum Ko
Multim. Tools Appl.5
2022 Capsule GAN for prostate MRI super-resolution
Mahdiyar Molahasani Majdabadi, Younhee Choi, S. Deivalakshmi, Seok-Bum Ko
Multim. Tools Appl.4
2022 Joint restoration convolutional neural network for low-quality image super resolution
Gadipudi Amaranageswarao, S. Deivalakshmi, Seok-Bum Ko
Vis. Comput.3
2021 Identifying Useful Features in Multispectral Images with Deep Learning for Optimizing Wheat Yield Prediction
abstract
Since unmanned aerial vehicles have been utilized in plant phenotyping, they have revolutionarily improved its accuracy. In this paper, we introduce a deep learning based approach for optimizing the yield prediction process of spring wheat (triticum aestivum), using multispectral images. We assessed both the temporal features to find the most valuable time to take images, as well as the contribution of spectral bands. We processed full stage multispectral images from four site-years (two sites during two years) of a wheat breeding project, and determined the prediction accuracy of the image-based predicted yields and compared them to the harvested yields taken in the field. The results compared the wheat images throughout the season and validated the most crucial flying times for acquiring images were at late-heading, late-flowering, dough-development, and harvesting stages. The two most useful colour-bands for yield prediction were red and red-edge. We found that removing these bands significantly decreased the prediction correctness. The results of this research could be a tool for the development of more efficient sensors and strategies for data collection in plant phenotyping.
Julio Torres-Tello, Seok-Bum Ko
ISCAS2
2021 Efficient Multiple-Precision Posit Multiplier
abstract
Posit number system has been recently widely applied in many fields of applications. For different applications, the precision requirements are usually different. In addition, the transprecision computing paradigm, which is proposed for energy efficient computation, even requires different precision in each computation step. To support computations of various precision in a single hardware architecture, in this paper, a unified architecture of multiple-precision posit multiplier is proposed. The proposed posit multiplier supports the commonly used Posit(8, 0), Posit(16, 1), and Posit(32, 2) formats, where one Posit(32, 2), or two parallel Posit(16, 1), or four parallel Posit(8, 0) multiplications can be accomplished each time. Each module of the proposed posit multiplier is carefully tailored for resource sharing among three supported precision formats. Compared to the Posit(32, 2) multiplier, the proposed multiple- precision multiplier adds the support for parallel low-precision posit multiplications with only 12.8% more area and 15.4% more power. The proposed architecture can be used in posit-enabled general-purpose processor designs.
Hao Zhang 0041, Seok-Bum Ko
ISCAS2
2021 Energy efficient spiking neural network processing using approximate arithmetic units and variable precision weights
Yi Wang 0064, Hao Zhang 0041, Kwang-Il Oh, Jae-Jin Lee, Seok-Bum Ko
J. Parallel Distributed Comput.5
2021 Enhancing the Utilization of Processing Elements in Spatial Deep Neural Network Accelerators
abstract
Equipping mobile platforms with deep learning applications is very valuable. Providing healthcare services in remote areas, improving privacy, and lowering needed communication bandwidth are the advantages of such platforms. Designing an efficient computation engine enhances the performance of these platforms while running deep neural networks (DNNs). Energy-efficient DNN accelerators use skipping sparsity and early negative output feature detection to prune the computations. Spatial DNN accelerators in principle can support computation-pruning techniques compared to other common architectures, such as systolic arrays. These accelerators need a separate data distribution fabric like buses or trees with support for high bandwidth to run the mentioned techniques efficiently and avoid network on chip (NoC)-based stalls. Spatial designs suffer from divergence and unequal work distribution. Therefore, applying computation-pruning techniques into a spatial design, which is even equipped with an NoC that supports high bandwidth for the processing elements (PEs), still causes stalls inside the computation engine. In a spatial architecture, the PEs that perform their tasks earlier have a slack time compared to others. In this article, we propose an architecture with a negligible area overhead based on sharing the scratchpads in a novel way between the PEs to use the available slack time caused by applying computation-pruning techniques or the used NoC format. With the use of our dataflow, a spatial engine can benefit from computation-pruning and data reuse techniques more efficiently. When compared to the reference design, our proposed method achieves a speedup of ×1.24 and an energy efficiency of ×1.18 per inference.
Mohammadreza Asadikouhanjani, Seok-Bum Ko
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 A Real-Time Architecture for Pruning the Effectual Computations in Deep Neural Networks
abstract
Integrating Deep Neural Networks (DNNs) into the Internet of Thing (IoT) devices could result in the emergence of complex sensing and recognition tasks that support a new era of human interactions with surrounding environments. However, DNNs are power-hungry, performing billions of computations in terms of one inference. Spatial DNN accelerators in principle can support computation-pruning techniques compared to other common architectures such as systolic arrays. Energy-efficient DNN accelerators skip bit-wise or word-wise sparsity in the input feature maps (ifmaps) and filter weights which means ineffectual computations are skipped. However, there is still room for pruning the effectual computations without reducing the accuracy of DNNs. In this paper, we propose a novel real-time architecture and dataflow by decomposing multiplications down to the bit level and pruning identical computations in spatial designs while running benchmark networks. The proposed architecture prunes identical computations by identifying identical bit values available in both ifmaps and filter weights without changing the accuracy of benchmark networks. When compared to the reference design, our proposed design achieves an average per layer speedup of$\times 1.4$and an energy efficiency of$\times 1.21$per inference while maintaining the accuracy of benchmark networks.
Mohammadreza Asadikouhanjani, Hao Zhang 0041, Gopalakrishnan Lakshminarayanan, Seok-Bum Ko
IEEE Trans. Circuits Syst. I Regul. Pap.5
2021 Area-Efficient Nano-AES Implementation for Internet-of-Things Devices
abstract
Due to the fast-growing number of connected tiny devices to the Internet of Things (IoT), providing end-to-end security is vital. Therefore, it is essential to design the cryptosystem based on the requirement of resource-constrained IoT devices. This article presents a lightweight advanced encryption standard (AES), a high-secure symmetric cryptography algorithm, implementation on field-programmable gate array (FPGA) and 65-nm technology for resource-constrained IoT devices. The proposed architecture includes 8-bit datapath and five main blocks. We design two specified register banks, Key-Register and State-Register, for storing the plain text, keys, and intermediate data. To reduce the area, Shift-Rows is embedded inside the State-Register. To adapt the Mix-Column to 8-bit datapath, we design an optimized 8-bit block for Mix-Columns with four internal registers, which accept 8-bit and send back 8-bit. Also, a shared optimized Sub-Bytes is employed for the key expansion phase and encryption phase. To optimize Sub-Bytes, we merge and simplify some parts of the Sub-Bytes. To reduce power consumption, we apply the clock gating technique to the design. Application-specific integrated circuit (ASIC) implementation results show a respective improvement in the area over the previous similar works from 35% to 2.4%. Based on the results, the proposed design is a suitable cryptosystem for tiny IoT devices.
Karim Shahbazi, Seok-Bum Ko
IEEE Trans. Very Large Scale Integr. Syst.2
2020 Epipolar Geometry on Drones Cameras for Swarm Robotics Applications
abstract
Nowadays, Convolutional Neural Networks are commonly used in classification of images due to their accuracy and performance. In robotics, cameras are used as the main sensor to be able to gather information of the environment; however, image processing could require great amounts of resources. An approach for object detection for 3D environment navigation was implemented, using onboard 2D cameras of mini-drones. A TensorFlow network was fine-tuned for specific object classification and epipolar geometry used to obtain measures of distance between three drones and the detected items. An average measurement accuracy of 0.6094 was obtained, with an average projection error of the cameras of 0.08779. The processing time for object prediction was approximately 0.02 seconds, which correspond to the 37% of the total time needed for a Node's iteration.
Andres Erazo, Eduardo Tayupanta, Seok-Bum Ko
ISCAS3
2020 Automated Teeth Extraction from Dental Panoramic X-Ray Images using Genetic Algorithm
abstract
Dental x-ray imaging helps dentists and radiologists to diagnose dental diseases and to provide patients with treatment plannings. In many cases, dental diseases are hard to detect by relying only on visual inspection. Therefore, automating the diagnosis process has been a topic of interest for dental problems. Teeth extraction is the basic task needed for nearly all dentistry decision support systems relying on radiographic images as the inputs. The most challenging type of image to perform extraction on is the panoramic image since it includes other parts of the patient's mouth, and structures lack explicit boundaries. The proposed method in this paper is the first automated teeth extraction system from dental panoramic images using evolutionary algorithms. First, the jaw is extracted from the main image. Then, upper and lower jaws are separated, followed by a genetic algorithm to detect teeth gap valleys. The method is assessed applying to 42 images, where the perceived accuracy is 81.14% for upper jaws and 73.63% for lower jaws, which is comparable with previous methods used on more straightforward image types.
Arman Haghanifar, Mahdiyar Molahasani Majdabadi, Seok-Bum Ko
ISCAS3
2020 Ensemble Learning for Improving Generalization in Aeroponics Yield Prediction
abstract
Agriculture plays a crucial role in economy of several countries and yield prediction is essential for production management and operation planning. Machine Learning (ML) is a growing trend in determining yield as a complex function of multiple input variables. Aeroponics is one of the efficient sustainable farming methods and allows all season farming despite hostile outdoors growing environment. In this paper, yield prediction in aeroponics is studied using ML. We have compared and analyzed three popular supervised ML methods - Dense Neural Network (DNN), Random Forest based on decision trees (RF) and Support Vector Regression (SVR). Air quality and water quality measurements including temperature, humidity, CO2, pH and Total Dissolved Solids (TDS) are used for yield prediction. Other static inputs such as number of days before and after transplant are also used. Six crops are studied (garlic chives, basil, red chard, rainbow chard, arugula, and mint). DNN performs particularly well with the prediction. The root mean square error (MSE), mean absolute error (MAE) and coefficient of determination (R2) are calculated to estimate the efficiency of the method. Mean square error and R2score of DNN are 0.10 and 0.67, RF follows DNN correctness with MSE and R2of 0.12 and 0.62, and SVR achieves 0.18 and 0.45 respectively, all of these values over the validation dataset. In addition to individual models, the two top performing models are combined as an ensemble model to improve overall performance, which shows an average R2score over the whole dataset divided by crop of 0.81.
Julio Torres-Tello, Suganthi Venkatachalam, Lyman Moreno, Seok-Bum Ko
ISCAS4
2020 Residual learning based densely connected deep dilated network for joint deblocking and super resolution
Gadipudi Amaranageswarao, S. Deivalakshmi, Seok-Bum Ko
Appl. Intell.3
2020 Wavelet based medical image super resolution using cross connected residual-in-dense grouped convolutional neural network
Gadipudi Amaranageswarao, S. Deivalakshmi, Seok-Bum Ko
J. Vis. Commun. Image Represent.3
2020 Capsule GAN for robust face super resolution
Mahdiyar Molahasani Majdabadi, Seok-Bum Ko
Multim. Tools Appl.2
2020 Blind compression artifact reduction using dense parallel convolutional neural network
Gadipudi Amaranageswarao, S. Deivalakshmi, Seok-Bum Ko
Signal Process. Image Commun.3
2020 Approximate Restoring Dividers Using Inexact Cells and Estimation From Partial Remainders
abstract
Approximate computing can be used in error-resilient applications to reduce power consumption and increase overall circuit performance. This article introduces two approximate dividers with restoring array-based architecture that achieve substantial hardware savings while maintaining high accuracy when compared to existing approximate designs. The first design replaces exact restoring divider cells with a proposed approximate cell in a column-wise fashion. The second design uses several rows of exact architecture to compute a partial remainder and then rounds and encodes the divisor and this partial remainder so that they may be used to express approximate outputs. A comprehensive accuracy and performance evaluation are performed for the proposed dividers as well as other state-of-the-art designs. When compared to an exact design, the proposed dividers have a reduced area and power consumption of 46 and 57 percent respectively while introducing minimal error. Furthermore, the trade-off between accuracy and improved performance is explored for various approximate dividers in order to determine which designs achieve the best compromise. The accuracy of the proposed dividers is then demonstrated using two image processing applications.
Elizabeth Adams, Suganthi Venkatachalam, Seok-Bum Ko
IEEE Trans. Computers3
2020 New Flexible Multiple-Precision Multiply-Accumulate Unit for Deep Neural Network Training and Inference
abstract
In this paper, a new flexible multiple-precision multiply-accumulate (MAC) unit is proposed for deep neural network training and inference. The proposed MAC unit supports both fixed-point operations and floating-point operations. For floating-point format, the proposed unit supports one 16-bit MAC operation or sum of two 8-bit multiplications plus a 16-bit addend. To make the proposed MAC unit more versatile, the bit-width of exponent and mantissa can be flexibly exchanged. By setting the bit-width of exponent to zero, the proposed MAC unit also supports fixed-point operations. For fixed-point format, the proposed unit supports one 16-bit MAC or sum of two 8-bit multiplications plus a 16-bit addend. Moreover, the proposed unit can be further divided to support sum of four 4-bit multiplications plus a 16-bit addend. At the lowest precision, the proposed MAC unit supports accumulating of eight 1-bit logic AND operations to enable the support of binary neural networks. Compared to the standard 16-bit half-precision MAC unit, the proposed MAC unit provides more flexibility with only 21.8 percent area overhead. Compared to a standard 32-bit single-precision MAC unit, the proposed MAC unit requires much less hardware cost but still provides 8-bit exponent in the numerical format to maintain large dynamic range for deep learning computing.
Hao Zhang 0041, Dongdong Chen 0002, Seok-Bum Ko
IEEE Trans. Computers3
2020 Stride 2 1-D, 2-D, and 3-D Winograd for Convolutional Neural Networks
abstract
Convolutional neural networks (CNNs) have been widely adopted for computer vision applications. CNNs require many multiplications, making their use expensive in terms of both computational complexity and hardware. An effective method to mitigate the number of required multiplications is via the Winograd algorithm. Previous implementations of CNNs based on Winograd use the 2-D algorithm F(2 × 2,3 × 3), which reduces computational complexity by a factor of 2.25 over regular convolution. However, current Winograd implementations only apply when using a stride (shift displacement of a kernel over an input) of 1. In this article, we presented a novel method to apply the Winograd algorithm to a stride of 2. This method is valid for one, two, or three dimensions. We also introduced new Winograd versions compatible with a kernel of size 3, 5, and 7. The algorithms were successfully implemented on an NVIDIA K20c GPU. Compared to regular convolutions, the implementations for stride 2 are 1.44 times faster for a 3 × 3 kernel, 2.04× faster for a 5 × 5 kernel, 2.42× faster for a 7 × 7 kernel, and 1.73× faster for a 3 × 3 × 3 kernel. Additionally, a CNN accelerator using a novel processing element (PE) performs two 2-D Winograd stride 1, or one 2-D Winograd stride 2, and operations per clock cycle was implemented on an Intel Arria-10 field-programmable gate array (FPGA). We accelerated the original and our proposed modified VGG-16 architectures and achieved digital signal processor (DSP) efficiencies of 1.22 giga operations per second (GOPS)/DSPs and 1.33 GOPS/DSPs, respectively.
Juan Yepez, Seok-Bum Ko
IEEE Trans. Very Large Scale Integr. Syst.2
2019 Energy-Efficient Approximate MAC Unit
abstract
Inexact computing generally involves trading a reduction in accuracy for an improvement in circuit area and power-consumption. The multiply-accumulate (MAC) operation is used extensively in convolutional neural networks and such applications stand to benefit greatly from the introduction of approximation to the MAC operation. This paper introduces an unsigned approximate MAC unit architecture in which approximation is introduced to both the multiplication and accumulation stages. Four variations of the proposed design are implemented using the TSMC 65 nm technology and are used in an image smoothing application. The proposed architecture is compared to that of the exact MAC unit and is shown to reduce circuit area and power-consumption by 67% and 49% respectively. When compared to other approximate MAC architectures, the proposed design improves area-power product by up to 66%.
Elizabeth Adams, Suganthi Venkatachalam, Seok-Bum Ko
ISCAS3
2019 Low-Cost 2-D Map Generation System for a Mobile Robot
abstract
This work is a proof of concept for an artificial vision system, which allows a robot to create two-dimensional maps of the area over which it moves, while detecting potentially dangerous objects. The proposed system uses a Kinect sensor for object detection, and an IP camera mounted on the roof of the test area for spatial location. Images from both sources are digitally processed on a Raspberry Pi 3 board, in order to locate the robot and potentially dangerous objects within the test area in real time. In this paper, an efficient and low cost system based on simple hardware (low cost, reduced size), with precise and real time distance measurement is proposed, as an aid for robot navigation.
Patricio Pérez, Julio Torres-Tello, Seok-Bum Ko
ISCAS3
2019 Design of Approximate Restoring Dividers
abstract
In this paper, two approximation models are proposed for restoring divider. In the first design, approximation is performed at circuit level, where exact restoring divider cells are replaced by approximate restoring divider cells by simplifying the logic equations. In the second model, restoring divider is analysed strategically and number of restoring divider cells are reduced by finding the portions of divisor and dividend with significant information. An approximation factor p is used in both designs. In model 1, the design with p = 8 has a 75% reduction in both area and power consumption compared to exact design, with a Q-MRED of 1.909 × 10-2 and Q-NMED of 0.449 × 10-2. The second model with an approximation factor p = 4 has 52% area savings and 61% power savings compared to exact design. The proposed models are found to have better error metrics compared to existing approximate designs, with better area and power savings at similar error values. A change detection image processing application is used for real time assessment of proposed and existing approximate dividers and one of the models achieves a PSNR of 54.27 dB.
Suganthi Venkatachalam, Elizabeth Adams, Seok-Bum Ko
ISCAS3
2019 Efficient Posit Multiply-Accumulate Unit Generator for Deep Learning Applications
abstract
The recently proposed posit number system is more accurate and can provide a wider dynamic range than the conventional IEEE754-2008 floating-point numbers. Its nonuniform data representation makes it suitable in deep learning applications. Posit adder and posit multiplier have been well developed recently in the literature. However, the use of posit in fused arithmetic unit has not been investigated yet. In order to facilitate the use of posit number format in deep learning applications, in this paper, an efficient architecture of posit multiply-accumulate (MAC) unit is proposed. Unlike IEEE754-2008 where four standard binary number formats are presented, the posit format is more flexible where the total bitwidth and exponent bitwidth can be any number. Therefore, in this proposed design, bitwidths of all datapath are parameterized and a posit MAC unit generator written in C language is proposed. The proposed generator can generate Verilog HDL code of posit MAC unit for any given total bitwidth and exponent bitwidth. The code generated by the generator is a combinational design, however a 5-stage pipeline strategy is also presented and analyzed in this paper. The worst case delay, area, and power consumption of the generated MAC unit under STM-28nm library with different bitwidth choices are provided and analyzed.
Hao Zhang 0041, Jiongrui He, Seok-Bum Ko
ISCAS3
2019 Design and Analysis of Area and Power Efficient Approximate Booth Multipliers
abstract
Approximate computing is an emerging technique in which power-efficient circuits are designed with reduced complexity in exchange for some loss in accuracy. Such circuits are suitable for applications in which high accuracy is not a strict requirement. Radix-4 modified Booth encoding is a popular multiplication algorithm which reduces the size of the partial product array by half. In this paper, three Approximate Booth Multiplier Models (ABM-M1, ABM-M2, and ABM-M3) are proposed in which approximate computing is applied to the radix-4 modified Booth algorithm. Each of the three designs features a unique approximation technique that involves both reducing the logic complexity of the Booth partial product generator and modifying the method of partial product accumulation. The proposed approximate multipliers are demonstrated to have better performance than existing approximate Booth multipliers in terms of accuracy and power. Compared to the exact Booth multiplier, ABM-M1 achieves up to a 23 percent reduction in area and 15 percent reduction in power with a Mean Relative Error Distance (MRED) value of 7:9 × 10-4. ABM-M2 has area and power savings of up to 51 and 46 percent respectively with a MRED of 2:7 × 10-2. ABM-M3 has area savings of up to 56 percent and power savings of up to 46 percent with a MRED of 3:4 × 10-3. The proposed designs are compared with the state-of-the-art existing multipliers and are found to outperform them in terms of area and power savings while maintaining high accuracy. The performance of the proposed designs are demonstrated using image transformation, matrix multiplication, and Finite Impulse Response (FIR) filtering applications.
Suganthi Venkatachalam, Elizabeth Adams, Seok-Bum Ko
IEEE Trans. Computers4
2019 Efficient Multiple-Precision Floating-Point Fused Multiply-Add with Mixed-Precision Support
abstract
In this paper, an efficient multiple-precision floating-point fused multiply-add (FMA) unit is proposed. The proposed FMA supports not only single-precision, double-precision, and quadruple-precision operations, as some previous works do, but also half-precision operations. The proposed FMA architecture can execute one quadruple-precision operation, or two parallel double-precision operations, or four parallel single-precision operations, or eight parallel half-precision operations every clock cycle. In addition to the support of normal FMA operations, the proposed FMA also supports mixed-precision FMA operations and mixed-precision dot-product operations. Specifically, the products of two lower precision multiplications can be accumulated to a higher precision addend. By setting the operands of one multiplication to zeros, the proposed FMA can also perform mixed-precision FMA operations. Support for mixed-precision FMA and mixed-precision dot-product is newly added but it only consumes 6.5 percent more area compared to a normal multiple-precision FMA unit. Compared to the state-of-the-art multiple-precision FMA design, the proposed FMA supports more floating-point operations such as half-precision FMA operations and mixed-precision operations with only 10.6 percent larger area.
Hao Zhang 0041, Dongdong Chen 0002, Seok-Bum Ko
IEEE Trans. Computers3
2018 Power Efficient Approximate Booth Multiplier
abstract
Power consumption is an important constraint in multimedia and deep learning applications. Approximate computing offers efficient approach to reduce power consumption. In this paper, novel approximation is proposed for radix-4 booth multiplication. Approximation is introduced in partial product generation and partial product accumulation circuits. Radix-4 partial product generation and accumulation approximation is proposed which remarkably enhances the performance. The proposed approximate booth multiplier achieves 41% area reduction and 49% power reduction compared to an exact booth multiplier. Also, it has better area, power and error metrics compared to existing works on approximate multipliers. The proposed multiplier is evaluated with an image processing application-in Discrete Cosine Transform (DCT) encoding part of JPEG compression and found to perform almost similar to exact multiplication unit.
Suganthi Venkatachalam, Seok-Bum Ko
ISCAS3
2018 An FPGA-based Closed-loop Approach of Angular Displacement for a Resolver-to-Digital-Converter
abstract
A closed-loop approach is proposed for the resolver-to-digital converter to obtain the accurate angular displacement. The proposed method can replace the analog conversion by efficient digital conversion with better accuracy. The approach is based on Co-ordinate Rotation Digital Computer (CORDIC) algorithm for fast calculation of trigonometric functions. This method employs two iterative strategies-direct iteration and continuous iteration to make the detection of the angular displacement fast. We analyze the discrete model using the iterative process, then implement and test on the FPGA Virtex-4QV. The design incorporates a photoelectric encoder and an FPGA which is feasible in high-temperature engineering applications while providing a better accuracy and high speed.
Juan Yepez, Seok-Bum Ko
ISCAS3
2018 Efficient Fixed/Floating-Point Merged Mixed-Precision Multiply-Accumulate Unit for Deep Learning Processors
abstract
Deep learning is getting more and more attentions in recent years. Many hardware architectures have been proposed for efficient implementation of deep neural network. The arithmetic unit, as a core processing part of the hardware architecture, can determine the functionality of the whole architecture. In this paper, an efficient fixed/floating-point merged multiply-accumulate unit for deep learning processor is proposed. The proposed architecture supports 16-bit half-precision floating-point multiplication with 32-bit single-precision accumulation for training operations of deep learning algorithm. In addition, within the same hardware, the proposed architecture also supports two parallel 8-bit fixed-point multiplications and accumulating the products to 32-bit fixed-point number. This will enable higher throughput for inference operations of deep learning algorithms. Compared to a half-precision multiply-accumulate unit (accumulating to single-precision), the proposed architecture has only 4.6% area overhead. With the proposed multiply-accumulate unit, the deep learning processor can support both training and high-throughput inference.
Hao Zhang 0041, Seok-Bum Ko
ISCAS3
2018 Approximate Sum-of-Products Designs Based on Distributed Arithmetic
Suganthi Venkatachalam, Seok-Bum Ko
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Design of Power and Area Efficient Approximate Multipliers
abstract
Approximate computing can decrease the design complexity with an increase in performance and power efficiency for error resilient applications. This brief deals with a new design approach for approximation of multipliers. The partial products of the multiplier are altered to introduce varying probability terms. Logic complexity of approximation is varied for the accumulation of altered partial products based on their probability. The proposed approximation is utilized in two variants of 16-bit multipliers. Synthesis results reveal that two proposed multipliers achieve power savings of 72% and 38%, respectively, compared to an exact multiplier. They have better precision when compared to existing approximate multipliers. Mean relative error figures are as low as 7.6% and 0.02% for the proposed approximate multipliers, which are better than the previous works. Performance of the proposed multipliers is evaluated with an image processing application, where one of the proposed models achieves the highest peak signal to noise ratio.
Suganthi Venkatachalam, Seok-Bum Ko
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Floating-Point Butterfly Architecture Based on Binary Signed-Digit Representation
abstract
Fast Fourier transform (FFT) coprocessor, having a significant impact on the performance of communication systems, has been a hot topic of research for many years. The FFT function consists of consecutive multiply add operations over complex numbers, dubbed as butterfly units. Applying floating-point (FP) arithmetic to FFT architectures, specifically butterfly units, has become more popular recently. It offloads compute-intensive tasks from general-purpose processors by dismissing FP concerns (e.g., scaling and overflow/underflow). However, the major downside of FP butterfly is its slowness in comparison with its fixed-point counterpart. This reveals the incentive to develop a high-speed FP butterfly architecture to mitigate FP slowness. This brief proposes a fast FP butterfly unit using a devised FP fused-dot-product-add (FDPA) unit, to compute AB ± CD ± E, based on binary-signed-digit (BSD) representation. The FP three-operand BSD adder and the FP BSD constant multiplier are the constituents of the proposed FDPA unit. A carry-limited BSD adder is proposed and used in the three-operand adder and the parallel BSD multiplier so as to improve the speed of the FDPA unit. Moreover, modified Booth encoding is used to accelerate the BSD multiplier. The synthesis results show that the proposed FP butterfly architecture is much faster than previous counterparts but at the cost of more area.
Amir Kaivani, Seok-Bum Ko
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Scalable Elliptic Curve Cryptosystem FPGA Processor for NIST Prime Curves
abstract
The architecture and the implementation of a high-performance scalable elliptic curve cryptography processor (ECP) are presented. The proposed ECP is able to support all five prime field elliptic curves recommended by the National Institute of Standards and Technology (NIST). The design takes advantage of the high-performance capabilities of the DSP48E slices available in Xilinx field-programmable gate arrays (FPGAs) to achieve high speed and low hardware resource utilization. The proposed design parallelizes the underlying prime field operations to reduce the latency of the elliptic curve point multiplication (ECPM) operation. Prime field inversion is performed efficiently using the same arithmetic blocks as the ones used for prime field multiplication and addition/subtraction. To the best of the authors' knowledge, the proposed scalable ECP is the fastest and smallest ECP that can support all five NIST recommended prime curves without the need to reconfigure the hardware. It can compute the ECPM between 1.709 and 28.04 ms using a Xilinx Virtex-5 FPGA.
K. C. Cinnati Loi, Seok-Bum Ko
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Bandwidth-aware routing and admission control for efficient video streaming over MANETs
Chhagan Lal, Vijay Laxmi, Manoj Singh Gaur, Seok-Bum Ko
Wirel. Networks4
2014 Highly adaptive and congestion-aware routing for 3D NoCs
abstract
In this paper, we propose a novel highly adaptive and congestion aware routing algorithm 3D meshes which is equally applicable to 2D meshes as well. The proposed algorithm allows cyclic dependencies in channel dependency graph (CDG) providing higher degree of adaptiveness. The algorithm uses congestion-aware channel selection strategy that results balanced distribution of traffic flows across the network. A packet follows non-minimal paths only when minimal paths are congested at the neighboring channels. The deadlock avoidance methodology adopted by our algorithm remains cost-efficient as it uses one extra virtual channel along each of Y and Z dimensions to achieve deadlock freedom.
Manoj Kumar 0001, Vijay Laxmi, Manoj Singh Gaur, Masoud Daneshtalab, Seok-Bum Ko, Mark Zwolinski
ACM Great Lakes Symposium on VLSI5
2014 High-speed FFT processors based on redundant number systems
abstract
Fast Fourier Transform (FFT) processors, having a significant impact on the performance of communication systems, have been a hot topic of research for many years. FFT function consists of consecutive multiply-add operations over complex numbers, dubbed as butterfly units. Use of redundant number systems is a way of increasing the speed of FFT coprocessors. It eliminates carry-propagation and hence permits latency reduction of each stage of the pipelined FFT architecture. This paper proposes a high-speed FFT processor using the devised fused-dot-product-add (FDPA) unit, to compute AB ± CD ± E, based on Binary-Signed-Digit (BSD) representation. Three-operand BSD adder and BSD constant multiplier are the constituents of the proposed FDPA unit. A carry-limited BSD adder is proposed and used in the three-operand adder and in the parallel BSD multiplier, so as to improve the speed of the FDPA unit. Moreover, modified-booth encoding is used to accelerate the BSD multiplier. Synthesis results show that the proposed design is about two times faster than the best previous work; but at cost of more area/power consumption.
Amir Kaivani, Seok-Bum Ko
ISCAS2
2014 FPGA implementation of low latency scalable Elliptic Curve Cryptosystem processor in GF(2m)
abstract
This paper presents the architecture of a scalable elliptic curve cryptography (ECC) processor (ECP). Two versions of scalable ECPs are presented, one for binary field pseudo-random curves and one for binary field Koblitz curves. The implementations of these designs are able to support all 5 key sizes of pseudo-random or Koblitz curves recommended by the National Institute of Standards and Technology (NIST) without reconfiguring the hardware. The paper proposes an architecture of a finite field multiplier that uses the Karatsuba-Ofman algorithm in order to reduce the latency of the finite field multiplication for larger key sizes. As a result, the latency of the overall elliptic curve point multiplication (ECPM) is reduced compared to previous designs of the scalable ECPs. To the authors' best knowledge, the proposed scalable ECPs are the fastest ECPs that can support all 5 pseudo-random or Koblitz curves recommended by NIST.
K. C. Cinnati Loi, Sen An, Seok-Bum Ko
ISCAS3
2014 A novel non-minimal/minimal turn model for highly adaptive routing in 2D NoCs
abstract
Networks-on-Chip (NoCs) are emerging as a promising communication paradigm to overcome bottleneck of traditional bus-based interconnects for current micro-architectures (MCSoC and CMP). One of the current issues in NoC routing is the use of acyclic Channel Dependency Graph (CDG) for deadlock freedom. This requirement forces certain routing turns to be prohibited, thus, reducing the degree of adaptiveness. In this paper, we propose a novel non-minimal turn model which allows cycles in CDG provided that Extended Channel Dependency Graph (ECDG) remains acyclic. The proposed turn model reduces number of restrictions on routing turns, hence able to provide path diversity through additional minimal and non-minimal routes between source and destination.
Manoj Kumar 0001, Vijay Laxmi, Manoj Singh Gaur, Masoud Daneshtalab, Pankaj Kumar Srivastava, Seok-Bum Ko, Mark Zwolinski
NOCS6
2013 High performance scalable elliptic curve cryptosystem processor in GF(2m)
abstract
The implementation of a scalable elliptic curve cryptography (ECC) processor is presented in this paper. The proposed ECC processor supports all 5 pseudo-random curves recommended by the National Institute of Standards and Technology (NIST) without the need to reconfigure the FPGA. The paper proposes a finite field arithmetic unit (FFAU) that reduces the number of clock cycles required to compute the elliptic curve point multiplication (ECPM) operation for ECC. The paper also presents a Lopez-Dahab algorithm with modified instructions to take advantage of the novel FFAU architecture. The completed scalable ECC processor (ECP) is implemented in hardware and a comparison analysis to the state-of-the-art designs is also discussed.
K. C. Cinnati Loi, Seok-Bum Ko
ISCAS2
2013 High-Speed Parallel Decimal Multiplication with Redundant Internal Encodings
abstract
The decimal multiplication is one of the most important decimal arithmetic operations which have a growing demand in the area of commercial, financial, and scientific computing. In this paper, we propose a parallel decimal multiplication algorithm with three components, which are a partial product generation, a partial product reduction, and a final digit-set conversion. First, a redundant number system is applied to recode not only the multiplier, but also multiples of the multiplicand in signed-digit (SD) numbers. Furthermore, we present a multioperand SD addition algorithm to reduce the partial product array. Finally, a digit-set conversion algorithm with a hybrid prefix network to decrease the number of the logic gates on the critical path is discussed. An analysis of the timing delay and an HDL model synthesized under 90 nm technology show that by considering the tradeoff of designs among three components, the overall delay of the proposed 16 × 16-digit multiplier takes about 11 percent less timing delay with 2 percent less area compared to the current fastest design.
Liu Han, Seok-Bum Ko
IEEE Trans. Computers2
2012 High-frequency sequential decimal multipliers
abstract
Multiplication, as one of the four basic operations embedded in arithmetic processors, is nowadays experiencing being spotlighted by the hardware designers involved in the revived decimal arithmetic. The decimal hardware units usually employ the sequential implementation for this operation, due to the high area cost of the parallel decimal multipliers. However, the main drawback of this iterative method is in regard to its high latency. This paper, with the intention of ameliorating this problem, proposes a high-frequency sequential decimal multiplier. The cycle time of the proposed multiplier is determined by a decimal carry-save adder which is about 22% less than that of the fastest previous design.
Amir Kaivani, Seok-Bum Ko
ISCAS3
2012 A low-power subsample-based image compression algorithm for capsule endoscopy
abstract
This paper presents an efficient sub-sample based image compression algorithm targeted to the endoscopic application. Endoscopic images are converted from RGB to YCgCo plane; the non-significant color components are then sub-sampled to obtain better compression ratio without heavily affecting the reconstruction quality. The algorithm uses simple integer-based Discrete Cosine Transform followed by a division-free quantization stage that results in low-cost implementation. The scheme is applied to both the traditional wide band images (WBI), as well as the narrow band images (NBI) for the performance assessment. The overall compression ratio and PSNR for the WBI and NBI are 84.53% and 82.36%, and 40.64 dB and 41.24 dB respectively. The hardware implementation is also presented that shows that the proposed scheme results in longer battery life compared to other existing schemes.
Atahar Mostafa, Khan A. Wahid, Seok-Bum Ko
ISCAS3
2012 Dynamic partial reconfigurable FFT/IFFT pruning for OFDM based Cognitive radio
abstract
Cognitive Radio is an application in which Spectrum utilization can be improved by allowing secondary users to use the spectrum when it is not used by licensed primary users. An adaptive OFDM system for Cognitive radio has the ability to nullify unnecessary individual carriers and avoid interference to licensed primary users. A Fast Fourier Transform (FFT) block forms the core of OFDM design. But, the zero valued inputs outnumber the non-zero valued inputs in the FFT block making the standard FFT algorithms computationally inefficient due to wasted operation on zero values. To overcome this problem, several pruning algorithms have been developed. But many of them are architecturally inefficient for FPGA implementation due to complexity of the overhead operations. Moreover, these algorithms are not suitable for applications like Cognitive radio which has zero inputs in arbitrary distributions making hardware implementation to be complex. This paper presents a novel and efficient dynamically partial reconfigurable (DPR) Transform Decomposition (TD) FFT and Radix 2 based IFFT pruning for OFDM based Cognitive Radio on FPGA. Tested FPGA results on XC2VP30 for the DPR method show the configuration time improvement, good area and power efficiency.
C. Vennila, Kumar Palaniappan CT, Kodati Vamsi Krishna, Gopalakrishnan Lakshminarayanan, Seok-Bum Ko
ISCAS5
2012 Design and implementation of a Radix-100 division unit
abstract
This paper presents a Radix-100 divider based on decimal non-restoring and selection by truncation method. Two decimal quotient digits can be selected in each iteration, which can reduce half of the iteration cycles. Initialization is required to scale the divisor into a pre-calculated range, and also used for generating some multiples of the scaled divisor. Implemented with STM 90-nm standard cells library, the proposed architecture takes 14 clock cycles, which is 373 FO4 to reach the desired accuracy. The latency is much shorter than Radix-10 dividers.
Liu Han, Seok-Bum Ko
ISCAS3
2012 Improved Decimal Floating-Point Logarithmic Converter Based on Selection by Rounding
abstract
This paper presents the algorithm and architecture of the decimal floating-point (DFP) logarithmic converter, based on the digit-recurrence algorithm with selection by rounding. The proposed approach can compute faithful DFP logarithm results for any one of the three DFP formats specified in the IEEE 754-2008 standard. In order to optimize the latency for the proposed design, we mainly integrate the following novel features: 1) using the redundant carry-save representation of the data path; 2) reducing the number of iterations by determining the number of initial iteration; and 3) retiming and balancing the delay of the proposed architecture. The proposed architecture is synthesized with STM 90-nm standard cell library and the results show that the critical path delay and the number of clock cycles of the proposed Decimal64 logarithmic converter are 1.55 ns (34.4 FO4) and 19, respectively, and the total hardware complexity is 43,572 NAND2 gates. The delay estimation results of the proposed architecture show that its latency is close to that of the binary radix-16 logarithmic converter, and that it has a significant decrease on latency compared with a recently published high performance CORDIC implementation.
Dongdong Chen 0002, Liu Han, Younhee Choi, Seok-Bum Ko
IEEE Trans. Computers4
2011 Nonspeculative decimal signed digit adder
abstract
Decimal floating point (DFP) arithmetic has been paid more attention in recent years, since it is superior to the binary counterpart in the financial and commercial computing including currency conversion, billing system, banking and tex calculation. Many DFP arithmetic units, such as addition, multiplication, division and fused-multi ply-add are not possible to achieve the high performance without a fast decimal fixed point adder. In this paper, the conventional four steps carry free signed digit addition algorithm is discussed. Furthermore, to improve the speed, we proposed a new method for the decimal SD addition and subtraction in digit set [-9,9]. To evaluate the design, a VHDL model is provided and synthesized in STM 90 nm technology. The result shows that our design has a better performance on timing delay and area compared with previous designs in the same digit set.
Liu Han, Dongdong Chen 0002, Khan A. Wahid, Seok-Bum Ko
ISCAS4
2011 Lossless implementation of Daubechies 8-tap wavelet transform
abstract
A new mapping scheme and its hardware implementation to error-freely compute the Daubechies 8-tap wavelet transform is presented. The multidimensional technique maps the irrational transform basis coefficients with integers and results in considerable reduction in hardware and power consumption. When implemented in Xilinx FPGA, the scheme costs 518 logic cells, 186 registers and runs at a frequency of 71MHz. While comparing with finite-precision architecture, the proposed scheme yields a reduction of 15% in hardware and 41% in power consumption for similar image reconstruction, and noticeable improvement in image reconstruction quality.
Khan A. Wahid, Seok-Bum Ko
ISCAS3
2010 A novel scalable parallel architecture for biological neural simulations
abstract
This paper presents a scalable hierarchical architecture for accelerating simulations of large-scale biological neural systems on FPGA-based platforms. The architecture provides a high degree of flexibility to optimize the parallelization ratio based on available hardware resources and model specifications such as complexity of dendritic trees. The proposed addressing scheme, design modularity and data process localization allowing the whole system to extend over multiple FPGA platforms to simulate a very large biological neural system. Compartmental approach and Hodgkin-Huxley methods are used as simulation models in our studies. The architecture is verified in MATLAB and implemented based on four types of hardware modules, with two modules synthesized on Xilinx XC5VLX110T-1 devices.
Peyman Pourhaj, Daniel H.-Y. Teng, Khan A. Wahid, Seok-Bum Ko
ISCAS4
2010 A high performance pseudo-multi-core ECC processor over GF(2163)
abstract
In this paper, we propose a high performance processor for elliptic curve cryptography (ECC) over GF(2163) by using polynomial presentation. It has three finite field (FF) RISC cores and a main controller to achieve instruction-level parallelism (ILP) with pipeline so that the largely parallelized algorithm for elliptic curve point multiplication can be well suited on this platform. Instructions for combined FF operation are proposed to decrease clock cycles in the instruction set. The interconnection among three FF cores and the main controller is obtained by analyzing the data dependency in the parallelized algorithm. The whole design is implemented on Xilinx XC4VLX80 FPGA device, and it can reach 185 MHz with 20,807 slices. The total time required for one ECC point scalar operation is 7.7μs in 1428 cycles.
Dongdong Chen 0002, Younhee Choi, Seok-Bum Ko
ISCAS5
2009 A 32-bit Decimal Floating-Point Logarithmic Converter
abstract
This paper presents a new design and implementation of a 32-bit decimal floating-point (DFP) logarithmic converter based on the digit-recurrence algorithm. The converter can calculate accurate logarithms of 32-bit DFP numbers which are defined in the IEEE 754-2008 standard. Redundant digit e1is obtained by look-up table in the first iteration and the rest redundant digits ejare selected by rounding the scaled remainder during the succeeding iterations. The sequential architecture of the proposed 32-bit DFP logarithmic converter is implemented on Xilinx Virtex-II Pro P30 FPGA device and then synthesized with TMSC 0.18-um standard cell library. The implementation results indicate that the maximum frequency of the proposed architecture is 47.7 MHz in FPGA and 107.9 MHz in TMSC 0.18-um technology. The faithful 32-bit DFP logarithm results can be obtained in 18 cycles.
Dongdong Chen 0002, Younhee Choi, Moon Ho Lee, Seok-Bum Ko
IEEE Symposium on Computer Arithmetic5
2009 A New Decimal Antilogarithmic Converter
abstract
This paper presents a new design and implementation of a 32-bit decimal floating-point (DFP) antilogarithmic converter based on the digit-recurrence algorithm with selection by rounding. The converter can calculate the accurate antilogarithm (10dec) of the 32-bit DFP numbers which are defined in the IEEE 754-2008 standard. The sequential architecture of the proposed 32-bit DFP antilogarithmic converter is implemented on Xilinx Virtex-II Pro P30 FPGA device. The proposed architecture occupies 2, 315 out of 13696(16%) slices and can obtain a faithful 32-bit DFP antilogarithm in 11 clock cycles running at 51.5 MHz. The 7-digit decimal fixed-point (FXP) antilogarithmic converter is an essential operational part of the 32-bit DFP antilogarithmic converter. We transform it to a 7-digit decimal exponential converter to compare with a 24-bit binary FXP exponential converter. The compared results show that the 7-digit decimal exponential converter occupies 2.18 times more area and 1.66 times slower than the 24-bit binary FXP exponential converter.
Dongdong Chen 0002, Daniel Teng, Khan A. Wahid, Moon Ho Lee, Seok-Bum Ko
ISCAS6
2009 Efficient Hardware Implementation of Hybrid Cosine-fourier-wavelet Transforms on a Single FPGA
abstract
This paper presents an efficient hardware implementation of a hybrid architecture to compute three 8-point transforms - the Discrete Cosine Transform, the Discrete Fourier Transform, and the Discrete Wavelet Transform on a single FPGA. The architecture is based on an element-wise matrix factorization and row-permutation algorithm, where the forward basis transformation matrices are decomposed into multiple sub-matrices and the common units are shared among them. The hardware implementation is parallel, pipelined and multiplication-free; it costs only 2,073 logic cells, 1,476 registers and runs at maximum frequency of 118 MHz with a very high process throughput of 944 Megabits/sec when synthesized onto an Altera FPGA device. The synthesized results for other FPGA technologies are also presented for performance assessment.
Khan A. Wahid, Samia Shimu, Daniel Teng, Moon Ho Lee, Seok-Bum Ko
ISCAS6
2008 Efficient hardware implementation of an image compressor for wireless capsule endoscopy applications
abstract
The paper presents an area- and power-efficient implementation of an image compressor for wireless capsule endoscopy application. The architecture uses a direct mapping to compute the two-dimensional Discrete Cosine Transform which eliminates the need of transpose operation and results in reduced area and low processing time. The algorithm has been modified to comply with the JPEG standard and the corresponding quantization tables have been developed and the architecture is implemented using the CMOS 0.18um technology. The processor costs less than 3.5k cells, runs at a maximum frequency of 150 MHz, and consumes 10 mW of power. The test results of several endoscopic colour images show that higher compression ratio (over 85%) can be achieved with high quality image reconstruction (over 30 dB).
Khan A. Wahid, Seok-Bum Ko, Daniel Teng
IJCNN2
2008 A novel decimal-to-decimal logarithmic converter
abstract
This paper presents a novel design and implementation of a 7-digit fixed-point decimal-to-decimal logarithmic converter. Two approaches, binary-based decimal approximation algorithm (Algorithm 1) and decimal linear approximation algorithm (Algorithm 2), are proposed and investigated. It shows that decimal linear approximation algorithm (Algorithm 2) is error-free in conversion between decimal and binary formats and also able to reduce maximum absolute error from binary-based Algorithm 1’s 0.00399 (integer cases) and 0.0483 (fraction cases) to 0.000994 (both cases). The Algorithm 2 is modeled in VHDL and implemented using combinational logic only in a Xilinx Virtex-II Pro P30 FPGA device. The logarithms results can be obtained in a single clock cycle, running at 50.9 MHz.
Dongdong Chen 0002, Younhee Choi, Daniel Teng, Khan A. Wahid, Seok-Bum Ko
ISCAS6
2004 Area Minimization of Exclusive-OR Intensive Circuits in FPGAs
Seok-Bum Ko
J. Electron. Test.1
2004 Efficient Realization of Parity Prediction Functions in FPGAs
Seok-Bum Ko, Jien-Chung Lo
J. Electron. Test.1
2002 Efficient Decomposition Techniques for FPGAs
Seok-Bum Ko, Jien-Chung Lo
HiPC1
2002 Studies of the SEMATECH IDDq test data
Seok-Bum Ko, Yu-Yau Guo, Jien-Chung Lo
J. Syst. Archit.1