Ioannis Kouretas

dblp:19/3565 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0002-8574-8469ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 7 first-author · 3 since 2021Theory of computation · 4 · 2 since 2021
YearPublicationVenuePosition
2025 A Mixed-Precision RNS DNN Accelerator
abstract
The Residue Number System (RNS) has been used for the design of Deep Neural Network (DNN) processing architectures due to its efficient implementation of the multiply-accumulate (MAC) operation. Prior-art RNS DNN accelerators have demonstrated notable benefits compared to conventional fixed-point (FXP) representations for arithmetic precisions of at least 8 bits. However, advanced quantization techniques have recently enabled accurate ultra-low-precision FXP DNN inference. Thus, it remains an open research question whether RNS can still outperform FXP representations for smaller precisions and especially in mixed-precision (MXP) quantization settings, where optimal bit-width configurations with respect to overall accuracy drop constraints are sought. This work addresses this gap by presenting an RNS-based MXP DNN accelerator that supports 3–8-bit quantization and consistently achieves superior model performance vs. hardware cost tradeoffs for various DNN models, resulting in up to 1.2× energy efficiency improvements compared to the FXP counterpart. Synthesized on a 22-nm technology, the RNS MXP accelerator achieves 6.93–14.58 TOPS/W, outperforming the state-of-the-art uniform-precision RNS accelerator by 1.4× while maintaining the original model accuracy, as well as mixed-precision FXP accelerators.
Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis
ISCAS3
2023 Improving Residue-Level Sparsity in RNS-based Neural Network Hardware Accelerators via Regularization
abstract
Residue Number System (RNS) has recently attracted interest for the hardware implementation of inference in machine-learning systems as it provides promising trade-offs in the area, time, and power dissipation space. In this paper we introduce a technique that utilizes regularization during training, and increases the percentage of residues which are zero, when the parameters of an artificial neural network (ANN) are expressed in an RNS. The proposed technique can also be used as a post-processing stage, allowing the optimization of pre-trained models for RNS implementation. By increasing the number of residues being zero, i.e., residue-level sparsity, the proposed technique facilitates new hardware architectures for RNS-based inference, allowing new trade-offs and improving performance over prior art without practically compromising accuracy. The introduced method increases residue sparsity by a factor of 4× to 6× in certain cases.
Emmanouil Kavvousanos, Vasilis Sakellariou, Ioannis Kouretas, Vassilis Paliouras, Thanos Stouraitis
ARITH3
2023 A multiplier-Free RNS-Based CNN accelerator exploiting bit-Level sparsity
abstract
In this work, a Residue Numbering System (RNS)-based Convolutional Neural Network (CNN) accelerator utilizing a multiplier-free distributed-arithmetic Processing Element (PE) is proposed. A method for maximizing the utilization of the arithmetic hardware resources is presented. It leads to an increase of the system's throughput, by exploiting bit-level sparsity within the weight vectors. The proposed PE design takes advantage of the properties of RNS and Canonical Signed Digit (CSD) encoding to achieve higher energy efficiency and effective processing rate, without requiring any compression mechanism or introducing any approximation. An extensive design space exploration for various parameters (RNS base, PE micro-architecture, encoding) using analytical models as well as experimental results from CNN benchmarks is conducted and the various trade-offs are analyzed. A complete end-to-end RNS accelerator is developed based on the proposed PE. The introduced accelerator is compared to traditional binary and RNS counterparts as well as to other state-of-the-art systems. Implementation results in a 22-nm process show that the proposed PE can lead to 1.85× and 1.54× more energy-efficient processing compared to binary and conventional RNS, respectively, with a 1.88× maximum increase of effective throughput for the employed benchmarks. Compared to a state-of-the-art, all-digital, RNS-based system, the proposed accelerator is 8.87× and 1.11× more energy- and area-efficient, respectively.
Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis
ARITH3
2022 A High-performance RNS LSTM block
abstract
The Residue Number System (RNS) has been proposed as an alternative to conventional binary representations for use in AI hardware accelerators. While it has been successfully utilized in applications targeting Convolutional Neural Networks (CNNs), its usage in other network models such as Recurrent Neural Networks (RNNs) has been set back due to the difficulty of implementing more complex activations functions like tanh and sigmoid ($\sigma$) in the RNS domain. In this paper, we seek to extend its usage in such models, and in particular LSTM networks, by providing efficient RNS implementations of the activation functions. To this aim, we derive improved accuracy piecewise linear approximations of the tanh and $\sigma$ functions using the minimax approach and propose a fully RNS-based hardware realization. We show that our approximations can effectively mitigate accuracy degradation in LSTM networks compared to naive approximations, while the RNS LSTM block can be up to 40% more efficient in terms of performance per area unit compared to a binary counterpart, when used in high performance-targeted accelerators.
Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis
ISCAS3
2021 Simplified Hardware Implementation of Memoryless Dot Product for Neural Network Inference
abstract
In this paper a simplified hardware implementation of a dot product arithmetic operation with constant coefficients is presented. The proposed methodology exploits a combination of distributed arithmetic and common subexpression sharing techniques. An algorithm is introduced for identifying the common sub partial sums systematically. Subsequently, a hardware architecture is proposed and the obtained circuits are synthesized in a 90-nm 1.0 V CMOS standard-cell library using Synopsys Design Compiler. Comparisons reveal significant reduction of 52% and 23% in area and power respectively for 1.5 ns delay over a regular dot product constant multiplier.
Ioannis Kouretas, Vassilis Paliouras
ISCAS1
2020 Implementing the Residue Logarithmic Number System Using Interpolation and Cotransformation
abstract
The Residue Logarithmic Number System (RLNS) offers fast multiplication and division, but poses challenges for implementing addition and subtraction because the underlying integer Residue Number System (RNS) has slow sign detection. The conventional Binary Logarithmic Number Systems (BLNS) has benefited from interpolation and cotransformation. We propose a dual-path ALU that speculates about the sign detection to adapt interpolation and cotransformation to the limitations of RLNS. Synthesis shows for the same precision and technology, the area of the proposed RLNS circuit is similar to BLNS and much smaller than prior RLNS methods. We also compare against Floating Point (FP).
Mark G. Arnold, Vassilis Paliouras, Ioannis Kouretas
IEEE Trans. Computers3
2019 Under- and Overflow Detection in the Residue Logarithmic Number System
abstract
The Residue Number System (RNS) offers fast and cheap carry-free integer arithmetic but has slow and expensive overflow detection. The Logarithmic Number System (LNS) offers fast real multiplication, division and powers with floating-point-like relative precision. The Residue Logarithmic Number System (RLNS) is a combination of the two systems that offers advantages for moderate-precision real applications where a-priori analysis allows under-and overflow to be ignored. An arithmetic hardware generator is essential because of the mathematical obscurity of combining RNS and LNS. Unfortunately, real applications often underflow. We consider options to deal with under-and overflow using the RLNSTool generator as a foundation.
Mark G. Arnold, Ioannis Kouretas, Vassilis Paliouras, John R. Cowles
ARITH2
2016 Dynamic delay variation behaviour of RNS multiply-add architectures
abstract
In this paper we investigate the impact of intra- and inter-die variations on the delay sensitivity of certain Residue Number System (RNS) arithmetic circuits in comparison to ordinary binary arithmetic logic. The timing yield of systems that contain multiply-add units (MAC) is of great importance since they dominate important applications such as digital signal processing. Specifically, we employ two different delay models for the estimation of delay distributions of RNS and binary MAC architectures. Our analysis quantitatively proves that RNS MAC architectures that use bases of the form {2n- 1, 2n, 2n+ 1} demonstrate better normalized delay variation than binary MAC architectures to characterize both their static timing behaviour and the timing behaviour taking into account the sensitizable paths. Furthermore, it is shown that certain simplified RNS MAC architectures outperform conventional RNS MAC architectures in terms of the μ + α · σ delay variation metric.
Kleanthis Papachatzopoulos, Ioannis Kouretas, Vassilis Paliouras
ISCAS2
2013 Delay-variation-tolerant FIR filter architectures based on the Residue Number System
abstract
This paper investigates the use of the Residue Number System (RNS) in the hardware design of VLSI FIR filters implemented in nano-scale technologies prone to process variation effects. It is here shown that the RNS substantially reduces the filter sensitivity to delay variations, when compared to digital filter designs that use conventional positional number systems, such as the widely-used two's-complement representation. The inherent tolerance of the introduced RNS architectures to the delay variations, is here shown to allow to circumvent the use of large design parameter margins. Therefore, we demonstrate that the use of RNS can achieve a high timing yield without resorting to costly over-design, which may unnecessarily increase system complexity. The particular benefit comes in addition to area, time and power benefits achieved due to the use of the RNS. The quantitative digital filter design space exploration reported in the paper takes into consideration the filter order as well as criteria related to the filter output signal quality such as the signal-to-noise ratio (SNR) and it demonstrates that the proposed architectures offer effective solutions for hardware design using modern and future nano-scale processes, for filter cases of practical interest.
Ioannis Kouretas, Vassilis Paliouras
ISCAS1
2013 Low-Power Logarithmic Number System Addition/Subtraction and Their Impact on Digital Filters
abstract
This paper presents techniques for low-power addition/subtraction in the logarithmic number system (LNS) and quantifies their impact on digital filter VLSI implementation. The impact of partitioning the look-up tables required for LNS addition/subtraction on complexity, performance, and power dissipation of the corresponding circuits is quantified. Two design parameters are exploited to minimize complexity, namely the LNS base and the organization of the LNS word. A roundoff noise model is used to demonstrate the impact of base and word length on the signal-to-noise ratio of the output of finite impulse response (FIR) filters. In addition, techniques for the low-power implementation of an LNS multiply accumulate (MAC) units are investigated. Furthermore, it is shown that the proposed techniques can be extended to cotransformation-based circuits that employ interpolators. The results are demonstrated by evaluating the power dissipation, complexity and performance of several FIR filter configurations comprising one, two or four MAC units. Simulations of placed and routed VLSI LNS-based digital filters using a 90-nm 1.0 V CMOS standard-cell library reveal that significant power dissipation savings are possible by using optimized LNS circuits at no performance penalty, when compared to linear fixed-point two's-complement equivalents.
Ioannis Kouretas, Charalambos Basetas, Vassilis Paliouras
IEEE Trans. Computers1
2012 Residue arithmetic for designing multiply-add units in the presence of non-gaussian variation
abstract
In this paper the utilization of Residue Number System (RNS) is investigated as a tool for variation-tolerant design. In particular circuits using various RNS bases are compared to the equivalent binary structures in terms of their sensitivity to the variation of process parameters. Furthermore, RNS advantages are quantitatively illustrated by considering a timing model with two non-gaussian distributions. It is shown that for bases where all moduli channels are candidates to contain the critical path of the RNS circuit, the delay variation is significantly reduced when compared to the equivalent binary structures.
Ioannis Kouretas, Vassilis Paliouras
ISCAS1
2011 Towards a Quaternion Complex Logarithmic Number System
abstract
The well-known generalization of real to complex arithmetic (two reals) extends further to more obscure quaternion arithmetic (four reals), which has applications in signal processing, aerospace, graphics and virtual reality. Quaternion multiplication implements 3D rotation, but is expensive (usually 16 floating-point multiplications and 12 additions). This paper proposes an alternative quaternion representation using logarithms to reduce multiplication cost. The real Logarithmic Number System (LNS) allows fast and inexpensive multiplication and division in embedded and FPGA-based systems. Recent advances in the Complex LNS (CLNS) have made fast log-polar complex representation affordable. Although the quaternion logarithm function is also well-defined, it is not useful to simplify multiplication (in the same way real and complex logarithms are) because quaternion multiplication is not commutative but quaternion addition is. To overcome this, we propose a novel Quaternion Complex (QCLNS) representation using a pair of CLNS numbers. This representation implements quaternion multiplication using only the theoretical minimum, of 8 LNS multipliers (i.e., fixed-point adders) and two CLNS adders. Because CLNS numbers are more compact than ordinary rectangular complex representation, single-precision QCLNS occupies 10.9 percent less memory than conventional quaternion representation. Extrapolating conventional LNS and floating-point synthesis data from Fu et al., QCLNS saves on average 10 percent of FPGA resources for precisions between 13 and 45 bits.
Mark G. Arnold, John R. Cowles, Vassilis Paliouras, Ioannis Kouretas
IEEE Symposium on Computer Arithmetic4
2011 A Residue Logarithmic Number System ALU using interpolation and cotransformation
abstract
The Residue Logarithmic Number System (RLNS) uses the Residue Number System (RNS) to represent logarithms that represent real values. Multiplication and division are easy; reasonable-precision addition and subtraction have not been economical because of RNS sign-detection. This paper adapts novel interpolation (for addition) and cotransformation (for subtraction) to fit the sign-detection limits. Smaller than prior RLNS hardware, the novel ALU defers sign detection with two speculative datapaths.
Mark G. Arnold, Ioannis Kouretas, Vassilis Paliouras
ASAP2
2010 Residue arithmetic bases for reducing delay variation
abstract
In this paper the utilization of Residue Number System (RNS) is investigated as a tool for variation-tolerant design. In particular circuits using various RNS bases are compared in terms of their sensitivity to the variation of process parameters. Furthermore, RNS advantages are quantitatively illustrated by considering a timing model. It is shown that for bases where all moduli channels are candidates to contain the critical path of the RNS circuit, the delay variation is reduced upto 86% when compared to the equivalent binary structures.
Ioannis Kouretas, Vassilis Paliouras
ISCAS1
2009 Variation-tolerant Design Using Residue Number System
abstract
In this paper the use of residue arithmetic is proposed as a technique to reduce delay variation in adders. It is found that the use of residue arithmetic offers significant delay variation reduction when compared to adders of the literature. Therefore this technique can be used to control variance of critical paths delay and efficiently meet timing constraints and thus improve timing yield. Experiments conducted span several values of intra-die and die-to-die variance, so that cases of practical interest for various nanoscale technologies are covered.
Ioannis Kouretas, Vassilis Paliouras
DSD1
2008 Low-power logarithmic number system addition/subtraction and their impact on digital filters
abstract
This paper discusses techniques for low-power addition/subtraction in the logarithmic number system (LNS) and evaluates their impact on digital filter implementation. Initially, the impact of partitioning the look-up tables (LUT) required for addition/subtraction on complexity, performance, and power dissipation is studied. Subsequently techniques for the low-power implementation of an LNS multiply- accumulate (MAC) unit are investigated. The obtained LNS MACs are used for the design of digital filters. Synthesis of LNS-based digital filters using a 0.18 mum 1.8 V CMOS standard-cell library, reveal that significant power dissipation savings are possible at no performance penalty, when compared to linear two's-complement equivalent.
Ioannis Kouretas, Charalambos Basetas, Vassilis Paliouras
ISCAS1