EDBT 2026 Demo / reviewers in the wild / expert
Thanos Stouraitis
dblp:44/3113 · also Athanasios Stouraitis
· DBLP profile ↗
60ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-3696-4958ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-authorComputer networks · 4 · 2 since 2021Theory of computation · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient On-the-Fly Twiddle Factor Generation for Falcon PQC NTT/INTT
Ghada Alsuhli, Hani Saleh, Mahmoud Al-Qutayri, Baker Mohammad, Thanos Stouraitis |
ISCAS | 5 |
| 2026 | Casting Ventricular Arrhythmia Detection as Anomaly Detection via One-Class Meta-LearningabstractVentricular arrhythmia detection is a critical yet challenging task in cardiac healthcare due to the rarity of abnormal episodes and the high inter-patient variability in cardiac signals. These challenges are further exacerbated in implantable cardioverter-defibrillators, which operate under stringent memory and computational constraints. In this paper, we analyze inter- and intra-patient variability using dimensionality reduction and divergence metrics, and leverage these observations to formulate ventricular arrhythmia detection as a deployment-aligned one-class meta-learning problem. Accordingly, we adopt a one-class formulation of model-agnostic meta-learning (OC-MAML) with a clinically grounded task design that reflects real-world deployment conditions. Specifically, patient-disjoint support and query sets are used to simulate realistic distribution shifts and inter-patient variability. By training primarily on normal intracardiac electrogram segments, the OC-MAML-based framework learns a task-agnostic initialization that rapidly adapts to new patients using only a few normal samples, thereby substantially reducing dependence on labeled arrhythmic data. Compared to conventionalvanillamodel-agnostic meta-learning (MAML), our proposed OC-MAML-based framework achieves a relative improvement of +14.1% in sensitivity, +2.5% in balanced accuracy, and +4.1% inF1-score, while reducing adaptation time by 5× and maintaining comparable memory efficiency. These results underscore the framework’s potential for scalable, nearly label-free deployment in edge-based cardiac monitoring systems. The code for the proposed framework is publicly available at https://github.com/jaradat/VAD-OC-MAML. Abeer A. Jaradat, Hani Saleh, Omar Alhussein, Ghada Alsuhli, Thanos Stouraitis |
IEEE Internet Things J. | 5 |
| 2025 | Enhanced CNN Performance without Retraining Via Weight Approximation and Data ReuseabstractThis paper introduces an efficient CNN algorithm to address key limitations in Deep Neural Networks (DNNs) used for image recognition, focusing particularly on model size and retraining time. Traditional methods often require significant training durations; however, applying approximation techniques during retraining can exacerbate these time demands. We present an approach that enhances approximation techniques while eliminating the need for model retraining, thus enabling DNN compression with minimal accuracy loss. The proposed method integrates three core strategies: weight arrangement, approximation, and data reuse. The DNN weights are initially arranged in ascending order to optimize subsequent operations. During inference, the approximation is applied to reduce the model size and minimize computational complexity by reducing the number of operations required for each multiply-accumulate (MAC) unit. Then, the original weights are replaced with the approximated values, enabling the reuse of computations and data across different sets of weights. As a result, the method significantly reduces memory access, computational demands, and energy consumption. Experimental results on the CIFAR-10 and TinyImageNet datasets demonstrate that our method achieves a model reduction rate of approximately 198.6× while maintaining a minimal loss in accuracy. The proposed technique bypasses the need for retraining, offering a practical solution to the growing complexity of DNN models in modern applications. Mohamed F. Tolba 0002, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri, Thanos Stouraitis |
ISCAS | 5 |
| 2025 | A Mixed-Precision RNS DNN AcceleratorabstractThe Residue Number System (RNS) has been used for the design of Deep Neural Network (DNN) processing architectures due to its efficient implementation of the multiply-accumulate (MAC) operation. Prior-art RNS DNN accelerators have demonstrated notable benefits compared to conventional fixed-point (FXP) representations for arithmetic precisions of at least 8 bits. However, advanced quantization techniques have recently enabled accurate ultra-low-precision FXP DNN inference. Thus, it remains an open research question whether RNS can still outperform FXP representations for smaller precisions and especially in mixed-precision (MXP) quantization settings, where optimal bit-width configurations with respect to overall accuracy drop constraints are sought. This work addresses this gap by presenting an RNS-based MXP DNN accelerator that supports 3–8-bit quantization and consistently achieves superior model performance vs. hardware cost tradeoffs for various DNN models, resulting in up to 1.2× energy efficiency improvements compared to the FXP counterpart. Synthesized on a 22-nm technology, the RNS MXP accelerator achieves 6.93–14.58 TOPS/W, outperforming the state-of-the-art uniform-precision RNS accelerator by 1.4× while maintaining the original model accuracy, as well as mixed-precision FXP accelerators. Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis |
ISCAS | 5 |
| 2025 | Efficient NTT/INTT processor for FALCON post-quantum cryptographyabstractFALCON is a lattice-based post-quantum cryptographic (PQC) digital signature standard known for its compact signatures and resistance to quantum attacks. Since its recent standardization, its hardware implementation remains an open challenge, particularly for key generation, which is significantly more complex than the simple and well-studied signature verification process. In this paper, targeting edge devices with constrained resources, we present an energy-efficient and area-optimized NTT/INTT architecture tailored to the specific requirements of FALCON key generation. By leveraging NTT-friendly primes and reducing the size of the multipliers in the Montgomery reduction algorithm — optimized for ASIC implementation — our design minimizes hardware complexity, achieving the lowest power and area consumption compared to state-of-the-art Montgomery reduction implementations. The proposed hardware architecture features a processing element array, distributed SRAMs, and ROMs, with three levels of reconfigurability, supporting both NTT and INTT operations. Designed using the Global Foundries’ 22 nm FD-SOI process, an Application-Specific Integrated Circuit (ASIC) is estimated to occupy 0.04 mm 2 and consume 18.2 mW at 1 GHz. The proposed processor achieves 700 times greater energy efficiency and performs computations 200 times faster than software implementations on the ARM Cortex-M4. It also achieves the lowest area–time product and highest energy efficiency among state-of-the-art NTT/INTT hardware accelerators. By carefully balancing power consumption and computational speed, this design offers an efficient solution for deploying FALCON key generation on devices with limited resources. Ghada Alsuhli, Hani Saleh, Mahmoud Al-Qutayri, Baker Mohammad, Thanos Stouraitis |
J. Inf. Secur. Appl. | 5 |
| 2025 | A Survey and Comparative Analysis of Number Systems for Deep Neural NetworksabstractDeep neural networks (DNNs) are indispensable in various artificial intelligence (AI) applications. However, their inherent complexity presents significant challenges, particularly when deploying them on resource-constrained devices. To overcome these hurdles, academia and industry are actively seeking ways to accelerate and optimize DNN implementations. A significant area of research revolves around discovering more effective methods to represent the enormous data volumes processed by DNNs. Traditional number systems (NSs) have proven nonoptimal for this task, prompting extensive exploration into alternative and bespoke systems for DNNs. This survey aims to comprehensively discuss various NSs utilized to efficiently represent DNN data. These systems are categorized mainly based on their impact on DNN performance and hardware implementation. This survey offers an overview of these categorized NSs and delves into different subsystems within each, outlining their effect on DNN performance and hardware design. Furthermore, these systems are compared quantitatively and qualitatively concerning their expected quantization error, memory utilization, and computational requirements. This survey also emphasizes the challenges linked with each system and the diverse proposed solutions to address them. Insights into the utilization of these NSs for sophisticated DNNs are also presented in this survey. Readers will acquire a deeper understanding of the importance of efficient NSs for DNNs, explore commonly used systems, comprehend the tradeoffs between these systems, delve into design considerations influencing their impact on DNN performance, and discover recent trends and potential research avenues in this field. Ghada Alsuhli, Vasilis Sakellariou, Hani Saleh, Mahmoud Al-Qutayri, Baker Mohammad, Thanos Stouraitis |
Proc. IEEE | 6 |
| 2024 | Containerized Microservices: A Survey of Resource Management FrameworksabstractThe growing adoption of microservice architectures (MSAs) has led to major research and development efforts to address their challenges and improve their performance, reliability, and robustness. Important aspects of MSA that are not sufficiently covered in the open literature include efficient cloud resource allocation and optimal power management. Other aspects of MSA remain widely scattered in the literature, including cost analysis, service level agreements (SLAs), and demand-driven scaling. In this article, we examine recent cloud frameworks for containerized microservices with a focus on efficient resource utilization using auto-scaling. We classify these frameworks on the basis of their resource allocation models and underlying hardware resources. We highlight current MSA trends and identify workload-driven resource sharing within microservice meshes and SLA streamlining as two key areas for future microservice research. Lamees M. Al Qassem, Thanos Stouraitis, Ernesto Damiani, Ibrahim M. Elfadel |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2023 | Improving Residue-Level Sparsity in RNS-based Neural Network Hardware Accelerators via RegularizationabstractResidue Number System (RNS) has recently attracted interest for the hardware implementation of inference in machine-learning systems as it provides promising trade-offs in the area, time, and power dissipation space. In this paper we introduce a technique that utilizes regularization during training, and increases the percentage of residues which are zero, when the parameters of an artificial neural network (ANN) are expressed in an RNS. The proposed technique can also be used as a post-processing stage, allowing the optimization of pre-trained models for RNS implementation. By increasing the number of residues being zero, i.e., residue-level sparsity, the proposed technique facilitates new hardware architectures for RNS-based inference, allowing new trade-offs and improving performance over prior art without practically compromising accuracy. The introduced method increases residue sparsity by a factor of 4× to 6× in certain cases. Emmanouil Kavvousanos, Vasilis Sakellariou, Ioannis Kouretas, Vassilis Paliouras, Thanos Stouraitis |
ARITH | 5 |
| 2023 | A multiplier-Free RNS-Based CNN accelerator exploiting bit-Level sparsityabstractIn this work, a Residue Numbering System (RNS)-based Convolutional Neural Network (CNN) accelerator utilizing a multiplier-free distributed-arithmetic Processing Element (PE) is proposed. A method for maximizing the utilization of the arithmetic hardware resources is presented. It leads to an increase of the system's throughput, by exploiting bit-level sparsity within the weight vectors. The proposed PE design takes advantage of the properties of RNS and Canonical Signed Digit (CSD) encoding to achieve higher energy efficiency and effective processing rate, without requiring any compression mechanism or introducing any approximation. An extensive design space exploration for various parameters (RNS base, PE micro-architecture, encoding) using analytical models as well as experimental results from CNN benchmarks is conducted and the various trade-offs are analyzed. A complete end-to-end RNS accelerator is developed based on the proposed PE. The introduced accelerator is compared to traditional binary and RNS counterparts as well as to other state-of-the-art systems. Implementation results in a 22-nm process show that the proposed PE can lead to 1.85× and 1.54× more energy-efficient processing compared to binary and conventional RNS, respectively, with a 1.88× maximum increase of effective throughput for the employed benchmarks. Compared to a state-of-the-art, all-digital, RNS-based system, the proposed accelerator is 8.87× and 1.11× more energy- and area-efficient, respectively. Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis |
ARITH | 5 |
| 2023 | OSμS: An Open-Source Microservice Prototyping PlatformabstractOne major advantage of microservice cloud architectures is the agility with which microservices can be replicated to help improve the overall quality of service and meet service-level contracts. Their challenge is to carefully balance the horizontal microservice replicas with the vertical resources of CPU, memory, and IO that are allocated to each microservice. The objective of such balancing act is, of course, to avoid both service bottlenecks and resource wastage. In this paper, we present OSμS, a new open-source microservice prototyping platform that has been developed and instrumented from the ground up with the objective of collecting fine-grained, non-proprietary metrology on microservice mesh performance. We will illustrate the use of OSμS for developing and evaluating machine-learning algorithms for the horizontal and vertical autoscaling of microservice architectures. A hybrid algorithm based on decision-tree learning will be implemented on OSμS and compared with the academic state of the art and existing cloud-provider solutions. The advantages of such algorithm in improving horizontal and vertical resource utilization will be highlighted. Lamees M. Al Qassem, Thanos Stouraitis, Ernesto Damiani, Ibrahim M. Elfadel |
CloudCom | 2 |
| 2023 | EACNN: Efficient CNN Accelerator Utilizing Linear Approximation and Computation ReuseabstractThis paper proposes an efficient hardware accelerator named EACNN for use in Convolution Neural Networks. EACNN is an efficient CNN architecture that is based on co-optimization of algorithms and hardware. The proposed approach is based on linear approximation of the weights for pre-trained networks with low loss of accuracy. Furthermore, a weight substitution and remapping technique adopts linear approximation coefficients to replace CNN weights. That leads to a repetition of the weight values across different kernels and enables the reuse of CNN computations for various output feature maps. The input activations corresponding to the same linear co-efficient can be multiplied and accumulated first and then reused to generate multiple output feature maps. This computational reuse method reduces the number of multiplication and addition operations and memory accesses, which is efficiently supported by a dedicated element in the proposed EACNN. Experimental results on CIFAR 10 and CIFAR 100 datasets show that the proposed method eliminates around 61% of the multiplications in the network without significant loss of accuracy$(< 3\%)$. As a demonstration, a hardware accelerator based on EACNN was implemented on Xilinx FPGA Artix 7 and achieved a 50% reduction in the FPGA hardware resources. Mohamed F. Tolba 0002, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri, Thanos Stouraitis |
ISCAS | 5 |
| 2022 | A High-performance RNS LSTM blockabstractThe Residue Number System (RNS) has been proposed as an alternative to conventional binary representations for use in AI hardware accelerators. While it has been successfully utilized in applications targeting Convolutional Neural Networks (CNNs), its usage in other network models such as Recurrent Neural Networks (RNNs) has been set back due to the difficulty of implementing more complex activations functions like tanh and sigmoid ($\sigma$) in the RNS domain. In this paper, we seek to extend its usage in such models, and in particular LSTM networks, by providing efficient RNS implementations of the activation functions. To this aim, we derive improved accuracy piecewise linear approximations of the tanh and $\sigma$ functions using the minimax approach and propose a fully RNS-based hardware realization. We show that our approximations can effectively mitigate accuracy degradation in LSTM networks compared to naive approximations, while the RNS LSTM block can be up to 40% more efficient in terms of performance per area unit compared to a binary counterpart, when used in high performance-targeted accelerators. Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis |
ISCAS | 5 |
| 2022 | FPGAaaS: A Survey of Infrastructures and SystemsabstractThe popularity of cloud computing services for delivering and accessing infrastructure on demand has significantly increased over the last few years. Concurrently, the usage of FPGAs to accelerate compute-intensive applications has become more widespread in different computational domains due to their ability to achieve high throughput and predictable latency while providing programmability and improved energy efficiency. Computationally intensive applications such as big data analytics, machine learning, and video processing have been accelerated by FPGAs. With the exponential workload increase in data centers, major cloud service providers have made FPGAs and their capabilities available as cloud services. However, enabling FPGAs in the cloud is not a trivial task due to incompatibilities with existing cloud infrastructure and operational challenges related to abstraction, virtualization, partitioning, and security. In this article, we survey recent frameworks for offering FPGA hardware acceleration as a cloud service, classify them based on their virtualization mode, tenancy model, communication interface, software stack, and hardware infrastructure. We further highlight current FPGAaaS trends and identify FPGA resource sharing, security, and microservicing as important areas for future research. Lamees M. Al Qassem, Thanos Stouraitis, Ernesto Damiani, Ibrahim M. Elfadel |
IEEE Trans. Serv. Comput. | 2 |
| 2020 | A Remote FPGA Laboratory as a Cloud MicroserviceabstractIn this paper, we propose a web-based microservice for programming FPGAs remotely in the cloud without requiring additional hardware or third party software. Docker and Docker-compose are used to implement a scalable, secure environment for FPGA design and operation. The paper further describes in detail the architectural design, user interface, tool flow, and hardware back-end of the FPGA cloud service. The Docker files are made publicly available at GitHub. Lamees M. Al Qassem, Thanos Stouraitis, Ernesto Damiani, Ibrahim M. Elfadel |
ISCAS | 2 |
| 2018 | Performance Analysis of Single Carrier Coherent and Noncoherent Modulation under I/Q ImbalanceabstractIn-phase/quadrature-phase Imbalance (IQI) is considered a major performance-limiting impairment in direct-conversion transceivers. Its effects become even more pronounced at higher carrier frequencies such as the millimeter-wave frequency bands considered for 5G systems. In this work, we quantify the effects of IQI on the performance of different modulations under multipath fading channels. This is realized by developing a comprehensive framework for the symbol error rate (SER) analysis of coherent phase shift keying (PSK), noncoherent differential phase shift keying (DPSK) and noncoherent frequency shift keying (FSK) under IQI effects. In this context, the moment generating function of the signal-to-interference-plus-noise-ratio is first derived for single-carrier systems suffering from transmitter (TX) IQI only, receiver (RX) IQI only and joint TX/RX IQI. Capitalizing on this, we derive analytic expressions for the SER of the different modulation schemes considered. These expressions are corroborated with simulation results and they provide insights into the dependence of IQI on the system parameters. We further demonstrate that, while in some cases, IQI can cause a slight degradation of the SER performance and, hence, it can be neglected, in other cases it should be compensated in order to achieve a reliable communication link. Bassant Selim, Sami Muhaidat, Paschalis C. Sofotasios, Bayan S. Sharif, Thanos Stouraitis, George K. Karagiannidis, Naofal Al-Dhahir |
VTC Spring | 5 |
| 2018 | Outage probability of single carrier NOMA systems under I/Q imbalanceabstractNon-orthogonal multiple access (NOMA) has been recently proposed as a viable technology that has the potential to improve the spectral efficiency of fifth generation (5G) wireless networks and beyond. However, in practical communication scenarios, transceiver architectures inevitably suffer from radio-frequency (RF) front-end related impairments that can lead to non-negligible degradation of the overall system performance. In this context, in-phase/quadrature-phase imbalance (IQI) constitutes a major impairment in direct-conversion transceivers. Based on this, the present contribution quantifies the effects of IQI on the performance of NOMA based systems under multipath fading conditions. This is realized by first deriving novel analytic expressions for the signal-to-interference-plus-noise ratio and the outage probability of NOMA systems subject to IQI at the transmitter and/or the receiver sites. Capitalizing on these results, we demonstrate that the effects of IQI differ considerably between the different NOMA users and depending on the considered system's parameters. Bassant Selim, Sami Muhaidat, Paschalis C. Sofotasios, Bayan S. Sharif, Thanos Stouraitis, George K. Karagiannidis, Naofal Al-Dhahir |
WCNC | 5 |
| 2016 | Editorial for QShine 2014 Special Issue
Victor C. M. Leung, Jiangchuan Liu, Edith C. H. Ngai, Jianping Pan 0001, Thanos Stouraitis |
Mob. Networks Appl. | 5 |
| 2016 | A High-Speed FPGA Implementation of an RSD-Based ECC ProcessorabstractIn this paper, an exportable application-specific instruction-set elliptic curve cryptography processor based on redundant signed digit representation is proposed. The processor employs extensive pipelining techniques for Karatsuba-Ofman method to achieve high throughput multiplication. Furthermore, an efficient modular adder without comparison and a high-throughput modular divider, which results in a short datapath for maximized frequency, are implemented. The processor supports the recommended NIST curve P256 and is based on an extended NIST reduction scheme. The proposed processor performs single-point multiplication employing points in affine coordinates in 2.26 ms and runs at a maximum frequency of 160 MHz in Xilinx Virtex 5 (XC5VLX110T) field-programmable gate array. Hamad Marzouqi, Mahmoud Al-Qutayri, Khaled Salah 0001, Dimitrios M. Schinianakis, Thanos Stouraitis |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2014 | An RNS barrett modular multiplication architectureabstractExisting RNS implementations of modular multiplication employ Montgomery's technique, especially in cryptography where consecutive multiplications of large integers are used for exponentiation. This work deviates from these approaches and an RNS architecture of Barrett's modular multiplication algorithm is presented. An algorithmic and architectural comparison with the state-of-the-art solutions shows that the proposed algorithm, although requiring more modular multiplications, may achieve competitive total delay. Dimitrios M. Schinianakis, Thanos Stouraitis |
ISCAS | 2 |
| 2013 | Hardware-fault attack handling in RNS-based Montgomery multipliersabstractHardware-fault attacks have become a prominent threat against secure cipher implementations. Faults are deliberately introduced during the operation of cryptographic hardware so that, based on the faulty outputs, secret keys may be recovered. This work focuses on the RSA-CRT algorithm, which, although famous and widely exploited, is known to be vulnerable to hardware-fault attacks. Most of the counter measures, proposed in the literature for this algorithm, are based on number theory techniques that apply at a protocol level. In these cases, security is offered at the cost of extra operations in the RSA-CRT protocol. Unlike these solutions, this work examines the security potential offered by hardware implementations. It attempts to prove that the use of a well-designed, residue-arithmetic, Montgomery multiplier overcomes hardware-fault attack threats, with no need to alter the basic RSA-CRT protocol. Dimitrios M. Schinianakis, Thanos Stouraitis |
ISCAS | 2 |
| 2013 | Efficient RNS Implementation of Elliptic Curve Point Multiplication Over ${\rm GF}(p)$abstractElliptic curve point multiplication (ECPM) is one of the most critical operations in elliptic curve cryptography. In this brief, a new hardware architecture for ECPM over GF(p) is presented, based on the residue number system (RNS). The proposed architecture encompasses RNS bases with various word-lengths in order to efficiently implement RNS Montgomery multiplication. Two architectures with four and six pipeline stages are presented, targeted on area-efficient and fast RNS Montgomery multiplication designs, respectively. The fast version of the proposed ECPM architecture achieves higher speeds and the area-efficient version achieves better area-delay tradeoffs compared to state-of-the-art implementations. Mohammad Esmaeildoust, Dimitrios M. Schinianakis, Hamid Javashi, Thanos Stouraitis, Keivan Navi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | GF(2n) Montgomery multiplication using Polynomial Residue ArithmeticabstractA methodology for incorporating Polynomial Residue Arithmetic (PRA) in the Montgomery multiplication algorithm for polynomials in GF(2n) is presented in this paper. The mathematical conditions that need to be satisfied, in order for this incorporation to be valid are examined and performance results are given in terms of the field characteristic n, the number of moduli elements L, and the moduli word-length w. The proposed architecture is highly parallelizable and flexible, as it supports Polynomial-to-PRA and PRA-to-Polynomial conversions, Chinese Remainder Theorem (CRT) for polynomials, Montgomery multiplication, and Montgomery exponentiation in the same hardware. Dimitrios M. Schinianakis, Alexander Skavantzos, Thanos Stouraitis |
ISCAS | 3 |
| 2011 | A RNS Montgomery multiplication architectureabstractA novel algorithm and VLSI architecture for Residue Number System (RNS) Montgomery modular multiplication are presented in this paper. An analysis of binary-to- RNS and RNS-to-binary conversions along with the proposed RNS Montgomery multiplication reveals common datapaths and a unified add/multiply architecture is derived that supports all aforementioned operations in the same hardware. The proposed algorithm is fully executed in RNS and the cost of the input/output conversions is compensated by the speed up of operations due to the inherent parallelism of RNS. If used repeatedly, the proposed architecture supports modular exponentiation and modular inversion as well, thus forming an end-to-end alternative for cryptographic implementations. Dimitrios M. Schinianakis, Thanos Stouraitis |
ISCAS | 2 |
| 2010 | Comparison of time and frequency domain interpolation implementations for MB-OFDM UWB transmitters
Eleni Fotopoulou, Dorina Thanou, Thanos Stouraitis |
ISCAS | 3 |
| 2010 | Modeling and exploiting spatial locality trade-offs in wavelet-based applications under varying resource requirementsabstractFuture dynamic applications will require new mapping strategies to deliver power-efficient performance. Fully static design-time mappings will not be able to optimally address the unpredictably varying application characteristics and system resource requirements. Instead, the platforms will not only need to be programmable in terms of instruction set processors, but also at least partial reconfigurability will be required, while the applications themselves will need to exploit this increased freedom at runtime to adapt to the dynamism. In this context, it is important for applications to optimally exploit the memory hierarchy under varying memory availability. This article presents an analysis of spatial locality trade-offs in wavelet-based applications, to be used in dynamic execution environments: Depending on the encountered runtime conditions, the execution switches to different memory optimized instantiations or localizations, optimally exploiting temporal and spatial locality under these conditions. This is enabled by systematic mapping guidelines, indicating how the miss-rate behavior of a localization is influenced by a specific execution condition, under which conditions a certain localization is optimal and which miss-rate gains may be obtained by switching to that localization. Bert Geelen, Vissarion Ferentinos, Francky Catthoor, Gauthier Lafruit, Diederik Verkest, Rudy Lauwereins, Thanos Stouraitis |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2009 | Spatial locality exploitation for runtime reordering of JPEG2000 wavelet data layoutsabstractExploitation of spatial locality is essential for memories to increase the access bandwidth and to reduce the access-related latency and energy per word. Spatial locality exploitation of a kernel can be improved by modifying placement of data in memory, but this may be felt not only by the kernel itself, but also in other application components accessing the same data. Thus care is needed to avoid global miss-rate improvements are thwarted by miss-rate increases in other application components. This article examines application-level miss-rate increases due to handling modified Wavelet Transform data layouts by explicitly reordering at runtime, exploiting the execution order freedom within a reordering buffer when the layout of surrounding components is known. For the JPEG2000 application, taking into account the reordering costs still results in 80% net WT miss-rate gains. Bert Geelen, Vissarion Ferentinos, Francky Catthoor, Gauthier Lafruit, Diederik Verkest, Rudy Lauwereins, Thanos Stouraitis |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2007 | Content-Adaptive Wavelet-Based Scalable Video CodingabstractMotion-compensated temporal filtering (MCTF) is an essential ingredient of recently developed wavelet-based scalable video coding schemes. Lifting implementation of these decompositions represents a versatile tool for spatio-temporal optimizations and several improvements have already been proposed in this framework. Several coding parameters affect the performance of the scalable video coding scheme, such as the number of temporal levels and the interpolation filter used for sub-pixel accuracy. We show that the influence of these parameters depends on the video content. Thus, we present an adaptive way of choosing the value of these parameters, based on the video content. Experimental results show that the proposed method not only significantly improves the performance, but reduces the complexity of the coding procedure. Dionisis Athanasopoulos, Thanos Stouraitis |
ISCAS | 2 |
| 2006 | Execution time comparison of lifting-based 2D wavelet transforms implementations on a VLIW DSPabstractSeveral input-traversal schedules have been proposed for the computation of the 2D discrete wavelet transform (DWT). In this paper, the row-column, the line-based and the block-based schedules for the 2D DWT computation are compared with respect to their execution time on a very long instruction word (VLIW) digital signal processor (DSP). Implementations of the wavelet transform according to the considered schedules have been developed. They are parameterized with respect to filter pair, image size, and number of decomposition levels. All implementations have been mapped on a VLIW DSP. Performance metrics for the implementations for a complete set of parameters have been obtained and compared. The experimental results show that each implementation performs better for different points of the parameter space Kostas Masselos, Yiannis Andreopoulos, Thanos Stouraitis |
ISCAS | 3 |
| 2006 | An RNS architecture of an Fp elliptic curve point multiplierabstractAn elliptic curve point multiplier (ECPM) is the main part of all elliptic curve cryptography (ECC) systems and its performance is decisive for the performance of the overall cryptosystem. A VLSI residue number system (RNS) architecture of an ECPM is presented in this paper. In the proposed approach, the necessary mathematical conditions that need to be satisfied, in order to replace typical finite field circuits with RNS ones, are investigated. It is shown that such an application is feasible and that it leads to a significant improvement in the execution time of a scalar point multiplication Dimitrios M. Schinianakis, Apostolos P. Fournaris, Athanasios Kakarountas, Thanos Stouraitis |
ISCAS | 4 |
| 2005 | Hidden messages in heavy-tails: DCT-domain watermark detection using alpha-stable modelsabstractThis paper addresses issues that arise in copyright protection systems of digital images, which employ blind watermark verification structures in the discrete cosine transform (DCT) domain. First, we observe that statistical distributions with heavy algebraic tails, such as the alpha-stable family, are in many cases more accurate modeling tools for the DCT coefficients of JPEG-analyzed images than families with exponential tails such as the generalized Gaussian. Motivated by our modeling results, we then design a new processor for blind watermark detection using the Cauchy member of the alpha-stable family. The Cauchy distribution is chosen because it is the only non-Gaussian symmetric alpha-stable distribution that exists in closed form and also because it leads to the design of a nearly optimum detector with robust detection performance. We analyze the performance of the new detector in terms of the associated probabilities of detection and false alarm and we compare it to the performance of the generalized Gaussian detector by performing experiments with various test images. Alexia Briassouli, Panagiotis Tsakalides, Thanos Stouraitis |
IEEE Trans. Multim. | 3 |
| 2003 | Power efficient data path synthesis of sum-of-products computationsabstractTechniques for the power efficient data path synthesis of sum-of-products computations between data and coefficients are presented. The proposed techniques exploit specific features of this type of computations. Efficient heuristics for the scheduling and assignment tasks, based on the concept of the Traveling Salesman's Problem, are described. Different cost functions are proposed to drive the synthesis tasks. The proposed cost functions target the power consumption either in the interconnect buses or in the functional units. Experimental results from different relevant digital signal processing algorithmic kernels prove that the proposed synthesis techniques lead to significant power savings. Kostas Masselos, Panagiotis Merakos, Spyros Theoharis, Thanos Stouraitis, Constantinos E. Goutis |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2003 | New power-of-2 RNS scaling scheme for cell-based IC designabstractPrevious scaling schemes are based on the conversion of the unpositional residue number system (RNS) digits into a positional number system via Chinese remainder theorem (CRT) or mixed-radix-conversion (MRC) and the back conversion into RNS with an associated size and speed penalty in cell-based integrated circuit (CBIC) designs. This paper presents a new scaling approach, which allows faster and more efficient schemes, because the scaling uses only RNS operations within the small word length channels. Uwe Meyer-Bäse, Thanos Stouraitis |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2002 | Multi-voltage low power convolvers using the polynomial residue number systemabstractA novel approach for the reduction of the power dissipated in a signal processing application is introduced in this paper. By exploiting the properties of the Polynomial Residue Number System (PRNS) and of the arithmetic modulo (2n+1), the power dissipation of implementing cyclic convolution is reduced up to four times. Furthermore, the corresponding power x delay product is reduced up to 2.4 times, while a simultaneous reduction of area cost is achieved. The particular performance improvement becomes possible by introducing a way to minimize the forward and inverse conversion overhead associated with PRNS. The introduced minimization exploits the fact that for the conversions for particular lengths of data sequences and particular moduli, only multiplications with powers of two and additions are required, thus leading to low implementation complexity. In addition multiple supply voltages are utilized to further reduce power dissipation by more than 30% for particular cases. Formulas that return the applicable supply voltage values per PRNS channel are derived in this paper. Vassilis Paliouras, Alexander Skavantzos, Thanos Stouraitis |
ACM Great Lakes Symposium on VLSI | 3 |
| 2001 | Low-Power Properties of the Logarithmic Number SystemabstractThe potential of reducing power dissipation in a digital system using the logarithmic number system (LNS) is investigated. To provide a quantitative measure of power savings, the equivalence of an LNS to a linear fixed-point system is initially explored. The bit assertion activity of an LNS encoded signal is studied for both uniform and correlated Gaussian inputs. It is shown that LNS reduces the average bit assertion probability by more than 50%, in certain cases, over an equivalent linear representation. Finally, the impact of LNS on the hardware architecture and, thus, to power dissipation, is discussed. It is found that the average number of logic transitions is reduced by several times, for certain arithmetic operations and word lengths, thus compensating the power-dissipation overhead due to the unavoidable linear-to-logarithmic and logarithmic-to-linear conversion. Vassilis Paliouras, Thanos Stouraitis |
IEEE Symposium on Computer Arithmetic | 2 |
| 2001 | A wavelet-tree image coding system with efficient memory utilizationabstractThis paper describes an efficient implementation of an image coding system based on the independent wavelet-tree coding concept. The system consists of a transform and a (de)coding engine that operate in a pipelined fashion. The main focus of this paper is on the encoding part since, due to the system architecture, the decoder has identical memory utilization. Experimental results prove that the proposed system achieves comparable coding performance to the state-of-the-art, while it localizes the memory accesses to small memory modules and uses minimal computational resources. Yiannis Andreopoulos, Peter Schelkens, Nikolaos D. Zervas, Thanos Stouraitis, Constantinos E. Goutis, Jan Cornelis 0001 |
ICASSP | 4 |
| 2001 | A local wavelet transform implementation versus an optimal row-column algorithm for the 2D multilevel decompositionabstractA new method for the implementation of the binary-tree decomposition of the convolution-based wavelet transform, called the local wavelet transform (LWT) has been recently proposed in the literature. While it produces exactly the same results as the classical row-column implementation of the transform, it has many implementation benefits. This fact is shown experimentally for the first time for a general-purpose processor-based architecture, by comparing our C implementation of the LWT with an optimal C implementation of the lifting-scheme row-column algorithm. The comparisons are made for the forward multilevel binary-tree decomposition using the 9/7 filter pair, in the typical Intel Pentium processor family. Yiannis Andreopoulos, Nikolaos D. Zervas, Gauthier Lafruit, Peter Schelkens, Thanos Stouraitis, Constantinos E. Goutis, Jan Cornelis 0001 |
ICIP (3) | 5 |
| 2001 | Operation-Saving VLSI Architectures for 3D Geometrical TransformationsabstractTwo VLSI architectures for the computationally efficient implementation of the elementary 3D geometrical transformations are introduced. The first one is based on a single floating-point multiply/add unit, while the other one comprises a four processing-element vector unit. By exploiting the structure of the elementary transformation matrices, some of the elements of which are ones and zeros, the proposed architectures avoid full-matrix multiplication for the matrix multiplications involved in the calculation of the transformation matrix by treating them as updates of specific elements, the new values of which are obtained by scalar operations in the case of the single-processor architecture or by simple vector operations in the case of the processor array. Thus, the floating-point operation count and the number of memory accesses required by a transformation are reduced and, therefore, the performance of the circuit which computes the transformation matrix, in terms of execution time, is improved at minimal hardware cost. Furthermore, a circuit is proposed which, for each sequence of transformations, selects the most appropriate direction for computing the product of the matrices in the corresponding stack of transformation matrices in order to further reduce the number of floating-point operations compared to the case where the direction of the computation of the successive matrix products is predetermined. The proposed single-processor architecture is suitable for low-cost applications, while the parallel execution scheme implemented by the introduced parallel processor may be implemented by any four-PE processor with small overhead. Konstantina Karagianni, Vassilis Paliouras, George Diamantakos, Thanos Stouraitis |
IEEE Trans. Computers | 4 |
| 2000 | A hybrid image compression algorithm based on fractal coding and wavelet transformabstractRecently, in order to achieve satisfactory image and video quality and fast transmission in sometimes very low bandwidth channels, the demand for high compression ratios and fast speed in the coding and decoding procedure has been increased. A novel algorithm for very high compression of images is proposed in this paper. First of all the image is decomposed through a wavelet transform. Then the low frequency part of the decomposed image is coded by using a near lossless method while the rest of the image is coded by using a fractal coding technique. In addition, two classification methods are applied sequentially, increasing the speed of the image compression algorithm while preserving the good results with respect to PSNR and compression ratio. Its performance is compared to existing standard algorithms achieving superior results in compression ratio as well as in speed. loannis Andreopoulos, Yorgos A. Karayiannis, Thanos Stouraitis |
ISCAS | 3 |
| 2000 | Low power synthesis of sum-of-products computation (poster session)abstractNovel techniques for the power efficient synthesis of sum-of-product computations are presented. Simple and efficient heuristics for scheduling and assignment are described. Different partly static cost functions are proposed to drive the synthesis tasks. The proposed cost functions target the power consumption either in the buses connecting the functional units with the storage elements or inside the functional units. The partly static nature of the proposed cost functions reduces the time of the synthesis procedure. Experimental results from different relevant digital signal processing algorithmic kernels prove that the proposed synthesis techniques lead to significant power savings. Kostas Masselos, Spyros Theoharis, Panagiotis Merakos, Thanos Stouraitis, Constantinos E. Goutis |
ISLPED | 4 |
| 2000 | Low power architectures for digital signal processing
Kostas Masselos, Panagiotis Merakos, Thanos Stouraitis, Constantinos E. Goutis |
J. Syst. Archit. | 3 |
| 1999 | Novel techniques for bus power consumption reduction in realizations of sum-of-product computationabstractNovel techniques for power-efficient implementation of sum of product computation are presented. The proposed techniques aim at reducing the switching activity required for the successive evaluation of the partial products, in the busses connecting the storage elements where data and coefficients are stored to the functional units. This is achieved through reordering the sequence of evaluation of the partial products. Heuristics based on the traveling salesman problem are proposed to perform the reordering for different categories of algorithms. Information related to both data (dynamic) and coefficients (static) is used to drive the reordering. Experimental results from the application of the proposed techniques on several signal-processing algorithms have proven that significant switching activity savings can be achieved. Kostas Masselos, Panagiotis Merakos, Thanos Stouraitis, Constantinos E. Goutis |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1998 | Multiple-Valued Logic Voltage-Mode Storage Circuits Based On True-Single-Phase Clocked LogicabstractA number of novel voltage-mode multiple-valued logic circuits are introduced. Adopting the main features of the true single-phase clocked logic, efficient quaternary logic dynamic and pseudo-static latches, dynamic and static master-slave storage units, and uni-signal controlled pass gates are proposed. These circuits use two kinds of MOS transistors, i.e., enhancement and depletion mode, each of which has two threshold voltages. The proposed circuits exhibit regular, modular, and iterative structure, which means that the MVL circuits are VLSI implementable and can be easily re-designed for any radix of an arithmetic system. Since we use only clock signal, the derived circuits have low power dissipation. Comparisons with existing circuits prove substantial improvements in terms of speed, power consumption, and transistor count. I. Thoidis, Dimitrios Soudris, Ioannis Karafyllidis, Adonios Thanailakis, Thanos Stouraitis |
Great Lakes Symposium on VLSI | 5 |
| 1998 | Novel codebook generation algorithms for vector quantization image compressionabstractNovel algorithms for vector quantization codebook design are presented. Two basic techniques are proposed. The first technique takes into consideration specific characteristics of the blocks of the training sequence during the generation of the initial codebook. In this way a representative initial codebook is generated. Starting from a high quality initial codebook the iterative optimization procedure converges fast to representative final codebook which in turn leads to a high output image quality. The second proposed technique extends small codebooks computationally. The main idea is the application of simple transformations on the codewords. This technique reduces the memory requirements of the traditional vector quantization making it useful for applications requiring low-power consumption. Kostas Masselos, Thanos Stouraitis, Constantinos E. Goutis |
ICASSP | 2 |
| 1998 | A novel algorithm for low-power image and video codingabstractA novel scheme for low-power image and video coding and decoding is presented. It is based on vector quantization, and reduces its memory requirements, which form a major disadvantage in terms of power consumption. The main innovation is the use of small codebooks, and the application of simple but efficient transformations to the codewords during coding to compensate for the quality degradation introduced by the small codebook size. In this way, the small codebooks are computationally extended, and the coding task becomes computation based rather than memory based, leading to significant power consumption reduction. The parameters of the transformations depend on the image block under coding, and thus the small codebooks are dynamically adapted each time to this specific image block, leading to image qualities comparable to or better than those corresponding to classical vector quantization. The algorithm leads to power savings of a factor of 10 in coding and of a factor of 3 in decoding at least, in comparison to classical full-search vector quantization. Both image quality and power consumption highly depend on the size of the codebook that is used. Kostas Masselos, Panagiotis Merakos, Thanos Stouraitis, Constantinos E. Goutis |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 1997 | An operation-saving VLSI geometry engine coreabstractA floating point geometry engine core is introduced in this paper. The proposed core is optimized for performing the 3-D geometrical transformations, including the hardware evaluation of sin x and cos x functions. The architecture exploits the structure of the transformation matrices, thus reducing the number of floating point operations required per transformation. VLSI chip implementation issues for the specific architecture are also discussed. Konstantina Karagianni, George Diamantakos, Vassilis Paliouras, Thanos Stouraitis |
ICASSP | 4 |
| 1995 | Alternative Architectures for the 2-D DCT AlgorithmabstractRecently, a new fast algorithm has been proposed for the computation of the 2-D N/spl times/N-point Discrete Cosine Transform, where N is decomposed into two mutually prime numbers N/sub 1/ and N/sub 2/. Using Prime-Factor Decomposition (PFD) and appropriate index mappings, the algorithm results in fewer multiplications than other fast 2-D DCT algorithms. In this paper, a methodology for systematic mapping of this algorithm onto hardware is presented. The proposed methodology reveals the existence of a few connected components in the signal-flow graph (SFG) of the algorithm and leads to architectures that can be systematically derived and exhibit varying throughput and hardware complexity. Chrissavgi Dre, Anna Tatsaki, Thanos Stouraitis, Constantinos E. Goutis |
ISCAS | 3 |
| 1995 | A Novel Algorithm for Multi-Operand Logarithmic Number System Addition and Subtraction Using Polynominal ApproximationabstractIn this paper, a novel algorithm for multi-operand Logarithmic Number System (LNS) addition and subtraction is presented. In particular, the computation of the nonlinear functions required for logarithmic addition and subtraction is decomposed into computing 2/sup -x/, log/sub 2/(1+x), some additions, and some shifts. The error behaviour of the algorithm is analyzed, upper bounds of the computational error are provided and it is shown that the introduced Propagation Error Cancellation (PEG) technique and the Error Spectrum Shaping can significantly narrow the error distribution. The inherent parallelism of the proposed algorithm and the pipelinability that exists in the computation of 2/sup -x/ and log/sub 2/(1+x) by using polynomials are exploited by simple VLSI architectures that exhibit important speed-up over the equivalent ROM-based designs. Also, a rule for the determination of the optimal number of pipeline stages is suggested. I. Orginos, Vassilis Paliouras, Thanos Stouraitis |
ISCAS | 3 |
| 1995 | On the computation of the prime factor DST
Anna Tatsaki, Chrissavgi Dre, Thanos Stouraitis, Constantinos E. Goutis |
Signal Process. | 3 |
| 1994 | Systematic Design of Multi-Modulus/Multi-Function Residue Number System ProcessorsabstractA methodology for the design of novel Residue Number System (RNS) processors is presented. It results in ROM-less processors, which perform basic residue arithmetic algorithms in more than one moduli channel, either serially or concurrently. Moreover, the proposed architectures achieve area savings, while operating at a high throughput rate. Criteria for selecting the most appropriate moduli of operation are presented. The derived architectures are compared to memory- and full adder-based designs.> Vassilis Paliouras, Thanos Stouraitis |
ISCAS | 2 |
| 1993 | Systematic design of full adder-based architectures for convolution
Dimitrios Soudris, Vassilis Paliouras, Thanos Stouraitis, Alexander Skavantzos, Constantinos E. Goutis |
ICASSP (1) | 3 |
| 1993 | Full Adder-based Inner Product Step Processors for Residue and Quadratic Residue Number Systems
Seon Wook Kim, Thanos Stouraitis, Alexander Skavantzos |
ISCAS | 2 |
| 1993 | Methodology for the Design of Signed-digit DSP Processors
Vassilis Paliouras, Dimitrios Soudris, Thanos Stouraitis |
ISCAS | 3 |
| 1993 | Borrow: A Fault-Tolerance Scheme for Wavefront Array ProcessorsabstractA hardware fault-tolerance mechanism for wavefront array processors is proposed and evaluated. A faulty processing element borrows a working one to reperform the computations and achieve reliable results. The performance level is a stronger function of the fault pattern rather than the actual number of faults. The scheme adopts simple localized or step-wise localized control for its implementation.> Thanos Stouraitis |
IEEE Trans. Computers | 1 |
| 1992 | Systematic development of architectures for multidimensional DSP using the residue number systemabstractA systematic methodology for mapping multidimensional algorithms onto array processor architectures based on the quadratic residue number system is presented. A class of algorithms with separable functions, which can be reduced to the computation of the circular convolution is considered. The array architecture results systematically from a directed graph using partitioning techniques and consists of identical processing elements called inner product step processors. Moreover, due to various graph partitions, many alternative array architectures in terms of I/O constraints, throughput, and hardware complexity can be derived.> Dimitrios Soudris, Vassilis Paliouras, Thanos Stouraitis |
ICASSP | 3 |
| 1992 | Decomposition of Complex Multipliers Using Polynomial EncodingabstractA method for complex multiplication that relies on encoding 2n-bit complex numbers as polynomials of degree 7 in the ring of polynomials modulo x/sup 8/-1 with n/4-bit coefficients is introduced. Complex multiplication can then be performed with an 8-point cyclic convolution plus some conversion overhead and, with care, this can be done without introducing any errors. The technique is suitable for designs using systolic arrays.> Alexander Skavantzos, Thanos Stouraitis |
IEEE Trans. Computers | 2 |
| 1989 | A complex DSP processor using polynomial encodingabstractThe design of a high-speed complex signal processor is presented. It is based on a novel multiplier whose hardware implementation is shown to be characterized by simplicity, a high degree of parallelism, regularity, and modularity. The multiplier design is made possible by recent advances in the theory of performing polynomial multiplication in modular rings with reduced complexity. This latter development is based on the polynomial residue number system (PNRS). While traditional parallel complex multiplication requires four real multiplications, the proposed scheme is based on decomposing the process into eight smaller concurrent processes. If p is the performance of each of the four processors of the traditional technique and h is the investment in hardware required for its realization, the performance of each of the processors of the proposed method is 4p and its hardware investment is h/16.> Alexander Skavantzos, Zarir B. Sarkari, Thanos Stouraitis |
ICASSP | 3 |
| 1989 | A hybrid floating-point/logarithmic number system digital signal processorabstractThe implementation of a novel hybrid floating-point processor is discussed. Additions are performed without any need for exponent alignment, and multiplications and divisions are performed in less time than that required for fixed-point additions. The processor is based on a combination of the logarithmic number system (LNS) representation with the signal digit (SD) representation. The SD number system offers parallelism at the digit level for the implementation of the various operations. A novel technique for parallel conversion of SD to sign-magnitude numbers is developed to enhance the overall design. The proposed processor compares favorably to previously developed hybrid floating-point processor designs. It is at least ten gate delays faster per addition/subtraction and eight gate delays faster per multiplication/division.> Thanos Stouraitis |
ICASSP | 1 |
| 1988 | Floating-point to logarithmic encoder error analysisabstractThe logarithmic number (LNS), which supports high-speed, high-precision arithmetic, is envisioned as a possible arithmetic coprocessor attachment to a floating-point (FLP) processor. An error analysis of an FLP-to-LNS encoder is presented. Analytic expressions for the probability density function of the encoding error are derived for a number of cases, according to the memory word lengths used for the encoding. Simulation has verified the theoretical results.> Thanos Stouraitis, Fred J. Taylor |
IEEE Trans. Computers | 1 |
| 1985 | A reconfigurable systolic primitive processor for signal processingabstractRecent developments in the area of logarithmic number systems have demonstrated that when properly configured, floating point precision can be achieved at fixed point speeds. To overcome the 12-bit historical address space limit of high speed lookup memory devices, a new Adaptive Radix Processor is proposed and an error analysis is performed. Interconnecting a number of identical ARPs, fast and compact DSP systems can be designed operating on a low error budget over a large dynamic range. Examples are presented using a "shared memory design. Finally, due to identical processors used in the designs reconfiguration can be used to fully utilize the available hardware and/or introduce a degree of fault tolerance. Thanos Stouraitis, Fred J. Taylor |
ICASSP | 1 |
| 1985 | A Radix-4FFT Using Complex RNS ArithmeticabstractRecent advancements in residue arithmetic have given rise to a complex number system variant which better than halves RNS multiplication complexity. This advantage is applied to the problem of implementing a high-speed radix-4 RNS FFT. It is shown that a significant improvement in both complexity and speed can be achieved. Fred J. Taylor, George Papadourakis, Alexander Skavantzos, Thanos Stouraitis |
IEEE Trans. Computers | 4 |