VLDB 2026 Research / reviewers in the wild / expert
Vassilis Paliouras
dblp:76/5675
· DBLP profile ↗
48ranked-venue papers
4as first author
19since 2021 · last 2025
0000-0002-1414-7500ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 3 first-author · 12 since 2021Theory of computation · 7 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Computer networks · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Mixed-Precision RNS DNN AcceleratorabstractThe Residue Number System (RNS) has been used for the design of Deep Neural Network (DNN) processing architectures due to its efficient implementation of the multiply-accumulate (MAC) operation. Prior-art RNS DNN accelerators have demonstrated notable benefits compared to conventional fixed-point (FXP) representations for arithmetic precisions of at least 8 bits. However, advanced quantization techniques have recently enabled accurate ultra-low-precision FXP DNN inference. Thus, it remains an open research question whether RNS can still outperform FXP representations for smaller precisions and especially in mixed-precision (MXP) quantization settings, where optimal bit-width configurations with respect to overall accuracy drop constraints are sought. This work addresses this gap by presenting an RNS-based MXP DNN accelerator that supports 3–8-bit quantization and consistently achieves superior model performance vs. hardware cost tradeoffs for various DNN models, resulting in up to 1.2× energy efficiency improvements compared to the FXP counterpart. Synthesized on a 22-nm technology, the RNS MXP accelerator achieves 6.93–14.58 TOPS/W, outperforming the state-of-the-art uniform-precision RNS accelerator by 1.4× while maintaining the original model accuracy, as well as mixed-precision FXP accelerators. Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis |
ISCAS | 2 |
| 2025 | Optimizing Indoor RIS-Aided Physical Layer Security: A Codebook-Generation Methodology and Measurement-Based AnalysisabstractSixth-Generation (6G) wireless networks aim to support innovative Internet-of-Things (IoT) applications that demand faster and more secure data transmission. While higher Open Systems Interconnection (OSI) layers employ measures like encryption and secure protocols to address data security, Physical-Layer Security (PLS) focuses on preventing information leakage to EavesDroppers (EDs) and mitigating the effects of jammers and spoofing attacks. In this context, the emerging technology of Reconfigurable Intelligent Surfaces (RISs) can play an instrumental role, enhancing PLS by intelligently reflecting electromagnetic waves to benefit Legitimate Users (LUs) while obstructing EDs. This paper presents practical indoor measurements to evaluate the capability of an RIS to enhance PLS, focusing on a varactor-based RIS technology designed for the FR1 band at 3.55 GHz. A comparative analysis of state-of-the-art RIS-aided secrecy optimization algorithms together with a novel approach designed in this paper, which relies on a newly generated RIS phase configuration codebook, highlight the potential of RISs to improve both data rates for LUs as well as secrecy against EDs in real-world indoor multipath environments. The results also demonstrate the frequency selectivity of the RIS, providing practical insights on the optimization of the technology. Dimitris Kompostiotis, Dimitris Vordonis, Vassilis Paliouras, George C. Alexandropoulos |
PIMRC | 3 |
| 2025 | Theoretical and Experimental Evaluation of AoA Estimation in Single-Anchor 5G Uplink PositioningabstractAs we move towards 6G, the demand for high-precision, cost-effective positioning solutions becomes increasingly critical. Single-anchor positioning offers a promising alternative to traditional multi-anchor approaches, particularly in complex propagation environments where infrastructure costs and deployment constraints present significant challenges. This paper provides a comprehensive evaluation of key algorithmic choices in the development of a single-anchor 5G uplink positioning testbed. Our developed testbed uses angle of arrival (AoA) estimation combined with range measurements from an ultra-wideband pair, to derive the position. The simulations conducted assess the impact of the selected algorithms on channel order and AoA estimation, while the influence of antenna calibration errors on AoA estimation is also examined. Finally, we compare simulations and results obtained from our developed platform. Thodoris Spanos, Fran Fabra, José A. López-Salcedo, Gonzalo Seco-Granados, Nikos Kanistras, Ivan Lapin, Vassilis Paliouras |
PIMRC | 7 |
| 2025 | Evaluating Beam Sweeping for AoA Estimation with an RIS Prototype: Indoor/Outdoor Field TrialsabstractReconfigurable Intelligent Surfaces (RISs) have emerged as a promising technology to enhance wireless communication systems by enabling dynamic control over the propagation environment. However, practical experiments are crucial towards the validation of the theoretical potential of RISs while establishing their real-world applicability, especially since most studies rely on simplified models and lack comprehensive field trials. In this paper, we present an efficient method for configuring a 1-bit RIS prototype at sub-6 GHz, resulting in a codebook oriented for beam sweeping; an essential protocol for initial access and Angle of Arrival (AoA) estimation. The measured radiation patterns of the RIS validate the theoretical model, demonstrating consistency between the experimental results and the predicted beamforming behavior. Furthermore, we experimentally prove that RIS can alter channel properties and by harnessing the diversity it provides, we evaluate beam sweeping as an AoA estimation technique. Finally, we investigate the frequency selectivity of the RIS and propose an approach to address indoor challenges by leveraging the geometry of environment. Dimitris Vordonis, Dimitris Kompostiotis, Vassilis Paliouras, George C. Alexandropoulos, Florin Grec |
WCNC | 3 |
| 2024 | SECURED for Health: Scaling Up Privacy to Enable the Integration of the European Health Data SpaceabstractIn this paper, we present the SECURED project11Funded in part by the European Union (EU), Grant Agreement no. 10109571. Views and opinions expressed are those of the authors and do not necessarily reflect those of the EU or the Health and Digital Executive Agency. Neither the EU nor the granting authority are responsible for them., aimed at improving privacy-preserving processing of data in the health domain. The technologies developed in the project will be demonstrated in four health-related use cases and with the involvement of SME's selected through an open funding call. Francesco Regazzoni 0001, Gergely Ács, Albert Zoltan Aszalos, Christos Avgerinos, Nikolaos Bakalos, Josep Lluís Berral, Joppe W. Bos, Marco Brohet, Andrés G. Castillo, Gareth T. Davies, Stefanos Florescu, Pierre-Elisée Flory, Alberto Gutierrez-Torre, Evangelos Haleplidis, Alice Héliou, Sotiris Ioannidis, Alexander El-Kady, Katarzyna Kapusta, Konstantina Karagianni, Pieter Kruizinga, Kyrian Maat, Zoltán Ádám Mann, Kalliopi Mastoraki, SeoJeong Moon, Maja Nisevic, Balazs Pejo, Kostas Papagiannopoulos, Vassilis Paliouras, Paolo Palmieri 0001, Francesca Palumbo, Juan Carlos Pérez Baun, Péter Pollner, Eduard Porta-Pardo, Luca Pulina, Muhammad Ali Siddiqi, Daniela Spajic, Christos Strydis, George Tasopoulos, Vincent Thouvenot, Christos Tselios, Apostolos P. Fournaris |
DATE | 28 |
| 2023 | Improving Residue-Level Sparsity in RNS-based Neural Network Hardware Accelerators via RegularizationabstractResidue Number System (RNS) has recently attracted interest for the hardware implementation of inference in machine-learning systems as it provides promising trade-offs in the area, time, and power dissipation space. In this paper we introduce a technique that utilizes regularization during training, and increases the percentage of residues which are zero, when the parameters of an artificial neural network (ANN) are expressed in an RNS. The proposed technique can also be used as a post-processing stage, allowing the optimization of pre-trained models for RNS implementation. By increasing the number of residues being zero, i.e., residue-level sparsity, the proposed technique facilitates new hardware architectures for RNS-based inference, allowing new trade-offs and improving performance over prior art without practically compromising accuracy. The introduced method increases residue sparsity by a factor of 4× to 6× in certain cases. Emmanouil Kavvousanos, Vasilis Sakellariou, Ioannis Kouretas, Vassilis Paliouras, Thanos Stouraitis |
ARITH | 4 |
| 2023 | A multiplier-Free RNS-Based CNN accelerator exploiting bit-Level sparsityabstractIn this work, a Residue Numbering System (RNS)-based Convolutional Neural Network (CNN) accelerator utilizing a multiplier-free distributed-arithmetic Processing Element (PE) is proposed. A method for maximizing the utilization of the arithmetic hardware resources is presented. It leads to an increase of the system's throughput, by exploiting bit-level sparsity within the weight vectors. The proposed PE design takes advantage of the properties of RNS and Canonical Signed Digit (CSD) encoding to achieve higher energy efficiency and effective processing rate, without requiring any compression mechanism or introducing any approximation. An extensive design space exploration for various parameters (RNS base, PE micro-architecture, encoding) using analytical models as well as experimental results from CNN benchmarks is conducted and the various trade-offs are analyzed. A complete end-to-end RNS accelerator is developed based on the proposed PE. The introduced accelerator is compared to traditional binary and RNS counterparts as well as to other state-of-the-art systems. Implementation results in a 22-nm process show that the proposed PE can lead to 1.85× and 1.54× more energy-efficient processing compared to binary and conventional RNS, respectively, with a 1.88× maximum increase of effective throughput for the employed benchmarks. Compared to a state-of-the-art, all-digital, RNS-based system, the proposed accelerator is 8.87× and 1.11× more energy- and area-efficient, respectively. Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis |
ARITH | 2 |
| 2023 | Received Power Maximization with Practical Phase-Dependent Amplitude Response in RIS-Aided OFDM Wireless CommunicationsabstractA Reconfigurable Intelligent Surface (RIS) acts as an intelligent scatterer, reflecting the incoming signal to the desired destination by adjusting a multitude of phase-shifts in real-time. Most research works consider full signal reflection; i.e., the reflection coefficient at each element is assumed to have unit amplitude, despite the phase-dependent amplitude response observed in practice. In this paper, based on a practical phase-shift model, we formulate an optimization problem to maximize the power received by the user subject to a frequency-selective fading channel under Orthogonal Frequency Division Multiplexing (OFDM) transmission. An iterative algorithm is proposed, applying alternating optimization (AO) techniques. Both continuous and discrete phase shifters are examined, while the mutual coupling effect due to inter-element spacing is also considered. Simulation results reveal that the proposed reflection optimization algorithm outperforms, in terms of achievable data rate, state-of-the-art passive beamforming methods, which are based on the conventional ideal phase shift model. Dimitris Kompostiotis, Dimitris Vordonis, Vassilis Paliouras |
ICASSP | 3 |
| 2023 | Invited Paper: Dilithium Hardware-Accelerated Application Using OpenCL-Based High-Level SynthesisabstractPost-quantum cryptography (PQC) has been gaining attention in the last few years due to the evolution of quantum computers and the need to replace traditional, quantum-attack-insecure cryptography schemes with quantum-attack-resistant schemes. Lattice-based cryptography (LBC) constitutes a highly promising post-quantum solution (Quantum Resistant), but implementations in software or hardware are challenging due to the use of operations based on large-size polynomials. LBC selected schemes for standardization by the National Institute of Standards and Technology (NIST), rely on matrix-to-matrix multiplications of high-order polynomials, having performance bottlenecks that are solved using the Number-Theoretic Transform (NTT). One of the NIST-selected schemes for digital signature (DS) is the CRYSTALS-Dilithium scheme, which uses$N=256$degee polynomials. In this paper a Hardware/Software (HW/SW) co-design solution is proposed for all security levels of Dilithium, utilizing the OpenCL framework and Vitis High-Level Synthesis (HLS) tool. In our work, the HW/SW co-design OpenCL mechanisms are analyzed extensively and communication overheads between the hardware kernel and an ARM processor are identified, while appropriate techniques are proposed in order to bypass the I/O time-bottleneck on a real-world application. The proposed implementation runs on the ARM Processing System (PS) of a Xilinx Multi-Processor System on Chip (MPSoC) system, utilizing the MPSoC FPGA Programmable Logic (PL) in order to accelerate the calculations relative to the heavy matrix-multiplication operation. Finally, the proposed HW/SW codesigned solution is realized as a real-world Linux-based Dilithium DS executable and manages to achieve realistic performance gain, in terms of time execution, versus a CPU-only execution ranging from 2-23% (depending on the utilized CPU Clock Frequency). Alexander El-Kady, Apostolos P. Fournaris, Vassilis Paliouras |
ICCAD | 3 |
| 2022 | Sensitivity to Threshold Voltage Variations of Exact and Incomplete Prefix Addition TreesabstractProcess variations have emerged as a severe performance bottleneck for advanced technology nodes. This paper investigates the delay behavior of a range of exact and incomplete parallel-prefix addition trees under variations, including speculative/approximate trees and trees with duplicated prefix nodes. The performance of a range of incomplete adder variants is investigated comparatively with conventional exact architectures at a 16-nm technology node. In order to capture variations generated from a range of process-dependent sources and their impact on delay characteristics, the analysis considers threshold voltage variations employing Spice-level simulations. Under nominal voltage and in the presence of threshold variations, incomplete architectures are found to still offer a smaller worst-case delay than their exact counterparts; similarly for a low-voltage scenario, where the supply voltage is reduced to 0.8 V. However, focusing on normalized delay variation, it is here found that incomplete-tree adders are more susceptible to variations as they show a wider delay spread than exact architectures around their respective mean values. As a remedy, it is shown that the number of stages in an incomplete adder can be used as a design parameter to investigate accuracy vs. variability-tolerance trade-offs. Furthermore, by duplicating delay-critical paths of prefix trees, worst-case delay and standard deviation reduce compared to the exact architectures. Kleanthis Papachatzopoulos, Vassilis Paliouras |
ISCAS | 2 |
| 2022 | A High-performance RNS LSTM blockabstractThe Residue Number System (RNS) has been proposed as an alternative to conventional binary representations for use in AI hardware accelerators. While it has been successfully utilized in applications targeting Convolutional Neural Networks (CNNs), its usage in other network models such as Recurrent Neural Networks (RNNs) has been set back due to the difficulty of implementing more complex activations functions like tanh and sigmoid ($\sigma$) in the RNS domain. In this paper, we seek to extend its usage in such models, and in particular LSTM networks, by providing efficient RNS implementations of the activation functions. To this aim, we derive improved accuracy piecewise linear approximations of the tanh and $\sigma$ functions using the minimax approach and propose a fully RNS-based hardware realization. We show that our approximations can effectively mitigate accuracy degradation in LSTM networks compared to naive approximations, while the RNS LSTM block can be up to 40% more efficient in terms of performance per area unit compared to a binary counterpart, when used in high performance-targeted accelerators. Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis |
ISCAS | 2 |
| 2022 | High-Level Synthesis design approach for Number-Theoretic MultiplierabstractLattice-based cryptography (LBC) performs polynomial multiplication using the Number Theoretic Transform (NTT), in order to reduce the polynomial multiplication complexity from O(n2) to O(n log n). Although NTT-based multipliers offer the fastest way to compute a polynomial multiplication product for high-degree polynomials (with non-trivial bit-length coefficients), they constitute a significant part of the overall LBC scheme delay thus becoming the main LBC efficiency bottleneck. Therefore, the need to optimize the NTT-based multiplication in an easy, automatic yet efficient manner is significant. High-Level synthesis (HLS) tools offer such a capability since they can hide the Register Transfer Level (RTL)-based design complexity (typically realized by hardware description languages) using high level descriptions in C, C++ or openCL. However, this design approach requires careful modifications for high-level description code like loop reordering, loop flattening, removing dependencies, loop pipelining and loop unrolling in order to produce through an HLS tool a design with performance comparable to RTL hand-crafted designs. In this paper, extending the work in [1] we propose a complete NTT-based polynomial multiplier that combines an HLS optimized Cooley-Tukey (CT) NTT design with a proposed, HLS optimized, Gentleman-Sande (GS) Inverse-NTT design to create a highly efficient multiplier design that can benefit from the HLS flexibility yet still achieve significant high speed. More specifically, in the paper, the read and write access of the NTT processing elements (PE) to the memory is significantly increased though appropriate code redesign and the use of the dependence HLS pragma is proposed in order to reduce the dependencies between PEs. The proposed work has been evaluated by introducing the proposed NTT multiplier in the LBC Dilithium digital-signature scheme (polynomial degree n = 256, coefficient modulus Q = 8380417) and managed to achieve significantly higher speed compared to other similar works. Alexander El-Kady, Apostolos P. Fournaris, Evangelos Haleplidis, Vassilis Paliouras |
VLSI-SoC | 4 |
| 2021 | Sum Propagate AddersabstractPublished in "IEEE Transactions on Emerging Topics in Computing, Volume: 9, Issue: 3, JulySeptember 2021" and orally presented at ARITH 2021. Giorgos Dimitrakopoulos, Kleanthis Papachatzopoulos, Vassilis Paliouras |
ARITH | 3 |
| 2021 | Optimizing Deep Learning Decoders for FPGA ImplementationabstractRecently, Deep Learning (DL) methods have been proposed for use in the decoding of linear block codes. While novel DL decoders show promising error correcting performance, they suffer from computational complexity issues, which prevent their usage with large block codes and make their implementation in digital hardware inefficient. The subject of the presented doctoral research is the design of DL decoding methods with low computational complexity and resource requirements, by applying compression and approximation techniques to the employed Neural Networks. Efficient hardware architectures are expected to be designed for these optimized DL decoders on FPGA devices, which will overcome the current performance limitations. Emmanouil Kavvousanos, Vassilis Paliouras |
FPL | 2 |
| 2021 | A Novel Stochastic Polar Architecture for All-Digital TransmissionabstractA novel architecture of an all-digital transmitter is proposed in this paper, introducing stochastic computation to a single-bit polar topology. The proposed architecture exploits the benefits of stochastic computing to simplify the design of an all-digital transmitter and minimize hardware complexity. Complexity reduction is achieved by implementing a digital oscillator as a cosine function in the stochastic domain instead of resorting to LUT-based or CORDIC-based implementations. An evaluation of the introduced all-digital polar architecture shows that the power spectral density of the transmitted passband signal is similar to that of a conventional transmitter. Synthesis of the proposed stochastic-enabled transmitter at a 28-nm FDSOI technology node reveals its minimal complexity requirements and 87.13% area reduction compared to conventional implementations. Christos Andriakopoulos, Kleanthis Papachatzopoulos, Vassilis Paliouras |
ISCAS | 3 |
| 2021 | Simplified Hardware Implementation of Memoryless Dot Product for Neural Network InferenceabstractIn this paper a simplified hardware implementation of a dot product arithmetic operation with constant coefficients is presented. The proposed methodology exploits a combination of distributed arithmetic and common subexpression sharing techniques. An algorithm is introduced for identifying the common sub partial sums systematically. Subsequently, a hardware architecture is proposed and the obtained circuits are synthesized in a 90-nm 1.0 V CMOS standard-cell library using Synopsys Design Compiler. Comparisons reveal significant reduction of 52% and 23% in area and power respectively for 1.5 ns delay over a regular dot product constant multiplier. Ioannis Kouretas, Vassilis Paliouras |
ISCAS | 2 |
| 2021 | An FPGA Accelerator for Spiking Neural Network Simulation and TrainingabstractSpiking Neural Networks (SNNs) have recently been employed to solve a number of machine learning problems traditionally addressed by classical Artificial Neural Networks (ANNs). SNNs are different than ANNs as they incorporate time in their computations and information is encoded in the exact timing or frequency of discrete events, spikes. SNNs promise to deliver a higher energy efficiency than ANNs, when implemented in neuromorphic hardware, due to their event-driven nature. Naturally, this different computational model they introduce, creates different design challenges for the implementation of large-scale networks and needs to be addressed by different architectures. The spiking accelerator presented here aims to facilitate and accelerate the process of developing SNNs for ML applications that are traditionally addressed by ANNs and help bridge the accuracy gap between them. It achieves a significant speedup of up to 800× for inference and up to 500× for training compared to software SNN simulations for certain set-ups. Vasilis Sakellariou, Vassilis Paliouras |
ISCAS | 2 |
| 2021 | High-Level Synthesis design approach for Number-Theoretic Transform ImplementationsabstractLattice-based cryptography performs polynomial multiplication using the Number Theoretic Transform (NTT), in order to reduce the polynomial multiplication complexity from $O\left(n^{2}\right)$ to $O(n \log n)$. NTT has been in the center of investigation in cryptography space, as it is applied in many cryptography schemes such as hash functions, homomorphic encryption, key-encapsulation mechanisms, and digital signatures. A common approach for rapid production of hardware designs commences from semi-automatic software production, as supported by the Xilinx High-Level Synthesis (HLS) toolchain or similar tools. Most of the times this approach requires careful modifications (e.g. code modification, loop reordering, loop flattening, removing dependencies, loop pipelining, loop unrolling) in order to achieve a design with performance comparable to a Register-Transfer Level (RTL) hand-crafted design. In this paper a design solution is proposed that solves the data and loop-carry dependencies of the Cooley-Tukey NTT algorithm, by assisting the HLS synthesizer to produce efficient designs, in terms of latency and resources. The proposed work has been evaluated using the Dilithium digital-signature scheme NTT version ($n=256, Q$ of 23 bits), and is shown to achieve a 20-50 % improvement in terms of latency (without really affecting the resources) compared to other existing HLS-based NTT solutions in the literature. Alexander El-Kady, Apostolos P. Fournaris, Thanasis Tsakoulis, Evangelos Haleplidis, Vassilis Paliouras |
VLSI-SoC | 5 |
| 2021 | A Multirate Fully Parallel LDPC Encoder for the IEEE 802.11n/ac/ax QC-LDPC Codes Based on Reduced Complexity XOR TreesabstractThis article proposes an encoding method based on a two-step encoding algorithm for the 12 quasi-cyclic (QC)-low-density parity-check (LDPC) (QC-LDPC) codes specified in the IEEE 802.11n/ac/ax standards. The proposed approach jointly considers all codes of the particular set, instead of targeting each code separately. The proposed algorithm performs multiplication by inverse matrices. The complexity of the multiplications is significantly reduced by the introduced encoding method. It allows the implementation of full-parallel architectures that execute the encoding process within a single clock cycle, or more for pipelined implementations, for any of the supported codes. A corresponding VLSI encoding architecture based on XOR-gate trees is also proposed. The proposed solution exploits the structure and features of the involved matrices to extract common subexpressions (CSs) using common sub-expression sharing techniques (CSST). Such expressions result due to common features of the original matrices and the corresponding inverses, identified in this article. Innovative subexpression extraction procedures that target the specific codes as a set are introduced here. Furthermore, illustrative single-clock hardware encoders derived by the proposed technique are integrated into 90- and 45-nm technologies at 1 GHz occupying 125 and 107 KGates, respectively, achieving throughput rates up to 1.62 Tbps. Ahmed Mahdi, Nikos Kanistras, Vassilis Paliouras |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | Maximum Delay Models for Parallel-Prefix Adders in the Presence of Threshold Voltage VariationsabstractThis paper introduces a delay modeling formulation for several Parallel-Prefix Adders in the presence of threshold voltage variability. A path-based model is derived for the delay variability of Kogge-Stone, Knowles, Sklansky, Brent-Kung, Han-Carlson, Ladner-Fischer, and New Adder architectures. The delay model accuracy is evaluated for the specific adders on the basis of SPICE Monte-Carlo Simulations at 45 nm and 16 nm nodes. The presented analysis reveals that the proposed path-based model estimates the maximum delay Probability Density Function of the particular adder architectures with sufficient accuracy, assuming 3σ intra-die threshold voltage variations as high as 10% of nominal value. Delay yield estimations produced by the proposed model are found to agree with those of Monte-Carlo Simulations for a number of highly probable critical paths, presenting an error less than 2%. For the particular adders and technology nodes, an approximately 10-fold reduction in simulation time is obtained when exploiting the proposed model. The particular observation indicates that the computational time for delay yield estimation of Parallel-Prefix Adders can be exponentially reduced with negligible accuracy loss when the analysis focuses solely on the Nominal-Maximum Delay critical path. Finally, a quantitative comparison of prefix adders to the Borrow-Save Adder is offered, in terms of complexity and susceptibility to variations. Kleanthis Papachatzopoulos, Vassilis Paliouras |
ARITH | 2 |
| 2020 | Novel Noise-Shaping Stochastic-Computing Converters for Digital FilteringabstractStochastic computing introduces massive parallelism in several practical applications by utilizing minimal-complexity processing elements, and provides inherent fault-tolerant features. However stochastic computation systems require long bit streams to achieve sufficient performance in terms of Signal to Noise Ratio. This paper proposes a first- and a second-order Noise-Shaping Binary-to-Stochastic Converter (NSBSC) for bipolar format. The proposed architecture schemes are compared with a baseline Binary-to-Stochastic Converter (BSC) in terms of Signal-to-Quantization-Noise Ratio (SQNR). It is shown that for certain test cases, the proposed architecture leads to 15.276 dB improved SQNR for the same bit stream length and, furthermore, achieves the same SQNR as the conventional converter using as much as 93.75% shorter stream lengths. Furthermore, the analysis includes area and power figures for the introduced hardware architectures for a 28-nm FDSOI technology. Finally, achieved NSBSC gains are shown to propagate at the output of a stochastic FIR filter, proving that the stochastic properties of the derived stream are maintained. Kleanthis Papachatzopoulos, Christos Andriakopoulos, Vassilis Paliouras |
ISCAS | 3 |
| 2020 | Implementing the Residue Logarithmic Number System Using Interpolation and CotransformationabstractThe Residue Logarithmic Number System (RLNS) offers fast multiplication and division, but poses challenges for implementing addition and subtraction because the underlying integer Residue Number System (RNS) has slow sign detection. The conventional Binary Logarithmic Number Systems (BLNS) has benefited from interpolation and cotransformation. We propose a dual-path ALU that speculates about the sign detection to adapt interpolation and cotransformation to the limitations of RLNS. Synthesis shows for the same precision and technology, the area of the proposed RLNS circuit is similar to BLNS and much smaller than prior RLNS methods. We also compare against Floating Point (FP). Mark G. Arnold, Vassilis Paliouras, Ioannis Kouretas |
IEEE Trans. Computers | 2 |
| 2019 | Under- and Overflow Detection in the Residue Logarithmic Number SystemabstractThe Residue Number System (RNS) offers fast and cheap carry-free integer arithmetic but has slow and expensive overflow detection. The Logarithmic Number System (LNS) offers fast real multiplication, division and powers with floating-point-like relative precision. The Residue Logarithmic Number System (RLNS) is a combination of the two systems that offers advantages for moderate-precision real applications where a-priori analysis allows under-and overflow to be ignored. An arithmetic hardware generator is essential because of the mathematical obscurity of combining RNS and LNS. Unfortunately, real applications often underflow. We consider options to deal with under-and overflow using the RLNSTool generator as a foundation. Mark G. Arnold, Ioannis Kouretas, Vassilis Paliouras, John R. Cowles |
ARITH | 3 |
| 2018 | A Reconfigurable LDPC Decoder Optimized for 802.11n/ac ApplicationsabstractThis paper presents a high data-rate low-density parity-check (LDPC) decoder, suitable for the 802.11n/ac (WiFi) standard. The innovative features of the proposed decoder relate to the decoding algorithms and the interconnection between the processing elements. The reduction of the hardware complexity of decoders based on the min-sum (MS) algorithms comes at the cost of performance degradation, especially at high-noise regions. We introduce more accurate approximations of the log-sum-product algorithm that also operate well for low signal-to-noise ratio values. Telecommunication standards, including WiFi, support more than one quasi-cyclic LDPC codes of different characteristics, such as codeword length and code rate. A proposed design technique derives networks, capable of supporting a variety of codes and efficiently realizing connectivity between a variable number of processing units, with a relatively small hardware overhead over the single-code case. As a demonstration of the proposed technique, we implemented a reconfigurable network based on barrel rotators, suitable for LDPC decoders compatible with WiFi standard. Our approach achieves low complexity and high clock frequency, compared with related prior works. A 90-nm application-specified integrated circuit implementation of the proposed high-parallel WiFi decoder occupies 4.88 mm2and achieves an information throughput rate of 4.5 Gbit/s at a clock frequency of 555 MHz. Ioannis Tsatsaragkos, Vassilis Paliouras |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Work in progress: An introduction to computing course using a Python-based experiential approachabstractThis paper discusses our experience with designing and implementing a new Introduction to Computing course that aims both to introduce students to programming and at the same time to cover key concepts of Computer Science using a hands-on experiential approach. The re-design of the course has put more emphasis on using Python as a tool for studying computer science. It was found that the students enjoyed this approach and responded positively to the challenge. Examples are given on how concepts of binary arithmetic, computer architecture, operating systems, networking have been introduced through Python coding examples and exercises, while the project work introduced the students to the challenges of software development. Nikolaos M. Avouris, Kyriakos N. Sgarbas, Vassilis Paliouras, Michalis Koukias |
EDUCON | 3 |
| 2017 | Logarithmic number system addition-subtraction using fractional normalizationabstractThis paper presents a technique for addition/subtraction in the Logarithmic Number System (LNS) which is based on Floating Point addition/subtraction, specifically on fractional normalization (FN). The proposed technique is analyzed and compared with two other methods for addition/subtraction using 11- and 17-bit word lengths. The results are demonstrated by evaluating complexity and performance of synthesized modules of the above addition/subtraction methods. Modules are synthesized using a 65-nm 0.9V UMC CMOS library. Giorgos Tsiaras, Vassilis Paliouras |
ISCAS | 2 |
| 2016 | Dynamic delay variation behaviour of RNS multiply-add architecturesabstractIn this paper we investigate the impact of intra- and inter-die variations on the delay sensitivity of certain Residue Number System (RNS) arithmetic circuits in comparison to ordinary binary arithmetic logic. The timing yield of systems that contain multiply-add units (MAC) is of great importance since they dominate important applications such as digital signal processing. Specifically, we employ two different delay models for the estimation of delay distributions of RNS and binary MAC architectures. Our analysis quantitatively proves that RNS MAC architectures that use bases of the form {2n- 1, 2n, 2n+ 1} demonstrate better normalized delay variation than binary MAC architectures to characterize both their static timing behaviour and the timing behaviour taking into account the sensitizable paths. Furthermore, it is shown that certain simplified RNS MAC architectures outperform conventional RNS MAC architectures in terms of the μ + α · σ delay variation metric. Kleanthis Papachatzopoulos, Ioannis Kouretas, Vassilis Paliouras |
ISCAS | 3 |
| 2016 | Application-Specific Low-Power MultipliersabstractIn this paper we propose techniques for decreasing the power consumption of multipliers. The proposed techniques exploit data statistics. Specifically we propose two architectures for two-input signed multipliers, namely selective activation multiplier and partitioned multiplier. To illustrate the benefits achieved by the proposed architectures, their application in the implementation of multipliers used in DCT and DWT is detailed. The power reduction reported is up to$20.7$percent for cases of practical interest. The proposed multipliers can be easily implemented, adapted, and combined with prior-art low-power techniques with a moderate area and time overhead. Panagiotis Sakellariou, Vassilis Paliouras |
IEEE Trans. Computers | 2 |
| 2013 | Delay-variation-tolerant FIR filter architectures based on the Residue Number SystemabstractThis paper investigates the use of the Residue Number System (RNS) in the hardware design of VLSI FIR filters implemented in nano-scale technologies prone to process variation effects. It is here shown that the RNS substantially reduces the filter sensitivity to delay variations, when compared to digital filter designs that use conventional positional number systems, such as the widely-used two's-complement representation. The inherent tolerance of the introduced RNS architectures to the delay variations, is here shown to allow to circumvent the use of large design parameter margins. Therefore, we demonstrate that the use of RNS can achieve a high timing yield without resorting to costly over-design, which may unnecessarily increase system complexity. The particular benefit comes in addition to area, time and power benefits achieved due to the use of the RNS. The quantitative digital filter design space exploration reported in the paper takes into consideration the filter order as well as criteria related to the filter output signal quality such as the signal-to-noise ratio (SNR) and it demonstrates that the proposed architectures offer effective solutions for hardware design using modern and future nano-scale processes, for filter cases of practical interest. Ioannis Kouretas, Vassilis Paliouras |
ISCAS | 2 |
| 2013 | Low-Power Logarithmic Number System Addition/Subtraction and Their Impact on Digital FiltersabstractThis paper presents techniques for low-power addition/subtraction in the logarithmic number system (LNS) and quantifies their impact on digital filter VLSI implementation. The impact of partitioning the look-up tables required for LNS addition/subtraction on complexity, performance, and power dissipation of the corresponding circuits is quantified. Two design parameters are exploited to minimize complexity, namely the LNS base and the organization of the LNS word. A roundoff noise model is used to demonstrate the impact of base and word length on the signal-to-noise ratio of the output of finite impulse response (FIR) filters. In addition, techniques for the low-power implementation of an LNS multiply accumulate (MAC) units are investigated. Furthermore, it is shown that the proposed techniques can be extended to cotransformation-based circuits that employ interpolators. The results are demonstrated by evaluating the power dissipation, complexity and performance of several FIR filter configurations comprising one, two or four MAC units. Simulations of placed and routed VLSI LNS-based digital filters using a 90-nm 1.0 V CMOS standard-cell library reveal that significant power dissipation savings are possible by using optimized LNS circuits at no performance penalty, when compared to linear fixed-point two's-complement equivalents. Ioannis Kouretas, Charalambos Basetas, Vassilis Paliouras |
IEEE Trans. Computers | 3 |
| 2012 | Residue arithmetic for designing multiply-add units in the presence of non-gaussian variationabstractIn this paper the utilization of Residue Number System (RNS) is investigated as a tool for variation-tolerant design. In particular circuits using various RNS bases are compared to the equivalent binary structures in terms of their sensitivity to the variation of process parameters. Furthermore, RNS advantages are quantitatively illustrated by considering a timing model with two non-gaussian distributions. It is shown that for bases where all moduli channels are candidates to contain the critical path of the RNS circuit, the delay variation is significantly reduced when compared to the equivalent binary structures. Ioannis Kouretas, Vassilis Paliouras |
ISCAS | 2 |
| 2011 | Towards a Quaternion Complex Logarithmic Number SystemabstractThe well-known generalization of real to complex arithmetic (two reals) extends further to more obscure quaternion arithmetic (four reals), which has applications in signal processing, aerospace, graphics and virtual reality. Quaternion multiplication implements 3D rotation, but is expensive (usually 16 floating-point multiplications and 12 additions). This paper proposes an alternative quaternion representation using logarithms to reduce multiplication cost. The real Logarithmic Number System (LNS) allows fast and inexpensive multiplication and division in embedded and FPGA-based systems. Recent advances in the Complex LNS (CLNS) have made fast log-polar complex representation affordable. Although the quaternion logarithm function is also well-defined, it is not useful to simplify multiplication (in the same way real and complex logarithms are) because quaternion multiplication is not commutative but quaternion addition is. To overcome this, we propose a novel Quaternion Complex (QCLNS) representation using a pair of CLNS numbers. This representation implements quaternion multiplication using only the theoretical minimum, of 8 LNS multipliers (i.e., fixed-point adders) and two CLNS adders. Because CLNS numbers are more compact than ordinary rectangular complex representation, single-precision QCLNS occupies 10.9 percent less memory than conventional quaternion representation. Extrapolating conventional LNS and floating-point synthesis data from Fu et al., QCLNS saves on average 10 percent of FPGA resources for precisions between 13 and 45 bits. Mark G. Arnold, John R. Cowles, Vassilis Paliouras, Ioannis Kouretas |
IEEE Symposium on Computer Arithmetic | 3 |
| 2011 | A Residue Logarithmic Number System ALU using interpolation and cotransformationabstractThe Residue Logarithmic Number System (RLNS) uses the Residue Number System (RNS) to represent logarithms that represent real values. Multiplication and division are easy; reasonable-precision addition and subtraction have not been economical because of RNS sign-detection. This paper adapts novel interpolation (for addition) and cotransformation (for subtraction) to fit the sign-detection limits. Smaller than prior RLNS hardware, the novel ALU defers sign detection with two speculative datapaths. Mark G. Arnold, Ioannis Kouretas, Vassilis Paliouras |
ASAP | 3 |
| 2011 | A flexible high-throughput hardware architecture for a gaussian noise generatorabstractIn this paper a flexible, high-throughput, low-complexity additive white gaussian noise (AWGN) channel generator is presented. The proposed generator employs a Mersemie-Twister to generate a long random number uniformly distributed sequence and a Box-Muller transformation implementation to derive gaussian noise samples. Emphasis is given on developing a high-throughput approximation unit for die elementary functions required for die transformation. The proposed techniques are shown to lead to solutions that provide four samples per clock, which in turn can sustain throughputs of 584MSps, with moderate clock frequency and hardware complexity. Ioannis Paraskevakos, Vassilis Paliouras |
ICASSP | 2 |
| 2010 | Residue arithmetic bases for reducing delay variationabstractIn this paper the utilization of Residue Number System (RNS) is investigated as a tool for variation-tolerant design. In particular circuits using various RNS bases are compared in terms of their sensitivity to the variation of process parameters. Furthermore, RNS advantages are quantitatively illustrated by considering a timing model. It is shown that for bases where all moduli channels are candidates to contain the critical path of the RNS circuit, the delay variation is reduced upto 86% when compared to the equivalent binary structures. Ioannis Kouretas, Vassilis Paliouras |
ISCAS | 2 |
| 2009 | Variation-tolerant Design Using Residue Number SystemabstractIn this paper the use of residue arithmetic is proposed as a technique to reduce delay variation in adders. It is found that the use of residue arithmetic offers significant delay variation reduction when compared to adders of the literature. Therefore this technique can be used to control variance of critical paths delay and efficiently meet timing constraints and thus improve timing yield. Experiments conducted span several values of intra-die and die-to-die variance, so that cases of practical interest for various nanoscale technologies are covered. Ioannis Kouretas, Vassilis Paliouras |
DSD | 2 |
| 2008 | Low-power logarithmic number system addition/subtraction and their impact on digital filtersabstractThis paper discusses techniques for low-power addition/subtraction in the logarithmic number system (LNS) and evaluates their impact on digital filter implementation. Initially, the impact of partitioning the look-up tables (LUT) required for addition/subtraction on complexity, performance, and power dissipation is studied. Subsequently techniques for the low-power implementation of an LNS multiply- accumulate (MAC) unit are investigated. The obtained LNS MACs are used for the design of digital filters. Synthesis of LNS-based digital filters using a 0.18 mum 1.8 V CMOS standard-cell library, reveal that significant power dissipation savings are possible at no performance penalty, when compared to linear two's-complement equivalent. Ioannis Kouretas, Charalambos Basetas, Vassilis Paliouras |
ISCAS | 3 |
| 2006 | A novel technique for low-power D/A conversion based on PAPR reductionabstractLarge dynamic range of the signal is a significant drawback for orthogonal frequency division multiplexing (OFDM) systems since it restricts the efficiency of the RF part of a wireless transmitter. This paper quantifies the impact of the partial transmit sequence (PTS) peak to average power ratio (PAPR) reduction algorithm on the complexity of a digital to analog converter (DAC). In particular, experimental results for an OFDM system with a large number of subcarriers demonstrate that the required DAC resolution is reduced by one bit, while its power consumption is reduced by a factor of two. In addition, this paper proposes two new sets of weighting factors which reduce the computational complexity of PTS by even 87.5%, while maintaining exactly the same performance Theodoros Giannopoulos, Vassilis Paliouras |
ISCAS | 2 |
| 2004 | An efficient computational method and a VLSI architecture for digital filtering of CP-OFDM signalsabstractThis paper introduces an efficient computational technique for orthogonal frequency division multiplexing (OFDM) based modem design with digital filters and discusses the corresponding VLSI architecture issues. General conditions are introduced, under which the proposed technique is applicable. The proposed technique is demonstrated by detailing the implementation of a 64-point, radix-4 pipelined FFT in combination with a parallel digital filter architecture. By exploiting the redundancy into the cyclic prefix part of the OFDM symbol, the computational load of the transmitter is reduced by 20% for cases of practical interest. The radix-r N-point FFT case is examined. Eleni Fotopoulou, Vassilis Paliouras |
GLOBECOM | 2 |
| 2002 | Multi-voltage low power convolvers using the polynomial residue number systemabstractA novel approach for the reduction of the power dissipated in a signal processing application is introduced in this paper. By exploiting the properties of the Polynomial Residue Number System (PRNS) and of the arithmetic modulo (2n+1), the power dissipation of implementing cyclic convolution is reduced up to four times. Furthermore, the corresponding power x delay product is reduced up to 2.4 times, while a simultaneous reduction of area cost is achieved. The particular performance improvement becomes possible by introducing a way to minimize the forward and inverse conversion overhead associated with PRNS. The introduced minimization exploits the fact that for the conversions for particular lengths of data sequences and particular moduli, only multiplications with powers of two and additions are required, thus leading to low implementation complexity. In addition multiple supply voltages are utilized to further reduce power dissipation by more than 30% for particular cases. Formulas that return the applicable supply voltage values per PRNS channel are derived in this paper. Vassilis Paliouras, Alexander Skavantzos, Thanos Stouraitis |
ACM Great Lakes Symposium on VLSI | 1 |
| 2001 | Low-Power Properties of the Logarithmic Number SystemabstractThe potential of reducing power dissipation in a digital system using the logarithmic number system (LNS) is investigated. To provide a quantitative measure of power savings, the equivalence of an LNS to a linear fixed-point system is initially explored. The bit assertion activity of an LNS encoded signal is studied for both uniform and correlated Gaussian inputs. It is shown that LNS reduces the average bit assertion probability by more than 50%, in certain cases, over an equivalent linear representation. Finally, the impact of LNS on the hardware architecture and, thus, to power dissipation, is discussed. It is found that the average number of logic transitions is reduced by several times, for certain arithmetic operations and word lengths, thus compensating the power-dissipation overhead due to the unavoidable linear-to-logarithmic and logarithmic-to-linear conversion. Vassilis Paliouras, Thanos Stouraitis |
IEEE Symposium on Computer Arithmetic | 1 |
| 2001 | Operation-Saving VLSI Architectures for 3D Geometrical TransformationsabstractTwo VLSI architectures for the computationally efficient implementation of the elementary 3D geometrical transformations are introduced. The first one is based on a single floating-point multiply/add unit, while the other one comprises a four processing-element vector unit. By exploiting the structure of the elementary transformation matrices, some of the elements of which are ones and zeros, the proposed architectures avoid full-matrix multiplication for the matrix multiplications involved in the calculation of the transformation matrix by treating them as updates of specific elements, the new values of which are obtained by scalar operations in the case of the single-processor architecture or by simple vector operations in the case of the processor array. Thus, the floating-point operation count and the number of memory accesses required by a transformation are reduced and, therefore, the performance of the circuit which computes the transformation matrix, in terms of execution time, is improved at minimal hardware cost. Furthermore, a circuit is proposed which, for each sequence of transformations, selects the most appropriate direction for computing the product of the matrices in the corresponding stack of transformation matrices in order to further reduce the number of floating-point operations compared to the case where the direction of the computation of the successive matrix products is predetermined. The proposed single-processor architecture is suitable for low-cost applications, while the parallel execution scheme implemented by the introduced parallel processor may be implemented by any four-PE processor with small overhead. Konstantina Karagianni, Vassilis Paliouras, George Diamantakos, Thanos Stouraitis |
IEEE Trans. Computers | 2 |
| 1997 | An operation-saving VLSI geometry engine coreabstractA floating point geometry engine core is introduced in this paper. The proposed core is optimized for performing the 3-D geometrical transformations, including the hardware evaluation of sin x and cos x functions. The architecture exploits the structure of the transformation matrices, thus reducing the number of floating point operations required per transformation. VLSI chip implementation issues for the specific architecture are also discussed. Konstantina Karagianni, George Diamantakos, Vassilis Paliouras, Thanos Stouraitis |
ICASSP | 3 |
| 1995 | A Novel Algorithm for Multi-Operand Logarithmic Number System Addition and Subtraction Using Polynominal ApproximationabstractIn this paper, a novel algorithm for multi-operand Logarithmic Number System (LNS) addition and subtraction is presented. In particular, the computation of the nonlinear functions required for logarithmic addition and subtraction is decomposed into computing 2/sup -x/, log/sub 2/(1+x), some additions, and some shifts. The error behaviour of the algorithm is analyzed, upper bounds of the computational error are provided and it is shown that the introduced Propagation Error Cancellation (PEG) technique and the Error Spectrum Shaping can significantly narrow the error distribution. The inherent parallelism of the proposed algorithm and the pipelinability that exists in the computation of 2/sup -x/ and log/sub 2/(1+x) by using polynomials are exploited by simple VLSI architectures that exhibit important speed-up over the equivalent ROM-based designs. Also, a rule for the determination of the optimal number of pipeline stages is suggested. I. Orginos, Vassilis Paliouras, Thanos Stouraitis |
ISCAS | 2 |
| 1994 | Systematic Design of Multi-Modulus/Multi-Function Residue Number System ProcessorsabstractA methodology for the design of novel Residue Number System (RNS) processors is presented. It results in ROM-less processors, which perform basic residue arithmetic algorithms in more than one moduli channel, either serially or concurrently. Moreover, the proposed architectures achieve area savings, while operating at a high throughput rate. Criteria for selecting the most appropriate moduli of operation are presented. The derived architectures are compared to memory- and full adder-based designs.> Vassilis Paliouras, Thanos Stouraitis |
ISCAS | 1 |
| 1993 | Systematic design of full adder-based architectures for convolution
Dimitrios Soudris, Vassilis Paliouras, Thanos Stouraitis, Alexander Skavantzos, Constantinos E. Goutis |
ICASSP (1) | 2 |
| 1993 | Methodology for the Design of Signed-digit DSP Processors
Vassilis Paliouras, Dimitrios Soudris, Thanos Stouraitis |
ISCAS | 1 |
| 1992 | Systematic development of architectures for multidimensional DSP using the residue number systemabstractA systematic methodology for mapping multidimensional algorithms onto array processor architectures based on the quadratic residue number system is presented. A class of algorithms with separable functions, which can be reduced to the computation of the circular convolution is considered. The array architecture results systematically from a directed graph using partitioning techniques and consists of identical processing elements called inner product step processors. Moreover, due to various graph partitions, many alternative array architectures in terms of I/O constraints, throughput, and hardware complexity can be derived.> Dimitrios Soudris, Vassilis Paliouras, Thanos Stouraitis |
ICASSP | 2 |