VLDB 2026 Research / reviewers in the wild / expert
Vasilis Sakellariou
dblp:296/1062 · also Vasileios Sakellariou
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-2581-3193ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Theory of computation · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Mixed-Precision RNS DNN AcceleratorabstractThe Residue Number System (RNS) has been used for the design of Deep Neural Network (DNN) processing architectures due to its efficient implementation of the multiply-accumulate (MAC) operation. Prior-art RNS DNN accelerators have demonstrated notable benefits compared to conventional fixed-point (FXP) representations for arithmetic precisions of at least 8 bits. However, advanced quantization techniques have recently enabled accurate ultra-low-precision FXP DNN inference. Thus, it remains an open research question whether RNS can still outperform FXP representations for smaller precisions and especially in mixed-precision (MXP) quantization settings, where optimal bit-width configurations with respect to overall accuracy drop constraints are sought. This work addresses this gap by presenting an RNS-based MXP DNN accelerator that supports 3–8-bit quantization and consistently achieves superior model performance vs. hardware cost tradeoffs for various DNN models, resulting in up to 1.2× energy efficiency improvements compared to the FXP counterpart. Synthesized on a 22-nm technology, the RNS MXP accelerator achieves 6.93–14.58 TOPS/W, outperforming the state-of-the-art uniform-precision RNS accelerator by 1.4× while maintaining the original model accuracy, as well as mixed-precision FXP accelerators. Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis |
ISCAS | 1 |
| 2025 | A Survey and Comparative Analysis of Number Systems for Deep Neural NetworksabstractDeep neural networks (DNNs) are indispensable in various artificial intelligence (AI) applications. However, their inherent complexity presents significant challenges, particularly when deploying them on resource-constrained devices. To overcome these hurdles, academia and industry are actively seeking ways to accelerate and optimize DNN implementations. A significant area of research revolves around discovering more effective methods to represent the enormous data volumes processed by DNNs. Traditional number systems (NSs) have proven nonoptimal for this task, prompting extensive exploration into alternative and bespoke systems for DNNs. This survey aims to comprehensively discuss various NSs utilized to efficiently represent DNN data. These systems are categorized mainly based on their impact on DNN performance and hardware implementation. This survey offers an overview of these categorized NSs and delves into different subsystems within each, outlining their effect on DNN performance and hardware design. Furthermore, these systems are compared quantitatively and qualitatively concerning their expected quantization error, memory utilization, and computational requirements. This survey also emphasizes the challenges linked with each system and the diverse proposed solutions to address them. Insights into the utilization of these NSs for sophisticated DNNs are also presented in this survey. Readers will acquire a deeper understanding of the importance of efficient NSs for DNNs, explore commonly used systems, comprehend the tradeoffs between these systems, delve into design considerations influencing their impact on DNN performance, and discover recent trends and potential research avenues in this field. Ghada Alsuhli, Vasilis Sakellariou, Hani Saleh, Mahmoud Al-Qutayri, Baker Mohammad, Thanos Stouraitis |
Proc. IEEE | 2 |
| 2023 | Improving Residue-Level Sparsity in RNS-based Neural Network Hardware Accelerators via RegularizationabstractResidue Number System (RNS) has recently attracted interest for the hardware implementation of inference in machine-learning systems as it provides promising trade-offs in the area, time, and power dissipation space. In this paper we introduce a technique that utilizes regularization during training, and increases the percentage of residues which are zero, when the parameters of an artificial neural network (ANN) are expressed in an RNS. The proposed technique can also be used as a post-processing stage, allowing the optimization of pre-trained models for RNS implementation. By increasing the number of residues being zero, i.e., residue-level sparsity, the proposed technique facilitates new hardware architectures for RNS-based inference, allowing new trade-offs and improving performance over prior art without practically compromising accuracy. The introduced method increases residue sparsity by a factor of 4× to 6× in certain cases. Emmanouil Kavvousanos, Vasilis Sakellariou, Ioannis Kouretas, Vassilis Paliouras, Thanos Stouraitis |
ARITH | 2 |
| 2023 | A multiplier-Free RNS-Based CNN accelerator exploiting bit-Level sparsityabstractIn this work, a Residue Numbering System (RNS)-based Convolutional Neural Network (CNN) accelerator utilizing a multiplier-free distributed-arithmetic Processing Element (PE) is proposed. A method for maximizing the utilization of the arithmetic hardware resources is presented. It leads to an increase of the system's throughput, by exploiting bit-level sparsity within the weight vectors. The proposed PE design takes advantage of the properties of RNS and Canonical Signed Digit (CSD) encoding to achieve higher energy efficiency and effective processing rate, without requiring any compression mechanism or introducing any approximation. An extensive design space exploration for various parameters (RNS base, PE micro-architecture, encoding) using analytical models as well as experimental results from CNN benchmarks is conducted and the various trade-offs are analyzed. A complete end-to-end RNS accelerator is developed based on the proposed PE. The introduced accelerator is compared to traditional binary and RNS counterparts as well as to other state-of-the-art systems. Implementation results in a 22-nm process show that the proposed PE can lead to 1.85× and 1.54× more energy-efficient processing compared to binary and conventional RNS, respectively, with a 1.88× maximum increase of effective throughput for the employed benchmarks. Compared to a state-of-the-art, all-digital, RNS-based system, the proposed accelerator is 8.87× and 1.11× more energy- and area-efficient, respectively. Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis |
ARITH | 1 |
| 2022 | A High-performance RNS LSTM blockabstractThe Residue Number System (RNS) has been proposed as an alternative to conventional binary representations for use in AI hardware accelerators. While it has been successfully utilized in applications targeting Convolutional Neural Networks (CNNs), its usage in other network models such as Recurrent Neural Networks (RNNs) has been set back due to the difficulty of implementing more complex activations functions like tanh and sigmoid ($\sigma$) in the RNS domain. In this paper, we seek to extend its usage in such models, and in particular LSTM networks, by providing efficient RNS implementations of the activation functions. To this aim, we derive improved accuracy piecewise linear approximations of the tanh and $\sigma$ functions using the minimax approach and propose a fully RNS-based hardware realization. We show that our approximations can effectively mitigate accuracy degradation in LSTM networks compared to naive approximations, while the RNS LSTM block can be up to 40% more efficient in terms of performance per area unit compared to a binary counterpart, when used in high performance-targeted accelerators. Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis |
ISCAS | 1 |
| 2021 | An FPGA Accelerator for Spiking Neural Network Simulation and TrainingabstractSpiking Neural Networks (SNNs) have recently been employed to solve a number of machine learning problems traditionally addressed by classical Artificial Neural Networks (ANNs). SNNs are different than ANNs as they incorporate time in their computations and information is encoded in the exact timing or frequency of discrete events, spikes. SNNs promise to deliver a higher energy efficiency than ANNs, when implemented in neuromorphic hardware, due to their event-driven nature. Naturally, this different computational model they introduce, creates different design challenges for the implementation of large-scale networks and needs to be addressed by different architectures. The spiking accelerator presented here aims to facilitate and accelerate the process of developing SNNs for ML applications that are traditionally addressed by ANNs and help bridge the accuracy gap between them. It achieves a significant speedup of up to 800× for inference and up to 500× for training compared to software SNN simulations for certain set-ups. Vasilis Sakellariou, Vassilis Paliouras |
ISCAS | 1 |