Dimitrios Stathis 0001

dblp:137/9408 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-5697-4272ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SIMBRAIN: A nonidealities-aware simulation framework for spiking neural networks based on memristor crossbars
Jiawei Xu 0002, Ruisi Shen, Dimitrios Stathis 0001, Lirong Zheng 0001, Zhuo Zou, Ahmed Hemani
Neurocomputing7
2025 MemMIMO: A Simulation Framework for Memristor-Based Massive MIMO Acceleration
abstract
Memristor-based crossbar architectures have proven highly effective for matrix vector multiplication (MVM) operations, making them a promising solution for accelerating the MVMs widely used in precoding algorithms for multiple input multiple output (MIMO) wireless communication systems. However, real-world implementation of memristor-based computing systems face challenges due to commonly observed non-idealities in both the devices themselves and the circuits they’re built into. To facilitate a rapid design flow and investigate the impact of non-idealities, an integrated open-source simulation framework MemMIMO is developed. The simulation framework estimates the accuracy and hardware performance of the computing system, offering a variety of flexible design options. MemMIMO integrates a behavioral model of the mix-signal architecture with a digital front-end. There are three major building blocks in MemMIMO: the device fitting block, the mapping block, and the performance estimation block. These blocks work together to map the complex MVMs in precoding algorithms for MIMO systems to crossbar-based architectures that incorporate memristor models characterized by physical device behavior. Using two typical use cases targeting six-generation (6G) massive MIMO communication as case studies, MemMIMO is used to model different memristor devices, explore the impact of non-idealities on system accuracy, and benchmark circuit-level performance metrics including area, speed, and power.
Jiawei Xu 0002, Dimitrios Stathis 0001, Ruisi Shen, Lirong Zheng 0001, Zhuo Zou, Ahmed Hemani
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 FPGA-Based HPC for Associative Memory System
abstract
Associative memory plays a crucial role in the cognitive capabilities of the human brain. The Bayesian Confidence Propagation Neural Network (BCPNN) is a cortex model capable of emulating brain-like cognitive capabilities, particularly associative memory. However, the existing GPU-based approach for BCPNN simulations faces challenges in terms of time overhead and power efficiency. In this paper, we propose a novel FPGA-based high performance computing (HPC) design for the BCPNN-based associative memory system. Our design endeavors to maximize the spatial and timing utilization of FPGA while adhering to the constraints of the available hardware resources. By incorporating optimization techniques including shared parallel computing units, hybrid-precision computing for a hybrid update mechanism, and the globally asynchronous and locally synchronous (GALS) strategy, we achieve a maximum network size of $150 \times 10$ and a peak working frequency of 100 MHz for the BCPNN-based associative memory system on the Xilinx Alveo U200 Card. The tradeoff between performance and hardware overhead of the design is explored and evaluated. Compared with the GPU counterpart, the FPGA-based implementation demonstrates significant improvements in both performance and energy efficiency, achieving a maximum latency reduction of $33.25 \times$, and a power reduction of over $6.9 \times$, all while maintaining the same network configuration.
Yu Yang 0020, Dimitrios Stathis 0001, Ahmed Hemani, Anders Lansner, Jiawei Xu 0002, Lirong Zheng 0001, Zhuo Zou
ASPDAC4
2024 Exploration of Custom Floating-Point Formats: A Systematic Approach
abstract
The remarkable advancements in AI algorithms over the past three decades have been paralleled by an exponential growth in their complexity, with parameter counts soaring from 60,000 in LeNet during the late 1980s to a staggering 175 billion in ChatGPT 3.0. To mitigate this surge in memory footprint, approximate computing has emerged as a promising strategy, focusing on deploying the minimal resolution necessary to maintain acceptable accuracy. Yet, current practices are hindered by two major challenges: a) the process of identifying the optimal resolution and representation format for each tensor remains a manual, ad hoc task, and b) the representation, typically in floating point (FP) format, is confined to standardized norms predominantly supported by commercial-off-the-shelf (COTS) products like GPUs. This paper tackles these issues by introducing a systematic approach to exploring the FP representation design space to find the ideal FP format for each tensor, thereby leveraging the full potential of FP quantization techniques. It is designed for custom hardware, enabling access to arbitrary FP formats, but also allows users to limit their exploration to standard FP formats, making it compatible with COTS. Additionally, the proposed method explores the Block Floating-Point (BFP) and automatically decides on the size of the blocks. A heuristic-based search method is proposed to handle the large design space. The proposed approach is general, and the heuristic is not biased towards any specific category of algorithms. We apply this method to a Self-Organizing Map (SOM) for bacterial genome identification and LeNet-5 neural network, demonstrating a significant reduction in memory footprint by around 94% and 96%, respectively, compared to the conventional 32-bit FP baseline.
Saba Yousefzadeh, Yu Yang 0020, Astile Peter, Dimitrios Stathis 0001, Ahmed Hemani
DSD4
2022 Reducing the Configuration Overhead of the Distributed Two-level Control System
abstract
With the growing demand for more efficient hardware accelerators for streaming applications, a novel Coarse-Grained Reconfigurable Architecture (CGRA) that uses a Dis-tributed Two-Level Control (D2LC) system has been proposed in the literature. Even though the highly distributed and parallel structure makes it fast and energy-efficient, the single-issue instruction channel between the level-l and level-2 controller in each D2LC cell becomes the bottleneck of its performance. In this paper, we improve its design to mimic a multi-issued architecture by inserting shadow instruction buffers between the level-l and level-2 controllers. Together with a zero-overhead hardware loop, the improved D2LC architecture can enable efficient overlap between loop iterations. We also propose a complete constraint programming based instruction scheduling algorithm to support the above hardware features. The experiment result shows that the improved D2LC architecture can achieve up to 25% of reduction on the instruction execution cycles and 35% reduction on the energy-delay product.
Yu Yang 0020, Dimitrios Stathis 0001, Ahmed Hemani
DATE2
2022 MOHAQ: Multi-Objective Hardware-Aware Quantization of recurrent neural networks
abstract
The compression of deep learning models is of fundamental importance in deploying such models to edge devices. The selection of compression parameters can be automated to meet changes in the hardware platform and application. This article introduces a Multi-Objective Hardware-Aware Quantization (MOHAQ) method, which considers hardware performance and inference error as objectives for mixed-precision quantization. The proposed method feasibly evaluates candidate solutions in a large search space by relying on two steps. First, post-training quantization is applied for fast solution evaluation (inference-only search). Second, we propose the ”beacon-based search” to retrain selected solutions only and use them as beacons to estimate the effect of retraining on other solutions. We use speech recognition models on TIMIT dataset. Experimental evaluations show that Simple Recurrent Unit (SRU)-based models can be compressed up to 8x by post-training quantization without any significant error increase. On SiLago, we found solutions that achieve 97% and 86% of the maximum possible speedup and energy saving, with a minor increase in error on an SRU-based model. On Bitfusion, the beacon-based search reduced the error gain of the inference-only search on SRU-based models and Light Gated Recurrent Unit (LiGRU)-based model by up to 4.9 and 3.9 percentage points, respectively.
Nesma M. Rezk, Tomas Nordström, Dimitrios Stathis 0001, Zain Ul-Abdin, Eren Erdal Aksoy, Ahmed Hemani
J. Syst. Archit.3
2021 Approximate computation of post-synaptic spikes reduces bandwidth to synaptic storage in a model of cortex
Dimitrios Stathis 0001, Yu Yang 0020, Ahmed Hemani, Anders Lansner
DATE1
2021 Synthesis of predictable global NoC by abutment in synchoros VLSI design
abstract
Synchoros VLSI design style has been proposed as an alternative to the standard cell-based design style; the word synchoros is derived from the Greek word choros for space. Synchoricity discretises space with a virtual grid, the way synchronicity discretises time with clock ticks. SiLago (Silicon Lego) blocks are atomic synchoros building blocks like Lego bricks. SiLago blocks absorb all metal layer details, i.e., all wires, to enable composition by abutment of valid; valid in the sense of being technology design rules compliant, timing clean and OCV ruggedized. Effectively, composition by abutment eliminates logic and physical synthesis for the end user. Like Lego system, synchoricity does need a finite number of SiLago block types to cater to different types of designs. Global NoCs are important system level design components. In this paper, we show, how with a small library of SiLago blocks for global NoCs, it is possible to automatically synthesize arbitrary global NoCs of different types, dimensions, and topology. The synthesized global NoCs are not only valid VLSI designs, but their cost metrics (area, latency, and energy) are known with post-layout accuracy in linear time. We argue that this is essential to be able to do chip-level design space exploration. We show how the abstract timing model of such global NoC SiLago blocks can be built and used to analyse the timing of global NoC links with post layout accuracy and in linear time. We validate this claim by subjecting the same VLSI designs of global NoC to commercial EDA's static timing analysis and show that the abstract timing analysis enabled by synchoros VLSI design gives the same results as the commercial EDA tools.
Jordi Altayó González, Dimitrios Stathis 0001, Ahmed Hemani
NOCS2
2020 NACU: A Non-Linear Arithmetic Unit for Neural Networks
abstract
Reconfigurable architectures targeting neural networks are an attractive option. They allow multiple neural networks of different types to be hosted on the same hardware, in parallel or sequence. Reconfigurability also grants the ability to morph into different micro-architectures to meet varying power-performance constraints. In this context, the need for a reconfigurable non-linear computational unit has not been widely researched. In this work, we present a formal and comprehensive method to select the optimal fixed-point representation to achieve the highest accuracy against the floating-point implementation benchmark. We also present a novel design of an optimised reconfigurable arithmetic unit for calculating non-linear functions. The unit can be dynamically configured to calculate the sigmoid, hyperbolic tangent, and exponential function using the same underlying hardware. We compare our work with the state-of-the-art and show that our unit can calculate all three functions without loss of accuracy.
Guido Baccelli, Dimitrios Stathis 0001, Ahmed Hemani, Maurizio Martina
DAC2
2016 Alternative Architectures Toward Reliable Memristive Crossbar Memories
abstract
Resistive random access memory (ReRAM), referred to as memristor, is an emerging memory technology to potentially replace conventional memories, which will soon be facing serious design challenges related to continued scaling. Memristor-based crossbar architecture has been shown to be the best implementation for ReRAM. However, it faces a major challenge related to the sneak current (current sneak paths) flowing through unselected memory cells, which significantly reduces the voltage read margins. In this paper, five alternative architectures (topologies) are applied to minimize the impact of sneak current; the architectures are based on the introduction of insulating junctions within the crossbar. Simulations that were performed while considering different memory accessing aspects, such as bit reading versus word reading, stored data background distribution, crossbar dimensions, etc., showed that read margins can be increased significantly (up to 4×) as compared with standard crossbar architectures. In addition, the proposed architectures eliminate the requirement for extra select devices at each cross point and have no operational complexity overhead.
Ioannis Vourkas, Dimitrios Stathis 0001, Georgios Ch. Sirakoulis, Said Hamdioui
IEEE Trans. Very Large Scale Integr. Syst.2
2015 XbarSim: An educational simulation tool for memristive crossbar-based circuits
abstract
Simulation is expected to become an indispensable educational and research tool for memristive circuits and architectures. To this end, this paper presents a novel, self-contained, platform-independent, GUI-based design and simulation tool for standard/alternative memristive crossbar architectures, targeting memory and/or logic applications. It permits the exploration of the crossbar-based memristive circuit design-space and allows for logic-in-memory computations.
Ioannis Vourkas, Dimitrios Stathis 0001, Georgios Ch. Sirakoulis
ISCAS2
2015 Live demonstration: XbarSim: An educational simulation tool for memristive crossbar-based circuits
abstract
This Live Demonstration is about an interactive software tool developed by the present authors. The tool will run on a personal laptop which the demonstrator will be responsible to bring to the conference site. There are no further special requirements and the mentioned provisions in the presentation booths, i.e. a power plug, a table, and a pin wall, are sufficient.
Ioannis Vourkas, Dimitrios Stathis 0001, Georgios Ch. Sirakoulis
ISCAS2
2013 Improved read voltage margins with alternative topologies for memristor-based crossbar memories
abstract
Memories based on hysteretic resistive materials are expected to have superior properties such as nonvolatility, low power consumption, as well as very high capacity. Crossbar arrays are considered very attractive for future ultimately scaled memories. In this paper, the memristor-based passive crossbar geometry is studied and a set of different topological patterns, which introduce insulating junctions within the memory array, is presented. In the worst-case reading scenario the simulations revealed significantly improved sensed voltage margins (up to > 4×) which alleviate the rigorous requirement for large and highperformance CMOS sensing circuits in passive crossbar memory systems.
Ioannis Vourkas, Dimitrios Stathis 0001, Georgios Ch. Sirakoulis
VLSI-SoC2