EDBT 2026 Demo / reviewers in the wild / expert
Athanasios Vasilopoulos
dblp:300/8910
· DBLP profile ↗
8ranked-venue papers
1as first author
6since 2021 · last 2025
0009-0001-9081-6139ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Mode Borderguard Controllers for Efficient On-Chip Communication in Heterogeneous Digital/Analog Neural Processing UnitsabstractDriven by the growing demand for data-intensive parallel computation, particularly for Matrix-Vector Multiplications (MVMs), and the pursuit of high energy efficiency, Analog In-Memory Computing (AIMC) has garnered significant attention. AIMC addresses the data movement bottleneck by performing MVMs directly within memory, significantly reducing latency and enhancing energy efficiency. Integrating AIMC with digital units for non-MVM operations yields heterogeneous Neural Processing Units (NPUs) that can be combined in a tiled architecture to deliver promising solutions for end-to-end AI inference. Besides powerful heterogeneous NPUs, an efficient on-chip communication infrastructure is also pivotal for inter-node data transmission and efficient AI model execution. This paper introduces the Borderguard Controller (BG-CTRL), a multi-mode, path-through routing controller designed to support three distinct operating modes-time-scheduling, data-driven, and time-sliced data-driven (TSDD)-each offering varying levels of routing flexibility and energy efficiency depending on the data flow patterns and AI model complexity. To demonstrate the design, BG-CTRLs are integrated into a 9-node system of heterogeneous NPUs, arranged in a 3x3 grid and connected using a 2D mesh topology. The system is synthesized using STM 28nm FD-SOI technology. Experimental results show that the BG-CTRL cluster achieves an aggregate throughput of 983 Gb/s, with an energy efficiency of up to 0.41 pJ/B/hop at 0.64 GHz, and a minimal area overhead of 204 kGE. Hong Pang, Carmine Cappetta, Riccardo Massa, Athanasios Vasilopoulos, Elena Ferro, Gamze Islamoglu, Angelo Garofalo, Francesco Conti 0001, Luca Benini, Irem Boybat, Thomas Boesch |
DATE | 4 |
| 2025 | Live Demonstration: Automated DNN Deployment on the IBM HERMES Project ChipabstractFor this demonstration, we will showcase the operation of a software stack capable of automatically deploying Matrix-Vector Matrix (MVM) operations of diverse deep learning workloads in a pipelined-manner on a phase-change memory-based analog in-memory computing chip with high-accuracy. For a real chip, each deployment step will be highlighted for a transformer-based network trained to perform an organic chemical reaction prediction task. Additionally, using an emulated mode of operation, these steps will also be highlighted for a Resnet-based network, which has been trained to perform image classification, and a hybrid CNN/LSTM network trained to infer nucleotide sequences from sequences of amplitude values measured from a sequencing device. Corey Lammie, Julian Büchel, Athanasios Vasilopoulos, Giacomo Camposampiero, Lionel Noussi, William Andrew Simon, Manuel Le Gallo, Abu Sebastian |
ISCAS | 3 |
| 2025 | NIMA: Near In-Memory High-Precision Accumulation Unit for Heterogeneous Analog/Digital Deep Learning AccelerationabstractAnalog In-Memory Computing (AIMC) crossbars often face size limitations that hinder mapping entire neural network layers onto a single AIMC tile. To overcome this, tall layers are typically distributed across multiple tiles or within a single tile by multiplexing different sets of weights at various time intervals. However, this generates partial vector-matrix multiplication (VMM) results that need to be accumulated to produce the final output, underscoring the need for integrated accumulation capabilities in AIMC systems. In this work, we propose a Near In-Memory High-Precision Accumulation Unit (NIMA) with built-in internal and external tile accumulation functionality, positioned at the periphery of the AIMC crossbar. This unit leverages the affine correction capabilities of existing systems and ensures high precision during partial VMM accumulation. We have physically implemented NIMA in a 14nm CMOS technology, providing comprehensive performance and area evaluations and comparisons with prior art. We investigate and compare different layer mappings and accumulation schemes supported by the proposed unit, evaluating both latency and Mean Absolute Error (MAE). We further assess the precision of the proposed unit on large layers of ResNet9 and ResNet32 for image classification on CIFAR10/CIFAR100 datasets. Irem Sanli, Elena Ferro, Athanasios Vasilopoulos, Thomas Boesch, Abu Sebastian, Irem Boybat |
ISCAS | 3 |
| 2025 | A framework for analog-digital mixed-precision neural network training and inferenceabstractRecent advancements in AI hardware highlight the potential of mixed-signal accelerators, which integrate analog computation for matrix multiplications with reduced-precision digital operations, to achieve superior performance and energy efficiency. In this paper, we present a framework designed to perform hardware-aware training and inference evaluation of neural networks (NNs) on such accelerators. This framework extends an existing toolkit, the IBM Analog Hardware Acceleration Kit (AIHWKit), using a quantization library, enabling flexible layer-wise deployment in either analog or digital units, the latter with configurable precision and quantization options. Our combined framework supports simultaneous quantization-and analog-aware training as well as post-training calibration routines. It can also evaluate the accuracy of NNs when deployed on mixed-signal accelerators. We demonstrate the need of such a framework through ablation studies on a ResNet-based vision model and a BERT-based language model, highlighting the importance of its functionality for maximizing accuracy during deployment. Our contribution is open-sourced as part of the core code of AIHWKit [1]. Athanasios Vasilopoulos, Emma Boulharts, Corey Lammie, Julian Büchel, Hadjer Benmeziane, Manuel Le Gallo, Abu Sebastian |
ISCAS | 1 |
| 2024 | A Precision-Optimized Fixed-Point Near-Memory Digital Processing Unit for Analog In-Memory ComputingabstractAnalog In-Memory Computing (AIMC) is an emerging technology for fast and energy-efficient Deep Learning (DL) inference. However, a certain amount of digital post-processing is required to deal with circuit mismatches and non-idealities associated with the memory devices. Efficient near-memory digital logic is critical to retain the high area/energy efficiency and low latency of AIMC. Existing systems adopt Floating Point 16 (FP16) arithmetic with limited parallelization capability and high latency. To overcome these limitations, we propose a Near-Memory digital Processing Unit (NMPU) based on fixed-point arithmetic. It achieves competitive accuracy and higher computing throughput than previous approaches while minimizing the area overhead. Moreover, the NMPU supports standard DL activation steps, such as ReLU and Batch Normalization. We perform a physical implementation of the NMPU design in a 14 nm CMOS technology and provide detailed performance, power, and area assessments. We validate the efficacy of the NMPU by using data from an AIMC chip and demonstrate that a simulated AIMC system with the proposed NMPU outperforms existing FP16- based implementations, providing 139 × speed-up, 7.8 × smaller area, and a competitive power consumption. Additionally, our approach achieves an inference accuracy of 86.65 %/65.06 %, with an accuracy drop of just 0.12 %/0.4 % compared to the FP16 baseline when benchmarked with ResNet9/ResNet32 networks trained on the CIFAR10/CIFAR100 datasets, respectively. Elena Ferro, Athanasios Vasilopoulos, Corey Lammie, Manuel Le Gallo, Luca Benini, Irem Boybat, Abu Sebastian |
ISCAS | 2 |
| 2024 | Improving the Accuracy of Analog-Based In-Memory Computing Accelerators Post-TrainingabstractAnalog-Based In-Memory Computing (AIMC) inference accelerators can be used to efficiently execute Deep Neural Network (DNN) inference workloads. However, to mitigate accuracy losses, due to circuit and device non-idealities, Hardware-Aware (HWA) training methodologies must be employed. These typically require significant information about the underlying hardware. In this paper, we propose two Post-Training (PT) optimization methods to improve accuracy after training is performed. For each crossbar, the first optimizes the conductance range of each column, and the second optimizes the input, i.e, Digital-to-Analog Converter (DAC), range. It is demonstrated that, when these methods are employed, the complexity during training, and the amount of information about the underlying hardware can be reduced, with no notable change in accuracy (≤0.1%) when finetuning the pretrained RoBERTa transformer model for all General Language Understanding Evaluation (GLUE) benchmark tasks. Additionally, it is demonstrated that further optimizing learned parameters PT improves accuracy. Corey Lammie, Athanasios Vasilopoulos, Julian Büchel, Giacomo Camposampiero, Manuel Le Gallo, Malte J. Rasch, Abu Sebastian |
ISCAS | 2 |
| 2006 | Calculating distortion in active CMOS mixers using Volterra seriesabstractIn this paper, an intermodulation distortion analysis for active CMOS mixers based on Volterra series theory is presented. As an outcome of the analysis, a software tool that facilitates a fast optimization of the mixer's linearity performance has been developed. Results from the tool confirm excellent match with those obtained by popular commercial SPICE-like simulators, whereas simulation time is lowered by an order of magnitude Gerasimos Theodoratos, Athanasios Vasilopoulos, Georgios Vitzilaios, Yannis Papananos |
ISCAS | 2 |
| 2006 | A low-voltage CMOS LNA with multiple magnetic feedback for WLAN applicationsabstractA CMOS low-noise-amplifier (LNA) topology that utilizes multiple monolithic transformer magnetic feedback is presented. The proposed topology permits negative and positive feedback to be applied constructively, in order to simultaneously neutralize the gate-drain overlap capacitance of the amplifying transistor and achieve high gain at high frequencies when driving an on-chip capacitance. It also allows for a stable design with adequate gain and large reverse isolation without noise figure (NF) degradation. Simulation results indicate voltage conversion gain of 17 dB, NF of 1.6 dB and third-order input intercept point (IIP3) of 13 dBm. The design is being implemented in a 0.13 mum CMOS technology Georgios Vitzilaios, Yannis Papananos, Gerasimos Theodoratos, Athanasios Vasilopoulos |
ISCAS | 4 |