Maria Spiropulu

dblp:211/7680 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0001-8172-7081ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2026 HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference
abstract
Lookup-table (LUT) based neural networks can deliver ultra-low latency and excellent hardware efficiency on FPGAs by mapping arithmetic operations directly onto the logic primitives. However, state-of-the-art LUT-aware training (LAT) approaches remain difficult to use in practice: they are often orders of magnitude slower to train than conventional networks, require non-trivial manual tuning for hardware efficiency, and lack an end-to-end work flow. This work presents HGQ-LUT1, a new LAT approach that achieves state-of-the-art hardware efficiency while accelerating training by over 100 times on modern GPUs. HGQ-LUT introduces LUT-Dense and LUT-Conv layers that are implemented with regular, accelerator-Efficient tensor operations during training, which are then compiled into logic LUTs for hardware. By combining these layers with ne-grained, element-wise heterogeneous quantization (including zero-bit pruning) and a LUT-aware resource surrogate, HGQ-LUT enables the automatic exploration of accuracy-resource trade-offs without manual bit-width tuning. We further integrate HGQ-LUT into open-source toolchains, enabling unified design, compilation, and bit-exact verification of hybrid architectures that mix LUT-based with conventional arithmetic blocks. These features make LAT-based DNNs practical for real-world deployment, such as at the CERN Large Hadron Collider’s experiments.
Zhiqiang Que, Bakhtiar Zadeh, Qibin Liu, Kevin H. Alvarez, Wayne Luk, Maria Spiropulu
FCCM7
2026 HGQ: High Granularity Quantization for Real-time Neural Networks on FPGAs
abstract
Model size and inference speed at deployment time, are major challenges in many deep learning applications. A promising strategy to overcome these challenges is quantization. However, a straightforward uniform quantization to very low precision can result in significant accuracy loss. Mixed-precision quantization, based on the idea that certain parts of the network can accommodate lower precision without compromising performance compared to other parts, offers a potential solution. In this work, we present High Granularity Quantization (HGQ), an innovative quantization-aware training method designed to fine-tune the per-weight and per-activation precision in an automatic way for ultra-low latency and low power neural networks which are to be deployed on FPGAs. We demonstrate that HGQ can outperform existing methods by a substantial margin, achieving resource reduction by up to a factor of 20 and latency improvement by a factor of 5 while preserving accuracy.
Zhiqiang Que, Thea Aarrestad, Vladimir Loncar, Jennifer Ngadiuba, Wayne Luk, Maria Spiropulu
FPGA7
2026 da4ml: Distributed Arithmetic for Real-time Neural Networks on FPGAs
abstract
Neural networks with a latency requirement on the order of microseconds, like the ones used at the CERN Large Hadron Collider, are typically deployed on FPGAs fully unrolled and pipelined. A bottleneck for the deployment of such neural networks is area utilization, which is directly related to the required constant matrix-vector multiplication (CMVM) operations. In this work, we propose an efficient algorithm for implementing CMVM operations with distributed arithmetic on FPGAs that simultaneously optimizes for area consumption and latency. The algorithm achieves resource reduction similar to state-of-the-art algorithms while being significantly faster to compute. The proposed algorithm is open sourced and integrated into the hls4ml library, a free and open source library for running real-time neural network inference on FPGAs. We show that the proposed algorithm can reduce on-chip resources by up to a third for realistic, highly quantized neural networks while simultaneously reducing latency, enabling the implementation of previously infeasible networks.
Zhiqiang Que, Vladimir Loncar, Wayne Luk, Maria Spiropulu
ACM Trans. Reconfigurable Technol. Syst.5
2021 Diolkos: improving ethernet throughput through dynamic port selection
abstract
In large networked systems, a sudden increase in traffic could slowdown the network significantly, impacting network quality for multiple users. We present Diolkos, a system that leverages smart switches to dynamically re-reroute data flows in response to drops in performance. In contrast to other techniques, our tool predicts the future throughput at each port in a switch if a data flow were to be sent through it, and updates which port should be taken to maximize throughput. We use several techniques to predict network switch performance on a software defined network (SDN) mimicking topologies commonly found in datacenters. Experimentally, we demonstrate the effectiveness of choosing a port to send flows through based on predicted performance. We found that using a distributed predictive technique achieves a 24% improvement over using a traditional heuristic technique.
Oceane Bel, Joosep Pata, Jean-Roch Vlimant, Nathan R. Tallent, Justas Balcas, Maria Spiropulu
CF6