Konstantinos Balaskas

dblp:226/4009 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0003-2886-6896ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 2 first-author · 11 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Design and Optimization of Mixed-Kernel Mixed-Signal SVMs for Flexible Electronics
abstract
Flexible Electronics (FE) have emerged as a promising alternative to silicon-based technologies, offering on-demand low-cost fabrication, conformality, and sustainability. However, their large feature sizes severely limit integration density, imposing strict area and power constraints, thus prohibiting the realization of Machine Learning (ML) circuits, which can significantly enhance the capabilities of relevant near-sensor applications. Support Vector Machines (SVMs) offer high accuracy in such applications at relatively low computational complexity, satisfying FE technologies’ constraints. Existing SVM designs rely solely on linear or Radial Basis Function (RBF) kernels, forcing a trade-off between hardware costs and accuracy. Linear kernels, implemented digitally, minimize overhead but sacrifice performance, while the more accurate RBF kernels are prohibitively large in digital, and their analog realization contains inherent functional approximation. In this work, we propose the first mixed-kernel and mixed-signal SVM design in FE, which unifies the advantages of both implementations and balances the cost/accuracy trade-off. To that end, we introduce a co-optimization approach that trains our mixed-kernel SVMs and maps binary SVM classifiers to the appropriate kernel (linear/RBF) and domain (digital/analog), aiming to maximize accuracy whilst reducing the number of costly RBF classifiers. Our designs deliver 7.7% higher accuracy than state-of-the-art single-kernel linear SVMs, and reduce area and power by 108× and 17× on average compared to digital RBF implementations.
Florentia Afentaki, Maha Shatta, Konstantinos Balaskas, Georgios Panagopoulos, Georgios Zervakis 0001, Mehdi Baradaran Tahoori
DATE3
2025 Late Breaking Results: Energy-Efficient Printed Machine Learning Classifiers with Sequential SVMs
abstract
Printed Electronics (PE) provide a mechanically flexible and cost-effective solution for machine learning (ML) circuits, compared to silicon-based technologies. However, due to large feature sizes, printed classifiers are limited by high power, area, and energy overheads, which restricts the realization of battery-powered systems. In this work, we design sequential printed bespoke Support Vector Machine (SVM) circuits that adhere to the power constraints of existing printed batteries while minimizing energy consumption, thereby boosting battery life. Our results show 6.5x energy savings while maintaining higher accuracy compared to the state of the art.
Spyridon Besias, Ilias Sertaridis, Florentia Afentaki, Konstantinos Balaskas, Georgios Zervakis 0001
DATE4
2025 Late Breaking Results: Leveraging Approximate Computing for Carbon-Aware DNN Accelerators
abstract
The rapid growth of Machine Learning (ML) has increased demand for DNN hardware accelerators, but their embodied carbon footprint poses significant environmental challenges. This paper leverages approximate computing to design sustainable accelerators by minimizing the Carbon Delay Product (CDP). Using gate-level pruning and precision scaling, we generate area-aware approximate multipliers and optimize the accelerator design with a genetic algorithm. Results demonstrate reduced embodied carbon while meeting performance and accuracy requirements.
Aikaterini Maria Panteleaki, Konstantinos Balaskas, Georgios Zervakis 0001, Hussam Amrouch, Iraklis Anagnostopoulos
DATE2
2025 Computing with Printed and Flexible Electronics
Mehdi Baradaran Tahoori, Georgios Zervakis 0001, Konstantinos Balaskas, Priyanjana Pal
ETS4
2025 R2T-Tiny: Runtime-Reconfigurable Throughput-Optimized TinyML for Hybrid Inference Acceleration on FPGA SoCs
abstract
The emergence of tiny machine learning (TinyML) has represented a paradigm shift toward energy efficient and on-device inference, with TinyML research primarily focusing on low-cost and energy-efficiency microcontroller units (MCUs). However, the low computational capabilities of MCUs greatly limit the performance of TinyML applications, especially in the case of throughput driven tasks. Field programmable gate arrays (FPGAs), are therefore a promising alternative to MCUs, offering high parallelism and low energy requirements. FPGA-based accelerators typically follow one of two directions: high-throughput resource-intensive pipelined designs, or low-throughput sequential systolic arrays. Incorporating both approaches into a hybrid streaming/sequential accelerator strikes a balance between resource efficiency and throughput, but incurs resource contention within a tiny FPGA. In this work, we address the aforementioned limitations and propose a throughput-driven hybrid acceleration methodology for TinyML. We introduce R2T-Tiny, an adaptive framework that brings layer-wise customizability into throughput-driven inference. By leveraging runtime partial reconfiguration, R2T-Tiny dynamically adjusts the accelerator type and applies tailored approximations per layer, achieving high throughput while adhering to the tight resource constraints of tiny embedded FPGAs. Our comprehensive evaluation on popular TinyML benchmarks showcases the capabilities of our framework in achieving high throughput inference on the PYNQ-Z2 FPGA board, increasing throughput by an average of 1.6x across 3 popular deep neural networks (DNNs) in the TinyML domain, when compared to DNNDK, a systolic array based accelerator from Xilinx, while incurring less than 1% accuracy loss.
Georgios Mentzos, Valentin Alexander Frey, Konstantinos Balaskas, Georgios Zervakis 0001, Jörg Henkel
ICCAD3
2025 Invited Paper: Feature-to-Classifier Co-Design for Mixed-Signal Smart Flexible Wearables for Healthcare at the Extreme Edge
abstract
Flexible Electronics (FE) offer a promising alternative to rigid silicon-based hardware for wearable healthcare devices, enabling lightweight, conformable, and low-cost systems. However, their limited integration density and large feature sizes impose strict area and power constraints, making ML-based healthcare systems–integrating analog frontend, feature extraction and classifier–particularly challenging. Existing FE solutions often neglect potential system-wide solutions and focus on the classifier, overlooking the substantial hardware cost of feature extraction and Analog-to-Digital Converters (ADCs)–both major contributors to area and power consumption. In this work, we present a holistic mixed-signal feature-to-classifier co-design framework for flexible smart wearable systems. To the best of our knowledge, we design the first analog feature extractors in FE, significantly reducing feature extraction cost. We further propose an hardware-aware NAS-inspired feature selection strategy within ML training, enabling efficient, application-specific designs. Our evaluation on healthcare benchmarks shows our approach delivers highly accurate, ultra-area-efficient flexible systems–ideal for disposable, low-power wearable monitoring.
Maha Shatta, Konstantinos Balaskas, Paula L. Duarte, Georgios Panagopoulos, Mehdi Baradaran Tahoori, Georgios Zervakis 0001
ICCAD2
2025 Compact Yet Highly Accurate Printed Classifiers Using Sequential Support Vector Machine Circuits
abstract
Printed Electronics (PE) technology has emerged as a promising alternative to silicon-based computing. It offers attractive properties such as on-demand ultra-low-cost fabrication, mechanical flexibility, and conformality. However, PE are governed by large feature sizes, prohibiting the realization of complex printed Machine Learning (ML) classifiers. Leveraging PE’s ultra-low non-recurring engineering and fabrication costs, designers can fully customize hardware to a specific ML model and dataset, significantly reducing circuit complexity. Despite significant advancements, state-of-the-art solutions achieve area efficiency at the expense of considerable accuracy loss. Our work mitigates this by designing area- and power-efficient printed ML classifiers with little to no accuracy degradation. Specifically, we introduce the first sequential Support Vector Machine (SVM) classifiers, exploiting the hardware efficiency of bespoke control and storage units and a single Multiply-Accumulate compute engine. Our SVMs yield on average 6x lower area and 4.6% higher accuracy compared to the printed state of the art.
Ilias Sertaridis, Spyridon Besias, Florentia Afentaki, Konstantinos Balaskas, Georgios Zervakis 0001
ISCAS4
2025 Exploration of Low-Power Flexible Stress Monitoring Classifiers for Conformal Wearables
abstract
Conventional stress monitoring relies on episodic, symptom-focused interventions, missing the need for continuous, accessible, and cost-efficient solutions. State-of-the-art approaches use rigid, silicon-based wearables, which, though capable of multitasking, are not optimized for lightweight, flexible wear, limiting their practicality for continuous monitoring. In contrast, flexible electronics (FE) offer flexibility and low manufacturing costs, enabling real-time stress monitoring circuits. However, implementing complex circuits like machine learning (ML) classifiers in FE is challenging due to integration and power constraints. Previous research has explored flexible biosensors and ADCs, but classifier design for stress detection remains underexplored. This work presents the first comprehensive design space exploration of low-power, flexible stress classifiers. We cover various ML classifiers, feature selection, and neural simplification algorithms, with over 1200 flexible classifiers. To optimize hardware efficiency, fully customized circuits with low-precision arithmetic are designed in each case. Our exploration provides insights into designing real-time stress classifiers that offer higher accuracy than current methods, while being low-cost, conformable, and ensuring low power and compact size.
Florentia Afentaki, Sri Sai Rakesh Nakkilla, Konstantinos Balaskas, Paula L. Duarte, Shiyi Jiang, Georgios Zervakis 0001, Farshad Firouzi, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori
ISLPED3
2025 Energy-Aware Heterogeneous Federated Learning via Approximate DNN Accelerators
abstract
In Federated Learning (FL), devices that participate in the training usually have heterogeneous resources, i.e., energy availability. In current deployments of FL, devices that do not fulfill certain hardware requirements are often dropped from the collaborative training. However, dropping devices in FL can degrade training accuracy and introduce bias or unfairness. Several works have tackled this problem on an algorithm level, e.g., by letting constrained devices train a subset of the server neural network (NN) model. However, it has been observed that these techniques are not effective w.r.t. accuracy. Importantly, they make simplistic assumptions about devices’ resources via indirect metrics, such as multiply accumulate (MAC) operations or peak memory requirements. We observe that memory access costs (that are currently not considered in simplistic metrics) have a significant impact on the energy consumption. In this work, for the first time, we consider on-device accelerator design for FL with heterogeneous devices. We utilize compressed arithmetic formats and approximate computing, targeting to satisfy limited energy budgets. Using a hardware-aware energy model, we observe that, contrary to the state of the art’s moderate energy reduction, our technique allows for lowering the energy requirements (by$4\times $) while maintaining higher accuracy.
Kilian Pfeiffer, Konstantinos Balaskas, Kostas Siozios, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Variability-Aware Approximate Circuit Synthesis via Genetic Optimization
abstract
One of the major barriers that CMOS devices face at nanometer scale is increasing parameter variation due to manufacturing imperfections. Process variations severely inhibit the reliable operation of circuits, as the operational frequency at the nominal process corner is insufficient to suppress timing violations across the entire variability spectrum. To avoid variability-induced timing errors, previous efforts impose pessimistic and performance-degrading timing guardbands atop the operating frequency. In this work, we employ approximate computing principles and propose a circuit-agnostic automated framework for generating variability-aware approximate circuits that eliminate process-induced timing guardbands. Variability effects are accurately portrayed with the creation of variation-aware standard cell libraries, fully compatible with standard EDA tools. The underlying transistors are fully calibrated against industrial measurements from Intel 14nm FinFET in which both electrical characteristics of transistors and variability effects are accurately captured. In this work, we explore the design space of approximate variability-aware designs to automatically generate circuits of reduced variability and increased performance without the need for timing guardbands. Experimental results show that by introducing negligible functional error of merely$\boldsymbol {5.3 \times 10^{-3}}$, our variability-aware approximate circuits can be reliably operated under process variations without sacrificing the application performance.
Konstantinos Balaskas, Florian Klemme, Georgios Zervakis 0001, Kostas Siozios, Hussam Amrouch, Jörg Henkel
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 Automated Design Approximation to Overcome Circuit Aging
abstract
Transistor aging phenomena manifest themselves as degradations in the main electrical characteristics of transistors. Over time, they result in a significant increase of cell propagation delay, leading to errors due to timing violations, since the operating frequency becomes unsustainable as the circuit ages. Conventional techniques employ timing guardbands to mitigate aging-induced delay increase, which leads to considerable performance losses from the beginning of the circuit’s lifetime. Leveraging the inherent error resilience of a vast number of application domains, approximate computing was recently introduced as an aging mitigation mechanism. In this work, we present the first automated framework for generatingaging-aware approximate circuits. Our framework, by applying directed gate-level netlist approximation, induces a small functional error and recovers the delay degradation due to aging. As a result, our optimized circuits eliminate aging-induced timing errors. Experimental evaluation over a variety of arithmetic circuits and image processing benchmarks demonstrates that for an average error of merely$5\times 10^{-3}$, our framework completely eliminates aging-induced timing guardbands. Compared to the respective baseline circuits without timing guardbands (i.e., iso-performance evaluation), the error of the circuits generated by our framework is$1208\times $smaller.
Konstantinos Balaskas, Georgios Zervakis 0001, Hussam Amrouch, Jörg Henkel, Kostas Siozios
IEEE Trans. Circuits Syst. I Regul. Pap.1