Alexey Serdyuk

dblp:323/3333 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0009-4064-6491ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Multi-Partner Project: CeCaS Accelerator Design for Efficient Supercomputing in Automotive Systems
abstract
Modern vehicles integrate an increasing amount of computational functionality, driven by the growing complexity of in-vehicle applications. At the same time, automotive system architectures are becoming more centralized, requiring powerful HPC platforms at the core. These platforms must deliver the performance needed for ADAS, AI, and autonomous driving, while also meeting stringent energy efficiency and safety requirements.The CeCaS project addresses these challenges across a wide range of topics and domains of expertise, including processor design in advanced FinFET technology, the transformation of the E/E architecture, and advanced packaging for automotive supercomputing platforms. Within CeCaS, our work focuses on application-specific accelerator design to enable efficient processing of compute-intensive workloads. In this paper, we present our contributions in this area, including the design of hardware accelerators for both conventional and neuromorphic AI workloads, the development and evaluation of representative AI benchmarks, and the use of virtual platforms for early design-space exploration and hardware/software co-design.
Annina Gutermann, Alexey Serdyuk, Fabian Lesniak, Julian Höfer, Hella Toto-Kiesa, Tanja Harbaum, Jürgen Becker 0001, Brian Pachideh, Sven Nitzsche, Moritz Neher, Carmen Weigelt, Jann Krausse, Victor Pazmino Betancourt, Klaus Knobloch, Lukas Groth, Andrija Neskovic, Saleh Mulhem, Mladen Berekovic
DATE2
2025 DSEParted: Co-Optimization of Embedded NPU Architectures and Neural Network Partitioning
abstract
Convolutional Neural Networks (CNNs) have become an essential tool in the domain of vision processing. However, dedicated accelerators are needed for energy-efficient execution of these networks, especially for embedded devices with tight energy constraints. Integrating multiple of these accelerators via chiplets promises a way to scale up the performance of these emerging systems by partitioning a neural network across multiple accelerators. This approach enables the execution of different layers on an accelerator with the best-suited dataflow. However, partitioning a neural network is a non-trivial task, especially when different accelerator architectures must be considered. In this paper, we propose our framework DSEParted, which automates the co-design of network partitioning and hardware architecture optimization. It employs a hierarchical optimization approach to gradually reduce the number of design candidates until an optimal system configuration is found for a partitioned computation of a neural network. We demonstrate that our framework can design a system that, for GoogLeNet, reduces latency by 22.5% when optimizing for latency. In addition, when optimizing for energy, it reduces the system area by 7.9%, with no impact on latency or energy compared to a baseline system. Further, we show that our partitioning-aware pruning strategy can reduce the EDP of the system by up to 49.7% in the case of ResNeXt-50, compared to a strategy that only optimizes the accelerators individually. Through the provided information, designers receive feedback on the efficiency of the full system at an early development stage. Our work is available open source1.1https://github.com/itiv-kit/cnn-parted
Patrick Schmidt 0003, Fabian Kreß, Alexey Serdyuk, Matthias Stammler, Tanja Harbaum, Jürgen Becker 0001
DSD3
2024 KIHT: Kaligo-Based Intelligent Handwriting Teacher
abstract
Kaligo-based Intelligent Handwriting Teacher (KIHT) is a bi-nationally funded research project. The aim of this joint project is to develop an intelligent learning device for automated handwriting, composed of existing components, which can be made available to as many students as possible. With KIHT, we specifically address the challenging task of using inertial sensors to retrace the trajectory of a pen without relying on external reference systems. The nearly unlimited freedom to let the pen glide over the paper has not yet provided a satisfactory solution to this challenge in the state-of-the-art methods, even with sophisticated algorithms and AI approaches. The final phase of the project is now being launched and together with partners from industry and academia, we are taking a holistic approach by considering the entire chain of components, from the pen to the embedded processing system, the algorithms and the app.
Tanja Harbaum, Alexey Serdyuk, Fabian Kreß, Tim Hamann, Jens Barth, Peter Kämpf, Florent Imbert, Yann Soullard, Romain Tavenard, Éric Anquetil, Jessica Delahaie
DATE2
2023 ATLAS: An Approximate Time-Series LSTM Accelerator for Low-Power IoT Applications
abstract
Enabling the use of Deep Neural Networks (DNNs) for time-series-based applications on low-power devices such as wearables opens up a wide range of new features and services. However, inference requires an enormous amount of operations to be performed by the computing platform. In addition, Long Short-Term Memory (LSTM)-based networks require memory to store the internal cell state for future calculations. In this paper, we therefore propose a hardware/software co-design based low-power LSTM hardware accelerator architecture for Internet of Things (IoT) applications called ATLAS. The design is based on approximate computing techniques to reduce the power consumption and inference latency by achieving high accuracy. Exemplary, we investigate the impact of applying our proposed architecture to a DNN for handwriting recognition. Thereby, we can show that the accuracy decreases only slightly when the inference is executed on ATLAS. The low power consumption is achieved by a minimal design requiring 173 LUTs, 67 FFs, one DSP, and one BRAM on a Xilinx FPGA. As a result, ATLAS enables the efficient use of LSTM-based DNNs in IoT devices.
Fabian Kreß, Alexey Serdyuk, Micha Hiegle, Disnebio Waldmann, Tim Hotfilter, Julian Höfer, Tim Hamann, Jens Barth, Peter Kämpf, Tanja Harbaum, Jürgen Becker 0001
DSD2
2023 A Unified SoC Lab Course: Combined Teaching of Mixed Signal Aspects, System Integration, Software Development and Documentation
abstract
University courses for System-on-Chip (SoC) design mostly focus on particular aspects. Whereas this can provide detailed understanding of these aspects, it neglects system integration specific topics such as crossing digital and analog domains. In addition, many courses skip practical issues and do not teach Electronic Design Automation (EDA) tools. This is reasonable in the context of a specialized course, but omitting these techniques often prevents students from making active use of the learned knowledge in their own projects. In the following, we present our technological platform to teach SoC design in a holistic lab course: The course takes students from writing the first line of Verilog code to advanced digital and analog design. It introduces simulation of digital and analog systems, debugging methods for software and hardware, CPU bus architecture, custom peripherals and driver development. This work is prototyped using Field Programmable Gate Arrays (FPGAs) and later transferred to an ASIC target, covering standard cell synthesis and analog layout. At the end of the semester, students finalize the project with technical documentation writing. The course is built on the design of an audio peripheral, combining all topics in a single real-world system. It enables students to apply theoretical aspects from various lectures in the SoC curriculum in practice and equips them with the skills needed to dive deeper into each of the involved topics on their own.
Johannes Pfau, Richard Leys, Marc Neu, Alexey Serdyuk, Ivan Peric, Jürgen Becker 0001
ISCAS4