EDBT 2026 Demo / reviewers in the wild / expert
Patrick Schmidt 0003
dblp:50/5159-3
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-0931-1230ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DSEParted: Co-Optimization of Embedded NPU Architectures and Neural Network PartitioningabstractConvolutional Neural Networks (CNNs) have become an essential tool in the domain of vision processing. However, dedicated accelerators are needed for energy-efficient execution of these networks, especially for embedded devices with tight energy constraints. Integrating multiple of these accelerators via chiplets promises a way to scale up the performance of these emerging systems by partitioning a neural network across multiple accelerators. This approach enables the execution of different layers on an accelerator with the best-suited dataflow. However, partitioning a neural network is a non-trivial task, especially when different accelerator architectures must be considered. In this paper, we propose our framework DSEParted, which automates the co-design of network partitioning and hardware architecture optimization. It employs a hierarchical optimization approach to gradually reduce the number of design candidates until an optimal system configuration is found for a partitioned computation of a neural network. We demonstrate that our framework can design a system that, for GoogLeNet, reduces latency by 22.5% when optimizing for latency. In addition, when optimizing for energy, it reduces the system area by 7.9%, with no impact on latency or energy compared to a baseline system. Further, we show that our partitioning-aware pruning strategy can reduce the EDP of the system by up to 49.7% in the case of ResNeXt-50, compared to a strategy that only optimizes the accelerators individually. Through the provided information, designers receive feedback on the efficiency of the full system at an early development stage. Our work is available open source1.1https://github.com/itiv-kit/cnn-parted Patrick Schmidt 0003, Fabian Kreß, Alexey Serdyuk, Matthias Stammler, Tanja Harbaum, Jürgen Becker 0001 |
DSD | 1 |
| 2025 | A Pixel Histogram-Based Safety Mechanism and Fault Detection Methodology for a Robust Image Signal Processor
Julian Höfer, Patrick Schmidt 0003, Hella Toto-Kiesa, Sebastian Höfer, Gregor Schewior, Dietmar Engelke, Karl-Heinz Eickel, Darius Grantz, Tanja Harbaum, Jürgen Becker 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | ZuSE-KI-Mobil: AI Chip Design Platform for Automotive and Industrial Applications
Shaown Mojumder, Simon Friedrich, Emil Matús, Matthias Lüders, Martin Friedrich, Oliver Renke, Holger Blume, Markus Kock, Gregor Schewior, Darius Grantz, Jens Benndorf, Julian Höfer, Patrick Schmidt 0003, Jürgen Becker 0001, Nael Fasfous, Pierpaolo Morì, Hans-Jörg Vögel, Samira Ahmadifarsani, Leonidas Kontopoulos, Ulf Schlichtmann, Yun-Jin Li, Gerhard P. Fettweis |
IEEE Trans. Very Large Scale Integr. Syst. | 13 |
| 2024 | EMDRIVE Architecture: Embedded Distributed Computing and Diagnostics from Sensor to EdgeabstractFuture automotive architectures are expected to transition from a network-centric to a domain-centered architecture featuring central compute units. Powerful domain controllers or smart sensors alleviate the load on these central units and communication systems. These controllers execute tasks with varying criticalities on heterogeneous multicore processors, and are ideally capable of dynamically balancing the computing load between the central unit and sensors. Here, Artificial Intelligence (AI) capabilities playa crucial role, as it is in high demand for such an automotive architecture. However, AI still requires specialized accelerators to improve their computation performance. Task-oriented distributed computing with criticalities up to ASIL-D necessitates the development and utilization of specialized methodologies, such as safety, through the isolation and abstraction of low-level hardware concepts. Meanwhile, online monitoring and diagnostics become vital features to detect errors during operation. The EMDRIVE architecture includes methods, components, and strategies to enhance the performance, safety, and security of such distributed computing platforms. The nationally funded EMDRIVE project connects its twelve partners from academia and industry and is currently in its intermediate stage. Patrick Schmidt 0003, Iuliia Topko, Matthias Stammler, Tanja Harbaum, Jürgen Becker 0001, Rico Berner, Omar Ahmed, Jakub Jagielski, Thomas Seidler, Markus Abel, Marius Kreutzer, Maximilian Kirschner, Victor Pazmino Betancourt, Robin Sehm, Lukas Groth, Andrija Neskovic, Rolf Meyer, Saleh Mulhem, Mladen Berekovic, Matthias Probst, Manuel Brosch, Georg Sigl, Thomas Wild, Matthias Ernst, Andreas Herkersdorf, Florian Aigner, Stefan Hommes, Sebastian Lauer, Maximilian Seidler, Thomas Raste, Gasper Skvarc Bozic, Ibai Irigoyen Ceberio, Albrecht Mayer |
DATE | 1 |
| 2024 | Ph.D. Project: Compiler-Driven Hardware/Software Co- Design for Embedded AIabstractAhstract- The increasing computational complexity of AI work-loads has led to the introduction of numerous accelerator archi-tectures. However, these designs often neglect the software tooling necessary to generate optimized mappings. Additionally, future compute architectures are expected to become more heteroge-neous, resulting in additional challenges regarding mapping of tasks to compute units. In this Ph.D. project, a novel methodology to generate compilers with an optimization pipeline for custom hardware accelerators is presented. The foundation of the work is a new Architecture Description Language (ADL) used to capture architectural details necessary for the compilation process. Based on this, simulators on both core and system level are generated that assists in the design of the underlying hardware. Finally, to assist with compilation, a flow to automatically register the new hardware design to a retargetable compiler will be designed. For the optimization pipeline, we focus on two specific aspects: Tensorization to map instruction sequences to the accelerator, and automated scheduling to optimize the execution time and energy. Patrick Schmidt 0003, Jürgen Becker 0001 |
FCCM | 1 |
| 2023 | CNNParted: An open source framework for efficient Convolutional Neural Network inference partitioning in embedded systems
Fabian Kreß, Vladimir Sidorenko, Patrick Schmidt 0003, Julian Höfer, Tim Hotfilter, Iris Fürst-Walter, Tanja Harbaum, Jürgen Becker 0001 |
Comput. Networks | 3 |