Marvin Damschen

dblp:155/9822 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
1since 2021 · last 2025
0000-0002-6236-5799ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-authorArtificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 26% Parallel and multicore computing · 21% GPUs and heterogeneous computing · 13%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
0.312018
Co-Scheduling on Fused CPU-GPU Architectures With Shared Last Level Caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Parallel and multicore computing › parallel scheduling
coscheduling
0.312018
Co-Scheduling on Fused CPU-GPU Architectures With Shared Last Level Caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
fused CPU-GPU architecture
0.312018
Co-Scheduling on Fused CPU-GPU Architectures With Shared Last Level Caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Memory systems › cache
shared last-level cache
0.312018
Co-Scheduling on Fused CPU-GPU Architectures With Shared Last Level Caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Processor architecture and microarchitecture › instruction set architecture › instruction set extension
custom instruction selection
0.212016
Extending the WCET Problem to Optimize for Runtime-Reconfigurable Processors · ACM Trans. Archit. Code Optim. 2016
Reconfigurable computing and FPGAs › dynamic reconfiguration
runtime reconfigurable processor
0.212016
Extending the WCET Problem to Optimize for Runtime-Reconfigurable Processors · ACM Trans. Archit. Code Optim. 2016
Electronic design automation
timing analysis
0.212016
Extending the WCET Problem to Optimize for Runtime-Reconfigurable Processors · ACM Trans. Archit. Code Optim. 2016
Embedded and real-time systems
worst-case execution time analysis
0.212016
Extending the WCET Problem to Optimize for Runtime-Reconfigurable Processors · ACM Trans. Archit. Code Optim. 2016
Parallel and multicore computing › parallel scheduling
heterogeneous scheduling
0.112018
Co-Scheduling on Fused CPU-GPU Architectures With Shared Last Level Caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Parallel and multicore computing
work distribution
0.112018
Co-Scheduling on Fused CPU-GPU Architectures With Shared Last Level Caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018

Methods — techniques the papers use, named apart from their topics

shared virtual memory · 0.3host-side profiling · 0.3OpenCL 2.0 · 0.3program path analysis · 0.2integer linear programming · 0.2greedy heuristic · 0.2
YearPublicationVenuePosition
2025 SAFE-COLOR: Color Fidelity Benchmarks and Thresholds for Safety-Critical Object Detection
abstract
Color fidelity is often overlooked in simulation-based validation for autonomous vehicles, yet even minor color mismatches can undermine the reliability of AI-driven perception systems. In this paper, we systematically examine how controlled deviations in color reproduction-quantified by$\Delta E$-affect object detection accuracy across 32 variants of YOLO. Using a Macbeth ColorChecker, we derive calibrations for key color transforms (brightness, contrast, hue, gamma, saturation and color bias) and apply these to the COCO validation set. Our evaluations demonstrate that increasing$\Delta E$yields significant drops in detection metrics, especially for safety-critical categories such as pedestrians and cyclists. Based on these findings, we propose$\Delta E$thresholds that define acceptable color fidelity in camera simulations (e.g.,$\Delta E\leq 3$for$\Delta \mathbf{mAP}\leq 1\%)$. Further-more, we contribute these transformed datasets and scripts as a publicly available benchmark, enabling reproducible comparisons and guiding future research on color-based vulnerabilities in automated driving and other safety-critical domains.
Marvin Damschen, Ramana Reddy Avula, Mazen Mohamad
IV1
2019 WCET Guarantees for Opportunistic Runtime Reconfiguration
abstract
Time-critical systems need to be analyzable for timing guarantees. There is an increasing demand for predictable performance that modern processor architectures fail to provide since they focus on average-case performance only. Recent work has demonstrated that runtime reconfiguration of hardware accelerators via an FPGA is a viable way to achieve high performance for optimized worst-case execution time (WCET) guarantees. Since execution of the worst-case path is highly improbable, configuring accelerators for this path costs reconfigurable area that could better be used to accelerate more probable paths. This work presents the first approach that comprises (1) an online average-case execution time (ACET) optimization while (2) maintaining the optimized WCET guarantee utilizing reconfigurable accelerators. We achieve this by a new design-time technique which determines the runtime slack bounds that allow speculative reconfiguration of accelerators that benefit the ACET. Combined with an online slack monitoring approach that introduces negligible overheads by using a performance counter, we show a runtime reduction of up to 10.4% for a complex and real-world application on top of an already-optimized WCET guarantee.
Marvin Damschen, Lars Bauer, Jörg Henkel
ICCAD1
2018 Co-Scheduling on Fused CPU-GPU Architectures With Shared Last Level Caches
abstract
Fused CPU-GPU architectures integrate a CPU and general-purpose GPU on a single die. Recent fused architectures even share the last level cache (LLC) between CPU and GPU. This enables hardware-supported byte-level coherency. Thus, CPU and GPU can execute computational kernels collaboratively, but novel methods to co-schedule work are required. This paper contributes three dynamic co-scheduling methods. Two of our methods implement workers that autonomously acquire work from a common set of independent work items (similar to bag- f-tasks scheduling). The third method, host-side profiling, uses a fraction of the total work of a kernel to determine a ratio of how to distribute work to CPU and GPU based on profiling. The resulting ratio is used for the following executions of the same kernel. Our methods are realized using OpenCL 2.0, which introduces fine-grained shared virtual memory (SVM) to allocate coherent memory between CPU and GPU. We port the Rodinia Benchmark Suite, a standard suite for heterogeneous computing, to fine-grained SVM and fused CPU-GPU architectures (Rodinia-SVM). We evaluate the overhead of fine-grained SVM and analyze the suitability of OpenCL 2.0's new features for co-scheduling. Our host-side profiling method performs competitively to the optimal choice of executing kernels either on CPU or GPU (hypothetical xor-Oracle). On average, it achieves 97% of xor-Oracle's performance and a 1.43× speedup over using the GPU alone (standard in Rodinia). We show, however, that in most cases it is not beneficial to split the work of a kernel between CPU and GPU compared to exclusively running it on the most suitable single compute device. For a fixed amount of work per device, cache-related stalls can increase by up to 1.75× when both devices are used in parallel instead of exclusively while cache misses remain the same. Thus, not the cost of cache conflicts, but inefficient cache coherence is a major performance bottleneck for current fused CPU-GPU Intel architectures with shared LLC.
Marvin Damschen, Frank Mueller 0001, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 Preemption of the Partial Reconfiguration Process to Enable Real-Time Computing With FPGAs
abstract
To improve computing performance in real-time applications, modern embedded platforms comprise hardware accelerators that speed up the task’s most compute-intensive parts. A recent trend in the design of real-time embedded systems is to integrate field-programmable gate arrays (FPGA) that are reconfigured with different accelerators at runtime, to cope with dynamic workloads that are subject to timing constraints. One of the major limitations when dealing with partial FPGA reconfiguration in real-time systems is that the reconfiguration port can only perform one reconfiguration at a time: if a high-priority task issues a reconfiguration request while the reconfiguration port is already occupied by a lower-priority task, the high-priority task has to wait until the current reconfiguration is completed (a phenomenon known as priority inversion ), unless the current reconfiguration is aborted (introducing unbounded delays in low-priority tasks, a phenomenon known as starvation ). This article shows how priority inversion and starvation can be solved by making the reconfiguration process preemptive —that is, allowing it to be interrupted at any time and resumed at a later time without restarting it from scratch. Such a feature is crucial for the design of runtime reconfigurable real-time systems but not yet available in today’s platforms. Furthermore, the trade-off of achieving a guaranteed bound on the reconfiguration delay for low-priority tasks and the maximum delay induced for high-priority tasks when preempting an ongoing reconfiguration has been identified and analyzed. Experimental results on the Xilinx Zynq-7000 platform show that the proposed implementation of preemptive reconfiguration introduces a low runtime overhead, thus effectively solving priority inversion and starvation.
Enrico Rossi, Marvin Damschen, Lars Bauer, Giorgio C. Buttazzo, Jörg Henkel
ACM Trans. Reconfigurable Technol. Syst.2
2017 Timing Analysis of Tasks on Runtime Reconfigurable Processors
abstract
Real-time embedded systems need to be analyzable for timing guarantees. Despite significant scientific advances, however, timing analysis lags years behind current microarchitectures with out-of-order scheduling pipelines, several hardware threads, and multiple (shared) cache layers. To satisfy the increasing performance demands, analyzable performance features are required. We propose a novel timing analysis approach to introduce runtime reconfigurable instruction set processors as one way to escape the scarcity of analyzable performance while preserving the flexibility of the system. We introduce extensions to the state-of-the-art Integer linear programming (ILP)-based program path analysis for computing precise worst case time bounds in the presence of the widely used technique to continue processor execution during reconfiguration by emulating not yet reconfigured custom instructions (CIs) in software. We identify and safely bound a timing anomaly of runtime reconfiguration, where executing faster than worst case time during reconfiguration extends the execution time of the whole program. Stalling the processor during reconfiguration (easier to analyze but not state-of-the-art for reconfigurable processors) is not required in our approach. Finally, we show the precision of our analysis on a complex multimedia application with multiple reconfigurable CIs for several hardware parameters and give advice on how to deal with reconfiguration delay under timing guarantees.
Marvin Damschen, Lars Bauer, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Extending the WCET Problem to Optimize for Runtime-Reconfigurable Processors
abstract
The correctness of a real-time system does not depend on the correctness of its calculations alone but also on the non-functional requirement of adhering to deadlines. Guaranteeing these deadlines by static timing analysis, however, is practically infeasible for current microarchitectures with out-of-order scheduling pipelines, several hardware threads, and multiple (shared) cache layers. Novel timing-analyzable features are required to sustain the strongly increasing demand for processing power in real-time systems. Recent advances in timing analysis have shown that runtime-reconfigurable instruction set processors are one way to escape the scarcity of analyzable processing power while preserving the flexibility of the system. When moving calculations from software to hardware by means of reconfigurable custom instructions (CIs)—additional to a considerable speedup—the overestimation of a task’s worst-case execution time (WCET) can be reduced. CIs typically implement functionality that corresponds to several hundred instructions on the central processing unit (CPU) pipeline. While analyzing instructions for worst-case latency may introduce pessimism, the latency of CIs—executed on the reconfigurable fabric—is precisely known. In this work, we introduce the problem of selecting reconfigurable CIs to optimize the WCET of an application. We model this problem as an extension to state-of-the-art integer linear programming (ILP)-based program path analysis. This way, we enable optimization based on accurate WCET estimates with integration of information about global program flow, for example, infeasible paths. We present an optimal solution with effective techniques to prune the search space and a greedy heuristic that performs a maximum number of steps linear in the number of partitions of reconfigurable area available. Finally, we show the effectiveness of optimizing the WCET on a reconfigurable processor by evaluating a complex multimedia application with multiple reconfigurable CIs for several hardware parameters.
Marvin Damschen, Lars Bauer, Jörg Henkel
ACM Trans. Archit. Code Optim.1
2015 Transparent offloading of computational hotspots from binary code to Xeon Phi
Marvin Damschen, Heinrich Riebler, Gavin Vaz, Christian Plessl
DATE1