Martin Geier 0001

dblp:65/2971-1 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
2since 2021 · last 2021
0000-0003-2481-9873ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 5 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Embedded and real-time systems · 100%
Computer networks
1 paper
Wireless networking · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Embedded and real-time systems › automotive embedded systems
automotive e/e architectures
0.212013
Let's put the car in your phone! · DAC 2013
Embedded and real-time systems
cyber-physical system platforms
0.212013
Let's put the car in your phone! · DAC 2013
Wireless networking › wireless transmission
radio link
0.012013
Let's put the car in your phone! · DAC 2013
Wireless networking
wireless link
0.012013
Let's put the car in your phone! · DAC 2013
YearPublicationVenuePosition
2021 Timing-Predictable Vision Processing for Autonomous Systems
abstract
Vision processing for autonomous systems today involves implementing machine learning algorithms and vision processing libraries on embedded platforms consisting of CPUs, GPUs and FPGAs. Because many of these use closed-source proprietary components, it is very difficult to perform any timing analysis on them. Even measuring or tracing their timing behavior is challenging, although it is the first step towards reasoning about the impact of different algorithmic and implementation choices on the end-to-end timing of the vision processing pipeline. In this paper we discuss some recent progress in developing tracing, measurement and analysis infrastructure for determining the timing behavior of vision processing pipelines implemented on state-of-the-art FPGA and GPU platforms.
Tanya Amert, Michael Balszun, Martin Geier 0001, F. Donelson Smith, James H. Anderson, Samarjit Chakraborty
DATE3
2021 Insert & Save: Energy Optimization in IP Core Integration for FPGA-based Real-time Systems
abstract
Today, many industrial, automotive and autonomous systems like robots are deployed in high-temperature and battery-powered environments. Due to cooling and runtime, this limits the energy consumption and makes the design of such embedded real-time systems even more challenging. Though Field Programmable Gate Arrays (FPGAs) offer the required performance, their static and - load-dependent - dynamic energy consumptions continue to prevent a widespread adoption. The existing methods for dynamic power reduction (like clock gating) are either limited in savings or require disruptive changes to well-established FPGA design flows. Whilst the former is caused by optimizing on fabric level only, the latter is due to the lack of support for a more efficient (but not yet mature and standardized) high-level design entry in current tools. In this paper, we thus explore an optimization methodology based on an existing, but not-fully-utilized intermediate level of abstraction that emerges in the IP core integration phase of the design. To this end, we exploit the fact that the vast majority of FPGA-based real-time processing pipelines is not exclusively assembled using a single type of design entry - i.e., neither entirely hand-written nor high-level synthesis only. Instead, suitable IP cores (from a variety of sources) are integrated via standardized bus interfaces such as AXI, Avalon or Wishbone. To facilitate effort- and power-efficient clock gating on integration-level, we present two “insert and save” IP cores that harness application information extracted from current AXI3 and AXI4-Stream interfaces. Based thereon, both cores precisely control the clock signals of every downstream processing stage for maximum energy savings. This approach not only nicely integrates with today's predominantly AXI-based designs but also results in clock gating structures that are particularly suitable for current FPGAs - as demonstrated by experimental evaluations on a Zynq-based Visual Servoing System with energy savings of 26%.
Martin Geier 0001, Marian Brändle, Samarjit Chakraborty
RTAS1
2020 Predictable Vision for Autonomous Systems
abstract
In this perspective cum case-study paper, we argue the need for designing timing-predictable vision processing algorithms for autonomous systems. Many core functions in systems like autonomous vehicles involve computer vision within a control loop. Designing such closed-loop controllers and guaranteeing their performance requires the vision processing to be predictable. But this is challenging given the multitude of choices when implementing vision processing algorithms, and the heterogeneity of the architectures (involving GPUs and FPGAs) on which such algorithms are implemented. Towards this, we report a tracing and measurement infrastructure we have been building and illustrate its potential utility using a case study.
Michael Balszun, Martin Geier 0001, Samarjit Chakraborty
ISORC2
2020 Debugging FPGA-accelerated Real-time Systems
abstract
The high computation/communication requirements along with reliability needs and limited power budgets necessitate complex processing platforms for emerging autonomous systems. Due to the current focus on performance, however, such platforms are increasingly difficult to predict and analyze. This holds true in terms of both performance (e.g., behavioral and temporal) aspects and power consumption. Ensuring functional safety thus requires new techniques to analyze performance, predictability and power. In this paper, we thus propose a novel hybrid tracing methodology to monitor (and, subsequently, optimize) temporal, functional and energy-related properties of Real-time Systems (RTSs). We target current Programmable SoCs (pSoCs) integrating a fixed-function System-on-Chip (SoC) with flexible Field Programmable Gate Array (FPGA) fabric. Although such heterogeneous systems are well suited for high-end, mixed-hardware/software real-time pipelines, they also offer more complex performance/energy trade-offs than software-only platforms. To systematically exploit this complexity, we present a resource-efficient trace IP core for the pSoC’s fabric and an external measurement/interface system – jointly capturing hybrid power/state traces for subsequent (i.e., offline) analysis. By fusing state data from our IP core with events-of-interest gathered from power traces of pSoC and co-monitored I/O components, we gain a holistic view on temporal RTS aspects. Events and synchronized multi-rail power data jointly extend the debugging coverage via automated identification of processing phases, computation of energy baselines, and estimation of potential savings. Our solution thus integrates functional, temporal and energy monitoring into a single, unified workflow, which, in contrast to traditional separate tools, delivers valuable new insights helpful during debugging and reduces both cost and effort. Experimental evaluations on a Zynq-based Visual Servoing System show the method’s various benefits.
Martin Geier 0001, Marian Brändle, Dominik Faller, Samarjit Chakraborty
RTAS1
2019 Cost-Effective Energy Monitoring of a Zynq-Based Real-Time System Including Dual Gigabit Ethernet
abstract
Recent FPGA architectures integrate various power management features already established in CPU-driven SoCs to reach more energy-sensitive application domains such as, e.g., automotive and robotics. This also qualifies hybrid Programmable SoCs (pSoCs) that combine fixed-function SoCs with configurable FPGA fabric for heterogeneous Real-time Systems (RTSs), which operate under predefined latency and power constraints in safety-critical environments. Their complex application-specific computation and communication (incl. I/O) architectures result in highly varying power consumption, which requires precise voltage and current sensing on all relevant supply rails to enable dependable evaluation of available and novel power management techniques. In this paper, we propose a low-cost 18-channel 16-bit-resolution measurement system capable of over 200 kSPS (kilo-samples per second) for instrumentation of current pSoC development boards. In addition, we propose to include crucial I/O components such as Ethernet PHYs into the power monitoring to gain a holistic view on the RTS's temporal behavior covering not only computation on FPGA and CPUs, but also communication in terms of, e.g., reception of sensor values and transmission of actuation signals. We present an FMC-sized implementation of our measurement system combined with two Gigabit Ethernet PHYs and one HDMI input. Paired with Xilinx' ZC702 development board, we are able to synchronously acquire power traces of a Zynq pSoC and the two PHYs precise enough to identify individual Ethernet frames.
Martin Geier 0001, Dominik Faller, Marian Brändle, Samarjit Chakraborty
FCCM1
2018 Hardware-accelerated data acquisition and authentication for high-speed video streams on future heterogeneous automotive processing platforms
abstract
With the increasing use of Ethernet-based communication backbones in safety-critical real-time domains, both efficient and predictable interfacing and cryptographically secure authentication of high-speed data streams are becoming very important. Although the increasing data rates of in-vehicle networks allow the integration of more demanding (e.g., camera-based) applications, processing speeds and, in particular, memory bandwidths are no longer scaling accordingly. The need for authentication, on the other hand, stems from the ongoing convergence of traditionally separated functional domains and the extended connectivity both in- (e.g., smart-phones) and outside (e.g., telemetry, cloud-based services and vehicle-to-X technologies) current vehicles. The inclusion of cryptographic measures thus requires careful interface design to meet throughput, latency, safety, security and power constraints given by the particular application domain. Over the last decades, this has forced system designers to not only optimize their software stacks accordingly, but also incrementally move interface functionalities from software to hardware. This paper discusses existing and emerging methods for dealing with high-speed data streams ranging from software-only via mixed-hardware/software approaches to fully hardware-based solutions. In particular, we introduce two approaches to acquire and authenticate GigE Vision Video Streams at full line rate of Gigabit Ethernet on Programmable SoCs suitable for future heterogeneous automotive processing platforms.
Martin Geier 0001, Fabian Franzen, Samarjit Chakraborty
ICCAD1
2013 Let's put the car in your phone!
abstract
Today high-end cars have extremely complex E/E architectures -- with 50--100 electronic control units (ECUs), connected by communication buses like CAN, FlexRay and Ethernet. They are used to run several (control) applications with many million lines of code. We propose a radically new architecture where all these applications are instead run on a mobile phone being carried by the driver. The car now has a considerably simpler architecture with few or no ECUs, using RF links to connect sensors and actuators to the mobile phone with a powerful multicore processor. We discuss the advantages and challenges and describe a small prototype implementation with an adaptive cruise control application.
Martin Geier 0001, Martin Becker 0001, Daniel Yunge, Benedikt Dietrich, Reinhard Schneider 0001, Dip Goswami, Samarjit Chakraborty
DAC1
2010 High-level timing analysis of concurrent applications on MPSoC platforms using memory-aware trace-driven simulations
abstract
Due to the growing complexity of multiprocessor systems-on-chip (MPSoCs), there is an increasing demand on efficient design space exploration techniques. In addition to the analysis of diverse hardware architectures, these techniques should assist the designer in the flexible evaluation of various scheduling policies and application mappings while taking effects of the shared on-chip communication infrastructure into account. Most available simulation approaches are either unable to cover all these aspects jointly or have poor simulation performance. In this paper, we present a framework for timing analysis of MPSoC architectures using abstract and yet accurate traces. The traces capture both precise processing latencies and memory access patterns and represent application- and OS-related workload. Performance estimation is performed by an interleaved execution of the traces on a highly configurable multiprocessor platform modeled in our trace-driven SystemC TLM simulator. Using the flexible scheduler model presented in this paper, various mappings and scheduling policies can be rapidly evaluated while considering on-chip interconnect contention and usage of shared resources. Due to the abstraction of the trace-driven simulations, the proposed framework allows for both fast and accurate explorations of MPSoC design alternatives.
Roman Plyaskin, Alejandro Masrur, Martin Geier 0001, Samarjit Chakraborty, Andreas Herkersdorf
VLSI-SoC3