Joachim Falk

dblp:26/1092 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
5since 2021 · last 2024
0009-0006-0834-3237ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 1 since 2021Computer networks · 3 · 1 since 2021Theory of computation · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 Self-Powering Dataflow Networks - Concepts and Implementation
abstract
Dataflow networks play a vital role in modeling and analyzing stream-processing systems in an analytic way, including digital signal and image processing systems. In this paper, we first present a system-level approach to synthesize such dataflow networks automatically to systems of communicating hardware actors connected by FIFO buffers. Although such data-triggered networks of (internally clocked) actors can achieve very high throughputs, the potential to power actors down in times of unavailability of data has not been addressed so far in any research. Here, we show that by refinement of the firing state machine of each actor in a given network, we enable the design of self-powering dataflow networks while exploiting either clock gating or power gating as a means to save power in times of inactivity of each individual actor in a network. The gains of self-powering dataflow networks in terms of power and energy savings when powering down and up actors dynamically is shown for different data arrival patterns and rates in detailed experiments for multiple IoT system applications. These systems are often working in normally-off mode and woken up only upon the availability of data. For these, drastic energy savings are reported.
Abrarul Karim, Joachim Falk, Dennis Schmidt, Jürgen Teich
MEMOCODE2
2022 Grant Prediction-based Dynamic Power Management for 5G to Reduce Mobile Device Energy Consumption
abstract
Reducing the energy consumption of mobile phones is an essential design goal. Whereas Dynamic Power Management (DPM) techniques have been proposed for cellular modems supporting the 5G NR protocol standard, these are purely reactive in nature. For the LTE protocol standard, there also exist approaches that predict grant-free idle intervals during which the modem can be switched off. In this paper, we investigate predictive DPM techniques for the novel 5G protocol standard by exploiting two particular opportunities: First, we may expect micro sleep states can be activated more often in comparison to reactive DPM techniques. Second, the 5G protocol standard offers additional data scheduling techniques using cross-slot and multi-slot scheduling. To select an appropriate prediction model for the resulting complex traffic scenarios, we evaluate six classifiers (Random Forest, Naive Bayes, Decision Tree, K-Nearest Neighbors, Feed-Forward Neural Network, and AdaBoost) for the grant prediction problem. Subsequently, each classifier is evaluated in terms of prediction accuracy and overall energy savings, taking into account the classifier's own energy requirement and timing constraints. As a result, Random Forest proved to be the superior classifier with achievable energy savings of up to 49 % for same-slot scheduling, up to 28 % for cross-slot scheduling, and up to 48 % for multi-slot scheduling, with False Negative Rates (FNRs) ranging between 0.001 - 0.12.
Peter Brand, Benjamin Hackenberg, Joachim Falk, Jürgen Teich
IWCMC3
2021 Multi-Step Ahead Grant Prediction for Dynamic Power Management in Cellular Modems
abstract
Reducing the energy consumption is an important objective in the development of cellular modems for LTE and 5G standards. In addition to hardware optimizations, Dynamic Power Management (DPM) techniques that aim at powering down idle system components are crucial for achieving this goal. However, most DPM techniques proposed so far for mobile devices are purely reactive. More promising are recently proposed proactive approaches that predict whether communication with the cellular base station will occur in the next time-slot or not. Due to transition times between power states, however, such single-step approaches may not exploit deeper power-saving states of components. As a remedy, this paper proposes multi-step ahead prediction techniques that forecast periods of no communication lasting multiple time-slots. In this context, we define a formal fine-grained power model and introduce novel predictive power management policies. Furthermore, we explore the impact of the forecast period–the number of predicted time-slots–in terms of false negative error rate as well as achievable energy savings compared to a standard LTE-compliant reactive DPM. Finally, it is shown that this overall energy saving potential of the proposed multi-step predictive approach may be higher by up to a factor of 3 compared to a single-step predictive approach without incurring a higher false negative error rate. Alternatively, the multi-step approach may allow a reduction of the false negative error rate by a factor of 4.1 compared to the single-step approach without reduction of expectable energy savings. In fact, the observed low false negative error rate of 0.0083 may facilitate a realization in future modem solutions.
Peter Brand, Joachim Falk, Eduard Potwigin, Jürgen Teich
ISNCC2
2021 Adaptive Predictive Power Management for Mobile LTE Devices
abstract
Reducing the energy consumption of mobile phones is a crucial design goal for cellular modem solutions for LTE and 5G NR standards. Most dynamic power management techniques targeting mobile devices proposed so far, however, are purely reactive in powering down and up system components. Promising approaches extend this, by predicting information from the cell and the communication protocol to take decisions proactively. In this paper, we present a complete proactive power management approach for the modem based on on-line grant prediction. In this context, we define proactive policies that allow a mobile device to go to sleep states more often compared to reactive power management systems, e.g., in time slots of predicted transmission inactivity in a cell. Furthermore, we propose and compare two algorithmic solutions to this proactive grant prediction problem, one a feed-forward neural network and one a SARSA-λ reinforcement agent. As the implementation of these machine learning techniques also creates additional energy and resource costs, both approaches are carefully designed, optimized, and evaluated not only in terms of prediction accuracy, but also in terms of overall energy savings. Notably, our predictor implementations are able to achieve up to 17 percent in overall energy savings on real-world traces.
Peter Brand, Joachim Falk, Jonathan Ah Sue, Johannes Brendel, Ralph Hasholzner, Jürgen Teich
IEEE Trans. Mob. Comput.2
2021 Multi-objective Optimization of Mapping Dataflow Applications to MPSoCs Using a Hybrid Evaluation Combining Analytic Models and Measurements
abstract
Dataflow modeling is well suited for a large variety of applications for modern multi-core architectures, e.g., from the signal processing and the control domain. Furthermore, Design Space Exploration (DSE) can be used to explore mappings of tasks to hardware resources (cores of an MPSoC) and their scheduling to obtain optimized trade-off solutions between throughput and resource costs. However, the throughput evaluation of an implementation candidate via compilation-in-the-loop or simulation-based approaches can be extremely time-consuming. Such a deficiency is very detrimental, because a typical DSE run needs to evaluate thousands of implementation candidates. As a remedy, we propose a hybrid-adaptive DSE where a max-plus algebra-based analytic throughput calculation method is used in the initial DSE phase to enable a fast progress of the search space exploration. However, as this analysis may be inaccurate as neglecting some real-world effects like cache and scheduling overhead, throughput measurements are taken later in the DSE. Moreover, we explore the trade-off between scheduling efficiency of implementation candidates—in favor of reducing concurrency—and exploiting concurrency to a large extent for parallel execution of the application. To find solutions of highest achievable throughput, it is shown that not only highly scheduling efficient implementation candidates but also highly parallel implementation candidates are essential when determining the initial population. In this realm, we contribute a method for diversity-based population initialization. For a representative set of benchmarks, it is shown that the combination of the two major contributions allows us to find much higher throughput multi-core solutions within a given exploration time compared to a state-of-the-art DSE approach.
Martín Letras, Joachim Falk, Tobias Schwarzer, Jürgen Teich
ACM Trans. Design Autom. Electr. Syst.2
2020 Clustering-Based Scenario-Aware LTE Grant Prediction
abstract
Reducing the energy consumption of mobile phones is a crucial design goal for cellular modem solutions for LTE and 5G standards. Recent approaches for dynamic power management incorporate traffic prediction to power down components of the modem as often as possible. These predictive approaches have been shown to still provide substantial energy savings, even if trained purely on-line. However, a higher prediction accuracy could be achieved when performing predictor training off-line. Additionally, having pre-trained predictors opens up the ability to successfully employ predictive techniques also in less favorable situations such as short intervals of stable traffic patterns. For this purpose, we introduce a notion of similarity, based on which a clustering is performed to identify similar traffic patterns. For each resulting cluster, i.e., an identified traffic scenario, one predictor is designed and trained off-line. At run time, the system selects the pre-trained predictor with the lowest average short-term false negative rate allowing for energy-efficient and highly accurate on-line prediction. Through experiments, it is shown that the presented mixed static/dynamic approach is able to improve the prediction accuracy and energy savings compared to a state-of-the-art approach by factors of up to 2 and up to 1.9, respectively.
Peter Brand, Muhammad Sabih, Joachim Falk, Jonathan Ah Sue, Jürgen Teich
WCNC3
2019 On the Analytic Evaluation of Schedules via Max-Plus Algebra for DSE of Multi-Core Architectures
abstract
Dataflow modeling is well suited for a wide variety of applications for multi-core architectures, e.g. signal processing and control domain. Additionally, Design Space Exploration (DSE) can be used to explore the distribution of tasks to resources and their scheduling to obtain optimized trade-off solutions between throughput and resource costs. However, the performance evaluation of an implementation candidate in particular via compilation and throughput measurement on the target hardware is prohibitively time-consuming. Thus, we propose to use a max-plus algebra-based analytic throughput calculation method in the initial DSE phase where a fast evaluation with low accuracy is sufficient to guide the search through the design space. However, this analysis neglects some real-world concerns like cache effects and scheduling overhead. Thus, a hybrid DSE is proposed where throughput measurements are taken later in the DSE to get more accurate throughput results for real-world platforms. Results show that our approach is able to find much higher throughput multi-core solutions within a given exploration time compared to a state-of-the-art DSE approach.
Martín Letras, Joachim Falk, Tobias Schwarzer, Jürgen Teich
SCOPES2
2019 Compilation of Dataflow Applications for Multi-Cores using Adaptive Multi-Objective Optimization
abstract
State-of-the-art system synthesis techniques employ meta-heuristic optimization techniques for Design Space Exploration (DSE) to tailor application execution, e.g., defined by a dataflow graph, for a given target platform. Unfortunately, the performance evaluation of each implementation candidate is computationally very expensive, in particular on recent multi-core platforms, as this involves compilation to and extensive evaluation on the target hardware. Applying heuristics for performance evaluation on the one hand allows for a reduction of the exploration time but on the other hand may deteriorate the convergence of the optimization technique toward performance-optimal solutions with respect to the target platform. To address this problem, we propose DSE strategies that are able to dynamically trade off between (i) approximating heuristics to guide the exploration and (ii) accurate performance evaluation, i.e., compilation of the application and subsequent performance measurement on the target platform. Technically, this is achieved by introducing a set of additional, but easily computable guiding objective functions, and varying the set of objective functions that are evaluated during the DSE adaptively. One major advantage of these guiding objectives is that they are generically applicable for dataflow models without having to apply any configuration techniques to tailor their parameters to the specific use case. We show this for synthetic benchmarks as well as a real-world control application. Moreover, the experimental results demonstrate that our proposed adaptive DSE strategies clearly outperform a state-of-the-art DSE approach known from literature in terms of the quality of the gained implementations as well as exploration times. Amongst others, we show a case for a two-core implementation where after about 3 hours of exploration time one of our proposed adaptive DSE strategies already obtains a 60% higher performance value than obtained by the state-of-the-art approach. Even when the state-of-the-art approach is given a total exploration time of more than 2 weeks to optimize this value, the proposed adaptive DSE strategy features a 20% higher performance value after a total exploration time of about 4 days.
Tobias Schwarzer, Joachim Falk, Martín Letras, Christian Heidorn, Stefan Wildermann, Jürgen Teich
ACM Trans. Design Autom. Electr. Syst.2
2018 Reinforcement Learning for Power-Efficient Grant Prediction in LTE
abstract
Reducing the energy consumption of mobile phones is a major concern in the design of cellular modem solutions for LTE and 5G standards. Apart from optimizing hardware for power efficiency, dynamic power management, i.e., powering down idle system components, is a crucial means to achieve this goal. The techniques proposed so far, however, are reactive rather than proactive. This leads to the inability to exploit a significant amount of opportunities to power down components, as the opportunity is recognized too late. We propose a dynamic power management technique that is capable of exploiting said opportunities through the application of reinforcement learning prediction techniques for proactive power management. However, the additional computational effort for prediction algorithms must be carefully analyzed and taken into account. Therefore, we investigate which conditions have to be met in order to achieve net energy savings. The proposed technique has been implemented and evaluated for potential savings on simulated traces of LTE data. The resulting predictor is designed to be trained online, without any prior system knowledge. For a fair evaluation and comparison, the power consumption of the training phase is also considered in the analysis. It is shown that energy savings of up to 23.9 % may be obtained on a modem for scenarios such as HTTP streaming.
Peter Brand, Joachim Falk, Jonathan Ah Sue, Johannes Brendel, Ralph Hasholzner, Jürgen Teich
SCOPES2
2018 A predictive dynamic power management for LTE-Advanced mobile devices
abstract
Power consumption is a key challenge for LTE-Advanced or future 5G mobile devices and current power management systems successfully achieve significant power savings. However, these systems are driven by static rules and provide a posteriori responses to traffic and context changes. In this paper, we propose a smart dynamic power management system for cellular modems, extending existing power saving mechanisms by using machine learning-based traffic prediction. With the a priori knowledge of specific scheduling messages, internal device parameters can be finely tuned to improve the modem power consumption. In order to accurately estimate the power saving potential of several LTE use cases, we build a relevant data set of live network modem traces, as well as a power model of the baseband physical layer and radio frequency components. Subsequently, we propose an evaluation methodology and apply it to analyze the predictive power management performance in terms of error rate and global power consumption outcome. Our analysis results in maximal power savings of 12% for meaningful traffic scenarios as well as the identification of variables of interest to improve the proposed power manager.
Jonathan Ah Sue, Peter Brand, Johannes Brendel, Ralph Hasholzner, Joachim Falk, Jürgen Teich
WCNC5
2017 Exploiting Predictability in Dynamic Network Communication for Power-Efficient Data Transmission in LTE Radio Systems
abstract
In embedded systems powered by batteries, power is undoubtedly a critical resource making power management an important topic in the design phase. Even though power management is a heavily researched topic, most approaches focus on improving the way the power manager reacts to outside control events. In this paper, we propose techniques that not only react but rather try to predict these outside control events in advance, thus, broadening the capabilities of any employed power manager by allowing for superior transition decisions and even saving redundant calculations. We present results on employing a predictive power management system that couples a classic dynamic power manager with a machine learning subsystem in the context of a mobile device in a Long Term Evolution (LTE) system, with emphasis on evaluating the potential of saving power as well as the handling of the induced prediction uncertainty. First, we examine the LTE communication protocol and showcase certain control data that has to be received periodically, but may contain no information for the receiver. Finally, we show a proof-of-concept based on real LTE traces and hardware simulation, that prediction of this information can be leveraged to allow for a far superior decision process compared to a non-predicting system. Here, we achieve a theoretical best case power saving of 15 % for an idealized prediction with 100 % accuracy and no additional power consumption.
Peter Brand, Jonathan Ah Sue, Johannes Brendel, Joachim Falk, Ralph Hasholzner, Jürgen Teich, Stefan Wildermann
SCOPES4
2017 Automatic Conversion of Simulink Models to SysteMoC Actor Networks
abstract
Simulink has gained a lot of acceptance due to its intuitive through block-based algorithm design, simulation, and rapid prototyping capabilities for signal processing as well as control applications. However, automatic code generation for heterogeneous architectures is currently not supported by Simulink. In the literature, there exist automatic translation toolchains for generation of C or C++ code from Simulink models, which then are used for implementation or validation purposes. But few of them approach the generation of models that can be used in well-established Electronic System Level (ESL) design methodologies and tools. In order to address this issue, we present a methodology to extract an executable specification based on Data Flow Graphs (DFGs) from a given Simulink model. Such a specification can then be used by ESL tools to perform a Design Space Exploration (DSE) and generate code for hardware/software partitions directly from the ESL model. In a case study from signal processing, we validate the equivalence of the results of the simulation in Simulink and the results obtained by simulation of the DFG fully automatically generated from the Simulink model in the SystemC-based actor language SysteMoC.
Martín Letras, Joachim Falk, Stefan Wildermann, Jürgen Teich
SCOPES2
2015 Throughput-optimizing Compilation of Dataflow Applications for Multi-Cores using Quasi-Static Scheduling
abstract
Application modeling using dynamic dataflow graphs is well-suited for multi-core platforms. However, there is often a mismatch between the fine granularity of the application and the platform. Tailoring this granularity to the platform promises performance gains by (a) reducing dynamic scheduling overhead and (b) exploiting compiler optimizations. In this paper, we propose a throughput-optimizing compilation approach that uses Quasi-Static Schedules (QSSs) to combine actors of static dataflow subgraphs. Our proposed approach combines core allocation, QSSs, and actor binding in a Design Space Exploration (DSE), optimizing the throughput for a number of available cores. During the DSE, each implementation candidate is compiled to and evaluated on the target hardware---here an Intel i7 and an ARM Cortex-A9. Experimental results including synthetic benchmarks as well as a real-world control application show that our proposed holistic compilation approach outperforms classic DSEs that are agnostic of QSS as well as a DSE that employs QSS as a post-processing step. Amongst others, we show a case where the compilation approach obtains a speedup of 9.91 x for a 4-core implementation, while a classic DSE only obtains a speedup of 2.12 x.
Tobias Schwarzer, Joachim Falk, Michael Glaß, Jürgen Teich, Christian Zebelein, Christian Haubelt
SCOPES2
2014 Model-based actor multiplexing with application to complex communication protocols
abstract
We propose a dynamic scheduling approach for the concurrent execution of logical actor instances on a single synthesized actor instance. Based on a formal dataflow model of computation, the proposed approach can be applied to a wide range of applications in a model-based design flow. As case-study, we evaluate a bus-cycle-accurate SystemC RTL model based on an InfiniBand network adapter in a PCI Express system.
Christian Zebelein, Christian Haubelt, Joachim Falk, Tobias Schwarzer, Jürgen Teich
DATE3
2014 Communication-Driven Automatic Virtual Prototyping for Networked Embedded Systems
abstract
Today, parts of an ESL model can be automatically synthesized to a low-level implementation, e. g., via high-level synthesis. However, to build a complete working virtual prototype directly from a given ESL model, one still has to perform several design steps manually. The work-at-hand tackles this problem by introducing bridge components already in the ESL model. These components influence Design Space Exploration (DSE) by adding their characteristics like cost and latency into evaluation. The complete system is divided into several subsystems connected through bridges, we call this process communication-driven decomposition. Once an optimized implementation solution is found by DSE and selected by the designer, every subsystem of this ESL model is handed over to the individual synthesis tool. Here, if two subsystems will be synthesized by different tools, the bridge connecting these two subsystems will be automatically duplicated into two instances and assigned to each subsystem. Then, synthesis tools generate code for each subsystem (including the bridge inside each subsystem). In the last step, the system integration process merges the corresponding bridge pairs together to build a complete virtual prototype. To automate the proposed design flow, we have developed a framework that automatically divides an ESL model into subsystems and synthesizes the interfaces for all bridges which strongly simplifies system integration. The designer is therefore free from the interface realization. Hence, the overall design development cycle is shortened. As a proof of concept, a distributed control application is presented to give evidence of the proposed technique's applicability and the achieved productivity gain.
Liyuan Zhang 0001, Joachim Falk, Tobias Schwarzer, Michael Glaß, Jürgen Teich
DSD2
2013 Representing mapping and scheduling decisions within dataflow graphs
Christian Zebelein, Christian Haubelt, Joachim Falk, Tobias Schwarzer, Jürgen Teich
FDL3
2013 A rule-based quasi-static scheduling approach for static islands in dynamic dataflow graphs
abstract
In this article, an efficient rule-based clustering algorithm for static dataflow subgraphs in a dynamic dataflow graph is presented. The clustered static dataflow actors are quasi-statically scheduled , in such a way that the global performance in terms of latency and throughput is improved compared to a dynamically scheduled execution, while avoiding the introduction of deadlocks as generated by naive static scheduling approaches. The presented clustering algorithm outperforms previously published approaches by a faster computation and more compact representation of the derived quasi-static schedule. This is achieved by a rule-based approach, which avoids an explicit enumeration of the state space. A formal proof of the correctness of the presented clustering approach is given. Experimental results show significant improvements in both, performance and code size, compared to a state-of-the-art clustering algorithm.
Joachim Falk, Christian Zebelein, Christian Haubelt, Jürgen Teich
ACM Trans. Embed. Comput. Syst.1
2011 A rule-based static dataflow clustering algorithm for efficient embedded software synthesis
abstract
In this paper, an efficient embedded software synthesis approach based on a generalized clustering algorithm for static dataflow subgraphs embedded in general dataflow graphs is proposed. The clustered subgraph is quasi-statically scheduled, thus improving performance of the synthesized software in terms of latency and throughput compared to a dynamically scheduled execution. The proposed clustering algorithm outperforms previous approaches by a faster computation and a more compact representation of the derived quasi-static schedules. This is achieved by a rule-based approach, which avoids an explicit enumeration of the state space. Experimental results show significant improvements in both performance and code size when compared to a state-of-the-art clustering algorithm.
Joachim Falk, Christian Zebelein, Christian Haubelt, Jürgen Teich
DATE1
2010 Efficient High-Level modeling in the networking domain
abstract
Starting Electronic System Level (ESL) design flows with executable High-Level Models (HLMs) has the potential to sustainability improve productivity. However, writing good HLMs for complex systems is still a challenging task. In the context of network controller design, modeling complexity has two major sources: (1) the functionality to handle a single connection, and (2) the number of connections to be handled in parallel. In this paper, we will propose an efficient actor-oriented modeling approach for complex systems by (1) integrating hierarchical FSMs into dynamic dataflow models, and (2) providing new channel types to allow concurrent processing of multiple connections. We will show the applicability of our proposed modeling approach to real-world system designs by presenting results from modeling and simulating a network controller for the Parallel Sysplex architecture used in IBM System z mainframes.
Christian Zebelein, Joachim Falk, Christian Haubelt, Jürgen Teich, Rainer Dorsch
DATE2
2010 Analysis of SystemC actor networks for efficient synthesis
abstract
Applications in the signal processing domain are often modeled by dataflow graphs. Due to heterogeneous complexity requirements, these graphs contain both dynamic and static dataflow actors. In previous work, we presented a generalized clustering approach for these heterogeneous dataflow graphs in the presence of unbounded buffers. This clustering approach allows the application of static scheduling methodologies for static parts of an application during embedded software generation for multiprocessor systems. It systematically exploits the predictability and efficiency of the static dataflow model to obtain latency and throughput improvements. In this article, we present a generalization of this clustering technique to dataflow graphs with bounded buffers, therefore enabling synthesis for embedded systems without dynamic memory allocation. Furthermore, a case study is given to demonstrate the performance benefits of the approach.
Joachim Falk, Christian Zebelein, Joachim Keinert, Christian Haubelt, Jürgen Teich, Shuvra S. Bhattacharyya
ACM Trans. Embed. Comput. Syst.1
2009 SystemCoDesigner - an automatic ESL synthesis approach by design space exploration and behavioral synthesis for streaming applications
abstract
With increasing design complexity, the gap from ESL (Electronic System Level) design to RTL synthesis becomes more and more crucial to many industrial projects. Although several behavioral synthesis tools exist to automatically generate synthesizable RTL code from C/C++/SystemC-based input descriptions and software generation for embedded processors is automated as well, an efficient ESL synthesis methodology combining both is still missing. This article presents SystemCoDesigner, a novel SystemC-based ESL tool to automatically optimize a hardware/software SoC (System on Chip) implementation with respect to several objectives. Starting from a SystemC behavioral model, SystemCoDesigner automatically extracts the mathematical model, performs a behavioral synthesis step, and explores the multiobjective design space using state-of-the-art multiobjective optimization algorithms. During design space exploration, a single design point is evaluated by simulating highly accurate performance models, which are automatically generated from the SystemC behavioral model and the behavioral synthesis results. Moreover, SystemCoDesigner permits the automatic generation of bit streams for FPGA targets from any previously optimized SoC implementation. Thus SystemCoDesigner is the first fully automated ESL synthesis tool providing a correct-by-construction generation of hardware/software SoC implementations. As a case study, a model of a Motion-JPEG decoder was automatically optimized and implemented using SystemCoDesigner. Several synthesized SoC variants based on this model show different tradeoffs between required hardware costs and achieved system throughput, ranging from software-only solutions to pure hardware implementations that reach real-time performance for QCIF streams on a 50MHz FPGA.
Joachim Keinert, Martin Streubühr, Thomas Schlichter, Joachim Falk, Jens Gladigau, Christian Haubelt, Jürgen Teich, Michael Meredith
ACM Trans. Design Autom. Electr. Syst.4
2008 A generalized static data flow clustering algorithm for mpsoc scheduling of multimedia applications
abstract
Abstract—In this paper, an efficient embedded software synthesis approach based on a generalized clustering algorithm for static dataflow subgraphs embedded in general dataflow graphs is proposed. The clustered subgraph is quasi-statically scheduled, thus improving performance of the synthesized software in terms of latency and throughput compared to a dynamically scheduled execution. The proposed clustering algorithm outperforms previous approaches by a faster computation and a more compact representation of the derived quasi-static schedules. This is achieved by a rule-based approach, which avoids an explicit enumeration of the state space. Experimental results show significant improvements in both performance and code size when compared to a state-of-the-art clustering algorithm.
Joachim Falk, Joachim Keinert, Christian Haubelt, Jürgen Teich, Shuvra S. Bhattacharyya
EMSOFT1
2008 Classification of General Data Flow Actors into Known Models of Computation
abstract
Applications in the signal processing domain are often modeled by data flow graphs which contain both dynamic and static data flow actors due to heterogeneous complexity requirements. Thus, the adopted notation to model the actors must be expressive enough to accommodate dynamic data flow actors. On the other hand, treating static data flow actors like dynamic ones hinders design tools in applying domain-specific optimization methods to static parts of the model, e.g., static scheduling. In this paper, we present a general notation and a methodology to classify an actor expressed by means of this notation into the synchronous and cyclo-static dataflow models of computation. This enables the use of a unified descriptive language to express the behavior of actors while still retaining the advantage to apply domain-specific optimization methods to parts of the system. In experiments we could improve both latency and throughput of a general data flow graph application using our proposed automatic classification in combination with a static single-processor scheduling approach by 57%.
Christian Zebelein, Joachim Falk, Christian Haubelt, Jürgen Teich
MEMOCODE2
2006 Task-accurate performance modeling in SystemC for real-time multi-processor architectures
abstract
We propose a framework, called virtual processing components (VPC) that permits the modeling and simulation of multiple processors running arbitrary scheduling strategies in SystemC. The granularity is given by task accuracy that guarantees a small simulation overhead
Martin Streubühr, Joachim Falk, Christian Haubelt, Jürgen Teich, Rainer Dorsch, Thomas Schlipf
DATE2
2006 Efficient Representation and Simulation of Model-Based Designs
Joachim Falk, Christian Haubelt, Jürgen Teich
FDL1