Gunar Schirner

dblp:88/4231 · DBLP profile ↗
← Back
50ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0002-5408-8496ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 45 · 9 first-author · 11 since 2021Software engineering, systems software and programming languages · 12 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1
YearPublicationVenuePosition
2026 CommILP: Synthesizing Communication Infrastructure for Domain Computing Platforms
abstract
Domain-specific accelerator-rich platforms are promising for high-throughput, low-power edge systems. Prior domain design-space exploration (Domain-DSE) methods [1] efficiently allocate processing elements (PEs) and bind applications, but largely assume simplified, centralized communication within the hardware tile, which limits scalability and wastes resources. Novel research work is needed that takes advantage of designtime known communication patterns to synthesize custom intratile interconnects that minimize area and energy while preserving application throughput. This paper introduces CommILP for synthesizing intra-tile communication infrastructure given a platform allocation and PE-to-PE communication graphs for multiple applications. CommILP instantiates Communication Elements (CEs), determines a minimal interconnect topology, and assigns routing paths for each app and connection. It guarantees that application throughput is preserved while minimizing area and energy. CommILP introduces a hierarchical ILP formulation across platform, application, and hop-stack levels, and supports both App-Level and Domain-Aggregated modeling modes to balance solution quality and runtime. Evaluated on ASIC and FPGA backends using both real (OpenVX-40) and synthetic (RNDComm-100) domains, CommILP achieves up to 69.4% area and 73% energy savings over centralized baselines, while remaining scalable to over 100 applications. It complements existing platform DSE tools by bridging the gap between computation binding and hardware-efficient communication synthesis.
Qucheng Jiang, Jacob Ginesin, Oscar Kellner, Tianrui Ma, Gunar Schirner
ASP-DAC5
2025 Enabling ILP-Based DSE for Multigranularity, Unified Domain Platforms With DmTSAR-ILP
abstract
Domain-specific HWACC-rich platforms blend high performance and efficiency presenting an opportunity to recover nonrecurring engineering costs through wider deployment for many applications. However, the design of such platforms is immensely challenging, partly due to the design space size. MG-DmDSE (Zhang et al., 2023) allocates domain platforms catering to multiple applications. To cope with complexity, it employs a genetic algorithm and heuristics. While this is a typical approach in design space exploration (DSE), it cannot guarantee optimality. Exact solutions could be obtained with integer linear programming (ILP). Yet, the multiapplication DSE presents aggregation challenges as applications might have widely different performance, leading to nonlinear aggregation techniques in previous work that ILPs cannot employ. This work introduces DmTSAR-ILP, an ILP-based platform allocation method that simultaneously considers all applications in a domain. DmTSAR-ILP captures the same underlying analytical model as MG-DmDSE and can explore its domain application set. To enable our linear domain-level formulation, we introduce the average performance achievement metric (APA), offering an ILP-friendly, fair aggregation across domain applications. To highlight benefits of DmTSAR-ILP, a domain platform for 40 OpenVX apps is generated (across varying area budgets). For all area budgets, DmTSAR-ILP platforms perform better or identical to MG-DmDSE even in their original metric. The proposed APA aggregation is fair, allowing more applications (55% vs 37.5% of applications in MG-DmDSE) to be fully accelerated on the generated platform. For the$0.1~\mathbf {mm}^{2}$area budget, DmTSAR-ILP’s platform increases application throughput by 22.5% over MG-DmDSE’s platform. DmTSAR-ILP is 70x faster on average than the heuristic-based MG-DmDSE by leveraging an ILP formulation that considers all applications simultaneously and recent advances in mixed-integer programming solver performance.
Bruno Morais, Qucheng Jiang, Gunar Schirner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 Grasp-HGN: Grasping the Unexpected
abstract
For transradial amputees, robotic prosthetic hands promise to regain the capability to perform daily living activities. To advance next-generation prosthetic hand control design, it is crucial to address current shortcomings in robustness to out of lab artifacts, and generalizability to new environments. Due to the fixed number of object to interact with in existing datasets, contrasted with the virtually infinite variety of objects encountered in the real world, current grasp models perform poorly on unseen objects, negatively affecting users’ independence and quality of life. To address this: (i) we define semantic projection, the ability of a model to generalize to unseen object types and show that conventional models like YOLO, despite 80% training accuracy, drop to 15% on unseen objects. (ii) We propose Grasp-LLaVA, a Grasp Vision Language Model enabling human-like reasoning to infer the suitable grasp type estimate based on the object’s physical characteristics resulting in a significant 50.2% accuracy over unseen object types compared to 36.7% accuracy of an SOTA grasp estimation model. Lastly, to bridge the performance-latency gap, we propose Hybrid Grasp Network (HGN), an edge-cloud deployment infrastructure enabling fast grasp estimation on edge and accurate cloud inference as a fail-safe, effectively expanding the latency vs. accuracy Pareto. HGN with confidence calibration (DC) enables dynamic switching between edge and cloud models, improving semantic projection accuracy by 5.6% (to 42.3%) with 3.5× speedup over the unseen object types. Over a real-world sample mix, it reaches 86% average accuracy (12.2% gain over edge-only), and 2.2× faster inference than Grasp-LLaVA alone.
Mehrshad Zandigohar, Mallesham Dasari, Gunar Schirner
ACM Trans. Embed. Comput. Syst.3
2024 Loco-Manipulation with Nonimpulsive Contact-Implicit Planning in a Slithering Robot
abstract
Object manipulation has been extensively studied in the context of fixed base and mobile manipulators. However, the overactuated locomotion modality employed by snake robots allows for a unique blend of object manipulation through locomotion, referred to as loco-manipulation. The following work presents an optimization approach to solving the loco-manipulation problem based on non-impulsive implicit contact path planning for our snake robot COBRA. We present the mathematical framework and show high-fidelity simulation results and experiments to demonstrate the effectiveness of our approach.
Adarsh Salagame, Kruthika Gangaraju, Harin Kumar Nallaguntla, Eric Sihite, Gunar Schirner, Alireza Ramezani
IROS5
2024 Heading Control for Obstacle Avoidance using Dynamic Posture Manipulation during Tumbling Locomotion
abstract
Passive tumbling structures are energy efficient, but often sacrifice control authority due to their under actuated nature. Unlike many passive tumbling robots, Northeastern University’s COBRA is a snake robot with eleven articulated joints that transforms into a wheel-like structure with a high degree of posture control during tumbling, and using this posture manipulation, COBRA can control its forward velocity and heading angle while tumbling. This paper presents a mathematical framework that describes the dynamics of posture manipulation during tumbling and identifies two types of control actions that allow it to control its movement. This is validated in hardware testing to demonstrate obstacle avoidance during passive tumbling using only posture manipulation.
Adarsh Salagame, Kruthika Gangaraju, Eric Sihite, Gunar Schirner, Alireza Ramezani
IROS4
2023 DmTSAR-ILP: Allocating a Unified Domain Platform for Streaming Applications
abstract
Domain-specific HWACC-rich platforms blend high performance & efficiency presenting an opportunity to recover non-recurring engineering costs through wider deployment for many applications. However, the design of such platforms is immensely challenging, partly due to the design space size.MG-DmDSE [1] allocates domain platforms catering to multiple applications. To cope with complexity, it employs a genetic algorithm and heuristics. While this is a typical approach in DSE, it cannot guarantee optimality. Exact solutions could be obtained with integer linear programming (ILP). Yet, the multi-application DSE presents aggregation challenges as applications might have widely different performance, leading to non-linear aggregation techniques in previous work that ILPs cannot employ.This work introduces DmTSAR-ILP, an ILP-based platform allocation method that simultaneously considers all applications in a domain. DmTSAR-ILP captures the same underlying analytical model as MG-DmDSE and can explore its domain application set. To enable our linear domain-level formulation, we introduce the average performance achievement metric (APA), offering an ILP-friendly, fair aggregation across domain applications.To highlight benefits of DmTSAR-ILP, a domain platform for 40 OpenVX apps is generated (across varying area budgets). For all area budgets, DmTSAR-ILP platforms perform better or identical to MG-DmDSE even in their original metric. The proposed APA aggregation is fair, allowing more applications (55% vs 37.5% of applications in MG-DmDSE) to be fully accelerated on the generated platform. For the 0.1 mm 2 area budget, DmTSAR-ILP’s platform increases application throughput by 22.5% over MG-DmDSE’s platform. DmTSAR-ILP is 70x faster on average than the heuristic-based MG-DmDSE by leveraging an ILP formulation that considers all applications simultaneously and recent advances in MIP solver performance.
Bruno Morais, Gunar Schirner
DAC2
2023 TSAR-ILP: Tile-Based, Synchronization-AwaRe ILP Allocating Heterogeneous Platforms for Streaming Applications
abstract
Automatic design space exploration (DSE) is key in hardware-software (HW/SW) co-design. To cope with the large design space, explorations are often heuristic-based and/or approximate yielding potentially locally optimal solutions. Without knowing the globally optimal solution, strong assertions about performance upper/lower bounds cannot be made. In contrast, integer linear programming (ILP) formulations can produce exact (optimal) solutions. Previous ILP-based formulations, however, lack support for tile-based architectures and realistic synchronization models, limiting their DSE capabilities. This work introduces a tile-based, synchronization-aware ILP (TSAR-ILP) formulation that overcomes previous limitations. With TSAR-ILP, the allocation/binding problems are introduced and formalized, attaining optimal solutions for mapping streaming applications onto template platforms. Using TSAR-ILP, this work explores a hardware accelerator-rich (HWACC-rich) platform with direct HWACC-to-HWACC communication under HW area constraints for 40 OpenVX applications. To illustrate design opportunities given by: 1) the ILP formulation and 2) direct HWACC-to-HWACC communication, this article analyzes the impact of job size. Results show that selecting smaller job sizes yields performance improvements and less area usage at the cost of slightly increased synchronization overhead. A job size reduction from 1 kB to 256 bytes gives$3.51\times $average performance increase across 40 applications. Finally, DSE with TSAR-ILP is shown not to be prohibitive through scalability analysis using a set of 5000 synthetic applications with varying size (10–125 nodes), with 94.3% of applications successfully achieving optimal solutions under 60 s.
Bruno Morais, Jinghan Zhang 0001, Gunar Schirner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Generating Unified Platforms Using Multigranularity Domain DSE (MG-DmDSE) Exploiting Application Similarities
abstract
Heterogeneous accelerator-rich (ACC-rich) platforms combining general-purpose cores and specialized HW accelerators (ACCs) promise high-performance and low-power streaming application deployments in a variety of domains, such as video analytics and software-defined radio. In order to benefit a domain of applications, a domain platform exploration tool must take advantage of structural and functional similarities across applications by allocating a common set of ACCs. A previous approach proposed a genetic domain exploration tool (GIDE) that applied a restrictive binding algorithm that mapped applications functions to monolithic accelerators. This approach suffered from a low average application throughput across and reduced platform generality. This article introduces a multigranularity-based domain design space exploration tool (MG-DmDSE) to improve both average application throughput as well as platform generality. The key contributions of MG-DmDSE are: 1) applying a multigranular decomposition of coarse-grained application functions into more granular compute kernels; 2) examining compute similarity between functions in order to provide more generic functions; 3) configuring monolithic ACCs by selectively bypassing compute elements within them during DSE to expose more functionality; and 4) speeding up MG-DmDSE platform allocation exploration through a greedy guided mutation (GGM) algorithm. To assess MG-DmDSE, both GIDE and MG-DmDSE were applied to applications in the OpenVX library. MG-DmDSE achieves an average$2.84\times $greater application throughput compared to GIDE. Additionally, 87.5% of applications benefited from running on the platform produced by MG-DmDSE versus 50% from GIDE, which indicated increased platform generality. The generated MG-DmDSE platforms achieve an average of 61.8% logarithmic throughput improvement for unknown applications over GIDE. GGM results in saving 84.8% of the exploration time in MG-DmDSE with only 0.23% performance loss.
Jinghan Zhang 0001, Aly Sultan, Mehrshad Zandigohar, Gunar Schirner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 NetCut: Real-Time DNN Inference Using Layer Removal
abstract
Deep Learning plays a significant role in assisting humans in many aspects of their lives. As these networks tend to get deeper over time, they extract more features to increase accuracy at the cost of additional inference latency. This accuracy-performance trade-off makes it more challenging for Embedded Systems, as resource-constrained processors with strict deadlines, to deploy them efficiently. This can lead to selection of networks that can prematurely meet a specified deadline with excess slack time that could have potentially contributed to increased accuracy. In this work, we propose: (i) the concept of layer removal as a means of constructing TRimmed Networks (TRNs) that are based on removing problem-specific features of a pretrained network used in transfer learning, and (ii) NetCut, a methodology based on an empirical or an analytical latency estimator, which only proposes and retrains TRNs that can meet the application's deadline, hence reducing the exploration time significantly. We demonstrate that TRNs can expand the Pareto frontier that trades off latency and accuracy to provide networks that can meet arbitrary deadlines with potential accuracy improvement over off-the-shelf networks. Our experimental results show that such utilization of TRNs, while transferring to a simpler dataset, in combination with NetCut, can lead to the proposal of networks that can achieve relative accuracy improvement of up to 10.43% among existing off-the-shelf neural architectures while meeting a specific deadline, and 27x speedup in exploration time.
Mehrshad Zandigohar, Deniz Erdogmus, Gunar Schirner
DATE3
2021 MG-DmDSE: Multi-Granularity Domain Design Space Exploration Considering Function Similarity
abstract
Heterogeneous accelerator-rich (ACC-rich) platforms combining general-purpose cores and specialized HW accelerators (ACCs) promise high-performance and low-power streaming application deployments in a variety of domains such as video analytics and software-defined radio. In order to benefit a domain of applications, a domain platform exploration tool must take advantage of structural and functional similarities across applications by allocating a common set of ACCs. A previous approach [1] proposed a GenetIc Domain Exploration tool (GIDE) that applied a restrictive binding algorithm that mapped applications functions to monolithic accelerators. This approach suffered from lower average application throughput across and reduced platform generality. This paper introduces a Multi-Granularity based Domain Design Space Exploration tool (MG-DmDSE) to improve both average application throughput as well as platform generality. The key contributions of MG-DmDSE are: (1) Applying a multi-granular decomposition of coarse grain application functions into more granular compute kernels. (2) Examining compute similarity between functions in order to produce more generic functions. (3) Configuring monolithic ACCs by selectively bypassing compute elements within them during DSE to expose more functionality. To assess MG-DmDSE, both GIDE and MG-DmDSE were applied to applications in the OpenVX library. MG-DmDSE achieves an average 2.84x greater application throughput compared to GIDE. Additionally, 87.5% of applications benefited from running on the platform produced by MG-DmDSE vs 50% from GIDE, which indicated increase platform generality.
Jinghan Zhang 0001, Aly Sultan, Hamed Tabkhi, Gunar Schirner
DATE4
2021 RDP3: Rapid Domain Platform Performance Prediction for Design Space Exploration
abstract
Heterogeneous Accelerator-rich (ACC-rich) platforms combining general-purpose cores and specialized HW Accelerators (ACCs) promise high-performance and low-power deployment of streaming applications, e.g. for video analytics, software-defined radio, and radar. In order to recover Non-Recurring Engineering (NRE) cost, a unified domain platform for a set of applications can be exploited, especially when applications have functional and structural similarities, which can benefit from common ACCs. However, identifying the most beneficial set of common ACCs is challenging, and current Design Space Exploration (DSE) methods for domain platform allocation suffer from a long exploration time bottleneck. In particular, compared to a traditional DSE, evaluating the performance of a platform for a domain of applications is much more time-consuming as binding exploration and evaluation for each application in the domain is required. Thus, a rapid domain performance evaluation is needed to speed up the exploration of the platform allocation.This paper introduces Rapid Domain Platform Performance Prediction (RDP3) methods to speed up the exploration in domain DSE. Key contributions are: (1) analyzing current domain DSE flow and its exploration time bottleneck; (2) introducing four RDP3methods to speedup the evaluation of different platform allocations: Heuristic Processing (HP) estimation, Linear Regression (LR), Decision Tree Regression (DTR), and Multi-Layer Perceptron (MLP) predictions; (3) comparing the performance of these predictions and integrating the prediction into the current domain DSE. To evaluate the efficacy of RDP3, we explore 10K platforms capable of processing OpenVX domain applications. We demonstrate that RDP3-MLP as the most promising method can achieve a speedup of 17.5K times with only 0.001 mean square error compared to the current platform evaluation using the analytical model. Integrating RDP3-MLP into the existing domain DSE method GIDE [1] can save 80.8% exploration time while still resulting in the same output platform design.
Jinghan Zhang 0001, Mehrshad Zandigohar, Gunar Schirner
ICCD3
2020 Allocating One Common ACC-Rich Platform for Many Streaming Applications
abstract
Many demanding streaming applications share functional and structural similarities with other apps in their respective domain, e.g., video analytics, software-defined radio, and radar. This opens the opportunity for specialization (e.g., heterogeneous computing) to achieve the needed efficiency and/or performance. However, current design space exploration (DSE) focuses on an individual application in isolation (e.g., one particular vision flow), but not a set of similar applications. Hence, optimizations that occur due to considering multiple applications simultaneously are missed. New DSE methodologies and tools are needed with a broader scope of application sets instead of individual applications. This article introduces a novel domain-specific DSE (DS-DSE) approach focusing on streaming applications. Key contributions are: 1) a formalized method to extract the functional and structural similarities of domain applications; 2) a rapid platform performance estimation and comparison at two abstraction levels: domain score (DS) and analytic performance estimation (APE) model; 3) two novel algorithms, dynamic score selection (DSS), and GenetIc domain exploration (GIDE), for hardware/software partitioning of a domain-specific platform to maximize the throughput across domain applications (under certain constraints); and 4) a methodology to evaluate a platform's benefit for a set of applications. We demonstrate DSS's and GIDE's benefits using OpenVX applications and synthetic domains. The DSS and GIDE generated domain-specific platforms improve performance over application-specific platforms by 58% and 75% for OpenVX, as well as by 23% and 48% for synthetic applications. GIDE's platforms reach 99.8% (OpenVX) and 97.6% (synthetic) throughput of the domain optimal platform obtained through exhaustive search.
Jinghan Zhang 0001, Hamed Tabkhi, Gunar Schirner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 Mitigating Application Diversity for Allocating a Unified ACC-Rich Platform
abstract
Heterogeneous accelerator-rich (ACC-rich) platforms combining general-purpose cores and specialized HW accelerators (ACCs) promise high-performance and low-power streaming application (app) deployments, e.g. for video analytics, software-defined radio, and radar. In order to recover NRE, a unified platform for a set of applications (apps) is desirable. When apps have functional and structural similarities, they can benefit from common ACCs. Identifying the most beneficial set of common ACCs is challenging. However, current allocation strategies mostly focus on one app in isolation. Automatically allocating a unified platform requires simultaneously considering many apps, an efficient design space traversal and a fair evaluation across diverse apps. This paper introduces a Unified ACC-rich Platform Allocation (UPA) methodology for sets of data flow apps. Key contributions are: (1) a genetic algorithm (GA) guided by a fair and efficient evaluation to allocate one unified platform for many apps, (2) defining relative efficiency for fair comparison across diverse apps, and (3) defining metrics to quantify many app platform efficiency. This paper demonstrates UPA's benefits using OpenVX apps. A 12-ACCs-UPA improves average efficiency 4.59x over app-dedicated platforms. The UPA platform enables more apps (55% of OpenVX apps) to be efficiently deployed (≥ 60% of optimal app-dedicated platform). The benefits increase even further with increasing ACC budget.
Jinghan Zhang 0001, Hamed Tabkhi, Gunar Schirner
ICCD3
2019 Alleviating Scalability Limitation of Accelerator-Based Platforms
abstract
Accelerator-based chip multiprocessors (ACMPs), which combine application-specific HW accelerators (ACCs) with host processor core(s), are promising architectures for high-performance and power-efficient computing. However, ACMPs with many ACCs have scalability limitations. The ACCs' performance benefits can be overshadowed by bottlenecks on shared resources of processor core(s), communication fabric/DMA, and on-chip memory. Primarily, this is rooted in the ACCs' data access and the orchestration dependency. Due to very loosely defined ACC communication semantics, and relying on general architectures, the resources bottlenecks hamper performance. This paper explores and alleviates the scalability limitations of ACMPs. To this end, this paper first proposes ACMPerf, an analytical model to capture the impact of the resources bottlenecks on the achievable ACCs' benefits. Then, this paper identifies and formalizes ACC communication semantics which paves the path toward a more scalable integration of ACCs. The semantics describe four primary aspects: 1) data access; 2) data granularity; 3) data marshalling; and 4) synchronization. Finally, this paper proposes a novel architecture of transparent self-synchronizing accelerators (TSS). TSS efficiently realizes our identified communication semantics of direct ACC-to-ACC connections often occurring in streaming applications. TSS delivers more of the ACCs' benefits than conventional ACMP architectures. Given the same set of ACCs, TSS has up to 130× higher throughput and 78× lower energy consumption, mainly due to reducing the load on shared architectural resources by 78.3×.
Nasibeh Teimouri, Hamed Tabkhi, Gunar Schirner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 DS-DSE: Domain-specific design space exploration for streaming applications
abstract
Domain-specific computing is promising for high-performance low-power execution of applications with similar functionality. In particular, streaming applications with significant functional and structural similarities can tremendously benefit. However, current Design Space Exploration (DSE) focuses on individual applications in isolation. Hence, much of the domain optimization opportunities are missed. DSE methodologies need to broaden the scope from individual applications in isolation to optimizing across applications within a domain. This paper introduces a novel Domain-Specific DSE (DS-DSE) approach for domain-specific computing with a focus on streaming applications. Key contributions are: (1) a formalized method to extract the functional and structural similarities of domain applications, (2) a novel algorithm for hardware/software partitioning of a domain-specific platform to maximize the throughput across domain applications (under certain constraints) and (3) a methodology to evaluate a domain platform. This paper demonstrates the benefits using 4 domains: OpenVX (vision processing), and 3 synthetic domains (with greater complexity). Our experiments demonstrate a performance improvement (average throughput) of 36.8% for OpenVX and 46.2% for synthetic domains of the DS-DSE generated platform compared to an application-specific platform.
Jinghan Zhang 0001, Hamed Tabkhi, Gunar Schirner
DATE3
2016 Framework for Rapid Development of Embedded Human-in-the-Loop Cyber-Physical Systems
abstract
Human-in-the-Loop Cyber-Physical Systems (HiLCPS) offers assistive technology that augments human interaction with the physical world, such as self-feeding, communication and mobility for functionally locked-in individuals. HiLCPS applications are typically implemented as networked embedded systems interfacing both human and the physical environment. Developing HiLCPS applications is challenging due to interfacing with hardware with different specifications and physical location (local/remote). Also, while algorithm designers prototype applications in MATLAB benefiting from an algorithm design environment, the gap from prototyping MATLAB application to embedded solution traditionally requires significant manual implementation. In this paper, we propose a HiLCPS Framework for the rapid development of embedded HiLCPS applications. The framework groups similar hardware types to classes, unifying their access and with this offering both hardware and location transparent access. The framework furthermore incorporates a domain-specific synthesis tool, called Hsyn. Hsyn empowers algorithm designers to prototype a portable, hardware-agnostic application in MATLAB while offering an automatic path to embedded deployment without requiring embedded knowledge. We demonstrate the benefit of the framework with a brain-controlled wheelchair application prototyped in MATLAB that transparently accesses a variety of EEG acquisition systems with local or remote connections. Then, by using Hsyn, the application is automatically deployed to a BeagleBone Black equipped with a custom-designed electrophysiological acquisition cape. Hsyn shows six orders of magnitude of productivity gain compared to manual embedded deployment. The wheelchair performs stepwise navigated based on human intent inference with 91% accuracy at 0.9 confidence threshold every 4 seconds on average over 9 users.
Shen Feng, Fernando Quivira, Gunar Schirner
BIBE3
2016 Improving scalability of CMPs with dense ACCs coverage
Nasibeh Teimouri, Hamed Tabkhi, Gunar Schirner
DATE3
2016 Guiding Power/Quality Exploration for Communication-Intense Stream Processing
abstract
In this paper, we explore the power/quality trade-off for streaming applications with a shift from the computation to the communication aspects of the design. The paper proposes a systematic exploration methodology to formulate and traverse power/quality trade-off for the class of adaptive streaming applications. The formalization enables to procedurally transition from a set of design requirements to architecture goals. The architecture goals can then be realized through design choices yielding system designs that meet the initial requirements. The reported results are based on an actual implementation of Mixture of Gaussian (MoG) background subtraction on Xilinx Zynq platform.
Hamed Tabkhi, Majid Sabbagh, Gunar Schirner
ACM Great Lakes Symposium on VLSI3
2016 Hardware thread reordering to boost OpenCL throughput on FPGAs
abstract
Availability of OpenCL for FPGAs has raised new questions about the efficiency of massive thread-level parallelism on FPGAs. The general trend is toward creating deep pipelining and in-order execution of many OpenCL threads across a shared data-path. While this can be a very effective approach for regular kernels, its efficiency significantly diminishes for irregular kernels with runtime-dependent control flow. We need to look for new approaches to improve execution efficiency of FPGAs when targeting irregular OpenCL kernels. This paper proposes a novel solution, called Hardware Thread Reordering (HTR), to boost the throughput of the FPGAs when executing irregular kernels possessing non-deterministic runtime control flow. The key insight of HRT is out-of-order OpenCL thread execution over a shared data-path to achieve significantly higher throughput. The thread reordering is performed at a basic-block level granularity. The synthesized basic-blocks are extended with independent pipeline control signals and context registers to bypass the live values of reordered threads. We demonstrate the efficiency of our proposed solution on three parallel irregular kernels. For the experiments, we utilize the LegUp tool to compare the baseline (in-order) data-path with HTR-enhanced data-path. Our RTL simulation results demonstrate that HTR-enhanced data-path achieves up to 11× increase in kernels throughput at a very low overhead (less than 2× increase in FPGA resources).
Amir Momeni, Hamed Tabkhi, Gunar Schirner, David R. Kaeli
ICCD3
2016 EEGu2: an embedded device for brain/body signal acquisition and processing
abstract
Brain/Body Computer Interface (BBCI) technology facilitates research in human cognition and assistive technologies. BBCI acquires and analyzes physiological signals from human body/brain such as electroencephalography (EEG) to observe human physiological states and potentially enable external control. BBCI devices require accurate data acquisition systems with sufficient dynamic range for various brain/body signals. Also, embedded processing is desirable for real-time interaction and flexible deployment. However, most off-the-shelf BBCI devices are very costly, e.g. g.USBamp at $15K and do not offer embedded processing. Hence, an open embedded device for BBCI acquisition and processing is needed to foster the BBCI research.
Shen Feng, Mian Tang, Fernando Quivira, Tim Dyson, Filip Cuckov, Gunar Schirner
RSP6
2016 Communication and cooling aware job allocation in data centers for communication-intensive workloads
Eduard Llamosí, Fulya Kaplan, Chulian Zhang, Jiayi Sheng, Martin C. Herbordt, Gunar Schirner, Ayse K. Coskun
J. Parallel Distributed Comput.7
2015 An efficient architecture solution for low-power real-time background subtraction
abstract
Embedded vision is a rapidly growing market with a host of challenging algorithms. Among vision algorithms, Mixture of Gaussian (MoG) background subtraction is a frequently used kernel involving massive computation and communication. Tremendous challenges need to be reolved to provide MoG's high computation and communication demands with minimal power consumption allowing its embedded deployment. This paper proposes a customized architecture for power-efficient realization of MoG background subtraction operating at Full-HD resolution. Our design process benefits from system-level design principles. An SLDL-captured specification (result of high-level explorations) serves as a specification for architecture realization and hand-crafted RTL design. To optimize the architecture, this paper employs a set of optimization techniques including parallelism extraction, algorithm tuning, operation width sizing and deep pipelining. The final MoG implementation consists of 77 pipeline stages operating at 148.5 MHz implemented on a Zynq-7000 SoC. Furthermore, our background subtraction solution is flexible allowing end users to adjust algorithm parameters according to scene complexity. Our results demonstrate a very high efficiency for both indoor and outdoor scenes with 145 mW on-chip power consumption and more than 600× speedup over software execution on ARM Cortex A9 core.
Hamed Tabkhi, Majid Sabbagh, Gunar Schirner
ASAP3
2015 Reducing Dynamic Dispatch Overhead (DDO) of SLDL-synthesized embedded software
abstract
System-Level Design Languages (SLDL) allow component-oriented specifications, e.g. for separating computation and communication. This separation allows for a flexible model composition, refinement and explorations. This flexibility, however, requires dynamic dispatch during execution that degrades the simulation performance. After synthesized to a target platform, the model re-composition is no longer required. Then, the involved Dynamic Dispatch Overhead (DDO) only limits performance without providing benefits. Thus, approaches are needed for software synthesis to analyze model connectivity and eliminate the DDO wherever possible. This paper introduces a static dispatch type analysis as part of the DDO-aware embedded C code synthesis from SLDL models. Our DDO-aware software (SW) synthesis emits faster, more readable static dispatch code whenever a static connectivity is determinable. By replacing virtual functions with direct function calls, the DDO can be totally eliminated allowing for aggressive inlining optimizations by the compiler. We demonstrate the benefits of the improved SW synthesis on a JPEG encoder, which runs up to 16% faster with DDO-reduction on an ARM9-based HW/SW platform. Our approach combines the flexibility benefits in specification modeling with efficient execution when synthesized to embedded targets.
Jiaxing Zhang 0003, Sanyuan Tang, Gunar Schirner
ASP-DAC3
2015 Revisiting accelerator-rich CMPs: challenges and solutions
abstract
Heterogeneous Chip Multiprocessors (CMP)s, which combine processor cores with specialized HW accelerators, are one main approach to high-performance low-power computing. While it is promising for few accelerators, the scalability is a major challenge with increasing number of accelerators. Resources including memory, communication fabric and processor turn into bottlenecks and result in accelerator under-utilization and cripple the performance.
Nasibeh Teimouri, Hamed Tabkhi, Gunar Schirner
DAC3
2015 Bridging Architecture and Programming for Throughput-Oriented Vision Processing (Abstract Only)
abstract
With the expansion of OpenCL support across many heterogeneous devices (including FPGAs, GPUs and CPUs), the programmability of these systems has been significantly increased. At the same time, new questions arise about which device should be targeted for each OpenCL software kernel. Once we select a device, then we are left to customize the application, selecting the right granularity of parallelism and frequency of host-to-device communication. In this paper, we study the impact of source-level decisions on the overall execution time when developing OpenCL program across different heterogeneous devices. We focus on two mainstream architecture classes (GPUs and FPGAs), and consider throughput-oriented advanced vision processing. To guide this exploration, we propose a new vertical classification for selecting the grain of parallelism for advanced vision processing applications. To carry out this study we have selected the Mean-shift object tracking algorithm as a representative candidate of advanced vision algorithms. Overall, our evaluation demonstrates that fine-grained parallelism can greatly benefit FPGA execution (up to a 4X speed-up), while a combination of coarse-grained and fine-grained parallelism achieves the best performance on a GPU (up to a 6X speed-up). Also, there can be a large benefit if we can execute both the parallel and serial parts of the program on a FPGA (up to a 21X speed-up).
Amir Momeni, Hamed Tabkhi, Gunar Schirner, David R. Kaeli
FPGA3
2015 Optimization of energy efficient relay position for galvanic coupled intra-body communication
abstract
Implanted medical sensors and actuators within the human body will enable remote data gathering, diagnosis, and the ability to directly control drug delivery actuators. To establish the communication links through the body tissues, we adopt galvanic coupling that uses low frequency electrical signals of weak amplitude. In this paper, we propose a topology management strategy using Weiszfeld algorithm that attempts to minimize the transmission power of the body nodes by reducing the distance from the source nodes to pick-up points or relays that gather and forward the received information. It takes into account the unique propagation model of the electrical signals within the body at various tissue layers, which is completely different from over the air RF. Our algorithm considers separately the constraints of on-skin nodes and the implanted nodes, especially in terms of minimizing the energy for the latter, which cannot be easily retrieved and re-charged. It also considers the difference in specific bandwidth requirements for the applications running within the nodes, by moving relays closer towards the high data rate demanding regions. We show that by optimizing the position of the relay node, the energy consumption can be significantly improved to extend the lifetime of the intra-body network up to several years.
Meenupriya Swaminathan, Gunar Schirner, Kaushik R. Chowdhury
WCNC2
2015 A Joint SW/HW Approach for Reducing Register File Vulnerability
abstract
The Register File (RF) is a particularly vulnerable component within processor core and at the same time a hotspot with high power density. To reduce RF vulnerability, conventional HW-only approaches such as Error Correction Codes (ECCs) or modular redundancies are not suitable due to their significant power overhead. Conversely, SW-only approaches either have limited improvement on RF reliability or require considerable performance overhead. As a result, new approaches are needed that reduce RF vulnerability with minimal power and performance overhead. This article introduces Application-guided Reliability-enhanced Register file Architecture (ARRA), a novel approach to reduce RF vulnerability of embedded processors. Taking advantage of uneven register utilization, ARRA mirrors, guided by a SW instrumentation, frequently used active registers into passive registers. ARRA is particularly suitable for control applications, as they have a high reliability demand with fairly low (uneven) RF utilization. ARRA is a cross-layer joint HW/SW approach based on an ARRA-extended RF microarchitecture, an ISA extension, as well as static binary analysis and instrumentation. We evaluate ARRA benefits using an ARRA-enhanced Blackfin processor executing a set of DSPBench and MiBench benchmarks. We quantify the benefits using RF Vulnerability Factor (RFVF) and Mean Work To Failure (MWTF). ARRA significantly reduces RFVF from 35% to 6.9% in cost of 0.5% performance lost for control applications. With ARRA’s register mirroring, it can also correct Multiple Bit Upsets (MBUs) errors, achieving an 8x increase in MWTF. Compared to a partially ECC-protected RF approach, ARRA demonstrates higher efficiency by achieving comparable vulnerability reduction at much lower power consumption.
Hamed Tabkhi, Gunar Schirner
ACM Trans. Archit. Code Optim.2
2014 Function-Level Processor (FLP): Raising efficiency by operating at function granularity for market-oriented MPSoC
abstract
The exponential growth in computation demand drives chip vendors to heterogeneous architectures combining Instruction-Level Processors (ILPs) and custom HW Accelerators (HWACCs) in an attempt to provide the needed processing capabilities while meeting power/energy requirements. ILPs, on one hand, are highly flexible, but power inefficient. Custom HWACCs, on the other hand, are inflexible (focusing on dedicated kernels), but highly power efficient. Since, designing HWACCs for every application is cost prohibitive, large portions of applications still run inefficiently on ILPs. New processing architectures are needed that combine the power efficiency of HWACCs while still retaining sufficient flexibility to realize applications across targeted market segments. This paper introduces Function-Level Processors (FLPs) to fill the gap between ILPs and dedicated HWACCs. FLPs are comprised of configurable Function Blocks (FBs) implementing selected functions which are then interconnected via programmable point-to-point connections constructing an extensible/configurable macro data-path. An FLP raises programming abstraction to a Function-Set Architecture (FSA) controlling FBs allocation, configuration and scheduling. We demonstrate FLP benefits with an industry example of the Pipeline-Vision Processor (PVP). We highlight the gained flexibility by mapping 10 embedded vision applications entirely to the FLP-PVP offering up to 22.4 GOPs/s with average power of 120 mW. The results also demonstrate that our FLP-PVP solution consumes 14×-18× less power than an ILP and 5x less power than a hybrid ILP+HWACCs solution.
Hamed Tabkhi, Robert Bushey, Gunar Schirner
ASAP3
2014 SIROM3 - A Scalable Intelligent Roaming Multi-modal Multi-sensor Framework
abstract
Understanding the future transportation infrastructure performance demands a smart Cyber-Physical Systems (CPS) approach integrating heterogeneous sensors, versatile computing systems, and mobile agents. However, due to sensor versatility and computing intricacy, designing such systems faces challenges of immense complexity in mobile sensor fusion, big data handling, system scalability, and integration. This paper introduces SIROM3, a Scalable Intelligent Roaming Multi-Modal Multi-Sensor framework, for next generation transportation infrastructure performance inspection. SIROM3offers a scalable and expandable framework through orthogonally abstracting software / hardware structures in a layered Run-Time Environment (RTE), which facilities sensor fusion, distributed computing, communication and mobile services. A Heterogeneous Stream File-system Overlay (HSFO) and a flexible plug-in system (PLEX) are embedded in SIROM3to simplify big data storage, processing, and correlation. To evaluate the scalability of SIROM, we implemented a mobile sensing system of 30 heterogeneous sensors and 5 computing platforms coordinated by 1 data center. SIROM's expandability is highlighted by adding an advanced radar platform which required less than 50 lines of C++ code for integration. Over 20 terabytes of data covering 300 miles have been collected, aggregated, and fused using SIROM3for comprehending the pavement dynamics of the entire city of Brockton, MA. SIROM3offers a unified solution and ideal research platform for rapid, intelligent and comprehensive evaluation of tomorrow's transportation infrastructure performance using heterogeneous systems.
Jiaxing Zhang 0003, Hanjiao Qiu, Salar Shahini Shamsabadi, Ralf Birken, Gunar Schirner
COMPSAC5
2014 Exploring the Heterogeneous Design Space for both Performance and Reliability
abstract
As we move into a new era of heterogeneous multi-core systems, our ability to tune the performance and understand the reliability of both hardware and software becomes more challenging. Given the multiplicity of different design trade-offs in hardware and software, and the rate of introduction of new architectures and hardware/software features, it becomes difficult to properly model emerging heterogeneous platforms.
Rafael Ubal, Dana Schaa, Perhaad Mistry, Yash Ukidave, Zhongliang Chen, Gunar Schirner, David R. Kaeli
DAC7
2014 Automatic specification granularity tuning for design space exploration
abstract
Algorithm Design Environments (ADE), such as Simulink, have been shown to be efficient for development, analysis, and evaluation of algorithms. Recent tools propose to facilitate algorithm / architecture co-design by bridging the gap from ADE to System-Level Design Environments (SLDE) through automatic synthesis from algorithm models to SLDL specifications. With the wide range of block characteristic (from simple logic functions to complex kernels) in the algorithm model, however, it is challenging to select a suitable compositional granularity for SLD Language (SLDL) blocks in the synthesized specification. A high volume of SLDL blocks of little computation will increase the number of mapping possibilities, whereas large blocks with heavy computation on the other hand allow inter-block fusion reducing the computational demands in the overall specification yet sacrificing the mapping flexibility. In this paper, we introduce an automatic specification granularity tuning mechanism to determine the granularity in the synthesized specification model hierarchy guided by the computational demands of algorithm blocks. Our granularity selection significantly simplifies the early design space exploration as only a meaningful block decomposition is exposed in the synthesized specification. It leads to an overall system with less computational demands by leveraging the block fusion capabilities in the ADE. At the same time our granularity decision ensures that sufficient flexibility remains in the system for exploring heterogeneous mapping of the algorithm. Our results on real world examples show that specification models can be synthesized with 80% efficiency through block fusion with 70–90% fewer but coarser grained blocks.
Jiaxing Zhang 0003, Gunar Schirner
DATE2
2014 A Power-Efficient FPGA-Based Mixture-of-Gaussian (MoG) Background Subtraction for Full-HD Resolution
abstract
This short paper briefly describes an FPGA-based realization of MoG background subtraction operating at fullHD frame resolution. Our HW hand-crafted MoG consists of 77 pipeline stages operating at 148.5 MHz implemented on a Zynq-7000 SoC. The results very high efficiency with a power consumption of less than 500 mW which is 600X more efficient than an embedded software solution.
Hamed Tabkhi, Majid Sabbagh, Gunar Schirner
FCCM3
2014 A GPU-Based Algorithm-Specific Optimization for High-Performance Background Subtraction
abstract
Background subtraction is an essential first stage in many vision applications differentiating foreground pixels from the background scene, with Mixture of Gaussians (MoG) being a widely used implementation choice. MoG's high computation demand renders a real-time single threaded realization infeasible. With it's pixel level parallelism, deploying MoG on top of parallel architectures such as a Graphics Processing Unit (GPU) is promising. However, MoG poses many challenges having a significant control flow (potentially reducing GPU efficiency) as well as a significant memory bandwidth demand. In this paper, we propose a GPU implementation of Mixture of Gaussians (MoG) that surpasses real-time processing for full HD (1080p 60 Hz). This paper describes step-wise optimizations starting from general GPU optimizations (such as memory coalescing, computation & communication overlapping), via algorithm-specific optimizations including control flow reduction and register usage optimization, to windowed optimization utilizing shared memory. For each optimization, this paper evaluates the performance potential and identifies architectural bottlenecks. Our CUDA-based implementation improves performance over sequential implementation by 57×, 97× and 101× through general, algorithm-specific, and windowed optimizations respectively, without impact to the output quality.
Chulian Zhang, Hamed Tabkhi, Gunar Schirner
ICPP3
2014 Application-Guided Power Gating Reducing Register File Static Power
abstract
Power and energy efficiency are on the top priority list in embedded computing. Embedded processors taped out in deep submicron technology have a high contribution of static power to overall power consumption. At the same time, current embedded processors often include a large register file (RF) to increase performance. However, a larger RF aggravates the static power issues associated with technology shrinking. Therefore, approaches to improve static power consumption of large RFs are in high demand. In this paper, we introduce an application-guided function-level register file power-gating (AFReP) approach to efficiently manage and reduce the RF's static power consumption. The AFReP is an interplay of automatic binary analysis and instrumentation at function-level granularity supported by instruction-set architecture and microarchitecture extensions. The AFReP enables runtime power-gating of registers during unutilized periods, whereas applications can fully benefit from a large RF during utilized periods. To demonstrate the AFReP's potential for reducing static power consumption, we have enhanced a Blackfin processor with the AFReP technology. Using the AFReP, the RF static power is reduced on average by 64% and 39% for control and DSP applications, respectively. At the same time, the AFReP only induces a very minimal overhead of 0.4% and 0.6%.
Hamed Tabkhi, Gunar Schirner
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Quantifying the energy efficiency of FFT on heterogeneous platforms
abstract
Heterogeneous computing using Graphic Processing Units (GPUs) has become an attractive computing model given the available scale of data-parallel performance and programming standards such as OpenCL. However, given the energy issues present with GPUs, some devices can exhaust power budgets quickly. Better solutions are needed to effectively exploit the power efficiency available on heterogeneous systems. In this paper we evaluate the power-performance trade-offs of different heterogeneous signal processing applications. More specifically, we compare the performance of 7 different implementations of the Fast Fourier Transform algorithms. Our study covers discrete GPUs and shared memory GPUs (APUs) from AMD (Llano APUs and the Southern Islands GPU), Nvidia (Fermi) and Intel (Ivy Bridge). For this range of platforms, we characterize the different FFTs and identify the specific architectural features that most impact power consumption. Using the 7 FFT kernels, we obtain a 48% reduction in power consumption and up to a 58% improvement in performance across these different FFT implementations. These differences are also found to be target architecture dependent. The results of this study will help the signal processing community identify which class of FFTs are most appropriate for a given platform. More important, we have demonstrated that different algorithms implementing the same fundamental function (FFT) can perform vastly different based on the target hardware and associated programming optimizations.
Yash Ukidave, Amir Kavyan Ziabari, Perhaad Mistry, Gunar Schirner, David R. Kaeli
ISPASS4
2012 Application-specific power-efficient approach for reducing register file vulnerability
abstract
This paper introduces a power efficient approach for improving reliability of heterogeneous register files in embedded processors. The approach is based on the fact that control applications have high demands in reliability, while many special-purpose register are unused in a considerable portion of execution. The paper proposes a static application binary analysis which is applied at function-level granularity and offers a systematic way to manage the RF's protection by mirroring the content of used registers into unused ones. The simulation results on an enhanced Blackfin processor demonstrate that Register File Vulnerability Factor (RFVF) is reduced from 35% to 6.9% in cost of 1% performance lost on average for control applications from Mibench suite.
Hamed Tabkhi, Gunar Schirner
DATE2
2012 AFReP: Application-guided Function-level Registerfile power-gating for embedded processors
abstract
With shrinking CMOS feature size, static power is growing significantly and power density has emerged as an increasing concern. At the same time, one trend of embedded processors is toward larger Register Files (RFs) which further increases static power dissipation and aggravating the issue. This paper introduces an Application-guided Function-level Register file Power-gating (AFReP) that reduces static power of RFs in embedded processors. Our AFReP approach is based on a automatic analysis of register lifetime in the application binary, followed an automatic binary instrumentation for runtime RF power-gating. The instrumented code executes on a processor with ISA and micro-architecture extension for power-gating control over individual registers. Our application binary analysis/instrumentation operates at function-level granularity, automatically gating the registers that do not contribute to program outcome. Our experimental results using an AFReP-enhanced Blackfin processor demonstrate average RF static power reduction by 60% and 52% for control and DSP applications from Mibench and DSPstone suites, respectively. The added instructions for run-time power-gating increase execution time by only 1% on average.
Hamed Tabkhi, Gunar Schirner
ICCAD2
2012 ARRA: Application-guided reliability-enhanced registerfile architecture for embedded processors
Hamed Tabkhi, Gunar Schirner
VLSI-SoC2
2010 Platform modeling for exploration and synthesis
abstract
Ever increasing complexity and heterogeneity of system platforms drive the need for a move to higher levels of abstraction accompanied by corresponding design automation tools. The basis for any automated flow are well-defined design models. In this paper, we present an overview and taxonomy of platform modeling at various levels. Experiments demonstrate the benefits of fast yet accurate intermediate models at varying levels for rapid, early design space exploration. Furthermore, paired with automatic model generation and hardware/software synthesis, an automated path from specification to implementation becomes possible.
Andreas Gerstlauer, Gunar Schirner
ASP-DAC2
2010 System-level development of embedded software
abstract
Embedded software plays an increasingly important role in implementing modern embedded systems. Development of embedded software, and of hardware-dependent software in particular, is challenging due to the tight integration with the underlying hardware architecture. In this paper, we describe our system-level design approach that allows designers to develop software in form of a platform-agnostic specification. Our design environment enables exploration of different architectural alternatives and subsequently generates the software implementation. It generates the application code, communication drivers, and an adaptation to a chosen RTOS. It completes the process by producing the final target binary for each processor. Our experimental results demonstrate the automatic generation of the binaries for five control and media oriented applications.
Gunar Schirner, Andreas Gerstlauer, Rainer Dömer
ASP-DAC1
2010 Accurate timed RTOS model for transaction level modeling
abstract
In this paper, we present an accurate timed RTOS model within transaction level models (TLMs). Our RTOS model, implemented on top of system level design language (SLDL), incorporates two key features: RTOS behavior model and RTOS overhead model. The RTOS behavior model provides dynamic scheduling, inter-process communication (IPC), and external communication for timing annotated user applications. While the RTOS behavior model is running, all RTOS events, such as context switch and interrupt handling, are passed to RTOS over-head model to adopt the overhead during system execution. Our RTOS overhead model has processor- and RTOS-specific pre-characterized overhead information to provide cycle approximate estimation. We demonstrate the applicability of our model using a multi-core platform executing a JPEG encoder. Experimental results show that the proposed RTOS model provides the high accuracy, 7% off compared to on-board measurements while simulating at speeds close to the reference C code.
Yonghyun Hwang, Gunar Schirner, Samar Abdi, Daniel Gajski
DATE2
2010 Fast and accurate processor models for efficient MPSoC design
abstract
With growing system complexity and ever-increasing software content, the development of embedded software for upcoming MPSoC architectures is a tremendous challenge. Traditional ISS-based validation becomes infeasible due to the large complexity. Addressing the need for flexible and fast simulating models, we introduce in this article our approach of abstract processor modeling in the context of multiprocessor architectures. We combine modeling of computation on processors with an abstract RTOS and accurate interrupt handling into a versatile, multifaceted processor model with several levels of features. Our processor models are utilized in a framework allowing designers to develop a system in a top-down manner using automatic model generation and compilation down to a given MPSoC architecture. During generation, instances of our processor models are integrated into a system model combining software, hardware, and bus communication. The generated system model serves for rapid design space exploration and a fast and accurate system validation. Our experimental results show the benefits of our processor modeling using an actual multiprocessor mobile phone baseband platform. Our abstract models of this complex system reach a simulation speed of 300MCycles/s within a high accuracy of less than 3% error. In addition, our results quantify the speed/accuracy trade-off at varying abstraction levels of our models to guide future processor model designers.
Gunar Schirner, Andreas Gerstlauer, Rainer Dömer
ACM Trans. Design Autom. Electr. Syst.1
2009 Hardware-dependent software synthesis for many-core embedded systems
abstract
This paper presents synthesis of hardware dependent software (HdS) for multicore and many-core designs using embedded system environment (ESE). ESE is a tool set, developed at UC Irvine, for transaction level design of multicore embedded systems. HdS synthesis is a key component of ESE back-end design flow. We follow a design process that starts with an application model consisting of C processes communicating via abstract message passing channels. The application model is mapped to a platform net-list of SW and HW cores, buses and buffers. A high speed transaction level model (TLM) is generated to validate abstract communication between processes mapped to different cores. The TLM is further refined into a pin-cycle accurate model (PCAM) for board implementation. The PCAM includes C code for all the HdS layers including routing, packeting, synchronization and bus transfer. The generated HdS methods provide a library of application level services to the C processes on individual SW cores. Therefore, the application developer does not need to write low level HdS for board implementation. Synthesis results for an multi-core MP3 decoder design, using ESE, show that the HdS is generated in order of seconds, compared to hours of manual coding. The quality of synthesized code is comparable to manually written code in terms of performance and code size.
Samar Abdi, Gunar Schirner, Ines Viskic, Hansu Cho, Yonghyun Hwang, Lochi Yu, Daniel Gajski
ASP-DAC2
2008 Automatic generation of hardware dependent software for MPSoCs from abstract system specifications
abstract
Increasing software content in embedded systems and SoCs drives the demand to automatically synthesize software binaries from abstract models. This is especially critical for Hardware dependent Software (HdS) due to the tight coupling. In this paper, we present our approach to automatically synthesize HdS from an abstract system model. We synthesize driver code, interrupt handlers and startup code. We furthermore automatically adjust the application to use RTOS services. We target traditional RTOS-based multi-tasking solutions, as well as a pure interrupt-based implementation (without any RTOS). Our experimental results show the automatic generation of final binary images for six real-life target applications and demonstrate significant productivity gains due to automation. Our HdS synthesis is an enabler for efficient MPSoC development and rapid design space exploration.
Gunar Schirner, Andreas Gerstlauer, Rainer Dömer
ASP-DAC1
2008 Introducing Preemptive Scheduling in Abstract RTOS Models using Result Oriented Modeling
abstract
With the increasing SW content of modern SoC designs, modeling and development of Hardware Dependent Software (HDS) become critical. Previous work addressed this by introducing abstract RTOS modeling, which exposes dynamic scheduling effects early in the system design flow. However, such models insufficiently capture preemption. In particular, the accuracy of preemption depends on the granularity of the timing annotation. For an accurately modeled interrupt response time, very fine-grained timing annotation is necessary, which contradicts the RTOS abstraction idea and is detrimental to simulation performance. In this paper, we eliminate the granularity dependency by applying the Result Oriented Modeling (ROM) technique previously used only for communication modeling. Our ROM approach allows precise preemptive scheduling, while retaining all the benefits of abstract RTOS modeling. Our experimental results demonstrate tremendous improvements. While the traditional model simulated an interrupt response time with a severe inaccuracy (12x longer in average and 40x longer for 96thpercentile), our ROM- based model was accurate within 8% (average and 50thpercentile) using identical timing annotations.
Gunar Schirner, Rainer Dömer
DATE1
2008 Quantitative analysis of the speed/accuracy trade-off in transaction level modeling
abstract
The increasing complexity of embedded systems requires modeling at higher levels of abstraction. Transaction level modeling (TLM) has been proposed to abstract communication for high-speed system simulation and rapid design space exploration. Although being widely accepted for its high performance and efficiency, TLM often exhibits a significant loss in model accuracy. In this article, we systematically analyze and quantify the speed/accuracy trade-off in TLM. To this end, we provide a classification of TLM abstraction levels based on model granularity and define appropriate metrics and test setups to quantitatively measure and compare the performance and accuracy of such models. Addressing several classes of embedded communication protocols, we apply our analysis to three common bus architectures, the industry-standard AMBA advanced high-performance bus (AHB) as an on-chip parallel bus, the controller area network (CAN) as an off-chip serial bus, and the Motorola ColdFire Master Bus as an example for a custom embedded processor bus. Based on the analysis of these individual busses, we then generalize our results for a broader conclusion. The general TLM trade-off offers gains of up to four orders of magnitude in simulation speed, generally however, at the price of low accuracy. We conclude further that model granularity is the key to efficient TLM abstraction, and we identify conditions for accuracy of abstract models. As a result, this article provides general guidelines that allow the system designer to navigate the TLM trade-off effectively and choose the most suitable model for the given application with fast and accurate results.
Gunar Schirner, Rainer Dömer
ACM Trans. Embed. Comput. Syst.1
2007 Abstract, Multifaceted Modeling of Embedded Processors for System Level Design
abstract
Embedded software is playing an increasing role in todays SoC designs. It allows a flexible adaptation to evolving standards and to customer specific demands. As software emerges more and more as a design bottleneck, early, fast, and accurate simulation of software becomes crucial. Therefore, an efficient modeling of programmable processors at high levels of abstraction is required. In this article, we focus on abstraction of computation and describe our abstract modeling of embedded processors. We combine the computation modeling with task scheduling support and accurate interrupt handling into a versatile, multi-faceted processor model with varying levels of features. Incorporating the abstract processor model into a communication model, we achieve fast co-simulation of a complete custom target architecture for a system level design exploration. We demonstrate the effectiveness of our approach using an industrial strength telecommunication example executing on a Motorola DSP architecture. Our results indicate the tremendous value of abstract processor modeling. Different feature levels achieve a simulation speedup of up to 6600 times with an error of less than 8% over a ISS based simulation. On the other hand, our full featured model exhibits a 3% error in simulated timing with a 1800 times speedup.
Gunar Schirner, Andreas Gerstlauer, Rainer Dömer
ASP-DAC1
2007 Result-Oriented Modeling - A Novel Technique for Fast and Accurate TLM
abstract
Efficient communication modeling is a critical task in system-on-chip design and exploration. In particular, fast and accurate communication is needed to predict the performance of a system. Recently, transaction level modeling is used to speed up communication simulation at the cost of accuracy. This paper proposes a novel modeling technique, called result-oriented modeling (ROM), which removes the inaccuracy drawback of transaction level models (TLMs) in many cases. Using ROM, simulation models yield nearly the same speed as their traditional TLM counterparts, yet are still 100% accurate in timing. ROM utilizes the fact that internal states in the communication channel are not observable by the caller. Hence, ROM omits the internal states entirely and optimistically predicts the end result. Retroactively, the outcome of the prediction is checked, and if necessary, corrective measures are taken to maintain the accuracy of the model. We have applied the ROM concept to two examples: the industry standard AMBA AHB and the controller area network. To validate the proposed ROM approach, we have analyzed the models in detail for performance and accuracy. Our experimental results show the clear advantages of the ROM concept. For both bus systems, ROM achieves 100% accuracy and highest speeds. In essence, ROM eliminates the TLM tradeoff for a wide range of platforms. It frees the system designer from having multiple models for different purposes and extends the TLM idea to applications that require timing accurate simulation, such as real-time communication.
Gunar Schirner, Rainer Dömer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2006 Quantitative analysis of transaction level models for the AMBA bus
abstract
The increasing complexity of embedded systems pushes system designers to higher levels of abstraction. Transaction level modeling (TLM) has been proposed to model communication in systems in an abstract manner Although being widely accepted, TLMs have not been analyzed for their loss in accuracy. This paper will analyze and quantify the speed-accuracy tradeoff of TLM using a case study on AMBA, an industry bus standard. It shows the results of modeling the advanced high-performance bus (AHB) of AMBA using a set of models at different abstraction levels. The analysis of the simulation speed shows improvements of two orders of magnitude for each TLM abstraction, while the timing in the model remains accurate for many applications. As a result, the paper will classify the different models towards their applicability in typical modeling situations, allowing the system designer to achieve fast and accurate simulation of communication
Gunar Schirner, Rainer Dömer
DATE1
2006 Fast and accurate transaction level models using result oriented modeling
abstract
Efficient communication modeling is a critical task in SoC design and exploration. In particular, fast and accurate communication is needed to predict the performance of a system. Recently, Transaction Level Modeling (TLM) is used to speedup communication simulation at the cost of accuracy. This paper proposes a novel modeling technique called Result Oriented Modeling (ROM) which removes the accuracy drawback of TLM. Using ROM, models yield the same speed as their TLM counterparts, yet still are 100 % accurate in timing. ROM utilizes the fact that internal states in the communication channel are not observable by the caller. Hence, ROM omits the internal states entirely and optimistically predicts the end result. Retroactively, the outcome is checked and, if necessary, corrective measures are taken to maintain the accuracy of the model. In this paper, we apply ROM to the AMBA AHB bus architecture. Our experimental results show that ROM exhibits the same high simulation performance as traditional TLM, yet it retains the same accuracy as the bus functional model. Thus, the proposed ROM approach eliminates the speed/accuracy tradeoff exhibited by traditional TLM. 1.
Gunar Schirner, Rainer Dömer
ICCAD1