EDBT 2026 Demo / reviewers in the wild / expert
Jinghan Zhang 0001
dblp:218/1130-1
· DBLP profile ↗
7ranked-venue papers
6as first author
4since 2021 · last 2023
0000-0003-2583-395XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 6 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | TSAR-ILP: Tile-Based, Synchronization-AwaRe ILP Allocating Heterogeneous Platforms for Streaming ApplicationsabstractAutomatic design space exploration (DSE) is key in hardware-software (HW/SW) co-design. To cope with the large design space, explorations are often heuristic-based and/or approximate yielding potentially locally optimal solutions. Without knowing the globally optimal solution, strong assertions about performance upper/lower bounds cannot be made. In contrast, integer linear programming (ILP) formulations can produce exact (optimal) solutions. Previous ILP-based formulations, however, lack support for tile-based architectures and realistic synchronization models, limiting their DSE capabilities. This work introduces a tile-based, synchronization-aware ILP (TSAR-ILP) formulation that overcomes previous limitations. With TSAR-ILP, the allocation/binding problems are introduced and formalized, attaining optimal solutions for mapping streaming applications onto template platforms. Using TSAR-ILP, this work explores a hardware accelerator-rich (HWACC-rich) platform with direct HWACC-to-HWACC communication under HW area constraints for 40 OpenVX applications. To illustrate design opportunities given by: 1) the ILP formulation and 2) direct HWACC-to-HWACC communication, this article analyzes the impact of job size. Results show that selecting smaller job sizes yields performance improvements and less area usage at the cost of slightly increased synchronization overhead. A job size reduction from 1 kB to 256 bytes gives$3.51\times $average performance increase across 40 applications. Finally, DSE with TSAR-ILP is shown not to be prohibitive through scalability analysis using a set of 5000 synthetic applications with varying size (10–125 nodes), with 94.3% of applications successfully achieving optimal solutions under 60 s. Bruno Morais, Jinghan Zhang 0001, Gunar Schirner |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Generating Unified Platforms Using Multigranularity Domain DSE (MG-DmDSE) Exploiting Application SimilaritiesabstractHeterogeneous accelerator-rich (ACC-rich) platforms combining general-purpose cores and specialized HW accelerators (ACCs) promise high-performance and low-power streaming application deployments in a variety of domains, such as video analytics and software-defined radio. In order to benefit a domain of applications, a domain platform exploration tool must take advantage of structural and functional similarities across applications by allocating a common set of ACCs. A previous approach proposed a genetic domain exploration tool (GIDE) that applied a restrictive binding algorithm that mapped applications functions to monolithic accelerators. This approach suffered from a low average application throughput across and reduced platform generality. This article introduces a multigranularity-based domain design space exploration tool (MG-DmDSE) to improve both average application throughput as well as platform generality. The key contributions of MG-DmDSE are: 1) applying a multigranular decomposition of coarse-grained application functions into more granular compute kernels; 2) examining compute similarity between functions in order to provide more generic functions; 3) configuring monolithic ACCs by selectively bypassing compute elements within them during DSE to expose more functionality; and 4) speeding up MG-DmDSE platform allocation exploration through a greedy guided mutation (GGM) algorithm. To assess MG-DmDSE, both GIDE and MG-DmDSE were applied to applications in the OpenVX library. MG-DmDSE achieves an average$2.84\times $greater application throughput compared to GIDE. Additionally, 87.5% of applications benefited from running on the platform produced by MG-DmDSE versus 50% from GIDE, which indicated increased platform generality. The generated MG-DmDSE platforms achieve an average of 61.8% logarithmic throughput improvement for unknown applications over GIDE. GGM results in saving 84.8% of the exploration time in MG-DmDSE with only 0.23% performance loss. Jinghan Zhang 0001, Aly Sultan, Mehrshad Zandigohar, Gunar Schirner |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | MG-DmDSE: Multi-Granularity Domain Design Space Exploration Considering Function SimilarityabstractHeterogeneous accelerator-rich (ACC-rich) platforms combining general-purpose cores and specialized HW accelerators (ACCs) promise high-performance and low-power streaming application deployments in a variety of domains such as video analytics and software-defined radio. In order to benefit a domain of applications, a domain platform exploration tool must take advantage of structural and functional similarities across applications by allocating a common set of ACCs. A previous approach [1] proposed a GenetIc Domain Exploration tool (GIDE) that applied a restrictive binding algorithm that mapped applications functions to monolithic accelerators. This approach suffered from lower average application throughput across and reduced platform generality. This paper introduces a Multi-Granularity based Domain Design Space Exploration tool (MG-DmDSE) to improve both average application throughput as well as platform generality. The key contributions of MG-DmDSE are: (1) Applying a multi-granular decomposition of coarse grain application functions into more granular compute kernels. (2) Examining compute similarity between functions in order to produce more generic functions. (3) Configuring monolithic ACCs by selectively bypassing compute elements within them during DSE to expose more functionality. To assess MG-DmDSE, both GIDE and MG-DmDSE were applied to applications in the OpenVX library. MG-DmDSE achieves an average 2.84x greater application throughput compared to GIDE. Additionally, 87.5% of applications benefited from running on the platform produced by MG-DmDSE vs 50% from GIDE, which indicated increase platform generality. Jinghan Zhang 0001, Aly Sultan, Hamed Tabkhi, Gunar Schirner |
DATE | 1 |
| 2021 | RDP3: Rapid Domain Platform Performance Prediction for Design Space ExplorationabstractHeterogeneous Accelerator-rich (ACC-rich) platforms combining general-purpose cores and specialized HW Accelerators (ACCs) promise high-performance and low-power deployment of streaming applications, e.g. for video analytics, software-defined radio, and radar. In order to recover Non-Recurring Engineering (NRE) cost, a unified domain platform for a set of applications can be exploited, especially when applications have functional and structural similarities, which can benefit from common ACCs. However, identifying the most beneficial set of common ACCs is challenging, and current Design Space Exploration (DSE) methods for domain platform allocation suffer from a long exploration time bottleneck. In particular, compared to a traditional DSE, evaluating the performance of a platform for a domain of applications is much more time-consuming as binding exploration and evaluation for each application in the domain is required. Thus, a rapid domain performance evaluation is needed to speed up the exploration of the platform allocation.This paper introduces Rapid Domain Platform Performance Prediction (RDP3) methods to speed up the exploration in domain DSE. Key contributions are: (1) analyzing current domain DSE flow and its exploration time bottleneck; (2) introducing four RDP3methods to speedup the evaluation of different platform allocations: Heuristic Processing (HP) estimation, Linear Regression (LR), Decision Tree Regression (DTR), and Multi-Layer Perceptron (MLP) predictions; (3) comparing the performance of these predictions and integrating the prediction into the current domain DSE. To evaluate the efficacy of RDP3, we explore 10K platforms capable of processing OpenVX domain applications. We demonstrate that RDP3-MLP as the most promising method can achieve a speedup of 17.5K times with only 0.001 mean square error compared to the current platform evaluation using the analytical model. Integrating RDP3-MLP into the existing domain DSE method GIDE [1] can save 80.8% exploration time while still resulting in the same output platform design. Jinghan Zhang 0001, Mehrshad Zandigohar, Gunar Schirner |
ICCD | 1 |
| 2020 | Allocating One Common ACC-Rich Platform for Many Streaming ApplicationsabstractMany demanding streaming applications share functional and structural similarities with other apps in their respective domain, e.g., video analytics, software-defined radio, and radar. This opens the opportunity for specialization (e.g., heterogeneous computing) to achieve the needed efficiency and/or performance. However, current design space exploration (DSE) focuses on an individual application in isolation (e.g., one particular vision flow), but not a set of similar applications. Hence, optimizations that occur due to considering multiple applications simultaneously are missed. New DSE methodologies and tools are needed with a broader scope of application sets instead of individual applications. This article introduces a novel domain-specific DSE (DS-DSE) approach focusing on streaming applications. Key contributions are: 1) a formalized method to extract the functional and structural similarities of domain applications; 2) a rapid platform performance estimation and comparison at two abstraction levels: domain score (DS) and analytic performance estimation (APE) model; 3) two novel algorithms, dynamic score selection (DSS), and GenetIc domain exploration (GIDE), for hardware/software partitioning of a domain-specific platform to maximize the throughput across domain applications (under certain constraints); and 4) a methodology to evaluate a platform's benefit for a set of applications. We demonstrate DSS's and GIDE's benefits using OpenVX applications and synthetic domains. The DSS and GIDE generated domain-specific platforms improve performance over application-specific platforms by 58% and 75% for OpenVX, as well as by 23% and 48% for synthetic applications. GIDE's platforms reach 99.8% (OpenVX) and 97.6% (synthetic) throughput of the domain optimal platform obtained through exhaustive search. Jinghan Zhang 0001, Hamed Tabkhi, Gunar Schirner |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Mitigating Application Diversity for Allocating a Unified ACC-Rich PlatformabstractHeterogeneous accelerator-rich (ACC-rich) platforms combining general-purpose cores and specialized HW accelerators (ACCs) promise high-performance and low-power streaming application (app) deployments, e.g. for video analytics, software-defined radio, and radar. In order to recover NRE, a unified platform for a set of applications (apps) is desirable. When apps have functional and structural similarities, they can benefit from common ACCs. Identifying the most beneficial set of common ACCs is challenging. However, current allocation strategies mostly focus on one app in isolation. Automatically allocating a unified platform requires simultaneously considering many apps, an efficient design space traversal and a fair evaluation across diverse apps. This paper introduces a Unified ACC-rich Platform Allocation (UPA) methodology for sets of data flow apps. Key contributions are: (1) a genetic algorithm (GA) guided by a fair and efficient evaluation to allocate one unified platform for many apps, (2) defining relative efficiency for fair comparison across diverse apps, and (3) defining metrics to quantify many app platform efficiency. This paper demonstrates UPA's benefits using OpenVX apps. A 12-ACCs-UPA improves average efficiency 4.59x over app-dedicated platforms. The UPA platform enables more apps (55% of OpenVX apps) to be efficiently deployed (≥ 60% of optimal app-dedicated platform). The benefits increase even further with increasing ACC budget. Jinghan Zhang 0001, Hamed Tabkhi, Gunar Schirner |
ICCD | 1 |
| 2018 | DS-DSE: Domain-specific design space exploration for streaming applicationsabstractDomain-specific computing is promising for high-performance low-power execution of applications with similar functionality. In particular, streaming applications with significant functional and structural similarities can tremendously benefit. However, current Design Space Exploration (DSE) focuses on individual applications in isolation. Hence, much of the domain optimization opportunities are missed. DSE methodologies need to broaden the scope from individual applications in isolation to optimizing across applications within a domain. This paper introduces a novel Domain-Specific DSE (DS-DSE) approach for domain-specific computing with a focus on streaming applications. Key contributions are: (1) a formalized method to extract the functional and structural similarities of domain applications, (2) a novel algorithm for hardware/software partitioning of a domain-specific platform to maximize the throughput across domain applications (under certain constraints) and (3) a methodology to evaluate a domain platform. This paper demonstrates the benefits using 4 domains: OpenVX (vision processing), and 3 synthetic domains (with greater complexity). Our experiments demonstrate a performance improvement (average throughput) of 36.8% for OpenVX and 46.2% for synthetic domains of the DS-DSE generated platform compared to an application-specific platform. Jinghan Zhang 0001, Hamed Tabkhi, Gunar Schirner |
DATE | 1 |