Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhuoran Zhao 0001

dblp:137/1412-1 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0003-1603-2712ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Performance modeling and evaluation · 62% Hardware reliability and fault tolerance · 18% Embedded and real-time systems · 11%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Computer networks
1 paper
Edge and fog computing · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed inference
0.312018
DeepThings: Distributed Adaptive Deep Learning Inference on Resource-Constrained IoT Edge Clusters · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Machine learning › Efficient and distributed learning › resource allocation
load balancing
0.312018
DeepThings: Distributed Adaptive Deep Learning Inference on Resource-Constrained IoT Edge Clusters · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Performance modeling and evaluation
performance prediction
0.312017
Source-Level Performance, Energy, Reliability, Power and Thermal (PERPT) Simulation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Performance modeling and evaluation
simulation
0.312017
Source-Level Performance, Energy, Reliability, Power and Thermal (PERPT) Simulation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Hardware reliability and fault tolerance
soft errors
0.112017
Source-Level Performance, Energy, Reliability, Power and Thermal (PERPT) Simulation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Energy-efficient computing › thermal modeling
thermal estimation
0.112017
Source-Level Performance, Energy, Reliability, Power and Thermal (PERPT) Simulation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Hardware reliability and fault tolerance › reliability analysis
vulnerability estimation
0.112017
Source-Level Performance, Energy, Reliability, Power and Thermal (PERPT) Simulation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017

Methods — techniques the papers use, named apart from their topics

work scheduling · 1.0fused tile partitioning · 1.0distributed work stealing · 1.0static back-annotation · 0.3host-compiled simulation · 0.3cache modeling · 0.3
YearPublicationVenuePosition
2020 Network-level Design Space Exploration of Resource-constrained Networks-of-Systems
abstract
Driven by recent advances in networking and computing technologies, distributed application scenarios are increasingly deployed on resource-constrained processing platforms. This includes networked embedded and cyber-physical systems as well as edge computing in mobile applications and the Internet of Things (IoT). In such resource-constrained Networks-of-Systems (NoS), computation and communication workloads need to be carefully co-optimized yet are tightly coupled. How to optimally partition and schedule application tasks among an appropriately designed NoS architecture requires a simultaneous consideration of design parameters from applications and processing platforms all the way to network configurations. Traditionally, however, systems and networks are designed in isolation and combined in an ad hoc manner, which ignores joint effects and optimization opportunities. To systematically explore and optimize NoS design spaces, a higher level of design abstraction on top of traditional system and network design is required. In this article, we propose a novel network-level design methodology for resource-constrained NoS optimization and design space exploration. A key component in such a design flow is fast yet accurate network/system co-simulation to rapidly evaluate NoS parameters with high fidelity. We first introduce a novel NoS simulator (NoSSim) that integrates source-level simulation models of applications with a host-compiled system simulation platform and a reconfigurable network simulation backplane to accurately capture system and network interactions. The co-simulation platform is further combined with model generation tools and a multi-objective genetic search algorithm to provide a comprehensive and fully automated NoS design space exploration framework. Finally, we apply our network-level design flow on several state-of-art IoT/mobile design case studies. Results show that NoSSim can achieve more than 86% simulation accuracy on average as compared to a real-world edge device cluster, where sensitivities to various design parameters are faithfully captured with high fidelity. When applying our network-level design space exploration methodology, design decisions are automatically optimized, where non-obvious NoS configurations are discovered outperforming manually designed solutions by more than 45%.
Zhuoran Zhao 0001, Kamyar Mirzazad Barijough, Andreas Gerstlauer
ACM Trans. Embed. Comput. Syst.1
2019 Quality/Latency-Aware Real-time Scheduling of Distributed Streaming IoT Applications
abstract
Embedded systems are increasingly networked and distributed, often, such as in the Internet of Things (IoT), over open networks with potentially unbounded delays. A key challenge is the need for real-time guarantees over such inherently unreliable and unpredictable networks. Generally, timeouts are used to provide timing guarantees while trading off data losses and quality. The schedule of distributed task executions and network timeouts thereby determines a fundamental latency-quality trade-off that is, however, not taken into account by existing scheduling algorithms. In this paper, we propose an approach for scheduling of distributed, real-time streaming applications under quality-latency goals. We formulate this as a problem of analytically deriving a static worst-case schedule of a given distributed dataflow graph that minimizes quality loss while meeting guaranteed latency constraints. Towards this end, we first develop a quality model that estimates SNR of distributed streaming applications under given network characteristics and an overall linearity assumption. Using this quality model, we then formulate and solve the scheduling of distributed dataflow graphs as a numerical optimization problem. Simulation results with random graphs show that quality/latency-aware scheduling improves SNR over a baseline schedule by 50% on average. When applied to a distributed neural network application for handwritten digit recognition, our scheduling methodology can improve classification accuracy by 10% over a naive distribution under tight latency constraints.
Kamyar Mirzazad Barijough, Zhuoran Zhao 0001, Andreas Gerstlauer
ACM Trans. Embed. Comput. Syst.2
2018 DeepThings: Distributed Adaptive Deep Learning Inference on Resource-Constrained IoT Edge Clusters
abstract
Edge computing has emerged as a trend to improve scalability, overhead, and privacy by processing large-scale data, e.g., in deep learning applications locally at the source. In IoT networks, edge devices are characterized by tight resource constraints and often dynamic nature of data sources, where existing approaches for deploying Deep/Convolutional Neural Networks (DNNs/CNNs) can only meet IoT constraints when severely reducing accuracy or using a static distribution that cannot adapt to dynamic IoT environments. In this paper, we propose DeepThings, a framework for adaptively distributed execution of CNN-based inference applications on tightly resource-constrained IoT edge clusters. DeepThings employs a scalable Fused Tile Partitioning (FTP) of convolutional layers to minimize memory footprint while exposing parallelism. It further realizes a distributed work stealing approach to enable dynamic workload distribution and balancing at inference runtime. Finally, we employ a novel work scheduling process to improve data reuse and reduce overall execution latency. Results show that our proposed FTP method can reduce memory footprint by more than 68% without sacrificing accuracy. Furthermore, compared to existing work sharing methods, our distributed work stealing and work scheduling improve throughput by 1.7 × -2.2× with multiple dynamic data sources. When combined, DeepThings provides scalable CNN inference speedups of 1.7×-3.5× on 2-6 edge devices with less than 23 MB memory each.
Zhuoran Zhao 0001, Kamyar Mirzazad Barijough, Andreas Gerstlauer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2017 Source-Level Performance, Energy, Reliability, Power and Thermal (PERPT) Simulation
abstract
With ever increasing design complexities, traditional cycle-accurate or instruction-set simulations are often too slow or too inaccurate for system prototyping in early design stages. As an alternative, host-compiled or source-level software simulation has been proposed, but existing approaches have largely focused on timing simulation only. In this paper, we propose a novel source-level simulation infrastructure that provides a full range of performance, energy, reliability, power and thermal (PERPT) estimation. Using a fully automated, retargetable back-annotation framework, intermediate representation code is statically annotated with timing, energy and resource accesses information obtained from low-level references at basic block granularity. The annotated model is natively compiled and combined with a cache model and occupancy analyzer to provide target performance, energy, soft-error vulnerability and power estimations. Finally, generated power traces are fed into thermal models for further temperature estimation. Comprehensive evaluations of our source-level models for PERPT estimations are performed. We applied our approach to PowerPC targets running various industry benchmark suites. source-level simulations are evaluated for different PERPT metrics and with cache models at various levels of detail to explore the speed and accuracy tradeoffs. More than 90% accuracy can be achieved for timing, energy, reliability and power estimation, and an average error of 0.05 K exists in steady-state thermal estimation. Simulation speeds range from 180 to 5740 MIPS for different types of metrics at different abstraction levels.
Zhuoran Zhao 0001, Andreas Gerstlauer, Lizy Kurian John
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1