Zhe Zhang 0006

dblp:87/5809-6 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
abstract
Large Language Models (LLMs) with Mixture-of-Expert (MoE) architectures achieve superior model performance with reduced computation costs, but at the cost of high memory capacity and bandwidth requirements. Near-Memory Processing (NMP) accelerators that stack memory directly on the compute through hybrid bonding have demonstrated high bandwidth with high energy efficiency, becoming a promising architecture for MoE models. However, as NMP accelerators comprise distributed memory and computation, how to map the MoE computation directly determines the LLM inference efficiency. Existing parallel mapping strategies, including Tensor Parallelism (TP) and Expert Parallelism (EP), suffer from either high communication costs or unbalanced computation utilization, leading to inferior efficiency. The dynamic routing mechanism of MoE LLMs further aggravates the efficiency challenges. Therefore, in this paper, we propose HD-MoE to automatically optimize the MoE parallel computation across an NMP accelerator. HD-MoE features an offline automatic hybrid parallel mapping algorithm and an online dynamic scheduling strategy to reduce the communication costs while maximizing the computation utilization. With extensive experimental results, we demonstrate that HD-MoE achieves a speedup ranging from 1.1× to 1.8× over TP, 1.1× to 1.5× over EP, and 1.0× to 1.4× over the baseline Hybrid TP-EP with Compute-Balanced parallelism strategies.
Haochen Huang, Shuzhang Zhong, Zhe Zhang 0006, Shuangchen Li, Dimin Niu, Hongzhong Zheng, Runsheng Wang, Meng Li 0004
ICCAD3
2022 Enabling High-Quality Uncertainty Quantification in a PIM Designed for Bayesian Neural Network
abstract
Uncertainty quantification measures the prediction uncertainty of a neural network facing out-of-training-distribution samples. Bayesian Neural Networks (BNNs) can provide high-quality uncertainty quantification by introducing specific noise to the weights during inference. To accelerate BNN inference, ReRAM processing-in-memory (PIM) architecture is a competitive solution to provide both high-efficient computing and in-situ noise generation at the same time. However, there normally exists a huge gap between the generated noise in PIM hardware and that required by a BNN model. We demonstrate that the quality of uncertainty quantification is substantially degraded due to this gap. To solve this problem, we propose a holistic framework called W2W-PIM. We first introduce an efficient method to generate noise in ReRAM PIM design according to the demand of a BNN model. In addition, the PIM architecture is carefully modified to enable the noise generation and evaluate uncertainty quality. Moreover, a calibration unit is further introduced to reduce the noise gap caused by imperfection of the noise model. Comprehensive evaluation results demonstrate that W2W-PIM framework can achieve high-quality uncertainty quantification and high energy-efficiency at the same time.
Bingzhe Wu, Guangyu Sun 0003, Zhe Zhang 0006, Zhihang Yuan, Runsheng Wang, Ru Huang 0001, Dimin Niu, Hongzhong Zheng, Zhichao Lu, Meng-Fan Chang, Tianchan Guan, Xin Si
HPCA4
2022 Hyperscale FPGA-as-a-service architecture for large-scale distributed graph neural network
abstract
Graph neural network (GNN) is a promising emerging application for link prediction, recommendation, etc. Existing hardware innovation is limited to single-machine GNN (SM-GNN), however, the enterprises usually adopt huge graph with large-scale distributed GNN (LSD-GNN) that has to be carried out with distributed in-memory storage. The LSD-GNN is very different from SM-GNN in terms of system architecture demand, workflow and operators, and hence characterizations.
Shuangchen Li, Dimin Niu, Yuhao Wang 0002, Zhe Zhang 0006, Tianchan Guan, Yijin Guan, Linyong Huang, Zhaoyang Du, Yuanwei Fang, Hongzhong Zheng, Yuan Xie 0001
ISCA5
2022 EPQuant: A Graph Neural Network compression approach based on product quantization
Linyong Huang, Zhe Zhang 0006, Zhaoyang Du, Shuangchen Li, Hongzhong Zheng, Yuan Xie 0001, Nianxiong Tan
Neurocomputing2
2021 BlockGNN: Towards Efficient GNN Acceleration Using Block-Circulant Weight Matrices
abstract
In recent years, Graph Neural Networks (GNNs) appear to be state-of-the-art algorithms for analyzing non-euclidean graph data. By applying deep-learning to extract high-level representations from graph structures, GNNs achieve extraordinary accuracy and great generalization ability in various tasks. However, with the ever-increasing graph sizes, more and more complicated GNN layers, and higher feature dimensions, the computational complexity of GNNs grows exponentially. How to inference GNNs in real time has become a challenging problem, especially for some resource-limited edge-computing platforms.To tackle this challenge, we propose BlockGNN, a software-hardware co-design approach to realize efficient GNN acceleration. At the algorithm level, we propose to leverage block-circulant weight matrices to greatly reduce the complexity of various GNN models. At the hardware design level, we propose a pipelined CirCore architecture, which supports efficient block-circulant matrices computation. Basing on CirCore, we present a novel BlockGNN accelerator to compute various GNNs with low latency. Moreover, to determine the optimal configurations for diverse deployed tasks, we also introduce a performance and resource model that helps choose the optimal hardware parameters automatically. Comprehensive experiments on the ZC706 FPGA platform demonstrate that on various GNN tasks, BlockGNN achieves up to 8.3× speedup compared to the baseline HyGCN architecture and 111.9× energy reduction compared to the Intel Xeon CPU platform.
Zhe Zhou 0002, Bizhao Shi, Zhe Zhang 0006, Yijin Guan, Guangyu Sun 0003, Guojie Luo
DAC3
2020 Reliability-Enhanced Circuit Design Flow Based on Approximate Logic Synthesis
abstract
With the downscaling of CMOS technology, the circuit design margin becomes more and more tight due to wider guardband, which is required to counteract the severer transistor aging and variations. Thus, reliability-enhanced circuit design is urgently needed to reduce the guardband. In this paper, a reliability-enhanced design framework based on approximate synthesis is proposed to completely eliminate the aging guardband. It mainly includes two key parts: first, a forward reliability simulation flow supporting statistical static timing analysis (SSTA) is performed to estimate the path failure rates after aging; if the timing constraints are not satisfied, then a backward delay-driven approximate logic synthesis flow will perform approximate local changes on the critical paths to reduce the delay until the reliability requirement is finally satisfied and no aging guardband is needed. The results show that the approximate circuit has a smaller aged delay than the original circuit, so that the path failure rates are significantly decreased. It indicates that the proposed design flow can convert the timing errors that have fatal impact on applications, into negligible error on low-significance bits to improve the resilience of circuits, which provides a new perspective of reliability-enhanced design at nanoscale.
Zuodong Zhang, Runsheng Wang, Zhe Zhang 0006, Ru Huang 0001, Chang Meng, Weikang Qian
ACM Great Lakes Symposium on VLSI3
2018 PIRVS: An Advanced Visual-Inertial SLAM System with Flexible Sensor Fusion and Hardware Co-Design
abstract
In this paper, we present the PerceptIn Robotics Vision System (PIRVS), a visual-inertial computing hardware with embedded simultaneous localization and mapping (SLAM) algorithm. The PIRVS hardware is equipped with a multi-core processor, a global-shutter stereo camera, and an IMU with precise hardware synchronization. The PIRVS software features a flexible sensor fusion approach to not only tightly integrate visual measurements with inertial measurements and also to loosely couple with additional sensor modalities. It runs in real-time on both PC and the PIRVS hardware. We perform a thorough evaluation of the proposed system using multiple public visual-inertial datasets. Experimental results demonstrate that our system reaches comparable accuracy of state-of-the-art visual-inertial algorithms on PC, while being more efficient on the PIRVS hardware.
Zhe Zhang 0006, Shaoshan Liu, Grace Tsai, Hongbing Hu, Chen-Chi Chu
ICRA1
2018 π-SoC: Heterogeneous SoC Architecture for Visual Inertial SLAM Applications
abstract
In recent years, we have observed a clear trend in the rapid rise of autonomous vehicles and robotics. One of the core technologies enabling these applications, Simultaneous Localization And Mapping (SLAM), imposes two main challenges: first, these workloads are computationally intensive and they often have real-time requirements; second, these workloads run on battery-powered mobile devices with limited energy budget. Hence, performance should be improved while simultaneously reducing energy consumption, two rather contradicting goals by conventional wisdom. Previous attempts to optimize SLAM performance and energy efficiency usually involve optimizing one function and fail to approach the problem systematically. In this paper, we first study the characteristics of visual inertial SLAM workloads on existing heterogeneous SoCs. Then based on the initial findings, we propose π-SoC, a heterogeneous SoC design that systematically optimize the IO interface, the memory hierarchy, as well as the the hardware accelerator. We implemented this system on a Xilinx Zynq UltraScale MPSoC and was able to deliver over 60 FPS performance with average power less than 5 W.
Jie Tang 0003, Bo Yu 0014, Shaoshan Liu, Zhe Zhang 0006, Weikang Fang
IROS4
2018 Trifo-VIO: Robust and Efficient Stereo Visual Inertial Odometry Using Points and Lines
abstract
In this paper, we present the Trifo Visual Inertial Odometry (Trifo-VIO), a tightly-coupled filtering-based stereo VIO system using both points and lines. Line features help improve system robustness in challenging scenarios when point features cannot be reliably detected or tracked, e.g. low-texture environment or lighting change. In addition, we propose a novel lightweight filtering-based loop closing technique to reduce accumulated drift without global bundle adjustment or pose graph optimization. We formulate loop closure as EKF updates to optimally relocate the current sliding window maintained by the filter to past keyframes. We also present the Trifo Ironsides dataset, a new visual-inertial dataset, featuring high-quality synchronized stereo camera and IMU data from the Ironsides sensor [3] with various motion types and textures and millimeter-accuracy groundtruth. To validate the performance of the proposed system, we conduct extensive comparison with state-of-the-art approaches (OKVIS, VINS-MONO and S-MSCKF) using both the public EuRoC dataset and the Trifo Ironsides dataset.
Grace Tsai, Zhe Zhang 0006, Shaoshan Liu, Chen-Chi Chu, Hongbing Hu
IROS3
2008 Robot-assisted intelligent 3D mapping of unknown cluttered search and rescue environments
abstract
In this paper a unique landmark identification method is proposed for identifying large distinguishable landmarks for 3D visual simultaneous localization and mapping (SLAM) in unknown cluttered urban search and rescue (USAR) environments. The novelty of the method is the utilization of both 3D (i.e., depth images) and 2D images. By utilizing a scale invariant feature transform (SIFT)-based approach and incorporating 3D depth imagery, we can achieve more reliable and robust recognition and matching of landmarks from multiple images for 3D mapping of the environment. Preliminary experiments utilizing the proposed methodology verify: (i) its ability to identify clusters of SIFT keypoints in both 3D and 2D images for representation of potential landmarks in the scene, and (ii) the use of the identified landmarks in constructing a 3D map of unknown cluttered USAR environments.
Zhe Zhang 0006, Goldie Nejat
IROS1
2007 Finding Disaster Victims: A Sensory System for Robot-Assisted 3D Mapping of Urban Search and Rescue Environments
abstract
In this paper the first application of utilizing a unique 3D real-time mapping sensor for sequential 3D map building within a visual simultaneous localization and mapping (SLAM) framework in unknown cluttered urban search and rescue (USAR) environments is proposed. The sensor utilizes a digital fringe projection and phase shifting technique to provide real-time 2D and 3D sensory information of the environment. The proposed sensor is unique over current technologies, in that it can directly map rubble in 3D and in real-time at a frame rate of up to 60 fps. Furthermore, we propose the development of a novel 3D visual SLAM method utilizing both 2D and 3D images taken by the sensor for robust and reliable landmark identification, mapping and localization algorithms utilizing a scale invariant feature transform (SIFT)-based approach. Preliminary experiments show the potential of the proposed 3D real-time sensory system for such unknown cluttered USAR environments.
Zhe Zhang 0006, Goldie Nejat, Peisen Huang
ICRA1
2006 Finding Disaster Victims: Robot-Assisted 3D Mapping of Urban Search and Rescue Environments via Landmark Identification
abstract
In this paper a landmark identification method is proposed for identifying large distinguishable landmarks for 3D visual simultaneous localization and mapping (SLAM) in a search and rescue environment. The novelty of the method is the utilization of both 3D (i.e., depth images) and 2D images. By utilizing a scale invariant feature transform (SIFT)-based approach and incorporating 3D depth imagery, we can use more reliable and robust recognition and matching between landmarks from multiple images for 3D mapping of the environment. Preliminary experiments utilizing the proposed method verified its ability to identify clusters of SIFT keypoints in the images for representation of potential landmarks in the scene
Goldie Nejat, Zhe Zhang 0006
ICARCV2