Guangda Zhang

dblp:135/7964 · DBLP profile ↗
← Back
23ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0003-4732-9674ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Energy Efficient Vector Gather Based on Data Serialization and Pulse Selection
Bingxi Pei, Guangda Zhang, Yande Jiang, Zhiyong Zhong
ISCAS5
2026 SRepair: Symbolic Regression-Based Repair for Hardware Design Code
abstract
Fixing bugs in hardware design code has become a challenging task due to the increasing complexity of modern circuit designs. As a result, automated program repair techniques have been proposed to synthesize patches for bugs in hardware designs and achieved promising results. However, existing techniques are still limited in synthesizing expressions for complex bugs. In this work, we explore the possibility of addressing complex bugs by proposing SREPAIR, a novel symbolic regression-based repair technique. The key novelty of SREPAIR lies in three aspects: 1) we propose a novel expression modification encoding that enables fine-grained adjustments to buggy expressions. 2) we introduce expression synthesis-based templates that allow for flexible and expressive repairs. 3) we develop a novel symbolic regression network-based synthesis algorithm that effectively synthesizes complex expressions. Experimental results on the four peer-reviewed datasets demonstrate that SREPAIR correctly fixes 56 bugs out of 112 bugs, which achieves 43.6% and 194.7% improvement over the previous state-of-the-art RTL-REPAIR (39 bugs) and CIRFIX (19 bugs). To evaluate the generalizability of SREPAIR, we further construct an augmented dataset of 282 bugs by mutating hardware designs. SREPAIR shows its better generalizability by correctly fixing 127 bugs, reaching 217.5% improvement over the best approach.
Zizhen Liu, Deheng Yang, Xiaoguang Mao, Jiayu He, Guangda Zhang, Yan Lei 0005, Jiang Wu 0017
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 Flex8: A Flexible Precision Co-design for 8-bit Neural Network
abstract
The rapid growth of neural network parameters poses a significant challenge for resource-constrained training, particularly in terms of storage and computation.While 8-bit quantization techniques have shown promise, their practical application remains limited, and mixed-precision methods have yet to achieve optimal results.This paper systematically evaluates the strengths and limitations of FP8 quantization by analyzing the impact of network architecture, exponent width, and mantissa width on FP8 training performance.Additionally, we investigate the effects of low-precision quantization across different layers, parameter types, and training stages.Based on these insights, we propose Flex8, a hardware-software co-design framework for mixed-precision optimization that integrates fine-grained quantization strategies, tunable exponent-width instructions, and a flexible 8-bit FPU design.Flex8 significantly improves FP8 training accuracy to levels comparable to singleprecision, offering a novel and effective approach to low-precision training.
Lu Wang 0019, Guangda Zhang, Xia Zhao 0004, Shiqing Zhang
CF3
2025 NearFetch: Saving Inter-Module Bandwidth in Many-Chip-Module GPUs
abstract
As Graphics Processing Units (GPUs) face increasing computing demands that surpass single-module capabilities due to transistor scaling and lithography constraints, the necessity for expanding the module count within GPUs grows. This escalation faces a significant challenge: the total inter-module bandwidth in many-chip-module GPUs is limited by manufacturing constraints in organic substrates or silicon interposers. Unlike Central Processing Units (CPUs), which are latency-sensitive, GPUs leverage their high thread-level parallelism to effectively hide memory access latency through simultaneous multithreading. This attribute makes GPUs inherently sensitive to bandwidth constraints, making the efficient exploitation of available inter-module bandwidth important. In this paper, we identify that fetching data from faraway memory in many-chip-module GPUs can easily cause bandwidth contention which degrades the real achieved data bandwidth per GPU module compared to fetching data from nearby memory. To further analyze this problem, we introduce the Inter-Module Bandwidth per Access (IBPA) metric for quantifying bandwidth usage and finding the network hop count directly impacts the IBPA and network contention. Next, we propose NearFetch, a routing-based solution to reduce IBPA. NearFetch works due to the fact that GPU modules along the routing path are typically much closer to the source GPU module while these GPU modules can supply $29.1 \%$ of data for high-sharing applications. NearFetch consists of two primary components: a data forwarding scheme, enabling data forwarding when the data resides in a remote GPU module, and a topology-aware Miss Status Handling Register (MSHR) coalescing scheme, responsible for recording the memory address information for future use in case of a data miss. By leveraging the data locality among various GPU modules, NearFetch substantially minimizes inter-module bandwidth usage, eliminating the need to fetch data from distant memory partitions. Our evaluation of NearFetch within the context of many-chip-module GPUs, across applications exhibiting diverse degrees of data locality, reveals that it reduces IBPA by $4 2. 6 \%$ and enhances performance by an average of $52.2 \%$ (with up to $9 8. 1 \%$ improvement) for high-sharing workloads.
Guangda Zhang, Shiqing Zhang, Huadong Dai
HPCA2
2025 VIFMM: Custom RISC-V Vector Instructions with INT-FP Mixed-Precision Computing for Accelerating LLM Inference
abstract
Weight-only quantization has been demonstrated to be an effective method to reduce memory and storage requirements for large language model (LLMs) deployed on resourceconstrained edge devices. However, it introduces frequent mixed precision operations between interger (INT) weights and floatingpoint (FP) activations. Conventional solutions typically quantize FP activations or dequantize INT weights during inference, which can degrade numerical accuracy and incur additional software overhead, ultimately slowing down inference. To eliminate this overhead, we propose VIFMM, an instruction set architecture (ISA) solution based on RISC-V vector (RVV) extension, supporting INT-FP mixed-precision computing, especially for operations between INT weights and FP activations, avoiding non-negligible data conversion in software. We implement VIFMM based on Ara, an open-source RVV core, along with a hardwarebased precision compensation mechanism to reuse the integer multiply accumulate calculation (MAC) units, to reduce hardware consumption and increase data precision. Experimental results show that VIFMM improves inference speed by 4%, with only 1.13% increase in hardware resource and 0.67% increase in power consumption, compared to conventional vector instructions. These findings illustrate that architecture-level support for INT-FP mixed-precision operations can significantly enhance LLM inference efficiency without sacrificing accuracy.
Guangda Zhang, Bingxi Pei, Tianbo Zhang
HPCC3
2025 UGPU: Dynamically Constructing Unbalanced GPUs for Enhanced Resource Efficiency
abstract
Different GPU generations have various numbers of SMs but still keep the balanced idea during the manufacture, i.e., the proportion of compute and memory resources within a single physical GPU is similar.Although GPU applications have different characteristics, it is still uncommon and uneconomic to build unbalanced physical GPUs for customers.With their powerful computational capabilities, GPUs are widely used in the cloud to accelerate diverse workloads from multiple users, creating opportunities to explore the unbalanced GPU concept in multitasking environments.In this paper, we take the first step in exploring the feasibility and performance benefits of building unbalanced GPUs.Specifically, these unbalanced GPUs, referred to as GPU slices, are dynamically constructed with dedicated compute and memory resources from a single physical GPU to effectively address the diverse demands of co-executing applications, achieving high performance during the execution.However, there are two challenges that must to be solved.First, determining the size of unbalanced GPU slices during execution is challenging, as predicting GPU performance under varying resource allocations is inherently difficult.Second, reallocating memory resources after partitioning requires extensive data migration, with traditional methods leading to unacceptable performance degradation.To address the first challenge, UGPU employs a demand-aware resource partitioning algorithm that partitions resources dynamically without relying on a complex or inaccurate performance model.For the second challenge, UGPU introduces PageMove, a novel mechanism for efficient page migration between different memory dies within an HBM stack.Our key insight is that all memory channels already have physical connections to all through-silicon via (TSV) within a DRAM stack, while different bank groups can transfer data at the same time.PageMove slightly modifies DRAM architecture, uses a customized memory address mapping, designs a new parallel page migration mode (PPMM) and updates the virtual memory management scheme.By doing this,
Xia Zhao 0004, Guangda Zhang, Lu Wang 0019, Huadong Dai
ISCA2
2025 Dual-Branch CNN-Transformer network for robust Zero-Watermarking of medical images
Jingyou Li, Rongle Wei, Xiaotian Xi, Guangda Zhang, Zixin Yang, Fengshan Zhang
Inf. Sci.4
2025 Rtl design flaws revisited: a data-driven study of systematic bug patterns in Verilog code
Xiankai Meng, Guangda Zhang, Jiayu He, Deheng Yang, Fangshu Chen, Chengcheng Yu, Xinlin Zhao, Jiang Wu 0017
J. Supercomput.3
2024 OLSATM: Online Learning Based State-Aware Task Migration on S-NUCA Many-Cores
abstract
Task migration maximizes performance while maintaining thermal safety in many-cores systems. Existing techniques exploit offline learning which requires tremendous training data and fixed-cycle migration which causes threads to miss the optimal migration timing. This paper presents Online Learning based State-Aware Task Migration (OLSATM). It pretrains a neural network (NN) with a small set of data and updates the model online to substitute the laborious data collection and model training of offline learning. It is state-aware and detects the timing when migration is needed, overcoming the shortcomings of periodical migration. Experimental results show that OLSATM enhances the performance by 3.7% and reduces the number of migration judgments by 26 % on average compared to the state-of-the-art task migration.
Yandong He, Guangda Zhang, Hengzhu Liu, Renzhi Chen
ICCD2
2024 ChameSC: Virtualizing Superscalar Core of a SIMD Architecture for Vector Memory Access
abstract
In modern computing, the persistent issue of the memory wall significantly inhibits performance efficiency, especially in Single Instruction, Multiple Data (SIMD) architectures processing data-parallel workloads. Contemporary applications, including machine learning, data analytics, and computer vision, often feature complex chains of dependent, indirect, stride, and indexed vector memory accesses. These intricate access patterns place substantial pressure on vector memory units, leading to underutilization of SIMD resources and overall performance degradation. This study introduces ChameSC11A combination of Chameleon and Superscalar Core., a novel, adaptive SIMD architecture addressing the escalating memory wall issue. ChameSC exploits the typically idle memory units of the superscalar core during data-parallel workload execution by dynamically virtualizing the core to function as a precise data prefetcher, fetching data exactly as needed by the SIMD architecture in advance. The L1 cache of the superscalar core serves as a proactive cache to store prefetched data, which can be directly accessed by the vector memory unit. This strategic utilization of idle resources significantly enhances memory-level parallelism and effective memory bandwidth. Evaluations of ChameSC demonstrate substantial performance enhancements, achieving an average speedup of 34.3% compared to conventional SIMD architectures across a range of representative data-parallel applications with various memory access behaviors. Compared to state-of-the-art data prefetching techniques such as Berti and IPCP, ChameSC offers better performance without introducing extra storage overhead.
Zhongzhu Pu, Guangda Zhang, Tiejian Zhang, Youhui Zhang
ICCD2
2024 AdCoalescer: An Adaptive Coalescer to Reduce the Inter-Module Traffic in MCM-GPUs
abstract
The demand for greater computing power has driven the development of Multi-chip-module GPUs (MCM-GPUs), which greatly improve parallel processing capabilities. Unfortunately, MCM-GPUs have encountered a notable challenge, the performance bottleneck caused by remote accesses through the inter-module network. In this work, we found significant data access redundancy among SMs within a GPU module which can be coalesced to reduce network pressure. However, how to design the coalescing scheme to identify memory addresses with high data locality is still not clear.
Xu Zhang 0086, Guangda Zhang, Lu Wang 0019, Shiqing Zhang, Xia Zhao 0004
ICPP2
2024 A Fast and Safe Neuromorphic Approach for Obstacle Avoidance of Unmanned Aerial Vehicle
abstract
Obstacle avoidance is a crucial task in unmanned aerial vehicles (UAV) motion planning. The accuracy and consistency of real-time visual information affect the gener-ation of obstacle avoidance commands, raising higher safety demands for obstacle avoidance. The neuromorphic computing-based obstacle avoidance solution can address these challenges. Dynamic vision sensors (DVS) exhibit low latency, low power consumption, and high dynamic range as novel neuromorphic sensors. Spiking neural networks (SNN) also leverage the same mechanism to efficiently process asynchronous and sparse event data generated by DVS, offering latency and energy efficiency advantages. Additionally, the optimal estimation method effectively mitigates the impact of noise and interference within the system, reducing the influence of errors on the algorithm and enhancing safety. Based on these considerations, this paper proposes a fast and safe obstacle avoidance framework. DVS is used to acquire event data from the environment, and a hardware-compatible lightweight SNN is employed to extract dynamic obstacle position information from the data. Compared to baseline methods, this approach reduces latency by 85%. Furthermore, two estimation methods are used to predict the movement of obstacles, ensuring flight safety by generating different UAV obstacle avoidance actions based on confidence intervals, even in the presence of obstacle information errors and omissions.
Zhong Wan, Xun Xiao, Jingyue Zhao, Junbo Tie, Renzhi Chen, Guangda Zhang, Huadong Dai
SMC8
2024 Cluster-aware scheduling in multitasking GPUs
Xia Zhao 0004, Huiquan Wang, Anwen Huang, Guangda Zhang
Real Time Syst.5
2023 NUBA: Non-Uniform Bandwidth GPUs
abstract
The parallel execution model of GPUs enables scaling to hundreds of thousands of threads, which is a key capability that many modern high-performance applications exploit. GPU vendors are hence increasing the compute and memory resources with every GPU generation — resulting in the need to efficiently stitch together a plethora of Symmetric Multiprocessors (SMs), Last-Level Cache (LLC) slices and memory controllers while maximizing bandwidth and keeping power consumption and design complexity in check. Conventional GPUs are Uniform Bandwidth Architectures (UBAs) as they provide equal bandwidth between all SMs and all LLC slices. UBA GPUs require a uniform high-bandwidth Network-on-Chip (NoC), and our key observation is that provisioning a NoC to match the LLC slice bandwidth incurs a hefty power and complexity overhead. We propose the Non-Uniform Bandwidth Architecture (NUBA), a GPU system architecture aimed at fully utilizing LLC slice bandwidth. A NUBA GPU consists of partitions that each feature a few SMs and LLC slices as well as a memory controller — hence exposing the complete LLC bandwidth to the SMs within a partition since they can be connected with point-to-point links — and a NoC between partitions — to enable access to remote data.Exploiting the potential of NUBA GPUs however requires carefully co-designing system software, the compiler and architectural policies. The critical system software component is our Local-And-Balanced (LAB) page placement policy which enables the GPU driver to place data in local partitions while avoiding load imbalance. Moreover, we propose Model-Driven Replication (MDR) which identifies read-only shared data with data-flow analysis at compile time. At run time, MDR leverages an architectural mechanism that replicates read-only shared data across LLC slices when this can be done without pressuring cache capacity. With LAB and MDR, our NUBA GPU improves average performance by 23.1% and 22.2% (and up to 183.9% and 182.4%) compared to iso-resource memory-side and SM-side UBA GPUs, respectively. When the NUBA concept is leveraged to reduce overhead while maintaining similar performance, NUBA reduces NoC power consumption by 12.1× and 9.4×, respectively.
Xia Zhao 0004, Magnus Jahre, Yuhua Tang, Guangda Zhang, Lieven Eeckhout
ASPLOS (2)4
2023 Dynamic Obstacle Avoidance for Unmanned Aerial Vehicle Using Dynamic Vision Sensor
Junbo Tie, Jingyue Zhao, Zhong Wan, Guangda Zhang, Lei Wang 0011
ICANN (10)11
2023 Contrastive Hierarchical Gating Networks for Rating Prediction
Jiahui Wen, Mingyang Zhong, Guangda Zhang
ICONIP (3)6
2021 Joint aspect terms extraction and aspect categories detection via multi-task learning
Youcai Wei, Jian Fang 0004, Jiahui Wen, Guangda Zhang
Expert Syst. Appl.6
2020 Speculative text mining for document-level sentiment classification
Jiahui Wen, Guangda Zhang, Wei Yin 0002
Neurocomputing2
2020 A unified model for recommendation with selective neighborhood modeling
Jiahui Wen, Mingyang Zhong, Guangda Zhang, Xue Li 0001
Inf. Process. Manag.5
2020 Hierarchical text interaction for rating prediction
Jiahui Wen, Hongkui Tu, Mingyang Zhong, Guangda Zhang, Wei Yin 0002, Jian Fang 0004
Knowl. Based Syst.5
2017 Handling Physical-Layer Deadlock Caused by Permanent Faults in Quasi-Delay-Insensitive Networks-on-Chip
abstract
Networks-on-Chip (NoCs) are promising fabrics to provide scalable and efficient on-chip communication for large-scale many-core systems. In place of the well-studied synchronous NoCs, the event-driven asynchronous ones have emerged as promising replacement thanks to their strong timing robustness especially when implemented in quasi-delay-insensitive (QDI) circuits. However, their fault tolerance has rarely been studied. The QDI NoCs show complicated failure scenarios and behave differently from synchronous ones. As the scaling semiconductor technology is expected with the accelerated aging process, permanent faults become more likely to happen at runtime. These faults can break the handshake, leading to physical-layer deadlocks which can spread and paralyze the whole QDI NoC. This physical-layer deadlock cannot be resolved using conventional fault-tolerant or deadlock management techniques. This paper systematically studies the impact of permanent faults on QDI NoCs, and presents novel deadlock detection and recovery techniques to handle the fault-caused physical-layer deadlock. The proposed detection technique has been implemented to protect the NoC data paths that occupy~90% of the logic. Employing the detection and recovery techniques to protect interrouter links (~60% of the logic), a permanently faulty link is precisely located and the network function can be recovered with graceful performance degradation.
Guangda Zhang, Wei Song 0002, Jim D. Garside, Javier Navaridas, Zhiying Wang 0003
IEEE Trans. Very Large Scale Integr. Syst.1
2014 On-line detection of the deadlocks caused by permanently faulty links in quasi-delay insensitive networks on chip
abstract
Asynchronous networks on chip (NoCs) are promising candidates for supporting the enormous communication needed by future many-core systems due to their low-energy and high-speed. Similar to synchronous NoCs, asynchronous NoCs are vulnerable to faults but their fault-tolerance is not studied adequately, especially the quasi-delay insensitive (QDI) NoCs. One of the key issues neglected by most designers is that permanent faults in QDI NoCs cause deadlocks, which cripples the traditional fault-tolerant techniques using redundant codes. A novel detection method has been proposed to locate the faulty link in a QDI NoC according to a common pattern shared by all fault-related deadlocks. It is shown that this method introduces low hardware overhead and reports permanently faulty links with a short delay and guaranteed accuracy.
Wei Song 0002, Guangda Zhang, Jim D. Garside
ACM Great Lakes Symposium on VLSI2
2013 Transient Fault Tolerant QDI Interconnects Using Redundant Check Code
abstract
Asynchronous logic is a promising technology for building the chip-level interconnect of multi-core systems. However, asynchronous circuits are vulnerable to faults. This paper presents a novel scheme to improve the robustness of asynchronous systems. Our first contribution is a fault tolerant delay-insensitive redundant check coding scheme named DIRC. Using DIRC in 4-phase 1-of-n quasi-delay-insensitive (QDI) interconnects, all 1-bit and some multi-bit transient faults can be tolerated. The DIRC and the basic 4-phase 1-of-n pipeline stages are mutually exchangeable so that arbitrary basic stages can be replaced by DIRC stages to strengthen the fault-tolerance of long wires. Our second contribution, RPA, is a redundant technique to protect the acknowledge wires from transient faults - an issue that has long been disregarded by the community. The DIRC pipelines (using DIRC plus RPA) were simulated using the UMC 0.13μm standard cell library and compared with the basic pipelines. Detailed experimental results show that the 128-bit DIRC 1-of-4 pipeline is only 13% slower than the basic one but increases fault-tolerance hundred-folds when multi-bit transient faults are considered.
Guangda Zhang, Wei Song 0002, Jim D. Garside, Javier Navaridas, Zhiying Wang 0003
DSD1