EDBT 2026 Demo / reviewers in the wild / expert
Guangda Zhang
dblp:135/7964
· DBLP profile ↗
23ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0003-4732-9674ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Energy Efficient Vector Gather Based on Data Serialization and Pulse Selection
Bingxi Pei, Guangda Zhang, Yande Jiang, Zhiyong Zhong |
ISCAS | 5 |
| 2026 | SRepair: Symbolic Regression-Based Repair for Hardware Design CodeabstractFixing bugs in hardware design code has become a challenging task due to the increasing complexity of modern circuit designs. As a result, automated program repair techniques have been proposed to synthesize patches for bugs in hardware designs and achieved promising results. However, existing techniques are still limited in synthesizing expressions for complex bugs. In this work, we explore the possibility of addressing complex bugs by proposing SREPAIR, a novel symbolic regression-based repair technique. The key novelty of SREPAIR lies in three aspects: 1) we propose a novel expression modification encoding that enables fine-grained adjustments to buggy expressions. 2) we introduce expression synthesis-based templates that allow for flexible and expressive repairs. 3) we develop a novel symbolic regression network-based synthesis algorithm that effectively synthesizes complex expressions. Experimental results on the four peer-reviewed datasets demonstrate that SREPAIR correctly fixes 56 bugs out of 112 bugs, which achieves 43.6% and 194.7% improvement over the previous state-of-the-art RTL-REPAIR (39 bugs) and CIRFIX (19 bugs). To evaluate the generalizability of SREPAIR, we further construct an augmented dataset of 282 bugs by mutating hardware designs. SREPAIR shows its better generalizability by correctly fixing 127 bugs, reaching 217.5% improvement over the best approach. Zizhen Liu, Deheng Yang, Xiaoguang Mao, Jiayu He, Guangda Zhang, Yan Lei 0005, Jiang Wu 0017 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | Flex8: A Flexible Precision Co-design for 8-bit Neural NetworkabstractThe rapid growth of neural network parameters poses a significant challenge for resource-constrained training, particularly in terms of storage and computation.While 8-bit quantization techniques have shown promise, their practical application remains limited, and mixed-precision methods have yet to achieve optimal results.This paper systematically evaluates the strengths and limitations of FP8 quantization by analyzing the impact of network architecture, exponent width, and mantissa width on FP8 training performance.Additionally, we investigate the effects of low-precision quantization across different layers, parameter types, and training stages.Based on these insights, we propose Flex8, a hardware-software co-design framework for mixed-precision optimization that integrates fine-grained quantization strategies, tunable exponent-width instructions, and a flexible 8-bit FPU design.Flex8 significantly improves FP8 training accuracy to levels comparable to singleprecision, offering a novel and effective approach to low-precision training. Lu Wang 0019, Guangda Zhang, Xia Zhao 0004, Shiqing Zhang |
CF | 3 |
| 2025 | NearFetch: Saving Inter-Module Bandwidth in Many-Chip-Module GPUsabstractAs Graphics Processing Units (GPUs) face increasing computing demands that surpass single-module capabilities due to transistor scaling and lithography constraints, the necessity for expanding the module count within GPUs grows. This escalation faces a significant challenge: the total inter-module bandwidth in many-chip-module GPUs is limited by manufacturing constraints in organic substrates or silicon interposers. Unlike Central Processing Units (CPUs), which are latency-sensitive, GPUs leverage their high thread-level parallelism to effectively hide memory access latency through simultaneous multithreading. This attribute makes GPUs inherently sensitive to bandwidth constraints, making the efficient exploitation of available inter-module bandwidth important. In this paper, we identify that fetching data from faraway memory in many-chip-module GPUs can easily cause bandwidth contention which degrades the real achieved data bandwidth per GPU module compared to fetching data from nearby memory. To further analyze this problem, we introduce the Inter-Module Bandwidth per Access (IBPA) metric for quantifying bandwidth usage and finding the network hop count directly impacts the IBPA and network contention. Next, we propose NearFetch, a routing-based solution to reduce IBPA. NearFetch works due to the fact that GPU modules along the routing path are typically much closer to the source GPU module while these GPU modules can supply $29.1 \%$ of data for high-sharing applications. NearFetch consists of two primary components: a data forwarding scheme, enabling data forwarding when the data resides in a remote GPU module, and a topology-aware Miss Status Handling Register (MSHR) coalescing scheme, responsible for recording the memory address information for future use in case of a data miss. By leveraging the data locality among various GPU modules, NearFetch substantially minimizes inter-module bandwidth usage, eliminating the need to fetch data from distant memory partitions. Our evaluation of NearFetch within the context of many-chip-module GPUs, across applications exhibiting diverse degrees of data locality, reveals that it reduces IBPA by $4 2. 6 \%$ and enhances performance by an average of $52.2 \%$ (with up to $9 8. 1 \%$ improvement) for high-sharing workloads. Guangda Zhang, Shiqing Zhang, Huadong Dai |
HPCA | 2 |
| 2025 | VIFMM: Custom RISC-V Vector Instructions with INT-FP Mixed-Precision Computing for Accelerating LLM InferenceabstractWeight-only quantization has been demonstrated to be an effective method to reduce memory and storage requirements for large language model (LLMs) deployed on resourceconstrained edge devices. However, it introduces frequent mixed precision operations between interger (INT) weights and floatingpoint (FP) activations. Conventional solutions typically quantize FP activations or dequantize INT weights during inference, which can degrade numerical accuracy and incur additional software overhead, ultimately slowing down inference. To eliminate this overhead, we propose VIFMM, an instruction set architecture (ISA) solution based on RISC-V vector (RVV) extension, supporting INT-FP mixed-precision computing, especially for operations between INT weights and FP activations, avoiding non-negligible data conversion in software. We implement VIFMM based on Ara, an open-source RVV core, along with a hardwarebased precision compensation mechanism to reuse the integer multiply accumulate calculation (MAC) units, to reduce hardware consumption and increase data precision. Experimental results show that VIFMM improves inference speed by 4%, with only 1.13% increase in hardware resource and 0.67% increase in power consumption, compared to conventional vector instructions. These findings illustrate that architecture-level support for INT-FP mixed-precision operations can significantly enhance LLM inference efficiency without sacrificing accuracy. Guangda Zhang, Bingxi Pei, Tianbo Zhang |
HPCC | 3 |
| 2025 | UGPU: Dynamically Constructing Unbalanced GPUs for Enhanced Resource EfficiencyabstractDifferent GPU generations have various numbers of SMs but still keep the balanced idea during the manufacture, i.e., the proportion of compute and memory resources within a single physical GPU is similar.Although GPU applications have different characteristics, it is still uncommon and uneconomic to build unbalanced physical GPUs for customers.With their powerful computational capabilities, GPUs are widely used in the cloud to accelerate diverse workloads from multiple users, creating opportunities to explore the unbalanced GPU concept in multitasking environments.In this paper, we take the first step in exploring the feasibility and performance benefits of building unbalanced GPUs.Specifically, these unbalanced GPUs, referred to as GPU slices, are dynamically constructed with dedicated compute and memory resources from a single physical GPU to effectively address the diverse demands of co-executing applications, achieving high performance during the execution.However, there are two challenges that must to be solved.First, determining the size of unbalanced GPU slices during execution is challenging, as predicting GPU performance under varying resource allocations is inherently difficult.Second, reallocating memory resources after partitioning requires extensive data migration, with traditional methods leading to unacceptable performance degradation.To address the first challenge, UGPU employs a demand-aware resource partitioning algorithm that partitions resources dynamically without relying on a complex or inaccurate performance model.For the second challenge, UGPU introduces PageMove, a novel mechanism for efficient page migration between different memory dies within an HBM stack.Our key insight is that all memory channels already have physical connections to all through-silicon via (TSV) within a DRAM stack, while different bank groups can transfer data at the same time.PageMove slightly modifies DRAM architecture, uses a customized memory address mapping, designs a new parallel page migration mode (PPMM) and updates the virtual memory management scheme.By doing this, Xia Zhao 0004, Guangda Zhang, Lu Wang 0019, Huadong Dai |
ISCA | 2 |
| 2025 | Dual-Branch CNN-Transformer network for robust Zero-Watermarking of medical images
Jingyou Li, Rongle Wei, Xiaotian Xi, Guangda Zhang, Zixin Yang, Fengshan Zhang |
Inf. Sci. | 4 |
| 2025 | Rtl design flaws revisited: a data-driven study of systematic bug patterns in Verilog code
Xiankai Meng, Guangda Zhang, Jiayu He, Deheng Yang, Fangshu Chen, Chengcheng Yu, Xinlin Zhao, Jiang Wu 0017 |
J. Supercomput. | 3 |
| 2024 | OLSATM: Online Learning Based State-Aware Task Migration on S-NUCA Many-CoresabstractTask migration maximizes performance while maintaining thermal safety in many-cores systems. Existing techniques exploit offline learning which requires tremendous training data and fixed-cycle migration which causes threads to miss the optimal migration timing. This paper presents Online Learning based State-Aware Task Migration (OLSATM). It pretrains a neural network (NN) with a small set of data and updates the model online to substitute the laborious data collection and model training of offline learning. It is state-aware and detects the timing when migration is needed, overcoming the shortcomings of periodical migration. Experimental results show that OLSATM enhances the performance by 3.7% and reduces the number of migration judgments by 26 % on average compared to the state-of-the-art task migration. Yandong He, Guangda Zhang, Hengzhu Liu, Renzhi Chen |
ICCD | 2 |
| 2024 | ChameSC: Virtualizing Superscalar Core of a SIMD Architecture for Vector Memory AccessabstractIn modern computing, the persistent issue of the memory wall significantly inhibits performance efficiency, especially in Single Instruction, Multiple Data (SIMD) architectures processing data-parallel workloads. Contemporary applications, including machine learning, data analytics, and computer vision, often feature complex chains of dependent, indirect, stride, and indexed vector memory accesses. These intricate access patterns place substantial pressure on vector memory units, leading to underutilization of SIMD resources and overall performance degradation. This study introduces ChameSC11A combination of Chameleon and Superscalar Core., a novel, adaptive SIMD architecture addressing the escalating memory wall issue. ChameSC exploits the typically idle memory units of the superscalar core during data-parallel workload execution by dynamically virtualizing the core to function as a precise data prefetcher, fetching data exactly as needed by the SIMD architecture in advance. The L1 cache of the superscalar core serves as a proactive cache to store prefetched data, which can be directly accessed by the vector memory unit. This strategic utilization of idle resources significantly enhances memory-level parallelism and effective memory bandwidth. Evaluations of ChameSC demonstrate substantial performance enhancements, achieving an average speedup of 34.3% compared to conventional SIMD architectures across a range of representative data-parallel applications with various memory access behaviors. Compared to state-of-the-art data prefetching techniques such as Berti and IPCP, ChameSC offers better performance without introducing extra storage overhead. Zhongzhu Pu, Guangda Zhang, Tiejian Zhang, Youhui Zhang |
ICCD | 2 |
| 2024 | AdCoalescer: An Adaptive Coalescer to Reduce the Inter-Module Traffic in MCM-GPUsabstractThe demand for greater computing power has driven the development of Multi-chip-module GPUs (MCM-GPUs), which greatly improve parallel processing capabilities. Unfortunately, MCM-GPUs have encountered a notable challenge, the performance bottleneck caused by remote accesses through the inter-module network. In this work, we found significant data access redundancy among SMs within a GPU module which can be coalesced to reduce network pressure. However, how to design the coalescing scheme to identify memory addresses with high data locality is still not clear. Xu Zhang 0086, Guangda Zhang, Lu Wang 0019, Shiqing Zhang, Xia Zhao 0004 |
ICPP | 2 |
| 2024 | A Fast and Safe Neuromorphic Approach for Obstacle Avoidance of Unmanned Aerial VehicleabstractObstacle avoidance is a crucial task in unmanned aerial vehicles (UAV) motion planning. The accuracy and consistency of real-time visual information affect the gener-ation of obstacle avoidance commands, raising higher safety demands for obstacle avoidance. The neuromorphic computing-based obstacle avoidance solution can address these challenges. Dynamic vision sensors (DVS) exhibit low latency, low power consumption, and high dynamic range as novel neuromorphic sensors. Spiking neural networks (SNN) also leverage the same mechanism to efficiently process asynchronous and sparse event data generated by DVS, offering latency and energy efficiency advantages. Additionally, the optimal estimation method effectively mitigates the impact of noise and interference within the system, reducing the influence of errors on the algorithm and enhancing safety. Based on these considerations, this paper proposes a fast and safe obstacle avoidance framework. DVS is used to acquire event data from the environment, and a hardware-compatible lightweight SNN is employed to extract dynamic obstacle position information from the data. Compared to baseline methods, this approach reduces latency by 85%. Furthermore, two estimation methods are used to predict the movement of obstacles, ensuring flight safety by generating different UAV obstacle avoidance actions based on confidence intervals, even in the presence of obstacle information errors and omissions. Zhong Wan, Xun Xiao, Jingyue Zhao, Junbo Tie, Renzhi Chen, Guangda Zhang, Huadong Dai |
SMC | 8 |
| 2024 | Cluster-aware scheduling in multitasking GPUs
Xia Zhao 0004, Huiquan Wang, Anwen Huang, Guangda Zhang |
Real Time Syst. | 5 |
| 2023 | NUBA: Non-Uniform Bandwidth GPUsabstractThe parallel execution model of GPUs enables scaling to hundreds of thousands of threads, which is a key capability that many modern high-performance applications exploit. GPU vendors are hence increasing the compute and memory resources with every GPU generation — resulting in the need to efficiently stitch together a plethora of Symmetric Multiprocessors (SMs), Last-Level Cache (LLC) slices and memory controllers while maximizing bandwidth and keeping power consumption and design complexity in check. Conventional GPUs are Uniform Bandwidth Architectures (UBAs) as they provide equal bandwidth between all SMs and all LLC slices. UBA GPUs require a uniform high-bandwidth Network-on-Chip (NoC), and our key observation is that provisioning a NoC to match the LLC slice bandwidth incurs a hefty power and complexity overhead. We propose the Non-Uniform Bandwidth Architecture (NUBA), a GPU system architecture aimed at fully utilizing LLC slice bandwidth. A NUBA GPU consists of partitions that each feature a few SMs and LLC slices as well as a memory controller — hence exposing the complete LLC bandwidth to the SMs within a partition since they can be connected with point-to-point links — and a NoC between partitions — to enable access to remote data.Exploiting the potential of NUBA GPUs however requires carefully co-designing system software, the compiler and architectural policies. The critical system software component is our Local-And-Balanced (LAB) page placement policy which enables the GPU driver to place data in local partitions while avoiding load imbalance. Moreover, we propose Model-Driven Replication (MDR) which identifies read-only shared data with data-flow analysis at compile time. At run time, MDR leverages an architectural mechanism that replicates read-only shared data across LLC slices when this can be done without pressuring cache capacity. With LAB and MDR, our NUBA GPU improves average performance by 23.1% and 22.2% (and up to 183.9% and 182.4%) compared to iso-resource memory-side and SM-side UBA GPUs, respectively. When the NUBA concept is leveraged to reduce overhead while maintaining similar performance, NUBA reduces NoC power consumption by 12.1× and 9.4×, respectively. Xia Zhao 0004, Magnus Jahre, Yuhua Tang, Guangda Zhang, Lieven Eeckhout |
ASPLOS (2) | 4 |
| 2023 | Dynamic Obstacle Avoidance for Unmanned Aerial Vehicle Using Dynamic Vision Sensor
Junbo Tie, Jingyue Zhao, Zhong Wan, Guangda Zhang, Lei Wang 0011 |
ICANN (10) | 11 |
| 2023 | Contrastive Hierarchical Gating Networks for Rating Prediction
Jiahui Wen, Mingyang Zhong, Guangda Zhang |
ICONIP (3) | 6 |
| 2021 | Joint aspect terms extraction and aspect categories detection via multi-task learning
Youcai Wei, Jian Fang 0004, Jiahui Wen, Guangda Zhang |
Expert Syst. Appl. | 6 |
| 2020 | Speculative text mining for document-level sentiment classification
Jiahui Wen, Guangda Zhang, Wei Yin 0002 |
Neurocomputing | 2 |
| 2020 | A unified model for recommendation with selective neighborhood modeling
Jiahui Wen, Mingyang Zhong, Guangda Zhang, Xue Li 0001 |
Inf. Process. Manag. | 5 |
| 2020 | Hierarchical text interaction for rating prediction
Jiahui Wen, Hongkui Tu, Mingyang Zhong, Guangda Zhang, Wei Yin 0002, Jian Fang 0004 |
Knowl. Based Syst. | 5 |
| 2017 | Handling Physical-Layer Deadlock Caused by Permanent Faults in Quasi-Delay-Insensitive Networks-on-ChipabstractNetworks-on-Chip (NoCs) are promising fabrics to provide scalable and efficient on-chip communication for large-scale many-core systems. In place of the well-studied synchronous NoCs, the event-driven asynchronous ones have emerged as promising replacement thanks to their strong timing robustness especially when implemented in quasi-delay-insensitive (QDI) circuits. However, their fault tolerance has rarely been studied. The QDI NoCs show complicated failure scenarios and behave differently from synchronous ones. As the scaling semiconductor technology is expected with the accelerated aging process, permanent faults become more likely to happen at runtime. These faults can break the handshake, leading to physical-layer deadlocks which can spread and paralyze the whole QDI NoC. This physical-layer deadlock cannot be resolved using conventional fault-tolerant or deadlock management techniques. This paper systematically studies the impact of permanent faults on QDI NoCs, and presents novel deadlock detection and recovery techniques to handle the fault-caused physical-layer deadlock. The proposed detection technique has been implemented to protect the NoC data paths that occupy~90% of the logic. Employing the detection and recovery techniques to protect interrouter links (~60% of the logic), a permanently faulty link is precisely located and the network function can be recovered with graceful performance degradation. Guangda Zhang, Wei Song 0002, Jim D. Garside, Javier Navaridas, Zhiying Wang 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | On-line detection of the deadlocks caused by permanently faulty links in quasi-delay insensitive networks on chipabstractAsynchronous networks on chip (NoCs) are promising candidates for supporting the enormous communication needed by future many-core systems due to their low-energy and high-speed. Similar to synchronous NoCs, asynchronous NoCs are vulnerable to faults but their fault-tolerance is not studied adequately, especially the quasi-delay insensitive (QDI) NoCs. One of the key issues neglected by most designers is that permanent faults in QDI NoCs cause deadlocks, which cripples the traditional fault-tolerant techniques using redundant codes. A novel detection method has been proposed to locate the faulty link in a QDI NoC according to a common pattern shared by all fault-related deadlocks. It is shown that this method introduces low hardware overhead and reports permanently faulty links with a short delay and guaranteed accuracy. Wei Song 0002, Guangda Zhang, Jim D. Garside |
ACM Great Lakes Symposium on VLSI | 2 |
| 2013 | Transient Fault Tolerant QDI Interconnects Using Redundant Check CodeabstractAsynchronous logic is a promising technology for building the chip-level interconnect of multi-core systems. However, asynchronous circuits are vulnerable to faults. This paper presents a novel scheme to improve the robustness of asynchronous systems. Our first contribution is a fault tolerant delay-insensitive redundant check coding scheme named DIRC. Using DIRC in 4-phase 1-of-n quasi-delay-insensitive (QDI) interconnects, all 1-bit and some multi-bit transient faults can be tolerated. The DIRC and the basic 4-phase 1-of-n pipeline stages are mutually exchangeable so that arbitrary basic stages can be replaced by DIRC stages to strengthen the fault-tolerance of long wires. Our second contribution, RPA, is a redundant technique to protect the acknowledge wires from transient faults - an issue that has long been disregarded by the community. The DIRC pipelines (using DIRC plus RPA) were simulated using the UMC 0.13μm standard cell library and compared with the basic pipelines. Detailed experimental results show that the 128-bit DIRC 1-of-4 pipeline is only 13% slower than the basic one but increases fault-tolerance hundred-folds when multi-bit transient faults are considered. Guangda Zhang, Wei Song 0002, Jim D. Garside, Javier Navaridas, Zhiying Wang 0003 |
DSD | 1 |