Lu Wang 0019

dblp:49/3800-19 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0001-5759-6544ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Flex8: A Flexible Precision Co-design for 8-bit Neural Network
abstract
The rapid growth of neural network parameters poses a significant challenge for resource-constrained training, particularly in terms of storage and computation.While 8-bit quantization techniques have shown promise, their practical application remains limited, and mixed-precision methods have yet to achieve optimal results.This paper systematically evaluates the strengths and limitations of FP8 quantization by analyzing the impact of network architecture, exponent width, and mantissa width on FP8 training performance.Additionally, we investigate the effects of low-precision quantization across different layers, parameter types, and training stages.Based on these insights, we propose Flex8, a hardware-software co-design framework for mixed-precision optimization that integrates fine-grained quantization strategies, tunable exponent-width instructions, and a flexible 8-bit FPU design.Flex8 significantly improves FP8 training accuracy to levels comparable to singleprecision, offering a novel and effective approach to low-precision training.
Lu Wang 0019, Guangda Zhang, Xia Zhao 0004, Shiqing Zhang
CF2
2025 UGPU: Dynamically Constructing Unbalanced GPUs for Enhanced Resource Efficiency
abstract
Different GPU generations have various numbers of SMs but still keep the balanced idea during the manufacture, i.e., the proportion of compute and memory resources within a single physical GPU is similar.Although GPU applications have different characteristics, it is still uncommon and uneconomic to build unbalanced physical GPUs for customers.With their powerful computational capabilities, GPUs are widely used in the cloud to accelerate diverse workloads from multiple users, creating opportunities to explore the unbalanced GPU concept in multitasking environments.In this paper, we take the first step in exploring the feasibility and performance benefits of building unbalanced GPUs.Specifically, these unbalanced GPUs, referred to as GPU slices, are dynamically constructed with dedicated compute and memory resources from a single physical GPU to effectively address the diverse demands of co-executing applications, achieving high performance during the execution.However, there are two challenges that must to be solved.First, determining the size of unbalanced GPU slices during execution is challenging, as predicting GPU performance under varying resource allocations is inherently difficult.Second, reallocating memory resources after partitioning requires extensive data migration, with traditional methods leading to unacceptable performance degradation.To address the first challenge, UGPU employs a demand-aware resource partitioning algorithm that partitions resources dynamically without relying on a complex or inaccurate performance model.For the second challenge, UGPU introduces PageMove, a novel mechanism for efficient page migration between different memory dies within an HBM stack.Our key insight is that all memory channels already have physical connections to all through-silicon via (TSV) within a DRAM stack, while different bank groups can transfer data at the same time.PageMove slightly modifies DRAM architecture, uses a customized memory address mapping, designs a new parallel page migration mode (PPMM) and updates the virtual memory management scheme.By doing this,
Xia Zhao 0004, Guangda Zhang, Lu Wang 0019, Huadong Dai
ISCA3
2024 AdCoalescer: An Adaptive Coalescer to Reduce the Inter-Module Traffic in MCM-GPUs
abstract
The demand for greater computing power has driven the development of Multi-chip-module GPUs (MCM-GPUs), which greatly improve parallel processing capabilities. Unfortunately, MCM-GPUs have encountered a notable challenge, the performance bottleneck caused by remote accesses through the inter-module network. In this work, we found significant data access redundancy among SMs within a GPU module which can be coalesced to reduce network pressure. However, how to design the coalescing scheme to identify memory addresses with high data locality is still not clear.
Xu Zhang 0086, Guangda Zhang, Lu Wang 0019, Shiqing Zhang, Xia Zhao 0004
ICPP3
2020 MDM: The GPU Memory Divergence Model
abstract
Analytical models enable architects to carry out early-stage design space exploration several orders of magnitude faster than cycle-accurate simulation by capturing first-order performance phenomena with a set of mathematical equations. However, this speed advantage is void if the conclusions obtained through the model are misleading due to model inaccuracies. Therefore, a practical analytical model needs to be sufficiently accurate to capture key performance trends across a broad range of applications and architectural configurations.In this work, we focus on analytically modeling the performance of emerging memory-divergent GPU-compute applications which are common in domains such as machine learning and data analytics. The poor spatial locality of these applications leads to frequent L1 cache blocking due to the application issuing significantly more concurrent cache misses than the cache can support, which cripples the GPU’s ability to use Thread-Level Parallelism (TLP) to hide memory latencies. We propose the GPU Memory Divergence Model (MDM) which faithfully captures the key performance characteristics of memory-divergent applications, including memory request batching and excessive NoC/DRAM queueing delays. We validate MDM against detailed simulation and real hardware, and report substantial improvements in (1) scope: the ability to model prevalent memory-divergent applications in addition to non-memory divergent applications; (2) practicality: 6.1× faster by computing model inputs using binary instrumentation as opposed to functional simulation; and (3) accuracy: 13.9% average prediction error versus 162% for the state-of-the-art GPUMech model.
Lu Wang 0019, Magnus Jahre, Almutaz Adileh, Lieven Eeckhout
MICRO1
2019 Intra-Cluster Coalescing and Distributed-Block Scheduling to Reduce GPU NoC Pressure
abstract
GPUs continue to boost the number of streaming multiprocessors (SMs) to provide increasingly higher compute capabilities. To construct a scalable crossbar network-on-chip (NoC) that connects the SMs to the memory controllers, a cluster structure is introduced in modern GPUs in which several SMs are grouped together to share a network port. Because of network port sharing, clustered GPUs face severe NoC congestion, which creates a critical performance bottleneck. In this paper, we target redundant network traffic to mitigate GPU NoC congestion. In particular, we observe that in many GPU-compute applications, different SMs in a cluster access shared data. Sending redundant requests to access the same memory location wastes valuable NoC bandwidth-we find on average 19 percent (and up to 48 percent) of the requests to be redundant. To remove redundant NoC traffic, we propose distributed-block scheduling, intra-cluster coalescing (ICC) and the coalesced cache (CC) to coalesce L1 cache misses within and across SMs in a cluster, respectively. Our evaluation results show that distributed-block scheduling, ICC and CC are complementary and improve both performance and energy consumption. We report an average performance improvement of 15 percent (and up to 67 percent) while at the same time reducing system energy by 6 percent (and up to 19 percent) and improving the energy-delay product (EDP) by 19 percent on average (and up to 53 percent), compared to state-of-the-art distributed CTA scheduling.
Lu Wang 0019, Xia Zhao 0004, David R. Kaeli, Zhiying Wang 0003, Lieven Eeckhout
IEEE Trans. Computers1
2018 Intra-Cluster Coalescing to Reduce GPU NoC Pressure
abstract
GPUs continue to increase the number of streaming multiprocessors (SMs) to provide increasingly higher compute capabilities. To construct a scalable crossbar network-on-chip (NoC) that connects the SMs to the memory controllers, a cluster structure is introduced in modern GPUs in which several SMs are grouped together to share a network port. Because of network port sharing, clustered GPUs face severe NoC congestion, which creates a critical performance bottleneck. In this paper, we target redundant network traffic to mitigate GPU NoC congestion. In particular, we observe that in many GPU-compute applications, different SMs in a cluster access shared data. Issuing redundant requests to access the same memory location wastes valuable NoC bandwidth - we find on average 19.4% (and up to 48%) of the requests to be redundant. To reduce redundant NoC traffic, we propose intracluster coalescing (ICC) to merge memory requests from different SMs in a cluster. Our evaluation results show that ICC achieves an average performance improvement of 9.7% (and up to 33%) over a conventional design.
Lu Wang 0019, Xia Zhao 0004, David R. Kaeli, Zhiying Wang 0003, Lieven Eeckhout
IPDPS1
2017 A high performance reliable NoC router
Lu Wang 0019, Sheng Ma, Chen Li 0015, Wei Chen 0009, Zhiying Wang 0003
Integr.1
2016 A high performance reliable NoC router
abstract
Aggressive scaling of CMOS process technology allows the fabrication of highly integrated chips, and enables the design of multiprocessors system-on-chip connected by the network-on-chip (NoC). However, it brings about widespread reliability challenges. Aiming to tackle the permanent faults on the router components, we propose a high performance, high reliability and low cost router design based on a generic 2-stage router. Four fault tolerant strategies are added in our reliable router. We exploit a double routing strategy for the routing computation(RC) failure, a default winner strategy for the virtual channel allocation (VA), a runtime arbiter selection strategy for the switch allocation (SA) failure and a double bypass bus strategy for the crossbar failure. Different from previous reliable routers, our design leverages the feature of pipeline optimization and routing algorithm to maintain the performance in fault tolerance especially under heavy network loads. Besides, our proposed router provides higher reliability with lower hardware consumption than previous reliable router designs.
Lu Wang 0019, Sheng Ma, Zhiying Wang 0003
ASP-DAC1
2016 DLL: A dynamic latency-aware load-balancing strategy in 2.5D NoC architecture
abstract
As the 3D stacking technology still faces several challenges, the 2.5D stacking technology gains better application prospects nowadays. With the silicon interposer, the 2.5D stacking can improve the bandwidth and capacity of the memory system. To satisfy the communication requirements of the integrated memory system, the free routing resources in the interposer should be explored to implement an additional network. Yet, the performance is strongly limited by the unbalanced loads between the CPU-layer network and the interposer-layer network. In this paper, to address this issue, we propose a dynamic latency-aware load-balancing (DLL) strategy. Our key innovations are detecting congestion of the network layer via the average latency of recent packets and making the network layer selection at each source node. We leverage the free routing resources in the interposer to implement a latency propagation ring. With the ring, the latency information tracked at destination nodes is propagated back to source nodes. We achieve load-balance by using these information. Experimental results show that compared with the baseline design, a destination-detection strategy and a buffer-aware strategy, our DLL strategy achieves 45%, 14.9% and 6.5% of average throughput improvements with minor overheads.
Chen Li 0015, Sheng Ma, Lu Wang 0019, Zicong Wang, Xia Zhao 0004, Yang Guo 0003
ICCD3