VLDB 2026 Research / reviewers in the wild / expert
Wu Zhou 0007
dblp:14/7096-7
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-1688-1151ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Identifying Optimal Workload Offloading Partitions for CPU-PIM Graph Processing AcceleratorsabstractThe integrated architecture that features both in-memory logic and host processors, or so-called “processing-in-memory” (PIM) architecture, is an emerging and promising solution to bridge the performance gap between the memory and host processors. In spite of the considerable potential of PIM, the workload offloading policy, which partitions the program and determines where code snippets are executed, is still a main challenge in PIM. In order to determine the best PIM offloading partitions, existing methods require in-depth program profiling to create the control flow graph (CFG) and then transform it into a graph-cut problem. These CFG-based solutions depend on detailed profiling of a crucial element, the execution time of basic blocks, to accurately assess the benefits of PIM offloading. The issue is that these execution times can change significantly in PIM, leading to inaccurate offloading decisions. To tackle this challenge, we present a novel PIM workload offloading framework called “RDPIM” for CPU-PIM graph processing accelerators, which systematically considers the variations in the execution time of basic blocks. By analyzing the relationship between data dependencies among workloads and the connectivity of input graphs, we identified three key features that can lead to variations in execution time. We developed a novel reuse distance (RD)-based model to predict the exact performance of basic blocks for optimal offloading decisions. We evaluate RDPIM using real-world graphs and compare it with some state-of-the-art PIM offloading approaches. Experiments have demonstrated that our method achieves an average speedup of$2\times $compared to CPU-only executions and up to$1.6\times $compared to state-of-the-art PIM offloading schemes. Le Luo 0002, Wu Zhou 0007, Xiaoming Chen 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | RHT_NoC: A Reconfigurable Hybrid Topology Architecture for Chiplet-Based Multicore SystemabstractChiplet-based system-on-chip (SoC) architectures, leveraging 2.5-D/3-D integration technologies, provide scalable solutions for a wide range of applications. Achieving high performance and cost-effectiveness in these systems relies heavily on optimizing die-to-die interconnect topologies and designs, which are essential for seamless interchiplet communication. This article introduces a reconfigurable hybrid topology (RHT) architecture designed for chiplet-based multicore systems. RHT achieves high performance and energy efficiency by dynamically reconfiguring the network topology to traffic variations, adaptively selecting transport subnets, and optimizing link bandwidth allocation, thereby minimizing congestion and maximizing packet throughput. Furthermore, RHT leverages global traffic information to dynamically combine Torus loops, maximizing opportunities for rapid packet transmission delivery while guaranteeing minimal hop counts. Moreover, RHT accelerates packet transmission via bufferless combined loops, extending the continuous sleeping periods of routers, improves power gating efficiency, and significantly reduces static power consumption. Simulation results indicate that the Mesh-DyRing achieves over a 40% reduction in network latency and more than a 20% decrease in power consumption overhead compared to the baseline design. When compared to WiNoC, an advanced hybrid wired-wireless topology design, the Mesh-DyRing-PG configuration reduces power consumption by 56.2% while maintaining equivalent average network latency. Wu Zhou 0007, Zhengfeng Huang, Huaguo Liang, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | Reconfigurable Fault-Tolerant Link With Bandwidth Expansion for 2.5-D Chiplet-Based SystemsabstractThe 2.5-D chiplet-based systems offer a promising path toward higher performance and integration density, but the reliability of interchiplet links, particularly vertical links (VLs), poses a significant challenge. This article proposes reconfigurable bidirectional link (ReBL), a novel architecture providing robust fault tolerance and enhanced bandwidth for these critical interconnects. ReBL features three key innovations: 1) a robust fault-tolerant mechanism leveraging ReBLs that dynamically adapts to and mitigates permanent link failures; 2) a dynamic bandwidth expansion technique that significantly enhances the system efficiency by utilizing idle links to optimize resource allocation and throughput; and 3) a virtual channel (VC) allocation strategy that guarantees deadlock-free operations through strategic channel partitioning and assignment. Evaluations using synthetic traffic and PARSEC benchmarks demonstrate ReBL’s significant advantages under high fault rates. Compared with the state-of-the-art reliable and deadlock-free routing (ReD) approach, ReBL achieves an average reduction of 21.3% in packet latency and 4.9% in application execution time across the evaluated benchmarks. These benefits are achieved with only a 6.06% area overhead over baseline. Wu Zhou 0007, Le Luo 0002, Fulong Chen 0002, Tianming Ni, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | DBU-PG: energy-efficient noc design using dual-buffering power gating
Wu Zhou 0007, Zhengfeng Huang, Huaguo Liang |
J. Supercomput. | 4 |
| 2023 | Dynamic detection of wireless interface faults and fault-tolerant routing algorithm in WiNoC
Chang Qian, Wu Zhou 0007, Qi Wang 0027, Huaguo Liang |
Integr. | 4 |
| 2023 | Improving power and performance of on-chip network through virtual channel sharing and power gating
Wu Zhou 0007, Huaguo Liang |
Integr. | 3 |
| 2023 | A transparent virtual channel power gating method for on-chip network routers
Wu Zhou 0007, Jianhua Li 0003 |
Integr. | 1 |
| 2023 | RMC_NoC: A Reliable On-Chip Network Architecture With Reconfigurable Multifunctional ChannelabstractAs chip fabrication has advanced to the nano level, the increased link density has heightened the risk of failures. The potential performance drawbacks resulting from these link failures have become a critical challenge in the design of reliable network-on-chip (NoC) systems. Fault-tolerant routing algorithms have proven to be effective strategies for handling this issue by diverting packets away from failed links to prevent congestion. However, these algorithms often result in excessive packet diversion, especially in the presence of a higher failure rate, which can significantly constrain the network’s behavior. This article introduces a novel NoC design with reconfigurable multifunctional channels (RMC_NoC). This design dynamically adapts the channel functions in response to network conditions to ensure that packets from failed links follow their original paths. In addition, it presents a channel buffer bubble flow control mechanism that can resolve congestion by redistributing congested traffic within the channel buffer. The evaluation results demonstrate that our approach ensures superior network communication even in the presence of permanent link failures, with minimal area overhead and power consumption. Moreover, our system exhibits lower latency and higher throughput compared to state-of-the-art fault-tolerant methods across various link failure rates. Notably, even at a severe failure rate of 30%, RMC_NoC exhibits only a 16.3% increase in latency compared to an ideal failure-free environment (Baseline) while still maintaining system communication capabilities to a considerable extent. Wu Zhou 0007, Zhengfeng Huang, Huaguo Liang, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | Energy-Efficient Multiple Network-on-Chip Architecture With Bandwidth ExpansionabstractAs technology feature sizes diminish to the nanometer regime, the leakage power crisis has become a major challenge in network-on-chip (NoC) design. Power gating (PG) is used to mitigate growing leakage power as an effective static power-saving technique. Applying PG in a multiple NoC (Multi-NoC) rather than a traditional NoC is a promising solution. However, limited by the channel width of the subnets, the increase in packet length will bring a severe serialization issue and performance loss. Previous Multi-NoC schemes have to wake up more subnets to minimize the performance loss, which also sacrifices their energy efficiency. In this article, we introduce an architecture, namely, BandExp, which allows subnets to expand their bandwidth by utilizing the idle physical links of other subnets. More bandwidth helps subnets mitigate the serialization issue and reduce the performance loss. Meanwhile, other subnets gain longer sleep cycles and thus save more energy. Evaluation results indicate that compared to the state-of-the-art Catnap, the proposed architecture reduces the average packet latency and execution time of different benchmarks by 19.3% and 3.2%, respectively. Also, the net static energy of the network is reduced by 23.2% on average, while the incurred area overhead is only 1.3%. Wu Zhou 0007, Zhengfeng Huang, Huaguo Liang, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |