VLDB 2026 Research / reviewers in the wild / expert
Qinzhe Wu
dblp:220/5329
· DBLP profile ↗
13ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-7988-1431ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Performance Implications of Pipelining the Data Transfer in CPU-GPU Heterogeneous SystemsabstractDriven by the increasing demands of machine learning, heterogeneous systems combining CPUs and GPUs have emerged as the dominant architecture for parallel computing in recent years. To optimize memory management and data transfer between CPUs and GPUs, Nvidia GPUs have introduced unified virtual memory ( UVM ) and pinned memory ( PM ) over the last decade. UVM can avoid explicit memory copies and potentially overlap GPU kernel computations with CPU-GPU data transfer. PM ensures that data with high locality remains in the main memory, preventing it from being paged out. In addition to these two techniques, asynchronous memory copy ( Async Memcpy ) was introduced recently in Nvidia GPUs to improve the CPU-GPU pipeline further. By utilizing Async Memcpy , the data transfer from GPU global memory to shared memory can be overlapped with GPU computations, adding an additional stage to the CPU-GPU data transfer pipeline. A thorough performance analysis of how Async Memcpy affects the current UVM and PM CPU-GPU data transfer scheme is desired. In this article, we provide performance implications of the combined effect of UVM , PM , and Async Memcpy , exploring which applications benefit from which combination of these features. We implement all these features on a suite of 25 workloads, including microbenchmarks and realworld applications. We observe an average performance gain of 24% when utilizing UVM and a 34% gain when employing PM on realworld applications, compared to not applying any data transfer optimization techniques. The performance benefits of Async Memcpy vary across different workloads. For workloads featuring extensive shared memory usage and high compute density (e.g., kmeans and lud ), Async Memcpy delivers around a 20% performance improvement over using UVM or PM alone. In other workloads like knn , we note a 20% performance degradation when using Async Memcpy . Furthermore, we conduct an in-depth investigation of the GPU kernel using performance counters to uncover the root causes of performance differences among various data transfer models. We also perform sensitivity analyses to examine how the number of blocks and threads, as well as the L1-cache/shared memory partitioning, impact performance. We explore future research directions aimed at enhancing the data transfer pipeline by overlapping memory allocation with data transfer and computation across GPU kernels. Ruihao Li 0002, Bagus Hanindhito, Sanjana Yadav, Qinzhe Wu, Krishna M. Kavi, Gayatri Mehta, Neeraja J. Yadwadkar, Lizy Kurian John |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | BLQ: Light-Weight Locality-Aware Runtime for Blocking-Less QueuingabstractMessage queues are used widely in parallel processing systems for worker thread synchronization. When there is a throughput mismatch between the upstream and downstream tasks, the message queue buffer will often exist as either empty or full. Polling on an empty or full queue will affect the performance of upstream or downstream threads, since such polling cycles could have been spent on other computation. Non-blocking queue is an alternative that allow polling cycles to be spared for other tasks per applications’ choice. However, application programmers are not supposed to bear the burden, because a good decision of what to do upon blocking has to take many runtime environment information into consideration. Qinzhe Wu, Ruihao Li 0002, Jonathan C. Beard, Lizy Kurian John |
CC | 1 |
| 2023 | NextGen-Malloc: Giving Memory Allocator Its Own Room in the HouseabstractMemory allocation and management have a significant impact on performance and energy of modern applications. We observe that performance can vary by as much as 72% in some applications based on which memory allocator is used. Many current allocators are multi-threaded to support concurrent allocation requests from different threads. However, such multi-threading comes at the cost of maintaining complex metadata that is tightly coupled and intertwined with user data. When memory management functions and other user programs run on the same core, the metadata used by management functions may pollute the processor caches and other resources. Ruihao Li 0002, Qinzhe Wu, Krishna M. Kavi, Gayatri Mehta, Neeraja J. Yadwadkar, Lizy Kurian John |
HotOS | 2 |
| 2023 | Spectral Feature Learning for Anomaly Change Detection in Hyperspectral ImageabstractRecently, anomaly change detection has attracted considerable interest in hyperspectral image (HSI) analysis. Hyperspectral anomaly change detection (HACD) focuses more on detecting anomalies between multi-temporal hyperspectral images. This paper presents a novel anomaly change detection algorithm that has the following innovations. First, using the variational autoencoder to reconstruct the background of input, forcing its latent features to obey the Gaussian distribution, which can help to reconstruct the background better. Second, in order to obtain a better quality reconstruction of the background map, the spectral angular distance is introduced as a model constraint. Third, impose constraints on the encoder to better learn the background information of the input and improve the reconstruction effect of the network. The proposed algorithm has superior performance over some other algorithms on "Viareggio 2013" datasets. Wen Ren, Qinzhe Wu, Hongyue Sun |
IGARSS | 3 |
| 2023 | Hyperspectral Image Classification Via Multi-Scale Residual Attention NetworkabstractThe application of convolutional neural network (CNN) in the field of hyperspectral image (HSI) classification has been pervasive. Since hyperspectral images contain numerous complicated spectral-spatial information, only using a single-channel CNN is difficult to fully extract information from the images. As the increasing number of network layers, the network classification accuracy will decrease, and the number of parameters will increase greatly. In order to avoid the above defects, we propose a multi-scale residual attention network to extract features from HSI more efficiently. In addition, the proposed network lowers the number of parameters by introducing depthwise separable convolution (DSC). Experimental studies on two commonly used hyperspectral image datasets verify that the classification performance of the proposed network outperforms some state-of-the-art networks. Qinzhe Wu, Wen Ren |
IGARSS | 2 |
| 2023 | Hyperspectral Image Classification Via 3D-CNN MHSA Fusion TransformerabstractAlthough Convolutional Neural Network (CNN) and Vision Transformer (ViT) have excellent performance in hyperspectral image (HSI) classification. However, due to the inherent network limitations, CNN cannot fully mine the spectral feature information well in HSI, and ViT could not effective extract the local spatial feature of HSI. In order to solve the above problems, we propose a new network which is 3D-CNN Multi-Head Self-Attention Fusion Transformer (3DMFT), which is combined with Transformer and CNN. 3DMFT fuses 3D-CNN and QKV to learn the deep and shallow features in HSI. Moreover, the local spatial-spectral position encoding obtains spatial-spectral position information between elements , and then induct pyramid model to transfer image feature from shallow layer to deep layer. Experimental results show that 3DMFT can obtain global context dependencies and local delicate feature well. Compared with some state-of-the-art methods, the proposed 3DMFT network is more efficient. Hongyue Sun, Qinzhe Wu |
IGARSS | 4 |
| 2022 | SPAMeR: Speculative Push for Anticipated Message Requests in Multi-Core SystemsabstractWith increasing core counts and multiple levels of cache memories, scaling multi-threaded and task-level parallel workloads is continuously becoming a challenge. A key challenge to scaling the number of communicating tasks (or threads) is the rate at which existing communication mechanisms scale (in terms of latency and bandwidth). Architectures with hardware accelerated queuing operations have the potential to reduce the latency and improve scalability of moving data between processing elements, reducing synchronization penalties, and thereby improving the performance of task-level parallel workloads. While hardware queues reduce synchronization penalties, they cannot fully hide load-to-use latency, i.e., perfect pipelines often are not realized. There is the potential, however, for better overlap. If the inter-processor communication latency is equal to or less than the time spent processing a message at the consumer, any and all latency may be overlapped while the consumer is processing. We exploit this property to speedup parallel applications above and beyond existing hardware queues. Qinzhe Wu, Ashen Ekanayake, Ruihao Li 0002, Jonathan C. Beard, Lizy Kurian John |
ICPP | 1 |
| 2021 | Virtual-Link: A Scalable Multi-Producer Multi-Consumer Message Queue Architecture for Cross-Core CommunicationabstractCross-core communication is increasingly a bottleneck as the number of processing elements increase per system-on-chip. Typical hardware solutions to cross-core communication are often inflexible; while software solutions are flexible, they have performance scaling limitations. A key problem, as we will show, is that of shared state in software-based message queue mechanisms. This paper proposes Virtual-Link (VL), a novel light-weight communication mechanism with hardware support to facilitate M:N lock-free data movement. VL reduces the amount of coherent shared state, which is a bottleneck for many approaches, to zero. VL provides further latency benefit by keeping data on the fast path (i.e., within the onchip interconnect). VL enables directed cache-injection (stashing) between PEs on the coherence bus, reducing the latency for core-to-core communication. VL is particularly effective for fine-grain tasks on streaming data. Evaluation on a full system simulator with 7 benchmarks shows that VL achieves a 2.09x speedup over state-of-the-art software-based communication mechanisms, while reducing memory traffic by 61%. Qinzhe Wu, Jonathan Beard, Ashen Ekanayake, Andreas Gerstlauer, Lizy Kurian John |
IPDPS | 1 |
| 2020 | Accelerating Force-directed Graph Layout with Processing-in-Memory ArchitectureabstractIn the big data domain, the visualization of graph systems provides users more intuitive experiences, especially in the field of social networks, transportation systems, and even medical and biological domains. Processing-in-Memory (PIM) has been a popular choice for deploying emerging applications as a result of its high parallelism and low energy consumption. Furthermore, memory cells of PIM platforms can serve as both compute units and storage units, making PIM solutions able to efficiently support visualizing graphs at different scales. In this paper, we focus on using the PIM platform to accelerate the Force-directed Graph Layout (FdGL) algorithm, which is one of the most fundamental algorithms in the field of visualization. We fully explore the parallelism inside the FdGL algorithm and integrate an algorithm level optimization strategy into our PIM system. In addition, we use programmable instruction sets to achieve more flexibility in our PIM system. Our PIM architecture can achieve 8.07× speedup compared with a GPU platform of the same peak throughput. Compared with state-of-the-art CPU and GPU platforms, our PIM system can achieve an average of 13.33× and 2.14× performance speedup with 74.51× and 14.30× energy consumption reduction on six real world graphs. Ruihao Li 0002, Shuang Song 0007, Qinzhe Wu, Lizy Kurian John |
HiPC | 3 |
| 2020 | Demystifying the MLPerf Training Benchmark SuiteabstractMLPerf, an emerging machine learning benchmark suite, strives to cover a broad range of machine learning applications. We present a study on the characteristics of MLPerf benchmarks and how they differ from previous deep learning benchmarks such as DAWNBench and DeepBench. MLPerf benchmarks are seen to exhibit moderately high memory transactions per second and moderately high compute rates, while DAWNBench creates a high-compute benchmark with low memory transaction rate, and DeepBench provides low compute rate benchmarks. We also observe that the various MLPerf benchmarks possess unique features that allow unveiling various bottlenecks in systems. We also observe variation in scaling efficiency across the MLPerf models. The variation exhibited by the different models highlight the importance of smart scheduling strategies for multi-GPU training. Another observation is that dedicated low latency interconnect between GPUs in multi-GPU systems is crucial for optimal distributed deep learning training. Furthermore, host CPU utilization increases with an increase in the number of GPUs used for training. Corroborating prior work, we also observe and quantify improvements possible by mixed-precision training using Tensor Cores. Snehil Verma, Qinzhe Wu, Bagus Hanindhito, Gunjan Jha, Eugene John, Ramesh Radhakrishnan, Lizy Kurian John |
ISPASS | 2 |
| 2020 | Demystifying graph processing frameworks and benchmarks
Junyong Deng, Qinzhe Wu, Shuang Song 0007, Joseph Dean, Lizy Kurian John |
Sci. China Inf. Sci. | 2 |
| 2019 | A Study of Core Utilization and Residency in Heterogeneous Smart Phone ArchitecturesabstractIn recent years, the smart phone platform has seen a rise in the number of cores and the use of heterogeneous clusters as in the Qualcomm Snapdragon, Apple A10 and the Samsung Exynos processors. This paper attempts to understand characteristics of mobile workloads, with measurements on heterogeneous multicore phone platforms with big and little cores. It answers questions such as the following: (i) Do smart phones need multiple cores of different types (eg: big or little)? (ii) Is it energy-efficient to operate with more cores (with less time) or fewer cores even if it might take longer? (iii)What are the best frequencies to operate the cores considering energy efficiency? (iv) Do mobile applications need out-of-order speculative execution cores with complex branch prediction? (v) Is IPC a good performance indicator for early design tradeoff evaluation while working on mobile processor design? Joseph Whitehouse, Qinzhe Wu, Shuang Song 0007, Eugene John, Andreas Gerstlauer, Lizy Kurian John |
ICPE | 2 |
| 2018 | Start Late or Finish Early: A Distributed Graph Processing System with Redundancy ReductionabstractGraph processing systems are important in the big data domain. However, processing graphs in parallel often introduces redundant computations in existing algorithms and models. Prior work has proposed techniques to optimize redundancies for out-of-core graph systems, rather than distributed graph systems. In this paper, we study various state-of-the-art distributed graph systems and observe root causes for these pervasively existing redundancies. To reduce redundancies without sacrificing parallelism, we further propose SLFE, a distributed graph processing system, designed with the principle of "start late or finish early". SLFE employs a novel preprocessing stage to obtain a graph's topological knowledge with negligible overhead. SLFE's redundancy-aware vertex-centric computation model can then utilize such knowledge to reduce the redundant computations at runtime. SLFE also provides a set of APIs to improve programmability. Our experiments on an 8-machine high-performance cluster show that SLFE outperforms all well-known distributed graph processing systems with the inputs of real-world graphs, yielding up to 75x speedup. Moreover, SLFE outperforms two state-of-the-art shared memory graph systems on a high-end machine with up to 1644x speedup. SLFE's redundancy-reduction schemes are generally applicable to other vertex-centric graph processing systems. Shuang Song 0007, Xu Liu 0001, Qinzhe Wu, Andreas Gerstlauer, Tao Li 0006, Lizy Kurian John |
Proc. VLDB Endow. | 3 |