VLDB 2026 Research / reviewers in the wild / expert
Mengchi Zhang
dblp:220/1698
· DBLP profile ↗
7ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 first-author · 2 since 2021Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | InfecBlock: Investigating the Effects of a Tower-Defense Serious Game for Increasing Epidemic-Related Health LiteracyabstractSerious game can potentially improve social awareness and health literacy related to the epidemic, where interactivity, such as strategic game elements, could play a crucial role in increasing learning engagement and motivation. In this paper, we present the user study of InfecBlock, a tower-defense game designed to facilitate individual users to acquire public health knowledge related to coronavirus disease and epidemic prevention. We employed a between-subject experiment design and collected a variety of quantitative data to examine players’ learning outcomes, engagement, and emotional responses. Our results confirmed the effectiveness of InfecBlock in improving learning performance and highlighted its potential to facilitate a more engaged and enjoyable learning experience. We discussed a set of implications for designing tower-defense serious games for supporting the improvements of public health literacy. Xiaoqing Sun, Kexin Miao, Mengchi Zhang, Xipei Ren |
Int. J. Hum. Comput. Interact. | 4 |
| 2024 | Concurrency-Aware Register Stacks for Efficient GPU Function CallsabstractSince the early days of computers, dividing a program into functions or subroutines has been a common way to manage complexity. Functions make programs easier to read, facilitate code reuse, and provide clean interfaces for separate compilation. However, function calls incur runtime overhead. We quantify the impact of this runtime overhead on GPUs and demonstrate that the register spills/fills required to maintain the function call application binary interface place significant bandwidth and capacity pressure on shared resources. To alleviate this overhead, we introduce Concurrency-Aware Register Stacks (CARS), a hardware mechanism that re-purposes segments of the GPU register file as a software-controlled hardware stack. CARS exploits the regularity in function prologue/epilogues to rename registers pushed to the stack with linear base + offset addressing, similar to the baseline GPU. Informed by lightweight call graph analysis and dynamic function behavior, CARS balances the space devoted to register stacks with the concurrency required to hide latency in GPUs. Without harming function-free programs, CARS improves the performance and energy efficiency of 22 function-calling applications by 26% and 28%, respectively, outperforming idealized GPUs with impractical resources. Ni Kang, Ahmad Alawneh, Mengchi Zhang, Timothy G. Rogers |
MICRO | 3 |
| 2021 | Judging a type by its pointer: optimizing GPU virtual functionsabstractProgrammable accelerators aim to provide the flexibility of traditional CPUs with significantly improved performance. A well-known impediment to the widespread adoption of programmable accelerators, like GPUs, is the software engineering overhead involved in porting the code. Existing support for C++ on GPUs allows programmers to port polymorphic code with little effort. However, the overhead from the virtual functions introduced by polymorphic code has not been well studied or mitigated on GPUs. Mengchi Zhang, Ahmad Alawneh, Timothy G. Rogers |
ASPLOS | 1 |
| 2021 | Characterizing Massively Parallel PolymorphismabstractGPU computing has matured to include advanced C++ programming features. As a result, complex applications can potentially benefit from the continued performance improvements made to contemporary GPUs with each new generation. Tighter integration between the CPU and GPU, including a shared virtual memory space, increases the usability of productive programming paradigms traditionally reserved for CPUs, like object-oriented programming. Programmers are no longer forced to restructure both their code and data for GPU acceleration. However, the implementation and performance implications of advanced C++ on massively multithreaded accelerators have not been well studied. In this paper, we study the effects of runtime polymorphism on GPUs. We first detail the implementation of virtual function calls in contemporary GPUs using microbenchmarking. We then propose Parapoly, the first open-source polymorphic GPU benchmark suite. Using Parapoly, we further characterize the overhead caused by executing dynamic dispatch on GPUs using massively scaled CPU workloads. Our characterization demonstrates that the optimization space for runtime polymorphism on GPUs is fundamentally different than for CPUs. Where indirect branch prediction and ILP extraction strategies have dominated the work on CPU polymorphism, GPUs are fundamentally limited by excessive memory system contention caused by virtual function lookup and register spilling. Using the results of our study, we enumerate several pitfalls when writing polymorphic code for GPUs and suggest several new areas of system and architecture research that can help alleviate overhead. Mengchi Zhang, Ahmad Alawneh, Timothy G. Rogers |
ISPASS | 1 |
| 2019 | POSTER: Quantifying the Direct Overhead of Virtual Function Calls on Massively Parallel ArchitecturesabstractProgrammable accelerators aim to provide the flexibility of traditional CPUs, with greatly improved performance and energy-efficiency. Arguably, the greatest impediment to the widespread adoption of programmable accelerators, like GPUs, is the software engineering overhead involved in getting the code to execute correctly. To help combat this issue, GPGPU computing has matured from its origins as a collection of graphics API hacks to include advanced programming features, including object-oriented programming. This level of support, in combination with a shared virtual memory space between the CPU and GPU, make it possible for rich object-oriented frameworks to execute on GPUs with little porting effort. However, executing this type of flexible code on a massively parallel accelerator introduces overhead that has not been well studied. In this poster, we analyze the direct overhead of virtual function calls on contemporary GPUs. Using the latest GPU architectures and compilers, this poster performs the analysis of how virtual function calls are implemented on GPUs. We quantify the direct overhead incurred from contemporary implementations and show that the massively multithreaded nature of GPUs creates deficiencies and opportunities not found in CPU implementations of virtual function calls. Mengchi Zhang, Roland N. Green, Timothy G. Rogers |
PACT | 1 |
| 2019 | Analyzing Machine Learning Workloads Using a Detailed GPU SimulatorabstractMachine learning (ML) has recently emerged as an important application driving future architecture design. Traditionally, architecture research has used detailed simulators to model and measure the impact of proposed changes. However, current open-source, publicly available simulators lack support for running a full ML stack like PyTorch. High-confidence, cycle-accurate simulations are crucial for architecture research and without them, it is difficult to rapidly prototype new ideas. In this paper, we describe changes we made to GPGPU-Sim, a popular, widely used GPU simulator, to run ML applications that use cuDNN and PyTorch, two widely used frameworks for running Deep Neural Networks (DNNs). This work has the potential to enable significant microarchitectural research into GPUs for DNNs. Our results show that the modified simulator, which has been made publicly available with this paper1Source code available at https://github.com/gpgpu-sim/gpgpu-sim_distribution (dev branch), provides execution time results within 18% of real hardware. We further use it to study other ML workloads and demonstrate how the simulator identifies opportunities for architectural optimization that prior tools are unable to provide. Jonathan S. Lew, Deval Shah, Suchita Pati, Shaylin Cattell, Mengchi Zhang, Amruth Sandhupatla, Christopher Ng, Negar Goli, Matthew D. Sinclair, Timothy G. Rogers, Tor M. Aamodt |
ISPASS | 5 |
| 2018 | Characterizing the Runtime Effects of Object-Oriented Workloads on GPUsabstractModern GPGPU programming extensions like OpenCL and CUDA have supported object-oriented workloads on GPUs for several generations. However, no analysis of object-oriented workloads running on massively parallel accelerators has been investigated. This extended abstract presents a performance analysis of object-oriented workloads on a PASCAL Titan X GPU. Our characterization demonstrates that GPUs have different performance trade-offs when running object-oriented code than traditional CPUs. Where CPUs are sensitive to the misprediction of indirect branches that result from virtual function calls, GPUs are more sensitive to the additional memory system pressure that comes from loading pointers and virtual function table entries. Mengchi Zhang, Roland N. Green, Timothy G. Rogers |
ISPASS | 1 |