VLDB 2026 Research / reviewers in the wild / expert
Mao Lin
dblp:62/10619
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PASTA: A Modular Program Analysis Tool Framework for AcceleratorsabstractThe increasing complexity and diversity of hardware accelerators in modern computing systems demand flexible, low-overhead program analysis tools. We present PASTA, a low-overhead and modular Program AnalysiS Tool Framework for Accelerators. PASTA abstracts over low-level profiling APIs and diverse deep learning frameworks, offering users a unified interface to capture and analyze runtime events at multiple levels. Its extensible design enables researchers and practitioners to rapidly prototype custom tools with minimal overhead. We demonstrate the utility of PASTA by developing several analysis tools, including tools for deep learning workload characterization and UVM optimization. Through extensive evaluations on mainstream deep learning workloads tested on NVIDIA and AMD GPUs under both single- and multi-GPU scenarios, we demonstrate PASTA’s broad applicability. On NVIDIA GPUs, we further show that PASTA provides detailed performance insights with significantly lower overhead (up to 1.3×104faster) than conventional analysis tools, thanks to its GPU-accelerated backend. PASTA strikes a practical balance between usability, extensibility, and efficiency, making it well-suited for modern accelerator-based computing environments. Mao Lin, Hyeran Jeon, Keren Zhou 0001 |
CGO | 1 |
| 2025 | Forest: Access-aware GPU UVM Management
Mao Lin, Guilherme Cox, Hyeran Jeon |
ISCA | 1 |
| 2023 | DrGPUM: Guiding Memory Optimization for GPU-Accelerated ApplicationsabstractGPUs are widely used in today’s computing platforms to accelerate applications in various domains. However, scarce GPU memory resources are often the dominant limiting factor in strengthening the applicability of GPU computing. In this paper, we propose DrGPUM, the first profiler that systematically investigates patterns of memory inefficiencies in GPU-accelerated applications. The strength of DrGPUM, when compared to a large class of existing GPU profilers, is its ability to (1) correlate problematic memory usage with data objects and GPU APIs, (2) identify and categorize object-level and intra-object memory inefficiencies, and (3) provide rich insights to guide memory optimization. Mao Lin, Keren Zhou 0001, Pengfei Su 0001 |
ASPLOS (3) | 1 |
| 2023 | Ecological network evolution analysis in collective intelligence design ecosystem
Zhong-Lin Fu, Wei Guo 0032, Lei Wang 0189, Li-Wen Shi, Mao Lin |
Adv. Eng. Informatics | 6 |
| 2023 | A Comprehensive Memory Management Framework for CPU-FPGA Heterogenous SoCsabstractEfficient utilization of restrained memory resources is of paramount importance in CPU-FPGA heterogeneous multiprocessor system-on-chip (HMPSoC)-based system design for memory-intensive applications. State-of-the-art high level synthesis (HLS) tools rely on the system programmers to manually determine the data placement within the complex memory hierarchy. Different data placement policies may lead to different system performance, and finding an optimal data placement policy is a nontrivial problem. For instance, we show counter-intuitive results that traditional frequency and locality-based data placement strategy designed for CPU architecture leads to nonoptimal system performance in CPU-FPGA HMPSoCs. In this work, we first propose an automatic data placement framework for field programmable gate array (FPGA) kernels to determine whether each array object should be accessed via the on-chip BRAM, shared CPU L2-cache, or DDR memory to achieve the optimal performance. Moreover, we find that when the CPU kernel and the FPGA kernel are executed in parallel, memory contentions may degrade the performance and the optimal data placement policy designed for the FPGA kernel alone will not achieve the optimal overall system performance. In this article, we proposed to use cache partitioning to alleviate the impact brought by memory contentions. We extend the framework designed for FPGA by adding the cross-layer memory contentions analysis to automatically generate an optimal data placement policy and cache partitioning mechanism for the parallel executing kernels. The proposed data placement framework can be seamlessly integrated with the commercial Vivado HLS. The experimental results on the Zedboard platform show an average$1.5\times $performance speedup for FPGA kernels compared with a greedy-based allocation strategy. When FPGA kernels and CPU kernels are executed in parallel, the FPGA kernel and the CPU kernel have a performance speedup of$1.62\times $and$1.10\times $on average, respectively. Zelin Du, Qianling Zhang, Mao Lin, Shiqing Li, Xin Li 0137, Lei Ju 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Formal Interoperability Models of Sensor Networks Based on Logical Workflow NetsabstractWith the advent of Internet of things and cyber physical system, sensor networks become more and more important. To better accomplish a task, multiple sensors or even multiple sensor networks need to interoperate. Logical workflow nets and cooperative logical workflow nets are introduced to formally model and analyze interoperability of sensor networks. Independent feasibility and Interoperable feasibility are important properties for ensuring correct execution and interoperability of sensor networks. Complete path nets, possible path nets, cooperative complete path nets and cooperative possible path nets are presented to decide independent feasibility and interoperable feasibility of logical workflow nets and cooperative logical workflow nets denoting interoperability of sensor networks. Wei Liu 0051, Mao Lin, Chun Yan |
Int. J. Softw. Eng. Knowl. Eng. | 2 |