VLDB 2026 Research / reviewers in the wild / expert
Pengfei Su 0001
dblp:80/2828-1
· DBLP profile ↗
11ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0001-7035-1998ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRACE4J: A Lightweight, Flexible, and Insightful Performance Tracing Tool for JavaabstractJava is often considered a superior programming language choice owing to its high portability, strong memory safety, and rapid development cycle. However, this superiority comes with increased complexity within Java software stacks, driven by the extensive use of layered libraries, rising levels of abstraction, and the combination of interpretation and just-in-time compilation. This complexity disjoins source code and its execution details on the underlying hardware, making it challenging to write efficient Java code. Performance tracing is key to bridging this gap by providing detailed, temporally ordered insights into a program’s runtime behavior. Existing tracing approaches generally fall into two categories: (1) instrumentation, which enjoys high accuracy but incurs significant overhead, and (2) sampling, which enjoys low overhead but sacrifices accuracy.We introduce TRACE4J, a novel performance tracing tool for Java that overcomes the limitations of existing approaches. TRACE4J intelligently integrates CPU hardware facilities (performance monitoring units and breakpoints), the JVM tool interface, the Linux perf_event interface, and instruction decoding to deliver lightweight, flexible, and insightful performance tracing. It applies to unmodified Java programs, runs on standard JVMs and commodity CPUs, and provides both end-to-end and on-demand tracing, making it suitable for production environments. Through evaluation, we demonstrate TRACE4J’s ability to deliver actionable performance insights with low overhead (no more than 5% time and memory impact). Using these insights, we were able to optimize several Java benchmarks and real-world applications, achieving substantial performance gains. Haide He, Pengfei Su 0001 |
CGO | 2 |
| 2026 | TenProf: A Tensor-Centric Profiler for Deep Learning Workload Analysis and Optimization
Xingjian Ding, Keren Zhou 0001, Yueming Hao, Pengfei Su 0001 |
ICS | 4 |
| 2024 | Centimani: Enabling Fast AI Accelerator Selection for DNN Training with a Novel Performance Predictor
Murali Emani, Xiaodong Yu 0001, Dingwen Tao, Xin He 0054, Pengfei Su 0001, Keren Zhou 0001, Venkatram Vishwanath |
USENIX ATC | 6 |
| 2023 | DrGPUM: Guiding Memory Optimization for GPU-Accelerated ApplicationsabstractGPUs are widely used in today’s computing platforms to accelerate applications in various domains. However, scarce GPU memory resources are often the dominant limiting factor in strengthening the applicability of GPU computing. In this paper, we propose DrGPUM, the first profiler that systematically investigates patterns of memory inefficiencies in GPU-accelerated applications. The strength of DrGPUM, when compared to a large class of existing GPU profilers, is its ability to (1) correlate problematic memory usage with data objects and GPU APIs, (2) identify and categorize object-level and intra-object memory inefficiencies, and (3) provide rich insights to guide memory optimization. Mao Lin, Keren Zhou 0001, Pengfei Su 0001 |
ASPLOS (3) | 3 |
| 2023 | DJXPerf: Identifying Memory Inefficiencies via Object-Centric Profiling for JavaabstractJava is the “go-to” programming language choice for developing scalable enterprise cloud applications. In such systems, even a few percent CPU time savings can offer a significant competitive advantage and cost savings. Although performance tools abound for Java, those that focus on the data locality in the memory hierarchy are rare. Pengfei Su 0001, Milind Chabbi, Shuyin Jiao, Xu Liu 0001 |
CGO | 2 |
| 2023 | MicroProf: Code-level Attribution of Unnecessary Data Transfer in Microservice ApplicationsabstractThe microservice architecture style has gained popularity due to its ability to fault isolation, ease of scaling applications, and developer’s agility. However, writing applications in the microservice design style has its challenges. Due to the loosely coupled nature, services communicate with others through standard communication APIs. This incurs significant overhead in the application due to communication protocol and data transformation. An inefficient service communication at the microservice application logic can further overwhelm the application. We perform a grey literature review showing that unnecessary data transfer is a real challenge in the industry. To the best of our knowledge, no effective tool is currently available to accurately identify the origins of unnecessary microservice communications that lead to significant performance overhead and provide guidance for optimization. To bridge the knowledge gap, we propose MicroProf , a dynamic program analysis tool to detect unnecessary data transfer in Java-based microservice applications. At the implementation level, MicroProf proposes novel techniques such as remote object sampling and hardware debug registers to monitor remote object usage. MicroProf reports the unnecessary data transfer at the application source code level. Furthermore, MicroProf pinpoints the opportunities for communication API optimization. MicroProf is evaluated on four well-known applications involving two real-world applications and two benchmarks, identifying five inefficient remote invocations. Guided by MicroProf , API optimization achieves an 87.5% reduction in the number of fields within REST API responses. The empirical evaluation further reveals that the optimized services experience a speedup of up to 4.59×. Syed Salauddin Mohammad Tariq, Lance Menard, Pengfei Su 0001, Probir Roy |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | OJXPERF: Featherlight Object Replica Detection for Java ProgramsabstractMemory bloat is an important source of inefficiency in complex production software, especially in software written in managed languages such as Java. Prior approaches to this problem have focused on identifying objects that outlive their life span. Few studies have, however, looked into whether and to what extent myriad objects of the same type are identical. A quantitative assessment of identical objects with code-level attribution can assist developers in refactoring code to eliminate object bloat, and favor reuse of existing object(s). The result is reduced memory pressure, reduced allocation and garbage collection, enhanced data locality, and reduced re-computation, all of which result in superior performance. Hao Xu 0048, Qidong Zhao, Pengfei Su 0001, Milind Chabbi, Shuyin Jiao, Xu Liu 0001 |
ICSE | 4 |
| 2019 | Redundant loads: a software inefficiency indicatorabstractModern software packages have become increasingly complex with millions of lines of code and references to many external libraries. Redundant operations are a common performance limiter in these code bases. Missed compiler optimization opportunities, inappropriate data structure and algorithm choices, and developers' inattention to performance are some common reasons for the existence of redundant operations. Developers mainly depend on compilers to eliminate redundant operations. However, compilers' static analysis often misses optimization opportunities due to ambiguities and limited analysis scope; automatic optimizations to algorithmic and data structural problems are out of scope. We develop LoadSpy, a whole-program profiler to pinpoint redundant memory load operations, which are often a symptom of many redundant operations. The strength of LoadSpy exists in identifying and quantifying redundant load operations in programs and associating the redundancies with program execution contexts and scopes to focus developers' attention on problematic code. LoadSpy works on fully optimized binaries, adopts various optimization techniques to reduce its overhead, and provides a rich graphic user interface, which make it a complete developer tool. Applying LoadSpy showed that a large fraction of redundant loads is common in modern software packages despite highest levels of automatic compiler optimizations. Guided by LoadSpy, we optimize several well-known benchmarks and real-world applications, yielding significant speedups. Pengfei Su 0001, Shasha Wen, Hailong Yang 0002, Milind Chabbi, Xu Liu 0001 |
ICSE | 1 |
| 2019 | Lightweight hardware transactional memory profilingabstractPrograms that use hardware transactional memory (HTM) demand sophisticated performance analysis tools when they suffer from performance losses. We have developed TxSampler---a lightweight profiler for programs that use HTM. TxSampler measures performance via sampling and provides a structured performance analysis to guide intuitive optimization with a novel decision-tree model. TxSampler computes metrics that drive the investigation process in a systematic way. It not only pinpoints hot transactions with time quantification of transactional and fallback paths, but also identifies causes of transaction aborts such as data contention, capacity overflow, false sharing, and problematic instructions. TxSampler associates metrics with full call paths that are even deeply embedded inside transactions and maps them to the program's source code. Our evaluation of more than 30 HTM benchmarks and applications shows that TxSampler incurs ~4% runtime overhead and negligible memory overhead for its insightful analyses. Guided by TxSampler, we are able to optimize several HTM programs and obtain nontrivial speedups. Qingsen Wang, Pengfei Su 0001, Milind Chabbi, Xu Liu 0001 |
PPoPP | 2 |
| 2019 | Pinpointing performance inefficiencies via lightweight variance profilingabstractExecution variance among different invocation instances of the same procedure is often an indicator of performance losses. On the one hand, instrumentation-based tools can insert calipers around procedures and identify execution variance; however, they can introduce high overheads. On the other hand, sampling-based tools insert no instrumentation and have low overheads; however, they cannot synchronize samples with procedure entry and exit. Pengfei Su 0001, Shuyin Jiao, Milind Chabbi, Xu Liu 0001 |
SC | 1 |
| 2019 | Pinpointing performance inefficiencies in JavaabstractMany performance inefficiencies such as inappropriate choice of algorithms or data structures, developers' inattention to performance, and missed compiler optimizations show up as wasteful memory operations. Wasteful memory operations are those that produce/consume data to/from memory that may have been avoided. We present, JXPerf, a lightweight performance analysis tool for pinpointing wasteful memory operations in Java programs. Traditional byte code instrumentation for such analysis (1) introduces prohibitive overheads and (2) misses inefficiencies in machine code generation. JXPerf overcomes both of these problems. JXPerf uses hardware performance monitoring units to sample memory locations accessed by a program and uses hardware debug registers to monitor subsequent accesses to the same memory. The result is a lightweight measurement at the machine code level with attribution of inefficiencies to their provenance --- machine and source code within full calling contexts. JXPerf introduces only 7% runtime overhead and 7% memory overhead making it useful in production. Guided by JXPerf, we optimize several Java applications by improving code generation and choosing superior data structures and algorithms, which yield significant speedups. Pengfei Su 0001, Qingsen Wang, Milind Chabbi, Xu Liu 0001 |
ESEC/SIGSOFT FSE | 1 |