VLDB 2026 Research / reviewers in the wild / expert
Pengbo Nie
dblp:278/0656
· DBLP profile ↗
9ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-3759-5242ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JOSer: Just-In-Time Object Serialization for Heavy Java Serialization WorkloadsabstractObject serialization is critical in Java, which preserves objects in memory and transfers them among software systems if needed. However, the serialization techniques of modern Java systems are usually inflexible and inefficient under heavy serialization workloads, as they rely on manually-defined object schemas or omni-functional serializers. To tackle the above problem, we reveal a novel, serialization-specific optimization opportunity in Java. Based on it, we develop JOSer (Just-in-time Object SERializer), an efficient, Just-in-Time (JIT) object serialization technique. At runtime, JOSer generates a set of class-specific, JIT-friendly object serializers (i.e., serialization code), and then continuously optimizes them with the JIT compiler of Java Virtual Machine (JVM). JOSer also shares the metadata of objects under serialization and the serializers under optimization. We evaluate JOSer against six Java serialization techniques including OpenJDK's built-in serialization technique. Overall, JOSer improves the throughput by up to 20~83× in serialization and 43~229× in deserialization. JOSer has been successfully deployed in real-world products, reducing serialization CPU usage of Flink by 35.32~41.14% and latency of search recommendations by 30+ ms. JOSer grounds Apache Fory#8482;, an open-source serialization framework available at https://fory.apache.org/. Chaokun Yang, Pengbo Nie, Qianwei Yu, Chengcheng Wan 0001, He Jiang 0001, Yuting Chen 0001 |
ASPLOS (2) | 2 |
| 2026 | Orchestrating optimization passes of machine learning compiler for reducing memory footprints of computation graphs
Qianwei Yu, Pengbo Nie, Chengcheng Wan 0001, He Jiang 0001, Jianjun Zhao 0001, Lei Qiao 0002, Yuting Chen 0001 |
J. Syst. Archit. | 2 |
| 2024 | OPASS: Orchestrating TVM's Passes for Lowering Memory Footprints of Computation GraphsabstractDeep learning (DL) compilers, such as TVM and TensorFlow, encompass a variety of passes for optimizing computation graphs (i.e., DL models). Despite the efforts on developing optimization passes, it remains a challenge in arranging these passes - most compilers employ fixed pass sequences that do not fit with computation graphs of diverse structures; on the other hand, optimization passes have cascade effects, making the structures of graphs under compilation volatile and as well making it difficult to generate optimal sequences for graphs. Inspired by recent progresses on static computing memory footprints (i.e., memory usages) of computation graphs, we introduce in this paper OPASS, a novel approach to orchestrating TVM's optimization passes for lowering memory footprints of computation graphs, and finally allowing the graphs to run on memory-constrained devices. The key idea is, given a computation graph$G$, to optimize the graph heuristically and iteratively: OPASS learns the effects of passes on the graph; it then optimizes$G$iteratively - each iteration picks up a pass by the reduction of the memory footprint of$G$and as well the implicit effects of the pass for further optimizations, letting the pass be applied. We evaluate OPASS on Rebench (a suite of computation graphs) and two real-world models (Transformer and ResNet). The results clearly show the strength of OPASS: it outperforms TVM's default sequence by$1.77\times$in reducing graphs' memory footprints, with affordable costs; it also offers extra memory reductions of$5\sim 12\%$by catching the implicit effects of passes. Furthermore, OPASS helps analyze positive/negative effects of passes to graphs' memory footprints, providing TVM developers with best practices for designing optimization pass sequences. Pengbo Nie, Chengcheng Wan 0001, He Jiang 0001, Jianjun Zhao 0001, Yuting Chen 0001 |
ICSME | 1 |
| 2024 | Detecting Numerical Deviations in Deep Learning Models Introduced by the TVM CompilerabstractDeep learning (DL) compilers are crucial for deploying DL models and speeding up their inferences. Meanwhile, they may introduce numerical deviations, and finally undefined or unexpected behaviours, into DL models. Many efforts have been spent on studying DL compilers’ logic bugs, whilst researchers often overlook numerical deviations introduced by DL compilers. This paper studies hazards and root causes of numerical deviations introduced by Apache’s TVM, a state-of-the-art, open-sourced DL compiler. This paper further proposes TracNe, an approach composed of an MEGA searcher and a tracer, to reveal and isolate compiler-introduced numerical deviations. Given a DL model, the MEGA searcher searches for deviation-triggering inputs and checks whether the model suffers from numerical deviations. The tracer performs a semantic-based matching between the models before and after compilation, isolating an erroneous scope in the compiled model. We evaluate TracNe on 60 synthesis and 9 industrial-edge models. The results show that TracNe reveals 5.6× more deviation-prone models than two typical search algorithms (MCMC and DEMC); it also localizes 64% more deviations than PLiner, a state-of-the-art isolation technique, while reducing 94.6% of isolation time of Pliner. Zichao Xia, Pengbo Nie |
ISSRE | 3 |
| 2023 | Lejacon: A Lightweight and Efficient Approach to Java Confidential Computing on SGXabstractIntel's SGX is a confidential computing technique. It allows key functionalities of C/C++/native applications to be confidentially executed in hardware enclaves. However, numerous cloud applications are written in Java. For supporting their confidential computing, state-of-the-art approaches deploy Java Virtual Machines (JVMs) in enclaves and perform confidential computing on JVMs. Meanwhile, these JVM-in-enclave solutions still suffer from serious limitations, such as heavy overheads of running JVMs in enclaves, large attack surfaces, and deep computation stacks. To mitigate the above limitations, we for-malize a Secure Closed-World (SCW) principle and then propose Lejacon, a lightweight and efficient approach to Java confidential computing. The key idea is, given a Java application, to (1) separately compile its confidential computing tasks into a bundle of Native Confidential Computing (NCC) services; (2) run the NCC services in enclaves on the Trusted Execution Environment (TEE) side, and meanwhile run the non-confidential code on a JVM on the Rich Execution Environment (REE) side. The two sides interact with each other, protecting confidential computing tasks and as well keeping the Trusted Computing Base (TCB) size small. We implement Lejacon and evaluate it against OcclumJ (a state-of-the-art JVM-in-enclave solution) on a set of benchmarks using the BouncyCastle cryptography library. The evaluation results clearly show the strengths of Lejacon: it achieves compet-itive performance in running Java confidential code in enclaves; compared with OcclumJ, Lejacon achieves speedups by up to 16.2x in running confidential code and also reduces the TCB sizes by 90+% on average. Xinyuan Miao, Sanhong Li, Pengbo Nie, Yuting Chen 0001, Beijun Shen, He Jiang 0001 |
ICSE | 7 |
| 2023 | GenCoG: A DSL-Based Approach to Generating Computation Graphs for TVM TestingabstractTVM is a popular deep learning (DL) compiler. It is designed for compiling DL models, which are naturally computation graphs, and as well promoting the efficiency of DL computation. State-of-the-art methods, such as Muffin and NNSmith, allow developers to generate computation graphs for testing DL compilers. However, these techniques are inefficient — their generated computation graphs are either type-invalid or inexpressive, and hence not able to test the core functionalities of a DL compiler. Pengbo Nie, Xinyuan Miao, Yuting Chen 0001, Chengcheng Wan 0001, Lei Bu, Jianjun Zhao 0001 |
ISSTA | 2 |
| 2023 | Coverage-directed Differential Testing of X.509 Certificate Validation in SSL/TLS ImplementationsabstractSecure Sockets Layer (SSL) and Transport Security (TLS) are two secure protocols for creating secure connections over the Internet. X.509 certificate validation is important for security and needs to be performed before an SSL/TLS connection is established. Some advanced testing techniques, such as frankencert , have revealed, through randomly mutating Internet accessible certificates, that there exist unexpected, sometimes critical, validation differences among different SSL/TLS implementations. Despite these efforts, X.509 certificate validation still needs to be thoroughly tested as this work shows. This article tackles this challenge by proposing transcert , a coverage-directed technique to much more effectively test real-world certificate validation code. Our core insight is to (1) leverage easily accessible Internet certificates as seed certificates and (2) use code coverage to direct certificate mutation toward generating a set of diverse certificates. The generated certificates are then used to reveal discrepancies, thus potential flaws, among different certificate validation implementations. We implement transcert and evaluate it against frankencert , NEZHA , and RFCcert (three advanced fuzzing techniques) on five widely used SSL/TLS implementations. The evaluation results clearly show the strengths of transcert : During 10,000 iterations, transcert reveals 71 unique validation differences, 12×, 1.4×, and 7× as many as those revealed by frankencert , NEZHA , and RFCcert , respectively; it also supplements RFCcert in conformance testing of the SSL/TLS implementations against 120 validation rules, 85 of which are exclusively covered by transcert -generated certificates. We identify 17 root causes of validation differences, all of which have been confirmed and 11 have never been reported previously. The transcert -generated X.509 certificates also reveal that the primary goal of certificate chain validation is stated ambiguously in the widely adopted public key infrastructure standard RFC 5280. Pengbo Nie, Chengcheng Wan 0001, Jiayu Zhu, Yuting Chen 0001, Zhendong Su 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2021 | ApproxiFuzzer: Fuzzing towards Deep Code Snippets in Java ProgramsabstractA real-world, complex software system can contain a number of code snippets. Many snippets are deep, surrounded by complicated triggering conditions and/or hidden in functions less frequently invoked. Fuzzing and symbolic execution are two mainstreams for exploring input spaces and increasing code coverage of complicated software systems. Meanwhile, it remains a challenge to determine whether a deep code snippet is reachable, and if it is reachable, which test(s) can reach it.This paper presents ApproxiFuzzer, an effective, demand-driven approach to fuzzing towards deep code snippets in Java programs. Given a program P, a target deep code snippet tcs, and a set of seeding test inputs, the key idea behind ApproxiFuzzer is to selectively mutate the test inputs and collect their execution traces such that the execution traces gradually approximate tcs; several measures are designed for measuring the distances between execution traces and the code snippet and directing the fuzzing process towards generating test inputs reaching tcs.We have implemented ApproxiFuzzer and evaluated it against Kelinci (an AFL-based fuzzer) and JDart (a concolic execution tool) on a set of real-world benchmarks. The evaluation clearly demonstrates the strengths of ApproxiFuzzer—ApproxiFuzzer outperforms Kelinci by 36× in efficiently generating test inputs, obtaining up to 18.2% higher code coverage; ApproxiFuzzer also outperforms JDart by 46.2∼96.2% in hitting deep code snippets. Xintian Yu, Enze Ma, Pengbo Nie, Beijun Shen, Yuting Chen 0001 |
COMPSAC | 3 |
| 2020 | Guided, Deep Testing of X.509 Certificate Validation via Coverage Transfer GraphsabstractSSL and TLS are two secure protocols for creating secure connections over the Internet. X.509 certificate validation is important for security and needs to be performed before an SSL/TLS connection is established. However, state-of-the-art testing techniques, such as frankencert and mucert, have revealed, through randomly mutating Internet accessible certificates, that there exist unexpected, sometimes critical, validation differences among different SSL/TLS implementations. Despite these strong efforts, certificate validation is still not thoroughly tested and more effective techniques are needed as this work shows. To this end, this paper introduces transcert, a novel approach for effectively guiding fuzzing to perform deep testing of X.509 certificate validation. The goal of transcert is to generate certificates that trigger diverse executions; it achieves this goal by introducing the concept of a coverage transfer graph to efficiently, precisely abstract program executions. In particular, it records the execution of how a given certificate is validated by a reference SSL/TLS implementation. It then constructs a coverage transfer graph to model the coverage transfer from a test certificate (seed) to its mutated certificates (mutants), and explores the coverage transfer graph by iteratively sampling and mutating certificates. We have implemented transcert and evaluated it against frankencert and mucert on four state-of-the-art SSL/TLS implementations. The evaluation results clearly show the strengths of transcert- during 10,000 iterations, transcert has revealed 3,469 validation differences, 8× as many as those revealed by frankencert and mucert. We have identified 11 root causes of validation differences, all of which have been confirmed and five have never been reported previously. We also found that the primary goal of certificate chain validation is stated ambiguously in the widely-adopted PKI standard RFC 5280. Jiayu Zhu, Chengcheng Wan 0001, Pengbo Nie, Yuting Chen 0001, Zhendong Su 0001 |
ICSME | 3 |