VLDB 2026 Research / reviewers in the wild / expert
Khanh Nguyen 0001
dblp:53/6791-1
· DBLP profile ↗
14ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0003-0400-1070ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 6 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Program Analysis with Deterministic Distinguishable Calling ContextabstractCalling context is crucial for improving the precision of program analyses in various use cases (clients), such as profiling, debugging, optimization, and security checking. Often the calling context is encoded using a numerical value. We have observed that many clients benefit not only from a deterministic but also globally distinguishable value across runs to simplify bookkeeping and guarantee complete uniqueness. However, existing work only guarantees determinism, not global distinguishability. Clients need to develop auxiliary helpers, which incurs considerable overhead to distinguish encoded values among all calling contexts. In this paper, we propose Deterministic Distinguishable Calling Context Encoding () that can enable both properties of calling context encoding natively. The key idea of is leveraging the static call graph and encoding each calling context as the running call path count. Thereby, a mapping is established statically and can be readily used by the clients. Our experiments with two client tools show that has a comparable overhead compared to two state-of-the-art encoding schemes, PCCE and PCC, and further avoids the expensive overheads of collision detection, up to 2.1× and 50%, for Splash-3 and SPEC CPU 2017, respectively. Sungkeun Kim, Khanh Nguyen 0001, Chia-Che Tsai, Abdullah Muzahid, Eun Jung Kim 0001 |
CC | 2 |
| 2024 | CacheIt: Application-Agnostic Dynamic Caching for Big Data AnalyticsabstractApache Spark arguably is the most prominent Big Data processing framework tackling the scalability challenge of a wide variety of modern workloads. A key to its success is caching critical data in memory, thereby eliminating wasteful computations of regenerating intermediate results. While critical to performance, caching is not automated. Instead, developers have to manually handle such a data management task via APIs, a process that is error-prone and labor-intensive, yet may still yield sub-optimal performance due to execution complexities. Existing optimizations rely on expensive profiling steps and/or application-specific cost models to enable a postmortem analysis and a manual modification to existing applications.This paper presents CacheIt, built to take the guesswork off the users while running applications as-is. CacheIt analyzes the program’s workflow, extracting important features such as dependencies and access patterns, using them as an oracle to detect high-value data candidates and guide the caching decisions at run time. CacheIt liberates users from low-level memory management requirements, allowing them to focus on the business logic instead. CacheIt is application-agnostic and requires no profiling or a cost model. A thorough evaluation with a broad range of Spark applications on real-world datasets shows that CacheIt is effective in maintaining satisfactory performance, incurring only marginal slowdown compared to the manually well-tuned counterparts. Muhammad Rafid, Nathanael Santoso, Khanh Nguyen 0001 |
IEEE Big Data | 4 |
| 2021 | Snicket: Query-Driven Distributed TracingabstractIncreasing application complexity has caused applications to be refactored into smaller components known as microservices that communicate with each other using RPCs. Distributed tracing has emerged as an important debugging tool for such microservice-based applications. Distributed tracing follows the journey of a user request from its starting point at the application's front-end, through RPC calls made by the front-end to different microservices recursively, all the way until a response is constructed and sent back to the user. To reduce storage costs, distributed tracing systems sample traces before collecting them for subsequent querying, affecting the accuracy of queries on the collected traces. Jessica Berg, Fabian Ruffy, Khanh Nguyen 0001, Nicholas Yang, Anirudh Sivaraman, Ravi Netravali, Srinivas Narayana |
HotNets | 3 |
| 2021 | Adaptive huge-page subrelease for non-moving memory allocators in warehouse-scale computersabstractModern C++ server workloads rely on 2 MB huge pages to improve memory system performance via higher TLB hit rates. Huge pages have traditionally been supported at the kernel level, but recent work has shown that user-level, huge page-aware memory allocators can achieve higher huge page coverage and thus performance. These memory allocators deal with a trade-off: 1) allocate memory from the operating system (OS) at the granularity of a huge page, achieve high performance, but potentially waste memory due to fragmentation, or 2) limit fragmentation by breaking up huge pages into smaller 4 KB pages and returning them to the OS, but reduce performance due to lower huge page coverage. For example, the state-of-the-art TCMalloc allocator handles this trade-off by releasing memory to the OS at a configurable release rate, breaking up huge pages as necessary. Martin Maas 0001, Chris Kennelly, Khanh Nguyen 0001, Darryl Gove, Kathryn S. McKinley |
ISMM | 3 |
| 2020 | Semeru: A Memory-Disaggregated Managed Runtime
Chenxi Wang 0005, Yuanqi Li, Zhenyuan Ruan, Khanh Nguyen 0001, Michael D. Bond, Ravi Netravali, Miryung Kim, Guoqing Harry Xu |
OSDI | 6 |
| 2019 | Gerenuk: thin computation over big native data using speculative program transformationabstractBig Data systems are typically implemented in object-oriented languages such as Java and Scala due to the quick development cycle they provide. These systems are executed on top of a managed runtime such as the Java Virtual Machine (JVM), which requires each data item to be represented as an object before it can be processed. This representation is the direct cause of many kinds of severe inefficiencies. Christian Navasca, Cheng Cai, Khanh Nguyen 0001, Brian Demsky, Shan Lu 0001, Miryung Kim, Guoqing Harry Xu |
SOSP | 3 |
| 2018 | Skyway: Connecting Managed Heaps in Distributed Big Data SystemsabstractManaged languages such as Java and Scala are prevalently used in development of large-scale distributed systems. Under the managed runtime, when performing data transfer across machines, a task frequently conducted in a Big Data system, the system needs to serialize a sea of objects into a byte sequence before sending them over the network. The remote node receiving the bytes then deserializes them back into objects. This process is both performance-inefficient and labor-intensive: (1) object serialization/deserialization makes heavy use of reflection, an expensive runtime operation and/or (2) serialization/deserialization functions need to be hand-written and are error-prone. This paper presents Skyway, a JVM-based technique that can directly connect managed heaps of different (local or remote) JVM processes. Under Skyway, objects in the source heap can be directly written into a remote heap without changing their formats. Skyway provides performance benefits to any JVM-based system by completely eliminating the need (1) of invoking serialization/deserialization functions, thus saving CPU time, and (2) of requiring developers to hand-write serialization functions. Khanh Nguyen 0001, Lu Fang 0003, Christian Navasca, Guoqing Harry Xu, Brian Demsky, Shan Lu 0001 |
ASPLOS | 1 |
| 2018 | Calling-to-reference context translation via constraint-guided CFL-reachabilityabstractA calling context is an important piece of information used widely to help developers understand program executions (e.g., for debugging). While calling contexts offer useful control information, information regarding data involved in a bug (e.g., what data structure holds a leaking object), in many cases, can bring developers closer to the bug's root cause. Such data information, often exhibited as heap reference paths, has already been needed by many tools. Cheng Cai, Qirun Zhang, Zhiqiang Zuo 0002, Khanh Nguyen 0001, Guoqing Harry Xu, Zhendong Su 0001 |
PLDI | 4 |
| 2018 | Understanding and Combating Memory Bloat in Managed Data-Intensive SystemsabstractThe past decade has witnessed increasing demands on data-driven business intelligence that led to the proliferation of data-intensive applications. A managed object-oriented programming language such as Java is often the developer’s choice for implementing such applications, due to its quick development cycle and rich suite of libraries and frameworks. While the use of such languages makes programming easier, their automated memory management comes at a cost. When the managed runtime meets large volumes of input data, memory bloat is significantly magnified and becomes a scalability-prohibiting bottleneck. This article first studies, analytically and empirically, the impact of bloat on the performance and scalability of large-scale, real-world data-intensive systems. To combat bloat, we design a novel compiler framework, called F acade , that can generate highly efficient data manipulation code by automatically transforming the data path of an existing data-intensive application. The key treatment is that in the generated code, the number of runtime heap objects created for data classes in each thread is (almost) statically bounded , leading to significantly reduced memory management cost and improved scalability. We have implemented F acade and used it to transform seven common applications on three real-world, already well-optimized data processing frameworks: GraphChi, Hyracks, and GPS. Our experimental results are very positive: the generated programs have (1) achieved a 3% to 48% execution time reduction and an up to 88× GC time reduction, (2) consumed up to 50% less memory, and (3) scaled to much larger datasets. Khanh Nguyen 0001, Kai Wang 0029, Yingyi Bu, Lu Fang 0003, Guoqing Harry Xu |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2016 | Yak: A High-Performance Big-Data-Friendly Garbage Collector
Khanh Nguyen 0001, Lu Fang 0003, Guoqing Harry Xu, Brian Demsky, Shan Lu 0001, Sanazsadat Alamian, Onur Mutlu |
OSDI | 1 |
| 2015 | FACADE: A Compiler and Runtime for (Almost) Object-Bounded Big Data ApplicationsabstractThe past decade has witnessed the increasing demands on data-driven business intelligence that led to the proliferation of data-intensive applications. A managed object-oriented programming language such as Java is often the developer's choice for implementing such applications, due to its quick development cycle and rich community resource. While the use of such languages makes programming easier, their automated memory management comes at a cost. When the managed runtime meets Big Data, this cost is significantly magnified and becomes a scalability-prohibiting bottleneck. This paper presents a novel compiler framework, called Facade, that can generate highly-efficient data manipulation code by automatically transforming the data path of an existing Big Data application. The key treatment is that in the generated code, the number of runtime heap objects created for data types in each thread is (almost) statically bounded, leading to significantly reduced memory management cost and improved scalability. We have implemented Facade and used it to transform 7 common applications on 3 real-world, already well-optimized Big Data frameworks: GraphChi, Hyracks, and GPS. Our experimental results are very positive: the generated programs have (1) achieved a 3%--48% execution time reduction and an up to 88X GC reduction; (2) consumed up to 50% less memory, and (3) scaled to much larger datasets. Khanh Nguyen 0001, Kai Wang 0029, Yingyi Bu, Lu Fang 0003, Jianfei Hu, Guoqing Harry Xu |
ASPLOS | 1 |
| 2015 | Interruptible tasks: treating memory pressure as interrupts for highly scalable data-parallel programsabstractReal-world data-parallel programs commonly suffer from great memory pressure, especially when they are executed to process large datasets. Memory problems lead to excessive GC effort and out-of-memory errors, significantly hurting system performance and scalability. This paper proposes a systematic approach that can help data-parallel tasks survive memory pressure, improving their performance and scalability without needing any manual effort to tune system parameters. Our approach advocates interruptible task (ITask), a new type of data-parallel tasks that can be interrupted upon memory pressure---with part or all of their used memory reclaimed---and resumed when the pressure goes away. Lu Fang 0003, Khanh Nguyen 0001, Guoqing Harry Xu, Brian Demsky, Shan Lu 0001 |
SOSP | 2 |
| 2015 | Speculative region-based memory management for big data systemsabstractMost real-world Big Data systems are written in managed languages. These systems suffer from severe memory problems due to the massive volumes of objects created to process input data. Allocating and deallocating a sea of objects puts a severe strain on the garbage collector, leading to excessive GC efforts and/or out-of-memory crashes. Region-based memory management has been recently shown to be effective to reduce GC costs for Big Data systems. However, all existing region-based techniques require significant user annotations, resulting in limited usefulness and practicality. This paper reports an ongoing project, aiming to design and implement a novel speculative region-based technique that requires only minimum user involvement. In our system, objects are allocated speculatively into their respective regions and promoted into the heap if needed. We develop an object promotion algorithm that scans regions for only a small number of times, which will hopefully lead to significantly improved memory management efficiency. We also present an OpenJDK-based implementation plan and an evaluation plan. Khanh Nguyen 0001, Lu Fang 0003, Guoqing Harry Xu, Brian Demsky |
PLOS@SOSP | 1 |
| 2013 | Cachetor: detecting cacheable data to remove bloatabstractModern object-oriented software commonly suffers from runtime bloat that significantly affects its performance and scalability. Studies have shown that one important pattern of bloat is the work repeatedly done to compute the same data values. Very often the cost of computation is very high and it is thus beneficial to memoize the invariant data values for later use. While this is a common practice in real-world development, manually finding invariant data values is a daunting task during development and tuning. To help the developers quickly find such optimization opportunities for performance improvement, we propose a novel run-time profiling tool, called Cachetor, which uses a combination of dynamic dependence profiling and value profiling to identify and report operations that keep generating identical data values. The major challenge in the design of Cachetor is that both dependence and value profiling are extremely expensive techniques that cannot scale to large, real-world applications for which optimizations are important. To overcome this challenge, we propose a series of novel abstractions that are applied to run-time instruction instances during profiling, yielding significantly improved analysis time and scalability. We have implemented Cachetor in Jikes Research Virtual Machine and evaluated it on a set of 14 large Java applications. Our experimental results suggest that Cachetor is effective in exposing caching opportunities and substantial performance gains can be achieved by modifying a program to cache the reported data. Khanh Nguyen 0001, Guoqing Harry Xu |
ESEC/SIGSOFT FSE | 1 |