VLDB 2026 Research / reviewers in the wild / expert
Ning Jia 0004
dblp:25/6775-4
· DBLP profile ↗
17ranked-venue papers
2as first author
15since 2021 · last 2026
0009-0003-8246-4713ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Practical and Scalable RDMA Connection Sharing for HPC WorkloadabstractRDMA is a fundamental communication infrastructure in high-performance computing (HPC). However, as the number of RDMA connections increases, system performance rapidly declines and memory consumption increases sharply. Previous research demonstrates that sharing RDMA connections among processes is necessary and effective to address the scalability problem. Unfortunately, previous work shares connections in software, thus incurring substantial overhead to each packet operation, and fails to comprehensively explore control policies to achieve superior sharing decisions. Yuejie Wang, Tuo Fang, Biyu Peng, Xin Sun 0027, Chengchao Xu, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu, Yunfei Du 0001, Guyue Liu |
EuroSys | 9 |
| 2026 | Accelerating Model Loading in LLM Inference by Programmable Page Cache
Hongbo Li 0007, Xiaojia Huang, Yongfeng Wang, Hanjun Guo, Yuxin Ren 0001, Ning Jia 0004 |
FAST | 8 |
| 2026 | Bridging the Memory Hotness Gap in Edge Systems with Hotness-Segregated Object AllocationabstractKernel operations in resource-constrained edge systems, such as memory swapping and deduplication, use the access frequency (hotness) of memory pages to guide page placement and reclamation. However, these operations suffer from page-hotness skew: a page may contain a mix of highly accessed and infrequently accessed objects, which causes inaccurate page-level classification, wasted DRAM capacity, and expensive I/O. We attribute this skewness to a cross-layer mismatch: the kernel manages memory at page granularity, whereas user-level allocators place objects without considering access hotness. Ruizhe Huang, Jiahua Wang, Qihang Xu, Peng Jiang 0007, Zhida An, Ding Li 0001, Yao Guo 0001, Xiangqun Chen, Yuxin Ren 0001, Ning Jia 0004 |
LCTES | 10 |
| 2025 | Enabling Efficient Mobile Tracing with BTraceabstractWith the growing complexity of smartphone systems, effective tracing becomes vital for enhancing their stability and optimizing the user experience. Unfortunately, existing tracing tools are inefficient in smartphone scenarios. Their distributed designs (with either per-core or per-thread buffers) prioritize performance but lead to missing crucial clues with high probability. While these problems can be overlooked in previous scenarios (e.g., servers), they drastically limit the usefulness of tracing on smartphones. Arnau Casadevall-Saiz, Diogo Behrens, Ming Fu, Ning Jia 0004, Hermann Härtig, Haibo Chen 0001 |
ASPLOS (2) | 7 |
| 2025 | Cost-Efficient Cloud Infrastructure with Hugepage-aware Memory DeduplicationabstractOptimizing memory cost-efficiency is the top demand for many cloud computing scenarios. Memory deduplication and hugepage are both essential techniques for reducing memory cost and improving efficiency. However, the simultaneous use of memory deduplication and hugepages faces a dillema. Existing approaches either split hugepages into small pages to achieve efficient memory deduplication or ignore redundant portions within hugepages to maintain hugepage performance. Ruizhe Huang, Xinyu Wang 0043, Zhida An, Hanwen Lei, Peng Jiang 0007, Ziqi Zhang 0017, Ding Li 0001, Yao Guo 0001, Xiangqun Chen, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu |
SoCC | 13 |
| 2025 | FlacIO: Flat and Collective I/O for Container Image Service
Hongbo Li 0007, Mingrui Liu 0005, Rui Jing, Hanjun Guo, Yuxin Ren 0001, Ning Jia 0004 |
FAST | 9 |
| 2025 | Towards Rack-as-a-Computer in Memory Interconnect Era with Coordinated Operating System SharingabstractEmerging memory interconnect (such as CXL and HCCS) promises rack-scale machine to become a reality, as the interconnect enables load/store accessible memory shared across the entire rack. However, the rack-scale shared memory poses two unique challenges on the operating system, primarily because of synchronization bottleneck and reliability issue. First, hardware cache coherence is not guaranteed, thus existing lock-based approach is ineffective to synchronize cross-node memory access. Second, memory faults significantly increase, and additional interconnect hops and switches expand fault surface and radius. As a result, current systems cannot efficiently leverage in-rack shared memory and instead manage rack resource in a disaggregated way, suffering from unnecessary networking/RDMA transmission overhead and redundant data copies. Yuxin Ren 0001, Mingrui Liu 0005, Hongbo Li 0007, Chang Liao, Xiaojia Huang, Hanjun Guo, Ning Jia 0004 |
HotStorage | 9 |
| 2025 | DPUaudit: DPU-assisted Pull-based Architecture for Near-Zero Cost System AuditingabstractSystem auditing frameworks are crucial for modern data center security, as they record system events to detect intrusions. However, existing software-based auditing frameworks are limited by their high runtime overhead. To address the limitations of software-based frameworks, researchers had proposed a hardware-based auditing framework that offloads log processing to isolated hardware. However, despite using powerful specialized hardware, this approach still suffers from high runtime overhead, which contradicts their efficiency goal. We have identified that the high overhead is due to the pushbased architecture, which involves operating a log sender on the monitored host. Consequently, the existing approach requires heavy software protection mechanisms to secure the log sender, resulting in high runtime overhead.In this paper, we propose a new DPU-assisted pull-based architecture called DPUaudit for hardware-based auditing, which achieves near-zero runtime overhead. Instead of using a log sender, DPUaudit utilizes DPU to actively pull system events from the monitored host. This eliminates the need for heavy mechanisms to handle and safeguard the log sender, achieving highly efficient system auditing. Experimental results show that, on average, DPUaudit only slows down applications on the monitored host by 2.1% for six mainstream data center applications under different workloads, which is at least one order of magnitude smaller than existing approaches, while still ensuring the integrity of audit logs. Peng Jiang 0007, Hanlin Jiang, Ruizhe Huang, Hanwen Lei, Zhineng Zhong, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu, Yao Guo 0001, Xiangqun Chen, Ding Li 0001 |
HPCA | 8 |
| 2025 | A-Tune-Online: Efficient and QoS-Aware Online Configuration Tuning for Dynamic WorkloadsabstractAutomatic configuration tuning of online services with dynamic workloads has attracted increasing interest. Effective online tuning ensures configurations adapt to workload changes over time to maintain optimal online service performance. To be practical, online tuning must satisfy the dynamicity, efficiency, and Quality of Service (QoS) requirements. However, existing online tuning approaches fail to meet these requirements due to the inability to eliminate negative effects from historical observations. In this paper, we propose A-Tune-Online, an online configuration tuning system that tackles dynamic workloads, delivering superior tuning efficiency, and QoS guarantee simultaneously to a wide range of online scenarios. We identify that restarting the optimization based on explicit workload shift detection is necessary and critical to eliminate negative historical observations. First, to invoke optimization restarts appropriately, we design a multi-stage multi-indicator detection strategy based on heuristic rules and configuration replays. Then, to avoid initial efficiency drop after re-optimization, A-Tune-Online utilizes a similarity-based dual warm start scheme that transfers knowledge from similar historical workloads effectively. Finally, to prevent transient performance degradation from violating QoS guarantee after optimization restart, we leverage lower confidence bound to construct a safety region where each configuration is expected to perform better than the QoS requirement. Empirical study on five tuning scenarios showcases the superiority of A-Tune-Online compared with state-of-art tuning systems. A-Tune-Online achieves an average speedup of 2.90x and 1.72x compared with OnlineTune and DDPG+, respectively. We provide a version of our system in https://github.com/PKU-DAIR/A-Tune-Online. Yu Shen 0003, Beicheng Xu, Yupeng Lu, Huaijun Jiang, Zhipeng Xie, Senbo Fu, Nan Zhang 0004, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu, Bin Cui 0001 |
ICDE | 10 |
| 2025 | Predictable and Secure System Auditing for Real-Time SystemsabstractSystem auditing frameworks are essential for operating system security as they record system events to support intrusion detection, compliance verification and attack reconstruction. However, existing auditing frameworks fail to meet the stringent requirements of real-time systems, which demand security, predictability, and efficiency. Though current solutions are optimized for security or performance, they do not focus on bounding the worst-case execution time (WCET) and incorporating into response-time analysis (RTA). This paper presents RT-NODROP, a secure and predictable auditing framework tailored for real-time systems. RT-NODROP employs a lightweight threadlet-based architecture to isolate audit events processing, periodically invoking threadlets to simultaneously bound WCET and event residence time. By integrating with real-time schedule, RT-NODROP ensures no event dropping, system efficiency, and predictability. We further develop an overhead-aware RTA and a period selection algorithm to balance security, performance, and schedulability. The evaluations demonstrate that RT-NODROP is superior over state-of-the-art frameworks (Sysdig, OMNILOG, Ellipsis), improving schedulability by$\mathbf{8 0. 1 1 \%, ~} \mathbf{1 1 7. 9 \%}$and$\mathbf{5 1. 0 5 \%}$, respectively. For the latency-intensive application Redis, RT-NODROP achieves up to$\mathbf{7 5. 1 \%}(\mathbf{1 3 8. 8 6 \%}, \mathbf{3 2 4. 6 \%})$higher throughput and$\mathbf{2. 1 9} \times$(3.07x, 5.02x) lower 99.9th percentile tail latency than Sysdig (OMNILOG, Ellipsis) while maintaining a minimum event residence time around 10 ms without event dropping. Peng Jiang 0007, Fanhang Hu, Ruizhe Huang, Shuomin Xue, Zhaomeng Deng, Yuxin Ren 0001, Ning Jia 0004, Yao Guo 0001, Xiangqun Chen, Ding Li 0001 |
RTSS | 7 |
| 2025 | How to Copy Memory? Coordinated Asynchronous Copy as a First-Class OS ServiceabstractIn modern systems, memory copy remains a critical performance bottleneck across various scenarios, playing a pervasive role in system-wide execution such as syscalls, IPC, and user-mode applications. Numerous efforts have aimed at optimizing copy performance, including zero-copy with page remapping and hardware-accelerated copy. However, they typically target specific use cases, such as Linux zero-copy send() for messages of ≥10KB. This paper argues for copy as a first-class OS service, offering three key benefits: (1) with the asynchronous copy abstraction provided by the service, applications can overlap their execution with copy; (2) the service can effectively utilize hardware capabilities to enhance copy performance; (3) the service's global view of copies further enables holistic optimization. To this end, we introduce Copier, a new OS service of coordinated asynchronous copy, to serve both user-mode applications and OS services. We build Copier-Linux to demonstrate Copier's ability to improve performance for diverse use cases, including Redis, Protobuf, network stack, proxy, etc. Evaluations show that Copier achieves up to a 1.8 × speedup for real-world applications like Redis and a 1.6 × improvement over zIO, the state-of-the-art in optimizing copy efficiency. To further facilitate adoption, we develop a toolchain to ease the use of Copier. We also integrate Copier into a commercial smartphone OS (HarmonyOS 5.0), achieving promising results. Jingkai He, Yunpeng Dong, Dong Du 0003, Mo Zou, Zhitai Yu, Yuxin Ren 0001, Ning Jia 0004, Yubin Xia, Haibo Chen 0001 |
SOSP | 7 |
| 2025 | ECStore: Achieving Efficient and Compressible Indexing on Outsourced Encrypted DatabasesabstractEncrypted Databases (EDBs) are essential for protecting sensitive data outsourced to public clouds, enabling diverse index-based queries over encrypted data. However, existing EDB indexes often incur high storage overhead and performance degradation, primarily due to the poor compressibility of pseudorandom encrypted values, which leads to frequent accesses to slower persistent storage as indexes outgrow main memory. We introduceECStore, the first EDB that supports compressible and efficient indexing. Observing that EDB indexes are used solely for lookups and never decrypted, we designECTree, a cryptographic hash-based index structure in which each node is a compressible bit-string identifier that conceals plaintext keys.ECTreeenables logarithmic-time encrypted search via a novel membership testing mechanism. To address false positives arising in dynamic workloads, we introduceDirected View Check(DVC), which detects inaccuracies and avoids redundant traversals. Additionally,ECTree's Merkle-tree-like structure supports encrypted query authentication, resisting server compromise. Extensive evaluations show thatECStorecan achieve up to 94.7% lower latency and 10.5x higher throughput on popular benchmarks compared to notable EDBs. Tianxiang Shen, Ji Qi 0002, Ning Jia 0004, Haoze Song, Xiapu Luo, Sen Wang 0004, Heming Cui |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Toward Private, Trusted, and Fine-Grained Inference Cloud Outsourcing with on-Device Obfuscation and VerificationabstractMore and more AI-enabled applications are deployed on numerous devices, which rely on cloud platforms to provide complicated AI models and execution. The untrusted cloud environment raises serious security concerns about user data privacy and remote computation integrity. However, existing approaches rely on expensive cryptographic or hardware methods which are rarely adopted by the cloud. This paper proposes a fine-grained outsourcing strategy that only outsources composite reversible computation inside inference tasks to the cloud. Composite reversible computation allows us to convert the results calculated using obfuscated data back to actual results. The cloud is not required to be trusted, and we employ obfuscation and verification at the device to guarantee secure and accurate inference execution. Thanks to the reversibility of the computation, processing on the obfuscated data does not lose model accuracy and the reversed results allow the device to verify the computation quality. We conduct a feasibility analysis and the preliminary result shows the on-device obfuscation and verification still have performance benefits in the case of fine-grained outsourced computation. Haili Bai, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu |
ICNP | 5 |
| 2024 | Microkernel Goes General: Performance and Compatibility in the HongMeng Production Microkernel
Haibo Chen 0001, Xie Miao, Ning Jia 0004, Fei Wang 0130, Hongyang Yang, Fengwei Xu |
OSDI | 3 |
| 2024 | Interference-free Operating System: A 6 Years' Experience in Mitigating Cross-Core Interference in LinuxabstractReal-time operating systems employ spatial and temporal isolation to guarantee predictability and schedulability of real-time systems on multi-core processors. Any unbounded and uncontrolled cross-core performance interference poses a significant threat to system time safety. However, the current Linux kernel has a number of interference issues and represents a primary source of interference. Unfortunately, existing research does not systematically and deeply explore the cross-core performance interference issue within the OS itself. This paper presents our industry practice for mitigating crosscore performance interference in Linux over the past 6 years. We have fixed dozens of interference issues in different Linux subsystems. Compared to the version without our improvements, our enhancements reduce the worst-case jitter by a factor of 8.7, resulting in a maximum $11.5 x$ improvement over system schedulability. For the worst-case latency in the Core Flight System and the Robot Operating System 2, we achieve a 1.6x and $1.64 x$ reduction over RT-Linux. Based on our development experience, we summarize the lessons we learned and offer our suggestions to system developers for systematically eliminating cross-core interference from the following aspects: task management, resource management, and concurrency management. Most of our modifications have been merged into Linux upstream and released in commercial distributions. Zhaomeng Deng, Ziqi Zhang 0017, Ding Li 0001, Yao Guo 0001, Yunfeng Ye, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu |
RTSS | 7 |
| 2014 | SPTU: Improving Dynamic Binary Translation through Software Prediction with Target UpdatingabstractIn dynamic translation system, handling indirect branch is a major source of performance overhead, because it must perform an on-the-fly address translation at each indirect branch execution. The translation systems usually adopt software prediction to reduce the overhead of address translation, but the low prediction accuracy restricts the performance improvement. Ning Jia 0004, Xu Cheng 0001 |
SYSTOR | 1 |
| 2013 | SPIRE: improving dynamic binary translation through SPC-indexed indirect branch redirectingabstractDynamic binary translation system must perform an address translation for every execution of indirect branch instructions. The procedure to convert Source binary Program Counter (SPC) address to Translated Program Counter (TPC) address always takes more than 10 instructions, becoming a major source of performance overhead. This paper proposes a novel mechanism called SPc-Indexed REdirecting (SPIRE), which can significantly reduce the indirect branch handling overhead. SPIRE doesn't rely on hash lookup and address mapping table to perform address translation. It reuses the source binary code space to build a SPC-indexed redirecting table. This table can be indexed directly by SPC address without hashing. With SPIRE, the indirect branch can jump to the originally SPC address without address translation. The trampoline residing in the SPC address will redirect the control flow to related code cache. Only 2-6 instructions are needed to handle an indirect branch execution. As part of the source binary would be overwritten, a shadow page mechanism is explored to keep transparency of the corrupt source binary code page. Online profiling is adopted to reduce the memory overhead. Ning Jia 0004, Jing Wang 0055, Dong Tong 0001 |
VEE | 1 |