EDBT 2026 Demo / reviewers in the wild / expert
Xinwei Hu
dblp:54/7910
· DBLP profile ↗
12ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Practical and Scalable RDMA Connection Sharing for HPC WorkloadabstractRDMA is a fundamental communication infrastructure in high-performance computing (HPC). However, as the number of RDMA connections increases, system performance rapidly declines and memory consumption increases sharply. Previous research demonstrates that sharing RDMA connections among processes is necessary and effective to address the scalability problem. Unfortunately, previous work shares connections in software, thus incurring substantial overhead to each packet operation, and fails to comprehensively explore control policies to achieve superior sharing decisions. Yuejie Wang, Tuo Fang, Biyu Peng, Xin Sun 0027, Chengchao Xu, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu, Yunfei Du 0001, Guyue Liu |
EuroSys | 10 |
| 2026 | FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
Junbin Kang, Mingkai Dong 0002, Shaohong Guo, Ziyan Qiu, Mingzhen You, Ziyi Tian, Anqi Yu, Tianhong Ding, Xinwei Hu, Haibo Chen 0001 |
NSDI | 12 |
| 2025 | Cost-Efficient Cloud Infrastructure with Hugepage-aware Memory DeduplicationabstractOptimizing memory cost-efficiency is the top demand for many cloud computing scenarios. Memory deduplication and hugepage are both essential techniques for reducing memory cost and improving efficiency. However, the simultaneous use of memory deduplication and hugepages faces a dillema. Existing approaches either split hugepages into small pages to achieve efficient memory deduplication or ignore redundant portions within hugepages to maintain hugepage performance. Ruizhe Huang, Xinyu Wang 0043, Zhida An, Hanwen Lei, Peng Jiang 0007, Ziqi Zhang 0017, Ding Li 0001, Yao Guo 0001, Xiangqun Chen, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu |
SoCC | 14 |
| 2025 | DPUaudit: DPU-assisted Pull-based Architecture for Near-Zero Cost System AuditingabstractSystem auditing frameworks are crucial for modern data center security, as they record system events to detect intrusions. However, existing software-based auditing frameworks are limited by their high runtime overhead. To address the limitations of software-based frameworks, researchers had proposed a hardware-based auditing framework that offloads log processing to isolated hardware. However, despite using powerful specialized hardware, this approach still suffers from high runtime overhead, which contradicts their efficiency goal. We have identified that the high overhead is due to the pushbased architecture, which involves operating a log sender on the monitored host. Consequently, the existing approach requires heavy software protection mechanisms to secure the log sender, resulting in high runtime overhead.In this paper, we propose a new DPU-assisted pull-based architecture called DPUaudit for hardware-based auditing, which achieves near-zero runtime overhead. Instead of using a log sender, DPUaudit utilizes DPU to actively pull system events from the monitored host. This eliminates the need for heavy mechanisms to handle and safeguard the log sender, achieving highly efficient system auditing. Experimental results show that, on average, DPUaudit only slows down applications on the monitored host by 2.1% for six mainstream data center applications under different workloads, which is at least one order of magnitude smaller than existing approaches, while still ensuring the integrity of audit logs. Peng Jiang 0007, Hanlin Jiang, Ruizhe Huang, Hanwen Lei, Zhineng Zhong, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu, Yao Guo 0001, Xiangqun Chen, Ding Li 0001 |
HPCA | 9 |
| 2025 | A-Tune-Online: Efficient and QoS-Aware Online Configuration Tuning for Dynamic WorkloadsabstractAutomatic configuration tuning of online services with dynamic workloads has attracted increasing interest. Effective online tuning ensures configurations adapt to workload changes over time to maintain optimal online service performance. To be practical, online tuning must satisfy the dynamicity, efficiency, and Quality of Service (QoS) requirements. However, existing online tuning approaches fail to meet these requirements due to the inability to eliminate negative effects from historical observations. In this paper, we propose A-Tune-Online, an online configuration tuning system that tackles dynamic workloads, delivering superior tuning efficiency, and QoS guarantee simultaneously to a wide range of online scenarios. We identify that restarting the optimization based on explicit workload shift detection is necessary and critical to eliminate negative historical observations. First, to invoke optimization restarts appropriately, we design a multi-stage multi-indicator detection strategy based on heuristic rules and configuration replays. Then, to avoid initial efficiency drop after re-optimization, A-Tune-Online utilizes a similarity-based dual warm start scheme that transfers knowledge from similar historical workloads effectively. Finally, to prevent transient performance degradation from violating QoS guarantee after optimization restart, we leverage lower confidence bound to construct a safety region where each configuration is expected to perform better than the QoS requirement. Empirical study on five tuning scenarios showcases the superiority of A-Tune-Online compared with state-of-art tuning systems. A-Tune-Online achieves an average speedup of 2.90x and 1.72x compared with OnlineTune and DDPG+, respectively. We provide a version of our system in https://github.com/PKU-DAIR/A-Tune-Online. Yu Shen 0003, Beicheng Xu, Yupeng Lu, Huaijun Jiang, Zhipeng Xie, Senbo Fu, Nan Zhang 0004, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu, Bin Cui 0001 |
ICDE | 11 |
| 2024 | Optimizing File Systems on Heterogeneous Memory by Integrating DRAM Cache with Virtual Memory Management
Yuxin Ren 0001, Mingrui Liu 0005, Hongbo Li 0007, Hanjun Guo, Xie Miao, Xinwei Hu, Haibo Chen 0001 |
FAST | 7 |
| 2024 | Toward Private, Trusted, and Fine-Grained Inference Cloud Outsourcing with on-Device Obfuscation and VerificationabstractMore and more AI-enabled applications are deployed on numerous devices, which rely on cloud platforms to provide complicated AI models and execution. The untrusted cloud environment raises serious security concerns about user data privacy and remote computation integrity. However, existing approaches rely on expensive cryptographic or hardware methods which are rarely adopted by the cloud. This paper proposes a fine-grained outsourcing strategy that only outsources composite reversible computation inside inference tasks to the cloud. Composite reversible computation allows us to convert the results calculated using obfuscated data back to actual results. The cloud is not required to be trusted, and we employ obfuscation and verification at the device to guarantee secure and accurate inference execution. Thanks to the reversibility of the computation, processing on the obfuscated data does not lose model accuracy and the reversed results allow the device to verify the computation quality. We conduct a feasibility analysis and the preliminary result shows the on-device obfuscation and verification still have performance benefits in the case of fine-grained outsourced computation. Haili Bai, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu |
ICNP | 6 |
| 2024 | Interference-free Operating System: A 6 Years' Experience in Mitigating Cross-Core Interference in LinuxabstractReal-time operating systems employ spatial and temporal isolation to guarantee predictability and schedulability of real-time systems on multi-core processors. Any unbounded and uncontrolled cross-core performance interference poses a significant threat to system time safety. However, the current Linux kernel has a number of interference issues and represents a primary source of interference. Unfortunately, existing research does not systematically and deeply explore the cross-core performance interference issue within the OS itself. This paper presents our industry practice for mitigating crosscore performance interference in Linux over the past 6 years. We have fixed dozens of interference issues in different Linux subsystems. Compared to the version without our improvements, our enhancements reduce the worst-case jitter by a factor of 8.7, resulting in a maximum $11.5 x$ improvement over system schedulability. For the worst-case latency in the Core Flight System and the Robot Operating System 2, we achieve a 1.6x and $1.64 x$ reduction over RT-Linux. Based on our development experience, we summarize the lessons we learned and offer our suggestions to system developers for systematically eliminating cross-core interference from the following aspects: task management, resource management, and concurrency management. Most of our modifications have been merged into Linux upstream and released in commercial distributions. Zhaomeng Deng, Ziqi Zhang 0017, Ding Li 0001, Yao Guo 0001, Yunfeng Ye, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu |
RTSS | 8 |
| 2023 | Auditing Frameworks Need Resource Isolation: A Systematic Study on the Super Producer Threat to System Auditing and Its Mitigation
Peng Jiang 0007, Ruizhe Huang, Ding Li 0001, Yao Guo 0001, Xiangqun Chen, Jianhai Luan, Yuxin Ren 0001, Xinwei Hu |
USENIX Security Symposium | 8 |
| 2022 | From Dynamic Loading to Extensible Transformation: An Infrastructure for Dynamic Library Transformation
Yuxin Ren 0001, Jianhai Luan, Yunfeng Ye, Shiyuan Hu, Wenqin Zheng, Wenfeng Zhang, Xinwei Hu |
OSDI | 9 |
| 2021 | A PDRVIO Loosely coupled Indoor Positioning System via Robust Particle FilterabstractIn recent years, the Visual Inertial Odometry (VIO) technology has attracted attention as a support technology that improves medical experience and management efficiency. However, the performance of most VIO systems will drop drastically when the light intensity changes significantly or there are few texture features from the images. This paper designs a visual-inertial fusion-based navigation Indoor Positioning system to deal with the challenging scenario. It loosely coupled an inertial sensor-based pedestrian dead reckoning (PDR) model with the VIO model via a robust particle filter. The state estimation of the particle filter is based on the PDR model. The VIO model is used for the measurements of the particle filter. It compensates the gross errors of the VIO with a visual error propagation model which is established according to the posterior observation residuals of visual feature points. It is verified through experiments that the PDR/VIO fusion indoor positioning system based on the robust particle filter implemented in this paper has improved positioning accuracy and strong ability to deal with complex scenes. Xinwei Hu, Weilong Huang, Lingxiang Zheng, Ao Peng, Huiru Zheng, Haiying Wang 0001 |
BIBM | 1 |
| 2009 | A Novel Coupling-Based Model for Wideband MIMO ChannelabstractIn this paper, a novel analytical model structure for wideband multiple-input multiple-output (MIMO) channel is presented. It is based on the power coupling between direction of departure (DoD), direction of arrival (DoA) and delay domain. As its realizations, firstly the singular value decomposition (SVD) based model is introduced and the virtual presentation model is extended to the wideband situation. Then a hybrid model is given on basis of the power coupling between the transmit eigenmodes, receive eigenmodes and frequency steering vectors. The hybrid model can provide the tradeoff between the complexity and accuracy. At last, the novel coupling-based model structure is summarized. With a 3.52 GHz wideband MIMO sounder, measurements are carried out in different indoor scenarios. Monte Carlo simulations are used to generate the channel realizations according to these proposed wideband models. Good agreements are achieved between the discussed models and measured data. Yan Zhang 0009, Xinwei Hu, Yuanzhi Jia, Xiang Chen 0007, Jing Wang 0001 |
GLOBECOM | 2 |