Yibin Shen

dblp:223/3172 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 5 since 2021Software engineering, systems software and programming languages · 7 · 5 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 ZOC: Elastic and Cost-Efficient Virtual SmartNIC Architecture for Cloud Physical Machines
Naixuan Guan, Xiaokang Hu, Yisheng Xie, Xishi Qiu, Chaojie Liu, Yuchao Cao, Banghao Ying, Dianchen Tian, Yangzeyu Zhang, Hujun Ge, Yibin Shen, Jiesheng Wu
NSDI12
2026 Law: Towards Consistent Low Latency in 802.11 Home Networks
Yibin Shen, Zili Meng
NSDI1
2026 Cacheman: A Comprehensive Last-Level Cache Management System for Multi-tenant Clouds
abstract
Competition for the last-level cache (LLC) is a long-standing issue in multi-tenant cloud environments, often leading to severe performance interference among co-located virtual machines. LLC management in the cloud faces unique challenges, including unpredictable tenant workloads, misaligned performance metrics, and the need to ensure fairness under service level agreements (SLAs). Existing LLC allocation methods fall short in addressing these challenges. We present Cacheman, a comprehensive LLC management system designed from real-world cloud deployment experience. Cacheman introduces a novel gradient-based sharing mechanism for LLC ways, enabling smooth LLC allocation adjustments that simultaneously improve fairness and utilization efficiency. Its real-time allocation algorithm promptly detects and mitigates unfair LLC allocation, adapting to dynamic workloads with second-scale responsiveness. Additionally, Cacheman supports performance consistency for tenants running distributed applications by enforcing negotiated upper bounds on cache usage. Extensive experiments demonstrate that Cacheman effectively achieves its multi-dimensional goals, and long-term production deployment further shows that it significantly reduces SLA violations caused by LLC contention.
Xiaokang Hu, Yuchao Cao, Naixuan Guan, Yifan Wu 0037, Xishi Qiu, Shengdong Dai, Ben Luo, Sanchuan Cheng, Fudong Qiu, Yibin Shen, Jiesheng Wu
PPoPP10
2026 Spillway: Orchestrating DPU and Host into a Unified vSwitching Fabric
abstract
The transition to Data Processing Unit (DPU)-centric architectures has become the de-facto standard in modern cloud networks, enabling infrastructure offload and improved host resource utilization. However, the fixed hardware limits of DPUs increasingly fail to keep pace with the rapid growth of host compute density and network-intensive workloads. As a result, when DPU resources are saturated, host compute capacity often remains underutilized due to insufficient network provisioning.
Xiaochong Jiang, Yilong Lv, Naixuan Guan, Qiming Zhao, Sihan Fu, Xuyang Ge, Denghui Wu, Yibin Shen, Guochun Hong, Yijian Dong, Yiquan Chen, Shaoliang An, Zhixiong Guo, Yisong Qiao, Hongwei Ding 0004, Shize Zhang, Rong Wen, Yang Song 0031, Zhigang Zong, Xing Li 0007, Chengkun Wei, Shunmin Zhu, Wenzhi Chen
SIGCOMM11
2026 Robustness evaluation and enhancement of LLMs in code generation: an empirical study
Senrong Xu, Yuan Yao 0001, Yibin Shen, Ping Yu 0011, Feng Xu 0007, Xiaoxing Ma
Empir. Softw. Eng.4
2026 EIDS: A Cloud Intrusion Detection System with High Performance and Maintainability
abstract
Intrusion Detection Systems (IDSes) are widely employed to identify potential attacks in guest virtual machines (VMs). Nonetheless, traditional IDSes fall short of the demands of high-performance clouds. First, monitoring VM events increases the tail latency of guest services. Second, the throughput of traditional IDSes cannot meet high-performance cloud requirements, leading to event loss and reduced detection accuracy. Finally, cloud providers typically run complex IDS tools within the VM. Updating IDS functionality requires modifying guest VMs, which hurts maintainability. To overcome these challenges, this article presents EIDS, a cloud IDS framework with high performance and good maintainability. We observe that the main bottleneck is collecting VM status, and the collected status can be divided into fundamental and supplementary status. EIDS then splits the status collection procedure spatially and temporally. First, we provide a status monitor with a separate architecture that isolates the status collection logic in a microVM, thus minimizing the code in guest VMs and improving maintainability. Second, EIDS introduces a two-phase status collection method to handle multiple events in batches, asynchronously, for high IDS throughput. A tiny tracer, implemented with eBPF, operates inside the user VM to collect fundamental status. The complex status collector runs in an isolated microVM. It utilizes Virtual Machine Introspection (VMI) to gather supplementary status, using the fundamental status to bridge the semantic gap. The status collector batches the collection for multiple events to amortize the fixed overhead of microVM switching and improve event tracing throughput. Finally, to minimize tail latency overhead, a fine-grained and workload-aware scheduler executes IDS logic with small time slices during user VM idle periods. We implemented a prototype of EIDS in Linux-KVM and conducted a comprehensive evaluation. We compared EIDS’s performance with Falco, an open-source IDS widely used by Kubernetes and AWS for runtime security monitoring. The results demonstrate that, compared to Falco, EIDS reduces the 99 th -percentile latency overhead by 97% and achieves a 13.8X improvement in IDS event handling throughput.
Xiaokang Hu, Zhichao Hua 0001, Naixuan Guan, Yibin Shen, Yang Yu 0002, Zeyu Mi, Yubin Xia, Jiesheng Wu
ACM Trans. Comput. Syst.5
2025 Fault Escaping: Improving Robustness of DPU Enhanced Platform with Mutual Assisted VM Recovery
abstract
Modern cloud servers achieve significant performance improvements by exploiting data processing units (DPUs) to offload virtualization overhead.Unlike traditional monolithic hypervisors, this offloading approach splits VM state and distributes hypervisor functions across the DPU and the Host, transforming the cloud server from a single system into a sophisticated orchestration of multiple self-managed processing units.However, the failure rate of these systems has significantly increased, as errors from either the DPU or the Host can crash the entire machine.All these changes necessitate a comprehensive revamp of current VM fault tolerance and isolation mechanisms.This paper explores a novel approach to VM fault tolerance by treating the split hypervisors on the DPU SoC and the Host as redundant peers.We propose a fault-escaping scheme that enables VMs to escape from a failing SoC or Host.With hybrid synchronization, the entire VM state becomes fully accessible on either the SoC or the Host with minimum synchronization overhead (in kilobytes).Instead of recovering the failed hypervisor, this scheme migrates the verified VM state, allowing the VM to escape from the failure
Chao Zhang 0115, Tao Xu 0057, Pai Liu, Zhilang Xu, Jinhu Li, Wenhui Shu, Feifei Fan, Yibin Shen, Jianming Song, Jiesheng Wu, Jian Li 0021
ASPLOS (3)13
2025 To PRI or Not To PRI, That's the question
Yun Wang 0039, Xianting Tian, Ben Luo, Zhixiang Wei, Zhibai Huang, Kailiang Xu, Kaihuan Peng, Kaijie Guo, Guangjian Wang, Shengdong Dai, Yibin Shen, Jiesheng Wu, Zhengwei Qi
OSDI14
2025 Effectively Virtual Page Prefetching via Spatial-Temporal Patterns for Memory-intensive Cloud Applications
abstract
In today's data-driven era, the explosive growth of global data volume has led to an increasing consumption of computing and storage resources. Effective management of virtual machines (VMs) memory usage is critical for cloud vendors to optimize system performance and resource utilization. Existing memory prefetching methods often slow down system performance, creating a difficult balance between maintaining service quality and optimizing resource use. For instance, Leap, which primarily utilizes address information, performs poorly in VM environments. The main issue is the performance drop caused by the reuse of memory resources in virtualized environments, a common situation in public clouds.
Yun Wang 0039, Tianmai Deng, Ben Luo, Yibin Shen, Zhixiang Wei, Yixiao Xu, Minglang Huang, Zhengwei Qi
PPoPP5
2025 Tai Chi: A General High-Efficiency Scheduling Framework for SmartNICs in Hyperscale Clouds
abstract
Cloud service providers increasingly adopt SmartNICs to offload data-plane services (e.g., DPDK and SPDK) and control-plane tasks (such as disk and NIC initialization). Our analysis of production environments reveals that data-plane services statically provision CPUs for peak load, resulting in 67.5% idle CPU cycles during 99% of their runtime in IaaS clouds, leading to wasted CPU resources. On the other hand, control-plane tasks fail to meet critical Service Level Objectives (SLOs), such as virtual machine startup time. Unfortunately, achieving control-plane SLO improvements through co-scheduling with idle data-plane services remains highly challenging, due to the combined effects of intrinsic scheduling latency and the substantial architectural complexity inherent to control-plane ecosystems.
Bang Di, Kaijie Guo, Yibin Shen, Sanchuan Cheng, Fudong Qiu, Xiaokang Hu, Naixuan Guan, Dongdong Huang, Jinhu Li, Yi Wang 0004, Yifang Yang, Yilong Lv, Zhenwei Lu, Jiesheng Wu
SOSP4
2024 vCrypto: a Unified Para-Virtualization Framework for Heterogeneous Cryptographic Resources
abstract
Transport Layer Security (TLS) connections involve costly cryptographic operations which incur significant resource consumption in the cloud. Hardware accelerators are affordable substitutes of expensive CPU cores to accommodate with the constantly increasing security requirements of datacenters. Existing accelerators virtualization mainly relies on passthrough of Single Root I/O Virtualization (SR-IOV) devices. However, deficiency of service accessibility, functionality and availability make device passthrough not an optimal solution for heterogeneous accelerators with different capabilities. To make up the gap, we propose vCrypto, a unified para-virtualization framework for heterogeneous cryptographic resources. vCrypto supports stateful crypto requests offloading and result retrieval with session lifecycle management and event driven notification. vCrypto transparently integrates virtual crypto device capabilities into the OpenSSL framework to benefit existing applications that are based on crypto library APIs without modification. Multiple physical resources can be partitioned flexibly and scheduled cooperatively to enhance the functionality, performance and robustness of virtual crypto service. Finally, vCrypto achieves an optimized performance with two layers polling and memory sharing mechanism. The comprehensive experiments show that with the same cryptographic resources used, vCrypto framework can provide 2.59x to 3.36x higher AES-CBC-HMAC-SHA1 throughput compared to passthrough SR-IOV device.
Chao Zhang 0115, Zongpu Zhang, Hubin Zhang, Weigang Li 0002, Yibin Shen, Jian Li 0021, Haibing Guan
INFOCOM9
2024 VPRI: Efficient I/O Page Fault Handling via Software-Hardware Co-Design for IaaS Clouds
abstract
Device pass-through has been widely adopted by cloud service providers to achieve near bare-metal I/O performance in virtual machines (VMs). However, this approach requires static pinning of VM memory, making on-demand paging unavailable. The hardware device I/O page fault (IOPF) capability offers an optimal solution to this limitation. Current IOPF approaches, using either standard IOMMU capabilities (ATS+PRI) or devices with independent IOMMU implementations, have not gained widespread adoption in public Infrastructure-as-a-Service clouds. This is due to high costs, platform dependency, and significant impacts on performance and service level objectives (SLOs). We present the Virtualized Page Request Interface (VPRI), a novel IOPF system developed through software-hardware collaboration. VPRI is not only platform-independent, free from address translation complexities, but also cost-effective, and designed to minimize SLO impact. Our work enables large-scale deployment of IOPF capability in Alibaba Cloud with negligible impact on SLOs. When integrated with memory management software, it significantly enhances memory utilization in public IaaS clouds, effectively overcoming the static memory pinning restriction associated with pass-through devices.
Kaijie Guo, Dingji Li, Ben Luo, Yibin Shen, Kaihuan Peng, Ning Luo 0003, Shengdong Dai, Jianming Song, Zeyu Mi
SOSP4
2023 Self-Supervised Interest Transfer Network via Prototypical Contrastive Learning for Recommendation
abstract
Cross-domain recommendation has attracted increasing attention from industry and academia recently. However, most existing methods do not exploit the interest invariance between domains, which would yield sub-optimal solutions. In this paper, we propose a cross-domain recommendation method: Self-supervised Interest Transfer Network (SITN), which can effectively transfer invariant knowledge between domains via prototypical contrastive learning. Specifically, we perform two levels of cross-domain contrastive learning: 1) instance-to-instance contrastive learning, 2) instance-to-cluster contrastive learning. Not only that, we also take into account users' multi-granularity and multi-view interests. With this paradigm, SITN can explicitly learn the invariant knowledge of interest clusters between domains and accurately capture users' intents and preferences. We conducted extensive experiments on a public dataset and a large-scale industrial dataset collected from one of the world's leading e-commerce corporations. The experimental results indicate that SITN achieves significant improvements over state-of-the-art recommendation methods. Additionally, SITN has been deployed on a micro-video recommendation platform, and the online A/B testing results further demonstrate its practical value. Supplement is available at: https://github.com/fanqieCoffee/SITN-Supplement.
Yibin Shen, Sijin Zhou, Xiang Chen 0017, Hongyan Liu 0001, Chunming Wu 0001, Chenyi Lei, Xianhui Wei, Fei Fang 0002
AAAI2
2023 MOEF: Modeling Occasion Evolution in Frequency Domain for Promotion-Aware Click-Through Rate Prediction
Xiaofeng Pan, Yibin Shen, Jing Zhang 0037, Hong Wen 0002, Chengjun Mao
DASFAA (2)2
2023 Efficient Memory Overcommitment for I/O Passthrough Enabled VMs via Fine-grained Page Meta-data Management
Ben Luo, Yibin Shen
USENIX ATC3
2022 TTPNet: A Neural Network for Travel Time Prediction Based on Tensor Decomposition and Graph Embedding
abstract
Travel time prediction of a given trajectory plays an indispensable role in intelligent transportation systems. Although many prior researches have struggled for accurate prediction results, most of them achieve inferior performance due to insufficient feature extraction of travel speed and road network structure from the trajectory data, which confirms the challenges involved in this topic. To overcome those issues, we propose a novel neuralNetworkforTravelTimePredictionbased on tensor decomposition and graph embedding, namedTTPNet, which can extract travel speed and representation of road network structure effectively from historical trajectories, as well as predict the travel time with better accuracy. Specifically,TTPNetconsists of three components: the first module (Travel Speed Features Layer) leverages non-negative tensor decomposition to restore travel speed distributions on different roads in the previous hour, and integrates a CNN-RNN model to extract both long-term and short-term travel speed features of the query trajectory; the second module (Road Network Structure Features Layer) utilizes graph embedding to generate the representation of local and global road network structure; the last module (Deep LSTM Prediction Layer) completes the final predicting task. Empirical results over two real-world large-scale datasets show that our proposedTTPNetmodel can achieve significantly better performance and remarkable robustness.
Yibin Shen, Cheqing Jin, Jiaxun Hua, Dingjiang Huang
IEEE Trans. Knowl. Data Eng.1
2020 High-density Multi-tenant Bare-metal Cloud
abstract
Virtualization is the cornerstone of the infrastructure-as-a-service (IaaS) cloud, where VMs from multiple tenants share a single physical server. This increases the utilization of data-center servers, allowing cloud providers to provide cost-efficient services. However, the multi-tenant nature of this service leads to serious security concerns, especially in regard to side-channel attacks. In addition, virtualization incurs non-negligible overhead in the performance of CPU, memory, and I/O. To this end, the bare-metal cloud has become an emerging type of service in the public clouds, where a cloud user can rent dedicated physical servers. The bare-metal cloud provides users with strong isolation, full and direct access to the hardware, and more predicable performance. However, the existing single-tenant bare-metal service has poor scalability, low cost efficiency, and weak adaptability because it can only lease entire physical servers to users and have no control over user programs after the server is leased. In this paper, we propose the design of a new high-density multi-tenant bare-metal cloud called BM-Hive. In BM-Hive, each bare-metal guest runs on its own compute board, a PCIe extension board with the dedicated CPU and memory modules. Moreover, BM-Hive features a hardware-software hybrid virtio I/O system that enables the guest to directly access the cloud network and storage services. BM-Hive can significantly improve the cost efficiency of the bare-metal service by hosting up to 16 bare-metal guests in a single physical server. In addition, BM-Hive strictly isolates the bare-metal guests at the hardware level for better security and isolation. We have deployed BM-Hive in one of the largest public cloud infrastructures. It currently serves tens of thousands of users at the same time. Our evaluation of BM-Hive demonstrates its strong performance over VMs.
Zhi Wang 0004, Yibin Shen
ASPLOS5
2020 Solving Math Word Problems with Multi-Encoders and Multi-Decoders
abstract
Math word problems solving remains a challenging task where potential semantic and mathematical logic need to be mined from natural language.Although previous researches employ the Seq2Seq technique to transform text descriptions into equation expressions, most of them achieve inferior performance due to insufficient consideration in the design of encoder and decoder.Specifically, these models only consider input/output objects as sequences, ignoring the important structural information contained in text descriptions and equation expressions.To overcome those defects, a model with multi-encoders and multi-decoders is proposed in this paper, which combines sequence-based encoder and graph-based encoder to enhance the representation of text descriptions, and generates different equation expressions via sequence-based decoder and tree-based decoder.Experimental results on the dataset Math23K show that our model outperforms existing state-of-the-art methods.
Yibin Shen, Cheqing Jin
COLING1
2020 A privacy-enhancing scheme against contextual knowledge-based attacks in location-based services
Jiaxun Hua, Yibin Shen, Xiuxia Tian, Yifeng Luo, Cheqing Jin
Frontiers Comput. Sci.3
2019 Fast and Scalable VMM Live Upgrade in Large Cloud Infrastructure
abstract
High availability is the most important and challenging problem for cloud providers. However, virtual machine monitor (VMM), a crucial component of the cloud infrastructure, has to be frequently updated and restarted to add security patches and new features, undermining high availability. There are two existing live update methods to improve the cloud availability: kernel live patching and Virtual Machine (VM) live migration. However, they both have serious drawbacks that impair their usefulness in the large cloud infrastructure: kernel live patching cannot handle complex changes (e.g., changes to persistent data structures); and VM live migration may incur unacceptably long delays when migrating millions of VMs in the whole cloud, for example, to deploy urgent security patches.
Zhi Wang 0004, Qi Li 0002, Junkang Fu, Yang Zhang 0016, Yibin Shen
ASPLOS7
2019 Road Intersection Detection Based on Direction Ratio Statistics Analysis
abstract
Large collections of GPS trajectory data provide us unprecedented opportunity to detect the road intersection automatically. However, in the real-world scenarios, the precision of existing detection methods cannot be guaranteed due to severe challenges including (i) low-quality raw GPS trajectory data and (ii) the difficulty of differentiating intersections from nonintersections. To tackle above issues, we propose a novel twophase road intersection detection framework, called as RIDF, which is comprised of trajectory quality improving and intersection extracting. More importantly, through extracting candidate cells based on direction statistic analysis and refining the locations of intersections using hybrid clustering strategy, our approach can effectively detect road intersections of different size. An experimental evaluation on two real data sets extensively assesses the quality of RIDF method by comparing it with state-of-theart methods. Experimental results demonstrate that our proposal can overcome the limitations of existing methods and thus have better accuracy than the existing work.
Min Pu, Jiali Mao, Yuntao Du 0002, Yibin Shen, Cheqing Jin
MDM4