Zhe Wang 0015

dblp:75/3158-15 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0001-5832-0574ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 11 · 6 first-author · 6 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HybridSkipList+: Rethinking Distributed Skiplist With Hybrid RDMA and Caching
abstract
Remote Direct Memory Access (RDMA) offers high performance through OS kernel bypass and has become a key technology in modern data centers. By exploiting memorysemantic operations, RDMA-based data structures can achieve high scalability and significantly reduce CPU utilization compared with traditional Ethernet-based systems. However, classic sorted indexes such asSkiplistsuffer from low throughput under RDMA memory semantics due to frequent and costly remote accesses. To achieve both scalability and high throughput for a lock-based concurrent Skiplist, this paper introduces a hybrid paradigm that combines one-sided and two-sided RDMA operations. This proposal is built upon a re-evaluation of core design choices in RDMA, including transport modes, caching strategies, and memory management. Building on this paradigm, we designHybridSkipList+, a distributed Skiplist system that integrates a client-side coherent cache with pull-based synchronization to reduce expensive network round trips, and a semi-continuous memory allocator to enhance RDMA access locality. We implementHybridSkipList+ on an eight-machine RDMA cluster and conduct extensive evaluations. Results show that HybridSkipList+ outperforms two baseline systems by up to 4.41× and 3.45× under typical workload conditions.
Yilei Lu 0002, Teng Ma 0006, Dongbiao He, Zhe Wang 0015, Cédric Westphal, Linghe Kong
IEEE Trans. Computers5
2026 Mosaic: Towards Efficient and Low-Latency Multi-NeRF Complex Scene Cloud Rendering
abstract
Neural Radiance Fields (NeRF) offers superior reconstruction accuracy over traditional methods like point clouds. This is particularly advantageous for applications in retail, VR, and navigation. However, when representing a complex scene with multiple semantically meaningful objects, especially those containing high-frequency details, the conventional single NeRF approach struggles to accurately capture these intricate details. Furthermore, the state-of-the-art multi-NeRF solutions, though effectively capturing details by representing each object as a separate NeRF, are too resource-intensive to be run on mobile devices. Existing NeRF rendering acceleration studies focus on single NeRF data representation and its framework improvements, without resorting to the visibility information for further optimization and considering the practical multi-NeRF rendering scenario. In this paper, we present Mosaic, a multi-NeRF cloud rendering system that facilitates high-fidelity, and low-latency cloud rendering for complex scenes. We first design a multi-resolution controller to dynamically determine the resolution for each pixel based on its contextual details and view direction. It then incorporates a visibility-aware region pruner to effectively eliminate unneeded pixels to avoid being rendered. The extensive experiments on various scenes show that Mosaic can achieve high-fidelity rendering with 43% lower transmission cost and up to 70% smaller computation consumption when compared with other state-of-the-art benchmarks.
Zhe Wang 0015, Yifei Zhu 0001, Linghe Kong
IEEE Trans. Mob. Comput.1
2026 Loss-Tolerant RDMA Network Over Commodity Devices
abstract
This paper proposes the concept of a “loss-tolerant” RDMA network, instantiating as NüWa. It reveals the fundamental issues under a lossy fabric — packet losses and repetitive retransmission timeouts (RTOs) cause severe performance degradation and even service interruption. The loss-tolerant RDMA must avoid “important” packet losses that trigger RTOs. However, existing loss-protection mechanisms fail to identify these packets precisely. They either generate massive misprotection or ignore selective repeat loss recovery, resulting in buffer overflows and failures of RTO protection. To tackle these issues, NüWa thoroughly analyzes distinct loss-recovery schemes and RTO reasons for commodity NICs. It designsswitch modeandNIC modeto accurately identify and protect all important packets. The switch mode inherently supports the widely deployed non-programmable NICs, while the NIC mode offloads identification complexity to advanced programmable NICs. With effective RTO avoidance, it improves flow completion time (FCT) by 2 ∼ 10× compared to state-of-the-art (SOTA) solutions. In severe incast and large-scale networks, it reduces FCT by 100× compared with a lossless fabric. In storage applications, N¨uWa improves IOPS by ∼ 300% compared to vanilla lossy fabric. For typical AI Workloads, it accelerates AllReduce/AlltoAll communication by 5.5 ∼ 13.7× compared to SOTA lossy network solutions.
Likai Wang 0013, Zhe Wang 0015, Yimu Yuan, Shuhan Tian, Linghe Kong, Qiao Xiang, Shizhen Zhao, Di Qu, Hexiang Song, Yashar Ganjali, Guihai Chen
IEEE Trans. Netw.2
2025 NeRFlex: Resource-aware Real-time High-quality Rendering of Complex Scenes on Mobile Devices
abstract
Neural Radiance Fields (NeRF) is a cutting-edge neural network-based technique for novel view synthesis in 3D reconstruction. However, its significant computational demands pose challenges for deployment on mobile devices. While mesh-based NeRF solutions have shown potential in achieving real-time rendering on mobile platforms, they often fail to deliver high-quality reconstructions when rendering practical complex scenes. Additionally, the non-negligible memory overhead caused by pre-computed intermediate results complicates their practical application. To overcome these challenges, we present NeRFlex, a resource-aware, high-resolution, real-time rendering framework for complex scenes on mobile devices. NeRFlex integrates mobile NeRF rendering with multi-NeRF representations that decompose a scene into multiple sub-scenes, each represented by an individual NeRF network. Crucially, NeRFlex considers both memory and computation constraints as first-class citizens and redesigns the reconstruction process accordingly. NeRFlex first designs a detail-oriented segmentation module to identify sub-scenes with high-frequency details. For each NeRF network, a lightweight profiler, built on domain knowledge, is used to accurately map configurations to visual quality and memory usage. Based on these insights and the resource constraints on mobile devices, NeRFlex presents a dynamic programming algorithm to efficiently determine configurations for all NeRF representations, despite the NP-hardness of the original decision problem. Extensive experiments on real-world datasets and mobile devices demonstrate that NeRFlex achieves real-time, high-quality rendering on commercial mobile devices.
Zhe Wang 0015, Yifei Zhu 0001
ICDCS1
2025 LLM-assisted Industrial-Scale Differential Testing of Package Incompatibilities in Linux Distributions
abstract
An open source Linux distribution often undergoes version upgrades and migrations, which is prone to incompatibility issues especially when it comes to large-scale software changes. Although differential testing has been widely used in software testing, it is still challenging to apply it for detecting such incompatibilities in the context of industrial settings. In this paper, we report our experience in leveraging LLMs to address the challenges faced by the Linux distribution community. Specifically, we develop an LLM-based differential testing method called Versify to assist maintainers of Linux distributions in locating incompatibilities during version upgrades and migrations. Its trial operation period within the Linux distribution community shows that it uncovered 8,489 instances of differing behavior, of which 644 were prioritized for attention by developers. After deduplication and filtering, 39 unique compatibility reports were identified. Feedback from Linux distributions developers indicates that our reports have provided valuable recommendations for package selection in future OS releases.
Chijin Zhou, Runzhe Wang, Weibo Zhang, Yuheng Shen, Xiaohai Shi, Tao Ma 0006, Zhe Wang 0015, Heyuan Shi
ASE9
2024 Diagnosing Application-network Anomalies for Millions of IPs in Production Clouds
Zhe Wang 0015, Huanwu Hu, Linghe Kong, Xinlei Kang, Qiao Xiang, Peihao Yang, Jiejian Wu, Yong Yang 0013, Tao Ma 0006, Zheng Liu 0022, Xianlong Zeng, Dennis Cai, Guihai Chen
USENIX ATC1
2024 Zero+: Monitoring Large-Scale Cloud-Native Infrastructure Using One-Sided RDMA
abstract
Cloud services have shifted from monolithic designs to microservices running on cloud-native infrastructure with monitoring systems to ensure service level agreements (SLAs). However, traditional monitoring systems no longer meet the demands of cloud-native monitoring. In Alibaba’s “double eleven” shopping festival, it is observed that the monitor occupies resources of the monitored infrastructure and even disrupts services. In this paper, we propose a novel monitoring system named for cloud-native monitoring. achieves zero overhead in collecting raw metrics using one-sided remote direct memory access (RDMA) and remedies network congestion by adopting a receiver-driven flow control scheme. also features a priority queue mechanism to meet different quality of service requirements and an efficient batch processing design to relieve CPU occupation. has been deployed and evaluated in four different clusters with heterogeneous RDMA NIC devices and architectures in Alibaba Cloud. Results show that achieves no CPU occupation at the monitored host and supports$1\sim10k$hosts with$0.1\sim1s$sampling interval using a single thread for network I/O. significantly relieves the incast issue and maintains$80\sim95\%$of bandwidth utilization in several clusters when monitoring$1k$hosts. also ensures services with high priority accomplish collecting metrics earlier than low priority ones by at least$400 \mu s$when monitoring$1k$hosts.
Jiejian Wu, Teng Ma 0006, Zhe Wang 0015, Linghe Kong, Zhenzao Wen, Yong Yang 0013, Tao Ma 0006, Zheng Liu 0022, Guihai Chen
IEEE/ACM Trans. Netw.4
2023 OrthZig: Concurrent Transmissions Based on Waveform Orthogonality in ZigBee
abstract
As one of the key technologies in the Internet of Things (IoT), Zigbee is widely used in industrial, agricultural, or medical scenarios due to its short distance, low delay, and high reliability of communications. However, the intensive deployment and concurrent transmissions of devices in ZigBee networks lead to severe collision and interference, reducing the throughput of the entire wireless network. Effective collision resolutions have been explored in the literature, but most of them use chip sequence as unit to decompose collision through waveform comparison and subtraction, which will cause issues such as error propagation, bit errors due to scarce samples, and high computational complexity. To make up for the above deficiencies, we propose a new physical layer design called OrthZig to achieve high-precision collision resolution in ZigBee networks without central controller. The key idea of OrthZig is to utilize the Direct Sequence Spread Spectrum adopted by ZigBee physical layer to achieve concurrent transmissions of multiple users based on the orthogonality of signal waveforms and distributed coordination. Additionally, we reduce the requirement for synchronization accuracy by partitioning the sync pair in the spread sequence, ensuring orthogonality for correct decoding. Extensive simulations are conducted from multiple dimensions to simulate the overall performance. Evaluation results demonstrate that OrthZig has lower average bit error rate and higher network throughput than the state-of-the-art.
Zhe Wang 0015, Linghe Kong, Ying Shao, Shahid Mumtaz
GLOBECOM2
2023 LigBee: Symbol-Level Cross-Technology Communication from LoRa to ZigBee
abstract
Low-power wide-area networks (LPWAN) evolve rapidly with advanced communication primitives (e.g., coding, modulation) being continuously invented. This rapid iteration on LPWAN, however, forms a communication barrier between legacy wireless sensor nodes deployed years ago (e.g., ZigBee-based sensor node) with their latest competitor running a different communication protocol (e.g., LoRa-based IoT node): they work on the same frequency band but share different MAC- and PHY-layer regulations and thus cannot talk to each other directly. To break this barrier, we propose LigBee, a cross-technology communication (CTC) solution that enables symbol-level communication from the latest LPWAN LoRa node to legacy ZIGBEE node. We have implemented LigBee on both software-defined radios and commercial-off-the-shelf (COTS) LoRa and ZigBee nodes, and demonstrated that LigBee builds a reliable CTC link from LoRa node to ZigBee node on both platforms. Our experimental results show that i) LigBee achieves a bit error rate (BER) in the order of 10−3with 70 ∼ 80% frame reception ratio (FRR), ii) the range of LigBee link is over 300m, which is 6 ∼ 7.5× the typical range of legacy ZigBee and state-of-the-art solution, and iii) the throughput of LigBee link is maintained on the order of kbps, which is close to the LoRa’s throughput.
Zhe Wang 0015, Linghe Kong, Longfei Shangguan, Liang He 0002, Kangjie Xu, Yifeng Cao, Qiao Xiang, Jiadi Yu, Teng Ma 0006, Zheng Liu 0022, Guihai Chen
INFOCOM1
2023 Embracing Channel Estimation in Multi-Packet Reception of ZigBee
abstract
As a low-power and low-cost wireless protocol, the promising ZigBee has been widely used in sensor networks and cyber-physical systems. Since ZigBee based networks usually adopt tree or cluster topology, the convergecast scenarios are common in which multiple transmitters send packets to one receiver, leading to the severe collision problem. The conventional ZigBee adopts carrier sense multiple access with collisions avoidance to avoid collisions, which introduces additional time/energy overhead. The state-of-the-art methods resolve collisions instead of avoidance, in which mZig decomposes a collision by the collision itself and reZig decodes a collision by comparing with reference waveforms. However, mZig falls into high decoding errors only exploiting the signal amplitudes while reZig incurs high computational complexity for waveform comparison. In this paper, we propose CmZig to embrace channel estimation in multiple-packet reception (MPR) of ZigBee, which effectively improves MPR via lightweight computing used for channel estimation and collision decomposition. First, CmZig enables accurate collision decomposition with low computational complexity, which uses the estimated channel parameters modeling both signal amplitudes and phases. Second, CmZig adopts reference waveform comparison only for collisions without chip-level time offsets, instead of the complex machine learning based method. We implement CmZig on USRP-N210 and establish a six-node testbed. Results show that CmZig achieves a bit error rate in the order of$10^{-3}$and more than 80% packet reception rate, which outperforms state-of-the-art mZig.
Zhe Wang 0015, Linghe Kong, Xue (Steve) Liu, Guihai Chen
IEEE Trans. Mob. Comput.1
2022 Zero Overhead Monitoring for Cloud-native Infrastructure using RDMA
Zhe Wang 0015, Teng Ma 0006, Linghe Kong, Zhenzao Wen, Guihai Chen, Wei Cao 0006
USENIX ATC1
2020 Online Concurrent Transmissions at LoRa Gateway
abstract
Long Range (LoRa) communication, thanks to its wide network coverage and low energy operation, has attracted extensive attentions from both academia and industry. However, existing LoRa-based Wide Area Network (LoRaWAN) suffers from severe inter-network interference, due to the following two reasons. First, the densely-deployed LoRa ends usually share the same network configurations, such as spreading factor (SF), bandwidth (BW) and carrier frequency (CF), causing interference when operating in the vicinity. Second, LoRa is tailored for low-power devices, which excludes LoRaWAN from using the listen-before-talk (LBT) mechanisms commonly used in wireless communication technologies, such as WiFi and ZigBee -LoRaWAN has to use the duty-cycled medium access policy and thus being incapable of channel sensing or collision avoidance. To mitigate the inter-network interference, we propose a novel solution achieving the online concurrent transmissions at LoRa gateway, called OCT, which recovers collided packets at the gateway and thus improves LoRaWAN's throughput. Moreover, OCT achieves the online concurrent transmission using only LoRa's (de)modulation information, thus can be easily deployed at LoRa gateway. We have implemented and evaluated OCT on USRP platform and commodity LoRa ends, showing OCT achieves: (i) >90% packet reception rate (PRR), (ii) 3 × 10-3bit error rate (BER), (iii) 2x and 3x throughput in the scenarios of two- and three- packet collisions respectively, and (iv) reducing 67% latency compared with state-of-the-art.
Zhe Wang 0015, Linghe Kong, Kangjie Xu, Liang He 0002, Kaishun Wu, Guihai Chen
INFOCOM1
2020 Reference Waveforms Forward Concurrent Transmissions in ZigBee Communications
abstract
The number of Internet of Things is growing exponentially, among which the ZigBee devices are being widely deployed, incurring severe collision problem in ZigBee networks. Instead of collision avoidance or packet retransmissions which introduce extra time/energy overhead, existing methods try to decompose multi-packet collision directly. For example, state-of-the-art mZig exploits collision-free chips to decompose the collided chips iteratively, however, suffers from the high bit error rate and low frame reception rate which limit the practical applications. Toward this end, we observe three major issues of existing solutions: 1) all existing solutions adopt the priori-chip-dependent decomposition pattern, leading to the error propagation; 2) the available samples for chip decoding can be scarce, resulting in severe scarce-sample errors; 3) existing solutions assume the consistent frequency offset for consecutive packets, leading to inaccurate frequency offset estimation. To solve these issues in collision decomposition, we propose FORWARD, a novel physical layer design to enable accurate collision decoding in ZigBee. The key idea is to generate all possible overlapping combinations as reference waveforms. The decomposition is determined by comparing the collided signal with the reference waveforms. Such a priori-chip-independent design has the advantages to eliminate the error propagation. To ensure sufficient samples for decoding, FORWARD always choose the longest segment as reference. Furthermore, the real-time channel estimation and frequency offset calibration ensure the accurate collision decoding. We implement FORWARD on USRP platform and evaluate its performance. Experimental results demonstrate that FORWARD reduces bit error rate by order of magnitude and increases frame reception rate by 10% ~ 50% compared with the state-of-the-art.
Zhe Wang 0015, Yifeng Cao, Linghe Kong, Guihai Chen, Jiadi Yu, Shaojie Tang 0001, Yingying Chen 0001
IEEE/ACM Trans. Netw.1
2019 Forward the Collision Decomposition in ZigBee
abstract
As wireless communication is tailored for low-power devices while the number of Internet of Things is growing exponentially, the collision problem in ZigBee is worsen. The classical approaches of solving collision problems lie in collision avoidance and packet retransmission, which could incur considerable overhead. The new trend is to decompose multipacket collision directly, however, the high bit error rate limits its practical applications. Toward this end, we observe three major issues in the existing solutions: 1) all existing solutions adopt the priori-chip-dependent decomposition pattern, leading to the error propagation; 2) the available samples for chip decoding can be scarce, resulting in severe scarce-sample errors; 3) existing solutions assume the consistent frequency offset for consecutive packets, leading to inaccurate frequency offset estimation. To solve the issues of collision decomposition in ZigBee, we propose FORWARD, a novel physical layer design to enable highly accurate collision decomposition in ZigBee. The key idea is to generate all possible collided combinations as reference waveforms. The decomposition is determined by comparing the collided signal with the reference waveforms. Such a priori-chip-independent design has the advantages to eliminate the cumulative errors incurred from error propagation. When decoding, FORWARD always choose the longest segment to ensure sufficient samples for decoding. Furthermore, the recursive calibration design is approaching the real-time frequency offset and dynamically compensates the reference waveform. We implement FORWARD on USRP based testbed and evaluate its performance. Experimental results demonstrate that FORWARD reduces bit error rate by 4.96× and increases throughput 1.46~2.8× compared with the state-of-the-art mZig.
Yifeng Cao, Zhe Wang 0015, Linghe Kong, Guihai Chen, Jiadi Yu, Shaojie Tang 0001, Yingying Chen 0001
ICNP2
2019 reZig: Decompose a Collision via Reference Waveform in ZigBee
abstract
With the proliferation of IoTs, massive deployments of ZigBee devices intensify the collision, a fundamental problem in wireless communication. Conventional solutions exploit the protocol based on CSMA to avoid collision, which incurs considerable overhead. Recently, new solutions which focus on concurrent transmission via collision decomposition are proposed instead of collision avoidance. However, existing approaches to collision decomposition in ZigBee suffer the problems of error propagation and scarce-sample error. To make up the weakness of existing solutions, we propose a novel solution, reZig, which leverages the generated reference waveform in the receiver to infer each single waveform in the collision. By inferring the chip independently, reZig is error-propagation-free. As there are multiple segments provided to decode one chip, reZig can select the segment with most raw samples to ensure the decoding reliability. We implement reZig on the USRP N210 testbed. Experiments demonstrate that reZig achieves the 1.46~ 2.8× throughput compared to mZig, the state-of-the-art collision-decomposition approach to ZigBee.
Yifeng Cao, Zhe Wang 0015, Linghe Kong, Guihai Chen
IPCCC2
2018 PPM: Preamble and Postamble Based Multi-Packet Reception for Green ZigBee Communication
abstract
ZigBee, a low-power wireless communication technology, has been used in various applications such as smart health/home/buildings. The proliferation of ZigBee-based applications (and thus devices), however, makes the concurrent transmissions - i.e., multiple transmitters send packets to the same receiver at the same time - common in practice, leading to inevitable collisions. To facilitate the concurrent transmissions of ZigBee, we design Pre/Post-amble based Multi-packet reception (PPM), a method that recovers the collided ZigBee messages by exploiting their collision-free chips and the overlapped chips in their pre/post-ambles. Such a collision recovery of PPM reduces the retransmissions caused due to collisions, facilitating the realization green ZigBee. We have prototyped and evaluated PPM with USRP, showing PPM recovers the collided messages with bit-error-rates in the order of 10-6, which is magnitudes lower than state-of-the-art methods.
Zhe Wang 0015, Linghe Kong, Guihai Chen, Liang He 0002
GLOBECOM1