Sheng Chen 0015

dblp:97/10424-15 · DBLP profile ↗
← Back
52ranked-venue papers
9as first author
47since 2021 · last 2026
0000-0001-7038-4407ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 27 · 3 first-author · 24 since 2021Systems, architecture and hardware · 19 · 4 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RespLoc: Static Device-Free Human Localization With Wi-Fi Respiration Signal
abstract
Device-free Wi-Fi localization is a promising technology to localize users who do not carry smart devices. The basic idea is to separate and analyze the signals reflected off human body from the multi-path signals. However, previous works could only localize moving users, because they can not distinguish the signals reflected off static users or objects like walls and furniture. This paper presents the Respiration Localization system,RespLoc, which for the first time enables device-free Wi-Fi localization system for static users. To recognize static users, the key insight is that people can breathe but objects cannot. However, it is non-trivial to extract user’s location from the respiration signal, because the respiration signal is significantly weaker than the regular motion signal. To this end, we propose the equivalent analysis method. Instead of using traditional signal separation, which suffers from severe noise due to residual signal components, we propose to construct an equivalent signal with the following properties: First, the equivalent signal follows the same variation law as the respiration signal; Second, the equivalent signal is not affected by irrelevant static signals. Based on this equivalent signal, we are able to resolve location features from the multi-path signals directly without separating them. We implementRespLocon commodity Wi-Fi devices, and extensive experimental results demonstrate thatRespLoccan localize static users with a median error of 0.89 meters.
Jiancheng Chen, Weiping Ge, Renrui Tan, Sheng Chen 0015, Xinyu Tong 0001, Keqiu Li
IEEE Internet Things J.4
2026 ARGUS: Cross-Antenna Channel Estimation and Intelligent Antenna Selection for Massive MIMO
abstract
Massive MIMO has emerged as a cornerstone technology for 5G-Advanced and future 6G networks, yet its practical deployment remains limited by hardware cost and power consumption. Switch-based architectures, which share a small number of RF chains among many antenna elements, provide a scalable alternative, but create a new bottleneck: only a subset of antennas is observable at any given moment, leaving the channel state of the remaining elements unknown. Lacking this information prevents the system from exploiting advanced physical-layer functions such as digital beamforming or multi-stream MIMO. In this paper, we present ARGUS, a generative channel reconstruction framework that infers the CSI of unobserved antennas from partial observations. The key idea is that all antenna responses are governed by the same underlying wireless propagation environment, enabling the task to be formulated as a generative inference problem. We employ a variational autoencoder to capture the latent spatial structure and reconstruct unobserved channels through sampling. Extensive experiments show that our reconstructed CSI incurs less than 2.5% achievable rate loss, and real-world measurements demonstrate a more than 90% antenna-selection match rate, confirming the practicality of the proposed approach.
Qibai Chen, Jianbo Hou, Haobo Gao, Jingyu Tong, Sheng Chen 0015, Xinyu Tong 0001, Xin Xie 0001, Xiulong Liu 0001, Keqiu Li
IEEE Internet Things J.6
2026 ArmNet: Robust Arm Motion Tracking for IoT Interaction Using a Single IMU
abstract
This paper presents ArmNet, a mobile sensing system for capturing the trajectory of the wrist using measurements from wrist-worn IMU devices, specifically targeting robust interaction within Internet of Things (IoT) ecosystems. Unlike existing solutions that directly map IMU data to joint positions, ArmNet integrates physical kinematic constraints with neural network modeling. This hybrid approach is crucial for resource-constrained IoT devices where computational overhead and sensor limitations are primary concerns. Specifically, kinematic priors capture spatial dependencies between the elbow and wrist, generating physics-guided intermediate features that reduce the solution space. These features, together with raw IMU data, are fed into a recurrent neural network to learn joint displacement vectors, which are then integrated into continuous trajectories. Extensive experiments show that ArmNet achieves robust and accurate arm tracking, generalizing well across users and motion patterns, thereby enabling a new modality for seamless human-computer interaction in smart environments.
Qinglin Jia, Xin Xie 0001, Xiulong Liu 0001, Xiaoyi Tao, Sheng Chen 0015, Keqiu Li
IEEE Internet Things J.6
2026 Physical-Semantic-Aware Multimodal Facial Expression Recognition for Human-Centric IoT
abstract
Facial Expression Recognition (FER) serves as a foundational sensory interface for Human-Centric IoT, supporting applications such as smart healthcare monitoring and affective intelligent environments. However, real-world performance is often hindered by theSemantic Gap, where models confuse visually similar expressions that arise from fundamentally different physiological muscle movements. To bridge this gap, we propose the Physical-Semantic-Aware Multimodal Framework (PSM-FER), which introduces 3D Blendshape (BS) coefficients as explicit physical priors to encode high-level muscle motion semantics. Our framework utilizes two synergistic pathways:Direct Physical Gating(DPG) for robust feature modulation andSemantic-Guided Spatial Attention(SGSA) for anatomical spatial recalibration. Additionally, an auxiliary physical regression task enforces anatomical consistency by regularizing the latent features to follow underlying biomechanical laws. Extensive experiments on the RAF-DB dataset demonstrate that PSM-FER achieves an accuracy of 92.37%, establishing a robust and interpretable foundation for affective sensing in complex IoT ecosystems.
Xin Xie 0001, Xiaoyi Tao, Xiulong Liu 0001, Sheng Chen 0015, Keqiu Li
IEEE Internet Things J.7
2026 Sequence-level watermarking for large language models
Runnan Si, Xin Xie 0001, Xiulong Liu 0001, Xiaoyi Tao, Xinyu Tong 0001, Sheng Chen 0015, Heng Qi, Keqiu Li
Knowl. Based Syst.7
2026 Bandwidth on a Budget: Real-Time Configuration for Edge Video Analysis
abstract
In an era marked by technological innovation, visual applications have become ubiquitous in everyday life. Harnessing the power of computer vision, these applications process and interpret video data from edge cameras, facilitating tasks such as object detection and vehicle counting. Yet, implementing complex deep learning models on cameras with limited computational capacity poses significant challenges. Furthermore, the bandwidth constraints and fluctuating nature of wide-area networks present substantial difficulties for video analysis systems dependent on cloud computing. This paper first characterizes the relationship between different parameter combinations (such as frame rate and resolution) and video analysis accuracy through offline analysis. It proposes a video stream analysis configuration selection scheme, SPStream, for slowly changing scenes, and a configuration file switching strategy, SPStream+, for rapidly changing scenes. These strategies use idle resources at the camera edge end to select the optimal configuration in real-time, adjust video encoding quality, and dynamically switch configuration files based on the changing states of object motion. Finally, a real-time video stream analysis system for vehicle counting and pedestrian detection suitable for both scenarios is designed, which saves bandwidth to the greatest extent while meeting the accuracy requirements of users and achieving high accuracy of video analysis.
Sheng Chen 0015, Xiaoyi Tao, Xin Xie 0001, Renrui Tan, Tu Hong, Xiulong Liu 0001
IEEE Trans. Computers1
2026 Hybrid Relay Architecture: A Decentralized and Semantically-Agnostic Interoperability Framework for Blockchain
abstract
The rapidly expanding blockchain ecosystem faces a critical interoperability crisis—evidenced by siloed total value locked and bridge losses. Existing interoperability solutions are trapped in a paradigm dilemma: centralized relays offer performance but introduce single points of failure, while decentralized relay-chains provide security but suffer from poor scalability and rigidity. To overcome these limitations, we introduce theHybrid Relay Architecture, a new decentralized interoperability paradigm that reconciles the efficiency of centralized relays with the security guarantees of decentralized systems via dynamic role stratification and probabilistic trust coordination. We instantiate this paradigm asCelestial, which realizes hybrid relaying through (1) a unified interchain data unit for coordination, (2) a geo-aware topology that enables adaptive role assignment, and (3) a proof of cross-chain transmission protocol that incentivizes honest node participation. Implemented in 12K+ LoC and evaluated across 90+ nodes spanning three continents,Celestialachieves 3,340 TPS with sub-600ms P99 latency, recovers from faults 6.4× faster than state-of-the-art systems, and reduces verification cost by 63% under adaptive attacks.
Sheng Chen 0015, Yiran Lv, Boyue Luan, Xiulong Liu 0001, Keqiu Li
IEEE Trans. Computers2
2026 LIBS: Instructional Action Quality Assessment via Supervoxel-Based Fine-Grained Attribution
abstract
The lack of actionable guidance is a fundamental limitation in Action Quality Assessment (AQA), as traditional methods provide overall scores without offering specific insights for improvement. Moreover, existing interpretable approaches often rely on expensive supervised spatial annotations or yield noisy, unsigned saliency maps. To address these challenges, we propose Learning Interpretability Based Supervoxels (LIBS), a novel framework for generating instructional feedback. Distinguishing itself from fully supervised methods, LIBS employs an unsupervised soft-clustering mechanism to segment videos into coherent supervoxels without requiring pixel-level mask annotations. This allows for scalable, fine-grained spatio-temporal analysis while preserving action continuity. Furthermore, we introduce a sensitivity propensity analysis to quantify the contribution of each supervoxel. Unlike traditional attribution methods, this mechanism explicitly decomposes the quality score into positive (strengths) and negative (flaws) components, enabling the system to decode abstract scores into concrete, actionable instructions. Experimental validation across multiple datasets demonstrates that LIBS achieves superior interpretability and efficiency compared to state-of-the-art baselines, marking an improvement from diagnostic to instructional AQA applications.
Xiaoyi Tao, Dongxu Ma, Liangzhi Li 0001, Manisha Verma, Lei Chen 0091, Xin Xie 0001, Sheng Chen 0015, Wenxin Li 0001, Jien Kato, Bing Zhang 0015, Xiulong Liu 0001
IEEE Trans. Computers7
2026 CLBP: A Cross-Modal Loss-Tolerant Beam Prediction Framework for V2V mmWave Communications
abstract
Millimeter-wave (mmWave) 5G-V2X communications face significant challenges in real-time beam alignment within high-mobility vehicular networks. While environmentaware beam prediction methods mitigate channel estimation overhead, their efficacy is severely compromised by modality data loss stemming from lighting variations, adverse weather, or sensor failures. To address this issue, we propose a Cross-modal Losstolerant Beam Prediction model (CLBP). CLBP robustly fuses RGB camera and LiDAR data, employing a novel cross-modal attention mechanism to achieve resilient feature alignment across these heterogeneous modalities. Furthermore, a Branch Features Dynamic Fusion (BFDF) module adaptively reweights modality features, suppressing noise from degraded inputs and promoting effective information propagation to enhance resilience. To facilitate realistic evaluation, we introduce a Data-Conditioned Missingness Mechanism (DCMM), which augments the DeepSense 6G V2V dataset with meticulously simulated sensor failure scenarios. Experimental results demonstrate CLBP's superior performance, achieving 94.48% Top-5 beam prediction accuracy even under 10% modality loss, and a 29% reduction in average power loss compared to baseline methods. These findings demonstrate CLBP's significant robustness in dynamic vehicular environments and its capacity to maintain consistent, high-performance beam prediction despite challenging data imperfections.
Xin Xie 0001, Xiulong Liu 0001, Zhe Peng, Xiaoyi Tao, Xinyu Tong 0001, Chaokun Zhang, Jiancheng Chen, Sheng Chen 0015, Keqiu Li
IEEE Trans. Mob. Comput.10
2026 Optimizing Timeliness for Distributed Stream Processing via Coflow Transmission
abstract
Distributed stream processing has recently gained much interest due to the need of extracting meaningful results from continuous data stream. To keep the extracted results fresh, the underlying network flows are often required to transmit packets continuously. Otherwise, these results will become stale, and their staleness is determined by the slowest flow. At this point,coflowscan be semantically comprised. Hence, efficient coflow transmission is critical for streaming applications. However, prior coflow-based solutions have significant limitations. They use a one-shot performance metric—CCT (coflow completion time), which cannot continuously reflect the staleness of the output results for a streaming application. To this end, we propose a new performance metric—coflow age(CA), for coflows generated by distributed streaming applications. The CA tracks thelongest time-since-last-serviceamong all flows in a coflow. In such a context, we consider a data center network with multiple coflows that continuously transmit packets between their source-destination pairs and address the problem of minimizing the average long-term CA while simultaneously satisfying the throughput constraints from the coflows. To solve this problem efficiently, we design a randomized algorithm and a drift-plus-age algorithm, and show that they can make the average long-term CA to achieve nearly two times and arbitrarily close to the optimal value, respectively. Through extensive simulations, we further demonstrate that both of the proposed algorithms can significantly reduce the CA of coflows, without violating the throughput requirement of any coflow, when compared to the state-of-the-art solution in both scenario with the packet arrival probability being known and unknown a prior.
Sheng Chen 0015, Wenxin Li 0001, Xu Yuan 0001, Keqiu Li, Heng Qi, Xiaobo Zhou 0003, Renhai Xu
IEEE Trans. Netw.1
2025 Fork: A Dual Congestion Control Loop for Small and Large Flows in Datacenters
abstract
Many existing transport designs aim to deliver ultra-low latency and high bandwidth for applications in high-speed datacenter networks. However, almost all of them intertwine the control of small and large flows using the same control entity (e.g., sender or receiver) and congestion feedback signal (e.g., ECN or credit), thus bringing significant performance impairments. By contrast, we seek to decouple the rate control of small flows from that of large ones.
Wenxin Li 0001, Yulong Li 0001, Lide Suo, Xuan Gao 0001, Xin Xie 0001, Sheng Chen 0015, Ziqi Fan, Wenyu Qu, Guyue Liu
EuroSys7
2025 SmartCache: Two-Dimensional KV-Cache Similarity for Efficient Long-Context LLM Decoding
abstract
Large language models (LLMs) achieve state-of-the-art performance in many NLP tasks but incur prohibitive memory-access and compute costs when processing very long contexts due to linearly growing KV Cache. Existing static sparsification methods rely on fixed heuristics, while dynamic schemes incur substantial runtime overhead. To address this trade-off, we propose SmartCache, a sparse inference system that exploits two-dimensional KV Cache similarity across adjacent decoding iterations and neighboring layers. SmartCache combines a similarity-driven dual-path selection algorithm, which adaptively reuses TopK KV entries from both the previous iteration and the preceding layer with a rolling-array cache index manager that reduces index storage complexity from$O(L \cdot k)$to$O(k)$. We analyze the layer and iterative sparse patterns of KV Cache in long context LLM decoding and show that SmartCache maintains semantic consistency while drastically reducing redundant computation and memory traffic. Extensive experiments on Llama-3-8B-Instruct-Gradient-1048k, Qwen2.5-7B-Instruct-1M, and glm-4-9b-chat-1m across four long-context benchmarks report up to$\mathbf{3 0. 5 \%}$end-to-end latency reduction and$\mathbf{1 5} \boldsymbol{\%} \mathbf{- 2 3 \%}$average latency reduction, with inference accuracy degradation constrained within 2% and occasional slight improvements. These results indicate that SmartCache offers a practical, high-accuracy solution for scalable long-sequence LLM inference.
Kaining Hui, Yitao Hu, Sheng Chen 0015, Xiulong Liu 0001, Keqiu Li
HPCC7
2025 SuperSpec: Enhanced Verification and Sampling for End-to-End LLM Speculative Decoding
abstract
Modern LLM decoding has the drawbacks of high cost and slow speed, and speculative decoding has been shown to be an effective solution to this problem. However, the inference latency still poses a significant challenge to maintaining service level objectives (SLOs) in systems that employ multiple draft models for speculative decoding. The verification phase in such systems if reliant on tree attention often constitutes a bottleneck especially when draft sequences lack common prefixes and substantially underutilizes GPU parallelism while increasing end-to-end latency. We introduce SuperSpec, an end-to-end speculative decoding system designed to co-optimize verification, sampling and draft generation. SuperSpec integrates three pivotal innovations: an Efficient Batch Verifier, which substitutes treebased flattening with batch parallel validation and layer-wise KV Cache replication; a Global Optimal Sampler, which assesses all candidate sequences within a batch to ascertain the longest valid path, thereby circumventing the local optima frequently encountered in tree-based rejection sampling; and a Dynamic Adaptive Multi-Drafter, which dynamically modulates the speculative length (K) for each drafter predicated on real-time idleness metrics and acceptance rates. Empirical evaluations of Qwen2.5-72B and the OPT-66B on various datasets show that SuperSpec improves average acceptance rate by 6.4% to 30.2%, and the end-to-end inference acceleration ratio by 7.12% to 62.06%, when compared to the state-of-the-art tree-based speculative decoding system SpecInfer. These improvements were achieved without compromising the quality of text generation, making SuperSpec an effective solution for accelerating LLM inference.
Yitao Hu, Sheng Chen 0015, Xiulong Liu 0001, Keqiu Li
HPCC7
2025 MoEoM: Joint Compute and Memory-Aware Balancing for Fast MoE Inference
abstract
Mixture-of-Experts (MoE) architectures have emerged as a scalable and efficient alternative to dense Transformer models by activating only a subset of experts per layer. However, deploying MoE models in multi-GPU environments faces severe challenges due to expert load imbalance and the resulting inefficient GPU utilization. Existing static replication strategies fail to adapt to dynamic token distributions, while dynamic rebalancing methods incur excessive communication and memory-access overheads, often outweighing the benefits of load balancing. This paper presents MoEoM, an inference system that enhances the efficiency of MoE models by innovatively taking memory-access costs into consideration. MoEoM integrates two complementary techniques: (i) a load-aware offline expert deployer, which symmetrically groups experts across GPUs and selectively replicates high-load experts, and (ii) an I/O-aware online token reallocator, which dynamically redistributes tokens among original and backup experts to minimize the maximum latency across GPUs. Experimental evaluations on state-of-theart MoE models demonstrate that MoEoM reduces end-toend inference latency by up to 16.6 %, improves throughput by 11–20 % in the prefill stage and 11–23 % in the decode stage, and decreases cross-GPU imbalance by 65–80 % (about 75 % on average), compared to prior MoE inference baselines. These results highlight the importance of incorporating both computation and memory-access overhead into expert placement and token scheduling for efficient large-scale MoE deployment.
Ziqi Gong, Yitao Hu, Sheng Chen 0015, Wenxin Li 0001, Keqiu Li
ICPADS3
2025 Lark: A Buffer-aware Building Block for Programmable Packet Scheduling in Datacenters
abstract
Programmable packet scheduling enables users to customize scheduling algorithms flexibly without designing new ASICs. Existing schemes prefre to approximate optimal Push-In First-Out (PIFO) using First-In First-Out (FIFO) queues in commodity programmable switches. Despite its availability, these schemes suffer performance degradation due to the unawareness of available switch buffer. To be specific, when the port buffer is drained, existing schemes discard all incoming packets, even though these packets have higher priorities than the enqueued packets. In this paper, we reveal that the problem's key culprit is the lack of coordination between buffer management and packet scheduling in the switch. To fill this gap, we present Lark, a buffer-aware building block for programmable scheduling schemes designed to solve the above problem. Its key idea is to proactively drop the low-priority packets when the allocated buffer is to be drained, thereby admitting the later-arriving high-priority packets. Lark contains two modules, a lightweight gradient-based online prediction module and a simple priority-based decision module. Lark relies the former module to identify whether the allocated buffer is to be drained and uses the later to determine whether to drop the incoming packet. We have integrated Lark into two representative schemes, SP-PIFO and AIFO. Our large-scale evaluations over three realistic workloads show that Lark can significantly optimize their key metrics without sacrificing throughnut.
Song Zhang 0008, Wenxin Li 0001, Yulong Li 0001, Lide Suo, Sheng Chen 0015, Yitao Hu, Laiping Zhao, Keqiu Li
INFOCOM6
2025 EVQ: Enabling Verifiable Blockchain Keyword Query in Federated-Storage Edge Computing
abstract
Due to the exponential growth of blockchain ledger sizes, federated-storage which enables multiple devices to jointly store data, has emerged as a promising solution for secure data storage in edge computing. However, how to achieve verifiable queries in such decentralized storage remains underexplored. Existing broadcast-based query methods lack a verification mechanism for query results, making it impossible to ensure their correctness and completeness. Meanwhile, authenticated data structure based (ADS-based) query strategies are constrained by the full ledger data and cannot provide verifiable query services for users in a federated-storage environment. To this end, this paper takes the lead to propose EVQ, a verifiable blockchain keyword query scheme tailored for federated-storage edge computing. We propose a split keyword-based ADS as the core structure of our framework which ensures that users can verify the correctness and completeness of query results while alleviating storage pressure of edge devices. Specifically, the proposed ADS is constructed through a two-phase process: top-bottom keyword index tree construction and bottomtop RSA accumulator integration. Splitting the ADS based on keywords enables distributed data storage and the generation of corresponding ADS for the stored data. To reduce the query costs incurred by edge devices during query processing, we formulate the Keyword Allocation Optimization (KAO) problem and propose a gain-ratio-based keyword allocation mechanism to determine the splitting scheme of the ADS. The experiments are conducted based on the Foursquare dataset, which contains approximately 18 months of global check-in data collected from Foursquare. The experimental results show that, compared to the merkle tree strategy that also integrates the RSA accumulator, our EVQ improves query performance by 24.77 x.
Baochao Chen, Xiulong Liu 0001, Hao Xu 0025, Sheng Chen 0015, Keqiu Li
IWQoS4
2025 BSSN: Enabling Adjustable Blockchain Storage for Resource-Constrained IoT Scenarios
abstract
Blockchain, with its immutability and decentralization, drives innovation in finance and supply chain, but the growing data volume makes storing complete ledger replicas impractical for users, especially in the resource-constrained Internet of Thing (IoT) scenarios. Existing solutions focus on nodes storing only a partial ledger to alleviate storage burdens. Nonetheless, these approaches prioritize storage optimization by minimizing the query cost and lack control over storage cost. Furthermore, these approaches overlook the relationships between network users, thus failing to fully measure the future query cost. Thus, this article proposes BSSN, a blockchain storage technology based on social networks. The combined use of storage cost and query cost is introduced for the first time to formulate the node allocation optimization (NAO) problem, and the multipopulation genetic ant colony (MGAC) algorithm will be employed to derive node allocation strategies. Specifically, we address three technical challenges: 1) to predict the transactions that nodes will participate in the future, we employ the social ties to obtain the access frequencies among users; 2) to strike a balance between the storage cost and query cost, we jointly model the two costs as a multiobjective optimization problem to formulate the NAO problem; and 3) to solve the NP-hard NAO problem, we use the MGAC algorithm, where the storage and query populations collaboratively search for solutions based on four operations. Extensive experiments indicate that compared with existing work, BSSN can reduce the average query cost to 67% with its adjustable storage cost, ensuring a balanced data storage among users.
Baochao Chen, Xiulong Liu 0001, Hao Xu 0025, Sheng Chen 0015, Keqiu Li
IEEE Internet Things J.4
2025 SmartGlove: Robust Sign Language Recognition With Cross-Domain Generation
abstract
Sign Language recognition is practically important in various scenarios such as smart home, medical rehabilitation, and intelligent industry. Compared with wireless sensing and computer vision methods, data glove-based methods have gained a plenty of attention, because they can perform well even in the environments with multi-path noise or visual occlusion. However, existing data glove-based methods usually require complex calibration and laborious dataset collection, and suffer from accumulated error. To address these challenges, we introduce a robust sign language recognition system with cross-domain generation, called SmartGlove, the first approach to achieve robust sign language recognition. To avoid complex calibration process, we propose a customized feature set that can enable user-insensitive and unintentional system calibration. To avoid the labor cost in training data collection, we propose a cross-domain data transformation technique to generate training data in target domain. To eliminate the accumulated error of sentence recognition, we utilize a context-based calibration method considering correlation among adjacent words. We implement SmartGlove with COTS devices, and extensive experiments reveal that SmartGlove achieves accuracy exceeding 97.11% for 30 sign language words, with an average recognition time of 47 milliseconds per word. Furthermore, the system recognizes 30 common sign language sentences with accuracy of 97.17%.
Mingli Feng, Xiulong Liu 0001, Jiancheng Chen, Jiuwu Zhang, Yuesen Liu, Sheng Chen 0015, Xiaoyi Tao, Xinyu Tong 0001, Xin Xie 0001, Keqiu Li
IEEE Internet Things J.7
2025 AirBFT: An Efficient and Robust Consensus Mechanism for Large-Scale Drone Collaboration
abstract
The application scenarios of drone collaboration are rapidly expanding, such as the low-altitude economy and wildfire protection. Blockchain-based drone collaboration requires a consensus mechanism to ensure efficient and secure consistency among large-scale distributed nodes. However, the existing consensus mechanism has problems with poor fault tolerance of topology and rigid proposal concurrency. To this end, this paper proposes AirBFT, an efficient and robust consensus mechanism for large-scale drone collaboration. First, this paper designs a new four-layer network topology, using upper-member and lower-member communication, while ensuring the maximum 1/3 resilience and fanout of √N. Secondly, this paper proposes a dynamic pipelining algorithm to adjust the parallelism of proposals according to the real-time network status. Finally, this paper proposes a committee sampling technology based on the EigenTrust algorithm to reduce the impact of the malicious behavior of Byzantine nodes. Experiments based on the public consensus framework show that compared with Kauri and HotStuff, the proposed AirBFT reduces transaction confirmation delay by 58%, the throughput is increased by 1.9 times, and it can ensure efficient operation with a 1/3 Byzantine node ratio.
Zhongju Yan, Chenyu Zhang 0008, Yiran Lv, Hao Xu 0025, Xiulong Liu 0001, Song Zhang 0008, Sheng Chen 0015, Xiaoyi Tao, Keqiu Li
IEEE Internet Things J.8
2025 TightLLM: Maximizing Throughput for LLM Inference via Adaptive Offloading Policy
abstract
Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks, largely due to their substantial model size. However, this also results in significant GPU memory demands during inference. To address these challenges on hardware with limited GPU memory, existing approaches employ offloading techniques that offload unused tensors to CPU memory, thereby reducing GPU memory usage. Since offloading involves data transfer between GPU and CPU, it introduces transfer overhead. To mitigate this, prior works typically overlap data transfer with GPU computation using a fixed pipelining strategy applied uniformly across all inference iterations, referred to asstaticoffloading. However, static offloading policies fail to maximize inference throughput because they cannot adapt to the dynamically changing transfer overhead during the inference process, leading to increasing GPU idleness and reduced inference throughput.We propose that offloading policies should beadaptiveto the varying transfer overhead across inference iterations to maximize inference throughput. To this end, we design and implement an adaptive offloading-based inference system called TightLLM with two key innovations. First, its key-value (KV) distributor employs atrade-compute-for-transferstrategy to address growing transfer overhead by dynamically recomputing portions of the KV cache, effectively overlapping data transfer with computation and minimizing GPU idleness. Second, TightLLM’s weight loader slices model weights and distributes the loading processacross multiple batches, amortizing the excessive weight loading overhead and significantly improving throughput. Evaluation across various combinations of GPU hardware and LLM models shows that TightLLM achieves 1.3 to 23 times higher throughput during the decoding phase and 1.2 to 22 times higher throughput in the prefill phase compared to state-of-the-art offloading systems. Due to the higher throughput in prefill and decoding phases, TightLLM can reduce the completion time for large-scale tasks, which involve processing and generating a substantial number of tokens, by 59.6% to 94.9%.
Yitao Hu, Xiulong Liu 0001, Guotao Yang, Sheng Chen 0015, Laiping Zhao, Wenxin Li 0001, Keqiu Li
IEEE Trans. Computers7
2024 FUYAO: DPU-enabled Direct Data Transfer for Serverless Computing
abstract
Serverless computing typically relies on the third-party forwarding method to transmit data between functions. This method couples control flow and data flow together, resulting in significantly slow data transmission speeds. This challenge makes it difficult for the serverless computing paradigm to meet the low-latency requirements of web services.
Laiping Zhao, Zhaolin Duan, Sheng Chen 0015, Yitao Hu, Zhiyuan Su, Wenyu Qu
ASPLOS (3)5
2024 Enabling 6D Pose Tracking on Your Acoustic Devices
abstract
The ubiquity of acoustic devices and the fine-grained sensing of acoustic signals have made acoustic device tracking a popular option. We propose to expand the use of commercial devices with microphones as an extension of the audio system to support intelligent applications, such as VR/AR. This paper introduces a novel 6D acoustic pose estimation system. To realize device-based pose estimation, most existing systems deploy multiple speakers. However, due to limited inaudible bandwidth, concurrent transmissions with multiple speakers pose challenges in balancing resolution and frame rate. To address this problem, we design 2×Track, a band multiplexing signal model that doubles the availability of limited bandwidth by utilizing a unique encoding strategy for concurrent transmissions. We also propose solutions to enhance signal feature estimation and implement a 6DoF pose tracking scheme tailored for distributed systems. The prototype is deployed on a typical circular microphone array, and experimental results show that 2×Track achieves a median position and orientation error of 7.6mm and 4.1°, respectively, in a 4-speaker setup. Our extended applications on commercial devices also showcase the versatility of our system, particularly in face orientation detection, air mouse and drone tracking.
Sheng Chen 0015, Xuanqi Meng, Xinyu Tong 0001, Xiulong Liu 0001, Xin Xie 0001, Wenyu Qu
MobiSys2
2024 VoiceMap: Autonomous Mapping of Microphone Array for Voice Localization
abstract
Voice command systems have been widely deployed on many smart devices for remote control. To further enrich the intelligence of these smart devices, the location of sound plays an important role in context-aware acoustic services. Despite initial steps made toward reliable voice localization, the state of the arts rely on prior knowledge of device location, device orientation and an indoor electronic map. To mitigate this additional cost, this paper presents VoiceMap, an autonomous mapping system of acoustic devices for voice localization. The insight behind VoiceMap is to explore the cooperation of sweeping robots and voice devices. Specifically, the sweeping robot is responsible for exploring the electronic map of the environment, while the microphone array is responsible for localizing the sweeping robot, so that we can establish the positional relationship between them. The core challenges are how to accurately locate the continuously moving robot, and how to synchronize the coordinate systems of the sweeping robot and the voice devices. To this end, we first design an inertial-based super-resolution method to estimate the angle of arrival (AoA) with respect to the robot. Then, we develop an effective coordinate synchronization mechanism, so that VoiceMap can automatically locate the voice devices on the electronic map generated by the robot. Finally, we implement a prototype system using commercial devices, and conduct comprehensive experiments to verify the proposed system. The experimental results show that we can realize a median error of 0.12m in terms of device localization.
Sheng Chen 0015, Renrui Tan, Xinyu Tong 0001, Keqiu Li
IEEE Internet Things J.1
2024 PosMonitor: Fine-Grained Sleep Posture Recognition With mmWave Radar
abstract
Sleep posture recognition is practically important in various scenarios such as sleep healthcare, bedridden patient care, and chronic disease diagnosis. With concerns of user privacy preserving, we prefer the wireless sensing methods to computer vision methods when dealing with sleep posture recognition. However, the existing wireless sensing methods suffer from at least one of the following major limitations: (i) difficult to deploy in practice; (ii) few posture categories; (iii) insufficient accuracy; (iv) poor generalization ability. In this paper, we use commercial-off-the-shelf (COTS) mmWave radar to implement a sleep posture recognition system called PosMonitor. When designing the PosMonitor system, we need to address the following challenging issues. First, we propose an angle purification method based on multi-frame joint analysis to alleviate the sparsity and instability of the point cloud. Then, we endow the point cloud with respiratory features to enhance its representation of the sleep posture. Further, to make the system applicable to different users, we extract relative respiratory features by normalization to overcome individual differences. Extensive experimental results show that our PosMonitor system can achieve 98% accuracy on average in recognizing 6 typical sleep postures and has good reliability across different conditions.
Xiulong Liu 0001, Sheng Chen 0015, Xin Xie 0001, Hankai Liu, Qixuan Cai, Xinyu Tong 0001, Wenyu Qu
IEEE Internet Things J.3
2024 A Wireless Signal Correlation Learning Framework for Accurate and Robust Multi-Modal Sensing
abstract
Wireless signal analytics in IoT systems can enable various promising wireless sensing applications such as localization, anomaly detection, and human activity recognition. As a matter of fact, there are significant correlations in terms of dimension, spatial and temporal aspects among wireless signals from multiple sensors. However, none of the wireless sensing research currently in use directly incorporates or exploits the signal correlations. Therefore, there is still substantial scope for improvement in regards to accuracy and robustness. We are introducing a novel framework called Signal Correlation Learning (SCL). This framework utilizes a directed graph to explicitly represent the signal correlation across various wireless sensors. We use signal embedding to depict the correlation features of a multi-dimensional sensor that arise from a multi-sensor system. Then, we perform Kullback-Leibler (KL) divergence on embedding vectors of any pair of sensors in the system to construct a subgraph at a given time point, which can measure the spatial signal correlation of sensors. Subsequently, several subgraphs spanning a specific time frame are fused into a coherent universal graph based on the small-world theory. This universal graph represents the three types of signal correlation simultaneously. A signal correlation aggregation structure is utilized to extract the features from the universal graph. These features can be used to address target sensing problems. We implement SCL in real RFID, Bluetooth, WIFI, and Zigbee systems, and evaluate its performance in three common wireless sensing problems including localization, anomaly detection, and human activity recognition. Extensive experiments demonstrate that our SCL framework significantly outperforms state-of-the-art wireless sensing algorithms by increasing$80\%\sim 190\%$in terms of accuracy, and by increasing$160\%\sim 220\%$in terms of robustness.
Xiulong Liu 0001, Bojun Zhang 0001, Sheng Chen 0015, Xin Xie 0001, Xinyu Tong 0001, Tao Gu 0001, Keqiu Li
IEEE J. Sel. Areas Commun.3
2024 Fine-Grained Recognition of Manipulation Activities on Objects via Multi-Modal Sensing
abstract
Fine-grained recognition of human manipulation activities on objects is crucial in the era of human-computer-object integration. However, there is a lack of solutions for simultaneous recognition of human identity, manipulation activities (including drawing and rotation), and manipulated objects. Therefore, we propose an RF-Camera system that combines RFID and computer vision techniques to address this challenge in multi-person and multi-object scenarios. In RF-Camera, we employ a skeleton-assisted method to extract facial images of target individuals, enabling precise recognition of their identities. To identify manipulation activities, we analyze the 3D hand trajectory and fingertip vector angle, differentiating drawing and rotation manipulation activities. Additionally, we model target person?s hand movements to predict phase data of the target tag, enabling the determination of person-object relationships. Implementing RF-Camera using COTS RFID and Kinect devices involves overcoming challenges such as extracting effective data from noisy streams, predicting virtual phase data considering hand-tag offset, and ensuring high tag reading rates in tag-dense scenarios. We conducted experiments involving six participants performing object manipulation activities, including drawing letters/symbols and rotating movements. Extensive experimental results show that RF-Camera achieves over 90% accuracy in recognizing person identity, manipulation activities, and person-object matching in most conditions.
Xiulong Liu 0001, Bojun Zhang 0001, Lizhang Wang, Sheng Chen 0015, Xin Xie 0001, Xinyu Tong 0001, Tao Gu 0001, Keqiu Li
IEEE Trans. Mob. Comput.4
2024 Toward Robust RFID Localization via Mobile Robot
abstract
A wide range of scenarios, such as warehousing, and smart manufacturing, have used RFID mobile robots for the localization of tagged objects. The state-of-the-art RFID-robot based localization works are based on the premise of stable speed. However, in reality this assumption can hardly be guaranteed because Commercial-Off-The-Shelf (COTS) robots typically have inconsistent moving speeds, and a small speed inconsistency will cause a large localization error. To this end, we propose a Speed Inconsistency-Immune approach to mobile RFID robot Localization (SILoc) system, which accurately locates targets when the robot moving speed varies or is even unknown. We propose an optimized unwrapping method to maximize the use of data, and a lightweight algorithm to calculate the locations in both 2D and 3D spaces. By utilizing the characteristics of tag-antenna distance and combining the phase data from multiple antennas, SILoc can effectively eliminate the side effects of speed inconsistency. To increase the flexibility, we further optimize the system and propose SILoc$+$, which enables the system to achieve localization with part of the data, keeping speed inconsistency-immune. Extensive experiments demonstrate that SILoc and SILoc$+$can achieve a centimeter-level localization accuracy in the scenario with an inconsistent or unknown robot moving speed.
Jiuwu Zhang, Xiulong Liu 0001, Sheng Chen 0015, Xinyu Tong 0001, Tao Gu 0001, Keqiu Li
IEEE/ACM Trans. Netw.3
2023 vHSFC: Generic and Agile Verification of Service Function Chain with parallel VNFs
abstract
With the advent of network function virtualization (NFV) and mobile edge computing (MEC), outsourcing network functions (NFs, i.e., firewall) to the MEC is becoming popular among network service providers. Notably, NF outsourcing raises an essential security concern about whether these outsourced NFs and associated service function chains (SFCs) are correctly implemented according to enterprises’ specifications. In particular, SFC with parallel VNFs, which take advantage of parallelism, have been conducted to reduce the traffic delay of traditional sequential SFC, called the hybrid SFC in this paper. Nevertheless, how to ensure correct behaviors and discover runtime mistakes for hybrid SFC remains an open problem.In this paper, we propose vHSFC, a verification scheme for hybrid SFC, enabling enterprises to verify the correctness of SFC enforcement in real-time. vHSFC achieves its goal with a lightweight verified routing protocol, which detects various hybrid SFC violations and attacks, i.e., packet modification, incompliant forwarding path, etc. To demonstrate the feasibility and performance of vHSFC, we have implemented the prototype on top of several containers and conducted extensive experiments with real traffic. The experimental results show that our vHSFC can continuously ensure proper enforcement and discover unexpected violations while incurring sensible overhead.
Sheng Chen 0015, Baochao Chen, Deke Guo, Keqiu Li
CSCWD1
2023 Generative Adversarial Network Based Asymmetric Deep Cross-Modal Unsupervised Hashing
Yuan Cao 0005, Yaru Gao, Jiacheng Lin, Sheng Chen 0015
ICA3PP (1)5
2023 Deep Hash Learning of Feature-Invariant Representation for Single-Label and Multi-label Retrieval
Yuan Cao 0005, Xinzheng Shang, Chengzhi Qian, Sheng Chen 0015
ICA3PP (1)5
2023 A Game Theory Based Task Offloading Scheme for Maximizing Social Welfare in Edge Computing
Sheng Chen 0015, Baochao Chen, Tu Hong, Renrui Tan, Xiaoyi Tao
ICA3PP (6)1
2023 Efficient Storage and Retrieval of Similar Data in Edge Computing Systems
abstract
Edge computing is migrating services from remote clouds to the network edge, where a vast amount of data is also flowing into edge nodes. In this context, the Edge Data-Sharing System (EDSS) enhances service quality by enabling edge nodes to cooperate. However, the EDSS is suitable for precise search and faces the existing high overhead when many users retrieve similar data. To solve the obstacle, this paper proposes a similarity-based edge storage system, SESS, which leverages the software-defined edge network to realize efficient storage and retrieval of similarity data. We first design RealminHash, a core module of SESS, for efficient signature and and index for each data. Then, SESS calculates the storage strategy based on the similarity between data. Importantly, SESS adjusts this strategy using periodic network information to ensure load balancing. Experimental results demonstrate that SESS realizes the nearest-neighbor storage while maintaining load balancing. SESS outperforms the well-known k-means and spectral clustering methods in terms of accuracy and latency and supports millisecond similar queries.
Yuanfeng Liu, Hanlong Liao, Sheng Chen 0015, Xiulong Liu 0001, Deke Guo
ICPADS4
2023 MiddleCache: Accelerating TCP based In-memory Key-value Stores using eBPF
abstract
In-memory key-value stores are widely used in modern web services to support large-scale user requests by caching popular data. Their performance is critical, and BMC, the state-of-the-art work, builds an in-kernel cache and processes requests before the stack using eBPF to reduce the overhead of the kernel network stack. However, BMC fails to support stateful protocol TCP because pre-stack processing creates TCP state bias between the client and server.TCP is widely used by in-memory key-value stores, is even the only choice for some applications (e.g., Redis), and also suffers from performance issues. In this work, we present MiddleCache, a TCP-enabled in-memory key-value store acceleration design. Our key observation is that the TCP state bias of the client and server can be inferred and eliminated with packet length. The design of MiddleCache has two key parts: (i) A compact TCP state maintenance mechanism that accumulates packet lengths and applies corrections to the packet header, which realize TCP support within the constrains of eBPF. (ii) Lock-free accumulation counters that support high-performance concurrent access by utilizing Receive Side Scaling (RSS). Our experiments show that, compared with Memcached, MiddleCache reduces 56% processing latency on cache hit and achieves a 3.8× throughput improvement on Facebook-like small-size requests workload.
Yiren Pang, Sheng Chen 0015, Wenxin Li 0001, Yulong Li 0001, Xin He 0043, Song Zhang 0008, Zewei Guan, Lide Suo
ICPADS2
2023 Sublessor: A Cost-Saving Internet Transit Mechanism for Cooperative MEC Providers in Industrial Internet of Things
abstract
Mobile edge computing (MEC) is becoming increasingly popular due to its remarkable computing capacities in close proximity to end users or devices. With the widespread use of Industrial Internet of Things, more and more cloud service providers move their services to the edge of the network for a better quality of service and become MEC providers. These MEC providers require to rent wide area network (WAN) connections to transfer industrial data, which is a considerable expense. In this article, we propose a framework calledSublessorto reduce the WAN transmission cost for a group of cooperative MEC providers. The key idea ofSublessoris allowing some specific MEC providers to act as Internet transit brokers, transmitting not only their own network traffic but also the traffic of their partners under a reasonable reselling price. This article formulates the problem as a mixed-integer programming and finds the most suitable broker number and corresponding reselling price without damaging the profit of both brokers and partners by a deep-reinforcement-learning-based algorithm. Experimental results show that our algorithm can significantly reduce the traffic transmission cost by up to 35%.
Sheng Chen 0015, Qihang Zhang, Xiaodong Dong, Xiaoyi Tao, Keqiu Li, Tie Qiu 0001, Ivan Lee 0001
IEEE Trans. Ind. Informatics1
2023 Availability-aware Provision of Service Function Chains in Mobile Edge Computing
abstract
With the advent of Network Function Virtualization (NFV) and Mobile Edge Computing (MEC), outsourcing network functions (NFs) to the MEC is becoming popular among network service providers (NSPs), since it brings the scalability and flexibility for NF deployment and maintenance. Each user’s request will go through a service function chain (SFC), which consists of several virtual network functions (VNFs, software substitutions for traditional hardware-based middleboxes) in a specific order, and then get a response. Unlike conventional hardware-based middleboxes, VNFs are not very reliable due to potential software faults and host malfunctions. Thus, a sensible way is to add redundancy for the primary VNFs of an SFC to enhance its availability. Nevertheless, which MEC node to place each VNF, and how many backup instances are enough to ensure the availability requirement of each SFC? These issues have not yet been resolved. In this article, we present the availability-aware provision of SFC (APoS) in the MEC environment with the primary goal of maximizing the number of served requests while meeting the requirements and reliability expectations of SFCs. For the APoS, we have primarily addressed the following two fundamental challenges: (i) First, how to efficiently map these primary and backup VNFs to meet the availability requirements of SFCs? At this point, we formulate it as an integer nonlinear programming (INLP) under the limitation of each MEC node’s resources. This issue is NP-hard, and a novel binary N-back search method is proposed to derive the optimal solution for the primary and backup VNFs mapping; (ii) Second, how can we reduce the latency for users to access their desired SFCs? Then, we investigate how to minimize the average delay for all requests in each time slot. To solve this problem, we design an online service switching (OSS) method, which jointly considers the queuing delay, communication delay, and switching delay. It achieves the optimal solution with a theoretical guarantee. Finally, we evaluate the proposed methods with real-world datasets. The results demonstrate that, compared with the benchmarks, our practices can achieve approximately 20% request acceptance gain and up to 30% delay reduction, on average.
Deke Guo, Sheng Chen 0015
ACM Trans. Sens. Networks4
2023 Low-cost crossed probing path planning for network failure localization
Hongyun Gao 0002, Laiping Zhao, Sheng Chen 0015, Keqiu Li
World Wide Web (WWW)3
2022 Efficient collision-slot utilization for missing tags identification in RFID system
Kaimin Guo, Xin Xie 0001, Sheng Chen 0015, Heng Qi, Keqiu Li
Comput. Commun.3
2022 An online dynamic pricing framework for resource allocation in edge computing
Sheng Chen 0015, Baochao Chen, Xiaoyi Tao, Xin Xie 0001, Keqiu Li
J. Syst. Archit.1
2022 Efficient Online Scheduling for Coflow-Aware Machine Learning Clusters
abstract
Distributed machine learning (DML) is an increasingly important workload. In a DML job, each communication phase can comprise acoflow, and there are dependencies among its coflows. Thus, efficient coflow scheduling becomes critical for DML jobs. However, the majority of existing solutions focus on scheduling single-stage coflows with no dependencies. While there are a few studies schedule dependent coflows of multi-stage jobs, they suffer from either practical or theoretical issues. Motivated by this situation, we study how to schedule dependent coflows of multiple DML jobs to minimize the total JCT in a shared cluster. We present a formal mathematical formulation for this problem and prove its NP-hardness. To solve this problem without job size information, we present an online coflow-aware optimization framework calledParrot. The core idea inParrotis to infer the job with the shortest remaining processing time (SRPT) each time and dynamically control the inferred job's bandwidth based on how confident it is an SRPT job while being mindful of not starving any other job. Specifically, in the design ofParrot, we present a least per-coflow attained service (LPCAS) policy to infer the SRPT job. We further propose a dynamic job weight assignment mechanism and a linear program (LP) based weighted bandwidth scaling strategy for sharing bandwidth among DML jobs. We have proved thatParrotalgorithm has a non-trivial competitive ratio. The results from large-scale trace-driven simulations further demonstrate that ourParrotcan reduce the total JCT by up to 58.4 percent, compared to the state-of-the-art Aalo solution.
Wenxin Li 0001, Sheng Chen 0015, Keqiu Li, Heng Qi, Renhai Xu, Song Zhang 0008
IEEE Trans. Cloud Comput.2
2022 Trading Cost and Throughput in Geo-Distributed Analytics With A Two Time Scale Approach
abstract
In the era of global-scale services, analytical queries are performed on datasets that span multiple data centers (DCs). Such geo-distributed queries generate a large amount of inter-DC data transfers at run time. Due to the expensive inter-DC bandwidth, various methods have been proposed to reduce the traffic cost in geo-distributed data analytics. However, current methods do not attempt to address the throughput issue in geo-distributed analytics. In this article, we target at characterizing and optimizing a cost-throughput tradeoff problem in geo-distributed data analytics. Our objectives are two-fold: (1) we minimize the inter-DC traffic cost when serving geo-distributed analytics with uncertain query demand, and (2) we maximize the system throughput, in terms of the number of query requests that can be successfully served with guaranteed queuing delay. Specifically, we formulate a stochastic optimization problem that seamlessly combines these two objectives. To solve this problem, we take advantage of Lyapunov optimization techniques to design and analyze a two-timescale online control framework. Without prior knowledge of future query requests, this framework makes online decisions on input data placement and admission control of query requests. Rigorous theoretical analyses show that our framework can achieve a near-optimal solution and maintain system stability and robustness as well. Extensive trace-driven simulation results further demonstrate that our framework is capable of reducing inter-DC traffic cost, improving system throughput, and guaranteeing a maximum delay for each query request.
Xinping Xu, Wenxin Li 0001, Renhai Xu, Heng Qi, Keqiu Li, Xiaobo Zhou 0003, Sheng Chen 0015
IEEE Trans. Cloud Comput.7
2022 Hash Learning With Variable Quantization for Large-Scale Retrieval
abstract
Approximate Nearest Neighbor(ANN) search is the core problem in many large-scale machine learning and computer vision applications such as multimodal retrieval. Hashing is becoming increasingly popular, since it can provide efficient similarity search and compact data representations suitable for handling such large-scale ANN search problems. Most hashing algorithms concentrate on learning more effective projection functions. However, the accuracy loss in the quantization step has been ignored and barely studied. In this paper, we analyse the importance of various projected dimensions, distribute them into several groups and quantize them with two types of values which can both better preserve the neighborhood structure among data. One is Variable Integer-based Quantization (VIQ) that quantizes each projected dimension with integer values. The other is Variable Codebook-based Quantization (VCQ) that quantizes each projected dimension with corresponding codebook values. We conduct experiments on five common public data sets containing up to one million vectors. The results show that the proposed VCQ and VIQ algorithms can both achieve much higher accuracy than state-of-the-art quantization methods. Furthermore, although VCQ performs better than VIQ, ANN search with VIQ provides much higher search efficiency.
Yuan Cao 0005, Sheng Chen 0015, Jie Gui, Heng Qi, Zhiyang Li 0001, Chao Liu 0008
IEEE Trans. Circuits Syst. Video Technol.2
2021 Localization of Tagged Objects on Shelf via a Portable Camera-augmented RFID Reader
abstract
Localization of target tagged objects on the shelf is of great significance in RFID-enabled warehousing scenarios. Compared with the RFID localization systems that use fixed reader antennas or mobile RFID-robot, the portable reader-based methods are much more cost-effective. Hence, this paper focuses on reader-portable RFID localization. However, the existing reader-portable localization systems suffer from the following limitations: (i) reader antenna is required to pass by the target tags. Thus, the tags in the corner can never be located; (ii) many reference tags need to be deployed on the shelf in advance, which considerably increases the manpower; (iii) specialized antenna is required, which limits the promotion potential. To this end, this paper proposes a Waving action-driven RFID Localization (WRL) system, which enables tag localization with a portable camera-augmented reader. In the WRL system, a user only needs to wave the camera-augmented reader before locating the target tags. Specifically, we first use a classical camera pose estimation method named PnP to recover the antenna’s movement trajectory in a pixel coordinate system. Then, WRL constructs a gridded hologram, in which camera data and RFID phase data are jointly used to calculate a probability for each grid. Intuitively, the higher probability a grid has, the more possible the target tag lies in the corresponding grid. Based on this idea, WRL calculates the target tag’s location on the shelf. We use the Commercial-Off-The-Shelf (COTS) RFID and camera devices to implement the WRL system. Extensive experiments have been conducted, and the results demonstrate that the mean localization error of WRL is less than 20cm with a confidence of about 95%.
Yazhe Tian, Sheng Chen 0015, Jiuwu Zhang, Zijuan Liu, Xiulong Liu 0001, Keqiu Li
ICCCN2
2021 Anomaly Detection of Network Streams via Dense Subgraph Discovery
abstract
We consider cyber security as one of the most significant technical challenges in current times. One of the main tasks is to detect anomalous patterns in the network streams as soon as they appear. In order to solve the above problem, previous propositions use statistical or machine learning-based methods to detect anomalous patterns in the network streams. However, these solutions incur significant low efficiency and precision due to the frequent recomputation of the results from scratch and unreasonable assumptions. In graph theory, dense subgraphs can be used to model the anomalous patterns if we abstract the network streams as a dynamic graph. This motivates us to explore dense subgraph discovery under the scenario where the network is updating. In this paper, we propose a graph-based framework, referred to as SAD, towards continuous dense subgraph discovery over network streams. In specific, we design an auxiliary data structure that is a concise representation of intermediate results, and its execution model allows a fast incremental maintenance strategy. In this way, we can detect anomalous patterns in the network streams in near real-time. Experiments demonstrate that SAD can not only get a higher accuracy of 90.2% but also faster than $11.4\times$ times compared to the state-of-the-art anomaly detection algorithms.
Qianzhen Zhang, Deming Mao, Ziyue Lu, Deke Guo, Sheng Chen 0015
ICCCN6
2021 An Intelligent Game Theory Framework for Detecting Advanced Persistent Threats
abstract
The advanced persistent threat (APT) is a stealthy cyber attack perpetrated by a group that gains unauthorized access to a computer network and remains undiscovered to steal specific data and resources. Fast detection and defense of APT attacks are critical tasks in cyber security. Previous works use simple feature extraction and classification methods to distinguish APT information flow from the normal one. However, APT attacks are latent, with very little flow and mixed in many normal information flows. Moreover, APT attacks can adjust their behavior according to the environment, making it challenging to be discovered and extract features. Meanwhile, dynamic information flow tracking (DIFT) is a tool for tracking information flow, which can also adjust the marking strategy according to the environment and is often used to track and detect APT information flow. On the other hand, game theory is a mathematical model that expresses the game of two or more parties. Therefore, this motivates us to model a game theory to solve the above challenge. In this paper, to solve the above obstacles, we propose an intelligent game theory framework named DPS, which models the strategic interaction between APTs and DIFT and aims to get a high reward for DIFT. Our proposed DPS framework utilizes deep reinforcement learning to find the Nash equilibrium. The game model is a nonzero-sum, average reward stochastic game. Specifically, we design a subgraph pruning strategy and deep Q-network to guide the player in exploring new strategies in the information flow graph. Finally, we implement our framework to compute an optimal defender strategy to defend cyber security. Based on 2 real-world datasets, the experiment results demonstrate that the DPS framework can delay APT intrusions under equilibrium in 3 epochs and get a better reward than the Uniform policy.
Qianzhen Zhang, Ziyue Lu, Sheng Chen 0015, Deke Guo
ICPADS5
2021 SILoc: A Speed Inconsistency-Immune Approach to Mobile RFID Robot Localization
abstract
Mobile RFID robots have been increasingly used in warehousing and intelligent manufacturing scenarios to pinpoint the locations of tagged objects. The accuracy of state-of-the-art RFID robot localization systems depends much on the stability of robot moving speed. However, in reality this assumption can hardly be guaranteed because a Commercial-Off-The-Shelf (COTS) robot typically has an inconsistent moving speed, and a small speed inconsistency will cause a large localization error. To this end, we propose a Speed Inconsistency-Immune approach to mobile RFID robot Localization (SILoc) system, which can accurately locate RFID tagged targets when the robot moving speed varies or is even unknown. SILoc employs multiple antennas fixed on the mobile robot to collect the phase data of target tags. We propose an optimized unwrapping method to maximize the use of the phase data, and a lightweight algorithm to calculate the locations in both 2D and 3D spaces based on the unwrapped phase profile. By utilizing the characteristics of tag-antenna distance and combining the phase data from multiple antennas, SILoc can effectively eliminate the side effects of moving speed inconsistency. Extensive experimental results demonstrate that SILoc can achieve a centimeter-level localization accuracy in the scenario with an inconsistent or unknown robot moving speed.
Jiuwu Zhang, Xiulong Liu 0001, Tao Gu 0001, Xinyu Tong 0001, Sheng Chen 0015, Keqiu Li
INFOCOM5
2021 Joint Service Placement for Maximizing the Social Welfare in Edge Federation
abstract
Mobile Edge Computing (MEC) is a promising cloud-network convergence paradigm which provides computational resources close to end devices at the network edge. There exist multiple Edge Infrastructure Providers (EIPs) in MEC which independently manage edges and provide services to customers. Due to the exponentially increasing data generated by end devices, it is almost impossible for a single EIP to accommodate offloaded data. Moreover, when considering that multiple EIPs provide services through federation, an urgent challenge is how to ensure the sustainability of federation. Most of the existing work improves the service provision capabilities of MEC by optimizing service placement without considering the existence of multiple EIPs. In this paper, we design the horizontal collaboration of edge federation, which integrates all edges of all EIPs. First, we model the service placement problem as a programming problem, towards the goal of maximizing social welfare. Then, we propose two dynamic pricing methods for EIPs to determine typical price for customers and insourcing price for other EIPs. The evaluation results based on two real-world data sets demonstrate that our proposed service placement model can increase the total gain of EIPs by up to 24.5% with a decrease of 35.5% in total delay.
Sheng Chen 0015, Baochao Chen, Xiulong Liu 0001, Deke Guo, Keqiu Li
IWQoS1
2021 Mobile Semantic-Aware Trajectory for Personalized Location Privacy Preservation
abstract
Synthesizing a fake trajectory with consistent lifestyle and meaningful mobility as the actual one is the most popular way to protect the location privacy in trajectory sharing. Recent location privacy preservation shows a strong personalized requirement from the mobile semantics between users and locations. However, the existing techniques cannot fully satisfy such personalized requirements, resulting in either overprotection or underprotection. It remains open to characterize and quantify the personalized requirement for location privacy preservation. In this article, we propose a mobile semantic-aware privacy model, named MSP. Specifically, we first characterize a new kind of user-related mobile semantic on-location set by constructing a hierarchical semantic tree, according to the user’s roles at locations. Then, a dedicated approach is proposed to evaluate the location’s privacy sensitivity and integrate it into the user-related mobile semantic. Finally, an adaptive privacy-preserving mechanism, MSP, is developed, fully considering the personalized requirement from both the user and the location. With this model in place, mobile semantic-aware synthetic trajectories are constructed adaptively. Extensive experiments with a real-world data set demonstrate that our MSP model can achieve an effective and flexible balance between the personalized privacy preservation and the data availability of synthetic trajectories.
Guoying Qiu, Deke Guo, Yulong Shen 0001, Guoming Tang, Sheng Chen 0015
IEEE Internet Things J.5
2020 Load Balance Awared Data Sharing Systems In Heterogeneous Edge Environment
abstract
Edge computing has become the de facto method for delay-sensitive applications, in which the computation and storage resources are placed at the edge of network. The main responsibility of edge computing is to carry data from the Cloud downlinks and terminal uplinks, and organize these data well on the edge side. This is the basis for subsequent analysis and processing of data. Therefore, this brings about a question as to how those data should be organized on the edge and how to store and retrieve them. In response to this demand, some methods have been proposed to solve the related problems of how to build data storage and retrieval services on the edge side. Those methods propose three different solutions: structured, unstructured, and hybrid schemes. However, the data storage and retrieval services for the heterogeneous edge environment is still lack of research. It is still not considered an important design test load balancing when the data is stored on the edge side. In this paper, we design and implement w-strategy, a load balance approach that implements the appropriate load balance among the heterogeneous edge nodes by using the weighted Voronoi diagram. Our solution utilizes the software defined networking paradigm to support a virtual-space based distributed hash tables (DHTs) to distribute data. Evaluation results show that w-strategy achieves better load balancing among the heterogeneous edge nodes compared to the existing methods, GRED and Chord. And, the w-strategy improves the average underutilization of the resources by 20%.
Sheng Chen 0015, Siyuan Gu, Baochao Chen, Deke Guo
ICPADS1
2020 Efficient Coflow Transmission for Distributed Stream Processing
abstract
Distributed streaming applications require the underlying network flows to transmit packets continuously to keep their output results fresh. These results will become stale if no updates come, and their staleness is determined by the slowest flow. At this point, coflows can be semantically comprised. Hence, efficient coflow transmission is critical for streaming applications. However, prior coflow-based solutions have significant limitations. They use a one-shot performance metric-CCT (coflow completion time), which cannot continuously reflect the staleness of the output results for a streaming application.To this end, we propose a new performance metric-coflow age (CA), for coflows generated by distributed streaming applications. The CA tracks the longest time-since-last-service among all flows in a coflow. In such a context, we consider a data center network with multiple coflows that continuously transmit packets between their source-destination pairs and address the problem of minimizing the average long-term CA while simultaneously satisfying the throughput constraints from the coflows. To solve this problem efficiently, we design a randomized algorithm and a drift-plus-age algorithm, and show that they can make the average long-term CA to achieve nearly two times and arbitrarily close to the optimal value, respectively. Through extensive simulations, we further demonstrate that both of the proposed algorithms can significantly reduce the CA of coflows, without violating the throughput requirement of any coflow, when compared to the state-of-the-art solution.
Wenxin Li 0001, Xu Yuan 0001, Wenyu Qu, Heng Qi, Xiaobo Zhou 0003, Sheng Chen 0015, Renhai Xu
INFOCOM6
2020 Fast and Accurate Detection of Unknown Tags for RFID Systems - Hash Collisions are Desirable
abstract
Unknown RFID tags appear when tagged items are not scanned before being moved into a warehouse, which can even cause serious security issues. This paper studies the practically important problem of unknown tag detection. Existing solutions either require low-cost tags to perform complex operations or beget a long detection time. To this end, we propose the Collision-Seeking Detection (CSD) protocol, in which the server finds out a collision-seed to make massive known tags hash-collide in the last $N$ slots of a time frame with size $f$ . Thus, all the leading ${f-N}$ pre-empty slots become useful for detection of unknown tags. A challenging issue is that, computation cost for finding the collision-seed is very huge. Hence, we propose a supplementary protocol called Balanced Group Partition (BGP), which divides tag population into $n$ small groups. The group number $n$ is able to trade off between communication cost and computation cost. We also give theoretical analysis to investigate the parameters to ensure the required detection accuracy. The major advantages of our CSD+BGP are two-fold: (i) it only requires tags to perform lightweight operations, which are widely used in classical framed slotted Aloha algorithms. Thus, it is more suitable for low-cost tags; (ii) it is more time-efficient to detect the unknown tags. Simulation results reveal that CSD+BGP can ensure the required detection accuracy, meanwhile achieving $1.7\times $ speedup in the single-reader scenarios and $3.9\times $ speedup in the multi-reader scenarios than the state-of-the-art detection protocol.
Xiulong Liu 0001, Sheng Chen 0015, Jia Liu 0008, Wenyu Qu, Fengjun Xiao, Alex X. Liu, Jiannong Cao 0001, Jiangchuan Liu
IEEE/ACM Trans. Netw.2
2018 How to Set Timeout: Achieving Adaptive Load Balance in Asymmetric Topology Based on Flowlet Switching
abstract
Traditional schemes achieving load balancing in asymmetric topology, which need to maintain global or local congestion information, turn out to be complicated to implement. One recent research has verified that flowlet switching is more simple and efficient to achieve adaptive load balancing in asymmetric topology. Nevertheless, one tricky problem lies in determining the flowlet timeout value, δ. Setting it too small would risk reordering issue while setting it too large would reduce flowlet opportunities. In this paper, by formulating the timeout setting problem with a stationary distribution of Markov chain, we give a theoretical reference for setting an appropriate timeout value in flowlet switching based load balancing scheme. Then, we implement a flowlet switching based load balancing scheme, called EasyLB, by extending OpenFlow protocol. Experiment results show that, by setting timeout value following the preceding theoretical reference, EasyLB is adaptive to asymmetric topology and achieves fast convergence of load balancing after link failures.
Zhiqiang Guo, Xiaodong Dong, Sheng Chen 0015, Xiaobo Zhou 0003, Keqiu Li
IPCCC3
2018 More Requests, Less Cost: Uncertain Inter-Datacenter Traffic Transmission with Multi-Tier Pricing
Xiaodong Dong, Sheng Chen 0015, Laiping Zhao, Xiaobo Zhou 0003, Heng Qi, Keqiu Li
J. Comput. Sci. Technol.2