Ju Ren 0001

dblp:00/468-1 · DBLP profile ↗
← Back
225ranked-venue papers
8as first author
168since 2021 · last 2026
0000-0003-2782-183XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 143 · 6 first-author · 108 since 2021Systems, architecture and hardware · 30 · 25 since 2021Security and privacy · 15 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MSCFL: Model Structure-Aware Clustered Federated Learning for System Heterogeneity and Data Drift
abstract
Federated Learning (FL) faces significant challenges arising from both data and system heterogeneity. While Clustered Federated Learning (CFL) mitigates data heterogeneity by grouping clients with similar data distributions, it remains vulnerable to system heterogeneity, which can slow convergence due to performance disparities among clients. Moreover, data drift may degrade clustering accuracy and training efficiency over time. In this work, we propose a Model Structure-aware Clustered Federated Learning (MSCFL) framework that simultaneously addresses the issues of data heterogeneity, system heterogeneity, and data drift. MSCFL incorporates model pruning (MP) into the CFL framework to enhance training efficiency under system heterogeneity. To enable this integration, we address the key challenge of performing effective clustering based on heterogeneous, pruned local models with varying structures. To this end, we design a model structure-based similarity computation algorithm to integrate CFL with MP. To effectively address data drift, we propose a dynamic cluster migration strategy that efficiently monitors model structures via Hamming Distance and triggers re-clustering only when necessary. Extensive experimental results show that MSCFL improves the accuracy and convergence speed of cluster models, outperforming traditional CFL in various settings.
Yang Xu 0013, Zifeng Xu, Cheng Zhang 0035, Ju Ren 0001, Yaoxue Zhang
AAAI5
2026 Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
abstract
Deploying Large Language Models (LLMs) on mobile devices faces the challenge of insufficient performance in smaller models and excessive resource consumption in larger ones. This paper highlights that mobile Neural Processing Units (NPUs) have underutilized computational resources, particularly their matrix multiplication units, during typical LLM inference. To leverage this wasted compute capacity, we propose applying parallel test-time scaling techniques on mobile NPUs to enhance the performance of smaller LLMs. However, this approach confronts inherent NPU challenges, including inadequate hardware support for fine-grained quantization and low efficiency in general-purpose computations. To overcome these, we introduce two key techniques: a hardware-aware tile quantization scheme that aligns group quantization with NPU memory access patterns, and efficient LUT-based replacements for complex operations such as Softmax and dequantization. We design and implement an end-to-end inference system that leverages the NPU's compute capability to support test-time scaling on Qualcomm Snapdragon platforms. Experiments show our approach brings significant speedups: up to 19.0× for mixed-precision GEMM and 2.2× for Softmax. More importantly, we demonstrate that smaller models using test-time scaling can match or exceed the accuracy of larger models, achieving a new performance-cost Pareto frontier.
Zixu Hao, Jianyu Wei, Tuowei Wang, Minxing Huang, Huiqiang Jiang, Shiqi Jiang 0002, Ting Cao 0003, Ju Ren 0001
EuroSys8
2026 CSVAR: Enhancing Visual Privacy in Federated Learning via Adaptive Shuffling Against Overfitting
Zhenya Ma, Yan Zhang 0073, Donghua Cai, Qiushi Li 0002, Yongheng Deng, Ye Zhang 0033, Ju Ren 0001, Xuemin Shen
ICC10
2026 EarAuth: Towards Practical Cardiac Vibration Authentication on COTS Wireless Earbuds
Yongjian Fu 0004, Wenpeng Zhu, Yingjun Wu, Hao Pan 0003, Guanbo Wang, Yongheng Deng, Yaoxue Zhang, Ju Ren 0001
INFOCOM11
2026 VoLLM: Smoothness-aware Serving of LLM-powered Voice Q&A via Adaptive Preemption
Yuanchun Li 0003, Ju Ren 0001, Lichen Pang, Shansong Yang, Yunxin Liu 0001
IPDPS3
2026 Revisiting Redundancy in Diffusion Transformers: A Temporal-Spatial Joint Caching Strategy for Efficient Sampling
abstract
Diffusion Transformers (DiTs) achieve impressive generative performance but suffer from significant inference latency. Feature caching–based acceleration methods reduce total computation by reusing results from earlier timesteps, but they largely ignore that temporal redundancy is dynamic and inconsistent across timesteps. Our analysis reveals this variability. More crucially, we identify a previously underexplored form of efficiency, namely spatial redundancy, characterized by high similarity between adjacent transformer blocks within the same timestep. Motivated by this dual-dimensional redundancy, we propose Temporal-Spatial Joint Cache, a training-free inference acceleration strategy that dynamically determines optimal reuse operations across temporal and spatial dimensions. Our approach features a redundancy-guided operation selector that estimates local feature stability using second-order divided differences, enabling fine-grained decisions between full computation, temporal cache, and spatial cache. Furthermore, we use interpolation-based feature prediction to capture local feature evolution for more accurate reuse. In addition, we propose a bounded cache distance control mechanism to mitigate error accumulation from excessive reuse. Together, these components allow our method to deliver substantial inference speedups without retraining or compromising generation fidelity, offering a new perspective on efficiency in diffusion transformer inference.
Chenxi Du, Yongheng Deng, Ju Ren 0001, Yaoxue Zhang
KDD (1)3
2026 AgentProg: Empowering Long-Horizon GUI Agents with Program-guided Context Management
abstract
The rapid development of mobile GUI agents has stimulated growing research interest in long-horizon task automation. However, building agents for these tasks faces a critical bottleneck: the reliance on ever-expanding interaction history incurs substantial context overhead. Existing context management and compression techniques often fail to preserve vital semantic information, leading to degraded task performance. We propose AgentProg, a program-guided approach for agent context management that reframes the interaction history as a program with variables and control flow. By organizing information according to the structure of program, this structure provides a principled mechanism to determine which information should be retained and which can be discarded. We further integrate a global belief state mechanism inspired by Belief MDP framework to handle partial observability and adapt to unexpected environmental changes. Experiments on AndroidWorld and our extended long-horizon task suite demonstrate that AgentProg has achieved state-of-the-art success rates on these benchmarks. More importantly, it maintains robust performance on long-horizon tasks while baseline methods experience catastrophic degradation. Our system is open-sourced at https://github.com/MobileLLM/AgentProg.
Shizuo Tian, Hao Wen 0004, Shanhui Zhao, Guohong Liu 0002, Ju Ren 0001, Yunxin Liu 0001, Yuanchun Li 0003
MobiSys7
2026 Bridging Storage and Execution: A Semantic Virtual Bus for On-Demand Application Streaming
Yaoxue Zhang, Ju Ren 0001
NSDI4
2026 RAN-Aware Delay Compensation for Delay-Sensitive Protocols in Cellular Networks
abstract
Delay-based protocols rely on end-to-end delay measurements to detect network congestion. However, in cellular networks, Radio Access Network (RAN) buffers introduce significant delays unrelated to congestion, fundamentally challenging these protocols’ assumptions. We identify two major types of RAN buffers - retransmission buffers and uplink scheduling buffers - that can introduce delays comparable to congestion-induced delays, severely degrading protocol performance. We present CellNinjia, a software-based system providing real-time visibility into RAN operations, and Gandalf, which leverages this visibility to systematically handle RAN-induced delays. Unlike existing approaches that treat these delays as random noise, Gandalf identifies specific RAN operations and compensates for their effects. Our evaluation in commercial 4G LTE and 5G networks shows that Gandalf enables substantial performance improvements - up to 7.49 × for Copa and 9.53 × for PCC Vivace - without modifying the protocols’ core algorithms, demonstrating that delay-based protocols can realize their full potential in cellular networks.
Tianyang Zhang 0012, Ju Ren 0001, Kyle Jamieson, Yaxiong Xie
SenSys4
2026 PeriNet: Periodic Deep Learning Framework for Modality-Agnostic Privacy Preserving
Yan Zhang 0104, Yihong Song, Manzhou Li, Qiushi Li 0002, Qi Li 0002, Ju Ren 0001
WWW6
2026 A Server-Side Model Intellectual Property Protection Method for Federated Learning Against Model Theft
abstract
Federated Learning (FL) has gained significant attention for enabling collaborative model training while preserving data privacy. However, protecting the intellectual property (IP) of models in FL, particularly against model theft by malicious clients, remains a critical challenge. Existing works often employ watermarking techniques to embed watermarks into models for ownership verification, but most of them can only verify ownership after the model has been stolen and cannot proactively defend against malicious clients attempting to steal the model. To address this limitation, this paper proposes FedLock, a novel server-side watermarking mechanism designed to resist model theft and safeguard the IP of global models. Specifically, FedLock leverages Split Federated Learning (SFL) to partition the model, effectively preventing malicious clients from accessing the complete set of global model parameters. To enhance security, FedLock incorporates an autoencoder for label encoding, safeguarding the server-side model from reconstruction attacks. Furthermore, FedLock introduces an additional backdoor client to embed a black-box watermark into the global model, enabling remote verification of model ownership. Experimental results demonstrate that FedLock achieves robust watermarking with minimal impact on model performance, effectively resisting various attacks, including model theft, pruning, and fine-tuning.
Wenxiong Chen, Xuantao Tang, Dan Wang 0031, Ju Ren 0001
IEEE Internet Things J.4
2026 LiBre: Toward Motion-Resilient Contactless Respiration Monitoring Using Mobile LiDAR
abstract
In this paper, we present LiBre, a LiDAR-based system for real-time respiration monitoring that remains accurate and continuous even under device motion. LiBre addresses the core challenge of disentangling large-scale device movement from subtle thoracoabdominal motions by integrating three key components: (i) an Object-Centric Feature Extraction module that produces clean, geometrically consistent human point clouds and enables multi-target sensing with minimal environmental interference; (ii) a Contrastive Registration framework that combines standard Iterative Closest Point (ICP) and masked Iterative Closest Point (MICP) to decouple device motion from respiration-induced displacements; and (iii) a Directional Residual Projection strategy that automatically estimates the RoI and projects residual motion along the dominant respiratory axis, eliminating the need for manual annotation. Following a brief autonomous stationary initialization phase to establish the respiratory RoI and optimized weights, we implement a complete prototype and validate its real-time performance. Experiments with 10 participants demonstrate that LiBre achieves respiration monitoring with < 1 BPM error at sensing distances up to 4 m, under device motion speeds up to 30 cm/s, and at orientation angles up to 60°, while supporting multi-person scenarios. The system processes each frame within 120 ms, meeting the requirements for real-time mobile health monitoring in practical applications.
Junying Hu, Yongjian Fu 0004, Xinyi Li 0005, Yaoxue Zhang, Ju Ren 0001
IEEE Internet Things J.7
2026 Nappa: NNA-Compatible and Privacy-Preserving DNN Training Framework via Vector Decomposition
abstract
How to preserve the data privacy during the training of deep neural network (DNN) is a key security concern in the artificial intelligence era. However, most existing solutions based on homomorphic encryption and Trusted Execution Environment (TEE) are incompatible with heterogeneous Neural Network Accelerators (NNAs), leading to significant performance loss. We propose a novel method based on vector decomposition to allocate operators across different NNAs, ensuring both throughput and privacy simultaneously. Furthermore, based on this approach, we have designed a compiler that automatically converts front-end model descriptions into backend encrypted computation graphs, which is running securely over trusted and untrusted hardware. This compiler heuristically determines the allocation scheme based on hardware affinity and cross-hardware communication costs, significantly reducing additional overhead. Experimental results demonstrate that our method does not incur extra accuracy costs and achieves a throughput significantly higher than existing methods. Deploying our approach at scale on a platform with a billion users, we have verified its negligible impact on real-world operations while ensuring the privacy protection capability for cross-domain data.
Yan Zhang 0002, Qiushi Li 0002, Ju Ren 0001, Yiqiao Liao, Jin Ouyang, Chengru Song, Honghuan Wu, Kaiqiao Zhan, Ben Wang 0006, Xu Chen 0004, Yaoxue Zhang
IEEE Trans. Dependable Secur. Comput.3
2026 A Fine-Tuning Data Recovery Attack on Generative Language Models via Backdooring
abstract
Generative language models (GLMs) are increasingly integrated into modern intelligent applications to power intelligent functionalities. Developers often fine-tune open-source GLMs on proprietary data and deploy them in real-world applications. In this paper, we reveal a novel model supply chain attack that exploits this workflow: by injecting backdoors into the source code of an open-source GLM, an adversary can induce the model to memorize fine-tuning data and later regenerate it via crafted prompts. We propose LURE, a new backdoor-based data recovery attack that exploits memorization capabilities of fine-tuned models. During fine-tuning, LURE stealthily injects unique and attacker-enumerable hash prompts, and incorporates a Position-Decay Weighted Aligned Cross-Entropy Loss into the original fine-tuning loss, strengthening the association between injected prompts and corresponding data samples for effective data recovery. To achieve stealthy and transparent attack injection, LURE employs a stealthy backdoor within the model’s source code, enabling automatic injection of hash prompts during fine-tuning and thus maintaining the user’s original fine-tuning workflow. LURE also proposes several optimizations to maintain minimal impact on the performance of the original task and external training state. Extensive evaluations demonstrate the remarkable efficacy of LURE, achieving a 45%-68% data recovery rate while maintaining the attack’s transparency, stealthiness, and showcasing its ability to evade existing defenses.
Zhenya Ma, Yongheng Deng, Ziqing Qiao, Quan Zhang 0003, Chijin Zhou, Fan Wu 0014, Yaoxue Zhang, Ju Ren 0001
IEEE Trans. Inf. Forensics Secur.8
2026 SubLoRa: High-Throughput LoRa Backscatter Communication
abstract
Ambient LoRa backscatter enables long-range communication due to its long-period symbol. Most of the existing works struggle to balance range and throughput: systems with symbol-level modulation offers long transmission range at the cost of low data rate, while systems with high modulation efficiency suffer from limited transmission distance due to weak signals. We propose SubLoRa, which significantly improves throughput while maintaining long-range communication. SubLoRa achieves the high-rate modulation by the proposed Subchirp Frequency Offset Modulation (SFOM), which divides a chirp into multiple subchirps each being shifted by frequency. We propose a prewaveform sampling strategy that enables SFOM with low power. For decoding, we propose a Frequency-Difference Recombination of Chirp (FDRC) demodulation based on time-domain correlation, which enables reliable decoding of low-power signals in long-range links. We implement SubLoRa and conduct extensive evaluation. The results show SubLoRa can achieve up to 29.66× throughput gain and 7.54× throughput gain compared with the State-Of-The-Art (SOTA) LoRa backscatter system PLoRa and Pacim, respectively.
Jingyi Bai, Caihui Du, Jihong Yu, Ju Ren 0001, Haipeng Yao
IEEE Trans. Mob. Comput.4
2026 Mobile and Multi-Device Wireless Charging
abstract
Wireless charging is a cornerstone technology for next-generation mobile and ubiquitous computing. However, its practical deployment has long been constrained by short range, poor flexibility, and lack of support for dynamic multi-device scenarios. In this paper, we propose ChargeX—a system that enables long-range and mobility-resilient wireless charging for multiple small devices. ChargeX pioneers the integration of metasurface-assisted magnetic beamforming, a high-frequency compact transceiver design, and a real-time closed-loop feedback-control mechanism. It further advances the field by introducing a joint optimization framework for dynamically allocating energy across mobile receivers with heterogeneous priorities and spatial-temporal demands. Experimental results demonstrate that it achieves meter-level charging distance, real-time response to device movement, and efficient coordination among multiple receivers, significantly outperforming state-of-the-art prototypes.
Bozhong Yu, Yongjian Fu 0004, Ju Ren 0001, Hao Pan 0003, Jeremy Gummeson, Ling Wang 0007, Yaoxue Zhang
IEEE Trans. Mob. Comput.4
2026 Learning Based Versatile Voice Eavesdropping Prevention for Mobile Devices
abstract
Voice-enabledmobile applications(apps) are exploding in popularity as they could be manipulated with voice commands to achieve convenient man-machine interaction. These voice-enabled apps also raise security and privacy concerns about whether they would maliciously invoke microphones to realize voice eavesdropping. To explore this issue, in this work, we design baleful apps to access the microphone covertly, the results of test studies demonstrate that covert eavesdropping attacks can bypass existing device detection schemes as well as are unnoticeable to human users. To prevent the covert voice eavesdropping attack, we propose a versatilemicrophone icon detection(MicID) scheme inspired by the groundtruth that authorization of the voice function requires the user to touch the specific microphone icon in most of voice-based apps. Specifically, we devise a deep learning model,lightweight YOLO(L-YOLO), to locate the microphone icon on the screen quickly and accurately. By determining whether the located microphone icon is touched by the user, we can judge whether the current microphone access belongs to the app's normal operation or illegal eavesdropping. Finally, we conduct extensive experiments by deploying the scheme on real devices and collecting dataset. The evaluation results show that the proposed MicID scheme achieves more than 99% accuracy with low computation cost.
Wenbin Huang 0003, Ju Ren 0001, Hangcheng Cao, Hongbo Jiang 0001, Panlong Yang, Zhangjie Fu 0001
IEEE Trans. Mob. Comput.2
2026 Blockchain-Enabled Storage Resource Trading for Collaborative Edges
abstract
As edge devices grow smarter and application scenarios become more diverse, users' demands for lower latency and higher efficiency in data storage and processing have risen sharply. Individual edge devices and nodes are no longer sufficient to meet these expanding storage requirements. Consequently, developing efficient, low-latency, and cost-effective solutions for collaborative storage across edge devices and nodes has become a critical challenge. In this paper, we present a framework for the transaction and pricing of storage resources in an edge computing environment involving multiple edge service providers, to address trust and incentive issues in storage resource collaboration. Firstly, we propose a secure and decentralized storage resource trading mechanism by leveraging blockchain technology and smart contracts. We introduce Proof of Transaction Expectation (PoTE), an efficient, reliable, and lightweight consensus mechanism, to ensure transaction transparency, openness, and non-repudiation. Secondly, we introduce a game theory-based storage resource pricing model, where a leader interacts with multiple followers to optimize profits while maintaining service quality. To address dynamic pricing and storage resource allocation problems under incomplete information, we propose the Stackelberg Game Approach based on Multi-Agent Reinforcement Learning (SGA-MARL), which formulates the optimal pricing and trading share decisions in the two-stage Stackelberg game as a stochastic Markov Decision Process (MDP). Simulations and prototype testing validate the effectiveness of the proposed system, with results showing that the PoTE consensus achieves up to 40% higher throughput than Proof-of-Work while reducing latency by over 50% compared to PBFT, and the SGA-MARL algorithm improves leader profit by approximately 30% and resource satisfaction rates by over 80% compared to baseline methods like MA-PPO and DQN.
Weimin Li 0002, Zhengmao Yan, Zeqiang Chen, Fan Wu 0014, Wenxiong Chen, Jianxun Liu 0001, Ju Ren 0001
IEEE Trans. Mob. Comput.8
2026 MagGuard: Detecting Mobile Eavesdropping via Built-In Magnetometers With Contrastive Learning
abstract
Protecting privacy-sensitive hardware usage on mobile devices is crucial. Although mobile operating systems (OSs) and smartphone manufacturers have set the permission settings, attackers can evade these defenses using covert methods, enabling malicious camera recording, microphone eavesdropping, and screen capture. Electronic devices emit unique yet weak electromagnetic interference (EMI) signals when accessing privacy-sensitive hardware. But, these signals are easily affected by foreground application activities and geomagnetic fluctuations caused by device movement. Our prior work showed that supervised learning can extract EMI features correlated with privacy hardware states from complex magnetometer readings, but it requires substantial labeled data, limiting practical deployment to new device models or OS versions. To eliminate this reliance on labeled data, this paper proposes a multimodal contrastive learning framework that leverages the device's built-in magnetometer and synchronized system logs as dual-modal inputs. Through self-supervised training, the framework can learn the intrinsic associations between EMI features and the operating states of privacy-sensitive hardware. Building on this, we design an EMI-based eavesdropping classifier that can analyze a user device's magnetometer readings offline to detect covert eavesdropping activities. Experimental results show that the proposed method can effectively identify eavesdropping behavior related to access to camera, microphone, and screen recording data. Testing across ten diverse mobile devices achieved an average classification accuracy of 89.1% on Android devices and 88.5% on iOS devices for identifying the specific hardware being eavesdropped upon.
Hao Pan 0003, Lanqing Yang, Yongjian Fu 0004, Yi-Chao Chen 0001, Guangtao Xue, Ju Ren 0001
IEEE Trans. Mob. Comput.6
2026 Two-Dimensional Stackelberg Game-Based Incentive Mechanism for Differential Private Federated Learning With Non-IID Data
abstract
Incentive mechanisms are essential for boosting client engagement in differential private federated learning (DP-FL). However, existing Stakelberg games-based incentive mechanisms typically assume that client decisions are one-dimension and that data is independent and identically distributed (IID) across clients. In reality, data distributions are often non-IID and clients have two-dimensional resources decisions, including data quantity and privacy. Therefore, in this paper, we present a novel two-dimensional Stackelberg game-based incentive mechanism for DP-FL with non-IID data, aiming to maximize the total utility of clients and server by seeking a balance between the clients' two-dimensional decisions and the server's payment. Specifically, we first formulate the utility functions of both server and clients under two-dimensional decisions and then model the interactions between server and clients as a single-leader-multiple-followers Stackelberg game. To derive the optimal decisions that maximize their utilities, we theoretically prove the existence of a Stackelberg equilibrium between server and clients. Due to the difficulty to directly calculate the Stackelberg equilibrium, we propose a bi-level multi-agent reinforment learning algorithm to learn the optimal decisions for both server and clients by trial and error. Extensive simulation results demonstrate that our proposed method outperforms the baselines in terms of total utility.
Dan Wang 0031, Xiaoyi Pang, Jiahui Hu 0001, Sheng Yue 0001, Ju Ren 0001
IEEE Trans. Mob. Comput.5
2026 MobiFuse: A High-Precision On-Device Depth Perception System With Multi-Data Fusion
abstract
We present MobiFuse, a high-precision depth perception system on mobile devices that combines dual RGB and Time-of-Flight (ToF) cameras. To achieve this, we leverage physical principles from various environmental factors to propose the Depth Error Indication (DEI) modality, characterizing the depth error of ToF and stereo-matching. Furthermore, we employ a progressive fusion strategy, merging geometric features from ToF and stereo depth maps with depth error features from the DEI modality to create precise depth maps. Additionally, we create a new ToF-Stereo depth dataset,RealToF, to train and validate our model. Our experiments demonstrate that MobiFuse excels over baselines by significantly reducing depth measurement errors by up to 77.7%. It also showcases strong generalization across diverse datasets and proves effectiveness in two downstream tasks: 3D reconstruction and 3D segmentation. The demo video of MobiFuse in real-life scenarios is available at the de-identified YouTube link.
Tingting Long, Ju Ren 0001, Yunxin Liu 0001, Yudong Zhao, Yaoxue Zhang, Youngki Lee 0001
IEEE Trans. Mob. Comput.5
2026 SRDrone: LLM-Driven Self-Refinement for Embodied Drone Task Planning
abstract
We introduceSRDrone, a novel system designed for self-refinement task planning in industrial-grade embodied drones.SRDroneincorporates two key technical contributions: First, it employs a continuous state evaluation methodology to robustly and accurately determine task outcomes and provide explanatory feedback. This approach supersedes conventional reliance on single-frame final-state assessment for continuous, dynamic drone operations. Second,SRDroneimplements a hierarchical Behavior Tree (BT) modification model. This model integrates multi-level BT plan analysis with a constrained strategy space to enable structured reflective learning from experience. Experimental results demonstrate thatSRDroneachieves a 44.87% improvement in Success Rate (SR) over baseline methods. Furthermore, real-world deployment utilizing an experience base optimized through iterative self-refinement attains a 96.25% SR. By embedding adaptive task refinement capabilities within an industrial-grade BT planning framework,SRDroneeffectively integrates the general reasoning intelligence of Large Language Models (LLMs) with the stringent physical execution constraints inherent to embodied drones. Code is available athttps://github.com/ZXiiiC/SRDrone.
Tingting Long, Xunhua Dai, Yongjian Fu 0004, Ju Ren 0001, Yaoxue Zhang
IEEE Trans. Mob. Comput.8
2026 Toward Communication-Efficient and Data-Free Collaborative Fine-Tuning Between Small and Large Language Models
abstract
While large language models (LLMs) exhibit impressive general capabilities, their performance on domainspecific tasks often requires fine-tuning with private data that cannot be shared due to privacy constraints. Directly deploying LLMs on resource-constrained clients for local fine-tuning is impractical due to their significant computation and communication costs. In addition, pre-trained LLMs are valuable intellectual property, and model owners are reluctant to distribute full model weights. To address these challenges, we proposeCoT-LM, a communication-efficient, computation-light, and data-free framework for collaborative fine-tuning between small (SLMs) and large language models (LLMs). InCoT-LM, clients fine-tune lightweight SLMs locally without uploading models or private data. These SLMs provide task-specific feedback to guide server-side LLM enhancement via an efficient communication protocol that exchanges only lightweight synthetic data and feedback. The framework supports both synchronous and asynchronous collaboration and enables mutual enhancement: the LLM improves its task-specific capabilities, while clients benefit from refined synthetic data or distilled knowledge. Extensive experiments demonstrate thatCoT-LMsignificantly boosts natural language understanding (NLU, up to 18.3% for LLMs and 8.0% for SLMs) and natural language generation (NLG, up to 31.7% for LLMs) performance across diverse tasks while preserving data privacy, model intellectual property, and generalization capabilities, achieving significant reductions in computation and communication overhead.
Zhenya Ma, Yongheng Deng, Ziqing Qiao, Yongjian Fu 0004, Sheng Yue 0001, Ju Ren 0001
IEEE Trans. Netw.6
2026 FOVA: Offline Federated Reinforcement Learning With Mixed-Quality Data
abstract
Offline Federated Reinforcement Learning (FRL), a marriage of federated learning and offline reinforcement learning, has attracted increasing interest recently. Albeit with some advancement, we find that the performance of most existing offline FRL methods drops dramatically when provided with mixed-quality data, that is, the logging behaviors (offline data) are collected by policies with varying qualities across clients. To overcome this limitation, this paper introduces a new vote-based offline FRL framework, named FOVA. It exploits avote mechanismto identify high-return actions during local policy evaluation, alleviating the negative effect of low-quality behaviors from diverse local learning policies. Besides, building on advantage-weighted regression (AWR), we construct consistent local and global training objectives, significantly enhancing the efficiency and stability of FOVA. Further, we conduct an extensive theoretical analysis and rigorously show that the policy learned by FOVA enjoys strict policy improvement over the behavioral policy. Extensive experiments corroborate the significant performance gains of our proposed algorithm over existing baselines on widely used benchmarks.
Nan Qiao 0008, Sheng Yue 0001, Ju Ren 0001, Yaoxue Zhang
IEEE Trans. Netw.3
2025 ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries
abstract
Existing research on human-centric video understanding typically focuses on analyzing specific moments or entire videos. However, many applications require higher precision at the frame level. In this work, we propose a novel task, BestShot, which aims to locate highlight frames within human-centric videos through language queries. This task requires not only a deep semantic understanding of human actions but also precise temporal localization. To support this task, we introduce the BestShot Benchmark. The benchmark is meticulously constructed by combining human-annotated highlight frames, duration labels and detailed textual descriptions. These descriptions cover three critical elements: (1) Visual content; (2) Fine-grained actions; and (3) Human pose descriptions. Together, these elements provide the necessary precision to identify the exact highlight frames in videos. To tackle this problem, we have collected two distinct datasets: (i) ShotGPT4o Dataset, which is algorithmically generated by GPT-4o and (ii) Image-SMPLText Dataset, which features large-scale and accurate per-frame pose descriptions using PoseScript and existing pose estimation datasets. Based on these datasets, we present a strong baseline model, ShotVL, fine-tuned from InternVL, specifically for BestShot. We highlight the impressive zero-shot capabilities of our model and offer comparative analyses with existing state-of-the-art (SOTA) models. ShotVL demonstrates a significant 64% improvement over InternVL on the BestShot Benchmark and a notable 68% improvement on the THUMOS14 Benchmark, while maintaining SOTA performance in general image classification and retrieval.
Wangyu Xue, Chen Qian 0006, Wentao Liu 0002, Ju Ren 0001, Siming Fan, Yaoxue Zhang
AAAI6
2025 Neuralink: Fast on-Device LLM Inference with Neuron Co-Activation Linking
abstract
Large Language Models (LLMs) have achieved remarkable success across various domains, yet deploying them on mobile devices remains an arduous challenge due to their extensive computational and memory demands.While lightweight LLMs have been developed to fit mobile environments, they suffer from degraded model accuracy.In contrast, sparsitybased techniques minimize DRAM usage by selectively transferring only relevant neurons to DRAM while retaining the full model in external storage, such as flash.However, such approaches are critically limited by numerous I/O operations, particularly on smartphones with severe IOPS constraints.In this paper, we propose Neuralink, a novel approach that accelerates LLM inference on smartphones by optimizing neuron placement in flash memory.Neuralink leverages the concept of Neuron Co-Activation, where neurons frequently activated together are linked to facilitate continuous read access and optimize I/O efficiency.Our approach incorporates a two-stage solution: an offline stage that reorganizes neuron placement based on co-activation patterns, and an online stage that employs tailored data access and caching strategies to align well with hardware characteristics.Evaluations conducted on a variety of smartphones and LLMs demonstrate that Neuralink achieves on average 1.49× improvements in end-to-end latency compared to the state-of-the-art.As the first solution to optimize storage placement under sparsity, Neuralink explores a new * Both authors contributed equally to this research.
Tuowei Wang, Ruwen Fan, Minxing Huang, Zixu Hao, Kun Li 0016, Ting Cao 0003, Youyou Lu, Yaoxue Zhang, Ju Ren 0001
ASPLOS (3)9
2025 ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning
abstract
Ziqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang, Lai Wei, Guanbo Wang, Fandong Meng, Jie Zhou, Ju Ren, Yaoxue Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Ziqing Qiao, Yongheng Deng, Jiali Zeng, Guanbo Wang, Fandong Meng, Jie Zhou 0016, Ju Ren 0001, Yaoxue Zhang
EMNLP9
2025 Malva: A Jitter-Aware Online Pruning Framework for DNN Inference Tasks
abstract
In fields like autonomous driving, strict constraints are imposed on the computing latency of deep neural network (DNN) inference tasks on edge servers. However, it is typical for edge servers to execute multiple tasks in parallel to serve multiple users, causing severe latency jitter due to resource competition, which seriously affects timeliness. Existing works ignore the computing jitter and regard computing latency as a deterministic value, failing to meet the timeliness requirement. To address this issue, we propose Malva, a framework for finegrained online pruning for DNN tasks, allowing flexible pruning at runtime based on jitter conditions. Specifically, Malva first partitions the DNN model into blocks and applies early exiting and pruning methods to create block variants. Then, the Malva scheduler flexibly selects the variant to be executed or exits early according to urgency and jitter conditions. Moreover, we propose a novel urgency-aware prediction strategy to estimate the accuracy impact of variants with incomplete pathway information during scheduling. Stress testing shows Malva can strictly maintain a zero deadline miss rate and significantly increase the stress required to cause the first deadline miss while still outperforming state-of-the-art methods in accuracy.
Ziyan Fu 0001, Yongheng Deng, Yingjun Wu, Zhibo Wang 0001, Su Yao, Yaoxue Zhang, Ju Ren 0001
IWQoS8
2025 FedAF: Alignment-Augmented Fusion for Federated Multimodal Learning with Small Labels
abstract
Federated multimodal learning is an emerging advancement in artificial intelligence, enabling the integration of data from diverse modalities while preserving data privacy. However, limited labeled data and modality heterogeneity on the clients pose significant challenges for effective federated multimodal model training. To address these challenges, this paper introduces FedAF, a novel alignment-augmented fusion framework tailored for federated multimodal learning. FedAF extracts unbiased and complementary information from multiple modalities with small data, enabling effective modality fusion and feature alignment for improving system performance. The framework introduces a three-stage strategy. First, FedAF utilizes labeled data to create unbiased anchor points, addressing disparities in client feature distributions. Second, FedAF employs a weighted enhancement contrast fusion scheme to improve feature clustering and reduce feature overlap. Finally, a multimodal semisupervised algorithm mitigates data heterogeneity and overfitting. Extensive experiments demonstrate that FedAF significantly outperforms baseline methods, showcasing its effectiveness in federated multimodal learning scenarios.
Guanbo Wang, Yongheng Deng, Yingjun Wu, Xinyi Li 0005, Tuowei Wang, Yaoxue Zhang, Ju Ren 0001
IWQoS9
2025 CGMM: Non-Invasive Continuous Glucose Monitoring in Wearables Using Metasurfaces
abstract
Non-invasive continuous glucose monitoring for diabetes patients remains challenging despite ongoing interest. This paper presents CGMM, a novel non-invasive wireless glucose monitoring system integrated into wearable devices. It features a specially designed metasurface that couples with the wearable's antenna and tissue fluid beneath the skin, amplifying frequency response changes caused by subtle glucose concentration variations. To address individual tissue variability and optimize the passive metasurface design, we develop a tunable metasurface and a one-shot calibration method to obtain the impedance for optimal resonance in glucose sensing environments with unknown parameters. The calibrated impedance is then used for the inverse design and fabrication of an economical passive metasurface. We implement prototypes of CGMM and conduct extensive experimental evaluations. In human experiments involving ten participants using the prototype with LibreVNA, the overall performance is quantified with relative errors ranging from -5.02% to 6.93% and an RMSE of 9.65 mg/dL.
Hao Pan 0003, Yezhou Wang, Jiting Liu, Ruichun Ma, Lili Qiu, Yi-Chao Chen 0001, Guangtao Xue, Ju Ren 0001
MobiCom8
2025 WDNN: Weighted Diffractive Neural Network for Physical-layer RF Signal Processing
abstract
Diffractive neural networks (NNs) have garnered attention for directly implementing wireless signal processing at the physical layer. However, they are limited by a constrained weight learning space and activation functions, which restricts their data processing capabilities. To address this, we propose an RF circuit-based weighted diffraction NN (WDNN) that rivals digital NNs in processing ability. We design a weighted asymmetric RF coupler unit that, when stacked into a network, enables diffractive propagation with arbitrary connection weights. Additionally, an activation module is introduced that utilizes RF amplifiers operating in their nonlinear regions. We validate the effectiveness of the proposed WDNN through three tasks: 32-level amplitude modulated (AM) signal decoding, 31-class angle of arrival (AoA) estimation, and 2-class Wi-Fi based fall detection. After training, WDNN achieves the accuracy of 98.5%, 93.7%, and 90.8% in the AM decoding, AoA estimation, and fall detection tasks, respectively; while the diffractive NN SOTA achieves only 21.6%, 16.9%, and 63.3%. We also implement the prototypes of WDNN and SOTA, and real-world experimental results demonstrate that our method achieves an average accuracy improvement of up to 76.85% across various tasks compared to SOTA.
Yezhou Wang, Yongjian Fu 0004, Hao Pan 0003, Qinyun Hu, Lili Qiu, Yi-Chao Chen 0001, Guangtao Xue, Ju Ren 0001
MobiCom8
2025 FocusX: All-in-Focus Image Synthesis for Dynamic Scenes on Mobile Devices
abstract
We propose FocusX, the first mobile-deployable system achieving artifact-free all-in-focus synthesis in dynamic scenes. Our approach introduces three key innovations: 1) For focal stack acquisition, our depth prior-based dynamic focusing method that adaptively selects focus distances using real-time scene depth distribution analysis and depth-of-field constrained spatial clustering, reducing redundant captures while ensuring full depth coverage; 2) To reduce pixel misalignment caused by lens breathing, we adopt a one-time offline calibration to map the relationship between field-of-view and focus distance, aligning the images by cropping accordingly; 3) We design the Diff-MotionAIFNet, a conditional diffusion-based model that decouples moving-static components for artifact-free AIF reconstruction in dynamic scene while preserving scene fidelity. We further contribute DynaAIFSet, containing 5,500 dynamic scenes (120K images) for training and evaluation. Experiments show FocusX achieves state-of-the-art performance, outperforming baselines up by 59.6% in SSIM and 49.1% in PSNR, respectively. The deployment latency of FocusX is 4.8s on Honor Magic7 Pro. This work bridges computational photography theory with mobile implementation constraints, delivering practical AIF enhancement for user-generated content.
Pengkai Li, Fengzu Li, Wei Gao 0006, Sheng Yue 0001, Yaoxue Zhang, Ju Ren 0001
MobiCom8
2025 Towards Distance-Adaptive Wireless Charging
abstract
Wireless charging holds significant promise for IoT devices and transportation networks by facilitating convenient and autonomous power supply. Traditional wireless charging technologies have typically adhered to a singular approach, choosing between near-field coupling or far-field radiation. However, our investigations uncover that each method outperforms the other at specific distances. This insight leads us to integrating the advantages of both to enable rapid wireless charging across any distance within the charging range. For this vision, we poses an intriguing question: "Can we develop a system that supports both near-field and far-field charging simultaneously?"
Shuning Wang, Linghui Zhong, Yongjian Fu 0004, Sheng Yue 0001, Ju Ren 0001, Yaoxue Zhang
MobiSys7
2025 CrossLM: A Data-Free Collaborative Fine-Tuning Framework for Large and Small Language Models
abstract
While large language models (LLMs) are endowed with broad knowledge, their task-specific performance is often suboptimal. Fine-tuning LLMs with task-specific data from diverse nodes is necessary, but this data is typically safeguarded and not shared publicly due to privacy concerns. A common solution involves downstream nodes downloading the LLM locally and fine-tuning it with their proprietary data. However, owners often regard pre-trained LLMs as valuable assets and are reluctant to share them. Additionally, the significant computational resources required by LLMs make local fine-tuning impractical for many nodes. To mitigate these problems, this paper proposes CrossLM, a data-free collaborative fine-tuning framework for large and small language models. CrossLM enables resource-constrained nodes to train smaller language models (SLMs) using their private task-specific data. These SLMs are subsequently leveraged to promote the task-specific natural language generation and understanding capabilities of the LLMs. Simultaneously, the SLMs of nodes also benefit from enhancement by the fine-tuned LLMs. In this way, CrossLM avoids sharing private data and proprietary LLMs, and also reduces the resource requirements of nodes. Through extensive experiments across a range of benchmark tasks and popular language models, we demonstrate that CrossLM significantly boosts the task-specific performance of both LLMs and SLMs while preserving the generalization capabilities of LLMs.
Yongheng Deng, Ziqing Qiao, Ye Zhang 0033, Zhenya Ma, Yang Liu 0165, Ju Ren 0001
MobiSys6
2025 MetaGen: LLM-Driven Generative Framework for Intelligent Metasurface Element
abstract
Metasurfaces are a transformative class of artificial electromagnetic materials with significant potential in communication, sensing, and security. However, existing design methods require detailed physical properties as input and lack flexibility under complex constraints, limiting their applicability. In this paper, we propose MetaGen, a general, efficient, and user-friendly generation framework for intelligent metasurface elements. MetaGen employs a fine-tuned large language model to translate natural language instructions into formatted physical properties and integrates a diffusion-based model to generate metasurface elements. Furthermore, we develop a metasurface element dataset with granular frequency sampling and extended geometric parameters to enable MetaGen to learn the complex relationships between metasurface element geometries and electromagnetic responses. Experimental results demonstrate that MetaGen effectively satisfies complex constraints, achieving electromagnetic responses closely aligned with target specifications.
Xinyi Li 0005, Yue-Jiang Dong, Ju Ren 0001, Yaoxue Zhang
MobiSys5
2025 Gains: Fine-grained Federated Domain Adaptation in Open Set
abstract
Conventional federated learning (FL) assumes a closed world with a fixed total number of clients. In contrast, new clients continuously join the FL process in real-world scenarios, introducing new knowledge. This raises two critical demands: detecting new knowledge, i.e., knowledge discovery, and integrating it into the global model, i.e., knowledge adaptation. Existing research focuses on coarse-grained knowledge discovery, and often sacrifices source domain performance and adaptation efficiency. To this end, we propose a fine-grained federated domain adaptation approach in open set (Gains). Gains splits the model into an encoder and a classifier, empirically revealing features extracted by the encoder are sensitive to domain shifts while classifier parameters are sensitive to class increments. Based on this, we develop fine-grained knowledge discovery and contribution-driven aggregation techniques to identify and incorporate new knowledge. Additionally, an anti-forgetting mechanism is designed to preserve source domain performance, ensuring balanced adaptation. Experimental results on multi-domain datasets across three typical data-shift scenarios demonstrate that Gains significantly outperforms other baselines in performance for both source-domain and target-domain clients. Code is available at: https://github.com/Zhong-Zhengyi/Gains.
Zhengyi Zhong, Wenzheng Jiang, Weidong Bao 0001, Ji Wang 0002, Cheems Wang, Guanbo Wang, Yongheng Deng, Ju Ren 0001
NeurIPS8
2025 MODepth: Benchmarking Mobile Multi-frame Monocular Depth Estimation with Optical Image Stabilization
abstract
This paper presents MODepth, a multi-frame monocular depth estimation system based on the controlled motion of an optical image stabilization (OIS) module. By actively injecting acoustic signals, we induce regular translational movements of the OIS lens, resulting in controllable camera pose changes and simplifying inter-frame pose estimation. Leveraging multi-frame images captured under OIS-controlled lens movements, we design a high-precision depth estimation network, MODNet, and introduce the principal point offset estimation module and pose estimation modules to fully exploit geometric information across frames. To validate the effectiveness of our approach, we collect a new dataset MODdata with 1100 samples in nearly 220 indoor scenarios and benchmark our model as an OIS-based multi-frame depth estimation method, comparing it to ground truth obtained from a depth sensor and other state-of-the-art monocular depth estimation algorithms. Our method achieves competitive or superior performance compared to fully supervised baselines, reaching an RMSE of 0.439, which outperforms all evaluated methods, demonstrating that self-supervised fine-tuning with OIS-induced parallax is a viable alternative to ground-truth supervision. Code and dataset are available at: https://github.com/liangjindeamo-yuer/MODEPTH
Yu Lu 0022, Hao Pan 0003, Dian Ding, Jiatong Ding, Yongjian Fu 0004, Yi-Chao Chen 0001, Ju Ren 0001, Guangtao Xue
SIGGRAPH Asia7
2025 UltraPoser: Pushing the Limits of IMU-based Full-Body Pose Estimation with Ultrasound Sensing on Consumer Wearables
abstract
Figure 1: UltraPoser enables ubiquitous full-body pose estimation by integrating ultrasound sensing and IMU using commodity wearable devices.In addition to measuring IMU data, a smartphone and smartwatch are used to transmit and receive ultrasound signals.The extracted ultrasound features capture motions from joints without any attached devices and offer drift-free range measurements to complement IMU data for more accurate pose estimation.
Shuning Wang, Yongjian Fu 0004, Ju Ren 0001, Xinyu Zhang 0003, Akshay Gadre, Ke Sun 0012
UIST6
2025 JENGA: Enhancing LLM Long-Context Fine-tuning with Contextual Token Sparsity
Tuowei Wang, Kun Li 0016, Ting Cao 0003, Ju Ren 0001, Yaoxue Zhang
USENIX ATC5
2025 Cooperative and Adaptive Service Function Chain Deployment in UAV Swarm Networks
abstract
The rapid advancement of UAV swarm networks has enabled their widespread application across various domains, including disaster relief, environmental monitoring, and intelligent transportation. Collaboration among UAVs within a swarm is vital for efficient resource utilization and optimal performance across these diverse applications. To address diverse service demands, deploying service function chains (SFC) in UAV swarm networks facilitates the real-time implementation of services through efficient resource allocation and UAV cooperation, thereby enhancing network reliability and efficiency. However, traditional SFC deployment strategies struggle to achieve reliability and efficiency due to dynamic topology and limited resources. Additionally, Stochastic Network Calculus (SNC) derives end-to-end latency, guaranteeing quality of service (QoS) in UAV swarm networks. To navigate this issue, we propose a cooperative dynamic SFC deployment algorithm that combines hierarchical proximal policy optimization (HPPO) with an edge-enhanced dynamic graph attention network (EDGAT) for real-time network state extraction. The simulation results validate the effectiveness of our proposed algorithm, showcasing improvements in deployment success rate and long-term average revenue.
Fuchang Xu, Haipeng Yao, Ju Ren 0001, Jihong Yu, Zunliang Wang, Tianle Mai, Chenlang Jin
VTC2025-Fall3
2025 Visually secure image encryption: Exploring deep learning for enhanced robustness and flexibility
Wei Chen 0155, Wenjiang Ji, Yichuan Wang 0003, Ju Ren 0001, Guanglei Sheng, Xinhong Hei 0001
Expert Syst. Appl.4
2025 NC-Load: On-Demand Program Loading and Running for Computing Sharing Among IoT Devices
abstract
The number of Internet of Things (IoT) devices has increased rapidly in recent years, but lack effective methods to integrate their computational power. In this article, we propose NC-Load, which couples IoT devices into a multiprocessor system, allowing process scheduling across different devices to share their computing power and improve overall throughput. Specifically, NC-Load consists of three key designs, i.e., remote page fault (RPF), lightweight program cropping, and identical memory layout migration, contributing to three merits compared to existing systems: 1) high storage efficiency: the target device launches the program with a locally stored lightweight icon and leverages RPFs to retrieve the required code/data from the source device; 2) on-demand memory loading: only the required memory portions are transmitted when scheduling programs across different devices, which ensures quick recovery of the program; and 3) consistent memory layout: to ensure consistency of addresses after program offloading, the virtual memory area layout of the source device is migrated to the target device. We implement NC-Load on Linux 6.1 and conduct performance evaluation using unmodified programs and the N-Queens cases. The results demonstrate that NC-Load can achieve superior performance in terms of storage efficiency, program performance, memory usage, and throughput.
Yanhao Dong, Sijing Duan, Feng Lyu 0001, Yongmin Zhang, Ju Ren 0001, Yaoxue Zhang
IEEE Internet Things J.6
2025 FuzzyPR: Efficient Person Retrieval Using Fuzzy Semantic Descriptions Under Surveillance Scenario
abstract
In visual Internet of Things(VIoT), visualized sensors like surveillance cameras play as a key component in smart cities, generating a large amount of recorded data in real time. Under this scenario, semantic person retrieval aims to locate certain person from real-world surveillance images based on semantic descriptions. Most previous works were based on the assumption that the semantic description can provide enough details to locate a target person, namely “precise person retrieval”. However, this assumption cannot be satisfied in many real-world applications, where we only have fuzzy semantic descriptions and expect to pick out a set of targets. As the “fuzzy person retrieval” task has not been deeply explored by previous works, we propose a novel efficient one-stage method FuzzyPR. In our work, we perform multi-head visual-semantic feature alignment to against the asymmetry between the text and image information. To improve the model’s ability of ’inference and associative’ during the fuzzy retrieval process, we design a multi-granular semantic retrieval proxy task to improve the associative ability of the localization module. Experimental results demonstrate that FuzzyPR achieves the best retrieval accuracy and efficiency on fuzzy semantic retrieval task.
Chuanwen Luo, Ju Ren 0001, Yaoxue Zhang
IEEE Internet Things J.3
2025 StreamSys: A Lightweight Executable Delivery System for Edge Computing
abstract
Edge computing brings several challenges when it comes to data movement. First, moving large data from edge devices to the server is likely to waste bandwidth. Second, complex data patterns (e.g., traffic cameras) on devices require flexible handling. An ideal approach is to move code to data instead. However, since only a small portion of code is required, moving the executable as well as their libraries to the devices can be an overkill. While loading code on demand from remote such as NFS can be a stopgap, but on the other hand leads to low efficiency for irregular access patterns. This article presentsStreamSys, a lightweight executable delivery system that loads code on demand by redirecting the local disk IO to the server through optimized network IO. We employ a Markov-based prefetch mechanism on the server side. It learns the access pattern of code and predicts the block sequence for the client to reduce the network round trip. Meanwhile, server-sideStreamSysasynchronously prereads the block sequence from the disk to conceal disk IO latency beforehand. Evaluation shows that the latency ofStreamSysis up to 71.4% lower than the native Linux file system based on SD card and up to 62% lower than NFS in wired environments.
Zhenya Ma, Yinggang Gao, Sheng Yue 0001, Ju Ren 0001, Yaoxue Zhang
IEEE Trans. Cloud Comput.5
2025 Pushing Seamless Wireless Communication With Cross-Band Metasurfaces
Bozhong Yu, Ju Ren 0001, Jeremy Gummeson, Yaoxue Zhang
IEEE Trans. Commun.5
2025 MAML-RAL: Learning Domain-Invariant HOI Rules for Real-Time Video Matting
abstract
Real-time video matting is essential for applications like online video conferencing but faces challenges in human-object interaction (HOI) scenarios, known as the HOI-matting problem. This problem is challenging due to its open-recognition nature, where no dataset can cover the wide range of potential HOI cases, making it difficult for feature-learning-based methods to generalize effectively. To address this issue, we present an HOI-matting dataset and introduce a Model-Agnostic Meta-Learning-based rule-aware learning approach (MAML-RAL). MAML-RAL combines transfer learning and meta-learning to capture domain-invariant HOI rules, complemented by a fast local adaptation strategy to counter domain shifts and background interference. Our method achieves a mean intersection-over-union (mIoU) of 92.3%, outperforming current algorithms, with local adaptation further boosting performance to a remarkable mIoU of 95.84%.
Jiang Xin, Sheng Yue 0001, Ju Ren 0001, Feng Qian 0001, Yaoxue Zhang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Truthful and Dual-Direction Combinatorial Multi-Armed Bandit Scheme to Maximize Profit for Mobile Crowd Sensing
abstract
Nowadays, Mobile Crowd Sensing (MCS) has become a popular paradigm for large-scale data collection using ubiquitous mobile sensing devices. However, most existing works do not consider that requester's payments are unknown prior, and assume that workers are honest, which may not be true in practice. To address these problems, we propose a novel Truthful and Dual-direction Combinatorial Multi-Armed Bandit (TD-CMAB) scheme, which maximizes the total profit of the dual-direction platform for both the worker side and the requester side. Specifically, for the worker side, to overcome the problem that the platform is not clear whether sensed data are true, we propose a worker recruitment strategy that identifies and recruits honest workers at low cost through the Upper Confidence Bound (UCB) algorithm based on truth data discovery. For the requester side, where requesters’ payments are unknown prior, we model requester selection as a CMAB problem and solve it by the proposed adaptive UCB algorithm. Furthermore, we theoretically prove the worst regret bound of the TD-CMAB. Finally, we evaluate the effectiveness of the TD-CMAB scheme through extensive experiments using the Beijing taxi dataset.
Xiangwan Fu, Saiqin Long, Anfeng Liu, Ju Ren 0001, Bin Guo 0001, Zhetao Li
IEEE Trans. Dependable Secur. Comput.4
2025 Towards Fair Federated Learning via Unbiased Feature Aggregation
abstract
Federated learning (FL) is a distributed machine learning framework that enables multiple clients to collaboratively train models without raw data exchange. Prior studies on FL mainly focus on optimizing learning performance, enhancing privacy preservation, and improving attack resilience. However, little work studies how to mitigate the unfairness of federated trained models while unfair models would make discriminatory decisions toward certain groups or populations (e.g., favoring males over females), leading to serious ethical concerns. Thus, it is crucial to mitigate model unfairness in FL, yet challenging as this requires centralized access to each data point's fairness-sensitive information (e.g., race, gender), which is prohibited in FL. In this work, we propose a novel fair FL framework FedUFA, where the server can aggregate clients’ learned knowledge in an unbiased manner, to obtain fair and high-usability federated trained models. Specifically, to unearth the bias in clients’ local data and account for potentially heterogeneous local models, we propose a knowledge distillation-based FL scheme, where clients’ knowledge of learned features on a public dataset is amalgamated to the server for aggregation. We train an unbiased feature mapper at the server to remove fairness-sensitive latent features and extract fair representations from clients’ submitted raw features. In particular, we design an adversarial training method to train the mapper, which involves apredictoraiming to maximize the prediction accuracy on the FL task and adiscriminatorintending to help identify fairness-sensitive features. Extensive experiments on real-world datasets demonstrate the effectiveness of FedUFA.
Zeqing He, Zhibo Wang 0001, Xiaowei Dong, Peng Sun 0003, Ju Ren 0001, Kui Ren 0001
IEEE Trans. Dependable Secur. Comput.5
2025 MoiréComm: Secure Screen-Camera Communication Based on Moiré Cryptography
abstract
Quick Response (QR) codes have become increasingly popular for screen-camera communication due to their swift readability and widespread smartphone use. Nevertheless, they are vulnerable to privacy invasions from unauthorized photography. Addressing this, we propose a novel Moiré encryption technique-based secure screen-camera communication system, named MoiréComm. The Moiré encryption can enhance security by using distinct spatial frequency patterns for camouflage. The original QR code is revealed as a Moiré pattern only when the camera in a designated position, e.g., directly in front and 30 cm from the screen. From any other positions, only the camouflaged QR code can be seen. Decryption schemes are customized for different scenarios. The multi-frame approach achieves a decryption success of over 98.6% within 13.2 frames in handheld scenarios. Conditional generative adversarial network (cGAN)-based decryption method decodes the Moiré QR code images with a 98.8% success rate in 0.02 s within three frames and is also applicable in handheld scenarios. For fixed screen-camera setups, our fast decryption scheme achieves 99.4% success within two frames, with average 0.4 s latency. Significantly, the decryption rate plunges to 0% for surveillance cameras displaced by 20$^\circ$or more than$\ge$10 cm from the target position, demonstrating MoiréComm's resilience against attacks.
Hao Pan 0003, Yongjian Fu 0004, Yu Lu 0022, Feitong Tan, Yi-Chao Chen 0001, Ju Ren 0001
IEEE Trans. Dependable Secur. Comput.6
2025 Mitigating Voice Assistant Eavesdropping via Event Source Review on Mobile Devices
abstract
Voice assistants have been widely adopted for their ability to provide non-touch human-computer interaction. However, while they offer convenience, their continuous listening for specific wake-up words raises privacy concerns, as it may lead to eavesdropping on user conversations. To investigate this issue, we devised covert eavesdropping attacks by perturbing and replaying events generated during the user’s normal activation of the voice assistant. The results demonstrate the feasibility and harmfulness of such eavesdropping attacks. To counter these covert voice eavesdropping attacks, we propose an effective defense scheme called CrossUnwind. This scheme leverages the groundtruth that voice assistant wake-up requires hardware to generate and send wake-up events. Specifically, we designed a novel tombstone file parsing process and an accurate event discrimination algorithm to obtain detailed call station information of the wake-up event without compromising the system. This allows us to determine whether the current wake-up event was generated by hardware. We deployed CrossUnwind on real devices and compared it to well-known machine learning and deep learning methods. The results demonstrate that CrossUnwind can achieve high accuracy in eavesdropping detection with faster speeds and lower resource utilization.
Wenbin Huang 0003, Ju Ren 0001, Hangcheng Cao, Hongbo Jiang 0001, Zhangjie Fu 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Efficient Subcarrier-Level OFDM Backscatter Communications
abstract
Most of the existing OFDM backscatter systems adopt phase-modulated schemes to embed tag data, suffering from symbol-level modulation limitation, heavy synchronization accuracy reliance, and small tolerability to symbol time offset (STO) / carrier frequency (CFO) offset. We introduce SubScatter, the first subcarrier-level frequency-modulated OFDM backscatter which is able to tolerate bigger synchronization errors, STO, and CFO. The unique feature of SubScatter is our subcarrier shift keying (SSK) modulation. This method pushes the modulation granularity to the subcarrier by encoding and mapping tag data into different subcarrier patterns. We also design a tandem frequency shift (TFS) scheme that enables SSK with low cost and low power. Furthermore, we design SubScatter+ that shows these advantages while providing an even higher throughput without requiring more subcarrier patterns. We prototype and test SubScatter and SubScatter+, and the results show that our systems outperforms prior works in terms of effectiveness and robustness. Specifically, SubScatter has 743 kbps throughput that is 3.1 times and 14.9 times higher than RapidRider and MOXcatter, respectively. It also has a lower BER under noise and interferences which is over 6 times better than RapidRider or MOXcatter. Moreover, our proposed SubScatter+ could increase the throughput of SubScatter by 30%.
Caihui Du, Jihong Yu, Zhenyu Yan 0002, Ju Ren 0001, Yun Li 0001
IEEE Trans. Mob. Comput.4
2025 MagSpy: Revealing User Privacy Leakage via Magnetometer on Mobile Devices
abstract
Various characteristics of mobile applications (apps) and associated in-app services can reveal potentially-sensitive user information; however, privacy concerns have prompted third-party apps to restrict access to data related to mobile app usage. This paper outlines a novel approach to extracting detailed app usage information by analyzing electromagnetic (EM) signals emitted from mobile devices during app-related tasks. The proposed system, MagSpy, recovers user privacy information from magnetometer readings that do not require access permissions. This EM leakage becomes complex when multiple apps are used simultaneously and is subject to interference from geomagnetic signals generated by device movement. To address these challenges, MagSpy employs multiple techniques to extract and identify signals related to app usage. Specifically, the geomagnetic offset signal is canceled using accelerometer and gyroscope sensor data, and a Cascade-LSTM algorithm is used to classify apps and in-app services. MagSpy also uses CWT-based peak detection and a Random Forest classifier to detect PIN inputs. A prototype system was evaluated on over 50 popular mobile apps with 30 devices. Extensive evaluation results demonstrate the efficacy of MagSpy in identifying in-app services (96% accuracy), apps (93.5% accuracy), and extracting PIN input information (96% top-3 accuracy).
Yongjian Fu 0004, Lanqing Yang, Hao Pan 0003, Yi-Chao Chen 0001, Guangtao Xue, Ju Ren 0001
IEEE Trans. Mob. Comput.6
2025 MASA: Multimodal Federated Learning Through Modality-Aware and Secure Aggregation
abstract
As a promising paradigm, federated learning has been applied to multimodal sensing tasks due to its deployment convenience. However, the recent advances in multimodal federated learning emphasize learning a high-quality multimodal model but overlook the model usage requirements of massive unimodal clients. Moreover, the privacy risk in model sharing and client data heterogeneity impact the efficacy of federated learning. In this paper, we propose a novel multimodal federated learning system named MASA. As a departure from existing approaches, MASA simultaneously enhances the model learning efficiency of both multimodal and unimodal clients while ensuring their data privacy. First, we employ a gated cross-modal distillation scheme to achieve performance-aware knowledge transfer across modality-heterogeneous clients. To enhance the system security, MASA integrates a lightweight split-shuffle mechanism to realize the anonymization and encryption of model aggregation. Moreover, to reach personalized collaboration while protecting privacy, MASA features an attention-based spontaneous client clustering mechanism to form client cluster structures securely and distributedly. We evaluate our MASA on four public multimodal datasets for human activity recognition. The results show that our MASA outperforms leading multimodal federated learning methods on the model performance of both multimodal and unimodal clients.
Jialin Guo, Yongjian Fu 0004, Zhiwei Zhai, Xinyi Li 0005, Yongheng Deng, Sheng Yue 0001, Hao Pan 0003, Ju Ren 0001
IEEE Trans. Mob. Comput.9
2025 FedCRAC: Improving Federated Classification Performance on Long-Tailed Data via Classifier Representation Adjustment and Calibration
abstract
Federated learning has been a popular distributed training paradigm that enables to train a shared model with data privacy protection. However, non-Independent Identically Distribution and long-tailed data distribution characteristics across mobile devices results in evident performance degradation, especially for classification tasks. Although plenty of research studies devote to alleviating classification performance degradation caused by highly-skewed data distribution, they still cannot improve the distinguishability of model representation on hard-to-learn tail classes, and face obvious divergence of local classifiers in FL setting. To this end, we propose Federated Classifier Representation Adjustment and Calibration to improve the representation distinguishability of tail classes and achieve inter-client representation alignment with acceptable resource consumption on attaching operations. We first design a Class Similarity-Aware Margin matrix to enlarge class representation discrepancy and improve local classifier discriminability on tail classes during client-side local training process. To mitigate the divergence of local classifiers across clients, we further propose the Self Distillation Classifier Calibration to achieve the aggregated global classifier calibration with the assistance of generated pseudo representation samples via self-distillation manner. We conduct various experiments under wide-range long-tailed and heterogeneous data settings. Experimental results show that FedCRAC outperforms state-of-the-art methods in terms of accuracy and resource consumption.
Xujing Li, Min Liu 0001, Ju Ren 0001, Xuefeng Jiang 0001, Tianliu He
IEEE Trans. Mob. Comput.4
2025 MagicWrite: One-Dimensional Acoustic Tracking-Based Air Writing System
abstract
Air writing technology enhances text input for IoT, VR, and AR devices, offering a spatially flexible alternative to physical keyboards. Addressing the demand for such innovation, this paper presents MagicWrite, a novel system utilizing acoustic-based 1D tracking, which is suitable for mobile devices with existing speaker and microphone infrastructure. Compared to 2D or 3D tracking of the finger, 1D tracking eliminates the need for multiple microphones and/or speakers and is more universally applicable. However, challenges emerge when using 1D tracking for recognizing handwritten letters due to trajectory loss and inter-user writing variability. To address this, we develop a general conversion technique that transforms image-based text datasets (e.g., MNIST) into 1D tracking trajectory data, generating artificial datasets of tracking traces (referred to asTrackMNISTs) to bolster system robustness and scalability. These tracking datasets facilitate the creation of personalized user databases that align with individual writing habits. Combined with a kNN classifier, our proposed MagicWrite ensures high accuracy and robustness in text input recognition while simultaneously reducing computational load and energy consumption. Extensive experiments validate that our proposed MagicWrite achieves exceptional classification accuracy for unseen users and inputs in five languages, marking it as a robust solution for air writing.
Hao Pan 0003, Yongjian Fu 0004, Ye Qi, Yi-Chao Chen 0001, Ju Ren 0001
IEEE Trans. Mob. Comput.5
2025 Squeezer: Efficient Multi-DNN Inference for Edge Video Analytics via Cross-Model Scheduling
abstract
Video analytics at the edge is becoming increasingly prevalent in many scenarios, such as smart campuses and intelligent factories. These applications often consist of multiple subtasks, which necessitates the optimization for multi-DNN (Deep Neural Network) inference. Due to limited consideration over cross-model scheduling, current practices cannot fully leverage available computing resources, leading to suboptimal performance. To address this, we propose Squeezer, a multiDNN serving framework that holistically schedules multiple DNN models on an edge server with a single GPU. Squeezer decouples the cross-model scheduling into a two-layered approach, which involves (1) balanced operator grouping which partitions operators of multiple DNN models into groups, significantly reducing the scheduling complexity and (2) kernel scheduler which orchestrates parallel execution within each group by considering the interplay among kernels running in parallel, thereby enabling cross-model optimizations in multi-DNN inference. Performance evaluation results demonstrate that Squeezer outperforms state-of-the-art baselines, achieving up to 1.91× improvement in system throughput.
Lingxiao Ma, Ziyan Fu 0001, Yuanchun Li 0003, Ju Ren 0001, Yaoxue Zhang, Yunxin Liu 0001
IEEE Trans. Mob. Comput.6
2025 Blockchain-Enabled Multiple Sensitive Task-Offloading Mechanism for MEC Applications
abstract
As mobile devices proliferate and mobile applications diversify, Mobile Edge Computing (MEC) has become widely adopted to efficiently allocate computing resources at the network edge and alleviate network congestion. In the MEC initial phase, the absence of vital information presents challenges in devising task-offloading policies, and identifying malicious devices responsible for providing inaccurate feedback is complex. To fill in such gaps, we introduce a consortium blockchain-enabledCommitteeVoting basedTaskOffloadingModel (CVTOM) to collaboratively formulate resource allocation policies and establish deterrence against malicious servers producing erroneous results intentionally. Different voting principle mechanisms of each committee member are first designed in a Blockchain-enabled system which helps to represent the system's resource status. Additionally, we propose a Multi-armed Bandits relatedThompsonSampling basedAdaptivePreferenceOptimization (TSAPO) algorithm for task-offloading policy, enhancing the timely identification of potent edge servers to improve computing resource utilization which first considers dynamic edge server space and parallel computing scenarios. The solid proof process greatly contributes to the theoretical analysis of the TSAPO. The simulation experiments demonstrate the delay and budget can be reduced by around 25% and 10% respectively, showcasing the superior performance of our approach.
Yang Xu 0013, Hangfan Li, Cheng Zhang 0035, Zhiqing Tang, Xiaoxiong Zhong, Ju Ren 0001, Hongbo Jiang 0001, Yaoxue Zhang
IEEE Trans. Mob. Comput.6
2025 Towards Privacy-Enhanced and Robust Clustered Federated Learning
abstract
Clustered federated learning (CFL) leverages data distribution similarities to cluster clients, facilitating personalized model training under data heterogeneity. However, most existing CFL schemes pose potential privacy risks for clients (e.g., gradient inversion attacks) as they rely on individual gradients for clustering. This also renders them incompatible with secure aggregation mechanisms that are widely employed in federated learning for privacy protection. Moreover, CFL introduces the risk of malicious clients dominating several clusters and conducting poisoning attacks therein, thereby threatening secure model training. To address these issues, we propose ProCFL, a Privacy-Enhanced and Robust CFL framework incorporating gradient-free clustering and peer validation. Specifically, we first design a new protocol for measuring data distribution similarity among clients without using their gradient information. Then, we transform the client clustering process into a weighted set covering problem and introduce a diversity-optimized clustering algorithm to achieve near-optimal clustering results while eliminating any need for prior knowledge. Furthermore, we develop a post-hoc detection mechanism that employs peer validation to identify and discard malicious client models. Extensive experimental evaluation of ProCFL validates its superior model robustness and accuracy performance compared to existing schemes.
Yang Xu 0013, Yunlin Tan, Cheng Zhang 0035, Peng Sun 0003, Yibang Zhang, Ju Ren 0001, Hongbo Jiang 0001, Yaoxue Zhang
IEEE Trans. Mob. Comput.6
2025 Multi-Agent Reinforcement Learning for Task Offloading in Crowd-Edge Computing
abstract
The Crowd-edge (CE) computing paradigm facilitates the utilization of the computational resources through simultaneously relying the edge computing and the collaboration among various mobile devices (MDs). Most existing works, focusing on offloading tasks from device to edge servers by centralized solutions, are unable to distribute tasks to massive MDs in CE. Meanwhile, designing a decentralized task offloading solution enabling task subscribers to individually make offloading decisions can be challenging given the randomness of crowd resource provisioning and limited knowledge of global status variations. In this paper, we propose a decentralized crowd-edge task offloading solution that enables users to optimally offload tasks to the CE in a distributed manner. Specifically, we formulate the corresponding problem as a stochastic optimization with partially observable status. By observing network and process delays at the crowd side, we further reform the optimization forms and provide a novel approximation policy, enabling users to optimize their offloading strategy based on local observations without interaction with each other. We then solve this task offloading problem by developing a Mixed Multi-Agent Proxy Policy Optimization algorithm (mixed MAPPO). Extensive testing, including numerical and system-level simulations, was conducted to validate the performance of the proposed algorithm in terms of task delay (including the processing delay and transmission delay), load rate, and resource utilization.
Su Yao, Ju Ren 0001, Weiqiang Wang 0002, Ke Xu 0002, Mingwei Xu 0001, Hongke Zhang
IEEE Trans. Mob. Comput.3
2025 DualRec: A Collaborative Training Framework for Device and Cloud Recommendation Models
abstract
Recommendation systems (RS) play a vital role in various domains. However, under recent data regulations like General Data Protection Regulation (GDPR), traditional RS that rely on collecting user's interaction data centrally face significant challenges. Federated learning (FL) enables collaborative model training among users while keeping their private data locally. Yet, the constrained resources of devices often limit the size of the learned model, resulting in suboptimal recommendation performance. To overcome the dilemma of data accessibility and model size, we propose DualRec, a novel collaborative training framework for device and cloud recommendation models. In DualRec, users train lightweight models on devices to harness their local private data, while a larger model is simultaneously trained on the cloud server to exploit its substantial resources. Devices and the cloud server collaboratively train their models, compensating for individual limitations of model size and data availability, enabling mutual empowerment and benefits. Specifically, we introduce an efficient aggregation mechanism for recommendation models to boost the collaborative training performance of device models. With the learned device models, we propose to generate pseudo user interaction data to train the server model. To enhance the training performance of the server model, we design an automated denoising mechanism to mitigate the negative impact of noisy samples in the generated pseudo dataset. Finally, the learned knowledge of the server model is distilled to device models for enhanced on-device recommendation performance. Extensive experiments demonstrate the superior performance of DualRec compared to state-of-the-art baselines.
Ye Zhang 0033, Yongheng Deng, Sheng Yue 0001, Qiushi Li 0002, Ju Ren 0001
IEEE Trans. Mob. Comput.5
2025 Collaborative Edge and Cloud Computing: Optimal Configuration and Computation Management
abstract
Mobile Edge Computing (MEC) plays an increasingly important role in the rapidly increasing mobile applications by providing high-quality computing services. The majority of current research has focused on designing efficient computing task offloading schemes to ensure the effectiveness of the MEC system. However, the configuration and resource management of the MEC system, which are crucial for its scattered feature, have not received due attention. This paper investigates the configuration and computation resource management problem for the MEC system by formulating a profit maximization problem. To address this problem, we first analyze the relationship among mobile users' offloading decisions, the configuration and computation management of the MEC system, and the service quality. Then, we design an optimal configuration and computation management scheme of the MEC system, which can not only maintain the efficiency of computing processes but also make a good trade-off between the profitability and the service quality. In such a way, the total expected profit of the MEC system can be maximized. Numerical evaluations show that the proposed optimal configuration and computation management scheme can efficiently improve the total profit of the MEC system.
Yongmin Zhang, Wei Wang 0343, Junfan Zhou, Yang Xu 0013, Ju Ren 0001, Yaoxue Zhang
IEEE Trans. Mob. Comput.6
2025 Momentum-Based Contextual Federated Reinforcement Learning
abstract
Federated Reinforcement Learning (FRL) is an attractive edge learning paradigm for decision-making applications, which has garnered significant interest recently. However, owing to the inherent spatio-temporal non-stationarity of local state-action distributions, current FRL approaches typically suffer from high interaction and communication costs. In this paper, we introduce a new FRL method, which incorporates momentum, importance sampling, and server-side adjustments, capable of controlling the gradient shifts induced by the non-stationary data. We prove that by proper selection of momentum parameters and interaction frequency, it can achieve$\tilde {\mathcal {O}}(H N^{-1}\epsilon ^{-3/2})$and$\tilde {\mathcal {O}}(\epsilon ^{-1})$interaction and communication complexities (N represents the agent number), where the interaction complexity achieves linear speedup with the number of agents, and the communication complexity aligns with the best achievable among existing first-order FL algorithms. Further, we leverage attention-based contextual representation extraction to enable the learning policy to adapt to heterogeneous tasks and environments. Extensive experiments demonstrate that our proposed method significantly outperforms existing baselines on a range of complex, high-dimensional single-task and multi-task benchmarks.
Sheng Yue 0001, Xingyuan Hua, Yongheng Deng, Ju Ren 0001, Yaoxue Zhang
IEEE Trans. Netw.5
2025 AugFL: Augmenting Federated Learning With Pretrained Models
abstract
Federated Learning (FL) has garnered widespread interest in recent years. However, owing to strict privacy policies or limited storage capacities of training participants such as IoT devices, its effective deployment is often impeded by the scarcity of training data in practical decentralized learning environments. In this paper, we study enhancing FL with the aid of (large) pre-trained models (PMs), that encapsulate wealthy general/domain-agnostic knowledge, to alleviate the data requirement in conducting FL from scratch. Specifically, we consider a networked FL system formed by a central server and distributed clients. First, we formulate the PM-aided personalized FL as a regularization-based federated meta-learning problem, where clients join forces to learn a meta-model with knowledge transferred from a private PM stored at the server. Then, we develop an inexact-ADMM-based algorithm, AugFL, to optimize the problem with no need to expose the PM or incur additional computational costs to local clients. Further, we establish theoretical guarantees for AugFL in terms of communication complexity, adaptation performance, and the benefit of knowledge transfer in general non-convex cases. Extensive experiments corroborate the efficacy and superiority of AugFL over existing baselines.
Sheng Yue 0001, Zerui Qin, Yongheng Deng, Ju Ren 0001, Yaoxue Zhang, Junshan Zhang
IEEE Trans. Netw.4
2025 MUCVR: Edge Computing-Enabled High-Quality Multi-User Collaboration for Interactive MVR
abstract
Mobile Virtual Reality (MVR), which aims to provide high-quality VR services to mobile devices of end users, has become the latest trend in virtual reality developments. The current MVR solution is to remotely render frame data from a cloud server, while the potential of edge computing in MVR is underexploited. In this paper, we propose a new approach named MUCVR to achieve high-quality interactive MVR collaboration for multiple users by exploiting edge computing. Firstly, we design “vertical” edge–cloud collaboration for VR task rendering, in which foreground interaction is offloaded to an edge server for rendering, while the background environment is rendered by the cloud server. Correspondingly, the VR device of a user is only responsible for decoding and displaying. Secondly, we propose the “horizontal” multi-user collaboration based on edge–edge cooperation, which synchronizes the data among edge servers. Finally, we implement the proposed MUCVR on an MVR device and the Unity VR application engine. The results show that MUCVR can effectively reduce the MVR service latency, improve the rendering performance, reduce the computing load on the VR device, and, ultimately, improve users' quality of experience.
Weimin Li 0002, Weihong Tian, Jie Gao 0002, Fan Wu 0014, Jianxun Liu 0001, Ju Ren 0001
IEEE Trans. Parallel Distributed Syst.7
2025 $C^{2}D$C2D: Context-Aware Concept Decomposition for Personalized Text-to-Image Synthesis
abstract
Concept decomposition is a technique for personalized text-to-image synthesis which learns textual embeddings of subconcepts from images that depicting an original concept. The learned subconcepts can then be composed to create new images. However, existing methods fail to address the issue of contextual conflicts when subconcepts from different sources are combined because contextual information remains encapsulated within the subconcept embeddings. To tackle this problem, we propose a Context-aware Concept Decomposition ($C^{2}D$C2D) framework. Specifically, we introduce a Similarity-Guided Divergent Embedding (SGDE) method to obtain subconcept embeddings. Then, we eliminate the latent contextual dependence between the subconcept embeddings and reconstruct the contextual information using an independent contextual embedding. This independent context can be combined with various subconcepts, enabling more controllable text-to-image synthesis based on subconcept recombination. Extensive experimental results demonstrate that our method outperforms existing approaches in both image quality and contextual consistency.
Jiang Xin, Xiaonan Fang 0001, Xueling Zhu, Ju Ren 0001, Yaoxue Zhang
IEEE Trans. Vis. Comput. Graph.4
2025 Privacy-aware Real-Time Target Person Matting in Multi-Person Scenes Using Dual Encoder-Decoder Networks
Jiang Xin, Xiaonan Fang 0001, Xueling Zhu, Ruyi Dai, Ju Ren 0001, Wenzhen Yue, Yaoxue Zhang
Vis. Comput.5
2024 RISiren: Wireless Sensing System Attacks via Metasurface
abstract
After over a decade of intensive research, wireless sensing technology is nearing commercialization. However, the inherent openness of the wireless medium exposes this technology to security flaws and vulnerabilities. In this paper, we introduce RISiren to reveal the risk. RISiren is a pioneering end-to-end black-box attack system leveraging programmable metasurface with a high level of stealthiness. The key insight of RISiren lies in its ability to generate malicious multipath using metasurface, thereby disrupting wireless channel metrics influenced by genuine human activities and facilitating malicious attacks. To ensure the effectiveness of RISiren, we propose a novel metasurface configuration strategy aiming at creating human-like activities that stem from a comprehensive analysis of how human activities impact wireless signal propagation. We have implemented and validated RISiren using commercial Wi-Fi devices. Our evaluation involved testing our attack strategies against five state-of-the-art systems (including five different types of recognition frameworks) representative of the current landscape. The experimental results show that the adversarial wireless signals generated by RISiren achieve over 90% attack success rate on average, and remain robust and effective across different environments and deployment setups, including through wall attack scenarios.
Chenghan Jiang, Jinjiang Yang, Xinyi Li 0005, Qi Li 0002, Xinyu Zhang 0003, Ju Ren 0001
CCS6
2024 Intelligent Hybrid Memory Scheduling Based on Page Pattern Recognition
abstract
Hybrid memory systems exhibit disparities in their heterogeneous memory components' access speeds. Dynamic page scheduling to ensure memory access predominantly occurs in the faster memory components is essential for optimizing the performance of hybrid memory systems. Recent works attempt to optimize page scheduling by predicting their hotness using neural network models. However, they face two crucial challenges: the page explosion problem and the new pages problem. We propose an intelligent hybrid memory scheduler driven by page pattern recognition to address these two challenges. Experimental results demonstrate that our approach outperforms state-of-the-art intelligent schedulers regarding effectiveness and cost.
Yanjie Zhen, Weining Chen, Wei Gao 0006, Ju Ren 0001, Kang Chen 0001, Yu Chen 0004
DATE4
2024 Flexible and Effective Cellular Traffic Data Synthesis with Large Language Model
abstract
Cellular traffic data hold significant potential for applications such as network planning, traffic prediction, mobility modeling, and personalized recommendations. However, limited data accessibility hinders more open data-driven research. Previous studies have explored data synthesis, while exhibiting flexible limitations in supporting conditional traffic synthesis, and are vulnerable to multidimensional data modeling. In this paper, we present LLMCell, a flexible and effective framework that leverages the arbitrary conditioning and contextual understanding capabilities of the large language model (LLM) to generate high-quality synthetic cellular traffic data. The LLMCell comprises three key components: i) a textual encoder for converting raw cellular traffic data into textual representations, ii) a generative model learner to fine-tune pre-trained LLM based on encoded textual representation for cellular traffic generation, and iii) a synthetic data sampling module for final synthetic data sampling and textual-to-data transformation. Experiments conducted on a large-scale dataset demonstrate the superior fidelity and utility of LLMCell over state-of-the-art baselines, and the synthetic data can effectively preserve user privacy. We release our synthetic dataset to the public to benefit future research in the wireless network community1.
Sijing Duan, Feng Lyu 0001, Jinfeng Cen, Ju Ren 0001, Peng Yang 0004, Yaoxue Zhang
GLOBECOM4
2024 SecSCS: A User-Centric Secure Smart Camera System Based on Blockchain
abstract
Smart cameras have gained immense popularity in commercial markets for their safety and security capabilities. Yet, the prevalent design of these intelligent camera systems often compels users to cede control of their data to poten-tially untrusted service providers, such as cloud services. This relinquishment can lead to unauthorized data access by these intermediaries, posing significant security and privacy risks. The conventional solutions have been to employ privacy-enhancing technologies to bypass these intermediaries, but at the cost of increased overhead for video streaming and sharing. In our study, we introduce SecSCS, a user-centric, blockchain-based secure camera system that incorporates essential features like video streaming, sharing, deletion, and permission restoration. SecSCS integrates a blockchain-enabled user login protocol with a secure device pairing mechanism that combines visual authorization with blockchain to flexibly manage the device ownership. We utilize blockchain to provide integrity protection for the video clips stored remotely, ensuring the video data remains tamper-proof. Furthermore, we present a video frame compression and a fast video encryption method aimed at boosting the efficiency of smart camera systems. Our evaluations show that, in comparison to the leading decentralized scheme, CaCTUs, SecSCS improves the computational and communication overhead for live streaming by a factor of 12.58 and 11.29, respectively, at a frame rate of 24 fps and a resolution of 720p.
Xinyuan Qian 0002, Hongwei Li 0001, Haoyong Wang, Guowen Xu, Shengmin Xu, Ju Ren 0001
ICDCS6
2024 EdgeVPR: Transformer-Based Real-Time Video Person Re-Identification at the Edge
abstract
Person re-identification (Re-ID) aims to search for a target person through non-overlapping cameras. With the rapid development of computing and storage capacity of edge sensors, performing person Re-Idon edge devices has become more and more popular in recent years. Since the raw data recorded by edge devices does not have to be transmitted to the server directly, this scenario greatly improves data privacy and security. Edge-based person Re-ID also reduces the computation and transmission pressure of central servers. In this paper, we take the first step in performing video-based person Re-ID on edge devices with limited computing and storage resources. To deal with person tracklets extracted from video recordings, we design EdgeVPR, a novel lightweight real-time video person Re-ID model based on Transformer architecture. We use multi-level knowledge distillation to learn lightweight models from server-side large models. For the lightweight model, we propose a multi-scale spatio-temporal attention module (MSTA) to replace the original multi-head self-attention (MSA) layers in Transformer. Our MSTA module can not only capture both spatial and temporal information from tracklets but also greatly reduces the computation compared with MSA layers. To deal with the challenge caused by occlusion or mis-classification in generating person tracklets, we perform patch transformation during the teacher model training process and use contrastive learning methods to enhance the model's robustness. A pluggable environment adapter is designed for the lightweight student model environment-oriented fine-tuning since edge sensors often face different shooting environments and angles. We perform experiments on MARS dataset [1] and DukeMTMC-VideoReID dataset [2]. Results show that EdgeVPR gets significantly better results compared with prior edge-based person Re-ID work.
Ju Ren 0001, Yaoxue Zhang
ICDCS2
2024 How to Leverage Diverse Demonstrations in Offline Imitation Learning
abstract
Offline Imitation Learning (IL) with imperfect demonstrations has garnered increasing attention owing to the scarcity of expert data in many real-world domains. A fundamental problem in this scenario is *how to extract positive behaviors from noisy data*. In general, current approaches to the problem select data building on state-action similarity to given expert demonstrations, neglecting precious information in (potentially abundant) *diverse* state-actions that deviate from expert ones. In this paper, we introduce a simple yet effective data selection method that identifies positive behaviors based on their *resultant states* - a more informative criterion enabling explicit utilization of dynamics information and effective extraction of both expert and beneficial diverse behaviors. Further, we devise a lightweight behavior cloning algorithm capable of leveraging the expert and selected data correctly. In the experiments, we evaluate our method on a suite of complex and high-dimensional offline IL benchmarks, including continuous-control and vision-based tasks. The results demonstrate that our method achieves state-of-the-art performance, outperforming existing methods on **20/21** benchmarks, typically by **2-5x**, while maintaining a comparable runtime to Behavior Cloning (BC).
Sheng Yue 0001, Jiani Liu 0005, Xingyuan Hua, Ju Ren 0001, Sen Lin 0001, Junshan Zhang, Yaoxue Zhang
ICML4
2024 OLLIE: Imitation Learning from Offline Pretraining to Online Finetuning
abstract
In this paper, we study offline-to-online Imitation Learning (IL) that pretrains an imitation policy from static demonstration data, followed by fast finetuning with minimal environmental interaction. We find the naive combination of existing offline IL and online IL methods tends to behave poorly in this context, because the initial discriminator (often used in online IL) operates randomly and discordantly against the policy initialization, leading to misguided policy optimization and *unlearning* of pretraining knowledge. To overcome this challenge, we propose a principled offline-to-online IL method, named OLLIE, that simultaneously learns a near-expert policy initialization along with an *aligned discriminator initialization*, which can be seamlessly integrated into online IL, achieving smooth and fast finetuning. Empirically, OLLIE consistently and significantly outperforms the baseline methods in **20** challenging tasks, from continuous control to vision-based domains, in terms of performance, demonstration efficiency, and convergence speed. This work may serve as a foundation for further exploration of pretraining and finetuning in the context of IL.
Sheng Yue 0001, Xingyuan Hua, Ju Ren 0001, Sen Lin 0001, Junshan Zhang, Yaoxue Zhang
ICML3
2024 BR-DeFedRL: Byzantine-Robust Decentralized Federated Reinforcement Learning with Fast Convergence and Communication Efficiency
abstract
In this paper, we propose Byzantine-Robust Decentralized Federated Reinforcement Learning (BR-DeFedRL), an innovative framework that effectively combats the harmful influence of Byzantine agents by adaptively adjusting communication weights, thereby significantly enhancing the robustness of the learning system. By leveraging decentralized learning, our approach eliminates the dependence on a central server. Striking a harmonious balance between communication round count and sample complexity, BR-DeFedRL achieves efficient convergence with a rate of $\mathcal{O}\left( {\frac{1}{{TN}}} \right)$, where T denotes the communication rounds and N represents the local steps related to variance reduction. Notably, each agent attains an ϵ-approximation with a state-of-the-art sample complexity of $\mathcal{O}\left( {\frac{1}{{\varepsilon N}} + \frac{1}{\varepsilon }} \right)$. Extensive experimental validations further affirm the efficacy of BR-DeFedRL, making it a promising and practical solution for Byzantine-robust decentralized federated reinforcement learning.
Jing Qiao, Zuyuan Zhang, Sheng Yue 0001, Yuan Yuan 0014, Zhipeng Cai 0001, Xiao Zhang 0015, Ju Ren 0001, Dongxiao Yu
INFOCOM7
2024 AeroRec: An Efficient On-Device Recommendation Framework using Federated Self-Supervised Knowledge Distillation
abstract
Modern recommendation systems operate entirely on the basis of central servers, requiring users to upload their behavior data from mobile devices to these servers. This practice has raised concerns about data privacy among users. Federated learning (FL), a machine learning technique designed to protect privacy, is becoming the standard solution. By combining federated learning with recommendation systems, Federated Recommendation Systems (FRS) allow users to collaboratively train a shared recommendation model without transmitting their original behavior data. However, existing federated learning solutions disregard the inherent limitations of resource-constrained mobile devices, which include limited storage space, computational overhead, and communication bandwidth. Although deploying a lightweight recommendation model can address the constraints of mobile devices, achieving satisfactory accuracy is difficult even with substantial communication overhead incurred during multiple rounds of federated learning training due to the limitations of the lightweight model’s capabilities and the sparsity of recommendation data. To address this issue, we propose AeroRec, a n e fficient on-device Feder ated Recommendation Framework. In AeroRec, we use federated self-supervised distillation to enhance the global model after each parameter aggregation in each round. This approach not only accelerates the convergence rate but also surpasses the upper limit of the lightweight model’s capability, thereby enabling a higher accuracy. We demonstrate that AeroRec outperforms several state-of-the-art FRS frameworks regarding recommendation accuracy and convergence speed through extensive experiments on three real-world datasets.
Tengxi Xia, Ju Ren 0001, Wei Rao 0003, Qin Zu, Yaoxue Zhang
INFOCOM2
2024 Momentum-Based Federated Reinforcement Learning with Interaction and Communication Efficiency
abstract
Federated Reinforcement Learning (FRL) has garnered increasing attention recently. However, due to the intrinsic spatio-temporal non-stationarity of data distributions, the current approaches typically suffer from high interaction and communication costs. In this paper, we introduce a new FRL algorithm, named MFPO, that utilizes momentum, importance sampling, and additional server-side adjustment to control the shift of stochastic policy gradients and enhance the efficiency of data utilization. We prove that by proper selection of momentum parameters and interaction frequency, MFPO can achieve $\widetilde {\mathcal{O}}\left({H{N^{ - 1}}{\varepsilon ^{ - 3/2}}}\right)$ and $\widetilde {\mathcal{O}}\left({{\varepsilon ^{ - 1}}}\right)$ interaction and communication complexities (N represents the number of agents), where the interaction complexity achieves linear speedup with the number of agents, and the communication complexity aligns the best achievable of existing first-order FL algorithms. Extensive experiments corroborate the substantial performance gains of MFPO over existing methods on a suite of complex and high-dimensional benchmarks.
Sheng Yue 0001, Xingyuan Hua, Ju Ren 0001
INFOCOM4
2024 Federated Offline Policy Optimization with Dual Regularization
abstract
Federated Reinforcement Learning (FRL) has been deemed as a promising solution for intelligent decision-making in the era of Artificial Internet of Things. However, existing FRL approaches often entail repeated interactions with the environment during local updating, which can be prohibitively expensive or even infeasible in many real-world domains. To overcome this challenge, this paper proposes a novel offline federated policy optimization algorithm, named DRPO, which enables distributed agents to collaboratively learn a decision policy only from private and static data without further environmental interactions. DRPO leverages dual regularization, incorporating both the local behavioral policy and the global aggregated policy, to judiciously cope with the intrinsic two-tier distributional shifts in offline FRL. Theoretical analysis characterizes the impact of the dual regularization on performance, demonstrating that by achieving the right balance thereof, DRPO can effectively counteract distributional shifts and ensure strict policy improvement in each federative learning round. Extensive experiments validate the significant performance gains of DRPO over baseline methods.
Sheng Yue 0001, Zerui Qin, Xingyuan Hua, Yongheng Deng, Ju Ren 0001
INFOCOM5
2024 RelayRec: Empowering Privacy-Preserving CTR Prediction via Cloud-Device Relay Learning
abstract
Click-through rate (CTR) prediction holds paramount importance across numerous applications, profoundly impacting user experience and business profitability. The freshness of a CTR prediction model significantly influences its performance, since users’ needs and interests may be changing over time, thereby requiring the model to be updated frequently. However, stringent data protection regulations have constrained the collection of users’ personal data, posing challenges to traditional model refreshing strategies that rely on centralized data collection. On-device learning techniques, such as federated learning (FL), offer a viable solution by enabling model training on devices without compromising user privacy. Nevertheless, the scarcity of training data with diverse distributions among devices presents considerable obstacles to on-device learning effectiveness. To address these challenges, we introduce RelayRec, a cloud-device relay learning framework designed for privacy-preserving CTR prediction. To establish competent initial models for devices, RelayRec categorizes pre-regulation cloud data into user preference groups, training preference-specific models for devices. Furthermore, a cloud-based automated model selector is developed to identify suitable initial models for devices. To elevate the relay learning performance of these initial models, we incorporate a personalized collaborative learning mechanism that aggregates device models based on user preferences. Extensive experimental evaluations underscore RelayRec’s superior performance compared to state-of-the-art benchmarks, affirming its efficacy in privacy-preserving CTR prediction.
Yongheng Deng, Guanbo Wang, Sheng Yue 0001, Wei Rao 0003, Qin Zu, Ju Ren 0001, Yaoxue Zhang
IPSN8
2024 Pushing Wireless Charging from Station to Travel
abstract
Wireless charging has achieved promising progress in recent years. However, the severe bottlenecks are the small charging range and poor flexibility. This paper presents ChargeX to enable smart and long-range wireless charging for small mobile devices. ChargeX incorporates emerging smart metasurface into the magnetic resonance coupling-based wireless charging to extend the charging range and accommodates the mobility of charging device. Unlike previous endeavors in metasurface-assisted wireless charging that focused on simulation, ChargeX makes efforts across software and hardware to meet three crucial requirements for a practical wireless charging system: (i) realize high-freedom and accurate metasurface control under the premise of low loss; (ii) obtain real-time feedback from the receiver and make effective manipulation for transmitted magnetic flux; and (iii) generate a proper AC signal source at the desired frequency band. We developed a prototype of ChargeX, and evaluated its performance through controlled experiments and real-world phone charging. Extensive experiments demonstrate the great potential of ChargeX for long-range and flexible wireless charging with a compact receiver design.
Bozhong Yu, Yongjian Fu 0004, Ju Ren 0001, Hao Pan 0003, Jeremy Gummeson, Yaoxue Zhang
MobiCom4
2024 RFMagus: Programming the Radio Environment With Networked Metasurfaces
abstract
The complexity and volatility of real-world radio environments often hamper wireless networks from achieving optimal performance. Recently, intelligent metasurfaces have been explored to dynamically reshape the radio propagation environment. However, existing systems are limited to standalone metasurfaces, only enabling one-time signal redirection/reshaping effects within their direct line-of-sight. They cannot effectively scale to cover larger areas. In this paper, we propose RFMagus, which employs a network of metasurfaces to overcome the limitation. We carefully optimize the configurations of the networked metasurfaces so that they can cooperatively and coherently propagate the analog signals towards the target regions. We have implemented the networked metasurfaces and deployed them in a variety of real-world environments. Experimental results demonstrate that RFMagus can effectively expand the coverage, improve the throughput, and operate transparently to different wireless standards.
Xinyi Li 0005, Gaoteng Zhao, Xinyu Zhang 0003, Ju Ren 0001
MobiCom5
2024 AutoMS: Automated Service for mmWave Coverage Optimization using Low-cost Metasurfaces
abstract
mmWave networks offer wide bandwidth for high-speed wireless communication but suffer from limited range and susceptibility to blockage. Existing coverage provisioning solutions not only incur high costs but also require significant expert knowledge and manual efforts. In this paper, we present AutoMS, an automated service framework to optimize mmWave coverage by strategically designing and placing low-cost passive metasurfaces. Our approach consists of three key components: (1) joint optimization of metasurface phase configurations and placement as well as access point beamforming codebooks. (2) a fast 3D ray-tracing simulator for accelerated large-scale metasurface channel modeling. (3) a metasurface design amenable to ultra-low-cost hot stamping fabrication, featuring high reflectivity, near 2π phase control, and wideband support. Simulation and testbed experiments show that AutoMS can increase the median received signal strength by 11 dB in target rooms and over 20 dB at previous blind spots, and improve the median throughput by over 3× in real-world scenarios.
Ruichun Ma, Shicheng Zheng, Hao Pan 0003, Lili Qiu, Liangyu Liu, Yihong Liu 0003, Ju Ren 0001
MobiCom9
2024 GPMS: Enabling Indoor GNSS Positioning using Passive Metasurfaces
abstract
Global Navigation Satellite System (GNSS) is extensively utilized for outdoor positioning and navigation. However, achieving high-precision indoor positioning is challenging due to the significant attenuation of GNSS signals indoors. To address this issue, we propose an innovative indoor GNSS positioning system called GPMS, which uses passive metasurface technology to redirect GNSS signals from outdoors into indoor spaces. These passive metasurfaces are strategically optimized for indoor coverage by steering and scattering the GNSS signals across a wide range of incident angles. We further develop a novel localization algorithm that can determine which metasurface the signal goes through and localize the user using the set of metasurfaces as anchor points. A distinct advantage of our localization algorithm is that it can be implemented on existing mobile devices without any hardware modifications. We implement the prototype of GPMS, and deploy six metasurfaces in two indoor environments, a 10×50 m2 office floor and a 15×20 m2 lecture room, to evaluate system performance. In terms of coverage, our GPMS increases the C/N0 from 9.1 dB-Hz to 23.2 dB-Hz and increases the number of visible satellites from 3.6 to 21.5 in the office floor. In terms of indoor positioning accuracy, our proposed system decreases the absolute positioning error from 30.6 m to 3.2 m in the office floor, and from 11.2 m to 2.7 m in the lecture room, demonstrating the feasibility and benefits of metasurface-assisted GNSS for indoor positioning.
Yezhou Wang, Hao Pan 0003, Lili Qiu, Linghui Zhong, Jiting Liu, Ruichun Ma, Yi-Chao Chen 0001, Guangtao Xue, Ju Ren 0001
MobiCom9
2024 Adaptive Metasurface-Based Acoustic Imaging using Joint Optimization
abstract
Acoustic imaging is attractive due to its ability to work under occlusion, different lighting conditions, and privacy-sensitive environments. Existing acoustic imaging methods require large transceiver arrays or device movement, which makes it challenging to use in many scenarios. In this paper, we develop a novel acoustic imaging system for low-cost devices with few speakers and microphones without any device movement. To achieve this goal, we leverage a 3D-printed passive acoustic metasurface to significantly enhance the diversity of the measurement data, thereby improving the imaging quality. Specifically, we jointly design the transmission signal, transceivers' beamforming weights, metasurface, and imaging algorithm to minimize the imaging reconstruction error in an end-to-end manner. We further develop a scheme to dynamically adapt the imaging resolution based on the distance to the target. We implement a system prototype. Using extensive experiments, we show that our system yields high-quality images across a wide range of scenarios.
Yongjian Fu 0004, Yongzhao Zhang, Yu Lu 0022, Lili Qiu, Yi-Chao Chen 0001, Yezhou Wang, Yijie Li 0002, Ju Ren 0001, Yaoxue Zhang
MobiSys9
2024 Empowering In-Browser Deep Learning Inference on Edge Through Just-In-Time Kernel Optimization
abstract
Web is increasingly becoming the primary platform to deliver AI services onto edge devices, making in-browser deep learning (DL) inference more prominent. Nevertheless, the heterogeneity of edge devices, combined with the underdeveloped state of Web hardware acceleration practices, hinders current in-browser inference from achieving its full performance potential on target devices.
Fucheng Jia, Shiqi Jiang 0002, Ting Cao 0003, Tianrui Xia, Yuanchun Li 0003, Qipeng Wang 0001, Ju Ren 0001, Yunxin Liu 0001, Lili Qiu, Mao Yang 0004
MobiSys10
2024 You Can Use But Cannot Recognize: Preserving Visual Privacy in Deep Neural Networks
Qiushi Li 0002, Yan Zhang 0104, Ju Ren 0001, Qi Li 0002, Yaoxue Zhang
NDSS3
2024 Orthcatter: High-throughput In-band OFDM Backscatter with Over-the-Air Code Division
Caihui Du, Jihong Yu, Ju Ren 0001, Jianping An
NSDI4
2024 Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
abstract
The adaptation of pre-trained large language models (LLMs) to diverse downstream tasks via fine-tuning is critical for numerous applications. However, the inefficiency of parameterefficient fine-tuning (PEFT) techniques presents significant challenges in terms of time investments and operational costs. In this paper, we first introduce a nuanced form of sparsity, termed Shadowy Sparsity, which is distinctive in fine-tuning and has not been adequately addressed for acceleration. Under Shadowy Sparsity, we propose Long Exposure1, an efficient system to accelerate PEFT for LLMs. Long Exposure comprises three key components: Shadowy-sparsity Exposer employs a prolonged sensing range to capture more sparsity details under shadowy sparsity; Sequence-oriented Predictor provides efficient yet accurate predictions to handle large sequence inputs and constantly-evolving parameters; and Dynamic-aware Operator facilitates more structured computational patterns and coalesced memory accesses, addressing dynamic sparse operations. Extensive evaluations show that Long Exposure outperforms state-of-the-arts with up to a $2.49 \times$ speedup in end-to-end fine-tuning, offering promising advancements in accelerating PEFT for LLMs.1Long Exposure is available at https://github.com/HPHEX/LongExposure.
Tuowei Wang, Kun Li 0016, Zixu Hao, Donglin Bai, Ju Ren 0001, Yaoxue Zhang, Ting Cao 0003, Mao Yang 0004
SC5
2024 M3Cam: Extreme Super-resolution via Multi-Modal Optical Flow for Mobile Cameras
abstract
The demand for ultra-high-resolution imaging in mobile phone photography is continuously increasing. However, the image resolution of mobile devices is typically constrained by the size of the CMOS sensor. Although deep learning-based super-resolution (SR) techniques have the potential to overcome this limitation, existing SR neural network models require large computational resources, making them unsuitable for real-time SR imaging on current mobile devices. Additionally, cloud-based SR systems pose privacy leakage risks. In this paper, we propose M3Cam, an innovative and lightweight SR imaging system for mobile phones. M3Cam can ensure high-quality 16× SR image (4× in both height and width) visualization with almost negligible latency. In detail, we utilize an optical image stabilization (OIS) module for lens control and introduce a new modality of data, namely gyroscope readings, to achieve high-precision and compact optical flow estimation modules. Building upon this concept, we design a multi-frame-based SR model utilizing the Swin Transformer. Our proposed system can generate a 16× SR image from four captured low-resolution images in real-time, with low computational load, low inference latency, and minimal reliance on runtime RAM. Through extensive experiments, we demonstrate that our proposed multi-modal optical flow model significantly enhances pixel alignment accuracy between multiple frames and delivers outstanding 16× SR imaging results under various shooting scenarios. Code and dataset are available at: https://github.com/liangjindeamo-yuer/M3CAM
Yu Lu 0022, Dian Ding, Hao Pan 0003, Yongjian Fu 0004, Feitong Tan, Yi-Chao Chen 0001, Guangtao Xue, Ju Ren 0001
SenSys10
2024 Dependent Task Offloading in Edge Computing Using GNN and Deep Reinforcement Learning
abstract
Task offloading is a widely used technology in Edge Computing (EC), which declines the makespan of user task with the aid of resourceful edge servers. How to solve the competition for computation and communication resources among tasks is a fundamental issue in task offloading. Besides, real-life user tasks often comprise multiple interdependent subtasks. Dependencies among subtasks significantly raises the complexity of task offloading, and makes it difficult to propose generalized approaches for scenarios of different size. In this paper, we study the Dependent Task Offloading (DTO) problem within both single-user single-edge and multi-user multi-edge scenario. First, we use Directed Acyclic Graph (DAG) to model dependent task, where nodes and directed edges represent the subtasks and their interdependencies respectively. Then, we propose a task scheduling method based on Graph Attention Network (GAT) and Deep Reinforcement Learning (DRL) to minimize the makespan of user tasks. More specifically, our method introduces a multi-discrete action DRL scheduler that simultaneously determines which subtask to consider and whether it should be offloaded at each step, and employs GAT to encode the graph-based state representation. To stabilize and speed up DRL scheduler training, we pretrain GAT encoder with unsupervised learning. Extensive experiments demonstrate that our proposed approach can be applied to various environments and outperforms prior methods.
Zequn Cao, Xiaoheng Deng, Sheng Yue 0001, Ping Jiang 0001, Ju Ren 0001, Jinsong Gui
IEEE Internet Things J.5
2024 MBSeg: Real-Time Contour-Based Method Instance Segmentation for Egocentric Vision
abstract
The instance segmentation task provides perceptual intelligence for Internet of Things (IoT) devices by identifying various objects in complex environments. However, achieving real-time performance on resource-constrained edge IoT devices poses a challenge due to the complexity of instance segmentation tasks. In this paper, we present a contour-based segmentation approach utilizing macroblocks(MBs), termed MBSeg. We utilize novel data, the MBs of H.265 videos, to guide object segmentation. Thanks to the prevalence of video encoding and decoding chips, this data is lightweight, fast, and easily accessible. MBSeg is a two-stage method. First, a lightweight object detection network acquires object positions and extracts features. Subsequently, we generate rough object contours from MBs, which are input into MBSnake, an active contour model optimized for edge devices, for further deformation. We introduce angle value evaluation of vertices to balance MBSnake’s accuracy and speed. We have implemented and validated MBSeg on commercial Android devices. Results demonstrate MBSeg achieves 30 FPS throughput on self-centric videos with multiple objects, a 6.4× speedup over YolactEdge, an edge instance segmentation method, with only a 13.8% drop in accuracy. This advancement provides a valuable foundation for self-centric visual wearable IoT.
Bangwen He, Wang Sun, Ju Ren 0001
IEEE Internet Things J.7
2024 Unishyper: A Rust-based unikernel enhancing reliability and efficiency of embedded systems
Keyang Hu, Wang Huang, Lei Wang 0126, Ce Mo, Runxiang Wang, Yu Chen 0004, Ju Ren 0001, Bo Jiang 0001
J. Syst. Archit.7
2024 PatternS: An intelligent hybrid memory scheduler driven by page pattern recognition
Yanjie Zhen, Weining Chen, Wei Gao 0006, Ju Ren 0001, Kang Chen 0001, Yu Chen 0004
J. Syst. Archit.4
2024 BatOpt: Optimizing GPU-Based Deep Learning Inference Using Dynamic Batch Processing
abstract
Deep learning (DL) has been applied in billions of mobile devices due to its astonishing performance in image, text, and audio processing. However, limited by the computing capability of mobile devices, a large amount of DL inference tasks need to be offloaded to edge or cloud servers, which makes powerful GPU servers are struggling to ensure the quality of service(QoS). To better utilize the highly parallel computing architecture of GPU to improve the QoS, we propose BatOpt, a framework that uses dynamic batch processing to strike a good balance between service latency and GPU memory usage in DL inference services. Specifically, BatOpt innovatively models the DL inference service as a$M/G(a,b)/1/N$queue, with the consideration of stochastic task arrivals, which enables it to predict the service latency accurately in different system states. Furthermore, we propose an optimization algorithm to trade off the service latency and GPU memory usage in different system states by analyzing the queueing model. We have implemented BatOpt on Pytorch and evaluated it on an RTX 2080 GPU using real DL models. BatOpt brings up to 31x and 4.3x times performance boost in terms of service latency, compared to single-input and fixed-batch-size strategies, respectively. And BatOpt's maximum GPU memory usage is only 0.3x that of greedy-dynamic-batch-size strategy on the premise of the same service latency.
Yunzhen Luo, Yanbo Wang 0003, Xiaoyan Kui, Ju Ren 0001
IEEE Trans. Cloud Comput.5
2024 On AoI of Grant-Free Access With HARQ
abstract
For mission-critical URLLC applications, timely status updates are essential. This paper investigates the age of information (AoI) of the three HARQ schemes specified in 5G R16, targeting to provide guidelines for future grant-free access design in 5G-Advanced and beyond. Specifically, we analyze two packet management policies: First-come-first-serve (FCFS) and preemption policy where new packets always preempt the buffer. We also study the AoI in a latency-sensitive scenario where expired packets are discarded. We derive exact expressions of AoI and peak AoI for all schemes and their lower bounds, revealing that the number of the maximum consecutive transmissions is critical for information freshness. Simulation results validate the theoretical analysis and show that Proactive HARQ scheme outperforms K-repetition HARQ scheme unconditionally and Reactive HARQ scheme with moderate system load or above. And discarding expired packets enhances system robustness for overload systems but yields larger AoI.
Jiwen Wang, Ju Ren 0001, Fangxin Wang 0001, Shuai Wang 0013, Jihong Yu
IEEE Trans. Commun.3
2024 UltraSR: Silent Speech Reconstruction via Acoustic Sensing
abstract
Silent Speech Interfaces (SSI) have been developed to convert silent articulatory gestures into speech, aiding communication in public spaces and assisting individuals with aphasia. Previous SSIs, which rely on wearable devices or cameras, often pose issues like prolonged contact or privacy risks. Recent advancements in acoustic sensing present new opportunities for gesture sensing, but they typically focus on content classification rather than reconstructing audible speech. This results in the loss of crucial speech characteristics such as rate, intonation, and emotion.In this paper, we propose UltraSR, a novel sensing system designed for accurate audible speech reconstruction by analyzing the disturbance of tiny articulatory gestures on reflected ultrasound signals. UltraSR employs a multi-scale feature extraction scheme to aggregate information from multiple views and introduces a new model that maps ultrasound to speech signals, enabling the reconstruction of audible speech from silent gestures.Instead of the laborious collection of massive training data, UltraSR constructs an inverse task to generate virtual gestures from widely available audio (e.g., phone calls) for efficient model training. Additionally, it incorporates a finetuning mechanism using unlabeled data for user adaptation.We implemented UltraSR on a portable smartphone and evaluated it in various environments. Results show that UltraSR can achieve a Character Error Rate (CER) as low as 5.22% and reduce the CER from 80.13% to 6.31% for new users with only 1 hour of ultrasound data, outperforming state-of-the-art acoustic-based approaches while preserving rich speech information.
Yongjian Fu 0004, Shuning Wang, Linghui Zhong, Ju Ren 0001, Yaoxue Zhang
IEEE Trans. Mob. Comput.5
2024 Manipulating Voice Assistants Eavesdropping via Inherent Vulnerability Unveiling in Mobile Systems
abstract
Numerous mobile devices are equipped with voice assistants to facilitate contactless user-device interaction. However, the widespread availability of voice assistants also raises security and privacy concerns, as they can be maliciously triggered to perform voice eavesdropping. Although diverse attacks have been taken to manipulate voice assistants for eavesdropping, they exhibit deficiencies of limited attack scopes and conspicuous attack behaviors because they target specific voice assistants or require extra voice commands to activate them. To manipulate arbitrary voice assistants for covert eavesdropping attack, we conduct a comprehensive analysis of voice assistant implementation in the Android system and refine a universal workflow. Through meticulous analysis and experimental verification, we uncover an inherent vulnerability that in voice assistants across device types that can be awakened by an artificial faking Intent. Building on this significant discovery, we propose an attack termed VoiceEar. It leverages a malicious event generation file and a first-in-first-out Intent generation algorithm to trigger voice assistants within the normal workflow for eavesdropping, without voice commands. Finally, we deploy the VoiceEar attacks on 25 mainstream mobile devices, and invite 95 volunteers for eavesdropping activity perception testing. The results unequivocally demonstrate the seamless execution of VoiceEar attacks, with neither users nor devices awareness.
Wenbin Huang 0003, Hangcheng Cao, Ju Ren 0001, Hongbo Jiang 0001, Zhangjie Fu 0001, Yaoxue Zhang
IEEE Trans. Mob. Comput.4
2024 CoralDB: A Collaborative Database for Data Sharing Based on Permissioned Blockchain
abstract
Systems that integrate distributed databases and existing blockchain platforms have recently emerged, which conveniently leverage their respective strengths to build efficient, secure, and usable data sharing and collaboration environments for different organizations. However, the performance of such systems can be limited by the native blockchain platforms due to the high latency of transactions. In this paper, we present CoralDB, a bottom-up fully redesigned hybrid system of blockchain and database, aimed at enabling untrusted organizations to collaborate and share data efficiently and securely at the database level. The storage layer of CoralDB ensures data security and system throughput through key modules such as customized block structure, consensus mechanism, and transaction pool. On top of the storage layer, a database layer is introduced, which extends the blockchain of the storage layer by incorporating connection pools, collaborative tables, and query interfaces, to enhance the usability and efficiency of data collaboration and sharing. Extensive experimental results demonstrate that CoralDB provides security assurances at the level of blockchain and enables efficient decentralized data collaboration and sharing.
Weimin Li 0002, Weihong Tian, Zhengmao Yan, Jie Gao 0002, Fan Wu 0014, Jianxun Liu 0001, Wenxiong Chen, Ju Ren 0001
IEEE Trans. Mob. Comput.9
2024 Efficient Resource Management and Expansion Scheme for Collaborative Edge-Cloud Computing
abstract
Integrating the advantages of both the edge and the cloud, the edge-cloud computing system emerges to provide high-quality computing services for mobile users. To improve system efficiency, we investigate a hybrid mode of resource collaboration and expansion for the edge-cloud computing system, in which edge servers not only can collaborate with the cloud by purchasing high-priority computation resources temporarily but also can expand their local computation resources permanently. In such a way, the edge server can maximize its long-term profit by making a trade-off between the purchasing cost and the expanding cost. By formulating the resource management problem as a long-term profit maximization one, we first analyze the relationships among the expected minimal purchasing cost, the computation delay, and the available computation resources. Then, we design an efficient resource reserving and expanding scheme to determine the optimal expected amounts of reserving resources and expansion resources. Next, we propose an efficient real-time resource purchasing scheme to obtain the optimal amount of real-time purchasing resources dynamically. Finally, simulation results show that the proposed efficient resource collaboration and expanding scheme can maximize the long-term profit while guaranteeing the computation delay.
Wei Wang 0343, Yongmin Zhang, Ju Ren 0001, Feng Lyu 0001, Yaoxue Zhang
IEEE Trans. Mob. Comput.4
2024 Age-Efficient Random Access With Load Adaptation
abstract
The lightweight and energy-efficient Frame Slotted Aloha (FSA) protocol has become a promising MAC protocol in large-scale IoT systems. Existing work on minimizing the age of information (AoI) of FSA protocol cannot significantly benefit from frequent packet generations when the packet generation rate$\lambda$exceeds its throughput$e^{-1}$. To fill this gap, this paper proposes two age threshold-based algorithms to reduce the AoI of FSA systems for$\lambda > e^{-1}$, namely TF and TF+. Their core ideas are to only allow the nodes with age gain over the configured thresholds to send their packets so that the FSA systems are slimmed to a stable one with$\lambda < e^{-1}$and a polling system, respectively. Technically, we design the threshold configuration rules for the two algorithms and characterize the normalized average AoI. We also conduct simulation and the results show that TF and TF+ achieve lower AoI than the prior works.
Jiwen Wang, Jihong Yu, Ju Ren 0001, Yun Li 0001
IEEE Trans. Mob. Comput.4
2024 Characterizing Internet Card User Portraits for Efficient Churn Prediction Model Design
abstract
Cellular Internet card (IC) as a new business model emerges, which penetrates rapidly and holds the potential to foster a great business market. However, with the explosive growth of IC users, the user churn problem becomes severe, affecting the IC business significantly, while there is lacking appropriate techniques in the literature to deal with the issue. In this article, we take the lead to study one large-scale data set from a provincial network operator of China, which contains about 4 million IC users and 22 million traditional card (TC) users. We first justify the IC user churn issue with data, and categorize the user churning reasons. Then, we shed light on understanding user portraits, which is the building block to enable efficient model design. Particularly, we conduct a systematical analytics on usage data by studying the difference of two types of users, examining the impact of user properties, and characterizing the user Internet using behaviors. Finally, by using the IC user portraits and usage patterns, we propose anICuserChurnPrediction model, namedICCP, which consists of a feature extraction component and a learning-based churn prediction architecture design. For feature extraction, both the static portrait features and temporal sequential features are captured. In the learning architecture, we devise the principal component analysis (PCA) block and the embedding/transformer layers to learn the respective information of two types of features, which are collectively fed into the classification multilayer perceptron layer (MPL) for churn prediction. A reference implementation ofICCPis conducted within the telecom system and extensive experiments corroborate the efficiency ofICCP.
Fan Wu 0014, Feng Lyu 0001, Ju Ren 0001, Peng Yang 0004, Shijie Gao, Yaoxue Zhang
IEEE Trans. Mob. Comput.3
2024 HiMoDepth: Efficient Training-Free High-Resolution On-Device Depth Perception
abstract
High-resolution depth estimation, with a minimum resolution of$1280\times 960$, is essential for achieving more immersive experiences in on-device 3D vision applications. However, implementing high-resolution solutions on resource-limited mobile devices presents significant challenges, such as the need for additional expensive depth sensors, computation-intensive machine learning models requiring large-scale datasets, or the need for device motion while the target object remains stationary. In this study, we propose HiMoDepth, an efficient training-free high-resolution depth estimation system that utilizes widely-available on-device dual cameras. HiMoDepth consists of two modules: 1) homogenizing the on-device heterogeneous cameras by iteratively cropping the Field-of-Views to make the focal length of the cameras equal and filtering out the out-of-sync frames based on time stamps, and 2) designing a hierarchical mobile GPU-friendly stereo matching method that effectively reduces the latency of stereo matching with high-resolution depth maps by using efficient data layout, reducing the number of memory accesses, and searching the corresponding pixel over a coarse-to-fine hierarchy. We implement HiMoDepth on multiple commodity mobile devices and conduct comprehensive evaluations. Experimental results show that HiMoDepth significantly outperforms the baselines in both accuracy and running speed on mobile devices that support high-resolution depth maps.
Ju Ren 0001, Bangwen He, Youngki Lee 0001, Ting Cao 0003, Yuanchun Li 0003, Yaoxue Zhang, Yunxin Liu 0001
IEEE Trans. Mob. Comput.3
2024 A Communication-Efficient Hierarchical Federated Learning Framework via Shaping Data Distribution at Edge
abstract
Federated learning (FL) enables collaborative model training over distributed computing nodes without sharing their privacy-sensitive raw data. However, in FL, iterative exchanges of model updates between distributed nodes and the cloud server can result in significant communication cost, especially when the data distributions at distributed nodes are imbalanced with requiring more rounds of iterations. In this paper, with our in-depth empirical studies, we disclose that extensive cloud aggregations can be avoided without compromising the learning accuracy if frequent aggregations can be enabled at edge network. To this end, we shed light on the hierarchical federated learning (HFL) framework, where a subset of distributed nodes can play as edge aggregators to support edge aggregations. Under the HFL framework, we formulate a communication cost minimization (CCM) problem to minimize the total communication cost required for model learning with a target accuracy by making decisions on edge aggragator selection and node-edge associations. Inspired by our data-driven insights that the potential of HFL lies in the data distribution at edge aggregators, we propose ShapeFL, i.e., SHaping dAta distRibution at Edge, to transform and solve the CCM problem. In ShapeFL, we divide the original problem into two sub-problems to minimize the per-round communication cost and maximize the data distribution diversity of edge aggregator data, respectively, and devise two light-weight algorithms to solve them accordingly. Extensive experiments are carried out based on several opened datasets and real-world network topologies, and the results demonstrate the efficacy of ShapeFL in terms of both learning accuracy and communication efficiency.
Yongheng Deng, Feng Lyu 0001, Tengxi Xia, Yue-Zhi Zhou, Yaoxue Zhang, Ju Ren 0001, Yuanyuan Yang 0001
IEEE/ACM Trans. Netw.6
2024 SESAME: A Resource Expansion and Sharing Scheme for Multiple Edge Services Providers
abstract
As a potential computing solution for fast-growing mobile and IoT applications, edge computing has been developed rapidly. However, due to the relatively limited resources of each edge node, it is difficult for edge nodes to provide quality-guaranteed services to dynamic and massive computation tasks individually. To address this challenge, this paper proposes a two-stage resource expansion and sharing scheme, named SESAME, to enable resource sharing among the edge nodes within/across multiple edge service providers (ESPs), to improve the overall efficiency of the edge computing system. To facilitate the operation and reduce complexity, the resource management scheme has both long-term and short-term decision periods. During the long-term period, an optimal conservative estimation-based resource expansion and pricing strategy has been designed to ensure the system stability and the interests of ESPs. During the short-term period, a resource-sharing strategy considering the internal and external behaviors of ESPs has been proposed to reduce resource-sharing costs while fully utilizing internal resources. In such a way, the resources from different ESPs can collaborate efficiently. Extensive experiments on real datasets show that our algorithm can effectively reduce ESP costs and improve system stability.
Jiani Liu 0005, Ju Ren 0001, Yongmin Zhang, Sheng Yue 0001, Yaoxue Zhang
IEEE/ACM Trans. Netw.2
2024 PoPeC: PAoI-Centric Task Offloading With Priority Over Unreliable Channels
abstract
Freshness-aware computation offloading has garnered increasing attention recently in the realm of edge computing, driven by the need to promptly obtain up-to-date information and mitigate the transmission of outdated data. However, most of the existing works assume that channels are reliable, neglecting the intrinsic fluctuations and uncertainty in wireless communication. More importantly, offloading tasks typically have diverse freshness requirements. Accommodation of various task priorities in the context of freshness-aware task scheduling and resource allocation remains an open and unresolved problem. To overcome these limitations, we cast the freshness-aware task offloading problem as a multi-priority optimization problem, considering the unreliability of wireless channels, prioritized users, and the heterogeneity of edge servers. Building upon the nonlinear fractional programming and the ADMM-Consensus method, we introduce a joint resource allocation and task offloading algorithm to solve the original problem iteratively. In addition, we devise a distributed asynchronous variant for the proposed algorithm to further enhance its communication efficiency. We rigorously analyze the performance and convergence of our approaches and conduct extensive simulations to corroborate their efficacy and superiority over the existing baselines.
Nan Qiao 0008, Sheng Yue 0001, Yongmin Zhang, Ju Ren 0001
IEEE/ACM Trans. Netw.4
2024 CODE$^{+}$+: Fast and Accurate Inference for Compact Distributed IoT Data Collection
abstract
In distributed IoT data systems, full-size data collection is impractical due to the energy constraints and large system scales. Our previous work has investigated the advantages of integrating matrix sampling and inference for compact distributed IoT data collection, to minimize the data collection cost while guaranteeing the data benefits. This paper further advances the technology by boosting fast and accurate inference for those distributed IoT data systems that are sensitive to computation time, training stability, and inference accuracy. Particularly, we proposeCODE$^{+}$+, i.e.,Compact Distributed IOTData CollEction Plus, which features a cluster-based sampling module and a Convolutional Neural Network (CNN)-Transformer Autoencoders-based inference module, to reduce cost and guarantee the data benefits. The sampling component employs a cluster-based matrix sampling approach, in which data clustering is first conducted and then a two-step sampling is performed in accordance with the number of clusters and clustering errors. The inference component integrates a CNN-Transformer Autoencoders-based matrix inference model to estimate the full-size spatio-temporal data matrix, which consists of a CNN-Transformer encoder that extracts the underlying features from the sampled data matrix and a lightweight decoder that maps the learned latent features back to the original full-size data matrix. We implementCODE$^{+}$+under three operational large-scale IoT systems and one synthetic Gaussian distribution dataset, and extensive experiments are provided to demonstrate its efficiency and robustness. With a 20% sampling ratio,CODE$^{+}$+achieves an average data reconstruction accuracy of 94% across four datasets, outperforming our previous version of 87% and state-of-the-art baseline of 71%.
Huali Lu, Feng Lyu 0001, Ju Ren 0001, Huaqing Wu, Conghao Zhou, Zhongyuan Liu, Yaoxue Zhang, Xuemin Shen
IEEE Trans. Parallel Distributed Syst.3
2023 Learning from Limited Heterogeneous Training Data: Meta-Learning for Unsupervised Zero-Day Web Attack Detection across Web Domains
abstract
Recently unsupervised machine learning based systems have been developed to detect zero-day Web attacks, which can effectively enhance existing Web Application Firewalls (WAFs). However, prior arts only consider detecting attacks on specific domains by training particular detection models for the domains. These systems require a large amount of training data, which causes a long period of time for model training and deployment. In this paper, we propose RETSINA, a novel meta-learning based framework that enables zero-day Web attack detection across different domains in an organization with limited training data. Specifically, it utilizes meta-learning to share knowledge across these domains, e.g., the relationship between HTTP requests in heterogeneous domains, to efficiently train detection models. Moreover, we develop an adaptive preprocessing module to facilitate semantic analysis of Web requests across different domains and design a multi-domain representation method to capture semantic correlations between different domains for cross-domain model training. We conduct experiments using four real-world datasets on different domains with a total of 293M Web requests. The experimental results demonstrate that RETSINA outperforms the existing unsupervised Web attack detection methods with limited training data, e.g., RETSINA needs only 5-minute training data to achieve comparable detection performance to the existing methods that train separate models for different domains using 1-day training data. We also conduct real-world deployment in an Internet company. RETSINA captures on average 126 and 218 zero-day attack requests per day in two domains, respectively, in one month.
Ye Wang 0002, Qi Li 0002, Zhuotao Liu, Ke Xu 0002, Ju Ren 0001, Ruilin Lin
CCS6
2023 Privacy-Preserving DNN Training with Prefetched Meta-Keys on Heterogeneous Neural Network Accelerators
abstract
The embedded software may migrate the collected data to the server for DNN computation acceleration, which may compromise privacy. We propose a DNN computation framework that combines TEE and NNA to address the privacy leakage problem. We design an NNA-friendly encryption method that enables NNA to correctly compute the encrypted linear input. Facing the overhead of TEE-NNA interaction, we design a pipeline-based prefetch mechanism that can reduce the TEE interaction overhead. Experimentally, our approach proves to be compatible with a wide range of NPUs and TPUs, and improves the performance by 8-19 times over the TEE scheme.
Qiushi Li 0002, Ju Ren 0001, Yan Zhang 0104, Chengru Song, Yiqiao Liao, Yaoxue Zhang
DAC2
2023 CLARE: Conservative Model-Based Reward Learning for Offline Inverse Reinforcement Learning
Sheng Yue 0001, Guanbo Wang, Wei Shao 0006, Sen Lin 0001, Ju Ren 0001, Junshan Zhang
ICLR6
2023 A Hierarchical Knowledge Transfer Framework for Heterogeneous Federated Learning
abstract
Federated learning (FL) enables distributed clients to collaboratively learn a shared model while keeping their raw data private. To mitigate the system heterogeneity issues of FL and overcome the resource constraints of clients, we investigate a novel paradigm in which heterogeneous clients learn uniquely designed models with different architectures, and transfer knowledge to the server to train a larger server model that in turn helps to enhance client models. For efficient knowledge transfer between client models and server model, we propose FedHKT, a Hierarchical Knowledge Transfer framework for FL. The main idea of FedHKT is to allow clients with similar data distributions to collaboratively learn to specialize in certain classes, then the specialized knowledge of clients is aggregated to a super knowledge covering all specialties to train the server model, and finally the server model knowledge is distilled to client models. Specifically, we tailor a hybrid knowledge transfer mechanism for FedHKT, where the model parameters based and knowledge distillation (KD) based methods are respectively used for client-edge and edge-cloud knowledge transfer, which can harness the pros and evade the cons of these two approaches in learning performance and resource efficiency. Besides, to efficiently aggregate knowledge for conducive server model training, we propose a weighted ensemble distillation scheme with server-assisted knowledge selection, which aggregates knowledge by its prediction confidence, selects qualified knowledge during server model training, and uses selected knowledge to help improve client models. Extensive experiments demonstrate the superior performance of FedHKT compared to state-of-the-art baselines.
Yongheng Deng, Ju Ren 0001, Feng Lyu 0001, Yang Liu 0165, Yaoxue Zhang
INFOCOM2
2023 Towards Seamless Wireless Link Connection
abstract
Sub-6GHz and mmWave complement each other in the next generation of wireless communications for wide coverage and high capacity. However, there is still a gap between current network technology and seamless connection in indoor environments due to the inevitable occlusions that particularly affect higher frequency bands like Wi-Fi 5GHz and mmWave. To overcome this gap, economical tunable metasurfaces offer a promising solution by redirecting the beam direction of incoming waves to bypass blockages. However, existing metasurface technologies focus on a single frequency band and lack a theoretical framework to guide surface design for multiple desired bands, limiting their potential for diverse applications.
Bozhong Yu, Ju Ren 0001, Jeremy Gummeson, Yaoxue Zhang
MobiSys3
2023 FedINC: An Exemplar-Free Continual Federated Learning Framework with Small Labeled Data
abstract
Federated learning (FL) has shown great promise for privacy-preserving learning by enabling collaborative training on decentralized clients. However, in realistic FL scenarios, clients often collect new data continuously, join or exit learning dynamically. As a result, the global model tends to forget old knowledge while learning new knowledge. Meanwhile, labeling the continuously arriving data in real-time is usually challenging. Therefore, the catastrophic forgetting problem intertwined with the label deficiency issue poses significant challenges for both learning new knowledge and consolidating old knowledge. To address these challenges, we develop a novel exemplar-free continual federated learning framework named FedINC, to learn a global incremental model with limited labeled data. We begin by excavating the cause of catastrophic forgetting via in-depth empirical studies. Based on that, we introduce targeted mechanisms for FedINC, including a hybrid contrastive learning mechanism to efficiently learn new knowledge with limited labeled data, a plastic feature regularization mechanism to preserve old task's representation space, a prototype-guided regularization mechanism to mitigate feature overlap between old and new classes while aligning the features of non-iid clients, and a prototype evolution mechanism for flexible and efficient incremental classification. Extensive experiments demonstrate the superior performance of FedINC in terms of both convergence speed and accuracy of the global model.
Yongheng Deng, Sheng Yue 0001, Tuowei Wang, Guanbo Wang, Ju Ren 0001, Yaoxue Zhang
SenSys5
2023 Auction-Based Dependent Task Offloading for IoT Users in Edge Clouds
abstract
The rapid proliferation of latency-sensitive Internet of Things (IoT) applications boosts the frequency of offloading compute-intensive tasks from IoT users to mobile edge computing (MEC) due to the limitation resources of IoT devices. It is inevitably for IoT users to compete for the computing resources of the MEC, especially when the computation tasks are dependent and have hard deadline constraints. However, most existing dependent task offloading schemes may not well consider the resource competition issues among IoT users, and possibly lead to limited system performance in multiuser scenario. To address this issue, we intend to design an auction-based dependent task-offloading mechanism to improve the efficiency of task offloading for multiple IoT users. First, we formulate the dependent task offloading as a valuation maximization problem in the trade of computing resources satisfying users’ latency requirements, which has been proved to be NP-hard. Then, by jointly considering the task graph structure and the current status of the MEC, we propose a truthful auction mechanism, named greedy winner selection strategy, in which a heuristic dependent task assignment for winners is designed to improve the efficiency of the task offloading. By conducting extensive simulations, we validate that the performance of the proposed dependent task offloading strategy is superior to existing competition algorithms, in terms of total valuations, average makespans, and success rates.
Jiagang Liu, Yongmin Zhang, Ju Ren 0001, Yaoxue Zhang
IEEE Internet Things J.3
2023 FL-AMM: Federated Learning Augmented Map Matching With Heterogeneous Cellular Moving Trajectories
abstract
Map matching is a fundamental component for location-based services (LBSs), such as vehicle mobility analysis, navigation services, traffic scheduling, etc. In this paper, we investigate federated learning augmented map matching based on heterogeneous cellular moving trajectories from different operator systems, the goal of which is to improve matching accuracy without violating the user privacy. First, we develop a data collection platform with one Android-based application, and conduct rigorous data collection campaigns. Second, we perform systematic data analytics to reveal the data-driven technical challenges, including the impact of sampling rate, high location error of cellular moving data, and poor heterogeneous matching performance. Third, we propose an augmented map matching model, named FL-AMM, i.e.,FederatedLearningAugmentedMapMatching, in which we i) adopt the vertical federated learning framework to achieve data collaboration and privacy protection for heterogeneous operators; ii) devise a data augmentation component to enhance the capability of representing the raw cellular data; and iii) design a map matching model to further learn the mapping function from cellular trajectory points to road segments. Finally, we conduct extensive data-driven experiments to corroborate the efficiency and robustness of the proposed FL-AMM.
Huali Lu, Feng Lyu 0001, Huaqing Wu, Ju Ren 0001, Yaoxue Zhang, Xuemin Shen
IEEE J. Sel. Areas Commun.5
2023 VeLP: Vehicle Loading Plan Learning from Human Behavior in Nationwide Logistics System
abstract
For a nationwide logistics transportation system, it is critical to make the vehicle loading plans (i.e., given many packages, deciding vehicle types and numbers) at each sorting and distribution center. This task is currently completed by dispatchers at each center in many logistics companies and consumes a lot of workloads for dispatchers. Existing works formulate such an issue as a cargo loading problem and solve it by combinatorial optimization methods. However, it cannot work in some real-world nationwide applications due to the lack of accurate cargo volume information and effective model design under complicated impact factors as well as temporal correlation. In this paper, we explore a new opportunity to utilize large-scale route and human behavior data (i.e., dispatchers' decision process on planning vehicles) to generate vehicle loading plans (i.e., plans). Specifically, we collect a five-month nationwide operational dataset from JD Logistics in China and comprehensively analyze human behaviors. Based on the data-driven analytics insights, we design a Vehicle Loading Plan learning model, named VeLP, which consists of a pattern mining module and a deep temporal cross neural network, to learn the human behaviors on regular and irregular routes, respectively. Extensive experiments demonstrate the superiority of VeLP, which achieves performance improvement by 35.8% and 50% for trunk and branch routes compared with baselines, respectively. Besides, we deployed VeLP in JDL and applied it in about 400 routes, reducing the time by approximately 20% in creating plans. It saves significant human workload and improves operational efficiency for the logistics company.
Sijing Duan, Feng Lyu 0001, Xin Zhu 0007, Yi Ding 0011, Haotian Wang 0008, Desheng Zhang 0002, Yaoxue Zhang, Ju Ren 0001
Proc. VLDB Endow.9
2023 Privacy-Enhanced Decentralized Federated Learning at Dynamic Edge
abstract
Decentralized Federated Learning (DeFL) plays a critical role in improving effectiveness of training and has been proved to give great scope to the development of edge computing. However, on the one hand, inaccessibility of private data and excessively exploiting the data throughout the learning process have become a public concern, and on the other hand the connections between server-less edge devices are always varying due to the mobility of edge intelligent devices. To address the above issues, we propose aPrivacy-Enhanced -Dynamic -Decentralized -Federated -Learning algorithm called PED$ ^{2}$FL in a dynamic edge environment. We design the PED$ ^{2}$FL under the analog transmission scheme, where mobile edge devices transmit privacy preserving data simultaneously and accomplish efficient information aggregation with doubly-stochastic adjacent matrices. With thorough analysis, it can be demonstrated that PED$ ^{2}$FL satisfies$(\epsilon,\delta)$-differential privacy while the per-device privacy budget decays exponentially with the number of the neighbors, which greatly improved the data utility compared to the fixed budget in the orthogonal transmission strategy. PED$ ^{2}$FL has the same convergence rate$\mathcal {O}(\sqrt{\frac{1}{KN}})$as the non-private decentralized learning algorithm D-PSGD without enhanced privacy protection, where$K$and$N$are the total iterations and the number of nodes, respectively. Extensive experiments show that algorithm PED$ ^{2}$FL also performs well with real-world settings.
Shuzhen Chen 0001, Dongxiao Yu, Ju Ren 0001, Cong'an Xu, Yanwei Zheng
IEEE Trans. Computers4
2023 PrivAim: A Dual-Privacy Preserving and Quality-Aware Incentive Mechanism for Federated Learning
abstract
Privacy protection and incentive mechanism are two fundamental problems in federated learning (FL), which aim at protecting the privacy of data owners and stimulating them to share more resources, respectively. Recent works have proposed differential privacy (DP) based privacy-preserving incentive mechanisms to solve both problems simultaneously. However, almost all of them took the privacy level as the only incentive item, without considering other factors, such as data quantity and quality. Moreover, an untrusted server can further infer sensitive information from the bids that reflect the true costs of data owners. To solve these problems, in this paper, we propose a dual-privacy preserving and quality-aware incentive mechanism, PrivAim, for federated learning. Specifically, it utilizes differential privacy to protect the local models and true costs against the untrusted parameter server, and carefully designs a multi-dimensional reverse auction mechanism to incentivize data owners with high quality and low cost to participate in FL without knowing the true bids. We theoretically prove that PrivAim satisfies$\Delta b$-truthfulness, individual rational, computational efficiency, and differential privacy. Extensive experiments show that PrivAim can effectively protect bid privacy, and achieve at least 21% and 6% improvement on social welfare and model accuracy, respectively, compared to the state-of-the-art.
Dan Wang 0031, Ju Ren 0001, Zhibo Wang 0001, Yichuan Wang 0003, Yaoxue Zhang
IEEE Trans. Computers2
2023 Towards Class-Balanced Privacy Preserving Heterogeneous Model Aggregation
abstract
Heterogeneous model aggregation (HMA) is an effective paradigm that integrates on-device trained models heterogeneous in architecture and target task into a comprehensive model. Recent works adopt knowledge distillation to amalgamate the knowledge of learned features and predictions from heterogeneous on-device models to realize HMA. However, most of them ignore that the disclosure of learned features exposes on-device models to privacy attacks. Moreover, the aggregated model may suffer from the imbalanced supervision caused by the uneven distribution of amalgamated knowledge about each class and show class bias. In this article, to address these issues, we propose a response-based class-balanced heterogeneous model aggregation mechanism, called CBHMA. It can effectively achieve HMA in a privacy-preserving manner and alleviate class bias in the aggregated model. Specifically, CBHMA aggregates on-device models by using only their response information to reduce their privacy leakage risk. To mitigate the impact of imbalanced supervision, CBHMA quantitatively measures the imbalanced supervision level for each class. Based on that, CBHMA customizes fine-grained misclassification costs for each class and utilizes such costs to adjust the importance of each class (more importance to classes with weaker supervision) in the response-based HMA algorithm. Extensive experiments on two real-world datasets demonstrate the effectiveness of CBHMA.
Xiaoyi Pang, Zhibo Wang 0001, Zeqing He, Peng Sun 0003, Meng Luo 0010, Ju Ren 0001, Kui Ren 0001
IEEE Trans. Dependable Secur. Comput.6
2023 TRUCON: Blockchain-Based Trusted Data Sharing With Congestion Control in Internet of Vehicles
abstract
The Internet of vehicles (IoV) has a substantial impact on traffic efficiency improvement and accidents avoidance. Due to restricted resources, vehicles must share observed data with RSUs and other vehicles to execute some time-tolerant computing tasks. However, data provided by vehicles cannot always be trusted due to the presence of attackers. Fake messages could have catastrophic ramifications, such as vehicle collisions. Furthermore, extensive data sharing might cause channel congestion, resulting in the loss of vital messages during delivery. To overcome the aforementioned issues, we propose TRUCON, a blockchain-based trusted data sharing mechanism with congestion control in IoV. Firstly, we propose a Kademlia algorithm-based traffic data forwarding method to control channel congestion state. By adjusting the bucket size and distance threshold, source vehicles can limit the number of reference vehicles forwarded. Secondly, we present a cuckoo filter-based traffic data deduplication and discrimination approach. To avoid repetitive sharing, vehicles and RSUs can check their local filters to verify if the current data report has been shared. Based on the foregoing, we propose a blockchain-based trust management mechanism with congestion control. RSUs serve as full nodes while vehicles are light nodes in the blockchain. Finally, we develop a trust management prototype system with congestion control that incorporates both on-chain and off-chain parts. It signifies that our scheme is both feasible and effective.
Mingyang Yuan, Yang Xu 0013, Cheng Zhang 0035, Yunlin Tan, Yichuan Wang 0003, Ju Ren 0001, Yaoxue Zhang
IEEE Trans. Intell. Transp. Syst.6
2023 Efficient Dependent Task Offloading for Multiple Applications in MEC-Cloud System
abstract
With the proliferation of versatile mobile applications, offloading compute-intensive tasks to the MEC/Cloud becomes a dramatic technique due to the limited resources and high user experience requirements at mobile devices. However, most existing works design their task offloading schemes without considering the dependence of tasks and the orchestration of the MEC and Cloud, and thus may limit the system performance. In this paper, we propose a dependent task offloading framework for multiple mobile applications, named COFE, where mobile devices can offload their compute-intensive tasks with dependent constraints to the MEC-Cloud system. It can assign the offloaded tasks to the MEC and Cloud adaptively to improve the user experience. Based on COFE, we formulate the task offloading problem as an average makespan minimization problem, which is proved to be NP-hard. Then, we propose a heuristic ranking-based algorithm to assign the offloaded tasks according to their bottom levels. Theoretical analysis proves the stability of the system under the proposed algorithm and extensive simulations validate that the proposed algorithm can significantly reduce the average makespan and deadline violation probabilities of offloaded applications.
Jiagang Liu, Ju Ren 0001, Yongmin Zhang, Xuhong Peng, Yaoxue Zhang, Yuanyuan Yang 0001
IEEE Trans. Mob. Comput.2
2023 Towards Privacy-Driven Truthful Incentives for Mobile Crowdsensing Under Untrusted Platform
abstract
Reverse auction-based incentive mechanisms have been commonly proposed to stimulate mobile users to participate in crowdsensing, where users submit bids to the platform to compete for interested tasks. Recent works pointed out that bid is a private information which can reveal sensitive information of users (e.g., location privacy), and proposed bidding-preserving mechanisms with differential privacy against inference attack. However, all these mechanisms rely on a trusted platform, and would fail in bid protection completely when the platform is untrusted. In this paper, we design novel privacy-preserving incentive mechanisms to protect users’ true bid information against the honest-but-curious platform while minimizing the social cost of winner selection. To this end, instead of uploading the true bid to the platform, a differentially private bid obfuscation function is designed with the exponential mechanism, which helps each user to obfuscate bids locally and submit obfuscated bids to the platform. Two solutions are proposed for the platform to solve the winner selection problem with the obfuscated information, which is proved to be NP-hard. Moreover, we further propose a novel task-bid pair protection truthful incentive mechanism to further prevent privacy leakage from the set of interested tasks, where each user encrypts his interested tasks via homomorphic encryption locally, and an encrypted task clustering method is proposed to group users with the same interested tasks into the same cluster for winner selection with users’ encrypted task-bid pairs. Both of theoretical analysis and extensive experiments demonstrate the effectiveness of proposed mechanisms against the untrusted platform.
Zhibo Wang 0001, Jingxin Li, Jiahui Hu 0001, Ju Ren 0001, Qian Wang 0002, Zhetao Li, Yanjun Li 0004
IEEE Trans. Mob. Comput.4
2023 Distributed Pricing and Bandwidth Allocation in Crowdsourced Wireless Community Networks
abstract
With the rapid growth of global mobile data traffic, Wi-Fi plays an increasingly important role in expanding network capacity. To overcome the geographical coverage limit of Wi-Fi APs, especially for mobile users, the crowdsourced wireless community network has emerged as a cost-effcient way for providing Internet access services. For instance, it is plausible to share their private residential Wi-Fi APs with each other by designing some tailored incentive/pricing mechanisms. Thus motivated, we propose a distributed pricing and bandwidth allocation scheme to maximize the profit of Wi-Fi providers and provide better Internet services to mobile users. Firstly, we study the stationary networks with incomplete information of users and propose distributed pricing and bandwidth allocation algorithms for single-AP regions and AP group regions, respectively. Then, we generalize the study to dynamic networks and explore distributed pricing based on the statistics of users mobility. Further, we design an online bandwidth allocation algorithm according to the real-time user information. Simulation results demonstrate that the proposed distributed pricing and bandwidth allocation scheme, comparing with the operators pricing scheme, has a better performance on both Wi-Fi APs profit and user experience.
Yongmin Zhang, Cenchen Ji, Nan Qiao 0008, Ju Ren 0001, Yaoxue Zhang, Yuanyuan Yang 0001
IEEE Trans. Mob. Comput.4
2023 MVPose: Realtime Multi-Person Pose Estimation Using Motion Vector on Mobile Devices
abstract
We present MVPose, a novel system designed to enable real-time multi-person pose estimation (PE) on commodity mobile devices, which consists of three novel techniques. First, MVPose takes a motion-vector-based approach to fast and accurately track the human keypoints across consecutive frames, rather than running expensive human-detection model and pose-estimation model for every frame. Second, MVPose designs a mobile-friendly PE model that uses lightweight feature extractors and multi-stage network to significantly reduce the latency of pose estimation without compromising the model accuracy. Third, MVPose leverages the heterogeneous computing resources of both CPU and GPU to execute the pose estimation model for multiple persons in parallel, which further reduces the total latency. We present extensive experiments to evaluate the effectiveness of the proposed tecniques by implemented the MVPose on five off-the-shelf commercial smartphones. Evaluation results show that MVPose achieves over30frames per second PE with4persons per frame, which significantly outperforms the state-of-the-art baseline, with a speedup of up to5.7×and3.8×in latency on CPU and GPU, respectively. Compared with baseline, MVPose achieves an improvement of10.1%in multi-person PE accuracy. Furthermore, MVPose achieves up to74.3%and57.6%energy-per-frame saving on average in comparison with the baseline on mobile CPU and GPU, respectively.
Yunxin Liu 0001, Ju Ren 0001, Xiaohui Xu, Fucheng Jia, Yaoxue Zhang
IEEE Trans. Mob. Comput.5
2023 Efficient Revenue-Based MEC Server Deployment and Management in Mobile Edge-Cloud Computing
abstract
With the explosive growth of mobile applications, the development of mobile edge computing (MEC) has been greatly promoted since it can ably improve the quality of service for mobile applications by providing low latency and high-quality computation services. Most existing works focus on improving the efficiency of MEC with an assumption that the MEC servers have already been deployed. However, without appropriate deployment of MEC servers, the profitability of the MEC system can be significantly restrained, which hinders the rapid promotion of the MEC. To address this issue, we formulate an MEC server deployment problem for the MEC operator as a revenue maximization problem. Firstly, we model and analyze the various factors that affect the revenue. Secondly, we formulate a revenue maximization problem, which is NP-hard, but it is proved to be convex with respect to the total available computation units. Based on this feature, we propose a three-layer optimization algorithm, named EDM, in which the location, the deployed computation units, and the wholesaled computation resources are determined gradually, to maximize the total revenue. Experimental results demonstrate that the proposed EDM algorithm has significant advantages on revenue improvement compared to competitive benchmarks.
Yongmin Zhang, Wei Wang 0343, Ju Ren 0001, Jinge Huang, Shibo He, Yaoxue Zhang
IEEE/ACM Trans. Netw.3
2022 HSFL: An Efficient Split Federated Learning Framework via Hierarchical Organization
abstract
Federated learning (FL) has emerged as a popular paradigm for distributed machine learning among vast clients. Unfortunately, resource-constrained clients often fail to participate in FL because they cannot pay for the memory resources required for model training due to their limited memory or bandwidth. Split federated learning (SFL) is a novel FL framework in which clients commit intermediate results of model training to a cloud server for client-server collaborative training of models, making resource-constrained clients also eligible for FL. However, existing SFL frameworks mostly require frequent communication with the cloud server to exchange intermediate results and model parameters, which results in significant communication overhead and elongated training time. In particular, this can be exacerbated by the imbalanced data distributions of clients. To tackle this issue, we propose HSFL, a hierarchical split federated learning framework that efficiently trains SFL model through hierarchical organization participants. Under the HSFL framework, we formulate a Cloud Aggregation Time Minimization (CATM) problem to minimize the global training time and design a light-weight client assignment algorithm based on dynamic programming to solve it. Moreover, we develop a self-adaption approach to cope with the dynamic computational resources of clients. Finally, we implement and evaluate HSFL on various real-world training tasks, elaborating on its effectiveness and superiority in terms of efficiency and accuracy compared to baselines.
Tengxi Xia, Yongheng Deng, Sheng Yue 0001, Ju Ren 0001, Yaoxue Zhang
CNSM5
2022 Privacy-Preserving DNN Model Authorization against Model Theft and Feature Leakage
abstract
Today’s intelligent services are built on well-trained deep neural network (DNN) models, which usually require large private datasets along with a high cost for model training. It consequently makes the model providers cherish the pre-trained DNN models and only distribute them to authorized users. However, malicious users can steal these valuable models for abuse, illegal copy and redistribution. Attackers can also extract private features from even authorized models to leak partial training datasets. They both violate privacy. Existing techniques from secure community attempt to avoid parameter leakage during model authorization but yet cannot solve privacy issues sufficiently. In this paper, we propose a privacy-preserving model authorization approach, AgAuth, to resist the aforementioned privacy threats. We devise a novel scheme called Information-Agnostic Conversion (IAC) for forwarding procedure to eliminate residual features in model parameters. Based on it, we then propose Inference-on-Ciphertext (CiFer) mechanism for DNN reasoning, which includes three stages in each forwarding. The Encrypt phase first converts the proprietary model parameters to demonstrate uniform distribution. The Forward stage per-forms forwarding function without decryption at authorized side. Specifically, this stage just computes over ciphertext. The Decrypt phase finally recovers the information-agnostic outputs to informative output tensor for real-world services. In addition, we implement a prototype and conduct extensive experiments to evaluate its performance. The qualitative and quantitative results demonstrate that our solution AgAuth is privacy-preserving to defend against model theft and feature leakage, without accuracy loss or notable performance decrease.
Qiushi Li 0002, Ju Ren 0001, Yue-Zhi Zhou, Yaoxue Zhang
ICC2
2022 Mobility-Aware Service Migration for Seamless Provision: A Reinforcement Learning Approach
abstract
Mobile Edge Computing (MEC) is a promising paradigm to support high-quality time-sensitive applications. In this paper, we investigate the service migration (i.e., whether, when, and where to migrate the services) to seamlessly serve mobile users in small-cell MEC systems. The service migration is formulated as an optimization problem to minimize the long-term system average delay that consists of queuing, communication, and migration delays. Considering the dynamic user mobility and network conditions, the formulated problem is non-convex and difficult to solve in real time. To this end, we propose a Mobility-aware Service Migration scheme, named MSM, to make real-time decisions on service migrations by utilizing reinforcement learning (RL) approaches. Specifically, we first design a user classification mechanism based on users’ mobility patterns to reduce the complexity of decision-making. We then formulate the service migration as a Markov decision process and devise an RL-based framework to make service migration decisions in real time in the dynamic MEC environment. Extensive data-driven experiments demonstrate the efficacy of MSM in reducing the system average delay.
Feng Lyu 0001, Fan Wu 0014, Huaqing Wu, Ju Ren 0001, Yaoxue Zhang, Xuemin Shen
ICC5
2022 ENIGMA: Low-Latency and Privacy-Preserving Edge Inference on Heterogeneous Neural Network Accelerators
abstract
Time-efficient artificial intelligence (AI) service has recently witnessed increasing interest from academia and industry due to the urgent needs in massive smart applications such as self-driving cars, virtual reality, high-resolution video streaming, etc. Existing solutions to reduce AI latency, like edge computing and heterogeneous neural-network accelerators (NNAs), face high risk of privacy leakage. To achieve both low-latency and privacy-preserving purposes on edge servers (e.g., NNAs), this paper proposes ENIGMA that can exploit the trusted execution environment (TEE) and heterogeneous NNAs of edge servers for edge inference. The low-latency is supported by a new ahead-of-time analysis framework for analyzing the linearity of multilayer neural networks, which automatically slices forward-graph and assigns sub-graphs to TEE or NNA. To avoid privacy leakage issue, we then introduce a pre-forwarded cipher generation (PFCG) scheme for computing linear sub-forward-graphs on NNA. The input data is encrypted to ciphertext that can be computed directly by linear sub-graphs, and the output can be decrypted to obtain the correct output. To enable non-linear computation of sub-graphs on TEE, we use ring-cache and automatic vectorization optimization to address the memory limitation of TEE. Qualitative analysis and quantitative experiments on GPU, NPU and TPU demonstrate that ENIGMA is not only compatible with heterogeneous NNAs, but also can avoid leakages of private features with latency as low as 50-milliseconds.
Qiushi Li 0002, Ju Ren 0001, Xinglin Pan, Yue-Zhi Zhou, Yaoxue Zhang
ICDCS2
2022 CODE: Compact IoT Data Collection with Precise Matrix Sampling and Efficient Inference
abstract
It is unpractical to conduct full-size data collection in ubiquitous IoT data systems due to the energy constraints of IoT sensors and large system scales. Although sparse sensing technologies have been proposed to infer missing data based on partial sampled data, they usually focus on data inference while neglecting the sampling process, restraining the inference efficiency. In addition, their inferring methods highly depend on data linearity correlations, which become less effective when data are not linearly correlated. In this paper, we propose, Compact IOT Data CollEction, namely CODE, to conduct precise data matrix sampling and efficient inference. Particularly, CODE integrates two major components, i.e., cluster-based matrix sampling and Generative Adversarial Networks (GAN)-based matrix inference, to reduce the data collection cost and guarantee the data benefits, respectively. In the sampling component, a cluster-based sampling approach is devised, in which data clustering is first conducted and then a two-step sampling is performed in accordance with the number of clusters and clustering errors. For the inference component, a GAN-based model is developed to estimate the full matrix, which consists of a generator network that learns to generate a fake matrix, and a discriminator network that learns to discriminate the fake matrix from the real one. A reference implementation of CODE is conducted under three operational large-scale IoT systems, and extensive data-driven experiment results are provided to demonstrate its efficiency and robustness.
Huali Lu, Feng Lyu 0001, Ju Ren 0001, Jiadi Yu, Fan Wu 0014, Yaoxue Zhang, Xuemin Shen
ICDCS3
2022 Kalmia: A Heterogeneous QoS-aware Scheduling Framework for DNN Tasks on Edge Servers
abstract
Motivated by the popularity of edge intelligence, DNN services have been widely deployed at the edge, posing significant performance pressure on edge servers. How to improve the QoS of edge DNN services becomes a crucial and challenging problem. Previous works, however, did not fully consider the heterogeneous QoS requirements on urgent and non-urgent tasks, causing frequent QoS violations. Meanwhile, our empirical study shows that severe task interference exists in concurrent DNN tasks, further degrading the timeliness of urgent tasks and throughput of non-urgent tasks. To address these issues, we propose Kalmia, a heterogeneous QoS-aware framework for DNN inference task scheduling on edge servers. Specifically, Kalmia includes an offline profiling stage and an online scheduling policy. In offline profiling, we build a regression model to predict the execution time of tasks. During online scheduling, we classify the tasks into urgent and non-urgent tasks and distribute them into two CUDA contexts. By a tailored scheduling strategy, non-urgent tasks can fully utilize the computing resources for throughput improvement, while the timeliness of urgent tasks can be guaranteed via preemption. Experimental results demonstrate that Kalmia can achieve up to 2.8× improvement in throughput and significantly reduce the deadline violation rate compared with state-of-the-art methods.
Ziyan Fu 0001, Ju Ren 0001, Yue-Zhi Zhou, Yaoxue Zhang
INFOCOM2
2022 Towards Online Privacy-preserving Computation Offloading in Mobile Edge Computing
abstract
Mobile Edge Computing (MEC) is a new paradigm where mobile users can offload computation tasks to the nearby MEC server to reduce their resource consumption. Some works have pointed out that the true amount of offloaded tasks may reveal the sensitive information (e.g., device usage pattern and location information) of users, and proposed several privacy-preserving offloading mechanisms. However, to the best of our knowledge, none of them can provide strict and provable privacy guarantee. In this paper, we focus on the privacy leakage issue in computation offloading in MEC with a honest-but-curious server, and propose a novel online privacy-preserving computation offloading mechanism, called OffloadingGuard, to generate efficient offloading strategies for users in real time, which provide strict user privacy guarantee while minimizing the total cost of task computation. To this end, we design a deep reinforcement learning-based offloading model which allows each user to adaptively determine the satisfactory perturbed offloading ratio according to the time-varying channel state at each time slot to achieve trade-off between user privacy and computation cost. In particular, to strictly protect the true amount of offloaded tasks and prevent the untrusted MEC server from revealing mobile users’ privacy, a range-constrained Laplace distribution is designed to obfuscate the original offloading ratio of each user and restrict the perturbed offloading ratio in a rational range. OffloadingGuard is proved to satisfy ϵ-differential privacy, and extensive experiments demonstrate its effectiveness.
Xiaoyi Pang, Zhibo Wang 0001, Jingxin Li, Ruiting Zhou, Ju Ren 0001, Zhetao Li
INFOCOM5
2022 An Efficient Two-Layer Task Offloading Scheme for MEC System with Multiple Services Providers
abstract
With the explosive growth of mobile and Internet of Things (IoT) applications, increasing Mobile Edge Computing (MEC) systems have been developed by diverse Edge Service Providers (ESPs), opening a new computing market with stiff competition. However, considering the spatiotemporally varying features of computation tasks, taking over all the received tasks alone may greatly degrade the service performance of the MEC system and lead to poor economical benefit. To this end, this paper proposes a two-layer collaboration model for ESPs. Each ESP can balance the computation workload among the internal edge nodes from the ESP and offload part of computation tasks to the ESP external edge servers from other ESPs. For internal load balancing, we propose a task balancing scheme based on the Alternating Direction Method of Multipliers (ADMM) to manage the computation tasks within the edge nodes of the ESP, such that the computation delay can be minimized. For external task offloading, we formulate a game-based pricing and task allocation scheme to derive the best game strategy, aiming at maximizing the total revenue of each ESP. Extensive simulation results demonstrate that the proposed schemes can achieve improved performance in terms of system revenue and stability, as well as computation delay.
Ju Ren 0001, Jiani Liu 0005, Yongmin Zhang, Feng Lyu 0001, Zhibo Wang 0001, Yaoxue Zhang
INFOCOM1
2022 Boosting Internet Card Cellular Business via User Portraits: A Case of Churn Prediction
abstract
Internet card (IC) as a new business model emerges, which penetrates rapidly and holds the potential to foster a great business market. However, the understanding of IC user portraits is insufficient, which is the building block to boost the IC business. In this paper, we take the lead to bridge the gap by studying one large-scale dataset collected from a provincial network operator of China, which contains about 4 million IC users and 22 million traditional card (TC) users. Particularly, we first conduct a systematical analysis on usage data by investigating the difference of two types of users, examining the impact of user properties, and characterizing the spatio-temporal networking patterns. After that, we shed light on one specific business case of churn prediction by devising an IC user Churn Prediction model, named ICCP, which consists of a feature extraction component and a learning architecture design. In ICCP, both the static portrait features and temporal sequential features are extracted, and one principal component analysis block and the embedding/transformer layers are devised to learn the respective information of two types of features, which are collectively fed into the classification multilayer perceptron layer for prediction. Extensive experiments corroborate the efficacy of ICCP.
Fan Wu 0014, Ju Ren 0001, Feng Lyu 0001, Peng Yang 0004, Yongmin Zhang, Yaoxue Zhang
INFOCOM2
2022 Target-oriented Semi-supervised Domain Adaptation for WiFi-based HAR
abstract
Incorporating domain adaptation is a promising solution to mitigate the domain shift problem of WiFi-based human activity recognition (HAR). The state-of-the-art solutions, however, do not fully exploit all the data, only focusing either on unlabeled samples or labeled samples in the target WiFi environment. Moreover, they largely fail to carefully consider the discrepancy between the source and target WiFi environments, making the adaptation of models to the target environment with few samples become much less effective. To cope with those issues, we propose a Target-Oriented Semi-Supervised (TOSS) domain adaptation method for WiFi-based HAR that can effectively leverage both labeled and unlabeled target samples. We further design a dynamic pseudo label strategy and an uncertainty-based selection method to learn the knowledge from both source and target environments. We implement TOSS with a typical meta learning model and conduct extensive evaluations. The results show that TOSS greatly outperforms state-of-the-art methods under comprehensive 1 on 1 and multi-source one-shot domain adaptation experiments across multiple real-world scenarios.
Feng Wang 0001, Jihong Yu, Ju Ren 0001, Zhi Wang 0001, Wei Gong 0001
INFOCOM4
2022 Dynamic Pricing Scheme for Edge Computing Services: A Two-layer Reinforcement Learning Approach
abstract
Edge computing servers (ECSs) have been widely deployed in large-scale mobile edge computing (MEC) systems, which can provide nearby computing services by charging users a price. Service pricing schemes can regulate user task offloading and affect the total revenue of service providers. Investigating how to maximize the revenue of service provider and improve the utilization of edge computing resources becomes crucial while is challenging, considering the users mobility and the uncertainty of users service requests. In this paper, we model the dynamic pricing process of ECS as a Markov decision process and propose a dynamic pricing approach based on Dueling Double Deep Q Network (D3QN) by using the current load conditions and user characteristics, the goal of which is to maximize the revenue of service provider. In addition, considering more ECSs in the MEC system, with the dynamic variations of ECSs loads and the different arrival rate of user tasks, we propose a joint scheduling approach based on D3QN (called RLJS) to collectively improve the total service revenue of service providers. Specifically, we first use a data-driven method to group the ECSs and then devise a D3QN-based task scheduling scheme to distribute tasks among ECS groups by considering the load and price conditions in real time. Simulation results demonstrate the efficacy of RLJS in improving the total revenue of the system provider and reducing the user delays.
Feng Lyu 0001, Xinyao Cai, Fan Wu 0014, Huali Lu, Sijing Duan, Ju Ren 0001
IWQoS6
2022 FastPR: One-stage Semantic Person Retrieval via Self-supervised Learning
abstract
Semantic person retrieval aims to locate a specific person in an image with the query of semantic descriptions, which has shown great significance in surveillance and security applications. Prior arts commonly adopt a two-stage method that first extracts the persons with a pretrained detector and then finds the target matching the descriptions optimally.However, existing works suffer from high computational complexity and low recall rate caused by error accumulation in the two-stage inference. To solve the problems, we propose FastPR, a one-stage semantic person retrieval method via self-supervised learning, to optimize the person localization and semantic retrieval simultaneously. Specifically, we propose a dynamic visual-semantic alignment mechanism which utilizes grid-based attention to fuse the cross-modal features, and employs a label prediction proxy task to constrain the attention process. To tackle the challenges that real-world surveillance images may suffer from low-resolution and occlusion, and the target persons may be within a crowd,we further propose a dual-granularity person localization module through designing an upsampling reconstruction proxy task to enhance the local feature of the target person in the fused features, followed by a tailored offset prediction proxy task to make the localization network capable of accurately identifying and distinguishing the target person in a crowd. Experimental results demonstrate that FastPR achieves the best retrieval accuracy compared to the state-of-the-art baseline methods, with over 15 times inference time reduction.
Ju Ren 0001, Xin Wang 0019, Wenwu Zhu 0001, Yaoxue Zhang
ACM Multimedia2
2022 MobiDepth: real-time depth estimation using on-device dual cameras
abstract
Real-time depth estimation is critical for the increasingly popular augmented reality and virtual reality applications on mobile devices. Yet existing solutions are insufficient as they require expensive depth sensors or motion of the device, or have a high latency. We propose MobiDepth, a real-time depth estimation system using the widely-available on-device dual cameras. While binocular depth estimation is a mature technique, it is challenging to realize the technique on commodity mobile devices due to the different focal lengths and unsynchronized frame flows of the on-device dual cameras and the heavy stereo-matching algorithm.
Ju Ren 0001, Bangwen He, Ting Cao 0003, Yuanchun Li 0003, Yaoxue Zhang, Yunxin Liu 0001
MobiCom3
2022 CoDL: efficient CPU-GPU co-execution for deep learning inference on mobile devices
abstract
Concurrent inference execution on heterogeneous processors is critical to improve the performance of increasingly heavy deep learning (DL) models. However, available inference frameworks can only use one processor at a time, or hardly achieve speedup by concurrent execution compared to using one processor. This is due to the challenges to 1) reduce data sharing overhead, and 2) properly partition each operator between processors.
Fucheng Jia, Ting Cao 0003, Shiqi Jiang 0002, Yunxin Liu 0001, Ju Ren 0001, Yaoxue Zhang
MobiSys6
2022 TailorFL: Dual-Personalized Federated Learning under System and Data Heterogeneity
abstract
Federated learning (FL) enables distributed mobile devices to collaboratively learn a shared model without exposing their raw data. However, heterogeneous devices usually have limited and different available resources, i.e., system heterogeneity, for model training and communicating, while the diverse data distribution among devices, i.e., data heterogeneity, may result in significant performance degradation. In this paper, we propose TailorFL, a dual-personalized FL framework, which tailors a submodel for each device with personalized structure for training and personalized parameters for local inference. To achieve this, we first excavate the personalization principle for data heterogeneous FL via in-depth empirical studies, and based on which, we propose a resource-aware and data-directed pruning strategy that makes each device's submodel structure match its resource capability and correlate with its local data distribution. To aggregate the submodels while preserving their dual personalization properties, we design a scaling-based aggregation strategy that scales parameters with the pruning rate of submodels and aggregates the overlapped parameters. Moreover, to further promote beneficial and restrain detrimental collaborations among devices, we propose a server-assisted model-tuning mechanism, which dynamically tunes device's submodel structure at the server side with the global view of device's data distribution similarities. Extensive experiments demonstrate that compared to the status quo approaches, TailorFL achieves an average of 22% increase in inference accuracy, and reduces the memory, computation, and communication costs for model training simultaneously.
Yongheng Deng, Weining Chen, Ju Ren 0001, Feng Lyu 0001, Yang Liu 0165, Yunxin Liu 0001, Yaoxue Zhang
SenSys3
2022 Hyperion: A Generic and Distributed Mobile Offloading Framework on OpenCL
abstract
Despite the significant development of mobile device SoCs, they are still inefficient in computing computation-intensive workloads, such as high-resolution image processing and AR/VR applications. Offloading offers a promising way to leverage cloud or edge servers for acceleration, but existing offloading is limited to specific tasks or specific hardware/software platforms, resulting in significant engineering overhead. To address this problem, we focus on the underlying layer of these applications (i.e., OpenCL) and propose Hyperion, a generic and distributed mobile offloading framework built on OpenCL. To achieve high-performance distributed execution for Hyperion, we first take a deep insight into the OpenCL data structures and design regularity-aware kernel analyzer to analyze the data dependency of work-groups and identify the essential data to offload. Then, context-aware execution time predictor is proposed to estimate the computing time of a given partitioned kernel workload that is highly impacted by many runtime factors. These techniques are integrated into pipeline-enabled and network-adaptive scheduler to make scheduling decisions, which coordinates the kernel partition and workload scheduling to form pipeline processing between data transmission and distributed execution with flexible adaptability to network dynamics. Extensive experimental results demonstrate that Hyperion achieves superior performance with an average 3.80× speedup compared with the best baseline and flexible adaptation to dynamic network conditions and available computing resources.
Ziyan Fu 0001, Ju Ren 0001, Yunxin Liu 0001, Ting Cao 0003, Yue-Zhi Zhou, Yaoxue Zhang
SenSys2
2022 SVoice: Enabling Voice Communication in Silence via Acoustic Sensing on Commodity Devices
abstract
Silent Speech Interface (SSI) has been proposed as a means of reconstructing audible speech from silent articulatory gestures for covert voice communication in public and voice assistance for the aphasic. Prior arts of SSI, either relying on wearable devices or cameras, may lead to extended contact requirements or privacy leakage risks. The recent advances in acoustic sensing have brought new opportunities for sensing gestures, but their original intention is to infer speech content for classification instead of audible speech reconstruction, resulting in the loss of some important speech information (e.g., speech rate, intonation, and emotion). In this paper, we propose, the first system that supports accurate audible speech reconstruction by analyzing the disturbance of tiny articulatory gestures on the reflected ultrasound signal. The design of introduces a new model that provides the unique mapping relationship between ultrasound and speech signals, so that the audible speech can be successfully reconstructed from the silent speech. However, establishing the mapping relationship depends on plenty of training data. Instead of the time-consuming collection of massive amounts of data for training, we construct an inverse task that constitutes a dual form with the original task to generate virtual gestures from widely available audio (e.g., phone calls) for facilitating model training. Furthermore, we introduce a fine-tuning mechanism using unlabeled data for user adaptation. We implement using a portable smartphone and evaluate it in various environments. The evaluation results show that can reconstruct speech with a (Character Error Rate) CER as low as 7.62%, and decrease the CER from 82.77% to 9.42% on new users with only 1 hour of ultrasound signals provided, which outperforms state-of-the-art acoustic-based approaches while preserving rich speech information.
Yongjian Fu 0004, Shuning Wang, Linghui Zhong, Ju Ren 0001, Yaoxue Zhang
SenSys5
2022 Online Market Mechanism for Mobile Data Rate Trading With Temporal Constraints
abstract
User-initiated mobile data trading, where mobile devices trade their mobile data quota via personal hotspots, is a promising approach to improve resource utilization. Most existing works only consider the data size, while ignoring the data rate and temporal requirements. To fill this void, we propose a novel data trading marketplace, where mobile users trade Internet access continuously for a time period with a specific data rate with neighboring mobile devices. Each request is characterized by an arrival time, departure time, the demanded data rate, and a value for getting services. To achieve the most system efficiency, we formulate an integer linear programming problem to maximize the total social welfare, which takes the data rate and temporal requirements into account. We next consider two request models: 1) a homogeneous request model and 2) a heterogeneous request model. In the homogeneous request model, all the requests demand the overall system lifetime, and we propose a computationally efficient auction that makes allocation decisions for all the requests simultaneously. In the heterogeneous request model, all the requests require different Internet access periods and dynamic arrive. Upon requests’ arrival, the system must make real-time allocations without the availability of future information. To jointly deal with requesters’ multidimensional private information (i.e., the arrival/departure time, demanded data rate, and the value), and the uncertainty about future arrival requests, we propose a multi-round online auction. Theoretical analysis shows that both the auctions satisfy the desired properties, including individual rationality, truthfulness, and computational efficiency. Simulation results show the efficiency of the proposed auctions.
Di Zhang 0010, Ju Ren 0001, Yue-Zhi Zhou, Yaoxue Zhang
IEEE Internet Things J.3
2022 RLSS: A Reinforcement Learning Scheme for HD Map Data Source Selection in Vehicular NDN
abstract
In the autonomous driving era, high-definition (HD) maps are an essential building block to enable fine-grained environmental perception, precise localization, and path planning. However, with rich multidimensional information, the size of HD map data is huge and cannot be stored onboard, where the dynamic map data need to be distributed in real time via vehicular networks and how to design the distribution mechanism (i.e., determining the data source for requests) becomes crucial. For the end-to-end communication protocols (i.e., TCP/IP), the main limitation is the vehicle mobility and high dynamic of the network topology, which can degrade the transmission performance dramatically. Therefore, in this article, we propose a reinforcement learning-based data source selection scheme, named RLSS, for efficient HD map distribution in vehicular named data networking (NDN) scenarios, which aims at seeking the best data source (roadside infrastructures or nearby vehicles) in accordance with the map data requests. Specifically, in RLSS, we adopt a deep reinforcement learning-based architecture to learn a neural network as an agent to make the decision of data source selection, which can work online after offline training based on historical selection action performance. In addition, to solve the “cold start” problem for a new vehicle, we propose a model aggregation algorithm and weight update approach to learn the model parameters from its nearby vehicles, which can guarantee the performance while saving the communication cost. Finally, we implement RLSS in NS-3 by adopting the tools of the ndnSIM, SUMO, and Gym. Extensive simulations demonstrate that RLSS can significantly improve the transmission performance in terms of delay, throughput, and packet loss when compared with state-of-the-art data source selection schemes.
Fan Wu 0014, Wang Yang 0002, Jialun Lu, Feng Lyu 0001, Ju Ren 0001, Yaoxue Zhang
IEEE Internet Things J.5
2022 Blockchain-Based Trustworthy Energy Dispatching Approach for High Renewable Energy Penetrated Power Systems
abstract
Renewable energy sources (RES) and low-carbon technology users play a vital role in modern power systems. However, RES generation is easily affected by the environment. Meanwhile, the load, such as electric vehicles (EVs) and prosumers, accounts for most low-carbon technology users. Their power is usually superimposed on peak loads without dispatching, which also exacerbates the instability of the power system. Current optimal dispatching mechanisms mainly rely on centralized organizations, while their dispatching process is not open and transparent. In this article, we propose a blockchain-based trustworthy dispatching approach for the distribution network in high renewable energy penetrated power systems. We first develop an optimal dispatching model considering EVs’ charging behavior and the prosumers’ economic benefits. With the model, prosumers can be dispatched to balance power and consume renewable energy, reducing the impact of disorderly charging on the grid and the abandonment of RES generation. An orderly charging iteration optimization (OCIO) algorithm is proposed to implement orderly EV charging while considering the charging cost and the period. We also propose a modified particle swarm optimization (mPSO) algorithm to publish dispatching tasks based on real-time power balance. Furthermore, blockchain is applied as an open and transparent ledger to record each entity’s power generation and consumption information, ensuring that the dispatching process is trustworthy. Finally, the effectiveness of the dispatching approach is verified in the modified IEEE 33-bus test system and Ethereum-based smart contracts.
Yang Xu 0013, Cheng Zhang 0035, Ju Ren 0001, Yaoxue Zhang, Xuemin Shen
IEEE Internet Things J.4
2022 Efficient Federated Meta-Learning Over Multi-Access Wireless Networks
abstract
Federated meta-learning (FML) has emerged as a promising paradigm to cope with the data limitation and heterogeneity challenges in today’s edge learning arena. However, its performance is often limited by slow convergence and corresponding low communication efficiency. In addition, since the available radio spectrum and IoT devices’ energy capacity are usually insufficient, it is crucial to control the resource allocation and energy consumption when deploying FML in practical wireless networks. To overcome the challenges, in this paper, we rigorously analyze the contribution of each device to the global loss reduction in each round and develop an FML algorithm (called NUFM) with a non-uniform device selection scheme to accelerate the convergence. After that, we formulate a resource allocation problem integrating NUFM in multi-access wireless systems to jointly improve the convergence rate and minimize the wall-clock time along with energy cost. By deconstructing the original problem step by step, we devise a joint device selection and resource allocation strategy to solve the problem with theoretical guarantees. Further, we show that the computational complexity of NUFM can be reduced from$O(d^{2})$to$O(d)$(with the model dimension$d$) via combining two first-order approximation techniques. Extensive simulation results demonstrate the effectiveness and superiority of the proposed methods in comparison with existing baselines.
Sheng Yue 0001, Ju Ren 0001, Jiang Xin, Yaoxue Zhang, Weihua Zhuang
IEEE J. Sel. Areas Commun.2
2022 A Blockchain-Based Multi-Cloud Storage Data Auditing Scheme to Locate Faults
abstract
Network storage services have benefited countless users worldwide due to the notable features of convenience, economy and high availability. Since a single service provider is not always reliable enough, more complex multi-cloud storage systems are developed for mitigating the data corruption risk. While a data auditing scheme is still needed in multi-cloud storage to help users confirm the integrity of their outsourced data. Unfortunately, most of the corresponding schemes rely on trusted institutions such as the centralized third-party auditor (TPA) and the cloud service organizer, and it is difficult to identify malicious service providers after service disputes. Therefore, we present a blockchain-based multi-cloud storage data auditing scheme to protect data integrity and accurately arbitrate service disputes. We not only introduce the blockchain to record the interactions among users, service providers, and organizers in data auditing process as evidence, but also employ the smart contract to detect service dispute, so as to enforce the untrusted organizer to honestly identify malicious service providers. We also use the blockchain network and homomorphic verifiable tags to achieve the low-cost batch verification without TPA. Theoretical analyses and experiments reveal that the scheme is effective in multi-cloud environments and the cost is acceptable.
Cheng Zhang 0035, Yang Xu 0013, Yupeng Hu 0004, Jiajing Wu, Ju Ren 0001, Yaoxue Zhang
IEEE Trans. Cloud Comput.5
2022 GBLinks: GNN-Based Beam Selection and Link Activation for Ultra-Dense D2D mmWave Networks
abstract
In this paper, we consider the problem of joint beam selection and link activation across a set of communication pairs to effectively control the interference between communication pairs via inactivating part communication pairs in ultra-dense device-to-device (D2D) mmWave communication networks. The resulting optimization problem is formulated as an integer programming problem that is nonconvex and NP-hard. Consequently, the global optimal solution, even the local optimal solution, cannot be generally obtained. To overcome this challenge, this paper resorts to design a deep learning architecture based on graph neural network to finish the joint beam selection and link activation, with taking the network topology information into account. Meanwhile, we present an unsupervised Lagrangian dual learning framework to train the parameters of the GBLinks model. Numerical results show that the proposed GBLinks model can converge to a stable point with the number of iterations increases, in terms of the weighted sum rate. Furthermore, the GBLinks model can reach near-optimal solutions through comparing with the exhaustive scheme in small-scale ultra-dense D2D mmWave communication networks and outperforms GreedyNoSched and the SCA-based method. It also shows that the GBLinks model can generalize to varying network densities and network coverage regions of ultra-dense D2D mmWave communication networks.
Shiwen He, Shaowen Xiong, Wei Zhang 0001, Yiting Yang, Ju Ren 0001, Yongming Huang 0001
IEEE Trans. Commun.5
2022 Service-Oriented Dynamic Resource Slicing and Optimization for Space-Air-Ground Integrated Vehicular Networks
abstract
In this paper, we study Space-Air-Ground integrated Vehicular Network (SAGVN), and propose an online control framework to dynamically slice the SAG spectrum resource for isolated vehicular services provisioning. In particular, at a given time slot, the system makes online decisions on the request admission and scheduling, UAV dispatching, and resource slicing for different services. To characterize the impact of those parameters, we construct a time-averaged queue stability criteria by taking queue backlogs of all services into consideration, and formulate a system revenue function which incorporates the time-averaged system throughput and UAV dispatching cost. The objective is to maximize the system revenue while stabilizing the time-averaged queue, which falls into the scope of Lyapunov optimization theory. By bounding the drift-plus-penalty, the original problem can be decoupled into four independent subproblems, each of which is readily solved. The merits of our control framework are three-fold: 1) the system is able to admit and process as many requests as possible (i.e., maximizing the time-averaged throughput); 2) the time-averaged UAV dispatching cost is minimized; and 3) service queues are stabilized in the long-term. Extensive simulations are carried out, and the results demonstrate that the control framework can effectively achieve the system revenue maximization and queueing stabilization. Moreover, it can balance the trade-off among system throughput, UAV dispatching cost, and queueing states via parameter tuning. Compared with the fixed slicing, our dynamic slicing can react to the vehicular environment rapidly and achieve an average 26% of throughput improvement.
Feng Lyu 0001, Peng Yang 0004, Huaqing Wu, Conghao Zhou, Ju Ren 0001, Yaoxue Zhang, Xuemin Shen
IEEE Trans. Intell. Transp. Syst.5
2022 Multitype Highway Mobility Analytics for Efficient Learning Model Design: A Case of Station Traffic Prediction
abstract
The provincial highway transportation system supports substantial cross-city transitions of people and logistics, where the prediction tasks in terms of station/road traffic, urban transitions, and individual traveling are crucial for boosting data intelligence. However, to achieve efficient prediction model design, the predictability analytics with data is the basis, but has not been sufficiently investigated in the existing literature yet. To bridge this gap, in this paper, we study one large-scale dataset collected from one provincial highway transportation system, which contains totally 21,685,765 vehicles and 351,766,743 transaction records, and conduct a comprehensive mobility analytics on its predictable performance. We first investigate the station traffic by mining its spatio-temporal correlations, then examine the multi-type urban transition flows (i.e., people flows and logistics) by demystifying the difference and similarity between the two types of behaviors, and finally analyze the uncertainty of individual traveling behaviors in terms of the destination and arriving time. After that, in accordance with the analytical findings, we cast a case study of data-driven model design for station traffic prediction. Specifically, a novel learning model is devised, named STAR, i.e., Spatio-Temporal Attention based pRediction model, which consists of station outflow/inflow temporal embedding components and spatio-temporal attention blocks to push the limit of prediction capability. Extensive experiments corroborate the efficacy of the proposed STAR.
Sijing Duan, Feng Lyu 0001, Ju Ren 0001, Peng Yang 0004, Desheng Zhang 0002, Yaoxue Zhang
IEEE Trans. Intell. Transp. Syst.3
2022 Privacy-Preserving Streaming Truth Discovery in Crowdsourcing With Differential Privacy
abstract
Differential privacy (DP) has gained popularity in truth discovery recently due to its strong privacy guarantee. However, existing DP mechanisms for streaming data publication are not suitable for truth discovery as they fail to consider the different reliabilities of individuals, while the DP-based approaches for truth discovery are not suitable for streaming data because they ignore the correlations between truths over time. Directly applying these existing methods to streaming crowdsourced data would lead to low accuracy of the discovered truth. To solve this problem, in this paper, we propose an edge computing based privacy-preserving truth discovery mechanism, named PrivSTD, for streaming crowdsourced data to realize high accuracy of discovered truth while protecting the privacy of workers. Specifically, edge servers are introduced between the untrusted cloud server and workers to securely calculate the local truths and workers’ reliabilities. A truth-dependent budget recycle mechanism is proposed for each edge server to adaptively determine the perturbed timestamp and allocate the privacy budget according to the changing pattern of local truths. Besides, a reliability-based perturbation mechanism is proposed to reduce the perturbation magnitude on the basis of worker's reliability. We theoretical analyze the data utility and computation cost of PrivSTD, and prove that PrivSTD can satisfy$w$-event ($\epsilon,\delta$)-differential privacy. Extensive experimental results on synthetic and real-world datasets demonstrate that PrivSTD achieves better utility than the state-of-the-art approaches.
Dan Wang 0031, Ju Ren 0001, Zhibo Wang 0001, Xiaoyi Pang, Yaoxue Zhang, Xuemin Shen
IEEE Trans. Mob. Comput.2
2022 Hear Sign Language: A Real-Time End-to-End Sign Language Recognition System
abstract
Sign language recognition (SLR) bridges the communication gap between the hearing-impaired and the ordinary people. However, existing SLR systems either cannot provide continuous recognition or suffer from low recognition accuracy due to the difficulty of sign segmentation and the insufficiency of capturing both finger and arm motions. The latest system, SignSpeaker, has a significant limit in recognizing two-handed signs with onlyonesmartwatch. To address these problems, this paper designs a novel real-time end-to-end SLR system, called DeepSLR, to translate sign language into voices to help people “hear” sign language. Specifically, two armbands embedded with an IMU sensor and multi-channel sEMG sensors are attached on the forearms to capture both coarse-grained arm movements and fine-grained finger motions. We propose an attention-based encoder-decoder model with a multi-channel convolutional neural network (CNN) to realize accurate, scalable, and end-to-end continuous SLR without sign segmentation. We have implemented DeepSLR on a smartphone and evaluated its effectiveness through extensive evaluations. The average word error rate of continuous sentence recognition is 10.8 percent, and it takes less than 1.1s for detecting signals and recognizing a sentence with 4 sign words, validating the recognition efficiency and real-time ability of DeepSLR in real-world scenarios.
Zhibo Wang 0001, Tengda Zhao, Jinxin Ma, Huajie Shao, Qian Wang 0002, Ju Ren 0001
IEEE Trans. Mob. Comput.8
2022 Towards Personalized Privacy-Preserving Truth Discovery Over Crowdsourced Data Streams
abstract
Truth discovery is an effective paradigm which could reveal the truth from crowdsouced data with conflicts, enabling data-driven decision-making systems to make quick and smart decisions. The increasing privacy concern promotes users to perturb or encrypt their private data before outsourcing, which poses significant challenges for truth discovery. Although several privacy-preserving truth discovery mechanisms have been proposed, none of them take personal privacy expectation into consideration. In this work, we propose a novel personalized privacy-preserving truth discovery (PPPTD) framework over crowdsourced data streams to achieve timely and accurate truth discovery while guaranteeing the protection of individual privacy. The key challenges of PPPTD lie in improving the accuracy of truth estimation from the perturbed streaming data with personalized protection level. To address these challenges, we first develop a personalized budget initialization mechanism to quantify each user’s privacy protection requirement, and allocate personalized privacy budgets to users according to their privacy requirements. Then we propose a deviation-aware weighted aggregation method to improve the accuracy of truth discovery from streaming data with varying degrees of perturbation. In order to achieve privacy-utility tradeoff, we further propose an influence-aware adaptive budget adjustment mechanism that adaptively re-allocates privacy budgets to users based on the evolution of their influence in the weighted aggregation. We prove that PPPTD can achieve$\epsilon $-differential privacy over the whole data generated by users and satisfy individual personalized privacy requirements. Extensive experiments on two real-world datasets demonstrate the effectiveness of PPPTD.
Xiaoyi Pang, Zhibo Wang 0001, Defang Liu, John C. S. Lui, Qian Wang 0002, Ju Ren 0001
IEEE/ACM Trans. Netw.6
2022 Improving Federated Learning With Quality-Aware User Incentive and Auto-Weighted Model Aggregation
abstract
Federated learning enables distributed model training over various computing nodes, e.g., mobile devices, where instead of sharing raw user data, computing nodes can solely commit model updates without compromising data privacy. The quality of federated learning relies on the model updates contributed by computing nodes training with their local data. However, with various factors (e.g., training data size, mislabeled data samples, skewed data distributions), the model update qualities of computing nodes can vary dramatically, while inclusively aggregating low-quality model updates can deteriorate the global model quality. To achieve efficient federated learning, in this paper, we propose a novel framework namedFAIR, i.e.,Federated leArning with qualIty awaReness. Particularly,FAIRintegrates three major components: 1) learning quality estimation: we adopt the model aggregation weight (learned in the third component) to reversely quantify the individual learning quality of nodes in a privacy-preserving manner, and leverage the historical learning records to infer the next-round learning quality; 2) quality-aware incentive mechanism: within the recruiting budget, we model a reverse auction problem to stimulate the participation of high-quality and low-cost computing nodes, and the method is proved to be truthful, individually rational, and computationally efficient; and 3) auto-weighted model aggregation: based on the gradient descent method, we devise an auto-weighted model aggregation algorithm to automatically learn the optimal aggregation weights to further enhance the global model quality. Based on real-world datasets and learning tasks, extensive experiments are conducted to demonstrate the efficacy ofFAIR.
Yongheng Deng, Feng Lyu 0001, Ju Ren 0001, Yi-Chao Chen 0001, Peng Yang 0004, Yue-Zhi Zhou, Yaoxue Zhang
IEEE Trans. Parallel Distributed Syst.3
2022 AUCTION: Automated and Quality-Aware Client Selection Framework for Efficient Federated Learning
abstract
The emergency of federated learning (FL) enables distributed data owners to collaboratively build a global model without sharing their raw data, which creates a new business chance for building data market. However, in practical FL scenarios, the hardware conditions and data resources of the participant clients can vary significantly, leading to different positive/negative effects on the FL performance, where the client selection problem becomes crucial. To this end, we proposeAUCTION, anAutomated and qUality-awareClient selecTIONframework for efficient FL, which can evaluate the learning quality of clients and select them automatically with quality-awareness for a given FL task within a limited budget. To designAUCTION, multiple factors such as data size, data quality, and learning budget that can affect the learning performance should be properly balanced. It is nontrivial since their impacts on the FL model are intricate and unquantifiable. Therefore,AUCTIONis designed to encode the client selection policy into a neural network and employ reinforcement learning to automatically learn client selection policies based on the observed client status and feedback rewards quantified by the federated learning performance. In particular, the policy network is built upon an encoder-decoder deep neural network with an attention mechanism, which can adapt to dynamic changes of the number of candidate clients and make sequential client selection actions to reduce the learning space significantly. Extensive experiments are carried out based on real-world datasets and well-known learning models to demonstrate the efficiency, robustness, and scalability ofAUCTION.
Yongheng Deng, Feng Lyu 0001, Ju Ren 0001, Huaqing Wu, Yue-Zhi Zhou, Yaoxue Zhang, Xuemin Shen
IEEE Trans. Parallel Distributed Syst.3
2022 TODG: Distributed Task Offloading With Delay Guarantees for Edge Computing
abstract
Edge computing has been an efficient way to provide prompt and near-data computing services for resource-and-delay sensitive IoT applications via computation offloading. Effective computation offloading strategies need to comprehensively cope with several major issues, including 1) the allocation of dynamic communication and computational resources, 2) delay constraints of heterogeneous tasks, and 3) requirements for computationally inexpensive and distributed algorithms. However, most of the existing works mainly focus on part of these issues, which would not suffice to achieve expected performance in complex and practical scenarios. To tackle this challenge, in this paper, we systematically study a distributed computation offloading problem with delay constraints, where heterogeneous computational tasks require continually offloading to a set of edge servers via a limiting number of stochastic communication channels. The task offloading problem is formulated as a delay-constrained long-term stochastic optimization problem under unknown prior statistical knowledge. To solve this problem, we first provide a technical path to transform and decompose it into several slot-level sub-problems. Then, we devise a distributed online algorithm, namely TODG, to efficiently allocate resources and schedule offloading tasks. Further, we present a comprehensive analysis for TODG in terms of the optimality gap, the worst-case delay, and the impact of system parameters. Extensive simulation results demonstrate the effectiveness and efficiency of TODG.
Sheng Yue 0001, Ju Ren 0001, Nan Qiao 0008, Yongmin Zhang, Hongbo Jiang 0001, Yaoxue Zhang, Yuanyuan Yang 0001
IEEE Trans. Parallel Distributed Syst.2
2021 The Invisible Shadow: How Security Cameras Leak Private Activities
abstract
This paper presents a new privacy threat, the Invisible Infrared Shadow Attack (IRSA), which leverages the inconspicuous infrared (IR) light emitted by indoor security cameras, to reveal in-home human activities behind opaque curtains. The key observation is that the in-home IR light source can project invisible shadows on the window curtains, which can be captured by an attacker outside using an IR-capable camera. The major challenge for IRSA lies in the shadow deformation caused by a variety of environmental factors involving the IR source position and curtain shape, which distorts the body contour. A two-stage attack scheme is proposed to circumvent the challenge. Specifically, a DeShaNet model performs accurate shadow keypoint detection through multi-dimension feature fusion. Then a scene constructor maps the 2D shadow keypoints to 3D human skeletons by iteratively reproducing the on-site shadow projection process in a virtual Unity 3D environment. Through comprehensive evaluation, we show that the proposed attack scheme can be successfully launched to recover 3D skeleton of the victims, even under severe shadow deformation. Finally, we propose potential defense mechanisms against the IRSA.
Xinyu Zhang 0003, Ju Ren 0001, Yaoxue Zhang
CCS3
2021 SHARE: Shaping Data Distribution at Edge for Communication-Efficient Hierarchical Federated Learning
abstract
Federated learning (FL) can enable distributed model training over mobile nodes without sharing privacy-sensitive raw data. However, to achieve efficient FL, one significant challenge is the prohibitive communication overhead to commit model updates since frequent cloud model aggregations are usually required to reach a target accuracy, especially when the data distributions at mobile nodes are imbalanced. With pilot experiments, it is verified that frequent cloud model aggregations can be avoided without performance degradation if model aggregations can be conducted at edge. To this end, we shed light on the hierarchical federated learning (HFL) framework, where a subset of distributed nodes are selected as edge aggregators to conduct edge aggregations. Particularly, under the HFL framework, we formulate a communication cost minimization (CCM) problem to minimize the communication cost raised by edge/cloud aggregations with making decisions on edge aggregator selection and distributed node association. Inspired by the insight that the potential of HFL lies in the data distribution at edge aggregators, we propose SHARE, i.e., SHaping dAta distRibution at Edge, to transform and solve the CCM problem. In SHARE, we divide the original problem into two sub-problems to minimize the per-round communication cost and mean Kullback-Leibler divergence of edge aggregator data, and devise two light-weight algorithms to solve them, respectively. Extensive experiments under various settings are carried out to corroborate the efficacy of SHARE.
Yongheng Deng, Feng Lyu 0001, Ju Ren 0001, Yongmin Zhang, Yue-Zhi Zhou, Yaoxue Zhang, Yuanyuan Yang 0001
ICDCS3
2021 FAIR: Quality-Aware Federated Learning with Precise User Incentive and Model Aggregation
abstract
Federated learning enables distributed learning in a privacy-protected manner, but two challenging reasons can affect learning performance significantly. First, mobile users are not willing to participate in learning due to computation and energy consumption. Second, with various factors (e.g., training data size/quality), the model update quality of mobile devices can vary dramatically, inclusively aggregating low-quality model updates can deteriorate the global model quality. In this paper, we propose a novel system named FAIR, i.e., Federated leArning with qualIty awaReness. FAIR integrates three major components: 1) learning quality estimation: we leverage historical learning records to estimate the user learning quality, where the record freshness is considered and the exponential forgetting function is utilized for weight assignment; 2) quality-aware incentive mechanism: within the recruiting budget, we model a reverse auction problem to encourage the participation of high-quality learning users, and the method is proved to be truthful, individually rational, and computationally efficient; and 3) model aggregation: we devise an aggregation algorithm that integrates the model quality into aggregation and filters out non-ideal model updates, to further optimize the global learning model. Based on real-world datasets and practical learning tasks, extensive experiments are carried out to demonstrate the efficacy of FAIR.
Yongheng Deng, Feng Lyu 0001, Ju Ren 0001, Yi-Chao Chen 0001, Peng Yang 0004, Yue-Zhi Zhou, Yaoxue Zhang
INFOCOM3
2021 Inexact-ADMM Based Federated Meta-Learning for Fast and Continual Edge Learning
abstract
In order to meet the requirements for performance, safety, and latency in many IoT applications, intelligent decisions must be made right here right now at the network edge. However, the constrained resources and limited local data amount pose significant challenges to the development of edge AI. To overcome these challenges, we explore continual edge learning capable of leveraging the knowledge transfer from previous tasks. Aiming to achieve fast and continual edge learning, we propose a platform-aided federated meta-learning architecture where edge nodes collaboratively learn a meta-model, aided by the knowledge transfer from prior tasks. The edge learning problem is cast as a regularized optimization problem, where the valuable knowledge learned from previous tasks is extracted as regularization. Then, we devise an ADMM based federated meta-learning algorithm, namely ADMM-FedMeta, where ADMM offers a natural mechanism to decompose the original problem into many subproblems which can be solved in parallel across edge nodes and the platform. Further, a variant of inexact-ADMM method is employed where the subproblems are 'solved' via linear approximation as well as Hessian estimation to reduce the computational cost per round to O(n). We provide a comprehensive analysis of ADMM-FedMeta, in terms of the convergence properties, the rapid adaptation performance, and the forgetting effect of prior knowledge transfer, for the general non-convex case. Extensive experimental studies demonstrate the effectiveness and efficiency of ADMM-FedMeta, and showcase that it substantially outperforms the existing baselines.
Sheng Yue 0001, Ju Ren 0001, Jiang Xin, Sen Lin 0001, Junshan Zhang
MobiHoc2
2021 FLAG: Flexible, Accurate, and Long-Time User Load Prediction in Large-Scale WiFi System Using Deep RNN
abstract
In this article, we proposeFLAGfor flexible, accurate, and long-time user load prediction in a large-scale WiFi system.FLAGenables prediction customization in both time granularity and prediction length. Under an operating WiFi system with more than 7000 APs, a reference implementation ofFLAGis developed, which consists of three major components. Fordata acquisition, we process 25 074 733 association records contributed by 55 809 users, to extract the ground truth of AP-level user load. Forfeature extraction, we perform a comprehensive data analytics to mine vital features to label each AP, which are extracted and classified into three categories, i.e., individual features, spatial features, and temporal features. For themodel design, we design a deep recurrent neural network (RNN) model, which contains two separate RNNs, i.e., the encoder RNN and decoder RNN. Particularly, the sequential feature vectors are injected into the encoder RNN to learn the “semantic” information, based on which the decoder RNN conducts sequential AP-level predictions. As the semantic vector is injected for each time step prediction, it can effectively reduce the accumulated prediction errors, which enable long period of time predictions. Real data set-based experiments corroborate the efficacy ofFLAG.
Wenxiong Chen, Feng Lyu 0001, Fan Wu 0014, Peng Yang 0004, Ju Ren 0001
IEEE Internet Things J.5
2021 A Survey of Millimeter-Wave Communication: Physical-Layer Technology Specifications and Enabling Transmission Technologies
abstract
Millimeter-wave (mmWave) frequency bands, which offer abundant underutilized spectral resources, have been explored and exploited in the past several years to meet the requirements of emerging wireless services highlighted by high data rates, ultrareliability, and ultralow delivery latency. Yet, the unique characteristics of mmWave, e.g., continuous wide bandwidth, large path, and penetration losses, along with hardware constraints, call for innovative technologies for mmWave communication. Recently, an extensive amount of work on mmWave communication has been carried out by researchers and practitioners from both academia and industry, and various technologies have been developed for mmWave communication systems to fulfill the full potential of mmWave frequency bands. In this article, we present a comprehensive survey of the standardization of mmWave communication, the latest progress and outcomes of the research on mmWave communication technologies, and the emerging applications of mmWave communication. In particular, we provide a timely and in-depth summary of the state-of-the-art technology specifications of mmWave communication with an emphasis on the physical (PHY) layer. Then, we elaborate on a number of well-established or promising antenna architectures in mmWave communication systems and investigate the enabling PHY layer transmission technologies. Finally, we show some existing and emerging applications of mmWave communication and discuss the potential open research issues.
Shiwen He, Yan Zhang 0073, Jiaheng Wang 0001, Jian Zhang 0048, Ju Ren 0001, Yaoxue Zhang, Weihua Zhuang, Xuemin Shen
Proc. IEEE5
2021 DeepNav: A scalable and plug-and-play indoor navigation system based on visual CNN
Ju Ren 0001, Yaoxue Zhang
Peer-to-Peer Netw. Appl.2
2021 Online Multi-Workflow Scheduling under Uncertain Task Execution Time in IaaS Clouds
abstract
Cloud has become an important platform for executing numerous deadline-constrained scientific applications generally represented by workflow models. It provides scientists a simple and cost-efficient method of running workflows on their rental Virtual Machines (VMs) anytime and anywhere. Since pay-as-you-go is a dominating pricing solution in clouds, extensive research efforts have been devoted to minimizing the monetary cost of executing workflows by designing tailored VM allocation mechanisms. However, most of them assume that the task execution time in clouds is static and can be estimated in advance, which is impractical in real scenarios due to performance fluctuation of VMs. In this paper, we propose an onliNe multi-workflOwSchedulingFramework, named NOSF, to schedule deadline-constrained workflows with random arrivals and uncertain task execution time. In NOSF, workflow scheduling process consists of three phases, including workflow preprocessing, VM allocation and feedback process. Built upon the new framework, a deadline-aware heuristic algorithm is then developed to elastically provision suitable VMs for workflow execution, with the objective of minimizing the rental cost and improving resource utilization. Simulation results demonstrate that the proposed algorithm significantly outperforms two state-of-the-art algorithms in terms of reducing VM rental costs and deadline violation probability, as well as improving the resource utilization efficiency.
Jiagang Liu, Ju Ren 0001, Pude Zhou, Yaoxue Zhang, Geyong Min, Noushin Najjari
IEEE Trans. Cloud Comput.2
2021 An Incentive-Aware Job Offloading Control Framework for Multi-Access Edge Computing
abstract
This paper considers a scenario in which an access point (AP) is equipped with a server of finite computing power, and serves multiple resource-hungry users by charging users a price. This price helps to regulate users' behavior in offloading jobs to the AP. However, existing works on pricing are based on abstract concave utility functions, giving no dependence on physical layer parameters. To that end, we first introduce a novel utility function, which measures the cost reduction by offloading as compared with executing jobs locally. Based on this utility function we then formulate two offloading games, with one maximizing individuals interest and the other maximizing the overall systems interest. We analyze the structural property of the games and admit in closed-form the Nash Equilibrium and the Social Equilibrium for the homogeneous user case, respectively. The proposed expressions are functions of user parameters such as the weights of time and energy, the distance from the AP, thus constituting an advancement over prior economic works that have considered only abstract functions. Finally, we propose an optimal price-based scheme, with which we prove that the interactive decision-making process with self-interested users converges to a Nash Equilibrium point equal to the Social Equilibrium point.
Lingxiang Li, Tony Q. S. Quek, Ju Ren 0001, Howard H. Yang, Zhi Chen 0002, Yaoxue Zhang
IEEE Trans. Mob. Comput.3
2021 LeaD: Large-Scale Edge Cache Deployment Based on Spatio-Temporal WiFi Traffic Statistics
abstract
Widespread and large-scale WiFi systems have been deployed in many corporate locations, while the backhual capacity becomes the bottleneck in providing high-rate data services to a tremendous number of WiFi users. Mobile edge caching is a promising solution to relieve backhaul pressure and deliver quality services by proactively pushing contents to access points (APs). However, how to deploy cache in large-scale WiFi system is not well studied yet quite challenging since numerous APs can have heterogeneous traffic characteristics, and future traffic conditions are unknown ahead. In this paper, given the cache storage budget, we explore the cache deployment in a large-scale WiFi system, which contains 8,000 APs and serves more than 40,000 active users, to maximize the long-term caching gain. Specifically, we first collect two-month user association records and conduct intensive spatio-temporal analytics on WiFi traffic consumption, gaining two major observations. First, per AP traffic consumption varies in a rather wide range and the proportion of AP distributes evenly within the range, indicating that the cache size should be heterogeneously allocated in accordance to the underlying traffic demands. Second, compared to a single AP, the traffic consumption of a group of APs (clustered by physical locations) is more stable, which means that the short-term traffic statistics can be used to infer the future long-term traffic conditions. We then propose our cache deployment strategy, named LeaD (i.e., Large-scale WiFi Edge cAche Deployment), in which we first cluster large-scale APs into well-sized edge nodes, then conduct the stationary testing on edge level traffic consumption and sample sufficient traffic statistics in order to precisely characterize long-term traffic conditions, and finally devise the TEG (Traffic-wEighted Greedy) algorithm to solve the long-term caching gain maximization problem. Extensive trace-driven experiments are carried out, and the results demonstrate that LeaD is able to achieve the near-optimal caching performance and can outperform other benchmark strategies significantly.
Feng Lyu 0001, Ju Ren 0001, Nan Cheng 0001, Peng Yang 0004, Minglu Li 0001, Yaoxue Zhang, Xuemin Shen
IEEE Trans. Mob. Comput.2
2021 SocialRecruiter: Dynamic Incentive Mechanism for Mobile Crowdsourcing Worker Recruitment With Social Networks
abstract
Worker recruitment is an important problem in mobile crowdsourcing (MCS), which aims to find sufficient and suitable participants to perform tasks. However, existing worker recruitment approaches mainly focus on how to select the most suitable workers for tasks from a large worker pool, while the recruitment problem under insufficient workers (e.g., a new MCS system) has not been well addressed. In this paper, we focus on the insufficient participation problem of MCS systems with limited number of workers, and propose to leverage social network to recruit workers for task completion as well as expanding the worker pool. To this end, we propose a dynamic incentive mechanism, called SocialRecruiter, to encourage workers on the MCS platform to propagate tasks through social networks, so that inviting friends to join in the MCS platform to further propagate and complete tasks. Motivated by the SIR epidemic model, we propose a novel task-specific epidemic model to characterize the status change of users for task propagation and completion through social networks. In order to encourage task completion and propagation, the propagating reward and completing reward are provided according to workers’ actions. In particular, in order to maximize the task completion within the financial budget, the propagating and completing rewards are dynamically updated at each cycle according to real-time worker recruitment progress. The extensive experimental results on two real-world datasets demonstrate that SocialRecruiter outperforms the state-of-the-art approaches in terms of worker recruitment and task completion.
Zhibo Wang 0001, Yuting Huang 0005, Xinkai Wang 0006, Ju Ren 0001, Qian Wang 0002
IEEE Trans. Mob. Comput.4
2021 Towards Personalized Task-Oriented Worker Recruitment in Mobile Crowdsensing
abstract
Worker recruitment in mobile crowdsensing systems aims to recruit the most suitable users to perform tasks with high quality and in real-time. Many worker recruitment or task matching mechanisms have been proposed, especially for crowdsourcing platforms, where content information of tasks from the implicit feedback of workers' attendance is extensively exploited to help workers find preferred tasks efficiently. Different from traditional crowdsourcing systems, tasks in mobile crowdsensing systems are usually time-sensitive and location-dependent which also play a crucial role in worker recruitment. However, these context information have not been effectively explored for user recruitment in mobile crowdsensing systems. In this paper, we propose a novel personalized task-oriented worker recruitment mechanism for mobile crowdsensing systems based on a careful characterization of workers' preference. In particular, we fully exploit the content information (e.g., task category, task description) together with the context information (e.g., task time, task location) from the implicit feedback of workers' attendance to accurately model workers' preference on tasks. Moreover, we regard the task-worker fitness prediction as a binary classification problem and utilize the Logit model to integrate the heterogeneous factors into a single framework to predict the matching probability of each task-worker pair. Finally, the workers with the highest matching probability are recruited proactively for each new task. Extensive experiments on real-world datasets demonstrate that the proposed mechanism achieves better performance than the benchmarks.
Zhibo Wang 0001, Jing Zhao 0011, Jiahui Hu 0001, Tianqing Zhu, Qian Wang 0002, Ju Ren 0001, Chao Li 0027
IEEE Trans. Mob. Comput.6
2021 NDN-MMRA: Multi-Stage Multicast Rate Adaptation in Named Data Networking WLAN
abstract
Named Data Networking (NDN) is considered as a prominent architecture towards future Wireless Local Area Networks (WLAN), and multicast plays an important role in data delivery such as media streaming, multipoint videoconferencing, etc. However, to achieve high-efficiency multicast in NDN WLAN is challenging for two significant reasons. First, without feedback mechanism in IEEE 802.11 standards, to guarantee reliability, the current multicast scheme transmits the multicast data with the basic rate (e.g., 1 Mbps for IEEE 802.11b), which inevitably increases the transmission delay for high-speed consumers. Second, as a NDN multicast group is constituted by consumers who are requesting the same content, multicast groups are easy to form and evolve rapidly, where a data rate adaptation scheme is requisite to accommodate differential multicast groups. In this paper, we propose a multi-stage multicast rate adaptation scheme for NDN WLAN, namedNDN-MMRA, to minimize the total transmission time with reliability guarantee for multicast group members. InNDN-MMRA, by checking the Pending Interest Table (PIT) status information, the number of consumers in each multicast group as well as their receiving capabilities are known ahead; with the available data rates in a specific 802.11 standard,NDN-MMRAdetermines: 1) how many transmission stages are required; and 2) in each stage, which data rate should be adopted. The merit is that with multi-stage transmissions, the data rate can be adapted in descending order to accommodate high-speed consumers with delay minimized, and low-speed consumers with reliability guaranteed. We implementNDN-MMRAin NS-3 by adopting the ndnSIM module, and conduct extensive experiments to demonstrate its efficacy under different IEEE 802.11 standards and various underlying WLAN topologies.
Fan Wu 0014, Wang Yang 0002, Ju Ren 0001, Feng Lyu 0001, Peng Yang 0004, Yaoxue Zhang, Xuemin Shen
IEEE Trans. Multim.3
2021 Multi-Path Selection and Congestion Control for NDN: An Online Learning Approach
abstract
In Named Data Networking (NDN) architecture, data can be obtained from multiple content sources (i.e., producers or caching nodes) with multiple paths, making the traditional end-to-end (i.e., TCP/IP) congestion control scheme invalid. In addition, the NDN multi-path discovery and management are still an open issue as the dynamic network topology changes. In this article, we propose a multi-path congestion control mechanism, named MPCC, which includes two major components, i.e., multi-path discovery and multi-path congestion control. Particularly, for multi-path discovery, we first devise apath tagto uniquely mark each sub-path in the forwarding process, and then propose a tag-aware forwarding strategy to discover and manage sub-paths. For multi-path selection and congestion control, we first integrate the metrics of packet loss, bandwidth, round trip time, and path centrality, for path assessment, based on which, we then leverage the Upper Confidence Bound (UCB) algorithm to select sub-paths in order to maximize the network throughput. In addition, for selected sub-paths, we have devised a sub-path window adaptation algorithm to avoid multi-path congestions. At last, we implement MPCC in ndnSIM and conduct extensive experiments for performance evaluation. Our results demonstrate that MPCC can discover all sub-paths in real-time for the multi-path scenario, and can effectively avoid multi-path congestions with improving throughput and reducing transmission time.
Fan Wu 0014, Wang Yang 0002, Muhua Sun, Ju Ren 0001, Feng Lyu 0001
IEEE Trans. Netw. Serv. Manag.4
2020 Attention-over-Attention Field-Aware Factorization Machine
abstract
Factorization Machine (FM) has been a popular approach in supervised predictive tasks, such as click-through rate prediction and recommender systems, due to its great performance and efficiency. Recently, several variants of FM have been proposed to improve its performance. However, most of the state-of-the-art prediction algorithms neglected the field information of features, and they also failed to discriminate the importance of feature interactions due to the problem of redundant features. In this paper, we present a novel algorithm called Attention-over-Attention Field-aware Factorization Machine (AoAFFM) for better capturing the characteristics of feature interactions. Specifically, we propose the field-aware embedding layer to exploit the field information of features, and combine it with the attention-over-attention mechanism to learn both feature-level and interaction-level attention to estimate the weight of feature interactions. Experimental results show that the proposed AoAFFM improves FM and FFM with large margin, and outperforms state-of-the-art algorithms on three public benchmark datasets.
Zhibo Wang 0001, Jinxin Ma, Qian Wang 0002, Ju Ren 0001, Peng Sun 0003
AAAI5
2020 SoSA: Socializing Static APs for Edge Resource Pooling in Large-Scale WiFi System
abstract
Large-scale WiFi system is gaining an increasing momentum rapidly in most corporate places. Enabling edge functions on the system is imperative to support unprecedented edge applications. However, building edge functionalities at each AP may incur frequent service migrations, low resource utilization, and inflexible resource provisioning. It is thus prospective to federate suitable APs to create a resource-pooled edge system such that users in association with federated APs can share the pooled resource. In this paper, we propose a novel architecture, named SoSA, to Socialize Static APs via user association transition activities for edge resource pooling. A reference implementation of SoSA is developed under an operating large-scale WiFi system in a campus area of 3.0925 km2. The novelty and contribution of SoSA lie in its three-layer design. In transition data feeding layer, we collect and process 25,074,733 association records of 55,809 users from 7,404 APs in a real WiFi system. In sociality construction and characterization layer, we construct an AP contact graph based on user transition statistics, under which we empirically study the sociality of APs and explore their evolving patterns. In edge resource pooling layer, by harnessing the AP sociality, we are able to customize resource pooling strategy to improve service provisioning performance. With adopting SoSA, we systematically investigate the performance of AP federation strategy in reducing service migration when users frequently transit among APs. Extensive data-driven experiments corroborate the efficacy of SoSA.
Feng Lyu 0001, Ju Ren 0001, Peng Yang 0004, Nan Cheng 0001, Yaoxue Zhang, Xuemin Shen
INFOCOM2
2020 Towards Pattern-aware Privacy-preserving Real-time Data Collection
abstract
Although time-series data collected from users can be utilized to provide services for various applications, they could reveal sensitive information about users. Recently, local differential privacy (LDP) has emerged as the state-of-art approach to protect data privacy by perturbing data locally before outsourcing. However, existing works based on LDP perturb each data point separately without considering the correlations between consecutive data points in time-series. Thus, the important patterns of each time-series might be distorted by existing LDP-based approaches, leading to severe degradation of data utility. In this paper, we focus on real-time data collection under a honest-but-curious server, and propose a novel pattern-aware privacy-preserving approach, called PatternLDP, to protect data privacy while the pattern of time-series can still be preserved. To this end, instead of providing the same level of privacy protection at each data point, each user only samples remarkable points in time-series and adaptively perturbs them according to their impacts on local patterns. In particular, we propose a pattern-aware sampling method based on Piecewise Linear Approximation (PLA) to determine whether to sample and perturb current data point. To reduce the utility loss caused by pattern change after perturbation, we propose an importance-aware randomization mechanism to adaptively perturb sampled data locally while achieving better trade-off between privacy and utility. A novel metric-based w-event privacy is introduced to measure the privacy protection degree for pattern-rich time-series. We prove that PatternLDP can provide the above privacy guarantee, and extensive experiments on real-world datasets demonstrate that PatternLDP outperforms existing mechanisms and can effectively preserve the important patterns.
Zhibo Wang 0001, Xiaoyi Pang, Ju Ren 0001, Zhe Liu 0001, Yongle Chen
INFOCOM4
2020 MobiPose: real-time multi-person pose estimation on mobile devices
abstract
Human pose estimation is a key technique for many vision-based mobile applications. Yet existing multi-person pose-estimation methods fail to achieve a satisfactory user experience on commodity mobile devices such as smartphones, due to their long model-inference latency. In this paper, we propose MobiPose, a system designed to enable real-time multi-person pose estimation on mobile devices through three novel techniques. First, MobiPose takes a motion-vector-based approach to fast locate the human proposals across consecutive frames by fine-grained tracking of joints of human body, rather than running the expensive human-detection model for every frame. Second, MobiPose designs a mobile-friendly model that uses lightweight multi-stage feature extractions to significantly reduce the latency of pose estimation without compromising the model accuracy. Third, MobiPose leverages the heterogeneous computing resources of both CPU and GPU to execute the pose estimation model for multiple persons in parallel, which further reduces the total latency. We have implemented the MobiPose system on off-the-shelf commercial smartphones and conducted comprehensive experiments to evaluate the effectiveness of the proposed techniques. Evaluation results show that MobiPose achieves over 20 frames per second pose estimation with 3 persons per frame, and significantly outperforms the state-of-the-art baseline, with a speedup of up to 4.5X and 2.8X in latency on CPU and GPU, respectively, and an improvement of 5.1% in pose-estimation model accuracy. Furthermore, MobiPose achieves up to 62.5% and 37.9% energy-per-frame saving on average in comparison with the baseline on mobile CPU and GPU, respectively.
Xiaohui Xu, Fucheng Jia, Yunxin Liu 0001, Xuanzhe Liu, Ju Ren 0001, Yaoxue Zhang
SenSys7
2020 Dynamic Spectrum Slicing and Optimization in SAG Integrated Vehicular Networks
abstract
In this paper, we propose an online control frame-work to dynamically slice the network resource for isolated service provisioning in Space-Air-Ground integrated Vehicular Network (SAGVN). In particular, at a given time slot, the system makes online decisions on the request admission and scheduling, UAV dispatching, and resource slicing for different services. To characterize the impact of those parameters, we construct a time-averaged queue stability criteria by taking queue backlogs of all services into consideration, and formulate a system revenue function which incorporates the time-averaged system throughput and UAV dispatching cost. The objective is to maximize the system revenue while stabilizing the time-averaged queue, which can be achieved via the Lyapunov optimization theory. By bounding the drift-plus-penalty, the problem then can be decoupled into four independent subproblems, which are readily solved. The merits of our control framework are three-fold: 1) the system can admit and process as many requests as possible; 2) the time-averaged UAV dispatching cost is minimized; and 3) service queues can be stabilized over time. Extensive simulations are carried out, and the results demonstrate that the control framework can effectively achieve the system revenue maximization and queueing stabilization. Moreover, it can balance the trade-off among system throughput, UAV dispatching cost, and queueing states via parameter tuning.
Feng Lyu 0001, Peng Yang 0004, Huaqing Wu, Conghao Zhou, Ju Ren 0001, Yaoxue Zhang, Xuemin Shen
VTC Fall5
2020 A trusted feature aggregator federated learning for distributed malicious attack detection
Xinhong Hei 0001, Xinyue Yin, Yichuan Wang 0003, Ju Ren 0001, Lei Zhu 0011
Comput. Secur.4
2020 An efficient privacy-enhanced attribute-based access control mechanism
abstract
Summary Owing to the rapid progress of network researching, attribute‐based access control (ABAC) has attracted more and more attention due to its appreciable expressiveness, flexibility, and scalability. Unfortunately, collecting user attributes is necessary to complete the standard ABAC decision process, which increases the risk of privacy disclosure. This problem increases public doubts about ABAC and hinders its popularization. In this paper, a privacy‐protected and efficient attribute‐based access control (EPABAC) scheme is proposed to prevent the privacy leakage of access subject in the decision‐making process of ABAC by introducing a novel hash‐based binary search tree. The analyses and experimental evaluations show that the EPABAC achieves user privacy protection in the decision‐making process with acceptable additional computing overhead.
Yang Xu 0013, Quanrun Zeng, Guojun Wang 0001, Cheng Zhang 0035, Ju Ren 0001, Yaoxue Zhang
Concurr. Comput. Pract. Exp.5
2020 Drive2friends: Inferring Social Relationships From Individual Vehicle Mobility Data
abstract
The number of vehicles has increased year by year, especially individual vehicles. In addition to meeting basic transportation needs, vehicles are expected to serve varied location-based services and applications for humans. However, it can constitute severe risks for privacy. In this article, we concentrate on one of the most sensitive information, namely, social relationships, that can be inferred from the vehicle mobility data. We propose a social relationship inference model, which provides a new perspective for privacy preservation in human mobility data. In particular, we extract discriminative features from both the spatial and temporal dimensions. Then, the heterogeneous features are being merged with a fusion model to improve the performance of inference. Extensive experiments on the real-world data set validate the effectiveness of the extracted features in estimating social connections and demonstrate that our method significantly outperforms the baseline models.
Jie Li 0058, Fanzi Zeng, Zhu Xiao, Hongbo Jiang 0001, Zhirun Zheng, Wenping Liu 0001, Ju Ren 0001
IEEE Internet Things J.7
2020 TrajData: On Vehicle Trajectory Collection With Commodity Plug-and-Play OBU Devices
abstract
For years, vehicle trajectory data have increasingly been important for a wide range of applications, from driver behavior investigation/classification, travel time/distance estimation, and routing in vehicular networks, to vehicle energy/emission evaluation. This article presents TrajData, the first systematic solution to reliable vehicle trajectory data collection, with only reliance on commercial-off-the-shelf (COTS) onboard unit (OBU) devices that utilize lightweight GPS modules and low-cost onboard diagnostics (OBD) readers. In the practical use of trajectory collection, GPS outages inevitably occur in urban environments thereby leading to large trajectory errors as well as missing vehicle location data. To resolve this, we propose a novel data-fusion-enabled deep learning approach with the purpose of achieving reliable vehicle trajectory collection in various urban road conditions. Specifically, we leverage motion information retrieved from OBD readers in TrajData to help reconstruct the trajectory data during GPS outages. By investigating the changes of direction angle from the OBD readings, we can identify different types of road sections. Furthermore, we integrate the neural arithmetic logic units (NALUs) into our trajectory reconstruction model to tame the challenges when GPS outages take place in various road sections. Experimental results from realistic data have demonstrated the effectiveness and reliability of the proposed method. In the road test, TrajData achieves an average position error below 15-m around a 60-s GPS outage, even in complex road sections, i.e., continuous turns and driving with accelerations/decelerations resulting in frequent changes of direction and speed.
Zhu Xiao, Fancheng Li, Ronghui Wu, Hongbo Jiang 0001, Yupeng Hu 0004, Ju Ren 0001, Chenglin Cai, Arun Iyengar
IEEE Internet Things J.6
2020 Privacy-preserving task recommendation with win-win incentives for mobile crowdsourcing
Wenjuan Tang, Kuan Zhang 0001, Ju Ren 0001, Yaoxue Zhang, Xuemin Shen
Inf. Sci.3
2020 Analyzing User-Level Privacy Attack Against Federated Learning
abstract
Federated learning has emerged as an advanced privacy-preserving learning technique for mobile edge computing, where the model is trained in a decentralized manner by the clients, preventing the server from directly accessing those private data from the clients. This learning mechanism significantly challenges the attack from the server side. Although the state-of-the-art attacking techniques that incorporated the advance of Generative adversarial networks (GANs) could construct class representatives of the global data distribution among all clients, it is still challenging to distinguishably attack a specific client (i.e., user-level privacy leakage), which is a stronger privacy threat to precisely recover the private data from a specific client. To analyze the privacy leakage of federated learning, this paper gives the first attempt to explore user-level privacy leakage by the attack from a malicious server. We propose a framework incorporating GAN with a multi-task discriminator, called multi-task GAN - Auxiliary Identification (mGAN-AI), which simultaneously discriminates category, reality, and client identity of input samples. The novel discrimination on client identity enables the generator to recover user specified private data. Unlike existing works interfering the federated learning process, the proposed method works “invisibly” on the server side. Furthermore, considering the anonymization strategy for mitigating mGAN-AI, we propose a beforehand linkability attack which re-identifies the anonymized updates by associating the client representatives. A novel siamese network fusing the identification and verification models is developed for measuring the similarity of representatives. The experimental results demonstrate the effectiveness of the proposed approaches and the superior to the state-of-the-art.
Mengkai Song, Zhibo Wang 0001, Yang Song 0013, Qian Wang 0002, Ju Ren 0001, Hairong Qi 0001
IEEE J. Sel. Areas Commun.6
2020 SDN/NFV-Empowered Future IoV With Enhanced Communication, Computing, and Caching
abstract
Internet-of-Vehicles (IoV) connects vehicles, sensors, pedestrians, mobile devices, and the Internet with advanced communication and networking technologies, which can enhance road safety, improve road traffic management, and support immerse user experience. However, the increasing number of vehicles and other IoV devices, high vehicle mobility, and diverse service requirements render the operation and management of IoV intractable. Software-defined networking (SDN) and network function virtualization (NFV) technologies offer potential solutions to achieve flexible and automated network management, global network optimization, and efficient network resource orchestration with cost-effectiveness and are envisioned as a key enabler to future IoV. In this article, we provide an overview of SDN/NFV-enabled IoV, in which SDN/NFV technologies are leveraged to enhance the performance of IoV and enable diverse IoV scenarios and applications. In particular, the IoV and SDN/NFV technologies are first introduced. Then, the state-of-the-art research works are surveyed comprehensively, which is categorized into topics according to the role that the SDN/NFV technologies play in IoV, i.e., enhancing the performance of data communication, computing, and caching, respectively. Some open research issues are discussed for future directions.
Weihua Zhuang, Qiang Ye 0002, Feng Lyu 0001, Nan Cheng 0001, Ju Ren 0001
Proc. IEEE5
2020 Guest Editorial: Special Section on End-Edge-Cloud Orchestrated Algorithms, Systems and Applications
abstract
This Special Section aims at providing a platform for sharing the state-of-the-art research and development on end-edge-cloud orchestration and publishing original research and peer-reviewed articles targeted to all readers of IEEE Transactions on Industrial Informatics. The content of the special issue focus on several topics that are recently concerned in the community, including the architectures and implementations, communication and networking protocols, computation offloading strategies, advanced machine learning and data analytical methods, performance modeling and optimization, and other enabling technologies for end-edge-cloud orchestrated systems and its industrial applications.
Hongbo Jiang 0001, Ju Ren 0001, John C. S. Lui, Schahram Dustdar
IEEE Trans. Ind. Informatics2
2020 Near-Optimal and Truthful Online Auction for Computation Offloading in Green Edge-Computing Systems
abstract
Utilizing the intelligence at the network edge, edge computing paradigm emerges to provide time-sensitive computing services for Internet of Things. In this paper, we investigate sustainable computation offloading in an edge-computing system that consists of energy harvesting-enabled mobile devices (MDs) and a dispatcher. The dispatcher collects computation tasks generated by IoT devices with limited computation power, and offloads them to resourceful MDs in exchange for rewards. We propose an online Rewards-optimal Auction (RoA) to optimize the long-term sum-of-rewards for processing offloaded tasks, meanwhile adapting to the highly dynamic energy harvesting (EH) process and computation task arrivals. RoA is designed based on Lyapunov optimization and Vickrey-Clarke-Groves auction, the operation of which does not require a prior knowledge of the energy harvesting, task arrivals, or wireless channel statistics. Our analytical results confirm the optimality of tasks assignment. Furthermore, simulation results validate the analytical analysis, and verify the efficacy of the proposed RoA.
Long Tan, Ju Ren 0001, Mohamad Khattar Awad, Shan Zhang 0001, Yaoxue Zhang, Peng-Jun Wan
IEEE Trans. Mob. Comput.3
2020 Efficient Computing Resource Sharing for Mobile Edge-Cloud Computing Networks
abstract
Both the edge and the cloud can provide computing services for mobile devices to enhance their performance. The edge can reduce the conveying delay by providing local computing services while the cloud can support enormous computing requirements. Their cooperation can improve the utilization of computing resources and ensure the QoS, and thus is critical to edge-cloud computing business models. This paper proposes an efficient framework for mobile edge-cloud computing networks, which enables the edge and the cloud to share their computing resources in the form of wholesale and buyback. To optimize the computing resource sharing process, we formulate the computing resource management problems for the edge servers to manage their wholesale and buyback scheme and the cloud to determine the wholesale price and its local computing resources. Then, we solve these problems from two perspectives: i) social welfare maximization and ii) profit maximization for the edge and the cloud. For i), we have proved the concavity of the social welfare and proposed an optimal cloud computing resource management to maximize the social welfare. For ii), since it is difficult to directly prove the convexity of the primal problem, we first proved the concavity of the wholesaled computing resources with respect to the wholesale price and designed an optimal pricing and cloud computing resource management to maximize their profits. Numerical evaluations show that the total profit can be maximized by social welfare maximization while the respective profits can be maximized by the optimal pricing and cloud computing resource management.
Yongmin Zhang, Xiaolong Lan, Ju Ren 0001, Lin Cai 0001
IEEE/ACM Trans. Netw.3
2020 Blockchain Empowered Arbitrable Data Auditing Scheme for Network Storage as a Service
abstract
The maturity of network storage technology drives users to outsource local data to remote servers. Since these servers are not reliable enough for keeping users' data, remote data auditing mechanisms are studied for mitigating the threat to data integrity. However, many traditional schemes achieve verifiable data integrity for users only without resolutions to data possession disputes, while others depend on centralized third-party auditors (TPAs) for credible arbitrations. Recently, the emergence of blockchain technology promotes inspiring countermeasures. In this article, we propose a decentralized arbitrable remote data auditing scheme for network storage service based on blockchain techniques. We use a smart contract to notarize integrity metadata of outsourced data recognized by users and servers on the blockchain, and also utilize the blockchain network as the self-recording channel for achieving non-repudiation verification interactions. We also propose a fairly arbitrable data auditing protocol with the support of the commutative hash technique, defending against dishonest provers and verifiers. Additionally, a decentralized adjudication mechanism is implemented by using the smart contract technique for creditably resolving data possession disputes without TPAs. The theoretical analysis and experimental evaluation reveal its effectiveness in undisputable data auditing and the limited requirement of costs.
Yang Xu 0013, Ju Ren 0001, Yan Zhang 0002, Cheng Zhang 0035, Bo Shen 0002, Yaoxue Zhang
IEEE Trans. Serv. Comput.2
2019 Big Data Analytics for User Association Characterization in Large-Scale WiFi System
abstract
Large-scale WiFi systems have been widely deployed in an increasing number of corporate places such as universities, big malls and companies, to provide fast Internet experience to users. However, user association patterns in such large-scale systems have not been well investigated, which is crucial for performance enhancement and intelligent system management. In this paper, we provide the analytics of a large-scale campus WiFi dataset, which includes more than 8,000 access points (APs) and 40,000 active users in the area of 3.0925 km2. By conducting extensive analysis on association patterns, we achieve several key insights as follows. First, user associations are highly dynamic as short association durations and frequent AP transitions prevail throughout the whole trace. Second, even though users may associate to many APs, they generally have a small preferable AP set in which they spend most of their WiFi connection time for data traffic; in addition, each user has distinct yet relatively fixed AP transition route, indicating that given its current associated AP, its next association AP is highly predictable. Third, diurnal association patterns are observed not only at single AP level, but also at the building and the system level, where the number of associated users and the data traffic vary periodically on a daily basis. These insights can provide valuable guidelines to numerous intelligent service provisions such as proactive service migration, edge content distribution, efficient network management.
Feng Lyu 0001, Ju Ren 0001, Nan Cheng 0001, Peng Yang 0004, Minglu Li 0001, Yaoxue Zhang, Xuemin Shen
ICC2
2019 Cutting Down Idle Listening Time: A NDN-Enabled Power Saving Mode Design for WLAN
abstract
The energy consumption for wireless interface is important for the power-constraint mobile and sensor devices. To improve energy efficiency in WLAN (such as Wi-Fi), power saving mode (PSM) is proposed, with an attempt to manage the time spent in idle listening (IL) state. The challenge is that the receiver has no knowledge about when the pending data will arrival under end-to-end communication protocols (TCP/IP); therefore each station has to spend more time in IL to wait for the pending data. To address this problem, we propose NDN-PSM, in which NDN communication architecture is leveraged to cut down unnecessary IL time. In particular, we introduce two new power states in NDN-PSM, i.e., light doze and deep doze. As stations can check pending interest table (PIT) information to predict data arrival precisely, they can switch to deep doze or light doze intelligently. The inherent receiver-driven patterns of NDN can make each station effectively go to deep doze state for power saving. We have implemented NDN-PSM in NS-3 through ndnSIM and the simulation results demonstrate that NDN-PSM can effectively reduce IL time as well as total power consumption and meanwhile retain low transmission delay. Specifically, compared to the PSM mechanism, NDN-PSM can reduce the average power consumption up to 56%.
Fan Wu 0014, Wang Yang 0002, Ju Ren 0001, Feng Lyu 0001, Peng Yang 0004, Yaoxue Zhang, Xuemin Shen
ICC3
2019 Optimal Design of Multiple Panel Arrays in LoS MIMO System
abstract
This paper investigates the optimal design of multiple panel arrays (MPAs) for line-of-sight (LoS) multiple-input multiple-output (MIMO) communication systems. We use the spherical wave channel model and give a geometric model to model the LoS channel, which allows the receive antenna arrays to have azimuth rotation, elevation angle rotation, up-down offset and left-right offset distance. Based on the geometric model, we derive the optimal antenna design conditions for achieving the maximum channel capacity and spatial multiplexing gain according to the effective degrees of freedom. The results show that the proposed antenna design can achieve high channel space freedom when the receive antennas have angle rotation and offset, and is suitable for the case where the receive antennas have a large left-right offset distance.
Ye Zhang 0033, Shiwen He, Yongming Huang 0001, Ju Ren 0001, Luxi Yang
ICC4
2019 Demystifying Traffic Statistics for Edge Cache Deployment in Large-Scale WiFi System
abstract
How to deploy cache in large-scale WiFi system is not well studied yet quite challenging since numerous Aps turn to be heterogeneous in terms of traffic consumption, and future traffic conditions are unknown ahead. In this paper, given the cache storage budge, we explore the cache deployment in a large-scale WiFi system which contains 8,000 APs and serves more than 40,000 active users, to maximize the long-term caching gain, i.e., the total reduced backhaul traffic. Specifically, we first collect enormous user association records and conduct intensive statistical analysis on the collected data, gaining two major observations. First, per AP traffic consumption varies in a rather wide range and the AP proportion distributes evenly within the range, which indicates that the cache size should be heterogeneously allocated in accordance to the underlying traffic demands. Second, compared to a single AP, the traffic consumption of a group of APs (clustered by physical locations) is more stable, which means that the short-term traffic statistics can be used to infer the future long-term traffic conditions. We then propose our cache deployment strategy, named LEAD (i.e., Large-scale wifi Edge cAche Deployment), in which we first cluster large-scale APs into well-sized edge nodes, then conduct the stationary testing on edge level traffic consumption and sample sufficient traffic statistics in order to precisely characterize future traffic conditions, and finally devise the TEG (Traffic-wEighted Greedy) algorithm to solve the long-term caching gain maximization problem. Extensive trace-driven simulations are carried out and simulation results demonstrate the efficacy of LEAD.
Feng Lyu 0001, Ju Ren 0001, Nan Cheng 0001, Peng Yang 0004, Minglu Li 0001, Yaoxue Zhang, Xuemin Shen
ICDCS2
2019 Towards Privacy-preserving Incentive for Mobile Crowdsensing Under An Untrusted Platform
abstract
Reverse auction-based incentive mechanisms have been commonly proposed to stimulate mobile users to participate in crowdsensing, where users submit bids to the platform to compete for tasks. Recent works pointed out that bid is a private information which can reveal sensitive information of users (e.g., location privacy), and proposed bid-preserving mechanisms with differential privacy against inference attack. However, all these mechanisms rely on a trusted platform, and would fail in bid protection completely when the platform is untrusted (e.g., honest-but-curious). In this paper, we focus on the bid protection problem in mobile crowdsensing with an untrusted platform, and propose a novel privacy-preserving incentive mechanism to protect users' true bids against the honest-but-curious platform while minimizing the social cost of winner selection. To this end, instead of uploading the true bid to the platform, a differentially private bid obfuscation function is designed with the exponential mechanism, which helps each user to obfuscate bids locally and submit obfuscated task-bid pairs to the platform. The winner selection problem with the obfuscated task-bid pairs is formulated as an integer linear programming problem and proved to be NP-hard. We consider the optimization problem at two different scenarios, and propose a solution based on Hungarian method for single measurement and a greedy solution for multiple measurements, respectively. The proposed incentive mechanism is proved to satisfy ε-differential privacy, individual rationality and γ-truthfulness. The extensive experiments on a real-world data set demonstrate the effectiveness of the proposed mechanism against the untrusted platform.
Zhibo Wang 0001, Jingxin Li, Jiahui Hu 0001, Ju Ren 0001, Zhetao Li, Yanjun Li 0004
INFOCOM4
2019 An adaptive and configurable protection framework against android privilege escalation threats
Yang Xu 0013, Guojun Wang 0001, Ju Ren 0001, Yaoxue Zhang
Future Gener. Comput. Syst.3
2019 Interference Cooperation via Distributed Game in 5G Networks
abstract
Nash noncooperative power game is an effective method to implement interference cooperation in downlink multiuser multiple-input multiple-output (MU-MIMO). Power equilibrium point of Nash noncooperative power game can achieve a satisfactory tradeoff between self-benefits of Internet of Things (IoT) users and interference between IoT users which largely enhance the edge IoT user throughput. However, either power strategy space, i.e., the enabled range of power allocation for IoT users, or overall BS transmit power in the existing Nash noncooperative power games is generally static. This limits the performance of systems, especially in IoT systems, etc., in 5G. As an effort to address these problems, we design a novel framework of Nash noncooperative game with iterative convergence for downlink MU-MIMO. We first decompose the MU-MIMO into multiple virtual single-antenna transmit-receive pairs with a stream analytical model. Afterwards, based on streams, we propose a noncooperative water-filling power game with pricing (WFPGP) where the power strategy space of each stream can be dynamically determined byiterative water-filling. We derive the sufficient condition for the existence and uniqueness of WFPGP game, in which the verification of the sufficient condition can be executed in a distributed manner. By simulations, we verify the performance of WFPGP compared to other Nash noncooperative games.
Shu Fu, Zhou Su 0001, Yunjian Jia, Yi Jin 0003, Ju Ren 0001, Bin Wu 0002, Kazi Mohammed Saidul Huq
IEEE Internet Things J.6
2019 Secure Data Aggregation of Lightweight E-Healthcare IoT Devices With Fair Incentives
abstract
With rapid development of e-healthcare systems, patients that are equipped with resource-limited e-healthcare devices (Internet of Things) generate huge amount of health data for health management. These health data possess significant medical value when aggregated from these distributed devices. However, efficient health data aggregation poses several security and privacy issues such as confidentiality disclosure and differential attacks, as well as patients may be reluctant to contribute their health data for aggregation. In this paper, we propose a privacy-preserving heath data aggregation scheme that securely collects health data from multiple sources and guarantee fair incentives for contributing patients. Specifically, we employ signature techniques to keep fair incentives for patients. Meanwhile, we add noises into the health data for differential privacy. Furthermore, we combine Boneh-Goh-Nissim cryptosystem and Shamir's secret sharing to keep data obliviousness security and fault tolerance. Security and privacy discussions show that our scheme can resist differential attacks, tolerate healthcare centers failures, and keep fair incentives for patients. Performance evaluations demonstrate cost-efficient computation, communication and storage overhead.
Wenjuan Tang, Ju Ren 0001, Yaoxue Zhang
IEEE Internet Things J.2
2019 EdgeSanitizer: Locally Differentially Private Deep Inference at the Edge for Mobile Data Analytics
abstract
Deep neural networks have been widely applied in various machine learning applications for mobile data analytics in cloud. However, this approach introduces significant data challenges, because the cloud operator can perform deep inferences on the available data. Recent advances in edge computing have paved the way to more efficient and private data processing at the edge of the network for simple tasks and lightweight models, but challenges still remain in building efficient complex models (e.g., deep learning) for edge computing. To tackle these issues, we propose EdgeSanitizer, a deep inference framework-based edge computing with local differential privacy for mobile data analytics. EdgeSanitizer leverages deep learning model to conduct data minimization and obfuscates the learned features by adaptively injecting noise, thereby forming a new protection layer against sensitive inference. We evaluate its performance in terms of data privacy and utility through theoretical analysis and experimental evaluation. The theoretical analysis proves that EdgeSanitizer can provide provable privacy guarantees with a large improvement in utility. And the experimental results demonstrate the robustness of our approach against sensitive inference, as well as its applicability on resource-constrained edge devices.
Chugui Xu, Ju Ren 0001, Liang She, Yaoxue Zhang, Zhan Qin, Kui Ren 0001
IEEE Internet Things J.2
2019 Two Time-Scale Resource Management for Green Internet of Things Networks
abstract
It is expected that billions of objects will be connected through sensors and embedded devices for pervasive intelligence in the coming era of Internet of Things (IoT). However, the performance of such ubiquitous interconnection highly depends on the supply of network resources in terms of both energy and spectrum. Librating IoT devices from the resource deficiency, we consider a green IoT network in which the IoT devices transmit data to a fusion node over multihop relaying. To achieve sustainable operation, IoT devices obtain energy from both ambient energy sources and power grid, while opportunistically access the licensed spectrum for data transmission. We formulate a stochastic problem to optimize the network utility minus the cost on on-grid energy purchasing. The problem formulation takes into account the different granularity in the changing of harvested energy, power price, and primary user activities. To address the problem, we propose a Lyapunov-based framework to decompose the problem into different time scales, based on which an online two time-scale resource allocation algorithm, is developed which determines the harvested and purchased energy in a large time scale, and the channel allocation and data collection in a small time scale. Furthermore, we analyze the required data buffer and energy buffer to support the proposed algorithm. Extensive simulation results validate the correctness of the analysis and the efficiency of the proposed algorithm.
Liang She, Ruyin Shen, Ju Ren 0001, Yaoxue Zhang
IEEE Internet Things J.5
2019 Cloud-Edge Coordinated Processing: Low-Latency Multicasting Transmission
abstract
Recently, edge caching and multicasting arise as two promising technologies to support high-data-rate and low-latency delivery in wireless communication networks. In this paper, we design three transmission schemes aiming to minimize the delivery latency for cache-enabled multigroup multicasting networks. In particular, full caching bulk transmission scheme is first designed as a performance benchmark for the ideal situation where the caching capability of each enhanced remote radio head (eRRH) is sufficient large to cache all files. For the practical situation where the caching capability of each eRRH is limited, we further design two transmission schemes, namely partial caching bulk transmission (PCBT) and partial caching pipelined transmission (PCPT) schemes. In the PCBT scheme, eRRHs first fetch the uncached requested files from the baseband unit (BBU) and then all requested files are simultaneously transmitted to the users. In the PCPT scheme, eRRHs first transmit the cached requested files while fetching the uncached requested files from the BBU. Then, the remaining cached requested files and fetched uncached requested files are simultaneously transmitted to the users. The design goal of the three transmission schemes is to minimize the delivery latency, subject to some practical constraints. Efficient algorithms are developed for the low-latency cloud-edge coordinated transmission strategies. Numerical results are provided to evaluate the performance of the proposed transmission schemes and show that the PCPT scheme outperforms the PCBT scheme in terms of the delivery latency criterion.
Shiwen He, Ju Ren 0001, Jiaheng Wang 0001, Yongming Huang 0001, Yaoxue Zhang, Weihua Zhuang, Xuemin Shen
IEEE J. Sel. Areas Commun.2
2019 Robust Multigroup Multicast Beamforming Design for Backhaul-Limited Cloud Radio Access Network
abstract
This letter investigates the robust beamforming design for multigroup multicast in a backhaul-limited cloud radio access network. Users requesting the same content form a multicast group, served by remote radio heads (RRHs) cooperatively. Each RRH acquires the requested contents from baseband unit via backhaul links. We first formulate the robust beamforming design as maximizing the sum of the minimum rate of users in each multicast group under the transmission power and backhaul constraints. Due to the introduction of the channel estimation error and inter-user interference, the considered problem becomes more complex and difficult to address directly. To overcome these difficulties, convex approximation methods are adopted to transform the original problem into convex one. Then, an effective optimization algorithm is developed to address the resulting problem. Numerical results demonstrate the effectiveness of the proposed robust beamforming design of multigroup multicast transmission.
Shiwen He, Yongming Huang 0001, Ju Ren 0001, Luxi Yang
IEEE Signal Process. Lett.4
2019 GANobfuscator: Mitigating Information Leakage Under GAN via Differential Privacy
abstract
By learning generative models of semantic-rich data distributions from samples, generative adversarial network (GAN) has recently attracted intensive research interests due to its excellent empirical performance as a generative model. The model is used to estimate the underlying distribution of a dataset and randomly generate realistic samples according to their estimated distribution. However, GANs can easily remember training samples due to the high model complexity of deep networks. When GANs are applied to private or sensitive data, the concentration of distribution may divulge some critical information. It consequently requires new technological advances to mitigate the information leakage under GANs. To address this issue, we propose GANobfuscator, a differentially private GAN, which can achieve differential privacy under GANs by adding carefully designed noise to gradients during the learning procedure. With GANobfuscator, analysts are able to generate an unlimited amount of synthetic data for arbitrary analysis tasks without disclosing the privacy of training data. Moreover, we theoretically prove that GANobfuscator can provide strict privacy guarantee with differential privacy. In addition, we develop a gradient-pruning strategy for GANobfuscator to improve the scalability and stability of data training. Through extensive experimental evaluation on benchmark datasets, we demonstrate that GANobfuscator can produce high-quality generated data and retain desirable utility under practical privacy budgets.
Chugui Xu, Ju Ren 0001, Yaoxue Zhang, Zhan Qin, Kui Ren 0001
IEEE Trans. Inf. Forensics Secur.2
2019 A Blockchain-Based Nonrepudiation Network Computing Service Scheme for Industrial IoT
abstract
Emerging network computing technologies extend the functionalities of industrial IoT (IIoT) terminals. However, this promising service-provisioning scheme encounters problems in untrusted and distributed IIoT scenarios because malicious service providers or clients may deny service provisions or usage for their own interests. Traditional nonrepudiation solutions fade in IIoT environments due to requirements of trusted third parties or unacceptable overheads. Fortunately, the blockchain revolution facilitates innovative solutions. In this paper, we propose a blockchain-based fair nonrepudiation service provisioning scheme for IIoT scenarios in which the blockchain is used as a service publisher and an evidence recorder. Each service is separately delivered via on-chain and off-chain channels with mandatory evidence submissions for nonrepudiation purpose. Moreover, a homomorphic-hash-based service verification method is designed that can function with mere on-chain evidence. And an impartial smart contract is implemented to resolve disputes. The security analysis demonstrates the dependability, and the evaluations reveal the effectiveness and efficiency.
Yang Xu 0013, Ju Ren 0001, Guojun Wang 0001, Cheng Zhang 0035, Jidian Yang, Yaoxue Zhang
IEEE Trans. Ind. Informatics2
2019 Efficient and Privacy-preserving Fog-assisted Health Data Sharing Scheme
abstract
Pervasive data collected from e-healthcare devices possess significant medical value through data sharing with professional healthcare service providers. However, health data sharing poses several security issues, such as access control and privacy leakage, as well as faces critical challenges to obtain efficient data analysis and services. In this article, we propose an efficient and privacy-preserving fog-assisted health data sharing (PFHDS) scheme for e-healthcare systems. Specifically, we integrate the fog node to classify the shared data into different categories according to disease risks for efficient health data analysis. Meanwhile, we design an enhanced attribute-based encryption method through combination of a personal access policy on patients and a professional access policy on the fog node for effective medical service provision. Furthermore, we achieve significant encryption consumption reduction for patients by offloading a portion of the computation and storage burden from patients to the fog node. Security discussions show that PFHDS realizes data confidentiality and fine-grained access control with collusion resistance. Performance evaluations demonstrate cost-efficient encryption computation, storage and energy consumption.
Wenjuan Tang, Ju Ren 0001, Kuan Zhang 0001, Yaoxue Zhang, Xuemin Shen
ACM Trans. Intell. Syst. Technol.2
2019 Flexible and Efficient Authenticated Key Agreement Scheme for BANs Based on Physiological Features
abstract
In Body Area Networks (BANs), bio-sensors can collect personal health information and cooperate with each other to provide intelligent health care services for medical users. Since personal health information is highly privacy-sensitive, the flourish of BANs still faces critical security challenges, especially secure communication between bio-sensors. In this paper, we propose a flexible and efficient authenticated key agreement scheme (PBAKA) to provide secure communication for BANs. Specifically, we employ a control unit (e.g., smart phone) to launch authentication based on physiological features collected from BANs, and integrate bilinear pairings to negotiate session keys for bio-sensors. Since physiological features can be collected from various kinds of bio-sensors in real time, PBAKA is flexible for adding new bio-sensors without pre-distributed keys. Meanwhile, PBAKA is computationally efficient by offloading authentication burden from resource-limited bio-sensors to the control unit. Security analysis demonstrates that PBAKA is provably secure under the decisional bilinear Diffie-Hellman assumption. Extensive experimental results validate efficient communication, computation and energy consumption of our scheme when compared with several existing solutions.
Wenjuan Tang, Kuan Zhang 0001, Ju Ren 0001, Yaoxue Zhang, Xuemin Shen
IEEE Trans. Mob. Comput.3
2019 Distributed and Efficient Object Detection via Interactions Among Devices, Edge, and Cloud
abstract
With the rapid development of Internet-of-Things and communication techniques, media transmission in surveillance applications is gradually relying on wireless networks. Meanwhile, the emergence of edge computing has pushed the media data analysis from the cloud to the edge of the network to achieve fast response for delay-sensitive media processing tasks. Object detection is a representative delay-sensitive image processing task in surveillance applications, but faces significant challenges in this context. For example, how to compress images for transmission in wireless environment without compromising the detection accuracy, and how to integrate and update local inference models online in an edge computing-based object detection system. In this paper, we propose an object detection architecture based on edge computing to achieve distributed and efficient object detection for surveillance applications. Under this architecture, we develop an adaptive Region-of-Interest-based image compression scheme for end devices to efficiently compress their captured images for wireless transmission but not to sacrifice the object detection accuracy of edge servers. Furthermore, we carefully design distributed and communication-efficient interactions among end devices, edge servers, and the cloud to dynamically optimize the object detection accuracy online. Extensive simulation results demonstrate that our proposed architecture not only achieves a competitive detection accuracy to traditional cloud-based objective detection solution with reduced response delay but also significantly improves the image transmission efficiency with adaptive image compression ratio.
Yun-Di Guo, Beiji Zou 0001, Ju Ren 0001, Yaoxue Zhang
IEEE Trans. Multim.3
2019 Enabling Trusted and Privacy-Preserving Healthcare Services in Social Media Health Networks
abstract
Social Media Health Networks provide a promising paradigm to attract patients to share and communicate their personal health status with other online patients, and consult healthcare services from online caregivers with social networks. Social Media Health Networks transform healthcare services from time-consuming offline hospital-centered paradigm to the convenient and efficient online paradigm through Internet, which can expand the traditional healthcare services and shorten the information gap between patients and caregivers. However, how to build the trust between patients and caregivers raises a challenging issue due to the openness of the social networks; meanwhile, the personal privacy may be disclosed when sharing personal health information with other patients and caregivers. In this paper, we propose a personalized and trusted healthcare service approach to enable trusted and privacy-preserving healthcare services in social media health networks, which can improve the trustiness between patients and caregivers through authentic ratings toward caregivers and guarantee the patients' privacy. Specifically, we employ the collaborative filtering model to seek appropriate personalized caregivers, bloom filter to extract and map the personal healthcare symptoms, and inner product to compute the similarity between patients for finding patients with similar health symptoms in a privacy-preserving way. Meanwhile, to guarantee authentic ratings and reviews toward caregivers, we develop a sybil attack detection scheme to find patients' fake ratings and reviews using different pseudonyms. Security analysis shows that our proposed approach can preserve the privacy of patients and prevent sybil attacks. Performance evaluation demonstrates that our approach can achieve prominent performance improvement, in terms of personalized caregivers finding and sybil attack resistance.
Wenjuan Tang, Ju Ren 0001, Yaoxue Zhang
IEEE Trans. Multim.2
2019 T-Gaming: A Cost-Efficient Cloud Gaming System at Scale
abstract
Cloud gaming (CG) system could pursue both high-quality gaming experience via intensive computing, and ultimate convenience anywhere at anytime through any energy-constrained mobile devices. Despite the abundance of efforts devoted, state-of-the-art CG systems still suffer from multiple key limitations: expensive deployment cost, high bandwidth consumption and unsatisfied quality of experience (QoE). As a result, existing works are not widely adopted in reality. This paper proposes a Transparent Gaming framework called T-Gaming that allows users to play any popular high-end desktop/console games on-the-fly over the Internet. T-Gaming utilizes the off-the-shelf consumer GPUs without resorting to the expensive proprietary GPU virtualization (vGPU) technology to reduce the deployment cost. Moreover, it enables prioritized video encoding based on the human visual feature to reduce the bandwidth consumption without noticeable visual quality degradation. Last but not least, T-Gaming adopts adaptive real-time streaming based on deep reinforcement learning (RL) to improve user's QoE. To evaluate the performance of T-Gaming, we implement and test a prototype system in the real world. Compared with the existing cloud gaming systems, T-Gaming not only reduces the expense per user by 75 percent hardware cost reduction and 14.3 percent network cost reduction, but also improves the normalized average QoE by 3.6-27.9 percent.
Hao Chen 0036, Xu Zhang 0006, Yiling Xu, Ju Ren 0001, Jingtao Fan, Zhan Ma 0001, Wenjun Zhang 0001
IEEE Trans. Parallel Distributed Syst.4
2019 Hybrid Precoder Design for Cache-Enabled Millimeter-Wave Radio Access Networks
abstract
In this paper, we study the design of a hybrid precoder, consisting of an analog and a digital precoder, for the delivery phase of downlink cache-enabled millimeter-wave (mm-wave) radio access networks (CeMm-RANs). In CeMm-RANs, enhanced remote radio heads (eRRHs), which are equipped with local cache and baseband signal processing capabilities in addition to the basic functionalities of conventional RRHs, are connected to the baseband processing unit via fronthaul links. Two different fronthaul information transfer strategies are considered, namely, hard fronthaul information transfer, where hard information of uncached requested files is transmitted via the fronthaul links to a subset of eRRHs, and soft fronthaul information transfer, where the fronthaul links are used to transmit quantized baseband signals of uncached requested files. The hybrid precoder is optimized for maximization of the minimum user rate under a fronthaul capacity constraint, an eRRH transmit power constraint, and a constant-modulus constraint on the analog precoder. The resulting optimization problem is non-convex, and hence, the global optimal solution is difficult to obtain. Therefore, convex approximation methods are employed to tackle the non-convexity of the achievable user rate, the fronthaul capacity constraint, and the constant modulus constraint on the analog precoder. Then, an effective algorithm with provable convergence is developed to solve the approximated optimization problem. The simulation results are provided to evaluate the performance of the proposed algorithms, where fully digital precoding is used as the benchmark. The results reveal that except for the case of a large fronthaul link capacity, soft fronthaul information transfer is preferable for CeMm-RANs. Furthermore, surprisingly, hybrid precoding outperforms fully digital precoding with soft fronthaul information transfer for medium-to-large file sizes and fronthaul capacity limited mm-wave cloud RANs.
Shiwen He, Yongpeng Wu 0001, Ju Ren 0001, Yongming Huang 0001, Robert Schober, Yaoxue Zhang
IEEE Trans. Wirel. Commun.3
2018 Joint Selection and Scheduling of Communication Requests in Multi-Channel Wireless Networks under SINR Model
Peng-Jun Wan, Huaqiang Yuan, Jiliang Wang, Ju Ren 0001, Yaoxue Zhang
INFOCOM4
2018 Resource allocation for hybrid energy powered cloud radio access network with battery leakage
abstract
This study proposes a resource allocation policy for hybrid energy powered cloud radio access network with battery leakage. To optimise the network utility, the authors first formulate a network utility maximisation problem while jointly considering multiple random processes, which include energy harvesting, data arrival, wireless channel condition, and grid energy prices. To tackle this problem, they exploit the Lyapunov optimisation technique to develop an online dynamic resource allocation framework, which contains four subproblems, i.e. data admission, hybrid energy management, power allocation and route scheduling. Based on the solutions of these subproblems, a network utility optimisation resource allocation (ORA) algorithm is proposed. Specifically, the ORA algorithm only needs to track the current system states without requiring a prior knowledge about channel and energy conditions. Theoretical performance analyses and simulation results verify that the proposed algorithm can achieve close‐to‐optimal utility with bounded data buffer and battery capacity.
Sijing Duan, Ju Ren 0001, Yaoxue Zhang
IET Commun.3
2018 BOAT: A Block-Streaming App Execution Scheme for Lightweight IoT Devices
abstract
The contradiction between the limited capability of lightweight Internet-of-Things (IoT) devices and ever-increasing user demands is a fundamental and challenging problem in the era of IoT. One of the most important challenges is that lightweight IoT devices are generally embedded with fixed applications on resource-constrained hardware, and cannot enable on-demand service provisioning. In this paper, we propose BOAT, a block-streaming application execution scheme based on transparent computing (TC), which can remotely retrieve the necessary parts of traditional applications from edge servers on demand and run them on IoT devices locally. Specifically, we first exploit TC to build a scalable IoT system, where the applications of lightweight IoT devices are stored in edge servers or cloud but can be dynamically loaded on IoT devices in a block-streaming way. Moreover, we present a code partition approach to split the codes of a whole service into numerous functional blocks at the edge side. Such that, IoT devices only need to load necessary blocks of an application to obtain the requested services without loading the whole application codes. A lightweight I/O virtualization mechanism and a fine-grained relocation technology are then developed to support the block-streaming service loading. Our experimental results on a lightweight wearable device demonstrate that the proposed scheme can efficiently achieve flexible service provisioning with improved scalability and reduced service loading delay and energy consumption.
Xuhong Peng, Ju Ren 0001, Liang She, Jie Li 0058, Yaoxue Zhang
IEEE Internet Things J.2
2018 A scalable and manageable IoT architecture based on transparent computing
Ju Ren 0001, Yaoxue Zhang, Junying Hu
J. Parallel Distributed Comput.2
2018 Guest editorial: special issue on transparent computing
Jiannong Cao 0001, Jingde Cheng, Jianhua Ma 0002, Ju Ren 0001
Peer-to-Peer Netw. Appl.4
2018 A Feasible Fuzzy-Extended Attribute-Based Access Control Technique
abstract
Attribute-based access control (ABAC) is a maturing authorization technique with outstanding expressiveness and scalability, which shows its overwhelmingly competitive advantage, especially in complicated dynamic environments. Unfortunately, the absence of a flexible exceptional approval mechanism in ABAC impairs the resource usability and business time efficiency in current practice, which could limit its growth. In this paper, we propose a feasible fuzzy-extended ABAC (FBAC) technique to improve the flexibility in urgent exceptional authorizations and thereby improving the resource usability and business timeliness. We use the fuzzy assessment mechanism to evaluate the policy-matching degrees of the requests that do not comply with policies, so that the system can make special approval decisions accordingly to achieve unattended exceptional authorizations. We also designed an auxiliary credit mechanism accompanied by periodic credit adjustment auditing to regulate expediential authorizations for mitigating risks. Theoretical analyses and experimental evaluations show that the FBAC approach enhances resource immediacy and usability with controllable risk.
Yang Xu 0013, Wuqiang Gao, Quanrun Zeng, Guojun Wang 0001, Ju Ren 0001, Yaoxue Zhang
Secur. Commun. Networks5
2018 Joint Load Scheduling and Voltage Regulation in the Distribution System With Renewable Generators
abstract
By equipping with the advanced smart meters and two-way communications infrastructure, smart grids, as a key component of future smart cities, are able to improve the energy efficiency and reduce the energy cost through real-time monitoring and customer load scheduling. However, the high penetration of intermittent renewable energy such as solar power may cause frequent overvoltage and undervoltage problems at certain buses, making the load scheduling face new challenges on voltage regulation. In this paper, we investigate the impact of voltage constraints on load scheduling by power flow analysis in a power distribution system with renewable generators. A voltage regulator (VR) is introduced to regulate the voltage of buses in the distribution system and assist load scheduling. To jointly minimize the cost and stabilize the voltages of the distribution system, we propose a grid-customer coordinated load scheduling strategy, which simultaneously determines the tap changes of the VR and scheduling of customer electricity loads in each time slot. Finally, we evaluate the performance of the proposed strategy based on realistic power demand and renewable energy generation datasets. Extensive numerical results demonstrate that the proposed strategy can remarkably reduce the energy cost and stabilize the voltage fluctuation of distribution systems.
Ju Ren 0001, Junying Hu, Ruilong Deng, Yaoxue Zhang, Xuemin Shen
IEEE Trans. Ind. Informatics1
2018 Towards Secure Network Computing Services for Lightweight Clients Using Blockchain
abstract
The emerging network computing technologies have significantly extended the abilities of the resource‐constrained IoT devices through the network‐based service sharing techniques. However, such a flexible and scalable service provisioning paradigm brings increased security risks to terminals due to the untrustworthy exogenous service codes loading from the open network. Many existing security approaches are unsuitable for IoT environments due to the high difficulty of maintenance or the dependencies upon extra resources like specific hardware. Fortunately, the rise of blockchain technology has facilitated the development of service sharing methods and, at the same time, it appears a viable solution to numerous security problems. In this paper, we propose a novel blockchain‐based secure service provisioning mechanism for protecting lightweight clients from insecure services in network computing scenarios. We introduce the blockchain to maintain all the validity states of the off‐chain services and edge service providers for the IoT terminals to help them get rid of untrusted or discarded services through provider identification and service verification. In addition, we take advantage of smart contracts which can be triggered by the lightweight clients to help them check the validities of service providers and service codes according to the on‐chain transactions, thereby reducing the direct overhead on the IoT devices. Moreover, the adoptions of the consortium blockchain and the proof of authority consensus mechanism also help to achieve a high throughput. The theoretical security analysis and evaluation results show that our approach helps the lightweight clients get rid of untrusted edge service providers and insecure services effectively with acceptable latency and affordable costs.
Yang Xu 0013, Guojun Wang 0001, Jidian Yang, Ju Ren 0001, Yaoxue Zhang, Cheng Zhang 0035
Wirel. Commun. Mob. Comput.4
2017 Lightweight and Privacy-Preserving Fog-Assisted Information Sharing Scheme for Health Big Data
abstract
With the advancements of electronic medical equipment, e-healthcare system becomes a promising paradigm to continuously monitor health conditions and remotely diagnose phenomena. Meanwhile, it generates a large volume of health data and poses several security challenges, such as access control and privacy leakage. In this paper, we propose a lightweight and privacy- preserving fog-assisted information sharing scheme (PFHD) for health big data. Specifically, we integrate fog computing into e-healthcare system to pre-process the raw health data and improve the efficiency of health data analysis. Furthermore, to prevent privacy leakage, we design a hierarchical attribute-based encryption method by encrypting the profile and health data with different access policies. In addition, we reduce the computation cost on devices by offloading health data encryption from devices to fog servers. Security discussions show that PFHD can achieve fine- grained health data sharing with privacy preservation. Performance evaluations demonstrate the efficiency of PFHD, especially in terms of encryption computation and storage costs.
Wenjuan Tang, Kuan Zhang 0001, Ju Ren 0001, Yaoxue Zhang, Xuemin Shen
GLOBECOM3
2017 DPPro: Differentially Private High-Dimensional Data Release via Random Projection
abstract
Releasing representative data sets without compromising the data privacy has attracted increasing attention from the database community in recent years. Differential privacy is an influential privacy framework for data mining and data release without revealing sensitive information. However, existing solutions using differential privacy cannot effectively handle the release of high-dimensional data due to the increasing perturbation errors and computation complexity. To address the deficiency of existing solutions, we propose DPPro, a differentially private algorithm for high-dimensional data release via random projection to maximize utility while guaranteeing privacy. We theoretically prove that DPPro can generate synthetic data set with the similar squared Euclidean distance between high-dimensional vectors while achieving (ϵ, δ)-differential privacy. Based on the theoretical analysis, we observed that the utility guarantees of released data depend on the projection dimension and the variance of the noise. Extensive experimental results demonstrate that DPPro substantially outperforms several state-of-the-art solutions in terms of perturbation error and privacy budget on high-dimensional data sets.
Chugui Xu, Ju Ren 0001, Yaoxue Zhang, Zhan Qin, Kui Ren 0001
IEEE Trans. Inf. Forensics Secur.2
2016 Indoor Temperature Control of Cost-Effective Smart Buildings via Real-Time Smart Grid Communications
abstract
Under the real-time electricity pricing environment in smart grid, building owners are faced with the indoor temperature control problem to minimize the daily electricity cost. Taking the heating scenario as an example, an intuitive strategy is to maintain the building's indoor temperature always at the lower bound of the predetermined comfort range. However, this strategy may not always achieve the lowest electricity bill, especially with the significant fluctuation of electricity prices. On the other hand, the cost minimization problem can be optimally solved one day ahead in a temporally- coupled manner, but the challenge lies in that the building owner needs to acquire the accurate information of electricity prices and outdoor temperatures of the next day, which may not be available. In this paper, we equivalently decouple the cost minimization problem into subproblems at each hour. Each subproblem can be temporally decoupled and optimally solved, only requiring the next-hour electricity price. Besides, the temporally-decoupled algorithm explicitly indicates when to take advantage of pre-heating/cooling for electricity cost reduction. It is demonstrated with the real data that our proposed algorithm could result in considerable economic savings compared with the intuitive strategy, paving the way towards practically applicable cost-effective smart buildings.
Ruilong Deng, Ju Ren 0001, Hao Liang 0002
GLOBECOM3
2016 Resource Allocation for Green Cloud Radio Access Networks Powered by Renewable Energy
abstract
In this paper, we investigate the sustainable resource allocation for green Cloud Radio Access Networks (C-RAN) powered by renewable energy. Specifically, the Base Station pool (BS pool) in the C-RAN distributes data to a set of remote radio heads (RRHs) with energy harvesting (EH) capability, and allocates sub-carriers to the selected RRHs for downlink transmissions, by jointly considering the user throughput and energy sustainability performance of RRHs. To this end, we formulate a utility optimization problem, characterizing the stochastic process of energy harvesting (EH) and wireless fading channel. Based on Lyapunov optimization techniques, we decompose the formulated problem into three sub-problems, including energy harvesting, data scheduling, and sub-carrier allocation. We then propose an efficient online algorithm to obtain the maximal aggregate user utility while ensuring the stability of the data buffers and sustainability of the energy buffers. Performance analysis demonstrates that the proposed algorithm can achieve a suboptimal performance with guaranteed upper bounds on data queue and energy queue lengths. Extensive simulations validate the effectiveness and efficiency of the proposed algorithm.
Zhigang Chen 0001, Lin X. Cai, Ju Ren 0001, Xuemin Shen
GLOBECOM5
2016 An imbalanced data classification method based on automatic clustering under-sampling
abstract
Classification of imbalanced datasets has become one of the most challenging problems in big data mining. Because the number of positive samples is far less than the negative samples, low accuracy and poor generalization performance and some other defects always go with learning process of traditional algorithms. Ensemble construction algorithm is an important method to handle this problem. Especially, the ensemble construction algorithm based on random under-sampling or clustering can effectively improve the performance of classification. However, the former causes information loss easily and the latter increases complexity. In this paper, we propose ACUS, an improved ensemble algorithm based on automatic clustering and under-sampling. ACUS conducts clustering first according to the weight of samples, and then it constructs balanced-distributed dataset which consists of a certain percentage of the majority class and all of the minority class from each cluster. With Adaboost algorithm construction, these datasets are used to get an ensemble classifier. Experimental results demonstrate the advantages of our proposed algorithm in terms of accuracy, simplicity and high stability.
Xiaoheng Deng, Weijian Zhong, Ju Ren 0001, Detian Zeng, Honggang Zhang 0003
IPCCC3
2016 Lifetime and Energy Hole Evolution Analysis in Data-Gathering Wireless Sensor Networks
abstract
Network lifetime is a crucial performance metric to evaluate data-gathering wireless sensor networks (WSNs) where battery-powered sensor nodes periodically sense the environment and forward collected samples to a sink node. In this paper, we propose an analytic model to estimate the entire network lifetime from network initialization until it is completely disabled, and determine the boundary of energy hole in a data-gathering WSN. Specifically, we theoretically estimate the traffic load, energy consumption, and lifetime of sensor nodes during the entire network lifetime. Furthermore, we investigate the temporal and spatial evolution of energy hole and apply our analytical results to WSN routing in order to balance the energy consumption and improve the network lifetime. Extensive simulation results are provided to demonstrate the validity of the proposed analytic model in estimating the network lifetime and energy hole evolution process.
Ju Ren 0001, Yaoxue Zhang, Kuan Zhang 0001, Anfeng Liu, Jianer Chen, Xuemin Shen
IEEE Trans. Ind. Informatics1
2016 Exploiting Secure and Energy-Efficient Collaborative Spectrum Sensing for Cognitive Radio Sensor Networks
abstract
Cognitive radio sensor network (CRSN) has emerged as a promising solution to address the spectrum scarcity problem in traditional sensor networks, by enabling sensor nodes to opportunistically access licensed spectrum. To protect the transmission of primary users and enhance spectrum utilization, collaborative spectrum sensing is generally adopted for improving spectrum sensing accuracy. However, as sensor nodes may be compromised by adversaries, these nodes can send false sensing reports to mislead the spectrum sensing decision, making CRSNs vulnerable to spectrum sensing data falsification (SSDF) attacks. Meanwhile, since the energy consumption of spectrum sensing is considerable for energy-limited sensor nodes, SSDF attack countermeasures should be carefully devised with the consideration of energy efficiency. To this end, we propose a secure and energy-efficient collaborative spectrum sensing scheme to resist SSDF attacks and enhance the energy efficiency in CRSNs. Specifically, we theoretically analyze the impacts of two types of attacks, i.e., independent and collaborative SSDF attacks, on the accuracy of collaborative spectrum sensing in a probabilistic way. To maximize the energy efficiency of spectrum sensing, we calculate the minimum number of sensor nodes needed for spectrum sensing to guarantee the desired accuracy of sensing results. Moreover, a trust evaluation scheme, named FastDtec, is developed to evaluate the spectrum sensing behaviors and fast identify compromised nodes. Finally, a secure and energy-efficient collaborative spectrum sensing scheme is proposed to further improve the energy efficiency of collaborative spectrum sensing, by adaptively isolating the identified compromised nodes from spectrum sensing. Extensive simulation results demonstrate that our proposed scheme can resist SSDF attacks and significantly improve the energy efficiency of collaborative spectrum sensing.
Ju Ren 0001, Yaoxue Zhang, Qiang Ye 0002, Kan Yang 0001, Kuan Zhang 0001, Xuemin Shen
IEEE Trans. Wirel. Commun.1
2016 Adaptive and Channel-Aware Detection of Selective Forwarding Attacks in Wireless Sensor Networks
abstract
Wireless sensor networks (WSNs) are vulnerable to selective forwarding attacks that can maliciously drop a subset of forwarding packets to degrade network performance and jeopardize the information integrity. Meanwhile, due to the unstable wireless channel in WSNs, the packet loss rate during the communication of sensor nodes may be high and vary from time to time. It poses a great challenge to distinguish the malicious drop and normal packet loss. In this paper, we propose a channel-aware reputation system with adaptive detection threshold (CRS-A) to detect selective forwarding attacks in WSNs. The CRS-A evaluates the data forwarding behaviors of sensor nodes, according to the deviation of the monitored packet loss and the estimated normal loss. To optimize the detection accuracy of CRS-A, we theoretically derive the optimal threshold for forwarding evaluation, which is adaptive to the time-varied channel condition and the estimated attack probabilities of compromised nodes. Furthermore, an attack-tolerant data forwarding scheme is developed to collaborate with CRS-A for stimulating the forwarding cooperation of compromised nodes and improving the data delivery ratio of the network. Extensive simulation results demonstrate that CRS-A can accurately detect selective forwarding attacks and identify the compromised sensor nodes, while the attack-tolerant data forwarding scheme can significantly improve the data delivery ratio of the network.
Ju Ren 0001, Yaoxue Zhang, Kuan Zhang 0001, Xuemin Shen
IEEE Trans. Wirel. Commun.1
2016 Dynamic Channel Access to Improve Energy Efficiency in Cognitive Radio Sensor Networks
abstract
Wireless sensor networks operating in the license-free spectrum suffer from uncontrolled interference as those spectrum bands become increasingly crowded. The emerging cognitive radio sensor networks (CRSNs) provide a promising solution to address this challenge by enabling sensor nodes to opportunistically access licensed channels. However, since sensor nodes have to consume considerable energy to support CR functionalities, such as channel sensing and switching, the opportunistic channel accessing should be carefully devised for improving the energy efficiency in CRSN. To this end, we investigate the dynamic channel accessing problem to improve the energy efficiency for a clustered CRSN. Under the primary users' protection requirement, we study the resource allocation issues to maximize the energy efficiency of utilizing a licensed channel for intra-cluster and inter-cluster data transmission, respectively. Moreover, with the consideration of the energy consumption in channel sensing and switching, we further determine the condition when sensor nodes should sense and switch to a licensed channel for improving the energy efficiency, according to the packet loss rate of the license-free channel. In addition, two dynamic channel accessing schemes are proposed to identify the channel sensing and switching sequences for intra-cluster and inter-cluster data transmission, respectively. Extensive simulation results demonstrate that the proposed channel accessing schemes can significantly reduce the energy consumption in CRSNs.
Ju Ren 0001, Yaoxue Zhang, Ning Zhang 0007, Xuemin Shen
IEEE Trans. Wirel. Commun.1
2015 SACRM: Social Aware Crowdsourcing with Reputation Management in mobile sensing
Ju Ren 0001, Yaoxue Zhang, Kuan Zhang 0001, Xuemin Shen
Comput. Commun.1
2014 Exploiting channel-aware reputation system against selective forwarding attacks in WSNs
abstract
Wireless sensor networks (WSNs) are vulnerable to selective forwarding attacks that selectively drop a subset of the forwarding packets to degrade network performances. Due to unstable wireless channels, the packet loss rate between sensor nodes might be high, especially in hostile environments. Therefore, it is difficult to distinguish the malicious drop and normal packet loss. In this paper, we propose a Channel-aware deputation System (CRS) to identify selective forwarding misbehaviours from normal packet losses caused by poor channel quality or medium access collision. Specifically, CRS is based on normal packet loss estimation and neighbour monitoring. Each node maintains a reputation table to evaluate forwarding behaviours of its neighbours. Reputation value is determined by the deviation of the monitored packet loss rate and estimated normal loss rate. The nodes with reputation below a threshold are identified as misbehaving nodes and isolated from data forwarding paths. Furthermore, we develop weighted reputation propagation and integration functions to improve detection efficiency. Through theoretical analysis and extensive simulations, we demonstrate that CRS can accurately detect selective forwarding attacks and significantly improve the network throughput.
Ju Ren 0001, Yaoxue Zhang, Kuan Zhang 0001, Xuemin Shen
GLOBECOM1
2012 DCFR: A novel Double Cost Function based Routing algorithm for wireless sensor networks
abstract
Cost function based routing has been widely studied in wireless sensor networks for energy efficiency and network lifetime elongation. Existing algorithms however have limited effects because they adopt a single cost function that does not fully capture nodal energy consumption situation. In this paper, we propose a novel Double Cost Function based Routing (DCFR) algorithm, which takes into account end-to-end energy consumption, nodal remaining energy, and energy consumption rate altogether. An extensive simulation indicates that DCFR can lead to more balanced and efficient energy usage among nodes than existing algorithms.
Anfeng Liu, Ju Ren 0001, Xu Li 0001, Zhigang Chen 0001, Xuemin Shen
ICC2
2012 Design principles and improvement of cost function based energy aware routing algorithms for wireless sensor networks
Anfeng Liu, Ju Ren 0001, Xu Li 0001, Zhigang Chen 0001, Xuemin Shen
Comput. Networks2