Wanli Xue

dblp:153/8037 · DBLP profile ↗
← Back
46ranked-venue papers
8as first author
34since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 10 since 2021Computer networks · 13 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 ADASign: Adaptive deformable visual attention for continuous sign language recognition
Xuyan Zhang, Huayu Ma, Wanli Xue, Leming Guo, Tiantian Yuan, Shengyong Chen
Neurocomputing3
2026 Frequency transform attack: a transferable adversarial framework for continuous sign language recognition
Yachao Lin, Wanli Xue, Leming Guo, Tiantian Yuan
Multim. Syst.3
2026 Instance-level geometric prompting for visual object tracking
Wanli Xue, Huayu Ma, Yangcan Wu, Shengyong Chen
Pattern Recognit.2
2026 Trustworthy Continuous Sign Language Recognition
abstract
Continuous sign language recognition (CSLR) uses visual cues (e.g., hands, face, mouth, and body) to automatically recognize the sign language of deaf people, helping them to actively communicate with hearing people. The effects of these visual cues change dynamically with the demonstration of sign language. However, previous CSLR methods usually model visual information from the entire frame or simple fused visual cues, and thus do not well describe such dynamic change among visual cues. Therefore, we propose the Trustworthy Fusion Network ( TFN) of visual cues for CSLR, which comprises two fundamental modules: Intra-cue Cross-modality Feature Fusion module ( IntraCFF) and Inter-cue Trustworthy Fusion module (InterTF). IntraCFF uses the calibrated joint-belief method to dynamically fuse cross-modality features of RGB and keypoint information, to obtain a robust visual cue feature. InterTF innovatively employs the Dempster-Shafer Theory (DST) to evaluate the uncertainty of different cues in expressing sign movements. Then, the trustworthy fusion via DST is used to adaptively weigh and credibly fuse the visual cues based on uncertainty. In addition, to address the semantic gap when fusing different cues, we design consistency fusion constraints during the training stage. These constraints enhance the semantic consistency of different cues with global sign movements. Experiments on publicly CSLR datasets validate the effectiveness of our TFN.
Yan Zhang 0154, Wanli Xue, Leming Guo, Yangcan Wu, Tiantian Yuan, Shengyong Chen
IEEE Trans. Multim.2
2025 TGSR: Template-Guided Semantic Resampling against Adversarial Tracking Attacks
abstract
Deep object tracking has made significant strides, demonstrating impressive accuracy across diverse visual scenarios. However, recent studies revealed that visual object trackers remain vulnerable to adversarial attacks specifically designed to disrupt tracking tasks. While image resampling—reconstructing images using resampled coordinates and bilinear interpolation—has shown promise in enhancing tracker robustness by disrupting adversarial patterns, this naive approach overlooks crucial semantic information and scale variations between frames relative to the initial object template. These variations typically arise from changes in the target object or background. To address this limitation, we propose a template-guided semantic resampling (TGSR) method to counter adversarial tracking attacks. Our approach comprises two key components: template-aware joint semantic and appearance implicit representation (T-SAIR) and template-aware predictive resampling (T-PRES). T-SAIR estimates semantic embeddings and corresponding pixel colors at arbitrary coordinates based on incoming frames and the object template, while T-PRES predicts pixel-wise coordinate shifts in response to scale changes relative to the template. The integration of these modules enables our method to effectively reconstruct frames, neutralizing adversarial perturbations while preserving semantic information relative to the object template. Extensive experimental evaluation against three attacks across typical tracking methods demonstrates the effectiveness of our approach.
Xuhong Ren, Jianlang Chen, Wanli Xue, Lei Ma 0003, Qing Guo 0005, Jianjun Zhao 0001, Shengyong Chen
ICME3
2025 SMANet: Sequence-enhanced multi-head attention network for robust neural semantic learning in noisy computational environments
Guo Jia, Jinqi Zhu, Weijia Feng, Wanli Xue
Neurocomputing7
2025 SCSC: The Super Compressed Semantic Communication Transmission Method for Images and Video Frames
abstract
Semantic communication is a transformative approach for efficient data transmission in bandwidth-constrained IoT networks, particularly for device-to-device (D2D) communication and real-time systems. To address the challenges of high-quality data reconstruction in low-bandwidth scenarios, such as IoT-enabled underwater or wireless networks, this article proposes a novel framework for ultraefficient image and video frame compression. The framework employs an optimized feature extraction method to capture essential semantic information from images or video frames, converting them into compact grayscale representations to minimize data volume for bandwidth-constrained devices. At the receiver, a dual-decoding strategy reconstructs structural details using lightweight semantic reconstruction techniques, followed by color attribute recovery, ensuring high visual quality in real-time applications. This approach enhances transmission efficiency, reduces latency, and maintains information integrity in resource-limited IoT environments. Experimental results on the Kodak dataset show a 7.02% compression ratio (PSNR = 13 dB), with a 9.76% improvement in bandwidth efficiency compared to existing methods. For 4K images and video frames, the framework achieves a 98.35% MS-SSIM retention rate under SNR conditions of 1–15 dB, demonstrating robust performance for real-time IoT communication.
Jinqi Zhu, Weijia Feng, Shuqing He, Wanli Xue
IEEE Internet Things J.6
2025 CORE: Multi-link graph attention network with inter-regional collaboration for continuous sign language recognition
Yan Zhang 0154, Wanli Xue, Tiantian Yuan, Shengyong Chen
Pattern Recognit.2
2025 Vehicle-Level Fairness-Oriented Constrained Multi-Agent Reinforcement Learning for Adaptive Traffic Signal Control
abstract
Multi-agent Reinforcement Learning (MARL) has shown considerable promise in enhancing the efficiency of adaptive traffic signal control (ATSC) systems. However, existing MARL approaches primarily focus on optimizing overall traffic flow, often overlooking the issue of fairness in vehicle waiting times. Considering that there is no need to strive for the ultimate fairness, this paper models the ATSC problem as a Constrained Partially Observable Markov Game (CPOMG), where fairness is modeled as a constraint on the maximum waiting time of vehicles on lanes of intersections instead of a reward term that pursues maximization. CPOMG aims to find a cooperative control policy with optimal traffic efficiency within the constrained solution space by multiple agents. On this basis, this paper proposes a new centralized training and decentralized execution cooperative MARL method, i.e., vehicle-level fairness multi-agent proximity policy optimization (VF-MAPPO). VF-MAPPO leverages a centralized trained global Critic Network to estimate the average vehicle traffic efficiency and vehicle maximum waiting time, and an Actor Network shared by all intersections for decentralized execution, which converts the optimization problem with constraints to an unconstrained optimization objective through the Lagrange multiplier method and adopts proximity policy optimization during training. Additionally, VF-MAPPO incorporates spatial-temporal graph attention in the Critic network to efficiently extract state representations in multi-intersection environments. We qualitatively analyzed the monotonic improvement guarantee of VF-MAPPO. Extensive experimental validation across two real-world and one synthetic scenarios substantiates that VF-MAPPO enhances vehicle-level fairness and maintains average traffic efficiency, surpassing state-of-the-art methods.
Wanting Liu, Chengwei Zhang 0001, Wanqing Fang, Kailing Zhou, Furui Zhan, Qi Wang 0044, Wanli Xue, Rong Chen 0003
IEEE Trans. Intell. Transp. Syst.8
2025 Sequential Decision MARL for Adaptive Traffic Signal Control With Different Intersections Priorities
abstract
Existing multi-agent reinforcement learning (MARL) in adaptive traffic signal control (ATSC) typically models cooperative control of multiple intersections as a cooperative Markov game, optimizing the average traffic efficiency of intersections with the same emphasis. However, it is insufficient to meet the requirements in real ATSC scenarios when all intersections are treated equally. To this end, this work proposes the Captain-Member Markov Game (CM-MG) that considers the different priorities between intersections. CM-MG categorizes intersections into special and ordinary intersections, controlled by captain agents and member agents to optimize the traffic efficiency of local and overall road networks, respectively. The cooperative requirements of CM-MG are achieved through a sequential decision-making principle. Captains have priority in choosing actions, and members make decisions sequentially, following the breadth-first traversal order in the road network after obtaining their precursors’ intentions. Then, a cooperative MARL algorithm, i.e, Sequential Decision Deep Graph Network (GNSD-Light), is proposed to learn the optimal joint policy that meets the learning goals of both captains and members. To be unrestricted by the scales of intersections, GNSD-Light adopts an autoregressive framework where all agents make decisions sequentially in a predetermined order by sharing the same decision model. In addition, to obtain sufficient state representation, two relative position encoding-based spatiotemporal representation modules are designed for GNSD-Light based on the characteristics of ATSC scenarios. Finally, through adequate experiments and qualitative analysis, we have confirmed that our method effectively balances traffic efficiency among both the overall road network and special intersections.
Wanting Liu, Chengwei Zhang 0001, Kailing Zhou, Furui Zhan, Wanli Xue, Rong Chen 0003
IEEE Trans. Intell. Transp. Syst.6
2025 Domain-Division Based Progressive Learning for Source-Free Domain Adaptation
abstract
With growing privacy and portability concerns, source-free domain adaptation requires only a source pre-trained model and an unlabeled target domain, allowing for effective adaptation to the target data. Most existing self-training methods focus on selecting and exploiting samples with reliable predictions, often neglecting others. Inspired by the finding that deep models learn clean samples faster than noisy ones, we propose a domain-division based progressive learning method named DPL. Specifically, our approach consists of two alternating stages, each beginning with the division of the target domain into easy-to-adapt and hard-to-adapt subdomains based on adaptation difficulty, followed by neighborhood-based pseudo label assignment. In stage one, we enhance classification accuracy through uncertainty-aware self-training and alignment of corresponding classes between subdomains. Stage two then applies tailored learning strategies to each subdomain, starting with consistency learning on the easy-to-adapt samples and progressing to utilizing local structural information for the more challenging ones, thereby mining the intrinsic properties of the target data. Extensive experiments on several widely used benchmarks validate the effectiveness of our approach, demonstrating superior performance compared to state-of-the-art methods.
Jing Li 0132, Meng Zhao 0001, Wanli Xue, Qinghua Hu, Shengyong Chen
IEEE Trans. Multim.4
2025 Learning Self-Corrective Network via Adaptive Self-Labeling and Dynamic NMS for High-Performance Long-Term Tracking
abstract
This article presents a self-corrective network-based long-term tracker (SCLT) including a self-modulated tracking reliability evaluator (STRE) and a self-adjusting proposal postprocessor (SPPP). The targets in the long-term sequences often suffer from severe appearance variations. Existing long-term trackers often online update their models to adapt the variations, but the inaccurate tracking results introduce cumulative error into the updated model that may cause severe drift issue. To this end, a robust long-term tracker should have the self-corrective capability that can judge whether the tracking result is reliable or not, and then it is able to recapture the target when severe drift happens caused by serious challenges (e.g., full occlusion and out-of-view). To address the first issue, the STRE designs an effective tracking reliability classifier that is built on a modulation subnetwork. The classifier is trained using the samples with pseudo labels generated by an adaptive self-labeling strategy. The adaptive self-labeling can automatically label the hard negative samples that are often neglected in existing trackers according to the statistical characteristics of target state, and the network modulation mechanism can guide the backbone network to learn more discriminative features without extra training data. To address the second issue, after the STRE has been triggered, the SPPP follows it with a dynamic NMS to recapture the target in time and accurately. In addition, the STRE and the SPPP demonstrate good transportability ability, and their performance is improved when combined with multiple baselines. Compared to the commonly used greedy NMS, the proposed dynamic NMS leverages an adaptive strategy to effectively handle the different conditions of in view and out of view, thereby being able to select the most probable object box that is essential to accurately online update the basic tracker. Extensive evaluations on four large-scale and challenging benchmark datasets including VOT2021LT, OxUvALT, TLP, and LaSOT demonstrate superiority of the proposed SCLT to a variety of state-of-the-art long-term trackers in terms of all measures. Source codes and demos can be found at https://github.com/TJUT-CV/SCLT.
Wanli Xue, Kaihua Zhang 0001, Bo Liu 0005, Chengwei Zhang 0001, Jingen Liu, Shengyong Chen
IEEE Trans. Neural Networks Learn. Syst.2
2025 Dual-stage temporal perception network for continuous sign language recognition
Wanli Xue, Jinlu Sun, Yazhou Wu, Tiantian Yuan, Shengyong Chen
Vis. Comput.2
2024 Dynamical semantic enhancement network for continuous sign language recognition
Suyang Wang, Leming Guo, Wanli Xue
Multim. Syst.3
2024 Enhanced decoupling graph convolution network for skeleton-based action recognition
Wanli Xue
Multim. Tools Appl.3
2024 Hunt-inspired Transformer for visual object tracking
Wanli Xue, Kaihua Zhang 0001, Shengyong Chen
Pattern Recognit.2
2024 Privacy-Preserving Probabilistic Data Encoding for IoT Data Analysis
abstract
The widespread integration of the Internet of Things (IoT) is crucial in advancing sustainable development. IoT service providers actively collect user data for analysis using sophisticated Deep Learning (DL) algorithms. This enables the extraction of valuable insights for business intelligence and improving service quality. However, as these datasets contain sensitive personal information, there is a risk of privacy breaches when DL models are employed. This vulnerability may result in Membership Inference Attacks (MIA), potentially leading to the unauthorized disclosure of highly sensitive data. Therefore, developing an efficient and privacy-preserving data analysis system for IoT is imperative. Recent research has highlighted the effectiveness of utilizing Bloom Filter (BF)-encoding in conjunction with Differential Privacy (DP) for safeguarding privacy during data analysis. Given its attributes of low complexity and high utility, this approach proves effective, particularly in resource-constrained IoT domains. With this in mind, we propose a novel framework for privacy-preserving IoT data analysis based on BF-encoded data. Our research introduces an innovative BF-encoding technique combined with Local Differential Privacy (LDP), capable of efficiently encoding various types of IoT data (such as facial images and smart-meter data) while maintaining privacy when integrated into DL algorithms for downstream analysis. Experimental results demonstrate that our BF-encoded data surpasses the utility of standard BF-encoded data when utilized in DL algorithms for downstream tasks, showcasing an approximate 30% improvement in classification accuracy. Furthermore, we assess the privacy of these DL models against MIA, revealing that attackers can only make random guesses with an accuracy of approximately 50%.
Zakia Zaman, Wanli Xue, Praveen Gauravaram, Wen Hu 0001, Jiaojiao Jiang 0001, Sanjay K. Jha
IEEE Trans. Inf. Forensics Secur.2
2024 Gloss Prior Guided Visual Feature Learning for Continuous Sign Language Recognition
abstract
Continuous sign language recognition (CSLR) is to recognize the glosses in a sign language video. Enhancing the generalization ability of CSLR's visual feature extractor is a worthy area of investigation. In this paper, we model glosses as priors that help to learn more generalizable visual features. Specifically, the signer-invariant gloss feature is extracted by a pre-trained gloss BERT model. Then we design a gloss prior guidance network (GPGN). It contains a novel parallel densely-connected temporal feature extraction (PDC-TFE) module for multi-resolution visual feature extraction. The PDC-TFE captures the complex temporal patterns of the glosses. The pre-trained gloss feature guides the visual feature learning through a cross-modality matching loss. We propose to formulate the cross-modality feature matching into a regularized optimal transport problem, it can be efficiently solved by a variant of the Sinkhorn algorithm. The GPGN parameters are learned by optimizing a weighted sum of the cross-modality matching loss and CTC loss. The experiment results on German and Chinese sign language benchmarks demonstrate that the proposed GPGN achieves competitive performance. The ablation study verifies the effectiveness of several critical components of the GPGN. Furthermore, the proposed pre-trained gloss BERT model and cross-modality matching can be seamlessly integrated into other RGB-cue-based CSLR methods as plug-and-play formulations to enhance the generalization ability of the visual feature extractor.
Leming Guo, Wanli Xue, Bo Liu 0005, Kaihua Zhang 0001, Tiantian Yuan, Dimitris N. Metaxas
IEEE Trans. Image Process.2
2023 Distilling Cross-Temporal Contexts for Continuous Sign Language Recognition
abstract
Continuous sign language recognition (CSLR) aims to recognize glosses in a sign language video. State-of-the-art methods typically have two modules, a spatial perception module and a temporal aggregation module, which are jointly learned end-to-end. Existing results in [9, 20, 25, 36] have indicated that, as the frontal component of the over-all model, the spatial perception module used for spatial feature extraction tends to be insufficiently trained. In this paper, we first conduct empirical studies and show that a shallow temporal aggregation module allows more thor-ough training of the spatial perception module. However, a shallow temporal aggregation module cannot well capture both local and global temporal context information in sign language. To address this dilemma, we propose a cross-temporal context aggregation (CTCA) model. Specifically, we build a dual-path network that contains two branches for perceptions of local temporal context and global temporal context. We further design a cross-context knowledge distil-lation learning objective to aggregate the two types of con-text and the linguistic prior. The knowledge distillation en-ables the resultant one-branch temporal aggregation mod-ule to perceive local-global temporal and semantic context. This shallow temporal perception module structure facili-tates spatial perception module learning. Extensive exper-iments on challenging CSLR benchmarks demonstrate that our method outperforms all state-of-the-art methods.
Leming Guo, Wanli Xue, Qing Guo 0005, Bo Liu 0005, Kaihua Zhang 0001, Tiantian Yuan, Shengyong Chen
CVPR2
2023 'Skimming-Perusal' Detection: A Simple Object Detection Baseline in GigaPixel-level Images
abstract
Object detection has achieved amazing performance in regular-sized images, but with the emergence of gigapixel-level images, even the most advanced object detection methods cannot be directly used to process them quickly and efficiently. Therefore, this paper proposes a simple baseline for gigapixel-level images object detection called Skimming-Perusal Detection (SPDet). The SPDet consists mainly of two parts, a skimming model and a perusal model. The skimming model is based on an efficient global-to-local search strategy to detect possible regions containing objects. Non-object regions are merged through a skimming iterative merging strategy to generate skimming patch candidates. The perusal model adaptive scales the skimming patch candidates guided by the coarse detection of the skimming model. Extensive evaluations on the PANDA dataset demonstrate that the SPDet boosts detection speed on gigapixel-level images by 6× while achieving better performance than a variety of state-of-the-art methods. The source code is released at https://github.com/TJUT-CV/SPDet.
Wanli Xue, Kaihua Zhang 0001, Shengyong Chen
ICME2
2023 Pistis: Replay Attack and Liveness Detection for Gait-Based User Authentication System on Wearable Devices Using Vibration
abstract
Wearable devices-based biometrics has become mainstream in the biometric domain, especially in mobile computing, due to its convenience, flexibility, and potentially high user acceptance. Among various modalities, wearable devices-based gait recognition has been recognized as an effective user authentication method and employed in various applications, such as automated entry systems for home, school, work, vehicles, and automated ticket payment/validation for public transport. However, how secure wearable gait remains an open research question. In this study, we conduct a comprehensive security analysis of the wearable gait. Then, we demonstrate that gait itself is not robust against some attacking methods, such as spoofing or forgery. Therefore, we argue that an anti-spoofing mechanism is important for enhancing the security of wearable gait biometric systems. To this end, we proposed a novel authentication protocol called$Pistis$that embedded gait biometrics and a liveness detection mechanism that is aiming to detect various attacks of gait authentication systems. Our extensive experiments based on 50 subjects demonstrate that$Pistis$is effective in liveness detection and authentication performance enhancement, providing 100% accuracy for human and nonhuman detection, and 99.53% accuracy for user authentication. Pistis can be used as a liveness detection method for wearable devices-based biometrics, significantly for wearable gait.
Hong Jia, Min Wang 0009, Yuezhong Wu, Wanli Xue, Chun Tung Chou, Jiankun Hu, Wen Hu 0001
IEEE Internet Things J.5
2022 Multi-level Temporal Relation Graph for Continuous Sign Language Recognition
Wanli Xue, Leming Guo, Tiantian Yuan, Shengyong Chen
PRCV (3)2
2022 A differential privacy-based classification system for edge computing in IoT
Wanli Xue, Yiran Shen 0001, Chengwen Luo 0001, Weitao Xu, Wen Hu 0001, Aruna Seneviratne
Comput. Commun.1
2022 Advertising Impression Resource Allocation Strategy with Multi-Level Budget Constraint DQN in Real-Time Bidding
Chengwei Zhang 0001, Kangjie Zheng, Wanli Xue, Tianpei Yang, Dou An, Yongqi Pi, Rong Chen 0003
Neurocomputing4
2022 PrivGait: An Energy-Harvesting-Based Privacy-Preserving User-Identification System by Gait Analysis
abstract
Smart space has emerged as a new paradigm that combines sensing, communication, and artificial intelligence technologies to offer various customized services. A fundamental requirement of these services is person identification. Although a variety of person-identification approaches has been proposed, they suffer from several limitations in practical applications, such as low energy efficiency, accuracy degradation, and privacy issue. This article proposes an energy-harvesting-based privacy-preserving gait recognition scheme for smart space, which is named PrivGait. In PrivGait, we extract discriminative features from 1-D gait signal and design an attention-based long short-term memory (LSTM) network to classify different people. Moreover, we leverage a novel Bloom filter-based privacy-preserving technique to address the privacy leakage problem. To demonstrate the feasibility of PrivGait, we design a proof-of-concept prototype using off-the-shelf energy-harvesting hardware. Extensive evaluation results show that the proposed scheme outperforms state of the art by 6%–10% and incurs low system cost while preserving user’s privacy.
Weitao Xu, Wanli Xue, Guohao Lan, Xingyu Feng 0001, Bo Wei 0003, Chengwen Luo 0001, Wei Li 0058, Albert Y. Zomaya
IEEE Internet Things J.2
2022 DIMY: Enabling privacy-preserving contact tracing
Regio A. Michelin, Wanli Xue, Guntur D. Putra, Sushmita Ruj, Salil S. Kanhere, Sanjay K. Jha
J. Netw. Comput. Appl.3
2022 DARTSRepair: Core-failure-set guided DARTS for network robustness to common corruptions
Xuhong Ren, Jianlang Chen, Felix Juefei-Xu, Wanli Xue, Qing Guo 0005, Lei Ma 0003, Jianjun Zhao 0001, Shengyong Chen
Pattern Recognit.4
2022 Stable Linear Structures and Seam Measurements for Parallax Image Stitching
abstract
Parallax tolerance is a fundamental problem in image stitching. To solve this problem, increasing effort has been devoted to the spatially varying and seam cutting. However, there still exist some issues that need to be adequately addressed. First, the implementation of the spatially varying warping requires a restricted premise that the overlapping region can be aligned: if this premise is not met, ghosting caused by misalignment will emerge. In addition, the spatially varying warping may cause distortion due to the issue of inconsistent homographies. Second, conventional seam cutting will lead to objects being cropped and duplicated. Therefore, in this paper, we propose a stable framework for stitching images: the framework consists of a uniform linear structure model that is able to mitigate the distortion of projection and perspective, while preserving the structures of objects in the non-overlapping region; and a stable hybrid actor-critic that estimates stable seam measurements in the overlapping region to diminish the parallax. Comparison experiments show that the proposed method is superior to some conventional methods with respect to mitigating ghosting and preserving structure.
Wanli Xue, Weilun Xie, Yao Zhang 0021, Shengyong Chen
IEEE Trans. Circuits Syst. Video Technol.1
2022 Object-Aware Ghost Identification and Elimination for Dynamic Scene Mosaic
abstract
Composite ghost is a common phenomenon that widely exists in dynamic scene image mosaic and significantly affects the naturalness of mosaic. To remove the ghost effectively and produce visually natural mosaic, we propose a novel image mosaic method by jointly identifying composite ghost and eliminating ghost regions without distorting, splitting, and duplicating objects. Specifically, our main contributions are three-fold:First, we propose themotion-awarecomposite ghost identification to localize the potential composite ghosts in the mosaic region (i.e., overlapping area between two images to be stitched) by detecting the salient-moving objects in two stitched images.Second, we design theobject-awarealternative region selection strategy to produce ghostless regions that can replace the localized composite ghosts while avoiding object distortion, object separation, and object repetition.Third, we realize theimage interpolation-basedcomposite ghost elimination that can generate natural stitched image by eliminating the composite ghost of the initial blending result with the selected image source. We validate the proposed method on challenging datasets and show that our method outperform the state-of-the-art methods.
Zhe Zhang 0039, Xuhong Ren, Wanli Xue, Chengwei Zhang 0001, Qing Guo 0005, Shengyong Chen
IEEE Trans. Circuits Syst. Video Technol.3
2022 Neighborhood Cooperative Multiagent Reinforcement Learning for Adaptive Traffic Signal Control in Epidemic Regions
abstract
Nowadays, multiagent reinforcement learning (MARL) have shared significant advances in the adaptive traffic signal control (ATSC) problems. For most of the researches, agents are all isomorphic, which disregards the situation in which isomerous intersections cooperative together in a real ATSC scenario, especially in epidemic regions where different intersections have quite different levels of importance. To this end, this paper models the ATSC problem as a networked Markov game (NMG), in which agents take into account information, including traffic conditions of it and its connected neighbors. A cooperative MARL framework named neighborhood cooperative hysteretic DQN (NC-HDQN) is proposed. Specifically, for each NC-HDQN agent in the NMG, first, the framework analyses correlation degrees with their connected neighbors and weighs observations and rewards by these correlations. Second, NC-HDQN agents independently optimize their strategies on the weighted information using hysteretic DQN (HDQN), which is designed to learn optimal joint strategies in cooperative multiagent games. Third, a rule-based NC-HDQN method and a Pearson correlation coefficient based NC-HDQN method, i.e., empirical NC-HDQN (ENC-HDQN) and Pearson NC-HDQN (PNC-HDQN), respectively, are designed. The first method maps the correlation degree between two connected agents according to vehicle numbers on roads between the two agents. In contrast, the second method uses the Pearson correlation coefficient to calculate the correlation degree adaptively. Our methods are empirically evaluated in both a synthetic scenario and two real-world traffic scenarios and give better performances in almost every standard test metric for ATSC.
Chengwei Zhang 0001, Wanli Xue, Xiaofei Xie, Tianpei Yang, Rong Chen 0003
IEEE Trans. Intell. Transp. Syst.4
2021 MARL for Traffic Signal Control in Scenarios with Different Intersection Importance
Liguang Luan, Wanqing Fang, Chengwei Zhang 0001, Wanli Xue, Rong Chen 0003, Chen Sang
DAI5
2021 Deepmix: Online Auto Data Augmentation for Robust Visual Object Tracking
abstract
Online updating of the object model via samples from historical frames is of great importance for accurate visual object tracking. Recent works mainly focus on constructing effective and efficient updating methods while neglecting the training samples for learning discriminative object models, which is also a key part of a learning problem. In this paper, we propose the DeepMix that takes historical samples’ embeddings as input and generates augmented embeddings online, enhancing the state-of-the-art online learning methods for visual object tracking. More specifically, we first propose the online data augmentation for tracking that online augments the historical samples through object-aware filtering. Then, we propose MixNet which is an offline trained network for performing online data augmentation within one-step, enhancing the tracking accuracy while preserving high speeds of the state-of-the-art online learning methods. The extensive experiments on three different tracking frameworks, i.e., DiMP, DSiam, and SiamRPN++, and three large-scale and challenging datasets, i.e., OTB-2015, LaSOT, and VOT, demonstrate the effectiveness and advantages of the proposed method.
Ziyi Cheng, Xuhong Ren, Felix Juefei-Xu, Wanli Xue, Qing Guo 0005, Lei Ma 0003, Jianjun Zhao 0001
ICME4
2021 InaudibleKey: Generic Inaudible Acoustic Signal based Key Agreement Protocol for Mobile Devices
abstract
Secure Device-to-Device (D2D) communication is becoming increasingly important with the ever-growing number of Internet-of-Things (IoT) devices in our daily life. To achieve secure D2D communication, the key agreement between different IoT devices without any prior knowledge is becoming desirable. Although various approaches have been proposed in the literature, they suffer from a number of limitations, such as low key generation rate and short pairing distance. In this paper, we present InaudibleKey, an inaudible acoustic signal based key generation protocol for mobile devices. Based on acoustic channel reciprocity, InaudibleKey exploits the acoustic channel frequency response of two legitimate devices as a common secret to generating keys. InaudibleKey employs several novel technologies to significantly improve its performance. We conduct extensive experiments to evaluate the proposed system in different real environments. Compared to state-of-the-art works, InaudibleKey improves key generation rate by 3 times, extends pairing distance by 3.2 times, and reduces information reconciliation counts by 2.5 times. Security analysis demonstrates that InaudibleKey is resilient to a number of malicious attacks. We also implement InaudibleKey on modern smartphones and resource-limited IoT devices. Results show that it is energy-efficient and can run on both powerful and resource-limited IoT devices without incurring excessive resource consumption.
Weitao Xu, Zhenjiang Li 0001, Wanli Xue, Xiaotong Yu, Bo Wei 0003, Jia Wang 0008, Chengwen Luo 0001, Wei Li 0058, Albert Y. Zomaya
IPSN3
2021 Towards a Compressive-Sensing-Based Lightweight Encryption Scheme for the Internet of Things
abstract
Internet of Things (IoT) is flourishing and has penetrated deeply into people's daily life. With the seamless connection to the physical world, IoT provides tremendous opportunities to a wide range of applications. However, potential risks exist when the IoT system collects sensor data and uploads it to the Cloud. The leakage of private data can be severe with curious database administrator or malicious hackers who compromise the Cloud. In this work, we propose Kryptein, a compressive-sensing-based lightweight encryption scheme for Cloud-enabled IoT systems to secure the interaction between the IoT devices and the Cloud. Kryptein supports random compressed encryption, statistical computation over cipher, and accurate raw data decryption. According to our evaluation based on two real datasets, Kryptein provides strong protection to the data. It is 250 times faster than other state-of-the-art systems and incurs 120 times less energy consumption. The performance of Kryptein is also measured on off-the-shelf IoT devices, and the result shows Kryptein can run efficiently on IoT devices. After comparing with other state-of-the-art lightweight ciphers on IoT (Simon and Speck), IoT system with Kryptein is expected to have a much more longevity with about 35 percent extended lifetime. Further, experiments illustrated IoT data variance will not affect Kryptein's accuracy in a long term usage, and Krpytein is also able to support basic analytics tasks like machine learning (e.g., classification).
Wanli Xue, Chengwen Luo 0001, Yiran Shen 0001, Rajib Rana, Guohao Lan, Sanjay K. Jha, Aruna Seneviratne, Wen Hu 0001
IEEE Trans. Mob. Comput.1
2020 SPARK: Spatial-Aware Online Incremental Attack Against Visual Tracking
Qing Guo 0005, Xiaofei Xie, Felix Juefei-Xu, Lei Ma 0003, Zhongguo Li, Wanli Xue, Wei Feng 0005, Yang Liu 0003
ECCV (25)6
2020 Few-Shot Guided Mix for DNN Repairing
abstract
Although deep neural networks (DNNs) achieve rather high performance in many cutting-edge applications (e.g., autonomous driving, medical diagnose), their trustworthiness on real-world scenarios still posts concerns, where some specific failure examples are often encountered during the real-world operational environment. With the limited failure examples collected during the practical operation, how to effectively leverage such failure cases to repair and enhance DNN so as to generalize to more potentially suspicious samples is challenging, but of great importance. In this paper, we formulate the failure-data-driven DNN repairing as a data augmentation problem, and design a novel augmentation-based repairing method, which to the best extent leverages limited failure cases. To realize the DNN repairing effects that generalize to specific failure examples, we originally propose few-shot guided mix (FSGMix) that augments training data with the guidance of failure examples. As a result, our method is able to achieve high generalization to the collected failure examples and other similar suspicious data. The preliminary evaluation on CIFAR-10 dataset demonstrates the potential of our proposed technique, which automatically learns to resolve the potential failure patterns in the DNN operational environment.
Xuhong Ren, Hua Qi, Felix Juefei-Xu, Zhuo Li 0013, Wanli Xue, Lei Ma 0003, Jianjun Zhao 0001
ICSME6
2020 Inaudible acoustic signal based key agreement system for IoT devices: poster abstract
abstract
Secure Device-to-Device (D2D) communication is becoming increasingly important with the ever-growing number of Internet-of-Things (IoT) devices in our daily life. To achieve secure D2D communication, the key agreement between different IoT devices without any prior knowledge is becoming desirable. Although various approaches have been proposed in the literature, they suffer from a number of limitations, such as low key generation rate and short pairing distance. In this paper, we present an inaudible acoustic signal based key generation protocol for mobile devices. Based on acoustic channel reciprocity, our system exploits channel frequency response of two legitimate devices as a common secret to generate keys. Extensive experiments are conducted to evaluate the proposed system in different real environments. Evaluation results show that the proposed system can generate the same secret key for two mobile devices with high probability.
Weitao Xu, Zhenjiang Li 0001, Wanli Xue, Xiaotong Yu, Jia Wang 0008, Chengwen Luo 0001, Wei Li 0058, Albert Y. Zomaya
SenSys3
2020 A multi-view CNN-based acoustic classification system for automatic animal species identification
Weitao Xu, Xiang Zhang 0012, Lina Yao 0001, Wanli Xue, Bo Wei 0003
Ad Hoc Networks4
2020 Sequence Data Matching and Beyond: New Privacy-Preserving Primitives Based on Bloom Filters
abstract
Bloom filter encoding has widely been used as an efficient masking technique for privacy-preserving matching functions. The existing matching techniques, however, are limited to relatively simple types such as string, categorical and signal numerical values. In this paper, we propose a new scheme that significantly extends the class of matching primitives that are based on privacy-preserving Bloom filter mechanism. These primitives include sequence data matching and popular distance-based machine learning algorithms such as KNN and SVM. Our scheme hash-maps a sequence data vector into the Bloom filter space while checking the similarity of the data points efficiently with negligible utility loss by adding a timestamp (bit) for each element in the data represented with its neighboring values. Furthermore, it includes a Laplace-like perturbation method on the constructed Bloom filters to address the weakness of deterministic probability led by encoding techniques. As a result, the proposed work guarantee the private data records are difficult to be discriminated due to collisions and differential privacy. The experimental results on three real-scenario based datasets illustrate that our method can achieve a significantly better trade-off between utility and privacy than the state-of-the-art differential privacy-based method by adding Laplace noise to the data directly.
Wanli Xue, Dinusha Vatsalan, Wen Hu 0001, Aruna Seneviratne
IEEE Trans. Inf. Forensics Secur.1
2019 SA-IGA: a multiagent reinforcement learning method towards socially optimal outcomes
Chengwei Zhang 0001, Xiaohong Li 0001, Jianye Hao, Siqi Chen 0001, Karl Tuyls, Wanli Xue, Zhiyong Feng 0002
Auton. Agents Multi Agent Syst.6
2019 Predictable Privacy-Preserving Mobile Crowd Sensing: A Tale of Two Roles
abstract
The rise of mobile crowd sensing has brought privacy issues into a sharp view. In this paper, our goal is to achieve the predictable privacy-preserving mobile crowd sensing, which we envision to have the capability to quantify the privacy protections, and simultaneously allowing application users to predict the utility loss at the same time. TheSalusalgorithm is first proposed to protect the private data against the data reconstruction attacks. To understand privacy protection, we quantify the privacy risks in terms of private data leakage under reconstruction attacks. To predict the utility, we provide accurate utility predictions for various crowd sensing applications using Salus. The risk assessments can be generally applied to different type of sensors on the mobile platform, and the utility prediction can also be used to support various applications that use data aggregators such as average, histogram, and classifiers. Finally, we propose and implement the$P^{3}$application framework. Both measurement results using online datasets and real-world case studies show that the$P^{3}$provides accurate risk assessments and utility estimations, which makes it a promising framework to support future privacy-preserving mobilecrowd sensing applications.
Chengwen Luo 0001, Wanli Xue, Yiran Shen 0001, Jianqiang Li 0001, Wen Hu 0001, Alex X. Liu
IEEE/ACM Trans. Netw.3
2018 HealCam: Energy-efficient and privacy-preserving human vital cycles monitoring on camera-enabled smart devices
Qing Yang 0009, Yiran Shen 0001, Fengyuan Yang 0001, Jianpei Zhang, Wanli Xue, Hongkai Wen 0001
Comput. Networks5
2018 Robust Visual Tracking via Multi-Scale Spatio-Temporal Context Learning
abstract
In order to tackle the incomplete and inaccurate of the samples in most tracking-by-detection algorithms, this paper presents an object tracking algorithm, termed as multi-scale spatio-temporal context (MSTC) learning tracking. MSTC collaboratively explores three different types of spatio-temporal contexts, named the long-term historical targets, the medium-term stable scene (i.e., a short continuous and stable video sequence), and the short-term overall samples to improve the tracking efficiency and reduce the drift phenomenon. Different from conventional multi-timescale tracking paradigm that chooses samples in a fixed manner, MSTC formulates a low-dimensional representation named fast perceptual hash algorithm to update long-term historical targets and the medium-term stable scene dynamically with image similarity. MSTC also differs from most tracking-by-detection algorithms that label samples as positive or negative, it investigates a fusion salient sample detection to fuse weights of the samples not only by the distance information, but also by the visual spatial attention, such as color, intensity, and texture. Numerous experimental evaluations with most state-of-the-art algorithms on the standard 50 video benchmark demonstrate the superiority of the proposed algorithm.
Wanli Xue, Chao Xu 0003, Zhiyong Feng 0002
IEEE Trans. Circuits Syst. Video Technol.1
2017 Kryptein: a compressive-sensing-based encryption scheme for the internet of things
abstract
Internet of Things (IoT) is flourishing and has penetrated deeply into people's daily life. With the seamless connection to the physical world, IoT provides tremendous opportunities to a wide range of applications. However, potential risks exist when the IoT system collects sensor data and uploads it to the cloud. The leakage of private data can be severe with curious database administrator or malicious hackers who compromise the cloud. In this work, we propose Kryptein, a compressive-sensing-based encryption scheme for cloud-enabled IoT systems to secure the interaction between the IoT devices and the cloud. Kryptein supports random compressed encryption, statistical decryption, and accurate raw data decryption. According to our evaluation based on two real datasets, Kryptein provides strong protection to the data. It is 250 times faster than other state-of-the-art systems and incurs 120 times less energy consumption. The performance of Kryptein is also measured on off-the-shelf IoT devices, and the result shows Kryptein can run efficiently on IoT devices.
Wanli Xue, Chengwen Luo 0001, Guohao Lan, Rajib Rana, Wen Hu 0001, Aruna Seneviratne
IPSN1
2016 CScrypt: A Compressive-Sensing-Based Encryption Engine for the Internet of Things: Demo Abstract
abstract
Internet of Things (IoT) have been connecting the physical world seamlessly and provides tremendous opportunities to a wide range of applications. However, potential risks exist when IoT system collects local sensor data and uploads to the Cloud. The private data leakage can be severe with curious database administrator or malicious hackers who compromise the Cloud. In this demo, we solve this problem of guaranteeing the user data privacy and security using compressive sensing based cryptographic method. We present CScrypt, a compressive-sensing-based encryption engine for the Cloud-enabled IoT systems to secure the interaction between the IoT devices and the Cloud. Our system exploits the fact that each individual's biometric data can be trained to a unique dictionary which can be used as an encryption key meanwhile to compress the original data. We will demonstrate a functioning prototype of our system using live data stream when attending the conference.
Wanli Xue, Chengwen Luo 0001, Rajib Rana, Wen Hu 0001, Aruna Seneviratne
SenSys1
2014 Arduface: An Embedded System Analysis Tool
Wanli Xue, Hyunsuk Chung, Soyeon Caren Han, Yang Sok Kim, Byeong Ho Kang 0001
PRICAI1