Xiaoyue Wan

dblp:203/9576 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Diffusion-based Personalized Pathology Disentanglement for Impaired Gait Analysis
abstract
In the context of global population aging, the prevalence of neurodegenerative diseases is rapidly increasing. Vision-based impaired gait analysis emerges as a promising alternative for automatic and non-invasive diagnosis. While prior efforts have advanced either accuracy or interpretability of gait analysis, few have effectively addressed both aspects in a unified framework. To bridge this gap, we propose DPPD, a Diffusion-based Personalized Pathology Disentanglement model that jointly performs quantitative gait scoring, dementia subtyping, and qualitative anomaly highlighting. Motivated by the observation that pathological gait features exhibit stronger inter-class separability across different gait severity than raw features, DPPD is proposed based on the subject-specific pathology disentanglement perspective. Specifically, it comprises three key components: (1) a 3DmotionBERT for encoding gait representation from 3D human pose sequences estimated, (2) a latent diffusion-based Gait Denoiser for generating personalized normal gait features, and (3) a Dual Pathology Disentanglement mechanism that captures both static pose and dynamic motion pathological representation from the residual between raw and normal gait features. These disentangled pathologies further enable quantitative classification and qualitative anomaly highlighting. Experiments on the PDGait and 3DGait datasets demonstrate that DPPD outperforms state-of-the-art methods in classification accuracy while providing reliable and interpretable visualizations of gait anomalies.
Xiaoyue Wan, Xu Zhao 0001
AAAI1
2025 RSB-Pose: Robust Short-Baseline Binocular 3D Human Pose Estimation With Occlusion Handling
abstract
In the domain of 3D Human Pose Estimation, which finds widespread daily applications, the requirement for convenient acquisition equipment continues to grow. To satisfy this demand, we focus on a short-baseline binocular setup that offers both portability and a geometric measurement capability that significantly reduces depth ambiguity. However, as the binocular baseline shortens, two serious challenges emerge: first, the robustness of 3D reconstruction against 2D errors deteriorates; second, occlusion reoccurs frequently due to the limited visual differences between two views. To address the first challenge, we propose the Stereo Co-Keypoints Estimation module to improve the view consistency of 2D keypoints and enhance the 3D robustness. In this module, the disparity is utilized to represent the correspondence of binocular 2D points, and the Stereo Volume Feature (SVF) is introduced to contain binocular features across different disparities. Through the regression of SVF, two-view 2D keypoints are simultaneously estimated in a collaborative way which restricts their view consistency. Furthermore, to deal with occlusions, a Pre-trained Pose Transformer module is introduced. Through this module, 3D poses are refined by perceiving pose coherence, a representation of joint correlations. This perception is injected by the Pose Transformer network and learned through a pre-training task that recovers iterative masked joints. Comprehensive experiments on H36M and MHAD datasets validate the effectiveness of our approach in the short-baseline binocular 3D Human Pose Estimation and occlusion handling.
Xiaoyue Wan, Zhuo Chen 0028, Xu Zhao 0001
IEEE Trans. Image Process.1
2024 FSGait: Fine-Grained Self-supervised Gait Abnormality Detection
Bingzhi Duan, Xiaoyue Wan, Xu Zhao 0001
ACCV (6)2
2024 Dual-Diffusion for Binocular 3D Human Pose Estimation
abstract
Binocular 3D human pose estimation (HPE), reconstructing a 3D pose from 2D poses of two views, offers practical advantages by combining multiview geometry with the convenience of a monocular setup. However, compared to a multiview setup, the reduction in the number of cameras increases uncertainty in 3D reconstruction. To address this issue, we leverage the diffusion model, which has shown success in monocular 3D HPE by recovering 3D poses from noisy data with high uncertainty. Yet, the uncertainty distribution of initial 3D poses remains unknown. Considering that 3D errors stem from 2D errors within geometric constraints, we recognize that the uncertainties of 3D and 2D are integrated in a binocular configuration, with the initial 2D uncertainty being well-defined. Based on this insight, we propose Dual-Diffusion specifically for Binocular 3D HPE, simultaneously denoising the uncertainties in 2D and 3D, and recovering plausible and accurate results. Additionally, we introduce Z-embedding as an additional condition for denoising and implement baseline-width-related pose normalization to enhance the model flexibility for various baseline settings. This is crucial as 3D error influence factors encompass depth and baseline width. Extensive experiments validate the effectiveness of our Dual-Diffusion in 2D refinement and 3D estimation. The code and models are available at https://github.com/sherrywan/Dual-Diffusion.
Xiaoyue Wan, Zhuo Chen 0028, Bingzhi Duan, Xu Zhao 0001
NeurIPS1
2024 Joint-Limb Compound Triangulation With Co-Fixing for Stereoscopic Human Pose Estimation
abstract
As a special subset of multi-view settings for 3D human pose estimation, stereoscopic settings show promising applications in practice since they are not ill-posed but could be as mobile as monocular ones. However, when there are only two views, the problems of occlusions and “double counting” (ambiguity between symmetric joints) pose greater challenges that are not addressed by previous approaches. On this concern, we propose a novel framework to detect limb orientations in field form and incorporate them explicitly with joint features. Two modules are proposed to realize the fusion. At 3D level, we designcompound triangulationas an explicit module that produces the optimal pose using 2D joint locations and limb orientations. The module is derived from reformulating triangulation in 3D space, and expanding it with the optimization of limb orientations. At 2D level, we propose a parameter-free module namedco-fixingto enable joint and limb features to fix each other to alleviate the impact of “double counting.” Features from both parts are first used to infer each other via simple convolutions and then fixed by the inferred ones respectively. We test our method on two public benchmarks, Human3.6M and Total Capture, and our method achieves state-of-the-art performance on stereoscopic settings and comparable results on common 4-view benchmarks.
Zhuo Chen 0028, Xiaoyue Wan, Yiming Bao, Xu Zhao 0001
IEEE Trans. Multim.2
2023 View consistency aware holistic triangulation for 3D human pose estimation
abstract
The rapid development of multi-view 3D human pose estimation (HPE) is attributed to the maturation of monocular 2D HPE and the geometry of 3D reconstruction. However, 2D detection outliers in occluded views due to neglect of view consistency, and 3D implausible poses due to lack of pose coherence, remain challenges. To solve this, we introduce a Multi-View Fusion module to refine 2D results by establishing view correlations. Then, Holistic Triangulation is proposed to infer the whole pose as an entirety, and anatomy prior is injected to maintain the pose coherence and improve the plausibility. Anatomy prior is extracted by PCA whose input is skeletal structure features, which can factor out global context and joint-by-joint relationship from abstract to concrete. Benefiting from the closed-form solution, the whole framework is trained end-to-end. Our method outperforms the state of the art in both precision and plausibility which is assessed by a new metric.
Xiaoyue Wan, Zhuo Chen 0028, Xu Zhao 0001
Comput. Vis. Image Underst.1
2022 Structural Triangulation: A Closed-Form Solution to Constrained 3D Human Pose Estimation
Zhuo Chen 0028, Xu Zhao 0001, Xiaoyue Wan
ECCV (5)3
2020 Reinforcement Learning-Based Mobile Offloading for Edge Computing Against Jamming and Interference
abstract
Mobile edge computing systems help improve the performance of computational-intensive applications on mobile devices and have to resist jamming attacks and heavy interference. In this paper, we present a reinforcement learning based mobile offloading scheme for edge computing against jamming attacks and interference, which uses safe reinforcement learning to avoid choosing the risky offloading policy that fails to meet the computational latency requirements of the tasks. This scheme enables the mobile device to choose the edge device, the transmit power and the offloading rate to improve its utility including the sharing gain, the computational latency, the energy consumption and the signal-to-interference-plus-noise ratio of the offloading signals without knowing the task generation model, the edge computing model, and the jamming/interference model. We also design a deep reinforcement learning based mobile offloading for edge computing that uses an actor network to choose the offloading policy and a critic network to update the actor network weights to improve the computational performance. We discuss the computational complexity and provide the performance bound that consists of the computational latency and the energy consumption based on the Nash equilibrium of the mobile offloading game. Simulation results show that this scheme can reduce the computational latency and save energy consumption.
Liang Xiao 0003, Xiaozhen Lu, Tangwei Xu, Xiaoyue Wan, Wen Ji 0003, Yanyong Zhang
IEEE Trans. Commun.4
2020 Reinforcement Learning-Based Downlink Interference Control for Ultra-Dense Small Cells
abstract
The dense deployment of small cells in 5G cellular networks raises the issue of controlling downlink inter-cell interference under time-varying channel states. In this paper, we propose a reinforcement learning based power control scheme to suppress downlink inter-cell interference and save energy for ultra-dense small cells. This scheme enables base stations to schedule the downlink transmit power without knowing the interference distribution and the channel states of the neighboring small cells. A deep reinforcement learning based interference control algorithm is designed to further accelerate learning for ultra-dense small cells with a large number of active users. Analytical convergence performance bounds including throughput, energy consumption, inter-cell interference, and the utility of base stations are provided and the computational complexity of our proposed scheme is discussed. Simulation results show that this scheme optimizes the downlink interference control performance after sufficient power control instances and significantly increases the network throughput with less energy consumption compared with a benchmark scheme.
Liang Xiao 0003, Hailu Zhang, Yilin Xiao 0001, Xiaoyue Wan, Sicong Liu 0002, Li-Chun Wang 0001, H. Vincent Poor
IEEE Trans. Wirel. Commun.4
2019 Reinforcement Learning with Safe Exploration for Network Security
abstract
Safe reinforcement learning is important for the safety critical applications especially network security, as the exploration of some dangerous actions can result in huge short-term losses such as network failure or large scale privacy leakage. In this paper, we propose a reinforcement learning algorithm with safe exploration and uses transfer learning to reduce the initial random exploration. A blacklist is maintained to record the most dangerous state-action pairs as a safety constraint. A safe deep reinforcement learning version uses a convolutional neural network to estimate the risk levels and thus further improves the safety of the exploration and accelerates the learning speed for the learning agent. As a case study, the proposed reinforcement learning with safe exploration is applied in the anti-jamming robot communications. Experimental results show that the proposed algorithms can improve the jamming resistance of the robot and reduce the outage rate to enter the most dangerous states compared with the benchmark algorithms.
Canhuang Dai, Liang Xiao 0003, Xiaoyue Wan, Ye Chen 0011
ICASSP3
2019 Learning-Based Privacy-Aware Offloading for Healthcare IoT With Energy Harvesting
abstract
Mobile edge computing helps healthcare Internet of Things (IoT) devices with energy harvesting provide satisfactory quality of experiences for computation intensive applications. We propose a reinforcement learning (RL)-based privacy-aware offloading scheme to help healthcare IoT devices protect both the user location privacy and the usage pattern privacy. More specifically, this scheme enables a healthcare IoT device to choose the offloading rate that improves the computation performance, protects user privacy, and saves the energy of the IoT device without being aware of the privacy leakage, IoT energy consumption, and edge computation model. This scheme uses transfer learning to reduce the random exploration at the initial learning process and applies a Dyna architecture that provides simulated offloading experiences to accelerate the learning process. A post-decision state learning method uses the known channel state model to further improve the offloading performance. We provide the performance bound of this scheme regarding the privacy level, the energy consumption, and the computation latency for three typical healthcare IoT offloading scenarios. Simulation results show that this scheme can reduce the computation latency, save the energy consumption, and improve the privacy level of the healthcare IoT device compared with the benchmark scheme.
Minghui Min, Xiaoyue Wan, Liang Xiao 0003, Ye Chen 0011, Minghua Xia, Di Wu 0001, Huaiyu Dai
IEEE Internet Things J.2
2018 Learning-Based Rogue Edge Detection in VANETs with Ambient Radio Signals
abstract
Edge computing for mobile devices in vehicular ad hoc networks (VANETs) has to address rogue edge attacks, in which a rogue edge node claims to be the serving edge in the vehicle to steal user secrets and help launch other attacks such as man-in-the-middle attacks. Rogue edge detection in VANETs is more challenging than the spoofing detection in indoor wireless networks due to the high mobility of onboard units (OBUs) and the large-scale network infrastructure with roadside units (RSUs). In this paper, we propose a physical (PHY)- layer rogue edge detection scheme for VANETs according to the shared ambient radio signals observed during the same moving trace of the mobile device and the serving edge in the same vehicle. In this scheme, the edge node under test has to send the physical properties of the ambient radio signals, including the received signal strength indicator (RSSI) of the ambient signals with the corresponding source media access control (MAC) address during a given time slot. The mobile device can choose to compare the received ambient signal properties and its own record or apply the RSSI of the received signals to detect rogue edge attacks, and determines test threshold in the detection. We adopt a reinforcement learning technique to enable the mobile device to achieve the optimal detection policy in the dynamic VANET without being aware of the VANET model and the attack model. Simulation results show that the Q-learning based detection scheme can significantly reduce the detection error rate and increase the utility compared with existing schemes.
Xiaozhen Lu, Xiaoyue Wan, Liang Xiao 0003, Yuliang Tang, Weihua Zhuang
ICC2
2018 PHY-Layer Authentication With Multiple Landmarks With Reduced Overhead
abstract
Physical (PHY)-layer authentication systems can exploit channel state information of radio transmitters to detect spoofing attacks in wireless networks. The use of multiple landmarks each with multiple antennas enhances the spatial resolution of radio transmitters, and thus improves the spoofing detection accuracy of PHY-layer authentication. Unlike most existing PHY-layer authentication schemes that apply hypothesis tests and rely on the known radio channel model, we propose a logistic regression-based authentication to remove the assumption on the known channel model, and thus be applicable to more generic wireless networks. The Frank-Wolfe algorithm is used to estimate the parameters of the logistic regression model, in which the convex problem under a ℓ1-norm constraint is solved for weight sparsity to avoid over-fitting in the learning process. We design a distributed Frank-Wolfe-based PHY-layer authentication to further reduce the communication overhead between the landmarks and the security agent. Then, we construct an incremental aggregated gradient-based scheme to provide online authentication with a higher accuracy and lower computation overhead. Simulation and experimental results validate the accuracy of the proposed authentication schemes, and show the reduced communication and computation overheads.
Liang Xiao 0003, Xiaoyue Wan, Zhu Han 0001
IEEE Trans. Wirel. Commun.2
2017 Reinforcement Learning Based Mobile Offloading for Cloud-Based Malware Detection
abstract
Cloud-based malware detection improves the detection performance for mobile devices that offload their malware detection tasks to security servers with much larger malware database and powerful computational resources. In this paper, we investigate the competition of the radio transmission bandwidths and the data sharing of the security server in the dynamic malware detection game, in which each mobile device chooses its offloading rate of the application traces to the security server. As the Q-learning technique has a slow learning rate in the game with high dimension, we have designed a mobile malware detection based on hotbooting-Q techniques, which initiates the quality values based on the malware detection experience. We propose an offloading strategy based on deep Q-network technique with a deep convolutional neural network to further improve the detection speed, the detection accuracy, and the utility. Preliminary simulation results verify the detection gain of the scheme compared with the Q- learning based strategy.
Xiaoyue Wan, Geyi Sheng, Yanda Li, Liang Xiao 0003, Xiaojiang Du
GLOBECOM1
2017 FHY-layer authentication with multiple landmarks with reduced communication overhead
abstract
In this paper, we propose a physical (PHY)-layer authentication system that exploits the channel state information of radio transmitters to detect spoofing attacks in wireless networks. By using multiple landmarks and multiple antennas in the channel estimation, this authentication system enhances the spatial resolution of the channel information and thus improves the spoofing detection accuracy. Unlike most existing hypothesis test based PHY-layer authentication schemes that rely on the known radio channel model, our proposed authentication system uses the logistic regression to remove the assumption on the known channel model and is applicable to more generic wireless networks. The Frank-Wolfe algorithm is then used to estimate the parameters of the logistic regression model, which solves the convex problem under a ℓ1-norm constraint for weight sparsity to avoid over-fitting in the learning process. The distributed Frank-Wolfe algorithm can further reduce the communication overhead between the landmarks and the security agent while keeping the spoofing detection accuracy. Simulation results can validate the accuracy of the proposed PHY-layer authentication with multiple landmarks and show the performance gain regarding the overall communication overhead.
Xiaoyue Wan, Liang Xiao 0003, Qiangda Li, Zhu Han 0001
ICC1