Wenbo Ding 0001

dblp:22/10126-1 · DBLP profile ↗
← Back
82ranked-venue papers
8as first author
59since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 30 since 2021Computer networks · 29 · 6 first-author · 13 since 2021Systems, architecture and hardware · 22 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
YearPublicationVenuePosition
2026 What You See Is What You Reach: Towards Spatial Navigation with High-Level Human Instructions
abstract
Embodied navigation is a fundamental capability that enables embodied agents to effectively interact with the physical world in various complex environments. However, a significant gap remains between current embodied navigation tasks and real-world requirements, as existing methods often struggle to integrate high-level human instructions with spatial understanding. To address this gap, we propose a new task of embodied navigation called spatial navigation, which encompasses two key components: spatial object navigation (SpON) for object-specific guidance and spatial area navigation (SpAN) for navigating to designated areas. Specifically, SpON guides agents to specific objects by leveraging spatial relationships and contextual understanding, while SpAN focuses on navigating to defined areas within complex environments. Together, these components significantly enhance agents’ navigation capabilities, enabling more effective interactions in real-world scenarios. To support this task, we have generated a spatial navigation dataset consisting of 10K trajectories within the simulator. This dataset includes high-level human instructions, detailed observations, and corresponding navigation actions, providing a comprehensive resource to enhance agent training and performance. Building on the spatial navigation dataset, we introduce SpNav, a hierarchical navigation framework. Specifically, SpNav employs vision-language model (VLM) to interpret high-level human instructions and accurately identify goal objects or areas within the observation range, achieving precise point-to-point navigation using a map and enhancing the agent’s ability to oper- ate effectively in complex environments by bridging the gap between perception and action. Extensive experiments show that SpNav achieves state-of-the-art (SOTA) performance in spatial navigation tasks across both simulated and real-world environments, validating the effectiveness of our method.
Haoxiang Fu, Xiaoshuai Hao, Qiang Zhang 0029, Long Chen 0015, Wenbo Ding 0001
AAAI8
2026 NavA³: Understanding Any Instruction, Navigating Anywhere, Finding Anything
abstract
Lingfeng Zhang, Xiaoshuai Hao, Yingbo Tang, Haoxiang Fu, Xinyu Zheng, Pengwei Wang, Zhongyuan Wang, Wenbo Ding, Shanghang Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiaoshuai Hao, Yingbo Tang, Haoxiang Fu, Pengwei Wang 0005, Zhongyuan Wang 0006, Wenbo Ding 0001, Shanghang Zhang
ACL (1)8
2026 Embodied Spatial Affordance: Spatial-Aware Affordance Learning for Embodied Navigation and Manipulation
abstract
Embodied navigation and manipulation are fundamental capabilities for embodied agents operating in physical environments. A key challenge in this process is understanding the spatial context and the affordances of the environment, which involves recognizing how objects can be interacted with (object affordance) and identifying suitable locations for movement and object placement (free space affordance). While Vision-Language Models (VLMs) have shown promise in high-level task planning, their ability to translate reasoning into precise executable actions remains limited, particularly in image-based spatial understanding and precise affordance localization-a critical gap in image processing for robotics. To bridge this gap, we propose EspA, a novel image-to-keypoint model that leverages spatial-aware affordance learning to predict actionable affordances directly from 2D image inputs. Built on a hierarchical vision-language architecture, EspA jointly reasons about object affordances and free space affordances, enabling pixel-level localization of both types of interactions. Crucially, EspA translates language instructions into precise 2D affordance keypoints from observed images, which are then projected into 3D actionable coordinates using depth information. To support this unified affordance reasoning, we introduce the Embodied Spatial Affordance (ESA) dataset, which captures both object-centric interactions and free space contexts. By jointly modeling these affordances in a shared representation space, EspA overcomes the limitations of prior works that treat them independently. The dataset's fine-grained annotations enable our model to learn the intricate relationship between object functionality and spatial feasibility, significantly enhancing the spatial understanding in embodied tasks. Extensive experimental results demonstrate that EspA outperforms existing state-of-the-art Vision-Language Models (VLMs), both open-source and closed-source, in object and free space affordance prediction. Furthermore, it exhibits superior performance in real-world embodied navigation and manipulation experiments. Our work advances the field of image-based spatial reasoning by providing a scalable solution for translating high-level instructions into low-level actionable affordances. We believe this work paves the way for more robust and versatile embodied agents capable of effectively interacting with complex environments. The dataset, benchmark, and evaluation code will be publicly available to facilitate future research. Project website: https://embodied-spatial-affordance.github.io/.
Xiaoshuai Hao, Yingbo Tang, Long Chen 0015, Wei Zhou 0021, Jungong Han, Wenbo Ding 0001, Xiao-Ping Zhang 0002
IEEE Trans. Image Process.7
2026 A Novel Integrated Sensing and Communication Scheme in UAVs-Enabled Vehicular Networks With MARL-Driven Adaptive Control
abstract
In this paper, we propose a novel integrated sensing and communication (ISAC) scheme tailored for UAVs-enabled vehicular networks, which leverages the information coverage capabilities of multiple UAVs and addresses critical challenges posed by multiple moving users. Unlike many traditional scheme, our scheme efficiently leverages ISAC signal echoes and real-time data uploads to provide communication services while achieving accurate sensing, thereby overcoming issues of resource waste and low operational efficiency. In the scheme, we aim to optimize both communication and sensing indicators, taking into account practical issues such as energy saving and collision avoidance for UAVs. However, the inherent complexity of multi-objective stochastic optimization in dynamic environments and limited communication resources render centralized UAV control inconvenient. To address the above challenges, we propose a novel multi-agent reinforcement learning (MARL) algorithm based on local information to realize the distributed adaptive control of motion decision, power selection, and channel allocation for UAVs. The algorithm combines random network distillation (RND) and dynamic data augmentation with multi-agent deep deterministic policy gradient (MADDPG) to encourage agents to explore effectively under sparse rewards and improve MADDPG's policy learning ability in finite data, thus approaching the global optimal solution. Experimental results demonstrate that the proposed algorithm can improve communication and sensing performance by more than 16.71% and 68.26% compared with other baselines and satisfy the set constraints. Furthermore, by adjusting hyperparameters, we can optimize the ISAC performance while achieving different energy savings levels for UAVs, proving that the designed scheme can reduce the waste of resources and improve the ISAC operation efficiency.
Ziyuan Wang 0002, Xiao-Ping Zhang 0002, Wenbo Ding 0001, Yuhan Dong, Xinlei Chen
IEEE Trans. Mob. Comput.3
2026 An Ensemble MARL Approach for Heterogeneous UAV Swarm Target Search in 3D Space
Changxu Wei, Ziyuan Wang 0002, Yixian Zhang, Wenbo Ding 0001, Xiao-Ping Zhang 0002
IEEE Trans. Mob. Comput.5
2026 SandWorm: Event-Based Visuotactile Perception With Active Vibration for Screw-Actuated Robot in Granular Media
abstract
Perception in granular media remains challenging due to unpredictable particle dynamics. To address this challenge, we present SandWorm, a biomimetic screw-actuated robot augmented by peristaltic motion to enhance locomotion, and SWTac, a novel event-based visuotactile sensor with an actively vibrated elastomer. The event camera is mechanically decoupled from vibrations by a spring isolation mechanism, enabling high quality tactile imaging of both dynamic and stationary objects. For algorithm design, we propose an IMU-guided temporal filter to enhance imaging consistency, improving MSNR by 24%. Moreover, we systematically optimize SWTac with vibration parameters, event camera settings and elastomer properties. Motivated by asymmetric edge features, we also implement contact surface estimation by U-Net. Experimental validation demonstrates SWTac's 0.2 mm texture resolution, 98% stone classification accuracy, and 0.15 N force estimation error, while SandWorm demonstrates versatile locomotion (up to 12.5 mm/s) in challenging terrains, successfully executes pipeline dredging and subsurface exploration in complex granular media (observed 90%success rate). Field experiments further confirm the system's practical performance.
Shoujie Li, Changqing Guo, Junhao Gong, Chenxin Liang, Wenhua Ding, Wenbo Ding 0001
IEEE Trans. Robotics6
2025 pFedGPA: Diffusion-based Generative Parameter Aggregation for Personalized Federated Learning
abstract
Federated Learning (FL) offers a decentralized approach to model training, where data remains local and only model parameters are shared between the clients and the central server. Traditional methods, such as Federated Averaging (FedAvg), linearly aggregate these parameters which are usually trained on heterogeneous data distributions, potentially overlooking the complex, high-dimensional nature of the parameter space. This can result in degraded performance of the aggregated model. While personalized FL approaches can mitigate the heterogeneous data issue to some extent, the limitation of linear aggregation remains unresolved. To alleviate this issue, we investigate the generative approach of diffusion model and propose a novel generative parameter aggregation framework for personalized FL, pFedGPA. In this framework, we deploy a diffusion model on the server to integrate the diverse parameter distributions and propose a parameter inversion method to efficiently generate a set of personalized parameters for each client. This inversion method transforms the uploaded parameters into a latent code, which is then aggregated through denoising sampling to produce the final personalized parameters. By encoding the dependence of a client's model parameters on the specific data distribution using the high-capacity diffusion model, pFedGPA can effectively decouple the complexity of the overall distribution of all clients' model parameters from the complexity of each individual client's parameter distribution. Our experimental results consistently demonstrate the superior performance of the proposed method across multiple datasets, surpassing baseline approaches.
Jiahao Lai, Jiaqi Li 0028, Jian Xu 0016, Yanru Wu, Boshi Tang, Wenbo Ding 0001, Yang Li 0104
AAAI8
2025 Mean-Field Aided QMIX: A Scalable and Flexible Q-Learning Approach for Large-Scale Agent Groups
abstract
Value decomposition methods are effective for multi-agent reinforcement learning (MARL), with QMIX being one of the most advanced. However, it struggles with scalability and flexibility in large-scale agent systems. The introduction of mean-field theory into MARL provides a potential solution to both scalability and flexibility challenges. In this paper, we propose Mean-Field Aided QMIX (MF-QMIX), a novel algorithm that addresses these challenges by incorporating mean-field theory into the QMIX framework. MF-QMIX significantly reduces the computational complexity from O(n) in QMIX to O(1), making it highly scalable and suitable for large-scale environments with many agents. Additionally, MF-QMIX is flexible, allowing the trained model to be applied across various group sizes without retraining, thus accommodating dynamic changes in the number of agents. Our extensive experiments demonstrate that MF-QMIX outperforms existing methods, both in computational efficiency and adaptability, across cooperative and competitive scenarios. This establishes MF-QMIX as a scalable and flexible solution for large-scale MARL problems.
Huaze Tang, Xiao-Ping Zhang 0002, Wenbo Ding 0001
ICASSP4
2025 Residual Kernel Policy Network: Enhancing Stability and Robustness in RKHS-Based Reinforcement Learning
abstract
Achieving optimal performance in reinforcement learning requires robust policies supported by training processes that ensure both sample efficiency and stability. Modeling the policy in reproducing kernel Hilbert space (RKHS) enables efficient exploration of local optimal solutions. However, the stability of existing RKHS-based methods is hindered by significant variance in gradients, while the robustness of the learned policies is often compromised due to the sensitivity of hyperparameters. In this work, we conduct a comprehensive analysis of the significant instability in RKHS policies and reveal that the variance of the policy gradient increases substantially when a wide-bandwidth kernel is employed. To address these challenges, we propose a novel RKHS policy learning method integrated with representation learning to dynamically process observations in complex environments, enhancing the robustness of RKHS policies. Furthermore, inspired by the advantage functions, we introduce a residual layer that further stabilizes the training process by significantly reducing gradient variance in RKHS. Our novel algorithm, the Residual Kernel Policy Network (ResKPN), demonstrates state-of-the-art performance, achieving a 30% improvement in episodic rewards across complex environments.
Yixian Zhang, Huaze Tang, Huijing Lin, Wenbo Ding 0001
ICLR4
2025 Chemistry3D: Robotic Interaction Toolkit for Chemistry Experiments
abstract
The advent of simulation engines has revolutionized learning and operational efficiency for robots, offering cost-effective and swift pipelines. However, the lack of a universal simulation platform tailored for chemical scenarios impedes progress in robotic manipulation and visualization of reaction processes. Addressing this void, we present Chemistry3D, an innovative toolkit that integrates extensive chemical and robotic knowledge. Chemistry3D not only enables robots to perform chemical experiments but also provides real-time visualization of temperature, color, and pH changes during reactions. Built on the NVIDIA Omniverse platform, Chemistry3D offers interfaces for robot operation, visual inspection, and liquid flow control, facilitating the simulation of special objects such as liquids and transparent entities. Leveraging this toolkit, we have devised RL tasks, object detection, and robot operation scenarios. Additionally, to discern disparities between the rendering engine and the real world, we conducted transparent object detection experiments using Sim2Real, validating the toolkit's exceptional simulation performance. The source code is available at https://github.com/huangyan28/Chemistry3D, and a related tutorial can be found at https://www.omni-chemistry.com.
Shoujie Li, Changqing Guo, Jiawei Zhang 0012, Linrui Zhang, Wenbo Ding 0001
ICRA7
2025 Bio-Inspired Soft Magnetic Swimming Robot for Flexible Motions
abstract
Bio-inspired soft robots have gained significant attention for their flexible design and adaptability to various environments, making them suitable for exploration and task execution in confined or hazardous areas. However, the deformation and motion of soft magnetic robots rely on both their structural design and magnetization, which complicates the guided movement and balance maintenance for aquatic environments. In this work, inspired by the flat and symmetrical body of rays, we design a soft magnetic fish-shaped robot capable of flexible motions and trajectory swimming on the water surface. This robot features the muscle made of magnetic elastomer, which connects with the acrylic skeleton and silicone film fins with a soft body. In the external magnetic field, the robot achieves hovering by flapping its fins, driven by the magnetically actuated deformation of its magnetic muscle. Besides, the robot's axial magnetization enables the rapid steering guided by a horizontal field. In experiments, the soft magnetic robot was tasked with performing a looping figure-eight trajectory movement on the water surface, guided by the field gradient generated by a dense planar electromagnetic coils' array. When moving, the onboard circuit board of the robot collected its inertial and temperature information, and sent these data to the host computer via Bluetooth in real-time for motion monitoring. Received data demonstrated that our robot performed the specified afloat swimming trajectory, exhibiting a good stability on its yaw angle during the continuous motion. The soft magnetic swimming robot shows its integrated functionalities in untethered actuation, on-robot sensing, and wireless communication, indicating a significant prospect on applications in inspection and cleaning within narrow pipelines and enclosed mechanical interior spaces.
Xiaosa Li, Zenan Lin, Wenbo Ding 0001
ICRA3
2025 MonoLDP: LED Assisted Indoor Mobile Bot Monocular Depth Prediction and Pose Estimation System
abstract
Multi-robot clusters are increasingly deployed in indoor environments, where effective communication and 3D perception are critical for coordinated operations. Monocular cameras, known for their lightweight design, cost-effectiveness, and versatility, present a promising solution for these tasks. However, relying solely on monocular cameras for comprehensive perception and communication presents significant challenges. To address this, we introduce MonoLDP, a novel system that leverages monocular cameras for depth estimation, mutual pose estimation, and visible light communication in indoor environments, providing an integrated framework to overcome these limitations. MonoLDP features a two-stage network: (1) a depth estimation module that infers depth from monocular images, and (2) a depth-guided 3D object recognition network for agent-relative localization and pose estimation. We created a custom dataset to validate the accuracy of MonoLDP. On our indoor dataset, MonoLDP outperforms the baseline by 43.39% in 3D detection and 42.39% in bird's-eye view detection, with an average localization error of 0.104 m and an orientation error of 1.66 degrees. Moreover, the depth estimation network demonstrates excellent performance on the NYU v2 dataset. Additionally, the system achieves a communication rate of 1.2 Kbps with a bit error rate below 10-2at a distance of up to 4 m using LED arrays. Our code will be released at https://github.com/RavenLiang1005/MonoLDP.git.
Chenxin Liang, Shoujie Li, Kit Wa Sou, Xinyu Luo, Wenbo Ding 0001
ICRA6
2025 PUGS: Zero-Shot Physical Understanding with Gaussian Splatting
abstract
Current robotic systems can understand the categories and poses of objects well. But understanding physical properties like mass, friction, and hardness, in the wild, remains challenging. We propose a new method that reconstructs 3D objects using the Gaussian splatting representation and predicts various physical properties in a zero-shot manner. We propose two techniques during the reconstruction phase: a geometryaware regularization loss function to improve the shape quality and a region-aware feature contrastive loss function to promote region affinity. Two other new techniques are designed during inference: a feature-based property propagation module and a volume integration module tailored for the Gaussian representation. Our framework is named as zero-shot physical understanding with Gaussian splatting, or PUGS. PUGS achieves new state-of-the-art results on the standard benchmark of ABO-500 mass prediction. We provide extensive quantitative ablations and qualitative visualization to demonstrate the mechanism of our designs. We show the proposed methodology can help address challenging real-world grasping tasks. Our codes, data, and models are available at https://github.com/EverNorif/PUGS
Yinghao Shuai, Yuantao Chen, Zijian Jiang, Nan Wang 0041, Jv Zheng, Jianzhu Ma, Meng Yang 0035, Zhicheng Wang 0022, Wenbo Ding 0001, Hao Zhao 0002
ICRA11
2025 Depth Restoration of Hand-Held Transparent Objects for Human-to-Robot Handover
abstract
Transparent objects are common in daily life, while their optical properties pose challenges for RGB-D cameras to capture accurate depth information. This issue is further amplified when these objects are hand-held, as hand occlusions further complicate depth estimation. For assistant robots, however, accurately perceiving hand-held transparent objects is critical to effective human-robot interaction. This paper presents a Hand-Aware Depth Restoration (HADR) method based on creating an implicit neural representation function from a single RGB-D image. The proposed method utilizes hand posture as an important guidance to leverage semantic and geometric information of hand-object interaction. To train and evaluate the proposed method, we create a highfidelity synthetic dataset named TransHand-$\mathbf{1 4 K}$with a real-tosim data generation scheme. Experiments show that our method has better performance and generalization ability compared with existing methods. We further develop a real-world human-to-robot handover system based on HADR, demonstrating its potential in human-robot interaction applications.
Haixin Yu, Shoujie Li, Ziwu Song, Wenbo Ding 0001
ICRA6
2025 Exo-ViHa: A Cross-Platform Exoskeleton System with Visual and Haptic Feedback for Efficient Dexterous Skill Learning
abstract
Imitation learning has emerged as a powerful paradigm for robot skills learning. However, traditional data collection systems for dexterous manipulation face challenges, including a lack of balance between acquisition efficiency, consistency, and accuracy. To address these issues, we introduce Exo-ViHa, an innovative 3D-printed exoskeleton system that enables users to collect data from a first-person perspective while providing real-time haptic feedback. This system combines a 3D-printed modular structure with a slam camera, a motion capture glove, and a wrist-mounted camera. Various dexterous hands can be installed at the end, enabling it to simultaneously collect the posture of the end effector, hand movements, and visual data. By leveraging the first-person perspective and direct interaction, the exoskeleton enhances the task realism and haptic feedback, improving the consistency between demonstrations and actual robot deployments. In addition, it has cross-platform compatibility with various robotic arms and dexterous hands. Experiments show that the system can significantly improve the success rate and efficiency of data collection for dexterous manipulation tasks. Webpage: https://exo-viha2025.github.io/.
Xintao Chao, Shilong Mu, Yushan Liu 0006, Shoujie Li, Chuqiao Lyu, Xiao-Ping Zhang 0002, Wenbo Ding 0001
IROS7
2025 UltraTac: Integrated Ultrasound-Augmented Visuotactile Sensor for Enhanced Robotic Perception
abstract
Visuotactile sensors provide high-resolution tactile information but are incapable of perceiving the material features of objects. We present UltraTac, an integrated sensor that combines visuotactile imaging with ultrasound sensing through a coaxial optoacoustic architecture. The design shares structural components and achieves consistent sensing regions for both modalities. Additionally, we incorporate acoustic matching into the traditional visuotactile sensor structure, enabling the integration of the ultrasound sensing modality without compromising visuotactile performance. Through tactile feedback, we can dynamically adjust the operating state of the ultrasound module to achieve more flexible functional coordination. Systematic experiments demonstrate three key capabilities: proximity sensing in the 3–8 cm range (R2= 0.99), material classification (average accuracy: 99.20%), and texture-material dual-mode object recognition achieves 92.11% accuracy on a 15-class task. Finally, we integrate the sensor into a robotic manipulation system to concurrently detect container surface patterns and internal content, which verifies its promising potential for advanced human-machine interaction and precise robotic manipulation.
Junhao Gong, Kit Wa Sou, Shoujie Li, Changqing Guo, Chuqiao Lyu, Ziwu Song, Wenbo Ding 0001
IROS8
2025 AirTouch: A Low-Cost Versatile Visuotactile Feedback System for Enhanced Robotic Teleoperation
abstract
Vision-based teleoperation systems are widely used due to their cost-effectiveness and intuitive operation. However, these systems often suffer from challenges such as hand occlusions, environmental variability, and the lack of tactile feedback, limiting their precision and applicability in complex tasks. To address these limitations, we present Air-Touch, a novel, low-cost visuotactile teleoperation system that integrates air pressure-based tactile feedback with lightweight hand pose estimation. AirTouch features an inflatable tactile bubble that provides adjustable feedback through closed-loop pneumatic control, enhancing the operator’s sense of interaction with remote environments. The system’s robust hand-tracking algorithm ensures accurate control even under dynamic and occlusion-prone conditions, while its hardware design eliminates the need for wearable devices, enabling intuitive operation. AirTouch supports a wide range of robotic end-effectors, including dexterous hands, parallel grippers, and suction cups, demonstrating versatility across multiple platforms. Extensive experiments validate AirTouch’s performance, achieving high precision in hand pose estimation and a 91% success rate in complex teleoperation tasks, all with a hardware cost as low as $39. These results highlight AirTouch as a scalable and practical solution for enhancing robotic teleoperation across industrial, medical, and hazardous scenarios.
Shoujie Li, Xingting Li, Ken Jiankun Zheng, Xueqian Wang 0001, Wenbo Ding 0001
IROS7
2025 Distributional Decision Transformer: Risk-Sensitive Offline RL via Quantile-Based Critics and Stochastic Return
abstract
Offline reinforcement learning faces a critical challenge in synthesizing high-reward trajectories from suboptimal datasets while robustly handling the stochasticity inherent in real-world decision-making. While combination of return-conditioned sequence models, such as Decision Transformers (DT), and dynamics programming critics shows great potential in trajectory synthesis, their deterministic action generation and scale Q value critic often fails to distinguish intentional behavioral variability from detrimental noise, leading to suboptimal policy collapse. To address this challenge, we propose the Distributional Decision Transformer (DDT), a novel framework that unifies probabilistic return distribution modeling with autoregressive action generation. DDT introduces two key innovations: (1) a Gaussian stochastic return mechanism that reparameterizes target returns as samplable distributions, enabling diverse action candidate generation; and (2) an Implicit Quantile Network (IQN) critic embedded within the deciding loop, which evaluates actions across the full spectrum of return distributions (quantiles τ ~ U(0, 1)). In D4RL benchmarks, DDT achieves state-of-the-art performance, achieving a 91.6 average normalized score in MuJoCo locomotion and 69.3 in sparse-reward settings. The results establish DDT as a principled solution for synthesis of risk-aware trajectory in offline RL.
Changxu Wei, Huaze Tang, Yixian Zhang, Chao Wang 0139, Xiao-Ping Zhang 0002, Wenbo Ding 0001
IROS6
2025 MuxHand: A Cost-Effective and Compact Dexterous Robotic Hand Using Time-Division Multiplexing Mechanism
abstract
The number of motors directly influences the dexterity, size, and cost of a robotic hand. In this paper, we present MuxHand, a robotic hand that utilizes a time-division multiplexing motor (TDMM) mechanism. This system enables independent control of 9 cables with just 4 motors, significantly reducing both cost and size while maintaining high dexterity. To enhance stability and smoothness during grasping and manipulation tasks, we integrate magnetic joints into the three 3D-printed fingers. These joints provide impact resistance, resetting capabilities. The three fingers together have a total of 30 degrees of freedom (DOF), 18 of which are passive DOF, allowing the hand to conform closely to the surface of an object during grasping. We conduct a series of experiments to assess the performance parameters of MuxHand, including its grasping and manipulation capabilities. The results show that the TDMM mechanism precisely controls each cable connected to the finger joints, enabling robust grasping and dexterous manipulation. Furthermore, compared to the traditional approach of assigning a motor to each active DOF, the cost is reduced by 42.06%. The maximum load of a single finger reaches 7.0 kg, the maximum load at the finger joint root is 12.0 kg, the maximum driving force at the joint root is 5.0 kg, and the maximum fingertip force is 10.0 N.
Jianle Xu, Shoujie Li, Houde Liu, Xueqian Wang 0001, Wenbo Ding 0001, Chongkun Xia
IROS6
2025 ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model
abstract
Multi-task robotic bimanual manipulation is becoming increasingly popular as it enables sophisticated tasks that require diverse dual-arm collaboration patterns. Compared to unimanual manipulation, bimanual tasks pose challenges to understanding the multi-body spatiotemporal dynamics. An existing method ManiGaussian [30] pioneers encoding the spatiotemporal dynamics into the visual representation via Gaussian world model for single-arm settings, which ignores the interaction of multiple embodiments for dual-arm systems with significant performance drop. In this paper, we propose ManiGaussian++, an extension of ManiGaussian framework that improves multi-task bimanual manipulation by digesting multi-body scene dynamics through a hierarchical Gaussian world model. To be specific, we first generate task-oriented Gaussian Splatting from intermediate visual features, which aims to differentiate acting and stabilizing arms for multi-body spatiotemporal dynamics modeling. We then build a hierarchical Gaussian world model with the leader-follower architecture, where the multi-body spatiotemporal dynamics is mined for intermediate visual representation via future scene prediction. The leader predicts Gaussian Splatting deformation caused by motions of the stabilizing arm, through which the follower generates the physical consequences resulted from the movement of the acting arm. As a result, our method significantly outperforms the current state-of-the-art bimanual manipulation techniques by an improvement of 20.2% in 10 simulated tasks, and achieves 60% success rate on average in 9 challenging real-world tasks. Our code is available at https://github.com/April-Yz/ManiGaussian_Bimanual.
Tengbo Yu, Guanxing Lu, Zaijia Yang, Haoyuan Deng, Season Si Chen, Jiwen Lu, Wenbo Ding 0001, Guoqiang Hu 0001, Yansong Tang, Ziwei Wang 0010
IROS7
2025 VET: A Visual-Electronic Tactile System for Immersive Human-Machine Interaction
abstract
In the pursuit of deeper immersion in human-machine interaction, achieving higher-dimensional tactile input and output on a single interface has become a key research focus. This study introduces the Visual-Electronic Tactile (VET) System, which builds upon vision-based tactile sensors (VBTS) and integrates electrical stimulation feedback to enable bidirectional tactile communication. We propose and implement a system framework that seamlessly integrates an electrical stimulation film with VBTS using a screen-printing preparation process, eliminating interference from traditional methods. While VBTS captures multi-dimensional input through visuotactile signals, electrical stimulation feedback directly stimulates neural pathways, preventing interference with visuotactile information. The potential of the VET system is demonstrated through experiments on finger electrical stimulation sensitivity zones, as well as applications in interactive gaming and robotic arm teleoperation. This system paves the way for new advancements in bidirectional tactile interaction and its broader applications.
Yisheng Yang, Shilong Mu, Chuqiao Lyu, Shoujie Li, Xinyue Chai, Wenbo Ding 0001
IROS7
2025 Poster: LightWalk: Passive Gait Recognition via Reflected VLC Signals
abstract
A visible light communication (VLC)-based sensing system is proposed for gait recognition and classification. Human-induced reflections are modeled using a time-varying channel representation, enabling the capture of gait dynamics without requiring wearable devices. A low-cost sensing module with embedded processing and multi-channel photodetectors is implemented. Filtered signals are transformed into spectral-spatial features and analyzed using deep learning models. Experiments involving ten participants across eight gait types demonstrate that contrastive and multi-scale models achieve over 98% accuracy, highlighting the potential of VLC-based sensing for unobtrusive, privacy-preserving, and real-time human gait recognition.
Jiarong Li 0004, Chihan Xu, Wenfeng Deng, Xiaojun Liang, Wenbo Ding 0001, Weihua Gui 0001
MobiCom5
2025 Universal Visuo-Tactile Video Understanding for Embodied Interaction
abstract
Tactile perception is essential for embodied agents to understand the physical attributes of objects that cannot be determined through visual inspection alone. While existing methods have made progress in visual and language modalities for physical understanding, they fail to effectively incorporate tactile information that provides crucial haptic feedback for real-world interaction. In this paper, we present VTV-LLM, the first multi-modal large language model that enables universal Visuo-Tactile Video (VTV) understanding, bridging the gap between tactile perception and natural language. To address the challenges of cross-sensor and cross-modal integration, we contribute VTV150K, a comprehensive dataset comprising 150,000 video frames from 100 diverse objects captured across three different tactile sensors (GelSight Mini, DIGIT, and Tac3D), annotated with four fundamental tactile attributes (hardness, protrusion, elasticity, and friction). We develop a novel three-stage training paradigm that includes VTV enhancement for robust visuo-tactile representation, VTV-text alignment for cross-modal correspondence, and text prompt finetuning for natural language generation. Our framework enables sophisticated tactile reasoning capabilities including feature assessment, comparative analysis, and scenario-based decision-making. Extensive experimental evaluations demonstrate that VTV-LLM achieves superior performance in tactile reasoning tasks, establishing a foundation for more intuitive human-machine interaction in tactile domains.
Shoujie Li, Xingting Li, Guangyu Chen, Fei Ma 0006, F. Richard Yu, Wenbo Ding 0001
NeurIPS8
2025 VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play
abstract
Robot sports, characterized by well-defined objectives, explicit rules, and dynamic interactions, present ideal scenarios for demonstrating embodied intelligence. In this paper, we present VolleyBots, a novel robot sports testbed where multiple drones cooperate and compete in the sport of volleyball under physical dynamics. VolleyBots integrates three features within a unified platform: competitive and cooperative gameplay, turn-based interaction structure, and agile 3D maneuvering.These intertwined features yield a complex problem combining motion control and strategic play, with no available expert demonstrations.We provide a comprehensive suite of tasks ranging from single-drone drills to multi-drone cooperative and competitive tasks, accompanied by baseline evaluations of representative reinforcement learning (RL), multi-agent reinforcement learning (MARL) and game-theoretic algorithms. Simulation results show that on-policy RL methods outperform off-policy methods in single-agent tasks, but both approaches struggle in complex tasks that combine motion control and strategic play.We additionally design a hierarchical policy which achieves 69.5% win rate against the strongest baseline in the 3 vs 3 task, demonstrating its potential for tackling the complex interplay between low-level control and high-level strategy.To highlight VolleyBots’ sim-to-real potential, we further demonstrate the zero-shot deployment of a policy trained entirely in simulation on real-world drones.
Zelai Xu, Ruize Zhang 0001, Chao Yu 0005, Huining Yuan 0002, Xiangmin Yi, Shilong Ji, Chuqi Wang, Wenbo Ding 0001, Xinlei Chen, Yu Wang 0002
NeurIPS10
2025 Stabilizing and improving federated learning with highly non-iid data and client dropout
Jian Xu 0016, Meilin Yang, Wenbo Ding 0001, Shao-Lun Huang
Appl. Intell.3
2025 AllTact Fin Ray: A Compliant Robot Gripper With Omni-Directional Tactile Sensing
abstract
Tactile sensing plays a crucial role in robot grasping and manipulation by providing essential contact information between the robot and the environment. In this paper, we present AllTact Fin Ray, a novel compliant gripper design with omni-directional and local tactile sensing capabilities. The finger body is unibody-casted using transparent elastic silicone, and a camera positioned at the base of the finger captures the deformation of the whole body and the contact face. Due to the global deformation of the adaptive structure, existing vision-based tactile sensing approaches that assume constant illumination are no longer applicable. To address this, we propose a novel sensing method where the global deformation is first reconstructed from the image using edge features and spatial constraints. Then, detailed contact geometry is computed from the brightness difference against a dynamically retrieved reference image. Extensive experiments validate the effectiveness of our proposed gripper design and sensing method in contact detection, force estimation, object grasping, and precise manipulation.
Siwei Liang, Jing Xu 0011, Hongyu Qian, Xiangjun Zhang, Dan Wu 0008, Wenbo Ding 0001, Rui Chen 0019
IEEE Trans Autom. Sci. Eng.7
2024 CoSTA: End-to-End Comprehensive Space-Time Entanglement for Spatio-Temporal Video Grounding
abstract
This paper studies the spatio-temporal video grounding task, which aims to localize a spatio-temporal tube in an untrimmed video based on the given text description of an event. Existing one-stage approaches suffer from insufficient space-time interaction in two aspects: i) less precise prediction of event temporal boundaries, and ii) inconsistency in object prediction for the same event across adjacent frames. To address these issues, we propose a framework of Comprehensive Space-Time entAnglement (CoSTA) to densely entangle space-time multi-modal features for spatio-temporal localization. Specifically, we propose a space-time collaborative encoder to extract comprehensive video features and leverage Transformer to perform spatio-temporal multi-modal understanding. Our entangled decoder couples temporal boundary prediction and spatial localization via an entangled query, boasting an enhanced ability to capture object-event relationships. We conduct extensive experiments on the challenging benchmarks of HC-STVG and VidSTG, where CoSTA outperforms existing state-of-the-art methods, demonstrating its effectiveness for this task.
Yaoyuan Liang, Yansong Tang, Zhao Yang 0002, Ziran Li, Jingang Wang, Wenbo Ding 0001, Shao-Lun Huang
AAAI7
2024 Beyond Aggregation: Efficient Federated Model Consolidation with Heterogeneity-Adaptive Weights Diffusion
abstract
As the Internet of Things (IoT) evolves, the need for enhanced data-sharing to improve edge device performance has led to the adoption of Federated Learning (FL) for data privacy and optimized data utilization. However, communication costs in FL remain a significant challenge. Traditional methods focus on client enhancements but overlook server-side aggregation, potentially increasing client computation loads. In response, we introduce a novel method, FedDiff, which utilizes diffusion models for generating model weights on FL servers, replacing traditional aggregation methods. Our approach, tailored for heterogeneous environments, significantly improves communication efficiency, achieving faster convergence and robust performance against weight noise in rigorous tests.
Jiaqi Li 0028, Xiaoyang Qu, Wenbo Ding 0001, Zihao Zhao 0001, Jianzong Wang
CIKM3
2024 A Self-adaptive Rotationally Invariant Particle Swarm Optimization for Global Optimization
abstract
Recently, the rotational invariance property has been introduced into the field of meta-heuristic algorithms, aiming to ensure consistent algorithm performance across arbitrarily rotated problems, thereby enhancing algorithmic universality. However, designing a practical rotationally invariant Particle Swarm Optimization (PSO) variant with superior performance across various optimization problems still remains a subject of in-depth research. In this paper, we propose a novel rotationally invariant PSO variant termed self-adaptive rotationally invariant PSO (sariPSO). Specifically, sariPSO incorporates a newly developed rotationally invariant solution updating equation, formulated using random vectors uniformly distributed within D-dimensional ellipsoids. Additionally, to further enhance the algorithm performance across various problems, sariPSO employs a self-adaptive approach to determine the axis length parameters of the ellipsoids, based on evaluation of performances of the solution updating equation under different axis length parameters in the last few iterations. Numerical experiments conducted on 10 well-known test problems demonstrate the rotational invariance property and superior performance of sariPSO compared to several existing PSO variants.
Haoxin Wang 0001, Wenbo Ding 0001, Libao Shi
GECCO3
2024 M3ARL: Moment-Embedded Mean-Field Multi-Agent Reinforcement Learning for Continuous Action Space
abstract
Mean-field theory offers a promising solution to the scalability issues encountered in multi-agent reinforcement learning (MARL) within large-scale systems. However, most existing MARL algorithms based on mean-field theory are typically constrained to discrete action space. In continuous action space, the conventional mean-field approximation of mean-field theory breaks down since the representation of mean-field action is not well-defined. In this paper, we propose Moment- Embedded Mean-Field Multi-Agent Reinforcement Learning (M3ARL) for continuous action space, embedding mean-field action into multi-order moments. Specifically, we analyze the mean-field approximation on Wasserstein space and derive the approximated form of Q-function. Furthermore, to derive a learn-able form of Q-function, we apply the cylindrical function and propose the moment-embedded representations. Based on the representations and proximal policy optimization (PPO) algorithm, we propose a novel algorithm, M3APPO. Finally, we validate the efficacy of the proposed algorithm through experiments of predator-prey environment on continuous action space scenarios.
Huaze Tang, Yuanquan Hu, Fanfan Zhao, Junji Yan, Wenbo Ding 0001
ICASSP6
2024 Optimizing Trading Strategies in Quantitative Markets Using Multi-Agent Reinforcement Learning
abstract
Quantitative markets are characterized by swift dynamics and abundant uncertainties, making the pursuit of profit-driven stock trading actions inherently challenging. Within this context, Reinforcement Learning (RL) — which operates on a reward-centric mechanism for optimal control — has surfaced as a potentially effective solution to the intricate financial decision-making conundrums presented. This paper delves into the fusion of two established financial trading strategies, namely the constant proportion portfolio insurance (CPPI) and the time-invariant portfolio protection (TIPP), with the multi-agent deep deterministic policy gradient (MADDPG) framework. As a result, we introduce two novel multi-agent RL (MARL) methods: CPPI-MADDPG and TIPP-MADDPG, tailored for probing strategic trading within quantitative markets. To validate these innovations, we implemented them on a diverse selection of 100 real-market shares. Our empirical findings reveal that the CPPI-MADDPG and TIPP-MADDPG strategies consistently outpace their traditional counterparts, affirming their efficacy in the realm of quantitative trading.
Hengxi Zhang, Zhendong Shi, Yuanquan Hu, Wenbo Ding 0001, Ercan E. Kuruoglu, Xiao-Ping Zhang 0002
ICASSP4
2024 Federated PAC-Bayesian Learning on Non-IID Data
abstract
Existing research has either adapted the Probably Approximately Correct (PAC) Bayesian framework for federated learning (FL) or used information-theoretic PAC-Bayesian bounds while introducing their theorems, but few consider the non-IID challenges in FL. Our work presents the first non-vacuous federated PAC-Bayesian bound tailored for non-IID local data. This bound assumes unique prior knowledge for each client and variable aggregation weights. We also introduce an objective function and an innovative Gibbs-based algorithm for the optimization of the derived bound. The results are validated on real-world datasets.
Zihao Zhao 0001, Yang Liu 0165, Wenbo Ding 0001, Xiao-Ping Zhang 0003
ICASSP3
2024 PowerGest: Self-Powered Gesture Recognition for Command Input and Robotic Manipulation
abstract
As human-computer interaction (HCI) advances, gesture recognition has emerged as a transformative technology for human-computer interaction. Traditional methods, often camera or glove-based, are restricted by various environmental conditions and user-specific demands, highlighting the need for more universal, non-intrusive, and sustainable solutions. Addressing this, we present PowerGest, a self-powered gesture recognition system based on a solar cell array. This innovative system leverages the dual functionalities of solar cells: energy harvesting and gesture sensing, providing an alternative to conventional methods. It integrates a designed low-powered data acquisition chip with a wireless transmission module and a user-friendly interface. PowerGest employs a series of signal processing methods and utilizes several machine learning algorithms, achieving over 97% accuracy for both numeric input and activity control gesture recognition tasks. With its broad applications in robotic control, text input, and more, PowerGest contributes to a more sustainable and intuitive HCI experience. Project demo: https://drive.google.com/drive/folders/10KEul8PAfvTUomi0JvZ8u411JyScCXQI?usp=sharing.
Jiarong Li 0004, Qinghao Xu, Zhancong Xu, Changshuo Ge, Liguang Ruan, Xiaojun Liang, Wenbo Ding 0001, Weihua Gui 0001, Xiao-Ping Zhang 0002
ICPADS7
2024 DeformNet: Latent Space Modeling and Dynamics Prediction for Deformable Object Manipulation
abstract
Manipulating deformable objects is a ubiquitous task in household environments, demanding adequate representation and accurate dynamics prediction due to the objects’ infinite degrees of freedom. This work proposes DeformNet, which utilizes latent space modeling with a learned 3D representation model to tackle these challenges effectively. The proposed representation model combines a PointNet encoder and a conditional neural radiance field (NeRF), facilitating a thorough acquisition of object deformations and variations in lighting conditions. To model the complex dynamics, we employ a recurrent state-space model (RSSM) that accurately predicts the transformation of the latent representation over time. Extensive simulation experiments with diverse objectives demonstrate the generalization capabilities of DeformNet for various deformable object manipulation tasks, even in the presence of previously unseen goals. Finally, we deploy DeformNet on an actual UR5 robotic arm to demonstrate its capability in real-world scenarios.
Chenchang Li, Zihao Ai, Xiaosa Li, Wenbo Ding 0001, Huazhe Xu
ICRA5
2024 Point-Wise Vibration Pattern Production via a Sparse Actuator Array for Surface Tactile Feedback
abstract
Surface vibration tactile feedback is capable of conveying various semantic information to humans via handheld electronic devices, such as smartphones, touch panels, and game controllers. However, covering the entire contacting surface of the device with a dense arrangement of actuators can affect its normal use. Determining how to produce desired vibration patterns at any contact point with only a few sparse actuators deployed on the surface of the handheld device remains a significant challenge. In this work, we develop a tactile feedback board in the size of a smartphone with only five actuators, and achieve the precise production of vibration patterns that can focus at any desired position on the board. Specifically, we investigate the vibration characteristics of a single passive coil actuator and construct its vibration pattern model for any position on the feedback board surface. Optimal phase and amplitude modulation, determined using the simulated annealing algorithm, is employed with five actuators in a sparse array. The vibration patterns from all actuators are superimposed linearly to synthetically generate different onboard vibration energy distributions for tactile sensing. Experiments demonstrated that point-wise vibration pattern production on our tactile board achieved an average level of about 0.9 in the Structural Similarity Index Measure (SSIM) evaluation, when compared to the ideal single-point-focused target vibration pattern. Four point-wise patterns focused on the top, bottom, left, and right parts of the tactile board were applied, to guide continuous directional movements without visual assistance, which shows significant implications for machine-assisted cognition based on vibration tactile feedback.
Xiaosa Li, Chengyue Lu, Wenbo Ding 0001
ICRA5
2024 Dual-modal Tactile E-skin: Enabling Bidirectional Human-Robot Interaction via Integrated Tactile Perception and Feedback
abstract
To foster an immersive and natural human-robot interaction (HRI), the implementation of tactile perception and feedback becomes imperative, effectively bridging the conventional sensory gap. In this paper, we propose a dual-modal electronic skin (e-skin) that integrates magnetic tactile sensing and vibration feedback for enhanced HRI. The dual-modal tactile e-skin offers multi-functional tactile sensing and programmable haptic feedback, underpinned by a layered structure comprised of flexible magnetic films, soft silicone elastomer, a Hall sensor and actuator array, and a microcontroller unit. The e-skin captures the magnetic field changes caused by subtle deformations through Hall sensors, employing deep learning for accurate tactile perception. Simultaneously, the actuator array generates mechanical vibrations to facilitate haptic feedback, delivering diverse mechanical stimuli. Notably, the dual-modal e-skin is capable of transmitting tactile information bidirectionally, enabling object recognition and fine-weighing operations. This bidirectional tactile interaction framework will enhance the immersion and efficiency of interactions between humans and robots.
Shilong Mu, Zenan Lin, Shoujie Li, Chenchang Li, Xiao-Ping Zhang 0002, Wenbo Ding 0001
ICRA8
2024 SATac: A Thermoluminescence Enabled Tactile Sensor for Concurrent Perception of Temperature, Pressure, and Shear
abstract
Most vision-based tactile sensors use elastomer deformation to infer tactile information, which can not sense some modalities, like temperature. As an important part of human tactile perception, temperature sensing can help robots better interact with the environment. In this work, we propose a novel multi-modal vision-based tactile sensor, SATac, which can simultaneously perceive information on temperature, pressure, and shear. SATac utilizes the thermoluminescence of strontium aluminate to sense a wide range of temperatures with exceptional resolution. Additionally, the pressure and shear can also be perceived by analyzing the Voronoi diagram. A series of experiments are conducted to verify the performance of our proposed sensor. We also discuss the possible application scenarios and demonstrate how SATac could benefit robot perception capabilities.
Ziwu Song, Kit Wa Sou, Shilong Mu, Dengfeng Peng, Xiao-Ping Zhang 0002, Wenbo Ding 0001
ICRA8
2024 VLocSense: Integrated VLC System for Indoor Passive Localization and Human Sensing
abstract
The demand for accurate and real-time indoor localization and human sensing is rising with the development of smart environments, with applications in security and smart homes. Effective systems enhance safety, energy efficiency, and user experience by leveraging existing infrastructure, reducing deployment costs, and integrating seamlessly. Traditional methods rely on dedicated hardware, while communication or lighting infrastructure can provide dual-purpose solutions. This research focuses on Visible Light Communication (VLC) technology, which uses visible light for data transmission. Our VLC system utilizes existing lighting infrastructure to transmit data, providing localization and human sensing functionalities. The system design strategically incorporates specific VLC transmitters and receivers to enhance sensing performance. The collected data is processed using advanced algorithms and machine learning models, ensuring robust, real-time, and cost-effective localization with an accuracy of 96.6% and human activity recognition with an accuracy of 98.3%. This multi-functionality system demonstrates the potential for VLC in healthcare monitoring and home automation applications.
Jiarong Li 0004, Changshuo Ge, Chihan Xu, Junhao Gong, Weihua Gui 0001, Xiaojun Liang, Wenbo Ding 0001
MobiCom8
2024 FormerReckoning: Physics Inspired Transformer for Accurate Inertial Navigation
abstract
Although modern localization methods have achieved remarkable accuracy with various sensors, there are still some circumstances where only proprioceptive sensing works (Inertial Navigation). However, localization and navigation using only IMU sensors (costing less than $1000) still face significant challenges such as low accuracy and large cumulative errors when using traditional filter methods. Furthermore, AI-based approaches, while promising, often yield unpredictable and unreliable outputs. This paper proposes FormerReckoning, an inertial localization estimation framework for wheeled robotics that incorporates physical prompts into a Transformer framework to enhance translation estimation accuracy. Our tests show that FormerReckoning not only reduces mean translation errors to 0.72% but also surpasses all baseline models in performance, demonstrating its potential to provide reliable and precise localization in a cost-effective manner.
Jiaqi Li 0028, Chenyu Zhao 0002, Yuzhu Mao, Xinlei Chen, Wenbo Ding 0001, Xiaoyang Qu, Jianzong Wang
MobiCom5
2024 HandySense: A Multimodal Collection System for Human Two-Handed Dexterous Manipulation
abstract
Humanoid robots with dexterous hands have gained significant attention due to their manipulation capabilities. Recent advancements are driven by large-scale real robot data and teleoperation technology, enabling precise operation demonstrations and smooth trajectories. Common methods like virtual reality devices, cameras, wearable gloves, and custom hardware face the inability to capture real information about human-object contact, such as tactile information. In this study, we present HandySense, a multimodal system integrating visual, tactile, motion, and spatial perception for robust and comprehensive two-handed manipulation tracking. HandySense includes RGB-D cameras, visual-inertial tracking cameras, and a motion capture glove with fingertip tactile sensors. Our framework achieved 99.45% accuracy in classifying 12 task stages, exhibiting the potential for large-scale human demonstration data collection and representing a pivotal step towards empowering humanoid robots to execute complex manipulations.
Shilong Mu, Xinyue Chai, Xingting Li, Wenbo Ding 0001
MobiCom6
2024 MAMGDT: Enhancing Multi-Agent Systems with Multi-Game Decision Transformer
abstract
Multi-agent systems (MAS) introduce the necessity about achieving collaborative goals and individual decision quality simultaneously. In the subfield of multi-agent reinforcement learning, transformer-based methods like MGDT enabled transportable utilization of temporal contexts in decision making for single agents. We introduces Multi-Agent Multi-Game Decision Transformer (MAMGDT) as a workaround for MAS tasks which presents a in-context learning approach to enhance decision-making in complex multi-agent environments. By incorporating causal masking and agent-wise interaction modeling, MAMGDT maintains temporal accuracy and captures the dynamics of agent interactions effectively. With the potential to optimize decision processes, MAMGDT is well-suited for various multi-agent tasks. We carried out demostration of MAMGDT on popular SMACv1 benchmarks with both collaborative and adversarial tasks, where MAMGDT outperformed other approaches.
Chao Wang 0139, Huaze Tang, Wenbo Ding 0001
MobiCom3
2024 Poster: Real-time Material and Texture Recognition Using Visible Light Communication
abstract
In response to the challenges presented by conventional material and texture recognition methods, our research introduces a system using visible light communication (VLC) technology. This approach provides a non-contact, non-destructive, dual-functional solution, overcoming the limitations of cost, safety, and environmental adaptability associated with traditional methods. Through a comprehensive design integrating hardware and software, our system utilizes VLC for precise and efficient recognition. Extensive testing confirms its effectiveness, achieving 97.7% accuracy in material identification and 93.8% in texture detection. This study highlights VLC's potential in enhancing automated recognition systems across various applications.
Jiarong Li 0004, Chenxin Liang, Xiaojun Liang, Wenbo Ding 0001, Jian Song 0004, Xiao-Ping Zhang 0002
MobiSys5
2024 Demo: SolarSense: A Self-powered Ubiquitous Gesture Recognition System for Industrial Human-Computer Interaction
abstract
SolarSense is a self-powered sensing system for gesture recognition using solar cell arrays, thereby offering a sustainable approach to human-computer interaction (HCI) within industrial settings. The system effectively employed the sensing and energy harvesting capabilities of solar cells, achieving over 97.0% accuracy in recognizing diverse gestures. The design incorporates a low-power wireless data acquisition chip, a signal processing framework, and a user interface to realize robotic control and text input applications. SolarSense enhances HCI with its eco-friendly and user-centric approach, which is suitable for Internet of things (IoT) scenarios. Demo: https://youtu.be/RmPolChw_c4.
Jiarong Li 0004, Qinghao Xu, Qingyang Zhu, Zhancong Xu, Changshuo Ge, Liguang Ruan, H. Y. Fu 0001, Xiaojun Liang, Wenbo Ding 0001, Weihua Gui 0001, Xiao-Ping Zhang 0002
MobiSys10
2024 PerFedRec++: Enhancing Personalized Federated Recommendation with Self-Supervised Pre-Training
abstract
Federated recommendation systems employ federated learning techniques to safeguard user privacy by transmitting model parameters instead of raw user data between user devices and the central server. Nevertheless, the current federated recommender system faces three significant challenges: (1) data heterogeneity: the heterogeneity of users’ attributes and local data necessitates the acquisition of personalized models to improve the performance of federated recommendation; (2) model performance degradation: the privacy-preserving protocol design in the federated recommendation, such as pseudo item labeling and differential privacy, would deteriorate the model performance; (3) communication bottleneck: the standard federated recommendation algorithm can have a high communication overhead. Previous studies have attempted to address these issues, but none have been able to solve them simultaneously. In this article, we propose a novel framework, named PerFedRec++ , to enhance the personalized federated recommendation with self-supervised pre-training. Specifically, we utilize the privacy-preserving mechanism of federated recommender systems to generate two augmented graph views, which are used as contrastive tasks in self-supervised graph learning to pre-train the model. Pre-training enhances the performance of federated models by improving the uniformity of representation learning. Also, by providing a better initial state for federated training, pre-training makes the overall training converge faster, thus alleviating the heavy communication burden. We then construct a collaborative graph to learn the client representation through a federated graph neural network. Based on these learned representations, we cluster users into different user groups and learn personalized models for each cluster. Each user learns a personalized model by combining the global federated model, the cluster-level federated model, and its own fine-tuned local model. Experiments on three real-world datasets show that our proposed method achieves superior performance over existing methods.
Sichun Luo, Yuanzhang Xiao, Yang Liu 0165, Wenbo Ding 0001, Linqi Song
ACM Trans. Intell. Syst. Technol.5
2024 SAFARI: Sparsity-Enabled Federated Learning With Limited and Unreliable Communications
abstract
Federated learning (FL) enables edge devices to collaboratively learn a model in a distributed fashion. Many existing researches have focused on improving communication efficiency of high-dimensional models and addressing bias caused by local updates. However, most FL algorithms are either based on reliable communications or assuming fixed and known unreliability characteristics. In practice, networks could suffer from dynamic channel conditions and non-deterministic disruptions, with time-varying and unknown characteristics. To this end, in this paper we propose a sparsity-enabled FL framework with both improved communication efficiency and bias reduction, termed as SAFARI. It makes use of similarity among client models to rectify and compensate for bias that results from unreliable communications. More precisely, sparse learning is implemented on local clients to mitigate communication overhead, while to cope with unreliable communications, a similarity-based compensation method is proposed to provide surrogates for missing model updates. With respect to sparse models, we analyze SAFARI under bounded dissimilarity. It is demonstrated that SAFARI under unreliable communications is guaranteed to converge at the same rate as the standard FedAvg with perfect communications. Implementations and evaluations on the CIFAR-10 dataset validate the effectiveness of SAFARI by showing that it can achieve the same convergence speed and accuracy as FedAvg with perfect communications, with up to 60% of the model weights being pruned and a high percentage of client updates missing in each round of model updates.
Yuzhu Mao, Zihao Zhao 0001, Meilin Yang, Le Liang, Yang Liu 0165, Wenbo Ding 0001, Tian Lan 0001, Xiao-Ping Zhang 0002
IEEE Trans. Mob. Comput.6
2024 AQUILA: Communication Efficient Federated Learning With Adaptive Quantization in Device Selection Strategy
abstract
The widespread adoption of Federated Learning (FL), a privacy-preserving distributed learning methodology, has been impeded by the challenge of high communication overheads, typically arising from the transmission of large-scale models. Existing adaptive quantization methods, designed to mitigate these overheads, operate under the impractical assumption of uniform device participation. Additionally, these methods are limited in their adaptability due to the necessity of manual quantization level selection and often overlook biases inherent in local devices' data, thereby affecting the robustness of the global model. In response, this paper introduces AQUILA (adaptivequantization in device selection strategy), a novel adaptive framework devised to effectively handle these issues, enhancing the efficiency and robustness of FL. AQUILA integrates a sophisticated device selection method that prioritizes the quality and usefulness of device updates. Utilizing the exact global model stored by devices enables a more precise device selection criterion, reduces model deviation, and limits the need for hyperparameter adjustments. Furthermore, AQUILA presents an innovative quantization criterion, optimized to improve communication efficiency while assuring model convergence. Our experiments demonstrate that AQUILA significantly decreases communication costs compared to existing methods, while maintaining comparable model performance across diverse non-homogeneous FL settings, such as Non-IID data and heterogeneous model architectures.
Zihao Zhao 0001, Yuzhu Mao, Zhenpeng Shi, Yang Liu 0165, Tian Lan 0001, Wenbo Ding 0001, Xiao-Ping Zhang 0002
IEEE Trans. Mob. Comput.6
2024 FedHAP: Federated Hashing With Global Prototypes for Cross-Silo Retrieval
abstract
Deep hashing has been widely applied in large-scale data retrieval due to its superior retrieval efficiency and low storage cost. However, data are often scattered in data silos with privacy concerns, so performing centralized data storage and retrieval is not always possible. Leveraging the approach of federated learning (FL) to perform deep hashing is a recent research trend. However, existing frameworks mostly rely on the aggregation of the local deep hashing models, which are trained by performing similarity learning with local skewed data only. Therefore, they cannot work well for non-IID clients in a real federated environment. To overcome these challenges, we propose a novel federated hashing framework that enables participating clients to jointly train a shared deep hashing model by leveraging the class-wise prototypical hash codes. Globally, sharing global prototypes with only one prototypical hash code per class assists in learning consistent code distributions across clients while minimizing the cost of communication and privacy. Locally, the use of the global prototypical hash codes are maximized by jointly training a discriminator network and the local hashing network. Extensive experiments on benchmark datasets are conducted to demonstrate that our method can significantly improve the performance of the deep hashing model in the federated environments with non-IID data distributions.
Meilin Yang, Jian Xu 0016, Wenbo Ding 0001, Yang Liu 0165
IEEE Trans. Parallel Distributed Syst.3
2024 M$^{3}$Tac: A Multispectral Multimodal Visuotactile Sensor With Beyond-Human Sensory Capabilities
abstract
To realize the exquisite interaction and precise manipulation for the robot, in this article, we propose a multispectral multimodal visuotactile sensor named M$^{3}$Tac, which combines visible, near-infrared, and mid-infrared imaging technologies for the first time and can exceed the sensing ability of human skin in terms of resolution (719 pixels/cm$^{2}$), temperature sensing range (−20–130$^\text{o}$C), etc. The M$^{3}$Tac cannot only realize high-quality sensing of deformation, texture, force, stickiness, and temperature comparable to human skin but also can realize proximity sensing that is lacking for human skin. To achieve this, we not only design a multispectral imaging system with an elastic film whose light penetrability can be regulated by the brightness of the light, but also develop corresponding algorithms, including the pixel-level force sensing with finite element method (accuracy:$\pm$0.023N), the proximity perception (accuracy:$\pm$3.8 mm), the 3-D reconstruction (accuracy: 0.33 mm), the super-resolution temperature sensing (accuracy:$\pm 0.3^\text{o}$C), the multimodal fusion classification (accuracy: 98%), and the stickiness recognition (accuracy: 98%). Finally, we conduct experiments to verify the effectiveness and application potential of our research.
Shoujie Li, Haixin Yu, Guoping Pan, Huaze Tang, Jiawei Zhang 0012, Linqi Ye, Xiao-Ping Zhang 0002, Wenbo Ding 0001
IEEE Trans. Robotics8
2023 STEV: Stretchable Triboelectric E-skin enabled Proprioceptive Vibration Sensing for Soft Robot
abstract
Vibration perception is essential for robotic sensing and dynamic control. Nevertheless, due to the rigorous demand for sensor conformability and stretchability, enabling soft robots with proprioceptive vibration sensing remains challenging. This paper proposes a novel liquid metal-based stretchable e-skin via a kirigami-inspired design to enable soft robot proprioceptive vibration sensing. The e-skin is fabricated into 0.1mm ultrathin thickness, ensuring its negligible influence on the overall stiffness of the soft robot. Moreover, the working mechanism of the e-skin is based on the ubiquitous triboelectrification effect, which transduces mechanical stimuli without external power supply. To demonstrate the practicability of the e-skin, we built a soft gripper consisting of three soft robotic fingers with proprioceptive vibration sensing. Our experiment shows that the gripper can accurately distinguish the grain category (six grains with the same mass, 99.9% accuracy) and the packaging quality (100% accuracy) by simply shaking the gripped bottle. In summary, a soft robotic proprioceptive vibration sensing solution is proposed; it helps soft robots to have a more comprehensive awareness of their self-state and may inspire further research on soft robotics.
Kai-Chong Lei, Huaze Tang, Shoujie Li, Yuan Dai, Wenbo Ding 0001, Xiao-Ping Zhang 0002
ICRA6
2023 Poster Abstract: TENG-enabled Self-powered Human-machine Interfaces for the Metaverse
abstract
Human-machine interface (HMI) of high degrees of freedom (DoF) is one of the most critical bases of the metaverse. The ideal HMI for the metaverse should be cheap, robust, customizable, and ergonomically friendly. In light of this, we propose a triboelectric nanogenerator (TENG)-based sensing system. We developed a low-cost, soft, light, and customizable TENG sensor to collect data from the human body. We then used an artificial neural network (ANN) to obtain the corresponding human motion from collected sensory data. The effectiveness of the proposed system is demonstrated with experiments of a working prototype.
Haoyang Wang 0012, Fanhang Man, Yuxuan Liu 0010, Xinlei Chen, Wenbo Ding 0001
IPSN6
2023 Visuotactile Sensor Enabled Pneumatic Device Towards Compliant Oropharyngeal Swab Sampling
abstract
Manual oropharyngeal (OP) swab sampling is an intensive and risky task. In this article, a novel OP swab sampling device of low cost and high compliance is designed by combining the visuotactile sensor and the pneumatic actuator-based gripper. Here, a concave visuotactile sensor called CoTac is first proposed to address the problems of high cost and poor reliability of traditional multi-axis force sensors. Besides, by imitating the doctor's fingers, a soft pneumatic actuator with a rigid skeleton structure is designed, which is demonstrated to be reliable and safe via finite element modeling and experiments. Furthermore, we propose a sampling method that adopts a compliant control algorithm based on the adaptive virtual force to enhance the safety and compliance of the swab sampling process. The effectiveness of the device has been verified through sampling experiments as well as in vivo tests, indicating great application potential. The cost of the device is around 30 US dollars and the total weight of the functional part is less than 0.1 kg, allowing the device to be rapidly deployed on various robotic arms.
Shoujie Li, Mingshan He, Wenbo Ding 0001, Linqi Ye, Xueqian Wang 0001, Junbo Tan, Jinqiu Yuan, Xiao-Ping Zhang 0002
IROS3
2023 Autonomous Swarm Robot Coordination via Mean-Field Control Embedding Multi-Agent Reinforcement Learning
abstract
The learning approaches of designing a controller to guide the collective behavior of swarm robots have gained significant attention in recent years. However, the scalability of swarm robots and their inherent stochasticity complicate the control problem due to increasing complexity, unpredictability, and non-linearity. Despite considerable progress made in swarm robotics, addressing these challenges remains a significant issue. In this work, we model the stochastic dynamics of a swarm robot system and then propose a novel control framework based on a mean-field control (MFC) embedding multi-agent reinforcement learning (MARL) approach named MF-MARL to deal with these challenges. While MARL is able to deal with stochasticity statistically, we integrate MFC, allowing MF-MARL to cope with large-scale robots. Moreover, we apply statistical moments of robots' state and control action to discretize continuous input and enable MF-MARL to be applied in continuous scenarios. To demonstrate the effectiveness of MF-MARL, we evaluate the performance of the robots on a specific swarm simulation platform. The experimental results show that our algorithm outperforms the traditional algorithms both in navigation and manipulation tasks. Finally, we demonstrate the adaptability of the proposed algorithm through the component failure test.
Huaze Tang, Hengxi Zhang, Zhenpeng Shi, Xinlei Chen, Wenbo Ding 0001, Xiao-Ping Zhang 0002
IROS5
2023 System Architecture and Evaluation Method for NOMA-Based Integrated Broadcast and Communication Networks
abstract
The non-orthogonal multiple access (NOMA) method is seen as one of the advantageous strategies for radio access that promises improved spectral efficiency in the fifth generation mobile networks (5G). This paper proposes an integrated broadcast and communication networks architecture based on NOMA, in which the physical channels of the cellular base stations are divided into two layers: one layer brings reinforced broadcast signals to the weakest areas of the broadcast tower and forms a single-frequency network transmission with it; the other layer provides diverse unicast services to users. In order to evaluate the spectral efficiency and coverage performance of the above system, an optimization problem based on the integrated spectrum revenue of the system is proposed. To maximize the objective function, the power allocation factors of the cellular base stations are optimized using the genetic algorithm. The simulation results validated that the system architecture is capable of integrating broadcast towers and cellular base stations to provide more reliable coverage of broadcast service and improve spectrum efficiency.
Linhui Chang, Fang Yang 0001, Jian Song 0004, Changshuo Ge, Wenbo Ding 0001
IWCMC5
2023 Every Parameter Matters: Ensuring the Convergence of Federated Learning with Dynamic Heterogeneous Models Reduction
abstract
Cross-device Federated Learning (FL) faces significant challenges where low-end clients that could potentially make unique contributions are excluded from training large models due to their resource bottlenecks. Recent research efforts have focused on model-heterogeneous FL, by extracting reduced-size models from the global model and applying them to local clients accordingly. Despite the empirical success, general theoretical guarantees of convergence on this method remain an open question. This paper presents a unifying framework for heterogeneous FL algorithms with online model extraction and provides a general convergence analysis for the first time. In particular, we prove that under certain sufficient conditions and for both IID and non-IID data, these algorithms converge to a stationary point of standard FL for general smooth cost functions. Moreover, we introduce the concept of minimum coverage index, together with model reduction noise, which will determine the convergence of heterogeneous federated learning, and therefore we advocate for a holistic approach that considers both factors to enhance the efficiency of heterogeneous federated learning.
Hanhan Zhou, Tian Lan 0001, Guru Venkataramani, Wenbo Ding 0001
NeurIPS4
2023 Mean-Field-Aided Multiagent Reinforcement Learning for Resource Allocation in Vehicular Networks
abstract
As one technique for autonomous driving, vehicular networks can achieve high efficiency with vehicle-and-infrastructure cooperation, bringing high safety and many value-added services. To achieve higher communication efficiency, much effort has been done to cope with the resource allocation issues for vehicular networks. Nevertheless, due to the strong nonconvexity and nonlinearity, the classical joint resource allocation problem in vehicular networks is typically NP-hard. The multiagent reinforcement learning (MARL) has emerged as a promising solution to tackle this challenge but its stability and scalability are not satisfactory when the amount of vehicles gets increased. In this article, we mainly investigate the issue of joint spectrum and power allocation in vehicular communication networks, and carefully consider the interactions between the vehicles and environment by incorporating the cooperative stochastic game theory with MARL, named complete-game MARL (CG-MARL), to achieve a better convergence and stability with the theoretical computational complexity$\mathcal {O}(n^{N})$with$n$denoting the dimension of action space and$N$denoting the number of V2X Vehicular. Furthermore, the mean-field game (MFG) theory is employed to further enhance the MARL for decreasing the horrible computing resource consumption caused by the CG-MARL to$\mathcal {O}(n^{2})$while maintaining an approximate performance. The simulation results demonstrate that the proposed mean-field-aided MARL (MF-MARL) for vehicular network resource allocation can achieve 95% near-optimal performance with much lower complexity, which indicates its significant potentials in the scenarios with massive and dense vehicles.
Hengxi Zhang, Chengyue Lu, Huaze Tang, Xiaoli Wei, Le Liang, Ling Cheng 0001, Wenbo Ding 0001, Zhu Han 0001
IEEE Internet Things J.7
2023 Visual-Tactile Fusion for Transparent Object Grasping in Complex Backgrounds
abstract
The grasping of transparent objects is challenging but of significance to robots. In this article, a visual–tactile fusion framework for transparent object grasping in complex backgrounds is proposed, which synergizes the advantages of vision and touch, and greatly improves the grasping efficiency of transparent objects. First, we propose a multiscene synthetic grasping dataset named SimTrans12 K together with a Gaussian-mask annotation method. Next, based on the TaTa gripper, we propose a grasping network named transparent object-grasping convolutional neural network for grasping position detection, which shows good performance in both synthetic and real scenes. Inspired by human grasping, a tactile calibration method and a visual–tactile fusion classification method are designed, which improve the grasping success rate by 36.7% compared with direct grasping and the classification accuracy by 39.1%. Furthermore, the tactile height sensing module and the tactile position exploration module are added to solve the problem of grasping transparent objects in irregular and visually undetectable scenes. The experimental results demonstrate the validity of the framework.
Shoujie Li, Haixin Yu, Wenbo Ding 0001, Houde Liu, Linqi Ye, Chongkun Xia, Xueqian Wang 0001, Xiao-Ping Zhang 0002
IEEE Trans. Robotics3
2022 Heterogeneous Mean-Field Multi-Agent Reinforcement Learning for Communication Routing Selection in SAGI-Net
abstract
The utilization of heterogeneous end devices such as the low earth orbit (LEO) satellite, unmanned aerial vehicles (UAVs) and ground users (GUs) deployed at different altitudes, known as the space-air-ground integrated network (SAGI-Net), can be quite promising towards a bunch of advanced applications. Whereas, the energy efficiency of the SAGI-Net communication system is a key criterion needed to be improved urgently in consideration that the inappropriate communication routing will undoubtedly cause a huge communication energy cost of the system especially with a large number of communication devices inside. In this paper, we proposed a novel communication routing selection model for the SAGI-Net system and established a heterogeneous multi-agent reinforcement learning (HMF-MARL) framework to optimize the communication energy efficiency of this system, where the mean-field theory was introduced to enhance the ability of classic MARL method while still maintaining a relatively low computational complexity. The experiment results show that the capacity of the heterogeneous multi-agent system has been improved by nearly 80% using the proposed HMF-MARL method compared with the classic MARL one, which hopefully shows the potential value on the implementation of the SAGI-Net system in the future.
Hengxi Zhang, Huaze Tang, Yuanquan Hu, Xiaoli Wei, Chenye Wu, Wenbo Ding 0001, Xiao-Ping Zhang 0002
VTC Fall6
2022 TCACNet: Temporal and channel attention convolutional network for motor imagery classification of EEG-based BCI
abstract
Brain–computer interface (BCI) is a promising intelligent healthcare technology to improve human living quality across the lifespan, which enables assistance of movement and communication, rehabilitation of exercise and nerves, monitoring sleep quality, fatigue and emotion. Most BCI systems are based on motor imagery electroencephalogram (MI-EEG) due to its advantages of sensory organs affection, operation at free will and etc. However, MI-EEG classification, a core problem in BCI systems, suffers from two critical challenges: the EEG signal’s temporal non-stationarity and the nonuniform information distribution over different electrode channels. To address these two challenges, this paper proposes TCACNet, a temporal and channel attention convolutional network for MI-EEG classification. TCACNet leverages a novel attention mechanism module and a well-designed network architecture to process the EEG signals. The former enables the TCACNet to pay more attention to signals of task-related time slices and electrode channels, supporting the latter to make accurate classification decisions. We compare the proposed TCACNet with other state-of-the-art deep learning baselines on two open source EEG datasets. Experimental results show that TCACNet achieves 11.4% and 7.9% classification accuracy improvement on two datasets respectively. Additionally, TCACNet achieves the same accuracy as other baselines with about 50% less training data. In terms of classification accuracy and data efficiency, the superiority of the TCACNet over advanced baselines demonstrates its practical value for BCI systems.
Rongye Shi, Qianxin Hui, Susu Xu, Shuai Wang 0049, Rui Na, Ying Sun 0012, Wenbo Ding 0001, Dezhi Zheng, Xinlei Chen
Inf. Process. Manag.8
2022 Communication-Efficient Federated Learning with Adaptive Quantization
abstract
Federated learning (FL) has attracted tremendous attentions in recent years due to its privacy-preserving measures and great potential in some distributed but privacy-sensitive applications, such as finance and health. However, high communication overloads for transmitting high-dimensional networks and extra security masks remain a bottleneck of FL. This article proposes a communication-efficient FL framework with an Adaptive Quantized Gradient (AQG), which adaptively adjusts the quantization level based on a local gradient’s update to fully utilize the heterogeneity of local data distribution for reducing unnecessary transmissions. In addition, client dropout issues are taken into account and an Augmented AQG is developed, which could limit the dropout noise with an appropriate amplification mechanism for transmitted gradients. Theoretical analysis and experiment results show that the proposed AQG leads to 18% to 50% of additional transmission reduction as compared with existing popular methods, including Quantized Gradient Descent (QGD) and Lazily Aggregated Quantized (LAQ) gradient-based methods without deteriorating convergence properties. Experiments with heterogenous data distributions corroborate a more significant transmission reduction compared with independent identical data distributions. The proposed AQG is robust to a client dropping rate up to 90% empirically, and the Augmented AQG manages to further improve the FL system’s communication efficiency with the presence of moderate-scale client dropouts commonly seen in practical FL scenarios.
Yuzhu Mao, Zihao Zhao 0001, Guangfeng Yan, Yang Liu 0165, Tian Lan 0001, Linqi Song, Wenbo Ding 0001
ACM Trans. Intell. Syst. Technol.7
2019 Resource Allocation in Heterogeneous Network With Visible Light Communication and D2D: A Hierarchical Game Approach
abstract
Utilizing the illuminating light-emitting diode (LED) for signal transmission, the visible light communication (VLC) offloads the congested traffic in licensed spectrum while providing users with high quality of service. However, in shaded areas, the performance of VLC is affected. In this paper, we apply Device-to-Device (D2D) technology on the VLC network, where some mobile users are able to relay the data transmission to nearby end mobile users. Considering the autonomous behaviors of mobile users as relays (RUs), cellular service provider (CSP) and VLC service provider (VLCSP), we propose a hierarchical game framework and analyze the distributive strategies for each individual. In the game, the data packet size, the price of licensed spectrum and data rates in possible data transmission routes are sequentially determined with equilibrium solutions, and each VLC transmitter (VLCT) determines the optimal data transmission route with stable matching. Simulations depict the high performance when combining VLC and D2D with the proposed strategies.
Huaqing Zhang 0001, Wenbo Ding 0001, Fang Yang 0001, Jian Song 0004, Zhu Han 0001
IEEE Trans. Commun.2
2018 Compressive Sensing over Graphs Based Inter-Community Detection Scheme in Mobile Social Networks
abstract
Recently, mobile social networks (MSNs) has been playing an increasingly large proportion in people's daily life, and consequently attracted enormous amount of researches in this area, including network data collection, user behavior analysis and so on. Many of them are based on the community structure of the MSNs, the detection of which has attracted academic attention. In this paper, we propose an inter-community detection scheme under the framework of the emerging compressive sensing (CS) over graphs. Firstly, we extract the social structure among users by calculating the probability of two users encountering with each other. Then, the proposed scheme utilizes the encounter probability and the edge-clustering coefficient to define a novel additive property to make the CS algorithm applicable. Simulation results demonstrate that the proposed detection scheme can detect inter- community links more accurately than the conventional random walk based method.
Tengjiao Wang 0001, Jingbo Tan, Wenbo Ding 0001, Yanru Zhang, Fang Yang 0001, Jian Song 0004, Zhu Han 0001
ICC3
2018 Intercommunity Detection Scheme for Social Internet of Things: Compressive Sensing Over Graphs Approach
abstract
As a cluster between Internet of Things (IoT) and mobile social networks (MSNs), social IoT allows close interaction between humans and things. Many applications of IoT and MSN are based on the community structure to realize efficient services, and thus the detection of the community structure has become a key problem and attracted much academic attention. In this paper, a compressive sensing (CS) over graphs-based intercommunity detection scheme is proposed for the social IoT. By exploiting the probability of two nodes encountering each other, the encounter probability and the edge-clustering coefficient are utilized to define a novel metric for each connection and construct the measurement matrix. Then a CS-based detection algorithm is proposed to detect the intercommunity links. Moreover, the edge-clustering coefficient is exploited as the prior information to further improve accuracy and reduce complexity. Simulations show that the proposed detection scheme outperforms the conventional random walk-based scheme in the social IoT.
Tengjiao Wang 0001, Jingbo Tan, Wenbo Ding 0001, Yanru Zhang, Fang Yang 0001, Jian Song 0004, Zhu Han 0001
IEEE Internet Things J.3
2017 Integrated power line and visible light communication system compatible with multi-service transmission
abstract
As the shortage of spectrum is becoming a serious problem to the wireless communication, visible light communication (VLC) is attracting widespread attention. However, the access to the backbone network is likely to completely modify the communication network, thus leading to a great cost. Fortunately, the technology of power line communication (PLC) is a perfect match to VLC, which provides power supply and connection for VLC to the backbone network. In this study, the authors take an integration of PLC and VLC into consideration and mainly focus on a communication framework with multi‐service. In the case of not making a big modification to the origin network, a direct re‐transmission concept is considered and three implementation schemes which differ in frequency domain, time domain, and bit division multiplexing (BDM) are proposed. The demonstration platform and the simulation results suggest that the proposed schemes have the ability to support multi‐service transmission with no big modification of current power line network. Moreover, the BDM scheme can accommodate different requirements of quality of service with better performance than conventional multi‐service schemes.
Xu Ma 0003, Junnan Gao, Fang Yang 0001, Wenbo Ding 0001, Hui Yang 0007, Jian Song 0004
IET Commun.4
2016 Impulsive Noise Cancellation for MIMO-OFDM PLC Systems: A Structured Compressed Sensing Perspective
abstract
In this paper, a novel impulsive noise (IN) cancel- lation scheme based on structured compressed sensing (SCS) for multiple input multiple output orthogonal frequency division multiplexing (MIMO-OFDM) power line communication systems is proposed. The SCS theory is introduced to IN recovery for the first time in this paper to the best of the authors' knowledge, and the gap of lack of research on the IN mitigation for MIMO PLC systems is filled. First, the measurements matrix of the IN is obtained, and the SCS optimization framework is formulated through the proposed spatially multiple measuring method, by fully exploiting the spatial correlation of the IN signals at different receive antennas. To efficiently reconstruct the IN signal, an enhanced SCS-based greedy algorithm, structured a priori aided sparsity adaptive matching pursuit (SPA-SAMP), is proposed, which significantly improves the accuracy and robustness compared with the state-of-art methods. Theoretical analysis and computer simulations validate that the proposed scheme outperforms conventional methods in the typical MIMO PLC system.
Sicong Liu 0002, Fang Yang 0001, Wenbo Ding 0001, Jian Song 0004, Zhu Han 0001
GLOBECOM3
2016 Dynamic Matching Based Distributed Spectrum Trading in Multi-Radio Multi-Channel CRNs
abstract
Spectrum trading not only improves spectrum utilization but also benefits both secondary users (SUs) with more accessing opportunities and primary users (PUs) with monetary gains. Although existing centralized designs consider the special features of spectrum trading (e.g., frequency reuse, interference mitigation, multi-radio multi- channel transmissions, etc.), they have to deploy new infrastructure, deal with extra control overhead, have scalability issues, and may miss many instantaneous opportunities. To address those issues, in this paper, we propose a novel dynamic matching based distributed spectrum trading (DMDST) scheme in multi-radio multi- channel cognitive radio (CR) networks. We employ conflict graph to characterize interference relationship among SUs with multiple CR radios, and formulate the centralized PUs' revenue maximization problem under multiple constrains. In view of the NP- hardness of solving the problem and no existence of centralized entity, we develop the DMDST algorithms based on conflict graph observed by PUs, solve the problem via dynamic matching with evolving preferences, and prove its stability. Through extensive simulations, we show that the results of proposed DMDST algorithm is close to the optimal one and outperforms other distributed algorithms without considering spectrum reuse.
Jingyi Wang 0002, Wenbo Ding 0001, Yuanxiong Guo, Chi Zhang 0001, Miao Pan, Jian Song 0004
GLOBECOM2
2016 A Hierarchical Game Approach for Visible Light Communication and D2D Heterogeneous Network
abstract
The visible light communication (VLC) technology is able to provide indoor users with high quality of service and low cost. However, the performance of the VLC is limited in shaded or strong sun- light areas. In this paper, we combine the VLC network with Device-to-Device (D2D) technology. We consider that some mobile users accessible to the VLC network is able to relay the transmitted data to the nearby mobile users. We suppose all the mobile users as relays (MUaRs), cellular service provider (CSP) and VLC service provider (VLCSP) are autonomous individuals, and propose a hierarchical game to analyze the optimal strategies for each of them. In the game, the VLCSP determines the data transmission route and data packet size first. Based on the behaviors of the VLCSP, the CSP determines the prices of wireless licensed spectrum. According to the behaviors of the VLCSP, the CSP and the MUaRs in the neighborhood, all MUaRs determines their transmission data rates for optimal utility. Therefore, there is a graphical game among all MUaRs. When all MUaRs consider the corresponding reactions of other MUaRs in the neighborhood and achieve Nash Equilibrium, the Stackelberg Equilibriums can be achieved in the Stackelberg games between the VLCSP and the CSP, between the CSP and all MUaRs and between the VLCSP and all MUaRs. Simulation results show the correctness of the analysis and the VLC network with multi-hop D2D relay is able to bring high profits for the CSP.
Huaqing Zhang 0001, Wenbo Ding 0001, Jian Song 0004, Zhu Han 0001
GLOBECOM2
2016 NBI cancellation for smart grid communications: A block sparse Bayesian learning perspective
abstract
A block sparse Bayesian learning (BSBL) based approach of narrowband interference (NBI) cancellation for cyclic prefixed orthogonal frequency division multiplexing based smart grid communications is proposed in this paper. The BSBL theory is firstly introduced to recover the practical block sparse NBI with a frequency offset compared with the sub-carriers. The block sparse representation of the NBI is constituted through the proposed temporal differential measuring approach. A BSBL based method, estimated partitioned BSBL, is proposed for NBI recovery. The intra-block correlation is firstly considered to facilitate the recovery of block sparse NBI. Reported simulation results demonstrate that the proposed methods are effective and significantly outperform conventional counterparts.
Sicong Liu 0002, Fang Yang 0001, Wenbo Ding 0001, Jian Song 0004
ICC3
2016 A cost-effective approach for ubiquitous broadband access based on hybrid PLC-VLC system
abstract
Visible light communication (VLC) using the light emitting diode (LED) will become an appealing alternative to the radio frequency communication technology for indoor wireless broadband access. However, VLC needs a ubiquitous network as its backbone to avoid becoming an information isolated island. Power line communication (PLC) systems could easily solve the informative problem of VLC while powering the LED lamps at the same time, which is considered as a good partner of VLC for the cost-effective implementation. In this paper, a novel and cost-effective framework of ubiquitous indoor broadband access based on deeply integrated VLC and PLC technology with only low-cost modification to the current infrastructure is therefore proposed. The broadband access network supports duplex transmission through each LED using the decode-and-forward (DF) working mode. This paper will present our recent research progress in this area, including a prototyping of duplex voice communications network based on hybrid PLC and VLC in our lab. Our research and development plan in this area for the near future will also be covered.
Jian Song 0004, Sicong Liu 0002, Guangxin Zhou, Bingyan Yu, Wenbo Ding 0001, Fang Yang 0001, Hongming Zhang 0010, Xun Zhang 0002, Amara Amara
ISCAS5
2016 Structured compressive sensing-based non-orthogonal time-domain training channel state information acquisition for multiple input multiple output systems
abstract
In practical multiple input multiple output (MIMO) systems, accurate knowledge of the channel state information (CSI) is a prerequisite to guarantee the system performance. The conventional CSI acquisition methods for MIMO system usually rely on the orthogonal (either time‐ or frequency‐domain) training sequences (TSs) to estimate the channel associated with each transmit–receive antenna pair, which is not spectrally efficient. This study proposes a non‐orthogonal time‐domain training‐based CSI acquisition approach for MIMO systems under the framework of structured compressive sensing. By exploiting the spatial–temporal correlations of the sparse MIMO channels, a spatially–temporally spARsity‐adaptivE‐simultaneous orthogonal matching pursuit algorithm is proposed, which could use the inter‐block interference free region of very small dimension within the received TS to recover the multiple channels. Furthermore, the proposed algorithm could utilise the priori channel partial common support to improve the recovery probability and reduce the complexity. Simulation results show that the proposed scheme has better performance and higher spectral efficiency than the conventional MIMO schemes, which might be an appealing solution for the future wireless communications.
Wenbo Ding 0001, Fang Yang 0001, Sicong Liu 0002, Jian Song 0004
IET Commun.1
2016 Spectrally Efficient CSI Acquisition for Power Line Communications: A Bayesian Compressive Sensing Perspective
abstract
Power line communication (PLC) techniques present a no extra wire solution for the communication purpose in a smart grid due to the ubiquity and low cost. Moreover, the through-the-grid property of PLC has naturally extended its possible applications, including but not limited to the automatic meter reading, line quality monitoring, online diagnostics, and network tomography. To guarantee the performance of communications as well as other applications in PLC systems, accurate channel state information (CSI) acquisition should be performed regularly. However, the conventional pilot-based CSI acquisition approaches in PLC systems have not made full use of the channel characteristics and hence suffer from a low spectral efficiency. In this paper, by exploiting the parametric sparsity and discretizing the electrical length in the well-known PLC channel model, we formulate the non-sparse (either time domain or frequency domain) PLC channel into a compressive sensing (CS) applicable problem. Furthermore, we propose a spectrally efficient CSI acquisition scheme under the framework of Bayesian CS and extend it to the multiple-input multiple-output PLC by investigating the channel spatial correlation. Compared with the existing sparse CSI acquisition schemes for PLC, such as the annihilating filter-based and the estimating signal parameters via rotational invariance technique-based ones, the proposed scheme has better mean square error performance and noise robustness.
Wenbo Ding 0001, Yang Lu 0007, Fang Yang 0001, Wei Dai 0001, Pan Li 0005, Sicong Liu 0002, Jian Song 0004
IEEE J. Sel. Areas Commun.1
2016 M3-STEP: Matching-Based Multi-Radio Multi-Channel Spectrum Trading With Evolving Preferences
abstract
Spectrum trading not only improves spectrum utilization but also benefits both secondary users (SUs) with more accessing opportunities and primary users (PUs) with monetary gains. Although the existing centralized designs consider the special features of spectrum trading (e.g., frequency reuse, interference mitigation, multi-radio multi-channel transmissions, and so on), they still have to face many practical but challenging issues, such as the new infrastructure deployment, the extra control overhead, and the scalability issues. To address those issues, in this paper, we propose a novel matching-based multi-radio multi-channel spectrum trading (M3-STEP) scheme in cognitive radio (CR) networks. We employ conflict graph to characterize the interference relationship among SUs with multiple CR radios, and formulate the centralized PUs' revenue maximization problem under multiple constrains. In view of the NP-hardness of solving the problem and no existence of centralized entity, we develop the M3-STEP algorithms based on conflict graph observed by PUs, solve the problem via dynamic matching with evolving preferences, and prove its pairwise stability. Simulation results show that the proposed M3-STEP algorithm achieves close to optimal performance and outperforms other distributed algorithms without considering spectrum reuse.
Jingyi Wang 0002, Wenbo Ding 0001, Yuanxiong Guo, Chi Zhang 0001, Miao Pan, Jian Song 0004
IEEE J. Sel. Areas Commun.2
2015 Sparse Bayesian Learning Based Symbol Detection for Generalised Spatial Modulation in Large-Scale MIMO Systems
abstract
Generalised spatial modulation (GSM) is extended from the concept of spatial modulation (SM). Due to its superior energy efficiency to classical multiple-input multiple-output (MIMO) techniques, GSM has attracted plenty of interest under the study of large-scale MIMO systems. In this paper, we propose a low-complexity symbol detector for GSM based on the framework of sparse Bayesian learning (SBL), which effectively recovers the GSM symbols by exploiting the inherent sparsity property of GSM. Compared to other sparsity-based symbol detectors, e.g., detectors based on the basis pursuit (BP) method and the orthogonal matching pursuit (OMP), the SBL-based detector is capable of achieving a superior reconstruction accuracy using less receive antennas. Numerical simulations are performed to substantiate the performance of the proposed detector.
Longzhuang He, Jintao Wang 0001, Wenbo Ding 0001, Jian Song 0004
GLOBECOM3
2015 Sparse channel state information acquisition for power line communications
abstract
Power line communication (PLC) systems present a “no new wires” solution for the telecommunication access with the additional advantages of ubiquitous availability, easy installation and cost effectiveness. In this paper, by exploiting the parametric sparsity of the PLC channels, we propose a robust sparse channel state information (CSI) acquisition scheme under the framework of Bayesian compressive sensing (CS), which could significantly reduce the pilot overhead and improve the spectral efficiency. Compared to the current sparse CSI acquisition schemes for PLC, including those based on annihilating filter and estimating signal parameters via rotational invariance techniques (ESPRIT) algorithm, the proposed scheme is numerically demonstrated to have better mean squared error (MSE) performance. Furthermore, the proposed method is an appealing solution for practical PLC system, since the performance degradation due to the parameter discretization is not significant.
Wenbo Ding 0001, Yang Lu 0007, Fang Yang 0001, Wei Dai 0001, Jian Song 0004
ICC1
2015 Novel Integrated power line and visible light communication system with bit division multiplexing
abstract
Visible light communication (VLC), as an appealing alternative to the classical radio frequency communication technologies, has a bright application prospect for future wireless communications. In this paper, we propose an integrated power line communication (PLC) and VLC system with bit division multiplexing (BDM). By the integration of VLC with PLC, the lamp can work as light source as well as access point, and the architecture of VLC network could be significantly simplified. Meanwhile, the BDM scheme could simultaneously support multi-service, e.g., data transmission and navigation services as global and local services, respectively, in a natural single frequency network (SFN) structure with low complexity, high spectral efficiency, and easy switching protocol as well. Simulation results demonstrate that this proposed scheme can accommodate different requirements of quality of service (QoS) and outperform conventional multi-service schemes with time-frequency channel resource allocation approaches.
Junnan Gao, Fang Yang 0001, Wenbo Ding 0001
IWCMC3
2015 A priori aided compressive sensing approach for impulsive noise reconstruction
abstract
In this paper, a novel impulsive noise (IN) cancellation scheme based on priori aided compressive sensing (CS) for OFDM-based communications systems is proposed. The IN is reconstructed from the frequency-domain measurements at the null sub-carriers based on the CS theory using greedy algorithms. With the aid of the a priori partial support obtained from the proposed time-domain thresholding method, we propose the enhanced greedy algorithm of priori aided sparsity adaptive matching pursuit (PA-SAMP) to improve the accuracy and robustness of the IN recovery. Theoretical analysis and computer simulations validate that the proposed method outperforms conventional CS-based and other classical IN mitigation methods for OFDM-based communications systems.
Sicong Liu 0002, Fang Yang 0001, Wenbo Ding 0001, Jian Song 0004
IWCMC3
2015 A positioning compatible multi-service transmission system based on the integration of VLC and PLC
abstract
As the shortage of spectrum is becoming a serious problem to the wireless communication, visible light communication (VLC) is attracting widespread attention. However, the accession to the backbone network is likely to completely modify the communication network, thus leading to a great cost. Fortunately, the technology of power line communication (PLC) is a perfect match to VLC, which provides power supply and connection for VLC to the backbone network. In this paper, we take an integration of VLC and PLC into consideration and mainly focus on a communication framework with multifunction. In the case of not making a big modification to the origin network, a direct re-transmission concept is considered and two implementation schemes which differ in frequency and time domain are proposed. A demonstration platform is built which suggests that both schemes have the ability to support positioning and multi-service transmission in a high data rate with no big modification of current power line network.
Xu Ma 0003, Wenbo Ding 0001, Fang Yang 0001, Hui Yang 0007, Jian Song 0004
IWCMC2
2015 Approach to suppress out-of-band emission for dual pseudo noise padded time-domain synchronous-orthogonal frequency division multiplexing systems
abstract
The dual pseudo noise padded (DPNP) time‐domain synchronous‐orthogonal frequency division multiplexing (TDS‐OFDM) which utilises the second pseudo noise (PN) sequence for channel estimation is able to reduce the complexity and improve the accuracy of channel estimation compared to the classical TDS‐OFDM. However, because of the duplicate PN sequence, DPNP TDS‐OFDM will suffer from severer out‐of‐band emission, which is a headache issue for the conventional TDS‐OFDM. In this study, a novel out‐of‐band suppression approach based on a new PN design criterion as well as a windowing operation is proposed for the DPNP TDS‐OFDM systems. The frame structure and system model are also modified to accommodate to the approach, while the overlap‐and‐add approach is adopted to guarantee the spectral efficiency. Both simulation and experimental results show that this approach could achieve satisfactory performance with much less complexity and no spectral efficiency loss compared to the classical DPNP TDS‐OFDM system. In addition, an iterative channel estimation method is also proposed for the modified frame structure to combat against the long channel delay spread.
Wenbo Ding 0001, Fang Yang 0001, Sicong Liu 0002, Jian Song 0004
IET Commun.1
2014 Simultaneous time-frequency channel estimation based on compressive sensing for OFDM system
abstract
Conventional time domain synchronous orthogonal frequency division multiplexing (TDS-OFDM) systems have the difficulty to support 256QAM or higher-order modulations and suffer from performance loss, especially under the severe fading channels with long delays. In this letter, a simultaneous time-frequency channel estimation method based on compressive sensing (CS) is proposed to solve this problem. First, the auxiliary channel information is obtained by exploiting the signal structure of TDS-OFDM. Then, a very few frequency-domain pilots in the OFDM block are used to acquire the accurate channel impulse response under the framework of CS. Besides, the auxiliary channel information is utilized to further reduce the complexity of the classical CS algorithm. Simulation results show that the proposed scheme can well support 256QAM under fading channels with long delays and have better performance than the conventional schemes.
Wenbo Ding 0001, Fang Yang 0001, Chao Zhang 0009, Linglong Dai, Jian Song 0004
GLOBECOM1
2014 Energy-efficient orthogonal frequency division multiplexing scheme based on time-frequency joint channel estimation
abstract
Time‐domain synchronous orthogonal frequency division multiplexing (TDS‐OFDM) enjoys the higher spectrum efficiency and faster synchronisation than the classical cyclic prefix OFDM (CP‐OFDM). However, TDS‐OFDM suffers from performance degradation especially under severely fading channels with long delays. To solve these problems, the authors propose an energy‐efficient OFDM scheme called time–frequency‐training orthogonal frequency division multiplexing (TFT‐OFDM) based on the time–frequency joint channel estimation under the framework of compressive sensing (CS). The power of the guard interval (GI) in the proposed scheme can be reduced to achieve higher energy efficiency, which is infeasible for CP‐OFDM. This method first utilises the time‐domain pseudo noise sequence to acquire partial support information of the channel, and then some frequency‐domain pilots are used for the exact channel estimation. Simulation results show that TFT‐OFDM with CS can achieve much higher energy efficiency than the classical CP‐OFDM, and outperforms the conventional OFDM schemes in both static and mobile environments. Moreover, for the channel with long delay spread, the TFT‐OFDM scheme with CS can demonstrate robustness and much better performance than the conventional OFDM schemes. In this way, the TFT‐OFDM scheme can use the same GI length for larger broadcasting coverage and hence further achieve higher energy efficiency.
Wenbo Ding 0001, Fang Yang 0001, Jian Song 0004, Zhisheng Niu
IET Commun.1
2013 Spectrum notch techniques for TDS-OFDM system
abstract
In this paper, the spectrum notch technique which combines the time-domain windowing and training sequence (TS) redesigning for the time domain synchronous orthogonal frequency division multiplexing (TDS-OFDM) system is proposed. An improved frame structure with overlapping between adjacent frames for the dual pseudo noise padded (DPNP) TDS-OFDM system is put forward to make the spectral efficiency undiminished. The simulation results show that this method can deepen the notches in the power spectrum density (PSD) compared to the conventional TDS-OFDM system by about 7 to 8 dB without any spectrum efficiency cost, which could be in accordance with the -30dB spectrum notch requirements in the spectral mask defined in HomePlugAV 1.1.
Wenbo Ding 0001, Fang Yang 0001, Jian Song 0004
IWCMC1
2012 Signaling-embedded preamble design for OFDM system with transmit diversity
abstract
In this paper, a signaling-embedded preamble design is proposed for the orthogonal frequency division multiplexing (OFDM) system with transmit diversity. Utilizing the transmit diversity, two identical training sequences (TSs) are embedded in two OFDM symbols separately in discrete Fourier transform (DFT) domain and transmitted by two different antennas simultaneously. The relative distance between the two TSs in DFT domain could be varied in order to indicate several bits of signaling information for the receiver to acquire the transmission parameters signaling (TPS) quickly. Furthermore, the preamble can also be used for frame synchronization, carrier frequency offset (CFO) estimation, and coarse channel estimation. The computational simulations were carried out under both additive white Gaussian noise (AWGN) and multipath channels, and the results demonstrated the design works accurately on the preamble detection, CFO estimation and signaling demodulation under both static and dynamic channels.
Wenbo Ding 0001, Fang Yang 0001, Jian Song 0004, Lifeng He
IWCMC1
2011 Measurement and prediction of DTMB reception quality in single frequency networks
abstract
This paper presents a method to measure and predict the reception quality of the digital terrestrial/television multimedia broadcasting (DTMB) system in a single frequency network (SFN) environment. Measurement of the signal-to-noise ratio (SNR) threshold at several market commercial receivers was carried out in the laboratory. After analyzing the channel capacity loss produced by the multiple transmitter reception, a prediction model of SNR threshold is presented. Simulation results are also performed to verify the accuracy of the prediction method. This model and method could be adopted for SFN planning to ensure a good service quality in the SFN environment.
Keqian Yan, Wenbo Ding 0001, Yanbin Yin, Fang Yang 0001, Changyong Pan
IWCMC2