Shaowu Yang

dblp:70/7915 · DBLP profile ↗
← Back
54ranked-venue papers
2as first author
45since 2021 · last 2026
0000-0002-8398-5612ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 1 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 10 since 2021Systems, architecture and hardware · 11 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MMG-VL: A Vision-Language Driven Approach for Multi-Person Motion Generation
abstract
Generating realistic and coordinated 3D human motion for multiple individuals within complex environments remains a significant challenge. Existing text-to-motion methods are often ``blind'' to the physical scene, leading to implausible motions, while scene-conditioned (HSI) approaches demand cumbersome full 3D data and largely neglect multi-person dynamics. To address these limitations, we introduce the VL2Motion paradigm and its embodiment, MMG-VL, a hierarchical framework that generates coordinated multi-person motions from the most accessible inputs: a single 2D image and natural language. MMG-VL first employs a Scene-Aware Intent Planner (SAIP) to interpret the visual context and decompose the user's command into a set of spatially-grounded, multi-person action blueprints. Subsequently, a Coordinated Motion Synthesizer (CMS) translates these blueprints into high-fidelity 3D motion sequences. The synergy between these stages is driven by two novel loss functions: a Spatial-Semantic Grounding Loss to ensure the planner's output is grounded in visual reality, and a Coordinated Environmental Realism Loss that enforces physical constraints and coherent group dynamics during synthesis. To facilitate this research, we introduce HumanVL, the first large-scale dataset featuring multi-person activities in multi-room scenes, providing aligned images, text, blueprints, 3D motions, and scene geometry. Extensive experiments demonstrate that MMG-VL significantly outperforms existing methods in generating spatially coherent, physically realistic, and coordinated multi-person motions, paving the way for more scalable and intuitive creation of dynamic virtual worlds.
Songyuan Yang, Wanrong Huang, Yinuo Liu, Kedi Zhang, Xihuai He, Shaowu Yang, Huibin Tan
AAAI6
2026 Autonomous Dynamic Target Search Using Decentralized Multi-robot Systems in Unknown Environments
Shaowu Yang, Kang Wen, Yinkun Wang, Kuang Hu
ICIC (15)1
2026 Trustworthy conflict-aware multi-view learning via bi-level evidence exploration
Guangyan Ji, Dian-xi Shi, Shaowu Yang, Zhiruo Zhang, Zichen Yao
Inf. Sci.3
2026 LCA-Med: A lightweight cross-modal adaptive feature processing module for detecting imbalanced medical image distribution
Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Hudan Pan, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001
Neural Networks4
2026 D3HRL: A distributed hierarchical reinforcement learning approach based on causal discovery and spurious correlation detection
Chenran Zhao, Dian-xi Shi, Mengzhu Wang, Jianqiang Xia, Huanhuan Yang, Songchang Jin, Shaowu Yang, Chunping Qiu
Neural Networks7
2026 TFGI: a continual learning framework for remaining useful life prediction based on temporal fusion model
Jianchao Tang, Shaowu Yang
J. Supercomput.3
2026 AnyUser: Translating Sketched User Intent Into Domestic Robots
abstract
We introduce AnyUser, a unified robotic instruction system for intuitive domestic task instruction via free-form sketches on camera images, optionally with language. AnyUser interprets multimodal inputs (sketch, vision, language) as spatial-semantic primitives to generate executable robot actions requiring no prior maps or models. Novel components include multimodal fusion for understanding and a hierarchical policy for robust action generation. Efficacy is shown via extensive evaluations: (1) Quantitative benchmarks on the large-scale dataset showing high accuracy in interpreting diverse sketch-based commands across various simulated domestic scenes. (2) Real-world validation on two distinct robotic platforms, a statically mounted 7-DoF assistive arm (KUKA LBR iiwa) and a dual-arm mobile manipulator (Realman RMC-AIDAL), performing representative tasks like targeted wiping and area cleaning, confirming the system's ability to ground instructions and execute them reliably in physical environments. (3) A comprehensive user study involving diverse demographics (elderly, simulated non-verbal, low technical literacy) demonstrating significant improvements in usability and task specification efficiency, achieving high task completion rates (85.7%-96.4%) and user satisfaction. AnyUser bridges the gap between advanced robotic capabilities and the need for accessible non-expert interaction, laying the foundation for practical assistive robots adaptable to real-world human environments.
Songyuan Yang, Huibin Tan, Kailun Yang 0001, Wenjing Yang 0002, Shaowu Yang
IEEE Trans. Robotics5
2025 Acting Beyond Learning: Imagination-Assisted Decision-Making in the Visual-based Multi-Agent Cooperative Scenarios
abstract
Learning optimal policies in multi-agent cooperative settings with visual observations is significant and challenging. Agents must first perform state representation learning for their image observations and then learn policies in the abstracted state space. Aiming at this problem, we propose a novel model-based MARL method named Contrastive Latent World for Policy Optimization (CLWPO). In CLWPO, we first design a state representation model to facilitate learning in the latent state space. With the support of this model, we construct the latent world and introduce a contrastive variational bound (CVB) to optimize it. Subsequently, we develop a heuristic policy optimization (HPO) scheme, incorporating model-free learning with model-based planning to obtain robust policies that predict future behaviors. In particular, in the planning, we maintain a queue of teammate models and calculate an adaptive rollout length for each agent to support their self-imagination and reduce the model-based return discrepancy. Finally, we conducted extensive experiments in the PettingZoo benchmark, and results show that CLWPO significantly enhances learning efficiency and improves agent performance compared to state-of-the-art MARL methods.
Huanhuan Yang, Dian-xi Shi, Songchang Jin, Guojun Xie, Chunping Qiu, Shaowu Yang
AAAI7
2025 HyperSDT: HyperNetwork Slide Decision Tree for Interpretable Tabular Learning
abstract
Recently, substantial progress has been achieved in leveraging deep learning models for tabular data learning. However, despite significant advancements, the predominant focus of these endeavors has been on augmenting the performance of contemporary deep learning models. Consequently, the interpretability of such models is frequently overlooked or rendered secondary, thereby posing a challenge in comprehending their underlying decision-making processes. In this work, we propose a novel HyperNetwork Slide Decision Tree (HyperSDT) approach to achieve interpretable deep learning for tabular data while maintaining a comparable accuracy to state-of-the-art methods. HyperSDT provides a comprehensive interpretable framework with interpretability by using Silde Decision Tree and Decision Transformer together. Our experimental results demonstrate that our framework is competitive with prior baselines under various tabular learning benchmarks while providing better interpretability. The code can be achieved via https://github.com/hunan-create/HyperSDT.
Xueqiong Li, Zhenhua Liang, Shaowu Yang, Ji Wang 0001
ICASSP5
2025 Multi-layer Network Disintegration via Deep Reinforcement Learning
abstract
Multi-layer networks (MLN) effectively model interactions across layers, and the network disintegration (ND) problem yields significant importance in the analysis of MLN. Unfortunately, previous advances in ND for single-layer networks exhibits inefficiency and lack of scalability when extended to MLN, as MLN involve complex inter-layer dependencies and interactions that are absent in single-layer networks. To bridge this gap, we propose a pioneer framework named Multi-layer Network Disintegration via Deep Reinforcement Learning (MNDRL), by re-formulating the disintegration process into a preprocessing-encoding-decoding pipeline. To be specific, MNDRL consists of two Graph Neural Network (GNN) models to effectively capture intra-layer and inter-layer representations separately. Empowered with DRL, our MNDRL achieves near-optimal results in the ND task with the NP-hard complexity by autonomously learning efficient disintegration strategies. Extensive experiments demonstrate that our MNDRL outperforms the results of baseline disintegration methods by over 30% on real-world datasets and by over 10% on synthetic datasets with varying node counts.
Zhenhua Liang, Xueqiong Li, Shaowu Yang, Hengzhu Liu
ICASSP5
2025 UniIVFT: Towards a Unified Framework for Infrared-Visible Fusion and Translation
abstract
Infrared-visible image fusion (IVF) and infrared-to-visible image translation (I2V) are two closely related tasks in multimodal image processing, both aimed at combining or transforming infrared and visible modalities to enhance image information content. Existing methods typically focus on either fusion or translation, often requiring redundant construction of similar components for each task, which limits the effective utilization of cross-modal interactions and feature encoding capabilities. Furthermore, these approaches are often hindered by their reliance on complex feature extract models, limiting their overall effectiveness and adaptability. In this paper, we introduce the Unified Multimodal Infrared-Visible Image Fusion and Translation (UniIVFT) framework, which integrates both fusion and translation tasks within a single architecture. We employ a vision transformer (ViT) encoder-decoder structure augmented with task-specific tokens and introduce a contrastive loss to effectively align infrared and visible image features before multimodal encoding. This alignment enhances the encoder’s ability to capture cross-modal interactions. In UniIVFT, both IVF and I2V tasks share a unified encoder architecture and use task-specific tokens to control model outputs, reducing redundant model construction and training. Extensive experiments demonstrate that UniIVFT achieves performance on par with that of SOTAs across multiple tasks while maintaining a lightweight architecture with fewer model parameters.
Xueqiong Li, Shaowu Yang, Huibin Tan, Yuhua Tang
ICASSP3
2025 Adaptive Distribution-Aware Modeling for Transformer Tracking
abstract
Adapting to changes in data distribution is a major challenge in visual object tracking. In Transformer-based tracking, Layer Normalization (LN) is often applied uniformly to both template and search features, limiting feature diversity. Additionally, models tend to converge to trivial solutions, and tracking samples are sensitive to distribution shifts, affecting robustness. To address these issues, we propose the Adaptive Distribution-Aware Transformer Tracker (ADAT), incorporating three key components: the Target-Aware Module (TAM), the Region-Aware Module (RAM), and the Self-Feedback-Aware Module (SFAM). TAM normalizes template and search features separately, preserving flexibility and enhancing target learning. RAM refines target perception by distinguishing between near and far target regions. SFAM filters out noisy samples and fine-tunes normalization parameters through self-feedback. While TAM and RAM regulate feature-level distribution, SFAM adjusts at the sample level. Extensive experiments show that ADAT outperforms existing methods, achieving superior performance on challenging benchmarks.
Mingyu Cao, Huibin Tan, Xueqiong Li, Wanrong Huang, Kedi Zhang, Yuhua Tang, Shaowu Yang
ICME7
2025 HBTP: Heuristic Behavior Tree Planning with Large Language Model Reasoning
abstract
Behavior Trees (BTs) are increasingly becoming a popular control structure in robotics due to their modularity, reactivity, and robustness. In terms of BT generation methods, BT planning shows promise for generating reliable BTs. However, the scalability of BT planning is often constrained by prolonged planning times in complex scenarios, largely due to a lack of domain knowledge. In contrast, pre-trained Large Language Models (LLMs) have demonstrated task reasoning capabilities across various domains, though the correctness and safety of their planning remain uncertain. This paper proposes integrating BT planning with LLM reasoning, introducing Heuristic Behavior Tree Planning (HBTP)-a reliable and efficient framework for BT generation. The key idea in HBTP is to leverage LLMs for task-specific reasoning to generate a heuristic path, which BT planning can then follow to expand efficiently. We first introduce the heuristic BT expansion process, along with two heuristic variants designed for optimal planning and satisficing planning, respectively. Then, we propose methods to address the inaccuracies of LLM reasoning, including action space pruning and reflective feedback, to further enhance both reasoning accuracy and planning efficiency. Experiments demonstrate the theoretical bounds of HBTP, and results from four datasets confirm its practical effectiveness in everyday service robot applications.
Yishuai Cai, Xinglin Chen, Yunxin Mao, Minglong Li, Shaowu Yang, Wenjing Yang 0002, Ji Wang 0001
ICRA5
2025 Spatio-Temporal Bidirectional Fusion RGB-T Tracking
abstract
The goal of RGB-T tracking is to fully utilize the information of the RGB and TIR modality sequences to enhance the robustness of object tracking. Existing RGB-T tracking methods usually simply fuse the features of the RGB and TIR search regions, resulting in insufficient information interaction and introducing unnecessary background noise. In addition, the utilization of temporal context information has not been fully explored. To address the above limitations, we propose an RGB-T tracking framework STTrack. We propose a Spatio-Temporal Bidirectional Adapter (STA), which is integrated into the ViT backbone. The STA conducts bidirectional information interaction between spatial and temporal context, so as to explore and fuse the synergistic and complementary information between modalities and the temporal context information during the tracking process more effectively. Simultaneously, we propose an Asynchronous Online Template Update (AOTU) strategy to update high-quality temporal context information for the tracker, to adapt to the appearance changes of the target over time. In addition, we introduce a Modal State Guided Fusion (MSGF) module to adaptively filter out unnecessary modality or background noise. Extensive experiments conducted on five popular RGB-T benchmark datasets demonstrate that STTrack outperforms existing state-of-the-art methods.
Dian-xi Shi, Jianqiang Xia, Jing Xie 0021, Shaowu Yang
IJCNN5
2025 A visual state space Model-Based Cross-Domain adaptive detection method for imbalanced medical image distribution
Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Hudan Pan, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001
Appl. Intell.4
2025 Dragon Boat Optimization: A Meta-Heuristic for Intelligent Systems
abstract
ABSTRACT Dragon boat racing, a popular aquatic folklore team sport, is traditionally held during the Dragon Boat Festival. Inspired by this event, we propose a novel human‐based meta‐heuristic algorithm called dragon boat optimization (DBO) in this paper. It models the unique behaviours of each crew member on the dragon boat during the race by introducing social psychology mechanisms (social loafing, social incentive). Throughout this process, the focus is on the interaction and collaboration among the crew members, as well as their decision‐making in various situations. During each iteration, DBO implements different state updating strategies. By accurately modelling the crew's behaviour and employing adaptive state update strategies, DBO consistently achieves high optimization performance, as validated by comprehensive testing on 29 benchmark functions and 2 structural design problems. Experimental results indicate that DBO outperforms 7 and 16 state‐of‐the‐art meta‐heuristic algorithms across these test functions and problems, respectively.
Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001
Expert Syst. J. Knowl. Eng.4
2025 GGF: Global Geometric Feature for Rotation-Invariant Point Cloud Understanding
Yunzhe Xiao, Haotian Wang 0001, Shaowu Yang
J. Comput. Sci. Technol.4
2025 Remote sensing image encryption algorithm based on DNA convolution
Jingxi Tian, Songchang Jin, Dian-xi Shi, Shaowu Yang
J. Supercomput.6
2024 Learning to Learn Better Visual Prompts
abstract
Prompt tuning provides a low-cost way of adapting vision-language models (VLMs) for various downstream vision tasks without requiring updating the huge pre-trained parameters. Dispensing with the conventional manual crafting of prompts, the recent prompt tuning method of Context Optimization (CoOp) introduces adaptable vectors as text prompts. Nevertheless, several previous works point out that the CoOp-based approaches are easy to overfit to the base classes and hard to generalize to novel classes. In this paper, we reckon that the prompt tuning works well only in the base classes because of the limited capacity of the adaptable vectors. The scale of the pre-trained model is hundreds times the scale of the adaptable vector, thus the learned vector has a very limited ability to absorb the knowledge of novel classes. To minimize this excessive overfitting of textual knowledge on the base class, we view prompt tuning as learning to learn (LoL) and learn the prompt in the way of meta-learning, the training manner of dividing the base classes into many different subclasses could fully exert the limited capacity of prompt tuning and thus transfer it power to recognize the novel classes. To be specific, we initially perform fine-tuning on the base class based on the CoOp method for pre-trained CLIP. Subsequently, predicated on the fine-tuned CLIP model, we carry out further fine-tuning in an N-way K-shot manner from the perspective of meta-learning on the base classes. We finally apply the learned textual vector and VLM for unseen classes.Extensive experiments on benchmark datasets validate the efficacy of our meta-learning-informed prompt tuning, affirming its role as a robust optimization strategy for VLMs.
Fengxiang Wang 0004, Wanrong Huang, Shaowu Yang, Long Lan
AAAI3
2024 Radar Recognition in the Wild: Enhancing Radar Emitter Recognition through Auto-Correlation Model-Agnostic Meta Learning
abstract
In Electronic Support Measure (ESM) systems, the recognition of radar emitters stands as a pivotal yet intricate task. The complex electromagnetic environments, however, often hinders the collection of clean radar signal data, and results in data with different noise levels. Consequently, formulating a robust recognition model with limited data becomes a big challenge, further compounded by the demand for generalizability across scenarios with different noise levels. While Model-Agnostic Meta Learning (MAML) has proven its effectiveness in solving few-shot learning problems in computer vision, its application in radar signal processing has remained unexplored deeply. This paper pioneers the incorporation of MAML and autocorrelation into radar signal processing. To fit MAML to radar signals, we introduce a novel loss function, termed AC-Loss, designed to facilitate learning effective signal representation by retaining the periodicity of the radar pulses which is the key feature for recognizing different Pulse Repetition Intervals (PRIs). This proposed Autocorrelation Model-Agnostic Meta Learning (AC-MAML) enhances its recognition capabilities while using only a sparse number of signal samples in both source and target domains. Empirical results show the superiority of AC-MAML, achieving an impressive average recognition accuracy of 90.4% across seven diverse target domain scenarios.
Yixian Luo, Shaowu Yang, Huibin Tan, Ruochun Jin, Hengzhu Liu, Xueqiong Li
ICASSP2
2024 CRNet: Cross-Reconstruction Network for Inconsistent Point Cloud Registration
abstract
Deep learning methods have made significant advancements in point cloud registration, achieving excellent performance on consistent point clouds. However, these methods face challenges when dealing with point clouds exhibiting inconsistent spatial distributions. To address this issue, we propose the Cross-Reconstruction Network (CRNet), a novel approach designed to register two inconsistent point clouds by reconstructing them into a consistent shape. CRNet utilizes a cross-learning framework that facilitates feature interaction between input point clouds at both global and point-wise levels. This interaction network enables the bidirectional generation of corresponding points to reconstruct consistent point clouds for transformation estimation. Furthermore, the transformation parameters can be refined by a regression network to achieve more accurate registration. The experimental results valuated on benchmark datasets demonstrate that CRNet outperforms state-of-the-art methods in inconsistent scenarios.
Yunzhe Xiao, Xueqiong Li, Shaowu Yang, Wenjing Yang 0002, Yong Dou
ICME3
2024 Task Allocation in Heterogeneous Multi-Robot Systems Based on Preference-Driven Hedonic Game
abstract
Multiple preferences between robots and tasks have been largely overlooked in previous research on Multi-Robot Task Allocation (MRTA) problems. In this paper, we propose a preference-driven approach based on hedonic game to address the task allocation problem of muti-robot systems in emergency rescue scenarios. We present a distributed framework considering various preferences between robots and tasks to determine the division of coalitions in such problems and evaluate the scalability and adaptability of our algorithm through relevant experiments. Furthermore, considering the strict communication limitations in emergency rescue scenarios, we have verified that our algorithm can efficiently converge to a Nash-stable coalition partition even in conditions of insufficient communication distance.
Liwang Zhang, Minglong Li, Wenjing Yang 0002, Shaowu Yang
ICRA4
2024 Coalition Formation Game Approach for Task Allocation in Heterogeneous Multi-Robot Systems under Resource Constraints
abstract
This paper studies a case of the multi-robot task allocation (MRTA) problem, where each unmanned aerial vehicle (UAV) is endowed with multiple but limited resources. Completing each task necessitates UAVs to combine different resources through coalition formation, which will incur various costs including flight cost, execution cost, and cooperation cost. To minimize the total cost while maximizing both task completion rate and resource utilization rate, we model the MRTA problem of the UAVs as a leader-follower coalition formation game. In this game, leader UAVs coordinate follower UAVs to fulfill task resource requisites. Meanwhile, follower UAVs select suitable coalitions to join based on the altruistic preference. Theoretical analysis confirms the existence of a Nash stable partition in the coalition formation game. To achieve this stable partition, we propose a coalition formation algorithm. Simulation experiments validate that the proposed algorithm outperforms existing methods for the MRTA problem under resource constraints in terms of both task completion rate and resource utilization rate.
Liwang Zhang, Minglong Li, Wenjing Yang 0002, Shaowu Yang
IROS5
2024 DBTN: An adaptive neural network for multiple-disease detection via imbalanced medical images distribution
Xiang Li 0089, Long Lan, Chang-Yong Sun, Shaowu Yang, Shuihua Wang, Wenjing Yang 0002, Heng Liu 0001, Yudong Zhang 0001
Appl. Intell.4
2024 Black-box adversarial patch attacks using differential evolution against aerial imagery object detectors
Guijian Tang, Wen Yao 0001, Chao Li 0076, Tingsong Jiang, Shaowu Yang
Eng. Appl. Artif. Intell.5
2024 EAFP-Med: An efficient adaptive feature processing module based on prompts for medical image detection
abstract
The rapid proliferation of medical imaging technologies presents a significant challenge for cross-domain adaptive image detection, as lesion representations can vary dramatically across technologies. To address this issue, we draw inspiration from large language models to propose EAFP-Med, an efficient adaptive feature processing module based on prompts for medical image detection. EAFP-Med incorporates a prompt-driven dynamic parameter update mechanism, empowering it to extract cross-domain multi-scale lesion features from medical images of diverse modalities adaptively. This exceptional flexibility liberates it from the constraints of any particular imaging technique, fostering great adaptability. Furthermore, EAFP-Med can also serve as a feature preprocessing module connected to any model front-end to enhance the lesion features in input images. Moreover, we propose a novel adaptive disease detection model named EAFP-Med ST, which utilizes the Swin Transformer V2 – Tiny (SwinV2-T) as its backbone and connects it to EAFP-Med. We have compared our method to nine state-of-the-art methods. Experimental results show that the overall accuracy of EAFP Med ST on chest X-ray, brain magnetic resonance imaging, and skin image datasets is 98.47%, 97.60%, and 99.06%, respectively, superior to all the compared state-of-the-art methods.
Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001
Expert Syst. Appl.4
2024 Boosting Few-shot Object Detection with Discriminative Representation and Class Margin
abstract
Classifying and accurately locating a visual category with few annotated training samples in computer vision has motivated the few-shot object detection technique, which exploits transfering the source-domain detection model to the target domain. Under this paradigm, however, such transferred source-domain detection model usually encounters difficulty in the classification of the target domain because of the low data diversity of novel training samples. To combat this, we present a simple yet effective few-shot detector, Transferable RCNN. To transfer general knowledge learned from data-abundant base classes to data-scarce novel classes, we propose a weight transfer strategy to promote model transferability and an attention-based feature enhancement mechanism to learn more robust object proposal feature representations. Further, we ensure strong discrimination by optimizing the contrastive objectives of feature maps via a supervised spatial contrastive loss. Meanwhile, we introduce an angle-guided additive margin classifier to augment instance-level inter-class difference and intra-class compactness, which is beneficial for improving the discriminative power of the few-shot classification head under a few supervisions. Our proposed framework outperforms the current works in various settings of PASCAL VOC and MSCOCO datasets; this demonstrates the effectiveness and generalization ability.
Shaowu Yang, Wenjing Yang 0002, Dian-xi Shi, Xuehui Li
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Chinese Medical Named Entity Recognition Based on Pre-training Model
Shaowu Yang, Yongjun Zhang 0006, Dian-xi Shi
GPC (1)2
2023 Memory-based Exploration-value Evaluation Model for Visual Navigation
abstract
We propose a hierarchical visual navigation solution, called Memory-based Exploration-value Evaluation Model (MEEM), to improve the agent's navigation performance. MEEM employs a hierarchical policy to tackle the challenge of sparse rewards, holds an episodic memory to store the historical information of the agent, and applies an Exploration-value Evaluation Model to calculate an exploration-value for action planning at each location in the observable area. We experimentally verify MEEM by navigation performance comparison on two datasets including the grid-map dataset and the 3D scenes Gibson dataset, where our approach achieves state-of-the-art performance on both. Specifically, the overall success rate of MEEM is 95% on the grid-map dataset while the best competitor reaches 68% only. As for the Gibson dataset, the success rate of ours and the best competitor SemExp are 69.8% and 54.4%, respectively. Ablation analysis on the tile-map dataset indicates that all three components of MEEM have positive effects.
Yongquan Feng, Minglong Li, Ruochun Jin, Shaowu Yang, Wenjing Yang 0002
ICRA6
2023 Collision-free Coverage Path Planning for the Variable-speed Curvature-constrained Robot
abstract
Dubins coverage has been extensively researched to address the coverage path planning (CPP) problem of a known environment for the curvature-constrained robot. However, its fixed-speed assumption prevents the robot from accelerating to reduce the time and limits its flexibility to avoid obstacles. Therefore, this paper presents a collision-free CPP approach (CFC) for the obstacle-constrained environment, which enhances time efficiency by constructing the variable-speed Dubins paths and ensures robot safety by building a risk potential surface for representing the possibility of collision. Furthermore, CFC models the CPP problem as an asymmetric traveling salesman problem (ATSP) and utilizes a graph pruning strategy to reduce the computational cost. Comparison tests with other Dubins coverage methods demonstrate that CFC provides shorter coverage times and better runtimes than the other Dubins coverage methods while preventing collision risk between the robot and obstacles. Physical experiments in a laboratory setting demonstrate the applicability of CFC to the physical robot.
Lin Li 0075, Dian-xi Shi, Songchang Jin, Yixuan Sun, Xing Zhou 0004, Shaowu Yang, Hengzhu Liu
ICRA6
2023 Task2Morph: Differentiable Task-Inspired Framework for Contact-Aware Robot Design
abstract
Optimizing the morphologies and the controllers that adapt to various tasks is a critical issue in the field of robot design, aka. embodied intelligence. Previous works typically model it as a joint optimization problem and use search-based methods to find the optimal solution in the morphology space. However, they ignore the implicit knowledge of task-to-morphology mapping which can directly inspire robot design. For example, flipping heavier boxes tends to require more muscular robot arms. This paper proposes a novel and general differentiable task-inspired framework for contact-aware robot design called Task2Morph. We abstract task features highly related to task performance and use them to build a task-to-morphology mapping. Further, we embed the mapping into a differentiable robot design process, where the gradient information is leveraged for both the mapping learning and the whole optimization. The experiments are conducted on three scenarios, and the results validate that Task2Morph outperforms DiffHand, which lacks a task-inspired morphology module, in terms of efficiency and effectiveness.
Yishuai Cai, Shaowu Yang, Minglong Li, Xinglin Chen, Yunxin Mao, Xiaodong Yi 0002, Wenjing Yang 0002
IROS2
2023 Multi-granularity knowledge distillation and prototype consistency regularization for class-incremental learning
Dian-xi Shi, Ziteng Qiao, Zhen Wang 0052, Shaowu Yang, Chunping Qiu
Neural Networks6
2022 Deep Reinforcement Learning for Multi-UAV Exploration Under Energy Constraints
Yating Zhou, Dian-xi Shi, Huanhuan Yang, Haomeng Hu, Shaowu Yang, Yongjun Zhang 0006
CollaborateCom (2)5
2022 Placement Optimization for UAV-Enabled Wireless Networks with Multi-Hop Backhauls in Urban Environments
abstract
In surveillance or search scenarios, exploiting unmanned aerial vehicles (UAVs) as relays to provide wireless data access for task-oriented ground robots (GRs) with remote base station have emerged as a promising application. This paper considers a UAV-enabled wireless network, where communication links could be line-of-sight (LoS) and non-line-of-sight (NLoS) due to obstacles in urban environments. Existing works typically adopted the free-space path loss model or the statistical channel model, which either ignored the impact of obstacles or assumed uniformly distributed obstacles and therefore might fail in practical NLoS scenarios. In this paper, taking the information of randomly distributed obstacles in environments into consideration, we aim to optimize the placement for the UAV-enabled multi-hop network to transfer more data collected by GRs and minimize the time delay in data transmission while satisfying the required communication quality. By reconstructing this complex non-convex optimization problem into two subprob-lems and solving them alternatively, we propose the multi-hop UAVs placement (mUP) method to get the solution, which contains the air-to-ground network formation (ATG-NF) algorithm and the communication quality-aware UAV placement (CQA-UP) algorithm. Simulation results show that in four types of typical urban environments or with different numbers of UAVs, the proposed mUP method achieves substantial performance gains in terms of communication quality and task performance compared to other placement approaches based on statistical channel models. We further discuss the robustness of the mUP method towards terrain measurement error.
Sining Yang, Dian-xi Shi, Yingxuan Peng, Shaowu Yang, Bo Zhang 0007, Wenjing Yang 0002
IPSN4
2022 Efficient Scale Divide and Conquer Network for Object Detection
Ziteng Qiao, Shaowu Yang, Dian-xi Shi
PRICAI (3)5
2022 FusionSeg: Motion Segmentation by Jointly Exploiting Frames and Events
Zhe Liu 0029, Shaowu Yang, Dian-xi Shi, Yongjun Zhang 0006
PRICAI (3)4
2022 Independent Multi-agent Reinforcement Learning Using Common Knowledge
abstract
Many recent multi-agent reinforcement learning algorithms used centralized training with decentralized execution (CTDE), which results in a training process that relies on global information and suffers from the dimensional explosion. The independent learning (IL) approaches are simple in structure and can be more easily deployed to a wider range of multi-agent scenarios, but they can only solve relatively simple problems due to environment non-stationarity and partially observable. With this motivation, we let IL agents compute common knowledge information and fuse it with observation to explicitly exploit common knowledge. In addition, we chose a suitable network structure according to the characteristics of IL, using convolutional layers and GRU layers. Based on the above two improvements, we implement two IL algorithms. In our experiments, the algorithms we implemented show significant performance improvements compared to original IL algorithms and further approach CTDE while outperforming multi-agent common knowledge reinforcement learning.
Haomeng Hu, Dian-xi Shi, Huanhuan Yang, Yingxuan Peng, Yating Zhou, Shaowu Yang
SMC6
2022 SCSE-E2VID: Improved event-based video reconstruction with an event camera
abstract
The recently emerging event camera has grown into a new type of sensor in the realm of vision, with benefits such as low power consumption, high dynamic range (HDR), microsecond time resolution, and no motion blur. While event cameras offer numerous advantages over conventional cameras, they only capture changes in intensity and give up lots of environmental details. This paper proposes an end-to-end UNet network called SCSE-E2VID to synthesize gray images from asynchronous events. We design an event fusion block to feed more related events to the encoder, allowing the network to extract more valuable features. The famous attention module called Spatial and Channel ‘Squeeze & Excitation’ Block (SCSE) is utilized to remove artifacts and better extract spatiotemporal features for the decoder. Besides, we add parallel convolutions in the upsampling block and refine the output features, which supplement content in reduced channels. In order to evaluate the performance of our proposed SCSE-E2VID, we implement quantitative and qualitative comparisons based on the public IJRR and HQF datasets. The results show that our method achieves better performance in terms of perceptual similarity and structural similarity when compared with state-of-art methods and demonstrates comparable performance in terms of squared error.
Dian-xi Shi, Ruihao Li 0001, Luoxi Jing, Shaowu Yang
SMC6
2022 E-HANet: Event-based Hybrid Attention Network for Optical Flow Estimation
abstract
Optical flow estimation is an essential task in computer vision. Standard cameras are prone to blurred images or over-saturated regions under extreme conditions. The event camera is a novel vision sensor inspired by the biological retina. It has the advantages of high time resolution, low delay and high dynamic range. We propose a new event representation and a novel deep learning network E-HANet (Event-based Hybrid Attention Network) for event-based dense optical flow estimation. To take full advantage of the complementarity between positive and negative events, we introduce the stacked positive and negative event slices as input. The feature extractor based on the channel attention is able to model the features from different event slices and fuse them by weight. We then present the hybrid attention weighting module to globally aggregate motion features. Compared to the event-based state-of-the-art, our approach reduces the average end-point error by 3% on the MVSEC dataset and 20% on the DSEC-Flow dataset.
Qimin Wang, Yongjun Zhang 0006, Shaowu Yang, Zhe Liu 0029, Luoxi Jing
SMC3
2022 Self-supervised representations for multi-view reinforcement learning
abstract
Learning policies from raw, pixel images are quite important for the real-world application of deep reinforcement learning (RL). Standard model-free RL algorithms focus on single-view settings and unify the representation learning and policy learning into an end-to-end training process. However, such a learning paradigm is sample-inefficiency and sensitive to hyper-parameters when supervised merely by the reward signals. Based on this, we present Self-Supervised Representations (S2R) for multi-view reinforcement learning, a sample-efficient representation learning method for learning features from high-dimensional images. In S2R, we introduce a representation learning framework and define a novel multi-view auxiliary objective based on the multi-view image states and Conditional Entropy Bottleneck (CEB) principle. We integrate S2R with the deep RL agent to learn robust representations that preserve task-relevant information while discarding task-irrelevant information and find optimal policies that maximize the expected return. Empirically, we demonstrate the effectiveness of S2R in the visual DeepMind Control (DMControl) suite and show its better performance on the default DMControl tasks and their variants by replacing the tasks’ default background with a random image or natural video.
Huanhuan Yang, Dian-xi Shi, Guojun Xie, Yingxuan Peng, Yantai Yang, Shaowu Yang
UAI7
2022 Multi actor hierarchical attention critic with RNN-based feature extraction
Dian-xi Shi, Chenran Zhao, Huanhuan Yang, Gongju Wang, Shaowu Yang, Yongjun Zhang 0006
Neurocomputing8
2021 ContriQ: Ally-Focused Cooperation and Enemy-Concentrated Confrontation in Multi-Agent Reinforcement Learning
abstract
Centralized training with decentralized execution (CTDE) is an important setting for cooperative multi-agent reinforcement learning (MARL) due to communication constraints during execution and scalability constraints during training, which has shown superior performance but still suffers from challenges. One branch is to understand the mutual interplay between agents. Due to the communication constraints in practice, agents cannot exchange perceptual information, and thus, many approaches use a centralized attention network with scalability constraints. Contrary to these common approaches, we propose to learn to cooperate in a decentralized way by applying attention mechanism on the local observation so that each agent could focus on allied agents with a decentralized model, and therefore promote understanding. Another branch is to model how agents cooperate and simplify the learning process. Previous approaches that focus on value decomposition have achieved innovative results but still suffer from problems. These approaches either limit the representation expressiveness of their value function classes or relax the IGM consistency to achieve scalability, which may lead to poor performance. We combine value composition with game abstraction by modeling the relationships between agents as a bi-level graph. We propose a novel value decomposition network based on it through a bi-level attention network, which indicates the contribution of allied agents attacking enemies and the priority of attacking each enemy under the situation of each time step, respectively. We show that our method substantially outperforms existing state-of-the-art methods on battle games in StarCraft Ⅱ, and attention analysis is also comprehensively discussed with sights.
Chenran Zhao, Dian-xi Shi, Huanhuan Yang, Shaowu Yang, Yongjun Zhang 0006
ACML5
2021 CIExplore: Curiosity and Influence-based Exploration in Multi-Agent Cooperative Scenarios with Sparse Rewards
abstract
Learning in a sparse-reward setting is a well-known challenge in RL (Reinforcement Learning). In the single-agent domain, this challenge can be addressed by introducing exploration bonuses driven by intrinsic motivation to encourage agents to visit unseen states. However, naively applying these methods in MARL (Multi-Agent Reinforcement Learning) cooperative settings with sparse rewards results in some inevitable problems: misunderstanding environmental knowledge and lack of collaboration among agents, etc. Based on this, in this paper, we propose the Curiosity and Influence-based Explore (CIExplore) method, which includes a new form of intrinsic reward and an internal counterfactual advantage function. Concretely, the intrinsic reward is a combination of joint curiosity reward and influence reward. The former is the variance of outputs across an ensemble of prediction models that take joint observations and actions of all agents as inputs to predict the next time's joint observations. And the latter quantifies the influence of one agent's behavior on other agents' state-value functions. Given that the joint curiosity reward is shared by all agents, we compute an internal counterfactual advantage function to address this intrinsic reward assignment problem. We demonstrate the efficacy of CIExplore in the multi-agent grid-world environments and show that it is compatible with both on-policy and off-policy MARL algorithms and be scalable to complex settings where agents' number or environment randomness increases.
Huanhuan Yang, Dian-xi Shi, Chenran Zhao, Guojun Xie, Shaowu Yang
CIKM5
2021 Attention-Aware Actor for Cooperative Multi-agent Reinforcement Learning
Chenran Zhao, Dian-xi Shi, Yaqianwen Su, Yongjun Zhang 0006, Shaowu Yang
CollaborateCom (2)6
2021 Identity-Based Data Augmentation via Progressive Sampling for One-Shot Person Re-identification
Runxuan Si, Shaowu Yang, Haoang Chi, Yuhua Tang
ICONIP (4)2
2019 Integrating Decision Sharing with Prediction in Decentralized Planning for Multi-Agent Coordination under Uncertainty
abstract
The performance of decentralized multi-agent systems tends to benefit from information sharing and its effective utilization. However, too much or unnecessary sharing may hinder the performance due to the delay, instability and additional overhead of communications. Aiming to a satisfiable coordination performance, one would prefer the cost of communications as less as possible. In this paper, we propose an approach for improving the sharing utilization by integrating information sharing with prediction in decentralized planning. We present a novel planning algorithm by combining decision sharing and prediction based on decentralized Monte Carlo Tree Search called Dec-MCTS-SP. Each agent grows a search tree guided by the rewards calculated by the joint actions, which can not only be sampled from the shared probability distributions over action sequences, but also be predicted by a sufficiently-accurate and computationally-cheap heuristics-based method. Besides, several policies including sparse and discounted UCT and DIY-bonus are leveraged for performance improvement. We have implemented Dec-MCTS-SP in the case study on multi-agent information gathering under threat and uncertainty, which is formulated as Decentralized Partially Observable Markov Decision Process (Dec-POMDP). The factored belief vectors are integrated into Dec-MCTS-SP to handle the uncertainty. Comparing with the random, auction-based algorithm and Dec-MCTS, the evaluation shows that Dec-MCTS-SP can reduce communication cost significantly while still achieving a surprisingly higher coordination performance.
Minglong Li, Wenjing Yang 0002, Zhongxuan Cai, Shaowu Yang, Ji Wang 0001
IJCAI4
2019 Scale-aware camera localization in 3D LiDAR maps with a monocular visual odometry
abstract
Abstract Localization information is essential for mobile robot systems in navigation tasks. Many visual‐based approaches focus on localizing a robot within prior maps acquired with cameras. It is critical where the Global Positioning System signal is unreliable. In contrast to conventional methods that localize a camera in an image‐based map, we propose a novel approach that localizes a monocular camera within a given three‐dimensional (3D) light detection and ranging (LiDAR) map. We employ visual odometry to reconstruct a semidense set of 3D points from the monocular camera images. These points are continuously matched against the 3D prior LiDAR map by a modified feature‐based point cloud registration method to track a full six‐degree‐of‐freedom camera pose. Since the monocular camera suffers from the scale‐drift problem due to the lack of depth information, the proposed method solves it by adopting updatable scale estimation. Experiments carried out on a publicly large‐scale data set demonstrate that the camera and LiDAR multimodal data matching problem is solved, and the localization accuracy of our method is comparable to state‐of‐the‐art approaches.
Manhui Sun, Shaowu Yang, Henzhu Liu
Comput. Animat. Virtual Worlds2
2018 Enhanced Visual Loop Closing for Laser-Based SLAM
abstract
Three-dimensional (3D) laser-based simultaneous localization and mapping (SLAM) can provide real-time pose information and construct accurate 3D map. However, detecting loop closures is a challenging task in the 3D laser-based SLAM because of the heavy computational overheads. In this paper, we propose a visual method to simultaneously detect and correct loop closures in the 3D laser-based SLAM based on prior work. In particular, we improve the experiments and evaluate our method by analyzing computational errors. The experimental results on the KITTI dataset prove that our method can efficiently reduce motion accumulation errors and successfully ensure the consistency performance of loop closure correction.
Zulun Zhu, Shaowu Yang, Huadong Dai
ASAP2
2018 Loop Detection and Correction of 3D Laser-Based SLAM with Visual Information
abstract
Three-dimensional (3D) laser-based simultaneous localization and mapping (SLAM) can provide real-time pose information and construct accurate 3D map. However, detecting loop closures is a challenge in the 3D laser-based SLAM for expensive computation of algorithms. In this paper, we propose a visual method to detect and correct loop closures. We introduce visual bags-of-words techniques for loop closure detection in the 3D laser-based SLAM. Time of computing similarities between points clouds can be saved. Our method maintains visual keyframes, each of which associates with its pose and segmentation of laser point clouds. Our experiments on KITTI dataset prove that our method can efficiently reduce motion accumulation errors and successfully ensure the real-time performance of loop closure correction.
Zulun Zhu, Shaowu Yang, Huadong Dai
CASA2
2018 Multi-UAV Collaborative Monocular SLAM Focusing on Data Sharing
Zhuoyue Yang, Dian-xi Shi, Yongjun Zhang 0006, Shaowu Yang, Ruoxiang Li
ICONIP (7)4
2017 CORB-SLAM: A Collaborative Visual SLAM System for Multiple Robots
Shaowu Yang, Xiaodong Yi 0002, Xuejun Yang
CollaborateCom2
2014 Visual SLAM for autonomous MAVs with dual cameras
abstract
This paper extends a monocular visual simultaneous localization and mapping (SLAM) system to utilize two cameras with non-overlap in their respective field of views (FOVs). We achieve using it to enable autonomous navigation of a micro aerial vehicle (MAV) in unknown environments. The methodology behind this system can easily be extended to multi-camera rigs, if the onboard computation capability allows this. We analyze the iterative optimizations for pose tracking and map refinement of the SLAM system in multicamera cases. This ensures the soundness and accuracy of each optimization update. Our method is more resistant to tracking failure than conventional monocular visual SLAM systems, especially when MAVs fly in complex environments. It also brings more flexibility to configurations of multiple cameras used onboard of MAVs. We demonstrate its efficiency with both autonomous flight and manual flight of a MAV. The results are evaluated by comparisons with ground truth data provided by an external tracking system.
Shaowu Yang, Sebastian A. Scherer, Andreas Zell
ICRA1
2010 Camera parameters auto-adjusting technique for robust robot vision
abstract
How to make vision system work robustly under dynamic light conditions is still a challenging research focus in computer/robot vision community. In this paper, a novel camera parameters auto-adjusting technique based on image entropy is proposed. Firstly image entropy is defined and its relationship with camera parameters is verified by experiments. Then how to optimize the camera parameters based on image entropy is proposed to make robot vision adaptive to the different light conditions. The algorithm is tested by using the omnidirectional vision in indoor RoboCup Middle Size League environment and the perspective camera in outdoor ordinary environment, and the results show that the method is effective and color constancy to some extent can be achieved.
Huimin Lu 0002, Hui Zhang 0053, Shaowu Yang, Zhiqiang Zheng 0002
ICRA3
2009 A Novel Camera Parameters Auto-adjusting Method Based on Image Entropy
Huimin Lu 0002, Hui Zhang 0053, Shaowu Yang, Zhiqiang Zheng 0002
RoboCup3