Dian-xi Shi

dblp:04/6023 · also Dianxi Shi · DBLP profile ↗
← Back
78ranked-venue papers
3as first author
54since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 2 first-author · 30 since 2021Human-computer interaction and ubiquitous computing · 18 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 14 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Computer networks · 4 · 4 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
abstract
While significant progress has been achieved in multimodal facial generation using semantic masks and textual descriptions, conventional feature fusion approaches often fail to enable effective cross-modal interactions, thereby leading to suboptimal generation outcomes. To address this challenge, we introduce MDiTFace—a customized diffusion transformer framework that employs a unified tokenization strategy to process semantic mask and text inputs, eliminating discrepancies between heterogeneous modality representations. The framework facilitates comprehensive multimodal feature interaction through stacked, newly designed multivariate transformer blocks that process all conditions synchronously. Additionally, we design a novel decoupled attention mechanism by dissociating implicit dependencies between mask tokens and temporal embeddings. This mechanism segregates internal computations into dynamic and static pathways, enabling caching and reuse of features computed in static pathways after initial calculation, thereby reducing additional computational overhead introduced by mask condition by over 94% while maintaining performance. Extensive experiments demonstrate that MDiTFace significantly outperforms other competing methods in terms of both facial fidelity and conditional consistency.
Yushe Cao, Dian-xi Shi, Xuechao Zou, Haikuo Peng, Chun Yu, Junliang Xing
AAAI2
2026 Dual-Pathway Diffusion for Hand Correction in Synthetic Portraits: Global Context Aware and Local Structure Refinement
abstract
Despite significant advancements in diffusion models for generating high-quality portrait images, the problem of malformed hands remains a critical and unresolved challenge. Existing methods primarily focus on structural restoration, often neglecting the semantic coherence between the corrected hand region and the overall image. To address this limitation, we propose a novel dual-pathway diffusion malformed hand correction method, named DD-MHC, which integrates a Global Context Aware Module (GCAM) and a Local Structure Refinement Module (LSRM) through a dual-path architecture. Under the synergy of cross-attention and spatial attention, these modules can effectively fuse global contextual features and local guiding cues, enabling precise restoration of the hand region. Comprehensive experiments demonstrate that DD-MHC significantly outperforms existing competing methods, particularly in enhancing the semantic consistency between the corrected hand region and the overall image. Additionally, to address the challenge of data scarcity in the research on correcting malformed hands within unconstrained scene portraits, we construct a brand-new General-Scene Portrait Dataset (GSPD), providing a standardized and reproducible data platform for subsequent related research.
Yushe Cao, Luoxi Jing, Yuanze Wang, Dian-xi Shi, Chun Yu, Junliang Xing
ICMR4
2026 Hierarchical area graph for object navigation with adaptive entropy-driven exploration
Jing Xie 0021, Dian-xi Shi, Junze Zhang, Yuetian Wang, Songchang Jin
Adv. Eng. Informatics2
2026 Trustworthy conflict-aware multi-view learning via bi-level evidence exploration
Guangyan Ji, Dian-xi Shi, Shaowu Yang, Zhiruo Zhang, Zichen Yao
Inf. Sci.2
2026 D3HRL: A distributed hierarchical reinforcement learning approach based on causal discovery and spurious correlation detection
Chenran Zhao, Dian-xi Shi, Mengzhu Wang, Jianqiang Xia, Huanhuan Yang, Songchang Jin, Shaowu Yang, Chunping Qiu
Neural Networks2
2025 Acting Beyond Learning: Imagination-Assisted Decision-Making in the Visual-based Multi-Agent Cooperative Scenarios
abstract
Learning optimal policies in multi-agent cooperative settings with visual observations is significant and challenging. Agents must first perform state representation learning for their image observations and then learn policies in the abstracted state space. Aiming at this problem, we propose a novel model-based MARL method named Contrastive Latent World for Policy Optimization (CLWPO). In CLWPO, we first design a state representation model to facilitate learning in the latent state space. With the support of this model, we construct the latent world and introduce a contrastive variational bound (CVB) to optimize it. Subsequently, we develop a heuristic policy optimization (HPO) scheme, incorporating model-free learning with model-based planning to obtain robust policies that predict future behaviors. In particular, in the planning, we maintain a queue of teammate models and calculate an adaptive rollout length for each agent to support their self-imagination and reduce the model-based return discrepancy. Finally, we conducted extensive experiments in the PettingZoo benchmark, and results show that CLWPO significantly enhances learning efficiency and improves agent performance compared to state-of-the-art MARL methods.
Huanhuan Yang, Dian-xi Shi, Songchang Jin, Guojun Xie, Chunping Qiu, Shaowu Yang
AAAI2
2025 Pano3R: Training Free Panoramic 3D Reconstruction
abstract
Panoramic 3D reconstruction is essential for immersive scene understanding in robotics, AR, and autonomous driving. However, most existing methods are designed for pinhole images and generalize poorly to 360° inputs due to the scarcity of panoramic training data and the high cost of retraining. We present Pano3R, the first training-free framework for panoramic 3D reconstruction that adapts existing pinhole-based models without any retraining. Pano3R consists of two stages. Specifically, the pre-processing stage applies a position-aware pairing strategy to decompose each panorama into a minimal set of perspective views. These views are selected to ensure sufficient co-visible regions while minimizing the number of projections. The test-time optimization stage incorporates a pose-prior-guided global alignment strategy to improve global consistency and mitigate accumulated errors. Our method enables accurate 360° reconstruction under both single- and multi-view input conditions. Extensive experiments demonstrate that Pano3R consistently improves reconstruction accuracy and pose estimation quality, establishing a strong and practical benchmark for training-free panoramic 3D reconstruction.
Shiming Song 0003, Yongjun Zhang 0006, Yuanze Wang, Mengzhu Wang, Yuetian Wang, Zhuojing Tian, Jinming Song, Dian-xi Shi
ECAI8
2025 Memory Prompt for Multi-Modal Visual Object Tracking
abstract
Multi-modal trackers have drawn widespread attention for robust tracking in challenging scenarios. However, existing multi-modal trackers often rely solely on spatial matching between the initial target template and the search regions, or incorporate only single-frame historical information, failing to fully exploit temporal correlations in tracking sequences. Additionally, most trackers that introduce temporal modeling require either retraining the entire network or designing specialized modules for temporal feature extraction, which incurs additional computational costs. To alleviate these limitations, inspired by human visual memory, we propose MPTrack, a novel tracker that directly reuses pre-extracted historical target features as memory prompts, establishing temporal dependencies without redundant feature extraction or specially designed temporal extraction networks. Our proposed Memory Prompt Fusion module effectively combines initial target templates with multiple historical memory cues to generate enhanced templates, enabling the perception of long-term appearance dynamics while mitigating potential interference from individual memory. Simultaneously, to avoid the computational cost of full-model training, we design a lightweight memory adapter that allows the frozen backbone network to efficiently adapt to the memory-enhanced template. Extensive experiments demonstrate that our method effectively incorporates temporal information and achieves promising results across different multi-modal tracking scenarios, including RGB+Thermal, RGB+Event, and RGB+Depth tracking tasks.
Yongjun Zhang 0006, Jianqiang Xia, Yushe Cao, Junze Zhang, Dian-xi Shi
ECAI8
2025 Multi-Agent Hierarchical Graph Attention Actor-Critic Reinforcement Learning
abstract
Multi-agent systems often face challenges such as elevated communication demands and intricate interactions. We propose an innovative hierarchical graph attention actor-critic reinforcement learning method to address the issues, which uses the hierarchical graph attention to capture the relationships of cooperation or competition among agents, and the agent enables a better understand of the dynamic environment. Specifically, we model the interaction among agents as a graph and encode the observations of the agents as a feature embedding vector with constant dimensionality to improve scalability. Through the "inter-agent" and "inter-group" attention layers, the embedding vector of each agent is updated into an information-condensed and contextualized state representation, which can adaptively extract the state-dependent relationship between agents, model the interaction at both the individual and group level, and thus learn more "advanced" strategies. Finally, we experiment on multiple multi-agent tasks to validate our proposed method’s effectiveness, stability, and scalability.
Tongyue Li, Dian-xi Shi, Songchang Jin, Zhen Wang 0052, Huanhuan Yang
ICASSP2
2025 Enhancing Visual Localization with Cross-Domain Image Generation
abstract
Visual localization aims to predict the absolute camera pose for a single query image. However, predominant methods focus on single-camera images and scenes with limited appearance variations, limiting their applicability to cross-domain scenes commonly encountered in real-world applications. Furthermore, the long-tail distribution of cross-domain datasets poses additional challenges for visual localization. In this work, we propose a novel cross-domain data generation method to enhance visual localization methods. To achieve this, we first construct a cross-domain 3DGS to accurately model photometric variations and mitigate the interference of dynamic objects in large-scale scenes. We introduce a text-guided image editing model to enhance data diversity for addressing the long-tail distribution problem and design an effective fine-tuning strategy for it. Then, we develop an anchor-based method to generate high-quality datasets for visual localization. Finally, we introduce positional attention to address data ambiguities in cross-camera images. Extensive experiments show that our method achieves state-of-the-art accuracy, outperforming existing cross-domain visual localization methods by an average of 59% across all domains. Project page: https://yzwang-sjtu.github.io/CDG-Loc.
Yuanze Wang, Yichao Yan, Shiming Song 0003, Songchang Jin, Yilan Huang, Xingdong Sheng, Dian-xi Shi
ICML7
2025 UniCT Depth: Event-Image Fusion Based Monocular Depth Estimation with Convolution-Compensated ViT Dual SA Block
abstract
Depth estimation plays a crucial role in 3D scene understanding and is extensively used in a wide range of vision tasks. Image-based methods struggle in challenging scenarios, while event cameras offer high dynamic range and temporal resolution but face difficulties with sparse data. Combining event and image data provides significant advantages, yet effective integration remains challenging. Existing CNN-based fusion methods struggle with occlusions and depth disparities due to limited receptive fields, while Transformer-based fusion methods often lack deep modality interaction. To address these issues, we propose UniCT Depth, an event-image fusion method that unifies CNNs and Transformers to model local and global features. We propose the Convolution-compensated ViT Dual SA (CcViT-DA) Block, designed for the encoder, which integrates Context Modeling Self-Attention (CMSA) to capture spatial dependencies and Modal Fusion Self-Attention (MFSA) for effective cross-modal fusion. Furthermore, we design the tailored Detail Compensation Convolution (DCC) Block to improve texture details and enhances edge representations. Extensive experiments show that UniCT Depth outperforms existing image, event, and fusion-based monocular depth estimation methods across key metrics.
Luoxi Jing, Dian-xi Shi, Zhe Liu 0029, Songchang Jin, Chunping Qiu, Ziteng Qiao, Jianqiang Xia
IJCAI2
2025 Gaussian Mixture Model for Graph Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) has been widely studied with the goal of transferring knowledge from a label-rich source domain to a related but unlabeled target domain. Most UDA techniques achieve this by reducing the feature discrepancies between the two domains to learn domain-invariant feature representations. While domain-invariant feature representations can reduce the differences between the source and target domains, excessively simplifying these differences may cause the model to overlook important domain-specific features, resulting in a decline in transfer learning effectiveness. To address this issue, this paper proposes a novel Gaussian Mixture Model for graph domain adaptation (GMM). This model effectively reduces the distributional bias between the source and target domains by modeling the distribution differences on a graph structure. GMM leverages the local structural information of the graph and the clustering capability of the Gaussian mixture model to automatically learn the latent mapping relationships between the source and target domains. To the best of our knowledge, this is the first work to introduce a Gaussian mixture model into UDA. Extensive experimental results on three standard benchmarks demonstrate that the proposed GMM algorithm outperforms state-of-the-art unsupervised domain adaptation methods in terms of performance.
Mengzhu Wang, Wenhao Ren, Yu Zhang 0268, Yanlong Fan, Dian-xi Shi, Luoxi Jing
IJCAI5
2025 Spatio-Temporal Bidirectional Fusion RGB-T Tracking
abstract
The goal of RGB-T tracking is to fully utilize the information of the RGB and TIR modality sequences to enhance the robustness of object tracking. Existing RGB-T tracking methods usually simply fuse the features of the RGB and TIR search regions, resulting in insufficient information interaction and introducing unnecessary background noise. In addition, the utilization of temporal context information has not been fully explored. To address the above limitations, we propose an RGB-T tracking framework STTrack. We propose a Spatio-Temporal Bidirectional Adapter (STA), which is integrated into the ViT backbone. The STA conducts bidirectional information interaction between spatial and temporal context, so as to explore and fuse the synergistic and complementary information between modalities and the temporal context information during the tracking process more effectively. Simultaneously, we propose an Asynchronous Online Template Update (AOTU) strategy to update high-quality temporal context information for the tracker, to adapt to the appearance changes of the target over time. In addition, we introduce a Modal State Guided Fusion (MSGF) module to adaptively filter out unnecessary modality or background noise. Extensive experiments conducted on five popular RGB-T benchmark datasets demonstrate that STTrack outperforms existing state-of-the-art methods.
Dian-xi Shi, Jianqiang Xia, Jing Xie 0021, Shaowu Yang
IJCNN2
2025 Excellent-student learning method for decentralized MARL with networked agents system
abstract
Abstract Multi-agent reinforcement learning has been widely applied in solving various sequential decision-making problems in recent years. However, a key challenge is sparse rewards, where agents receive meaningful reward signals only upon task completion. This issue leads to inefficient exploration and slow learning progress. To address this problem, we propose an excellent-student learning method supported by a decentralized distributed learning paradigm, drawing inspiration from academically diverse classrooms. Specifically, we design a novel excellent-student learning model, which suggests that agents mimic the learning behaviors of excellent students. This model requires agents to share knowledge with other agents and engage in individual exploration driven by curiosity. Next, to foster collaborative team learning behaviors, similarity measurement techniques are integrated to enhance knowledge sharing among agents. An intrinsic reward function is designed, combining individual exploration with whole-class sharing, providing additional motivation for discovering new actions and states. This reward function is seamlessly incorporated into the policy learning process. Finally, experiments conducted in various multi-agent particle environments demonstrate significant improvements in training efficiency and stability.
Dian-xi Shi, Huanhuan Yang, Tongyue Li, Zhen Wang 0052
Comput. J.2
2025 Remote sensing image encryption algorithm based on DNA convolution
Jingxi Tian, Songchang Jin, Dian-xi Shi, Shaowu Yang
J. Supercomput.5
2024 MEFusion: Unsupervised Mutual Enhancement for Multimodal Image Fusion
abstract
Image fusion aims to extract valuable information from each modality to create a fused image. Currently, state-of-the-art image fusion approaches tend to initially decompose each modality into distinct yet complementary features, and transfer beneficial information through carefully hand-crafted or learned fusion rules to the target. Nevertheless, previous approaches treat each modality in isolation before fusion, potentially under-utilising the complementary information available across modalities. In this word, we introduce a novel method called MEFusion that pioneers cross-modality mutual enhancement before feature decomposition. By harnessing the individual strengths of each modality, MEFusion elevates the overall quality and comprehensiveness of the fusion outcome. To facilitate a bidirectional enhancement for each feature across modalities, we have designed a pluggable co-attention mechanism that seamlessly integrates into a lightweight dual-path transformer. Furthermore, to enrich the details of each modality, we propose an unsupervised cross-modality mutual enhancement loss, which overcomes the limitations of requiring paired training data for enhancement tasks. Extensive experiments conducted on several benchmark datasets demonstrate the superiority of our proposed MEFusion method in terms of traditional fusion metrics and perceptual quality improvement of fused images.
Yushe Cao, Siwen Jiao, Penghao Sun, Baoyun Peng, Dian-xi Shi, Yuanchun Shi
ECAI5
2024 ICF-Loc: An Infrared-Based Coarse-to-Fine Approach for UAV Visual Geolocation under GPS-Denied Environments
abstract
Visual geolocation plays a crucial role when GPS is unavailable in the Unmanned Aerial Vehicles (UAVs). Many methods rely on visible light cameras, which may not perform well in low-light or foggy conditions. To address this issue, we propose an advanced UAV visual geolocation method called ICF-Loc, which utilizes infrared images in a coarse-to-fine approach. ICF-Loc consists of two stages: a retrieval-based coarse localization stage and a matching-based fine localization stage. The goal of the coarse localization stage is to identify the satellite image that is most similar to the UAV’s infrared image. To bridge the distribution gap between the visible and infrared domains, we propose a feature transfer module. In the fine localization stage, the UAV’s infrared image is matched with the satellite image obtained during coarse localization to estimate the UAV’s position and orientation accurately. We have designed a cross-modal image matching method based on the Fourier transform for precise estimation. The experimental results demonstrate the effectiveness of our proposed approach on both synthetic and real-world datasets.
Zhen Wang 0052, Dian-xi Shi, Chunping Qiu, Songchang Jin, Tongyue Li
ICME2
2024 Unified Single-Stage Transformer Network for Efficient RGB-T Tracking
Jianqiang Xia, Dian-xi Shi, Linna Song, Songchang Jin, Chenran Zhao, Yu Cheng 0009, Lei Jin 0003, Jianan Li 0001, Gang Wang 0031, Junliang Xing, Jian Zhao 0006
IJCAI2
2024 SVR-AVT: Scale Variation Robust Active Visual Tracking
abstract
Active Visual Tracking (AVT) is a significant research area with extensive applications in fields such as drones and autonomous driving. AVT involves controlling camera motion based on visual observations to track target object(s). In dynamic environments, especially with the presence of distractors, AVT faces the challenge of scale variation. Existing methods struggle to effectively handle these scale changes. To address this problem, this paper proposes a novel Scale Variation Robust Active Visual Tracking method (SVR-AVT). We first introduce a multi-scale multi-stage curriculum learning approach. By progressively increasing the complexity of tracking tasks, the tracker adapts to target of various scales. Secondly, we design a scale attention network, which adaptively extracts important scale features through multiple convolutional branches with different receptive fields and a scale attention mechanism. Moreover, we employ maximum position entropy learning to encourage the target to explore the environment more extensively. Experimental results in 3D environments demonstrate that SVR-AVT significantly outperforms existing methods in handling distraction and scale variation, and exhibits strong generalization capability in unseen environments.
Zhang Biao, Songchang Jin, Qianying Ouyang, Huanhuan Yang, Yuxi Zheng, Chunlian Fu, Dian-xi Shi
IJCNN7
2024 Improved Communication and Collision-Avoidance in Dynamic Multi-Agent Path Finding
abstract
Multi-Agent Path Finding (MAPF) is a classic problem with a wide range of applications. To cope with more complex situations in reality, Dynamic MAPF (DMAPF) has received much attention. The existing DMAPF definition lacks completeness or considers too simple situations. In this paper, we comprehensively model DMAPF based on realistic scenarios. Consequently, dynamic scenarios bring many problems. The dynamics of agent tasks bring the problem of more difficult coordination and cooperation of the multi-agent system, and the dynamics of obstacles bring the problem of increased collisions. To address these problems, this paper proposes a fully decentralised multi-agent reinforcement learning method CO3, which uses COmmon knowledge in selective COmmunication and proposes obstacle COllision avoidance mechanism. Firstly, common knowledge for communication improves cooperation between agents, which improves system performance and reduces collisions between agents. Secondly, the obstacle collision avoidance mechanism consists of a collision avoidance helper module and a critical region. The collision avoidance helper module improves the agents’ alertness to nearby obstacles, and the critical region gives an early warning to the agents to beware of distant obstacles. The obstacle collision avoidance mechanism can effectively reduce collisions between agents and obstacles. Finally, experiments show that CO3 can solve the DMAPF problem quite well, and the number of collisions is significantly lower than other learning-based methods in a dynamic environment.
Jing Xie 0021, Yongjun Zhang 0006, Qianying Ouyang, Huanhuan Yang, Dian-xi Shi, Songchang Jin
IJCNN6
2024 An anti-collision algorithm for robotic search-and-rescue tasks in unknown dynamic environments
abstract
This paper deals with the search-and-rescue tasks of a mobile robot with multiple interesting targets in an unknown dynamic environment. The problem is challenging because the mobile robot needs to search for multiple targets while avoiding obstacles simultaneously. To ensure that the mobile robot avoids obstacles properly, we propose a mixed-strategy Nash equilibrium based Dyna-Q (MNDQ) algorithm. First, a multi-objective layered structure is introduced to simplify the representation of multiple objectives and reduce computational complexity. This structure divides the overall task into subtasks, including searching for targets and avoiding obstacles. Second, a risk-monitoring mechanism is proposed based on the relative positions of dynamic risks. This mechanism helps the robot avoid potential collisions and unnecessary detours. Then, to improve sampling efficiency, MNDQ is presented, which combines Dyna-Q and mixed-strategy Nash equilibrium. By using mixed-strategy Nash equilibrium, the agent makes decisions in the form of probabilities, maximizing the expected rewards and improving the overall performance of the Dyna-Q algorithm. Furthermore, a series of simulations are conducted to verify the effectiveness of the proposed method. The results show that MNDQ performs well and exhibits robustness, providing a competitive solution for future autonomous robot navigation tasks.
Dian-xi Shi, Huanhuan Yang, Tongyue Li, Zhen Wang 0052
Frontiers Inf. Technol. Electron. Eng.2
2024 Learning cooperative strategies in multi-agent encirclement games with faster prey using prior knowledge
Tongyue Li, Dian-xi Shi, Zhen Wang 0052, Huanhuan Yang
Neural Comput. Appl.2
2024 JFDI: Joint Feature Differentiation and Interaction for domain adaptive object detection
Ziteng Qiao, Dian-xi Shi, Songchang Jin, Zhen Wang 0052, Chunping Qiu
Neural Networks2
2024 Sequence Matching for Image-Based UAV-to-Satellite Geolocalization
abstract
UAV-to-satellite geolocalization offers accurate drift-free navigation in the absence of external positioning signals. Increased deep-learning-based approaches have demonstrated their potential for high accuracy by framing the problem as a one-to-all retrieval task. However, in real-world scenario, the problem is not just a one-to-all retrieval task, which leads to a gap between research and applications. Based on this observation, We attempt to look closer to the problem instead of designing sophisticated network architectures or objective functions. In this study, we proposed a flexible and simple coarse-to-fine sequence-matching solution with targeted joint use of deep learning and classical machine learning approaches. Our goal is to improve geolocalization accuracy by matching UAV images with a few relevant reference image patches instead of all images. To this end, we first coarsely constructed a sequence of reference satellite image patches corresponding to the UAV trajectory, for which we proposed a deep feature- and manifold learning-based image-sorting method. Once the reference satellite patches are sorted and aligned with the UAV trajectory, the reference sequence is determined. Given a query UAV frame, the search area can be decreased from two to one dimensions. In particular, both deep-learning-based and classical image-matching algorithms can provide competitive accuracy when integrating sequence constraints. We demonstrate that classical manifold learning-based and image matching methods perform exceptionally well for UAV-to-satellite geolocalization when utilized jointly with suitable deep learning techniques. We validated the approach’s unique outperformance on two challenging and realistic UAV-to-satellite geolocalization datasets. Dataset, code and models are available for research purposes at https://seqmatch.geovisuallocalization.com/.
Zhen Wang 0052, Dian-xi Shi, Chunping Qiu, Songchang Jin, Tongyue Li, Zhe Liu 0029, Ziteng Qiao
IEEE Trans. Geosci. Remote. Sens.2
2024 Boosting Few-shot Object Detection with Discriminative Representation and Class Margin
abstract
Classifying and accurately locating a visual category with few annotated training samples in computer vision has motivated the few-shot object detection technique, which exploits transfering the source-domain detection model to the target domain. Under this paradigm, however, such transferred source-domain detection model usually encounters difficulty in the classification of the target domain because of the low data diversity of novel training samples. To combat this, we present a simple yet effective few-shot detector, Transferable RCNN. To transfer general knowledge learned from data-abundant base classes to data-scarce novel classes, we propose a weight transfer strategy to promote model transferability and an attention-based feature enhancement mechanism to learn more robust object proposal feature representations. Further, we ensure strong discrimination by optimizing the contrastive objectives of feature maps via a supervised spatial contrastive loss. Meanwhile, we introduce an angle-guided additive margin classifier to augment instance-level inter-class difference and intra-class compactness, which is beneficial for improving the discriminative power of the few-shot classification head under a few supervisions. Our proposed framework outperforms the current works in various settings of PASCAL VOC and MSCOCO datasets; this demonstrates the effectiveness and generalization ability.
Shaowu Yang, Wenjing Yang 0002, Dian-xi Shi, Xuehui Li
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Chinese Medical Named Entity Recognition Based on Pre-training Model
Shaowu Yang, Yongjun Zhang 0006, Dian-xi Shi
GPC (1)5
2023 Collision-free Coverage Path Planning for the Variable-speed Curvature-constrained Robot
abstract
Dubins coverage has been extensively researched to address the coverage path planning (CPP) problem of a known environment for the curvature-constrained robot. However, its fixed-speed assumption prevents the robot from accelerating to reduce the time and limits its flexibility to avoid obstacles. Therefore, this paper presents a collision-free CPP approach (CFC) for the obstacle-constrained environment, which enhances time efficiency by constructing the variable-speed Dubins paths and ensures robot safety by building a risk potential surface for representing the possibility of collision. Furthermore, CFC models the CPP problem as an asymmetric traveling salesman problem (ATSP) and utilizes a graph pruning strategy to reduce the computational cost. Comparison tests with other Dubins coverage methods demonstrate that CFC provides shorter coverage times and better runtimes than the other Dubins coverage methods while preventing collision risk between the robot and obstacles. Physical experiments in a laboratory setting demonstrate the applicability of CFC to the physical robot.
Lin Li 0075, Dian-xi Shi, Songchang Jin, Yixuan Sun, Xing Zhou 0004, Shaowu Yang, Hengzhu Liu
ICRA2
2023 Improved Event-Based Dense Depth Estimation via Optical Flow Compensation
abstract
Event cameras have the potential to overcome the limitations of classical computer vision in real-world applications. Depth estimation is a crucial step for high-level robotics tasks and has attracted much attention from the community. In this paper, we propose an event-based dense depth estimation architecture, Mixed-EF2DNet, which firstly predicts inter-grid optical flow to compensate for lost temporal information, and then estimates multiple contextual depth maps that are fused to generate a robust depth estimation map. To supervise the network training, we further design a smoothing loss function used to smooth local depth estimates and facilitate estimating reasonable depth for pixels without events. In addition, we introduce SE-resblocks in the depth network to enhance the network representation by selecting feature channels. Experimental evaluations on both real-world and synthetic datasets show that our method performs better in terms of accuracy when compared to state-of-the-art algorithms, especially in scene detail estimation. Besides, our method demonstrates excellent generalization in cross-dataset tasks.
Dian-xi Shi, Luoxi Jing, Ruihao Li 0001, Zhe Liu 0029, Huachi Xu
ICRA1
2023 NeRF-IBVS: Visual Servo Based on NeRF for Visual Localization and Navigation
abstract
Visual localization is a fundamental task in computer vision and robotics. Training existing visual localization methods requires a large number of posed images to generalize to novel views, while state-of-the-art methods generally require dense ground truth 3D labels for supervision. However, acquiring a large number of posed images and dense 3D labels in the real world is challenging and costly. In this paper, we present a novel visual localization method that achieves accurate localization while using only a few posed images compared to other localization methods. To achieve this, we first use a few posed images with coarse pseudo-3D labels provided by NeRF to train a coordinate regression network. Then a coarse pose is estimated from the regression network with PNP. Finally, we use the image-based visual servo (IBVS) with the scene prior provided by NeRF for pose optimization. Furthermore, our method can provide effective navigation prior, which enable navigation based on IBVS without using custom markers and depth sensor. Extensive experiments on 7-Scenes and 12-Scenes datasets demonstrate that our method outperforms state-of-the-art methods under the same setting, with only 5\% to 25\% training data. Furthermore, our framework can be naturally extended to the visual navigation task based on IBVS, and its effectiveness is verified in simulation experiments.
Yuanze Wang, Yichao Yan, Dian-xi Shi, Wenhan Zhu, Jianqiang Xia, Jeff Tan, Songchang Jin, Ke Gao 0012, Xiaokang Yang 0001
NeurIPS3
2023 Enhancing Active Visual Tracking Under Distractor Environments
Qianying Ouyang, Chenran Zhao, Jing Xie 0021, Zhang Biao, Tongyue Li, Yuxi Zheng, Dian-xi Shi
PRCV (3)7
2023 Open Self-Supervised Features for Remote-Sensing Image Scene Classification Using Very Few Samples
abstract
Big models, large datasets, and self-supervised learning (SSL) have recently gained substantial research interest due to their potential to alleviate our reliance on annotations. Considering the current high generalization ability of self-supervised models in literature, we explore in the letter how helpful SSL can be for a crucial task in remote sensing (RS), image scene classification, when forced to rely on only a few labeled samples. We proposed a simple prototype-based classification procedure without training and fine-tuning, which uses open self-supervised features from the contrastive language-image pre-training (CLIP). We test our method by exploiting ready-to-use open features on four diversified benchmark datasets, including red-green-blue (RGB) and multispectral (MS) images. Highly competitive accuracy has been obtained compared to work with similar settings, i.e., based on an exceedingly small number of labels. To the best of our knowledge, our model is the first to achieve such high accuracy in austere label conditions. We further analyze our approach from different perspectives, including its advantages and limitations, reasons for its astonishing performance, potential applications, and future improvements.
Chunping Qiu, Anzhu Yu, Xiaodong Yi 0002, Naiyang Guan, Dian-xi Shi, Xiaochong Tong
IEEE Geosci. Remote. Sens. Lett.5
2023 Learning Visual Representation Clusters for Cross-View Geo-Location
abstract
Cross-view geo-location is a crucial research field that determines the geographic location from images taken from different viewpoints. It is often studied as a retrieval task, where the query images are with unknown locations, and the database includes images with geo-tags from a different platform. Learning image representations by neural networks is an important step, and one typical training method is using a classification loss, where cross-view images of the same locations are considered the same category. However, existing methods only focus on pushing the representation distances of different categories while ignoring the intra-category representation distances of samples from different platforms. Considering that controlling the intra-category distance can help to guide the model to extract compact category-sharing representations from cross-view images, we propose a categorized cluster loss to learn separate and compact representation clusters. Categorized cluster loss can supervise the network to learn invariant information from samples of different platforms by constraining both the inter-category and intra-category feature distances. Meanwhile, we design a category-view-stratified sampling strategy, which samples balanced inputs in terms of both category and view in each batch during the learning process. We implemented our approach with a lightweight OSNet-based network and achieved higher accuracy with fewer parameters on a typical and challenging cross-view geo-location dataset than most state-of-the-art (SOTA) methods.
Haoshuai Song, Zhen Wang 0052, Dian-xi Shi, Xiaochong Tong, Yaxian Lei, Chunping Qiu
IEEE Geosci. Remote. Sens. Lett.4
2023 Multi-granularity knowledge distillation and prototype consistency regularization for class-incremental learning
Dian-xi Shi, Ziteng Qiao, Zhen Wang 0052, Shaowu Yang, Chunping Qiu
Neural Networks2
2023 UEFPN: Unified and Enhanced Feature Pyramid Networks for Small Object Detection
abstract
Object detection models based on feature pyramid networks have made significant progress in general object detection. However, small object detection is still a challenge for the existing models. In this paper, we think that two factors in the existing feature pyramid networks inhibit the performance of small object detection. The first one is that the different feature domains of shallow and deep layer features inhibit the model performance. The second one is that the accumulation of upper layer features leads to feature aliasing effect on the lower layer features, which interferes with the representations of small object features. Therefore, we propose Unified and Enhanced Feature Pyramid Networks (UEFPN) to improve the APs and ARs of small object detection. It has the following three characteristics: (1) Using the deep features of high-resolution image and original image to form the multi-scale features of unified domain. (2) In multi-scale features fusion, we learn the importance of upper layer features with the Channel Attention Fusion module (CAF) , to optimize feature aliasing effect and enhance the context information of shallow layer features. (3) UEFPN can be quickly applied to different models. The results of many experiments show that the models with UEFPN achieve significant performance improvement in small object detection compared with the baseline models.
Ziteng Qiao, Dian-xi Shi, Xiaodong Yi 0002, Yuhui Zhang 0002
ACM Trans. Multim. Comput. Commun. Appl.2
2022 Deep Reinforcement Learning for Multi-UAV Exploration Under Energy Constraints
Yating Zhou, Dian-xi Shi, Huanhuan Yang, Haomeng Hu, Shaowu Yang, Yongjun Zhang 0006
CollaborateCom (2)2
2022 A Cross-Modal Object-Aware Transformer for Vision-and-Language Navigation
abstract
Vision-and-language navigation (VLN) combines cross-modal object references and scene descriptions to provide a breadcrumb trail to a goal location. Whereas existing VLN approaches often do not take full advantage of cross-modal object information, this work proposes a transformer network with perceptual cross-modal object data that fuses and aligns the two cue features of object reference to help agents capture object features. In our method, linguistic object processing provides semantic-level contextual information for visual object features. With this design, our model is able to leverage object features to assist the agent in substantially improving performance on the R2R and R4R benchmarks. Through extensive experiments on R2R and R4R, we demonstrate the effectiveness of the proposed model, and our method improves the absolute 1.6% in SPL on R2R and 2.1% in CLS on R4R. Our analysis shows that the network performs better when focusing on longer heavily object-referenced navigation instructions, which also indicates that our approach is better able to use object features and align them to references in the instructions.
Han Ni, Dayong Zhu, Dian-xi Shi
ICTAI4
2022 Placement Optimization for UAV-Enabled Wireless Networks with Multi-Hop Backhauls in Urban Environments
abstract
In surveillance or search scenarios, exploiting unmanned aerial vehicles (UAVs) as relays to provide wireless data access for task-oriented ground robots (GRs) with remote base station have emerged as a promising application. This paper considers a UAV-enabled wireless network, where communication links could be line-of-sight (LoS) and non-line-of-sight (NLoS) due to obstacles in urban environments. Existing works typically adopted the free-space path loss model or the statistical channel model, which either ignored the impact of obstacles or assumed uniformly distributed obstacles and therefore might fail in practical NLoS scenarios. In this paper, taking the information of randomly distributed obstacles in environments into consideration, we aim to optimize the placement for the UAV-enabled multi-hop network to transfer more data collected by GRs and minimize the time delay in data transmission while satisfying the required communication quality. By reconstructing this complex non-convex optimization problem into two subprob-lems and solving them alternatively, we propose the multi-hop UAVs placement (mUP) method to get the solution, which contains the air-to-ground network formation (ATG-NF) algorithm and the communication quality-aware UAV placement (CQA-UP) algorithm. Simulation results show that in four types of typical urban environments or with different numbers of UAVs, the proposed mUP method achieves substantial performance gains in terms of communication quality and task performance compared to other placement approaches based on statistical channel models. We further discuss the robustness of the mUP method towards terrain measurement error.
Sining Yang, Dian-xi Shi, Yingxuan Peng, Shaowu Yang, Bo Zhang 0007, Wenjing Yang 0002
IPSN2
2022 Efficient Scale Divide and Conquer Network for Object Detection
Ziteng Qiao, Shaowu Yang, Dian-xi Shi
PRICAI (3)6
2022 FusionSeg: Motion Segmentation by Jointly Exploiting Frames and Events
Zhe Liu 0029, Shaowu Yang, Dian-xi Shi, Yongjun Zhang 0006
PRICAI (3)5
2022 Independent Multi-agent Reinforcement Learning Using Common Knowledge
abstract
Many recent multi-agent reinforcement learning algorithms used centralized training with decentralized execution (CTDE), which results in a training process that relies on global information and suffers from the dimensional explosion. The independent learning (IL) approaches are simple in structure and can be more easily deployed to a wider range of multi-agent scenarios, but they can only solve relatively simple problems due to environment non-stationarity and partially observable. With this motivation, we let IL agents compute common knowledge information and fuse it with observation to explicitly exploit common knowledge. In addition, we chose a suitable network structure according to the characteristics of IL, using convolutional layers and GRU layers. Based on the above two improvements, we implement two IL algorithms. In our experiments, the algorithms we implemented show significant performance improvements compared to original IL algorithms and further approach CTDE while outperforming multi-agent common knowledge reinforcement learning.
Haomeng Hu, Dian-xi Shi, Huanhuan Yang, Yingxuan Peng, Yating Zhou, Shaowu Yang
SMC2
2022 SCSE-E2VID: Improved event-based video reconstruction with an event camera
abstract
The recently emerging event camera has grown into a new type of sensor in the realm of vision, with benefits such as low power consumption, high dynamic range (HDR), microsecond time resolution, and no motion blur. While event cameras offer numerous advantages over conventional cameras, they only capture changes in intensity and give up lots of environmental details. This paper proposes an end-to-end UNet network called SCSE-E2VID to synthesize gray images from asynchronous events. We design an event fusion block to feed more related events to the encoder, allowing the network to extract more valuable features. The famous attention module called Spatial and Channel ‘Squeeze & Excitation’ Block (SCSE) is utilized to remove artifacts and better extract spatiotemporal features for the decoder. Besides, we add parallel convolutions in the upsampling block and refine the output features, which supplement content in reduced channels. In order to evaluate the performance of our proposed SCSE-E2VID, we implement quantitative and qualitative comparisons based on the public IJRR and HQF datasets. The results show that our method achieves better performance in terms of perceptual similarity and structural similarity when compared with state-of-art methods and demonstrates comparable performance in terms of squared error.
Dian-xi Shi, Ruihao Li 0001, Luoxi Jing, Shaowu Yang
SMC2
2022 Self-supervised representations for multi-view reinforcement learning
abstract
Learning policies from raw, pixel images are quite important for the real-world application of deep reinforcement learning (RL). Standard model-free RL algorithms focus on single-view settings and unify the representation learning and policy learning into an end-to-end training process. However, such a learning paradigm is sample-inefficiency and sensitive to hyper-parameters when supervised merely by the reward signals. Based on this, we present Self-Supervised Representations (S2R) for multi-view reinforcement learning, a sample-efficient representation learning method for learning features from high-dimensional images. In S2R, we introduce a representation learning framework and define a novel multi-view auxiliary objective based on the multi-view image states and Conditional Entropy Bottleneck (CEB) principle. We integrate S2R with the deep RL agent to learn robust representations that preserve task-relevant information while discarding task-irrelevant information and find optimal policies that maximize the expected return. Empirically, we demonstrate the effectiveness of S2R in the visual DeepMind Control (DMControl) suite and show its better performance on the default DMControl tasks and their variants by replacing the tasks’ default background with a random image or natural video.
Huanhuan Yang, Dian-xi Shi, Guojun Xie, Yingxuan Peng, Yantai Yang, Shaowu Yang
UAI2
2022 Multi actor hierarchical attention critic with RNN-based feature extraction
Dian-xi Shi, Chenran Zhao, Huanhuan Yang, Gongju Wang, Shaowu Yang, Yongjun Zhang 0006
Neurocomputing1
2022 PLC-VIO: Visual-Inertial Odometry Based on Point-Line Constraints
abstract
Visual–inertial odometry (VIO) is widely studied and used in autonomous robots. This article proposes a novel tightly coupled monocular VIO system based on point-line constraints (PLC-VIO). In the front end, PLC-VIO presents a line segment extraction and merging algorithm based on the EDLines method and achieves real-time feature tracking based on the geometric constraints between feature points and lines. In the back end, PLC-VIO reconstructs new 3-D landmarks of feature lines through points on the line and optimizes the states by minimizing a cost function that combines the preintegrated inertial measurement unit (IMU) error term together with the point and line reprojection error terms in a sliding window optimization framework. A loop closure module is also integrated, which enables relocalization and drift elimination. The corresponding experimental evaluations are conducted using public datasets to validate the effectiveness and robustness of the proposed system, and the results show that PLC-VIO can achieve good performance when compared with other state-of-the-art systems and, at the same time, with no compromise to real-time performance.Note to Practitioners—Visual–inertial odometry (VIO) can estimate the states of the rigid body (including position, attitude, and velocity) that is widely used in robotic navigation, autonomous driving, virtual reality (VR), and augmented reality (AR). Aiming at the problem of estimating the states of autonomous robots in the GPS-denied environment, this article proposes a novel VIO system based on the point-line constraints (PLC-VIO). PLC-VIO can not only achieve accurate pose estimation for robots due to the introduction of the line features but also make no concession to real-time performance. Furthermore, PLC-VIO can also enrich the texture features of the environment during the 3-D mapping construction. The corresponding experiments are implemented in public datasets to evaluate the effectiveness, efficiency, and robustness of the proposed system. We believe that PLC-VIO can be widely used in robotic navigation and AR/VR fields to provide accurate position and environment information in real time.
Zhe Liu 0029, Dian-xi Shi, Ruihao Li 0001, Yongjun Zhang 0006, Xiaoguang Ren
IEEE Trans Autom. Sci. Eng.2
2021 ContriQ: Ally-Focused Cooperation and Enemy-Concentrated Confrontation in Multi-Agent Reinforcement Learning
abstract
Centralized training with decentralized execution (CTDE) is an important setting for cooperative multi-agent reinforcement learning (MARL) due to communication constraints during execution and scalability constraints during training, which has shown superior performance but still suffers from challenges. One branch is to understand the mutual interplay between agents. Due to the communication constraints in practice, agents cannot exchange perceptual information, and thus, many approaches use a centralized attention network with scalability constraints. Contrary to these common approaches, we propose to learn to cooperate in a decentralized way by applying attention mechanism on the local observation so that each agent could focus on allied agents with a decentralized model, and therefore promote understanding. Another branch is to model how agents cooperate and simplify the learning process. Previous approaches that focus on value decomposition have achieved innovative results but still suffer from problems. These approaches either limit the representation expressiveness of their value function classes or relax the IGM consistency to achieve scalability, which may lead to poor performance. We combine value composition with game abstraction by modeling the relationships between agents as a bi-level graph. We propose a novel value decomposition network based on it through a bi-level attention network, which indicates the contribution of allied agents attacking enemies and the priority of attacking each enemy under the situation of each time step, respectively. We show that our method substantially outperforms existing state-of-the-art methods on battle games in StarCraft Ⅱ, and attention analysis is also comprehensively discussed with sights.
Chenran Zhao, Dian-xi Shi, Huanhuan Yang, Shaowu Yang, Yongjun Zhang 0006
ACML2
2021 CIExplore: Curiosity and Influence-based Exploration in Multi-Agent Cooperative Scenarios with Sparse Rewards
abstract
Learning in a sparse-reward setting is a well-known challenge in RL (Reinforcement Learning). In the single-agent domain, this challenge can be addressed by introducing exploration bonuses driven by intrinsic motivation to encourage agents to visit unseen states. However, naively applying these methods in MARL (Multi-Agent Reinforcement Learning) cooperative settings with sparse rewards results in some inevitable problems: misunderstanding environmental knowledge and lack of collaboration among agents, etc. Based on this, in this paper, we propose the Curiosity and Influence-based Explore (CIExplore) method, which includes a new form of intrinsic reward and an internal counterfactual advantage function. Concretely, the intrinsic reward is a combination of joint curiosity reward and influence reward. The former is the variance of outputs across an ensemble of prediction models that take joint observations and actions of all agents as inputs to predict the next time's joint observations. And the latter quantifies the influence of one agent's behavior on other agents' state-value functions. Given that the joint curiosity reward is shared by all agents, we compute an internal counterfactual advantage function to address this intrinsic reward assignment problem. We demonstrate the efficacy of CIExplore in the multi-agent grid-world environments and show that it is compatible with both on-policy and off-policy MARL algorithms and be scalable to complex settings where agents' number or environment randomness increases.
Huanhuan Yang, Dian-xi Shi, Chenran Zhao, Guojun Xie, Shaowu Yang
CIKM2
2021 Attention-Aware Actor for Cooperative Multi-agent Reinforcement Learning
Chenran Zhao, Dian-xi Shi, Yaqianwen Su, Yongjun Zhang 0006, Shaowu Yang
CollaborateCom (2)2
2021 Polar Loss for Event-Based Object Detection
abstract
Event cameras are bio-inspired sensors that respond to pixel-level brightness changes in the form of asynchronous and sparse events. Each event is a 4-D tuple, i.e., (timestamp, x, y, polarity). Recently, algorithms for object detection based on learning have made considerable strides for event cameras. These methods transform event sequences into a tensor-like representation which can be processed by deep learning methods. However, these conversion methods don’t make full use of the polarity information of event sequences. We discover that tensor-like representations of different polarities show different features. Thus, we propose Polar Loss, a loss function that can enhance the difference between tensor-like representations of different polarities. To evaluate the effectiveness of our loss function, we design a network architecture for object detection. Our results state that our model shows good flexibility and expressiveness on event-based object detection when trained with the Polar Loss.
Huachi Xu, Dian-xi Shi, Luoxi Jing
ICTAI2
2021 Two-stage Local Spatio-temporal Event Filter based on Adaptive Thresholds
abstract
Recently, the emerging event camera shows great potential for robotics and AR/VR application thanks to its advantages including low latency, high dynamic range, low power consumption, etc. However, the output of event cameras usually has a large amount of noise which will affect unlocking their potentials and obtaining wider applications. In this paper, we present a two-stage local spatio-temporal event filter (LSTEF) based on adaptive thresholds. Considering the spatio-temporal constraints of generated event streams, a local sliding window is adopted for the noise candidate selection stage and noise filtering stage. An adaptive thresholding mechanism is also introduced into the filter in order to improve the generalization performance. Corresponding experimental evaluations are performed on the public datasets to prove the efficiency of the proposed filter. Results show that the presented LSTEF can successfully achieve the event denoising, and at the same time, effectively preserve useful scene information.
Luoxi Jing, Jun Luo 0011, Dian-xi Shi, Ruihao Li 0001, Huachi Xu, Yuning Cui 0001
IJCNN3
2021 Optical Flow Estimation through Fusion Network based on Self-supervised Deep Learning
abstract
Compared with conventional cameras, dynamic vision sensors (namely event cameras) are especially suitable for high-speed and high-dynamic applications due to their advantages (high dynamic range, low latency, etc.). However, it still struggles when encountering low-speed and low-texture scenes. Aiming at the problem of optical flow estimation, we propose a novel unsupervised learning estimation method with both event data and gray image frames as the input. Two different fusion mechanisms are presented and discussed in our work, one is the Direct Fusion Network (DFEV-FlowNet), and the other one is Local Squeeze Extraction Network (LSENet). DFEV-FlowNet directly fuses synthesized event frames and gray image frames, and LSENet adopts a local squeeze extraction weights adaptive mechanism. In order to achieve self-supervised learning, photometric constraints between consecutive frames are used to drive the network training. We use the public event dataset MVSEC to evaluate the proposed optical flow estimation method qualitatively and quantitatively. The results show that our method demonstrate better performance in terms of the estimation accuracy.
Dian-xi Shi, Ruihao Li 0001, Huachi Xu
IJCNN2
2021 Coordinated Multi-UAV Adaptive Exploration Under Recurrent Connectivity Constraints
Yaqianwen Su, Dian-xi Shi, Jiachi Xu, Xionghui He
MobiQuitous2
2021 Joint Communication-Motion Planning for UAV Relaying in Urban Areas
abstract
In this paper, we consider a challenging surveillance scenario where there could exist line-of-sight (LOS) propagations and non-line-of-sight (NLOS) propagations in air-to-ground (ATG) channel and air-to-air (ATA) channel due to obstacles in urban areas, and a ground mobile robot is deployed to survey this area and transmit collected data to a remote base station via an unmanned aerial vehicle (UAV) relay. In this scenario, we aim to plan the optimal transmit power and trajectory of the UAV relay to minimize energy consumption while maintaining the communication quality. Existing works typically rely on the free-space path loss model and the statistical channel model, thus neglect the positions and shapes of obstacles and may fail in practical NLOS scenarios. In this paper, we first exploit the end-to-end packet error rate (PER)-based communication model, which captures the LOS propagation and NLOS propagation. Then, taken the information of obstacles in environment into consideration, we propose an UAV relay-assisted joint communication-motion planning (UAV-JCMP) method for minimizing the total energy consumption in urban areas. By decomposing the concave problem into two subproblems and dividing its domain into several convex subdomains according to LOS conditions, we further get the optimal solution. At last, numerical results demonstrate that substantial energy-efficient improvements can be achieved over methods that only optimize communication energy consumption and methods using statistical channel model. We further discuss the robustness of UAV-JCMP method towards terrain measurement error.
Sining Yang, Dian-xi Shi, Yingxuan Peng, Yongjun Zhang 0006
SECON2
2021 A Balanced Shadow-Following Coverage Path Planning Approach under Energy Constraints
abstract
Area coverage is one of the significant fundamental problems in mobile robotics. In this paper, we study the coverage path planning (CPP) problem of unmanned aerial vehicle(s) (UAVs) with limited energy. Most existing solutions for this problem adopted that the UAV returns to the charging station and replenishes energy, which ensures the complete coverage for the area of interest. However, to guarantee the integrity of the uncovered area, there is usually no gain to the cover task during the process of returning, which leads to poor efficiency. To address this problem, we propose a shadow-following (SF) strategy which performs the cover task when the UAV returns along the shadow path, improving the energy utilization and coverage efficiency. Then, we extend for multi-UAV scenarios and propose a load-balanced (LB) multi-UAV CPP algorithm, which minimizes the number of robots and assigns the balanced task to each UAV. Finally, we validate the results in experiments performed with a UAV and multiple UAVs. The results show that the SF and LB algorithm performs well in the coverage efficiency.
Dian-xi Shi, Lin Li 0075, Jiachi Xu
SMC2
2021 Multi-agent deep reinforcement learning with type-based hierarchical group communication
Dian-xi Shi, Gongju Wang, Yongjun Zhang 0006
Appl. Intell.2
2020 BiC-DDPG: Bidirectionally-Coordinated Nets for Deep Multi-agent Reinforcement Learning
Gongju Wang, Dian-xi Shi
CollaborateCom (2)2
2020 Distributed Reinforcement Learning with States Feature Encoding and States Stacking in Continuous Action Space
Dian-xi Shi, Xucan Chen
CollaborateCom (1)2
2020 Networked Multi-robot Collaboration in Cooperative-Competitive Scenarios Under Communication Interference
Dian-xi Shi, Yongjun Zhang 0006, Liujing Wang, Fujiang She
CollaborateCom (1)2
2020 IDNet: A Single-Shot Object Detector Based on Feature Fusion
abstract
This paper proposes a novel single shot network for object detection. The proposed network, termed IDNet, explores the strategies of the feature fusion to alleviate the scale variation problem in object detection. IDNet mainly consists of two feature fusion modules: an indirect feature fusion module (IF) and a direct feature fusion module (DF). The IF shares long-range dependencies within pyramidal layers and based on these information, IDNet learns to emphasize informative regions and suppress the less useful ones on each layer. The DF is a feature fusion strategy based on modified lateral connection inspired by feature pyramid networks (FPN). It utilizes the averaging operation to reduce the change of feature maps' order of magnitude during fusing features to further improve the performance for detecting small instances. Comprehensive experiments are performed and the results indicate the effectiveness of IDNet, which reaches 80.3 mAP on PASCAL VOC 2007 benchmark.
Yuning Cui 0001, Dian-xi Shi, Yongjun Zhang 0006, Qianchong Sun
ICTAI2
2020 Dual-template Siamese Network with Cross-Correlation Fusion for Tracking
abstract
Siamese network based trackers treat tracking as a maximum matching between the detection region and the template, in which the template is either fixed or updated. The template-fixed trackers degrade accuracy since the appearance variations of the object in the subsequent frames are lacked; while the template-updated trackers depress robustness once the tracking drift occurs. In this paper, we apply a dual-template tracking strategy in the Siamese region proposal network to compensate for the separate merits of fixed template and mutative template. As for the dual-template, the one template is fixed and the other template is constantly updated. Meanwhile, we integrate a learnable convolutional neural network to fuse two Cross-Correlation responses generated from two template tracking branches. Based on Siamese network, we try to utilize the complementary fusion of different template responses to promote tracking performances. Extensive experiments on five tracking benchmarks of OTB100, VOT2016, VOT2018, GOT-10k and UAV123 demonstrate that our approach achieves the state-of-the-art tracking performances.
Dian-xi Shi, Ying Kang, Zunlin Fan, Songchang Jin, Yongjun Zhang 0006
ICTAI2
2020 Multi-Agent Feature Learning and Integration for Mixed Cooperative and Competitive Environment
abstract
At present, most of the centralized training with decentralized execution (CTDE) multi-agent reinforcement learning (MARL) algorithms have good results in the research of homogeneous scenarios. Heterogeneous multi-agent scenarios with different roles, cooperation modeling and credit assignment problems lead difficulty to learn effective collective strategies. In this paper, we propose a method of feature learning and feature integration about cooperation. Specifically, in the aspect of feature learning, through graph attention network, the relationship between agents is simplified to graph adjacency matrix representation, so that their feature vectors have relationship attributes. At the same time, for feature integration, we use batch normalization (BN) method to concatenate trained feature. We expect that agent relations can be modeled by end-to-end design. Meanwhile, attention mechanism can enhance the communication between interrelated agents. Through the experiments, our method has a significant result on improving the cooperative-competitive scenario of heterogeneous multi-agent. Moreover, we can visualize the output to analyze the reasonable collaborative and emphases attack policy.
Dian-xi Shi, Yongjun Zhang 0006, Liujing Wang
ICTAI2
2020 Selective Feature Network for Object Detection
abstract
Scale variation is one of the important challenges in object detection. Many state-of-the-art objectors tackle this problem by utilizing the feature pyramids. However, the current methods of producing feature pyramids are still inefficient to integrate the semantic information from other layers. In this work, our motivation is to build a feature pyramid efficiently with the selected contextual feature by integrating the informative features and suppressing the useless ones. To achieve this goal, we propose a novel single-stage detection network termed Selective Feature Network(SFNet) which consists of a semantic-enhanced module and a selective feature module. The semantic-enhanced module improves the semantics of basic pyramids via a lightweight architecture. In conjunction with that, a selective feature module is employed to combine features across different channels and scales by attention mechanism. The resulting contextual feature is then injected into the pyramidal features. Comprehensive experiments are performed on PASCAL VOC and MS COCO datasets. Results demonstrate that, with a VGG16 based SFNet, our approach obtains significant improvements over the competitors without losing real-time processing speed.
Yuning Cui 0001, Dian-xi Shi, Yongjun Zhang 0006, Qianchong Sun
IJCNN2
2020 GHGC: Goal-based Hierarchical Group Communication in Multi-Agent Reinforcement Learning
abstract
In large-scale multi-agent systems, the existence of a large number of agents with different target tasks and connected by complex game relationships causes great difficulty for policy learning. Therefore, simplifying the learning process is an important issue. In multi-agent systems, agents with the same target tasks or attributes often interact more with each other and exhibit behaviors more similar. That means there are stronger collaborations between these agents. Most existing multi-agent reinforcement learning (MARL) algorithms expect to learn the collaborative strategies of all agents directly in order to maximize the common rewards. This causes the difficulty of policy learning to increase exponentially as the number and types of agents increase. To address this problem, we propose a goal-based hierarchical group communication (GHGC) algorithm. This algorithm divides the agents into different groups, and maintains the group's cognitive consistency through knowledge sharing. Subsequently, we introduce a group communication and value decomposition method to ensure cooperation between the various groups. Experiments demonstrate that our model outperforms state-of-the-art MARL methods on the widely adopted StarCraft II benchmarks across different scenarios, and also possesses potential value for large-scale real-world applications.
Dian-xi Shi, Gongju Wang, Yongjun Zhang 0006
SMC2
2020 Friend-or-Foe Deep Deterministic Policy Gradient
abstract
One of the toughest challenges in the multi-agent deep reinforcement learning (MADRL) is that when the opponents' policies change rapidly, the collaborative agents can't learn well to respond to the opponents' policies effectively. This may lead to a local optimum w.r.t. the learned policy of the collaborative agents may be only locally optimal to the opponents' current policies. To address this problem, we propose a novel algorithm termed Friend-or-Foe Deep Deterministic Policy Gradient (FD2PG), in which the cooperative agents can be trained more robust and have stronger cooperation ability in continuous action space. These collaborative agents can generalize easily and respond correctly, even if their opponents' policies alter. Inspired by the classic Friend-or-Foe Q-learning algorithm (FFQ), we introduce the idea of minimizing the foes and maximizing the friends into the centralized training distributed execution framework, multi-agent deep deterministic policy gradient algorithm (MADDPG), to enhance collaborative agents' robustness and cooperativity. Besides, we introduce a Minimax Multi-Agent Learning (MMAL) method to explore two special equilibriums (the adversarial equilibrium and the coordination equilibrium), which can guarantee the convergence of FD2PG and improve optimization. Extensive fine-grained experiments, including four representative scenario experiments and two scale-performance correlation experiments, were conducted to demonstrate the superior performance of FD2PG comparing with existing baselines.
Dian-xi Shi, Gongju Wang, Yongjun Zhang 0006
SMC2
2020 Ground-Based Target Localization Using Monocular Vision on a Quadcopter
abstract
Vision-based target localization in the air-to-ground scenes of quadcopter is a hot topic in the military and civilian applications. The current localization approaches mostly rely on the intrinsic parameters of camera which are needed to calibrate in advance. But if the calibration is inaccurate, the localization based on the focal length of camera will result in inaccuracy. Moreover, due to the autonomous flight of quadcopter, the pose changes of the quadcopter will lead to the erroneous locating results. In this paper, we propose an accurate ground-based target localization method by means of monocular vision on a quadcopter. Based on the field of view (FOV) angle of camera, we present a localization algorithm requiring no calibration. Meanwhile, we introduce the quadcopter pose superposition strategy and the IMU quaternion denoising method to eliminate the impact of quadcopter pose changes in flight. The simulation experimental results demonstrate the high performances of our locating method in both accuracy and robustness.
Dian-xi Shi, Ying Kang, Zunlin Fan
SMC2
2020 AHAC: Actor Hierarchical Attention Critic for Multi-Agent Reinforcement Learning
abstract
Deep reinforcement learning has made significant progress in multi-agent tasks in recent years. However, most previous studies focus on solving full cooperative tasks, which do not perform well in mixed tasks. In mixed tasks, the agent needs to comprehensively consider the information provided by friends and enemies to learn its strategy, and its strategy is sensitive to the received information. There is a great necessity to efficiently learn information representation for mixed tasks. To this end, we present an approach that conducts information representation learning for multiple agents using hierarchical attention mechanism. Our approach adopts the framework of centralized training and decentralized execution. It applies hierarchical attention to centrally computed critics, so critics process the received information more accurately and assist actors to choose better actions. The hierarchical attention critic uses two different attention levels, the agent-level and the group-level, to assign different weights to information of friends and enemies respectively and then summarize them at each time-step. It can achieve more effective and scalable learning in mixed tasks. In addition, our approach uses recurrent neural networks that process sequence input information more efficiently. Experimental results show that our approach is not only applicable to cooperative environments but also better in mixed environments. Especially in the predator-prey task, our approach receives twice as much reward as baselines.
Dian-xi Shi, Gongju Wang
SMC2
2020 Balanced Multi-Region Coverage Path Planning for Unmanned Aerial Vehicles
abstract
Nowadays, Unmanned Aerial Vehicles (UAVs) are playing increasingly important roles in agriculture, rescuing and surveillance due to their small size, low cost and strong adaptability. Coverage Path Planning (CPP) is a fundamental problem for UAV applications, which means to find a path covering all the targets or regions of interest. Researches on CPP in a single region have been studied for decades, but rare to be devoted to covering multiple scattered regions of multiple UAVs. This paper proposes an attempt to solve this problem in a short time and take the task balance of multiple UAVs into account meanwhile. This work constructs a model for multiple UAVs to cover scattered regions firstly. In view of the high computational complexity of the precise solution, we keep innovating a heuristic measure based on the model to make the solving process feasible. To settle the problem of imbalanced time consumption caused by super regions, we improve the previously built model furthermore. A series of experiments validated that the presented approaches exhibit the effectiveness and balance of time consumption in diverse scenarios.
Xiaoxiao Yu, Songchang Jin, Dian-xi Shi, Lin Li 0075, Ying Kang, Junbo Zou
SMC3
2019 Multi-Target Multi-Camera Tracking with Human Body Part Semantic Features
abstract
Recently, Multi-Target Multi-Camera Tracking (MTMCT) has gained more and more attention. It is a challenging task with major problems including occlusion, background clutter, poses and camera point of view variations. Compared to single camera tracking, which takes advantage of location information and strict time constraints, good appearance features are more important to MTMCT. This drives us to extract robust and discriminative features for MTMCT. We propose MTMCT\_HS which uses human body part semantic features to overcome the above challenges. We use a two-stream deep neural network to extract the global appearance features and human body part semantic maps separately, and employ aggregation operations to generate final features. We argue that these features are more suitable for affinity measurement, which can be seen as the average of appearance similarity weighted by the corresponding human body part similarity. Next, our tracker adopts a hierarchical correlation clustering algorithm, which combines targets' appearance feature similarity with motion correlation for data association. We validate the effectiveness of our MTMCT\_HS method by demonstrating its superiority over the state-of-the-art method on DukeMTMC benchmark. Experiments show that the extracted features with human body part semantics are more effective for MTMCT compared with the methods solely employing global appearance features.
Mingkun Wang, Dian-xi Shi, Naiyang Guan, Zunlin Fan
CIKM2
2019 Unsupervised Pedestrian Trajectory Prediction with Graph Neural Networks
abstract
Trajectory prediction can aid in target tracking, automatic system navigation, social behavior prediction, analysis and other computer vision tasks. When people walk in crowded spaces such as sidewalks, subway and airports, etc., they naturally adjust their walking style according to the scene context and follow common social etiquette, such as maintaining separation and avoiding collisions. Accurate prediction of trajectories is a big challenge in a crowded scenario where interaction between targets may cause complex societal dynamics. Unlike the prediction of a single person trajectory, it is difficult to capture the real motion of multiple people by only considering the historical positions of each individual separately. Benefit from the recent success of graph neural networks, we propose a model called GNN-TP for pedestrian trajectory prediction. GNN-TP is a purely data-driven model that simultaneously infers the interactions between pedestrians in an unsupervised way and predicts their future trajectories jointly in crowded scenes. On the one hand, GNN-TP infers the interactions employing the observed historical trajectories. We transfer pedestrians' information on the graph-structured data and classify the interaction type based on edges' features. On the other hand, it learns the dynamical model and predicts future trajectories based on the inferred interactions and the observations. Extensive experiments show that our trajectory prediction model achieves efficient and state-of-the-art performance on several public datasets.
Mingkun Wang, Dian-xi Shi, Naiyang Guan, Liujing Wang, Ruoxiang Li
ICTAI2
2019 FA-Harris: A Fast and Asynchronous Corner Detector for Event Cameras
abstract
Recently, the emerging bio-inspired event cameras have demonstrated potentials for a wide range of robotic applications in dynamic environments. In this paper, we propose a novel fast and asynchronous event-based corner detection method which is called FA-Harris. FA-Harris consists of several components, including an event filter, a Global Surface of Active Events (G-SAE) maintaining unit, a corner candidate selecting unit, and a corner candidate refining unit. The proposed G-SAE maintenance algorithm and corner candidate selection algorithm greatly enhance the real-time performance for corner detection, while the corner candidate refinement algorithm maintains the accuracy of performance by using an improved event-based Harris detector. Additionally, FA-Harris does not require artificially synthesized event-frames and can operate on asynchronous events directly. We implement the proposed method in C++ and evaluate it on public Event Camera Datasets. The results show that our method achieves approximately 8× speed-up when compared with previously reported event-based Harris detector, and with no compromise on the accuracy of performance.
Ruoxiang Li, Dian-xi Shi, Yongjun Zhang 0006, Kaiyue Li, Ruihao Li 0001
IROS2
2018 Multi-UAV Collaborative Monocular SLAM Focusing on Data Sharing
Zhuoyue Yang, Dian-xi Shi, Yongjun Zhang 0006, Shaowu Yang, Ruoxiang Li
ICONIP (7)2
2017 A novel orientation- and location-independent activity recognition method
Dian-xi Shi, Xiaoyun Mo
Pers. Ubiquitous Comput.1
2015 Auxo: an architecture-centric framework supporting the online tuning of software adaptivity
Huaimin Wang 0001, Bo Ding 0001, Dian-xi Shi, Jiannong Cao 0001, Alvin Chan Toong Shoon
Sci. China Inf. Sci.3
2012 LBSNRank: personalized pagerank on location-based social networks
abstract
Different from traditional social networks, the location-based social networks allow people to share their locations according to location-tagged user-generated contents, such as checkins, trajectories, text, photos, etc. In location-based social networks, which are based on users' checkins, people could share his or her location according to checkin while visiting around. However, people's locations change frequently and the rankings of people change dynamically too, which makes ranking on graphs a challenging work. To address this challenge, we propose the LBSNRank algorithm on graphs with nodes whose contents change dynamically. To validate our algorithm on real datasets, we have crawled and analyzed a dataset from the Dianping website. Experiments on this real dataset show that our LBSNRank algorithm performs better than traditional personalized PageRank in efficiency.
Zhaoyan Jin, Dian-xi Shi, Quanyuan Wu, Huining Yan
UbiComp2
2010 Taming software adaptability with architecture-centric framework
abstract
In many cases, we would like to enhance the predefined adaptability of a running application, for example, to enable it to cope with a strange environment. To make such kind of runtime modifications is a challenging task. In existing engineering practices, the online policy upgrade approach just focuses on the modification of adaptation decision logic and lacks system-level means to assess the validity of an upgrade. This paper proposes a framework for adaptive software that supports the online reconfiguration of each concern in the “sensing-decision-execution” adaptation loop. To achieve this goal, our framework supports an architecture style which encapsulates adaptation concerns as software architecture elements. And then, it maintains a runtime architecture model to enable the dynamic reconfiguration of those elements as well as help to ensure the validity of a change. A third party can selectively add, remove or replace part of this model to enhance the running application's adaptability. We validated this framework by two cases extracted from real life.
Bo Ding 0001, Huaimin Wang 0001, Dian-xi Shi, Jiannong Cao 0001
PerCom3
2009 Towards Unanticipated Adaptation: An Architecture-Based Approach
abstract
Over its lifetime, adaptive software may have to deal with the environment not anticipated during the original development. In such cases, we should introduce new adaptive code, for example, to detect the strange contexts or update the out-of-date adaptation decision logic. This paper proposes an engineering approach facilitates this kind of post-delivery modifications based on software architecture techniques. Our approach introduces a component model separates different adaptation concerns (sensing, decision and execution) as different types of software architecture elements. The clear separation lays the foundation for the independent maintenance of each concern. And then, with the aid of a container supports the instantiation and run-time modification of the software architecture model, those concerns can be bound together without recompiling the whole software, even while it is running. Our approach enables the fine-grained, low-cost modifications of delivered adaptive software in the case that an unanticipated environment emerges.
Bo Ding 0001, Huaimin Wang 0001, Dian-xi Shi, Xiang Rao
SERA3
2008 Towards Role Based Trust Management without Distributed Searching of Credentials
Gang Yin, Huaimin Wang 0001, Dian-xi Shi
ICICS5
2008 Component Based Context Model
abstract
Context awareness is one of the most fundamental issues in pervasive computing. In this paper, component based context model based on middleware architecture is proposed. Moreover, OWL-based context ontology for modelling context information to easily share and reuse context knowledge is presented. By giving fire alarm scenario for our prototype, the proposed component based context architecture can be deployed in different context-aware application and can provide a middleware support for context representation and knowledge sharing.
Bo Ding 0001, Huaimin Wang 0001, Dian-xi Shi
WAIM4
2004 An Authorization Framework Based on Constrained Delegation
Gang Yin, Meng Teng, Huaimin Wang 0001, Yan Jia 0001, Dian-xi Shi
ISPA5