VLDB 2026 Research / reviewers in the wild / expert
Shaorong Xie
dblp:76/4084
· DBLP profile ↗
118ranked-venue papers
7as first author
93since 2021 · last 2026
0000-0002-8016-9310ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 65 · 4 first-author · 56 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 7 since 2021Systems, architecture and hardware · 11 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 8 · 3 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Beyond Vision: Vision-Language Distillation and Edge-Aware Mix Diffusion in Semi-Supervised Semantic SegmentationabstractIn semi-supervised semantic segmentation (SSSS), segmentation performance is heavily constrained by the quality of pseudo labels. However, prevalent pseudo-label optimization approaches rely on the model’s internal self-correction. When the model fails to recognize or adequately represent certain classes, this self-enhancement mechanism amplifies initial mistakes, ultimately leading to poor semantic or spatial consistency. To address this limitation, we propose ViLaDiff to enhance pseudo-label quality. Specifically, ViLaDiff first employs a prompt-guided image captioning task to generate descriptive text for each input image, providing high-level semantic context. To our knowledge, this is the first attempt to introduce vision-language modeling into SSSS. We design a vision-language fusion module to enhance feature semantics and discriminative capability. It integrates cross-modal interactions with dual-path knowledge to ensure semantic consistency. Additionally, while language provides high-level semantic guidance, it is inherently limited in expressing fine-grained spatial structures. Therefore, we propose an edge-aware mixed-noise diffusion process. It simulates feature-level uncertainty through Gaussian perturbations and introduces class-flipping noise into the masks to model misclassification errors. To enhance boundary refinement, we apply a higher flipping probability along mask edges, enabling edge-aware modeling during denoising. Extensive experiments on public benchmarks validate that our method significantly improves pseudo-label quality and segmentation performance. Yuehua Liu, Xiaomao Li, Shaorong Xie |
AAAI | 5 |
| 2026 | Risk-Aware Safe Reinforcement Learning with CVaR-Based Constraints
Shaorong Xie, Tao Wang 0072, Xiangfeng Luo |
ICIC | 2 |
| 2026 | FKQG: Few-shot question generation from knowledge graph via large language model in-context learning
Ruishen Liu, Shaorong Xie, Xinzhi Wang 0001, Xiangfeng Luo, Hang Yu 0006 |
Data Knowl. Eng. | 2 |
| 2026 | Spatial-temporal clues based interactive message aggregation for multi-agent collaboration
Shaorong Xie, Xiangfeng Luo, Zhenyu Zhang 0013, Hang Yu 0006 |
Frontiers Comput. Sci. | 1 |
| 2026 | Scalable multi-agent reinforcement learning with group interaction-based role adaptation
Tao Wang 0072, Xiangfeng Luo, Shaorong Xie |
Neurocomputing | 5 |
| 2026 | TAPPS: Learning diverse policies via task-adaptive partial parameter sharing in heterogeneous multi-agent reinforcement learning
Zhenyu Zhang 0013, Xiangfeng Luo, Shaorong Xie |
Knowl. Based Syst. | 6 |
| 2026 | Dilated memory in hierarchical reinforcement learning for long-horizontal task
Zhenyu Zhang 0013, Shaorong Xie, Xiangfeng Luo |
Neural Networks | 2 |
| 2026 | Enhancing Reliability in Medical Image Classification of Imperfect ViewsabstractThe fusion of multi-view medical images through deep neural networks is essential for boosting diagnostic precision in the field of medical image analysis. However, the reliability of these diagnostic results is often compromised by imperfections in image views, manifested as noise, artifacts, and data deficits arising from inconsistent diagnostic frequencies. These issues introduce a significant risk when merging medical views in a clinical setting. To address these problems, we introduce the Reliability-Enhanced Multi-view Network (REMNet), a novel framework designed to tackle two critical challenges: 1) reducing misclassification and uncertainty from imperfect view integration, and 2) improving the reliability and interpretability of multi-view medical image predictions. Specifically, REMNet merges information from multiple views into a coherent evidence framework and incorporates a Dirichlet prior within our predictive model to more accurately estimate confidence in predictions. Coupled with a robust fusion strategy and a precise confidence calibration process, REMNet consolidates the diverse strengths of various medical imaging views, reduces the impact of view imperfections, and enhances the reliability of medical imaging diagnostics. The superiority of REMNet is validated through comprehensive theoretical analysis and empirical experiments on multi-view medical image datasets across different modalities. Wei Liu 0303, Yufei Chen 0002, Xiaodong Yue 0002, Changqing Zhang 0002, Shaorong Xie |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Sharper insights: Discrimination-enhanced semantic segmentation of similar morphology nuclei in breast cancer pathology imagesabstractSimilar Morphology Cells (SMCs) play a critical role in pathology diagnosis. This study identifies and defines two primary challenges for SMCs segmentation based on detailed investigation. At the macro level of slides, the sample distribution within pathology images is highly imbalanced due to varying cell functionalities. At the micro level of cells, differences between SMCs are imperceptible. These two challenges lead to confusion between visually similar cell classes in SMCs segmentation, reducing segmentation accuracy. To address these issues, we propose Sharper Insights (SI), a novel approach designed to enhance the model’s discriminative capabilities. Specifically, we introduce a Discrimination Enhancement Module (DEM) to increase the model’s sensitivity to subtle variations among SMCs. This module utilizes two attention mechanisms to sharpen the model’s focus on minute differences. Additionally, a novel loss function, Minority Categories Augmentation (MCA), is introduced to mitigate distribution imbalance by dynamically adjusting category-specific weights. Finally, to verify the effectiveness and generalization of the proposed method SI, we curate a comprehensive breast cancer dataset SimMorphBC. It primarily focuses on SMCs, which have not been explicitly proposed in previous datasets. Extensive experiments on breast cancer pathology images demonstrate the superior performance of the proposed approach. Guoxin Sun, Yuehua Liu, Liming Xin, Xiao Zou, Shaorong Xie |
Vis. Informatics | 7 |
| 2025 | Integrating Low-Level Visual Cues for Enhanced Unsupervised Semantic SegmentationabstractUnsupervised semantic segmentation algorithms aim to identify meaningful semantic groups without annotations. Recent approaches leveraging self-supervised transformers as pre-training backbones have successfully obtained high-level dense features that effectively express semantic coherence. However, these methods often overlook local semantic coherence and low-level features such as color and texture. We propose integrating low-level visual cues to complement high-level visual cues derived from self-supervised pre-training branches. Our findings indicate that low-level visual cues provide a more coherent recognition of color-texture aspects, ensuring the continuity of spatial structures within classes. This insight led us to develop IL2Vseg, an unsupervised semantic segmentation method that leverages the complementation of low-level visual cues. The core of IL2Vseg is a spatially-constrained fuzzy clustering algorithm based on color affinities, which preserves the intra-class affinity of spatially-adjacent and similarly-colored pixels in low-level visual cues. Additionally, to effectively couple low-level and high-level visual cues, we introduce a feature similarity loss function to optimize the feature representation of fused visual cues. To further enhance consistent feature learning, we incorporate contrast loss functions based on color invariance and luminosity invariance, which improve the learning of features from different semantic categories. Extensive experiments on multiple datasets, including COCO-Stuff-27, Cityscapes, Potsdam, and MaSTr1325, demonstrate that IL2Vseg achieves state-of-the-art results. Yuhao Qing, Dan Zeng 0001, Shaorong Xie, Kaer Huang, Yueying Wang |
AAAI | 3 |
| 2025 | STSR-Nav: Spatial-Temporal Scene Representation for Navigation Policy Learning Based on UAV-UGV Collaborative PerceptionabstractIn UAV-UGV collaboration, UGVs enhance their perception capabilities by utilizing shared observation data from UAVs to make effective decisions. However, learning navigation policies directly from multi-perspective observation data is challenging due to spatial inconsistencies and redundancy. Existing approaches either rely on low-level scene representations, which may reduce efficiency, or use mid-level semantic representations that lack essential geometric and temporal details, resulting in sub-optimal decisions. This paper presents STSR-Nav, a novel framework for navigation policy learning based on UAV-UGV collaborative perception. The framework first encodes semantic information and geometric relationships from UAV images using a Bird'View (BEV) scene graph. A learnable graph neural network then integrates the BEV scene graph with local perception to infer spatial features. Temporal learning is applied to these features via a memory network, generating spatial-temporal features that are embedded into the UGV's local perception space using a multi-head attention mechanism. Finally, the fused spatial-temporal features provide a high-level scene representation for navigation policy learning. Unity simulations demonstrate the framework's improved efficiency and generalization capabilities. Shaorong Xie, Xiangfeng Luo, Xinzhi Wang 0001, Zhenyu Zhang 0013, Tao Wang 0072 |
CSCWD | 2 |
| 2025 | Rethinking Decoding in Multi-intent Spoken Language UnderstandingabstractMulti-intent spoken language understanding (SLU) can handle multiple intent utterances in real-world scenarios, which has gained increasing research attention. Despite promising results achieved by existing joint models, they (1) perform utterance-level or token-level intent detection, resulting in suboptimal performance due to fixed thresholds or voting mechanisms; (2) incorporate all predicted intents into each slot hidden state and execute parallel slot decoding, which lacks precise intent-slot alignment and overlooks sequential dependencies between slot tokens. In this paper, we propose a new framework to tackle these two issues. For the first issue, we utilize a global pointer with auxiliary tasks to achieve span-based intent detection. For the second issue, we leverage span-based predicted intents for precise intent-slot guidance and introduce rotational position encoding in the interaction module to explicitly model sequential dependencies for precise slot filling. Experimental results on two benchmarks demonstrate the superiority of our framework. Zhen Xiong, Kefan Shen, Zhihong Zhu 0001, Shaorong Xie, Wei Liu 0027 |
ICASSP | 5 |
| 2025 | Zero-Shot Scene Graph Generation with Bias Correction and Unseen Space Optimization
Yinsai Guo, Liyan Ma, Shaorong Xie |
ICIC (6) | 4 |
| 2025 | HACL: A Hybrid Adaptive Curriculum Learning Framework for Multi-modal Sarcasm DetectionabstractMulti-modal sarcasm detection (MSD) aims to identify sarcasm by analyzing inconsistencies across image-text pairs. Despite promising results achieved, existing methods predominantly focus on model-level improvements, while overlooking data-level challenges. In this paper, we introduce HACL, a Hybrid Adaptive Curriculum Learning framework for MSD. Concretely, we first propose a Hybrid Difficulty Measurer (HDM) that quantifies sarcasm difficulty at both representation and prediction levels, where representation-level difficulty captures cross-modal semantic conflicts, and prediction-level difficulty reflects model learning progress via smoothed loss. Furthermore, an Adaptive Weight Optimizer (AWO) is proposed to dynamically balance these two difficulty metrics, prioritizing implicit sarcasm cues and hard samples. Combined with a competence-based curriculum scheduler, HACL progressively trains models from easy to hard samples. Experiments on two benchmarks demonstrate state-of-the-art performance, with ablation studies validating the necessity of HDM and AWO. Further analyses highlight HACL’s generalizability on multi-modal sentiment analysis (MSA) and its superiority over large vision-language models (LVLMs). Kefan Shen, Yukang Huang, Wenyao Wang, Shaorong Xie, Zhihong Zhu 0001, Wei Liu 0027 |
IJCNN | 4 |
| 2025 | ESSI-KG: Enhancing Structural Semantic Integration in KG-to-Text Pretraining ModelsabstractKnowledge graph-to-text (KG-to-text) generation involves generating fluent and faithful text from structured data within knowledge graphs. Existing methods typically linearize the graph data and use pre-trained models, but they often neglect the complexity of graph structures and relations, limiting their ability to capture information effectively. In this paper, we propose a hierarchical attention model named ESSI-KG, which utilizes positional embeddings and attention mechanisms to enhance the representation and learning of relational information within knowledge graphs. We design three types of positional visualization embedding layers to encode graph structure information into sequences, thereby more effectively capturing the local information within triples in the knowledge graph. Additionally, we introduce entity attention layers and relation attention layers to retain global graph structure information and focus on relations among triples in the knowledge graph. Extensive experiments conducted on four benchmark datasets provide compelling evidence that ESSI-KG outperforms existing state-of-the-art models, demonstrating enhanced logical consistency and producing more accurate and informative textual outputs. Our model attains a BLEU-4 score that exceeds other baselines by over 1.27 on WebNLG. Specifically, under the few-shot settings, ESSI-KG reaches baseline performance with only one-twentieth of the labeled examples typically required. Xin Tie, Ruishen Liu, Xiangfeng Luo, Shaorong Xie, Xinzhi Wang 0001 |
IJCNN | 4 |
| 2025 | HARFT: Role and Function-Based Cooperation Approach for Heterogeneous Multi-AgentabstractHeterogeneous multi-agent reinforcement learning algorithms have been introduced to tackle complex real-world tasks such as gaming, disaster rescue, and navigation. The ability of heterogeneous agents to collaborate effectively is crucial for solving these complex tasks. However, the vast joint action and state space makes achieving cooperation among heterogeneous agents particularly challenging. Existing heterogeneous agent algorithms struggle to facilitate effective cooperation, and current role-based methods are not applicable due to the differing observation and action dimensions of heterogeneous agents. This paper proposes a novel algorithm called the Heterogeneous Agent Role Function Tree (HARFT), which forms corresponding roles based on the different functions of heterogeneous agents, allowing them to cooperate effectively. Additionally, we introduce the Heterogeneous Agent Role Predictive Network (HARPN), which enhances cooperation by predicting the current roles of other agents and adjusting collaboration based on their different functions. Finally, the effectiveness of our algorithm is validated in both the SMAC environment and a heterogeneous agent cooperative environment. Shaorong Xie, Xiangfeng Luo |
IJCNN | 2 |
| 2025 | Feature Fusion Contrastive Learning and Adaptive Correction Framework for Unsupervised Graph Anomaly DetectionabstractUnsupervised graph anomaly detection methods have attracted significant research interest in recent years. However, existing graph neural network models face a key challenge: they unintentionally integrate the behavioral features of normal nodes into the features of abnormal nodes during the aggregation process. This problem misleads model training, weakens the generalization, and complicates the anomaly detection. To address this challenge, we propose a Feature Fusion Contrastive Learning and Adaptive Correction framework, called FFCAC. The framework introduces a multi-scale feature fusion and contrastive learning approach that decomposes raw node features into multiple subspaces and applies multi-scale convolution to capture higher-order behavioral representations of nodes. Through contrastive learning, FFCAC computes the anomaly scores of node features, replacing the potentially misleading message passing mechanism commonly used in traditional GNNs. In addition, a Transformer-based correction mechanism is used to refine the anomaly scores by modelling global structural relationships through learned weights. This design is able to adaptively aggregate neighbourhood information, which enhances the generalization ability of the model while effectively identifying camouflaged abnormal nodes. Experimental results on multiple open-source datasets demonstrate that FFCAC achieves significant improvements in unsupervised graph anomaly detection compared to existing methods. Shaorong Xie, Junquan Gu, Xiangfeng Luo, Hang Yu 0006 |
IJCNN | 2 |
| 2025 | Test Time Prompt Adaption for Domain GeneralizationabstractDomain generalization (DG) has achieved significant progress with the emergence of pre-trained vision-language models such as CLIP. Advanced works have explored prompt learning to adapt these models to specific target domains. However, most existing methods primarily focus on designing better prompts to improve performance on the target domain, overlooking the importance of dynamically adjusting the model based on the distribution of test data. In this work, we propose a novel Test-time Prompt Adaptation (TPA) method, which dynamically adjusts prompts during inference by leveraging both image-level and feature-level augmentations of test samples. Specifically, for each test instance, we generate multiple augmented views and select representations that exhibit consistency to construct a support set, which is then used to dynamically refine the prompts. During inference, we utilize a key-value cache model built from the support set to create adaptive weights, guiding predictions based on the similarity between the test sample and the support set in a fully non-parametric manner. Without any additional training, this efficient and effective approach achieves significant performance improvements. We conduct extensive experiments on ImageNet and 10 additional datasets, demonstrating the superiority of our method over existing approaches. Haiting Zheng, Shaorong Xie, Hang Yu 0006 |
IJCNN | 3 |
| 2025 | Unsupervised adaptive learning method for salient object detection under weak observation conditions
Ying Tong, Xiangfeng Luo, Liyan Ma, Shaorong Xie |
Appl. Intell. | 4 |
| 2025 | Embedding prescribed-time adaptive control protocol unveiling distributed consensus in multirobot systems via directed topology
Yonghao Xie, Xinru Ma, Shaorong Xie |
Sci. China Inf. Sci. | 4 |
| 2025 | Spatial-temporal intention representation with multi-agent reinforcement learning for unmanned surface vehicles strategies learning in asset guarding task
Yang Li 0151, Shaorong Xie, Hang Yu 0006, Zhenyu Zhang 0013, Xiangfeng Luo |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Saliency and correlation learning for co-salient object detection
Ying Tong, Xiangfeng Luo, Liyan Ma, Shaorong Xie |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | SRAD: Autonomous Decision-Making Method for UAV Based on Safety Reinforcement LearningabstractABSTRACT Unmanned aerial vehicles (UAVs) are increasingly vital across numerous sectors, from logistics and rescue operations to military endeavours and beyond. However, ensuring safety in the decision‐making processes surrounding UAV operations in real‐world settings has become an urgent and complex challenge. At present, the main methods to minimise the risk of drone decision‐making include utilising pre‐established control rules, expert prior knowledge and regularisation constraints. However, these methodologies require UAVs to meet demanding prerequisites, including the acquisition of extensive decision‐making experience and the establishment of comprehensive rules. Regrettably, these strict requirements often lead to frequent UAV crashes in uncertain environments and subsequent mission failures. In order to tackle these issues, we propose a self‐decision‐making method for quadcopter UAVs based on safe reinforcement learning. Our method utilises a multilevel cascading feature semantic space for reinforcement learning, integrating depth images, greyscale images, semantic segmentation images and object detection results as inputs. This approach aims to facilitate safe autonomous learning. Moreover, we integrate real offline labelled data to enhance the safety policy. Depending on the varying levels of risk encountered during the UAV's decision‐making process, we dynamically select different safety policies. Through this iterative process, the UAV progressively eliminates extreme actions and reverts to the UAV learning policy module. Experimental results indicate that our method not only ensures safe decision‐making for UAVs in uncertain environments but also exhibits superior safety decision‐making efficacy compared to certain baseline methods. Xiangfeng Luo, Shaorong Xie |
Expert Syst. J. Knowl. Eng. | 3 |
| 2025 | FL-Evo: Jointly modeling fact and logic evolution patterns for temporal knowledge graph reasoning
Ruishen Liu, Xinzhi Wang 0001, Shaorong Xie, Xiangfeng Luo, Huizhe Su |
Expert Syst. Appl. | 3 |
| 2025 | Beyond expression: Comprehensive visualization of knowledge triplet facts
Wei Liu 0027, Yixue He, Chao Wang 0095, Shaorong Xie, Weimin Li 0001 |
Inf. Process. Manag. | 4 |
| 2025 | Fuzzy knowledge inference-based dynamic task allocation method for multi-agent systems
Xinzhi Wang 0001, Xiangfeng Luo, Shaorong Xie |
Inf. Sci. | 5 |
| 2025 | Improving inference via rich path information and logic rules for document-level relation extraction
Huizhe Su, Shaorong Xie, Hang Yu 0006, Changsen Yuan, Xinzhi Wang 0001, Xiangfeng Luo |
Knowl. Inf. Syst. | 2 |
| 2025 | Contrastive zero-shot relational learning for knowledge graph completion
Zhiyi Fang, Hang Yu 0006, Changhua Xu, Zhuofeng Li, Shaorong Xie |
Knowl. Based Syst. | 6 |
| 2025 | Representation purification with semantic contrastive learning for hyper-relational knowledge graph
Xiangfeng Luo, Xinzhi Wang 0001, Shaorong Xie, Ruoxin Zheng |
Knowl. Based Syst. | 4 |
| 2025 | Intention-guided imitation learning methods under limited expert demonstration data
Xiangfeng Luo, Shaorong Xie |
Knowl. Based Syst. | 3 |
| 2025 | Co-saliency guided multi-modal learning for referring video object segmentation
Ying Tong, Xiangfeng Luo, Liyan Ma, Shaorong Xie |
Knowl. Based Syst. | 4 |
| 2025 | Efficient offline-to-online reinforcement learning with pre-reduced out-of-distribution Q-values
Tao Wang 0072, Xiangfeng Luo, Zhenyu Zhang 0013, Shaorong Xie |
Neural Networks | 4 |
| 2025 | High Specificity Guided Cross-Domain Few-Shot SegmentationabstractCross-domain few-shot segmentation (CD-FSS) is a challenging vision task that involves segmenting novel classes from unseen domains using only a few annotated examples. Recent methods typically rely on powerful fundamental models for feature extraction, combined with complex, parameter-based decoders for segmentation. Due to the scarcity of annotated samples in the target domain, training a large number of parameters from scratch only on the source domain can lead to overfitting, which harms generalization in cross-domain set tings. These methods often assume that the fundamental feature extractor provides a sufficiently robust feature space, and freeze it during training to reduce the number of parameters. However, there is an inherent discrepancy between the tasks of training the fundamental model and performing segmentation, leading to a feature space mismatch. This misalignment results in features that, while useful for general tasks, may not fully satisfy the specific requirements of the segmentation task. To address this issue, we propose a metric-based approach called the High Specificity Guided Prototype Network (HSGNet). Our method is lightweight and focuses on fine-tuning the embedding network to align the feature space for the segmentation task. Specifically, we introduce a novel Feature Enrichment Module with extremely few parameters, which enhances the embedding network's ability to better align with the segmentation requirements. Instead of using a parameter-based decoder, our approach employs a non-parameter, similarity-based, high-specificity segmentation strategy. Additionally, we introduce a non-parameter test-time refinement mechanism to further improve prediction accuracy. Extensive experiments on cross-domain benchmarks demonstrate that our method achieves state-of-the-art performance with minimal additional parameters. Pinzhuo Tian, Hang Yu 0006, Shaorong Xie |
IEEE Trans. Multim. | 4 |
| 2025 | Adaptive Dynamics-Based Prescribed-Time Control for Robots Formation Tracking in Task SpaceabstractThis article investigates the prescribed-time formation control in the task space of multirobot systems (MRSs), which is subject to the uncertain nonlinear dynamics and the position requirements. The strategy constructs a cascade system consisting of control and reference layers by connections of coupled prescribed-time control units, which separately guarantee the convergence of coordination errors, accuracy of states, and synchronization between layers. Meanwhile, the transformation from task space to joint space is built based on the pseudo-inverse of the Jacobi matrix, which avoids the singular value problem brought by computing the inverse kinematics. The adaptive control method utilizes parameter estimation to eliminate the inaccuracy problem brought by the pseudo-inverse of Jacobi matrix transformation. Then, this article provides solutions to the formation control and the formation along the trajectory control. Correspondingly, the Lyapunov analysis process proves the system’s stability and the parameter estimation’s boundedness, which confirms the sufficient conditions for realizing the prescribed-time formation control of the MRSs. Finally, this article presents examples of time-varying formation and along-trajectories formation, thereby demonstrating the effect of the controller. Xinru Ma, Yonghao Xie, Jun Liu 0007, Yan Peng 0001, Shaorong Xie, Jun Luo 0006 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2024 | Shadow-Enlightened Image OutpaintingabstractConventional image outpainting methods usually treat unobserved areas as unknown and extend the scene only in terms of semantic consistency, thus overlooking the hidden information in shadows cast by unobserved areas, such as the invisible shapes and semantics. In this paper, we propose to extract and utilize the hidden information of un-observed areas from their shadows to enhance image out-painting. To this end, we propose an end-to-end deep approach that explicitly looks into the shadows within the image. Specifically, we extract shadows from the input image and identify instance-level shadow regions cast by the un-observed areas. Then, the instance-level shadow representations are concatenated to predict the scene layout of each unobserved instance and outpaint the unobserved areas. Finally, two discriminators are implemented to enhance alignment between the extended semantics and their shadows. In the experiments, we show that our proposed approach provides complementary cues for outpainting and achieves considerable improvement on all datasets by adopting our approach as a plug-in module. Hang Yu 0006, Shaorong Xie, Jiayan Qiu |
CVPR | 3 |
| 2024 | DGLF: A Dual Graph-based Learning Framework for Multi-modal Sarcasm DetectionabstractZhihong Zhu, Kefan Shen, Zhaorun Chen, Yunyan Zhang, Yuyan Chen, Xiaoqi Jiao, Zhongwei Wan, Shaorong Xie, Wei Liu, Xian Wu, Yefeng Zheng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zhihong Zhu 0001, Kefan Shen, Zhaorun Chen, Yunyan Zhang, Yuyan Chen, Xiaoqi Jiao, Zhongwei Wan, Shaorong Xie, Wei Liu 0027, Xian Wu 0001, Yefeng Zheng 0001 |
EMNLP | 8 |
| 2024 | Deep Evidential Active Learning with Uncertainty-Aware Determinantal Point Process
Yuxian Zhou, Xiaodong Yue 0002, Yufei Chen 0002, Shaorong Xie |
ICPR (1) | 4 |
| 2024 | Efficient Compensation of Action for Reinforcement Learning Policies in Sim2RealabstractSimulation to reality (sim-to-real) transfer is a promising alternative for training behavioral policies in reinforcement learning (RL). However, many significant differences between the simulator and the real-world environment cause policies to make inconsistent actions in simulation and reality. So policies trained in simulation often perform poorly in real-world. Researchers have explored various methods to bridge this gap, including building highly realistic simulators, implementing automated domain randomization, and employing domain adaptation method. However, highly realistic simulators and automated domain randomization methods rely heavily on extensive real-world data and complex manual processes. Domain adaptation methods require the collection and annotation of high-precision real-world data, leading to low learning efficiency and restricting these methods to specific domains. To address these challenges, this paper proposes transforming the discrepancies in the transfer process into a problem of action sequence similarity. By enhancing the similarity between policy action sequences, we aim to reinforce the consistency of policy actions made in both simulation and reality. For the challenging issue of annotating real-world data, we employ a Generative Adversarial Network (GAN) framework to construct a sim-to-real consistency loss function, thus avoiding reliance on precise real-world data sampling and calibration. To avoid a large amount of real-world data sampling, we introduce Bayesian optimization to accurately and efficiently search for the optimal parameters of the compensation module. Through extensive experiments in multiple sim-to-sim scenarios as well as sim-to-real scenarios, we demonstrate that our method significantly reduces the precision and quantity requirements for real-world data sampling while maintaining high transfer performance. Shaorong Xie, Xiangfeng Luo, Tao Wang 0072 |
ICTAI | 2 |
| 2024 | Structure-Aware Adaptive Hybrid Interaction Modeling for Image-Text Matching
Wei Liu 0027, Chao Wang 0095, Yan Peng 0001, Shaorong Xie |
MMM (1) | 5 |
| 2024 | Improving Inference via Rich Path Information for Dialogue Relation Extraction
Huizhe Su, Hang Yu 0006, Yanghao Zhou, Changsen Yuan, Shaorong Xie, Xiangfeng Luo |
NLPCC (5) | 6 |
| 2024 | Geometric relation-based feature aggregation for 3D small object detection
Hang Yu 0006, Xiangfeng Luo, Shaorong Xie |
Appl. Intell. | 4 |
| 2024 | Adaptive coupled-sliding-variable-based finite-time control of composite formation for multi-robot systems
Xinru Ma, Jun Liu 0007, Yueying Wang, Shaorong Xie, Jun Luo 0006 |
Sci. China Inf. Sci. | 5 |
| 2024 | Combining prompt learning with contextual semantics for inductive relation prediction
Shaorong Xie, Qifei Pan, Xinzhi Wang 0001, Xiangfeng Luo, Vijayan Sugumaran |
Expert Syst. Appl. | 1 |
| 2024 | Hierarchical Knowledge-Enhancement Framework for multi-hop knowledge graph reasoning
Shaorong Xie, Ruishen Liu, Xinzhi Wang 0001, Xiangfeng Luo, Vijayan Sugumaran, Hang Yu 0006 |
Neurocomputing | 1 |
| 2024 | A Dual Decision-Making Continuous Reinforcement Learning Method Based on Sim2RealabstractContinuous reinforcement learning carries potential security risks when applied in real-world scenarios, which could have significant societal implications. While its field of application is expanding, the majority of applications still remain confined to virtual environments. If only a single continuous learning method is applied to an unmanned system, it will still forget previously learned experiences, and retraining will be required when it encounters unknown environments. This reduces the learning efficiency of the unmanned system. To address these issues, some scholars have suggested prioritizing the experience playback pool and using transfer learning to apply previously learned strategies to new environments. However, these methods only alleviate the speed at which the unmanned system forgets its experiences and do not fundamentally solve the problem. Additionally, they cannot prevent dangerous actions and falling into local optima. Therefore, we propose a dual decision-making continuous learning method based on simulation to reality (Sim2Real). This method employs a knowledge body to eliminate the local optimal dilemma, and corrects bad strategies in a timely manner to ensure that the unmanned system makes the best decision every time. Our experimental results demonstrate that our method has a 30% higher success rate than other state-of-the-art methods, and the model transfer to real scenes is still highly effective. Xinzhi Wang 0001, Xiangfeng Luo, Shaorong Xie |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2024 | Knowledge-guided communication preference learning model for multi-agent cooperation
Hang Yu 0006, Zhenyu Zhang 0013, Yang Li 0151, Shaorong Xie, Xiangfeng Luo |
Inf. Sci. | 7 |
| 2024 | DP-DDCL: A discriminative prototype with dual decoupled contrast learning method for few-shot object detection
Yinsai Guo, Liyan Ma, Xiangfeng Luo, Shaorong Xie |
Knowl. Based Syst. | 4 |
| 2024 | EGLR: Two-staged Explanation Generation and Language Reasoning framework for commonsense question answering
Wei Liu 0027, Chao Wang 0095, Yan Peng 0001, Shaorong Xie |
Knowl. Based Syst. | 5 |
| 2024 | Enhanced Incremental Image Stitching for Low-Altitude UAV Imagery With Depth EstimationabstractThis letter proposes Depth-Aided Incremental Image Stitching (DAIIS), an algorithm tailored for low-altitude unmanned aerial vehicle (UAV) images in urban environments. DAIIS integrates single-view depth estimation and plane fitting to identify ground feature points, enabling precise computation of transformation matrices and significantly reducing stitching errors caused by large depth variations. Unlike traditional methods, DAIIS effectively addresses misalignments and artifacts, producing seamless and accurate stitched images that closely resemble the real scene. Evaluated on three distinct datasets, DAIIS demonstrated superior performance, achieving a 70% reduction in mean square error (mse) and notable improvements in peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) compared to other methods. The results validate DAIIS as a robust solution for high-precision image stitching, enhancing accuracy and visual quality in challenging low-altitude scenarios, which offers reliable and high-quality results for various urban applications. Yusheng Yang, Shaorong Xie, Yangmin Xie |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | MVPN: Multi-granularity visual prompt-guided fusion network for multimodal named entity recognition
Wei Liu 0027, Aiqun Ren, Chao Wang 0095, Yan Peng 0001, Shaorong Xie, Weimin Li 0001 |
Multim. Tools Appl. | 5 |
| 2024 | Selective arguments representation with dual relation-aware network for video situation recognition
Wei Liu 0027, Chao Wang 0095, Yan Peng 0001, Shaorong Xie |
Neural Comput. Appl. | 5 |
| 2024 | Saliency information and mosaic based data augmentation method for densely occluded object recognition
Ying Tong, Xiangfeng Luo, Liyan Ma, Shaorong Xie, Yinsai Guo |
Pattern Anal. Appl. | 4 |
| 2024 | Prototype learning for adversarial domain adaptation
Yuchun Fang, Chen Chen 0114, Wei Zhang 0021, Zhaoxiang Zhang 0001, Shaorong Xie |
Pattern Recognit. | 6 |
| 2024 | DSCA: A Dual Semantic Correlation Alignment Method for domain adaptation object detection
Yinsai Guo, Hang Yu 0006, Shaorong Xie, Liyan Ma, Xinzhi Cao, Xiangfeng Luo |
Pattern Recognit. | 3 |
| 2024 | Enhanced Implicit Sentiment Understanding With Prototype Learning and Demonstration for Aspect-Based Sentiment AnalysisabstractIn the field of social computing, the task of aspect-based sentiment analysis (ABSA) aims to classify the sentiment polarity of a given aspect in a sentence. The absence of explicit opinion words in the implicit aspect sentiment expressions poses a greater challenge for capturing their sentiment features in the reviews from social media. Many recent efforts use dependency trees or attention mechanisms to model the association between the aspect and other contextual words. However, dependency tree-based methods are inefficient in constructing valuable associations for sentiment classification due to the lack of explicit opinion words. In addition, the use of attention mechanisms to obtain global semantic information easily leads to an undesired focus on irrelevant words that may have sentiments but are not directly related to the specific aspect. In this article, we propose a novel prototype-based demonstration (PD) model for the ABSA task, which contains prototype learning and PD stages. In the prototype learning stage, we employ mask-aware attention to capture the global sentiment feature of aspect and learn sentiment prototypes through contrastive learning. This allows us to acquire comprehensive central semantics of the sentiment polarity that contains the implicit sentiment features. In the PD stage, to provide explicit guidance for the latent knowledge within the T5 model, we utilize prototypes similar to the aspect sentiment as the neural demonstration. Our model outperforms others with a 1.68%/0.28% accuracy gain on the Laptop/Restaurant datasets, especially in the ISE slice, showing improvements of 1.17%/0.26%. These results confirm the superiority of our PD-ABSA in capturing implicit sentiment and improving classification performance. This provides a solution for implicit sentiment classification in social computing. Huizhe Su, Xinzhi Wang 0001, Shaorong Xie, Xiangfeng Luo |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | DIE-CDK: A Discriminative Information Enhancement Method With Cross-Modal Domain Knowledge for Fine-Grained Ship DetectionabstractDue to the overarching similarities of ships, subtle information is imperative for fine-grained ship detection. However, this information is easily lost in adverse weather (e.g., fog, rain, snow, and cloud) or occlusion scenarios. Experts can quickly and accurately recognize fine-grained objects because they have the domain knowledge to help them find the most discriminative information (e.g., edge, structure, texture, and class semantics); thus, they do not need a lot of information to make an identification. Motivated by it, we propose a discriminative information enhancement method with cross-modal domain knowledge (DIE-CDK) for fine-grained ship detection. The core idea behind DIE-CDK is to enhance the discriminative information about fine-grained ships by fusing cross-modal domain knowledge. The introduced cross-modal domain knowledge comprises local and global knowledge: 1) local knowledge is the knowledge of visual shape (e.g., edge contour) which is extracted from the image domain; and 2) global knowledge is the knowledge of the class semantics which is obtained from the common sense domain. In addition, to further study fine-grained ship detection, we introduce a Fine-grained ship dataset (called FgShips). Experiments show that our proposed DIE-CDK method achieves impressive gains in detection performance and outperforms state-of-the-art methods on fine-grained ship and public datasets. Yinsai Guo, Hang Yu 0006, Liyan Ma, Xiangfeng Luo, Shaorong Xie |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | CCDet: Confidence-Consistent Learning for Dense Object DetectionabstractModern detectors commonly employ classification scores to reflect the localization quality of detection results. However, there exists an inconsistency between them, misguiding the selection of high-quality predictions and providing unreliable results for downstream applications. In this paper, we find that the root of this confidence inconsistency lies in the inaccurate IoU estimation and the spatial misalignment of the learned features between the classification and localization tasks. Therefore, a Confidence-Consistent Detector (CCDet) which includes the Distribution-based IoU Prediction (DIP) and Consistency-aware label assignment (CLA), is proposed. DIP provides more stable and accurate IoU estimation by learning the probability distribution over the IoU range and employing the expectation as the predicted IoU. CLA adopts both the prediction performance and consistency degree of samples as assignment metrics to select positives, which guides the classification and localization tasks to promote similar feature distribution. Comprehensive experiments demonstrate that CCDet can effectively mitigate the confidence inconsistency between classification and localization, and achieve stable improvement across different baselines. On the test-dev set of MS COCO, CCDet acquires a single-model single-scale AP of 50.1%, surpassing most of the existing object detectors. Chang Liu 0082, Xiaomao Li, Weiping Xiao, Shaorong Xie |
IEEE Trans. Image Process. | 4 |
| 2024 | Concept Drift Adaptation by Exploiting Drift TypeabstractConcept drift is a phenomenon where the distribution of data streams changes over time. When this happens, model predictions become less accurate. Hence, models built in the past need to be re-learned for the current data. Two design questions need to be addressed in designing a strategy to re-learn models: which type of concept drift has occurred, and how to utilize the drift type to improve re-learning performance. Existing drift detection methods are often good at determining when drift has occurred. However, few retrieve information about how the drift came to be present in the stream. Hence, determining the impact of the type of drift on adaptation is difficult. Filling this gap, we designed a framework based on a lazy strategy called Type-Driven Lazy Drift Adaptor (Type-LDA). Type-LDA first retrieves information about both how and when a drift has occurred, then it uses this information to re-learn the new model. To identify the type of drift, a drift type identifier is pre-trained on synthetic data of known drift types. Furthermore, a drift point locator locates the optimal point of drift via a sharing loss. Hence, Type-LDA can select the optimal point, according to the drift type, to re-learn the new model. Experiments validate Type-LDA on both synthetic data and real-world data, and the results show that accurately identifying drift type can improve adaptation accuracy. Hang Yu 0006, Zhenyu Zhang 0013, Xiangfeng Luo, Shaorong Xie |
ACM Trans. Knowl. Discov. Data | 5 |
| 2024 | Type-LDD: A Type-Driven Lite Concept Drift Detector for Data StreamsabstractConcept drift is a phenomenon that the distribution of data streams changes with time. When this happens, model predictions become less accurate. Hence, concept drift needs to be detected and adapted. Existing drift detection methods are good at determining when drift has occurred, but few retrieve information about how the drift came to be present in the stream, i.e., what type of drift has occurred. Hence, discussing the impact of the type of drift on adaptation is a difficult thing. To fill this gap, we propose a pre-trained framework for training a drift detector called a type-driven lite concept drift detector (Type-LDD) that retrieves information about both when and how a drift has occurred. In our proposed pre-trained framework, the Type-LDD including a drift-type identifier and a drift-point locator was based on a synthetic dataset containing a range of drift types. When repurposing the pre-trained model for detecting new data streams, a knowledge distillation module fine-tunes the proposed Type-LDD to speed up inference and keep detection accuracy. The proposed Type-LDD is validated on both synthetic data and real-world data, and demonstrated that accurately identifying the type of drift that has occurred can improve adaptation accuracy. Hang Yu 0006, Jie Lu 0001, Yiliao Song, Shaorong Xie, Guangquan Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Event-Based Distributed Secure Control of Unmanned Surface Vehicles With DoS AttacksabstractThis study investigates the distributed secure control problem of multiple unmanned surface vehicles (USVs) in the presence of wave-induced disturbances and unified abrupt and incipient rudder angle faults in physical layer, and aperiodic Denial-of-Service (DoS) attacks in cyber layer. Multi-USVs with rudder angle fault and DoS attack modeling are first established. Then, the decentralized unknown input observer (UIO)-based fault estimation and distributed secure control approach is developed in a co-designed framework for multi-USVs with cyber–physical threats. Advantages of the proposed secure scheme are: 1) actuator faults, DoS attacks, and event-triggering strategies with varying action instants, durations, and locations are synchronously addressed and 2) the criteria of exponential consensus are derived by virtue of attack frequency and average dwelling time technique without prior knowledge of unknown wave-induced perturbation bounds and elimination of Zeno behavior in an event-based mechanism. Comparative simulations outline the performance and advantage of the proposed distributed secure control algorithm. Chun Liu 0006, Bin Jiang 0001, Xiao Fan Wang 0001, Youmin Zhang 0001, Shaorong Xie |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2023 | Safe Multi-View Deep ClassificationabstractMulti-view deep classification expects to obtain better classification performance than using a single view. However, due to the uncertainty and inconsistency of data sources, adding data views does not necessarily lead to the performance improvements in multi-view classification. How to avoid worsening classification performance when adding views is crucial for multi-view deep learning but rarely studied. To tackle this limitation, in this paper, we reformulate the multi-view classification problem from the perspective of safe learning and thereby propose a Safe Multi-view Deep Classification (SMDC) method, which can guarantee that the classification performance does not deteriorate when fusing multiple views. In the SMDC method, we dynamically integrate multiple views and estimate the inherent uncertainties among multiple views with different root causes based on evidence theory. Through minimizing the uncertainties, SMDC promotes the evidences from data views for correct classification, and in the meantime excludes the incorrect evidences to produce the safe multi-view classification results. Furthermore, we theoretically prove that in the safe multi-view classification, adding data views will certainly not increase the empirical risk of classification. The experiments on various kinds of multi-view datasets validate that the proposed SMDC method can achieve precise and safe classification results. Wei Liu 0303, Yufei Chen 0002, Xiaodong Yue 0002, Changqing Zhang 0002, Shaorong Xie |
AAAI | 5 |
| 2023 | Correlation and Foreground Attention to Improve Object DetectionabstractObject Detection (OD) can be viewed as a multi-objective task to achieve object localization and class recognition. With the rapid development of the deep neural networks (DNNs), on the one hand, the performance of OD has been significantly improved by relying on the high-quality feature extraction and representation of DNNs. On the other hand, it can be challenging to accurately detect and recognize objects with non-salient or confusing features. In this paper, we propose an efficient and pluggable OD method by using attention mechanism to solve these issues from two aspects. Firstly, we exploit the semantic relationship between objects as a prior knowledge to reduce the incorrect recognition of objects with confusing features, where the relationship is encoded as an attention map by using a graph convolutional network, and then this attention map is used to reweight the feature intensities of objects belonging to different classes. Then, based on the feature map extracted from DNNs, we extract a sub-feature map containing foreground information and use this map to generate foreground attention map to improve the feature saliency of the objects. The qualitative and quantitative experimental results well verify the effectiveness of our method. Yudi Dong, Xiaodong Yue 0002, Zhikang Xu, Shaorong Xie |
ICIP | 4 |
| 2023 | Multi-Dimensional Pruned Sparse Convolution for Efficient 3D Object DetectionabstractIn recent years, significant progress has been made in 3D object detection. The focus of research has primarily been on improving the detection accuracy of models, however, neglecting their efficiency during actual deployment. Aiming at this issue, in this paper, we propose a multi-dimensional pruning method from the perspectives of data and model. Specifically, given the input data represented by the voxel grid, we first measure the voxel importance and propose an importance-based sampling module to sparsify voxels while preserving informative ones. The model pruning is wrapped in the framework of weighted voxel distillation, where the student model is obtained by pruning the channels of teacher model and only the informative voxels in the teacher model are involved and transferred to the students. In addition, the proposed method can be seamlessly integrated into current voxel-based 3D detectors without any additional costs. Experimental results on the KITTI and ONCE datasets show that our method can achieve a reduction of over 80% in GFLOPs while maintaining superior performance. Linye Li, Xiaodong Yue 0002, Zhikang Xu, Shaorong Xie |
ICIP | 4 |
| 2023 | HCTA: Hierarchical Cooperative Task Allocation in Multi-Agent Reinforcement LearningabstractDespite significant advances in multi-agent reinforcement learning (MARL) in recent years, learning about cooperative policies for complex tasks remains a challenge. The concept of task allocation provides a valuable tool for designing and comprehending intricate multi-agent systems, as it enables agents with similar roles to collaborate on the same subtasks. However, existing research efforts have not yet provided clear guidelines on how task allocation can better adapt to dynamic environments. This paper introduces a novel framework called Hierarchical Cooperative Task Allocation (HCTA) that is based on long-term behavior analysis. The proposed framework is designed to solve the dynamic task allocation problem in fully cooperative tasks within multi-agent systems. Firstly, we train a task selector based on action chains to enhance the efficiency of task allocation. Secondly, we introduce a hierarchical cooperative mechanism to realize the parallel training of upper-level task selection and lower-level task decision making, so as to accelerate the learning of cooperative strategy. Experimental results show that HTCA significantly improves the learning efficiency of joint policy through dynamic task allocation, and enhances the success rate in complex cooperative tasks in StarCraft II micromanagement benchmark. Shaorong Xie, Xiangfeng Luo, Yang Li 0151, Hang Yu 0006 |
ICTAI | 2 |
| 2023 | Hierarchical Semantic Contrast for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) with image-level annotations has achieved great processes through class activation map (CAM). Since vanilla CAMs are hardly served as guidance to bridge the gap between full and weak supervision, recent studies explore semantic representations to make CAM fit for WSSS and demonstrate encouraging results. However, they generally exploit single-level semantics, which may hamper the model to learn a comprehensive semantic structure. Motivated by the prior that each image has multiple levels of semantics, we propose hierarchical semantic contrast (HSC) to ameliorate the above problem. It conducts semantic contrast from coarse-grained to fine-grained perspective, including ROI level, class level, and pixel level, making the model learn a better object pattern understanding. To further improve CAM quality, building upon HSC, we explore consistency regularization of cross supervision and develop momentum prototype learning to utilize abundant semantics across different images. Extensive studies manifest that our plug-and-play learning paradigm, HSC, can significantly boost CAM quality on both non-saliency-guided and saliency-guided baselines, and establish new state-of-the-art WSSS performance on PASCAL VOC 2012 dataset. Code is available at https://github.com/Wu0409/HSC_WSSS. Yuanchen Wu, Xiaoqiang Li 0002, Songmin Dai, Jide Li, Tong Liu 0001, Shaorong Xie |
IJCAI | 6 |
| 2023 | Multi-perspective Feature Fusion for Event-Event Relation Extraction
Wei Liu 0027, Zhangdeng Pang, Shaorong Xie, Weimin Li 0001 |
NLPCC (2) | 3 |
| 2023 | Event Contrastive Representation Learning Enhanced with Image Situational Information
Wei Liu 0027, Shaorong Xie, Weimin Li 0001 |
NLPCC (2) | 3 |
| 2023 | Object-Level Contrast Learning for 3D Sparse Object Detection in Ocean SceneabstractLiDAR-based 3D object detection provides the necessary high-precision environmental sensing information for the safe navigation of smart ships.However, relying on viewpoint projections, voxelized point clouds, or using inefficient point sampling methods, current LiDAR 3D object detection methods treat all objects uniformly and quantitatively while ignoring the specificity of sparse objects in the scene, which leaves less useful information about sparse objects.In this paper, we propose an end-to-end two-stage architecture, Object-Level Contrast Learning 3D Object Detection network (OCL), for better construction of sparse object features and improving the ability of model to detect sparse objects.In the first stage, the Contrast Learning based Sparse Object Feature Enhancement training strategy is proposed to decrease the feature discrepancy between sparse and regular objects in object-level.In the second stage, we use the Point-level Feature Multiple Aggregation Strategy to aggregate finer point-level features of sparse objects.Extensive experiments show that OCL achieves excellent performance on both Ship dataset and KITTI dataset.Furthermore, our work proposes a promising new idea for applying contrast learning to 3D object detection. Yuheng He, Xiangfeng Luo, Liyan Ma, Shaorong Xie |
SEKE | 5 |
| 2023 | DDCL: A Dual Decision-making Continuous Reinforcement Learning Method Based on Sim2RealabstractContinuous reinforcement learning carries potential security risks when applied in real-world scenarios, which could have significant societal implications.While its field of application is expanding, the majority of applications still remain confined to virtual environments.if only a single continuous learning method is applied to an unmanned system, it will still forget previously learned experiences, and retraining will be required when it encounters unknown environments.This reduces the learning efficiency of the unmanned system.To address these issues, some scholars have suggested prioritizing the experience playback pool and using transfer learning to apply previously learned strategies to new environments.However, these methods only alleviate the speed at which the unmanned system forgets its experiences and do not fundamentally solve the problem.Additionally, they cannot prevent dangerous actions and falling into local optima.Therefore, we propose a dual decision-making continuous learning method based on Simulation to Reality (Sim2Real).This method employs a knowledge body to eliminate the local optimal dilemma, and corrects bad strategies in a timely manner to ensure that the unmanned system makes the best decision every time.Our experimental results demonstrate that our method has a 30% higher success rate than other state-of-the-art methods, and the model transfer to real scenes is still highly effective. Xinzhi Wang 0001, Xiangfeng Luo, Shaorong Xie |
SEKE | 4 |
| 2023 | Feature semantic space-based sim2real decision model
Xiangfeng Luo, Shaorong Xie |
Appl. Intell. | 3 |
| 2023 | Recurrent prediction model for partially observable MDPs
Shaorong Xie, Zhenyu Zhang 0013, Hang Yu 0006, Xiangfeng Luo |
Inf. Sci. | 1 |
| 2023 | An end-to-end neural framework using coarse-to-fine-grained attention for overlapping relational triple extractionabstractAbstract In recent years, the extraction of overlapping relations has received great attention in the field of natural language processing (NLP). However, most existing approaches treat relational triples in sentences as isolated, without considering the rich semantic correlations implied in the relational hierarchy. Extracting these overlapping relational triples is challenging, given the overlapping types are various and relatively complex. In addition, these approaches do not highlight the semantic information in the sentence from coarse-grained to fine-grained. In this paper, we propose an end-to-end neural framework based on a decomposition model that incorporates multi-granularity relational features for the extraction of overlapping triples. Our approach employs an attention mechanism that combines relational hierarchy information with multiple granularities and pretrained textual representations, where the relational hierarchies are constructed manually or obtained by unsupervised clustering. We found that the different hierarchy construction strategies have little effect on the final extraction results. Experimental results on two public datasets, NYT and WebNLG, show that our mode substantially outperforms the baseline system in extracting overlapping relational triples, especially for long-tailed relations. Huizhe Su, Hao Wang 0097, Xiangfeng Luo, Shaorong Xie |
Nat. Lang. Eng. | 4 |
| 2023 | Mitigate the classification ambiguity via localization-classification sequence in object detection
Chang Liu 0082, Shaorong Xie, Xiaomao Li, Jiantao Gao, Weiping Xiao, Baojie Fan, Yan Peng 0001 |
Pattern Recognit. | 2 |
| 2023 | An Adversarial Meta-Training Framework for Cross-Domain Few-Shot LearningabstractMeta-learning provides a promising way for deep learning models to efficiently learn in few-shot learning. With this capacity, many deep learning systems can be applied in many real applications. However, many existing meta-learning based few-shot learning systems suffer from vulnerable generalization when new tasks are from unseen domains (a.k.a, cross-domain few-shot learning). In this work, we consider this problem from the perspective of designing a model-agnostic meta-training framework to improve the generalization of existing meta-learning methods in cross-domain few-shot learning. In this way, compared with focusing on elaborately designing modules for a specific meta-learning model, our method is endowed with the ability to be compatible with different meta-learning models in various few-shot problems. To achieve this goal, a novel adversarial meta-training framework is proposed. The proposed framework utilizes max-min episodic iteration. In the episode of maximization, our framework focuses on how to dynamically generate appropriate pseudo tasks which benefit learning cross-domain knowledge. In the episode of minimization, our method aims to solve how to help meta-learning model learn cross-task and robust meta-knowledge. To comprehensively evaluate our framework, experiments are conducted on two few-shot learning settings, three meta-learning models, and eight datasets. These results demonstrate that our method is applicable to various meta-learning models in different few-shot learning problems. The superiority of our method is verified compared with existing state-of-the-art methods. Pinzhuo Tian, Shaorong Xie |
IEEE Trans. Multim. | 2 |
| 2022 | Occlusion-robust Face Alignment using A Viewpoint-invariant Hierarchical Network ArchitectureabstractThe occlusion problem heavily degrades the localization performance of face alignment. Most current solutions for this problem focus on annotating new occlusion data, introducing boundary estimation, and stacking deeper models to improve the robustness of neural networks. However, the performance degradation of models remains under extreme occlusion (i.e. average occlusion of over 50%) because of missing a large amount of facial context information. We argue that exploring neural networks to model the facial hierarchies is a more promising method for dealing with extreme occlusion. Surprisingly, in recent studies, little effort has been devoted to representing the facial hierarchies using neural networks. This paper proposes a new network architecture called GlomFace to model the facial hierarchies against various occlusions, which draws inspiration from the viewpoint-invariant hierarchy of facial structure. Specifically, GlomFace is functionally divided into two modules: the part-whole hierarchical module and the whole-part hierarchical module. The former captures the part-whole hierarchical dependencies of facial parts to suppress multi-scale occlusion information, whereas the latter injects structural reasoning into neural networks by building the whole-part hierarchical relations among facial parts. As a result, GlomFace has a clear topological interpretation due to its correspondence to the facial hierarchies. Extensive experimental results indicate that the proposed GlomFace performs comparably to existing state-of-the-art methods, especially in cases of extreme occlusion. Models are available at https://github.com/zhuccly/GlomFace-Face-Alignment. Congcong Zhu, Xintong Wan, Shaorong Xie, Xiaoqiang Li 0002, Yinzheng Gu |
CVPR | 3 |
| 2022 | Prior Semantic Harmonization Network for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation(FSS) is intended to segment a foreground object from a query image with a novel object using only a few annotated support images. Although attracting the attention of many researchers, this challenging problem remains to be not well solved due to two critical issues: (1)The information mismatching between support and query features leads to model distraction. (2)The key feature of query images is not activated well. In this paper, we introduce the Prior Semantic Harmonization Network(PSHNet) to tackle these limitations. PSHNet is composed of three effective modules. The Semantic Harmonization Module(SHM) corrects the information matching between support and query images, while the Feature Activation Module(FAM) activates the key feature of query images. Furthermore, we introduce a Hierarchical Aggregation Module(HAM) to refine each output of the multi-scale module. Experiments show that our model achieves an excellent performance on both PASCAL-5iand COCO-20idatasets. Liyan Ma, Yan Peng 0001, Shaorong Xie |
ICIP | 5 |
| 2022 | Open-World Object Detection via Discriminative Class Prototype LearningabstractOpen-world object detection (OWOD) is a challenging problem that combines object detection with incremental learning and open-set learning. Compared to standard object detection, the OWOD setting is task to: 1) detect objects seen during training while identifying unseen classes, and 2) incrementally learn the knowledge of the identified unknown objects when the corresponding annotations is available. We propose a novel and efficient OWOD solution from a prototype perspective, which we call OCPL: Open-world object detection via discriminative Class Prototype Learning, which consists of a Proposal Embedding Aggregator (PEA), an Embedding Space Compressor (ESC) and a Cosine Similarity-based Classifier (CSC). All our proposed modules aim to learn the discriminative embeddings of known classes in the feature space to minimize the overlapping distributions of known and unknown classes, which is beneficial to differentiate known and unknown classes. Extensive experiments performed on PASCAL VOC and MS-COCO benchmark demonstrate the effectiveness of our proposed method. Jinan Yu, Liyan Ma, Yan Peng 0001, Shaorong Xie |
ICIP | 5 |
| 2022 | Learning from Suboptimal Demonstration via Trajectory-Ranked Adversarial ImitationabstractRobots trained by Imitation Learning(IL) are used in many tasks(e.g., autonomous vehicle manipulation). Generative Adversarial Imitation Learning (GAIL) assumes that the demonstration set used for training is of high quality. However, such demonstrations are difficult and expensive to obtain. GAIL-related methods fail to learn effective strategies if non-high quality demonstrations are used because the performance of agents trained by this method is limited by the demonstrator's operations. Our idea is to enable the agent to learn strategy with better performance than the demonstrator from a suboptimal demonstration set, which contains non-high quality demonstrations that are easier to obtain. Inspired by this, we propose the Trajectory-Ranked Adversarial Imitation Learning (TRAIL) method. First, for demonstration set processing, we introduce a ranking process and define the concept of Performance Relative Advantage of suboptimal demonstrations to specify the ranking order. Second, for model training, we reconstruct the objective function of GAIL and use an experience replay buffer, enabling the agent to learn implicit features and ranking information from the ranked suboptimal demonstration set and possess the ability to outperform the demonstrator. Experiments show that in Mujoco's tasks, our method can learn from a suboptimal demonstration set and can achieve better performance than baseline methods. Shaorong Xie, Hang Yu 0006, Xiangfeng Luo, Zhenyu Zhang 0013 |
ICTAI | 2 |
| 2022 | Offline Reinforcement Learning via Policy Regularization and Ensemble Q-FunctionsabstractOffline reinforcement learning aims to learn effective policies from a fixed set of data collected in advance and without further interaction with the environment during learning. This setting will promote the applications of reinforcement learning in the real world, in which interaction is costly or dangerous. However, existing off-policy algorithms can fail in offline settings due to the distributional shift between the learned policy and the policy that collected the dataset. To solve the problem, we develop a lightweight and effective algorithm, policy regularization with behavior model (PRBM). Firstly, PRBM trains a behavior model as the regularization term during policy optimization to avoid choosing out-of-distribution (OOD) actions. Secondly, to avoid the overestimation of OOD actions, PRBM trains multiple Q- functions and uses their min-max mixture to compute Q-values. Our experiments on different datasets from various continuous control tasks demonstrate that PRBM outperforms most baselines (especially on medium quality datasets) and requires only half of their training time. Tao Wang 0072, Shaorong Xie, Mingke Gao, Zhenyu Zhang 0013, Hang Yu 0006 |
ICTAI | 2 |
| 2022 | USVs-Sim: A general simulation platform for unmanned surface vessels autonomous learningabstractAbstract Unmanned surface vessels (USVs) have been fully used in the civilian and military fields in recent years, which dramatically expands protective capability and detection range. However, the marine environment's complexity and variability make that verification of various advanced USVs control algorithms face high costs and high risks. In this article, we present USVs‐Sim, a novel high‐fidelity general simulation platform for USVs autonomous navigation data generation and control strategy testing. USVs‐Sim is a collection of high‐level extensible modules that allows the rapid development and testing of USVs configurations and facilitates the construction of complex ocean scenarios. USVs‐Sim supports the steering or thrusting limits of USVs, as well as unique dynamics profiles. The platform can specify specific USVs sensor systems and change the time of day and weather conditions to generate robust data. USVs‐Sim facilitates training of deep‐learning algorithms by enabling data export from USVs sensors, including vision data, lidar, relative positions of ocean targets. Therefore, USVs‐Sim allows for the rapid prototyping, development, and testing of USVs autonomous control algorithms in a complex marine environment. In this article, we detail the general simulation platform and testing several representative USVs intelligent control algorithms on the platform. Wei Wang 0296, Yang Li 0151, Zhenyu Zhang 0013, Xiangfeng Luo, Shaorong Xie |
Concurr. Comput. Pract. Exp. | 6 |
| 2022 | Geometric relation based point clouds classification and segmentationabstractAbstract The inherent disorder and irregularity of 3D point clouds pose great challenges to classification and segmentation tasks. To tackle these problems, we propose a geometric relation based point clouds classification and segmentation network. Specifically, we design two novel modules named geometric relation based convolution (GRC) and relational attention interpolation (RAI) to infer the local relations of point clouds. In GRC module, the convolutional weights and local features are both reasoned from predefined local geometric relations between the central point of each local point clouds and its neighboring points. The global shape awareness is obtained by stacking convolutional layers of the GRC. In RAI, a relational attention interpolation approach is proposed for the segmentation task. The attentional weights of different neighboring points are inferred from local relations of geometry and features, which is capable of guiding RAI to pay more attention to the relevant points and be sensitive to the boundaries of segments. Experimental results show that the proposed method makes full use of the geometric relations between local points, and presents good performance on both classification and segmentation tasks. Suqin Sheng, Xiangfeng Luo, Shaorong Xie |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | MIFAS: Multi-source heterogeneous information fusion with adaptive importance sampling for link predictionabstractAbstract Link prediction plays an important role in constructing knowledge graph. Recently, graph representation learning models yield state‐of‐the‐art results. However, existing models concentrate merely on triples or graph structures and mostly ignore textual descriptions, resulting in incomplete or partial information. In this paper, we propose a novel graph representation learning model to address this challenge, namely multi‐source heterogeneous information fusion with adaptive importance sampling. Our model leverages multiple sources, such triple, graph structure and textual description, and generate rich‐attribute embeddings for entities, encapsulating relations simultaneously. We also propose an adaptive importance sampling algorithm to boost aggregation of useful features from local neighbours. Additionally, we also boost node aggregation of useful features from local neighbours by adaptive importance sampling algorithm in our model. Experimental results on two benchmark datasets show that our proposed model significantly outperforms state‐of‐the‐art methods. Tingting Jiang 0008, Hao Wang 0097, Xiangfeng Luo, Shaorong Xie, Jingchao Wang 0001 |
Expert Syst. J. Knowl. Eng. | 4 |
| 2022 | ET-HF: A novel information sharing model to improve multi-agent cooperation
Shaorong Xie, Hang Yu 0006, Yang Li 0151, Zhenyu Zhang 0013, Xiangfeng Luo |
Knowl. Based Syst. | 1 |
| 2022 | On Stability and Stabilization of T-S Fuzzy Systems With Time-Varying Delays via Quadratic Fuzzy Lyapunov MatrixabstractThis article proposes improved stability and stabilization criteria for Takagi–Sugeno (T–S) fuzzy systems with time-varying delays. First, a novel augmented fuzzy Lyapunov–Krasovskii functional (LKF) including the quadratic fuzzy Lyapunov matrix is constructed, which can provide much information of T–S fuzzy systems and help to achieve the lager allowable delay upper bounds. Then, improved delay-dependent stability and stabilization criteria are derived for the studied systems. Compared with the traditional methods, since the third-order Bessel–Legendre inequality and the extended reciprocally convex matrix inequality are well employed in the derivative of the constructed LKF to give tighter bounds of the single integral terms, the conservatism of derived criteria is further reduced. In addition, the quadratic fuzzy Lyapunov matrix introduced in LKF, which contains the quadratic membership functions, is also an important reason for obtaining less conservative results. Finally, numerical examples demonstrate that the proposed method is less conservative than some existing ones and the studied system can be well controlled by the designed controller. Chen Peng 0001, Xiangpeng Xie 0001, Shaorong Xie |
IEEE Trans. Fuzzy Syst. | 4 |
| 2022 | Distributed Dimensionality Reduction Fusion Estimation for Stochastic Uncertain Systems With Fading Measurements Subject to Mixed AttacksabstractIn this article, the distributed fusion estimation issue with the dimensionality reduction strategy under DoS attacks and deception attacks is investigated for a class of stochastic uncertain systems with fading measurements. The stochastic uncertainties existed in the system and measurement equations are represented by state-dependent noises. The fading measurements are depicted by stochastic variables with known statistics. Then, a novel attack and compensation model is proposed to display the randomly occurring behaviors of the DoS attacks and the deception attacks within a unified framework. Furthermore, a distributed multisensor fusion estimation (DMSFE) algorithm is presented. An explicit form of dimensionality reduction is designed against attacks. Stability conditions are derived such that the mean square errors (MSEs) of the proposed DMSFE are bounded. A sequential covariance intersection fusion estimator (SCIFE) is designed to prevent the cross fusion covariance matrices calculating, which owns lower accuracy by smaller computation cost than DMSFE. An illustrative example is provided to show the effectiveness and merits of the proposed algorithm. Sha Fan, Huaicheng Yan 0001, Hao Zhang 0008, Yueying Wang, Yan Peng 0001, Shaorong Xie |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2021 | Learning Heterogeneous Strategies via Graph-based Multi-agent Reinforcement LearningabstractIn a mixed cooperative-competitive environment, each agent needs to learn heterogeneous strategies. The complex game relationship between heterogeneous agents causes difficulties for strategy learning. In this paper, we propose a graph-based multi-agent reinforcement learning method to simplify the learning process, named hierarchical heterogeneous graph multi-agent actor-critic (H2G-MAAC). The method first uses a hierarchical heterogeneous graph to model the hierarchical relationships among multiple heterogeneous agents. Then, it conducts representation learning for multiple agents with hierarchical graph attention networks. Finally, it learns heterogeneous strategies with multi-agent actor-critic. We conduct experiments in Predator-Prey games. The results indicate that the proposed method can simplify the learning process and outperforms existing state-of-the-art methods. Yang Li 0151, Xiangfeng Luo, Shaorong Xie |
ICTAI | 3 |
| 2021 | Learning Discriminative Representations for Fine-Grained Diabetic Retinopathy GradingabstractDiabetic retinopathy is one of the leading causes of blindness. However, no specific symptoms of early DR lead to a delayed diagnosis, which results in disease progression in patients. To determine the disease severity levels, ophthalmologists need to focus on the discriminative parts of the retinal images. In recent years, deep learning has achieved great success in medical image analysis. However, most works directly employ algorithms based on convolutional neural networks (CNNs), which ignore the fact that the difference among classes is subtle and gradual. Hence, we consider automatic image grading of DR as a fine-grained classification task, and construct a bilinear model to identify the pathologically discriminative areas. In order to leverage the ordinal information among classes, we put the soft labels with ordinal information among classes into the loss function rather than the most commonly used one-hot labels for the diabetic retinopathy classification. In addition, other than only using a categorical loss to train our network, we also introduce the metric loss to learn a more discriminative feature space which is beneficial to locate the finer discriminative lesion parts. Experimental results demonstrate the superior performance of the proposed method on publicly available IDRiD, DeepDRiD and FGADR datasets. Liyan Ma, Zhijie Wen, Shaorong Xie, Yupeng Xu |
IJCNN | 4 |
| 2021 | Unmanned surface vessel obstacle avoidance with prior knowledge-based reward shapingabstractAbstract Autonomous obstacle avoidance control of unmanned surface vessels (USVs) in complex marine environments is always fundamental for its scientific search and detection. Traditional methods usually model USV motion and environments in a mathematical way that needs perceptual information. Unfortunately, it is difficult to provide sufficient perceptual information due to complex marine environments, resulting in inaccurate modeling. Reinforcement learning has recently enjoyed increasing popularity in the problem of obstacle avoidance since it can settle problems by partially observable environment information. However, the autonomous USV obstacle avoidance using reinforcement learning is still facing the challenge of designing appropriate reward functions under complex marine environments. To address these issues, we propose a prior knowledge‐based USV reinforcement learning obstacle avoidance algorithm. In this algorithm, an actor‐critic network is used as the main architecture of the algorithm, and prior knowledge‐based reward shaping used to design relevant reward function for USV obstacle avoidance. A standard USV based on a visual sensor is designed, and the state input of the algorithm is through USV's front vision sensor. We conducted simulation experiments and results prove that our algorithm can effectively converge, and USV achieves high velocity and low collision rate in the complex marine environment. Wei Wang 0296, Xiangfeng Luo, Yang Li 0151, Shaorong Xie |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | Learning adversarial policy in multiple scenes environment via multi-agent reinforcement learningabstractLearning adversarial policy aims to learn behavioural strategies for agents with different goals, is one of the most significant tasks in multi-agent systems. Multi-agent reinforcement learning (MARL), as a state-of-the-art learning-based model, employs centralised or decentralised control methods to learn behavioural strategies by interacting with environments. It suffers from instability and slowness in the training process. Considering that parallel simulation or computation is an effective way to improve training performance, we propose a novel MARL method called Multiple scenes multi-agent proximal Policy Optimisation (MPO) in this paper. In MPO, we first simulate multiple parallel scenes in the training environment. Multiple policies control different agents in the same scene, and each policy also controls several identical agents from multiple scenes. Then, we expand proximal policy optimisation (PPO) with an improved actor-critic network, ensuring the stability of training in multi-agent tasks. The actor network only uses local information for decision making, and the critic network uses global information for training. Finally, effective training trajectories are computed with two criteria from multiple parallel scenes rather than single to accelerate the learning process. We evaluate our approach in two simulated 3D environments, one of which is Unity's official open-source soccer game, and the other is unmanned surface vehicles (USVs) built by Unity. Experiments demonstrate that MPO converges more stable and faster than benchmark methods in model training, and demonstrates excellent adversarial policy compared with benchmark models. Yang Li 0151, Xinzhi Wang 0001, Wei Wang 0296, Zhenyu Zhang 0013, Jianshu Wang, Xiangfeng Luo, Shaorong Xie |
Connect. Sci. | 7 |
| 2021 | Real-Time Monocular Obstacle Detection Based on Horizon Line and Saliency Estimation for Unmanned Surface Vehicles
Jun Liu 0007, Shaorong Xie, Jun Luo 0006 |
Mob. Networks Appl. | 4 |
| 2021 | Fuzzy Control and Filtering for Nonlinear Singularly Perturbed Markov Jump Systemsabstractcontrol and filtering problems for Markov jump singularly perturbed systems approximated by Takagi-Sugeno fuzzy models. The underlying transition probabilities (TPs) are assumed to vary randomly in a finite set, which is characterized by a higher level TP matrix. The mode- and variation-dependent fuzzy static output-feedback controller (SOFC) and filter are designed, respectively, to fulfill the control and filtering purposes. To facilitate the fuzzy SOFC synthesis, the closed-loop system is transformed into a fuzzy piecewise-homogeneous Markov jump singularly perturbed descriptor system (MJSPDS) by descriptor representation. A rigorous proof of mean-square exponential admissibility for the resulting fuzzy MJSPDS is presented. The criterion ensuring the mean-square exponential stability of the fuzzy filtering error system is further formed based on similar procedures. By setting the specific forms of the related matrix variables, the solutions for the predesigned fuzzy SOFC and filter are furnished, respectively. Finally, feasibility and validities of the developed fuzzy control and filtering results are verified by two practical examples. Yueying Wang, Choon Ki Ahn, Huaicheng Yan 0001, Shaorong Xie |
IEEE Trans. Cybern. | 4 |
| 2021 | Sliding-Mode Control of Fuzzy Singularly Perturbed Descriptor SystemsabstractDue to the complicated model characteristics, only a few results focusing on stability analysis have appeared on singularly perturbed descriptor systems (SPDSs). This article instead proposes an integral sliding-mode control strategy for a kind of Takagi-Sugeno fuzzy approximation-based nonlinear SPDSs under time-varying nonlinear perturbation. An appropriate fuzzy integral switching manifold that fully accommodates the system features is designed to completely reject the matched perturbation without amplifying the unmatched one. To facilitate the synthesis of the high-level controller (HLC), the sliding-mode dynamics (SMD) is transformed into an augmented form. Thanks to the adoptions of a novel singular perturbation Lyapunov function, Finsler's lemma, as well as the fixed-point principle, the existence and uniqueness of the solution and the exponential admissibility for the augmented SMD are analyzed. A solution for the designed HLC is further provided. To guarantee the sliding motion, a fuzzy integral sliding-mode controller (FISMC) is synthesized by analyzing the sliding motion reachability. An adaptive FISMC is also given to deal with the unknown upper bounds of the matched perturbation. Finally, the applicability of the developed FISMC strategy is testified by a practical example. Yueying Wang, Xiangpeng Xie 0001, Mohammed Chadli, Shaorong Xie, Yan Peng 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2021 | Adaptive Sliding Mode Fault-Tolerant Fuzzy Tracking Control With Application to Unmanned Marine VehiclesabstractThis article presents a fault-tolerant tracking control strategy for Takagi–Sugeno fuzzy model-based nonlinear systems which combines integral sliding mode control with adaptive control technique. Two common actuator faults: 1) loss of effectiveness and 2) increased bias input, are considered simultaneously. The fuzzy tracking control system is first established by incorporating the integral term of the output tracking error. Then, an appropriate fuzzy integral switching surface is designed such that the corresponding sliding motion only suffers from the unamplified unmatched disturbance. The solution of the nominal tracking controller can be transformed into a to convex optimization problem. In particular, an adaptive fuzzy sliding mode tracking controller is synthesized to ensure the accessibility of the sliding motion despite the effect of actuator faults and unknown disturbances. Finally, the proposed tracking strategy is verified by applying it to the dynamic positioning control of unmanned marine vehicles. Yueying Wang, Bin Jiang 0001, Zhengguang Wu, Shaorong Xie, Yan Peng 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2020 | Cooperative Multi-Agent Reinforcement Learning with Hierarchical Relation Graph under Partial ObservabilityabstractCooperation among agents with partial observation is an important task in multi-agent reinforcement learning (MARL), aiming to maximize a common reward. Most existing cooperative MARL approaches focus on building different model frameworks, such as centralized, decentralized, and centralized training with decentralized execution. These methods employ partial observation of agents as input directly, but rarely consider the local relationship between agents. The local relationship can help agents integrate observation information among different agents in a local range, and then adopt a more effective cooperation policy. In this paper, we propose a MARL method based on spatial relationship called hierarchical relation graph soft actor-critic (HRG-SAC). The method first uses a hierarchical relation graph generation module to represent the spatial relationship between agents in local space. Second, it integrates feature information of the relation graph through the graph convolution network (GCN). Finally, the soft actor-critic (SAC) is used to optimize agents' actions in training for compliance control. We conduct experiments on the Food Collector task and compare HRG-SAC with three baseline methods. The results demonstrate that the hierarchical relation graph can significantly improve MARL performance in the cooperative task. Yang Li 0151, Xinzhi Wang 0001, Jianshu Wang, Wei Wang 0296, Xiangfeng Luo, Shaorong Xie |
ICTAI | 6 |
| 2020 | Realistic Style-Transfer Generative Adversarial Network With a Weight-Sharing StrategyabstractStyle transfer aims to generate images by combining the style of one image and the content of another. Though valuable efforts have been made in generating high-quality style transferred images, the resulting images are far from the distribution of real images. This greatly limits the application of style transfer such as improving the diversity of training set in computer vision task. We find that the reason of style transfer failing to generate realistic images is lack of reference targets and neglect of preserving data distribution. To solve the problem, we propose a Style-transfer Generative Adversarial Network with a weight-sharing strategy to make the stylized images be resemblance to the real images. The experimental results demonstrate that the proposed method can generate images with satisfying style transfers and high visual quality. Moreover, we apply our stylized images to augment the training set of object detection task, and improve the average precision faithfully. We believe that our method can enhance the performance of style transfer on computer vision tasks. Shixiong Zhu, Xiangfeng Luo, Liyan Ma, Shaorong Xie |
ICTAI | 4 |
| 2020 | Design and experiment of bio-inspired GER fluid damper
Huayan Pu, Yining Huang, Yi Sun 0002, Min Wang 0023, Shujin Yuan, Zhen Kong, Peipei Yang, Liufeng Chu, Yan Peng 0001, Shaorong Xie, Jun Luo 0006 |
Sci. China Inf. Sci. | 11 |
| 2020 | Diverse receptive field network with context aggregation for fast object detection
Shaorong Xie, Chang Liu 0082, Jiantao Gao, Xiaomao Li, Jun Luo 0006, Baojie Fan, Jiahong Chen, Huayan Pu, Yan Peng 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2020 | The autonomous navigation and obstacle avoidance for USVs with ANOA deep reinforcement learning method
Xing Wu 0001, Haolei Chen, Changgu Chen, Mingyu Zhong, Shaorong Xie, Yike Guo, Hamido Fujita |
Knowl. Based Syst. | 5 |
| 2020 | Data driven hybrid edge computing-based hierarchical task guidance for efficient maritime escorting with multiple unmanned surface vehicles
Jiajia Xie, Jun Luo 0006, Yan Peng 0001, Shaorong Xie, Huayan Pu, Xiaomao Li, Zhou Su 0001, Yuan Liu 0025 |
Peer-to-Peer Netw. Appl. | 4 |
| 2020 | Automated Parallel Electrical Characterization of Cells Using Optically-Induced DielectrophoresisabstractThis article reports an automated optically-induced dielectrophoresis (ODEP) system for characterizing the specific membrane capacitance (SMC) of individual cells. The simulation of cell motion is conducted to analyze the electrokinetic forces acting on the cell. A self-developed visual tracking algorithm for multicells is used to realize an automated process for determining the frequency-sweeping range, crossover frequencies, and cell radii. The SMC values of malignant bladder cancer cells (T24 and RT4) and normal urothelial cells (SV-HUC-1) were quantified using the automated system, demonstrating that the system has a measurement speed of ~1 cell/s, an accuracy of 1 kHz for the crossover frequency determination, and an accuracy of 0.2 μm for the cell radius measurement. Na Liu 0004, Yanbin Lin, Yan Peng 0001, Liming Xin, Tao Yue 0001, Changhai Ru, Shaorong Xie, Huayan Pu, Haige Chen, Wen J. Li, Yu Sun 0001 |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2019 | Proximal Policy Optimization with Mixed Distributed TrainingabstractInstability and slowness are two main problems in deep reinforcement learning. Even if proximal policy optimization (PPO) is the state of the art, it still suffers from these two problems. We introduce an improved algorithm based on proximal policy optimization, mixed distributed proximal policy optimization (MDPPO), and show that it can accelerate and stabilize the training process. In our algorithm, multiple different policies train simultaneously and each of them controls several identical agents that interact with environments. Actions are sampled by each policy separately as usual, but the trajectories for the training process are collected from all agents, instead of only one policy. We find that if we choose some auxiliary trajectories elaborately to train policies, the algorithm will be more stable and quicker to converge especially in the environments with sparse rewards. Zhenyu Zhang 0013, Xiangfeng Luo, Tong Liu 0001, Shaorong Xie, Jianshu Wang, Wei Wang 0296, Yang Li 0151, Yan Peng 0001 |
ICTAI | 4 |
| 2019 | The Collaborative Strategy of Multiple USVs with Deep Reinforcement Learning MethodabstractThe unmanned surface vehicle (USV) has been widely used to accomplish tasks that cannot be completed by ships with human drivers on certain sea areas. It is not only necessary but essential to obtain a robust strategy in order to ensure multiple USVs accomplish collaborative tasks successfully and efficiently. To meet the challenge, a deep reinforcement learning method is proposed, which is combined with an improved A star algorithm. A statistically promising collaborative strategy is achieved by the proposed method under the guidance from the unmanned aerial vehicles (UAVs). After the collaborative strategy is generated, the improved A star algorithm is used to navigate the USVs. To verify the proposed algorithm, several tasks are tested on a simulation platform. Experimental results demonstrate that the proposed method outperforms state-of-the-art reinforcement learning methods such as DQN and DeepSarsa. © 2019 The authors and IOS Press. All rights reserved. Xing Wu 0001, Mingyu Zhong, Guofei Feng, Shaorong Xie, Yike Guo |
SoMeT | 4 |
| 2018 | The Multiple Unmanned Surface Vehicles Cooperative Defense Based on PM-PSO and GA-PSO in the Sophisticated Sea EnvironmentabstractThe unmanned surface vehicles (USVs) have become a major trend in the construction of naval equipment and its flexibility and intelligence making it widely used in real-scenes. For cooperative defense with multiple USVs to intercept intruders, it is proposed that planning the path with obstacle avoidance and protecting the target by task allocation actions. The particle swarm optimization based on probe mechanism (PM-PSO) is proposed for pathing planning with obstacle avoidance. With the consideration of the constraints of different defense schemes such as the path cost, the interception loss, the defense income and so on, it is proposed that the dispersed particle swarm optimization based on genetic algorithm (GA-PSO) for the interception task allocation. Furthermore, the fitness function is proposed to evaluate the feasibility of the interception path and the quality of the allocation scheme. Extensive simulation experiments are conducted and demonstrated the effectiveness, rationality and superiority of the proposed methods. Yuan Liu 0025, Xing Wu 0001, Yike Guo, Shaorong Xie, Huayan Pu, Yan Peng 0001 |
SoMeT | 4 |
| 2018 | Topic detection model in a single-domain corpus inspired by the human memory cognitive processabstractSummary A corpus (eg, patents or news texts) is an important knowledge resource that contains various topics, such as specific technologies or social events. Topic detection models of corpus, eg, Latent Dirichlet Allocation and KeyGraph, provide an important basis for exploring the status quo and trends in science, technology, or social events. However, these models suffer from low retrieval performance as they only consider text own explicit semantics in a single‐domain corpus. In addition, many incremental models, such as online‐LDA, are based on time slices. In this paper, a new topic detection model is proposed to improve the topic detection performance of a single‐domain corpus, which is inspired by a human memory cognitive process (THC). First, to improve the accuracy, distributions over words and inter‐word relations across a corpus are utilized as background knowledge, which is a type of implicit semantics, and we can find a more semantic‐sensitive part of texts. Second, to realize online topic detection without time slices, we introduce a probability gain‐based dynamic probabilistic model to detect latent topics by learning a model based on the dynamic human memory cognitive process. These two steps constitute the framework of our model. The experimental results for four public datasets (Reuters‐R8, Reuters‐R52, WebKB, and Cade12) reveal that our model is approximately ten percent higher than other baselines (eg, KeyGraph and LDA) on the Adjusted Rand Index (ARI). Taotao Zhao, Xiangfeng Luo, Subin Huang, Shaorong Xie |
Concurr. Comput. Pract. Exp. | 5 |
| 2018 | Automated Non-Invasive Measurement of Single Sperm's Motility and MorphologyabstractMeasuring cell motility and morphology is important for revealing their functional characteristics. This paper presents automation techniques that enable automated, non-invasive measurement of motility and morphology parameters of single sperm. Compared to the status quo of qualitative estimation of single sperm's motility and morphology manually, the automation techniques provide quantitative data for embryologists to select a single sperm for intracytoplasmic sperm injection. An adapted joint probabilistic data association filter was used for multi-sperm tracking and tackled challenges of identifying sperms that intersect or have small spatial distances. Since the standard differential interference contrast (DIC) imaging method has side illumination effect which causes inherent inhomogeneous image intensity and poses difficulties for accurate sperm morphology measurement, we integrated total variation norm into the quadratic cost function method, which together effectively removed inhomogeneous image intensity and retained sperm's subcellular structures after DIC image reconstruction. In order to relocate the same sperm of interest identified under low magnification after switching to high magnification, coordinate transformation was conducted to handle the changes in the field of view caused by magnification switch. The sperm's position after magnification switch was accurately predicted by accounting for the sperm's swimming motion during magnification switch. Experimental results demonstrated an accuracy of 95.6% in sperm motility measurement and an error <10% in morphology measurement. Changsheng Dai, Zhuoran Zhang 0001, James Huang 0002, Xian Wang 0001, Changhai Ru, Huayan Pu, Shaorong Xie, Sergey Moskovtsev, Clifford Librach, Keith Jarvi, Yu Sun 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2017 | Irradiation test of the control system of a tracked robot for nuclear disaster responseabstractThe importance of mobile robots for the nuclear disaster response has been realized after Fukushima Dai-ichi nuclear power plant accident. In this paper, we propose a tracked robot for rescue and search in nuclear environment. A gamma-ray irradiation test of the robot's control system is conducted, in order to evaluate the performance of the robot in nuclear environment. The parallel test method is described and the test result is reported. Huayan Pu, Jun Luo 0006, Yang Yang 0044, Yi Sun 0002, Shaorong Xie |
IECON | 6 |
| 2017 | Modeling of lug-soil interaction forces acting on a single lug during rotational motion in sandy soilabstractTo improve the mobility of locomotive devices on loose, sandy terrain, protrusions or convex patterns called lugs (i.e., grousers) are attached to the surface of a locomotive modulus. Following our previous study, in which the effects of angular speed, lug sinkage length, and soil cumulative deformation on lug-soil interaction forces during the fixed-axis rotational motion were experimentally confirmed, this study proposed an approximation equation to formulize the relationship among the normal force, lug sinkage length, and lug rotational angle. Moreover, the measured tangential force is compared with values calculated from a conventional tangential force model for discussing its accuracy of predicting the tangential force. Conclusions from this study present the fundamental principles for understanding the lug-soil interaction mechanics for a lug that is performing arbitrary planar motion on sandy terrain. Yang Yang 0044, Jun Luo 0006, Shaorong Xie, Huayan Pu, Yi Sun 0002, Na Liu 0004 |
IECON | 4 |
| 2017 | The Cooperative Defense Strategy by Multi-USVsabstractBased on the multi-agents system control theory and technology, this paper presents the cooperative defense process of multiple unmanned surface vehicles (USVs) operating in complicated sea environment, and explains the quantification of the battle effectiveness, cooperative strategy, task allocation and finally describes in detail the cooperative strategies on random graph. Then we point out the problems in the current cooperative defense process and the future development direction. The cooperative defense research of USVs in the sea environment has a great significance to the effective promotion of social and military efficiency. Yuan Liu 0025, Xing Wu 0001, Yike Guo, Shaorong Xie, Huayan Pu, Yan Peng 0001 |
SoMeT | 4 |
| 2017 | The Combat of Unmanned Surface Vehicles Based on Wolves AttackabstractUnmanned combat system is one of the development trend of modern weapons and equipment and has applied to military affairs. The major goal of the Unmanned Surface Vehicles (USVs in short) is to destroy protected targets in the shortest time. This paper originates from biology, putting forward a new attack strategy—The USV combat Based on Wolves Attack with Weight. Namely, using the characteristics of the wolves attack to study the process of attacking. It also discusses the attack measures from weights, velocity and firepower of USVs through weight distribution and summarizes the advantages of wolves attack in USV combat. Juan Pu, Xing Wu 0001, Yike Guo, Shaorong Xie, Huayan Pu, Yan Peng 0001 |
SoMeT | 4 |
| 2017 | The Cooperative Defense System by Team of USVs in Complicated Sea EnvironmentabstractBased on the multi-agents system control theory and technology, this paper presents the construction of cooperative defense system of multiple unmanned surface vehicles (USVs) operating in complicated sea environment, develops the mathematical formula describing the trajectory equation when the USV intercepts the intruder, and explains the proposed coordination control method used in the corresponding defense system. Then we point out the future development requirement of the cooperative defense system. The cooperative defense system research of USVs in the sea environment has a great significance to the effective promotion of social and military efficiency, and it is the basis of the implement about cooperative strategies. Xing Wu 0001, Yuan Liu 0025, Yike Guo, Shaorong Xie, Huayan Pu, Yan Peng 0001 |
SoMeT | 4 |
| 2017 | A System for Automated Detection of Ampoule Injection ImpuritiesabstractAmpoule injection is a routinely used treatment in hospitals due to its rapid effect after intravenous injection. During manufacturing, tiny foreign particles can be present in the ampoule injection. Therefore, strict inspection must be performed before ampoule injections can be sold for hospital use. In the quality control inspection process, most ampoule enterprises still rely on manual inspection which suffers from inherent inconsistency and unreliability. This paper reports an automated system for inspecting foreign particles within ampoule injections. A custom-designed hardware platform is applied for ampoule transportation, particle agitation, and image capturing and analysis. Constructed trajectories of moving objects within liquid are proposed for use to differentiate foreign particles from air bubbles and random noise. To accurately classify foreign particles, multiple features including particle area, mean gray value, geometric invariant moments, and wavelet packet energy spectrum are used in supervised learning to generate feature vectors. The results show that the proposed algorithm is effective in classifying foreign particles and reducing false positive rates. The automated inspection system inspects over 150 ampoule injections per minute (versus ~ 12 ampoule injections per minute by technologist) with higher accuracy and repeatability. In addition, the automated system is capable of diagnosing impurity types while existing inspection systems are not able to classify detected particles. Ji Ge, Shaorong Xie, Yaonan Wang 0001, Jun Liu 0007, Hui Zhang 0023, Falu Weng, Changhai Ru, Chao Zhou 0002, Min Tan 0001, Yu Sun 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2016 | An automated system for investigating sperm orientation in fluid flowabstractMammalian sperms reorient against fluid flow in the female reproductive tract, known as rheotaxis. Compared to chemotaxis that provides short-distance guidance, rheotaxis provides long-distance guidance for a sperm to find the egg cell. However, only a low number of sperms are capable of rheotaxis and their tail behavior during reorientation is not yet known. We have developed an automated system to manipulate human sperm orientation in fluid flow and quantitatively reveal sperm behavior changes during rheotaxis. The system automatically detects multiple sperms, selects the sperm for analysis, controls fluid flow, and quantifies sperm tail behavior. Sperm head angle is used as feedback to control fluid flow and select reorienting sperms. High accuracy of head angle tracking and automated sperm selection enables the capturing of dynamic sperm turning behavior in a large sample size. Algorithms are developed to track sperm tail skeletons and quantify tail beating amplitude and asymmetry, based on which the first quantitative analysis of sperm tail behavior in rheotaxis is obtained. Experimental results reveal, for the first time, that the sperms that are capable of reorienting against fluid flow beat their tails more asymmetrically than those sperms that are unable to reorient against fluid flow while no significant difference was found in their tail beating amplitudes. Zhuoran Zhang 0001, Jun Liu 0007, Jim Meriano, Changhai Ru, Shaorong Xie, Jun Luo 0006, Yu Sun 0001 |
ICRA | 5 |
| 2016 | Studying of rectilinear locomotion for a two-segment system with anisotropic dry friction modelabstractThis paper contributes to the understanding of the fundamental properties of rectilinearly locomotion of an one-dimensional system travelling on the horizontal plane, where dry Coulomb friction acting between it and surface. We discuss an approximate steady-state motion on a simplified two-segment system, which is propelled by a periodic internal excitation. First, the analysis is presented of the sufficient and necessary conditions for the system to move from the state of rest. Then, the explicit equations to calculate the constant average velocity of the steady-state are found. In addition, the influences of different parameters on the average velocity of the steady-state motion are discussed. Finally, the obtained theoretical results are verified by the numerical simulations. Shaorong Xie, Jun Luo 0006 |
IROS | 2 |
| 2014 | Automated microrobotic characterization of cell-cell communicationabstractMost mammalian cells (e.g., cancer cells and cardiomyocytes) adhere to a culturing surface. Compared to robotic injection of suspended cells (e.g., embryos and oocytes), fewer attempts were made to automate the injection of adherent cells due to their smaller size, highly irregular morphology, small thickness (a few micrometers thick), and large variations in thickness across cells. This paper presents a recently developed robotic system for automated microinjection of adherent cells. The system is embedded with several new capabilities: automatically locating micropipette tips; robustly detecting the contact of micropipette tip with cell culturing surface and directly with cell membrane; and precisely compensating for accumulative positioning errors. These new capabilities make it practical to perform adherent cell microinjection truly via computer mouse clicking in front of a computer monitor, on hundreds and thousands of cells per experiment (vs. a few to tens of cells as state-of-the-art). System operation speed, success rate, and cell viability rate were quantitatively evaluated based on robotic microinjection of over 4,000 cells. This paper also reports the use of the new robotic system to perform cell-cell communication studies using large sample sizes. The gap junction function in a cardiac muscle cell line (HL-1 cells), for the first time, was quantified with the system. Jun Liu 0007, Vinayakumar Siragam, Clement Leung, Zhe Lu, Changhai Ru, Shaorong Xie, Jun Luo 0006, Robert M. Hamilton, Yu Sun 0001 |
ICRA | 8 |
| 2014 | Locating End-Effector Tips in Robotic MicromanipulationabstractIn robotic micromanipulation, end-effector tips must be first located under microscopy imaging before manipulation is performed. The tip of micromanipulation tools is typically a few micrometers in size and highly delicate. In all existing micromanipulation systems, the process of locating the end-effector tip is conducted by a skilled operator, and the automation of this task has not been attempted. This paper presents a technique to automatically locate end-effector tips. The technique consists of programmed sweeping patterns, motion history image end-effector detection, active contour to estimate end-effector positions, autofocusing and quad-tree search to locate an end-effector tip, and, finally, visual servoing to position the tip to the center of the field of view. Two types of micromanipulation tools (micropipette that represents single-ended tools and microgripper that represents multiended tools) were used in experiments for testing. Quantitative results are reported in the speed and success rate of the autolocating technique, based on over 500 trials. Furthermore, the effect of factors such as imaging mode and image processing parameter selections was also quantitatively discussed. Guidelines are provided for the implementation of the technique in order to achieve high efficiency and success rates. Jun Liu 0007, Kathryn Tang, Zhe Lu, Changhai Ru, Jun Luo 0006, Shaorong Xie, Yu Sun 0001 |
IEEE Trans. Robotics | 7 |
| 2013 | Automated Pick-Place of Silicon NanowiresabstractPick-place of single nanowires inside scanning electron microscopes (SEM) is useful for prototyping functional devices and characterizing nanowires's properties. Nanowire pick-place has been typically performed via teleoperation, which is time-consuming and highly skill-dependent. This paper presents an automated approach to the pick-place of single nanowires. Through SEM visual detection and vision-based motion control, the system automatically transferred individual silicon nanowires from their growth substrate to a microelectromechanical systems (MEMS) device that characterized the nanowires's electromechanical properties. The performance of the nanorobotic pick-up and placement procedures was experimentally quantified. Xutao Ye, Yong Zhang 0046, Changhai Ru, Jun Luo 0006, Shaorong Xie, Yu Sun 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2007 | Biomimetic control of pan-tilt-zoom camera for visual tracking based-on an autonomous helicopterabstractA novel control strategy of pan-tilt-zoom camera is described. Because the active camera is mounted on a moving autonomous helicopter in visual tracking system, and the tracked object is moving at same time, and there exists the vibration influence of the helicopter, image stabilization becomes poor, and all pixels are running. Therefore, a biomimetic control strategy of on-board pan-tilt-zoom camera is presented. In this paper, the biomimetic oculomotor control model is obtained based on physiological neural path of eye movement control. In order to validate the functions of the biomimetic control model, simulation experiments were done under the same condition as the physiological experiments in physiological researches. Then the biomimetic controller of onboard pan-tilt-zoom camera is developed. The results of flight tracking experiments show that the biomimetic controller can compensate the deflection caused by the flight platform, and enhance the visual tracking system performance. Shaorong Xie, Jun Luo 0006, Zhenbang Gong, Hairong Zou, Xiangguo Fu |
IROS | 1 |
| 2006 | Real-time Vision-based Object Tracking from a Moving Platform in the AirabstractGenerally the scene of the visual surveillance using one camera is confined within certain limits. If you want to extend the surveillance area, you need multiple cameras. It attempts to detect, recognize and track certain objects from image sequences. This paper presents a simple and effective approach segmenting and tracking a moving object from a moving platform in the air such as UAV or ballonet airship, etc. The moving platform can follow the tracked object during a long distance. It can be used to track the doubtful vehicle and strengthen the visual surveillance more. The proposed approach is efficient and robust. The object is searched only in a subwindow of the current frame. The position of sub-window is estimated by its position in the last two frames. The features of the object are updated in each frame. The approach produces good tracking results at a frame rate of 25 fps. It is available in tracking a moving object in various scenes Zhenbang Gong, Shaorong Xie, Hairong Zou |
IROS | 3 |