EDBT 2026 Demo / reviewers in the wild / expert
Shoukun Xu
dblp:73/7832
· DBLP profile ↗
38ranked-venue papers
8as first author
35since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Computer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Mamba-driven sifter for salient object detection
Yi Liu 0038, Dingwen Zhang, Shoukun Xu |
Expert Syst. Appl. | 6 |
| 2026 | Integrating foundation models with capsule networks for enhanced weakly-supervised semantic segmentation
Zhuang Yao, Gengshen Wu, Yi Liu 0038, Shoukun Xu |
Expert Syst. Appl. | 6 |
| 2026 | Dynamic routing towards few-shot point cloud semantic segmentation
Guangqi Jiang, Zhengyao Li, Gengshen Wu, Yi Liu 0038, Shoukun Xu |
Neural Networks | 5 |
| 2026 | Visual In-Context Learning for Underwater Image RestorationabstractUnderwater images often exhibit common visual degradations, such as color distortion, loss of details, and reduced sharpness, which inevitably compromise the effectiveness of underwater vision tasks. However, most underwater image restoration methods solely focus on learning degradation features from raw images, neglecting the incorporation of additional contextual information to guide restoration, which limits the capability of deep models to restore image quality. In this paper, we propose Visual In-Context Learning (VICL) for underwater image restoration, which leverages degradation information from context to improve image quality. In VICL, Degraded Context Extraction Block (DCEB) employs a self-attention mechanism to extract degradation information from context. In addition, Context Spatial Feature Fusion Block (CSFFB) consists of a Degraded Context Guidance Block (DCGB) and a Multi-Feature Fusion Block (MFFB). DCGB employs a cross-attention mechanism to fuse degraded context with spatial features for guiding underwater image restoration. MFFB replaces traditional encoder-decoder skip connections to better coordinate feature fusion. Extensive experiments on multiple underwater image benchmarks demonstrate that VICL outperforms state-of-the-art methods both quantitatively and visually. The code is available at:https://github.com/zhangao668/VICL. Guangqi Jiang, Yi Liu 0038, Huibing Wang, Shoukun Xu |
IEEE Signal Process. Lett. | 5 |
| 2025 | Gradient Balanced Part-Whole Relational Weakly Supervised Semantic Segmentation
Zhuang Yao, Guangqi Jiang, Lin Shi 0007, Gengshen Wu, Shoukun Xu, Yi Liu 0038 |
KSEM (1) | 5 |
| 2025 | Autonomous cyber defense for AIoT using Graph Attention Network-Enhanced reinforcement learning
Yihang Shi, Lin Shi 0007, Shoukun Xu |
Comput. Commun. | 4 |
| 2025 | Seamless Detection: Unifying Salient Object Detection and Camouflaged Object Detection
Yi Liu 0038, Dingwen Zhang, Shoukun Xu, Jungong Han |
Expert Syst. Appl. | 6 |
| 2025 | Part-Whole Relational Fusion Towards Multi-Modal Scene Understanding
Yi Liu 0038, Shoukun Xu, Jungong Han |
Int. J. Comput. Vis. | 3 |
| 2025 | Advancements in few-shot nested Named Entity Recognition: The efficacy of meta-learning convolutional approaches
ShuaiChen Zhu, Lin Shi 0007, Shoukun Xu |
Neurocomputing | 4 |
| 2025 | CAFNet: Circular Attention Fusion for medical image segmentation
Baohua Yuan, Lin Shi 0007, Mingjie Jiang, Juxiao Zhang, Qile Qin, Shoukun Xu |
Knowl. Based Syst. | 10 |
| 2025 | MVG-FD: Multi-Modal Visual Guidance and Feature Decomposition for Underwater Image RestorationabstractUnderwater images are frequently affected by light absorption and scattering, which lead to color distortion, reduced contrast, and blurred details, significantly degrading overall image quality. Most underwater image restoration methods are confined to the pixel space of the raw modality, overlooking the important role of other modalities and different frequency-domain features. As a result, the representational capacity of deep learning models is not fully realized, affecting the generation of high-quality images. To address the above issues, we propose Multi-modal Visual Guidance and Feature Decomposition (MVG-FD) method for underwater image restoration. Specifically, we introduce Modality Visual Guidance (MVG) module, which integrates the complementary information provided by depth modality features into the raw features to guide the model in restoring the color of underwater images. Meanwhile, we design Feature Decomposition (FD) module, which utilizes Learnable Wavelet Decomposition (LWD) to decompose and extract the high-frequency bands of the raw features to help restore the texture details of the image. MVG-FD significantly improves PSNR and SSIM on existing datasets. The code is available at:https://github.com/zhangao668/MVG-FD. Guangqi Jiang, Yi Liu 0038, Huibing Wang, Shoukun Xu |
IEEE Signal Process. Lett. | 5 |
| 2025 | Efficient network defense policies via GNN-enhanced reinforcement learning
Shoukun Xu, Yihang Shi, Lin Shi 0007 |
J. Supercomput. | 1 |
| 2025 | Capsule Networks With Residual Pose RoutingabstractCapsule networks (CapsNets) have been known difficult to develop a deeper architecture, which is desirable for high performance in the deep learning era, due to the complex capsule routing algorithms. In this article, we present a simple yet effective capsule routing algorithm, which is presented by a residual pose routing. Specifically, the higher-layer capsule pose is achieved by an identity mapping on the adjacently lower-layer capsule pose. Such simple residual pose routing has two advantages: 1) reducing the routing computation complexity and 2) avoiding gradient vanishing due to its residual learning framework. On top of that, we explicitly reformulate the capsule layers by building a residual pose block. Stacking multiple such blocks results in a deep residual CapsNets (ResCaps) with a ResNet-like architecture. Results on MNIST, AffNIST, SmallNORB, and CIFAR-10/100 show the effectiveness of ResCaps for image classification. Furthermore, we successfully extend our residual pose routing to large-scale real-world applications, including 3-D object reconstruction and classification, and 2-D saliency dense prediction. The source code has been released on https://github.com/liuyi1989/ResCaps. Yi Liu 0038, De Cheng, Dingwen Zhang, Shoukun Xu, Jungong Han |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | DPNet: A Dual-Path Network With Distance-Aware Attention for Medical Image SegmentationabstractAccurate and automated segmentation of medical images is crucial for enhancing the efficiency of disease diagnosis and treatment. In the past few years, there has been a proliferation of models based on convolutional neural networks (CNN), and they have achieved remarkable results in various medical image segmentation tasks. However, due to the limitations of the convolutional receptive field, it is impossible to effectively extract long-range contextual features. Although the receptive field theoretically increases gradually with the deepening of the neural network hierarchy, it still restricts the network’s ability to learn the long-range dependencies during the encoding process. To address these challenges, this paper proposes a dual-path network with Distance-Aware Attention for medical image segmentation (DPNet). Specifically, we open up an auxiliary path in the model encoder’s deep layer to learn the image’s global contextual information. The module for learning global feature information is called Distance-Aware Attention (DAA) Module, which determines the degree of attention between different parts of the image based on the distance between them, so that the attention is more targeted. In addition, we design a Channel-Spatial Attention Guidance (CSAG) module to let the extracted global features guide the local features to fully utilize the extracted global information and achieve the segmentation effect at a higher level. In order to verify the rationality of our model design, we conduct extensive experiments on the datasets ISIC2017 and ISIC2018, and the experimental results show that DPNet outperforms other state-of-the-art segmentation models. The code is available at https://github.com/w573352/DPNet. Shoukun Xu, Qile Qin, Baohua Yuan |
BIBM | 1 |
| 2024 | Autonomous Cyber Defense using Graph Attention Network Enhanced Reinforcement LearningabstractUbiquitous networked hosts and Internet of Things (IoT) devices have enabled critical network applications in both enterprise and industrial environments. However, network threats have proliferated, constantly challenging normal network operations and data security of IoT devices. In facing these challenges, defenders have increasingly adopted Deep Reinforcement Learning (DRL) approaches, aiming to leverage their self-learning and adaptability to bolster network security. Despite these efforts, existing approaches often exhibit significant performance bottlenecks when navigating the complexities of network scenarios. This paper introduces a novel algorithm, the Normalized Graph-Attention Proximal Policy Optimization (NGA-PPO), which synergistically integrates Graph Attention Networks (GAT) with the existing Proximal Policy Optimization (PPO) framework. By harnessing a weighted mechanism to amalgamate intricate network connectivity data, NGA-PPO empowers intelligent defense agents to dynamically decipher dependencies and interactions within complex network architectures, thereby facilitating precise and prompt defense actions. Comprehensive experiments within the Yawning Titan network security simulation platform reveal that NGA-PPO outperforms other methods, delivering substantial improvements in performance and exhibiting distinct advantages in robustness and generalization capabilities. Specifically, as network scenarios become more complex, existing algorithms face significant performance limitations, whereas NGA-PPO consistently demonstrates high performance, thereby affirming its efficacy and viability in mitigating complex network threats. Yihang Shi, Lin Shi 0007, Shoukun Xu |
ICPADS | 4 |
| 2024 | Spatial-Frequency Integration Network with Dual Prompt Learning for Few-shot Image ClassificationabstractFew-shot image classification is a challenging task that aims to recognize image classes based on only a few training images. However, existing methods face the following two main challenges: (1) Ignoring the frequency domain information during image feature extraction. (2) It does not take the semantic gap between multiple modalities into consideration, which limits the classification performance. To overcome these limitations, we propose a novel method named Spatial-Frequency Integration Network with Dual Prompt Learning for few-shot image classification. Firstly, we introduce a spatial-frequency integration module that combines spatial domain and low-frequency information to extract discriminative image features from the image modality. Secondly, we design a dual prompting module, which integrates learnable prompts and hand-crafted prompts to improve the generalization of applications to new classes. Thirdly, we propose an image-text interaction module to enhance inter-modal complementary and consistency. Both theoretical and experimental validations confirm the effectiveness of the proposed method in few-shot image classification. Yi Liu 0038, Shoukun Xu, Guangqi Jiang |
ISPA | 3 |
| 2024 | Pose Convolutional Routing Towards Lightweight Capsule NetworksabstractAn important branch of artificial intelligence systems and architectures is Capsule Networks (CapsNets) have been known extremely large amount of parameters and computation because of the complex capsule routing algorithm, making it difficult deep architectures in the era of deep learning. To address this challenge, in this paper, we propose a simple yet effective capsule routing algorithm. Specifically, we activate the pose of the entity using its activation probability. On top of that, a convolution on the activated pose matrix to learn the high-level capsules’ pose matrices. Activations of the high-level capsules can be digged from their pose matrices via convolution and activation. Such mechanism generates fewer network parameters and lightweight computation, which make it practitable a deep CapsNets architecture. Experiments on CIFAR-10/100, Small-NORB, MINIST and even large-scale benchmark PASCAL VOC 2007, demonstrate the effectiveness of the proposed method. Chengxin Lv, Guangqi Jiang, Shoukun Xu, Yi Liu 0038 |
ISPA | 3 |
| 2024 | LF-DETR: Lightweight Detection Network for Floating Surface Garbage Based on Adaptive Channel ShufflingabstractTo address the issues of low accuracy and efficiency in water surface target detection caused by uneven lighting and water ripples, this study proposes a channel modeling-based DETR-like model. Starting from efficient partial convolutions and inverted residual blocks, the concept of dynamic sparsity is introduced, and a Flexible Masked Residual Block (FMRB) for a lightweight backbone is designed. Based on the design principle of channel grouping, an adaptive channel shuffling strategy is derived. Our model achieved an mAP of 79.8% on the FGDD dataset and 61.7% on the Pascal VOC2007 dataset. Compared to the original model, the FLOPs of our LF-DETR model are reduced by 24.9%, down to 42.8G, and the number of parameters is reduced by 29.4%. Extensive experimental results demonstrate that our model exhibits an excellent balance between efficiency and accuracy across multiple benchmarks, highlighting its great potential in practical applications. Shoukun Xu, Gaochao Yang, Ning Li 0048 |
ISPA | 1 |
| 2024 | Federated Multi-Agent Reinforcement Learning for AoI Minimization in UAV-Enabled IoV-MEC SystemsabstractWith the advancement of Mobile Edge Computing (MEC), effective solutions for communication scenarios, including the industrial Internet of Things (IoT) and the Internet of Vehicles (IoV), are becoming increasingly feasible. Unmanned aerial vehicles (UAVs) can further enhance flexibility in delivering computational services within MEC contexts. Addressing the urgent need for information freshness, we propose a three-tier IoV-MEC system, supported by multiple UAVs and a cloud center, aiming to minimize the system's average age of information (AoI). We propose a heterogeneous multi-agent reinforcement learning algorithm based on the actor-critic framework, where vehicles act as data sources, UAVs serve as edge devices, and the cloud acts as a control center. All three tiers learn interaction strategies cooperatively based on the observations. To further enhance system performance, we implement an efficient federated learning method, allowing same-tier agents to share learning parameters, thus improving system performance and convergence speed. Extensive simulation results demonstrate that the proposed algorithm outperforms baseline algorithms in terms of average AoI and convergence speed. Shoukun Xu, Xueyuan Wang, Mustafa Cenk Gursoy |
ISPA | 1 |
| 2024 | SymbTQL: A Symbolic Temporal Logic Framework for Enhancing LLM-based TKGQAabstractTemporal Knowledge Graph Question Answering (TKGQA) deals with the intricate task of processing and understanding time-evolving facts, and responding to natural language queries that involve sophisticated temporal constraints. Current Large Language Models (LLM) exhibit notable limitations in converting natural language queries into temporal reasoning tasks within complex TKGQA contexts. To mitigate these limitations, this paper proposes a pioneering approach named Symbolic Temporal Query for LLM TKGQA (SymbTQL). This novel framework bolsters LLM capacity to tackle complex temporal challenges through a tripartite process: initially, LLM is fine-tuned with a limited set of annotated data to heighten its comprehension of temporal constraints within TKGQA tasks; this is followed by employing a symbolic logic chain reasoning framework that breaks down intricate natural language queries into a sequence of simpler sub-questions; ultimately, the accuracy of the queries is ensured by aligning the generated query statements with actual data using sentence similarity encoding. Experimental evidence indicates that SymbTQL outperforms numerous baselines by 5.1% and 3.5% on the MultiTQ and CronQuestions datasets, respectively, and excels by 15.4% in managing multi-step complex temporal challenges, with ablation studies confirming its strong generalizability and logical coherence. ShuaiChen Zhu, Lin Shi 0007, Shoukun Xu |
ISPA | 4 |
| 2024 | Enhancing the Transferability and Stealth of Deepfake Detection Attacks Through Latent Diffusion Models
Shoukun Xu |
PRCV (4) | 2 |
| 2024 | Deep unsupervised part-whole relational visual saliency
Yi Liu 0038, Dingwen Zhang, Shoukun Xu |
Neurocomputing | 4 |
| 2024 | FAUC-S: Deep AUC maximization by focusing on hard samples
Shoukun Xu, Yanrui Ding, Junru Luo |
Neurocomputing | 1 |
| 2024 | Multiview latent space learning with progressively fine-tuned deep features for unsupervised domain adaptation
Chenyang Zhu 0001, Qian Wang 0017, Yunxin Xie, Shoukun Xu |
Inf. Sci. | 4 |
| 2024 | DENS-YOLOv6: a small object detection model for garbage detection on water surface
Gaochao Yang, Baohua Yuan, Shoukun Xu |
Multim. Tools Appl. | 6 |
| 2024 | TCGNet: Type-Correlation Guidance for Salient Object DetectionabstractContrast and part-whole relations induced by deep neural networks like Convolutional Neural Networks (CNNs) and Capsule Networks (CapsNets) have been known as two types of semantic cues for deep salient object detection. However, few works pay attention to their complementary properties in the context of saliency prediction. In this paper, we probe into this issue and propose a Type-Correlation Guidance Network (TCGNet) for salient object detection. Specifically, a Multi-Type Cue Correlation (MTCC) covering CNNs and CapsNets is designed to extract the contrast and part-whole relational semantics, respectively. Using MTCC, two correlation matrices containing complementary information are computed with these two types of semantics. In return, these correlation matrices are used to guide the learning of the above semantics to generate better saliency cues. Besides, a Type Interaction Attention (TIA) is developed to interact semantics from CNNs and CapsNets for the aim of saliency prediction. Experiments and analysis on five benchmarks show the superiority of the proposed approach. Codes has been released on https://github.com/liuyi1989/TCGNet. Yi Liu 0038, Ling Zhou 0002, Gengshen Wu, Shoukun Xu, Jungong Han |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | A No Parameter Synthetic Minority Oversampling Technique Based on Finch for Imbalanced Data
Shoukun Xu, Zhibang Li, Baohua Yuan, Gaochao Yang, Xueyuan Wang |
ICIC (4) | 1 |
| 2023 | Enhancing Cybersecurity in Industrial Control System with Autonomous Defense Using Normalized Proximal Policy Optimization ModelabstractIndustrial control networks are frequent targets of cyber attacks, calling for autonomous defense strategies to combat threats effectively and promptly. We propose to leverage Reinforcement Learning (RL) to train defenders how to select the most appropriate actions in response to attackers’ behavior. We model the defender and attacker as agents in RL environments and define action space and reward function. We focus on evaluating value-based and policy gradient-based RL algorithms and propose enhancements to Proximal Policy Optimization (PPO) method through normalization. We then leverage CybORG, a simulation tool to study how the RL algorithms perform to achieve autonomous defense. Our extensive experiments reveal that the proposed normalized PPO outperforms other models in terms of stability, robustness, and average rewards. The results also demonstrate the effectiveness of the defender in protecting the operational server upon changing the attacker’s position. Additionally, increasing the number of attackers had a significant impact on policy-based algorithms, particularly PPO, with decreased reward values highlighting the increased vulnerability of the network. Shoukun Xu, Zihao Xie, Chenyang Zhu 0006, Xueyuan Wang, Lin Shi 0007 |
ICPADS | 1 |
| 2023 | DKTNet: Dual-Key Transformer Network for small object detection
Shoukun Xu, Jianan Gu, Yining Hua |
Neurocomputing | 1 |
| 2022 | Blind stereoscopic image quality assessment using 3D saliency selected binocular perception and 3D convolutional neural network
Lixia Yun, Shoukun Xu |
Multim. Tools Appl. | 3 |
| 2022 | Disentangled Capsule Routing for Fast Part-Object Relational SaliencyabstractRecently, the Part-Object Relational (POR) saliency underpinned by the Capsule Network (CapsNet) has been demonstrated to be an effective modeling mechanism to improve the saliency detection accuracy. However, it is widely known that the current capsule routing operations have huge computational complexity, which seriously limited the usability of the POR saliency models in real-time applications. To this end, this paper takes an early step towards a fast POR saliency inference by proposing a novel disentangled part-object relational network. Concretely, we disentangle horizontal routing and vertical routing from the original omnidirectional capsule routing, thus generating Disentangled Capsule Routing (DCR). This mechanism enjoys two advantages. On one hand, DCR that disentangles orthogonal 1D (i.e., vertical and horizontal) routing greatly reduces parameters and routing complexity, resulting in much faster inference than omnidirectional 2D routing adopted by existing CapsNets. On the other hand, thanks to the light POR cues explored by DCR, we could conveniently integrate the part-object routing process to different feature layers in CNN, rather than just applying it to the small-scaled one as in previous works. This helps to increase saliency inference accuracy. Compared to previous POR saliency detectors, DPORTNet infers visual saliency (5 ∼ 9 ) × faster, and is more accurate. DPORTNet is available under the open-source license at https://github.com/liuyi1989/DCR. Yi Liu 0038, Dingwen Zhang, Nian Liu 0002, Shoukun Xu, Jungong Han |
IEEE Trans. Image Process. | 4 |
| 2022 | Learning Discriminative Cross-Modality Features for RGB-D Saliency DetectionabstractHow to explore useful information from depth is the key success of the RGB-D saliency detection methods. While the RGB and depth images are from different domains, a modality gap will lead to unsatisfactory results for simple feature concatenation. Towards better performance, most methods focus on bridging this gap and designing different cross-modal fusion modules for features, while ignoring explicitly extracting some useful consistent information from them. To overcome this problem, we develop a simple yet effective RGB-D saliency detection method by learning discriminative cross-modality features based on the deep neural network. The proposed method first learns modality-specific features for RGB and depth inputs. And then we separately calculate the correlations of every pixel-pair in a cross-modality consistent way, i.e., the distribution ranges are consistent for the correlations calculated based on features extracted from RGB (RGB correlation) or depth inputs (depth correlation). From different perspectives, color or spatial, the RGB and depth correlations end up at the same point to depict how tightly each pixel-pair is related. Secondly, to complemently gather RGB and depth information, we propose a novel correlation-fusion to fuse RGB and depth correlations, resulting in a cross-modality correlation. Finally, the features are refined with both long-range cross-modality correlations and local depth correlations to predict salient maps. In which, the long-range cross-modality correlation provides context information for accurate localization, and the local depth correlation keeps good subtle structures for fine segmentation. In addition, a lightweight DepthNet is designed for efficient depth feature extraction. We solve the proposed network in an end-to-end manner. Both quantitative and qualitative experimental results demonstrate the proposed algorithm achieves favorable performance against state-of-the-art methods. Fengyun Wang, Jinshan Pan, Shoukun Xu, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Throughput-aware path planning for UAVs in D2D 5G networks
Lin Shi 0007, Zhongyi Jiang, Shoukun Xu |
Ad Hoc Networks | 3 |
| 2021 | Automatic detection of safety helmet wearing based on head region locationabstractAbstract In order to solve the problem of difficult and low precision in the detection of safety helmet wearing in the complex pose of construction worker, a detection method of safety helmet wearing based on pose estimation is proposed. In the pose estimation model of OpenPose, the residual network optimized feature extraction is introduced to obtain the skeletal point information of the construction worker, and then the pose of the construction worker is estimated based on the skeletal point information, three‐point localization method is proposed for the front and back pose, and skin colour detection method is proposed for the side pose, and then to determine the head region. The YOLO v4 is used to detect the safety helmet region, and then the construction worker's safety helmet wearing is judged according to whether the head region intersects the safety helmet region or not. Experimental results show that the detection accuracy of the method is higher than other methods, and the adaptability to the environment is stronger. Yuwan Gu, Lin Shi 0007, Lihua Zhuang, Shoukun Xu |
IET Image Process. | 6 |
| 2021 | An ensemble multi-scale residual attention network (EMRA-net) for image Dehazing
Jixiao Wang, Shoukun Xu |
Multim. Tools Appl. | 3 |
| 2020 | QoS-Aware UAV Coverage path planning in 5G mmWave network
Lin Shi 0007, Shoukun Xu, Zhongxu Zhan |
Comput. Networks | 2 |
| 2020 | Fabric defect inspection based on lattice segmentation and template statistics
Liang Jia, Chen Chen 0001, Shoukun Xu, Ju Shen |
Inf. Sci. | 3 |
| 2020 | Two stages double attention convolutional neural network for crowd counting
Zhao Zou, Yuhui Zheng, Shoukun Xu |
Multim. Tools Appl. | 4 |