Shaohua Dong

dblp:188/4523 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 OA-WGAN: A Clustering-Guided GAN for Disk Fault Augmentation
Shaohua Dong
ICA3PP (2)2
2025 Efficient and Accurate Low-Resolution Transformer Tracking
abstract
High-performance Transformer trackers have exhibited excellent results, yet they often bear a heavy computational load. Observing that a smaller input can immediately and conveniently reduce computations without changing the model, an easy solution is to adopt a low-resolution input for efficient Transformer tracking. Albeit faster, this hurts tracking accuracy much due to the information loss in low resolution tracking. In this paper, we aim to mitigate such information loss to boost performance of low-resolution Transformer tracking via dual knowledge distillation from a frozen high-resolution (but not a larger) Transformer tracker. The core lies in two simple yet effective distillation modules, including query-key-value knowledge distillation (QKV-KD) and discrimination knowledge distillation (Disc-KD), across resolutions. The former, from the global view, allows the low-resolution tracker to inherit features and interactions from the high-resolution tracker, while the later, from the target-aware view, enhances the target-background distinguishing capacity via imitating discriminative regions from its high-resolution counterpart. With dual knowledge distillation, our Low-Resolution Transformer Tracker, dubbed LoReTrack, enjoys not only high efficiency owing to reduced computation but also enhanced accuracy by distilling knowledge from the high-resolution tracker. In extensive experiments, LoReTrack with a 2562resolution consistently improves baseline with the same resolution, and shows competitive or better results compared to the 3842high-resolution Transformer tracker, while running 52% faster and saving 56% MACs. Moreover, LoReTrack is resolution-scalable. With a 1282resolution, it runs 25 fps on a CPU with SUC scores of 64.9%/46.4% on LaSOT/LaSOText, surpassing other CPU real-time trackers. Code is released at https://github.com/ShaohuaDong2021/LoReTrack.
Shaohua Dong, Yunhe Feng, James Liang, Qing Yang 0003, Yuewei Lin, Heng Fan 0001
IROS1
2025 Denoising of the magnetic flux leakage signal using dynamic feature fusion and a multi-scale autoencoder network
Lushuai Xu, Shaohua Dong, Haotian Wei, Pengkun Zhang, Cong Zuo 0002, Mingxing Guo, Penghui Liao
Eng. Appl. Artif. Intell.2
2024 Beyond MOT: Semantic Multi-object Tracking
Hao Wang 0093, Jiali Yao, Shaohua Dong, Heng Fan 0001, Libo Zhang 0001
ECCV (35)6
2024 Efficient Multimodal Semantic Segmentation via Dual-Prompt Learning
abstract
Multimodal (e.g., RGB-Depth/RGB-Thermal) fusion has shown great potential for improving semantic segmentation in complex scenes (e.g., indoor/low-light conditions). Existing approaches often fully fine-tune a dual-branch encoder-decoder framework with a complicated feature fusion strategy for achieving multimodal semantic segmentation, which is training-costly due to the massive parameter updates in feature extraction and fusion. To address this issue, we propose a surprisingly simple yet effective dual-prompt learning network (dubbed DPLNet) for training-efficient multimodal (e.g., RGBD/T) semantic segmentation. The core of DPLNet is to directly adapt a frozen pre-trained RGB model to multimodal semantic segmentation, reducing parameter updates. For this purpose, we present two prompt learning modules, comprising multimodal prompt generator (MPG) and multimodal feature adapter (MFA). MPG works to fuse the features from different modalities in a compact manner and is inserted from shallow to deep stages to generate the multi-level multimodal prompts that are injected into the frozen backbone, while MFA adapts prompted multimodal features in the frozen backbone for better multimodal semantic segmentation. Since both the MPG and MFA are lightweight, only a few trainable parameters (3.88M, 4.4% of the pre-trained backbone parameters) are introduced for multimodal feature fusion and learning. Using a simple decoder (3.27M parameters), DPLNet achieves new state-of-the-art performance or is on a par with other complex approaches on four RGB-D/T semantic segmentation datasets while satisfying parameter efficiency. Moreover, we show DPLNet is general and applicable to other multimodal segmentation tasks. Without special design, DPLNet outperforms many complicated models. The source code can be found at https://github.com/ShaohuaDong2021/DPLNet.
Shaohua Dong, Yunhe Feng, Qing Yang 0003, Yan Huang 0002, Dongfang Liu, Heng Fan 0001
IROS1
2024 VastTrack: Vast Category Visual Object Tracking
abstract
In this paper, we propose a novel benchmark, named VastTrack, aiming to facilitate the development of general visual tracking via encompassing abundant classes and videos. VastTrack consists of a few attractive properties: (1) Vast Object Category. In particular, it covers targets from 2,115 categories, significantly surpassing object classes of existing popular benchmarks (e.g., GOT-10k with 563 classes and LaSOT with 70 categories). Through providing such vast object classes, we expect to learn more general object tracking. (2) Larger scale. Compared with current benchmarks, VastTrack provides 50,610 videos with 4.2 million frames, which makes it to date the largest dataset in term of the number of videos, and hence could benefit training even more powerful visual trackers in the deep learning era. (3) Rich Annotation. Besides conventional bounding box annotations, VastTrack also provides linguistic descriptions with more than 50K sentences for the videos. Such rich annotations of VastTrack enable the development of both vision-only and vision-language tracking. In order to ensure precise annotation, each frame in the videos is manually labeled with multi-stage of careful inspections and refinements. To understand performance of existing trackers and to provide baselines for future comparison, we extensively evaluate 25 representative trackers. The results, not surprisingly, display significant drops compared to those on current datasets due to lack of abundant categories and videos from diverse scenarios for training, and more efforts are urgently required to improve general visual tracking. Our VastTrack, the toolkit, and evaluation results are publicly available at https://github.com/HengLan/VastTrack.
Junyuan Gao, Weihong Li 0002, Shaohua Dong, Heng Fan 0001, Libo Zhang 0001
NeurIPS5
2024 Intelligent identification of girth welds defects in pipelines using neural networks with attention modules
Lushuai Xu, Shaohua Dong, Haotian Wei, Donghua Peng, Weichao Qian, Qingying Ren, Luming Wang, Yundong Ma
Eng. Appl. Artif. Intell.2
2024 Multichannel Multimodal Emotion Analysis of Cross-Modal Feedback Interactions Based on Knowledge Graph
abstract
Abstract Multimodal sentiment analysis is a downstream branch task of sentiment analysis with high attention at present. Previous work in multimodal sentiment analysis have focused on the representation and fusion of modalities, capturing the underlying semantic relationships between modalities by considering contextual information. While this approach is feasible for simple contextual comments, more complex comments require the integration of external knowledge to obtain more accurate sentiment information. However, incorporating external knowledge into sentiment analysis to enhance information complementarity has not been thoroughly investigated. To address this, we propose a multichannel cross-modal feedback interaction model that incorporates the knowledge graph into multimodal sentiment analysis. Our proposed model consists of two main components: the cross-modal feedback recurrent interaction module and the external knowledge module for capturing latent information. The cross-modal interaction employs a self-feedback mechanism during network training, extracting feature representations of each modality and using these representations to mask sensory inputs, allowing the model to perform feedback-based feature masking. The external knowledge graph captures potential semantic information representations in the textual data through knowledge graph embedding. Finally, a global feature fusion module is employed for multichannel multimodal information integration. On two publicly available datasets, our method demonstrates good performance in terms of accuracy and F1 scores, compared to state-of-the-art models and several baselines.
Shaohua Dong, Xiaochao Fan, Xinchun Ma
Neural Process. Lett.1
2024 EGFNet: Edge-Aware Guidance Fusion Network for RGB-Thermal Urban Scene Parsing
abstract
Urban scene parsing is the core of the intelligent transportation system, and RGB–thermal urban scene parsing has recently attracted increasing research interest in the field of computer vision. However, most existing approaches fail to perform good boundary extraction for prediction maps and cannot fully use high-level features. In addition, these methods simply fuse the features from RGB and thermal modalities but are unable to obtain comprehensive fused features. To address these problems, an edge-aware guidance fusion network (EGFNet) was developed in this study for RGB–thermal urban scene parsing. First, a prior edge map generated using the RGB and thermal images were introduced to capture detailed information in the prediction map and then embed the prior edge cues into the feature maps. To fuse the RGB and thermal information effectively, a multimodal fusion module was designed that guarantees adequate cross-modal fusion. Considering the importance of high-level semantic information, global and semantic information modules were proposed to extract rich semantic information from the high-level features. For decoding, simple elementwise addition was utilized for cascaded feature fusion. Finally, to improve the parsing accuracy, multitask deep supervision was applied to the semantic and boundary maps. Extensive experiments were performed on benchmark datasets to demonstrate the effectiveness of the proposed EGFNet and its superior performance compared with the state-of-the-art methods.
Shaohua Dong, Wujie Zhou, Caie Xu, Weiqing Yan
IEEE Trans. Intell. Transp. Syst.1
2022 Edge-Aware Guidance Fusion Network for RGB-Thermal Scene Parsing
abstract
RGB–thermal scene parsing has recently attracted increasing research interest in the field of computer vision. However, most existing methods fail to perform good boundary extraction for prediction maps and cannot fully use high-level features. In addition, these methods simply fuse the features from RGB and thermal modalities but are unable to obtain comprehensive fused features. To address these problems, we propose an edge-aware guidance fusion network (EGFNet) for RGB–thermal scene parsing. First, we introduce a prior edge map generated using the RGB and thermal images to capture detailed information in the prediction map and then embed the prior edge information in the feature maps. To effectively fuse the RGB and thermal information, we propose a multimodal fusion module that guarantees adequate cross-modal fusion. Considering the importance of high-level semantic information, we propose a global information module and a semantic information module to extract rich semantic information from the high-level features. For decoding, we use simple elementwise addition for cascaded feature fusion. Finally, to improve the parsing accuracy, we apply multitask deep supervision to the semantic and boundary maps. Extensive experiments were performed on benchmark datasets to demonstrate the effectiveness of the proposed EGFNet and its superior performance compared with state-of-the-art methods. The code and results can be found at https://github.com/ShaohuaDong2021/EGFNet.
Wujie Zhou, Shaohua Dong, Caie Xu, Yaguan Qian
AAAI2
2022 GEBNet: Graph-Enhancement Branch Network for RGB-T Scene Parsing
abstract
RGB-T (red–green–blue and thermal) scene parsing has recently drawn considerable research attention. Although existing methods efficiently conduct RGB-T scene parsing, their performance remains limited by a small receptive field. Unlike methods that capture the global context by fusing multiscale features or using an attention mechanism, we propose a graph-enhancement branch network (GEBNet), which uses long-range dependencies obtained from the branch to refine a coarse semantic map generated by the decoder. Semantic and detail modules embedded in the graph-enhancement branch fuse high- and low-level features. Furthermore, inspired by the ability of graph neural networks to capture the global context, we integrate a novel graph-enhancement module into the network branch to obtain global information from both high-level semantic information and low-level details. Results from extensive experiments on the MFNet and PST900 datasets demonstrate the high performance of the proposed GEBNet and the contributions of its main components to the parsing performance.
Shaohua Dong, Wujie Zhou, Xiaohong Qian, Lu Yu 0003
IEEE Signal Process. Lett.1
2019 p-hub median location optimization of hub-and-spoke air transport networks in express enterprise
abstract
Summary Based on the single allocation p‐hub median location problem in hub‐and‐spoke networks, this paper studies the aviation hubs location optimization of express enterprise, primarily using p‐median model and greedy dropping heuristic algorithm to solve the initial candidate locations of the aviation hubs by Lingo. Then, the integer program model of the location and allocation of the aviation hubs was set up, with the upper limit of the distribution distance as the main constraint. The improved locations were calculated by immune algorithm on Matlab. The numerical example of A express enterprise was presented. After analyzing the status and the problems of aviation hub location in East China, the two best hub locations Wuxi and Hangzhou were finally determined by above model with the total cost reduced by 19.6%. The optimization results show the validation of the model.
Xinwei Zhang 0005, Shaohua Dong
Concurr. Comput. Pract. Exp.2