Kechen Song

dblp:171/5985 · DBLP profile ↗
← Back
50ranked-venue papers
7as first author
42since 2021 · last 2026
0000-0002-7636-3460ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Transfer graph reasoning network for misaligned visible-thermal object detection
Xiaotong Xue, Hongshu Chen, Kechen Song, Yunhui Yan, Baihua Li, Qinggang Meng
Knowl. Based Syst.3
2026 Efficient Six-Degrees of Freedom (6-DoF) Grasp Pose Detection in Cluttered Scenes via Multimodal Fusion and Object-Centric Receptive Fields
abstract
Efficient 6-DoF grasp pose detection is a fundamental and challenging task in robotic manipulation. Existing methods mainly rely on the geometric features of point clouds to synthesise 6-DoF grasp poses, as point clouds contain rich geometric information. However, the rich visual cues contained in images, such as object edges, textures, and other fine-grained details, can also serve as important references. In this paper, we propose a simple yet efficient multimodal 6-DoF grasp pose detection algorithm that also integrates object-centric receptive fields. Specifically, we design an early fusion strategy to integrate point cloud and image information, providing more discriminative multimodal features for grasp pose detection. Furthermore, an object-centric receptive fields module based on grounded segment anything model (Grounded-SAM) is introduced, which provides global references for each point during the grasp parameter decoding stage. Experimental results show that the proposed model achieves competitive performance on the GraspNet-1Billion dataset. Compared with baseline, the AP improves by 7.44, 8.42, and 2.37 in Seen, Similar, and Novel scenes, respectively. To test the model's robustness to image noise, we designed a series of noise experiments, demonstrating the model's strong anti-interference ability. Finally, we deployed the algorithm on a Franka robot and conducted extensive grasping experiments, achieving a high grasp success rate.
Xiaozheng Liu, Kechen Song, ZhongLei Liu, Zehao Xu, Yunhui Yan
IEEE Trans. Ind. Informatics2
2026 MCF-Det: A Cross-Domain Fusion Framework for Few-Shot Blade Defect Detection
abstract
Industrial aero-engine blades are safety-critical, yet labeled blade-defect images are scarce, whereas related steel-surface defects are well annotated. This setting results in cross-domain few-shot defect detection: transferring knowledge from steel to blades under limited target supervision. Such transfer is hindered by material bias, diluted defect cues in multiscale features, and unstable few-shot decision boundaries. We propose MCF-Det, a YOLOv10-based framework that addresses these issues via three integrated modules: material knowledge transfer module aligns material-aware features to suppress transfer bias; cross-domain feature adaptation module enhances defect-sensitive responses through spatial–channel reweighting; few-shot defect calibration module calibrates logits using a contrastively learned prototype bank. The framework enables end-to-end detection without altering YOLOv10 deployment interfaces and with bounded inference overhead. On a proprietary blade dataset under a unified few-shot protocol, multi-scale cross-feature fusion detection (MCF-Det) improves mean average precision (mAP) by 5.4 points over the strongest of 16 retrained baselines, while additional cross-material transfer and stress tests offer supplementary robustness analysis under larger appearance shifts. K-shot analysis and nonparameteric tests confirm consistent and significant gains in low-shot regimes. We release code, configurations, and scripts to reproduce all results reported on public benchmarks under the reported splits and seeds.
Xin Wen 0014, Yu He 0004, Haixu Yin, Kechen Song
IEEE Trans. Ind. Informatics5
2025 UAV applications in intelligent traffic: RGBT image feature registration and complementary perception
Yingying Ji, Kechen Song, Hongwei Wen, Xiaotong Xue, Yunhui Yan, Qinggang Meng
Adv. Eng. Informatics2
2025 Limited label-support pavement damage segmentation network with uniform rectification and intrinsic cross-dimensional constraint
Yunhui Yan, Yanyan Wang 0007, Kechen Song, Liming Huang
Adv. Eng. Informatics3
2025 An edge-guided defect segmentation network for in-service aerospace engine blades
Xianming Yang, Kechen Song, Shaoning Liu, Fuqi Sun, Jun Li 0119, Yunhui Yan
Eng. Appl. Artif. Intell.2
2025 Leveraging labelled data knowledge: A cooperative rectification learning network for semi-supervised 3D medical image segmentation
Yanyan Wang 0007, Kechen Song, Yuyuan Liu, Yunhui Yan, Gustavo Carneiro 0001
Medical Image Anal.2
2025 TRDM: A Two-Stage Real-Time Discrimination Method for Spiral Weld Defects Under Dynamic Distorted Imaging
abstract
Non-destructive testing of weld based on radio-graphic images is crucial for quality control of spiral steel pipes. However, accurately discriminating defect images from high-resolution and low-contrast weld images remains a challenging task. Moreover, the real-time imaging of radiography in industrial settings introduces dynamic distorted interferences to weld images, which has received limited attention in existing research. To address these issues, we propose a two-stage real-time discrimination method (TRDM) consisting of two carefully designed networks: an extremely lightweight weld region segmentation network and a semantic reconstruction-based discrimination network. Our method aims to adaptively reduce interference from invalid regions in the weld images to better exploit defect features from a semantic perspective and reconstruct them as defect-free regions as much as possible. To overcome the scarcity of weld defect image data, we construct two publicly available industrial weld datasets, including an image dataset and a video dataset. TRDM achieves competitive performance with state-of-the-art techniques, demonstrating its effectiveness in discriminating weld images. Furthermore, we integrate the TRDM into developed detection software for defect assessment, validating its efficacy in real welding scenarios. The source code and dataset are publicly available at https://github.com/Weld-det/TRDM.
Kechen Song, Yu Zhang 0213, Xiujian Jia, Xiaozheng Liu, Yunhui Yan
IEEE Trans Autom. Sci. Eng.2
2025 SDD-DETR: Surface Defect Detection for No-Service Aero-Engine Blades With Detection Transformer
abstract
Vision-based surface defect detection (SDD) for no-service aero-engine blades provides a fast and effective way to monitor product quality. Most existing detection algorithms for aero-engine blades are 1) based on CNN, including artificially designed non-maximum suppression (NMS) operations, and 2) focus on improving the detection accuracy rather than improving the inference speed and even ignoring the latter. To solve the above problems, we introduce a novel object detection paradigm, DEtection TRansformer (DETR), to design a novel network (SDD-DETR) with high accuracy for the SDD of aero-engine blades. To our knowledge, the paper is the first to introduce the DETR detector to SDD of aero-engine blades. While providing high accuracy, the inference speed of DETR remained slow due to self-attention operation and feed-forward network (FFN). Therefore, two lightweight modules have been designed for SDD of aero-engine blades: a progressive feature input multi-scale deformable attention module (PFI-MSDA) and a lightweight FFN (LW-FFN). PFI-MSDA hierarchically reduces the number of tokens input to the self-attention module, thereby reducing the time complexity of the self-attention layer. LW-FFN shrinks the complexity of multilayer perceptron. In addition, no parameter sharing of the detection head is utilized to compensate for the accuracy drop caused by the lightweight. Experiments verify that our method has the same AP and F1-score as DINO (a DETR-based detector), but our approach is lighter. Compared with DINO, the FLOPs are reduced by$113.4{G}$, the inference speed is increased by 42.4%, and the runtime memory usage is reduced by$5.9{G}$, which allows our method to be trained on low-end GPUs with more batch size, further improving the training efficiency. The code is available athttps://github.com/VDT-2048/SDD-DETR.Note to Practitioners—The motivation for this paper is to design a high-precision and high-inference speed visual detection method for the SDD of aero-engine blades. Most high-precision vision methods are based on the transformer framework. However, its high complexity and poor compatibility in deployment environments lead to slower detection speeds. Although the application object in this paper is aero-engine blades, it is also applicable in other fields of industry, such as rail detection, plate and strip steel detection, etc. However, the method proposed cannot be supported by deployment environments such as RKNN because of the deformable attention operator, so it takes a certain amount of time to be deployed and put into practical use. Currently, the frameworks of visual and language large models are based on transformers, consistent with the framework of our method, which makes extending our approach to large visual and multi-modal models more accessible.
Xiangkun Sun, Kechen Song, Xin Wen 0014, Yanyan Wang 0007, Yunhui Yan
IEEE Trans Autom. Sci. Eng.2
2025 CMAI-Det: Cross-Modal Alignment and Interaction for RGB-T Object Detection in Drone Scenes
abstract
RGB-T object detection is increasingly applied in drone surveillance, autonomous driving, and intelligent transportation due to its robustness in complex conditions. However, most existing methods assume well-aligned image pairs and often overlook misalignment between modalities, which arises from drone motion, viewpoint shifts, and sensor inconsistencies, thereby hindering detection accuracy. To address this limitation, we present CMAI-Det, a framework designed for object detection in unaligned RGB-T images. The approach employs a dual-stream extractor to strengthen the representational capacity of visible and thermal features. A modality-cooperative alignment module, using the thermal stream as reference, integrates multi-scale deformable convolutions and attention to align visible features. An adaptive fusion scheme is further introduced to balance modality contributions according to their reliability under varying conditions, enhancing feature robustness. Finally, a perception-driven deformable detection head improves discriminability by reinforcing spatial selectivity and structural adaptability, enabling precise modeling of diverse object appearances. Extensive experiments on the DVTOD dataset show that CMAI-Det surpasses 12 state-of-the-art RGB-T detectors in challenging drone scenarios, achieving an mAP of 86.1. It also performs strongly at standard thresholds, with AP50 of 90.0 and AP75 of 82.3, underscoring its robustness and effectiveness in complex detection tasks. The code is available at https://github.com/yinhaixu2000-coder/CMAI-Detection.
Xin Wen 0014, Haixu Yin, Wanying Nie, Jianxun Zhao, Kechen Song
IEEE Trans. Geosci. Remote. Sens.6
2025 SRPCNet: Self-Reinforcing Perception Coordination Network for Seamless Steel Pipes Internal Surface Defect Detection
abstract
Seamless steel pipes (SSPs) are vital material for industries. However, internal surface defects (ISDs) in SSPs are challenging to detect, and will significantly affect SSPs performance and lifespan. Existing detection methods are labor-intensive and have low visualization of detection results. Therefore, this article present a novel detection system comprising thePipelineAll-aspect internalSurface defectSpiral detecting robot and an interactive visualization software. After testing in the SSPs factory, the system achieves comprehensive, wireless and efficient detection and visualization for ISDs. In addition, we construct a dataset for ISDs in SSPs, named as SSP2000. The dataset contains 2000 images across nine defect categories, with many challenges in it. Furthermore, to accurately detect defects, we design the SRPCNet which can effectively address the challenges. Specifically, we first use the synergize perception augmentation module to enrich the feature space and to enhance the perception. Then, the hierarchical attention integrate module merges deep and shallow features using adaptive attention weights. Finally, the bilateral self-fusion module fully exploits intralayer features and produce prediction results. The proposed SRPCNet outperforms existing methods on eight evaluation metrics.
Hongshu Chen, Kechen Song, Yunhui Yan, Jun Li 0119
IEEE Trans. Ind. Informatics2
2024 Feature-based domain disentanglement and randomization: A generalized framework for rail surface defect segmentation in unseen scenarios
Kechen Song, Menghui Niu, Hongkun Tian, Yanyan Wang 0007, Yunhui Yan
Adv. Eng. Informatics2
2024 A novel multi-exposure fusion-induced stripe inpainting method for blade reflection-encoded images
Kechen Song, Chongyan Sun, Xin Wen 0014, Yunhui Yan
Adv. Eng. Informatics1
2024 Uncertainty inspired domain adaptation network for rail surface defect segmentation
Yunhui Yan, Kechen Song, Yanyan Wang 0007, Hongkun Tian, Jingbo Guo
Eng. Appl. Artif. Intell.3
2024 EAFNet: Extraction-amplification-fusion network for tiny cracks detection
Ziang Zhou, Wensong Zhao, Kechen Song, Yanyan Wang 0007, Jun Li 0119
Eng. Appl. Artif. Intell.3
2024 A visible-infrared clothes-changing dataset for person re-identification in natural scene
Xianbin Wei, Kechen Song, Wenkang Yang, Yunhui Yan, Qinggang Meng
Neurocomputing2
2024 A Rapid Screening Method for Suspected Defects in Steel Pipe Welds by Combining Correspondence Mechanism and Normalizing Flow
abstract
Nondestructive testing of welding images is still a significant challenge due to the imaging characteristics of radiographic images and the extremely random distribution of welding defect types. The practical application of supervised methods for welding nondestructive testing encounters significant challenges due to the limited availability of densely annotated samples and the absence of prior information regarding unknown defects. Hence, this article first proposes a rapid screening method using only defect-free image training to screen suspected defect images and normal images of welds in real time. Specifically, we first combine the correspondence mechanism and the representation mechanism, aiming to: 1) alleviate the smoothing reconstruction behavior caused by small defects and weak texture defects in weld; and 2) mitigate the offset learning behavior resulting from the differences between natural images and industrial weld images in the normalizing flow method. We propose a memory-aware transformer-based encoder, thus improving representations of complex defect-free images. Moreover, a dual-decoder strategy is introduced, which remaps the latent dependencies generated by the encoder through a semantic correspondence mechanism and reconstruction-guided normalizing flow, enabling effective learning of knowledge from weld images. We apply this framework to an industrial case of weld images, the experimental results demonstrate that our method outperforms other existing approaches.
Kechen Song, Yanyan Wang 0007, Guotong Lv, Yunhui Yan, Xingjie Li 0003
IEEE Trans. Ind. Informatics2
2024 MFANet: Multifeature Aggregation Network for Cross-Granularity Few-Shot Seamless Steel Tubes Surface Defect Segmentation
abstract
Defect segmentation on the inner surface of seamless steel tubes (SSTs) is a crucial technical means for evaluating product quality. However, both the category and quantity of defective samples of SSTs are sparse, limiting the generalization of traditional supervised learning and general few-shot defect segmentation (FSDS) methodologies. Moreover, the existing fine-grained segmentation method results in an arduous and time-consuming dataset-building process. Motivated by this, a novel defect segmentation paradigm called cross-granularity FSDS (CG-FSDS) is proposed. This paradigm aims to learn the defect segmentation capability on the coarse-grained labeled defect dataset and subsequently generalize it to segment fine-grained labeled defective samples of SSTs. The feasibility of CG-FSDS is evaluated by the proposed multifeature aggregation network (MFANet). To address the real challenge of defect segmentation in SSTs, we establish a cross-granularity benchmark called CGFSDS-9, which consists of six categories of inner surface defects in SSTs with fine-grained annotation and three categories of general metal surface defect samples with coarse-grained annotation. Our MFANet achieves superior results compared to other FSDS methods and showcases state-of-the-art performance on this benchmark.
Kechen Song, Hu Feng, Tonglei Cao, Yunhui Yan
IEEE Trans. Ind. Informatics1
2024 DCBFusion: an infrared and visible image fusion method through detail enhancement, contrast reserve and brightness balance
Shenghui Sun, Kechen Song, Yi Man, Hongwen Dong, Yunhui Yan
Vis. Comput.2
2023 SISG-Net: Simultaneous instance segmentation and grasp detection for robot grasp in clutter
Yunhui Yan, Ling Tong 0006, Kechen Song, Hongkun Tian, Yi Man, Wenkang Yang
Adv. Eng. Informatics3
2023 Modal complementary fusion network for RGB-T salient object detection
Kechen Song, Hongwen Dong, Hongkun Tian, Yunhui Yan
Appl. Intell.2
2023 RGB-T image analysis technology and application: A survey
Kechen Song, Ying Zhao 0040, Liming Huang, Yunhui Yan, Qinggang Meng
Eng. Appl. Artif. Intell.1
2023 Rotation adaptive grasping estimation network oriented to unknown objects based on novel RGB-D fusion strategy
Hongkun Tian, Kechen Song, Yunhui Yan
Eng. Appl. Artif. Intell.2
2023 Thermal images-aware guided early fusion network for cross-illumination RGB-T salient object detection
Han Wang 0048, Kechen Song, Liming Huang, Hongwei Wen, Yunhui Yan
Eng. Appl. Artif. Intell.2
2023 Data-driven robotic visual grasping detection for unknown objects: A problem-oriented review
Hongkun Tian, Kechen Song, Jing Xu 0016, Yunhui Yan
Expert Syst. Appl.2
2023 Antipodal-points-aware dual-decoding network for robotic visual grasp detection oriented to multi-object clutter scenes
Hongkun Tian, Kechen Song, Jing Xu 0016, Yunhui Yan
Expert Syst. Appl.2
2023 Informed anytime Bi-directional Fast Marching Tree for optimal motion planning in complex cluttered environments
Jing Xu 0016, Kechen Song, Yunhui Yan, Yihang Peng
Expert Syst. Appl.3
2023 Exploring the potential of Siamese network for RGBT object tracking
Liangliang Feng, Kechen Song, Junyi Wang 0003, Yunhui Yan
J. Vis. Commun. Image Represent.2
2023 MFS enhanced SAM: Achieving superior performance in bimodal few-shot segmentation
Ying Zhao 0040, Kechen Song, Yunhui Yan
J. Vis. Commun. Image Represent.2
2023 Cross-modality salient object detection network with universality and anti-interference
Hongwei Wen, Kechen Song, Liming Huang, Han Wang 0048, Yunhui Yan
Knowl. Based Syst.2
2023 Multiple Graph Affinity Interactive Network and a Variable Illumination Dataset for RGBT Image Salient Object Detection
abstract
Salient object detection (SOD) of images refers to simulating the attention mechanism of human vision to capture the most attractive objects in an image. Current SOD mainly relies on RGB images captured by optical cameras. However, existing optical cameras are not comparable to the human visual system, especially in poorly illumination scenes. Our human visual system is able to resolve scenes well in low light conditions, while optical cameras can barely image without enough illumination. To make machine vision closer to the imaging of the human eye, we propose to use thermal infrared (T) images to compensate RGB images and build a variable illumination RGBT dataset named VI-RGBT1500 for SOD. This dataset is collected under three different illumination conditions including sufficient illumination, uneven illumination and insufficient illumination to fully demonstrate the superiority of the RGBT image combination. Furthermore, we propose a multiple graph affinity interactive (MGAI) network to validate the proposed dataset. Our network structure is simple using only the MGAI to fuse the features of different modalities. Meanwhile, the MGAI model highlights valuable information during the interaction, which facilitates feature representation under variable illumination. The proposed VI-RGBT1500 dataset and three publicly available RGBT SOD datasets are used for the comparison experiments, and the results with the state-of-the-art methods prove that our VI-RGBT1500 dataset is valuable and the performance of the MGAI network is competitive. The VI-RGBT1500 dataset and the MGAI network are available at:https://github.com/huanglm-me/VI-RGBT1500.
Kechen Song, Liming Huang, Aojun Gong, Yunhui Yan
IEEE Trans. Circuits Syst. Video Technol.1
2023 Modality Registration and Object Search Framework for UAV-Based Unregistered RGB-T Image Salient Object Detection
abstract
UAVs are widely used in various industries, and various visual tasks under the perspective of the UAV have been widely studied. In particular, the RGB-T detection method based on UAVs has shown significant advantages. However, existing RGB-T methods are designed based on registration image pairs rather than detecting images directly acquired by UAVs. This detection process is limited by the accuracy of image registration. And image registration wastes a lot of time. To solve the above problems, we construct an unregistered RGB-T image salient object detection (SOD) dataset under the UAV perspective, known as UAV RGB-T 2400. The dataset includes many challenging scenes, and the images are not manually registered. Further, we construct a modality registration and object search (MROS) framework for unregistered RGB-T SOD. Firstly, a modality registration scheme is proposed to solve the unregistration problem of modal features. We successively perform pixel-level registration from a local perspective and semantic-level registration from a global perspective for different modal features. And we carry out the channel and spatial interaction for the different modal features in modality registration. Aiming at the interference problem in the UAV detection environment, we propose an object search scheme. The two high-level features are used to search the object location, and the three low-level features are used to refine the object and produce prediction results. Experimental results on the UAV RGB-T 2400 dataset show that MROS is effective compared with state-of-the-art methods. The code is available at: https://github.com/VDT-2048/UAV-RGB-T-2400.
Kechen Song, Hongwei Wen, Xiaotong Xue, Liming Huang, Yingying Ji, Yunhui Yan
IEEE Trans. Geosci. Remote. Sens.1
2023 Shape-Consistent One-Shot Unsupervised Domain Adaptation for Rail Surface Defect Segmentation
abstract
Deep neural networks have greatly improved the performance of rail surface defect segmentation when the test samples have the same distribution as the training samples. However, in practical inspection scenarios, the rail surface exhibits variations in appearance due to different service times and natural conditions. Conventional deep learning models show limited generalization in scenes with distribution differences. To address this problem, we propose a novel one-shot unsupervised domain adaptation framework. Specifically, we introduce a shape-consistent style transfer module that performs pixel-level distribution alignment between the training and test images. Based on the one-shot test image, the training image is reconstructed to have the same appearance as the test image. Meanwhile, we employ a multitask learning strategy to prevent content distortion of the reconstructed images. To improve the robustness of the model to distribution differences, we design an edge-aware defect segmentation model and train the model using the reconstructed training images. The experimental results show that our method effectively improves the robustness of the model to distribution differences and achieves satisfying results in the task of rail surface defect segmentation.
Kechen Song, Menghui Niu, Hongkun Tian, Yanyan Wang 0007, Yunhui Yan
IEEE Trans. Ind. Informatics2
2023 Normal-Knowledge-Based Pavement Defect Segmentation Using Relevance-Aware and Cross-Reasoning Mechanisms
abstract
Automatic pavement defect segmentation is a big challenge because of the class diversity and extremely random distribution of defects. Most existing approaches focus on supervised strategies to achieve decent performance. Due to the difficulty of getting massive densely annotated samples and the limited prior knowledge of potential defects, these methods have significant bottlenecks in the actual pavement settings. This paper proposes a relevance-aware and cross-reasoning network (RCN) for anomaly segmentation of pavement defects, which can segment defects using merely non-defective images for training. A relevance-aware transformer-based encoder is first devised to model intrinsic interdependencies across local features, thus improving representations of complex non-defective images. Next, a dual decoder strategy is proposed to remap the encoder-generated latent dependencies at the local semantic and global detailed levels, respectively. Specifically, a cross-reasoning refinement module is built in the local decoder to reason the cross-relationship between spatial and channel dimensions. Finally, a context-aware abnormal distillation measurement is developed to evaluate the semantic reconstruction deviations during the inference. Under the guidance of semantic affinity, this measurement allows our model to highlight defective areas adaptively. Extensive experimental results on four datasets indicate that RCN outperforms other leading anomaly segmentation methods.
Yanyan Wang 0007, Menghui Niu, Kechen Song, Peng Jiang 0021, Yunhui Yan
IEEE Trans. Intell. Transp. Syst.3
2022 Unidirectional RGB-T salient object detection with intertwined driving of encoding and fusion
Jie Wang 0095, Kechen Song, Yanqi Bao, Yunhui Yan, Yahong Han
Eng. Appl. Artif. Intell.2
2022 Multi-Graph Fusion and Learning for RGBT Image Saliency Detection
abstract
RGB and thermal infrared (RGBT) image saliency detection is a relatively new direction in the field of computer vision. Combining the advantages of RGB images and T images can significantly improve detection performance. Currently, there are only a few methods to work on RGBT saliency detection, and the number of image samples cannot meet the training requirements for deep learning, so it remains valuable to propose an effective unsupervised method. In this paper, we present an unsupervised RGBT saliency detection method based on multi-graph fusion and learning. Firstly, RGB images and T images are adaptively fused based on boundary information to produce more accurate superpixels. Next, a multi-graph fusion model is proposed to selectively learn useful information from multi-modal images. Finally, we implement the theory of finding good neighbors in the graph affinity and propose different algorithms for two stages of saliency ranking. Experimental results on three RGBT datasets show that the proposed method is effective compared with the state-of-the-art algorithms.
Liming Huang, Kechen Song, Jie Wang 0095, Menghui Niu, Yunhui Yan
IEEE Trans. Circuits Syst. Video Technol.2
2022 CGFNet: Cross-Guided Fusion Network for RGB-T Salient Object Detection
abstract
RGB salient object detection (SOD) has made great progress. However, the performance of this single-modal salient object detection will be significantly decreased when encountering some challenging scenes, such as low light or darkness. To deal with the above challenges, thermal infrared (T) image is introduced into the salient object detection. This fused modal is called RGB-T salient object detection. To achieve deep mining of the unique characteristics of single modal and the full integration of cross-modality information, a novel Cross-Guided Fusion Network (CGFNet) for RGB-T salient object detection is proposed. Specifically, a Cross-Scale Alternate Guiding Fusion (CSAGF) module is proposed to mine the high-level semantic information and provide global context support. Subsequently, we design a Guidance Fusion Module (GFM) to achieve sufficient cross-modality fusion by using single modal as the main guidance and the other modal as auxiliary. Finally, the Cross-Guided Fusion Module (CGFM) is presented and serves as the main decoding block. And each decoding block is consists of two parts with two modalities information of each being the main guidance, i.e., cross-shared Cross-Level Enhancement (CLE) and Global Auxiliary Enhancement (GAE). The main difference between the two parts is that the GFM using different modalities as the main guide. The comprehensive experimental results prove that our method achieves better performance than the state-of-the-art salient detection methods. The source code has released at:https://github.com/wangjie0825/CGFNet.git.
Jie Wang 0095, Kechen Song, Yanqi Bao, Liming Huang, Yunhui Yan
IEEE Trans. Circuits Syst. Video Technol.2
2022 Deep Metric Learning-Based for Multi-Target Few-Shot Pavement Distress Classification
abstract
Pavement distress detection is of great significance for road maintenance and to ensure road safety. At present, detection methods based on deep learning have achieved outstanding performance in related fields. However, these methods require large-scale training samples. For pavement distress detection, it is difficult to collect more images with pavement distress, and the types of pavement diseases are increasing with time, so it is impossible to ensure sufficient pavement distress samples to train the supervised deep model. In this article, we propose a new few-shot pavement distress detection method based on metric learning, which can effectively learn new categories from a few labeled samples. In this article, we adopt the backend network (ResNet18) to extract multilevel feature information from the base classes and then send the extracted features into the metric module. In the metric module, we introduce the attention mechanism to learn the feature attributes of “what” and “where” and focus the model on the desired characteristics. We also introduce a new metric loss function to maximize the distance between different categories while minimizing the distance between the same categories. In the testing stage, we calculate the cosine similarity between the support set and query set to complete novel category detection. The experimental results show that the proposed method significantly outperforms several benchmarking methods on the pavement distress dataset (the classification accuracies of 5-way 1-shot and 5-way 5-shot are 77.20% and 87.28%, respectively).
Hongwen Dong, Kechen Song, Qi Wang 0054, Yunhui Yan, Peng Jiang 0021
IEEE Trans. Ind. Informatics2
2022 Automatic Inspection and Evaluation System for Pavement Distress
abstract
Pavement distress detection is of significance for road maintenance and traffic safety. Manual pavement distress detection suffers from high workloads, inefficiency, low accuracy, and high cost. To replace the manual operations in the pre-filling detection with the aim to improve efficiency and reduce cost, this paper proposes a three-stage automatic inspection and evaluation system for pavement distress based on improved deep convolutional neural networks (CNNs). First, the system integrates multi-level context information from the CNN classification model to construct discriminative super-features to determine whether there is distress in the pavement image and the type of the distress, so as to achieve rapid detection of pavement distress. Then, the pavement images with distress are fed into the CNN segmentation model to highlight the distress region with pixel-wise. In the segmentation model, a novel pyramid feature extraction module and a novel guidance attention mechanism are introduced. Finally, we evaluate the degree of pavement damage according to the segmentation results of the CNN segmentation model. In the experiments, we compare our classification model and segmentation model with other state-of-the-art methods on two pavement distress datasets, and the results demonstrate that the proposed models achieve out-performance on different evaluation metrics.
Hongwen Dong, Kechen Song, Yanyan Wang 0007, Yunhui Yan, Peng Jiang 0021
IEEE Trans. Intell. Transp. Syst.2
2021 Visible and thermal images fusion architecture for few-shot semantic segmentation
Yanqi Bao, Kechen Song, Jie Wang 0095, Liming Huang, Hongwen Dong, Yunhui Yan
J. Vis. Commun. Image Represent.2
2021 Unsupervised Saliency Detection of Rail Surface Defects Using Stereoscopic Images
abstract
Visual information is increasingly recognized as a useful method to detect rail surface defects due to its high efficiency and stability. However, it cannot sufficiently detect a complete defect in the complex background information. The addition of surface profiles can effectively improve this by including a 3-D information of defects. However, in high-speed detection, the traditional 3-D profile acquisition is difficult and separate from the image acquisition, which cannot satisfy the above-mentioned requirements effectively. Therefore, an unsupervised stereoscopic saliency detection method based on a binocular line-scanning system is proposed in this article. This method can simultaneously obtain a highly precise image as well as profile information while also avoids the decoding distortion of the structured light reconstruction method. In our method, a global low-rank nonnegative reconstruction algorithm with a background constraint is proposed. Unlike the low-rank recovery model, the algorithm has a more comprehensive low rank and background clustering properties. Furthermore, outlier detection based on the geometric properties of the rail surface is also proposed in this method. Finally, the image saliency results and depth outlier detection results are associated with the collaborative fusion, and a dataset (RSDDS-113) containing the rail surface defects is established for the experimental verification. The experimental results demonstrate that our method can obtain a mean absolute error of 0.09 and area under the ROC curve of 0.94, better than 15 state-of-the-art algorithms.
Menghui Niu, Kechen Song, Liming Huang, Qi Wang 0054, Yunhui Yan, Qinggang Meng
IEEE Trans. Ind. Informatics2
2021 Two Deep Learning Networks for Rail Surface Defect Inspection of Limited Samples With Line-Level Label
abstract
Rail surface defect (RSD) inspection is an essential routine maintenance task. Computer vision testing is suitable for RSD inspection with its intuitiveness and rapidity. Deep learning techniques, which can extract deep semantic features, have been applied to inspect RSDs in recent years. However, these methods demand thousands of samples. And sample collection requires hard-working and costs high. To address the issue, a novel inspection scheme for RSDs is presented for limited samples with a line-level label, which regards defect images as sequence data and classifies pixel lines. Thousands of pixel lines are easy to be collected and labeling line-level is a simple task in labeling works. Then two methods OC-IAN and OC-TD are designed for inspecting express rail defects and common/heavy rail defects, respectively. OC-IAN and OC-TD both employ one-dimensional convolutional neural network (ODCNN) to extract features and long- and short-term memory (LSTM) network to extract context information. The main differences between OC-IAN and OC-TD are that OC-TD applies a double-branch structure and removes the attention module. Experimental results on RSDDs dataset demonstrate that our methods are effective and outperform the state-of-the-art methods on defect-level metrics (Type-I: Rec-0.9314, Pre-0.8421, F1-0.8845; Type-II: Rec-0.9427, Pre-0.9176, F1-0.9300).
Kechen Song, Qi Wang 0054, Yu He 0004, Xin Wen 0014, Yunhui Yan
IEEE Trans. Ind. Informatics2
2020 A batch informed sampling-based algorithm for fast anytime asymptotically-optimal motion planning in cluttered environments
Jing Xu 0016, Kechen Song, Hongwen Dong, Yunhui Yan
Expert Syst. Appl.2
2020 Learning discriminative update adaptive spatial-temporal regularized correlation filter for RGB-T tracking
Mingzheng Feng, Kechen Song, Yanyan Wang 0007, Jie Liu 0043, Yunhui Yan
J. Vis. Commun. Image Represent.2
2020 NERNet: Noise estimation and removal network for image denoising
Bingyang Guo, Kechen Song, Hongwen Dong, Yunhui Yan, Zhibiao Tu, Liu Zhu
J. Vis. Commun. Image Represent.2
2020 Learning object-centric complementary features for zero-shot learning
Jie Liu 0043, Kechen Song, Yu He 0004, Hongwen Dong, Yunhui Yan, Qinggang Meng
Signal Process. Image Commun.2
2020 RGB-T Saliency Detection via Low-Rank Tensor Learning and Unified Collaborative Ranking
abstract
Saliency detection is a significant research topic in the field of image processing and computer vision. Currently, most saliency detection methods are applied to RGB images, so that they may encounter adverse scenarios characterized by complex background, inclement weather, and low illumination. Fusing complementary advantages of RGB and thermal infrared (RGB-T) images can effectively boost saliency detection performance. Therefore, we propose a novel RGB-T saliency detection method in this letter. To this end, we first regard superpixels as graph nodes and calculate the affinity matrix for each feature. Then, we propose a low-rank tensor learning model for the graph affinity, which can suppress redundant information and improve the relevance of similar image regions. Finally, a novel ranking algorithm is proposed to jointly obtain the optimal affinity matrix and saliency values under a unified structure. Test results on two RGB-T datasets illustrate the proposed method performs well when against the state-of-the-art algorithms.
Liming Huang, Kechen Song, Aojun Gong, Yunhui Yan
IEEE Signal Process. Lett.2
2020 PGA-Net: Pyramid Feature Fusion and Global Context Attention Network for Automated Surface Defect Detection
abstract
Surface defect detection is a critical task in industrial production process. Nowadays, there are lots of detection methods based on computer vision and have been successfully applied in industry, they also achieved good results. However, achieving full automation of surface defect detection remains a challenge, due to the complexity of surface defect, in intraclass. While the defects between interclass contain similar parts, there are large differences in appearance of the defects. To address these issues, this article proposes a pyramid feature fusion and global context attention network for pixel-wise detection of surface defect, called PGA-Net. In the framework, the multiscale features are extracted at first from backbone network. Then the pyramid feature fusion module is used to fuse these features into five resolutions through some efficient dense skip connections. Finally, the global context attention module is applied to the fusion feature maps of adjacent resolution, which allows effective information propagate from low-resolution fusion feature maps to high-resolution fusion ones. In addition, the boundary refinement block is added to the framework to refine the boundary of defect and improve the result of the prediction. The final prediction is the fusion of the five resolutions fusion feature maps. The results of evaluation on four real-world defect datasets demonstrate that the proposed method outperforms the state-of-the-art methods on mean intersection of union and mean pixel accuracy (NEU-Seg: 82.15%, DAGM 2007: 74.78%, MT_defect: 71.31%, Road_defect: 79.54%).
Hongwen Dong, Kechen Song, Yu He 0004, Jing Xu 0016, Yunhui Yan, Qinggang Meng
IEEE Trans. Ind. Informatics2
2016 Noise robust image matching using adjacent evaluation census transform and wavelet edge joint bilateral filter in stereo vision
Kechen Song, Xin Wen 0014, Yunhui Yan
J. Vis. Commun. Image Represent.1
2015 Adjacent evaluation of local binary pattern for texture classification
Kechen Song, Yunhui Yan
J. Vis. Commun. Image Represent.1