Yang Zhang 0053

dblp:06/6785-53 · DBLP profile ↗
← Back
29ranked-venue papers
12as first author
25since 2021 · last 2026
0000-0002-4170-4798ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Prompt-Driven Lightweight Foundation Model for Instance Segmentation-Based Fault Detection in Freight Trains
abstract
Accurate visual fault detection in freight trains remains a critical challenge for intelligent transportation system maintenance, due to complex operational environments, structurally repetitive components, and frequent occlusions or contaminations in safety-critical regions. Conventional instance segmentation methods based on convolutional neural networks and Transformers often suffer from poor generalization and limited boundary accuracy under such conditions. To address these challenges, we propose a lightweight self-prompted instance segmentation framework tailored for freight train fault detection. Our method leverages the Segment Anything Model by introducing a self-prompt generation module that automatically produces task-specific prompts, enabling effective knowledge transfer from foundation models to domain-specific inspection tasks. In addition, we adopt a Tiny Vision Transformer backbone to reduce computational cost, making the framework suitable for real-time deployment on edge devices in railway monitoring systems. We construct a domain-specific dataset collected from real-world freight inspection stations and conduct extensive evaluations. Experimental results show that our method achieves$74.6~AP^{\text {box}}$and$74.2~AP^{\text {mask}}$on the dataset, outperforming existing state-of-the-art methods in both accuracy and robustness while maintaining low computational overhead. This work offers a deployable and efficient vision solution for automated freight train inspection, demonstrating the potential of foundation model adaptation in industrial-scale fault diagnosis scenarios. Project page:https://github.com/MVME-HBUT/SAM_FTI-FDet
Guodong Sun 0002, Qihang Liang, Xingyu Pan, Moyun Liu, Yang Zhang 0053
IEEE Trans. Intell. Transp. Syst.5
2026 Teacher-Student Instance-Level Adversarial Augmentation for Single Domain Generalized Medical Image Segmentation
abstract
Recently, single-source domain generalization (SDG) has gained popularity in medical image segmentation. As a prominent technique, adversarial image augmentation technique can generate synthetic training data that are challenging for the segmentation model to recognize. To avoid the over-augmentation problem, existing adversarial-based works often employ augmenters with relatively simple structures for medical images, typically operating at the image level, limiting the diversity of the augmented images. In this paper, we propose a Teacher-Student Instance-level Adversarial Augmentation (TSIAA) model for generalized medical image segmentation. The objective of TSIAA is to derive domain-generalizable representations by exploring out-of-source data distributions. First, we construct an Instance-level Image Augmenter (IIAG) using several Instance-level Augmentation Modules (IAMs), which are based on the learnable constrained Bèzier transformation function. Compared to image-level adversarial augmentation, instance-level adversarial augmentation breaks the uniformity of augmentation rules across different structures within an image, thereby providing greater diversity. Then, TSIAA conducts Teacher-Student (TS) learning through an adversarial approach, alternating novel image augmentation and generalized representation learning. The former delves into out-of-source and plausible data, while the latter continuously updates both the student and teacher to ensure the original and augmented features maintain consistent and generalized characteristics. By integrating both strategies, our proposed TSIAA model achieves significant improvements over state-of-the-art methods in four challenging SDG tasks. The code can be accessed at https://github.com/Wangzs0228/TSIAA.
Zhengshan Wang, Long Chen 0001, Xuelin Xie, Yang Zhang 0053, Yunpeng Cai, Weiping Ding 0001
IEEE Trans. Medical Imaging4
2025 Variable-Stiffness Nasotracheal Intubation Robot with Passive Buffering: A Modular Platform in Mannequin Studies
Ruoyi Hao, Jiewen Lai, Wenqi Zhong, Dihong Xie, Yang Zhang 0053, Catherine Po Ling Chan, Jason Ying-Kuen Chan, Hongliang Ren 0001
ICRA7
2025 3D reconstruction and precision evaluation of industrial components via Gaussian Splatting
Guodong Sun 0002, Dingjie Liu, Shaoran An, Yang Zhang 0053
Comput. Graph.5
2025 Efficient RGB-D scene understanding via multi-task adaptive learning and cross-dimensional feature guidance
Guodong Sun 0002, Gaoyang Zhang, Yang Zhang 0053
Knowl. Based Syst.5
2025 V²-SfMLearner: Learning Monocular Depth and Ego-Motion for Multimodal Wireless Capsule Endoscopy
abstract
Deep learning can predict depth maps and capsule ego-motion from capsule endoscopy videos, aiding in 3D scene reconstruction and lesion localization. However, the collisions of the capsule endoscopies within the gastrointestinal tract cause vibration perturbations in the training data. Existing solutions focus solely on vision-based processing, neglecting other auxiliary signals like vibrations that could reduce noise and improve performance. Therefore, we propose V2-SfMLearner, a multimodal approach integrating vibration signals into vision-based depth and capsule motion estimation for monocular capsule endoscopy. We construct a multimodal capsule endoscopy dataset containing vibration and visual signals, and our artificial intelligence solution develops an unsupervised method using vision-vibration signals, effectively eliminating vibration perturbations through multimodal learning. Specifically, we carefully design a vibration network branch and a Fourier fusion module, to detect and mitigate vibration noises. The fusion framework is compatible with popular vision-only algorithms. Extensive validation on the multimodal dataset demonstrates superior performance and robustness against vision-only algorithms. Without the need for large external equipment, our V2-SfMLearner has the potential for integration into clinical capsule robots, providing real-time and dependable digestive examination tools. The findings show promise for practical implementation in clinical settings, enhancing the diagnostic capabilities of doctors. Note to Practitioners—This paper is motivated by the problem of estimating the depth and ego-motion information for the wireless capsule endoscopy in the human gastrointestinal tract to realize accurate, efficient, robust, and real-time inspection. Our estimation method does not engage any external localization equipment. Instead, inspired by the existing research on integrating capsule endoscopy and inertial measurement units, we introduce vibration signals into vision-based depth and ego-motion estimation approaches, improving the accuracy and robustness of the estimation results based on multimodal learning methods. Research on capsule robots or computer vision can readily be combined with our framework for various clinical and industrial applications.
Long Bai 0008, Beilei Cui, Yanheng Li 0002, Shilong Yao, Sishen Yuan, Yanan Wu 0003, Yang Zhang 0053, Max Q.-H. Meng, Zhen Li 0026, Weiping Ding 0001, Hongliang Ren 0001
IEEE Trans Autom. Sci. Eng.8
2025 Learning Dynamic-Sensitivity Enhanced Correlation Filter With Adaptive Second-Order Difference Spatial Regularization for UAV Tracking
abstract
Discriminative correlation filter (DCF)-based tracking algorithms continue to advance in the field of UAV tracking due to their computational efficiency. The idea of integrating the advantages of historical information and response adjustments into the CF tracking framework is continuously being developed. However, maintaining the stability of mobile video tracking in highly dynamic environments is extremely challenging. This difficulty arises from frequent changes in targets and backgrounds, as well as the stochastic noise generated by the photon-counting process in sensors. In addition, the inconsistent rates of these changes are often overlooked and require further scrutiny. In this paper, we propose a dynamic sensitivity enhanced correlation filter with adaptive second-order difference spatial regularization to address the issue of inconsistent motion rates in dynamic videos. We use the non-local means algorithm to denoise template images before feature extraction, improving the discriminative power of target contours. Then, we incorporate the proposed dynamic-sensitivity error method into CF learning and employ a novel adaptive second-order difference spatial regularization to simultaneously optimize the filter coefficients and spatial regularization weights. This regularization effectively works in synergy with the dynamic-sensitivity error strategy. Furthermore, an additional ADMM optimizer is introduced to derive the solution, thereby improving the convergence and computational efficiency of the algorithm. This algorithm supports the adjustment of filter updates in dynamic environments by balancing consistency with previous filter templates and flexibility to accommodate rapid target changes. By conducting extensive experiments on three challenging UAV tracking databases, we compare the proposed model with existing models. The experimental results demonstrate our superior performance. Code is released at:https://github.com/Johnsonirene/LDECF.
Yu-Feng Yu 0001, Zhongsen Chen, Yang Zhang 0053, Chuanbin Zhang, Weiping Ding 0001
IEEE Trans. Intell. Transp. Syst.3
2025 An Adaptive Fuzzy Rough Neural Network and Its Application in Classification
abstract
Fuzzy rough set theory is an important approach for analyzing data uncertainty. However, the model lacks adaptive learning capabilities and cannot fit labeled data effectively in classification tasks. This study aims to introduce an adaptive learning mechanism into fuzzy rough set theory to enhance its data-fitting capability. To this end, this study seamlessly integrates fuzzy rough set theory with neural networks and proposes a novel fuzzy rough neural network model. This model adaptively learns fuzzy similarity relations in rough set models using the backpropagation algorithm. The proposed network model comprises five layers: the input, membership, fuzzy lower approximation, fully connected, and output layers. The fuzzy similarity relations between the input samples and training samples are computed in the membership layer. These relations are utilized in the fuzzy lower approximation layer to describe the degree to which the samples belong to different classes. The fuzzy rough lower approximations of the input samples are finally fused in the fully connected layer using feature weight coefficients. In the backpropagation stage, the gradient of the objective function is used to correct the fuzzy similarity relations and feature weight coefficients. This study theoretically proved that the proposed fuzzy rough network has a generalized function approximation property and can approximate any decision function. Experimental analysis showed that the proposed method is effective and performs better than most of the existing state-of-the-art algorithms.
Changzhong Wang, Yang Zhang 0053, Shuang An, Weiping Ding 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2024 FS-OreDet: Feature enhancement and relationship exploration for boosting few-shot object detector of ore images
Guodong Sun 0002, Yuting Peng 0001, Chengming Xu 0001, Yanwei Fu 0001, Yang Zhang 0053
Eng. Appl. Artif. Intell.8
2024 Efficient segmentation with texture in ore images based on box-supervised approach
Guodong Sun 0002, Delong Huang, Yuting Peng 0001, Yang Zhang 0053
Eng. Appl. Artif. Intell.6
2024 Multiple prior representation learning for self-supervised monocular depth estimation via hybrid transformer
Guodong Sun 0002, Mingxuan Liu 0002, Moyun Liu, Yang Zhang 0053
Eng. Appl. Artif. Intell.5
2024 A concise but high-performing network for image guided depth completion in autonomous driving
Moyun Liu, Youping Chen, Jingming Xie, Yang Zhang 0053, Joey Tianyi Zhou
Knowl. Based Syst.6
2024 Adaptive Relative Fuzzy Rough Learning for Classification
abstract
Fuzzy rough set theory offers a valuable approach for addressing data uncertainty; however, existing fuzzy rough sets often failed to capture data distribution in practical applications due to their limited adaptive learning capabilities. This article aims to enhance the data fitting ability of fuzzy rough set theory by incorporating an adaptive learning mechanism. Consequently, this study introduces a relative fuzzy rough set model with adaptive learning (ALRFRS), where the adaptive learning mechanism considers not only feature weights but also class variances. To describe the similarity between samples in data regions with significant differences in class density, this study introduces the concepts of relative distance and relative fuzzy similarity relation. Subsequently, the study defines relative fuzzy rough approximation operators for each feature and combines them based on varying feature weights to establish a relative fuzzy rough set model. In addition, the study conducts an analysis of some fundamental properties of the model. To enable adaptive fuzzy rough learning, the study formulates an objective function that incorporates feature weights and class variances and provides the corresponding optimization learning algorithm. Experimental results demonstrate that the adaptive learning mechanism can enhance the classification accuracy of relative fuzzy rough models, surpassing the performance of most existing excellent algorithms.
Yang Zhang 0053, Changzhong Wang, Yang Huang 0009, Weiping Ding 0001
IEEE Trans. Fuzzy Syst.1
2024 An Efficient MLP-Based Point-Guided Segmentation Network for Ore Images With Ambiguous Boundary
abstract
The precise segmentation of ore images is critical to the successful execution of the beneficiation process. Due to the homogeneous appearance of the ores, which leads to low contrast and unclear boundaries, accurate segmentation becomes challenging, and recognition becomes problematic. This article proposes a lightweight framework based on multilayer perceptron (MLP), which focuses on solving the problem of edge blurring. Specifically, we introduce a lightweight backbone better suited for efficiently extracting low-level features. Besides, we design a feature pyramid network consisting of two MLP structures that balance local and global information, thus, enhancing detection accuracy. Furthermore, we propose a novel loss function that guides the prediction points to match the instance edge points to achieve clear object boundaries. We have conducted extensive experiments to validate the efficacy of our proposed method. Our approach achieves a remarkable processing speed of over 27 frames per second with a model size of only 73 MB. Moreover, our method delivers a consistently high level of accuracy, with impressive performance scores of 60.4 and 48.9 in$AP_{50}^{\text{box}}$and$AP_{50}^{\text{mask}}$, respectively, as compared with the currently available state-of-the-art techniques, when tested on the ORE image dataset.
Guodong Sun 0002, Yuting Peng 0001, Mengya Xu, An Wang 0007, Hongliang Ren 0001, Yang Zhang 0053
IEEE Trans. Ind. Informatics8
2024 Efficient Visual Fault Detection for Freight Train via Neural Architecture Search With Data Volume Robustness
abstract
Deep learning-based fault detection methods have achieved significant success. In visual fault detection of freight trains, there exists a large characteristic difference between interclass components (scale variance) but intraclass on the contrary, which entails scale-awareness for detectors. Moreover, the design of task-specific networks heavily relies on human expertise. As a consequence, neural architecture search (NAS) that automates the model design process gains considerable attention because of its promising performance. However, NAS is computationally intensive due to the large search space and huge data volume. In this work, we propose an efficient NAS-based framework for visual fault detection of freight trains to search for the task-specific detection head with capacities of multiscale representation. First, we design a scale-aware search space for discovering an effective receptive field in the head. Second, we explore the robustness of data volume to reduce search costs based on the specifically designed search space, and a novel sharing strategy is proposed to reduce memory and further improve search efficiency. Extensive experimental results demonstrate the effectiveness of our method with data volume robustness, which achieves 46.8 and 47.9 mAP on the bottom view and side view datasets, respectively. Our framework outperforms the state-of-the-art approaches and linearly decreases the search costs with reduced data volumes.
Yang Zhang 0053, Mingying Li, Huilin Pan, Moyun Liu
IEEE Trans. Ind. Informatics1
2024 MENet: Multi-Modal Mapping Enhancement Network for 3D Object Detection in Autonomous Driving
abstract
To achieve more accurate perception performance, LiDAR and camera are gradually chosen to improve 3D object detection simultaneously. However, it is still a non-trivial task to build an effective fusion mechanism, and this is hindering the development of multi-modal based method. Especially, the mapping relationship construction between two modalities is far from fully explored. Canonical cross-modal mapping suffers from failure when the calibration matrix is incorrect, and it also greatly wastes the amount and density of RGB image information. This paper aims to extend the traditional one-to-one alignment relationship between LiDAR and camera. For all projected point clouds, we enhance their cross-modal mapping relationship through aggregating color-texture related feature and shape-contour related feature. Further, a mapping pyramid is proposed to leverage the semantic representation of the image feature at different stages. Based on the above mapping enhancement strategies, our method increases the engagement rate of image. Finally, we design a fusion module based on an attention mechanism to improve the point cloud feature with the auxiliary image feature. Extensive experiments on the KITTI dataset and SUN-RGBD dataset show that our model achieves satisfactory 3D object detection, especially for categories with sparse point clouds compared with other multi-modal fusion networks.
Moyun Liu, Youping Chen, Jingming Xie, Yang Zhang 0053, Zhenshan Bing, Genghang Zhuang, Kai Huang 0001, Joey Tianyi Zhou
IEEE Trans. Intell. Transp. Syst.5
2023 S2ME: Spatial-Spectral Mutual Teaching and Ensemble Learning for Scribble-Supervised Polyp Segmentation
An Wang 0007, Mengya Xu, Yang Zhang 0053, Mobarakol Islam, Hongliang Ren 0001
MICCAI (1)3
2023 Efficient visual fault detection for freight train braking system via heterogeneous self distillation in the wild
Yang Zhang 0053, Huilin Pan, Mingying Li, Guodong Sun 0002
Adv. Eng. Informatics1
2023 Channel Attentional Correlation Filters Learning With Second-Order Difference for UAV Tracking
abstract
Unmanned aerial vehicle (UAV) visual tracking has been a hot research topic in the field of remote sensing. Many filter-based UAV trackers have achieved excellent performance. However, existing methods do not distinguish the importance of different feature channels with semantic information and background information, which may hinder the tracker’s ability to adapt to changing environments. To deal with this problem, we propose a channel attentional correlation filters learning model (CACF). Specifically, we introduce the fuzzy C-means algorithm to pre-classify the extracted features and then perform weight penalty to feature channels with different membership degrees. In addition, the filter can adapt more effectively to the background’s rapid changes during the UAV tracking process by learning the second-order difference between adjacent three frame features. Finally, the comparative experiments are conducted on three mainstream UAV datasets, including DTB70, UAV123@10fps, and UAVDT. The experimental results demonstrate the effectiveness of the proposed method. The tracking performance of CACF surpasses that of other state-of-the-art trackers.
Yang Zhang 0053, Yu-Feng Yu 0001, Ke-Kun Huang, Yingxu Wang 0002
IEEE Geosci. Remote. Sens. Lett.1
2023 Faster OreFSDet: A lightweight and effective few-shot object detector for ore images
Yang Zhang 0053, Yuting Peng 0001, Chengming Xu 0001, Yanwei Fu 0001, Guodong Sun 0002
Pattern Recognit.1
2023 Adaptive fusion affinity graph with noise-free online low-rank representation for natural image segmentation
Yang Zhang 0053, Moyun Liu, Guodong Sun 0002, Jingwu He
Pattern Recognit.1
2023 Robust Correlation Filter Learning With Continuously Weighted Dynamic Response for UAV Visual Tracking
abstract
Unmanned Aerial Vehicles (UAV) visual tracking has always been a challenging task. Existing correlation filter tracking algorithms typically utilize the Histograms of Oriented Gradients (HOG) and Color Names (CN) method to directly incorporate the extracted target features into the model updating process. However, in low-resolution video quality, it leads to unstable target feature values. To address this limitation, we propose a novel preprocessing technique involving Gaussian denoising. This preprocessing step is designed to enhance the stability of the target’s feature values and make the target’s scale information clearer, thereby improving the tracker’s recognition capability for the target and effectively reducing noise interference. Furthermore, in contrast to other UAV trackers that rely on a singular representation of contextual information, this paper aims to enhance the utilization of historical information. Therefore, we introduce a context-based approach that integrates continuously weighted dynamic response maps from both temporal and spatial perspectives. Our tracker has the ability to adapt to rapid environmental changes during the tracking process while simultaneously reducing the potential risks of model overfitting and distortion. Extensive experiments are conducted on authoritative datasets, including DTB70, UAV123@10fps, and UAVDT, comparing our model against other advanced trackers. The experimental results validate the superior tracking performance and robustness of our tracker.
Yang Zhang 0053, Yu-Feng Yu 0001, Long Chen 0001, Weiping Ding 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Visual Fault Detection of Multiscale Key Components in Freight Trains
abstract
Fault detection for key components in the braking system of freight trains is critical for ensuring railway transportation safety. Despite the frequently employed methods based on deep learning, these fault detectors are extremely reliant on hardware resources and complex to implement. In addition, no train fault detectors consider the drop in accuracy induced by scale variation of fault parts. This article proposes a lightweight anchor-free framework to solve the above problems. Specifically, to reduce the amount of computation and model size, we introduce a lightweight backbone and adopt an anchor-free method for localization and regression. To improve detection accuracy for multiscale parts, we design a feature pyramid network to generate rectangular layers of different sizes to map parts with similar aspect ratios. Experiments on four fault datasets show that our framework achieves 98.44% accuracy while the model size is only 22.5 MB, outperforming the state-of-the-art detectors.
Yang Zhang 0053, Huilin Pan, Guodong Sun 0002
IEEE Trans. Ind. Informatics1
2022 Affinity Fusion Graph-Based Framework for Natural Image Segmentation
abstract
This paper proposes an affinity fusion graph framework to effectively connect different graphs with highly discriminating power and nonlinearity for natural image segmentation. The proposed framework combines adjacency-graphs and kernel spectral clustering based graphs (KSC-graphs) according to a new definition named affinity nodes of multi-scale superpixels. These affinity nodes are selected based on a better affiliation of superpixels, namely subspace-preserving representation which is generated by sparse subspace clustering based on subspace pursuit. Then a KSC-graph is built via a novel kernel spectral clustering to explore the nonlinear relationships among these affinity nodes. Moreover, an adjacency-graph at each scale is constructed, which is further used to update the proposed KSC-graph at affinity nodes. The fusion graph is built across different scales, and it is partitioned to obtain final segmentation result. Experimental results on the Berkeley segmentation dataset and Microsoft Research Cambridge dataset show the superiority of our framework in comparison with the state-of-the-art methods. The code is available athttps://github.com/Yangzhangcst/AF-graph.
Yang Zhang 0053, Moyun Liu, Jingwu He, Yanwen Guo 0001
IEEE Trans. Multim.1
2021 A Unified Light Framework for Real-Time Fault Detection of Freight Train Images
abstract
Real-time fault detection for freight trains plays a vital role in guaranteeing the security and optimal operation of railway transportation under stringent resource requirements. Despite the promising results for deep-learning-based approaches, the performance of these fault detectors on freight train images is far from satisfactory in both accuracy and efficiency. This article proposes a unified light framework to improve detection accuracy while supporting a real-time operation with a low-resource requirement. We first design a novel lightweight backbone (real-time fault detection network-RFDNet) to improve the accuracy and reduce computational cost. Then, we propose a multiregion proposal network using multiscale feature maps generated from the RFDNet to improve the detection performance. Finally, we present multilevel position-sensitive score maps and region of interest pooling to further improve accuracy with few redundant computations. Extensive experimental results on public benchmark datasets suggest that our RFDNet can significantly improve the performance of the baseline network with higher accuracy and efficiency. Experiments on six fault datasets show that our method is capable of real-time detection at over 38 frames/s and achieves competitive accuracy and lower computation than the state-of-the-art detectors.
Yang Zhang 0053, Moyun Liu, Yang Yang 0092, Yanwen Guo 0001
IEEE Trans. Ind. Informatics1
2019 PANet: A Context Based Predicate Association Network for Scene Graph Generation
abstract
Scene graph generation is widely studied in recent years, which tries to understand the interactions of different objects as a whole. The earlier researches only recognize a few relationships or model contexts among different relationships, neglecting the associations of predicates for each object pair. In this paper, we propose a two-stage framework named predicate association network (PANet) to properly extract contexts and model predicate association. In the first stage, instance-level and scene-level context are extracted for object classification and further used for predicate classification in the next stage. With a recurrent neural network, alignment technique and attention mechanism are combined to collect the associations of predicates in the second stage. The experiments on the Visual Genome dataset show that our method is effective and outperforms the state-of-the-art methods.
Yunian Chen, Yang Zhang 0053, Yanwen Guo 0001
ICME3
2019 Improving Open Set Domain Adaptation Using Image-to-Image Translation
abstract
The open set domain adaptation problem was rarely studied and its existing solutions are mostly based on learning a joint latent space which may encounter issues when the domains differ significantly from each other. This work is driven by the question whether or not it is beneficial to operate the source images to another image domain as close to the target as possible. We propose to address the open set domain adaptation problem by aligning sample at both feature space and pixel space. Our approach, called Open Set Translation and Adaptation Network (Ostan), consists of two main components: translation and adaptation. The translation model is a cycle-consistent generative adversarial network, which translates any source sample to the "style" of a target domain. The adaptation network is built upon OpenBP, an open set domain adaptation framework, and trained using both (labeled) translated source images and (unlabeled) target images. The proposed Ostan model significantly outperforms the state-of-the-art open set domain adaptation methods on multiple public datasets. Our experiment also demonstrates that an image-to-image translation component can further improve the decision boundaries for both known and unknown classes.
Hongjie Zhang 0002, Yang Zhang 0053, Yanwen Guo 0001
ICME5
2019 An Adaptive Affinity Graph with Subspace Pursuit for Natural Image Segmentation
abstract
Graph-based segmentation methods have become a major trend in computer vision. Due to the advantages of assimilating different graphs, a multi-scale fusion graph have a better performance than a single graph with single-scale. However, it is not reliable to determine a principle of graph combination. In this paper, we propose an adaptive affinity graph with subspace pursuit (AASP-graph) for natural image segmentation. The input image is first over-segmented into superpixels at different scales. An improved affinity propagation clustering method is proposed to select global nodes of these superpixels adaptively. Then, a L0-graph at each scale is obtained by a sparse representation of global nodes based on subspace pursuit. The adjacency-graph is finally built upon all superpixels of each scale and updated by the L0-graph. Experimental results on the Berkeley segmentation database show the effectiveness of the proposed AASP-graph in comparison with state-of-the-art approaches.
Yang Zhang 0053, Yanwen Guo 0001, Jingwu He
ICME1
2018 A Unified Framework for Fault Detection of Freight Train Images Under Complex Environment
abstract
This paper proposes a novel unified framework for fault detection of the freight train images based on convolutional neural network (CNN) under complex environment. Firstly, the multi region proposal networks (MRPN) with a set of prior bounding boxes are introduced to achieve high quality fault proposal generation. And then, we apply a linear non-maximum suppression method to retain the most suitable anchor while removing redundant boxes. Finally, a powerful multi-level region-of-interest (ROI) pooling is proposed for proposal classification and accurate detection. The experimental results indicate that the proposed method can achieve high performance on four typical fault benchmarks, substantially outperforming the state-of-the-art methods.
Yang Zhang 0053, Yanwen Guo 0001, Guodong Sun 0002
ICIP1