Chule Yang

dblp:193/6767 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-6548-5841ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SGFM-Net: A semantically guided feature mining network for fine-grained ship classification in remote sensing images
Chule Yang, Minming Ye, Longfei Su, Naiyang Guan
Knowl. Based Syst.2
2025 LPN: Language-Guided Prototypical Network for Few-Shot Classification
abstract
Few-shot classification aims to adapt to new tasks with limited labeled examples. Recent methods have explored various techniques for measuring the similarity between query and support images, along with meta-training and pre-training strategies, to leverage visual features more effectively. However, the potential of multi-modality information remains unexplored, presenting a promising avenue for improvement in few-shot classification. In this paper, we propose a Language-guided Prototypical Network (LPN) for few-shot classification without image-level captions. LPN leverages the complementarity of vision and language modalities through two parallel branches with pre-fusion and post-fusion. Firstly, we introduce the language modality by utilizing a pre-trained text encoder to extract class-level text features from class names. In the visual branch, we process images using a conventional image encoder and leverage the class-level features to align the visual features, effectively capturing more class-relevant visual information. In the text branch, we combine the class-level text features with the visual features using a language-guided decoder. This decoder generates image-specific text features for the pre-fusion step. Additionally, we utilize these class-level text features to refine the prototypical head, creating robust prototypes for subsequent measurements. Finally, to enhance overall performance, we aggregate the visual and text logits, adjusting for discrepancies between the modalities during the post-fusion process. Extensive experiments demonstrate that LPN outperforms several state-of-the-art methods on benchmark datasets, showcasing its effectiveness and robustness.
Kaihui Cheng, Chule Yang, Naiyang Guan
IEEE Trans. Circuits Syst. Video Technol.2
2024 Zero-Shot Sea-Land Segmentation Based on Edge Analysis and Automatic Prompt Point Generation
abstract
In this study, we tackle the critical challenge of sea-land segmentation for enhancing ship detection in remote sensing imagery. By focusing the search area, we significantly reduce the false alarm rate in ship detection. To overcome the limitations of scarce training data and annotation complexity, we introduce a novel zero-shot sea-land segmentation framework. This framework is integrated with a robust image segmentation model and employs an automatic prompt point generation strategy to facilitate sea-land discrimination. Our method is structured around three core components: image preprocessing, image block classification, and prompt point generation. The preprocessing module generates an edge feature map using downsampling and edge detection techniques. The classification module segments land and sea areas through a rasterization and classification process. The prompt point generation module leverages edge point clustering to create prompts on the land image block background. These prompts are then utilized by the image segmentation model to achieve precise sea segmentation. Our extensive experiments on public land-sea segmentation datasets show that our approach is robust across various parameter settings and outperforms current state-of-the-art methods.
Chule Yang, Naiyang Guan
ICARCV2
2024 Text-Guided Feature Mining for Fine-Grained Ship Classification in Optical Remote Sensing Images
abstract
Ship classification in remote sensing images holds paramount significance for surveillance and defense applications. However, the intricate nature of ship classes poses a challenge in accurately discriminating among diverse ship types. This challenge is exacerbated by the long-tailed data distribution, where certain classes are underrepresented by insufficient samples. Furthermore, the limited visual information available from single-perspective remote sensing imaging complicates the extraction of discriminative features. To address these challenges, we propose the Text-Guided Feature Mining Network (TGFMN). This approach aims to exploit strong textual knowledge to construct discriminative image features. We employ a unified classifier for pre-training to classify features from both modalities, ensuring that the dual modal features are integrated into a cohesive feature space, thereby maximizing the efficacy of textual information. The text-guided strategy is employed to enhance attention regions within the shallow spatial dimensions of visual features, facilitating the extraction of highly discriminative visual features during the fine-tuning process. The experimental results demonstrate the effectiveness of our method, achieving state-of-the-art performance on two benchmark datasets for fine-grained ship classification in remote sensing: FGSC-23 and FGSCR-42.
Chule Yang, Naiyang Guan
ICARCV3
2023 TeAw: Text-Aware Few-Shot Remote Sensing Image Scene Classification
abstract
The recent advance has shown that few-shot learning may be a promising way to alleviate the data reliance of remote sensing image scene classification. However, most existing works focus on extracting distinguishable features only from visual modality, while the problem of learning knowledge from multiple modalities has barely been visited. In this work, we propose a text-aware framework for few-shot remote sensing image scene classification (TeAw). Specifically, TeAw converts the class names to more detailed text descriptions and extracts text features using a pre-trained text encoder. Mean-while, TeAw obtains image features via an image encoder. Then we compute the correlation between the text and the image features, which helps the model grasp the core concept of the input image. Finally, TeAw calculates the similarity of local features between supports and queries to get the predictions. Extensive experiments show the outperformance of our TeAw compared with other SOTA methods.
Kaihui Cheng, Chule Yang, Zunlin Fan, Dayan Wu, Naiyang Guan
ICASSP2
2023 A Multi-Channel Aggregation Framework for Object Detection in Large-Scale SAR Image
abstract
Synthetic aperture radar (SAR) has gradually demonstrated its advantages in a variety of application fields. However, due to the complexity of the background, the simplicity of the texture, and the multi-scale of the target, object detection in large-scale SAR images is still a major challenge. This paper focuses on a multi-channel aggregation framework, and jointly considers the pre-processing and post-processing for algorithm optimization to improve the overall performance of object detection in large-scale SAR images. In this paper, multiple sets of slices of large-scale images are first sliced using slicers of various sizes. Feeding different groups of image slices into the detection network for feature extraction can achieve a greater degree of scaling of target features, thereby improving the ability of multi-scale target discovery. Then, a post-processing module is proposed to transform and fuse the results from multiple channels, and two sub-algorithms result fusion and result refinement are proposed to eliminate false alarms. Result fusion merges the detection boxes from each channel and proposes a weighted NMS strategy to update the confidence of the optimal box. Result refinement proposes an adaptive belief update strategy to filter the remaining boxes for eliminating those detection boxes with low beliefs. Qualitative and quantitative experiments were conducted to prove the effectiveness of the proposed method, which can reduce the false alarm rate while increasing the recall rate.
Chule Yang, Zunlin Fan, Zeting Yu, Qianchong Sun, Mengyuan Dai
ICASSP1
2023 AREA: Adaptive Reweighting via Effective Area for Long-Tailed Classification
abstract
Large-scale data from the real-world usually follow a long-tailed distribution (i.e., a few majority classes occupy plentiful training data, while most minority classes have few samples), making the hyperplanes heavily skewed to the minority classes. Traditionally, reweighting is adopted to make the hyperplanes fairly split the feature space, where the weights are designed according to the number of samples. However, we find that the number of samples in a class can not accurately measure the size of its spanned space, especially for the majority class, where the size of its spanned space is usually larger than the samples’ number because of the high diversity. Therefore, weights designed based on the samples’ number will still compress the space of minority classes. In this paper, we reconsider reweighting from a totally new perspective of analyzing the spanned space of each class. We argue that, besides statistical numbers, relations between samples are also significant for sufficiently depicting the spanned space. Consequently, we estimate the size of the spanned space for each category, namely effective area, by detailedly analyzing its samples’ distribution. By treating samples of a class as identically distributed random variables and analyzing their correlations, a simple and non-parametric formula is derived to estimate the effective area. Then, the weight simply calculated inversely proportional to the effective area of each class is adopted to achieve fairer training. Note that our weights are more flexible as they can be adaptively adjusted along with the optimizing features during training. Experiments on four long-tailed datasets show that the proposed weights outperform the state-of-the-art reweighting methods. Moreover, our method can also achieve better results on statistically balanced CIFAR-10/100. Code is available at https://github.com/xiaohua-chen/AREA.
Xiaohua Chen 0002, Yucan Zhou, Dayan Wu, Chule Yang, Bo Li 0063, Qinghua Hu, Weiping Wang 0005
ICCV4
2022 Clustering and Separating Similarities for Deep Unsupervised Hashing
abstract
The lack of supervised information is the pivotal problem in unsupervised hashing. Most methods leverage deep features extracted from pre-trained models to generate semantic similarities as supervised information. These fixed features are, however, neither designed originally for retrieval nor updated adaptively during training. In this paper, we propose a novel deep Unsupervised Cluster and Separate Hashing (UCSH) to address these issues. Specifically, we introduce a fully end-to-end deep hashing network with a binary latent Variational AutoEncoder (VAE), which enables hash codes capable of reconstructing deep features as well as preserving semantic relations. Moreover, a ‘Cluster and Separate’ scheme is proposed to jointly cluster deep features and separate semantic similarities. Both the implicit feature clustering and the explicit similarity separating loss encourage the separation of similar and dissimilar pairs, enabling the iteratively updated similarities to better excavate semantic relations. Experiments conducted on three benchmarks show the superiority of UCSH.
Wanqian Zhang, Dayan Wu, Chule Yang, Bo Li 0063, Weiping Wang 0005
ICASSP3
2022 Cross-view vehicle re-identification based on graph matching
Chao Zhang 0074, Chule Yang, Dayan Wu, Hongbin Dong
Appl. Intell.2
2022 Uncertainty-Aware and Multigranularity Consistent Constrained Model for Semi-Supervised Hashing
abstract
Recently, deep semi-supervised hashing methods have attracted increasing attention, which can significantly improve retrieval performance by leveraging abundant unlabeled data. These methods usually generate surrogate supervision signals to learn with unlabeled data, such as neighborhood information and augmentation invariant requirements. However, an essential issue of these methods is that the supervised signals are not always reliable, which may damage the performance. In this paper, we propose a novel Uncertainty-Aware and Multi-Granularity Consistent Constrained Semi-Supervised Hashing (UMCSH) method to alleviate the negative effects of noisy supervised signals and enlarge the inter-class distance. Specifically, our UMCSH mainly consists of an Uncertainty-Aware Instance-Level Consistency (UAILC) model and a Cluster-Based Class-Level Consistency (CBCLC) model. UAILC introduces an uncertainty estimation method to select reliable supervised signals to extract discriminative features for each unlabeled data. CBCLC establishes connections between labeled data and unlabeled data by encouraging each unlabeled sample to be close to the hash center (calculated with the labeled data) according to its pseudo-label. Extensive experimental results demonstrate the superior performance of our proposed approach compared with several state-of-the-art semi-supervised hashing methods.
Shuai Cheng 0002, Yucan Zhou, Wanqian Zhang, Dayan Wu, Chule Yang, Bo Li 0063, Weiping Wang 0005
IEEE Trans. Circuits Syst. Video Technol.5
2022 MSIF: Multisize Inference Fusion-Based False Alarm Elimination for Ship Detection in Large-Scale SAR Images
abstract
Ship detection in large-scale synthetic aperture radar (SAR) images has essential value in both military and civilian applications. However, due to the complexity of the background and the simplicity of the texture, ship detection in large-scale SAR images is prone to false alarms, such as similar-shaped reefs, islands, sea clutter, and inland buildings. This article proposes a multisize inference fusion framework to eliminate false alarms and improve the overall performance of ship detection in large-scale SAR images. In this framework, a multisize slicer is proposed to expand the scale range of image expression. Then, a detection model library is built to keep various types of models for different task scenarios and requirements. Finally, two subapproaches are proposed for false alarm elimination, namely, pixel feature filtering (FAE-pff) and multisource fusion (FAE-msf), to reduce false detection results in the output of the detection model. FAE-pff calculates how obvious each target is relative to the background and eliminates less obvious results. FAE-msf obtains bounding boxes and corresponding confidences from multiple inference sources and fuses them through weighting and updating them to achieve complementation and enhancement of information. Various experiments were conducted to evaluate the performance of each module qualitatively and quantitatively. It proves the effectiveness of the proposed framework, which can achieve more correct detections while greatly reducing erroneous detections.
Chao Zhang 0074, Chule Yang, Kaihui Cheng, Naiyang Guan, Hongbin Dong
IEEE Trans. Geosci. Remote. Sens.2
2020 Multi-Robot Collaborative Reasoning for Unique Person Recognition in Complex Environments
abstract
The discovery of unique or suspicious people is essential for active surveillance of security or patrol robots, and multi-robot collaboration and dynamic reasoning can further enhance their adaptability in large-scale environments. This paper proposes a hierarchical probabilistic reasoning framework for a multi-robot system to actively identify the unique person with distinct motion patterns in large-scale and dynamic environments. Linear and angular velocities are considered typical motion patterns, which are extracted by using heterogeneous sensors to detect and track people. First, single robot reasoning is performed, each robot judges the uniqueness of people by comparing their motion patterns based on local observations. Meanwhile, multi-robot reasoning is also performed, by fusing the perceptual information from each individual robot to form a global observation and then make another judgment based on it. Finally, each robot can decide which result should be adopted by comparing the beliefs of local and global judgments. Experimental results show that the method is feasible in various environments.
Chule Yang, Yufeng Yue, Mingxing Wen, Yuanzhe Wang
ICARCV1
2020 Day and Night Collaborative Dynamic Mapping in Unstructured Environment Based on Multimodal Sensors
abstract
Enabling long-term operation during day and night for collaborative robots requires a comprehensive understanding of the unstructured environment. Besides, in the dynamic environment, robots must be able to recognize dynamic objects and collaboratively build a global map. This paper proposes a novel approach for dynamic collaborative mapping based on multimodal environmental perception. For each mission, robots first apply heterogeneous sensor fusion model to detect humans and separate them to acquire static observations. Then, the collaborative mapping is performed to estimate the relative position between robots and local 3D maps are integrated into a globally consistent 3D map. The experiment is conducted in the day and night rainforest with moving people. The results show the accuracy, robustness, and versatility in 3D map fusion missions.
Yufeng Yue, Chule Yang, Jun Zhang 0042, Mingxing Wen, Zhenyu Wu 0001, Danwei Wang
ICRA2
2020 A Hierarchical Framework for Collaborative Probabilistic Semantic Mapping
abstract
Performing collaborative semantic mapping is a critical challenge for cooperative robots to maintain a comprehensive contextual understanding of the surroundings. Most of the existing work either focus on single robot semantic mapping or collaborative geometry mapping. In this paper, a novel hierarchical collaborative probabilistic semantic mapping framework is proposed, where the problem is formulated in a distributed setting. The key novelty of this work is the mathematical modeling of the overall collaborative semantic mapping problem and the derivation of its probability decomposition. In the single robot level, the semantic point cloud is obtained based on heterogeneous sensor fusion model and is used to generate local semantic maps. Since the voxel correspondence is unknown in collaborative robots level, an Expectation-Maximization approach is proposed to estimate the hidden data association, where Bayesian rule is applied to perform semantic and occupancy probability update. The experimental results show the high quality global semantic map, demonstrating the accuracy and utility of 3D semantic map fusion algorithm in real missions.
Yufeng Yue, Chule Yang, Jun Zhang 0042, Mingxing Wen, Yuanzhe Wang, Danwei Wang
ICRA4
2019 Probabilistic Reasoning for Unique Role Recognition Based on the Fusion of Semantic-Interaction and Spatio-Temporal Features
abstract
This paper deals with the problem of recognizing the unique role in dynamic environments. Different from social roles, the unique role refers to those who are unusual in their carrying items or movements in the scene. In this paper, we propose a hierarchical probabilistic reasoning method that relates spatial relationships between interested objects and humans with their temporal changes to recognize the unique individual. Two observation models, Object Existence Model (OEM) and Human Action Model (HAM), are established to support role inference by analyzing the corresponding semantic-interaction features and spatio-temporal features. Then, OEM and HAM results of each person are compared with the overall distribution in the scene, respectively. Finally, we can determine the role through the fusion of two observation models. Experiments are conducted in both indoor and outdoor environments concerning different settings, degrees of clutter, and occlusions. The results show that the proposed method can adapt to a variety of scenarios and outperforms other methods on accuracy and robustness, moreover, exhibiting stable performance even in complex scenes.
Chule Yang, Yufeng Yue, Jun Zhang 0042, Mingxing Wen, Danwei Wang
IEEE Trans. Multim.1
2018 Probabilistic Fusion Framework for Collaborative Robots 3D Mapping
abstract
Fusion of local 3D maps generated by individual robots to a globally consistent 3D map is one of the fundamental challenges in multi-robot mapping missions. In this paper, we propose a probabilistic mathematical formulation to address the integrated map fusion problem. More specifically, the problem of estimating fused map posterior can be factorized into a product of relative transformation posterior and the global map posterior, which enables us to solve map matching and map merging problems efficiently. In addition, a distributed communication strategy is employed to share map information among robots. The proposed approach is evaluated in indoor and mixed environments, which shows its utility in 3D map fusion for multi-robot mapping missions.
Yufeng Yue, P. G. C. N. Senarathne, Chule Yang, Jun Zhang 0042, Mingxing Wen, Danwei Wang
FUSION3
2018 A Two-step Method for Extrinsic Calibration between a Sparse 3D LiDAR and a Thermal Camera
abstract
To obtain the 6 DOF extrinsic parameters (rotation and translation matrix) between a 3D ranging sensor and a thermal camera, previous methods require a high-resolution 3D ranging sensor to reliably detect features. Although sparse 3D LiDARs are widely used on autonomous robots, to the best of our knowledge, the extrinsic calibration between a sparse 3D LiDAR (particularly Velodyne VLP-16) and a thermal camera has not been considered in the literature. In this paper, we present a two-step method to address the problem, where a monocular visual camera is used to assist the process. The proposed method decomposes the problem into two steps: extrinsic calibration between a sparse 3D LiDAR and a visual camera; extrinsic calibration between a visual camera and a thermal camera. Experiments are conducted to demonstrate the effectiveness of the proposed two-step method.
Jun Zhang 0042, Prarinya Siritanawan, Yufeng Yue, Chule Yang, Mingxing Wen, Danwei Wang
ICARCV4
2016 Organ-Based Facial Verification Using Thermal Camera
abstract
So far, most of the facial recognition methods focus on visual image texture and color information. Although they have worked well, most of them still fail to deal with severe illumination changes. In this paper, a novel approach is proposed for facial verification by analyzing thermal data from different organs of the human face. This new thermal facial pattern is free from illumination changes and can even work in a very dark place. In this experiment, facial thermal data were collected from 30 people with diverse genders, ages, and races. Three persons were tracked to verify the consistency of the thermal pattern in normal circumstances, changing light conditions and different physical conditions. A new distance function is introduced for pattern similarity measurement. The proposed approach successfully distinguished different persons with a high verification rate 91.26% according to F-measure and displayed the thermal pattern changes according to different physical conditions.
Chule Yang, Danwei Wang, Prarinya Siritanawan
ISM1