Yongguo Ling

dblp:276/3201 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0003-1582-0987ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CIA: Cluster-Instance Alignment for Unsupervised Day-Night Vehicle Re-Identification
abstract
Cross-time vehicle re-identification (Re-ID), especially across day and night conditions, remains a challenging problem due to drastic illumination variations that lead to significant domain shifts. While existing methods perform well under daytime scenarios, their effectiveness degrades severely in cross-domain settings, and fully supervised solutions demand costly annotations in both domains. In this paper, we introduce a new setting, Unsupervised Day-Night Vehicle Re-Identification (USL-DN-ReID), and propose a novel Cluster-Instance Alignment (CIA) framework to address it. CIA performs dual-level alignment: 1) at the cluster level, a Dictionary-Guided Graph Matching (DGM) module builds a cross-domain topological graph using soft similarities among cluster centers and solves global matching via the Hungarian algorithm; 2) at the instance level, a Multi-Factor Adaptive Alignment (MAA) module introduces a multi-factor adaptive weighting strategy that emphasizes high-confidence pairwise relations while suppressing noise. Together, these components enable robust and scalable cross-domain adaptation without requiring target-domain labels. Extensive experiments conducted on the DN-348 and DN-Wild benchmarks demonstrate the effectiveness and superiority of the proposed CIA framework, setting new state-of-the-art results on both datasets.
Yongguo Ling, Wenhao Shao
AAAI1
2026 SCF-Net: Spatial-Channel Fusion and Feature Refinement for Vessel Re-Identification
abstract
ABSTRACT Vessel re‐identification (ReID) plays a critical role in maritime surveillance by matching vessels across different camera views. Compared with person or vehicle ReID, vessel ReID faces unique challenges due to subtle interclass differences and large intraclass variations caused by viewpoint changes. These issues are further exacerbated by the highly similar appearances of vessels and the lack of fine‐grained identity cues commonly found in other ReID tasks. To address these challenges, we propose a spatial‐channel fusion network (SCF‐Net), a dual‐branch deep framework that integrates a spatial‐channel fusion (SCF) module and a feature refinement and alignment (FRA) module. The SCF module captures interdependent relationships between spatial and channel dimensions, enabling the network to emphasize discriminative regions while suppressing irrelevant background information. The FRA module refines high‐dimensional embeddings into a compact representation and enforces intraclass similarity via a learnable multilayer perceptron (MLP) and a supervised mean squared error (MSE) loss. By jointly optimizing the two branches and the FRA output, SCF‐Net effectively learns both interclass discrimination and intraclass compactness. Extensive experiments demonstrate that SCF‐Net achieves competitive performance on public vessel ReID benchmarks, highlighting its effectiveness in handling subtle interclass differences and large intraclass variations.
Gangzhu Lin, Yongguo Ling, Wenhao Shao, Shaozi Li, Hongfeng Xu
Concurr. Comput. Pract. Exp.2
2026 Identity-Compensated Style Distillation for Visible-Infrared Person Re-Identification
abstract
Visible-Infrared Person Re-Identification (VI-ReID) that matches pedestrian images across visible and infrared modalities suffers from substantial modality discrepancies and intra-class variations. While existing methods typically address the modality gap via style alignment, they often lose identity-relevant semantics and overlook fine-grained inter-class nuances, such as body part contours and structural cues around the head, shoulders, or feet. To tackle these challenges, we propose an Identity-Compensated Style Distillation (ICSD) network that enforces cross-modality style consistency and enhances the discriminative power of modality-invariant features. Specifically, ICSD comprises two core components: (1) a Style Knowledge Distillation (SKD) module, which integrates Style Discrepancy Reduction (SDR) and Identity Knowledge Compensation (IKC) to align modality styles while preserving identity-relevant semantics; (2) an Identity Discrimination Amplification (IDA) module, which captures and enhances subtle inter-class differences by refining identity-specific cues, thereby facilitating more accurate discrimination between different pedestrians. Extensive experiments on three public benchmarks-SYSU-MM01, RegDB, and LLCM-demonstrate that ICSD consistently outperforms state-of-the-art methods, validating the effectiveness and complementarity of its components.
Yongguo Ling, Zihao Hu, Nan Pu, Zhun Zhong, Xudong Jiang 0001
IEEE Trans. Image Process.1
2025 VAMN: View-Invariant Adaptive Multi-granularity Network for Unsupervised Vessel Re-identification
Yize Ma, Yongguo Ling, Gangzhu Lin
PRCV (2)2
2025 Text-Guided Multiround Learning for Unsupervised Visible-Infrared Person Re-Identification
abstract
Unsupervised Visible-Infrared Person Re-Identification (USVI-ReID) aims to match images of the same individual across visible and infrared modalities without relying on identity annotations. This task is critical in Visual Internet of Things (VIoT) systems. Existing methods typically generate pseudo-labels through clustering and cross-modality associations. However, the quality of these pseudo-labels is often compromised by unstable clustering performance, modality discrepancies, and unreliable matching strategies, resulting in suboptimal accuracy. Therefore, obtaining more reliable and robust pseudo-labels remains challenging in this domain. To address these challenges, we propose a Text-Guided Multi-Round Learning (TGMRL) framework that enhances the reliability and robustness of cross-modality pseudo-label associations. TGMRL comprises two core components: (1) the Text Dual-Cross Similarity Matching (TDSM) module, which facilitates cross-modality cluster alignment by constructing both visual and text-based cluster centers and integrating their similarity matrices for more accurate correspondence; and (2) the Multi-Round Pseudo-Label Guidance (MRPG) module, which enhances label consistency by imposing temporal regularization across clustering iterations through both similarity and Intersection-Over-Union (IoU) based measures. Extensive experiments on two benchmark datasets demonstrate that TGMRL significantly improves the reliability of pseudo-labels and achieves state-of-the-art performance on the USVI-ReID task.
Yongguo Ling, Zihao Hu, Wenhao Shao, Shaozi Li, Thomas Wu 0001
IEEE Internet Things J.2
2025 OTMA: Optimal transfer modality alignment for visible-thermal person re-identification
Yongguo Ling, Zihao Hu, Gangzhu Lin, Shaozi Li, Min Jiang 0005
Knowl. Based Syst.1
2025 Cross-modality average precision optimization for visible thermal person re-identification
Yongguo Ling, Zhiming Luo, Dazhen Lin, Shaozi Li, Min Jiang 0005, Nicu Sebe, Zhun Zhong
Pattern Recognit.1
2025 Dual-Modality-Shared Learning and Label Refinement for Unsupervised Visible-Infrared Person ReID
abstract
Unsupervised visible-infrared person re-identification (USVI-ReID) aims to match a person across two modalities without annotations. Current research primarily addresses the modality gap by establishing cross-modality correspondences through matching algorithms and utilizing memory banks for contrastive learning. However, the inherent noise in pseudo labels and neglect of hard samples often limit the efficacy of cross-modality learning. In this article, we propose a dual-modality-shared learning and label refinement (DLLR) algorithm for USVI-ReID. First, we leverage a cluster similarity matching (CSM) module and a cluster relationship-based label refinement (CRLR) algorithm to create and refine pseudo labels. Then, we adopt a weighted modality-shared memory (WMM) to construct memory banks by jointly considering sample distribution and feature differences, thereby enhancing the effectiveness of cross-modality learning. Extensive experiments on three publicly available datasets validate the effectiveness of our proposed method, which outperforms state-of-the-art methods. The code is available at https://github.com/CharRic/DLLR .
Licun Dai, Zhiming Luo, Yongguo Ling, Jiaxing Chai, Shaozi Li
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Bridge Gap in Pixel and Feature Level for Cross-Modality Person Re-Identification
abstract
Visible thermal person re-identification (VT-ReID) plays a vital role in intelligent surveillance systems, particularly in weak lighting environments. VT-ReID faces substantial challenges, including the cross-modality gap and intra-class variations. Existing methods address these challenges through either pixel-level image translation techniques or feature-level metric learning techniques. However, the former approaches require additional computational costs and often generate noisy images, making model training challenging. The latter methods focus on constraining the relations between individual instances or class centers, while often ignoring joint consideration of the relationship between the two aspects. In addition, these works do not fully investigate the mutual benefits at both pixel-level and feature-level. To address these limitations, we propose a unified Dual-level Smooth Gap (DSG) learning framework that simultaneously smooths the cross-modality gap at the pixel and feature levels. Specifically, on the one hand, we develop a parameter-free Class-aware Modality Mix (CMM) to smooth the cross-modality gap at the pixel level. CMM can capture and explore internal information between the two modalities by mixing images from different modalities belonging to the same class. On the other hand, we devise an efficient Center-guided Metric Learning (CML) to reduce the inter-modality discrepancy and intra-class variations at the feature level. CML enhances model discrimination and generalization by enforcing constraints on both class centers and instances. Experiments on two benchmark datasets demonstrate the mutual benefits of our proposed and show the superior performance of our method over state-of-the-art methods.
Yongguo Ling, Zhun Zhong, Zhiming Luo, Shaozi Li, Nicu Sebe
IEEE Trans. Circuits Syst. Video Technol.1
2023 Cross-Modality Earth Mover's Distance for Visible Thermal Person Re-identification
abstract
Visible thermal person re-identification (VT-ReID) suffers from inter-modality discrepancy and intra-identity variations. Distribution alignment is a popular solution for VT-ReID, however, it is usually restricted to the influence of the intra-identity variations. In this paper, we propose the Cross-Modality Earth Mover's Distance (CM-EMD) that can alleviate the impact of the intra-identity variations during modality alignment. CM-EMD selects an optimal transport strategy and assigns high weights to pairs that have a smaller intra-identity variation. In this manner, the model will focus on reducing the inter-modality discrepancy while paying less attention to intra-identity variations, leading to a more effective modality alignment. Moreover, we introduce two techniques to improve the advantage of CM-EMD. First, Cross-Modality Discrimination Learning (CM-DL) is designed to overcome the discrimination degradation problem caused by modality alignment. By reducing the ratio between intra-identity and inter-identity variances, CM-DL leads the model to learn more discriminative representations. Second, we construct the Multi-Granularity Structure (MGS), enabling us to align modalities from both coarse- and fine-grained levels with the proposed CM-EMD. Extensive experiments show the benefits of the proposed CM-EMD and its auxiliary techniques (CM-DL and MGS). Our method achieves state-of-the-art performance on two VT-ReID benchmarks.
Yongguo Ling, Zhun Zhong, Zhiming Luo, Fengxiang Yang, Donglin Cao, Yaojin Lin, Shaozi Li, Nicu Sebe
AAAI1
2023 Dual-Stream Transformer With Distribution Alignment for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification(VI-ReID) aims to match the person images captured by visible and infrared cameras and suffers from severe cross-modality discrepancy and intra-modality variations. Existing approaches mainly use convolution neural network (CNN)-based architectures to extract pedestrian features, which fail to capture the long-range dependencies within an image. In addition, previous works usually attempt to bridge the modality gap by using adversarial learning to generate style-consistent images or designing different feature-level metric learning constraints. However, few works consider the cross-modality disparity from the perspective of assessing overall distance distribution discrepancy. To address these problems, we design a pure Transformer-based Visible-Infrared (TransVI) network with a conventional two-stream structure, which can explicitly capture modality-specific representations and learn multi-modality sharable knowledge. TransVI can efficiently address the lack of global dependency in CNN-based architectures due to the multi-head self-attention modules in the transformer, which allows us to capture the long-range dependencies of pedestrian images. Furthermore, we introduce the Cross-Modality Dissimilarity-based Maximum Mean Discrepancy (CMD-MMD) constraint to handle the cross-modality discrepancy at the distance distribution level. Specifically, CMD-MMD leverages intra-modality distribution separability to guide inter-modality distribution separability learning, aligning pair-wise distance distributions of intra- and inter-modality for within-class and between-class, respectively. In this way, the distance distributions of intra- and inter-modality become more similar, significantly mitigating the cross-modality discrepancy and learning more modality invariant representations. Extensive experimental results on two public VI-ReID datasets confirm that our proposed framework can achieve state-of-the-art performance.
Zehua Chai, Yongguo Ling, Zhiming Luo, Dazhen Lin, Min Jiang 0005, Shaozi Li
IEEE Trans. Circuits Syst. Video Technol.2
2021 A Multi-Constraint Similarity Learning with Adaptive Weighting for Visible-Thermal Person Re-Identification
abstract
The challenges of visible-thermal person re-identification (VT-ReID) lies in the inter-modality discrepancy and the intra-modality variations. An appropriate metric learning plays a crucial role in optimizing the feature similarity between the two modalities. However, most existing metric learning-based methods mainly constrain the similarity between individual instances or class centers, which are inadequate to explore the rich data relationships in the cross-modality data. Besides, most of these methods fail to consider the importance of different pairs, incurring an inefficiency and ineffectiveness of optimization. To address these issues, we propose a Multi-Constraint (MC) similarity learning method that jointly considers the cross-modality relationships from three different aspects, i.e., Instance-to-Instance (I2I), Center-to-Instance (C2I), and Center-to-Center (C2C). Moreover, we devise an Adaptive Weighting Loss (AWL) function to implement the MC efficiently. In the AWL, we first use an adaptive margin pair mining to select informative pairs and then adaptively adjust weights of mined pairs based on their similarity. Finally, the mined and weighted pairs are used for the metric learning. Extensive experiments on two benchmark datasets demonstrate the superior performance of the proposed over the state-of-the-art methods.
Yongguo Ling, Zhiming Luo, Yaojin Lin, Shaozi Li
IJCAI1
2020 Class-Aware Modality Mix and Center-Guided Metric Learning for Visible-Thermal Person Re-Identification
abstract
Visible thermal person re-identification (VT-REID) is an important and challenging task in that 1) weak lighting environments are inevitably encountered in real-world settings and 2) the inter-modality discrepancy is serious. Most existing methods either aim at reducing the cross-modality gap in pixel- and feature-level or optimizing cross-modality network by metric learning techniques. However, few works have jointly considered these two aspects and studied their mutual benefits. In this paper, we design a novel framework to jointly bridge the modality gap in pixel- and feature-level without additional parameters, as well as reduce the inter- and intra-modalities variations by a center-guided metric learning constraint. Specifically, we introduce the Class-aware Modality Mix (CMM) to generate internal information of the two modalities for reducing the modality gap in pixel-level. In addition, we exploit the KL-divergence to further align modality distributions on feature-level. On the other hand, we propose an efficient Center-guided Metric Learning (CML) method for decreasing the discrepancy within the inter- and intra-modalities, by enforcing constraints on class centers and instances. Extensive experiments on two datasets show the mutual advantage of the proposed components and demonstrate the superiority of our method over the state of the art.
Yongguo Ling, Zhun Zhong, Zhiming Luo, Paolo Rota, Shaozi Li, Nicu Sebe
ACM Multimedia1