EDBT 2026 Demo / reviewers in the wild / expert
Yiming Yang 0001
dblp:25/1666-1
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2026
0009-0005-7466-1041ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous DrivingabstractIn autonomous driving, vision-centric 3D object detection recognizes and localizes 3D objects from RGB images. However, due to high annotation costs and diverse outdoor scenes, training data often fails to cover all possible test scenarios, known as the out-of-distribution (OOD) issue. Training-free image editing offers a promising solution for improving model robustness by training data enhancement without any modifications to pre-trained diffusion models. Nevertheless, inversion-based methods often suffer from limited effectiveness and inherent inaccuracies, while recent rectified-flow-based approaches struggle to preserve objects with accurate 3D geometry. In this paper, we propose DriveFlow, a Rectified Flow Adaptation method for training data enhancement in autonomous driving based on pre-trained Text-to-Image flow models. Based on frequency decomposition, DriveFlow introduces two strategies to adapt noise-free editing paths derived from text-conditioned velocities. 1) High-Frequency Foreground Preservation: DriveFlow incorporates a high-frequency alignment loss for foreground to maintain precise 3D object geometry. 2) Dual-Frequency Background Optimization: DriveFlow also conducts dual-frequency optimization for background, balancing editing flexibility and semantic consistency. Comprehensive experiments validate the effectiveness and efficiency of DriveFlow, demonstrating comprehensive performance improvements on all categories across OOD scenarios. Yiming Yang 0001, Chaoda Zheng, Yifan Zhang 0004, Shuaicheng Niu, Zilu Guo, Gui Gui, Shuguang Cui, Zhen Li 0026 |
AAAI | 2 |
| 2025 | Topo2Seq: Enhanced Topology Reasoning via Topology Sequence LearningabstractExtracting lane topology from perspective views (PV) is crucial for planning and control in autonomous driving. This approach extracts potential drivable trajectories for self-driving vehicles without relying on high-definition (HD) maps. However, the unordered nature and weak long-range perception of the DETR-like framework can result in misaligned segment endpoints and limited topological prediction capabilities. Inspired by the learning of contextual relationships in language models, the connectivity relations in roads can be characterized as explicit topology sequences. In this paper, we introduce Topo2Seq, a novel approach for enhancing topology reasoning via topology sequences learning. The core concept of Topo2Seq is a randomized order prompt-to-sequence learning between lane segment decoder and topology sequence decoder. The dual-decoder branches simultaneously learn the lane topology sequences extracted from the Directed Acyclic Graph (DAG) and the lane graph containing geometric information. Randomized order prompt-to-sequence learning extracts unordered key points from the lane graph predicted by the lane segment decoder, which are then fed into the prompt design of the topology sequence decoder to reconstruct an ordered and complete lane graph. In this way, the lane segment decoder learns powerful long-range perception and accurate topological reasoning from the topology sequence decoder. Notably, topology sequence decoder is only introduced during training and does not affect the inference efficiency. Experimental evaluations on the OpenLane-V2 dataset demonstrate the state-of-the-art performance of Topo2Seq in topology reasoning. Yiming Yang 0001, Yueru Luo, Bingkun He, Erlong Li, Zhipeng Cao 0002, Chao Zheng 0004, Shuqi Mei, Zhen Li 0026 |
AAAI | 1 |
| 2025 | VesSAM: Efficient Multi-Prompting for Segmenting Complex VesselabstractPrecise vessel segmentation is vital for clinical applications such as diagnosis and surgical planning but remains challenging due to thin, branching geometries and low texture contrast. Although foundation models such as the Segment Anything Model (SAM) show strong performance in general segmentation tasks, they remain suboptimal for vascular structures. In this work, we present VesSAM, a powerful and efficient framework tailored for 2D vessel segmentation. VesSAM integrates three core modules: a convolutional adapter that enhances local texture features, a multi-prompt encoder that fuses anatomical cues via hierarchical cross-attention, and a lightweight mask decoder that reduces jagged artifacts. We also introduce an automated pipeline to generate structured multi-prompt annotations, and curate a diverse benchmark dataset spanning 8 datasets across 5 imaging modalities. Extensive experiments show that VesSAM surpasses state-of-the-art PEFT-based SAM variants by over$\text{1 0 \%}$Dice and 13% IoU, while maintaining competitive accuracy to fully fine-tuned methods with far fewer parameters. VesSAM also generalizes well to out-of-distribution (OoD) settings, outperforming all baselines in average OoD Dice and IoU. Suzhong Fu, Jingqi Dong, Yiming Yang 0001, Yao Zhu 0003, Min Chang Jordan Ren, Delin Deng, Angelica I. Avilés-Rivero, Shuguang Cui, Zhen Li 0026 |
BIBM | 5 |
| 2025 | Extended Cross-Modality United Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised learning visible-infrared person re-identification (USL-VI-ReID) aims to learn modality-invariant features from unlabeled cross-modality data. However, existing approaches lack comprehensive cross-modality clustering or excessively pursue cluster-level association, which hinders reliable learning of modality-invariant features. To address these challenges, we propose an Extended Cross-Modality United Learning (ECUL) framework, which integrates Extended Modality-Camera Clustering (EMCC) and Two-Step Memory Updating Strategy (TSMem) modules. Specifically, we design ECUL to naturally unify intra-modality clustering, inter-modality clustering, and inter-modality instance selection, establishing compact and accurate cross-modality associations while reducing the introduction of noisy labels. Moreover, EMCC captures and filters neighborhood relationships by extending the encoding vector, which further promotes the learning of modality-invariant and camera-invariant knowledge in terms of the clustering algorithm. Finally, TSMem provides accurate and generalized proxy points for contrastive learning by updating memory in stages. Comprehensive experiments conducted on the SYSU-MM01 and RegDB datasets demonstrate that the proposed ECUL framework shows promising performance and even outperforms certain supervised methods. Ruixing Wu, Yiming Yang 0001, Jiakai He, Haifeng Hu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Part-Based Bi-Directional Enhancement Learning for Unsupervised Visible-Infrared Re-IdentificationabstractUnsupervised Learning Visible-Infrared Person Re-identification (USL-VI-ReID) aims to learn uniform feature representations for retrieving persons from unlabeled cross-modality data, which can accomplish 24-hour surveillance without expensive manual annotations. However, USL-VI-ReID is a cross-modality retrieval task that suffers from cross-modality label association and cross-modality feature discrepancy problems. To address these two problems, we propose a Part-based Bidirectional Enhancement (PBE) framework for learning a cross-modality uniform representation of USL-VI-ReID. The PBE consists of both forward and backward enhancements: 1) To alleviate the cross-modality label association problem, we propose a Part-based Label Forward Enhancement (PLFE) module. The PLFE module employs part features to complement global features during the label association process, thus generating higher-quality VI-associated pseudo-labels for the forward enhancement of the feature learning process. 2) To mitigate the cross-modality feature discrepancy problem, we propose a Part-based Feature Backward Enhancement (PFBE) module. The PFBE module utilizes part features to augment global features during the feature learning process, thus learning more robust uniform features for the backward enhancement of the label association process. Based on the part features, our PBE method achieves bi-directional enhancement during the label association and feature learning processes for robust recurrent training. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate that the proposed PBE framework outperforms existing USL-VI-ReID methods. Code is available at https://github.com/heqlin5/PBE. Qiaolin He, Yiming Yang 0001, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Progressive Cross-Modal Association Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identification (USL-VI-ReID) aims to explore the cross-modal associations and learn modality-invariant representations without manual labels. The field provides flexible and economical methods for person re-identification across light and dark scenes. Existing approaches utilize cluster-level strong association methods, such as graph matching and optimal transport, to correlate modal differences, which may result in mis-linking between clusters and introduce noise. To overcome this limitation and gradually acquire reliable cross-modal associations, we propose a Progressive Cross-modal Association Learning (PCAL) method for USL-VI-ReID. Specifically, our PCAL naturally integrates Triple-modal Adversarial Learning (TAL), Cross-modal Neighbor Expansion (CNE) and Modality-invariant Contrastive Learning (MCL) into a unified framework. TAL fully utilizes the advantage of Channel Augmented (CA) technique to reduce modal differences, which facilitates subsequent mining of cross-modal associations. Furthermore, we identify the modal bias problem in existing clustering methods, which hinders the effective establishment of cross-modal associations. To address this problem, CNE is proposed to balance the contribution of cross-modal neighbor information, linking potential cross-modal neighbors as much as possible. Finally, MCL is then introduced to refine the cross-modal associations and learn modality-invariant representations. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate the competitive performance of PCAL method. Code is available at https://github.com/YimingYang23/PCA USLVIReID. Yiming Yang 0001, Weipeng Hu, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Dynamic Modality-Camera-Invariant Clustering for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised learning visible-infrared person re-identification (USL-VI-ReID) offers a more flexible and cost-effective alternative compared to supervised methods. This field has gained increasing attention due to its promising potential. Existing methods simply cluster modality-specific samples and employ strong association techniques to achieve instance-to-cluster or cluster-to-cluster cross-modality associations. However, they ignore cross-camera differences, leading to noticeable issues with excessive splitting of identities. Consequently, this undermines the accuracy and reliability of cross-modal associations. To address these issues, we propose a novel dynamic modality-camera-invariant clustering (DMIC) framework for USL-VI-ReID. Specifically, our DMIC naturally integrates modality-camera-invariant expansion (MIE), dynamic neighborhood clustering (DNC), and hybrid modality contrastive learning (HMCL) into a unified framework, which eliminates both the cross-modality and cross-camera discrepancies in clustering. MIE fuses intermodal and intercamera distance coding to bridge the gaps between modalities and cameras at the clustering level. DNC employs two dynamic search strategies to refine the network's optimization objective, transitioning from improving discriminability to enhancing cross-modal and cross-camera generalizability. Moreover, HMCL is designed to optimize instance- and cluster-level distributions. Memories for intramodality and intermodality training are updated using randomly selected samples, facilitating real-time exploration of modality-invariant representations. Extensive experiments have demonstrated that our DMIC addresses the limitations present in current clustering approaches and achieves competitive performance, which significantly reduces the performance gap with supervised methods. Yiming Yang 0001, Weipeng Hu, Qiaolin He, Haifeng Hu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Unsupervised Visible-Infrared Person ReID via Modality-Camera Balance Label RefinementabstractUnsupervised Learning Visible-Infrared Person Re-Identification (USL-VI-ReID) focuses on developing a cross-modality retrieval model without the need for labels, minimizing the dependence on costly manual annotation across modalities. Recently, various approaches focus on reducing the cross-modality discrepancies. However, they ignore that USL-VI-ReID is also a task of solving discrepancies while exploring fine-grained information in hierarchical domains. In this article, we propose a hierarchical Modality-Camera Balance Label Refinement (MCBL) framework to balance the contributions of each camera-modality. Meanwhile, we explore the fine-grained features and refine the noise labels at each training stages. Specifically, our MCBL naturally combines Modality-Camera Balanced Label Mining (MBLM), Unreliable Pseudo-Label Re-align (UPR), and Hybrid Modality-Camera Contrastive Learning (HMCCL) into a unified framework, which balances the association information for each hierarchical domain through refining noise labels. Technically, MBLM filters cluster-level noise samples utilizing a modality-camera balance strategy, thereby ensuring that reliable samples are stored in memory for effective contrast learning. UPR refines the noise labels through the re-alignment methods at the instance level, thus improving the accuracy of labels and further enhancing the model’s generalization ability. Moreover, the key of HMCCL is optimizing the distribution at both the instance and cluster levels, which forces the sample to be close to its cluster proxy while being far from others in a real-time memory update phase. Extensive experiments have shown that our MCBL addresses the current limitations of camera discrepancy and achieves competitive performance. Jiakai He, Yiming Yang 0001, Haifeng Hu 0001, Ruixing Wu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Unsupervised NIR-VIS Face Recognition via Homogeneous-to-Heterogeneous Learning and Residual-Invariant EnhancementabstractNear-Infrared and Visible light (NIR-VIS) face recognition methods have achieved remarkable success in the fields of security surveillance, criminal investigation, and multimedia information retrieval. But the existing methods heavily rely on carefully annotated labels, leading to expensive manual labelling consumption and deployment flexibility. This motivates us to design unsupervised methods to address NIR-VIS recognition without relying on label information. To this end, we propose a novel homogeneous-to-HEterogeneous learning and Residual-invariant Enhancement (HERE) network for Unsupervised NIR-VIS Heterogeneous Face Recognition (NIR-VIS-UHFR). As the name suggests, the optimization of HERE follow a ”homogeneous-to-heterogeneous learning” strategy to fully explore complementary and common semantic information across different modalities. During the homogeneous learning phase, Modality-Adversarial Contrastive Learning (MACL) leverages the collaboration of modality contrastive learning and adversarial learning. On the one hand, MACL learns compact and discriminative intra-modal representations for NIR and VIS data, respectively. On the other hand, MACL guarantees that NIR-VIS data conform to the common feature distribution in a shared feature space, effectively reducing modal differences even in the absence of identity information between modalities. In the heterogeneous learning phase, K-reciprocal-Encoding-based Cross-modal Labeling (KECL) is introduced as robust pseudo label estimation to fully explore cross-modal relationships and group cross-modal features into clusters. With the pseudo labels provided by KECL, Refined cross-modal Contrastive Learning (RCL) is developed with modality-invariant averaging initialization and dynamic focus weighting strategies to extract modality-invariant features. Finally, Residual-invariant Representations Enhancement (RRE) mines partial features under the cross-modal face for robust matching. Compared to supervised methods, our unsupervised HERE demonstrates comparable performance on multiple datasets, greater scalability and practicality in deployment by reducing data acquisition requirements and costs. Yiming Yang 0001, Weipeng Hu, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Pseudo Label Association and Prototype-Based Invariant Learning for Semi-Supervised NIR-VIS Face RecognitionabstractRemarkable success of the existing Near-InfraRed and VISible (NIR-VIS) approaches owes to sufficient labeled training data. However, collecting and tagging data from different domains is a time-consuming and expensive task. In this paper, we tackle the NIR-VIS face recognition problem in a semi-supervised manner, termed as semi-supervised NIR-VIS Heterogeneous Face Recognition (NIR-VIS-sHFR). To cope with this problem, we propose a novel pseudo Label association and Prototype-based invariant Learning (LPL), consisting of three key components, i.e., Cross-domain pseudo Label Association (CLA), Intra-domain Compact Representation learning (ICR), and Prototype-based Inter-domain Invariant learning (PII). Firstly, the CLA iteratively builds inter-domain association graphs for pseudo-label association, subsequently facilitating cross-domain model development based on the generated pseudo-labels. Furthermore, the ICR is proposed to achieve the separation of in-domain features from different clusters and the aggregation of features from the same cluster, by performing cluster adaptation learning with prototype-based initialization. Finally, with the cross-domain pseudo-label training data produced by CLA, the PII explores potential domain-invariant and identity-related features, which employs cross-domain prototypes with identity-associated momentum updating to effectively guide inter-domain instances learning. The semi-supervised LPL method achieves comparable performance to recent supervised learning methods on multiple challenging NIR-VIS datasets, which demonstrates that the LPL is capable of learning robust cross-domain representations even without identity label information. Weipeng Hu, Yiming Yang 0001, Haifeng Hu 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | Syncretic Space Learning Network for NIR-VIS Face RecognitionabstractTo overcome the technical bottleneck of face recognition in low-light scenarios, Near-InfraRed and VISible (NIR-VIS) heterogeneous face recognition is proposed for matching well-lit VIS faces with poorly lit NIR faces. Current cross-modal synthesis methods visually convert the NIR modality to the VIS modality and then perform face matching in the VIS modality. However, using a heavyweight GAN network on unpaired NIR-VIS faces may lead to high synthesis difficulty, low inference efficiency, and other problems. To alleviate the above problems, we simultaneously synthesize NIR and VIS images into modality-independent syncretic images and propose a novel syncretic space learning (SSL) model to eliminate the modal gap. First, Syncretic Modality Generator (SMG) synthesizes NIR and VIS images into syncretic images using channel-level convolution with a shallow CNN. In particular, the discriminative structural information is well preserved and the face quality can be further improved with small modal variations in a self-supervised learning manner. Second, Modality-adversarial Syncretic space Learning (MSL) projects NIR and VIS images into the syncretic space by a syncretic-modality adversarial learning strategy with syncretic pattern guided objective, so the modal gap of NIR-VIS faces can be effectively reduced. Finally, the Syncretic Distribution Consistency (SDC) constructed by NIR-syncretic, syncretic-syncretic, and VIS-syncretic consistency can enhance the intra-class compactness and learn discriminative representations. Extensive experiments on three challenging datasets demonstrate the effectiveness of the SSL method. Yiming Yang 0001, Weipeng Hu, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Neutral Face Learning and Progressive Fusion Synthesis Network for NIR-VIS Face RecognitionabstractTo meet the strong demand for deploying face recognition systems in low-light scenarios, the Near-InfraRed and VISible (NIR-VIS) face recognition task is receiving increasing attention. However, heterogeneous faces have the characteristics of heterogeneity and non-neutrality. Heterogeneity refers to the fact that the matching images are in different modalities, and non-neutrality means that the matching images are significantly different in pose, expression, lighting, etc. Both situations pose challenges for NIR-VIS face matching. To address this problem, we propose a novel Neutral face Learning and Progressive Fusion synthesis (NLPF) network to disentangle the latent attributes of heterogeneous faces and learn neutral face representations. Our approach naturally integrates Identity-related Neutral face Learning (INL) and Attribute Progressive Fusion (APF) into a joint framework. Firstly, INL eliminates modal variations and residual variations by guiding the network to learn homogeneous neutral face feature representations, which tackles the challenge of heterogeneity and non-neutrality by mapping cross-modal images to a common neutral representation subspace. Besides, APF is presented to perform the disentanglement and reintegration of identity-related features, modality-related features and residual features in a progressive fusion manner, which helps to further purify identity-related features. Comprehensive evaluations are carried out on three mainstream NIR-VIS datasets to verify the robustness and effectiveness of the NLPF model. In particular, NLPF has competitive recognition performance on LAMP-HQ, the most challenging NIR-VIS dataset so far. Yiming Yang 0001, Weipeng Hu, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Robust Cross-Domain Pseudo-Labeling and Contrastive Learning for Unsupervised Domain Adaptation NIR-VIS Face RecognitionabstractNear-infrared and visible face recognition (NIR-VIS) is attracting increasing attention because of the need to achieve face recognition in low-light conditions to enable 24-hour secure retrieval. However, annotating identity labels for a large number of heterogeneous face images is time-consuming and expensive, which limits the application of the NIR-VIS face recognition system to larger scale real-world scenarios. In this paper, we attempt to achieve NIR-VIS face recognition in an unsupervised domain adaptation manner. To get rid of the reliance on manual annotations, we propose a novel Robust cross-domain Pseudo-labeling and Contrastive learning (RPC) network which consists of three key components, i.e., NIR cluster-based Pseudo labels Sharing (NPS), Domain-specific cluster Contrastive Learning (DCL) and Inter-domain cluster Contrastive Learning (ICL). Firstly, NPS is presented to generate pseudo labels by exploring robust NIR clusters and sharing reliable label knowledge with VIS domain. Secondly, DCL is designed to learn intra-domain compact yet discriminative representations. Finally, ICL dynamically combines and refines intrinsic identity relationships to guide the instance-level features to learn robust and domain-independent representations. Extensive experiments are conducted to verify an accuracy of over 99% in pseudo label assignment and the advanced performance of RPC network on four mainstream NIR-VIS datasets. Yiming Yang 0001, Weipeng Hu, Haiqi Lin, Haifeng Hu 0001 |
IEEE Trans. Image Process. | 1 |