Xuefeng Tao

dblp:250/1863 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 One-step multiview anchor graph clustering via semantic alignment
Jun Kong 0001, Min Jiang 0008, Xuefeng Tao
Neurocomputing4
2026 Unsupervised Person Re-Identification With Diffusion Model via Semantic-Aware Disentanglement Representation Learning
abstract
Unsupervised person re-identification (Re-ID) requires learning semantic representation without identity labels. Existing methods entangle identity-related person features with camera-related background features, hindering discriminative feature learning. Also, these methods often disrupt the semantic structure of the person, weakening the semantic representation. In this paper, we propose the Semantic-Aware Disentanglement Representation Learning (SDRL) framework with diffusion models for unsupervised person Re-ID. Firstly, to enhance feature learning, we propose the Disentanglement Aggregation Model (DAM). This model disentangles identity-related features from camera-related features to generate multi-view features. Secondly, to promote the consistency of multi-view features, we design the multi-view similarity consistency (MSC) loss to constrain intra-camera and cross-camera similarity distributions. Thirdly, to generate semantically meaningful patches, we propose the Semantic Spatial Diffusion Model (SSDM). This model operates on identity-related features to perform the denoising diffusion process over spatial transformer parameters. Finally, to further enhance the semantic representation of generated patches, we design the Semantic Decoupled Contrastive (SDC) loss to perceive the inherent semantic structure. Numerous experiments on three demanding datasets prove that our approach is superior to the current unsupervised Re-ID approaches. The source code will be publicly available at https://github.com/taoxuefong/SDRL-reid.
Xuefeng Tao, Jun Kong 0001, Min Jiang 0008, Jiayi Li 0003, Ajmal Mian
IEEE Trans. Circuits Syst. Video Technol.1
2025 AU-Net: Adaptive Unified Network for Joint Multi-Modal Image Registration and Fusion
abstract
Joint multi-modal image registration and fusion (JMIRF) typically follows a register-first, fuse-later paradigm. It has a registration module to align parallax images and a fusion module to fuse registered images. Existing research typically focuses on the mutual enhancement between the two modules, but this is essentially a straightforward combination rather than an efficient, unified network. Moreover, executing the two modules separately may cause inefficiency, as the total runtime is merely the sum of both steps without investigating potential shared structures. In this paper, we propose an Adaptive Unified Network (AU-Net) following a novel end-to-end paradigm called Feature-Level Joint Training (FLJT). Firstly, AU-Net learns registration and fusion within a unified network through shared structure and hierarchical semantic interaction. A multi-level dynamic fusion module is designed to adaptively fuse input features from different scales and modalities. Secondly, the image-to-image translation based on Denoising Diffusion Probabilistic Models (DDPMs) is introduced to train AU-Net using simple and reliable single-modal metrics. Unlike previous unidirectional translation, we explore bidirectional translation to provide additional implicit branch supervision. Furthermore, a cache-like scheme is proposed to elegantly circumvent the additional computational overhead caused by the iterative denoising of DDPMs. Finally, our method was validated on two publicly available datasets, demonstrating advantages over state-of-the-art methods in terms of qualitative evaluation, quantitative evaluation, and computational complexity analysis. The code will be publically available at https://github.com/luming1314/AU-Net.
Ming Lu 0008, Min Jiang 0008, Xuefeng Tao, Jun Kong 0001
IEEE Trans. Image Process.3
2025 OAFTracker: One-Stage Associative Multiple Object Tracking With Fine-Grained Orthogonal Representation
abstract
Multiple object tracking based on the tracking-by-detection paradigm relies on appearance information and motion information for trajectory association. Employing global re-identification features and two-stage association strategies can improve the utilization of both types of information for detections with different confidence scores. However, when targets are occluded, coarse-grained global representations can lead to false positive detections. Additionally, two-stage association strategies tend to prioritize matching high-confidence detections over more accurate low-confidence detections, leading to identity switch problems. To address these issues, we propose the OAFTracker framework, which focuses on local representations and a one-stage association strategy. Firstly, a Fine-grained Representation Orthogonal Fusion (FROF) network is designed to adaptively integrate local and global representations. Secondly, we propose a One-stage Association Matching (OAM) strategy. This strategy combines multiple distance constraints to ensure fairness in matching detections with different confidence scores to predicted trajectories. Additionally, we propose an Adaptive Variable Noise (AVN) Kalman filtering algorithm to dynamically update the state of predicted trajectories. Finally, extensive experiments conducted on two public datasets demonstrate the effectiveness of the OAFTracker method.
Jun Kong 0001, Min Jiang 0008, Xuefeng Tao
IEEE Trans. Multim.4
2024 NAORL: Network Feature Aware Offline Reinforcement Learning for Real Time Bandwidth Estimation
abstract
Bandwidth Estimation(BWE) is the most important and challenging problem for Real Time Communication(RTC) systems. The rule-based BWE is designed with hand-crafted rules, which mainly depend on human knowledge, and therefore is difficult to generalize to unknown scenarios. Learning-based BWE algorithms, especially online reinforcement learning-based algorithms, are proposed to explore new decisions in complex network environments adaptively. However, these algorithms require frequent interactions with the environment, which would cause catastrophic experience for RTC users.
Wei Zhang 0074, Xuefeng Tao
MMSys2
2024 Patch-based tendency camera multi-constraint learning for unsupervised person re-identification
Xuefeng Tao, Jun Kong 0001, Min Jiang 0008, Tianshan Liu
J. Vis. Commun. Image Represent.1
2024 Semantic Camera Self-Aware Contrastive Learning for Unsupervised Vehicle Re-Identification
abstract
Unsupervised vehicle re-identification (ReID) aims to retrieve vehicle images from different cameras without using identity labels. Patch features, which capture fine-grained semantic information of vehicles, are crucial for ReID. However, existing methods often fail to preserve the discriminative semantic structure of vehicles due to the non-uniformity of feature attributes across patches. Moreover, domain discrepancy among cameras also requires attention, as it can cause large intra-class variance and noisy clustering results. To tackle these problems, in this letter, we propose a novel Semantic Camera Self-Aware Contrastive Learning (SCSCL) framework for unsupervised vehicle ReID. Firstly, we design the Semantic Self-Aware Contrastive (SSC) loss to perceive the semantic attributes of vehicle images from spatial transformer parameters, thereby enhancing the semantic representation of patch features. Secondly, we design the Camera Self-Aware Contrastive (CSC) loss to perceive the cross-camera distance distributions to facilitate the exploration of instance constraints, thereby enabling cross-camera clustering-friendly representations. Finally, extensive experimental results on VeRi-776 and VehicleID datasets attest to the efficacy of our method over the state-of-the-art performance.
Xuefeng Tao, Jun Kong 0001, Min Jiang 0008
IEEE Signal Process. Lett.1
2024 Unsupervised Learning of Intrinsic Semantics With Diffusion Model for Person Re-Identification
abstract
Unsupervised person re-identification (Re-ID) aims to learn semantic representations for person retrieval without using identity labels. Most existing methods generate fine-grained patch features to reduce noise in global feature clustering. However, these methods often compromise the discriminative semantic structure and overlook the semantic consistency between the patch and global features. To address these problems, we propose a Person Intrinsic Semantic Learning (PISL) framework with diffusion model for unsupervised person Re-ID. First, we design the Spatial Diffusion Model (SDM), which performs a denoising diffusion process from noisy spatial transformer parameters to semantic parameters, enabling the sampling of patches with intrinsic semantic structure. Second, we propose the Semantic Controlled Diffusion (SCD) loss to guide the denoising direction of the diffusion model, facilitating the generation of semantic patches. Third, we propose the Patch Semantic Consistency (PSC) loss to capture semantic consistency between the patch and global features, refining the pseudo-labels of global features. Comprehensive experiments on three challenging datasets show that our method surpasses current unsupervised Re-ID methods. The source code will be publicly available at https://github.com/taoxuefong/Diffusion-reid.
Xuefeng Tao, Jun Kong 0001, Min Jiang 0008, Ming Lu 0008, Ajmal Mian
IEEE Trans. Image Process.1
2024 Learning Semantic Polymorphic Mapping for Text-Based Person Retrieval
abstract
Text-Based Person Retrieval (TBPR) aims to identify a particular individual within an extensive image gallery using text as the query. The principal challenge inherent in the TBPR task revolves around how to map cross-modal information to a potential common space and learn a generic representation. Previous methods have primarily focused on aligning singular text-image pairs, disregarding the inherent polymorphism within both images and natural language expressions for the same individual. Moreover, these methods have also ignored the impact of semantic polymorphism-based intra-modal data distribution on cross-modal matching. Recent methods employ cross-modal implicit information reconstruction to enhance inter-modal connections. However, the process of information reconstruction remains ambiguous. To address these issues, we propose the Learning Semantic Polymorphic Mapping (LSPM) framework, facilitated by the prowess of pre-trained cross-modal models. Firstly, to learn cross-modal information representations with better robustness, we design the Inter-modal Information Aggregation (Inter-IA) module to achieve cross-modal polymorphic mapping, fortifying the foundation of our information representations. Secondly, to attain a more concentrated intra-modal information representation based on semantic polymorphism, we design Intra-modal Information Aggregation (Intra-IA) module to further constrain the embeddings. Thirdly, to further explore the potential of cross-modal interactions within the model, we design the implicit reasoning module, Masked Information Guided Reconstruction (MIGR), with constraint guidance to elevate overall performance. Extensive experiments on both CUHK-PEDES and ICFG-PEDES datasets show that we achieve state-of-the-art results on Rank-1, mAP and mINP compared to existing methods.
Jiayi Li 0003, Min Jiang 0008, Jun Kong 0001, Xuefeng Tao
IEEE Trans. Multim.4
2024 Hierarchical Camera-Aware Contrast Extension for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (Re-ID) targets to learn discriminative representations without annotations. Recently, clustering-based methods have shown promising performance, which utilize clustering to generate identity pseudo labels for model optimization. Large intra-class variance mainly caused by domain discrepancy among cameras could lead to noisy clustering results. However, abundant camera-aware sample pairs relations have not been exploited fully to facilitate learning of features with comprehensive knowledge, so as to tackle this issue. In this paper, we propose hierarchical camera-aware contrast extension (HCACE) for unsupervised person Re-ID. Firstly, cognitive collaboration contrast scheme (CCCS) is introduced to explore hierarchical camera-aware relations at the proxy-level, so as to collaboratively promote model to learn representative knowledge. Secondly, aggregative instance contrast extension scheme (AICES) is proposed to promote the learning of potential fine-grained knowledge by aggregating refined camera-aware inter-instance relations. Especially in AICES, hard negative instance extension (HNIE) is designed to generate extended negative instances, so as to assist the exploration of transitional cross-camera inter-instance relations. Finally, extensive experiments on three benchmark datasets validate superior performance of proposed HCACE.
Min Jiang 0008, Jun Kong 0001, Xuefeng Tao
IEEE Trans. Multim.4
2023 Disentangled representation learning for collaborative filtering based on hyperbolic geometry
Meicheng Zhang, Min Jiang 0008, Xuefeng Tao, Jun Kong 0001
Knowl. Based Syst.3
2023 Weakly Supervised Distribution Discrepancy Minimization Learning With State Information for Person Re-Identification
abstract
Weakly supervised person re-identification (Re-ID) is appealing to handle real-world tasks by using state information that is available without manual annotation. At present, most methods perform unsupervised cross domain (UCD) learning by transferring the knowledge from the labeled source domain to the unlabeled target domain, which results in poor performance due to the severe shift. To address this problem, in this paper, we utilize the tracklet and camera information as weak supervision to propose a distribution discrepancy minimization learning (DDML) model for UCD person Re-ID. In addition to aligning data distributions from the perspective of domain adaptation learning, two losses are developed from the view of neighborhood invariance exploration to optimize matching results. Specifically, to bridge the gap between domains, we propose a camera-distribution-based (CDB) loss to align pair-wise distance distributions. Furthermore, to alleviate the biased search within the target domain, we propose a ranking-confidence-based (RCB) loss to perform the mined neighborhood for intra-camera and inter-camera separately to explore a high degree of confidence neighbor relations. Extensive experiments on three challenging datasets demonstrate that applying our method to unlabeled target domain outperforms current weakly supervised methods for person Re-ID.
Jun Kong 0001, Xuefeng Tao, Min Jiang 0008, Tianshan Liu
IEEE Trans. Multim.2
2022 Unsupervised Domain Adaptation by Multi-Loss Gap Minimization Learning for Person Re-Identification
abstract
Unsupervised domain adaptation (UDA) person re-identification (ReID) faces enormous challenges due to the severe shift between the source and target domains, as well as the dramatic variations within the target domain. In this paper, to address these issues, we propose a multi-loss gap minimization learning (MGML) approach for UDA person ReID. Firstly, we introduce the part model to learn discriminative patch features and design a Patch-based Part Ignoring (PPI) loss to select reliable instances for the efficient learning of the part model. Then, given the gap that typically occurs because of the inter-domain shift and intra-domain variations, a Gap-based Minimum Camera Discrepancy (G-MCD) loss is proposed. Specifically, in terms of the inter-domain, we propose to leverage the tracklet and camera information to label each distance vector, and accordingly align pair-wise distance distributions to bridge the inter-domain gap. As for the intra-domain, to alleviate the biased search, we propose to perform the mined neighborhood for intra-camera and inter-camera separately to optimize matching results by exploring neighborhood relations more deeply. Finally, experimental results on three challenging datasets demonstrate that applying our method to unlabeled target domain outperforms current UDA methods for person ReID.
Xuefeng Tao, Jun Kong 0001, Min Jiang 0008, Tianshan Liu
IEEE Trans. Circuits Syst. Video Technol.1
2020 CCAN: Constraint Co-Attention Network for Instance Grasping
abstract
Instance grasping is a challenging robotic grasping task when a robot aims to grasp a specified target object in cluttered scenes. In this paper, we propose a novel end-to-end instance grasping method using only monocular workspace and query images, where the workspace image includes several objects and the query image only contains the target object. To effectively extract discriminative features and facilitate the training process, a learning-based method, referred to as Constraint Co-Attention Network (CCAN), is proposed which consists of a constraint co-attention module and a grasp affordance predictor. An effective co-attention module is presented to construct the features of a workspace image from the extracted features of the query image. By introducing soft constraints into the co-attention module, it highlights the target object's features while trivializes other objects' features in the workspace image. Using the features extracted from the co-attention module, the cascaded grasp affordance interpreter network only predicts the grasp configuration for the target object. The training of the CCAN is totally based on simulated self-supervision. Extensive qualitative and quantitative experiments show the effectiveness of our method both in simulated and real-world environments even for totally unseen objects.
Junhao Cai, Xuefeng Tao
ICRA2