VLDB 2026 Research / reviewers in the wild / expert
Min Jiang 0008
dblp:35/994-8
· DBLP profile ↗
54ranked-venue papers
10as first author
44since 2021 · last 2026
0000-0003-3826-6405ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 9 first-author · 33 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Strengthening temporal action segmentation through diffusion models
Danfeng Zhuang, Min Jiang 0008, Hichem Arioui, Hedi Tabia |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | One-step multiview anchor graph clustering via semantic alignment
Jun Kong 0001, Min Jiang 0008, Xuefeng Tao |
Neurocomputing | 3 |
| 2026 | LLM guided counterfactual reasoning for zero-shot knowledge based visual question answering
Zhuhan Zhang, Min Jiang 0008, Jun Kong 0001, Jiayi Li 0003 |
Neurocomputing | 2 |
| 2026 | Auto-Feedback Semantic Interface Learning for Text-Based Person Retrieval
Jiayi Li 0003, Min Jiang 0008, Jun Kong 0001, Saeed Anwar, Ajmal Mian |
Pattern Recognit. | 2 |
| 2026 | Sample compression-based anchor strategy for tensor multi-view clustering
Jun Kong 0001, Min Jiang 0008, Yaozu Kan |
Pattern Recognit. | 3 |
| 2026 | Confidence-Guided Latent Chain-of-Thought Reasoning for Text-Based Person Retrieval
Min Jiang 0008, Jun Kong 0001 |
IEEE Signal Process. Lett. | 2 |
| 2026 | SPCL: Semantic Polymorphism and Commonality Learning for Text-Based Person RetrievalabstractText-Based Person Retrieval (TBPR) refers to identifying a specific target pedestrian image based on natural language descriptions. Most previous methods rely on one-to-one alignment between paired text-image data, ignoring the polymorphic nature of visual and linguistic information. Moreover, constrained by ID, earlier methods have shown limited exploration of intra-individual and inter-individual relations. This limitation confines them to exploring characteristics within individuals, making it challenging to uncover commonalities and invariants that extend across IDs (e.g., attributes). Recently, due to the lack of accurate annotations, exploring attribute-based cross-modal interactions and alignments has become a significant challenge in TBPR. To address these issues, we propose a Semantic Polymorphism and Commonality Learning (SPCL) framework. First, we present Relation-Sensitive Semantic Polymorphism Alignment (RSSPA) and ID-Based Semantic Polymorphism Alignment (IBSPA) to explore ID-limited Feature Redistribution. Second, we transcend the constraints of ID, leveraging ID-Free Attribute Alignment (IFAA) from a macro perspective to explore commonalities and invariants based on attribute features. Finally, from a micro perspective, we design Attribute Prior Fusion Reconstruction (APFR) to optimize the attention of our model, exploring the positive impact of attribute priors on cross-modal interaction. Experiments on CUHK-PEDES, ICFG-PEDES and RSTPReid show that our method achieves state-of-the-art performance on Rank-1, mAP and mINP. Jiayi Li 0003, Jun Kong 0001, Yunde Zhang, Ming Lu 0008, Min Jiang 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Unsupervised Person Re-Identification With Diffusion Model via Semantic-Aware Disentanglement Representation LearningabstractUnsupervised person re-identification (Re-ID) requires learning semantic representation without identity labels. Existing methods entangle identity-related person features with camera-related background features, hindering discriminative feature learning. Also, these methods often disrupt the semantic structure of the person, weakening the semantic representation. In this paper, we propose the Semantic-Aware Disentanglement Representation Learning (SDRL) framework with diffusion models for unsupervised person Re-ID. Firstly, to enhance feature learning, we propose the Disentanglement Aggregation Model (DAM). This model disentangles identity-related features from camera-related features to generate multi-view features. Secondly, to promote the consistency of multi-view features, we design the multi-view similarity consistency (MSC) loss to constrain intra-camera and cross-camera similarity distributions. Thirdly, to generate semantically meaningful patches, we propose the Semantic Spatial Diffusion Model (SSDM). This model operates on identity-related features to perform the denoising diffusion process over spatial transformer parameters. Finally, to further enhance the semantic representation of generated patches, we design the Semantic Decoupled Contrastive (SDC) loss to perceive the inherent semantic structure. Numerous experiments on three demanding datasets prove that our approach is superior to the current unsupervised Re-ID approaches. The source code will be publicly available at https://github.com/taoxuefong/SDRL-reid. Xuefeng Tao, Jun Kong 0001, Min Jiang 0008, Jiayi Li 0003, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Mitigating Inherent Bias of Answer Heuristic Based Frameworks in Knowledge-Based Visual Question AnsweringabstractKnowledge-based Visual Question Answering (KBVQA) aims to utilize external knowledge to answer image-related questions. Current KBVQA frameworks prompt large language models (LLMs) with answer heuristics, narrowing the focus to more relevant answers. However, these answer heuristic based frameworks exhibit the bias towards overemphasizing highest-scoring answers, neglecting potential answers with lower scores. This bias arises from the deficiencies of in-context example constructions and underperformances of multi-query ensemble strategy. In this paper, we propose MinBias, an approach designed tomitigateinherentbiasof answer heuristic based VQA framework. Firstly, to help LLMs learn dialectically, we propose Dialectical Learning Space Construction strategy (DLSC). This module mitigates bias via guaranteeing diversity in the selected examples. Secondly, to strengthen the connection between image captions and questions, we propose Large Language Model guided Visual Masking (LVM) algorithm. This module mitigates bias via enhancing visual clues in example content. Finally, we propose Visual Entailment based Answer Rerank module (VEAR). This module mitigates the bias arising from multi-query strategy via obtaining unbiased evaluations. Extensive experiments demonstrate that our method mitigates the bias of the answer heuristic based VQA framework and enhances the model performance. Min Jiang 0008, Jun Kong 0001, Danfeng Zhuang, Ming Lu 0008 |
IEEE Trans. Multim. | 2 |
| 2025 | Learning Distinct Semantic Information in Modality-Specific and Modality-Shared Features for Visible-Infrared Person Re-identification
Min Jiang 0008, Huang Luo, Jun Kong 0001, Ming Lu 0008 |
PRCV (17) | 1 |
| 2025 | MAD-DGTD: Multivariate time series Anomaly Detection based on Dynamic Graph structure learning with Time Delay
Jun Kong 0001, Meicheng Zhang, Min Jiang 0008, Tianshan Liu |
Neurocomputing | 4 |
| 2025 | Vari-FocalFuse: Sparse Attention With Adaptive Attention Spans for Image FusionabstractExisting image fusion methods introduce convolution or attention to perceive multi-scale targets in different source images. However, fixed operation spans, including convolution spans and attention spans, result in insufficient or redundant perception. In this work, we propose Vari-FocalFuse, which performs sparse attention with adaptive attention spans. Firstly, to perform local perception and capture intricate details, we propose Focal Attention (FAtten). It localizes attention spans to the nearest neighbor tokens like convolutions, modeling spatial local correlation. Secondly, we adapt the attention spans to multi-scale targets, and propose Memorization Content Awareness (MCAware). The attention spans are expanded until they can cover the corresponding targets. Thirdly, to perform global perception and minimize redundant perception, we propose Vari-Focal Attention (VFAtten) by combining FAtten and MCAware. The adaptive attention spans can capture multi-scale targets without taking irrelevant information into account. Finally, to mitigate noise caused by redundant perception, we propose Sandwich GLU (SwGLU). Spatial perception refined gating process is utilized to mitigate redundant information. Extensive experiments demonstrate that our Vari-FocalFuse achieves state-of-the-art (SOTA) performance on various image fusion tasks. Shengchen Zhu, Min Jiang 0008, Jun Kong 0001, Ming Lu 0008, Jiayi Li 0003 |
IEEE Signal Process. Lett. | 2 |
| 2025 | AU-Net: Adaptive Unified Network for Joint Multi-Modal Image Registration and FusionabstractJoint multi-modal image registration and fusion (JMIRF) typically follows a register-first, fuse-later paradigm. It has a registration module to align parallax images and a fusion module to fuse registered images. Existing research typically focuses on the mutual enhancement between the two modules, but this is essentially a straightforward combination rather than an efficient, unified network. Moreover, executing the two modules separately may cause inefficiency, as the total runtime is merely the sum of both steps without investigating potential shared structures. In this paper, we propose an Adaptive Unified Network (AU-Net) following a novel end-to-end paradigm called Feature-Level Joint Training (FLJT). Firstly, AU-Net learns registration and fusion within a unified network through shared structure and hierarchical semantic interaction. A multi-level dynamic fusion module is designed to adaptively fuse input features from different scales and modalities. Secondly, the image-to-image translation based on Denoising Diffusion Probabilistic Models (DDPMs) is introduced to train AU-Net using simple and reliable single-modal metrics. Unlike previous unidirectional translation, we explore bidirectional translation to provide additional implicit branch supervision. Furthermore, a cache-like scheme is proposed to elegantly circumvent the additional computational overhead caused by the iterative denoising of DDPMs. Finally, our method was validated on two publicly available datasets, demonstrating advantages over state-of-the-art methods in terms of qualitative evaluation, quantitative evaluation, and computational complexity analysis. The code will be publically available at https://github.com/luming1314/AU-Net. Ming Lu 0008, Min Jiang 0008, Xuefeng Tao, Jun Kong 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | OAFTracker: One-Stage Associative Multiple Object Tracking With Fine-Grained Orthogonal RepresentationabstractMultiple object tracking based on the tracking-by-detection paradigm relies on appearance information and motion information for trajectory association. Employing global re-identification features and two-stage association strategies can improve the utilization of both types of information for detections with different confidence scores. However, when targets are occluded, coarse-grained global representations can lead to false positive detections. Additionally, two-stage association strategies tend to prioritize matching high-confidence detections over more accurate low-confidence detections, leading to identity switch problems. To address these issues, we propose the OAFTracker framework, which focuses on local representations and a one-stage association strategy. Firstly, a Fine-grained Representation Orthogonal Fusion (FROF) network is designed to adaptively integrate local and global representations. Secondly, we propose a One-stage Association Matching (OAM) strategy. This strategy combines multiple distance constraints to ensure fairness in matching detections with different confidence scores to predicted trajectories. Additionally, we propose an Adaptive Variable Noise (AVN) Kalman filtering algorithm to dynamically update the state of predicted trajectories. Finally, extensive experiments conducted on two public datasets demonstrate the effectiveness of the OAFTracker method. Jun Kong 0001, Min Jiang 0008, Xuefeng Tao |
IEEE Trans. Multim. | 3 |
| 2024 | MIAFusion: Infrared and Visible Image Fusion via Multi-scale Spatial and Channel-Aware Interaction Attention
Teng Lin, Ming Lu 0008, Min Jiang 0008, Jun Kong 0001 |
PRCV (8) | 3 |
| 2024 | Patch-based tendency camera multi-constraint learning for unsupervised person re-identification
Xuefeng Tao, Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Text-Guided Prototype Generation for Occluded Person Re-IdentificationabstractOccluded person re-identification (ReID) focuses on identifying persons who are partially occluded, especially in multi-camera scenarios. The majority of methods employ the background to make artificial occlusions. However, simple artificial occlusions could not effectively simulate real-world occluded scenarios, due to its lack of semantic information and its limitation in disrupting the model's attention. In this paper, we present the Text-Guided Prototype Generation (TGPG) for occluded person ReID. On the one hand, to fully employ the potential of text as priori information, the Mask Prototype Generation (MPG) strategy is presented to generate the prototypes that could capture attention in the pretrained model, similar to the realistic occlusions. On the other hand, to create a relationship between holistic person features and occluded person features, the Intra-modality Spatial Consistency (ISC) loss is introduced, enhancing the consistency and representativeness of the generated mask prototypes. Comprehensive experiments conducted on the Occluded-Duke and Occluded-ReID datasets confirm our method's superiority over state-of-the-art approaches. Min Jiang 0008, Jun Kong 0001 |
IEEE Signal Process. Lett. | 1 |
| 2024 | Semantic Camera Self-Aware Contrastive Learning for Unsupervised Vehicle Re-IdentificationabstractUnsupervised vehicle re-identification (ReID) aims to retrieve vehicle images from different cameras without using identity labels. Patch features, which capture fine-grained semantic information of vehicles, are crucial for ReID. However, existing methods often fail to preserve the discriminative semantic structure of vehicles due to the non-uniformity of feature attributes across patches. Moreover, domain discrepancy among cameras also requires attention, as it can cause large intra-class variance and noisy clustering results. To tackle these problems, in this letter, we propose a novel Semantic Camera Self-Aware Contrastive Learning (SCSCL) framework for unsupervised vehicle ReID. Firstly, we design the Semantic Self-Aware Contrastive (SSC) loss to perceive the semantic attributes of vehicle images from spatial transformer parameters, thereby enhancing the semantic representation of patch features. Secondly, we design the Camera Self-Aware Contrastive (CSC) loss to perceive the cross-camera distance distributions to facilitate the exploration of instance constraints, thereby enabling cross-camera clustering-friendly representations. Finally, extensive experimental results on VeRi-776 and VehicleID datasets attest to the efficacy of our method over the state-of-the-art performance. Xuefeng Tao, Jun Kong 0001, Min Jiang 0008 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Unsupervised Learning of Intrinsic Semantics With Diffusion Model for Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) aims to learn semantic representations for person retrieval without using identity labels. Most existing methods generate fine-grained patch features to reduce noise in global feature clustering. However, these methods often compromise the discriminative semantic structure and overlook the semantic consistency between the patch and global features. To address these problems, we propose a Person Intrinsic Semantic Learning (PISL) framework with diffusion model for unsupervised person Re-ID. First, we design the Spatial Diffusion Model (SDM), which performs a denoising diffusion process from noisy spatial transformer parameters to semantic parameters, enabling the sampling of patches with intrinsic semantic structure. Second, we propose the Semantic Controlled Diffusion (SCD) loss to guide the denoising direction of the diffusion model, facilitating the generation of semantic patches. Third, we propose the Patch Semantic Consistency (PSC) loss to capture semantic consistency between the patch and global features, refining the pseudo-labels of global features. Comprehensive experiments on three challenging datasets show that our method surpasses current unsupervised Re-ID methods. The source code will be publicly available at https://github.com/taoxuefong/Diffusion-reid. Xuefeng Tao, Jun Kong 0001, Min Jiang 0008, Ming Lu 0008, Ajmal Mian |
IEEE Trans. Image Process. | 3 |
| 2024 | Learning Semantic Polymorphic Mapping for Text-Based Person RetrievalabstractText-Based Person Retrieval (TBPR) aims to identify a particular individual within an extensive image gallery using text as the query. The principal challenge inherent in the TBPR task revolves around how to map cross-modal information to a potential common space and learn a generic representation. Previous methods have primarily focused on aligning singular text-image pairs, disregarding the inherent polymorphism within both images and natural language expressions for the same individual. Moreover, these methods have also ignored the impact of semantic polymorphism-based intra-modal data distribution on cross-modal matching. Recent methods employ cross-modal implicit information reconstruction to enhance inter-modal connections. However, the process of information reconstruction remains ambiguous. To address these issues, we propose the Learning Semantic Polymorphic Mapping (LSPM) framework, facilitated by the prowess of pre-trained cross-modal models. Firstly, to learn cross-modal information representations with better robustness, we design the Inter-modal Information Aggregation (Inter-IA) module to achieve cross-modal polymorphic mapping, fortifying the foundation of our information representations. Secondly, to attain a more concentrated intra-modal information representation based on semantic polymorphism, we design Intra-modal Information Aggregation (Intra-IA) module to further constrain the embeddings. Thirdly, to further explore the potential of cross-modal interactions within the model, we design the implicit reasoning module, Masked Information Guided Reconstruction (MIGR), with constraint guidance to elevate overall performance. Extensive experiments on both CUHK-PEDES and ICFG-PEDES datasets show that we achieve state-of-the-art results on Rank-1, mAP and mINP compared to existing methods. Jiayi Li 0003, Min Jiang 0008, Jun Kong 0001, Xuefeng Tao |
IEEE Trans. Multim. | 2 |
| 2024 | Hierarchical Camera-Aware Contrast Extension for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) targets to learn discriminative representations without annotations. Recently, clustering-based methods have shown promising performance, which utilize clustering to generate identity pseudo labels for model optimization. Large intra-class variance mainly caused by domain discrepancy among cameras could lead to noisy clustering results. However, abundant camera-aware sample pairs relations have not been exploited fully to facilitate learning of features with comprehensive knowledge, so as to tackle this issue. In this paper, we propose hierarchical camera-aware contrast extension (HCACE) for unsupervised person Re-ID. Firstly, cognitive collaboration contrast scheme (CCCS) is introduced to explore hierarchical camera-aware relations at the proxy-level, so as to collaboratively promote model to learn representative knowledge. Secondly, aggregative instance contrast extension scheme (AICES) is proposed to promote the learning of potential fine-grained knowledge by aggregating refined camera-aware inter-instance relations. Especially in AICES, hard negative instance extension (HNIE) is designed to generate extended negative instances, so as to assist the exploration of transitional cross-camera inter-instance relations. Finally, extensive experiments on three benchmark datasets validate superior performance of proposed HCACE. Min Jiang 0008, Jun Kong 0001, Xuefeng Tao |
IEEE Trans. Multim. | 2 |
| 2023 | Action Text Diffusion Prior Network for Action SegmentationabstractAction segmentation is a challenging task that requires accurate parsing and labeling of each action. There are two types of methods for action segmentation. The first type primarily focuses on extracting high-quality features from videos, while the second type focuses on combining textual and perceptual features through multimodal fusion. However, both types of methods have their limitations. The first type is limited to a single modality and does not leverage multimodal information, while the second type, although promising, is restricted by the language used to describe the actions in the texts. To solve these problems, we propose in this paper, an Action Text Diffusion Prior Network (ATDPN) which simultaneously improves the quality of the extracted visual features (by introducing a Video-level Diffusion Prior Sampling) and integrates the textual information to fullest extent. This leads to superior action segmentation results. Our experiments performed on GTEA dataset demonstrate the effective feature extraction ability of ATDPN. Danfeng Zhuang, Min Jiang 0008, Hichem Arioui, Hedi Tabia |
CBMI | 2 |
| 2023 | Disentangled representation learning for collaborative filtering based on hyperbolic geometry
Meicheng Zhang, Min Jiang 0008, Xuefeng Tao, Jun Kong 0001 |
Knowl. Based Syst. | 2 |
| 2023 | Multi-stream ternary enhanced graph convolutional network for skeleton-based action recognition
Jun Kong 0001, Shengquan Wang, Min Jiang 0008, Tianshan Liu |
Neural Comput. Appl. | 3 |
| 2023 | Participants-based Synchronous Optimization Network for skeleton-based action recognition
Danfeng Zhuang, Min Jiang 0008, Jun Kong 0001 |
Pattern Recognit. Lett. | 2 |
| 2023 | Time-to-space progressive network using overlap skeleton contexts for action recognition
Danfeng Zhuang, Min Jiang 0008, Jun Kong 0001 |
Signal Process. | 2 |
| 2023 | CALTracker: Cross-Task Association Learning for Multiple Object TrackingabstractMultiple object tracking has recently achieved excellent performance based on the joint optimization of detection and re-identification tasks. However, joint optimization normally homogenizes the features in the detection and re-identification tasks, which weakens the representation of each task's inherent geometric and semantic information. Additionally, the stability of the tracking trajectory will be impacted by the feature misalignment of the associated information between different tasks. In this letter, we propose a Cross-task Association Learning Tracker (CALTracker) to trade off the inherent and associated information. We first design a Triplet Shrinkage Decoupling (TSD) module to ensure the independence of sub-task features in the optimization process, thereby minimizing the optimization conflicts caused by homogeneous features. Secondly, to improve the consistent representation of the associated information between subtasks, a Double Attention Cross-task Learning (DACL) strategy is designed to achieve cross-task feature alignment and mutual gain. Finally, extensive experimental results on MOT17 and MOT20 demonstrate the effectiveness of the proposed method over the state-of-the-art performance. Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
IEEE Signal Process. Lett. | 3 |
| 2023 | Weakly Supervised Distribution Discrepancy Minimization Learning With State Information for Person Re-IdentificationabstractWeakly supervised person re-identification (Re-ID) is appealing to handle real-world tasks by using state information that is available without manual annotation. At present, most methods perform unsupervised cross domain (UCD) learning by transferring the knowledge from the labeled source domain to the unlabeled target domain, which results in poor performance due to the severe shift. To address this problem, in this paper, we utilize the tracklet and camera information as weak supervision to propose a distribution discrepancy minimization learning (DDML) model for UCD person Re-ID. In addition to aligning data distributions from the perspective of domain adaptation learning, two losses are developed from the view of neighborhood invariance exploration to optimize matching results. Specifically, to bridge the gap between domains, we propose a camera-distribution-based (CDB) loss to align pair-wise distance distributions. Furthermore, to alleviate the biased search within the target domain, we propose a ranking-confidence-based (RCB) loss to perform the mined neighborhood for intra-camera and inter-camera separately to explore a high degree of confidence neighbor relations. Extensive experiments on three challenging datasets demonstrate that applying our method to unlabeled target domain outperforms current weakly supervised methods for person Re-ID. Jun Kong 0001, Xuefeng Tao, Min Jiang 0008, Tianshan Liu |
IEEE Trans. Multim. | 3 |
| 2023 | BGTracker: Cross-Task Bidirectional Guidance Strategy for Multiple Object TrackingabstractRecent works have shown that the joint-detection-and-embedding (JDE) paradigm has significantly enhanced the performance of multiple object tracking by simultaneously learning detection and re-identification features. These methods always utilize a weight-shared backbone network and two non-interactive branches for different tasks. This non-interactive multi-task learning strategy cannot make full use of geometric and semantic information between detection and re-identification tasks. And in the JDE paradigm, there exists a feature misalignment between detection and re-identification due to their different optimization directions. In this paper, BGTracker is proposed as a novel online tracking framework with a cross-task bidirectional guidance strategy between detection and re-identification. Firstly, we propose a Channel-based Decoupling module and Cross-direction Transformer to alleviate feature misalignment, which can obtain task-aligned embeddings and discriminative representations at the feature level. Then, we propose the bidirectional guidance strategy to link the two tasks by the prediction map's statistical information. In this strategy, two designed feature transformations are employed to utilize the advantages of each task for complementing each other at the task level. Finally, extensive experiments demonstrate that the proposed BGTracker outperforms various existing methods on the MOTChallenge benchmarks. Chen Zhou 0003, Min Jiang 0008, Jun Kong 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | AOH: Online Multiple Object Tracking With Adaptive Occlusion HandlingabstractMultiple object tracking has improved drastically in recent years due to one-shot tracking methods. These methods design joint-detection-and-tracking structures to achieve real-time tracking performance and introduce more powerful detectors to deal with missed objects. However, most of them perform poorly in crowded scenes because of frequent occlusions. Several previous works have attempted to alleviate the occlusion issue, but they hardly involve the essence of the problem. In this letter, we suppose that occlusion is closely related to crowd density, so the degree of occlusion can be estimated. Therefore, we propose a Potential Object Mining strategy to adaptively obtain occluded objects for reducing broken trajectories, which re-weights detections based on the predicted density map. Additionally, for the strategy, Dense Estimator is designed to predict the density of each region in an image by employing a Transformer-based structure. Combining them together forms our Adaptive Occlusion Handling (AOH) tracking framework. Extensive experiments on MOTChallenge benchmarks (MOT17 and MOT20) demonstrate that our AOH achieves the state-of-the-art performance, especially on the heavily occluded MOT20. Min Jiang 0008, Chen Zhou 0003, Jun Kong 0001 |
IEEE Signal Process. Lett. | 1 |
| 2022 | MTT: Multi-Scale Temporal Transformer for Skeleton-Based Action RecognitionabstractIn the task of skeleton-based action recognition, long-term temporal dependencies are significant cues for sequential skeleton data. State-of-the-art methods rarely have access to long-term temporal information, due to the limitations of their receptive fields. Meanwhile, most of the recent multiple branches methods only consider different input modalities but ignore the information in various temporal scales. To address the above issues, we propose a multi-scale temporal transformer (MTT) in this letter, for skeleton-based action recognition. Firstly, the raw skeleton data are embedded by graph convolutional network (GCN) blocks and multi-scale temporal embedding modules (MT-EMs), which are designed as multiple branches to extract features in various temporal scales. Secondly, we introduce transformer encoders (TE) to integrate embeddings and model the long-term temporal pattern. Moreover, we propose a task-oriented lateral connection (LaC) aiming to align semantical hierarchies. LaC distributes input embeddings to the downstream transformer encoders (TE), according to semantical levels. The classification headers aggregate results from TE and predict the action categories at last. The proposed method is shown efficiency and universality during experiments and achieves the state-of-the-art on three large datasets, NTU-RGBD 60, NTU-RGBD 120 and Kinetics-Skeleton 400. Jun Kong 0001, Yuhang Bian, Min Jiang 0008 |
IEEE Signal Process. Lett. | 3 |
| 2022 | SPRTracker: Learning Spatial-Temporal Pixel Aggregations for Multiple Object TrackingabstractRecently, multiple object tracking based on a single frame has achieved excellent performance. However, in crowded scenes, occlusion and motion blur will increase the difficulty of foreground object detection. In this letter, we propose a Spatial-temporal Pixel Resampling Tracker (SPRTracker) that introduces a novel cross-frame input framework to improve the anti-occlusion and anti-interference ability during tracking. We first propose a sampling mechanism in which Inter-frame Pixel Alignment (IPA) is designed to maintain spatial and temporal consistency in the propagation of pixel-level information between frames, aiming to improve the anti-interference ability during tracking target motion. Secondly, To improve the compensation effect of the gain information in the historical frame to the occluded target in the current frame, Similarity Reparameter Fusion (SRF) strategy is designed to fuse the features of the two frames. We evaluated the proposed method on two common benchmarks (MOT17 and MOT20), and the experimental results effectively demonstrated the superiority of our method. Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
IEEE Signal Process. Lett. | 3 |
| 2022 | MOTFR: Multiple Object Tracking Based on Feature RecodingabstractThe stable continuation of trajectories among different targets has always been the key to the tracking performance of multi-object tracking (MOT) tasks. If features of the target are aggregated and classified simply, the discriminant features of the target will be ignored. This will affect the robustness of the trajectory generated by the model. Meanwhile, many popular models are keen to execute detection and feature extraction tasks in parallel. But these two tasks will conflict with each other when optimized respectively. Therefore, we propose our tracker MOTFR to solve the above problems. In this paper, we propose a Locally Shared Information Decoupling Module (LSIDM) to reduce task optimization conflicts while ensuring the necessary information sharing. Meanwhile, a feature recoding module for deep extraction of identity discriminative features is proposed, which is called the Feature Purification Module (FPM). By combining LSIDM and FPM modules, the model utilizes the discriminative appearance features to guide the optimization of detection and further improves the performance of our model. To solve the problem of targets disappearing due to various abnormal occlusion, a Short-term Trajectory Online Complement Strategy (STOCS) is proposed to realize the trajectories continuation of these targets in the tracking stage. Through sufficient experiments, we demonstrate the superior performance of our MOTFR, which guarantees high-quality detection while achieving the stability of the target trajectory. Jun Kong 0001, Ensen Mo, Min Jiang 0008, Tianshan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Multidimensional Prototype Refactor Enhanced Network for Few-Shot Action RecognitionabstractFew-shot action recognition classifies new actions with only few training samples, of which the mainstream methods adopt class means to obtain prototypes as the representations of each category. However, affected by sample capacity and extreme samples, mean-of-class prototypes can’t well represent the average level of samples. In this paper, we enhance the prototypes from multiple dimensions for better classification. We firstly propose a novel similarity optimization mechanism where Prototype Aggregation Adaptive Loss (PAAL) is designed to deeply mine the similarity between samples and prototypes for enhancing the ability of inter-class differential detail identification. Secondly, for mitigating the impact of the samples on class prototypes, we refactor the prototype calculation formula with Cross-Enhanced Prototype (CEP) to narrow intra-class differences in which Reweighted Similarity Attention (RSA) is designed to update prototypes. Finally, Dynamic Temporal Transformation (DTT) is proposed to alleviate inconsistent distribution of temporal information for obtaining better video-level descriptors. Extensive experiments on standard benchmark datasets demonstrate that our proposed method achieves the state-of-the-art results. Shuwen Liu 0008, Min Jiang 0008, Jun Kong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Unsupervised Domain Adaptation by Multi-Loss Gap Minimization Learning for Person Re-IdentificationabstractUnsupervised domain adaptation (UDA) person re-identification (ReID) faces enormous challenges due to the severe shift between the source and target domains, as well as the dramatic variations within the target domain. In this paper, to address these issues, we propose a multi-loss gap minimization learning (MGML) approach for UDA person ReID. Firstly, we introduce the part model to learn discriminative patch features and design a Patch-based Part Ignoring (PPI) loss to select reliable instances for the efficient learning of the part model. Then, given the gap that typically occurs because of the inter-domain shift and intra-domain variations, a Gap-based Minimum Camera Discrepancy (G-MCD) loss is proposed. Specifically, in terms of the inter-domain, we propose to leverage the tracklet and camera information to label each distance vector, and accordingly align pair-wise distance distributions to bridge the inter-domain gap. As for the intra-domain, to alleviate the biased search, we propose to perform the mined neighborhood for intra-camera and inter-camera separately to optimize matching results by exploring neighborhood relations more deeply. Finally, experimental results on three challenging datasets demonstrate that applying our method to unlabeled target domain outperforms current UDA methods for person ReID. Xuefeng Tao, Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Sparse Attention Module for optimizing semantic segmentation performance combined with a multi-task feature extraction network
Min Jiang 0008, Fuhao Zhai, Jun Kong 0001 |
Vis. Comput. | 1 |
| 2021 | A novel deep learning model DDU-net using edge features to enhance brain tumor segmentation on MR images
Min Jiang 0008, Fuhao Zhai, Jun Kong 0001 |
Artif. Intell. Medicine | 1 |
| 2021 | Paralleled attention modules and adaptive focal loss for Siamese visual trackingabstractAbstract Recently, Siamese‐based trackers have drawn amounts of attention in visual tracking field because of their excellent performance on different tracking benchmarks. However, most Siamese‐based trackers encounter difficulties under circumstances such as similar objects interference and background clutters. Besides, there exists an extreme foreground–background data imbalance that weakens the performance during training but few loss functions pay attention to it. The authors intend to address the issues mentioned above by introducing a module named paralleled spatial and channel attention (PSCA) and adaptive focal loss (AFL). Firstly, paralleled spatial and channel attention is proposed to enhance the extracted features and eliminate the noise information from both spatial and channel aspects. Secondly, adaptive focal loss is proposed as the loss function to make the model focus on hard samples that contribute more to training process. Finally, paralleled spatial and channel attention and modified ResNet are combined for extracting more powerful features. Experimental results show that the authors' method achieves outstanding performance in multiple benchmarks while keeping a beyond‐real‐time frame rate. Yuyao Zhao, Min Jiang 0008, Jun Kong 0001 |
IET Image Process. | 2 |
| 2021 | Incorporating multi-level CNN and attention mechanism for Chinese clinical named entity recognition
Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
J. Biomed. Informatics | 3 |
| 2021 | Diverse Features Fusion Network for video-based action recognition
Haoyang Deng, Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Interactive information module for person re-identification
Jun Kong 0001, Min Jiang 0008 |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Dynamic Center Aggregation Loss With Mixed Modality for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a challenging cross-modality pedestrian retrieval task which aims to match person images between the visible and infrared modality of the same identity. Existing methods usually adopt two-stream network to solve cross-modality gap, but they ignore the pixel-level discrepancy between the visible and infrared images. Some methods introduce auxiliary modalities in the network, but they lack powerful constraints on the feature distribution of multiple modalities. In this letter, we propose a Dynamic Center Aggregation (DCA) loss with mixed modality for VI-ReID. Concretely, we employ a mixed modality as a bridge between the visible and infrared modality, reducing the difference of the two modalities at the pixel-level. The mixed modality is generated by a Dual-modality Feature Mixer (DFM), which combines the features of visible and infrared images. Moreover, we dynamically adjust the relative distance across multi-modality through DCA loss, which is conducive to explore the modality-invariant feature. We evaluate the proposed method on two public available VI-ReID datasets (SYSU-MM01 and RegDB). Experimental results demonstrate that our method achieves competitive performance. Jun Kong 0001, Qibin He 0003, Min Jiang 0008, Tianshan Liu |
IEEE Signal Process. Lett. | 3 |
| 2021 | Mutual Learning and Feature Fusion Siamese Networks for Visual Object TrackingabstractRecently Siamese-based trackers have shown their outstanding performance in visual object tracking community. But they seldom pay attention to the inter-branch interaction as well as intra-branch feature fusion from different convolution layers. In this paper, we build up a comprehensive Siamese network which consists of a mutual learning subnetwork (M-net) and a feature fusion subnetwork (F-net), to realize object tracking. Each of them is a Siamese network with special functions. M-net is designed to help the two branches mine the dependencies from each other, thus the object template is adaptively updated to a certain extent. F-net fuses different levels of convolutional features for full usage of spatial and semantic information. We also design a global-local channel attention (GLCA) module in F-net to capture the channel dependencies for a proper feature fusion. Our method takes ResNet as feature extractor and is trained offline in an end-to-end style. We evaluate our method in several famous benchmarks such as OTB2013, OTB2015, VOT2015, VOT2016, NFS and TC128. Extensive experimental results demonstrate our method achieves competitive results while maintaining a considerable real-time speed. Min Jiang 0008, Yuyao Zhao, Jun Kong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Symmetrical Enhanced Fusion Network for Skeleton-Based Action RecognitionabstractA novel method for skeleton-based action recognition by fusing multi-level spatial features and multi-level temporal features is proposed in this article. Recently, Graph Convolutional Network (GCN) for skeleton-based action recognition has attracted the eyes of many researchers and has a great performance in the field of action recognition. But most of them focus on changing architecture of single-stream network and only use simple methods like average fusion to fuse different forms of skeleton data. In this article, we shift the focus to the problem that insufficient interactions between the different forms of features for that networks are unable to fully capture efficient information from skeleton data. To tackle this problem, we propose a multi-stream network called Symmetrical Enhanced Fusion Network (SEFN). The network is composed of a spatial stream, a temporal stream and a fusion stream. The spatial stream extracts spatial features from skeleton data by GCN. The temporal stream is able to extract temporal features from skeleton data with the help of the embedded Motion Sequence Calculation Algorithm. The fusion stream provides an early fusion method and extra fusion information for the whole network. It gathers multi-level features from two feature extractions and fuses them with the Multi-perspective Attention Fusion Module (MPAFM) we propose. The MPAFM enables different forms of data to enhance each other and can strengthen feature extractions. In the final, we generalize the skeleton data from joint data to bone data and evaluate our network in three large-scale benchmarks: NTU-RGBD, NTU-RGBD 120 and Kinetics-Skeleton. Experiment results demonstrate that our method achieves competitive performance. Jun Kong 0001, Haoyang Deng, Min Jiang 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Collaborative model tracking with robust occlusion handlingabstractCurrently, the discriminative correlation filter‐based trackers have achieved higher tracking accuracy. However, visual tracking still faces challenges in terms of heavy occlusion, scale variation and so on. In this study, the authors intend to solve heavy occlusion by introducing collaborative model into classifier‐box. Firstly, they introduce complex colour features into correlation filter tracker to improve the effect of the tracker. Secondly, they introduce a multi‐scale method into their tracker to ease the scale problem. Thirdly, in order to solve the heavy occlusion in the tracking process, they adopt the locally weighted distance and classifier‐box. Their algorithm achieves distance precision rates of 81.7 and 77.4% on OTB2013 dataset and OTB2015 dataset, respectively. Their contribution focuses on solving heavy occlusion by using colour features, locally weighted distance and classifier‐box. The experimental results on OTB2013 and OTB2015 datasets demonstrate their algorithm to perform better than state‐of‐the‐art methods. Jun Kong 0001, Yitao Ding, Min Jiang 0008 |
IET Image Process. | 3 |
| 2020 | Cross-level reinforced attention network for person re-identification
Min Jiang 0008, Jun Kong 0001, Zhende Teng, Danfeng Zhuang |
J. Vis. Commun. Image Represent. | 1 |
| 2020 | Spatial-temporal saliency action mask attention network for action recognition
Min Jiang 0008, Na Pan, Jun Kong 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2020 | Multiple depth-levels features fusion enhanced network for action recognition
Shengquan Wang, Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Robust part-based visual tracking via adaptive collaborative modellingabstractDiscriminative correlation filter‐based tracking algorithms have recently shown impressive performance on benchmark data sets. However, visual tracking is still a challenging task in the case of partial occlusions, irregular deformations and so on. In this study, the authors intend to solve these issues by introducing the adaptive collaborative model into part‐based tracking. First, instead of a simple linear superposition, the collaborative strategy they proposed combines the template model and colour‐based model adaptively and relies on the strengths of both to promote the accuracy. Second, we utilise the voting strategy to figure out the final object position from reliable parts, and the motion information is used in evaluation for reliable parts to enable the tracker to be robust in various situations. Third, the authors utilise a discriminative multi‐scale estimate method to solve the problem of scale variations. Finally, they introduce a dimensionality reduction method to limit the computational complexity of the tracker. Abundant experiments demonstrate that the tracker performs superiorly against several advanced algorithms on both the Online Tracking Benchmark (OTB) 2013 and OTB2015 data sets while maintaining the high frame rates. Jun Kong 0001, Benxuan Wang, Min Jiang 0008 |
IET Image Process. | 3 |
| 2019 | Collaborative multimodal feature learning for RGB-D action recognition
Jun Kong 0001, Tianshan Liu, Min Jiang 0008 |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Regularisation learning of correlation filters for robust visual trackingabstractRecently, kernelised correlation filter (KCF)‐based trackers aroused increasing interest and achieved extremely compelling results in different competitions and benchmarks in the field of visual object tracking. However, the training mechanism of the KCF that exploits simple linear combinations of filter from the previous frame easily cause error accumulation. To overcome this problem, the authors propose a novel training strategy that utilises all of the previous training samples, and a sparsity‐related loss function regularised by the L 1 norm to deal with the problem of the fixed template size in KCF trackers, a separate scale filter is learned for scale estimation during the tracking process. Moreover, powerful features that include histogram of oriented gradients (HOG) and colour features are integrated to further improve the robustness of the authors’ tracking. Extensive experiments in various challenging situations demonstrate that the proposed method performs favourably against several state‐of‐the‐art tracking algorithms. Min Jiang 0008, Jianyu Shen, Jun Kong 0001, Hongtao Huo |
IET Image Process. | 1 |
| 2016 | Generalized ℓP-regularized representation for visual tracking
Jun Kong 0001, Chenhua Liu, Min Jiang 0008, Shengwei Tian, Hui-Cheng Lai |
Neurocomputing | 3 |
| 2015 | Informative joints based human action recognition using skeleton contexts
Min Jiang 0008, Jun Kong 0001, George Bebis, Hongtao Huo |
Signal Process. Image Commun. | 1 |
| 2014 | IR remote sensing image registration based on multi-scale feature extractionabstractInfrared remote sensing image has poor contrast and lower SNR so that real-time and robustness are not superior in image registration. In order to solve it, a novel registration based on Multi-scale feature extraction is proposed in this paper. This algorithm is designed in two aspects. Firstly, Gaussian convolution template size adjusts adaptively with the increasing of scale factors. Then the Multi-space is reconstructed. Secondly, feature points bidirectional matching based on the City-block distance is introduced into image registration. So the real-time performance and robustness are enhanced further. Finally, the experimental results showed that by this improved algorithm the infrared remote sensing images are registered more quickly and accurately than by traditional SIFT algorithm. Jun Kong 0001, Min Jiang 0008, Yi-Ning Sun |
IJCNN | 2 |