Canlong Zhang

dblp:78/10949 · DBLP profile ↗
← Back
105ranked-venue papers
6as first author
67since 2021 · last 2026
0000-0003-4375-1405ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 76 · 4 first-author · 49 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 1 first-author · 15 since 2021Databases, data management, data science and information retrieval · 12 · 6 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 DT-CC: Digital Twin-Driven Knowledge Engineering for Classical Chinese Historical Document
Shiyi Qin, Canlong Zhang, Wenhui Tang, Yanchao Ma
KSEM (2)2
2026 A novel dendritic neuron model enhanced by the synaptic-attention mechanism and fusion-dendritic layer
Runcong Ma, Yonghua Pang, Canlong Zhang, Xudong Luo 0003
Neurocomputing3
2026 Adaptive confidence-driven learning and cross-modal hard sample mining for unsupervised visible-infrared person re-identification
Canlong Zhang, Haifei Ma, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei
Inf. Process. Manag.2
2026 Spatio-temporal semantic alignment leveraging human structural priors for text-to-video person retrieval
Canlong Zhang, Sheng Xie, Yuanlin Shi, Xiaochun Lu
Inf. Sci.2
2026 Gradually target activating and self-supervised enhancing for visual tracking
Yanfang Deng, Canlong Zhang, Xiaochun Lu
Knowl. Based Syst.3
2026 Counterfactual causal inference for robust visual question answering
Wei Li 0233, Zhixin Li 0001, Fuyun Deng, Canlong Zhang
Neural Networks5
2026 Twin contrastive interventional-cause hashing for unsupervised cross-modal retrieval
Bo Li 0144, Zhixin Li 0001, Shuni Jiang, Canlong Zhang, Huifang Ma
Neural Networks4
2026 CMAG: Cross-Modal Attention and Graph-Enhanced Memory for Unsupervised Visible-Infrared Person Re-Identification
abstract
Unsupervised visible-infrared person re-identification (USL-VI-ReID) has garnered widespread attention due to its surveillance application value in complex environments. However, it faces four key challenges: modality discrepancy, batch training limitations, pseudo-label noise, and camera view bias. This paper proposes the CMAG (Cross-Modal Attention and Graph-enhanced Memory) framework, which innovatively combines circular topology structure with cross-modal attention mechanisms to address these challenges. CMAG introduces four core innovations: (1) applying circular topology structure to provide pseudo-label verification through detecting circular paths in feature space, effectively addressing the pseudo-label noise problem; (2) designing a cross-modal attention mechanism for Vision Transformers with residual fusion to balance modality-specific and shared information, solving the modality discrepancy issue; (3) constructing a graph-structured memory enhancement module with adaptive graph construction and multi-layer feature propagation to overcome batch training limitations; and (4) integrating camera-specific clustering with circular structure constraints to reduce camera background bias. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate the effectiveness of CMAG, achieving approximately 3.5% improvement in Rank-1 accuracy and 2.8% in mAP on average compared to state-of-the-art methods, validating our approach’s advantages in addressing key challenges in unsupervised cross-modal person re-identification.Code is available at https://github.com/hurryup186/CMAG.
Canlong Zhang, Junwei Tian, Haifei Ma, Zhixin Li 0001, Zhiwen Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Prototypical Graph Alignment for Text-based Person Search
abstract
Text-based Person Search is one of the downstream tasks of cross-modal retrieval. The key challenge is aligning features of two extremely irrelavant modalities into the same latent space. Recent works within Prototype Learning introduce a few learnable parameters to map heterogeneous features into the same prototype space, aligning them using weighted-sum prototype on global matching or independent prototypes on local matching. However, these matches both ignore the relationship between prototypes. Thus we propose a Prototypical Graph Alignment (PGA) architecture for Text-based Person Search. For a given mini-batch, our PGA first transforms image-text features of identical person to his corresponding prototype space, enabling unique individual to obtain customized unimodal prototypical graphs and cross-modal fused graphs, and each graph consists of prototype representation and comparable relationship. By minimizing the discrepancy between unimodal and cross-modal graphs, we can achieve node-level local alignment and edge-level relational alignments. Extensive experiments conducted on three primary TPS datasets demonstrate the effectiveness of the proposed method.
Canlong Zhang, Chunrong Wei
ICASSP2
2025 Uncertainty-Aware Prototype Semantic Decoupling for Text-Based Person Search in Full Images
Zengli Luo, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei
KSEM (2)2
2025 Hierarchical-enhanced graph convolutional networks leveraging causal inference for aspect-based sentiment analysis
Fengling Zhou, Zhixin Li 0001, Canlong Zhang, Huifang Ma
Appl. Intell.3
2025 PM2.5 Concentration Prediction Using CNN-LSTM Model Based on Multi-Feature Fusion
abstract
ABSTRACT In order to solve the problem that the existing PM 2.5 concentration prediction methods ignore the spatial and temporal influencing factors of PM 2.5 concentration, this paper constructs a spatial characteristic factor of PM 2.5 concentration based on the maximum information coefficient, and proposes a CNN‐LSTM combined prediction model based on multi‐feature fusion, which transforms the abstract spatial and temporal influencing factors into quantifiable features. The model has good feature extraction ability and strong ability to capture short‐term transient information and long‐range dependent information in time series data, which improves the prediction performance of the model. The experimental results show that the prediction accuracy of CNN‐LSTM model based on multi‐feature fusion is 87.21%, and MAPE is 6.25, 4.84, and 1.29 less than BP, SVR, and LightGBM, and 1.91 and 7.04 less than CNN and LSTM.
Jiexia Huang, Canlong Zhang
Concurr. Comput. Pract. Exp.5
2025 Dynamic window sampling strategy for image captioning
Zhixin Li 0001, Jiahui Wei, Tiantao Xian, Canlong Zhang, Huifang Ma
Eng. Appl. Artif. Intell.4
2025 Recursively learning fine-grained spatial-temporal features for video-based person Re-identification
Haifei Ma, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001
Eng. Appl. Artif. Intell.2
2025 Multi-scale Feature Refinement via Perspective Scaling and Adaptive Regularization for text-based person search
Sheng Xie, Canlong Zhang, Runcong Ma, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei
Eng. Appl. Artif. Intell.2
2025 A cross-modal collaborative guiding network for sarcasm explanation in multi-modal multi-party dialogues
Xingjie Zhuang, Zhixin Li 0001, Canlong Zhang, Huifang Ma
Eng. Appl. Artif. Intell.3
2025 Relevance-aware prompt-tuning method for multimodal social entity and relation extraction
Zhenbin Chen, Zhixin Li 0001, Mingqi Liu, Canlong Zhang, Huifang Ma
Neurocomputing4
2025 Unsupervised infrared-visible person re-identification by multi-level Dual-Stream Contrastive Learning
Canlong Zhang, Haifei Ma, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei
Neurocomputing2
2025 Fusing grid and adaptive region features for image captioning
Jiahui Wei, Zhixin Li 0001, Canlong Zhang, Huifang Ma
Image Vis. Comput.3
2025 Dynamic feature projection and grouped contrastive learning for text-to-image person re-identification
Shun He, Canlong Zhang, Xiaochun Lu, Zhixin Li 0001, Zhiwen Wang 0001
Knowl. Based Syst.2
2025 Temporal Motion and Spatial Enhanced Appearance with Transformer for video-based person ReID
abstract
For video-based person Re-Identification (Re-ID), how to efficiently extract temporal motion features and spatial appearance features from video sequences is a key issue. Conventional approaches focus on modelling the entire video spatio-temporal features, ignoring the inherent differences between temporal motion features (e.g., gait) that change over time and spatial appearance features (e.g., clothing) that are stable over time in terms of attributes. Because of their different sensitivities in real-world scenarios, conventional approaches often lose critical fine-grained features. To address these issues, we propose a T emporal M otion and spatial E nhanced A ppearance with T ransformer-based (T 2 MEA) framework for modelling spatial–temporal video discriminative representations. Specifically, (1) Dual-Branch Architecture: The content branch emphasises extracting the overall structure of the video using the spatial–temporal aggregation (STA) module from a global view, whereas the fovea branch focuses on gaining local fine-grained spatio-temporal features. (2) Zero-Parameter Design: the [CLS] Token Channel Shift Interaction (TCSI) module captures the dynamic features and static features between adjacent frames without additional parameters; the Spatial Patches Shift Enhancing (SPSE) module is introduced to enhance appearance features within frame to address occlusion and illumination changes without additional parameters. (3) Spatial–Temporal Interaction: The Cross-Attention Aggregation (CAA) module is proposed to interact between temporal and spatial features and further enrich the spatial–temporal feature representation for video sequences. Extensive experiments on three public Re-ID benchmarks (MARS, iLIDS-VID, and PRID-2011) demonstrate that the proposed framework outperforms several state-of-the-art methods.
Haifei Ma, Canlong Zhang, Enhao Ning, Chai Wen Chuah
Knowl. Based Syst.2
2025 DyCR-Net: A dynamic context-aware routing network for multi-modal sarcasm detection in conversation
Xingjie Zhuang, Zhixin Li 0001, Fengling Zhou, Jingliang Gu, Canlong Zhang, Huifang Ma
Knowl. Based Syst.5
2025 Joint feature augmentation and posture label for cloth-changing person re-identification
Liman Jiang, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei
Multim. Syst.2
2025 Enhancing robust VQA via contrastive and self-supervised learning
abstract
Visual Question Answering (VQA) aims to evaluate the reasoning abilities of an intelligent agent using visual and textual information. However, recent research indicates that many VQA models rely primarily on learning the correlation between questions and answers in the training dataset rather than demonstrating actual reasoning ability. To address this limitation, we propose a novel training approach called Enhancing Robust VQA via Contrastive and Self-supervised Learning (CSL-VQA) to construct a more robust VQA model. Our approach involves generating two types of negative samples to balance the biased data, using self-supervised auxiliary tasks to help the base VQA model overcome language priors, and filtering out biased training samples . In addition, we construct positive samples by removing spurious correlations in biased samples and perform auxiliary training through contrastive learning . Our approach does not require additional annotations and is compatible with different VQA backbones. Experimental results demonstrate that CSL-VQA significantly outperforms current state-of-the-art approaches, achieving an accuracy of 62.30% on the VQA-CP v2 dataset, while maintaining robust performance on the in-distribution VQA v2 dataset. Moreover, our method shows superior generalization capabilities on challenging datasets such as GQA-OOD and VQA-CE, proving its effectiveness in reducing language bias and enhancing the overall robustness of VQA models.
Runlin Cao, Zhixin Li 0001, Zhenjun Tang, Canlong Zhang, Huifang Ma
Pattern Recognit.4
2024 Gradually Spatio-Temporal Feature Activation for Target Tracking
abstract
Most existing transformer-based trackers use ViT [1] as the backbone to extract and fuse feature tokens of target templates and search region. Since both the target template and the search region contain background information, their tokens are prone to background interference in interaction that affects tracking performance. We propose a GFATrack that combines spatiotemporal information with prominent target features. The tracker mainly consists of the Dynamic Template Refinement branch and the Search Feature Enhancement branch. The former activates the target features in the dynamic template and provides temporal information. The latter enhances search features by aggregating spatial information from the initial template and temporal information from the dynamic template to achieve precise tracking. Our proposed Feature Activation Module can effectively fuse refined features with reference features and highlight high-similarity features among them. At the same time, we propose a concise and effective dynamic threshold update strategy to capture time context update dynamic templates from historical prediction results. Many experiments have verified the effectiveness and latest performance of the proposed method.
Yanfang Deng, Canlong Zhang, Zhixin Li 0001, Chunrong Wei, Zhiwen Wang 0001, Shuqi Pan
ICASSP2
2024 Mask-guided Salient Feature Mining for Cloth-Changing Person Re-identification
abstract
Cloth-changing person re-identification (CC-ReID) aims at retrieving pedestrians with changing clothes across multiple cameras. Most methods often focus on exploiting discriminative biometric features for resisting the clothing changes. Nevertheless, simply concatenating various features not only increases the computations, but also introduces redundant information. In this paper, we propose a Mask-guided Salient Feature Mining (MSFM) to learn cloth-irrelevant features. Specifically, we introduce human parsing results in the data pre-processing stage, and exploit region-specific cues by performing pixel-level mask on the parsing results. Besides, a Multi-local Attention (MLA) is proposed, where the model can focus on local cues in horizontal direction and obtain robust identity-related representations. Meanwhile, we introduce a part loss supervised by selective masking regions for capturing fine-grained features and constraining clothing features. Extensive experiments on two public cloth-changing datasets demonstrate our proposed MSFM can achieve superior performance over existing state-of-the-art methods.
Liman Jiang, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei
ICME2
2024 Person Re-identification utilizing Text to Search Video
abstract
Previous research in pedestrian re-identification can be broadly categorized into three classes: image-to-image, video-to-video, and text-to-image pedestrian re-identification. However, these paradigms exhibit certain limitations in practical applications. Hence, this paper introduces a novel task: utilizing natural language to retrieve pedestrians in videos. Specifically, given a textual description of a person, the model’s objective is to retrieve the pedestrian from a video dataset that best matches the provided text. Due to the absence of datasets specifically designed for text-to-video pedestrian re-identification, we undertook manual annotations on the video-based pedestrian re-identification dataset, MARS, to establish the T-MARS dataset. In this paper, we propose the Implicit Alignment Framework with Local Capture Attention Module(IAFL) for text-to-video pedestrian re-identification. The Local Capture Attention Module incorporates local priors while modeling inter-frame complementary relationships, thereby effectively transferring knowledge from text-to-image models to the task of text-to-video pedestrian re-identification. A plethora of text-to-video models are evaluated and compared on this dataset. We observe that the proposed Implicit Alignment Framework is more suitable for achieving cross-modal granularity alignment in pedestrian re-identification tasks. Additionally, IAFL establishes the state-of-the-art performance in pedestrian search.
Shunkai Zhou, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei
ICME2
2024 MLLM-Driven Semantic Enhancement and Alignment for Text-Based Person Search
Junwei Tian, Canlong Zhang, Chunrong Wei
ICONIP (7)2
2024 Unsupervised cross-modal hashing retrieval via Dynamic Contrast and Optimization
Xiumin Xie, Zhixin Li 0001, Bo Li 0144, Canlong Zhang, Huifang Ma
Eng. Appl. Artif. Intell.4
2024 GAP: A novel Generative context-Aware Prompt-tuning method for relation extraction
Zhenbin Chen, Zhixin Li 0001, Yufei Zeng, Canlong Zhang, Huifang Ma
Expert Syst. Appl.4
2024 SiamCA: Siamese visual tracking with customized anchor and target-aware interaction
Shuqi Pan, Canlong Zhang, Zhixin Li 0001, Liaojie Hu
Expert Syst. Appl.2
2024 Full-view salient feature mining and alignment for text-based person search
Sheng Xie, Canlong Zhang, Enhao Ning, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei
Expert Syst. Appl.2
2024 Similarity Graph-correlation Reconstruction Network for unsupervised cross-modal hashing
Zhixin Li 0001, Bo Li 0144, Canlong Zhang, Huifang Ma
Expert Syst. Appl.4
2024 Sentence salience contrastive learning for abstractive text summarization
Zhixin Li 0001, Zhenbin Chen, Canlong Zhang, Huifang Ma
Neurocomputing4
2024 A review on video person re-identification based on deep learning
Haifei Ma, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei
Neurocomputing2
2024 Text-based person search by non-saliency enhancing and dynamic label smoothing
Yonghua Pang, Canlong Zhang, Zhixin Li 0001, Chunrong Wei, Zhiwen Wang 0001
Neural Comput. Appl.2
2024 Mining core information by evaluating semantic importance for unpaired image captioning
Jiahui Wei, Zhixin Li 0001, Canlong Zhang, Huifang Ma
Neural Networks3
2024 Target-Aware Tracking With Spatial-Temporal Context Attention
abstract
Current trackers only rely on a fixed target template to localize the target in each frame, which is however prone to fail in case of fast appearance changes or the presence of distractor objects. Having some historical knowledge about the tracked targets as well as their surrounding scenes can be highly beneficial for robust tracking. This historical information can be propagated through the sequence and used to timely perceive the change in target appearance and explicitly avoid distractor objects. In this work, we propose a Spatial-Temporal Context Attention (STCA) model which utilizes the appearance and state information of previously tracked targets as well as their surrounding scenes to more accurately localize the real target in the current frame. We embed an improved position encoder into the STCA, which enables the target template, context template and search patch to perform extensive interactional fusion through simultaneously self-attention and cross-attention calculation. By embedding the STCA module into Transformer, we construct a target-aware based online tracking network (named TATrack) that has a backbone to extract features better suited to the tracking task, a neck to further suppress distractors and highlight target, and a classification-regression head to make the tracking scores consistently reflect the quality of the bounding boxes. In addition, we also design a simple yet effective online updating approach to select high-quality context templates. Our tracker reaches the latest level on several benchmarks, including LaSOT, TrackingNet, GOT10k, OTB100 and UAV123. The code and trained models are available at https://github.com/hekaijie123/TATrack.
Kaijie He, Canlong Zhang, Sheng Xie, Zhixin Li 0001, Zhiwen Wang 0001, Rui-Guo Qin
IEEE Trans. Circuits Syst. Video Technol.2
2023 Target-Aware Tracking with Long-Term Context Attention
abstract
Most deep trackers still follow the guidance of the siamese paradigms and use a template that contains only the target without any contextual information, which makes it difficult for the tracker to cope with large appearance changes, rapid target movement, and attraction from similar objects. To alleviate the above problem, we propose a long-term context attention (LCA) module that can perform extensive information fusion on the target and its context from long-term frames, and calculate the target correlation while enhancing target features. The complete contextual information contains the location of the target as well as the state around the target. LCA uses the target state from the previous frame to exclude the interference of similar objects and complex backgrounds, thus accurately locating the target and enabling the tracker to obtain higher robustness and regression accuracy. By embedding the LCA module in Transformer, we build a powerful online tracker with a target-aware backbone, termed as TATrack. In addition, we propose a dynamic online update algorithm based on the classification confidence of historical information without additional calculation burden. Our tracker achieves state-of-the-art performance on multiple benchmarks, with 71.1% AUC, 89.3% NP, and 73.0% AO on LaSOT, TrackingNet, and GOT-10k. The code and trained models are available on https://github.com/hekaijie123/TATrack.
Kaijie He, Canlong Zhang, Sheng Xie, Zhixin Li 0001, Zhiwen Wang 0001
AAAI2
2023 Customized Anchors Can Better Fit the Target in Siamese Tracking
Shuqi Pan, Canlong Zhang, Zhixin Li 0001, Liaojie Hu, Yanfang Deng
ICONIP (15)2
2023 Text-Based Person re-ID by Saliency Mask and Dynamic Label Smoothing
Yonghua Pang, Canlong Zhang, Zhixin Li 0001, Liaojie Hu
ICONIP (5)2
2023 Multi-scale Context Aggregation for Video-Based Person Re-Identification
Canlong Zhang, Zhixin Li 0001, Liaojie Hu
ICONIP (14)2
2023 Reinforced domain adaptation with attention and adversarial learning for unsupervised person Re-ID
Peiyi Wei, Canlong Zhang, Yanping Tang, Zhixin Li 0001, Zhiwen Wang 0001
Appl. Intell.2
2023 Mining graph-based dynamic relationships for object detection
Xiwei Yang, Zhixin Li 0001, Xinfang Zhong, Canlong Zhang, Huifang Ma
Eng. Appl. Artif. Intell.4
2023 Multi-scene image enhancement based on multi-channel illumination estimation
Runxing Zhao, Zhiwen Wang 0001, Wuyuan Guo, Canlong Zhang
Expert Syst. Appl.4
2023 Discriminative feature mining with relation regularization for person re-identification
Jing Yang 0046, Canlong Zhang, Zhixin Li 0001, Yanping Tang, Zhiwen Wang 0001
Inf. Process. Manag.2
2022 Multi-Level Relation Aware Network for Person Re-Identification
abstract
Person attribute or pose information has improved person re-identification performance, however, inaccurate pose or attribute module will damage the final identification performance. Based on this, we propose a multi-scale relation aware network (MSRA) for person re-identification. Specifically, we design an attribute relation mining module to construct an attribute map through constraint loss to learn the correlation between different attributes. Besides, we construct a multi-level Pose Pyramid based on the physical structure of the human body, so as to model the internal relationship between pose points. Finally, we designed a cross-scale graph convolution to infer the cooperative structural relation between different layers of components and fused it with the attribute relation module to reinforce the feature. Many experiments on three large-scale datasets verify the effectiveness and state-of-the-art performance of the proposed method.
Jing Yang 0046, Canlong Zhang, Zhixin Li 0001, Yanping Tang
ICASSP2
2022 SATNet: Captioning with Semantic Alignment and Feature Enhancement
Wenhui Bai, Canlong Zhang, Zhixin Li 0001, Peiyi Wei, Zhiwen Wang 0001
ICONIP (3)2
2022 Multi-level Network Based on Text Attention and Pose-Guided for Person Re-ID
Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001
ICONIP (7)2
2022 Graph Structure Guided Transformer for Semantic Segmentation
abstract
Segmentation is an essential operation of image processing, and utilizing long-range context information is the key for pixel-wise prediction tasks such as semantic segmentation. Convolutional Neural Networks (CNNs) are good at modeling local relationships through convolutional operations, but they are often inefficient in capturing global relationships between distant regions and require stacking multiple convolutional lay-ers. Utilizing the advantages of transformer in modeling long-range dependency, this paper proposes a novel Graph Structure Guided Transformer (GSGT) to realize semantic segmentation. Different from the previous methods that hard-divide the image in a regular grid manner, our graph projection method maps the two-dimensional feature map into a graph structure according to certain semantic relevance, so as to meet the data structure form required by the transformer. Meanwhile, to fully utilize the graph structure information, we also propose a graph embedding attention module, which utilizes the local topology of the graph structure to complement the global context of transformer. Moreover, GSGT is easy to be incorporated with various CNN backbones and transformer model variants to significantly improve the segmentation accuracy and convergence speed. Experiments on Cityscapes, VOC and ADE20K datasets demonstrate that the proposed method performs well in semantic seamentation task.
Luyang Qian, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001
ICTAI2
2022 Image Captioning According to User's Intention and Style
abstract
Exciting image captioning models are usually individuality-agnostic, and they cannot generate individual description according to the user's intention and language style. To address above problem, this paper proposes a personalized image captioning model by using fine-grained scene control graph and the style control factor to respectively represent user's intention and speaking style. More specifically, we first construct a scene control graph that consists of the object, its attribute and the relationship between it and other objects, and employe the graph flow attention to control the focus of the description. Secondly, propose a style generator to extract user's style pattern from his profile generated based on his gender, age, education level and other information. Finally, enter style factor in the style control module before generating sentences, which enables the language decoder to output a personalized caption for image. The experimental results on MSCOCO and FlickrStyle datasets show that the proposed method can generate personalized and diverse image caption sentences.
Canlong Zhang, Zhiwen Wang 0001, Zhixin Li 0001
IJCNN2
2022 Fast semantic segmentation network with attention gate and multi-layer fusion
Yanping Tang, Canlong Zhang, Qinghe Cheng, Zhixin Li 0001, Luyang Qian
Multim. Tools Appl.2
2022 Correction to: Fast semantic segmentation network with attention gate and multi-layer fusion
Yanping Tang, Canlong Zhang, Qinghe Cheng, Zhixin Li 0001, Luyang Qian
Multim. Tools Appl.2
2022 Fine-grained alignment network and local attention network for person re-identification
Dongming Zhou 0003, Canlong Zhang, Yanping Tang, Zhixin Li 0001
Multim. Tools Appl.2
2022 PAFM: pose-drive attention fusion mechanism for occluded person re-identification
Jing Yang 0046, Canlong Zhang, Yanping Tang, Zhixin Li 0001
Neural Comput. Appl.2
2022 Dual Global Enhanced Transformer for image captioning
Tiantao Xian, Zhixin Li 0001, Canlong Zhang, Huifang Ma
Neural Networks3
2021 Relation Also Need Attention: Integrating Relation Information Into Image Captioning
abstract
Image captioning methods with attention mechanism are leading this field, especially models with global and local attention. But there are few conventional models to integrate the relationship information between various regions of the image. In this paper, this kind of relationship features are embedded into the fused attention mechanism to explore the internal visual and semantic relations between different object regions. Besides, to alleviate the exposure bias problem and make the training process more efficient, we combine Generative Adversarial Network with Reinforcement Learning and employ the greedy decoding method to generate a dynamic baseline reward for self-critical training. Finally, experiments on MSCOCO datasets show that the model can generate more accurate and vivid image captioning sentences and perform better in multiple prevailing metrics than the previous advanced models.
Zhixin Li 0001, Tiantao Xian, Canlong Zhang, Huifang Ma
ACML4
2021 Improved Occluded Person Re-Identification with Multi-feature Fusion
Jing Yang 0046, Canlong Zhang, Zhixin Li 0001, Yanping Tang
ICANN (4)2
2021 Joint Scence Network and Attention-Guided for Image Captioning
abstract
Image captioning is an interesting and challenging task. The previously established image captioning approach is based mainly on the encoder-decoder architecture, but it suffers from problems such as inaccurate captioning information, and the generated captioning sentences are not sufficiently rich. This paper proposes a novel image captioning model that is based on a self-attention network and a scene graph relationship network. First, an improved self-attention network is added to the extraction of visual features to evaluate the effectiveness of image global information for image generation. Then, we design a visual intensity parameter to coordinate the strategies of visual features and language model for word generation. Finally, a graph convolutional network is designed to extract the relationships from the scene information to render the generated caption more exciting and to increase the accuracy of the fine-grained captioning. We demonstrated the satisfactory performance of the model on the MS-COCO and Flickr 30K datasets. The experimental results demonstrate that the proposed model realizes state-of-the-art performance.
Dongming Zhou 0003, Jing Yang 0046, Canlong Zhang, Yanping Tang
ICDM3
2021 Domain-Adaptation Person Re-Identification via Style Translation and Clustering
Peiyi Wei, Canlong Zhang, Zhixin Li 0001, Yanping Tang, Zhiwen Wang 0001
ICONIP (1)2
2021 Batch Part-mask Network for person re-identification
abstract
The combination of global feature and local feature is an important method to improve the recognition performance of person re- identification (re- ID). Exiting component-based methods mainly learn local representations by locating regions with specific semantics, which increases the difficulty of learning and lacks effectiveness and robustness for scenarios with large differences. In this paper, we propose Batch Part-mask Network, which takes ResNet-50 which add a self-attention mechanism to effectively improved pixel-level ground truth data as a baseline of the global branch and Batch Part-mask Branch. The global branch encodes the global feature representation, and the Batch Part-mask Branch composed of part 1 Branch and part 2 Branch learns the local detail feature. The network then links features from the two branches and provides a more comprehensive representation of spatial distribution features. Our method performs well in person reID. On the Duke dataset, its Rank-1 accuracy is 88.4%, which is improved by 1 %-2%. The mAP accuracy on the Market-1501 dataset was improved by 2%-3%, reaching 86.3%.
Songyu Chang, Canlong Zhang, Zhixin Li 0001
IJCNN2
2021 Adversarial Training for Image Captioning Incorporating Relation Attention
Zhixin Li 0001, Canlong Zhang, Huifang Ma
PRICAI (1)3
2021 Stable self-attention adversarial learning for semi-supervised semantic image segmentation
Zhixin Li 0001, Canlong Zhang, Huifang Ma
J. Vis. Commun. Image Represent.3
2021 Improve relation extraction with dual attention-guided graph convolutional networks
Zhixin Li 0001, Yaru Sun, Jianwei Zhu, Suqin Tang, Canlong Zhang, Huifang Ma
Neural Comput. Appl.5
2021 Joint deep separable convolution network and border regression reinforcement for object detection
Yu Quan, Zhixin Li 0001, Shengjia Chen, Canlong Zhang, Huifang Ma
Neural Comput. Appl.4
2021 A Semi-supervised Learning Approach Based on Adaptive Weighted Fusion for Automatic Image Annotation
abstract
To learn a well-performed image annotation model, a large number of labeled samples are usually required. Although the unlabeled samples are readily available and abundant, it is a difficult task for humans to annotate large numbers of images manually. In this article, we propose a novel semi-supervised approach based on adaptive weighted fusion for automatic image annotation that can simultaneously utilize the labeled data and unlabeled data to improve the annotation performance. At first, two different classifiers, constructed based on support vector machine and covolutional neural network, respectively, are trained by different features extracted from the labeled data. Therefore, these two classifiers are independently represented as different feature views. Then, the corresponding features of unlabeled images are extracted and input into these two classifiers, and the semantic annotation of images can be obtained respectively. At the same time, the confidence of corresponding image annotation can be measured by an adaptive weighted fusion strategy. After that, the images and its semantic annotations with high confidence are submitted to the classifiers for retraining until a certain stop condition is reached. As a result, we can obtain a strong classifier that can make full use of unlabeled data. Finally, we conduct experiments on four datasets, namely, Corel 5K, IAPR TC12, ESP Game, and NUS-WIDE. In addition, we measure the performance of our approach with standard criteria, including precision, recall, F-measure, N+, and mAP. The experimental results show that our approach has superior performance and outperforms many state-of-the-art approaches.
Zhixin Li 0001, Canlong Zhang, Huifang Ma, Weizhong Zhao, Zhi-Ping Shi 0002
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Integrating Scene Semantic Knowledge into Image Captioning
abstract
Most existing image captioning methods use only the visual information of the image to guide the generation of captions, lack the guidance of effective scene semantic information, and the current visual attention mechanism cannot adjust the focus intensity on the image. In this article, we first propose an improved visual attention model. At each timestep, we calculated the focus intensity coefficient of the attention mechanism through the context information of the model, then automatically adjusted the focus intensity of the attention mechanism through the coefficient to extract more accurate visual information. In addition, we represented the scene semantic knowledge of the image through topic words related to the image scene, then added them to the language model. We used the attention mechanism to determine the visual information and scene semantic information that the model pays attention to at each timestep and combined them to enable the model to generate more accurate and scene-specific captions. Finally, we evaluated our model on Microsoft COCO (MSCOCO) and Flickr30k standard datasets. The experimental results show that our approach generates more accurate captions and outperforms many recent advanced models in various evaluation metrics.
Haiyang Wei, Zhixin Li 0001, Feicheng Huang, Canlong Zhang, Huifang Ma, Zhongzhi Shi
ACM Trans. Multim. Comput. Commun. Appl.4
2020 Image Captioning with Internal and External Knowledge
abstract
Automatically generating a human-like description for a given image is a potential research in artificial intelligence, which has attracted a great of attention recently. Most of the existing attention methods explore the mapping relationships between words in sentence and regions in image, such unpredictable matching manner sometimes causes inharmonious alignments that may reduce the quality of generated captions. In this paper, we make our efforts to reason about more accurate and meaningful captions. We first propose word attention to improve the correctness of visual attention when generating sequential descriptions word-by-word. The special word attention emphasizes on word importance when focusing on different regions of the input image, and makes full use of the internal annotation knowledge to assist the calculation of visual attention. Then, in order to reveal those incomprehensible intentions that cannot be expressed straightforwardly by machines, we inject external knowledge extracted from knowledge graph into the encoder-decoder framework to facilitate meaningful captioning. We validate our model on two freely available captioning benchmarks: Microsoft COCO dataset and Flickr30k dataset. The results demonstrate that our approach achieves state-of-the-art performance and outperforms many of the existing approaches.
Feicheng Huang, Zhixin Li 0001, Shengjia Chen, Canlong Zhang, Huifang Ma
CIKM4
2020 Improving Object Detection with Relation Mining Network
abstract
Due to the deteriorated quality of feature in the propagation process of the neural network, it may be hard for traditional detector to identify a small object by just utilizing information within one region proposal. To overcome the limitation of the traditional object detector, we proposed a graph based relation mining network, to capture the relation information from labels and images. The semantic relation network is proposed to mine the global semantic relation in labels, and the spatial relation network is proposed to capture the local spatial relation in images. The feature representation is further improved by aggregating the outputs of the two networks. Instead of directly disseminating visual features in the network, the relation mining network explores more advanced feature information. Experiments on the PASCAL VOC and MS COCO datasets demonstrate that key relation information significantly improve the performance of object detection with better ability to detect small objects and reasonable bounding box. The results on COCO dataset demonstrate our method can detect objects robustly, increasing the detection performance of small objects from average precision and average recall by 4.7% and 7.6% respectively in performance relative to Faster R-CNN.
Shengjia Chen, Zhixin Li 0001, Feicheng Huang, Canlong Zhang, Huifang Ma
ICDM4
2020 Robust Adversarial Learning For Semi-Supervised Semantic Segmentation
abstract
The semi-supervised semantic segmentation adversarial learning well reduces the use of a large number of manually labeled labels. However, the convolution operator of the generator in the generative adversarial network (GAN) has a local receptive field, so it can only deal with long-range dependencies after passing through multiple convolutional layers. In order to solve this problem, we added two layers of self-attention modules to the GAN generator, and modeled the semantic dependency relationship in the spatial dimension. The self-attention module selectively aggregates the features at each location by weighting and summing the features at all locations. Some recent studies have shown that the adjustment of the discriminator affects the performance of GAN. In order to solve the problem of GAN training instability, we applied spectral normalization to the GAN discriminator and found that this improved the stability of the training. Our method has better performance than existing full/semi-supervised semantic image segmentation techniques.
Zhixin Li 0001, Canlong Zhang, Huifang Ma
ICIP3
2020 Object Detection Using Dual Graph Network
abstract
Most object detection methods focus only on the local information near the region proposal and ignore the object's global semantic relation and local spatial relation information, resulting in limited performance. To capture and explore these important relations, we propose a detection method based on a graph convolutional network (GCN). Two independent relation graph networks are used to obtain the global semantic information of the object in labels and the local spatial information in images. Semantic relation networks can implicitly acquire global knowledge, and by constructing a directed graph on the dataset, each node is represented by the word embedding of labels and then sent to the GCN to obtain high-level semantic representation. The spatial relation network encodes the relation by the positional relation module and the visual connection module, and enriches the object features through local key information from objects. The feature representation is further improved by aggregating the outputs of the two networks. Instead of directly disseminating visual features in the network, the dual-graph network explores more advanced feature information, giving the detector the ability to obtain key relations in labels and region proposals. Experiments on the PASCAL VOC and MS COCO datasets demonstrate that key relation information significantly improve the performance of detection with better ability to detect small objects and reasonable boduning box. The results on COCO dataset demonstrate our method obtains around 32.3% improvement on AP in terms of small objects.
Shengjia Chen, Zhixin Li 0001, Feicheng Huang, Canlong Zhang, Huifang Ma
ICPR4
2020 Cross-media Hash Retrieval Using Multi-head Attention Network
abstract
The cross-media hash retrieval method is to encode multimedia data into a common binary hash space, which can effectively measure the correlation between samples from different modalities. In order to further improve the retrieval accuracy, this paper proposes an unsupervised cross-media hash retrieval method based on multi-head attention network. First of all, we use a multi-head attention network to make better matching images and texts, which contains rich semantic information. At the same time, an auxiliary similarity matrix is constructed to integrate the original neighborhood information from different modalities. Therefore, this method can capture the potential correlations between different modalities and within the same modality, so as to make up for the differences between different modalities and within the same modality. Secondly, the method is unsupervised and does not require additional semantic labels, so it has the potential to achieve large-scale cross-media retrieval. In addition, batch normalization and replacement hash code generation functions are adopted to optimize the model, and two loss functions are designed, which make the performance of this method exceed many supervised deep cross-media hash methods. Experiments on three datasets show that the average performance of this method is about 5 to 6 percentage points higher than the state-of-the-art unsupervised method, which proves the effectiveness and superiority of this method.
Zhixin Li 0001, Chuansheng Xu, Canlong Zhang, Huifang Ma
ICPR4
2020 Reinforcement Learning with Dual Attention Guided Graph Convolution for Relation Extraction
abstract
To better learn the dependency relationship between nodes, we address the relationship extraction task by capturing rich contextual dependencies based on the attention mechanism, and using distributional reinforcement learning to generate optimal relation information representation. This method is called Dual Attention Graph Convolutional Network (DAGCN), to adaptively integrate local features with their global dependencies. Specifically, the samples are represented as nodes on the graph, and the relationships within and between nodes are studied. We consider the influence between node feature locations and associate each location information of the feature with other features. This allows the feature vector to contain a wider range of semantic information to enhance the ability of feature representation. We consider the information features of node dependence, use adjacent nodes to represent their own nodes, and encode the features of node relation, so as to enhance the global dependence between nodes. We sum the outputs of the two attention modules and use reinforcement learning to predict the classification of nodes relationship to further improve feature representation which contributes to more precise extraction results. The results on the common datasets show that the model can obtain more useful information for relational extraction tasks, and achieve better performances on various evaluation indexes.
Zhixin Li 0001, Yaru Sun, Suqin Tang, Canlong Zhang, Huifang Ma
ICPR4
2020 Object Detection Model Based on Scene-Level Region Proposal Self-Attention
abstract
In order to improve the performance of two-stage object detection and consider the importance of scene and semantic information for visual recognition, the neural network of object detection algorithm is studied and analyzed in this paper. The main research work of this paper includes: A scene level region proposal self-attention object detection model based on depth separable convolution is proposed. In order to obtain stronger semantic information and context information of the target scene, the scene-level region proposal self-attention module is reconstructed based on the process of region proposal recognition. The feature map of the output feature pyramid network is sent into three parallel branches: semantic segmentation module, candidate area network module and region proposal self-attention module. At the same time, for the overall performance of the model, a deep separable convolutional network module is constructed on the backbone network, which includes six stages. In the fifth to sixth stage of the network, the separable convolutional network module is integrated respectively. Finally, a object detection method based on border regression network enhancement is proposed to achieve accurate target location. In order to verify the effectiveness of each model, the experimental results of each model are analyzed.
Yu Quan, Zhixin Li 0001, Canlong Zhang, Huifang Ma
ICPR3
2020 Adaptive Graph Convolutional Networks with Attention Mechanism for Relation Extraction
abstract
In the relationship extraction task of NLP, how to effective use of the rich structural information on the dependency tree is a challenging research problem. To better learn the dependency relationship between nodes, we address the relationship extraction task by capturing rich contextual dependencies based on the attention mechanism, and using distributional reinforcement learning to generate optimal relation information representation. Unlike using an attention mechanism to effectively make use of relevant information, we propose a Dual Attention Graph Convolutional Network (DAGCN) to adaptively integrate local features with their global dependencies. Specifically, we append two types of attention modules on top of GCN, which model the semantic interdependencies in spatial and relational dimensions respectively. The position attention module selectively aggregates the feature at each position by a weighted sum of the features at all positions of nodes internal features. Similar features would be related to each other regardless of their distances. Meanwhile, the relation attention module selectively emphasizes interdependent node relations by integrating associated features among all nodes. We sum the outputs of the two attention modules and use reinforcement learning to predict the classification of nodes relationship to further improve feature representation which contributes to more precise extraction results. The results on the TACRED and SemEval datasets show that the model can obtain more useful information for relational extraction tasks, and achieve better performances on various evaluation indexes.
Zhixin Li 0001, Yara Sun, Suqin Tang, Canlong Zhang, Huifang Ma
IJCNN4
2020 Object Detection by Integrating Scene-Level Semantic Information and Border Regression Reinforcement
Yu Quan, Zhixin Li 0001, Canlong Zhang, Huifang Ma
IJCNN3
2020 Robust Semi-Supervised Semantic Segmentation Based on Self-Attention and Spectral Normalization
abstract
The application of adversarial learning for semi-supervised semantic image segmentation based on convolutional neural networks can effectively reduce the number of manually generated labels required in the training process. However, the convolution operator of the generator in the generative adversarial network (GAN) has a local receptive field, so that the long-range dependencies between different image regions can only be modeled after passing through multiple convolutional layers. The present work addresses this issue by introducing a self-attention mechanism in the generator of the GAN to effectively account for relationships between widely separated spatial regions of the input image with supervision based on pixel-level ground truth data. In addition, the adjustment of the discriminator has been demonstrated to affect the stability of GAN training performance. This is addressed by applying spectral normalization to the GAN discriminator during the training process. The proposed stable self-attention adversarial learning semi-supervised semantic image segmentation network is demonstrated to provide superior image segmentation performance compared with the results of current semi-supervised and fully-supervised semantic image segmentation techniques.
Zhixin Li 0001, Canlong Zhang, Huifang Ma
IJCNN3
2020 Multi-level Visual Fusion Networks for Image Captioning
abstract
Image captioning is a multi-modal complex task in machine learning. Traditional methods focus only on entities in visual strategy networks, and can't reason about the relationship between entities and attributes. There are problems of exposure bias and error accumulation in language strategy networks. To this end, this paper proposes a multi-level visual fusion network model based on reinforcement learning. In the visual strategy network, multi-level neural network modules are used to transform visual features into feature sets of visual knowledge. The fusion network generates function words that make the description more fluent, and is used for the interaction between the visual strategy network and the language strategy network. The self-criticism strategy gradient algorithm based on reinforcement learning in language strategy networks is used to achieve end-to-end optimization of visual fusion networks. We evaluated our model on the Flickr 30K and MS-COCO datasets, and verified the accuracy of the model and the diversity of model learning subtitles through experiments. Our model achieves better performance over state-of-the-art methods.
Dongming Zhou 0003, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001
IJCNN2
2020 Image Captioning Based on Visual and Semantic Attention
Haiyang Wei, Zhixin Li 0001, Canlong Zhang
MMM (1)3
2020 The synergy of double attention: Combine sentence-level and word-level attention for image captioning
Haiyang Wei, Zhixin Li 0001, Canlong Zhang, Huifang Ma
Comput. Vis. Image Underst.3
2020 Boost image captioning with knowledge reasoning
Feicheng Huang, Zhixin Li 0001, Haiyang Wei, Canlong Zhang, Huifang Ma
Mach. Learn.4
2020 Classify multi-label images via improved CNN model with adversarial network
Zhixin Li 0001, Canlong Zhang, Huifang Ma
Multim. Tools Appl.3
2019 Cross-Media Image-Text Retrieval Combined with Global Similarity and Local Similarity
abstract
In this paper, we study the problem of image-text matching in order to make the image and text have better semantic matching. In the previous work, people just simply used the pre-training network to extract image and text features and project directly into a common subspace, or change various loss functions on this basis, or use the attention mechanism to directly match the image region proposals and the text phrases. This is not a good match for the semantics of the image and the text. In this study, we propose a method of cross-media retrieval based on global representation and local representation. We constructed a cross-media two-level network to explore better semantic matching between images and text, which contains subnets that handle both global and local features. Specifically, we not only use the self-attention network to obtain a macro representation of the global image but also use the local fine-grained patch with the attention mechanism. Then, we use a two-level alignment framework to promote each other to learn different representations of cross-media retrieval. The innovation of this study lies in the use of more comprehensive features of image and text to design the two kinds of similarity and add them up in some way. Experimental results show that this method is effective in image-text retrieval. Experimental results on the Flickr30K and MS-COCO datasets show that this model has a better recall rate than many of the current advanced cross-media retrieval models.
Zhixin Li 0001, Canlong Zhang
DSAA3
2019 Cross-media Image-Text Retrieval Based on Two-Level Network
Zhixin Li 0001, Fengqi Zhang, Canlong Zhang
ICONIP (1)4
2019 Sentence-Level Semantic Features Guided Adversarial Network for Zhuang Language Part-of-Speech Tagging
abstract
The intelligent information processing of the standard Zhuang language spoken mainly in Southern China is presently in its infancy, and lacks a well-defined language corpus and automatic part-of-speech tagging methods. Therefore, this study proposes an adversarial part-of-speech tagging method based on reinforcement learning, which solves the problems associated with a lack of a language corpus, time-consuming laborious manual marking, and the low performance of machine marking. Firstly, we construct a markup dictionary based on the grammatical characteristics of standard Zhuang and the Penn Chinese Treebank. Secondly, a dependency syntax analysis is applied for constructing the semantic information feature vectors of sentences, and long short-term memory is adopted as the policy network architecture to enhance available information using recurrent memory, and a conditional random field is employed as the discriminant network to perform label inference with global normalization. Finally, we use reinforcement learning as the model framework, target parts of speech as the feedback of the environment, and then obtain the optimal policy through adversarial learning. The results show that the combination of reinforcement learning and adversarial network alleviates the dependence of the model on the training corpus to some extent, and can quickly and effectively expand the scale of the annotation dictionary for the Zhuang language, thereby obtaining better labeling results.
Zhixin Li 0001, Yaru Sun, Suqin Tang, Canlong Zhang, Huifang Ma
ICTAI4
2019 Sparse High-Level Attention Networks for Person Re-Identification
abstract
When extracting convolutional features from person images with low resolution, a large amount of available information will be lost due to the pooling, which will lead to the reduction of the accuracy of person classification models. This paper proposes a new classification model, which can effectively to reduce the loss of important information about the convolutional neural works. Firstly, the SE module in the Squeeze-and-Excitation Networks (SENet) is extracted and normalized to generate the Normalized Squeeze-and-Excitation (NSE) module. Then, 4 NSE modules are applied to the convolutional layers of ResNet. Finally, a Sparse Normalized Squeeze-and-Excitation Network (SNSENet) is constructed by adding 4 shortcut connections between the convolutional layers. The experimental results of Market-1501 show that the rank-1 of SNSE-ResNet-50 is 3.7% and 4.2% higher than that of SE-ResNet-50 and ResNet-50 respectively, it has done well in other person re-identification datasets.
Sheng Xie, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001
ICTAI2
2019 Automatic Image Annotation based on Co-Training
abstract
To learn a well-performed image annotation model, a large number of labeled samples are usually required. Although the unlabeled samples are readily available and abundant, it's a difficult task for humans to annotate large amounts of images manually. In this paper, we propose a novel semi-supervised approach based on co-training algorithm for automatic image annotation, which can utilize the labeled data and unlabeled data for the system simultaneously. Firstly, two different classifiers, namely the CNN (convolutional neural network) and the LDA-SVM, are constructed by all the labeled data. These two classifiers are independently represented as different feature views. Then, the most confident data with relevant pseudo-labels are chosen and amalgamated with the whole labeled dataset. After that, the two classifiers are retrained with the new labeled dataset until a stop condition is reached. In each iteration process, the unlabeled samples are labeled by high confidence pseudo-labels that are estimated by an adaptive weighted fusion method. Finally, we conduct experiments on two datasets, namely, IAPR TC-2 and NUS-WIDE, and measure the performance of the model with standard criteria, including precision, recall, F-measure, N+ and mAP. The experimental results show that our approach has superior annotation performance and outperforms many state-of-the-art automatic image annotation approaches.
Zhixin Li 0001, Canlong Zhang, Huifang Ma, Weizhong Zhao
IJCNN3
2019 Image Captioning Based On Sentence-Level And Word-Level Attention
abstract
Existing attention models of image captioning typically extract only word-level attention information. i.e., the attention mechanism extracts local attention information from the image to generate the current word. We propose an image captioning approach based on self-attention to utilize image features more effectively. The self-attention mechanism can extract sentence-level attention information with richer visual representation from images. Furthermore, we propose a double attention model. The model combines sentence-level and word-level attention information to better simulate human perception system. We implement supervision and optimization in the intermediate stage of the model to solve over-fitting and information interference problems, and we apply reinforcement learning to two-stage training to optimize the evaluation metrics of the model. Finally, we evaluate our model on MSCOCO dataset. The experimental results show that our approach can generate more accurate and richer captions, and outperforms many state-of-the-art image captioning approaches on various evaluation metrics.
Haiyang Wei, Zhixin Li 0001, Canlong Zhang, Yu Quan
IJCNN3
2019 Object Detection by Combining Deep Dilated Convolutions Network and Light-Weight Network
Yu Quan, Zhixin Li 0001, Canlong Zhang
KSEM (1)3
2019 Collaborating CNN and SVM for Automatic Image Annotation
abstract
To learn a well-performed image annotation model, a large number of labeled samples are usually required. In this paper, we propose a novel semi-supervised approach based on adaptive weighted fusion for automatic image annotation, which can utilize the labeled data and unlabeled data simultaneously. Firstly, two different classifiers, namely the CNN (convolutional neural network) and the LDA-SVM, are constructed by all the labeled data. These two classifiers are independently represented as different feature views. Then, the most confident data with relevant pseudo-labels are chosen and amalgamated with the whole labeled dataset. After that, the two classifiers are retrained with the new labeled dataset until a stop condition is reached. In each iteration process, the unlabeled samples are labeled by high confidence pseudo-labels that are estimated by an adaptive weighted fusion strategy. Finally, we conduct experiments on two datasets, namely IAPR TC12 and NUS-WIDE, and measure the performance of the model with standard criteria, including precision, recall, F-measure, N+ and mAP. The experimental results show that our approach outperforms many state-of-the-art automatic image annotation approaches.
Zhixin Li 0001, Canlong Zhang, Huifang Ma, Weizhong Zhao
ICMR3
2019 D_dNet-65 R-CNN: Object Detection Model Fusing Deep Dilated Convolutions and Light-Weight Networks
Yu Quan, Zhixin Li 0001, Fengqi Zhang, Canlong Zhang
PRICAI (3)4
2019 Joint spatiograms for multi-modality tracking with online update
Canlong Zhang, Yanping Tang, Zhixin Li 0001, Zhiwen Wang 0001
Pattern Recognit. Lett.1
2018 Parallel Connecting Deep and Shallow CNNs for Simultaneous Detection of Big and Small Objects
Canlong Zhang, Dongcheng He, Zhixin Li 0001, Zhiwen Wang 0001
PRCV (4)1
2018 An Improved Convolutional Neural Network Model with Adversarial Net for Multi-label Image Classification
Zhixin Li 0001, Canlong Zhang
PRICAI3
2018 Joint compressive representation for multi-feature tracking
Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001
Neurocomputing1
2018 A hybrid architecture based on CNN for cross-modal semantic instance annotation
Yongzhe Zheng, Zhixin Li 0001, Canlong Zhang
Multim. Tools Appl.3
2017 Analysis of Influencing Factors of Shooting Rate Based on Trajectory Prediction of the Basketball
abstract
The main factors affecting the shooting rate are the angle of shooting, the speed of shot, the height of the shot and the trajectory of the basketball under the influence of direction and speed of wind. In this paper, we propose the method to analyze the influence factors of shooting rate based on trajectory prediction. It can accurately get the motion trajectory of basketball through the forecast. We use forecast analysis with mechanical theory and motion trajectory to get shot height, angle and the release speed of the shooting of the impact, providing theoretical basis for shooting training. The experimental results show that the shooting parameters can be greatly improved by shooting the relevant parameters of the shooting point.
Zhiwen Wang 0001, Lianyuan Jiang, Canlong Zhang, Zhenghuan Hu
WISA5
2017 Automatic image annotation using fuzzy association rules and decision tree
Zhixin Li 0001, Lingzhi Li 0004, Kaobi Yan, Canlong Zhang
Multim. Syst.4
2016 A compressive tracking based on time-space Kalman fusion model
Xiao Yun, Zhongliang Jing, Bo Jin 0003, Canlong Zhang
Sci. China Inf. Sci.5
2016 Hybrid Histogram of Oriented Optical Flow for Abnormal Behavior Detection in Crowd Scenes
abstract
Abnormal behavior detection in crowd scenes has received considerable attention in the field of public safety. Traditional motion models do not account for the continuity of motion characteristics between frames. In this paper, we present a new feature descriptor, called the hybrid optical flow histogram. By importing the concept of acceleration, our method can indicate the change of speed in different directions of a movement. Therefore, our descriptor contains more information on the movement. We also introduce a spatial and temporal region saliency determination method to extract the effective motion area only for samples, which could effectively reduce the computational costs, and we apply a sparse representation to detect abnormal behaviors via sparse reconstruction costs. Sparse representation has a high rate of recognition performance and stability. Experiments involving the UMN datasets and the videos taken by us show that our method can effectively identify various types of anomalies and that the recognition results are better than existing algorithms.
Qiang Wang 0034, Qiao Ma, Chaohui Lu, Canlong Zhang
Int. J. Pattern Recognit. Artif. Intell.5
2016 A New multi-instance multi-label learning approach for image and text classification
Kaobi Yan, Zhixin Li 0001, Canlong Zhang
Multim. Tools Appl.3
2013 Locally discriminative stable model for visual tracking with clustering and principle component analysis
abstract
The challenge of visual tracking mainly comes from intrinsic appearance variations of the target and extrinsic environment changes around the target in a long duration, so the tracker that can simultaneously tolerate these variabilities is largely expected. In this study, the authors propose a new tracking approach based on discriminative stable regions (DSRs). The DSRs are obtained based on the criterion of maximal local entropy and spatial discrimination, which enables the tracker to handle well distractors and appearance variations. The collaborative tracking incorporated hierarchical clustering can tolerate motion noise and occlusions. In addition, as an efficient tool, the principle component analysis is used to discover the potential affine relation between DSR and the target, which timely adapts to the shape deformation of the target. Extensive experiments show that the proposed method achieves superior performance in many challenging target tracking tasks.
Canlong Zhang, Zhongliang Jing, Yanping Tang, Bo Jin 0003
IET Comput. Vis.1
2013 A sparse proximal Newton splitting method for constrained image deblurring
Han Pan, Zhongliang Jing, Rongli Liu, Bo Jin 0003, Canlong Zhang
Neurocomputing6
2013 Robust visual tracking using discriminative stable regions and K-means clustering
Canlong Zhang, Zhongliang Jing, Han Pan, Bo Jin 0003, Zhixin Li 0001
Neurocomputing1
2012 A dual-kernel-based tracking approach for visual target
Canlong Zhang, Zhongliang Jing, Bo Jin 0003, Zhixin Li 0001
Sci. China Inf. Sci.1