Zhong Zhang 0001

dblp:28/1568-1 · DBLP profile ↗
← Back
58ranked-venue papers
26as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 12 first-author · 8 since 2021Artificial intelligence and machine learning · 20 · 10 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 7 since 2021Computer networks · 9 · 4 first-author · 3 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Pedestrian trajectory prediction using multi-cue transformer
Yanlong Tian, Xiaoting Fan, Zhong Zhang 0001, Xinshan Zhu
J. Vis. Commun. Image Represent.5
2025 Cross-modal Shared Concept Learning for Text-to-Image Person Retrieval
abstract
Text-to-image person retrieval aims to match target pedestrian images based on a text query. Existing methods mainly learn feature alignment between texts and pedestrian images from global and local perspectives. However, they ignore the matching ambiguity problem caused by different modal characteristics during the alignment learning, thereby resulting in sub-optimal performance. In this paper, we propose a cross-modal Shared Concept Learning (SCL) method, which reformulates alignment as learning shared cross-modal concepts in order to improve the semantic consistency between modalities. Specifically, we design a Shared Concept Perception (SCP) module to capture shared cross-modal concepts from global-local perspectives using decoupled visual semantic concepts guided by text concepts within the identity. Furthermore, we propose a Visual-driven Concept Matching (VCM) module to learn the identity invariance of concepts in each text under the guidance of the visual information. Extensive experiments on three text-to-image person retrieval benchmarks demonstrate the effectiveness and superiority of SCL.
Di He 0008, Xinshan Zhu, Zhong Zhang 0001
ICME5
2025 Unsupervised Person Reidentification Using Stripe-Driven Fusion Transformer Network
abstract
In recent years, some methods utilize a transformer as the backbone to model the long‐range context dependencies, reflecting a prevailing trend in unsupervised person reidentification (Re‐ID) tasks. However, they only explore the global information through interactive learning in the framework of the transformer, which ignores the learning of the part information in the interaction process for pedestrian images. In this study, we present a novel transformer network for unsupervised person Re‐ID, a stripe‐driven fusion transformer (SDFT), designed to simultaneously capture the global interaction and the part interaction when modeling the long‐range context dependencies. Meanwhile, we present a stripe‐driven regularization (SDR) to constrain the part aggregation features and the global features by considering the consistency principle from the aspects of the features and the clusters, aiming to improve the representational capacity of the features. Furthermore, to investigate the relationships between local regions of pedestrian images, we present a stripe‐driven contrastive loss (SDCL) to learn discriminative part features from the perspectives of pedestrian identity and stripes. The proposed method has undergone extensive validations on publicly available unsupervised person Re‐ID benchmarks, and the experimental results confirm its superiority and effectiveness.
Zeyu Zang, Shuang Liu 0001, Zhong Zhang 0001, Xinshan Zhu
IET Softw.4
2025 Cross-Modal Alignment Enhancement Network for Text-to-Image Person Re-Identification
Di He 0008, Xinshan Zhu, Bin Li 0028, Shenglu Yue, Zhong Zhang 0001
IEEE Internet Things J.5
2025 Cross-domain person re-identification via learning Heterogeneous Pseudo Labels
Zhong Zhang 0001, Di He 0008, Shuang Liu 0001
Pattern Recognit.1
2025 Cloud-Type Classification Using Multimodal Integration Transformer Based on Cloud Images and Millimeter-Wave Cloud Radar Observations
abstract
The existing methods fail to simultaneously utilize the appearance information and the internal structure of clouds for cloud type classification, resulting in incomplete cloud representation. In this paper, we exploit cloud images and Millimeter-wave Cloud Radar (MMCR) observations for cloud type classification, and propose a novel Transformer network named Multi-modal Integration Transformer (MMITrans) to describe completed information of clouds. To this end, we design MMITrans as three subnetworks, i.e., vision subnetwork, MMCR subnetwork and multi-modal fusion network. Specifically, we extract the visual features from the cloud images through the vision subnetwork. Meanwhile, we first convert MMCR observations into several cloud-related indicators, and propose the Indicator-Tokenization to effectively tokenize them and obtain the indicator features using the MMCR subnetwork. Furthermore, we propose the Multi-modal Cross Attention in the multi-modal fusion network to sufficiently fuse the visual features and the indicator features in a multiple-input way. We perform a series of experiments on Cloud images and Millimeter-wave cloud radar observations Dataset, i.e., CMD-Beijing and CMD-Gansu, and the experimental results demonstrate the superiority of the proposed MMITrans.
Shuang Liu 0001, Zeyu Zang, Zhong Zhang 0001, Shuzhen Hu, Baihua Xiao
IEEE Trans. Geosci. Remote. Sens.3
2025 Completed Interaction Networks for Pedestrian Trajectory Prediction
abstract
The social and environmental interactions, as well as the pedestrian goal are crucial for pedestrian trajectory prediction. This is because they could learn both complex interactions in the scenes and the intentions of the pedestrians. However, most existing methods either learn the one-moment social interactions, or supervise the pedestrian trajectories using long-term goal, resulting in suboptimal prediction performances. In this paper, we propose a novel network named Completed Interaction Network (CINet) to simultaneously consider the social interactions in all moments, the environmental interactions and the short-term goal of pedestrians in a unified framework for pedestrian trajectory prediction. Specifically, we propose the Spatio-Temporal Transformer Layer (STTL) to fully mine the spatio-temporal information among historical trajectories of all pedestrians in order to obtain the social interactions in all moments. Additionally, we present the Gradual Goal Module (GGM) to capture the environmental interactions under the supervision of the short-term goal, which is beneficial to understanding the intentions of the pedestrian. Afterwards, we employ the cross-attention to effectively integrate the all-moment social and environmental interactions. The experimental results on three standard pedestrian datasets, i.e., ETH/UCY, SDD and inD demonstrate that our method achieves a new state-of-the-art performance. Furthermore, the visualization results indicate that our method could predict trajectories more reasonably in complex scenarios such as sharp turns, infeasible areas and so on.
Zhong Zhang 0001, Jianglin Zhou, Shuang Liu 0001, Baihua Xiao
IEEE Trans. Multim.1
2024 A comprehensive review of image retargeting
Xiaoting Fan, Zhong Zhang 0001, Baihua Xiao, Tariq S. Durrani
Neurocomputing2
2024 SSN: Shift Suppression Network for Endogenous Shift of Photovoltaic Defect Detection
abstract
Most of the existing photovoltaic (PV) defect detection methods are based on the assumption that the training and testing samples satisfy the independent identically distributed. However, in real PV scenarios, the endogenous shift problem exists widely, including background style shift and defect instance shift, which seriously affects the performance of detectors. In this article, we propose a novel network called shift suppression network (SSN) for the endogenous shift of PV defect detection, which consists of two core components, background style suppression (BSS) module and cross-layer graph reasoning (CGR) module. Specifically, BSS uses channel statistics matching alignment to adaptively suppress background style shifts without reliance on unknown domain data. CGR learns semantic dependencies between multiscale channel feature maps through cross-layer interaction, which improves the discriminative ability and localization ability in the face of defect instance shifts. To advance the study of endogenous shift, we provide the first electroluminescence (EL) endogenous shift dataset for PV modules, which creates by three groups of EL images collected at different times, totaling 16 323. The comprehensive evaluation results show that the SSN outperforms the state-of-the-art methods. Furthermore, we design an intelligent defect detection system for PV modules based on SSN. The system has higher detection accuracy and an intuitive visual interface, which can meet the actual detection function and requirements of the production line.
Shenshen Zhao, Haiyong Chen, Zhong Zhang 0001
IEEE Trans. Ind. Informatics4
2024 Completed Part Transformer for Person Re-Identification
abstract
Recently, part information of pedestrian images has been demonstrated to be effective for person re-identification (ReID), but the part interaction is ignored when using Transformer to learn long-range dependencies. In this article, we propose a novel transformer network named Completed Part Transformer (CPT) for person ReID, where we design the part transformer layer to learn the completed part interaction. The part transformer layer includes the intra-part layer and the part-global layer, where they consider long-range dependencies from the aspects of the intra-part interaction and the part-global interaction, simultaneously. Furthermore, in order to overcome the limitation of fixed number of the patch tokens in the transformer layer, we propose the Adaptive Refined Tokens (ART) module to focus on learning the interaction between the informative patch tokens in the pedestrian image, which improves the discrimination of the pedestrian representation. Extensive experimental results on four person ReID datasets, i.e., MSMT17, Market1501, DukeMTMC-reID, and CUHK03, demonstrate that the proposed method achieves a new state-of-the-art performance, e.g., it achieves 68.0% mAP and 84.6% Rank-1 accuracy on MSMT17.
Zhong Zhang 0001, Di He 0008, Shuang Liu 0001, Baihua Xiao, Tariq S. Durrani
IEEE Trans. Multim.1
2023 Integration graph attention network and multi-centre constrained loss for cross-modality person re-identification
abstract
Abstract Cross‐modality person re‐identification is a challenging task due to the large visual appearance difference between RGB and infrared images. Existing studies mainly focus on learning local features and ignore the correlation between local features. In this paper, the Integration Graph Attention Network is proposed to learn the completed correlation between local features via the graph structure. To this end, the authors learn the coarse‐fine attention weights to aggregate the local features by considering local detail and global information. Furthermore, the Multi‐Centre Constrained Loss is proposed to optimise the feature similarity by constraining the centres of modality and identity. It simultaneously utilises three kinds of centre constraints, that is intra‐identity centre constraint, modality centre constraint, and inter‐identity centre constraint, in order to reduce the influence of modality information explicitly. The proposed method is evaluated on two standard benchmark datasets, that is SYSU‐MM01 and RegDB, and the results demonstrate that the authors’ method achieves better performance than the state‐of‐the‐art methods, for example, surpassing NFS by 4.8% and 6.0% mAP on the single‐shot setting in All‐search and Indoor‐search modes, respectively.
Di He 0008, Jingrui Zhang, Zhong Zhang 0001, Shuang Liu 0001, Tariq S. Durrani
IET Comput. Vis.3
2023 Cross-modality person re-identification using hybrid mutual learning
abstract
Abstract Cross‐modality person re‐identification (Re‐ID) aims to retrieve a query identity from red, green, blue (RGB) images or infrared (IR) images. Many approaches have been proposed to reduce the distribution gap between RGB modality and IR modality. However, they ignore the valuable collaborative relationship between RGB modality and IR modality. Hybrid Mutual Learning (HML) for cross‐modality person Re‐ID is proposed, which builds the collaborative relationship by using mutual learning from the aspects of local features and triplet relation. Specifically, HML contains local‐mean mutual learning and triplet mutual learning where they focus on transferring local representational knowledge and structural geometry knowledge so as to reduce the gap between RGB modality and IR modality. Furthermore, Hierarchical Attention Aggregation is proposed to fuse local feature maps and local feature vectors to enrich the information of the classifier input. Extensive experiments on two commonly used data sets, that is, SYSU‐MM01 and RegDB verify the effectiveness of the proposed method.
Zhong Zhang 0001, Sen Wang 0007, Shuang Liu 0001, Baihua Xiao, Tariq S. Durrani
IET Comput. Vis.1
2023 Integration Transformer for Ground-Based Cloud Image Segmentation
abstract
Recently, convolutional neural network (CNN) dominates the ground-based cloud image segmentation task, but disregards the learning of long-range dependencies due to the limited size of filters. Although Transformer-based methods could overcome this limitation, they only learn long-range dependencies at a single scale, hence failing to capture multi-scale information of cloud image. The multi-scale information is beneficial to ground-based cloud image segmentation, because the features from small scales tend to extract detailed information while features from large scales have the ability to learn global information. In this paper, we propose a novel deep network named Integration Transformer (InTransformer), which builds long-range dependencies from different scales. To this end, we propose the Hybrid Multi-head Transformer Block (HMTB) to learn multi-scale long-range dependencies, and hybridize CNN and HMTB as the encoder at different scales. The proposed InTransformer hybridizes CNN and Transformer as the encoder to extract multi-scale representations, which learns both local information and long-range dependencies with different scales. Meanwhile, in order to fuse the patch tokens with different scales, we propose Mutual Cross-Attention Module (MCAM) for the decoder of InTransformer which could adequately interact multi-scale patch tokens in a bidirectional way. We have conducted a series of experiments on large ground-based cloud detection database TLCDD and SWIMSEG. The experimental results show that the performance of our method outperforms other methods, proving the effectiveness of the proposed InTransformer.
Shuang Liu 0001, Zhong Zhang 0001, Xiaozhong Cao, Tariq S. Durrani
IEEE Trans. Geosci. Remote. Sens.3
2023 Hybrid Contrastive Learning for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (Re-ID) aims to learn discriminative features without human-annotated labels. Recently, contrastive learning has provided a new prospect for unsupervised person Re-ID, and existing methods primarily constrain the feature similarity among easy sample pairs. However, the feature similarity among hard sample pairs is neglected, which yields suboptimal performance in unsupervised person Re-ID. In this paper, we propose a novel Hybrid Contrastive Model (HCM) to perform the identity-level contrastive learning and the image-level contrastive learning for unsupervised person Re-ID, which adequately explores feature similarities among hard sample pairs. Specifically, for the identity-level contrastive learning, an identity-based memory is constructed to store pedestrian features. Accordingly, we define the dynamic contrast loss to identify identity information with dynamic factor for distinguishing hard/easy samples. As for the image-level contrastive learning, an image-based memory is established to store each image feature. We design the sample constraint loss to explore the similarity relationship between hard positive and negative sample pairs. Furthermore, we optimize the two contrastive learning processes in one unified framework to make use of their own advantages as so to constrain the feature distribution for extracting potential information. Extensive experiments demonstrate that the proposed HCM distinctly outperforms existing methods.
Tongzhen Si, Fazhi He, Zhong Zhang 0001, Yansong Duan
IEEE Trans. Multim.3
2022 Contrastive Meta-Learning for Drug-Target Binding Affinity Prediction
abstract
Effective drug-target binding affinity (DTA) prediction is essential for drug discovery and development. The development of machine learning techniques considerably advances it. However, the cold-start problems in DTA prediction are still under-explored, which significantly degrades prediction performances on novel drugs and novel targets. In this paper, we propose a contrastive meta-learning (CML) framework to address these issues. We define drug-anchored tasks and target-anchored tasks, which enables the employment of meta-learning to accumulate common knowledge from various tasks so as to adapt to new tasks faster and better. Besides, we utilize a task inequality loss to measure task disparities and enhance model sensitivities to new tasks. We also propose a contrastive learning block (CLB) to explore correlations among drug-target pairs across tasks, which facilitates DTA prediction performance improvements. We compare CML with various baselines on two benchmarks and comparison results show that CML outperforms or achieves competitive results to its competitors.
Sihan Xu, Xiangrui Cai, Zhong Zhang 0001, Hua Ji
BIBM4
2022 Ground-Based Cloud Detection Using Multiscale Attention Convolutional Neural Network
abstract
Cloud detection plays a significant role in ground-based remote sensing observation, and it is quite challenging due to the variations in illumination and cloud form, and the vague boundaries between cloud and sky. In this letter, we propose a novel deep model named multiscale attention convolutional neural network (MACNN) for ground-based cloud detection, which possesses a symmetric encoder–decoder structure. For accurate cloud detection, we design the multiscale module in MACNN to obtain different receptive fields by using different hole rates for the filters, and meanwhile, we propose the attention module in MACNN to learn the attention coefficients in order to reflect different importance of pixels. Furthermore, we release the Tianjin Normal University (TJNU) cloud detection database (TCDD) to provide a comparative study for different methods, and to the best of our knowledge, it is the largest cloud detection database. We conduct a series of experiments on the TCDD, and the experimental results demonstrate that the proposed MACNN outperforms state-of-the-art methods in five quantitative evaluation criteria.
Zhong Zhang 0001, Shuzhen Yang, Shuang Liu 0001, Baihua Xiao, Xiaozhong Cao
IEEE Geosci. Remote. Sens. Lett.1
2022 Cross-Domain Person Re-Identification Using Heterogeneous Convolutional Network
abstract
Person re-identification (Re-ID) is a challenging task due to variations in pedestrian images, especially in cross-domain scenarios. The existing cross-domain person Re-ID approaches extract the feature from single pedestrian image, but they ignore the correlations among pedestrian images. In this paper, we propose Heterogeneous Convolutional Network (HCN) for cross-domain person Re-ID, which learns the appearance information of pedestrian images and the correlations among pedestrian images simultaneously. To this end, we first utilize Convolutional Neural Network (CNN) to extract the appearance features for pedestrian images. Then we construct a graph in the target dataset where the appearance features are treated as the nodes and the similarity represents the linkage between the nodes. Afterwards, we propose Dual Graph Convolution (DGConv) to explicitly learn the correlation information from the similar and dissimilar samples, which could avoid the over-smoothing caused by the fully connected graph. Furthermore, we design HCN as a multi-branch structure to mine the structural information of pedestrians. We conduct extensive evaluations for HCN on three datasets, i.e. Market-1501, DukeMTMC-reID and MSMT17, and the results demonstrate that HCN is superior to the state-of-the-art methods.
Zhong Zhang 0001, Shuang Liu 0001, Baihua Xiao, Tariq S. Durrani
IEEE Trans. Circuits Syst. Video Technol.1
2022 Ground-Based Remote Sensing Cloud Classification via Context Graph Attention Network
abstract
Most ground-based remote sensing cloud classification methods focus on learning representation features for cloud images while ignoring the correlations among cloud images. Recently, graph convolutional network (GCN) is applied to provide the correlations for ground-based remote sensing cloud classification, in which the graph convolutional layer aggregates information from the connected nodes of graph in a weighted way. However, the weights assigned by GCN cannot reflect the importance of connected nodes precisely, which declines the discrimination of the aggregated features (AFs). To overcome the limitation, in this article, we propose the context graph attention network (CGAT) for ground-based remote sensing cloud classification. Specifically, the context graph attention layer (CGA layer) of CGAT is proposed to learn the context attention coefficients (CACs) and obtain the AFs of nodes based on the CACs. We compute the CACs not only considering the two connected nodes but also their neighborhood nodes in order to stabilize the aggregation process. In addition, we propose to utilize two different transformation matrices to transform the node and its connected nodes into new feature spaces, which could enhance the discrimination of AFs. We concatenate the AFs with the deep features (DFs) as final representations for cloud classification. Since existing ground-based cloud data sets (GCDs) have limited cloud images, we release a new data set named GCD that is the largest one for ground-based cloud classification. We conduct a series of experiments on GCD, and the experimental results verify the effectiveness of CGAT.
Shuang Liu 0001, Linlin Duan, Zhong Zhang 0001, Xiaozhong Cao, Tariq S. Durrani
IEEE Trans. Geosci. Remote. Sens.3
2022 Ground-Based Remote Sensing Cloud Detection Using Dual Pyramid Network and Encoder-Decoder Constraint
abstract
Many methods for ground-based remote sensing cloud detection learn representation features using the encoder–decoder structure. However, they only consider the information from single scale, which leads to incomplete feature extraction. In this article, we propose a novel deep network named dual pyramid network (DPNet) for ground-based remote sensing cloud detection, which possesses an encoder–decoder structure with dual pyramid pooling module (DPPM). Specifically, we process the feature maps of different scales in the encoder through dual pyramid pooling. Then, we fuse the outputs of the dual pyramid pooling in the same pyramid level using the attention fusion. Furthermore, we propose the encoder–decoder constraint (EDC) to relieve information loss in the process of encoding and decoding. It constrains the values and the gradients of probability maps from the encoder and the decoder to be consistent. Since the number of cloud images in the publicly available databases for ground-based remote sensing cloud detection is limited, we release the TJNU Large-scale Cloud Detection Database (TLCDD) that is the largest database in this field. We conduct a series of experiments on TLCDD, and the experimental results verify the effectiveness of the proposed method.
Zhong Zhang 0001, Shuzhen Yang, Shuang Liu 0001, Xiaozhong Cao, Tariq S. Durrani
IEEE Trans. Geosci. Remote. Sens.1
2021 Learning Hybrid Relationships for Person Re-identification
abstract
Recently, the relationship among individual pedestrian images and the relationship among pairwise pedestrian images have become attractive for person re-identification (re-ID) as they effectively improve the ability of feature representation. In this paper, we propose a novel method named Hybrid Relationship Network (HRNet) to learn the two types of relationships in a unified framework that makes use of their own advantages. Specifically, for the relationship among individual pedestrian images, we take the features of pedestrian images as the nodes to construct a locally-connected graph, so as to improve the discriminative ability of nodes. Meanwhile, we propose the consistent node constraint to inject the identity information into the graph learning process and guide the information to propagate accurately. As for the relationship among pairwise pedestrian images, we treat the feature differences of pedestrian images as the nodes to construct a fully-connected graph so as to estimate robust similarity of nodes. Furthermore, we propose the inter-graph propagation to alleviate the information loss for the fully-connected graph. Extensive experiments on Market-1501, DukeMTMCreID, CUHK03 and MSMT17 demonstrate that the proposed HRNet outperforms the state-of-the-art methods.
Shuang Liu 0001, Wenmin Huang, Zhong Zhang 0001
AAAI3
2021 Person Re-Identification Using Heterogeneous Local Graph Attention Networks
abstract
Recently, some methods have focused on learning local relation among parts of pedestrian images for person re-identification (Re-ID), as it offers powerful representation capabilities. However, they only provide the intra-local relation among parts within single pedestrian image and ignore the inter-local relation among parts from different images, which results in incomplete local relation information. In this paper, we propose a novel deep graph model named Heterogeneous Local Graph Attention Networks (HLGAT) to model the inter-local relation and the intra-local relation in the completed local graph, simultaneously. Specifically, we first construct the completed local graph using local features, and we resort to the attention mechanism to aggregate the local features in the learning process of inter-local relation and intra-local relation so as to emphasize the importance of different local features. As for the inter-local relation, we propose the attention regularization loss to constrain the attention weights based on the identities of local features in order to describe the inter-local relation accurately. As for the intra-local relation, we propose to inject the contextual information into the attention weights to consider structure information. Extensive experiments on Market-1501, CUHK03, DukeMTMC-reID and MSMT17 demonstrate that the proposed HLGAT outperforms the state-of-the-art methods.
Zhong Zhang 0001, Haijia Zhang, Shuang Liu 0001
CVPR1
2021 Dynamically occluded samples via adversarial learning for person re-identification in sensor networks
Wenmin Huang, Shuang Liu 0001, Ruiling Luo, Tongzhen Si, Zhong Zhang 0001
Ad Hoc Networks5
2021 Visible Thermal Person Reidentification via Mutual Learning Convolutional Neural Network in 6G-Enabled Visual Internet of Things
abstract
Visible thermal person reidentification (VT Re-ID) in the 6G-enabled Visual Internet of Things (VIoT) is essential for implementing 24-h surveillance in wide-area space. The existing methods mainly focus on reducing the visual difference between visible and thermal visual streams (V-streams). However, they cannot fully exploit the information of visible and thermal V-streams, which leads to suboptimal representations. In this article, we propose a novel deep network termed the mutual learning convolutional neural network (MLCNN) to transfer useful information between visible and thermal V-streams for VT Re-ID in 6G-enabled VIoT. The proposed MLCNN consists of the feature generation module and the mutual learning module. The feature generation module aims to map the input samples into a more discriminative feature space in order to decrease the intramodality difference. The mutual learning module employs the identification loss to supervise the predictive identity accuracy and utilizes the Kullback–Leibler divergence to constitute the mutual learning loss for useful information transfer. Extensive experiments on two popular visible thermal data sets (SYSU-MM01 and RegDB) prove the effectiveness of the proposed MLCNN.
Zhong Zhang 0001, Sen Wang 0007
IEEE Internet Things J.1
2021 Part-guided graph convolution networks for person re-identification
Zhong Zhang 0001, Haijia Zhang, Shuang Liu 0001, Tariq S. Durrani
Pattern Recognit.1
2020 Deep tensor fusion network for multimodal ground-based cloud classification in weather station networks
Shuang Liu 0001, Zhong Zhang 0001
Ad Hoc Networks3
2020 Person re-identification using Hybrid Task Convolutional Neural Network in camera sensor networks
Shuang Liu 0001, Wenmin Huang, Zhong Zhang 0001
Ad Hoc Networks3
2020 Cross-domain person re-identification using Dual Generation Learning in camera sensor networks
Zhong Zhang 0001, Shuang Liu 0001
Ad Hoc Networks1
2020 Fuzzy Multilayer Clustering and Fuzzy Label Regularization for Unsupervised Person Reidentification
abstract
Unsupervised person reidentification has received more attention due to its wide real-world applications. In this paper, we propose a novel method named fuzzy multilayer clustering (FMC) for unsupervised person reidentification. The proposed FMC learns a new feature space using a multilayer perceptron for clustering in order to overcome the influence of complex pedestrian images. Meanwhile, the proposed FMC generates fuzzy labels for unlabeled pedestrian images, which simultaneously considers the membership degree and the similarity between the sample and each cluster. We further propose the fuzzy label regularization (FLR) to train the convolutional neural network (CNN) using pedestrian images with fuzzy labels in a supervised manner. The proposed FLR could regularize the CNN training process and reduce the risk of overfitting. The effectiveness of our method is validated on three large-scale person reidentification databases, i.e., Market-1501, DukeMTMC-reID, and CUHK03.
Zhong Zhang 0001, Meiyan Huang, Shuang Liu 0001, Baihua Xiao, Tariq S. Durrani
IEEE Trans. Fuzzy Syst.1
2020 Multimodal Ground-Based Remote Sensing Cloud Classification via Learning Heterogeneous Deep Features
abstract
Recently, multimodal cloud samples are utilized to learn completed feature representations for cloud classification. However, the existing methods neglect the related information from other multimodal cloud samples in the learning process, which leads to inadequate learning. In this article, we propose a novel deep model to learn heterogeneous deep features (HDFs) for multimodal ground-based remote sensing cloud classification. Specifically, we first design the convolutional neural network (CNN) extractor to combine the visual information and the multimodal information (MI) to obtain the CNN-based features of multimodal cloud samples. Afterward, we treat the CNN-based features of multimodal cloud samples as the nodes of graph, and utilize the similarity between nodes as the adjacency matrix. We feed the graph and the adjacency matrix into the graph convolutional network (GCN) extractor to obtain the GCN-based features that could capture correlations among multimodal cloud samples using graph convolutional layers. After obtaining CNN-based features and GCN-based features, we concatenate the two kinds of heterogeneous features to represent the multimodal cloud samples. As a result, the concatenated feature contains the visual information, the MI and the related information among multimodal cloud samples. We conduct a series of experiments on the multimodal ground-based cloud database (MGCD), and the experimental results verify that the proposed HDF outperforms state-of-the-art methods.
Shuang Liu 0001, Linlin Duan, Zhong Zhang 0001, Xiaozhong Cao, Tariq S. Durrani
IEEE Trans. Geosci. Remote. Sens.3
2019 Compact Triplet Loss for person re-identification in camera sensor networks
Tongzhen Si, Zhong Zhang 0001, Shuang Liu 0001
Ad Hoc Networks2
2019 Hybrid Cross Deep Network for Domain Adaptation and Energy Saving in Visual Internet of Things
abstract
Recently, Visual Internet of Things (VIoT) has become a fast-growing field based on various applications. In this paper, we focus on two critical challenges for applications in VIoT, i.e., domain adaptation and energy saving. The images captured by various visual sensors in VIoT appear quite different due to changes in visual sensor locations, visual sensor settings, image resolutions, and illuminations. Meanwhile, VIoT generates a number of images, and transmitting original images would take up much bandwidth. In order to effectively classify such images and save energy, we propose a novel deep model named hybrid cross deep network (HCDN), which could learn domain-invariant and discriminative features for images in VIoT. The proposed HCDN is designed to contain the cross regularization loss and the classification loss. Moreover, it is also trained with images from different visual sensors. Specifically, the cross regularization loss selects the triplet samples from the source domain and the target domain, and adopts the calibration parameter to align the difference between two domains. We employ the vector extracted from the proposed HCDN to represent each image, which requires a smaller storage capacity than the original images. Energy consumption will be reduced when we transmit such vectors to the intelligent visual label system for image classification in VIoT. The proposed HCDN is verified on two domain adaptation datasets, and the experimental results prove its effectiveness.
Zhong Zhang 0001, Donghong Li
IEEE Internet Things J.1
2018 Discriminative Structural Metric Learning for Person Reidentification in Visual Internet of Things
abstract
Recently, visual Internet of Things (VIoT) has been deployed in many critical missions that are related to society security, such as anti-terrorism, abnormal event detection, crisis monitoring surveillance, etc. In this paper, we focus on a key fundamental problem in VIoT, person reidentification, which aims to correlate people with appropriate labels. The same person captured by various visual sensors in VIoT appears different significantly. To effectively measure the similarity between image pairs, we propose a novel method named discriminative structural metric learning (DSML), which utilizes intraregion metric, weak extra-region metric and extra-region metric to fully mine the structural information of pedestrian in a local way. According to DSML, we obtain a vector where each element is a local similarity score between two subregions. In order to aggregate the local similarity scores into a global one, we further propose a novel aggregating similarity score method named discriminative subregion aggregating (DSA). The DSA could learn the discriminative subregion by assigning different weights for each subregion. The experimental results demonstrate that the proposed method achieves better performance than the state-of-the-art methods.
Zhong Zhang 0001, Meiyan Huang
IEEE Internet Things J.1
2017 Learning completed discriminative local features for texture classification
Zhong Zhang 0001, Shuang Liu 0001, Xing Mei, Baihua Xiao
Pattern Recognit.1
2016 Multiple Continuous Virtual Paths Based Cross-View Action Recognition
abstract
In this paper, we propose a novel method for cross-view action recognition via multiple continuous virtual paths which connect the source view and the target view. Each point on one virtual path is a virtual view which is obtained by a linear transformation of an action descriptor. All the virtual views are concatenated into an infinite-dimensional feature to characterize continuous changes from the source to the target view. To utilize these infinite-dimensional features directly, we propose a virtual view kernel (VVK) to compute the similarity between two infinite-dimensional features, which can be readily used to construct any kernelized classifiers. In addition, a constraint term is introduced to fully utilize the information contained in the unlabeled samples which are easier to obtain from the target view. The rationality behind the constraint is that any action video belongs to only one class. To further explore complementary visual information, we utilize multiple continuous virtual paths. The original source and target views are projected to different auxiliary source and target views using the random projection technique. Then we fuse all the VVKs generated from all pairs of auxiliary views. Our method is verified on the IXMAS and MuHAVi datasets, and the experimental results demonstrate that our method achieves better performance than the state-of-the-art methods.
Zhong Zhang 0001, Shuang Liu 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002
Int. J. Pattern Recognit. Artif. Intell.1
2016 Information integration for ground-based cloud classification using joint consistent sparse coding in heterogeneous sensor network
Shuang Liu 0001, Zhong Zhang 0001, Xiaozhong Cao
Signal Process.2
2016 Coupled principal component analysis based face recognition in heterogeneous sensor networks
Zhong Zhang 0001, Shuang Liu 0001
Signal Process.1
2016 Cross domain boosting for information fusion in heterogeneous sensor-cyber sources
Zhong Zhang 0001, Shuang Liu 0001
Signal Process.1
2015 Scene text recognition by learning co-occurrence of strokes based on spatiality embedded dictionary
abstract
Text information contained in scene images is very helpful for high‐level image understanding. In this study, the authors propose to learn co‐occurrence of local strokes for scene text recognition by using a spatiality embedded dictionary (SED). Unlike spatial pyramid partitioning images into grids to incorporate spatial information, the authors SED associates every codeword with a particular response region and introduces more precise spatial information for robust character recognition. After localised soft coding and max pooling of the first layer, a sparse dictionary is learned to model co‐occurrence of several local strokes, which further improves classification performance. Experimental results on two scene character recognition datasets ICDAR2003 and CHARS74 K demonstrate that their character recognition method outperforms state‐of‐the‐art methods. Besides, competitive word recognition results are also reported for four benchmark word recognition datasets ICDAR2003, ICDAR2011, ICDAR2013 and street view text when combining their character recognition method with a conditional random field language model.
Song Gao 0009, Chunheng Wang, Baihua Xiao, Cunzhao Shi, Wen Zhou 0002, Zhong Zhang 0001
IET Comput. Vis.6
2015 Ground-Based Cloud Detection Using Automatic Graph Cut
abstract
Ground-based cloud detection plays an essential role in meteorological research, and object segmentation techniques have recently been introduced to solve this issue. As a kind of object segmentation technique, interactive graph cut has emerged as a very powerful tool due to its effective segmentation ability. However, it requires users to provide labels for certain pixels as “object” or “background,” which inevitably prohibits automatic cloud detection in large-scale applications. In this letter, we focus on the issue of automatic cloud detection and propose a novel algorithm named as automatic graph cut. We treat clouds as a special kind of object and eliminate human labeling by two procedures. First, we adaptively compute the thresholds for each cloud image which automatically label some pixels as “cloud” or “clear sky” with high confidence. Then, those labeled pixels serve as hard constraint seeds for the following graph cut algorithm. The experimental results show that the proposed algorithm not only achieves better results than the state-of-the-art cloud detection algorithms but also achieves comparable results with the interactive segmentation algorithm.
Shuang Liu 0001, Zhong Zhang 0001, Baihua Xiao, Xiaozhong Cao
IEEE Geosci. Remote. Sens. Lett.2
2015 Automatic Cloud Detection for All-Sky Images Using Superpixel Segmentation
abstract
Cloud detection plays an essential role in meteorological research and has received considerable attention in recent years. However, this issue is particularly challenging due to the diverse characteristics of clouds. In this letter, a novel algorithm based on superpixel segmentation (SPS) is proposed for cloud detection. In our proposed strategy, a series of superpixels could be obtained adaptively by SPS algorithm according to the characteristics of clouds. We first calculate a local threshold for each superpixel and then determine a threshold matrix for the whole image. Finally, cloud can be detected by comparing with the obtained threshold matrix. Experimental results show that our proposed algorithm achieves better performance than the current cloud detection algorithms.
Shuang Liu 0001, Zhong Zhang 0001, Chunheng Wang, Baihua Xiao
IEEE Geosci. Remote. Sens. Lett.3
2015 Robust relative attributes for human action recognition
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001
Pattern Anal. Appl.1
2014 Learning co-occurrence strokes for scene character recognition based on spatiality embedded dictionary
abstract
Robust scene-text-extraction system can be used in lots of areas. In this work, we propose to learn co-occurrence of local strokes for robust character recognition by using a spatiality embedded dictionary (SED). Different from spatial pyramid partitioning images into grids to incorporate spatial information, our SED associates every codeword with a particular response region and introduces more precise spatial information for character recognition. After localized soft coding and max pooling of the first layer, a sparse dictionary is learned to model co-occurrence of several local strokes, which further improves classification performance. Experiment on benchmark datasets demonstrates the effectiveness of our method and the results outperform state-of-the-art algorithms.
Song Gao 0009, Chunheng Wang, Baihua Xiao, Cunzhao Shi, Wen Zhou 0002, Zhong Zhang 0001
ICIP6
2014 Stroke Bank: A High-Level Representation for Scene Character Recognition
abstract
Text information contained in scene images is very useful for image understanding. In this paper, we propose a high-level representation named stroke bank for scene character recognition. Inspired by the work of object bank, we train stroke detectors and use detectors' maximal output as features. Specifically, we collect training samples for stroke detectors based on labeled key points. We also propose to restrict classification areas of each stroke detector to particular local regions, which alleviates computation burden and retains discrimination power at the same time. Experiments on benchmark datasets demonstrate the effectiveness of our method and the results outperform state-of-the-art algorithms.
Song Gao 0009, Chunheng Wang, Baihua Xiao, Cunzhao Shi, Zhong Zhang 0001
ICPR5
2014 Human action recognition using weighted pooling
abstract
Pooling strategies, such as max pooling and sum pooling, have been widely used to obtain the global representations for action videos. However, these pooling strategies have several disadvantages. First, they are easily affected by unwanted background local features, the absence of discriminative local features and the times of actions periodically performed by actors. Second, most pooling strategies only use local features to build the global representation that captures little mid‐level features for action representation. In this study, the authors propose a novel weighted pooling strategy based on actionlets representation for action recognition. The actionlets are defined as the movements of large bodies such as legs, arms and head, which capture rich mid‐level features for action representation. Besides, the authors’ method also incorporates the distribution information of actionlets into pooling procedure. Specifically, a pooling weight, which determines the importance of actionlet on the final video representation, is assigned to each actionlet. To learn the weight, they propose a novel discriminative learning algorithm to capture the discriminative information for pooling operation. They evaluate their weighted pooling on three datasets: KTH actions dataset, UCF sports dataset and Youtube actions dataset. Experimental results show the effectiveness of the proposed method.
Wen Zhou 0002, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001
IET Comput. Vis.4
2014 Action recognition via structured codebook construction
Wen Zhou 0002, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001
Signal Process. Image Commun.4
2014 SLD: A Novel Robust Descriptor for Image Matching
abstract
Image matching based on local features is a challenging task because it is difficult to build a robust local descriptor which is invariant to large variations in scale, viewpoints, illumination and rotation. To address these issues, Scale Invariant Feature Transform (SIFT) descriptor has been proposed to build a robust and distinctive local descriptor. However, it is not fully affine invariant. In this letter, we propose a novel robust descriptor: Sampling based Local Descriptor (SLD) to perform reliable image matching under large variations in scale, viewpoints, illumination and rotation. We build the descriptor based on elliptical sampling which samples image pixels according to the elliptic equations. The main advantage of elliptical sampling is that two controllable parameters of elliptical sampling can generate descriptors with different viewpoints and rotations. Besides, the descriptor has two notable properties: 1) it is fully invariant to affine changes; 2) it enables fast matching process because we only need to search two controllable parameters for elliptical sampling, which is more efficient than other affine invariant descriptors. We test the proposed descriptor on standard benchmark for evaluation. Experimental results show the robustness of the proposed method under large variations in illumination, viewpoints and scale.
Wen Zhou 0002, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001
IEEE Signal Process. Lett.4
2014 Cross-View Action Recognition Using Contextual Maximum Margin Clustering
abstract
Recently, maximum margin clustering (MMC) has been proposed for a cross-view action recognition. However, such a method neglects the temporal relationship between contiguous frames in the same action video. In this paper we propose a novel method called contextual maximum margin clustering (CMMC) to tackle cross-view action recognition. In CMMC, we add temporal regularization to give a high penalty when the contiguous frames are dissimilar. Thus, the CMMC not only achieves the goal of finding maximum margin hyperplanes, but also explicitly considers the temporal information among contiguous frames. Our method is verified on the IXMAS dataset and the experimental results demonstrate that our method can achieve better performance than the state-of-the-art methods.
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2013 Scene Text Recognition Using Part-Based Tree-Structured Character Detection
abstract
Scene text recognition has inspired great interests from the computer vision community in recent years. In this paper, we propose a novel scene text recognition method using part-based tree-structured character detection. Different from conventional multi-scale sliding window character detection strategy, which does not make use of the character-specific structure information, we use part-based tree-structure to model each type of character so as to detect and recognize the characters at the same time. While for word recognition, we build a Conditional Random Field model on the potential character locations to incorporate the detection scores, spatial constraints and linguistic knowledge into one framework. The final word recognition result is obtained by minimizing the cost function defined on the random field. Experimental results on a range of challenging public datasets (ICDAR 2003, ICDAR 2011, SVT) demonstrate that the proposed method outperforms state-of-the-art methods significantly both for character detection and word recognition.
Cunzhao Shi, Chunheng Wang, Baihua Xiao, Song Gao 0009, Zhong Zhang 0001
CVPR6
2013 Cross-View Action Recognition via a Continuous Virtual Path
abstract
In this paper, we propose a novel method for cross-view action recognition via a continuous virtual path which connects the source view and the target view. Each point on this virtual path is a virtual view which is obtained by a linear transformation of the action descriptor. All the virtual views are concatenated into an infinite-dimensional feature to characterize continuous changes from the source to the target view. However, these infinite-dimensional features cannot be used directly. Thus, we propose a virtual view kernel to compute the value of similarity between two infinite-dimensional features, which can be readily used to construct any kernelized classifiers. In addition, there are a lot of unlabeled samples from the target view, which can be utilized to improve the performance of classifiers. Thus, we present a constraint strategy to explore the information contained in the unlabeled samples. The rationality behind the constraint is that any action video belongs to only one class. Our method is verified on the IXMAS dataset, and the experimental results demonstrate that our method achieves better performance than the state-of-the-art methods.
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001, Cunzhao Shi
CVPR1
2013 Tensor Ensemble of Ground-Based Cloud Sequences: Its Modeling, Classification, and Synthesis
abstract
Since clouds are one of the most important meteorological phenomena related to the hydrological cycle and affect Earth radiation balance and climate changes, cloud analysis is a crucial issue in meteorological research. Most researchers only consider the classification task of cloud images while less attention has been paid to the synthesis one. In addition, all the existing research on cloud identification from sky images is based on single cloud images. However, the cloud-measuring devices on the ground actually take one image of the clouds every few minutes and collect a series of cloud images. Thus, the existing methods neglect the temporal information exhibited by contiguous cloud images. To overcome this drawback, in this letter we treat ground-based cloud sequences (GCSs) as dynamic texture. We then propose the Tensor Ensemble of Ground-based Cloud Sequences (eTGCS) model which represents the ensemble of GCSs in a tensor manner. In the eTGCS model, all GCSs form a single tensor, and each GCS is a subtensor of the single tensor. There are two main characteristics of the eTGCS model: 1) All GCSs share an identical mode subspace, which makes the classification convenient, and 2) a new GCS can be synthesized as long as the parameters of the eTGCS model are used. Therefore, less storage space is required. Comprehensive experiments are conducted to prove the superiority of our eTGCS model. The classification accuracy achieves 92.31%, and the synthesized GCSs are similar to the original ones in visual appearance.
Shuang Liu 0001, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001, Xiaozhong Cao
IEEE Geosci. Remote. Sens. Lett.4
2013 Attribute Regularization Based Human Action Recognition
abstract
Recently, attributes have been introduced as a kind of high-level semantic information to help improve the classification accuracy. Multitask learning is an effective methodology to achieve this goal, which shares low-level features between attributes and actions. Yet such methods neglect the constraints that attributes impose on classes, which may fail to constrain the semantic relationship between the attributes and actions. In this paper, we explicitly consider such attribute-action relationship for human action recognition, and correspondingly, we modify the multitask learning model by adding attribute regularization. In this way, the learned model not only shares the low-level features, but also gets regularized according to the semantic constrains. In addition, since attribute and class label contain different amounts of semantic information, we separately treat attribute classifiers and action classifiers in the framework of multitask learning for further performance improvement. Our method is verified on three challenging datasets (KTH, UIUC, and Olympic Sports), and the experimental results demonstrate that our method achieves better results than that of previous methods on human action recognition.
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001
IEEE Trans. Inf. Forensics Secur.1
2012 Multi-scale Fusion of Texture and Color for Background Modeling
abstract
Background modeling from a stationary camera is a crucial component in video surveillance. Traditional methods usually adopt single feature type to solve the problem, while the performance is usually unsatisfactory when handling complex scenes. In this paper, we propose a multi-scale strategy, which combines both texture and color features, to achieve a robust and accurate solution. Our contributions are two folds: one is that we propose a novel textureoperator named Scale-invariant Center-symmetric Local Ternary Pattern, which is robust to noise and illumination variations, the other is that a multi-scale fusion strategy is proposed for the issue. Our method is verified on several complex real world videoswith illumination variation, soft shadows and dynamic backgrounds. We compare our method with four state-of-the-art methods, and the experimental results clearly demonstrate that our method achievesthe highest classification accuracy in complex real world videos.
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Shuang Liu 0001, Wen Zhou 0002
AVSS1
2012 Human Action Recognition with Attribute Regularization
abstract
Recently, attributes have been introduced to help object classification. Multi-task learning is an effective methodology to achieve this goal, which shares low-level features between attribute and object classifiers. Yet such a method neglects the constraints that attributes impose on classes which may fail to constrain the semantic relationship between the attribute and object classifiers. In this paper, we explicitly consider such attribute-object relationship, and correspondingly, we modify the multi-task learningmodel by adding attribute regularization. In this way, the learned model not only shares the low-level features, but also gets regularized according to the semantic constrains. Our method is verified on two challenging datasets (KTH and Olympic Sports), andthe experimental results demonstrate that our method achieves better results than previous methods in human action recognition.
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001
AVSS1
2012 Soft-signed sparse coding for ground-based cloud classification
Shuang Liu 0001, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001, Yunxue Shao
ICPR4
2012 Contextual Fisher kernels for human action recognition
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001
ICPR1
2012 Learning weighted features for human action recognition
Wen Zhou 0002, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001
ICPR4
2012 Human action recognition by bagging data dependent representation
Wen Zhou 0002, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001
ICPR4
2012 Action Recognition Using Context-Constrained Linear Coding
abstract
Although traditional bag-of-words model has shown promising results for action recognition, it takes no consideration of the relationship among spatio–temporal points; furthermore, it also suffers serious quantization error. In this letter, we propose a novel coding strategy called context-constrained linear coding (CLC) to overcome these limitations. We first calculate the contextual distance between local descriptors and each codeword by considering the spatio–temporal contextual information. Then, linear coding using contextual distance is adopted to alleviate the quantization error. Our method is verified on two challenging databases (KTH and UCF sports), and the experimental results demonstrate that our method achieves better results than previous methods in action recognition.
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001
IEEE Signal Process. Lett.1