Shuang Liu 0001

dblp:58/6609-1 · DBLP profile ↗
← Back
44ranked-venue papers
14as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 first-author · 5 since 2021Computer networks · 8 · 4 first-author · 2 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Text-prompt multi-scale integration and detail-aware enhancement network for homography estimation
abstract
As a fundamental task in computer vision, homography estimation plays a critical role in many applications, such as image stitching, augmented reality, and multi-view geometry. However, existing methods primarily rely on low-level image visual features to estimate homography, which may produce inevitable alignment distortions. Motivated by the remarkable success of vision-language models in various computer vision tasks, we propose a text-prompt multi-scale integration and detail-aware enhancement network (TMIDE-Net) for homography estimation, where the multi-modal information, i.e. , visual and text information, is introduced to compensate for the insufficient geometric cues in low-texture regions. Specifically, with the help of auxiliary semantic language knowledge extracted by a frozen pretrained CLIP, a text-prompt multi-scale feature integration module is designed to extract and fuse image features and text features. Then, we design a coarse-to-fine homography estimation module to improve homography alignment accuracy in low-texture and illumination-variant regions, where a detail-aware enhancement block is presented to enhance the fine-grained texture representation capability. Finally, a multi-constraint hybrid loss is applied to obtain robust homography estimation in complex scenes. Extensive experiments indicate that the proposed TMIDE-Net outperforms the state-of-the-arts both quantitatively and qualitatively, reducing the average error by approximately 13.5%.
Xiaoting Fan, Ronglu Wang, Shuang Liu 0001
J. Vis. Commun. Image Represent.3
2025 Unsupervised Person Reidentification Using Stripe-Driven Fusion Transformer Network
abstract
In recent years, some methods utilize a transformer as the backbone to model the long‐range context dependencies, reflecting a prevailing trend in unsupervised person reidentification (Re‐ID) tasks. However, they only explore the global information through interactive learning in the framework of the transformer, which ignores the learning of the part information in the interaction process for pedestrian images. In this study, we present a novel transformer network for unsupervised person Re‐ID, a stripe‐driven fusion transformer (SDFT), designed to simultaneously capture the global interaction and the part interaction when modeling the long‐range context dependencies. Meanwhile, we present a stripe‐driven regularization (SDR) to constrain the part aggregation features and the global features by considering the consistency principle from the aspects of the features and the clusters, aiming to improve the representational capacity of the features. Furthermore, to investigate the relationships between local regions of pedestrian images, we present a stripe‐driven contrastive loss (SDCL) to learn discriminative part features from the perspectives of pedestrian identity and stripes. The proposed method has undergone extensive validations on publicly available unsupervised person Re‐ID benchmarks, and the experimental results confirm its superiority and effectiveness.
Zeyu Zang, Shuang Liu 0001, Zhong Zhang 0001, Xinshan Zhu
IET Softw.3
2025 SDLFusion: A salient-aware differentiated learning network for infrared and visible image fusion
Xiaoting Fan, Shuang Liu 0001, Baihua Xiao
Knowl. Based Syst.3
2025 Cross-domain person re-identification via learning Heterogeneous Pseudo Labels
Zhong Zhang 0001, Di He 0008, Shuang Liu 0001
Pattern Recognit.3
2025 Cloud-Type Classification Using Multimodal Integration Transformer Based on Cloud Images and Millimeter-Wave Cloud Radar Observations
abstract
The existing methods fail to simultaneously utilize the appearance information and the internal structure of clouds for cloud type classification, resulting in incomplete cloud representation. In this paper, we exploit cloud images and Millimeter-wave Cloud Radar (MMCR) observations for cloud type classification, and propose a novel Transformer network named Multi-modal Integration Transformer (MMITrans) to describe completed information of clouds. To this end, we design MMITrans as three subnetworks, i.e., vision subnetwork, MMCR subnetwork and multi-modal fusion network. Specifically, we extract the visual features from the cloud images through the vision subnetwork. Meanwhile, we first convert MMCR observations into several cloud-related indicators, and propose the Indicator-Tokenization to effectively tokenize them and obtain the indicator features using the MMCR subnetwork. Furthermore, we propose the Multi-modal Cross Attention in the multi-modal fusion network to sufficiently fuse the visual features and the indicator features in a multiple-input way. We perform a series of experiments on Cloud images and Millimeter-wave cloud radar observations Dataset, i.e., CMD-Beijing and CMD-Gansu, and the experimental results demonstrate the superiority of the proposed MMITrans.
Shuang Liu 0001, Zeyu Zang, Zhong Zhang 0001, Shuzhen Hu, Baihua Xiao
IEEE Trans. Geosci. Remote. Sens.1
2025 Completed Interaction Networks for Pedestrian Trajectory Prediction
abstract
The social and environmental interactions, as well as the pedestrian goal are crucial for pedestrian trajectory prediction. This is because they could learn both complex interactions in the scenes and the intentions of the pedestrians. However, most existing methods either learn the one-moment social interactions, or supervise the pedestrian trajectories using long-term goal, resulting in suboptimal prediction performances. In this paper, we propose a novel network named Completed Interaction Network (CINet) to simultaneously consider the social interactions in all moments, the environmental interactions and the short-term goal of pedestrians in a unified framework for pedestrian trajectory prediction. Specifically, we propose the Spatio-Temporal Transformer Layer (STTL) to fully mine the spatio-temporal information among historical trajectories of all pedestrians in order to obtain the social interactions in all moments. Additionally, we present the Gradual Goal Module (GGM) to capture the environmental interactions under the supervision of the short-term goal, which is beneficial to understanding the intentions of the pedestrian. Afterwards, we employ the cross-attention to effectively integrate the all-moment social and environmental interactions. The experimental results on three standard pedestrian datasets, i.e., ETH/UCY, SDD and inD demonstrate that our method achieves a new state-of-the-art performance. Furthermore, the visualization results indicate that our method could predict trajectories more reasonably in complex scenarios such as sharp turns, infeasible areas and so on.
Zhong Zhang 0001, Jianglin Zhou, Shuang Liu 0001, Baihua Xiao
IEEE Trans. Multim.3
2024 Completed Part Transformer for Person Re-Identification
abstract
Recently, part information of pedestrian images has been demonstrated to be effective for person re-identification (ReID), but the part interaction is ignored when using Transformer to learn long-range dependencies. In this article, we propose a novel transformer network named Completed Part Transformer (CPT) for person ReID, where we design the part transformer layer to learn the completed part interaction. The part transformer layer includes the intra-part layer and the part-global layer, where they consider long-range dependencies from the aspects of the intra-part interaction and the part-global interaction, simultaneously. Furthermore, in order to overcome the limitation of fixed number of the patch tokens in the transformer layer, we propose the Adaptive Refined Tokens (ART) module to focus on learning the interaction between the informative patch tokens in the pedestrian image, which improves the discrimination of the pedestrian representation. Extensive experimental results on four person ReID datasets, i.e., MSMT17, Market1501, DukeMTMC-reID, and CUHK03, demonstrate that the proposed method achieves a new state-of-the-art performance, e.g., it achieves 68.0% mAP and 84.6% Rank-1 accuracy on MSMT17.
Zhong Zhang 0001, Di He 0008, Shuang Liu 0001, Baihua Xiao, Tariq S. Durrani
IEEE Trans. Multim.3
2023 Integration graph attention network and multi-centre constrained loss for cross-modality person re-identification
abstract
Abstract Cross‐modality person re‐identification is a challenging task due to the large visual appearance difference between RGB and infrared images. Existing studies mainly focus on learning local features and ignore the correlation between local features. In this paper, the Integration Graph Attention Network is proposed to learn the completed correlation between local features via the graph structure. To this end, the authors learn the coarse‐fine attention weights to aggregate the local features by considering local detail and global information. Furthermore, the Multi‐Centre Constrained Loss is proposed to optimise the feature similarity by constraining the centres of modality and identity. It simultaneously utilises three kinds of centre constraints, that is intra‐identity centre constraint, modality centre constraint, and inter‐identity centre constraint, in order to reduce the influence of modality information explicitly. The proposed method is evaluated on two standard benchmark datasets, that is SYSU‐MM01 and RegDB, and the results demonstrate that the authors’ method achieves better performance than the state‐of‐the‐art methods, for example, surpassing NFS by 4.8% and 6.0% mAP on the single‐shot setting in All‐search and Indoor‐search modes, respectively.
Di He 0008, Jingrui Zhang, Zhong Zhang 0001, Shuang Liu 0001, Tariq S. Durrani
IET Comput. Vis.4
2023 Cross-modality person re-identification using hybrid mutual learning
abstract
Abstract Cross‐modality person re‐identification (Re‐ID) aims to retrieve a query identity from red, green, blue (RGB) images or infrared (IR) images. Many approaches have been proposed to reduce the distribution gap between RGB modality and IR modality. However, they ignore the valuable collaborative relationship between RGB modality and IR modality. Hybrid Mutual Learning (HML) for cross‐modality person Re‐ID is proposed, which builds the collaborative relationship by using mutual learning from the aspects of local features and triplet relation. Specifically, HML contains local‐mean mutual learning and triplet mutual learning where they focus on transferring local representational knowledge and structural geometry knowledge so as to reduce the gap between RGB modality and IR modality. Furthermore, Hierarchical Attention Aggregation is proposed to fuse local feature maps and local feature vectors to enrich the information of the classifier input. Extensive experiments on two commonly used data sets, that is, SYSU‐MM01 and RegDB verify the effectiveness of the proposed method.
Zhong Zhang 0001, Sen Wang 0007, Shuang Liu 0001, Baihua Xiao, Tariq S. Durrani
IET Comput. Vis.4
2023 Integration Transformer for Ground-Based Cloud Image Segmentation
abstract
Recently, convolutional neural network (CNN) dominates the ground-based cloud image segmentation task, but disregards the learning of long-range dependencies due to the limited size of filters. Although Transformer-based methods could overcome this limitation, they only learn long-range dependencies at a single scale, hence failing to capture multi-scale information of cloud image. The multi-scale information is beneficial to ground-based cloud image segmentation, because the features from small scales tend to extract detailed information while features from large scales have the ability to learn global information. In this paper, we propose a novel deep network named Integration Transformer (InTransformer), which builds long-range dependencies from different scales. To this end, we propose the Hybrid Multi-head Transformer Block (HMTB) to learn multi-scale long-range dependencies, and hybridize CNN and HMTB as the encoder at different scales. The proposed InTransformer hybridizes CNN and Transformer as the encoder to extract multi-scale representations, which learns both local information and long-range dependencies with different scales. Meanwhile, in order to fuse the patch tokens with different scales, we propose Mutual Cross-Attention Module (MCAM) for the decoder of InTransformer which could adequately interact multi-scale patch tokens in a bidirectional way. We have conducted a series of experiments on large ground-based cloud detection database TLCDD and SWIMSEG. The experimental results show that the performance of our method outperforms other methods, proving the effectiveness of the proposed InTransformer.
Shuang Liu 0001, Zhong Zhang 0001, Xiaozhong Cao, Tariq S. Durrani
IEEE Trans. Geosci. Remote. Sens.1
2022 Ground-Based Cloud Detection Using Multiscale Attention Convolutional Neural Network
abstract
Cloud detection plays a significant role in ground-based remote sensing observation, and it is quite challenging due to the variations in illumination and cloud form, and the vague boundaries between cloud and sky. In this letter, we propose a novel deep model named multiscale attention convolutional neural network (MACNN) for ground-based cloud detection, which possesses a symmetric encoder–decoder structure. For accurate cloud detection, we design the multiscale module in MACNN to obtain different receptive fields by using different hole rates for the filters, and meanwhile, we propose the attention module in MACNN to learn the attention coefficients in order to reflect different importance of pixels. Furthermore, we release the Tianjin Normal University (TJNU) cloud detection database (TCDD) to provide a comparative study for different methods, and to the best of our knowledge, it is the largest cloud detection database. We conduct a series of experiments on the TCDD, and the experimental results demonstrate that the proposed MACNN outperforms state-of-the-art methods in five quantitative evaluation criteria.
Zhong Zhang 0001, Shuzhen Yang, Shuang Liu 0001, Baihua Xiao, Xiaozhong Cao
IEEE Geosci. Remote. Sens. Lett.3
2022 Cross-Domain Person Re-Identification Using Heterogeneous Convolutional Network
abstract
Person re-identification (Re-ID) is a challenging task due to variations in pedestrian images, especially in cross-domain scenarios. The existing cross-domain person Re-ID approaches extract the feature from single pedestrian image, but they ignore the correlations among pedestrian images. In this paper, we propose Heterogeneous Convolutional Network (HCN) for cross-domain person Re-ID, which learns the appearance information of pedestrian images and the correlations among pedestrian images simultaneously. To this end, we first utilize Convolutional Neural Network (CNN) to extract the appearance features for pedestrian images. Then we construct a graph in the target dataset where the appearance features are treated as the nodes and the similarity represents the linkage between the nodes. Afterwards, we propose Dual Graph Convolution (DGConv) to explicitly learn the correlation information from the similar and dissimilar samples, which could avoid the over-smoothing caused by the fully connected graph. Furthermore, we design HCN as a multi-branch structure to mine the structural information of pedestrians. We conduct extensive evaluations for HCN on three datasets, i.e. Market-1501, DukeMTMC-reID and MSMT17, and the results demonstrate that HCN is superior to the state-of-the-art methods.
Zhong Zhang 0001, Shuang Liu 0001, Baihua Xiao, Tariq S. Durrani
IEEE Trans. Circuits Syst. Video Technol.3
2022 Ground-Based Remote Sensing Cloud Classification via Context Graph Attention Network
abstract
Most ground-based remote sensing cloud classification methods focus on learning representation features for cloud images while ignoring the correlations among cloud images. Recently, graph convolutional network (GCN) is applied to provide the correlations for ground-based remote sensing cloud classification, in which the graph convolutional layer aggregates information from the connected nodes of graph in a weighted way. However, the weights assigned by GCN cannot reflect the importance of connected nodes precisely, which declines the discrimination of the aggregated features (AFs). To overcome the limitation, in this article, we propose the context graph attention network (CGAT) for ground-based remote sensing cloud classification. Specifically, the context graph attention layer (CGA layer) of CGAT is proposed to learn the context attention coefficients (CACs) and obtain the AFs of nodes based on the CACs. We compute the CACs not only considering the two connected nodes but also their neighborhood nodes in order to stabilize the aggregation process. In addition, we propose to utilize two different transformation matrices to transform the node and its connected nodes into new feature spaces, which could enhance the discrimination of AFs. We concatenate the AFs with the deep features (DFs) as final representations for cloud classification. Since existing ground-based cloud data sets (GCDs) have limited cloud images, we release a new data set named GCD that is the largest one for ground-based cloud classification. We conduct a series of experiments on GCD, and the experimental results verify the effectiveness of CGAT.
Shuang Liu 0001, Linlin Duan, Zhong Zhang 0001, Xiaozhong Cao, Tariq S. Durrani
IEEE Trans. Geosci. Remote. Sens.1
2022 Ground-Based Remote Sensing Cloud Detection Using Dual Pyramid Network and Encoder-Decoder Constraint
abstract
Many methods for ground-based remote sensing cloud detection learn representation features using the encoder–decoder structure. However, they only consider the information from single scale, which leads to incomplete feature extraction. In this article, we propose a novel deep network named dual pyramid network (DPNet) for ground-based remote sensing cloud detection, which possesses an encoder–decoder structure with dual pyramid pooling module (DPPM). Specifically, we process the feature maps of different scales in the encoder through dual pyramid pooling. Then, we fuse the outputs of the dual pyramid pooling in the same pyramid level using the attention fusion. Furthermore, we propose the encoder–decoder constraint (EDC) to relieve information loss in the process of encoding and decoding. It constrains the values and the gradients of probability maps from the encoder and the decoder to be consistent. Since the number of cloud images in the publicly available databases for ground-based remote sensing cloud detection is limited, we release the TJNU Large-scale Cloud Detection Database (TLCDD) that is the largest database in this field. We conduct a series of experiments on TLCDD, and the experimental results verify the effectiveness of the proposed method.
Zhong Zhang 0001, Shuzhen Yang, Shuang Liu 0001, Xiaozhong Cao, Tariq S. Durrani
IEEE Trans. Geosci. Remote. Sens.3
2021 Learning Hybrid Relationships for Person Re-identification
abstract
Recently, the relationship among individual pedestrian images and the relationship among pairwise pedestrian images have become attractive for person re-identification (re-ID) as they effectively improve the ability of feature representation. In this paper, we propose a novel method named Hybrid Relationship Network (HRNet) to learn the two types of relationships in a unified framework that makes use of their own advantages. Specifically, for the relationship among individual pedestrian images, we take the features of pedestrian images as the nodes to construct a locally-connected graph, so as to improve the discriminative ability of nodes. Meanwhile, we propose the consistent node constraint to inject the identity information into the graph learning process and guide the information to propagate accurately. As for the relationship among pairwise pedestrian images, we treat the feature differences of pedestrian images as the nodes to construct a fully-connected graph so as to estimate robust similarity of nodes. Furthermore, we propose the inter-graph propagation to alleviate the information loss for the fully-connected graph. Extensive experiments on Market-1501, DukeMTMCreID, CUHK03 and MSMT17 demonstrate that the proposed HRNet outperforms the state-of-the-art methods.
Shuang Liu 0001, Wenmin Huang, Zhong Zhang 0001
AAAI1
2021 Person Re-Identification Using Heterogeneous Local Graph Attention Networks
abstract
Recently, some methods have focused on learning local relation among parts of pedestrian images for person re-identification (Re-ID), as it offers powerful representation capabilities. However, they only provide the intra-local relation among parts within single pedestrian image and ignore the inter-local relation among parts from different images, which results in incomplete local relation information. In this paper, we propose a novel deep graph model named Heterogeneous Local Graph Attention Networks (HLGAT) to model the inter-local relation and the intra-local relation in the completed local graph, simultaneously. Specifically, we first construct the completed local graph using local features, and we resort to the attention mechanism to aggregate the local features in the learning process of inter-local relation and intra-local relation so as to emphasize the importance of different local features. As for the inter-local relation, we propose the attention regularization loss to constrain the attention weights based on the identities of local features in order to describe the inter-local relation accurately. As for the intra-local relation, we propose to inject the contextual information into the attention weights to consider structure information. Extensive experiments on Market-1501, CUHK03, DukeMTMC-reID and MSMT17 demonstrate that the proposed HLGAT outperforms the state-of-the-art methods.
Zhong Zhang 0001, Haijia Zhang, Shuang Liu 0001
CVPR3
2021 Dynamically occluded samples via adversarial learning for person re-identification in sensor networks
Wenmin Huang, Shuang Liu 0001, Ruiling Luo, Tongzhen Si, Zhong Zhang 0001
Ad Hoc Networks2
2021 Local Alignment Deep Network for Infrared-Visible Cross-Modal Person Reidentification in 6G-Enabled Internet of Things
abstract
In this article, we propose a novel deep framework termed local alignment deep network (LADN) for infrared-visible cross-modal person reidentification (IVCM ReID) in 6G-enabled IoT, which could meet the demands of all-day and real-time surveillance. The proposed LADN is designed as a two-stream structure, and it learns shallow interested feature maps and common subspace feature maps to reduce the gap between IR and RGB images. To overcome the challenge of pose and viewpoint variations of pedestrians, we learn the local features in the deep layers. We also propose the local alignment triplet loss (LAT) to align local features, which could capture the consistent local information via comparing noncorresponding local features in a certain range. Furthermore, we learn the global features to provide the global field of vision for the representation. The proposed LADN is optimized in an end-to-end way by combining different cross-modality losses. We evaluate the proposed method on two standard benchmark data sets, i.e., SYSU-MM01 and RegDB, and the results demonstrate the effectiveness of LADN.
Shuang Liu 0001, Jingrui Zhang
IEEE Internet Things J.1
2021 Part-guided graph convolution networks for person re-identification
Zhong Zhang 0001, Haijia Zhang, Shuang Liu 0001, Tariq S. Durrani
Pattern Recognit.3
2020 Deep tensor fusion network for multimodal ground-based cloud classification in weather station networks
Shuang Liu 0001, Zhong Zhang 0001
Ad Hoc Networks2
2020 Person re-identification using Hybrid Task Convolutional Neural Network in camera sensor networks
Shuang Liu 0001, Wenmin Huang, Zhong Zhang 0001
Ad Hoc Networks1
2020 Cross-domain person re-identification using Dual Generation Learning in camera sensor networks
Zhong Zhang 0001, Shuang Liu 0001
Ad Hoc Networks3
2020 Fuzzy Multilayer Clustering and Fuzzy Label Regularization for Unsupervised Person Reidentification
abstract
Unsupervised person reidentification has received more attention due to its wide real-world applications. In this paper, we propose a novel method named fuzzy multilayer clustering (FMC) for unsupervised person reidentification. The proposed FMC learns a new feature space using a multilayer perceptron for clustering in order to overcome the influence of complex pedestrian images. Meanwhile, the proposed FMC generates fuzzy labels for unlabeled pedestrian images, which simultaneously considers the membership degree and the similarity between the sample and each cluster. We further propose the fuzzy label regularization (FLR) to train the convolutional neural network (CNN) using pedestrian images with fuzzy labels in a supervised manner. The proposed FLR could regularize the CNN training process and reduce the risk of overfitting. The effectiveness of our method is validated on three large-scale person reidentification databases, i.e., Market-1501, DukeMTMC-reID, and CUHK03.
Zhong Zhang 0001, Meiyan Huang, Shuang Liu 0001, Baihua Xiao, Tariq S. Durrani
IEEE Trans. Fuzzy Syst.3
2020 Multimodal Ground-Based Remote Sensing Cloud Classification via Learning Heterogeneous Deep Features
abstract
Recently, multimodal cloud samples are utilized to learn completed feature representations for cloud classification. However, the existing methods neglect the related information from other multimodal cloud samples in the learning process, which leads to inadequate learning. In this article, we propose a novel deep model to learn heterogeneous deep features (HDFs) for multimodal ground-based remote sensing cloud classification. Specifically, we first design the convolutional neural network (CNN) extractor to combine the visual information and the multimodal information (MI) to obtain the CNN-based features of multimodal cloud samples. Afterward, we treat the CNN-based features of multimodal cloud samples as the nodes of graph, and utilize the similarity between nodes as the adjacency matrix. We feed the graph and the adjacency matrix into the graph convolutional network (GCN) extractor to obtain the GCN-based features that could capture correlations among multimodal cloud samples using graph convolutional layers. After obtaining CNN-based features and GCN-based features, we concatenate the two kinds of heterogeneous features to represent the multimodal cloud samples. As a result, the concatenated feature contains the visual information, the MI and the related information among multimodal cloud samples. We conduct a series of experiments on the multimodal ground-based cloud database (MGCD), and the experimental results verify that the proposed HDF outperforms state-of-the-art methods.
Shuang Liu 0001, Linlin Duan, Zhong Zhang 0001, Xiaozhong Cao, Tariq S. Durrani
IEEE Trans. Geosci. Remote. Sens.1
2019 Compact Triplet Loss for person re-identification in camera sensor networks
Tongzhen Si, Zhong Zhang 0001, Shuang Liu 0001
Ad Hoc Networks3
2019 Multimodal GAN for Energy Efficiency and Cloud Classification in Internet of Things
abstract
Efficient processing of large-scale multimodal sensor data is a key issue for applying the Internet of Things (IoT). Accurate cloud classification is critical for weather and climate monitoring, which are parts of IoT applications. In this paper, we propose a novel generative deep model named multimodal generative adversarial network (Multimodal GAN) to improve both the energy efficiency and the cloud classification accuracy in IoT. The proposed Multimodal GAN is composed of a discriminator and a generator, each of which is devised to a two-stream network. The branches of two-stream structure correspond to the cloud visual information and the cloud scalar information, respectively. Therefore, the Multimodal GAN is capable of generating the cloud visual information and cloud scalar information simultaneously. Afterward, the training set is extended by the generated multimodal cloud samples, and the deep multimodal cloud classification model is trained by the extended training set. As a result, the classification model possesses high generalization ability and is less prone to be over-fitting. Moreover, the feature representations extracted from the classification model reflect the salient information of raw multimodal cloud data, and therefore they can be stored and transmitted in IoT. The effectiveness of the proposed method in energy efficiency and cloud classification is validated on the multimodal cloud dataset.
Shuang Liu 0001
IEEE Internet Things J.1
2018 Class-Constrained Transfer LDA for Cross-View Action Recognition in Internet of Things
abstract
Internet of Things (IoT) is a fast-growing field based on different techniques and applications. In this paper, we focus on IoT in monitoring mission which is widely used in everyday life, such as security surveillance, health-care, independent living, etc. In order to overcome the challenges of various viewpoints and heterogenous sensors in IoT, we propose a novel method named class-constrained transfer linear discriminant analysis (CTLDA), which learns two projection matrices with different dimensionality for mapping the original features into a common subspace. The target of learning two projection matrices is to maximize the interclass data and minimize the intraclass data. Meanwhile, we propose the class-constrained regularization which enforces the neighboring samples with the same class to be still close to each other so as to further improve the discrimination of CTLDA. The class-constrained regularization possesses two strategies, i.e., hard constraint and soft constraint. The experimental results demonstrate that our method achieves better performance than the state-of-the-art methods.
Shuang Liu 0001
IEEE Internet Things J.1
2017 Learning completed discriminative local features for texture classification
Zhong Zhang 0001, Shuang Liu 0001, Xing Mei, Baihua Xiao
Pattern Recognit.2
2016 Multiple Continuous Virtual Paths Based Cross-View Action Recognition
abstract
In this paper, we propose a novel method for cross-view action recognition via multiple continuous virtual paths which connect the source view and the target view. Each point on one virtual path is a virtual view which is obtained by a linear transformation of an action descriptor. All the virtual views are concatenated into an infinite-dimensional feature to characterize continuous changes from the source to the target view. To utilize these infinite-dimensional features directly, we propose a virtual view kernel (VVK) to compute the similarity between two infinite-dimensional features, which can be readily used to construct any kernelized classifiers. In addition, a constraint term is introduced to fully utilize the information contained in the unlabeled samples which are easier to obtain from the target view. The rationality behind the constraint is that any action video belongs to only one class. To further explore complementary visual information, we utilize multiple continuous virtual paths. The original source and target views are projected to different auxiliary source and target views using the random projection technique. Then we fuse all the VVKs generated from all pairs of auxiliary views. Our method is verified on the IXMAS and MuHAVi datasets, and the experimental results demonstrate that our method achieves better performance than the state-of-the-art methods.
Zhong Zhang 0001, Shuang Liu 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002
Int. J. Pattern Recognit. Artif. Intell.2
2016 Information integration for ground-based cloud classification using joint consistent sparse coding in heterogeneous sensor network
Shuang Liu 0001, Zhong Zhang 0001, Xiaozhong Cao
Signal Process.1
2016 Coupled principal component analysis based face recognition in heterogeneous sensor networks
Zhong Zhang 0001, Shuang Liu 0001
Signal Process.2
2016 Cross domain boosting for information fusion in heterogeneous sensor-cyber sources
Zhong Zhang 0001, Shuang Liu 0001
Signal Process.3
2015 Ground-Based Cloud Detection Using Automatic Graph Cut
abstract
Ground-based cloud detection plays an essential role in meteorological research, and object segmentation techniques have recently been introduced to solve this issue. As a kind of object segmentation technique, interactive graph cut has emerged as a very powerful tool due to its effective segmentation ability. However, it requires users to provide labels for certain pixels as “object” or “background,” which inevitably prohibits automatic cloud detection in large-scale applications. In this letter, we focus on the issue of automatic cloud detection and propose a novel algorithm named as automatic graph cut. We treat clouds as a special kind of object and eliminate human labeling by two procedures. First, we adaptively compute the thresholds for each cloud image which automatically label some pixels as “cloud” or “clear sky” with high confidence. Then, those labeled pixels serve as hard constraint seeds for the following graph cut algorithm. The experimental results show that the proposed algorithm not only achieves better results than the state-of-the-art cloud detection algorithms but also achieves comparable results with the interactive segmentation algorithm.
Shuang Liu 0001, Zhong Zhang 0001, Baihua Xiao, Xiaozhong Cao
IEEE Geosci. Remote. Sens. Lett.1
2015 Automatic Cloud Detection for All-Sky Images Using Superpixel Segmentation
abstract
Cloud detection plays an essential role in meteorological research and has received considerable attention in recent years. However, this issue is particularly challenging due to the diverse characteristics of clouds. In this letter, a novel algorithm based on superpixel segmentation (SPS) is proposed for cloud detection. In our proposed strategy, a series of superpixels could be obtained adaptively by SPS algorithm according to the characteristics of clouds. We first calculate a local threshold for each superpixel and then determine a threshold matrix for the whole image. Finally, cloud can be detected by comparing with the obtained threshold matrix. Experimental results show that our proposed algorithm achieves better performance than the current cloud detection algorithms.
Shuang Liu 0001, Zhong Zhang 0001, Chunheng Wang, Baihua Xiao
IEEE Geosci. Remote. Sens. Lett.1
2015 Robust relative attributes for human action recognition
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001
Pattern Anal. Appl.5
2014 Cross-View Action Recognition Using Contextual Maximum Margin Clustering
abstract
Recently, maximum margin clustering (MMC) has been proposed for a cross-view action recognition. However, such a method neglects the temporal relationship between contiguous frames in the same action video. In this paper we propose a novel method called contextual maximum margin clustering (CMMC) to tackle cross-view action recognition. In CMMC, we add temporal regularization to give a high penalty when the contiguous frames are dissimilar. Thus, the CMMC not only achieves the goal of finding maximum margin hyperplanes, but also explicitly considers the temporal information among contiguous frames. Our method is verified on the IXMAS dataset and the experimental results demonstrate that our method can achieve better performance than the state-of-the-art methods.
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2013 Cross-View Action Recognition via a Continuous Virtual Path
abstract
In this paper, we propose a novel method for cross-view action recognition via a continuous virtual path which connects the source view and the target view. Each point on this virtual path is a virtual view which is obtained by a linear transformation of the action descriptor. All the virtual views are concatenated into an infinite-dimensional feature to characterize continuous changes from the source to the target view. However, these infinite-dimensional features cannot be used directly. Thus, we propose a virtual view kernel to compute the value of similarity between two infinite-dimensional features, which can be readily used to construct any kernelized classifiers. In addition, there are a lot of unlabeled samples from the target view, which can be utilized to improve the performance of classifiers. Thus, we present a constraint strategy to explore the information contained in the unlabeled samples. The rationality behind the constraint is that any action video belongs to only one class. Our method is verified on the IXMAS dataset, and the experimental results demonstrate that our method achieves better performance than the state-of-the-art methods.
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001, Cunzhao Shi
CVPR5
2013 Tensor Ensemble of Ground-Based Cloud Sequences: Its Modeling, Classification, and Synthesis
abstract
Since clouds are one of the most important meteorological phenomena related to the hydrological cycle and affect Earth radiation balance and climate changes, cloud analysis is a crucial issue in meteorological research. Most researchers only consider the classification task of cloud images while less attention has been paid to the synthesis one. In addition, all the existing research on cloud identification from sky images is based on single cloud images. However, the cloud-measuring devices on the ground actually take one image of the clouds every few minutes and collect a series of cloud images. Thus, the existing methods neglect the temporal information exhibited by contiguous cloud images. To overcome this drawback, in this letter we treat ground-based cloud sequences (GCSs) as dynamic texture. We then propose the Tensor Ensemble of Ground-based Cloud Sequences (eTGCS) model which represents the ensemble of GCSs in a tensor manner. In the eTGCS model, all GCSs form a single tensor, and each GCS is a subtensor of the single tensor. There are two main characteristics of the eTGCS model: 1) All GCSs share an identical mode subspace, which makes the classification convenient, and 2) a new GCS can be synthesized as long as the parameters of the eTGCS model are used. Therefore, less storage space is required. Comprehensive experiments are conducted to prove the superiority of our eTGCS model. The classification accuracy achieves 92.31%, and the synthesized GCSs are similar to the original ones in visual appearance.
Shuang Liu 0001, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001, Xiaozhong Cao
IEEE Geosci. Remote. Sens. Lett.1
2013 Attribute Regularization Based Human Action Recognition
abstract
Recently, attributes have been introduced as a kind of high-level semantic information to help improve the classification accuracy. Multitask learning is an effective methodology to achieve this goal, which shares low-level features between attributes and actions. Yet such methods neglect the constraints that attributes impose on classes, which may fail to constrain the semantic relationship between the attributes and actions. In this paper, we explicitly consider such attribute-action relationship for human action recognition, and correspondingly, we modify the multitask learning model by adding attribute regularization. In this way, the learned model not only shares the low-level features, but also gets regularized according to the semantic constrains. In addition, since attribute and class label contain different amounts of semantic information, we separately treat attribute classifiers and action classifiers in the framework of multitask learning for further performance improvement. Our method is verified on three challenging datasets (KTH, UIUC, and Olympic Sports), and the experimental results demonstrate that our method achieves better results than that of previous methods on human action recognition.
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001
IEEE Trans. Inf. Forensics Secur.5
2012 Multi-scale Fusion of Texture and Color for Background Modeling
abstract
Background modeling from a stationary camera is a crucial component in video surveillance. Traditional methods usually adopt single feature type to solve the problem, while the performance is usually unsatisfactory when handling complex scenes. In this paper, we propose a multi-scale strategy, which combines both texture and color features, to achieve a robust and accurate solution. Our contributions are two folds: one is that we propose a novel textureoperator named Scale-invariant Center-symmetric Local Ternary Pattern, which is robust to noise and illumination variations, the other is that a multi-scale fusion strategy is proposed for the issue. Our method is verified on several complex real world videoswith illumination variation, soft shadows and dynamic backgrounds. We compare our method with four state-of-the-art methods, and the experimental results clearly demonstrate that our method achievesthe highest classification accuracy in complex real world videos.
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Shuang Liu 0001, Wen Zhou 0002
AVSS4
2012 Human Action Recognition with Attribute Regularization
abstract
Recently, attributes have been introduced to help object classification. Multi-task learning is an effective methodology to achieve this goal, which shares low-level features between attribute and object classifiers. Yet such a method neglects the constraints that attributes impose on classes which may fail to constrain the semantic relationship between the attribute and object classifiers. In this paper, we explicitly consider such attribute-object relationship, and correspondingly, we modify the multi-task learningmodel by adding attribute regularization. In this way, the learned model not only shares the low-level features, but also gets regularized according to the semantic constrains. Our method is verified on two challenging datasets (KTH and Olympic Sports), andthe experimental results demonstrate that our method achieves better results than previous methods in human action recognition.
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001
AVSS5
2012 Soft-signed sparse coding for ground-based cloud classification
Shuang Liu 0001, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001, Yunxue Shao
ICPR1
2012 Contextual Fisher kernels for human action recognition
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001
ICPR5
2012 Action Recognition Using Context-Constrained Linear Coding
abstract
Although traditional bag-of-words model has shown promising results for action recognition, it takes no consideration of the relationship among spatio–temporal points; furthermore, it also suffers serious quantization error. In this letter, we propose a novel coding strategy called context-constrained linear coding (CLC) to overcome these limitations. We first calculate the contextual distance between local descriptors and each codeword by considering the spatio–temporal contextual information. Then, linear coding using contextual distance is adopted to alleviate the quantization error. Our method is verified on two challenging databases (KTH and UCF sports), and the experimental results demonstrate that our method achieves better results than previous methods in action recognition.
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001
IEEE Signal Process. Lett.5