Cuiwei Liu

dblp:122/2635 · DBLP profile ↗
← Back
32ranked-venue papers
16as first author
22since 2021 · last 2026
0000-0003-4279-4841ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 7 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Subgoal-Induced Reinforcement Learning for Cross-Scene Generalization
Huaijun Qiu, Cuiwei Liu
ICIC3
2026 SLNeRF: Joint Optimization of Structured Light and NeRF
abstract
Recently, based on multi-view stereo (MVS) methods, utilizing stereo prior information to guide novel-view synthesis has become an important approach to addressing the generalization problem of neural radiance fields (NeRF). However, this approach faces challenges in handling certain difficult scenarios, such as weak texture regions, where it struggles to effectively extract geometric features information. As a result, it encounters limitations in stereo image feature representation and inaccuracies in prior depth estimation. To tackle these problems, we propose the first generalizable novel-view synthesis method that jointly optimizes structured light and NeRF. Considering that active stereo methods based on structured light optimization can enhance geometric feature extraction capability in weak texture regions through the addition of a texture layer, we integrate active stereo vision into novel-view synthesis. To this end, we propose a novel framework, dubbed SLNeRF. First, we design a structured light generation scheme based on Fourier transform and establish a differentiable imaging model using geometric optics and the Lambertian model to generate active stereo images. Then, we obtain stereo image features, as well as prior depth information through the feature extractor, which are used to construct 3D feature volumes. Finally, we accomplish the novel-view synthesis task through a neural renderer. Compared to state-of-the-art generalizable NeRF methods, our method reports encouraging results on public datasets as well as in real-world scenarios.
Tong Jia 0001, Shuyang Lin, Dongyue Chen 0001, Ping Xiao, Cuiwei Liu
IEEE Trans. Multim.6
2025 Bias Mitigation in Federated Few-Shot Class-Incremental Learning via Multi-Prototype Collaboration
abstract
Federated learning aims to collaboratively train a shared global model from multiple clients while preserving data privacy. However, real-world applications often involve clients learning from limited and dynamically arriving data, requiring the global model to classify all encountered classes. This paper introduces federated few-shot class-incremental learning, enabling effective learning of new classes from scarce samples within a decentralized framework. Existing methods suffer from new class bias, where new classes are often misclassified as previously learned ones. Additionally, they face local bias due to non-IID data distribution, which leads client models to focus excessively on their specific local data characteristics. We propose a Decoupled Multi-Prototype Collaboration (DMPC) method to mitigate both biases. First, we introduce a Global Consistency Aggregation mechanism (GCA) that re-weights local prototypes based on their consistency, resulting in more representative global prototypes and effectively eliminating local bias. Second, we design a Multi-Prototype Testing strategy (MPT) that enhances classification accuracy by leveraging both local and global prototypes, thereby mitigating new class bias. More importantly, GCA and MPT exhibit significant synergistic effects. Extensive experiments on three widely used datasets demonstrate the robustness and superiority of our method in bias reduction.
Siang Xu, Huaijun Qiu, Zhuo Yan, Cuiwei Liu
IJCNN4
2025 Instance-Specific Learning for Skeleton-Based Action Recognition with Varying Data Quality
abstract
Skeleton-based action recognition technology has gained significant attention and made great progress in recent years. However, the performance of existing methods declines significantly when the quality of skeleton data extracted by pose estimation algorithms varies. To address this issue, this study proposes an instance-specific learning method aimed at enhancing the model’s ability to learn discriminative features when handling skeleton data of varying quality. We introduce a Dynamic Instance Discriminability Assessment (DIDA) mechanism and a Staged Instance Weighting (SIW) strategy. The DIDA mechanism dynamically evaluates the discriminability of instances by combining prior knowledge with feedback from the model during the training process. The SIW strategy adjusts the weights of instances at different training stages based on their discriminability. Notably, our method requires only a minimal increase in computational cost during training and incurs no additional computational overhead during testing compared to baseline models. We utilized Pifpaf and HR-Net pose estimation methods to extract skeleton data of varying quality from the NTU60, NTU120, and HMDB51 video datasets and conducted extensive experimental validation. The results indicate that the proposed method significantly enhances the action recognition performance while maintaining computational efficiency.
Huaijun Qiu, Zhuo Yan, Cuiwei Liu
IJCNN4
2025 IHGSL: Interpretable Heuristic Graph Structure Learning for Multi-Robot Autonomous Collaborative Systems
abstract
In multi-robot systems, capturing the complex and dynamic interaction relationships is essential for enhancing autonomous collaboration. However, existing learning-based approaches usually overlook the understanding of these relationships, leading to reliability issues and hindering their application to real-world scenarios. This paper proposes a novel approach called Interpretable Heuristic Graph Structure Learning (IHGSL) to better comprehend the complex collaborative relationships in multi-robot systems. We first construct a predicate space to define diverse predicates that express fundamental relationships. Then we employ the variational information bottleneck technique to acquire a latent representation of the current observation by aligning it with the historical trajectory. On this basis, the predicates that the robot should currently focus on the most are learned, and some interaction relationships are established accordingly. Thereby an interpretable relationship graph is generated heuristically to guide the achievement of multi-robot autonomous collaborative decision-making. Through experimental evaluation, we demonstrate the process of relationship inference, thus validating the interpretability of IHGSL. Compared with existing methods, IHGSL also achieves superior collaboration performance, which highlights the effectiveness of the learned heuristic graph structure.
Cuiwei Liu, Zhixiao Sun
IROS3
2025 Few-Shot Class-Incremental Learning With Non-IID Decentralized Data
abstract
Few-shot class-incremental learning is crucial for developing scalable and adaptive intelligent systems, as it enables models to acquire new classes with minimal annotated data while safeguarding the previously accumulated knowledge. Nonetheless, existing methods deal with continuous data streams in a centralized manner, limiting their applicability in scenarios that prioritize data privacy and security. To this end, this paper introduces federated few-shot class-incremental learning, a decentralized machine learning paradigm tailored to progressively learn new classes from scarce data distributed across multiple clients. In this learning paradigm, clients locally update their models with new classes while preserving data privacy, and then transmit the model updates to a central server where they are aggregated globally. However, this paradigm faces several issues, such as difficulties in few-shot learning, catastrophic forgetting, and data heterogeneity. To address these challenges, we present a synthetic data-driven framework that leverages replay buffer data to maintain existing knowledge and facilitate the acquisition of new knowledge. Within this framework, a noise-aware generative replay module is developed to fine-tune local models with a balance of new and replay data, while generating synthetic data of new classes to further expand the replay buffer for future tasks. Furthermore, a class-specific weighted aggregation strategy is designed to tackle data heterogeneity by adaptively aggregating class-specific parameters based on local models performance on synthetic data. This enables effective global model optimization without direct access to client data. Comprehensive experiments across three widely-used datasets underscore the effectiveness and preeminence of the introduced framework. We will release our code at: https://github.com/XuSiang1/F2SCIL-SDD.
Cuiwei Liu, Siang Xu, Huaijun Qiu, Zhi Liu 0002, Liang Zhao 0004
IEEE Internet Things J.1
2025 An Open-Set Domain Adaptation Framework for Hyperspectral Image Classification With Pixel-Aware Weighting and Decoupled Alignment
abstract
Recent studies have shown that deep domain adaptation techniques perform excellently in cross-domain hyperspectral image classification. However, these methods typically assume that the source domain and the target domain share the same class set, while in practice, the target domain may include unknown classes, and direct alignment can result in negative transfer. Moreover, in hyperspectral image classification based on deep learning, using the label of the central pixel to represent the label of the image patch may lead to feature bias due to the uncertainty of the labels of neighboring pixels, thereby reducing the generalization performance of the model. To address this, this paper proposes an open-set domain adaptation framework, including a Pixel-Aware Weight Learning (PAWL) module and a Decoupled Dual Alignment (DDA) strategy. The PAWL module effectively reduces the feature bias caused by inconsistency in neighboring pixel labels by analyzing the uncertainty of neighboring pixel labels and utilizing adaptive weight learning, thereby improving recognition performance in open-set environments. The DDA strategy decouples the features of the source domain and target domain into known and unknown classes and aligns them separately to mitigate negative transfer. Experiments on two cross-scene hyperspectral datasets validated the effectiveness of the method.
Zhaokui Li, Mingtai Qi, Yan Wang 0087, Xuewei Gong, Cuiwei Liu, Jinjun Wang
IEEE Geosci. Remote. Sens. Lett.5
2025 Hyperspectral Target Detection Using Diffusion Model and Convolutional Gated Linear Unit
abstract
Deep learning can effectively extract latent information from data to enhance target-background separation in hyperspectral target detection (HTD). However, these models typically require extensive labeled samples, while available target spectra in hyperspectral images (HSI) are scarce. Additionally, existing deep models struggle with target detection in complex backgrounds due to subtle spectral differences. To address these issues, we propose a novel HTD method based on diffusion model and convolutional gated linear unit (HTD-DMCG). First, the diffusion model is integrated with MixUp for data augmentation to generate a diverse and sufficiently large sample set. Next, a Transformer architecture utilizing a convolutional gated linear unit is designed to effectively capture global dependencies and local feature correlations, leading to more discriminative feature representations. Additionally, a new target aggregation and background separation loss is introduced, which emphasizes target sample aggregation while increasing the distance between targets and background samples to enhance separability. The HTD-DMCG method is compared against classical and state-of-the-art HTD methods on four real HSI datasets. Extensive experiments show that it can effectively outperform existing methods in target detection performance. The code is available at https://github.com/Li-ZK/HTD-DMCG.
Zhaokui Li, Xiaobin Zhao, Cuiwei Liu, Xuewei Gong, Wei Li 0032, Qian Du 0001, Bo Yuan 0013
IEEE Trans. Geosci. Remote. Sens.4
2025 A Novel EAGLe Framework for Robust UAV-View Geo-Localization
abstract
This paper addresses the UAV-view geo-localization task, which focuses on bi-directional retrieval between UAV-view and satellite-view images. Generally, existing methods aim to learn image representations that can distinguish between different locations while effectively mitigating the cross-view domain gap. However, these methods often struggle in noisy UAV flight environments, as they fail to account for environmental domain shifts caused by varying weather and lighting conditions. To this end, we propose a novel Environment-Agnostic Geo-Localization (EAGLe) framework, which integrates a dual-objective discriminator and a style mixture module into diverse UAV-view geo-localization networks to enhance their robustness in dynamic environments. Specifically, the dual-objective discriminator not only distinguishes between UAV and satellite views but also identifies various environmental styles in UAV-view images. Through adversarial learning, the dual-objective discriminator encourages the feature encoder to produce features that remain invariant to both viewpoint and environmental variations. Furthermore, the style mixture module is integrated into the feature encoder to extend diversity at the feature level, allowing EAGLe to learn a broader range of environmental styles beyond the training data. Extensive experiments on the University-1652 and SUES-200 datasets demonstrate that the proposed EAGLe significantly improves the reliability of UAV-view geo-localization networks under dynamic and unpredictable environmental conditions, while maintaining inference efficiency.
Cuiwei Liu, Shiting Peng, Shishen Li, Huaijun Qiu, Yuhao Xia, Zhaokui Li, Liang Zhao 0004
IEEE Trans. Geosci. Remote. Sens.1
2024 Fusion Attention Graph Convolutional Network with Hyperskeleton for UAV Action Recognition
Qin Dai, Cuiwei Liu, Xiangbin Shi
ICIC (12)4
2024 Spatio-Temporal Graph Learning for Enhanced Agent Collaboration in Multi-Aircraft Combat
abstract
In recent years, significant progress has been made in Multi-Agent Deep Reinforcement Learning (MADRL) for addressing cooperative decision-making challenges in multi-aircraft air combat tasks. This paper introduces a novel Spatio-Temporal Relationship Graph Structure Learning method (STRGSL), aimed at overcoming the challenges in capturing the complex and dynamic interactions between agents. The proposed STRGSL constructs a historical behavior graph based on past observations as well as a real-time interaction graph from current observations, providing a comprehensive consideration of both immediate and long-term agent relationships. Leveraging a Graph Neural Network (GNN), STRGSL generates agent representations that fuse current and historical relationships. By integrating STRGSL into a MADRL framework, we jointly optimize both the structure of relationship graphs and the cooperative policies of agents. Experiments carried out in an aircraft combat scenario and two multi-agent cooperative scenarios demonstrate that the proposed STRGSL promotes collaboration among multiple agents, thereby enhancing the overall performance across different scenarios.
Zhengchao Wang, Cuiwei Liu, Huaijun Qiu
IJCNN2
2024 Enhancing action recognition from low-quality skeleton data via part-level knowledge distillation
Cuiwei Liu, Youzhi Jiang, Chong Du, Zhaokui Li
Signal Process.1
2024 Adaptive Global Embedding Learning: A Two-Stage Framework for UAV-View Geo-Localization
abstract
This letter aims to deal with the UAV-view geolocalization problem, which is essentially to achieve bi-directional cross-view matching between UAV-view and satellite-view images. The existing studies have confirmed the importance of learning part-wise representations for this task. We go a step further by proposing a two-stage learning framework. The first stage focuses on extracting part-wise representations. In the second stage, a novel Adaptive Embedding Network (AEN) integrates these representations into a global embedding of the entire image to avoid an equal influence of all local parts on image similarity measures. Current mainstream methods typically employ CrossEntropy loss to learn location-dependent representations, aiming to push the distance between different locations in the learned representation space. Some approaches also utilize KL loss or Triplet loss to bring a pair of UAV-satellite images from the same location closer for learning view-invariant representations. However, they overlook a critical concern: a notable representation bias exists among UA-view images captured from the same location but at different viewpoints or heights. To address these issues, we devise a novel cross-view matching loss that narrows the distance between the global embeddings of a satellite-view image and the affinity-aware prototype of multiple true-matched UAV-view images. The experimental results on the University1652 dataset indicate that similarity measures in the learned embedding space exhibit excellent generalization to images from new locations, achieving superior cross-view matching performance compared to previous methods
Cuiwei Liu, Shishen Li, Chong Du, Huaijun Qiu
IEEE Signal Process. Lett.1
2023 View Distribution Alignment with Progressive Adversarial Learning for UAV Visual Geo-Localization
Cuiwei Liu, Huaijun Qiu, Zhaokui Li, Xiangbin Shi
KSEM (2)1
2023 A Transformer-Based Adaptive Semantic Aggregation Method for UAV Visual Geo-Localization
Shishen Li, Cuiwei Liu, Huaijun Qiu, Zhaokui Li
PRCV (4)2
2023 Design and Implementation of Mask Detection System Based on Improved YOLOv5s
abstract
In this paper, we propose a lightweight mask detection algorithm and implement an intelligent vehicle system. The algorithm uses YOLOv5s as the backbone network, and at the same time incorporates the SE attention mechanism to optimize the timeliness, and is finally deployed on an intelligent vehicle system with BCM2711 as the control platform. Experiments prove that the algorithm proposed in this paper reduces the detection time by 30% while ensuring a higher MAP, which has certain value for promotion.
Changyu Zhao, Zhuo Yan, Huangxin Xu, Xueliang Chen, Xinyu Zhong, Cuiwei Liu, Anyan Xiao, Xingyan Lv
TrustCom7
2023 Few-Shot Hyperspectral Image Classification With Self-Supervised Learning
abstract
Recently, few-shot learning (FSL) has been introduced for hyperspectral image (HSI) classification with few labeled samples. However, existing FSL-based HSI classification methods mainly focus on the meta-knowledge transfer between HSIs. Compared with HSIs, natural images have sufficient annotated data. To utilize natural images (base class data) to achieve accurate classification of HSIs (novel class data), we propose a novel few-shot classification framework with SSL (FSCF-SSL) for HSIs in this article. The orientation of objects in natural images is relatively unitary, whereas the objects of image patches for each pixel in HSIs have diverse orientations in the spatial domain. To make better use of base classes, we design an SSL with geometric transformations (SSLGTs), which sets rotation labels as supervision to extract low-level features that can better represent diverse orientations, and then conduct SSLGT and FSL on base classes to learn transferable spatial meta-knowledge. Next, a spectral-spatial feature extraction network is carefully designed to better utilize the spatial and spectral information of HSIs, where the weights of the first seven layers of the spatial part are initialized by the weights of the corresponding layers trained on base classes. Finally, to fully explore the few annotated data from novel classes, we design an SSL with contrastive learning (SSLCL) that can mine the category-invariant features contained in the novel class data itself, and then perform SSLCL and FSL on novel classes to learn more discriminative individual knowledge. Experimental results on four HSI datasets show that FSCF-SSL offers a significant improvement over state-of-the-art methods. The code is available athttps://github.com/Li-ZK/FSCF-SSL-2023.
Zhaokui Li, Yushi Chen 0002, Cuiwei Liu, Qian Du 0001, Zhuoqun Fang, Yan Wang 0087
IEEE Trans. Geosci. Remote. Sens.4
2022 A Graph Convolutional Network with Early Attention Module for Skeleton-based Action Prediction
abstract
This paper addresses the problem of skeleton-based action prediction, which aims to predict the action label when the skeleton sequence is partially observed. The action prediction task is more challenging compared to the after-the-fact action recognition since it needs to make decisions according to the beginning part of action executions. The existing methods improve the action prediction performance by taking advantage of the global action knowledge in full sequences, and some of them require the correspondence between a partial sequence and its associated full sequence. In this paper, we step towards a new direction by exploiting the discriminative information in early observations of actions as much as possible. We propose a Graph Convolutional Network with Early Attention Module (GCN-EAM), which employs a series of spatial-temporal graph convolution blocks to extract features from skeletons. In order to infer the action category as fast as possible, we introduce an early attention module to adaptively emphasize discriminative observations at the beginning stage of actions. The proposed method is evaluated on the large-scale NTU-RGB+D dataset and achieves excellent performance for action prediction.
Cuiwei Liu, Zhuo Yan, Youzhi Jiang, Xiangbin Shi
ICPR1
2022 A Novel Two-Stage Knowledge Distillation Framework for Skeleton-Based Action Prediction
abstract
This letter addresses the challenging problem of action prediction with partially observed sequences of skeletons. Towards this goal, we propose a novel two-stage knowledge distillation framework, which transfers prior knowledge to assist the early prediction of ongoing actions. In the first stage, the action prediction model (also referred to as the student) learns from a couple of teachers to adaptively distill action knowledge at different progress levels for partial sequences. Then the learned student acts as a teacher in the next stage, with the objective of optimizing a better action prediction model in a self-training manner. We design an adaptive self-training strategy from the perspective of undermining the supervision from the annotated labels, since this hard supervision is actually too strict for partial sequences without enough discriminative information. Finally, the action prediction models trained in the two stages jointly constitute a two-stream architecture for action prediction. Extensive experiments on the large-scale NTU RGB+D dataset validate the effectiveness of the proposed method.
Cuiwei Liu, Zhaokui Li, Zhuo Yan, Chong Du
IEEE Signal Process. Lett.1
2022 Dual-Channel Residual Network for Hyperspectral Image Classification With Noisy Labels
abstract
Hyperspectral image (HSI) classification has drawn increasing attention recently. However, it suffers from noisy labels that may occur during field surveys due to a lack of prior information or human mistakes. To address this issue, this article proposes a novel dual-channel residual network (DCRN) to resolve HSI classification with noisy labels. Currently, the influence of noisy labels is reduced by simply detecting and removing those anomalous samples. Different from such a specifically designed noise cleansing method, DCRN is easy to implement but highly effective. It enhances its model robustness to noisy labels to a great extent by employing a novel dual-channel structure and a noise-robust loss function. In this way, DCRN can mitigate influence from noisy labels while fully utilizing useful information from mislabeled samples for augmented training. Experiments are conducted on several hyperspectral data sets with manually generated noisy labels to demonstrate its excellent performance. The code is available athttps://github.com/Li-ZK/DCRN-2021.
Yimin Xu, Zhaokui Li, Wei Li 0032, Qian Du 0001, Cuiwei Liu, Zhuoqun Fang, Lin Zhai
IEEE Trans. Geosci. Remote. Sens.5
2021 A Novel Key Point Trajectory Model for Fall Detection from RGB-D Videos
abstract
This paper aims to address the problem of fall detection from RGB-D image sequences. Towards this goal, we propose a novel Key Point Trajectory Model which represents a fall action as a series of trajectory descriptors. In the proposed model, 16 key points including 14 skeleton points and 2 centers of body parts are extracted from each pair of RGB and depth images. Then a global trajectory descriptor is constructed on 16 trajectories that are obtained by connecting the key points across several frames in the RGB-D sequence. The trajectory descriptor incorporates the spatial, depth, and temporal context of key points and characterizes the global motion of human over a short period of time. A random forest is employed to learn the classifier of trajectory descriptors, and an integration rule is developed for detecting falls according to the classification results of all trajectory descriptors within a video. Experiments conducted on two fall detection datasets demonstrate that our method achieves better performance in comparison with state-of-the-art methods.
Cuiwei Liu, Jianxiong Lv, Zhaokui Li, Zhuo Yan, Xiangbin Shi
CSCWD1
2021 Action Prediction Network with Auxiliary Observation Ratio Regression
abstract
This paper focuses on predicting the category of an ongoing action with incomplete observations. We propose a novel Action Prediction Network with Auxiliary Observation Ratio regression (AORAP Net), which enhances the discriminative power of partial video features by encoding the prior knowledge of complete actions. The proposed AORAP Net consists of an encoder, an observation ratio regression module, and an action classifier. The encoder transfers global action information from full videos to construct enhanced partial video features. The observation ratio regression module is developed to achieve an auxiliary task of inferring the progress level of an ongoing action. We demonstrate that this module can guide the encoder to learn enhanced features similar to full video features and discriminative for action prediction. The classifier makes the final decision of the action category in terms of the enhanced features. Considering the fact that there is not enough discriminative information at the early stage of actions, a new classification loss is designed to adapt to the action prediction task and alleviate the over-fitting of training videos. Extensive experiments on the BIT-Interaction dataset and the UT-Interaction dataset validate the effectiveness of the proposed method.
Cuiwei Liu, Yiming Gao 0001, Zhaokui Li, Chong Du, Xiangbin Shi
ICME1
2018 A Deep Network Based on Multiscale Spectral-Spatial Fusion for Hyperspectral Classification
Zhaokui Li, Deyuan Zhang, Cuiwei Liu, Yan Wang 0087, Xiangbin Shi
KSEM (2)4
2018 A discriminative structural model for joint segmentation and recognition of human actions
Cuiwei Liu, Jingyi Hou, Xinxiao Wu, Yunde Jia
Multim. Tools Appl.1
2017 Action Prediction Using Unsupervised Semantic Reasoning
Cuiwei Liu, Yaguang Lu, Xiangbin Shi, Zhaokui Li, Liang Zhao 0004
ICONIP (3)1
2016 A Hierarchical Video Description for Complex Activity Understanding
Cuiwei Liu, Xinxiao Wu, Yunde Jia
Int. J. Comput. Vis.1
2016 Transfer Latent SVM for Joint Recognition and Localization of Actions in Videos
abstract
In this paper, we develop a novel transfer latent support vector machine for joint recognition and localization of actions by using Web images and weakly annotated training videos. The model takes training videos which are only annotated with action labels as input for alleviating the laborious and time-consuming manual annotations of action locations. Since the ground-truth of action locations in videos are not available, the locations are modeled as latent variables in our method and are inferred during both training and testing phrases. For the purpose of improving the localization accuracy with some prior information of action locations, we collect a number of Web images which are annotated with both action labels and action locations to learn a discriminative model by enforcing the local similarities between videos and Web images. A structural transformation based on randomized clustering forest is used to map the Web images to videos for handling the heterogeneous features of Web images and videos. Experiments on two public action datasets demonstrate the effectiveness of the proposed model for both action localization and action recognition.
Cuiwei Liu, Xinxiao Wu, Yunde Jia
IEEE Trans. Cybern.1
2015 Cross-View Action Recognition Over Heterogeneous Feature Spaces
abstract
In cross-view action recognition, what you saw in one view is different from what you recognize in another view, since the data distribution even the feature space can change from one view to another. In this paper, we address the problem of transferring action models learned in one view (source view) to another different view (target view), where action instances from these two views are represented by heterogeneous features. A novel learning method, called heterogeneous transfer discriminant-analysis of canonical correlations (HTDCC), is proposed to discover a discriminative common feature space for linking source view and target view to transfer knowledge between them. Two projection matrices are learned to, respectively, map data from the source view and the target view into a common feature space via simultaneously minimizing the canonical correlations of interclass training data, maximizing the canonical correlations of intraclass training data, and reducing the data distribution mismatch between the source and target views in the common feature space. In our method, the source view and the target view neither share any common features nor have any corresponding action instances. Moreover, our HTDCC method is capable of handling only a few or even no labeled samples available in the target view, and can also be easily extended to the situation of multiple source views. We additionally propose a weighting learning framework for multiple source views adaptation to effectively leverage action knowledge learned from multiple source views for the recognition task in the target view. Under this framework, different source views are assigned different weights according to their different relevances to the target view. Each weight represents how contributive the corresponding source view is to the target view. Extensive experiments on the IXMAS data set demonstrate the effectiveness of HTDCC on learning the common feature space for heterogeneous cross-view action recognition. In addition, the weighting learning framework can achieve promising results on automatically adapting multiple transferred source-view knowledge to the target view.
Xinxiao Wu, Han Wang 0001, Cuiwei Liu, Yunde Jia
IEEE Trans. Image Process.3
2014 Weakly Supervised Action Recognition and Localization Using Web Images
Cuiwei Liu, Xinxiao Wu, Yunde Jia
ACCV (5)1
2014 Learning a discriminative mid-level feature for action recognition
Cuiwei Liu, Mingtao Pei, Xinxiao Wu, Yu Kong 0001, Yunde Jia
Sci. China Inf. Sci.1
2013 Cross-View Action Recognition over Heterogeneous Feature Spaces
abstract
In cross-view action recognition, "what you saw" in one view is different from "what you recognize" in another view. The data distribution even the feature space can change from one view to another due to the appearance and motion of actions drastically vary across different views. In this paper, we address the problem of transferring action models learned in one view (source view) to another different view (target view), where action instances from these two views are represented by heterogeneous features. A novel learning method, called Heterogeneous Transfer Discriminantanalysis of Canonical Correlations (HTDCC), is proposed to learn a discriminative common feature space for linking source and target views to transfer knowledge between them. Two projection matrices that respectively map data from source and target views into the common space are optimized via simultaneously minimizing the canonical correlations of inter-class samples and maximizing the intraclass canonical correlations. Our model is neither restricted to corresponding action instances in the two views nor restricted to the same type of feature, and can handle only a few or even no labeled samples available in the target view. To reduce the data distribution mismatch between the source and target views in the common feature space, a nonparametric criterion is included in the objective function. We additionally propose a joint weight learning method to fuse multiple source-view action classifiers for recognition in the target view. Different combination weights are assigned to different source views, with each weight presenting how contributive the corresponding source view is to the target view. The proposed method is evaluated on the IXMAS multi-view dataset and achieves promising results.
Xinxiao Wu, Han Wang 0001, Cuiwei Liu, Yunde Jia
ICCV3
2012 Action recognition with discriminative mid-level features
Cuiwei Liu, Yu Kong 0001, Xinxiao Wu, Yunde Jia
ICPR1