Qiuxia Wu

dblp:68/10772 · DBLP profile ↗
← Back
36ranked-venue papers
9as first author
25since 2021 · last 2026
0000-0002-2284-7806ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 17 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 2 since 2021Security and privacy · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Pb4U-GNet: Resolution-Adaptive Garment Simulation via Propagation-before-Update Graph Network
abstract
Garment simulation is fundamental to various applications in computer vision and graphics, from virtual try-on to digital human modelling. However, conventional physics-based methods remain computationally expensive, hindering their application in time-sensitive scenarios. While graph neural networks (GNNs) offer promising acceleration, existing approaches exhibit poor cross-resolution generalisation, demonstrating significant performance degradation on higher-resolution meshes beyond the training distribution. This stems from two key factors: (1) existing GNNs employ fixed message-passing depth that fails to adapt information aggregation to mesh density variation, and (2) vertex-wise displacement magnitudes are inherently resolution-dependent in garment simulation. To address these issues, we introduce Propagation-before-Update Graph Network (Pb4U-GNet), a resolution-adaptive framework that decouples message propagation from feature updates. Pb4U-GNet incorporates two key mechanisms: (1) dynamic propagation depth control, adjusting message-passing iterations based on mesh resolution, and (2) geometry-aware update scaling, which scales predictions according to local mesh characteristics. Extensive experiments show that even trained solely on low-resolution meshes, Pb4U-GNet exhibits strong generalisability across diverse mesh resolutions, addressing a fundamental challenge in neural garment simulation.
Aoran Liu, Kun Hu 0008, Clinton Mo, Qiuxia Wu, Wenxiong Kang, Zhiyong Wang 0001
AAAI4
2026 Layered Evidence-Centric Graph Construction for Explainable Multi-hop Question Answering
Weiguo Zeng, Haoyang Xie, Zhenxuan Chao, Qiuxia Wu, Jianfeng Qu, Zhixu Li
DEXA (1)4
2025 RI-MAE: Rotation-Invariant Masked AutoEncoders for Self-Supervised Point Cloud Representation Learning
abstract
Masked point modeling methods have recently achieved great success in self-supervised learning for point cloud data. However, these methods are sensitive to rotations and often exhibit sharp performance drops when encountering rotational variations. In this paper, we propose a novel Rotation-Invariant Masked AutoEncoders (RI-MAE) to address two major challenges: 1) achieving rotation-invariant latent representations, and 2) facilitating self-supervised reconstruction in a rotation-invariant manner. For the first challenge, we introduce RI-Transformer, which features disentangled geometry content, rotation-invariant relative orientation and position embedding mechanisms for constructing rotation-invariant point cloud latent space. For the second challenge, a novel dual-branch student-teacher architecture is devised. It enables the self-supervised learning via the reconstruction of masked patches within the learned rotation-invariant latent space. Each branch is based on an RI-Transformer, and they are connected with an additional RI-Transformer predictor. The teacher encodes all point patches, while the student solely encodes unmasked ones. Finally, the predictor predicts the latent features of the masked patches using the output latent embeddings from the student, supervised by the outputs from the teacher. Extensive experiments demonstrate that our method is robust to rotations, achieving the state-of-the-art performance on various downstream tasks.
Kunming Su, Qiuxia Wu, Panpan Cai, Xiaogang Zhu 0001, Xuequan Lu, Zhiyong Wang 0001, Kun Hu 0008
AAAI2
2025 WaveLoss: An Adaptive Dynamic Loss for Deep Gait Recognition
abstract
Designing an appropriate loss function can enhance the discriminative power on gait recognition. However, previous research focuses on improving network structure and enriching input modalities but overlooks the loss functions. Although transferring loss functions from face recognition can address sample-level loss, additional design is needed for part-level loss. Therefore, we have designed a new loss function called Waveloss, aimed at adaptively and dynamically changing the preference for parts of different difficulties. First, the previous method treats the loss of different parts equally, which brings the problems of difficult convergence or susceptibility to noise interference, so we propose norm-fusion to adaptively learn samples of different difficulties. Additionally, since we find the exponential value represents preference for learning different samples, we introduce the Dynamic Learning Process, which dynamically adjusts the exponential value during iteration to focus on samples of varying difficulties at different training stages. Finally, as the changes of the exponential value leads to significant fluctuations in the gradient, we introduce the gradient truncation and normalization to avoid getting trapped in local optima and gradient vanishing or exploding by adaptively adjusting the gradient. Experimental results demonstrate that our proposed Waveloss achieves state-of-the-art performance on various gait recognition datasets and can improve the performance of different backbones as well.
Qiuxia Wu
AAAI2
2025 DC-PCN: Point Cloud Completion Network with Dual-Codebook Guided Quantization
abstract
Point cloud completion aims to reconstruct complete 3D shapes from partial 3D point clouds. With advancements in deep learning techniques, various methods for point cloud completion have been developed. Despite achieving encouraging results, a significant issue remains: these methods often overlook the variability in point clouds sampled from a single 3D object surface. This variability can lead to ambiguity and hinder the achievement of more precise completion results. Therefore, in this study, we introduce a novel point cloud completion network, namely Dual-Codebook Point Completion Network (DC-PCN), following an encder-decoder pipeline. The primary objective of DC-PCN is to formulate a singular representation of sampled point clouds originating from the same 3D surface. DC-PCN introduces a dual-codebook design to quantize point-cloud representations from a multilevel perspective. It consists of an encoder-codebook and a decoder-codebook, designed to capture distinct point cloud patterns at shallow and deep levels. Additionally, to enhance the information flow between these two codebooks, we devise an information exchange mechanism. This approach ensures that crucial features and patterns from both shallow and deep levels are effectively utilized for completion. Extensive experiments on the PCN, ShapeNet_Part, and ShapeNet34 datasets demonstrate the state-of-the-art performance of our method.
Qiuxia Wu, Kunming Su, Zhiyong Wang 0001, Kun Hu 0008
AAAI1
2025 SCDFormer: Spatial and Channel Denoising Transformer for Human Pose Estimation Using Millimeter-Wave Radar
abstract
The millimeter-wave radar-based human pose estimation technology has attracted significant attention due to its cost-effectiveness and non-intrusive nature. However, the existing methods focus on modeling spatial dependence but overlook channel connections, failing to highlight important channels. In addition, the millimeter-wave point cloud is noisy, which implies some point-pair connections are irrelevant and channel connections are noisy. To address the above issues, we propose two basic units named Spatial Denoising Self-Attention Layer (SDAL) and Channel Denoising Self-Attention Layer (CDAL). SDAL filters out the irrelevant point-pair connections while modeling spatial dependence. CDAL emphasizes the important channels and filters out the noisy channel connections. Based on the SDAL and CDAL, we propose a new network named Spatial and Channel Denoising Transformer (SCDFormer). Our SCDFormer can not only filter out the irrelevant point-pair connections, but also emphasize important channels and filter out the noisy channel connections. Experiments demonstrate that our SCDFormer achieves state-of-the-art on the Mars, mRI and MiliPoint datasets.
Qiuxia Wu, Panpan Cai, Wenxiong Kang
IJCB1
2025 MoTeNet: Motion-Temporal Network for Dynamic Hand Gesture Recognition on Point Clouds
abstract
As deep learning techniques are increasingly applied to gesture recognition, point cloud-based dynamic gesture recognition methods have attracted significant attention. However, many existing approaches overlook per-point motion features and the overall temporal dynamics of gestures, thereby failing to fully exploit motion and temporal cues embedded in point cloud sequences. To address this issue, we propose Motion-Temporal Network (MoTeNet), a novel framework for dynamic gesture recognition on point clouds, which preserves spatial structural information while modeling both global frame-level temporal changes and per-point motion. MoTeNet extracts spatial geometric features through a Hierarchical Graph Convolution (HGC) module, and incorporates a Motion Feature Encoding Module (MFEM) to encode point-level motion features across adjacent frames. This approach provides fine-grained dynamic information for subsequent temporal modeling. Furthermore, an Adaptive Temporal Feature Fusion (ATFF) module integrates convolutional neural networks (CNNs) and transformers to adaptively fuse short-term and long-term temporal dependencies, enabling comprehensive modeling of the dynamic evolution of gestures. Experimental results demonstrate that MoTeNet achieves state-of-the-art performance on point cloud gesture recognition benchmarks, including SHREC’17, DHG, and NVGesture, with ablation studies further validating the effectiveness of the proposed framework.
Qiuxia Wu, Xinran Xie, Sangni Xu, Wenxiong Kang
IJCB1
2025 T2WI-BCMIC: Non-Fat Saturated T2-Weighted Imaging Dataset for Bladder Cancer Muscle Invasion Classification
Han Huang 0002, Qiuxia Wu, Huanjun Wang, Qian Cai
MICCAI (13)3
2025 Physics based circuit compatible model for hybrid antiferroelectric random access memory
Qiuxia Wu, Wenwu Xiao, Wenxuan Ma 0008, Chunfu Zhang, Xiaohua Ma 0001, Yue Hao 0001
Sci. China Inf. Sci.1
2025 Music source separation via hybrid waveform and spectrogram based generative adversarial network
Qiuxia Wu, Haipeng Deng, Kun Hu 0008, Zhiyong Wang 0001
Multim. Tools Appl.1
2024 GSTNet: Gait Spatio-Temporal Network for Gait Recognition Using Millimeter-Wave Radar
abstract
The millimeter-wave (mmWave) radar-based gait recognition technology has attracted significant attention due to its cost-effectiveness and weather resilience. The majority of existing methods are primarily designed for the recognition of gait patterns on the fixed route, and they achieve significant performance. However, fewer advancements have been made in gait pattern recognition for the free route, primarily owing to the challenges posed by data sparsity and route diversity in such contexts. To take those issues into account, we present a novel framework named Gait Spatio-Temporal Network (GSTNet), designed for efficient human gait recognition from spatiotemporal features. The GSTNet is composed of Multi-Temporal Resolution DGCNN (MTRD) and Dynamic Feature Capturing Module (DFCM). The MTDR regroups frames before edge convolution, extracting not only the spatial features of each frame but also the dynamic features from neighboring frames to cope with relative data sparsity. The DFCM adaptively generates weights for each frame through temporal features to handle diverse gait samples. Experiments demonstrate our GSTNet achieves state-of-the-art on the mmGait and STPointGCN datasets. Additionally, we provide results from the ablation study to further validate the efficacy of the proposed framework.
Qiuxia Wu, Kunming Su, Sangni Xu
ICASSP1
2024 Fast Online Adaptation of Visual SLAM via Variational Information Transfer and Preservation
Sangni Xu, Hao Xiong 0001, Qiuxia Wu, Shlomo Berkovsky, Zhiyong Wang 0001
MMAsia3
2024 URINet: Unsupervised point cloud rotation invariant representation learning via semantic and structural reasoning
Qiuxia Wu, Kunming Su
Comput. Vis. Image Underst.1
2023 Material-Aware Self-Supervised Network for Dynamic 3D Garment Simulation
abstract
Dynamic 3D garment simulation has various applications in many domains. Recently, self-supervised learning for this task has been studied to reduce annotation costs. However, different material characteristics of garments have been rarely explored, limiting the generalization and flexibility of existing methods. Therefore, in this paper, a novel self-supervised deep learning architecture is proposed, namely Material-aware Self-supervised Network (MSN), as a material-aware approach for dynamically simulating garments with different materials. Specifically, a material-aware parameterized regressor is introduced based on the observation that material characteristics change continuously regarding the fabric parameters. As a result, MSN realises real-time garment simulation with various material properties without model re-training. Moreover, to simulate garments of different categories (e.g., t-shirts vs. dresses), a sampling-based linear skinning strategy is studied in MSN. Comprehensive experiments on the widely used AMASS dataset demonstrated the effectiveness of MSN both quantitatively and qualitatively.
Aoran Liu, Kun Hu 0008, Wenxi Yue, Qiuxia Wu, Zhiyong Wang 0001
ICME4
2023 Online Visual SLAM Adaptation against Catastrophic Forgetting with Cycle-Consistent Contrastive Learning
abstract
Visual SLAM (Simultaneous Localisation and Mapping) aims to simultaneously estimate camera poses and depth maps from navigation videos captured. While recent deep learning based methods have achieved great success on this task, they tend to work well on source domain data and suffer from performance degradation on the unseen data of target domain. Hence, we propose an online adaptation approach to continuously adapt a pre-trained visual SLAM model to changing environments in a self-supervised manner. To preserve pre-learned knowledge against catastrophic forgetting, we perform updating on a novel adapter proposed rather than fine-tuning the whole model for adaptation. The adapter includes a cross-domain feature translation module that translates pre-learned features into translated features suitable for adaptation. Ideally, the translated new features should not only contain pre-learned knowledge but also substantially distinct from pre-learned features since these two features represent different domains. We thus introduce cycle-consistent contrastive learning to maximize the dissimilarity between these two features by enlarging the distance between them in the feature space. Besides, our contrastive learning method exploiting cycle-consistency contraint enables the translated features to be transferred back to the pre-learned ones, which helps the translated features better preserve pre-learned knowledge. Comprehensive experiments on both synthetic and real-world datasets demonstrate superior adaptation performance of our proposed method over several state-of-the-art baselines.
Sangni Xu, Hao Xiong 0001, Qiuxia Wu, Zhihui Wang 0001, Zhiyong Wang 0001
ICRA3
2023 InvolutionGAN: lightweight GAN with involution for unsupervised image-to-image translation
Haipeng Deng, Qiuxia Wu, Han Huang 0002, Xiaowei Yang 0003, Zhiyong Wang 0001
Neural Comput. Appl.2
2023 Distinguishing and Matching-Aware Unsupervised Point Cloud Completion
abstract
Real-scanned point clouds are often incomplete due to occlusion, light reflection and limitations of sensor resolution, which impedes the related progress of downstream tasks, e.g., shape classification and object detection. Although there has been impressive research progress on the point cloud completion topic, they rely on the premise of extensive paired training data. However, collecting complete point clouds in some specified scenarios is labor-intensive and even impractical. To mitigate this problem, we propose DMNet, a distinguishing and matching-aware unsupervised point cloud completion network. Our work belongs to the group of unsupervised completion methods but goes beyond previous studies. Firstly, we propose a distinguishing-aware feature extractor to learn discriminable semantic information for different instances, simultaneously enhancing the robust invariant representation under noise disturbances. Secondly, we design a hierarchy-aware hyperbolic decoder to recover the complete geometry of point clouds, which not only can capture the implicit hierarchical relationships in data but also has an explicit extended nature. Finally, we develop a matching-aware refiner to eliminate noise points via aligning the topology structure of the input and predicted partial point clouds. Extensive experiments on MVP, Completion3D and KITTI datasets prove the effectiveness of our method, which performs favorably over state-of-the-art methods both quantitatively and qualitatively.
Haihong Xiao, Yuqiong Li, Wenxiong Kang, Qiuxia Wu
IEEE Trans. Circuits Syst. Video Technol.4
2023 Self-supervised monocular depth estimation via two mechanisms of attention-aware cost volume
Zhongcheng Hong, Qiuxia Wu
Vis. Comput.2
2022 Endowing rotation invariance for 3D finger shape and vein verification
Weili Yang, Qiuxia Wu, Wenxiong Kang
Frontiers Comput. Sci.3
2022 A Deep Clustering via Automatic Feature Embedded Learning for Human Activity Recognition
abstract
Traditional clustering algorithms are widely used for building bag-of-words (BOW) models to aggregate spatio-temporal feature points extracted from a video for human activity recognition problems. Their performances are restricted by the computational complexity which limits the number of feature points being used. In contrast, deep clustering yields good clustering performance without the limit of the number of feature points. Therefore, this work proposes a dual stacked autoencoders features embedded clustering (DSAFEC) and a BOW construction method based on the DSAFEC (B-DSAFEC) to reduce the computational complexity and to remove the selection restriction. The DSAFEC first transforms feature points extracted from a video to a learned feature space and then probabilities of cluster assignment of feature points are predicted to build BOWs for human activity recognition. A soft clustering is used by assigning each feature point to multiple clusters yielding the largest probabilities instead of only one in hard clustering. Experimental results on three benchmark human activity datasets show that the B-DSAFEC yields better performance compared to five reference methods which are developed based on either traditional clustering methods or deep clustering methods.
Ting Wang 0015, Wing W. Y. Ng, Jinde Li, Qiuxia Wu, Shuai Zhang 0001, Chris D. Nugent, Colin Shewell
IEEE Trans. Circuits Syst. Video Technol.4
2021 Self-supervised Multi-view Stereo via Effective Co-Segmentation and Data-Augmentation
abstract
Recent studies have witnessed that self-supervised methods based on view synthesis obtain clear progress on multi-view stereo (MVS). However, existing methods rely on the assumption that the corresponding points among different views share the same color, which may not always be true in practice. This may lead to unreliable self-supervised signal and harm the final reconstruction performance. To address the issue, we propose a framework integrated with more reliable supervision guided by semantic co-segmentation and data-augmentation. Specially, we excavate mutual semantic from multi-view images to guide the semantic consistency. And we devise effective data-augmentation mechanism which ensures the transformation robustness by treating the prediction of regular samples as pseudo ground truth to regularize the prediction of augmented samples. Experimental results on DTU dataset show that our proposed methods achieve the state-of-the-art performance among unsupervised methods, and even compete on par with supervised methods. Furthermore, extensive experiments on Tanks&Temples dataset demonstrate the effective generalization ability of the proposed method.
Yu Qiao 0001, Wenxiong Kang, Qiuxia Wu
AAAI5
2021 Learning Efficient Rotation Representation for Point Cloud via Local-Global Aggregation
abstract
Recently, there have been attempted to solve the problem of rotation perturbation in point cloud analysis. However, most of them fail to exploit the long-distance context and lose global location information. To address this issue, we propose a novel rotation-invariant network called LGANet, which is assembled with two key modules: local representation learning module and global alignment module. The local representation learning module is to capture local geometric features from K-nearest neighbors in both 3D Cartesian space and a latent space, while the global alignment module focuses on supplementing global location information with the adaptive selection mechanism. Extensive experiments on widely used datasets have demonstrated that our LGANet is superior to other state-of-the-art methods on the premise of ensuring rotation invariance in both classification and part segmentation.
Ruibin Gu, Qiuxia Wu, Wing W. Y. Ng, Zhiyong Wang 0001
ICME2
2021 PVLNet: Parameterized-View-Learning neural network for 3D shape recognition
Lvequan Wang, Qiuxia Wu, Wenxiong Kang
Comput. Graph.3
2021 Single-scale siamese network based RGB-D object tracking with adaptive bounding boxes
Qiuxia Wu, Han Huang 0002
Neurocomputing2
2021 ERINet: Enhanced rotation-invariant network for point cloud classification
Ruibin Gu, Qiuxia Wu, Wing W. Y. Ng, Zhiyong Wang 0001
Pattern Recognit. Lett.2
2020 Visual saliency based global-local feature representation for skin cancer classification
abstract
With the rapid increase in the cases of deadly skin cancer, the classification on different types of skin cancer has been emerging as one of the most significant issues in the field of medical image. Several approaches have been proposed to help in diagnosing the categories of the skin lesions by means of traditional features or leveraging the widely used deep learning models. However, there are lack of the integrated frameworks to combine the hand‐crafted traditional features and the deep Conv‐features. Furthermore, the effective way to extract global and local features is also conducive to distinguish the specific lesions from normal skin. Hence, in this study, the authors present an integrated model to acquire more representative global–local features including the traditional local binary pattern features and deep Conv‐features. In addition, several fusion strategies have conducted on the Global‐DNN and Local‐DNN for better performance. In order to extract more explicit features from the specific lesion areas, a target segmentation method based on visual saliency detection is employed to eliminate the background interference. Experimental results on ISIC‐2017 skin cancer dataset demonstrate that the proposed Global‐DNN and Global‐Local models can obtain more effective feature representation which achieve outperformed results for skin cancer classification.
Qiuxia Wu
IET Image Process.2
2019 Adaptive dual fractional-order variational optical flow model for motion estimation
abstract
Insufficient illumination and illumination variation in image sequences make it challenging for algorithms to obtain clear outlines for objects in motion. This study proposes a high‐performance adaptive dual fractional‐order variational optical flow model which could be used to resolve these issues. The proposed method revitalises the original dual fractional‐order optical flow model and adopts a fractional differential mask in both the data and smoothness terms of the traditional Horn–Schunck model. The main innovation of this work is to fit a flow field regional to a variety of fractional‐order differential masks. The domain of each region is determined adaptively. The order and size of the fractional‐order differential masks for each region are adjusted by image signal to noise ratio while the shape of the fractional‐order differential mask is regulated to prevent interference from surrounding regions. Adjusting the fractional‐order differential mask adaptively enables the proposed method to accurately segment motion objects in poor and variable illumination regions as well. The experimental results show that our algorithm outperforms the current state‐of‐the‐art algorithms on low‐light real scene videos and also achieves competitive results on the Middlebury, KITTI and MPI Sintel public benchmarks.
Bin Zhu 0009, Lianfang Tian, Qiliang Du, Qiuxia Wu, Farisi Zeyad Sahl, Yao Yeboah
IET Comput. Vis.4
2019 Feature covariance matrix-based dynamic hand gesture recognition
Linpu Fang, Guile Wu, Wenxiong Kang, Qiuxia Wu, Zhiyong Wang 0001, David Dagan Feng
Neural Comput. Appl.4
2018 A novel finger vein verification system based on two-stream convolutional network learning
Yuxun Fang, Qiuxia Wu, Wenxiong Kang
Neurocomputing2
2016 Real-time vehicle detection with foreground-based cascade classifier
abstract
The strategy based on Haar‐like features and the cascade classifier for vehicle detection systems has captured growing attention for its effectiveness and robustness; however, such a vehicle detection strategy relies on exhaustive scanning of an entire image with different sizes sliding windows, which is tedious and inefficient, since a vehicle only occupies a small part of the whole scene. Therefore, the authors propose a real‐time vehicle detection algorithm which is based on the improved Haar‐like features and combines motion detection with a cascade of classifiers. They adopt a visual background extractor, accompanied by morphological processing, to obtain foregrounds. These foregrounds retain vehicle features and provide the positions within images where vehicles are most likely to be located. Subsequently, vehicle detection is performed only at these positions by using a cascade of classifiers instead of a single strong classifier, which is able to improve the detection performance. The authors’ algorithm has been successfully evaluated on the public datasets, which demonstrates its robustness and real‐time performance.
Xiaobin Zhuang, Wenxiong Kang, Qiuxia Wu
IET Image Process.3
2015 Palm vein recognition based on multi-sampling and feature-level fusion
Xuekui Yan, Wenxiong Kang, Feiqi Deng, Qiuxia Wu
Neurocomputing4
2014 A new descriptor resistant to affine transformation and monotonic intensity change
Zeyi Huang, Wenxiong Kang, Qiuxia Wu
Comput. Vis. Image Underst.3
2014 Contactless Palm Vein Recognition Using a Mutual Foreground-Based Local Binary Pattern
abstract
Local binary pattern (LBP) is popular for the texture representation owing to its discrimination ability and computational efficiency, but when used to describe the sparse texture in palm vein images, the discrimination ability is diluted, leading to lower performance, especially for contactless palm vein matching. In this paper, an improved mutual foreground LBP method is presented for achieving a better matching performance for contactless palm vein recognition. First, the normalized gradient-based maximal principal curvature algorithm and k -means method are utilized for texture extraction, which can effectively suppress noise and improve accuracy and robustness. Then, an LBP matching strategy was adopted for similarity measurements on the basis of extracted palm veins and their neighborhoods, which include the vast majority of useful distinctive information for identification while eliminating interference by excluding the background. To further improve the LBP performance, the matched pixel ratio was adopted to determine the best matching region (BMR). Finally, the matching score obtained in the process of finding the BMR was fused with results of LBP matching at the score level to further improve the identification performance. A series of rigorous contrast experiments using the palm vein data set in the CASIA multispectral palmprint image database were conducted. The obtained low equal error rate (0.267%) and comparisons with the most state-of-the-art approaches demonstrate that our method is feasible and effective for contactless palm vein recognition.
Wenxiong Kang, Qiuxia Wu
IEEE Trans. Inf. Forensics Secur.2
2014 Pose-Invariant Hand Shape Recognition Based on Finger Geometry
abstract
In this paper, a pose-invariant hand shape recognition method based on the geometry of the fingers is proposed. Firstly, inspired by the segmentation method presented by Yoruk et al., we conduct a novel improvement on the segmentation for extracting the region of the fingers when the hand is in a natural pose. Secondly, Fourier descriptors and finger area functions are employed to extract the finger boundary curve features and region areas, respectively. Finally, score-level fusion based on a weighted sum is used to obtain matching results. Because the finger segmentation strategy and the feature extraction method are both rotation and translation invariant, the proposed method is more suitable for a naturally posed hand. Experiments using the Bogazici University Hand database show that the proposed method can achieve an equal error rate of 0.0369 for all data and 0.0273 for samples with an intragroup angle deviation of less than 45°. Thus, the proposed method is suitable for real-world applications.
Wenxiong Kang, Qiuxia Wu
IEEE Trans. Syst. Man Cybern. Syst.2
2013 Discriminative two-level feature selection for realistic human action recognition
Qiuxia Wu, Zhiyong Wang 0001, Feiqi Deng, Yong Xia 0001, Wenxiong Kang, David Dagan Feng
J. Vis. Commun. Image Represent.1
2013 Realistic Human Action Recognition With Multimodal Feature Selection and Fusion
abstract
Although promising results have been achieved for human action recognition under well-controlled conditions, it is very challenging to recognize human actions in realistic scenarios due to increased difficulties such as dynamic backgrounds. In this paper, we propose to take multimodal (i.e., audiovisual) characteristics of realistic human action videos into account in human action recognition for the first time, since, in realistic scenarios, audio signals accompanying an action generally provide a cue to the nature of the action, such as phone ringing to answering the phone . In order to cope with diverse audio cues of an action in realistic scenarios, we propose to identify effective features from a large number of audio features with the generalized multiple kernel learning algorithm. The widely used space-time interest point descriptors are utilized as visual features, and a support vector machine is employed for both audio- and video-based classifications. At the final stage, fuzzy integral is utilized to fuse recognition results of both audio and visual modalities. Experimental results on the challenging Hollywood-2 Human Action data set demonstrate that the proposed approach is able to achieve better recognition performance improvement than that of integrating scene context. It is also discovered how audio context influences realistic action recognition from our comprehensive experiments.
Qiuxia Wu, Zhiyong Wang 0001, Feiqi Deng, Zheru Chi, David Dagan Feng
IEEE Trans. Syst. Man Cybern. Syst.1