Qiuqi Ruan

dblp:57/4736 · DBLP profile ↗
← Back
80ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0001-8107-7365ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 9 since 2021Computer networks · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 GroupSD: Self-Distillation From Intermediate ViT Layers for Generalizable Person Re-Identification
Jieru Jia, Jianchao Yang, Chao Li 0070, Qiuqi Ruan
IEEE Trans. Circuits Syst. Video Technol.6
2025 EIRA: an explicit-implicit representation alignment for multimodal relation extraction
Gaoyun An, Zhaoqilin Yang, Xingyu Ren, Qiuqi Ruan
Multim. Syst.5
2025 Attention redirection transformer with semantic oriented learning for unbiased scene graph generation
Gaoyun An, Yi-Gang Cen, Qiuqi Ruan
Pattern Recognit.4
2024 GBC: Guided Alignment and Adaptive Boosting CLIP Bridging Vision and Language for Robust Action Recognition
abstract
The Contrastive Language-Image Pre-training (CLIP) model achieves strong generalization by using a large number of text-image pairs for contrastive learning. However, when it is transferred to action recognition, the following two questions remain to be solved: 1) How to guide the model to focus more on human-body-related regions to better align actions and text, and 2) How to make the model strengthen itself in a targeted manner to deal with difficult-to-classify categories. To solve these problems, a Guided alignment and adaptive Boosting CLIP (GBC) is proposed, which employs visual prior knowledge and benefits from both feature and decision aggregation in a boosting manner. During early training, visual prior knowledge related to human body is adopted, which enables the model to better align human actions with category text to be robust to distribution shift. At the later stage of training, the CLIP encoder is frozen, and multiple downstream feature & decision aggregation modules are sequentially generated and trained. In such way, the model is able to boost the performance from different perspectives in the Boosting manner and at a linearly increasing cost. Moreover, a class-adaptive re-weighting strategy is proposed to make the model focus more on optimizing categories that are difficult to classify. The effectiveness of our model is validated on six action recognition datasets (Kinetics-600, Kinetics-400, Jester, HMDB-51, UCF-101, and Mini-Kinetics-200), including both fully supervised and zero-shot experiments. Our model achieves superior results compared to state-of-the-art methods on all datasets.
Zhaoqilin Yang, Gaoyun An, Zhenxing Zheng, Shan Cao 0002, Qiuqi Ruan
IEEE Trans. Circuits Syst. Video Technol.5
2024 Hybrid Perturbation Strategy for Semi-Supervised Crowd Counting
abstract
A simple yet effective semi-supervised method is proposed in this paper based on consistency regularization for crowd counting, and a hybrid perturbation strategy is used to generate strong, diverse perturbations, and enhance unlabeled images information mining. The conventional CNN-based counting methods are sensitive to texture perturbation and imperceptible noises raised by adversarial attack, therefore, the hybrid strategy is proposed to combine a spatial texture transformation and an adversarial perturbation module to perturb the unlabeled data in the semantic and non-semantic spaces, respectively. Moreover, a cross-distribution normalization technique is introduced to address the model optimization failure caused by BN layer in the strong perturbation, and to stabilize the optimization of the learning model. Extensive experiments have been conducted on the datasets of ShanghaiTech, UCF-QNRF, NWPU-Crowd, and JHU-Crowd++. The results demonstrate that the proposed semi-supervised counting method performs better over the state-of-the-art methods, and it shows better robustness to various perturbations.
Xin Wang 0135, Yue Zhan, Yang Zhao 0014, Tangwen Yang, Qiuqi Ruan
IEEE Trans. Image Process.5
2024 TG-Pose: Delving Into Topology and Geometry for Category-Level Object Pose Estimation
abstract
Category-level 6D object pose estimation aims to estimate the pose and size of unseen objects with known categories. Existing methods mainly focus on capturing geometric features to handle shape variations, and are prone to failure in occlusion and noisy environments. In this paper, we propose TG-Pose, a unified pose estimation framework that delves into topology and geometry to deal with the above issues. To exploit topological properties, we first propose a topological feature predictor and a topological label generator to dig into the underlying structural details from encoded features using persistent homology. Then, the topological and geometric features are employed to facilitate the symmetry reconstruction of the original point cloud to obtain a reliable and coherent object shape, which, in turn, guides the pose estimation. For each object category, we construct geometric and topological templates by leveraging inherent intra-class similarities. These templates enhance the reliability of pose estimation and the completeness of object structure through geometric alignment and topological guidance, especially when handling incomplete objects. Moreover, a pose-aware enhancement strategy is designed to enhance the encoder in learning pose-sensitive features and robustness to noisy point clouds. Experimental results show that TG-Pose outperforms the state-of-the-art solutions on public benchmarks and achieves better generalization in real-world datasets. Project Page https://sites.google.com/view/tg-pose.
Yue Zhan, Xin Wang 0135, Lang Nie, Yang Zhao 0014, Tangwen Yang, Qiuqi Ruan
IEEE Trans. Multim.6
2023 PCR: A Large-Scale Benchmark for Pig Counting in Real World
Jieru Jia, Shuorui Zhang, Qiuqi Ruan
PRCV (4)3
2023 A dual-modal graph attention interaction network for person Re-identification
abstract
Abstract Person Re‐identification (Re‐ID) is a task of matching target pedestrians under cross‐camera surveillance. Learning discriminative feature representations is the main issue for person Re‐ID. A few recent methods introduce text descriptions as auxiliary information to enhance feature representations, as it offers richer semantic information and perspective consistency. However, these works usually process text and images separately, which leads to the absence of cross‐modal interactions. In this article, a Dual‐modal Graph Attention Interaction Network (Dual‐GAIN) is proposed to integrate visual features and textual features into a heterogeneous graph to model the relationship between them, simultaneously. The proposed Dual‐GAIN mainly consists of two components: a dual‐stream feature extractor and a Graph Attention Interaction Network (GAIN). Specifically, the two‐stream feature extractor is utilised to extract visual features and textual features respectively. Then, visual local features and textual features are treated as nodes to construct a multi‐modal graph. Cosine similarity constrained attention weights are introduced in GAIN, which is designed for cross‐modal interaction and feature fusion on this heterogeneous multi‐modal graph. Experiments on public large‐scale datasets, that is, Market‐1501, CUHK03 labelled, and CUHK03 detected, demonstrate our method achieves the state‐of‐the‐art performance.
Gaoyun An, Qiuqi Ruan
IET Comput. Vis.3
2023 SRI3D: Two-stream inflated 3D ConvNet based on sparse regularization for action recognition
abstract
Abstract Although most state‐of‐the‐art action recognition models have adopted a two‐stream 3D convolutional structure as a backbone network, few works have studied the impact of loss functions on action recognition models. In addition, sparsity is used as a key prior knowledge in many fields. However, as far as is known, no one has studied the influence of the sparsity of network output on the output of deep learning‐based action recognition models. Therefore, this paper proposes a novel two‐stream inflated 3D ConvNet based on the sparse regularization (SRI3D) model for action recognition. In order to allow the network to learn the sparsity of output, the ℓ 1 norm is embedded in the loss function in regularization form in a plug‐and‐play manner. It can make the classification result after the fusion of the two‐stream network only be the category with the highest confidence in one of the streams and not the other cases. The proposed loss function based on sparse regularization makes the output vector of the neural network as sparse as possible so that the classification results will not be ambiguous. Experimental results show that compared with other state‐of‐the‐art models, this SRI3D has a competitive advantage on Kinetics‐400, Something‐Something V2, UCF‐101 and HMDB‐51.
Zhaoqilin Yang, Gaoyun An, Zhenxing Zheng, Qiuqi Ruan
IET Image Process.5
2023 Semi-Supervised Crowd Counting With Spatial Temporal Consistency and Pseudo-Label Filter
abstract
Semi-supervised crowd counting (SSCC) aims to learn a crowd counting model with limited labeled images and a large number of unlabeled images. Previous works leverage unlabeled images by pseudo-labeling and spatial consistency regularization paradigms, which frequently adopt teacher-student frameworks. However, their performances are readily degraded due to the inconsistent and unreliable pseudo density map in complex crowd scenes. Here, we argue that the SSCC performance can be significantly improved by reducing the over-fitting of the incorrect pseudo labels, and a novel spatial-temporal consistency framework, named STC-Crowd, is proposed. Under different spatial perturbations, spatial consistency enables the counting model to output consistent predictions for the same crowd image. Temporal consistency generates similar feature embedding over adjacent training stages, alleviating the inconsistent issues of pseudo density maps generated for the same image over time. To store the temporal feature embedding of different density levels for temporal consistency, a dynamic temporal knowledge memory (DTKM) is deliberately designed, and considerably reduces the storage cost. Besides, a pseudo-label filter (PLF) mechanism is used to alleviate the negative impact of incorrect pseudo density maps, by reducing the supervision weights of unreliable pseudo labels with high uncertainty. Extensive experiments on four benchmark datasets show that our method obtains competent performance against leading SSCC methods, and especially works better on limited labeled images.
Xin Wang 0135, Yue Zhan, Yang Zhao 0014, Tangwen Yang, Qiuqi Ruan
IEEE Trans. Circuits Syst. Video Technol.5
2023 Collaborative and Multilevel Feature Selection Network for Action Recognition
abstract
The feature pyramid has been widely used in many visual tasks, such as fine-grained image classification, instance segmentation, and object detection, and had been achieving promising performance. Although many algorithms exploit different-level features to construct the feature pyramid, they usually treat them equally and do not make an in-depth investigation on the inherent complementary advantages of different-level features. In this article, to learn a pyramid feature with the robust representational ability for action recognition, we propose a novel collaborative and multilevel feature selection network (FSNet) that applies feature selection and aggregation on multilevel features according to action context. Unlike previous works that learn the pattern of frame appearance by enhancing spatial encoding, the proposed network consists of the position selection module and channel selection module that can adaptively aggregate multilevel features into a new informative feature from both position and channel dimensions. The position selection module integrates the vectors at the same spatial location across multilevel features with positionwise attention. Similarly, the channel selection module selectively aggregates the channel maps at the same channel location across multilevel features with channelwise attention. Positionwise features with different receptive fields and channelwise features with different pattern-specific responses are emphasized respectively depending on their correlations to actions, which are fused as a new informative feature for action recognition. The proposed FSNet can be inserted into different backbone networks flexibly, and extensive experiments are conducted on three benchmark action datasets, Kinetics, UCF101, and HMDB51. Experimental results show that FSNet is practical and can be collaboratively trained to boost the representational ability of existing networks. FSNet achieves superior performance against most top-tier models on Kinetics and all models on UCF101 and HMDB51.
Zhenxing Zheng, Gaoyun An, Shan Cao 0002, Dapeng Oliver Wu, Qiuqi Ruan
IEEE Trans. Neural Networks Learn. Syst.5
2022 PromptLearner-CLIP: Contrastive Multi-Modal Action Representation Learning with Context Optimization
Zhenxing Zheng, Gaoyun An, Shan Cao 0002, Zhaoqilin Yang, Qiuqi Ruan
ACCV (4)5
2021 Global and Local Knowledge-Aware Attention Network for Action Recognition
abstract
Convolutional neural networks (CNNs) have shown an effective way to learn spatiotemporal representation for action recognition in videos. However, most traditional action recognition algorithms do not employ the attention mechanism to focus on essential parts of video frames that are relevant to the action. In this article, we propose a novel global and local knowledge-aware attention network to address this challenge for action recognition. The proposed network incorporates two types of attention mechanism called statistic-based attention (SA) and learning-based attention (LA) to attach higher importance to the crucial elements in each video frame. As global pooling (GP) models capture global information, while attention models focus on the significant details to make full use of their implicit complementary advantages, our network adopts a three-stream architecture, including two attention streams and a GP stream. Each attention stream employs a fusion layer to combine global and local information and produces composite features. Furthermore, global-attention (GA) regularization is proposed to guide two attention streams to better model dynamics of composite features with the reference to the global information. Fusion at the softmax layer is adopted to make better use of the implicit complementary advantages between SA, LA, and GP streams and get the final comprehensive predictions. The proposed network is trained in an end-to-end fashion and learns efficient video-level features both spatially and temporally. Extensive experiments are conducted on three challenging benchmarks, Kinetics, HMDB51, and UCF101, and experimental results demonstrate that the proposed network outperforms most state-of-the-art methods.
Zhenxing Zheng, Gaoyun An, Dapeng Oliver Wu, Qiuqi Ruan
IEEE Trans. Neural Networks Learn. Syst.4
2020 Interactions Guided Generative Adversarial Network for unsupervised image captioning
Shan Cao 0002, Gaoyun An, Zhenxing Zheng, Qiuqi Ruan
Neurocomputing4
2020 View-specific subspace learning and re-ranking for semi-supervised person re-identification
Jieru Jia, Qiuqi Ruan, Yi Jin 0001, Gaoyun An, Shiming Ge
Pattern Recognit.2
2020 Aligned Dynamic-Preserving Embedding for Zero-Shot Action Recognition
abstract
Zero-shot learning (ZSL) typically explores a shared semantic space in order to recognize novel categories in the absence of any labeled training data. However, the traditional ZSL methods always suffer from serious domain shift problem in human action recognition. This is because: 1) existing ZSL methods are specifically designed for object recognition from static images, which do not capture the temporal dynamics of video sequences, and poor performances are always generated if those methods are directly applied to zero-shot action recognition; 2) these methods always blindly project the target data into a shared space using a semantic mapping obtained by the source data without any adaptation, in which the underlying structures of target data are ignored; and 3) severe inter-class variations exist in various action categories. The traditional ZSL methods do not take relationships across different categories into consideration. In this paper, we propose a novel aligned dynamic-preserving embedding (ADPE) model for zero-shot action recognition in a transductive setting. In our model, an adaptive embedding of target videos is learned, exploring the distributions of both the source and target data. An aligned regularization is further proposed to couple the centers of target semantic representations with their corresponding label prototypes in order to preserve the relationships across different categories. Most significantly, during our embedding, the temporal dynamics of video sequences are simultaneously preserved via exploiting the temporal consistency of video sequences and capturing the temporal evolution of successive segments of actions. Our model can effectively overcome the domain shift problem in zero-shot action recognition. The experiments on Olympic sports, HMDB51, and UCF101 datasets demonstrate the effectiveness of our model.
Yu Kong 0001, Qiuqi Ruan, Gaoyun An, Yun Fu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2019 Frustratingly Easy Person Re-Identification: Generalizing Person Re-ID in Practice
Jieru Jia, Qiuqi Ruan, Timothy M. Hospedales
BMVC2
2019 Residual Joint Attention Network with Graph Structure Inference for Object Detection
Chuansheng Xu, Gaoyun An, Qiuqi Ruan
ICIG (1)3
2019 Spatial-temporal pyramid based Convolutional Neural Network for action recognition
Zhenxing Zheng, Gaoyun An, Dapeng Oliver Wu, Qiuqi Ruan
Neurocomputing4
2019 FERLrTc: 2D+3D facial expression recognition via low-rank tensor completion
Yunfang Fu, Qiuqi Ruan, Ziyan Luo, Yi Jin 0001, Gaoyun An, Jun Wan 0001
Signal Process.2
2018 Hierarchical and Spatio-Temporal Sparse Representation for Human Action Recognition
abstract
In this paper, we present a novel two-layer video representation for human action recognition employing hierarchical group sparse encoding technique and spatio-temporal structure. In the first layer, a new sparse encoding method named locally consistent group sparse coding (LCGSC) is proposed to make full use of motion and appearance information of local features. LCGSC method not only encodes global layouts of features within the same video-level groups, but also captures local correlations between them, which obtains expressive sparse representations of video sequences. Meanwhile, two kinds of efficient location estimation models, namely an absolute location model and a relative location model, are developed to incorporate spatio-temporal structure into LCGSC representations. In the second layer, action-level group is established, where a hierarchical LCGSC encoding scheme is applied to describe videos at different levels of abstractions. On the one hand, the new layer captures higher order dependency between video sequences; on the other hand, it takes label information into consideration to improve discrimination of videos' representations. The superiorities of our hierarchical framework are demonstrated on several challenging datasets.
Yu Kong 0001, Qiuqi Ruan, Gaoyun An, Yun Fu 0001
IEEE Trans. Image Process.3
2017 Multiple metric learning with query adaptive weights and multi-task re-weighting for person re-identification
Jieru Jia, Qiuqi Ruan, Gaoyun An, Yi Jin 0001
Comput. Vis. Image Underst.2
2017 A sparse neighborhood preserving non-negative tensor factorization algorithm for facial expression recognition
Gaoyun An, Shuai Liu 0003, Qiuqi Ruan
Pattern Anal. Appl.3
2017 Multiview Hessian Semisupervised Sparse Feature Selection for Multimedia Analysis
abstract
Facing a large number of unlabeled data and a small number of labeled data, semisupervised sparse feature selection has received increasing attention in recent years. However, most semisupervised feature selection algorithms are developed for single-view data and cannot naturally handle multiview data. Moreover, most existing semisupervised sparse feature selection methods are based on Laplacian regularization, which is a lack of extrapolating power. To overcome the above-mentioned drawbacks, we present a multiview Hessian semi-supervised sparse feature selection (MHSFS) framework in this paper. MHSFS can directly accomplish multiview sparse feature selection by exploiting multiview learning to reveal and leverage the correlated and complemental information among different views. In addition, MHSFS can achieve better performance based on Hessian regularization, which favors functions whose values linearly vary with respect to geodesic distance and preserves the local manifold structure well. A simple yet efficient iterative method is proposed to solve the objective function, followed by convergence analysis. We apply the proposed method into different multimedia analysis tasks, such as image annotation, video concept detection, and 3D motion analysis. The results show that MHSFS outperforms the state-of-the-art sparse feature selection methods and achieves good performance.
Caijuan Shi, Gaoyun An, Ruizhen Zhao, Qiuqi Ruan, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.4
2016 Action Recognition Using Local Consistent Group Sparse Coding with Spatio-Temporal Structure
abstract
This paper presents a novel and efficient framework for human action recognition through integrating the local consistent group sparse representation with spatio-temporal structure of each video sequence. We firstly propose a sparse encoding scheme named local consistent group sparse coding (LCGSC) to generate the sparse representation of each video sequence. The novel encoding scheme takes global structural information of features belonging to one group into consideration as well as the local correlations between similar features. In order to incorporate the spatio-temporal structures, an average location (AL) model is proposed to describe the distribution of each visual word along the spatio-temporal coordinates on the basis of the obtained sparse codes. Eventually, each video sequence is jointly represented by the sparse representation and the spatio-temporal layouts which fully model its motion, appearance and spatio-temporal information. Our framework is computationally efficient and achieves comparable performance on the challenging datasets with state-of-the-art methods.
Qiuqi Ruan, Gaoyun An, Yun Fu 0001
ACM Multimedia2
2016 Geometric Preserving Local Fisher Discriminant Analysis for person re-identification
Jieru Jia, Qiuqi Ruan, Yi Jin 0001
Neurocomputing2
2016 Facial expression recognition using sparse local Fisher discriminant analysis
Qiuqi Ruan, Gaoyun An
Neurocomputing2
2015 Multiple strategies to enhance automatic 3D facial expression recognition
Xiaoli Li 0009, Qiuqi Ruan, Gaoyun An, Yi Jin 0001, Ruizhen Zhao
Neurocomputing2
2015 Sparse feature selection based on L2, 1/2-matrix norm for web image annotation
Caijuan Shi, Qiuqi Ruan
Neurocomputing2
2015 Context and locality constrained linear coding for human action recognition
Qiuqi Ruan, Gaoyun An, Wanru Xu
Neurocomputing2
2015 Semi-supervised sparse feature selection based on multi-view Laplacian regularization
Caijuan Shi, Qiuqi Ruan, Gaoyun An
Image Vis. Comput.2
2015 Projection-optimal local Fisher discriminant analysis for feature extraction
Qiuqi Ruan, Gaoyun An
Neural Comput. Appl.2
2015 Fully automatic 3D facial expression recognition using polytypic multi-block local binary patterns
Xiaoli Li 0009, Qiuqi Ruan, Yi Jin 0001, Gaoyun An, Ruizhen Zhao
Signal Process.2
2015 Coupled Discriminative Feature Learning for Heterogeneous Face Recognition
abstract
This paper presents a coupled discriminative feature learning (CDFL) method for heterogeneous face recognition (HFR). Different from most existing HFR approaches which use hand-crafted feature descriptors for face representation, our CDFL directly learns discriminative features from raw pixels for face representation. In particular, a couple of image filters is learned in CDFL to simultaneously exploit discriminative information and to reduce the appearance difference of face images captured across different modalities. With the help of the learned filters, CDFL can maximize the interclass variations and minimize the intraclass variations of the learned feature vectors, and meanwhile maximize the correlation of face images of the same person from different modalities by solving a generalized eigenvalue problem. Experimental results on three different heterogeneous face recognition applications show the effectiveness of our proposed approach.
Yi Jin 0001, Jiwen Lu, Qiuqi Ruan
IEEE Trans. Inf. Forensics Secur.3
2015 Hessian Semi-Supervised Sparse Feature Selection Based on ${L_{2, 1/2}}$ -Matrix Norm
abstract
Semi-supervised sparse feature selection, which can exploit the small number labeled data and large number unlabeled data simultaneously, has become an important technique in many applications on large-scale web image owing to its high efficiency and effectiveness. Recently, graph Laplacian-based semi-supervised sparse feature selection has obtained considerable attention, but it suffers with only few labeled data because Laplacian regularization is short of extrapolating power. In this paper we propose a novel semi-supervised sparse feature selection framework based on Hessian regularization and l2,1/2- matrix norm, namely Hessian sparse feature selection based on L2,1/2- matrix norm (HFSL). Hessian regularization favors functions whose values vary linearly with respect to geodesic distance and preserves the local manifold structure well, leading to good extrapolating power to boost semi-supervised learning, and then to enhance HFSL performance. The l2,1/2-matrix norm model makes HFSL select the most discriminative sparse features with good robustness. An efficient iterative algorithm is designed to optimize the objective function. We apply our algorithm into the image annotation task and conduct extensive experiments on two web image datasets. The results demonstrate that our algorithm outperforms state-of-the-art sparse feature selection methods and is promising for large-scale web image applications.
Caijuan Shi, Qiuqi Ruan, Gaoyun An, Ruizhen Zhao
IEEE Trans. Multim.2
2014 Complete discriminative feature learning: A new approach for heterogeneous face recognition
abstract
In this paper, we propose a new feature learning approach called complete discriminative feature learning (CDFL) for heterogeneous face recognition. Unlike most existing heterogeneous face recognition methods where hand-crafted feature descriptors are used for face representation, the proposed CD-FL aims to learn an optimal weighted discriminative image filter to improve learning discriminative filters, so that complete discriminative information is exploited and the feature difference between different modalities is effectively reduced, simultaneously. Experimental results shows that our approach consistently outperforms the state-of-the-art methods.
Yi Jin 0001, Jiwen Lu, Qiuqi Ruan, Yap-Peng Tan
ICME3
2014 Dimensionality reduction using graph-embedded probability-based semi-supervised discriminant analysis
Wei Li 0162, Qiuqi Ruan, Jun Wan 0001
Neurocomputing2
2014 Sparse feature selection based on graph Laplacian for web image annotation
Caijuan Shi, Qiuqi Ruan, Gaoyun An
Image Vis. Comput.2
2014 CSMMI: Class-Specific Maximization of Mutual Information for Action and Gesture Recognition
abstract
In this paper, we propose a novel approach called class-specific maximization of mutual information (CSMMI) using a submodular method, which aims at learning a compact and discriminative dictionary for each class. Unlike traditional dictionary-based algorithms, which typically learn a shared dictionary for all of the classes, we unify the intraclass and interclass mutual information (MI) into an single objective function to optimize class-specific dictionary. The objective function has two aims: 1) maximizing the MI between dictionary items within a specific class (intrinsic structure) and 2) minimizing the MI between the dictionary items in a given class and those of the other classes (extrinsic structure). We significantly reduce the computational complexity of CSMMI by introducing an novel submodular method, which is one of the important contributions of this paper. This paper also contributes a state-of-the-art end-to-end system for action and gesture recognition incorporating CSMMI, with feature extraction, learning initial dictionary per each class by sparse coding, CSMMI via submodularity, and classification based on reconstruction errors. We performed extensive experiments on synthetic data and eight benchmark data sets. Our experimental results show that CSMMI outperforms shared dictionary methods and that our end-to-end system is competitive with other state-of-the-art approaches.
Jun Wan 0001, Vassilis Athitsos, Pat Jangyodsuk, Hugo Jair Escalante, Qiuqi Ruan, Isabelle Guyon
IEEE Trans. Image Process.5
2013 Projection-optimal tensor local fisher discriminant analysis for image feature extraction
abstract
Tensor-based feature extraction approaches have been proved to be effective since they can solve the undersampled problem. In this paper, we propose a novel method called projection-optimal tensor local fisher discriminant analysis (PoTLFDA), which shares the character of local fisher discriminant analysis (LFDA). A novel affinity matrix is defined to effectively reflect the relationships of points in original tensor space and embedding space. The projection matrices are optimized by alternately solving the trace ratio problem. Convergence proof of the proposed algorithm is also given in this paper. Experiment results on face databases demonstrate the effectiveness of PoTLFDA.
Qiuqi Ruan, Zhenjiang Miao
ICIP2
2013 Fourier Spectral of PalmCode as Descriptor for Palmprint Recognition
Meiru Mu, Qiuqi Ruan, Luuk J. Spreeuwers, Raymond N. J. Veldhuis
ICPRAM2
2013 Enhancing sparsity via ℓp (0<p<1) minimization for robust face recognition
Qiuqi Ruan
Neurocomputing3
2013 Graph-preserving shortest feature line segment for dimensionality reduction
Wei Li 0162, Qiuqi Ruan, Jun Wan 0001
Neurocomputing2
2013 One-shot learning gesture recognition from RGB-D data using bag of features
Jun Wan 0001, Qiuqi Ruan, Wei Li 0162, Shuang Deng
J. Mach. Learn. Res.2
2013 A Mandarin edutainment system integrated virtual learning environments
Yue Ming 0001, Qiuqi Ruan, Guodong Gao
Speech Commun.2
2012 Activity Recognition from RGB-D Camera with 3D Local Spatio-temporal Features
abstract
Kinect, as a 3D digital capturing device, can collect the RGB and depth information of human activities rapidly. We study fusing the depth and RGB information for activity recognition. We introduce histogram color-based image thresholding to detect skin on human body, and use a GMM model to segment human hand areas. We design a new local descriptor, called a 3D Motion Scale-Invariant Feature Transform (3D MoSIFT), which can effectively detect interesting points based on both RGB and depth information, and consequently encode the visual and motion information from both to describe the interesting points. Experiments, based on a video dataset collected by a Kinect camera, show that adding depth information in the descriptor can distinctly improve the accuracy of human activity recognition. We introduce the F1-score measurement to evaluate and compare our performance with the other algorithms.
Yue Ming 0001, Qiuqi Ruan, Alex Hauptmann 0001
ICME2
2012 Similarity weighted sparse representation for classification
Qiuqi Ruan, Zhenjiang Miao
ICPR2
2012 A remarkable standard for estimating the performance of 3D facial expression features
Xiaoli Li 0009, Qiuqi Ruan, Yue Ming 0001
Neurocomputing2
2012 Orthogonal tensor rank one differential graph preserving projections with its application to facial expression recognition
Shuai Liu 0003, Qiuqi Ruan, Yi Jin 0001
Neurocomputing2
2012 Tensor rank one differential graph preserving analysis for facial expression recognition
Shuai Liu 0003, Qiuqi Ruan, Chuantao Wang, Gaoyun An
Image Vis. Comput.2
2012 Robust sparse bounding sphere for 3D face recognition
Yue Ming 0001, Qiuqi Ruan
Image Vis. Comput.2
2011 Shift and gray scale invariant features for palmprint identification using complex directional wavelet and local binary pattern
Meiru Mu, Qiuqi Ruan
Neurocomputing2
2011 Mean and Standard Deviation as Features for Palmprint Recognition Based on Gabor Filters
abstract
The two-dimensional (2D) Gabor function has been recognized as a very useful tool in feature extraction of image, due to its optimal localization properties in both spatial and frequency domain. This paper presents a novel palmprint feature extraction method based on the statistics of decomposition coefficients of the Gabor wavelet transform. It is experimentally found that the magnitude coefficients of the Gabor wavelet transform within each subband uniformly to approximate the Lognormal distribution. Based on this fact, we create the palmprint representation using two simple statistics (mean and standard deviation) as feature components after applying the logarithmic transformation of Gabor filtered magnitude coefficients for each subband with different orientations and scales. The optimum setting of the number of Gabor filters and orientation of each Gabor filter is experimentally determined. For palmprint recognition, the popularly used Fisher Linear Discriminant (FLD) analysis is further applied on the constructed feature vectors to extract discriminative features and reduce dimensionality. All experiments are both executed over the CCD-based HongKong PolyU Palmprint Database of 7752 images and the scanner-based BJTU_PalmprintDB (V1.0) of 3460 images. The results demonstrate the effectiveness of the proposed palmprint representation in achieving the improved recognition performance.
Meiru Mu, Qiuqi Ruan
Int. J. Pattern Recognit. Artif. Intell.2
2011 Region Covariance Matrices as Feature Descriptors for Palmprint Recognition Using Gabor Features
abstract
Region covariance matrices (RCMs) as feature descriptors have been developed due to the advantages of low dimensionality, being scale and illumination independent. How to define a feature mapping vector for the RCMs construction of strong discriminating ability is still an open issue. In this paper, there is a focus on finding a more efficient feature mapping vector for RCMs as palmprint descriptors based on Gabor magnitude and phase (GMP) information. Specially, Gabor magnitude (GM) features of each palmprint image approximate a lognormal distribution. For palmprint recognition, the logarithmic transformation of GM proves to be important for the discriminating ability of corresponding RCMs. All experiments are performed on the public Hong Kong Polytechnic University (PolyU) Palmprint Database of 7752 images. The results demonstrate the efficiency of our proposed method, and also show that adding pixel locations and intensity component to the feature mapping vector has a negative effect on palmprint recognition performance for our proposed Log_GMP based RCM method.
Meiru Mu, Qiuqi Ruan
Int. J. Pattern Recognit. Artif. Intell.2
2011 Orthogonal Tensor Neighborhood Preserving Embedding for facial expression recognition
Shuai Liu 0003, Qiuqi Ruan
Pattern Recognit.2
2011 Graph-Preserving Sparse Nonnegative Matrix Factorization With Application to Facial Expression Recognition
abstract
In this paper, a novel graph-preserving sparse nonnegative matrix factorization (GSNMF) algorithm is proposed for facial expression recognition. The GSNMF algorithm is derived from the original NMF algorithm by exploiting both sparse and graph-preserving properties. The latter may contain the class information of the samples. Therefore, GSNMF can be conducted as an unsupervised or a supervised dimension reduction method. A sparse representation of the facial images is obtained by minimizing the l(1)-norm of the basis images. Furthermore, according to the graph embedding theory, the neighborhood of the samples is preserved by retaining the graph structure in the mapped space. The GSNMF decomposition transforms the high-dimensional facial expression images into a locality-preserving subspace with sparse representation. To guarantee convergence, we use the projected gradient method to calculate the nonnegative solution of GSNMF. Experiments are conducted on the JAFFE database and the Cohn-Kanade database with unoccluded and partially occluded facial images. The results show that the GSNMF algorithm provides better facial representations and achieves higher recognition rates than nonnegative matrix factorization. Moreover, GSNMF is also more robust to partial occlusions than other tested methods.
Ruicong Zhi, Markus Flierl, Qiuqi Ruan, W. Bastiaan Kleijn
IEEE Trans. Syst. Man Cybern. Part B3
2010 Orthogonal Discriminant Neighborhood Preserving Embedding for facial expression recognition
abstract
In this paper, a new manifold learning algorithm called Orthogonal Discriminant Neighborhood Preserving Embedding (ODNPE) is proposed for facial expression recognition. The ODNPE pursues orthogonal projections vectors to preserve the local manifold within same classes and keep the separability between different classes. The obtained orthogonal projections vectors can keep the metric structure of the manifold embedded in high dimensional space such that the intrinsic dimensions of the manifold can be well learned. Furthermore, we design a novel penalty graph to describe the separability between pair-wise different classes. The proposed algorithm is compared with some other algorithms on two facial expression databases, and the experimental results show its effectivity.
Shuai Liu 0003, Qiuqi Ruan
ICIP2
2010 Learning effective features for 3D face recognition
abstract
3D images provide several advantages over 2D images for face recognition, especially when considering expression variations. In this paper, a novel framework is proposes for 3D-based face recognition. The key idea in the proposed algorithm is a representation of the facial surface, by what is called a Bending Invariant (BI), invariant to isometric deformations resulting from expressions and postures. In order to encode relationships in neighboring mesh nodes, Gaussian-Hermite moments are used for the obtained geometric invariant, which is a richer representation, due to their mathematical orthogonality and effectiveness in characterizing local details of the signal. The signature images are then decomposed into their principle components based on Spectral Regression Kernel Discriminate Analysis (SRKDA) resulting in a huge time saving. Our experiments are based on FRGC v2.0 face database. Experimental results show our framework provides better effectiveness and efficiency than many commonly used existing methods and handles variations in facial expression quite well.
Yue Ming 0001, Qiuqi Ruan
ICIP2
2010 An illumination normalization model for face recognition under varied lighting conditions
Gaoyun An, Jiying Wu, Qiuqi Ruan
Pattern Recognit. Lett.3
2009 Gait Recognition Using Procrustes Shape Analysis and Shape Context
Niqing Yang, Xiaojuan Wu, Qiuqi Ruan
ACCV (3)5
2009 Facial expression recognition based on graph-preserving sparse non-negative matrix factorization
abstract
In this paper, we present a novel algorithm for representing facial expressions. The algorithm is based on the non-negative matrix factorization (NMF) algorithm, which decomposes the original facial image matrix into two non-negative matrices, namely the coefficient matrix and the basis image matrix. We call the novel algorithm graph-preserving sparse non-negative matrix factorization (GSNMF). GSNMF utilizes both sparse and graph-preserving constraints to achieve a non-negative factorization. The graph-preserving criterion preserves the structure of the original facial images in the embedded subspace while considering the class information of the facial images. Therefore, GSNMF has more discriminant power than NMF. GSNMF is applied to facial images for the recognition of six basic facial expressions. Our experiments show that GSNMF achieves on average a recognition rate of 93.5% compared to that of discriminant NMF with 91.6%.
Ruicong Zhi, Markus Flierl, Qiuqi Ruan, W. Bastiaan Kleijn
ICIP3
2009 Discriminant sparse nonnegative matrix factorization
abstract
In this paper, a novel discriminant sparse non-negative matrix factorization (DSNMF) algorithm is proposed. We derive DSNMF method from original NMF algorithm by considering both sparseness constraint and discriminant information constraint. Furthermore, projected gradient method is used to solve the optimization problem. DSNMF makes use of prior class information which is important in classification, so it is a supervised method. Furthermore, by minimization l1-norm of the basis, we get a sparse representation of the facial images. Experiments are carried out for facial expression recognition. The experimental results obtained on Cohn-Kanade facial expression database indicate that DSNMF is efficient for facial expression recognition.
Ruicong Zhi, Qiuqi Ruan
ICME2
2009 Palmprint recognition using Gabor-based local invariant features
Qiuqi Ruan
Neurocomputing2
2009 Independent Gabor Analysis of Discriminant Features Fusion for Face Recognition
abstract
A discriminant feature fusion model is proposed for face recognition with large variations of pose, expression, lighting, etc. Discriminant features are extracted by the wavelet transform-based method from two source images. One source image is a holistic gray value image and the other is an illumination invariant geometric image. Face sample is reconstructed by the adaptive fused discriminant feature. Then a bank of Gabor filters is built to extract Gabor representations of the reconstructed samples. Finally higher-order statistical relationships among variables of samples are extracted for classifier. According to experiments, the model outperforms conventional algorithms under complex conditions (large variations of lighting, expression, accessory, etc.).
Jiying Wu, Gaoyun An, Qiuqi Ruan
IEEE Signal Process. Lett.3
2008 Gabor-based multi-scale Illumination Normalization model for face recognition
abstract
A novel Gabor-based multi-scale illumination normalization (GMSIN) model is proposed and applied to face recognition. GMSIN uses Total variation under different norm constraints. It removes the lighting effect in two scale parts of image and fuses the multi-scaled illumination invariant features. Then a bank of Gabor filters is built to extract lighting invariant Gabor face representations. Finally the higher-order statistical relationships among variables of samples are extracted for classifier. According to the experiments on the large scale CAS-PEAL face database, GMSIN could outperform conventional algorithms when they face most outliers (lighting, expression, masking etc.).
Jiying Wu, Gaoyun An, Qiuqi Ruan
ICIP3
2008 Discriminant spectral analysis for facial expression recognition
abstract
Spectral analysis is a recently proposed method for feature extraction. Studies show that the features extracted by spectral analysis can also be used to classification. In this paper, we propose a nonlinear feature extraction method called discriminant spectral analysis (DSA) algorithm for facial expression recognition. DSA takes both intra-locality and inter-locality structure of the data into account, and the features extracted by DSA have more discriminant power than traditional methods. Moreover, DSA is a nonlinear method which can effectively discover the intrinsic nonlinear manifold structure hidden in the data. Experimental results on Cohn-Kanade and JAFFE facial databases show the effectiveness of DSA algorithm.
Ruicong Zhi, Qiuqi Ruan
ICIP2
2008 Fuzzy discriminant projections for facial expression recognition
abstract
A linear projective map called fuzzy discriminant projections has been proposed in this paper. Fuzzy discriminant projection (FDP) is motivated by locality preserving projections which can optimally preserve the neighborhood structure of the data set. FDP utilizes the soft assignment method to weight pairs of samples with membership degree, and tries to find the optimal projective directions by maximizing the ratio of between-class distance against within-class distance. The resulting embedding subspace has more discriminant and robust power than that of traditional methods. Experiments on Cohn-Kanade databases show that FDP can effectively distinct the confusing facial expressions and obtain higher recognition accuracies than other subspacebased methods.
Ruicong Zhi, Qiuqi Ruan, Zhenjiang Miao
ICPR2
2008 Palmprint recognition using Gabor feature-based (2D)2PCA
Qiuqi Ruan
Neurocomputing2
2008 Facial expression recognition based on two-dimensional discriminant locality preserving projections
Ruicong Zhi, Qiuqi Ruan
Neurocomputing2
2008 Two-dimensional direct and weighted linear discriminant analysis for face recognition
Ruicong Zhi, Qiuqi Ruan
Neurocomputing2
2008 Fusing Global and Local Complete Linear Discriminant Features by Fuzzy Integral for Face Recognition
abstract
Face recognition becomes very difficult in a complex environment, and the combination of multiple classifiers is a good solution to this problem. A novel face recognition algorithm GLCFDA-FI is proposed in this paper, which fuses the complementary information extracted by complete linear discriminant analysis from the global and local features of a face to improve the performance. The Choquet fuzzy integral is used as the fusing tool due to its suitable properties for information aggregation. Experiments are carried out on the CAS-PEAL-R1 database, the Harvard database and the FERET database to demonstrate the effectiveness of the proposed method. Results also indicate that the proposed method GLCFDA-FI outperforms five other commonly used algorithms — namely, Fisherfaces, null space-based linear discriminant analysis (NLDA), cascaded-LDA, kernel-Fisher discriminant analysis (KFDA), and null-space based KFDA (NKFDA).
Qiuqi Ruan, Yi Jin 0001
Int. J. Pattern Recognit. Artif. Intell.2
2008 Palmprint recognition with improved two-dimensional locality preserving projections
Qiuqi Ruan
Image Vis. Comput.2
2008 Independent Gabor Analysis of Multiscale Total Variation-Based Quotient Image
abstract
A new algorithm for independent Gabor analysis of multiscale total variation-based quotient image is proposed and applied to face recognition with only one sample per subject here. With our preproposed multiscale TV-based quotient image (TVQI) model, the large-scale and small-scale features are firstly fused to produce the most expressive lighting invariant face. Then a bank of Gabor filters is built to extract lighting invariant Gabor face representations with specified scales and orientations. Last, an information maximization algorithm is adopted to extract higher-order statistical relationships among variables of samples for classifier. According to the experiments on the large-scale CAS-PEAL face database, our approach could outperform Gabor-based ICA, Gabor-based KPCA, and TVQI when they face most outliers (lighting, expression, masking, etc.).
Gaoyun An, Jiying Wu, Qiuqi Ruan
IEEE Signal Process. Lett.3
2007 A Novel Image Interpolation Method Based on Both Local and Global Information
Jiying Wu, Qiuqi Ruan, Gaoyun An
ICIC (1)2
2007 Gabor-Based Improved Locality Preserving Projections for Face Recognition
abstract
A novel Gabor-based improved locality preserving projections for face recognition is presented in this paper. This new algorithm is based on a combination of Gabor wavelets representation of face images and improved locality preserving projections for face recognition and it is robust to changes in illumination and facial expressions and poses. In this paper, Gabor filter is first designed to extract the features from the whole face images, and then a locality preserving projections, which is improved by two-directional 2DPCA to eliminate redundancy among Gabor features, is used to subject these feature vectors onto locality subspace projection. Experiments based on the ORL face database demonstrate the effectiveness and efficiency of the new method. Results show that our new algorithm outperforms the other popular approaches reported in the literature and achieves a much higher accurate recognition rate.
Yi Jin 0001, Qiuqi Ruan
ICIP (1)2
2007 An Improved 2DLPP Method on Gabor Features for Palmprint Recognition
abstract
We propose an improved 2DLPP method on Gabor features (I2DLPPG) for palmprint recognition in this paper. 2DPCA is first utilized for dimensionality reduction of Gabor feature space maintaining most prominent 2D information. Thus similarity matrix corresponding to elements is easily constructed and the followed 2DLPP can be implemented directly in the reduced feature space. The proposed method preserving more intrinsic manifold structure of feature matrices yields higher recognition accuracy than the existing 2DLPP which treats the Gabor feature matrices as a whole. Meanwhile, fewer coefficients are extracted for image representation and recognition owing to 2DLPP and 2DPCA in the row and column directions simultaneously. Euclidean distance and the nearest classifier are finally used for classification. The recognition accuracy of the proposed I2DLPPG can reach 99.5% with 15 x 5 features. Experiments results demonstrate the effectiveness of our proposed method in both recognition accuracy and speed.
Qiuqi Ruan
ICIP (2)2
2007 A feature-dependent fuzzy bidirectional flow for adaptive image sharpening
Shujun Fu, Qiuqi Ruan, Wenqia Wang, Fuzheng Gao, Heng-Da Cheng
Neurocomputing2
2006 A Novel Model for Gabor-Based Independent Radial Basis Function Neural Networks and Its Application to Face Recognition
Gaoyun An, Qiuqi Ruan
ICONIP (2)2
2005 A Scalable Overlay Multicast Congestion Control for Multimedia Streaming
abstract
This paper presents a scalable overlay multicast congestion control scheme (overlayTFMRC) for multimedia streaming over Internet. The scheme seeks to extend multicast scalability within our previous TCP-friendly unicast congestion control scheme streamTFRC using multiple time-scale prediction for multimedia QoS transport. Besides efficient QoS transport support based on our previous work, overlayTFMRC implements multicast scalability design into session rate control, feedback compression and data forward respectively by fully considering overlay network characteristics. Compared with TFMCC, simulation experiments illustrate that our scheme achieves not only better session scalability performance in terms of session throughput and its smoothness, but also better TCP-friendly performance in terms of intra-protocol & inter-protocol fairness and smoothness.
Qiuqi Ruan
LCN2
2005 Secure semi-blind watermarking based on iteration mapping and image features
Qiuqi Ruan, Heng-Da Cheng
Pattern Recognit.2