Ruicong Zhi

dblp:61/3350 · DBLP profile ↗
← Back
31ranked-venue papers
14as first author
20since 2021 · last 2026
0000-0002-6913-4248ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 BiPG-FER: Bi-intelligence probabilistic graph for facial expression inference drived by action units
Ruicong Zhi
Comput. Vis. Image Underst.2
2026 A weak-alignment multi-sensor fusion method for object detection using visible, thermal infrared, and radar data
Xiaoyuan Liang, Ruicong Zhi
Eng. Appl. Artif. Intell.2
2026 A diluted confounding factors model applied to causal link prediction
Yuewu Hou, Ruicong Zhi, Junxiong Zeng
Neurocomputing2
2026 Tensor decomposition based lightweight residual fusion for Multimodal sentiment understanding
Ruicong Zhi
J. Intell. Inf. Syst.2
2026 DMIC2: A dynamic modality importance and cascaded cross-attention framework for multimodal sentiment analysis
Jinming Ping, Ruicong Zhi, Shufan Guo, Yuewu Hou, Xiaoyuan Liang
Knowl. Based Syst.2
2026 DGPDL: Domain-Guided Prompt Distribution Learning for Generalizable Face Anti-Spoofing
abstract
The overfitting of domain signals results in poor domain generalization of face anti-spoofing. The current methods usually improve the diversity of source domains to alleviate this overfitting. However, this benefit is minimal, as even the most diverse domain signals will also be absent in the target domain. In this work, we propose a Domain-Guided Prompt Distribution Learning (DGPDL) built on Vision-Language Models like CLIP, which explores a unified representation of domain signals as a prompt across the source and target domain to alleviate the understanding bias caused by domain gaps. Specifically, we first define a learnable Domain-Specific Distribution (DSD) that covers as many domain elements as possible, such as image quality, color tone, camera settings, etc., which establish connections between different domains and linearly combinable prompt in any domain; Then, based on the style statistics of the given sample, we construct its optimal Domain-Specific Prompts (DSPs) from the defined DSD through the designed Prompt Assemble Attention (PAA) with the similarity matching; Finally, the assembled DSPs will act as carrier or agent to perform on both the vision and language branches, synergistically improving the model's recognition of domain signals. By using the prompt to represent domain signals uniformly, if the model can be robust to DSPs in the source domain, it should be applicable to target domain, as they share the same DSD. By representing domain signals as prompts rather than instantiation features, DGPDL effectively reduces the reliance on specific domain appearances. This design enables the model to dynamically adapt to unseen target domains without the need for retraining. Extensive experiments show that the DGPDL is effective and outperforms the state-of-the-art methods on several cross-domain benchmarks.
Ajian Liu 0001, Xun Lin, Ruicong Zhi, Yanyan Liang 0001, Xinshan Zhu, Zhanchuan Cai, Jun Wan 0001, Sergio Escalera, Zhen Lei 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 AU-EMO Correlation based zero-shot facial expression recognition with graph convolutional network
Ruicong Zhi, Jun Wan 0001
Pattern Recognit.2
2025 PMMTalk$:$ Speech-Driven 3D Facial Animation From Complementary Pseudo Multi-Modal Features
abstract
Speech-driven 3D facial animation has improved a lot recently while most related works only utilize acoustic modality and neglect the influence of visual and textual cues, leading to unsatisfactory results in terms of precision and coherence. We argue that visual and textual cues are not trivial information. Therefore, we present a novel framework, namely PMMTalk, using complementaryPseudoMulti-Modal features for improving the accuracy of facial animation. The framework entails three modules: PMMTalk encoder, cross-modal alignment module, and PMMTalk decoder. Specifically, the PMMTalk encoder employs the off-the-shelf talking head generation architecture and speech recognition technology to extract visual and textual information from speech, respectively. Following this, the cross-modal alignment module aligns the audio-image-text features at temporal and semantic levels. Subsequently, the PMMTalk decoder is employed to predict lip-syncing facial blendshape coefficients. Contrary to prior methods, PMMTalk only requires an additional random reference face image but yields more accurate results. Additionally, it is artist-friendly as it seamlessly integrates into standard animation production workflows by introducing facial blendshape coefficients. Finally, given the scarcity of 3D talking face datasets, we introduce a large-scale3DChineseAudio-VisualFacialAnimation (3D-CAVFA) dataset. Extensive experiments and user studies show that our approach outperforms the state of the art. Codes and datasets are available at PMMTalk.
Tianshun Han, Shengnan Gui, Baihui Li, Lijian Liu, Benjia Zhou, Ruicong Zhi, Yanyan Liang 0001, Jun Wan 0001
IEEE Trans. Multim.9
2025 Dual Balanced Class-Incremental Learning With im-Softmax and Angular Rectification
abstract
Owing to the superior performances, exemplar-based methods with knowledge distillation (KD) are widely applied in class incremental learning (CIL). However, it suffers from two drawbacks: 1) data imbalance between the old/learned and new classes causes the bias of the new classifier toward the head/new classes and 2) deep neural networks (DNNs) suffer from distribution drift when learning sequence tasks, which results in narrowed feature space and deficient representation of old tasks. For the first problem, we analyze the insufficiency of softmax loss when dealing with the problem of data imbalance in theory and then propose the imbalance softmax (im-softmax) loss to relieve the imbalanced data learning, where we re-scale the output logits to underfit the head/new classes. For another problem, we calibrate the feature space by incremental-adaptive angular margin (IAAM) loss. The new classes form a complete distribution in feature space yet the old are squeezed. To recover the old feature space, we first compute the included angle of normalized features and normalized anchor prototypes, and use the angle distribution to represent the class distribution, then we replenish the old distribution with the deviation from the new. Each anchor prototype is predefined as a learnable vector for a designated class. The proposed im-softmax reduces the bias in the linear classification layer. IAAM rectifies the representation learning, reduces the intra-class distance, and enlarges the inter-class margin. Finally, we seamlessly combine the im-softmax and IAAM in an end-to-end training framework, called the dual balanced class incremental learning (DBL), for further improvements. Experiments demonstrate the proposed method achieves state-of-the-art (SOTA) performances on several benchmarks, such as CIFAR10, CIFAR100, Tiny-ImageNet, and ImageNet-100.
Ruicong Zhi, Yicheng Meng, Junyi Hou, Jun Wan 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Intelligent Navigation System That Gives Trajectory Guidance in 3D Scenes
Ruicong Zhi
PRCV (10)1
2024 WDTtrack: tracking multiple objects with indistinguishable appearance and irregular motion
Zeyong Zhao, Ruicong Zhi
Appl. Intell.3
2024 MambaFormerSR: A Lightweight Model for Remote-Sensing Image Super-Resolution
abstract
In recent years, deep learning-based method have shown impressive performance in super-resolution (SR) reconstruction tasks for remote-sensing images. However, many existing methods require significant computational resources, posing a challenge for edge devices with limited computing capabilities. To address this issue, we introduce Mamba, an emerging novel state space model (SSM), which possesses a global receptive field and linear complexity, and apply it to remote-sensing SR tasks. Furthermore, to enhance global modeling capabilities, we integrate Mamba and Transformer to develop a lightweight remote-sensing SR model named MambaFormerSR. First, we design a state space and attention fusion module (SAFM) to capture long-range spatial dependencies and fully extract image features. Subsequently, we introduce an improved feed-forward network module, named convolutional Fourier transform feed-forward network (CTFFN), which utilizes convolutions to incorporate local information and employs a frequency-based attention module to enable the model to focus on informative frequency components, further refining the features. Finally, we add an auxiliary reconstruction branch to reconstruct coarse-grained image information, enabling the network to concentrate on extracting fine-grained image details and reconstructing high-quality images. Extensive experiments have demonstrated the superiority of our approach, with our model achieving a PSNR value of 30.28 dB on the UCMerced dataset for$\times 3$SR with 306k parameters.
Ruicong Zhi, Xiaopei Fan, Jingye Shi
IEEE Geosci. Remote. Sens. Lett.1
2024 A Multitask Network and Two Large-Scale Datasets for Change Detection and Captioning in Remote Sensing Images
abstract
Remote sensing change detection (RSCD) recognizes pixel-level change regions between images, while remote sensing change captioning (RSCC) describes the nature and properties of these changes in natural language. The past have studied two tasks individually. Resulting in a single interpretation content produced and limited application scenarios. And the complementarity of change regions and deep semantic information can further improve the performance and robustness of the model. Therefore, we try simultaneously to solve both RSCD and RSCC tasks (i.e., RS-CDC) under a multitask framework, namely a CNN-Transformer-based multitask network (CTMTNet). Specifically, we design multiattention feature enhancement module (MAFEM) and feature fusion block (FFB) to enhance local information and location perception of features from bitemporal images. The MAFEM weights the channel and space separately to capture local information more accurately and enhance location perception. The FFB fuses bitemporal features and uses multilevel residual connections to ensure that change information is not lost during transfer. Finally, we use two decoders to output the change maps (CMs) and change captioning, respectively. During training, we use an improved multitask loss function for CTMTNet to balance the two tasks. For exploring the RS-CDC task, we construct two large-scale datasets named LEVIR-CDC and WHU-CDC dataset. We benchmark the existing state-of-the-art (SOTA) change detection (CD) and change captioning methods on these two datasets and a newly publicized LEVIR-MCI dataset, and the results show that the proposed CTMTNet significantly outperforms comparative methods.
Jingye Shi, Mengge Zhang, Yuewu Hou, Ruicong Zhi, Jiqiang Liu
IEEE Trans. Geosci. Remote. Sens.4
2024 A Double-Head Global Reasoning Network for Object Detection of Remote Sensing Images
abstract
Object detection of remote sensing images is a fundamental and important task for a wide range of applications, such as agriculture, military, and geological exploration. However, it is still a challenging task due to complex background, heavy occlusion, and object ambiguities. Complex remote sensing scenes contain rich visual and spatial relationships between proposals, while most existing methods ignore these relationships. In this article, we introduce a novel double-head global reasoning network (DGRN), which endows an object detection model with the ability to adapt global reasoning by propagating visual and spatial embeddings between positive proposals (foreground) and negative proposals (background). Instead of propagating the features of positive proposals, we evolve high-level embeddings globally to explore potential information in all proposals. Specifically, our method first classifies and locates positive proposals and negative proposals simultaneously, and then adaptively learns two sparse region-to-region undirected graphs for classification and location, respectively. Finally, the graph reasoning module (GRM) conducts the propagation of proposals’ embedding to improve the performance of object classification and localization. Without any prior knowledge, our DGRN method explores reasoning graphs between proposals for object classification and location, respectively, and then uses graph convolutional neural network (GCN) to propagate information and achieve relational reasoning. Our method is light-weighted and flexible enough to enhance many object detection models, making object detection models the capability of relational reasoning. The experimental results on DOTA and DIOR showed that the proposed method had better detection performance than object detection networks without considering relationships.
Jingye Shi, Ruicong Zhi, Jingru Zhao, Yuewu Hou, Zeyong Zhao, Jiqiang Liu
IEEE Trans. Geosci. Remote. Sens.2
2023 Cross-dataset face analysis based on multi-task learning
Caixia Zhou, Ruicong Zhi
Appl. Intell.2
2023 Drop-relationship learning for semi-supervised facial action unit recognition
Ruicong Zhi, Caixia Zhou
Neurocomputing2
2022 Gaussian distribution-based facial expression feature extraction network
Ruicong Zhi
Pattern Recognit. Lett.2
2022 Micro-expression recognition with supervised contrastive learning
Ruicong Zhi
Pattern Recognit. Lett.1
2021 Binary Convolutional Neural Networks for Facial Action Unit Detection
Ruicong Zhi
ICIG (2)3
2021 Action unit analysis enhanced facial expression recognition by deep neural network evolution
Ruicong Zhi, Caixia Zhou, Shuai Liu 0003, Yi Jin 0001
Neurocomputing1
2020 A comprehensive survey on automatic facial action unit analysis
Ruicong Zhi
Vis. Comput.1
2018 Translation and scale invariants of Krawtchouk moments
Ruicong Zhi, Lianyu Cao, Gang Cao 0001
Inf. Process. Lett.1
2018 Efficient Group-n Encoding and Decoding for Facial Age Estimation
abstract
Different ages are closely related especially among the adjacent ages because aging is a slow and extremely non-stationary process with much randomness. To explore the relationship between the real age and its adjacent ages, an age group-n encoding (AGEn) method is proposed in this paper. In our model, adjacent ages are grouped into the same group and each age corresponds to n groups. The ages grouped into the same group would be regarded as an independent class in the training stage. On this basis, the original age estimation problem can be transformed into a series of binary classification sub-problems. And a deep Convolutional Neural Networks (CNN) with multiple classifiers is designed to cope with such sub-problems. Later, a Local Age Decoding (LAD) strategy is further presented to accelerate the prediction process, which locally decodes the estimated age value from ordinal classifiers. Besides, to alleviate the imbalance data learning problem of each classifier, a penalty factor is inserted into the unified objective function to favor the minority class. To compare with state-of-the-art methods, we evaluate the proposed method on FG-NET, MORPH II, CACD and Chalearn LAP 2015 databases and it achieves the best performance.
Zichang Tan, Jun Wan 0001, Zhen Lei 0001, Ruicong Zhi, Guodong Guo, Stan Z. Li
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Ensemble based extreme learning machine for cross-modality face matching
Yi Jin 0001, Jiuwen Cao, Ruicong Zhi
Multim. Tools Appl.4
2011 Graph-Preserving Sparse Nonnegative Matrix Factorization With Application to Facial Expression Recognition
abstract
In this paper, a novel graph-preserving sparse nonnegative matrix factorization (GSNMF) algorithm is proposed for facial expression recognition. The GSNMF algorithm is derived from the original NMF algorithm by exploiting both sparse and graph-preserving properties. The latter may contain the class information of the samples. Therefore, GSNMF can be conducted as an unsupervised or a supervised dimension reduction method. A sparse representation of the facial images is obtained by minimizing the l(1)-norm of the basis images. Furthermore, according to the graph embedding theory, the neighborhood of the samples is preserved by retaining the graph structure in the mapped space. The GSNMF decomposition transforms the high-dimensional facial expression images into a locality-preserving subspace with sparse representation. To guarantee convergence, we use the projected gradient method to calculate the nonnegative solution of GSNMF. Experiments are conducted on the JAFFE database and the Cohn-Kanade database with unoccluded and partially occluded facial images. The results show that the GSNMF algorithm provides better facial representations and achieves higher recognition rates than nonnegative matrix factorization. Moreover, GSNMF is also more robust to partial occlusions than other tested methods.
Ruicong Zhi, Markus Flierl, Qiuqi Ruan, W. Bastiaan Kleijn
IEEE Trans. Syst. Man Cybern. Part B1
2009 Facial expression recognition based on graph-preserving sparse non-negative matrix factorization
abstract
In this paper, we present a novel algorithm for representing facial expressions. The algorithm is based on the non-negative matrix factorization (NMF) algorithm, which decomposes the original facial image matrix into two non-negative matrices, namely the coefficient matrix and the basis image matrix. We call the novel algorithm graph-preserving sparse non-negative matrix factorization (GSNMF). GSNMF utilizes both sparse and graph-preserving constraints to achieve a non-negative factorization. The graph-preserving criterion preserves the structure of the original facial images in the embedded subspace while considering the class information of the facial images. Therefore, GSNMF has more discriminant power than NMF. GSNMF is applied to facial images for the recognition of six basic facial expressions. Our experiments show that GSNMF achieves on average a recognition rate of 93.5% compared to that of discriminant NMF with 91.6%.
Ruicong Zhi, Markus Flierl, Qiuqi Ruan, W. Bastiaan Kleijn
ICIP1
2009 Discriminant sparse nonnegative matrix factorization
abstract
In this paper, a novel discriminant sparse non-negative matrix factorization (DSNMF) algorithm is proposed. We derive DSNMF method from original NMF algorithm by considering both sparseness constraint and discriminant information constraint. Furthermore, projected gradient method is used to solve the optimization problem. DSNMF makes use of prior class information which is important in classification, so it is a supervised method. Furthermore, by minimization l1-norm of the basis, we get a sparse representation of the facial images. Experiments are carried out for facial expression recognition. The experimental results obtained on Cohn-Kanade facial expression database indicate that DSNMF is efficient for facial expression recognition.
Ruicong Zhi, Qiuqi Ruan
ICME1
2008 Discriminant spectral analysis for facial expression recognition
abstract
Spectral analysis is a recently proposed method for feature extraction. Studies show that the features extracted by spectral analysis can also be used to classification. In this paper, we propose a nonlinear feature extraction method called discriminant spectral analysis (DSA) algorithm for facial expression recognition. DSA takes both intra-locality and inter-locality structure of the data into account, and the features extracted by DSA have more discriminant power than traditional methods. Moreover, DSA is a nonlinear method which can effectively discover the intrinsic nonlinear manifold structure hidden in the data. Experimental results on Cohn-Kanade and JAFFE facial databases show the effectiveness of DSA algorithm.
Ruicong Zhi, Qiuqi Ruan
ICIP1
2008 Fuzzy discriminant projections for facial expression recognition
abstract
A linear projective map called fuzzy discriminant projections has been proposed in this paper. Fuzzy discriminant projection (FDP) is motivated by locality preserving projections which can optimally preserve the neighborhood structure of the data set. FDP utilizes the soft assignment method to weight pairs of samples with membership degree, and tries to find the optimal projective directions by maximizing the ratio of between-class distance against within-class distance. The resulting embedding subspace has more discriminant and robust power than that of traditional methods. Experiments on Cohn-Kanade databases show that FDP can effectively distinct the confusing facial expressions and obtain higher recognition accuracies than other subspacebased methods.
Ruicong Zhi, Qiuqi Ruan, Zhenjiang Miao
ICPR1
2008 Facial expression recognition based on two-dimensional discriminant locality preserving projections
Ruicong Zhi, Qiuqi Ruan
Neurocomputing1
2008 Two-dimensional direct and weighted linear discriminant analysis for face recognition
Ruicong Zhi, Qiuqi Ruan
Neurocomputing1