VLDB 2026 Research / reviewers in the wild / expert
Zhaoqiang Xia
dblp:122/2614
· DBLP profile ↗
51ranked-venue papers
10as first author
28since 2021 · last 2027
0000-0003-0630-3339ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 17 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Security and privacy · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | IPRec: A multimodal recommendation model with item-specific features and progressive knowledge distillation
Junmei Feng, Yaomin Zhao, Qiguang Miao, Zixiang Lu, Zhaoqiang Xia |
Expert Syst. Appl. | 6 |
| 2026 | Enhancing person-job fit through multi-temporal career trajectory modeling
Junmei Feng, Shuchun Li, Qiguang Miao, Zhaoqiang Xia |
Expert Syst. Appl. | 6 |
| 2026 | A comprehensive survey on contactless vital sign monitoring using vision-based, radio-based, and fusion approaches
Zichen Li, Xiaoting Wu, Constantino Álvarez Casado, Ville Lindholm, Kristina Mikkonen, Zhaoqiang Xia, Xiaoyi Feng, Miguel Bordallo López |
Neurocomputing | 6 |
| 2026 | HORNET: Fast and minimal adversarial perturbationsabstractFixed-budget attacks aim to generate adversarial examples—carefully crafted inputs designed to induce misclassifications during inference—while adhering to a predefined perturbation budget. These attacks maximize misclassification confidence and benefit from the transferability property, enabling the generated adversarial examples to remain effective even against multiple unknown models. However, to preserve their transferability, such attacks often yield perceptible perturbations, compromising the visual integrity of the adversarial examples. In this paper, we introduce HORNET, an extension of gradient-based fixed-budget attacks designed to minimize the perturbation magnitude of adversarial examples while maintaining their transferability against the target model. HORNET utilizes a distinct source model to craft the adversarial examples and employs a limited number of queries to the unknown target model to further minimize perturbation magnitude. We evaluate HORNET empirically by integrating it with 41 existing attack implementations and testing it against 9 different models, resulting in a total of 1700 unique configurations. Our results demonstrate that HORNET outperforms the state of the art in generating minimally perturbed yet highly transferable adversarial examples across all tested models. Code available at: https://github.com/louiswup/HORNET . Jiaping Wu, Antonio Emanuele Cinà, Francesco Villani, Zhaoqiang Xia, Luca Demetrio, Luca Oneto, Davide Anguita, Fabio Roli, Xiaoyi Feng |
Inf. Sci. | 4 |
| 2026 | Gender-independent kinship verification network via fuzzy disentangling and multi-metric inferenceabstractKinship verification aims to determine whether two individuals share a familial relationship based on facial information. Cross-gender relationships (i.e., Father-Daughter and Mother-Son) continue to face formidable challenges due to the diversity and uncertainty of genetic inheritance. Existing studies primarily focus on extracting robust features and measuring similarity, with limited attention given to the fuzziness of gender differences. To address this issue, this paper proposes a kinship verification framework based on a fuzzy neural network, which adaptively extracts gender-independent kinship features and handles relationship fuzziness to improve cross-gender verification performance. Specifically, the Swin Transformer, which has demonstrated excellent performance in facial analysis, is employed to extract initial features. A fuzzy neural network is then designed to disentangle gender and kinship features, with a gender recognition task introduced to further enhance this disentanglement and improve the gender independence of kinship features. Subsequently, a multi-metric fuzzy reasoning module is adopted to integrate kinship features, extract latent kinship cues, and leverage a contrastive loss function to effectively mine potential negative sample information, thereby significantly enhancing the model's robustness. Experimental results on three publicly available datasets demonstrate that the proposed method achieves state-of-the-art performance. Lei Li 0008, Shanshan Gao 0003, Chaoran Cui, Zhaoqiang Xia |
Neural Networks | 5 |
| 2026 | Hybrid-Supervised Hypergraph-Enhanced Transformer for Micro-Gesture Based Emotion RecognitionabstractMicro-gestures are unconsciously performed body gestures that can convey the emotion states of humans and start to attract more research attention in the fields of human behavior understanding and affective computing as an emerging topic. However, the modeling of human emotion based on micro-gestures has not been explored sufficiently. In this work, we propose to recognize the emotion states based on the micro-gestures by reconstructing behavioral patterns with a hypergraph-enhanced Transformer in a hybrid-supervised framework. In the framework, hypergraph Transformer based encoder and decoder are separately designed by stacking the hypergraph-enhanced self-attention and multiscale temporal convolution modules. Especially, to better capture the subtle motion of micro-gestures, we construct a decoder with additional upsampling operations for a reconstruction task in a self-supervised learning manner. We further propose a hypergraph-enhanced self-attention module where the hyperedges between skeleton joints are gradually updated to present the relationships of body joints for modeling the subtle local motion. Lastly, for exploiting the relationship between the emotion states and local motion of micro-gestures, an emotion recognition head from the output of encoder is designed with a shallow architecture and learned in a supervised way. The end-to-end framework is jointly trained in a one-stage way by comprehensively utilizing self-reconstruction and supervision information. The proposed method is evaluated on two publicly available datasets, namely iMiGUE and SMG, and achieves the best performance under multiple metrics, which is superior to the existing methods. The code is available on Github (https://github.com/xiazhaoqiang/H2OFormerMicroGestureRec). Zhaoqiang Xia, Haoyu Chen 0001, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2026 | Face Presentation Attack Detection by Exploiting Prior Knowledge of Region RelationshipsabstractFace recognition systems have been widely deployed in mobile devices for user authentication and payment applications. However, these biometric systems remain vulnerable to face presentation attacks, posing significant security risks. In recent years, numerous countermeasures have been proposed, with analysis of differences between bona fide and attack presentations being a commonly adopted strategy. Nevertheless, the variations in image attributes and region movements have not been thoroughly explored. In attack images, the textures of the facial region and the background tend to be more similar, while local regions often exhibit more consistent directions of movement compared to bona fide presentations. Motivated by this observation, we propose a novel face presentation attack detection method that leverages prior knowledge of region relationships. Specifically, each input face image sequence is first divided into small patches, which are then processed by a pre-trained$TimeSformer$network utilizing divided time and space attention mechanisms to extract deep features. Two metrics—$Cosine$similarity and mean squared error ($MSE$)—are subsequently employed to measure the texture similarity and movement relationships of the regions of interest. During the inference phase, these measurements are fused to distinguish bona fide from attack presentations. Extensive ablation and comparison experiments, conducted on six face presentation attack detection (PAD) databases (i.e., Idiap Replay-Attack, CASIA-MFSD, OULU-NPU, MSU-MFSD, 3DMAD, and HKBU-MARs V1+), demonstrate that our method achieves superior detection performance, significantly improving precision over state-of-the-art approaches in most experimental settings. Lei Li 0008, Shanshan Gao 0003, Zhaoqiang Xia, Fabio Roli, Yuanfeng Zhou |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | HandS3C: 3D Hand Mesh Reconstruction with State Space Spatial Channel Attention from RGB imagesabstractReconstructing the hand mesh from one single RGB image is a challenging task because hands are often occluded by other objects. Most previous works attempt to explore more additional information and adopt attention mechanisms for improving 3D reconstruction performance, while it would increase computational complexity simultaneously. To achieve a performance preserving architecture with high computational efficiency, in this work, we propose a simple but effective 3D hand mesh reconstruction network (i.e., HandS3C), which is the first time to incorporate state space model into the task of hand mesh reconstruction. In the network, we design a novel state-space spatial-channel attention module that extends the effective receptive field, extracts hand features in the spatial dimension, and enhances regional features of hands in the channel dimension. This helps to reconstruct a complete and detailed hand mesh. Extensive experiments conducted on well-known datasets facing heavy occlusions (such as FREIHAND, DEXYCB, and HO3D) demonstrate that our proposed HandS3C achieves state-of-the-art performance while maintaining minimal parameters. Code can be available in https://github.com/JiaoZixun/HandS3C. Zixun Jiao, Xihan Wang, Zhaoqiang Xia, Lianhe Shao, Quanli Gao |
ICASSP | 3 |
| 2025 | CI3Former: A Cross-Image Information Interaction Network for Kinship VerificationabstractKinship verification using facial information determines whether two faces share a familial relationship. Existing methods improve verification by leveraging negative sample information and addressing distribution differences but often extract independent features from parent and child images separately, ignoring variations in pairwise similarity. To overcome this, we propose CI3Former, a Swin-Transformer-based model that enables cross-image information interaction for joint feature extraction. By incorporating a Self-Attention based Interaction (SAI) module within each Swin-Transformer block, our method allows mutual querying between parent and child features, dynamically guiding region-level feature extraction and adaptively focusing on similar regions. Additionally, we introduce a Multi-metric Similarity based Interaction (MSI) module for feature fusion, which processes paired features through similarity measurements before final prediction. The model is trained with contrastive and binary cross-entropy losses to enhance coupled feature learning. Extensive experiments on four kinship verification datasets and a signature verification dataset demonstrate that CI3Former outperforms state-of-the-art methods, showcasing its effectiveness, robustness, and strong cross-task generalization. Lei Li 0008, Dong Huang 0003, Zhaoqiang Xia |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Face anti-spoofing via jointly modeling local texture and constructed depth
Lei Li 0008, Zhihao Yao 0007, Shanshan Gao 0003, Huijian Han, Zhaoqiang Xia |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | DeepFake detection based on high-frequency enhancement network for highly compressed content
Zhaoqiang Xia, Gian Luca Marcialis, Chen Dang, Xiaoyi Feng |
Expert Syst. Appl. | 2 |
| 2024 | SCPAD: An approach to explore optical characteristics for robust static presentation attack detection
Chen Dang, Zhaoqiang Xia, Lei Li 0008, Xiaoyi Feng |
Multim. Tools Appl. | 2 |
| 2024 | HandFormer: Hand pose reconstructing from a single RGB image
Zixun Jiao, Xihan Wang, Jingcao Li, Rongxin Gao, Jiao Liang, Zhaoqiang Xia, Quanli Gao |
Pattern Recognit. Lett. | 7 |
| 2024 | Information-Enhanced Network for Noncontact Heart Rate Estimation From Facial VideosabstractRemote photoplethysmography (rPPG) is a vital way of measuring heart rate (HR) to reflect human physical and mental health, which is useful for diagnosing cardiovascular and neurological diseases. Many non-contact HR estimation methods have been proposed gradually in recent years, but the majority of approaches are based on a single-modal HR information source, resulting in ineffective and unsatisfactory estimation results due to noise and insufficient information. This paper proposes a novel information-enhanced network for HR estimation based on multimodal (e.g., RGB and NIR) sources to address these problems. In the network, context and modal difference information are sequentially enhanced from spatiotemporal and modal views for accurately describing HR-aware features, while maximum frequency information is enhanced for inhibiting heartbeat noise. Specifically, a context-enhanced video Swin-Transformer (CET) module is exploited to extract useful rPPG signal features from facial visible-light and near-infrared videos. Then, a novel modal difference enhanced fusion (MDEF) module is designed to acquire a fused rPPG signal, which is taken as the input of the frequency-enhanced estimation (FEE) module to obtain the corresponding HR value. These three modules are integrated and jointly learned in an end-to-end way, and the multimodal combinations can provide highly complementary information for estimating HR value. Experimental and evaluation results on three multimodal datasets show that the proposed model achieves a superior effect compared to the state-of-the-art methods. Zhaoqiang Xia, Xiaobiao Zhang, Jinye Peng 0001, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Wooden spoon crack detection by prior knowledge-enriched deep convolutional network
Lei Li 0008, Huijian Han, Xiaoyi Feng, Fabio Roli, Zhaoqiang Xia |
Eng. Appl. Artif. Intell. | 7 |
| 2023 | Personalized recommendation with hybrid feedback by refining implicit data
Junmei Feng, Kunwei Wang, Qiguang Miao, Zhaoqiang Xia |
Expert Syst. Appl. | 5 |
| 2023 | Pain estimation with integrating global-wise and region-wise convolutional networksabstractAbstract Pain is a common phenomenon in clinical patients, which indicates patients are suffering from uncomfortable conditions for necessary treatments. So the assessment of pain status becomes a significant task in current medical institutions. Of late, various conventional hand‐crafted or deep learning methods on face images are presented to estimate pain intensity automatically. However, these approaches usually feed the whole face into the automatic estimation system and explore little information on the interdependencies of related regions during the formation of pain expression. In this paper, a hierarchical deep network (HDN) involving regional and holistic information simultaneously is proposed via two scale branches. In HDN, a region‐wise branch is designed to extract features from pain related regions of face images while a global‐wise branch explores the interdependencies of pain related regions. Besides, in global‐wise branch, a multi‐task learning method is employed to detect action units while estimating pain intensity. Finally, the pain estimation outputs of two branches are fused in a decision level. On current pain estimation benchmarks, it is empirically shown that the proposed HDN outperforms the existing methods and the essential components in HDN have key influences on final prediction. Dong Huang 0003, Zhaoqiang Xia, Lei Li 0008, Yupeng Ma |
IET Image Process. | 2 |
| 2023 | Hardening RGB-D object recognition systems against adversarial patch attacks
Luca Demetrio, Antonio Emanuele Cinà, Xiaoyi Feng, Zhaoqiang Xia, Xiaoyue Jiang, Ambra Demontis, Battista Biggio, Fabio Roli |
Inf. Sci. | 5 |
| 2023 | Why adversarial reprogramming works, when it fails, and how to tell the differenceabstractAdversarial reprogramming allows repurposing a machine-learning model to perform a different task. For example, a model trained to recognize animals can be reprogrammed to recognize digits by embedding an adversarial program in the digit images provided as input. Recent work has shown that adversarial reprogramming may not only be used to abuse machine-learning models provided as a service, but also beneficially, to improve transfer learning when training data is scarce. However, the factors affecting its success are still largely unexplained. In this work, we develop a first-order linear model of adversarial reprogramming to show that its success inherently depends on the size of the average input gradient, which grows when input gradients are more aligned, and when inputs have higher dimensionality. The results of our experimental analysis, involving fourteen distinct reprogramming tasks, show that the above factors are correlated with the success and the failure of adversarial reprogramming. Xiaoyi Feng, Zhaoqiang Xia, Xiaoyue Jiang, Ambra Demontis, Maura Pintor, Battista Biggio, Fabio Roli |
Inf. Sci. | 3 |
| 2023 | Stateful detection of adversarial reprogrammingabstractAdversarial reprogramming allows stealing computational resources by repurposing machine learning models to perform a different task chosen by the attacker. For example, a model trained to recognize images of animals can be reprogrammed to recognize medical images by embedding an adversarial program in the images provided as inputs. This attack can be perpetrated even if the target model is a black box, supposed that the machine-learning model is provided as a service and the attacker can query the model and collect its outputs. So far, no defense has been demonstrated effective in this scenario. We show for the first time that this attack is detectable using stateful defenses, which store the queries made to the classifier and detect the abnormal cases in which they are similar. Once a malicious query is detected, the account of the user who made it can be blocked. Thus, the attacker must create many accounts to perpetrate the attack. To decrease this number, the attacker could create the adversarial program against a surrogate classifier and then fine-tune it by making a few queries to the target model. In this scenario, the effectiveness of the stateful defense is reduced, but we show that it is still effective. Xiaoyi Feng, Zhaoqiang Xia, Xiaoyue Jiang, Maura Pintor, Ambra Demontis, Battista Biggio, Fabio Roli |
Inf. Sci. | 3 |
| 2023 | Hybrid attention network and center-guided non-maximum suppression for occluded face detection
Mingxin Jin, Huifang Li 0004, Zhaoqiang Xia |
Multim. Tools Appl. | 3 |
| 2023 | Micro-expression spotting with multi-scale local transformer in long videos
Xupeng Guo, Xiaobiao Zhang, Lei Li 0008, Zhaoqiang Xia |
Pattern Recognit. Lett. | 4 |
| 2023 | Demodulation Based Transformer for rPPG Generation and Heart Rate EstimationabstractAs a convenient, efficient and inexpensive medical technology, heart rate (HR) estimation from face videos has gradually become a hot topic and many methods have been exploited to learn a spatial-temporal map for HR estimation. However, these methods have poor performance under time-varying ambient lighting. The main reason is that the current methods neglect the optical modeling of extracting the contactless physiological signal from skin. In this study, we describe the signal extraction as a modulation process and draw the conclusion that light changing could produce amplitude modulation jamming after modeling analysis. Therefore, a demodulation-based Transformer is newly designed for rPPG signal purification. In addition, a Pwelch-based Softmax operation is incorporated for HR estimation to improve accuracy. Finally, the hybrid loss combined with the negative Pearson correlation coefficient and cross-entropy loss is introduced for entire network learning. The experimental results on two databases (COHFACE and PURE) are performed to verify the effectiveness of the proposed method. Xiaobiao Zhang, Zhaoqiang Xia, Xiaoyi Feng |
IEEE Signal Process. Lett. | 2 |
| 2023 | Non-Local Color Compensation Network for Intrinsic Image DecompositionabstractSingle image-based intrinsic image decomposition attempts to separate one input image into several intrinsic components, which is inherently an under-constrained problem. Some recent works have been proposed to estimate the intrinsic components using encoder-decoder structures. However, they generally lack exploration of the different component-oriented feature constraints and feature selection processes. In this paper, a non-local color compensation network (NCCNet) is proposed. Firstly, the hue and value channels of HSV color space are used as the complementary information for RGB images for the estimation of albedo and shading, respectively. The color space representation serves as an external constraint, which does not require expensive sensors or complicated computations. Secondly, an integrated non-local attention scheme is proposed to describe the relations of non-adjacent regions with a lower computational complexity compared to traditional methods. Then the non-local and local attention are combined to describe correlations among features and used as feature selectors between the encoder and decoder. Thirdly, the mutual constraint between albedo and shading is also explored in the network to further optimize the process. In order to train the network, a unified mutual exclusion loss function is proposed. Extensive experiments are conducted on several popular datasets, and the proposed NCCNet achieves improved performance with comparable computational cost compared to competing methods. Xiaoyue Jiang, Zhaoqiang Xia, Moncef Gabbouj, Jinye Peng 0001, Xiaoyi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Inclusive Consistency-Based Quantitative Decision-Making Framework for Incremental Automatic Target RecognitionabstractWhen new unknown samples are captured continually in the open-world environment, the concept diversity accumulation of existing classes and the identification/creation of new concept classes should be considered simultaneously. Since the initial training set of existent classes may be under-prepared, adhering to immediate decisions will inevitably lead to reduced open set recognition performance and higher costs of labeling/updating. Inspired by quantitative indicators in predictive reliability assessment and semi-supervised/active learning, the inclusive-consistency-based quantitative decision-making framework (ICQdm) is proposed for incremental automatic target recognition (ATR) to evaluate the identifiability and typicality of new unknown samples, which could give the decision-making guide of recognition and updating. For recognition decision-making, the first consistency indicator calculates the reliability of the unknown sample being included by one specific training class. The test samples with low reliability should be the wrongly classified samples and new unknown classes’ samples, which are difficult to be labeled by recognition models themselves. For updating decision-making, the second consistency indicator is designed to be the sample distribution density under the inclusive constraint, which could highlight the dense sample distributions of new unknown samples outside the known training distribution. Experiments verify that the proposed ICQdm outperforms other comparison methods on the open set recognition reliability evaluation and labeling/updating efficiency. Sihang Dang, Zhaoqiang Xia, Xiaoyue Jiang, Shuliang Gui, Xiaoyi Feng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Spatio-Temporal Pain Estimation Network With Measuring Pseudo Heart Rate GainabstractPain is a significant indicator that shows people are suffering from an unwell experience and its automatic estimation has attracted much interest in recent years. Of late, most estimation methods are designed to capture the dynamic pain information from visual signals while a few physiological-signal based methods can provide extra potential cues to analyze the pain more accurately. However, it is still challenging to capture the physiological data from patients as it requires contact devices and patients’ cooperation. In this paper, we propose to leverage the pseudo physiological information by generating new modal data from the original visual videos and jointly estimating the pain by an end-to-end network. To extract the representations from bi-modal data, we design a spatio-temporal pain estimation network, which employs a dual-branch framework for extracting pain-aware visual and pseudo physiological features separately and fuses the features in a probabilistic way. The inherent vital sign, i.e., heart rate gain (HRG), from pseudo physiological information can be utilized as an auxiliary signal and integrated with the visual pain estimation framework. Moreover, specially-designed 3D convolution filters and attention structures are employed to extract spatio-temporal features for both branches. To use the HRG as an auxiliary way for pain estimation, we propose a probabilistic inference model by jointly considering the visual branch and physiological branch, which makes our model estimate the pain comprehensively. Experiments on two publicly-available datasets show the effectiveness of introducing the pseudo modality, and the proposed method can outperform the state-of-the-art methods. Dong Huang 0003, Xiaoyi Feng, Haixi Zhang, Zitong Yu, Jinye Peng 0001, Guoying Zhao 0001, Zhaoqiang Xia |
IEEE Trans. Multim. | 7 |
| 2021 | RBPR: A hybrid model for the new user cold start problem in recommender systems
Junmei Feng, Zhaoqiang Xia, Xiaoyi Feng, Jinye Peng 0001 |
Knowl. Based Syst. | 2 |
| 2021 | CasQNet: Intrinsic Image Decomposition Based on Cascaded Quotient NetworkabstractIntrinsic image analysis plays an important role for image understanding, since it can provide accurate reflectance, shape and illumination information of the scene. However, intrinsic image analysis is an ill-posed problem which need to apply extra constrains for the decomposition of reflectance image and shading image from a single image. Recently deep neural networks are introduced for intrinsic image analysis, which can produce two intrinsic components simultaneously. In fact, the mutually exclusive relationship between reflectance image and shading image is not only a constraint for decomposition but also can improve the decomposition results. However, this relationship is always omitted in the current networks. In order to address this problem, we propose a novel deep network called as Cascaded Quotient Network (CasQNet) for intrinsic image decomposition. The CasQNet consists of two sub-networks: a Pyramid Mini-U-Net (PyNet) that specifically extracts the reflectance image in multi-scale and a Shading Optimization Network (SoNet) that optimizes the resulting shading. These two sub-networks are cascaded by a quotient operation, which directly enforces the mutually exclusive relationship between reflectance image and shading image in the network architecture. In PyNet, the task of reconstructing reflectance image is achieved by a series of nested multi-scale U-Nets, which simplified the learning task for each U-Net. SoNet is designed to address the unsmooth and blur problems of extreme points caused by the quotient operation. PyNet and SoNet are trained alternately and finally jointed in cascaded structure. Furthermore, we combine multiple loss functions, which consist of data loss, correlation loss and reconstruction loss, for improving the learning effectiveness. To evaluate our proposed algorithm, extensive experiments are performed on three datasets, i.e., ShapeNet, BOLD Surface and MIT Intrinsic Image datasets. Qualitative and quantitative results show that our model achieves the best performance compared to the state-of-the-art methods. Yupeng Ma, Xiaoyue Jiang, Zhaoqiang Xia, Moncef Gabbouj, Xiaoyi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Mix Dimension in Poincaré Geometry for 3D Skeleton-based Action RecognitionabstractGraph Convolutional Networks (GCNs) have already demonstrated their powerful ability to model the irregular data, e.g., skeletal data in human action recognition, providing an exciting new way to fuse rich structural information for nodes residing in different parts of a graph. In human action recognition, current works introduce a dynamic graph generation mechanism to better capture the underlying semantic skeleton connections and thus improves the performance. In this paper, we provide an orthogonal way to explore the underlying connections. Instead of introducing an expensive dynamic graph generation paradigm, we build a more efficient GCN on a Riemann manifold, which we think is a more suitable space to model the graph data, to make the extracted representations fit the embedding matrix. Specifically, we present a novel spatial-temporal GCN (ST-GCN) architecture which is defined via the Poincaré geometry such that it is able to better model the latent anatomy of the structure data. To further explore the optimal projection dimension in the Riemann space, we mix different dimensions on the manifold and provide an efficient way to explore the dimension for each ST-GCN layer. With the final resulted architecture, we evaluate our method on two current largest scale 3D datasets, i.e., NTU RGB+D and NTU RGB+D 120. The comparison results show that the model could achieve a superior performance under any given evaluation metrics with only 40% model size when compared with the previous best GCN method, which proves the effectiveness of our model. Wei Peng 0009, Jingang Shi, Zhaoqiang Xia, Guoying Zhao 0001 |
ACM Multimedia | 3 |
| 2020 | Infrared and visible image fusion using a shallow CNN and structural similarity constraintabstractIn recent years, image fusion methods based on deep networks have been proposed to combine infrared and visible images for achieving better fusion image. However, issues such as limited training data, scarce reference images and misalignment of multi‐source images, still limit the fusion performance. To address these problems, we propose an end‐to‐end shallow convolutional neural network with structural constraints, which has only one convolutional layer to fuse infrared and visible images. Different from other methods, our proposed model requires less training data and reference images and is more robust to the misalignment of a couple of images. More specifically, the infrared image and the visible image are first provided as inputs to a convolutional layer to extract the information that should be fused; then, all feature maps are concatenated together and fed into a convolutional layer with one channel to obtain the fused image; finally, a structural similarity loss between the fused image and the input infrared and visible images is computed to update the network parameters and eliminate the effects of pixel misalignment. Extensive experiments show the effectiveness of our proposed method on fusion of infrared and visible images with the performance that outperforms the state‐of‐the‐art methods. Lei Li 0008, Zhaoqiang Xia, Huijian Han, Guiqing He, Fabio Roli, Xiaoyi Feng |
IET Image Process. | 2 |
| 2020 | CompactNet: learning a compact space for face presentation attack detection
Lei Li 0008, Zhaoqiang Xia, Xiaoyue Jiang, Fabio Roli, Xiaoyi Feng |
Neurocomputing | 2 |
| 2020 | Pain-attentive network: a deep spatio-temporal attention model for pain estimation
Dong Huang 0003, Zhaoqiang Xia, Joshua Mwesigye, Xiaoyi Feng |
Multim. Tools Appl. | 2 |
| 2020 | Revealing the Invisible With Model and Data Shrinking for Composite-Database Micro-Expression RecognitionabstractComposite-database micro-expression recognition is attracting increasing attention as it is more practical for real-world applications. Though the composite database provides more sample diversity for learning good representation models, the important subtle dynamics are prone to disappearing in the domain shift such that the models greatly degrade their performance, especially for deep models. In this paper, we analyze the influence of learning complexity, including input complexity and model complexity, and discover that the lower-resolution input data and shallower-architecture model are helpful to ease the degradation of deep models in composite-database task. Based on this, we propose a recurrent convolutional network (RCN) to explore the shallower-architecture and lower-resolution input data, shrinking model and input complexities simultaneously. Furthermore, we develop three parameter-free modules (i.e., wide expansion, shortcut connection and attention unit) to integrate with RCN without increasing any learnable parameters. These three modules can enhance the representation ability in various perspectives while preserving not-very-deep architecture for lower-resolution data. Besides, three modules can further be combined by an automatic strategy (a neural architecture search strategy) and the searched architecture becomes more robust. Extensive experiments on the MEGC2019 dataset (composited of existing SMIC, CASME II and SAMM datasets) have verified the influence of learning complexity and shown that RCNs with three modules and the searched combination outperform the state-of-the-art approaches. Zhaoqiang Xia, Wei Peng 0009, Huai-Qian Khor, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Spatiotemporal Recurrent Convolutional Networks for Recognizing Spontaneous Micro-ExpressionsabstractRecently, the recognition task of spontaneous facial micro-expressions has attracted much attention with its various real-world applications. Plenty of handcrafted or learned features have been employed for a variety of classifiers and achieved promising performances for recognizing micro-expressions. However, the micro-expression recognition is still challenging due to the subtle spatiotemporal changes of micro-expressions. To exploit the merits of deep learning, we propose a novel deep recurrent convolutional networks based micro-expression recognition approach, capturing the spatiotemporal deformations of micro-expression sequence. Specifically, the proposed deep model is constituted of several recurrent convolutional layers for extracting visual features and a classificatory layer for recognition. It is optimized by an end-to-end manner and obviates manual feature design. To handle sequential data, we exploit two ways to extend the connectivity of convolutional networks across temporal domain, in which the spatiotemporal deformations are modeled in views of facial appearance and geometry separately. Besides, to overcome the shortcomings of limited and imbalanced training samples, two temporal data augmentation strategies as well as a balanced loss are jointly used for our deep network. By performing the experiments on three spontaneous micro-expression datasets, we verify the effectiveness of our proposed micro-expression recognition approach compared to the state-of-the-art methods. Zhaoqiang Xia, Xiaopeng Hong, Xingyu Gao 0001, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Corrections to "Spatiotemporal Recurrent Convolutional Networks for Recognizing Spontaneous Micro-Expressions"abstractPresents corrections to the author's information in the above named paper. Zhaoqiang Xia, Xiaopeng Hong, Xingyu Gao 0001, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | A Hybrid Pan-Sharpening Approach Using Nonnegative Matrix Factorization for WorldView Imageries
Guiqing He, Jiaqi Ji, Zhaoqiang Xia |
PRCV (2) | 4 |
| 2019 | Scene video text tracking based on hybrid deep text detection and layout constraint
Xihan Wang, Xiaoyi Feng, Zhaoqiang Xia |
Neurocomputing | 3 |
| 2019 | Trademark image retrieval via transformation-invariant deep hashing
Zhaoqiang Xia, Jie Lin 0001, Xiaoyi Feng |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | Replayed Video Attack Detection Based on Motion Blur AnalysisabstractFace presentation attacks are the main threats to face recognition systems, and many presentation attack detection (PAD) methods have been proposed in recent years. Although these methods have achieved significant performance in some specific intrusion modes, difficulties still exist in addressing replayed video attacks. That is because the replayed fake faces contain a variety of aliveness signals, such as eye blinking and facial expression changes. Replayed video attacks occur when attackers try to invade biometric systems by presenting face videos in front of the cameras, and these videos are often launched by a liquid-crystal display (LCD) screen. Due to the smearing effects and movements of LCD, videos captured from the real and replayed fake faces present different motion blurs, which are reflected mainly in blur intensity variation and blur width. Based on these descriptions, a motion blur analysis-based method is proposed to deal with the replayed video attack problem. We first present a 1D convolutional neural network (CNN) for motion blur intensity variation description in the time domain, which consists of a serial of 1D convolutional and pooling filters. Then, a local similar pattern (LSP) feature is introduced to extract blur width. Finally, features extracted from 1D CNN and LSP are fused to detect the replayed video attacks. Extensive experiments on two standard face PAD databases, i.e., relay-attack and OULU-NPU, indicate that our proposed method based on the motion blur analysis significantly outperforms the state-of-the-art methods and shows excellent generalization capability. Lei Li 0008, Zhaoqiang Xia, Abdenour Hadid, Xiaoyue Jiang, Haixi Zhang, Xiaoyi Feng |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | Face spoofing detection with local binary pattern network
Lei Li 0008, Xiaoyi Feng, Zhaoqiang Xia, Xiaoyue Jiang, Abdenour Hadid |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | No-reference image quality assessment with center-surround based natural scene statistics
Jun Wu 0022, Zhaoqiang Xia, Huifang Li 0004, Kezheng Sun, Ke Gu 0001, Hong Lu 0008 |
Multim. Tools Appl. | 2 |
| 2018 | Panchromatic and multi-spectral image fusion for new satellites based on multi-channel deep model
Guiqing He, Zhaoqiang Xia, Jianping Fan 0001 |
Mach. Vis. Appl. | 3 |
| 2017 | Similar Trademark Image Retrieval Integrating LBP and Convolutional Neural Network
Xiaoyi Feng, Zhaoqiang Xia, Shijie Pan, Jinye Peng 0001 |
ICIG (3) | 3 |
| 2017 | Intrinsic Image Decomposition: A Comprehensive Review
Yupeng Ma, Xiaoyi Feng, Xiaoyue Jiang, Zhaoqiang Xia, Jinye Peng 0001 |
ICIG (1) | 4 |
| 2017 | Multi-orientation Scene Text Detection Leveraging Background Suppression
Xihan Wang, Xiaoyi Feng, Zhaoqiang Xia, Jinye Peng 0001, Eric Granger |
ICIG (1) | 3 |
| 2017 | Face anti-spoofing via deep local binary patternsabstractConvolutional neural networks (CNNs) have achieved excellent performance in the field of pattern recognition when huge amount of training data is available. However, training a CNN model is less obvious when only a limited amount of data is given such as in the case of face anti-spoofing problem. It is indeed not easy to collect very large sets of fake faces. Especially for the fully-connected layers, tens of thousands of parameters need to be learned. To tackle this problem of lack of training data in face anti-spoofing, we propose to explore the incorporation of hand-crafted features in the CNN framework. In our proposed approach, the color local binary patterns (LBP) features are extracted from the convolutional feature maps, which are fine tuned based on the VGG-face model. These features are then fed into support vector machine (SVM) classifier. Extensive experiments are conducted on two benchmark and publicly available databases showing very interesting performance compared to state-of-the-art methods. Lei Li 0008, Xiaoyi Feng, Xiaoyue Jiang, Zhaoqiang Xia, Abdenour Hadid |
ICIP | 4 |
| 2017 | Deep convolutional hashing using pairwise multi-label supervision for large-scale visual search
Zhaoqiang Xia, Xiaoyi Feng, Jie Lin 0001, Abdenour Hadid |
Signal Process. Image Commun. | 1 |
| 2016 | Spontaneous micro-expression spotting via geometric deformation modeling
Zhaoqiang Xia, Xiaoyi Feng, Jinye Peng 0001, Xianlin Peng, Guoying Zhao 0001 |
Comput. Vis. Image Underst. | 1 |
| 2015 | A regularized optimization framework for tag completion and image retrieval
Zhaoqiang Xia, Xiaoyi Feng, Jinye Peng 0001, Jun Wu 0022, Jianping Fan 0001 |
Neurocomputing | 1 |
| 2015 | Automatic tag-to-region assignment via multiple instance learning
Zhaoqiang Xia, Yi Shen 0005, Xiaoyi Feng, Jinye Peng 0001, Jianping Fan 0001 |
Multim. Tools Appl. | 1 |
| 2013 | Multiple Instance Learning for Automatic Image Annotation
Zhaoqiang Xia, Jinye Peng 0001, Xiaoyi Feng, Jianping Fan 0001 |
MMM (2) | 1 |