VLDB 2026 Research / reviewers in the wild / expert
Xiaoyi Feng
dblp:04/3473
· DBLP profile ↗
75ranked-venue papers
7as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 25 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Security and privacy · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A comprehensive survey on contactless vital sign monitoring using vision-based, radio-based, and fusion approaches
Zichen Li, Xiaoting Wu, Constantino Álvarez Casado, Ville Lindholm, Kristina Mikkonen, Zhaoqiang Xia, Xiaoyi Feng, Miguel Bordallo López |
Neurocomputing | 7 |
| 2026 | HORNET: Fast and minimal adversarial perturbationsabstractFixed-budget attacks aim to generate adversarial examples—carefully crafted inputs designed to induce misclassifications during inference—while adhering to a predefined perturbation budget. These attacks maximize misclassification confidence and benefit from the transferability property, enabling the generated adversarial examples to remain effective even against multiple unknown models. However, to preserve their transferability, such attacks often yield perceptible perturbations, compromising the visual integrity of the adversarial examples. In this paper, we introduce HORNET, an extension of gradient-based fixed-budget attacks designed to minimize the perturbation magnitude of adversarial examples while maintaining their transferability against the target model. HORNET utilizes a distinct source model to craft the adversarial examples and employs a limited number of queries to the unknown target model to further minimize perturbation magnitude. We evaluate HORNET empirically by integrating it with 41 existing attack implementations and testing it against 9 different models, resulting in a total of 1700 unique configurations. Our results demonstrate that HORNET outperforms the state of the art in generating minimally perturbed yet highly transferable adversarial examples across all tested models. Code available at: https://github.com/louiswup/HORNET . Jiaping Wu, Antonio Emanuele Cinà, Francesco Villani, Zhaoqiang Xia, Luca Demetrio, Luca Oneto, Davide Anguita, Fabio Roli, Xiaoyi Feng |
Inf. Sci. | 9 |
| 2026 | 3D differential decomposition for video deepfake detection with identity suppressionabstractDetecting deepfake videos remains a challenging task, especially in scenarios involving unknown manipulation methods or unseen data distributions. Most existing video deepfake detection methods rely on high-level semantic features, which often lead to overfitting of facial identity information and poor transferability. In this work, we explore a novel perspective by modeling videos through 3D differential operations along temporal and spatial dimensions. To exploit the spatial–temporal variation information of the video content, the proposed approach decomposes videos into single-axis 1D differential signals, which are then transformed into 2D representations for efficient learning. This procedure enables the use of lightweight 2D CNNs while retaining directional forgery cues. Our experiments, aimed at analyzing whether these differential signals capture discriminative patterns useful for distinguishing real from fake content, show that the proposed method achieves strong intra-dataset performance and reveals complementary information across dimensions. These findings suggest that differential signals could potentially support generalization when integrated into broader detection frameworks. • We propose 3D Differential Decomposition modeling for deepfake video detection. • Multi-directional and multi-order differential operation are considered. • Optimization for differential order selection and fusion strategy are explored. Marco Micheletto, Giulia Orrù, Xiaoyi Feng, Gian Luca Marcialis |
Signal Process. Image Commun. | 4 |
| 2026 | Hybrid-Supervised Hypergraph-Enhanced Transformer for Micro-Gesture Based Emotion RecognitionabstractMicro-gestures are unconsciously performed body gestures that can convey the emotion states of humans and start to attract more research attention in the fields of human behavior understanding and affective computing as an emerging topic. However, the modeling of human emotion based on micro-gestures has not been explored sufficiently. In this work, we propose to recognize the emotion states based on the micro-gestures by reconstructing behavioral patterns with a hypergraph-enhanced Transformer in a hybrid-supervised framework. In the framework, hypergraph Transformer based encoder and decoder are separately designed by stacking the hypergraph-enhanced self-attention and multiscale temporal convolution modules. Especially, to better capture the subtle motion of micro-gestures, we construct a decoder with additional upsampling operations for a reconstruction task in a self-supervised learning manner. We further propose a hypergraph-enhanced self-attention module where the hyperedges between skeleton joints are gradually updated to present the relationships of body joints for modeling the subtle local motion. Lastly, for exploiting the relationship between the emotion states and local motion of micro-gestures, an emotion recognition head from the output of encoder is designed with a shallow architecture and learned in a supervised way. The end-to-end framework is jointly trained in a one-stage way by comprehensively utilizing self-reconstruction and supervision information. The proposed method is evaluated on two publicly available datasets, namely iMiGUE and SMG, and achieves the best performance under multiple metrics, which is superior to the existing methods. The code is available on Github (https://github.com/xiazhaoqiang/H2OFormerMicroGestureRec). Zhaoqiang Xia, Haoyu Chen 0001, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Exploring Facial Kinship Verification through Contactless Heart Activity AnalysisabstractFacial Kinship Verification (FKV) aims at automatically determining whether two subjects have a kinship relation based on human faces. It has potential applications in finding missing children and social media analysis. Traditional FKV faces challenges as it is vulnerable to spoof attacks and raises privacy issues. In this paper, we explore for the first time the FKV by analyzing cardiac activity through physiological signals, with a specific focus on remote Photoplethysmography (rPPG). rPPG signals are extracted from facial videos, resulting in a one-dimensional signal that measures the changes in visible light reflection emitted to and detected from the skin caused by the heartbeat. Specifically, in this paper, we employed a straightforward one-dimensional Convolutional Neural Network (1DCNN) with a 1DCNN-Attention module and kinship contrastive loss to learn the kin similarity from rPPGs. The network takes multiple rPPG signals extracted from various facial Regions of Interest (ROIs) as inputs. Additionally, the 1DCNN attention module is designed to learn and capture the discriminative kin features from feature embeddings. Finally, we demonstrate the feasibility of rPPG to detect kinship with the experiment evaluation on the UvANEMO Smile Database from different kin relations. Xiaoting Wu, Xiaoyi Feng, Constantino Álvarez Casado, Miguel Bordallo López |
ICASSP | 2 |
| 2025 | A Unified Framework for Industrial Cel-Animation Colorization with Temporal-Structural Awareness
Xiaoyi Feng, Tao Huang 0022, Peng Wang 0168, Zizhou Huang, Haihang Zhang, Yuntao Zou, Dagang Li 0001, Kaifeng Zou |
ICCV | 1 |
| 2025 | UMIS-YOLO: Underwater Multimodal Images Instance Segmentation With YOLOabstractUnderwater instance segmentation plays a pivotal role in various applications. Among them, coral instance segmentation is of great significance in the fields of marine biology and environmental monitoring, and is crucial for comprehensive understanding of coral reef ecosystems. Traditional methods for underwater instance segmentation predominantly rely on RGB images. However, the complex morphology of corals and strong background interference often result in poor segmentation outcomes. To tackle these problems, this study presents a novel multimodal instance segmentation method, termed UMIS-YOLO, which is grounded in the YOLO architecture. UMIS-YOLO incorporates a dual backbone network design that substantially enhances the feature extraction capabilities for both RGB images and depth images, thereby improving the effectiveness of instance segmentation. At the same time, we propose two innovative plug-and-play modules: the Frequency Domain Feature Enhancement Fusion (FDFEF) module and the Residual Feature Fusion (RFF) module. The FDFEF module leverages Fourier transform to enhance the features of both modalities in the frequency domain, employing learnable weights to enable the complementary integration of amplitude and phase information. While the RFF module utilizes a residual learning strategy to efficiently merge low-level and high-level features prior to the segmentation head, thereby improving pixel-level segmentation accuracy. Additionally, we introduce a challenging high-resolution dataset, UMIS-Coral, which comprises RGB images and depth images captured in complex coral environments. Meanwhile, we expand the depth images for the UIIS dataset to further verify the effectiveness of UMIS-YOLO. The experimental results indicate that the UMIS-YOLO model achieved mAP50 and mAP75 improvements of 2.3 and 3.0 on the UMIS-Coral dataset, as well as 3.9 and 2.8 on the UIIS dataset, respectively. Furthermore, the model is characterized by its lightweight architecture and rapid segmentation capabilities. The source code and the dataset are publicly accessible at https://github.com/zhangsanhulk/UMIS-YOLO. Yue Yang 0051, Xiaoyi Feng, Ming Li 0037, Xiangyun Hu, Jiangying Qin, Armin Gruen, DeRen Li, Jianya Gong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Intrinsic Image Decomposition Based on Quantized Prior CodebookabstractIntrinsic image decomposition is a low-level image processing task that extracts the reflectance and lighting components from an image. This process can improve the illumination robustness of perception tasks, such as object detection, recognition, and image understanding. Recently, deep image generation frameworks have been used to generate intrinsic images. However, the encoder and decoder lack prior knowledge constraints. This paper presents a quantized codebook for embedding intrinsic features that guide the extraction of intrinsic images. To enhance reconstruction accuracy, we propose a purification method to eliminate irrelevant elements from the codebook. Additionally, we propose self-attention and cross-attention modules to integrate the intrinsic features of the codebook into the input image features for reconstruction. The effectiveness of the algorithm is demonstrated through experiments conducted on several popular datasets. Fangzheng Yuan, Xiaoyue Jiang, Xiaoyi Feng, Moncef Gabbouj |
ICIP | 3 |
| 2024 | Diversified and Structure-Realistic Fundus Image Synthesis for Diabetic Retinopathy Lesion Segmentation
Xiaoyi Feng, Minqing Zhang, Mengxian He, Mengdi Gao, Wu Yuan 0001 |
MICCAI (12) | 1 |
| 2024 | Texture and artifact decomposition for improving generalization in deep-learning-based deepfake detectionabstractThe harmful utilization of DeepFake technology poses a significant threat to public welfare, precipitating a crisis in public opinion. Existing detection methodologies, predominantly relying on convolutional neural networks and deep learning paradigms, focus on achieving high in-domain recognition accuracy amidst many forgery techniques. However, overseeing the intricate interplay between textures and artifacts results in compromised performance across diverse forgery scenarios. This paper introduces a groundbreaking framework, denoted as Texture and Artifact Detector (TAD), to mitigate the challenge posed by the limited generalization ability stemming from the mutual neglect of textures and artifacts. Specifically, our approach delves into the similarities among disparate forged datasets, discerning synthetic content based on the consistency of textures and the presence of artifacts. Furthermore, we use a model ensemble learning strategy to judiciously aggregate texture disparities and artifact patterns inherent in various forgery types, thereby enabling the model’s generalization ability. Our comprehensive experimental analysis, encompassing extensive intra-dataset and cross-dataset validations along with evaluations on both video sequences and individual frames, confirms the effectiveness of TAD. The results from four benchmark datasets highlight the significant impact of the synergistic consideration of texture and artifact information, leading to a marked improvement in detection capabilities. Marco Micheletto, Giulia Orrù, Sara Concas, Xiaoyi Feng, Gian Luca Marcialis, Fabio Roli |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | DeepFake detection based on high-frequency enhancement network for highly compressed content
Zhaoqiang Xia, Gian Luca Marcialis, Chen Dang, Xiaoyi Feng |
Expert Syst. Appl. | 6 |
| 2024 | SCPAD: An approach to explore optical characteristics for robust static presentation attack detection
Chen Dang, Zhaoqiang Xia, Lei Li 0008, Xiaoyi Feng |
Multim. Tools Appl. | 6 |
| 2024 | Information-Enhanced Network for Noncontact Heart Rate Estimation From Facial VideosabstractRemote photoplethysmography (rPPG) is a vital way of measuring heart rate (HR) to reflect human physical and mental health, which is useful for diagnosing cardiovascular and neurological diseases. Many non-contact HR estimation methods have been proposed gradually in recent years, but the majority of approaches are based on a single-modal HR information source, resulting in ineffective and unsatisfactory estimation results due to noise and insufficient information. This paper proposes a novel information-enhanced network for HR estimation based on multimodal (e.g., RGB and NIR) sources to address these problems. In the network, context and modal difference information are sequentially enhanced from spatiotemporal and modal views for accurately describing HR-aware features, while maximum frequency information is enhanced for inhibiting heartbeat noise. Specifically, a context-enhanced video Swin-Transformer (CET) module is exploited to extract useful rPPG signal features from facial visible-light and near-infrared videos. Then, a novel modal difference enhanced fusion (MDEF) module is designed to acquire a fused rPPG signal, which is taken as the input of the frequency-enhanced estimation (FEE) module to obtain the corresponding HR value. These three modules are integrated and jointly learned in an end-to-end way, and the multimodal combinations can provide highly complementary information for estimating HR value. Experimental and evaluation results on three multimodal datasets show that the proposed model achieves a superior effect compared to the state-of-the-art methods. Zhaoqiang Xia, Xiaobiao Zhang, Jinye Peng 0001, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Audio-Visual Kinship Verification: A New Dataset and a Unified Adaptive Adversarial Multimodal Learning ApproachabstractFacial kinship verification refers to automatically determining whether two people have a kin relation from their faces. It has become a popular research topic due to potential practical applications. Over the past decade, many efforts have been devoted to improving the verification performance from human faces only while lacking other biometric information, for example, speaking voice. In this article, to interpret and benefit from multiple modalities, we propose for the first time to combine human faces and voices to verify kinship, which we refer it as the audio-visual kinship verification study. We first establish a comprehensive audio-visual kinship dataset that consists of familial talking facial videos under various scenarios, called TALKIN-Family. Based on the dataset, we present the extensive evaluation of kinship verification from faces and voices. In particular, we propose a deep-learning-based fusion method, called unified adaptive adversarial multimodal learning (UAAML). It consists of the adversarial network and the attention module on the basis of unified multimodal features. Experiments show that audio (voice) information is complementary to facial features and useful for the kinship verification problem. Furthermore, the proposed fusion method outperforms baseline methods. In addition, we also evaluate the human verification ability on a subset of TALKIN-Family. It indicates that humans have higher accuracy when they have access to both faces and voices. The machine-learning methods could effectively and efficiently outperform the human ability. Finally, we include the future work and research opportunities with the TALKIN-Family dataset. Xiaoting Wu, Xueyi Zhang 0001, Xiaoyi Feng, Miguel Bordallo López, Li Liu 0002 |
IEEE Trans. Cybern. | 3 |
| 2023 | Wooden spoon crack detection by prior knowledge-enriched deep convolutional network
Lei Li 0008, Huijian Han, Xiaoyi Feng, Fabio Roli, Zhaoqiang Xia |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | Hardening RGB-D object recognition systems against adversarial patch attacks
Luca Demetrio, Antonio Emanuele Cinà, Xiaoyi Feng, Zhaoqiang Xia, Xiaoyue Jiang, Ambra Demontis, Battista Biggio, Fabio Roli |
Inf. Sci. | 4 |
| 2023 | Why adversarial reprogramming works, when it fails, and how to tell the differenceabstractAdversarial reprogramming allows repurposing a machine-learning model to perform a different task. For example, a model trained to recognize animals can be reprogrammed to recognize digits by embedding an adversarial program in the digit images provided as input. Recent work has shown that adversarial reprogramming may not only be used to abuse machine-learning models provided as a service, but also beneficially, to improve transfer learning when training data is scarce. However, the factors affecting its success are still largely unexplained. In this work, we develop a first-order linear model of adversarial reprogramming to show that its success inherently depends on the size of the average input gradient, which grows when input gradients are more aligned, and when inputs have higher dimensionality. The results of our experimental analysis, involving fourteen distinct reprogramming tasks, show that the above factors are correlated with the success and the failure of adversarial reprogramming. Xiaoyi Feng, Zhaoqiang Xia, Xiaoyue Jiang, Ambra Demontis, Maura Pintor, Battista Biggio, Fabio Roli |
Inf. Sci. | 2 |
| 2023 | Stateful detection of adversarial reprogrammingabstractAdversarial reprogramming allows stealing computational resources by repurposing machine learning models to perform a different task chosen by the attacker. For example, a model trained to recognize images of animals can be reprogrammed to recognize medical images by embedding an adversarial program in the images provided as inputs. This attack can be perpetrated even if the target model is a black box, supposed that the machine-learning model is provided as a service and the attacker can query the model and collect its outputs. So far, no defense has been demonstrated effective in this scenario. We show for the first time that this attack is detectable using stateful defenses, which store the queries made to the classifier and detect the abnormal cases in which they are similar. Once a malicious query is detected, the account of the user who made it can be blocked. Thus, the attacker must create many accounts to perpetrate the attack. To decrease this number, the attacker could create the adversarial program against a surrogate classifier and then fine-tune it by making a few queries to the target model. In this scenario, the effectiveness of the stateful defense is reduced, but we show that it is still effective. Xiaoyi Feng, Zhaoqiang Xia, Xiaoyue Jiang, Maura Pintor, Ambra Demontis, Battista Biggio, Fabio Roli |
Inf. Sci. | 2 |
| 2023 | Demodulation Based Transformer for rPPG Generation and Heart Rate EstimationabstractAs a convenient, efficient and inexpensive medical technology, heart rate (HR) estimation from face videos has gradually become a hot topic and many methods have been exploited to learn a spatial-temporal map for HR estimation. However, these methods have poor performance under time-varying ambient lighting. The main reason is that the current methods neglect the optical modeling of extracting the contactless physiological signal from skin. In this study, we describe the signal extraction as a modulation process and draw the conclusion that light changing could produce amplitude modulation jamming after modeling analysis. Therefore, a demodulation-based Transformer is newly designed for rPPG signal purification. In addition, a Pwelch-based Softmax operation is incorporated for HR estimation to improve accuracy. Finally, the hybrid loss combined with the negative Pearson correlation coefficient and cross-entropy loss is introduced for entire network learning. The experimental results on two databases (COHFACE and PURE) are performed to verify the effectiveness of the proposed method. Xiaobiao Zhang, Zhaoqiang Xia, Xiaoyi Feng |
IEEE Signal Process. Lett. | 4 |
| 2023 | Non-Local Color Compensation Network for Intrinsic Image DecompositionabstractSingle image-based intrinsic image decomposition attempts to separate one input image into several intrinsic components, which is inherently an under-constrained problem. Some recent works have been proposed to estimate the intrinsic components using encoder-decoder structures. However, they generally lack exploration of the different component-oriented feature constraints and feature selection processes. In this paper, a non-local color compensation network (NCCNet) is proposed. Firstly, the hue and value channels of HSV color space are used as the complementary information for RGB images for the estimation of albedo and shading, respectively. The color space representation serves as an external constraint, which does not require expensive sensors or complicated computations. Secondly, an integrated non-local attention scheme is proposed to describe the relations of non-adjacent regions with a lower computational complexity compared to traditional methods. Then the non-local and local attention are combined to describe correlations among features and used as feature selectors between the encoder and decoder. Thirdly, the mutual constraint between albedo and shading is also explored in the network to further optimize the process. In order to train the network, a unified mutual exclusion loss function is proposed. Extensive experiments are conducted on several popular datasets, and the proposed NCCNet achieves improved performance with comparable computational cost compared to competing methods. Xiaoyue Jiang, Zhaoqiang Xia, Moncef Gabbouj, Jinye Peng 0001, Xiaoyi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Distribution Reliability Assessment-Based Incremental Learning for Automatic Target RecognitionabstractIn order to rapidly improve the automatic target recognition (ATR) system when new unknown samples are constantly captured, it is necessary to examine the existing training samples and recognition model so that the ATR system could autonomously assess new unknown samples with low predictive reliability during the recognition process and learn them preferentially. Incremental learning methods generally consider forming key exemplar set from existing known samples, but rarely managing updates of unknown samples. In this paper, an incremental samples’ evaluation and management method from the perspective of distribution reliability (DRaIL) is proposed, which realizes the retention of existent reliable exemplars and the predictive-reliability-assessment-based updating of new unknown samples simultaneously. DRaIL preserves the prior distribution in the high-density and overlap regions first, and then the classification reliability and “in-of-distribution" reliability of new unknown samples are evaluated based on the consistency between the new and the preserved distribution. Updating the new samples with low reliability using new labels could rapidly improve the classification surface and add new classes. Experimental results for the practical incremental learning scenario demonstrate the validity of the proposed DRaIL on representative exemplar selection and reliability ranking performance. Sihang Dang, Zongyong Cui, Zongjie Cao, Yiming Pi, Xiaoyi Feng |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Inclusive Consistency-Based Quantitative Decision-Making Framework for Incremental Automatic Target RecognitionabstractWhen new unknown samples are captured continually in the open-world environment, the concept diversity accumulation of existing classes and the identification/creation of new concept classes should be considered simultaneously. Since the initial training set of existent classes may be under-prepared, adhering to immediate decisions will inevitably lead to reduced open set recognition performance and higher costs of labeling/updating. Inspired by quantitative indicators in predictive reliability assessment and semi-supervised/active learning, the inclusive-consistency-based quantitative decision-making framework (ICQdm) is proposed for incremental automatic target recognition (ATR) to evaluate the identifiability and typicality of new unknown samples, which could give the decision-making guide of recognition and updating. For recognition decision-making, the first consistency indicator calculates the reliability of the unknown sample being included by one specific training class. The test samples with low reliability should be the wrongly classified samples and new unknown classes’ samples, which are difficult to be labeled by recognition models themselves. For updating decision-making, the second consistency indicator is designed to be the sample distribution density under the inclusive constraint, which could highlight the dense sample distributions of new unknown samples outside the known training distribution. Experiments verify that the proposed ICQdm outperforms other comparison methods on the open set recognition reliability evaluation and labeling/updating efficiency. Sihang Dang, Zhaoqiang Xia, Xiaoyue Jiang, Shuliang Gui, Xiaoyi Feng |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Facial Kinship Verification: A Comprehensive Review and OutlookabstractThe goal of Facial Kinship Verification (FKV) is to automatically determine whether two individuals have a kin relationship or not from their given facial images or videos. It is an emerging and challenging problem that has attracted increasing attention due to its practical applications. Over the past decade, significant progress has been achieved in this new field. Handcrafted features and deep learning techniques have been widely studied in FKV. The goal of this paper is to conduct a comprehensive review of the problem of FKV. We cover different aspects of the research, including problem definition, challenges, applications, benchmark datasets, a taxonomy of existing methods, and state-of-the-art performance. In retrospect of what has been achieved so far, we identify gaps in current research and discuss potential future research directions. Xiaoting Wu, Xiaoyi Feng, Xiaochun Cao, Xin Xu 0001, Dewen Hu, Miguel Bordallo López, Li Liu 0002 |
Int. J. Comput. Vis. | 2 |
| 2022 | Spatio-Temporal Pain Estimation Network With Measuring Pseudo Heart Rate GainabstractPain is a significant indicator that shows people are suffering from an unwell experience and its automatic estimation has attracted much interest in recent years. Of late, most estimation methods are designed to capture the dynamic pain information from visual signals while a few physiological-signal based methods can provide extra potential cues to analyze the pain more accurately. However, it is still challenging to capture the physiological data from patients as it requires contact devices and patients’ cooperation. In this paper, we propose to leverage the pseudo physiological information by generating new modal data from the original visual videos and jointly estimating the pain by an end-to-end network. To extract the representations from bi-modal data, we design a spatio-temporal pain estimation network, which employs a dual-branch framework for extracting pain-aware visual and pseudo physiological features separately and fuses the features in a probabilistic way. The inherent vital sign, i.e., heart rate gain (HRG), from pseudo physiological information can be utilized as an auxiliary signal and integrated with the visual pain estimation framework. Moreover, specially-designed 3D convolution filters and attention structures are employed to extract spatio-temporal features for both branches. To use the HRG as an auxiliary way for pain estimation, we propose a probabilistic inference model by jointly considering the visual branch and physiological branch, which makes our model estimate the pain comprehensively. Experiments on two publicly-available datasets show the effectiveness of introducing the pseudo modality, and the proposed method can outperform the state-of-the-art methods. Dong Huang 0003, Xiaoyi Feng, Haixi Zhang, Zitong Yu, Jinye Peng 0001, Guoying Zhao 0001, Zhaoqiang Xia |
IEEE Trans. Multim. | 2 |
| 2021 | Shadow Detection and Removal Based on Multi-task Generative Adversarial Networks
Xiaoyue Jiang, Zhongyun Hu, Yue Ni, Xiaoyi Feng |
ICIG (3) | 5 |
| 2021 | Non-contact Pain Recognition from Video Sequences with Remote Physiological Measurements PredictionabstractAutomatic pain recognition is paramount for medical diagnosis and treatment. The existing works fall into three categories: assessing facial appearance changes, exploiting physiological cues, or fusing them in a multi-modal manner. However, (1) appearance changes are easily affected by subjective factors which impedes objective pain recognition. Besides, the appearance-based approaches ignore long-range spatial-temporal dependencies that are important for modeling expressions over time; (2) the physiological cues are obtained by attaching sensors on human body, which is inconvenient and uncomfortable. In this paper, we present a novel multi-task learning framework which encodes both appearance changes and physiological cues in a non-contact manner for pain recognition. The framework is able to capture both local and long-range dependencies via the proposed attention mechanism for the learned appearance representations, which are further enriched by temporally attended physiological cues (remote photoplethysmography, rPPG) that are recovered from videos in the auxiliary task. This framework is dubbed rPPG-enriched Spatio-Temporal Attention Network (rSTAN) and allows us to establish the state-of-the-art performance of non-contact pain recognition on publicly available pain databases. It demonstrates that rPPG predictions can be used as an auxiliary task to facilitate non-contact automatic pain recognition. Ruijing Yang, Ziyu Guan, Zitong Yu, Xiaoyi Feng, Jinye Peng 0001, Guoying Zhao 0001 |
IJCAI | 4 |
| 2021 | RBPR: A hybrid model for the new user cold start problem in recommender systems
Junmei Feng, Zhaoqiang Xia, Xiaoyi Feng, Jinye Peng 0001 |
Knowl. Based Syst. | 3 |
| 2021 | CasQNet: Intrinsic Image Decomposition Based on Cascaded Quotient NetworkabstractIntrinsic image analysis plays an important role for image understanding, since it can provide accurate reflectance, shape and illumination information of the scene. However, intrinsic image analysis is an ill-posed problem which need to apply extra constrains for the decomposition of reflectance image and shading image from a single image. Recently deep neural networks are introduced for intrinsic image analysis, which can produce two intrinsic components simultaneously. In fact, the mutually exclusive relationship between reflectance image and shading image is not only a constraint for decomposition but also can improve the decomposition results. However, this relationship is always omitted in the current networks. In order to address this problem, we propose a novel deep network called as Cascaded Quotient Network (CasQNet) for intrinsic image decomposition. The CasQNet consists of two sub-networks: a Pyramid Mini-U-Net (PyNet) that specifically extracts the reflectance image in multi-scale and a Shading Optimization Network (SoNet) that optimizes the resulting shading. These two sub-networks are cascaded by a quotient operation, which directly enforces the mutually exclusive relationship between reflectance image and shading image in the network architecture. In PyNet, the task of reconstructing reflectance image is achieved by a series of nested multi-scale U-Nets, which simplified the learning task for each U-Net. SoNet is designed to address the unsmooth and blur problems of extreme points caused by the quotient operation. PyNet and SoNet are trained alternately and finally jointed in cascaded structure. Furthermore, we combine multiple loss functions, which consist of data loss, correlation loss and reconstruction loss, for improving the learning effectiveness. To evaluate our proposed algorithm, extensive experiments are performed on three datasets, i.e., ShapeNet, BOLD Surface and MIT Intrinsic Image datasets. Qualitative and quantitative results show that our model achieves the best performance compared to the state-of-the-art methods. Yupeng Ma, Xiaoyue Jiang, Zhaoqiang Xia, Moncef Gabbouj, Xiaoyi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | A drug information embedding method based on graph convolution neural networkabstractNew drug development is an extremely time-consuming and high-risk process. [1]It has been widely valued by the biomedical industry to fully explore the new uses of existing drugs and reorientate them. [2]How to find drug disease with potential therapeutic relationship from a large number of unproven relationship pairs is the research focus of drug reorientation. With the help of machine learning model, we can improve the enrichment degree of potential drug disease relationship pairs, and reduce the false positive rate of prediction. In the past few years, a series of graph based convolutional network models have been developed to calculate the information latent feature representation of nodes and links. Researchers at home and abroad have done a lot of research on network embedding technology based on biomedical data, and have achieved a series of important research results. Among them, the research methods used can be divided into two categories: one is the traditional machine learning algorithm based on artificial feature extraction, the other is the method based on deep learning. For example, kipf and welling [3]proposed a new graph convolution network (GCN) with parts of existing models, DeepDR [4] and DTINet [5] based on node characteristics and their connections, which can be used for node classification. Aiming at the problem of imbalance of drug information data samples, the invention provides a drug relocation method based on deep learning multi-source heterogeneous network. In order to avoid the limitations of traditional feature extraction methods, such as highly dependent on the experience and knowledge of medical staff, strong subjectivity, consuming a lot of time and energy to complete, and extracting high-quality features with distinguishing features often exists In this paper, with the help of graph convolution encoder model and variational auto encoder neural network, we can automatically learn the characteristics of multi-source and heterogeneous drug low-dimensional network, and complete the drug relocation of drug disease association prediction. Xiaoyi Feng, Shaoliang Peng, Fei Li 0040, Xiangxiang Zeng, Yunhao Liu 0001 |
HealthCom | 1 |
| 2020 | Are spoofs from latent fingerprints a real threat for the best state-of-art liveness detectors?abstractWe investigated the threat level of realistic attacks using latent fingerprints against sensors equipped with state-of-art liveness detectors and fingerprint verification systems which integrate such liveness algorithms. To the best of our knowledge, only a previous investigation was done with spoofs from latent prints. In this paper, we focus on using snapshot pictures of latent fingerprints. These pictures provide molds, that allows, after some digital processing, to fabricate high-quality spoofs. Taking a snapshot picture is much simpler than developing fingerprints left on a surface by magnetic powders and lifting the trace by a tape. What we are interested here is to evaluate preliminary at which extent attacks of the kind can be considered a real threat for state-of-art fingerprint liveness detectors and verification systems. To this aim, we collected a novel data set of live and spoof images fabricated with snapshot pictures of latent fingerprints. This data set provide a set of attacks at the most favourable conditions. We refer to this method and the related data set as “ScreenSpoof”. Then, we tested with it the performances of the best liveness detection algorithms, namely, the three winners of the LivDet competition. Reported results point out that the ScreenSpoof method is a threat of the same level, in terms of detection and verification errors, than that of attacks using spoofs fabricated with the full consensus of the victim. We think that this is a notable result, never reported in previous work. Roberto Casula, Giulia Orrù, Daniele Angioni, Xiaoyi Feng, Gian Luca Marcialis, Fabio Roli |
ICPR | 4 |
| 2020 | Generalized Operational Classifiers for Material IdentificationabstractMaterial is one of the intrinsic features of objects, and consequently material recognition plays an important role in image understanding. The same material may have various shapes and appearance, while keeping the same physical characteristic. This brings great challenges for material recognition. Besides suitable features, a powerful classifier also can improve the overall recognition performance. Due to the limitations of classical linear neurons, used in all shallow and deep neural networks, such as CNN, we propose to apply the generalized operational neurons to construct a classifier adaptively. These generalized operational perceptrons (GOP) contain a set of linear and nonlinear neurons, and possess a structure that can be built progressively. This makes GOP classifier more compact and can easily discriminate complex classes. The experiments demonstrate that GOP networks trained on a small portion of the data (4%) can achieve comparable performances to state-of-the-arts models trained on much larger portions of the dataset. Xiaoyue Jiang, Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Xiaoyi Feng |
MMSP | 6 |
| 2020 | Deep neural rejection against adversarial examplesabstractAbstract Despite the impressive performances reported by deep neural networks in different application domains, they remain largely vulnerable to adversarial examples, i.e., input samples that are carefully perturbed to cause misclassification at test time. In this work, we propose a deep neural rejection mechanism to detect adversarial examples, based on the idea of rejecting samples that exhibit anomalous feature representations at different network layers. With respect to competing approaches, our method does not require generating adversarial examples at training time, and it is less computationally demanding. To properly evaluate our method, we define an adaptive white-box attack that is aware of the defense mechanism and aims to bypass it. Under this worst-case setting, we empirically show that our approach outperforms previously proposed methods that detect adversarial examples by only analyzing the feature representation provided by the output network layer. Angelo Sotgiu, Ambra Demontis, Marco Melis, Battista Biggio, Giorgio Fumera, Xiaoyi Feng, Fabio Roli |
EURASIP J. Inf. Secur. | 6 |
| 2020 | Infrared and visible image fusion using a shallow CNN and structural similarity constraintabstractIn recent years, image fusion methods based on deep networks have been proposed to combine infrared and visible images for achieving better fusion image. However, issues such as limited training data, scarce reference images and misalignment of multi‐source images, still limit the fusion performance. To address these problems, we propose an end‐to‐end shallow convolutional neural network with structural constraints, which has only one convolutional layer to fuse infrared and visible images. Different from other methods, our proposed model requires less training data and reference images and is more robust to the misalignment of a couple of images. More specifically, the infrared image and the visible image are first provided as inputs to a convolutional layer to extract the information that should be fused; then, all feature maps are concatenated together and fed into a convolutional layer with one channel to obtain the fused image; finally, a structural similarity loss between the fused image and the input infrared and visible images is computed to update the network parameters and eliminate the effects of pixel misalignment. Extensive experiments show the effectiveness of our proposed method on fusion of infrared and visible images with the performance that outperforms the state‐of‐the‐art methods. Lei Li 0008, Zhaoqiang Xia, Huijian Han, Guiqing He, Fabio Roli, Xiaoyi Feng |
IET Image Process. | 6 |
| 2020 | CompactNet: learning a compact space for face presentation attack detection
Lei Li 0008, Zhaoqiang Xia, Xiaoyue Jiang, Fabio Roli, Xiaoyi Feng |
Neurocomputing | 5 |
| 2020 | Pain-attentive network: a deep spatio-temporal attention model for pain estimation
Dong Huang 0003, Zhaoqiang Xia, Joshua Mwesigye, Xiaoyi Feng |
Multim. Tools Appl. | 4 |
| 2020 | Revealing the Invisible With Model and Data Shrinking for Composite-Database Micro-Expression RecognitionabstractComposite-database micro-expression recognition is attracting increasing attention as it is more practical for real-world applications. Though the composite database provides more sample diversity for learning good representation models, the important subtle dynamics are prone to disappearing in the domain shift such that the models greatly degrade their performance, especially for deep models. In this paper, we analyze the influence of learning complexity, including input complexity and model complexity, and discover that the lower-resolution input data and shallower-architecture model are helpful to ease the degradation of deep models in composite-database task. Based on this, we propose a recurrent convolutional network (RCN) to explore the shallower-architecture and lower-resolution input data, shrinking model and input complexities simultaneously. Furthermore, we develop three parameter-free modules (i.e., wide expansion, shortcut connection and attention unit) to integrate with RCN without increasing any learnable parameters. These three modules can enhance the representation ability in various perspectives while preserving not-very-deep architecture for lower-resolution data. Besides, three modules can further be combined by an automatic strategy (a neural architecture search strategy) and the searched architecture becomes more robust. Extensive experiments on the MEGC2019 dataset (composited of existing SMIC, CASME II and SAMM datasets) have verified the influence of learning complexity and shown that RCNs with three modules and the searched combination outperform the state-of-the-art approaches. Zhaoqiang Xia, Wei Peng 0009, Huai-Qian Khor, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Spatiotemporal Recurrent Convolutional Networks for Recognizing Spontaneous Micro-ExpressionsabstractRecently, the recognition task of spontaneous facial micro-expressions has attracted much attention with its various real-world applications. Plenty of handcrafted or learned features have been employed for a variety of classifiers and achieved promising performances for recognizing micro-expressions. However, the micro-expression recognition is still challenging due to the subtle spatiotemporal changes of micro-expressions. To exploit the merits of deep learning, we propose a novel deep recurrent convolutional networks based micro-expression recognition approach, capturing the spatiotemporal deformations of micro-expression sequence. Specifically, the proposed deep model is constituted of several recurrent convolutional layers for extracting visual features and a classificatory layer for recognition. It is optimized by an end-to-end manner and obviates manual feature design. To handle sequential data, we exploit two ways to extend the connectivity of convolutional networks across temporal domain, in which the spatiotemporal deformations are modeled in views of facial appearance and geometry separately. Besides, to overcome the shortcomings of limited and imbalanced training samples, two temporal data augmentation strategies as well as a balanced loss are jointly used for our deep network. By performing the experiments on three spontaneous micro-expression datasets, we verify the effectiveness of our proposed micro-expression recognition approach compared to the state-of-the-art methods. Zhaoqiang Xia, Xiaopeng Hong, Xingyu Gao 0001, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | Corrections to "Spatiotemporal Recurrent Convolutional Networks for Recognizing Spontaneous Micro-Expressions"abstractPresents corrections to the author's information in the above named paper. Zhaoqiang Xia, Xiaopeng Hong, Xingyu Gao 0001, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Multim. | 4 |
| 2019 | Pulmonary DR Image Anomaly Detection Based on Deep Learning
Zhendong Song, Dong Huang 0003, Xiaoyi Feng |
ICIG (1) | 4 |
| 2019 | Scene video text tracking based on hybrid deep text detection and layout constraint
Xihan Wang, Xiaoyi Feng, Zhaoqiang Xia |
Neurocomputing | 2 |
| 2019 | Trademark image retrieval via transformation-invariant deep hashing
Zhaoqiang Xia, Jie Lin 0001, Xiaoyi Feng |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Discriminative Spatiotemporal Local Binary Pattern with Revisited Integral Projection for Spontaneous Facial Micro-Expression RecognitionabstractRecently, there have been increasing interests in inferring mirco-expression from facial image sequences. Due to subtle facial movement of micro-expressions, feature extraction has become an important and critical issue for spontaneous facial micro-expression recognition. Recent works used spatiotemporal local binary pattern (STLBP) for micro-expression recognition and considered dynamic texture information to represent face images. However, they miss the shape attribute of face images. On the other hand, they extract the spatiotemporal features from the global face regions while ignore the discriminative information between two micro-expression classes. The above-mentioned problems seriously limit the application of STLBP to micro-expression recognition. In this paper, we propose a discriminative spatiotemporal local binary pattern based on an integral projection to resolve the problems of STLBP for micro-expression recognition. First, we revisit an integral projection for preserving the shape attribute of micro-expressions by using robust principal component analysis. Furthermore, a revisited integral projection is incorporated with local binary pattern across spatial and temporal domains. Specifically, we extract the novel spatiotemporal features incorporating shape attributes into spatiotemporal texture features. For increasing the discrimination of micro-expressions, we propose a new feature selection based on Laplacian method to extract the discriminative information for facial micro-expression recognition. Intensive experiments are conducted on three availably published micro-expression databases including CASME, CASME2 and SMIC databases. We compare our method with the state-of-the-art algorithms. Experimental results demonstrate that our proposed method achieves promising performance for micro-expression recognition. Xiaohua Huang 0003, Xin Liu 0012, Guoying Zhao 0001, Xiaoyi Feng, Matti Pietikäinen |
IEEE Trans. Affect. Comput. | 5 |
| 2019 | Replayed Video Attack Detection Based on Motion Blur AnalysisabstractFace presentation attacks are the main threats to face recognition systems, and many presentation attack detection (PAD) methods have been proposed in recent years. Although these methods have achieved significant performance in some specific intrusion modes, difficulties still exist in addressing replayed video attacks. That is because the replayed fake faces contain a variety of aliveness signals, such as eye blinking and facial expression changes. Replayed video attacks occur when attackers try to invade biometric systems by presenting face videos in front of the cameras, and these videos are often launched by a liquid-crystal display (LCD) screen. Due to the smearing effects and movements of LCD, videos captured from the real and replayed fake faces present different motion blurs, which are reflected mainly in blur intensity variation and blur width. Based on these descriptions, a motion blur analysis-based method is proposed to deal with the replayed video attack problem. We first present a 1D convolutional neural network (CNN) for motion blur intensity variation description in the time domain, which consists of a serial of 1D convolutional and pooling filters. Then, a local similar pattern (LSP) feature is introduced to extract blur width. Finally, features extracted from 1D CNN and LSP are fused to detect the replayed video attacks. Extensive experiments on two standard face PAD databases, i.e., relay-attack and OULU-NPU, indicate that our proposed method based on the motion blur analysis significantly outperforms the state-of-the-art methods and shows excellent generalization capability. Lei Li 0008, Zhaoqiang Xia, Abdenour Hadid, Xiaoyue Jiang, Haixi Zhang, Xiaoyi Feng |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2018 | A Dynamic Generalized Opposition-Based Learning Fruit Fly Algorithm for Function OptimizationabstractAs a novel evolutionary algorithm, fruit fly optimization algorithm (FOA) has received great attentions and wide applications in recent years. However, existing literature have demonstrated that the basic FOA often risks getting prematurely stuck in the local optima. In this paper, an improved FOA, named as dynamic generalized opposition-based learning fruit fly optimization algorithm (DGOBL-FOA), is proposed to mitigate the aforementioned drawback hence improve the optimization performance. Three carefully designed operators are incorporated into the basic FOA, i.e., a cloud model based osphresis search is applied to enhance the local refinement search ability in the osphresis phase, then a generalized opposition-based learning operation is adopted to strengthen the global coarse search ability, meanwhile a dynamic shrinking parameter strategy is designed to adjust the learning intensity and narrow down the search space iteratively, which contributes to a good balance between the global exploration and local exploitation. To verify the effectiveness of the proposed algorithm, numerical experiments are conducted on 18 well-studied benchmark functions with dimension of 30. The computation results and statistical analysis indicate that the proposed DGOBL-FOA achieve significantly better performance comparing to other FOA variants and the state-of-the-art metaheuristics. Xiaoyi Feng, Ao Liu 0002, Weiliang Sun, Xiaofeng Yue, Bo Liu 0008 |
CEC | 1 |
| 2018 | Incorporating high-level and low-level cues for pain intensity estimationabstractPain is a transient physical reaction that exhibits on human faces. Automatic pain intensity estimation is of great importance in clinical and health-care applications. Pain expression is identified by a set of deformations of facial features. Hence, features are essential for pain estimation. In this paper, we propose a novel method that encodes low-level descriptors and powerful high-level deep features by a weighting process, to form an efficient representation of facial images. To obtain a powerful and compact low-level representation, we explore the way of using second-order pooling over the local descriptors. Instead of direct concatenation, we develop an efficient fusion approach that unites the low-level local descriptors and the high-level deep features. To the best of our knowledge, this is the first approach that incorporates the low-level local statistics together with the high-level deep features in pain intensity estimation. Experiments are evaluated on the benchmark databases of pain. The results demonstrate that the proposed low-to-high-level representation outperforms other methods and achieves promising results. Ruijing Yang, Xiaopeng Hong, Jinye Peng 0001, Xiaoyi Feng, Guoying Zhao 0001 |
ICPR | 4 |
| 2018 | BCH-LSH: a new scheme of locality-sensitive hashingabstractSimilarity searching of high‐dimensional data is fundamental in the multimedia research field. In recent years, the binary code indexing has achieved significant applications in the context of similarity searching. However, most of the existing binary coding methods adopt a random generation method in near neighbour cluster problems, which involve unnecessary computations and degrade similarity in object points. To avoid the uncertainty of random generation codes, in this study, the authors propose a new locality sensitive hashing (LSH) algorithm based on q ‐ary Bose–Chaudhuri–Hocquenghem (BCH) code. BCH–LSH algorithm utilises the characteristics of the designed distance of BCH codes and uses the BCH codes generator matrix as a transform basis of the hash function to map the source data into the hash space. The experiments show that the BCH–LSH algorithm is superior to the E2LSH algorithm in average precision, average recall ratio and running speed. Yuena Ma, Xiaoyi Feng, Yang Liu 0035, Shuhong Li |
IET Image Process. | 2 |
| 2018 | Face spoofing detection with local binary pattern network
Lei Li 0008, Xiaoyi Feng, Zhaoqiang Xia, Xiaoyue Jiang, Abdenour Hadid |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Team effectiveness based optimizationabstractDuring the past two decades, developing and improving intelligent optimization algorithms (IOAs) has been one of the hottest research topics in evolutionary computing. Most of the IOAs are inspired by natural phenomena, and shown to be highly effective in solving complex optimization problems. In this study, we propose a new population based metaheuristic algorithm named Team Effectiveness Based Optimization (TEBO), which is inspired from human society rather than natural world. The high-intelligence of human beings together with the research achievements in the field of team effectiveness motivate the development of this novel optimization algorithm. In this paper, we present how the members learn and cooperate in a team environment in order to improve the overall team performance. We investigate the proposed TEBO on a comprehensive set of 30 benchmark problems in CEC 2014 competition on Single Objective Real-Parameter Numerical Optimization. The results confirm the competitiveness of the proposed algorithm comparing to other state-of-the-art intelligent optimization algorithms, including invasive weed optimization (IWO), biogeography-based optimization (BBO), gravitational search algorithm (GSA), hunting search (HuS), bat algorithm (BA) and water wave optimization (WWO). Xiaoyi Feng, Mengchen Ji, Xinghua Qu, Bo Liu 0008 |
CEC | 1 |
| 2017 | OULU-NPU: A Mobile Face Presentation Attack Database with Real-World VariationsabstractThe vulnerabilities of face-based biometric systems to presentation attacks have been finally recognized but yet we lack generalized software-based face presentation attack detection (PAD) methods performing robustly in practical mobile authentication scenarios. This is mainly due to the fact that the existing public face PAD datasets are beginning to cover a variety of attack scenarios and acquisition conditions but their standard evaluation protocols do not encourage researchers to assess the generalization capabilities of their methods across these variations. In this present work, we introduce a new public face PAD database, OULU-NPU, aiming at evaluating the generalization of PAD methods in more realistic mobile authentication scenarios across three covariates: unknown environmental conditions (namely illumination and background scene), acquisition devices and presentation attack instruments (PAI). This publicly available database consists of 5940 videos corresponding to 55 subjects recorded in three different environments using high-resolution frontal cameras of six different smartphones. The high-quality print and video-replay attacks were created using two different printers and two different display devices. Each of the four unambiguously defined evaluation protocols introduces at least one previously unseen condition to the test set, which enables a fair comparison on the generalization capabilities between new and existing approaches. The baseline results using color texture analysis based face PAD method demonstrate the challenging nature of the database. Zinelabidine Boulkenafet, Jukka Komulainen, Lei Li 0008, Xiaoyi Feng, Abdenour Hadid |
FG | 4 |
| 2017 | Text Detection Based on Affine Transformation
Xiaoyue Jiang, Xiaoyi Feng |
ICIG (1) | 5 |
| 2017 | Similar Trademark Image Retrieval Integrating LBP and Convolutional Neural Network
Xiaoyi Feng, Zhaoqiang Xia, Shijie Pan, Jinye Peng 0001 |
ICIG (3) | 2 |
| 2017 | Intrinsic Image Decomposition: A Comprehensive Review
Yupeng Ma, Xiaoyi Feng, Xiaoyue Jiang, Zhaoqiang Xia, Jinye Peng 0001 |
ICIG (1) | 2 |
| 2017 | Multi-orientation Scene Text Detection Leveraging Background Suppression
Xihan Wang, Xiaoyi Feng, Zhaoqiang Xia, Jinye Peng 0001, Eric Granger |
ICIG (1) | 2 |
| 2017 | Face anti-spoofing via deep local binary patternsabstractConvolutional neural networks (CNNs) have achieved excellent performance in the field of pattern recognition when huge amount of training data is available. However, training a CNN model is less obvious when only a limited amount of data is given such as in the case of face anti-spoofing problem. It is indeed not easy to collect very large sets of fake faces. Especially for the fully-connected layers, tens of thousands of parameters need to be learned. To tackle this problem of lack of training data in face anti-spoofing, we propose to explore the incorporation of hand-crafted features in the CNN framework. In our proposed approach, the color local binary patterns (LBP) features are extracted from the convolutional feature maps, which are fine tuned based on the VGG-face model. These features are then fed into support vector machine (SVM) classifier. Extensive experiments are conducted on two benchmark and publicly available databases showing very interesting performance compared to state-of-the-art methods. Lei Li 0008, Xiaoyi Feng, Xiaoyue Jiang, Zhaoqiang Xia, Abdenour Hadid |
ICIP | 2 |
| 2017 | Variable neighborhood based memetic algorithm for Just-in-Time distributed assembly permutation flowshop schedulingabstractDistributed Assembly Permutation Flowshop Scheduling Problem (DAPFSP) represents a widely-applied manufacturing pattern, which is formed by two relatively independent stages, i.e., processing stage and assembly stage. In this paper, we introduce the Just-in-Time (JIT) constraint between the two stages of the conventional DAPFSP to form the Just-in-Time DAPFSP (JIT-DAPFSP), with respect to minimizing the maximum weighted earliness/tardiness cost. A variable neighborhood search based Memetic Algorithm (VNS-MA) is proposed by incorporating several novel neighborhoods manipulating the factory assignment. Computational tests are carried out on 10 small-scale and 5 large-scale benchmark problems to compare the results of different operators and algorithms. To the best of our knowledge, this is the first report to model JIT constraints and objective in DAPFSP. Keyao Wang, Wenzhe Duan, Xiaoyi Feng, Bo Liu 0008 |
SMC | 4 |
| 2017 | Deep convolutional hashing using pairwise multi-label supervision for large-scale visual search
Zhaoqiang Xia, Xiaoyi Feng, Jie Lin 0001, Abdenour Hadid |
Signal Process. Image Commun. | 2 |
| 2016 | A novel water wave optimization based memetic algorithm for flow-shop schedulingabstractRecently, Water Wave Optimization (WWO) has been presented as an optimization approach which takes inspiration from shallow wave model considering wind forcing, nonlinear wave interaction, and frictional dissipation. In contrast to WWO dedicated for continuous optimization problems, this paper extends WWO for solving combinatorial permutation flow-shop scheduling problem (PFSSP). In particular, three fundamental search operators in WWO, i.e., propagation, refraction, and breaking have been re-formulated to make WWO suitable for solving combinatorial optimization problem. To enhance the searching performance and efficiency, an improved NEH based algorithm has been applied. Simulation results based on well-known benchmarks and comparisons indicate the competitive advantage of the proposed WWO based memetic algorithm in which the global exploration and the local exploitation are well balanced, providing satisfactory solutions over the state-of-the-art (meta) heuristics such as genetic algorithm for PFSSP. Xin Yun, Xiaoyi Feng, Shou-Yang Wang, Bo Liu 0008 |
CEC | 2 |
| 2016 | Spontaneous micro-expression spotting via geometric deformation modeling
Zhaoqiang Xia, Xiaoyi Feng, Jinye Peng 0001, Xianlin Peng, Guoying Zhao 0001 |
Comput. Vis. Image Underst. | 2 |
| 2016 | Optimal targeting of nonlinear chaotic systems using a novel evolutionary computing strategy
Xiaoyi Feng, Bo Liu 0008 |
Knowl. Based Syst. | 2 |
| 2016 | Single image super-resolution by combining self-learning and example-based learning methods
Na Ai, Jinye Peng 0001, Xuan Zhu 0003, Xiaoyi Feng |
Multim. Tools Appl. | 4 |
| 2016 | Labeling faces with names based on the name semantic network
Xueping Su, Xiaoyi Feng |
Multim. Tools Appl. | 3 |
| 2016 | Parameter estimation of nonlinear chaotic system by improved TLBO strategy
Baozhu Li, Yuanhui Qin, Xiaoyi Feng, Bo Liu 0008 |
Soft Comput. | 5 |
| 2015 | Lighting Alignment for Image Sequences
Xiaoyue Jiang, Xiaoyi Feng |
ICIG (2) | 2 |
| 2015 | A regularized optimization framework for tag completion and image retrieval
Zhaoqiang Xia, Xiaoyi Feng, Jinye Peng 0001, Jun Wu 0022, Jianping Fan 0001 |
Neurocomputing | 2 |
| 2015 | SISR via trained double sparsity dictionaries
Na Ai, Jinye Peng 0001, Xuan Zhu 0003, Xiaoyi Feng |
Multim. Tools Appl. | 4 |
| 2015 | Automatic tag-to-region assignment via multiple instance learning
Zhaoqiang Xia, Yi Shen 0005, Xiaoyi Feng, Jinye Peng 0001, Jianping Fan 0001 |
Multim. Tools Appl. | 3 |
| 2014 | Cross-modality based celebrity face naming for news image collections
Xueping Su, Jinye Peng 0001, Xiaoyi Feng, Jun Wu 0022, Jianping Fan 0001 |
Multim. Tools Appl. | 3 |
| 2014 | Integrating bilingual search results for automatic junk image filtering
Chunlei Yang, Jinye Peng 0001, Xiaoyi Feng, Jianping Fan 0001 |
Multim. Tools Appl. | 3 |
| 2013 | Multiple Instance Learning for Automatic Image Annotation
Zhaoqiang Xia, Jinye Peng 0001, Xiaoyi Feng, Jianping Fan 0001 |
MMM (2) | 3 |
| 2013 | Multi-label multi-instance learning with missing object tags
Yi Shen 0005, Jinye Peng 0001, Xiaoyi Feng, Jianping Fan 0001 |
Multim. Syst. | 3 |
| 2011 | Efficient Large-Scale Image Data Set Exploration: Visual Concept Network and Image Summarization
Chunlei Yang, Xiaoyi Feng, Jinye Peng 0001, Jianping Fan 0001 |
MMM (2) | 2 |
| 2011 | Towards More Precise Social Image-Tag Alignment
Jinye Peng 0001, Xiaoyi Feng, Jianping Fan 0001 |
MMM (2) | 3 |
| 2007 | Radial Basis Function Neural Network Predictor for Parameter Estimation in Chaotic Noise
Hongmei Xie, Xiaoyi Feng |
ISNN (2) | 2 |
| 2006 | Eyes Location Using a Neural Network
Xiaoyi Feng, Li-ping Yang, Zhi Dang, Matti Pietikäinen |
ISNN (2) | 1 |
| 2005 | A Novel Real Time System for Facial Expression Recognition
Xiaoyi Feng, Matti Pietikäinen, Abdenour Hadid, Hongmei Xie |
ACII | 1 |