EDBT 2026 Demo / reviewers in the wild / expert
Xiyuan Hu
dblp:83/7614
· DBLP profile ↗
62ranked-venue papers
7as first author
34since 2021 · last 2026
0000-0002-7095-6986ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 29 · 1 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Theory of computation · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Milmer: a framework for multiple instance learning based multimodal emotion recognition
Zaitian Wang, Yu Liang 0003, Xiyuan Hu, Tianhao Peng 0002, Jiakai Wang, Weili Zhang, Shuang Niu, Xiaoyang Xie |
Neurocomputing | 4 |
| 2026 | Continuous Shape-to-Texture Face Aging With Flow-Based Prior Latent Age Modulation and Attentional Alignment StyleGANabstractAlbeit recent Generative Models have achieved notable progress in synthesizing realistic facial aging images, many of them, e.g., GAN-based methods, cannot accurately capture the continuous progression of age-related shape-to-texture changes over time. In this paper, we propose an innovative facial age transformation framework that enables the generation of continuous shape-to-texture aging facial images. Firstly, the Prior Latent Age Modulation (PLAM) is designed to leverage the advantages of continuous sampling in high-dimensional space by normalizing flows to achieve precise and reversible mapping between the age attribute variable distributions and the prior latent space, ensuring smooth transitions along with facial aging. Secondly, we introduce the Attentional Feature Fusion (AFF), which dynamically allocates weights to effectively fuse the age attribute features by the latent space manipulation with the content features in StyleGAN, thereby generating facial images that accurately depict facial characteristics from shape to texture corresponding to specific ages. Finally, through quantitative and qualitative analysis of existing datasets, we validate the effectiveness and superiority of our proposed method in facial aging tasks. Xiyuan Hu, Jinglei Qu, Chen Chen 0036, Yu Liang 0003 |
IEEE Trans. Image Process. | 1 |
| 2025 | Masked Spectrum ViT with Cross-Scale Fusion for Universal Deepfake DetectionabstractWith the rapid evolution of deepfake technologies, existing universal detection methods often fail to generalize when faced with new forgery patterns, as they tend to overfit specific artifacts found in the training data. To address this challenge, we propose a detection framework that integrates adaptive spectrum masking and cross-scale feature fusion, termed Masked Spectrum Vision Transformer with Cross-scale Fusion (MSViT-CF). Our approach introduces two key innovations: (1) Multi-band Artifact Enhancement Module (MAEM) reconstructs input images through wavelet decomposition and strategically perturbs high-frequency subbands via adaptive masking, amplifying subtle forgery traces while forcing the model to learn generalized artifact representations; (2) Dynamic Scale Fusion Transformer (DSFT) integrates multi-resolution frequency features through parallel convolutional-transformer pathways, dynamically weighting local spectrum anomalies and global structural inconsistencies. MAEM enhances artifact sensitivity through frequency-space discrepancy learning, while DSFT establishes cross-scale relationships between pixel-level irregularities and semantic-level inconsistencies via learnable attention gates. Experimental results demonstrate that MSViT-CF significantly outperforms existing state-of-the-art methods in detecting deepfake images generated by various GANs and diffusion models, exhibiting superior universality and robustness. Kaiwen Xu, Xiyuan Hu, Chen Chen 0036 |
ECAI | 2 |
| 2025 | A spatio-frequency cross fusion model for deepfake detection and segmentation
Junshuai Zheng, Ning Zhang 0033, Xiyuan Hu, Kaiwen Xu, Dongyang Gao, Zhenmin Tang |
Neurocomputing | 4 |
| 2025 | COD-SAM: Camouflage object detection using SAM
Dongyang Gao, Chen Chen 0036, Xiyuan Hu |
Pattern Recognit. | 5 |
| 2025 | Road Surface State Change Detection Based on Binocular Vision for Autonomous Driving SystemabstractRoad surface condition monitoring is crucial for enhancing transportation safety and efficiency, with applications in autonomous driving and urban infrastructure management. Existing methods often rely on single-camera setups or manual inspections, which are either insufficient for real-time monitoring or labor-intensive. This system focuses on two critical factors: road slope and surface damage, both significantly impacting driving safety and experience, highlighting the need for timely detection. To ensure accuracy and robustness, the system employs a binocular camera for detailed road environment insights and integrates urban sensing techniques. Its hardware deployment processes stereo vision data on embedded platforms, ensuring compatibility with urban IoT networks. This approach surpasses single-camera systems in detecting road surface variations. The research motivation stems from the pressing need to enhance road safety and driving conditions in urban areas. By analyzing binocular camera data and urban sensing technologies, the system offers real-time road condition analysis for effective decision-making. Regarding results, the system showed robust performance in detecting both road slope and surface damage. Slope detection achieved high accuracy with minimal error, and road damage detection reached an overall accuracy of 84%. The system remained stable across diverse conditions, including adverse weather and varying lighting. Liangtian Zhao, Xiangmin Xu 0001, Shanshan Pei, Xiyuan Hu, Qiwei Xie |
ACM Trans. Auton. Adapt. Syst. | 5 |
| 2025 | A Dynamic Zoning Cooperative Control Method for Urban Road Traffic Flow in a Connected Traffic EnvironmentabstractThe existing zoning collaborative control method cannot efficiently adapt to dynamic changes in traffic flow due to the fixed functional zoning of road segments. This study proposes the Dynamic Zoning COllaborative COntrol Method (DZCOM) for urban road traffic flow in a connected traffic environment. First, DZCOM dynamically divides the road segment into two functional zones, vehicle lane-changing and speed-adjustment zones, to integrate different driving behaviors. Second, to achieve rapid lane changes for multiple vehicles, a Multi-vehicle Cooperative Lane-changing Strategy based on Separate Lane-Speed Guidance (MCLS-SLSG) is designed, which rearranges the vehicles entering from the upstream intersection in longitudinal space, creating conditions for left- and right-turn vehicles to change lanes quickly. Finally, the trajectory of connected vehicles and the signal timing of intersections are jointly optimized based on the dynamic programming method to minimize average delay and ride comfort. The simulation results show that optimizing only the longitudinal trajectory of vehicles can reduce the average delay at intersections by 7.9% under moderate traffic volume. The joint optimization of vehicle trajectory and signal timing further reduced the average delay by 21.6%, demonstrating the effectiveness of DZCOM. Further research has shown that higher connected and automated vehicle penetration rates and speed limits can help reduce the average delay. Simultaneously, intersection spacing is another key factor restricting the operational effectiveness of the DZCOM; the method works best in the range of 400–600 m. Xiancai Jiang, Mengying Li, Xiyuan Hu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | C2lRec: Causal Contrastive Learning for User Cold-start Recommendation with Social VariablesabstractEmbedding-based recommender systems rely on historical interactions to model users, which poses challenges for recommending to new users, known as the user cold-start problem. Some approaches incorporate social networks to deduce preferences based on the social circles of cold-start users to solve the problem of sparse features. However, such methods have difficulty distinguishing between superficial correlations and causal relationships in social behaviors, leading to inaccuracies in predicting user preferences. To address the aforementioned issues, we propose the Causal Contrastive Learning Recommendation (C2lRec) framework. Specifically, we causally model the inference of hidden preferences from the feature and historical behavior of warm users and predict user interactions based on such preferences. The counterfactual inference is subsequently performed to intervene and extract interactions from historical behaviors of warm users that influence their preferences, designating as primary causal variables. Additionally, we utilize the primary causal variables from users within the social circle of cold-start users to substitute the missing historical interactions of cold-start users and employ a similar causal modeling approach to uncover hidden preferences as we do with warm users. Finally, we realize causal contrastive learning to enhance the distribution of cold-start users. Extensive experiments conducted on three public datasets demonstrate that the recommendation performance of C2lRec exceeds that of state-of-the-art methods. Xiaolong Xu 0001, Hongsheng Dong, Haolong Xiang, Xiyuan Hu, Xiaoyong Li 0002, Xiaoyu Xia 0001, Xuyun Zhang, Lianyong Qi, Wan-Chun Dou |
ACM Trans. Inf. Syst. | 4 |
| 2025 | Deep Learning to Hash for Time-Aware QoS Prediction Based on VQ-VAEabstractIn mobile edge computing (MEC) environments, with the increase in the number of web services possessing the same or similar functions, the prediction of nonfunctional quality of service (QoS) indicators turns increasingly vital in satisfying the diverse needs of different users. Despite the significant achievements of QoS prediction approaches such as collaborative filtering (CF), they often fail to obtain informative representations due to complex user-service-time QoS data. Moreover, changing network conditions and overlooked temporal factors lead to reduced accuracy. In response to these challenges, this paper proposes a novel deep learning to hash method for time-aware QoS prediction built on a VQ-VAE (namedPred$_{QoS}$QoS). In the proposedPred$_{QoS}$QoS, inspired by advanced CF approaches, we make QoS predictions at a specific moment for the target user through its similar timeslots. Specifically, we train a codebook to discretize the continuous latent representation obtained by the encoder based on vector quantization (VQ), which is motivated by the generative vector-quantized variational autoencoder (VQ-VAE) model, allowing us to derive compact binary codes representing the QoS data. Then, the similarities between QoS data are determined, helping to make effective and efficient predictions. Furthermore, the time-awarePred$_{QoS}$QoSapproach incorporates a temporal factor by training on multiple QoS data (including the user, service, and time dimensions). As a result, the discrete hash codes (i.e., QoS data representations) derived from the encoder can fully uncover the time factor’s dynamic impact on the QoS data, thereby yielding significantly improved prediction performance. In summary, thePred$_{QoS}$QoSapproach learns compact hash codes for original QoS data while taking time into account, enabling accurate predictions to be produced through similar timeslots for the target users. Finally, comprehensive experiments carried out on the real-world WS-DREAM dataset affirm the exceptional performance of thePred$_{QoS}$QoS. Lingzhen Kong, Xiyuan Hu, Lianyong Qi, Xiaolong Xu 0001, Yiwen Zhang 0001, Lina Yao 0001, Xuyun Zhang |
IEEE Trans. Serv. Comput. | 2 |
| 2024 | Global-Local Unified Cross-View Enhancing Framework for Person Re-IdentificationabstractIn Pedestrian Re-identification (ReID) tasks, extracting robust and discriminative features is a key challenge. Recent studies have mainly concentrated on the extraction of features from single images. However, limited research has been done on considering the relationships between different images, which is crucial in the ReID task. To effectively integrate features from multiple perspectives, we propose an explicitly feature enhancing strategy, Global-local unified cross-view enhancing(GLE) framework for person re-identification, based on Transformer among images sharing the same identity. Specifically, our method includes two aspects: (i) Multi-scale global feature enhancement(MSGE): It displays enhancing global features by interactively updating the features of different images sharing the same identity. (ii) Key local feature enhancement(KLFE): It aims to enhance local feature interaction in key areas by selecting and updating top-k patches based on attention scores derived from thermal maps in each training stage. Experimental results demonstrate that our method can extract more robust and discriminative features, achieving state-of-the-art performance on four pedestrian re-identification datasets. Tianran Chen, Chen Chen 0036, Xiyuan Hu |
DSAA | 4 |
| 2024 | Inter-Frame Multiscale Probabilistic Cross-Attention for Surveillance Object DetectionabstractAccurate and robust object detection for surveillance videos hold immense potential applications in the field of public security. However, because of the artifacts in the surveillance video, such as motion blur, noise, low illumination, etc., relying solely on single-frame object detection algorithms may not guarantee the trustiness of surveillance video data analysis. Nevertheless, surveillance videos also exhibit characteristics of stable background and high inter-frame correlation. Therefore, in this paper, we introduce a novel model for surveillance object detection based on vision transformer with inter-frame multiscale probabilistic cross-attention. This model leverages Inter-Frame Semantic Cross-Attention (IFSC) to capture dynamic spatio-temporal features, thereby improving detection performance in low-quality frames. Additionally, it employs Inter-Frame Probabilistic Sparse Cross-Attention (IFPSC) to highlight salient features and suppress background features, enhancing the robustness of surveillance object detection. The experimental results on the UA-DETRAC dataset have demonstrated that the proposed surveillance object detector outperforms other SOTA models and achieves an optimal balance between speed and accuracy. Huanhuan Xu, Xiyuan Hu |
DSAA | 2 |
| 2024 | Deepfake Detection With Combined Unsupervised-Supervised Contrastive LearningabstractThe malicious dissemination of fake images has caused a societal trust crisis, deepfake detection becomes a hot topic now. Through existing detection methods achieve good results in intra-dataset, their performance are poor for unknown manipulations or datasets. To deal with this problem, this paper proposes a new deepfake detection model with combined unsupervised-supervised contrastive learning. By combining unsupervised contrastive learning and supervised contrastive learning with deepfake detection together, the model can discover the essence of fake images from both individual and class features. In addition, a multi-scale attention fusion module is proposed, which helps to enhance the model stability by fusion global and local features of the image. Finally, lots of experiments prove that our method has good performance and generalization ability in intra-dataset, cross-dataset and cross-manipulation scenarios. Junshuai Zheng, Xiyuan Hu, Zhenmin Tang |
ICIP | 3 |
| 2024 | TS-SAM: Two Small Steps for SAM, One Giant Leap for Abnormal detectionsabstractThe advent of pretrained models has significantly advanced the field of artificial intelligence. The introduction of the Segment Anything Model (SAM) has emerged as a pivotal milestone in computer vision, despite its primary application in image segmentation tasks. However, its initial performance in abnormal detection, for instance, camouflaged object detection, shadow detection and forgery detection, was not impressive. To address this, we propose to take two small steps from SAM: an efficient latent high frequency (LHF) feature enhancement module, named "LHF-Adapter", and a robust segmentation metric named "sIoU-Loss", these augmentations are seamlessly integrated into SAM, resulting in the refined model denoted as TS-SAM, tailored specifically to bolster its performance in downstream abnormal detection tasks. Extensive experiments demonstrate that the proposed method outperforms not only SAM but also the most advanced SAM variations. Our method achieves a performance improvement of up to 4.5% in camouflage object detection, setting a new state-of-the-art benchmark. Index Terms—SAM, Transformer, Adapter Dongyang Gao, Chen Chen 0036, Haotian Zhang 0025, Xiyuan Hu |
ICME | 5 |
| 2024 | Illumination Enlightened Spatial-temporal Inconsistency for Deepfake Video DetectionabstractThe rapid advancement of facial manipulation techniques has greatly simplified the creation of deepfake videos, posing a major threat to social safety, public opinions and even political stability. Existing deepfake detection methods primarily concentrate on capturing spatial artifacts or extracting uniform temporal inconsistency, neglecting the potential of exploiting dynamic spatiotemporal inconsistency. To address these issues, this paper proposes a novel network that effectively leverages dynamic spatiotemporal inconsistency, termed DSTI, by integrating the sequential illumination features and intra/inter-frame clues. The proposed DSTI contains two branches: one branch employs a transformer encoder to perform inconsistency computation from sequential illumination representations derived from 3D facial models, including illumination coefficients, 3D normal vectors, and luminance values. The other branch utilizes a timesformer network to capture intra/inter-frame inconsistency from sampled videos. Extensive experimentation validates that the proposed method outperforms other competitive approaches. Kaiyue Tian, Chen Chen 0036, Xiyuan Hu |
ICME | 4 |
| 2024 | 🔭 WATCHER: Wavelet-Guided Texture-Content Hierarchical Relation Learning for Deepfake Detection
Chen Chen 0036, Ning Zhang 0033, Xiyuan Hu |
Int. J. Comput. Vis. | 4 |
| 2024 | Gender Classification Based on Spatio-Frequency Feature Fusion of OCT Fingerprint Images in the IoT EnvironmentabstractIn the rapidly evolving landscape of the Internet of Things (IoT), concerns about privacy and security have become significant as interconnected devices communicate and collaborate. Fingerprints, serving as unique biometric identifiers, play a crucial role in the authentication and identification processes within this interconnected and exchanged network. However, attention is often directed towards the disclosure of visible fingerprints, overlooking latent fingerprints. This is primarily due to the challenges involved in extracting latent fingerprints, especially those remaining on the adhesive side of tape. Traditional methods physically/chemically peel tape to extract these fingerprints, but cause irreversible damage to the tape, hindering accurate fingerprint extraction. In this context, our investigation reveals that Optical Coherence Tomography (OCT) technology allows for the extraction of high-quality OCT fingerprint images from the adhesive side of tape, yielding precise fingerprint recognition and gender classification results. Concretely, we build a novel type of robotic-arm spectral-domain OCT (SD-OCT), which is software-controlled for the movement of the sample arm, making sample scanning more flexible and efficient. Furthermore, we utilize a deep learning network to perform representation learning on OCT fingerprints for the purpose of gender classification. In the first branch, we input OCT fingerprints into an EfficientNet-B3 network to learn their spatial domain features. Simultaneously, in the second branch, we design a network that utilizes Discrete Cosine Transform (DCT) to extract frequency domain features from OCT fingerprints. Ultimately, we integrate the spatial and frequency domain features extracted from OCT fingerprint images to generate comprehensive features. Therefore, in this paper, we introduce a novel Gender Classification approach based on Spatio-Frequency Feature Fusion of OCT Fingerprint Images (named GenClassOCT-SF). The GenClassOCT-SF involves a robotic-arm SD-OCT system for superior-quality fingerprints acquisition and a deep learning network for spatial and frequency domain feature extraction. The fusion of these features enables highly accurate gender classification. Finally, we conduct gender classification experiments on the collected OCT fingerprint dataset to demonstrate the effectiveness of our proposed method. Lingzhen Kong, Kangkang Liu, Xiyuan Hu, Ning Zhang 0033, Lianyong Qi, Xiangrui Li, Xiaokang Zhou |
IEEE Internet Things J. | 3 |
| 2024 | A new deepfake detection model for responding to perception attacks in embodied artificial intelligence
Junshuai Zheng, Xiyuan Hu, Chen Chen 0036, Dongyang Gao, Zhenmin Tang |
Image Vis. Comput. | 2 |
| 2024 | Synchronizing Detection and Removal of Smoke in Endoscopic Images With Cyclic Consistency Adversarial NetsabstractSmoke removal is an important and meaningful issue for endoscopic surgery, which can enhance the visual quality of endoscopic images. Because it is practically impossible to construct a large training dataset of pair-matched endoscopic images with/without smoke, the Generative Adversarial Nets (GANs) based models are usually used for endoscopic image desmoke. But they have difficulties in either locating the accurate smoke area, or recovering realistic internal organ or tissue details. In this paper, we propose a new approach, called Desmoke-CycleGAN, which combined detection, estimation, and removal of smoke together, to improve the CycleGAN model for endoscopic image smoke removal. In addition, both pixel-level and perceptual-level consistency loss have been incorporated in the proposed model, which helps the model to be more stable and efficient for recovering realistic details in endoscopic images. The experimental results have demonstrated that this method outperforms other state-of-the-art smoke removal approaches with unpaired real endoscopic images. Zhisen Hu, Zuxing Xuan, Xiyuan Hu |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | START: Automatic Sleep Staging with Attention-based Cross-modal Learning TransformerabstractAutomatic sleep staging is vital to scale up sleep assessment and diagnosis to serve millions experiencing sleep deprivation and disorders and enable longitudinal sleep monitoring in home environments. However, how to learn from multi-channel raw physiological signal inputs (e.g., EEG and EOG) to capture the sleep stage and physiological signal relations remains a big challenge. In this paper, we propose a sleep staging model, named Sleep Staging Cross-modal Transformer (START), which is a transformer-only method for sleep stage classification. Our model is capable of learning a joint representation from both EEG and EOG signals by using a cross-modal fusion strategy. Experimental results show that our model outperforms the state-of-the-art methods on two public datasets. Furthermore, our model provides considerable reductions in parameters and training time compared to previous methods. Jingpeng Sun, Rongxiao Wang, Gangming Zhao, Chen Chen 0036, Yixiao Qu, Xiyuan Hu, Yizhou Yu |
BIBM | 7 |
| 2023 | Dynamic Graph Learning with Content-guided Spatial-Frequency Relation Reasoning for Deepfake DetectionabstractWith the springing up of face synthesis techniques, it is prominent in need to develop powerful face forgery detection methods due to security concerns. Some existing methods attempt to employ auxiliary frequency-aware information combined with CNN backbones to discover the forged clues. Due to the inadequate information interaction with image content, the extracted frequency features are thus spatially irrelavant, struggling to generalize well on increasingly realistic counterfeit types. To address this issue, we propose a Spatial-Frequency Dynamic Graph method to exploit the relation-aware features in spatial and frequency domains via dynamic graph learning. To this end, we introduce three well-designed components: 1) Content-guided Adaptive Frequency Extraction module to mine the content-adaptive forged frequency clues. 2) Multiple Domains Attention Map Learning module to enrich the spatial-frequency contextual features with multiscale attention maps. 3) Dynamic Graph Spatial-Frequency Feature Fusion Network to explore the high-order relation of spatial and frequency features. Extensive experiments on several benchmark show that our proposed method sustainedly exceeds the state-of-the-arts by a considerable margin. Chen Chen 0036, Xiyuan Hu, Silong Peng |
CVPR | 4 |
| 2023 | Mining Temporal Inconsistency with 3D Face Model for Deepfake Video Detection
Ziyi Cheng, Chen Chen 0036, Xiyuan Hu |
PRCV (7) | 4 |
| 2023 | DT-TransUNet: A Dual-Task Model for Deepfake Detection and Segmentation
Junshuai Zheng, Xiyuan Hu, Zhenmin Tang |
PRCV (7) | 3 |
| 2023 | Balanced knowledge distillation for long-tailed learning
Shaoyu Zhang 0001, Chen Chen 0036, Xiyuan Hu, Silong Peng |
Neurocomputing | 3 |
| 2023 | Fast and accurate super-resolution of MR images based on lightweight generative adversarial network
Zuxing Xuan, Jianpin Zhou, Xiyuan Hu |
Multim. Tools Appl. | 4 |
| 2023 | A landmark-free approach for automatic, dense and robust correspondence of 3D faces
Zhenfeng Fan, Xiyuan Hu, Chen Chen 0036, Xiaolian Wang, Silong Peng |
Pattern Recognit. | 2 |
| 2023 | A Binocular Vision Application in IoT: Realtime Trustworthy Road Condition Detection System in Passable AreaabstractThe structural information detection of road conditions, which is adopted for improving driving comfort, patrol inspection, road maintenance, and accident rescue. In order to improve the trustworthiness of road condition detection, a real-time artificial intelligence road detection system based on binocular vision sensors is investigated in this article. The system is deployed on the low-power edge computing platform, which can upload the processing results to the cloud through the Internet-of-Things devices. The authors use binocular disparity information and image-based lightweight deep segmentation network to enhance the detection robustness and accuracy in the industrial Internet-of-Things application scenarios. Considering the small training dataset, a special data labeling regularization and training strategy have also been proposed for training this network. In addition, we employ multiframes feature matching and measurement data filtering to enhance the measurement accuracy. The experimental results demonstrate that our monocular–binocular fusion framework is robust and efficient. Qiwei Xie, Xiyuan Hu, Lei Ren 0001, Lianyong Qi |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Auditory Receptive Field Net Based Automatic Snore Detection for Wearable DevicesabstractAlthough obstructive sleep apnea and hypopnea syndrome (OSAHS) is a common sleep disease, it is sometimes difficult to be detected in time because of the inconvenience of polysomnography (PSG) examination. Since snoring is one of the earliest symptoms of OSAHS, it can be used for early OSAHS prediction. With the recent development of wearable and IoT sensors, we proposed a deep learning-based accurate snore detection model for long-term home monitoring of snoring during sleep. To enhance the discriminability of features between snoring and non-snoring events, an auditory receptive field (ARF) net was proposed and integrated into the feature extraction network. Based on the feature maps derived by the feature extraction network, the detection model predicted a series of candidate boxes and corresponding confidence scores for each candidate box, which denoted whether the candidate box contained a snore event from the input sound waveforms. A snore detection dataset with a total duration of more than 4600 min was developed to evaluate the proposed model. The experimental results on this dataset revealed that the proposed model outperformed other traditional approaches and deep learning models. Xiyuan Hu, Jingpeng Sun, Jinping Dong, Xuyun Zhang |
IEEE J. Biomed. Health Informatics | 1 |
| 2023 | Progressive Context-Aware Graph Feature Learning for Target Re-IdentificationabstractThis paper aims at robust and discriminative feature learning for target re-identification (Re-ID). In addition to paying attention to the individual appearance information as in most Re-ID methods, we further utilize the abundant contextual information as additional clues to guide the feature learning. Graph as a format of structured data is used to represent the target sample with its context. It describes the first-order appearance information of the samples and the second-order topological relationship information among samples, based on which we compute the feature representation by learning a graph feature embedding. We provide a detailed analysis of graph convolutional network mechanism applied in target Re-ID and propose a novel progressive context-aware graph feature learning method, in which the message passing is dominated by a pre-defined adjacency relationship followed by a learned relationship in a self-adaptive way. The proposed method fully exploits and utilizes contextual information at a low cost for Re-ID. Extensive experiments on five Re-ID benchmarks demonstrate the state-of-the-art performance of the proposed method. Min Cao 0005, Cong Ding 0013, Chen Chen 0036, Hao Dou, Xiyuan Hu, Junchi Yan |
IEEE Trans. Multim. | 5 |
| 2022 | Improving Surveillance Object Detection with Adaptive Omni-Attention over Both Inter-frame and Intra-frame Context
Chen Chen 0036, Xiyuan Hu |
ACCV (2) | 4 |
| 2022 | Regularizing deep networks with label geometry for accurate object localization on small training datasets
Xiaolian Wang, Xiyuan Hu, Chen Chen 0036, Silong Peng |
Pattern Recognit. Lett. | 2 |
| 2021 | Bioinformatics and machine learning methodologies to identify the effects of central nervous system disorders on glioblastoma progressionabstractGlioblastoma (GBM) is a common malignant brain tumor which often presents as a comorbidity with central nervous system (CNS) disorders. Both CNS disorders and GBM cells release glutamate and show an abnormality, but differ in cellular behavior. So, their etiology is not well understood, nor is it clear how CNS disorders influence GBM behavior or growth. This led us to employ a quantitative analytical framework to unravel shared differentially expressed genes (DEGs) and cell signaling pathways that could link CNS disorders and GBM using datasets acquired from the Gene Expression Omnibus database (GEO) and The Cancer Genome Atlas (TCGA) datasets where normal tissue and disease-affected tissue were examined. After identifying DEGs, we identified disease-gene association networks and signaling pathways and performed gene ontology (GO) analyses as well as hub protein identifications to predict the roles of these DEGs. We expanded our study to determine the significant genes that may play a role in GBM progression and the survival of the GBM patients by exploiting clinical and genetic factors using the Cox Proportional Hazard Model and the Kaplan-Meier estimator. In this study, 177 DEGs with 129 upregulated and 48 downregulated genes were identified. Our findings indicate new ways that CNS disorders may influence the incidence of GBM progression, growth or establishment and may also function as biomarkers for GBM prognosis and potential targets for therapies. Our comparison with gold standard databases also provides further proof to support the connection of our identified biomarkers in the pathology underlying the GBM progression. Humayan Kabir Rana, Silong Peng, Xiyuan Hu, Chen Chen 0036, Julian M. W. Quinn, Mohammad Ali Moni |
Briefings Bioinform. | 4 |
| 2021 | Towards effective learning for face super-resolution with shape and pose perturbations
Xiyuan Hu, Zhenfeng Fan, Xu Jia 0012, Xuyun Zhang, Lianyong Qi, Zuxing Xuan |
Knowl. Based Syst. | 1 |
| 2021 | Accurate AM-FM signal demodulation and separation using nonparametric regularization method
Xiyuan Hu, Silong Peng, Baokui Guo |
Signal Process. | 1 |
| 2021 | Progressive Bilateral-Context Driven Model for Post-Processing Person Re-IdentificationabstractMost existing person re-identification methods compute pairwise similarity by extracting robust visual features and learning the discriminative metric. Owing to visual ambiguities, these content-based methods that determine the pairwise relationship only based on the similarity between them, inevitably produce a suboptimal ranking list. Instead, the pairwise similarity can be estimated more accurately along the geodesic path of the underlying data manifold by exploring the rich contextual information of the sample. In this paper, we propose a lightweight post-processing person re-identification method in which the pairwise measure is determined by the relationship between the sample and the counterpart's context in an unsupervised way. We translate the point-to-point comparison into the bilateral point-to-set comparison. The sample's context is composed of its neighbor samples with two different definition ways: the first order context and the second order context, which are used to compute the pairwise similarity in sequence, resulting in a progressive post-processing model. The experiments on four large-scale person re-identification benchmark datasets indicate that (1) the proposed method can consistently achieve higher accuracies by serving as a post-processing procedure after the content-based person re-identification methods, showing its state-of-the-art results, (2) the proposed lightweight method only needs about 6 milliseconds for optimizing the ranking results of one sample, showing its high-efficiency. Code is available at: https://github.com/123ci/PBCmodel. Min Cao 0005, Chen Chen 0036, Hao Dou, Xiyuan Hu, Silong Peng, Arjan Kuijper |
IEEE Trans. Multim. | 4 |
| 2020 | Illuminating Vehicles With Motion Priors For Surveillance Vehicle DetectionabstractVehicle detection in traffic surveillance videos is a special subtask in object detection, where desired objects are vehicles moving on the road while the background is still within a sequence. The disparity of speed within each frame, i.e. moving and static, is consistent with the vehicle and background semantic to some extent, thus motions can be extracted to enhance the appearance of foreground. In this paper, we propose a motion prior embedded parallel architecture for vehicle detection, aiming at illuminating vehicles and suppressing false positives in the background. We further implement extensive experiments on the UA-DETRAC dataset to validate the effectiveness of our approach, and achieve promising performance in both accuracy and speed. Xiaolian Wang, Xiyuan Hu, Chen Chen 0036, Zhenfeng Fan, Silong Peng |
ICIP | 2 |
| 2020 | Deep Top-rank Counter Metric for Person Re-identificationabstractIn the research field of person re-identification, deep metric learning that guides the efficient and effective embedding learning serves as one of the most fundamental tasks. Recent efforts of the loss function based deep metric learning methods mainly focus on the top rank accuracy optimization by minimizing the distance difference between the correctly matching sample pair and wrongly matched sample pair. However, it is more straightforward to count the occurrences of correct top-rank candidates and maximize the counting results for better top rank accuracy. In this paper, we propose a generalized logistic function based metric with effective practicalness in deep learning, namely the“deeptop-rankcountermetric”, to approximately optimize the counted occurrences of the correct top-rank matches. The properties that qualify the proposed metric as a well-suited deep re-identification metric have been discussed and a progressive hard sample mining strategy is also introduced for effective training and performance boosting. The extensive experiments show that the proposed top-rank counter metric outperforms other loss function based deep metrics and achieves the state-of-the-art accuracies. Chen Chen 0036, Hao Dou, Xiyuan Hu, Silong Peng |
ICPR | 3 |
| 2020 | Towards Low-Bit Quantization of Deep Neural Networks with Limited DataabstractRecent machine learning methods use increasingly large deep neural networks to achieve state-of-the-art results in various tasks. Network quantization can effectively reduce computation and memory costs without modifying network structures, facilitating the deployment of deep neural networks (DNNs) on cloud and edge devices. However, most of the existing methods usually need time-consuming training or fine-tuning and access to the original training dataset that may be unavailable due to privacy or security concerns. In this paper, we present a novel method to achieve low-precision quantization with limited data. Firstly, to reduce the complexity of per-channel quantization and degeneration of per-layer quantization, we introduce group quantization that separates the output channels into groups and processes each group independently. Secondly, to better distill knowledge from the pre-trained FP32 model with limited data, we introduce a two-stage knowledge distillation method that divides the optimization process into blockwise optimization and joint optimization to address the limitation of layer-wise supervision and global supervision. Extensive experiments on ImageNet2012 (ResNet18/50, ShuffleNetV2, and MobileNetV2) demonstrate that the proposed approach can significantly improve the quantization model's accuracy when only a few training samples are available. We further show that the method also extends to other computer vision architectures and tasks such as object detection. Chen Chen 0036, Xiyuan Hu, Silong Peng |
ICPR | 3 |
| 2020 | EvoQ: Mixed Precision Quantization of DNNs via Sensitivity Guided Evolutionary SearchabstractNetwork quantization can effectively reduce computation and memory costs without modifying network structures, facilitating the deployment of deep neural networks (DNNs) on edge devices. However, most of the existing methods usually need time-consuming training or fine-tuning and access to the original training dataset that may be unavailable due to privacy or security concerns. In this paper, we introduce a novel method named EvoQ that employs evolutionary search to achieve mixed precision quantization with limited data, which can optimize the resource allocation without adding computation consumption. Considering the shortage of samples and expensive search costs, we use 50 samples to measure the output difference between the quantization model and the pre-trained model for the evaluation of quantization policy, which can save the time obviously while maintaining high accuracy. To improve the search efficiency, we analyze the quantization sensitivity of each layer and utilize the results to optimize the mutation operation. At last, we calibrate the outputs and intermediate features of the quantization model using the selected 50 samples to improve the performance further. We implement extensive experiments on a diverse set of models, including ResNet18/50/101, SqueezeNet, ShuffleNetV2, and MobileNetV2 on ImageNet, as well as SSD-VGG and SSD-ResNet50 on PASCAL VOC. Our method can improve the performance apparently and outperforms the existing post-training quantization methods, demonstrating the effectiveness of EvoQ. Chen Chen 0036, Xiyuan Hu, Silong Peng |
IJCNN | 3 |
| 2020 | PCA-SRGAN: Incremental Orthogonal Projection Discrimination for Face Super-resolutionabstractGenerative Adversarial Networks (GANs) have been employed for face super resolution but they bring distorted facial details easily and still have weakness on recovering realistic texture. To further improve the performance of GAN-based models on super-resolving face images, we propose PCA-SRGAN which pays attention to the cumulative discrimination in the orthogonal projection space spanned by PCA projection matrix of face data. By feeding the principal component projections ranging from structure to details into the discriminator, the discrimination difficulty will be greatly alleviated and the generator can be enhanced to reconstruct clearer contour and finer texture, helpful to achieve the high perception and low distortion eventually. This incremental orthogonal projection discrimination has ensured a precise optimization procedure from coarse to fine and avoids the dependence on the perceptual regularization. We conduct experiments on CelebA and FFHQ face datasets. The qualitative visual effect and quantitative evaluation have demonstrated the overwhelming performance of our model over related works. Hao Dou, Chen Chen 0036, Xiyuan Hu, Zuxing Xuan, Zhisen Hu, Silong Peng |
ACM Multimedia | 3 |
| 2020 | Asymmetric CycleGAN for image-to-image translations with uneven complexities
Hao Dou, Chen Chen 0036, Xiyuan Hu, Libang Jia, Silong Peng |
Neurocomputing | 3 |
| 2019 | Boosting Local Shape Matching for Dense 3D Face CorrespondenceabstractDense 3D face correspondence is a fundamental and challenging issue in the literature of 3D face analysis. Correspondence between two 3D faces can be viewed as a non-rigid registration problem that one deforms into the other, which is commonly guided by a few facial landmarks in many existing works. However, the current works seldom consider the problem of incoherent deformation caused by landmarks. In this paper, we explicitly formulate the deformation as locally rigid motions guided by some seed points, and the formulated deformation satisfies coherent local motions everywhere on a face. The seed points are initialized by a few landmarks, and are then augmented to boost shape matching between the template and the target face step by step, to finally achieve dense correspondence. In each step, we employ a hierarchical scheme for local shape registration, together with a Gaussian reweighting strategy for accurate matching of local features around the seed points. In our experiments, we evaluate the proposed method extensively on several datasets, including two publicly available ones: FRGC v2.0 and BU-3DFE. The experimental results demonstrate that our method can achieve accurate feature correspondence, coherent local shape motion, and compact data representation. These merits actually settle some important issues for practical applications, such as expressions, noise, and partial data. Zhenfeng Fan, Xiyuan Hu, Chen Chen 0036, Silong Peng |
CVPR | 2 |
| 2019 | Asymmetric Cyclegan for Unpaired NIR-to-RGB Face Image TranslationabstractTranslating near-infrared (NIR) face into color (RGB) face, is helpful to improve the visual effect of images and the performance of face recognition. The model for unpaired image-to-image translation is suitable for this task due to the high cost of pixel-matched data. Because of the complexity difference between NIR and RGB image domains, the complexity inequality in bidirectional NIR-RGB translations is significant. We analyze the limitation of the original CycleGAN in asymmetric translation tasks, and propose an Asymmetric Cycle-GAN model with U-net-like generators of unequal sizes to adapt to the asymmetric need in NIR-RGB translations. The edge-retain loss between NIR and the generated RGB images is also introduced to enhance face visual quality. The qualitative visual evaluation and quantitative evaluation with face ID and skin color criteria show that our model achieves great improvements compared with state-of-the-art methods on three public datasets and a newly proposed dataset. Hao Dou, Chen Chen 0036, Xiyuan Hu, Silong Peng |
ICASSP | 3 |
| 2019 | Improving Object Detection with Consistent Negative Sample Mining
Xiaolian Wang, Xiyuan Hu, Chen Chen 0036, Zhenfeng Fan, Silong Peng |
ICONIP (2) | 2 |
| 2019 | TP-ADMM: An Efficient Two-Stage Framework for Training Binary Neural Networks
Chen Chen 0036, Xiyuan Hu, Silong Peng |
ICONIP (4) | 3 |
| 2019 | Plug-and-Play Based Optimization Algorithm for New Crime Density Estimation
Xiangchu Feng, Chen-ping Zhao, Silong Peng, Xiyuan Hu, Zhao-Wei Ouyang |
J. Comput. Sci. Technol. | 4 |
| 2019 | Towards fast and kernelized orthogonal discriminant analysis on person re-identification
Min Cao 0005, Chen Chen 0036, Xiyuan Hu, Silong Peng |
Pattern Recognit. | 3 |
| 2018 | Ranking Loss: A Novel Metric Learning Method for Person Re-identification
Min Cao 0005, Chen Chen 0036, Xiyuan Hu, Silong Peng |
ACCV (2) | 3 |
| 2018 | Dense Semantic and Topological Correspondence of 3D Faces without Landmarks
Zhenfeng Fan, Xiyuan Hu, Chen Chen 0036, Silong Peng |
ECCV (16) | 2 |
| 2018 | Region-specific Metric Learning for Person Re-identificationabstractPerson re-identification addresses the problem of matching individual images of the same person captured by different non-overlapping camera views. Distance metric learning plays an effective role in addressing the problem. With the features extracted on several regions of person image, most of distance metric learning methods have been developed in which the learnt cross-view transformations are region-generic, i.e all region-features share a homogeneous transformation. The spatial structure of person image is ignored and the distribution difference among different region-features is neglected. Therefore in this paper, we propose a novel region-specific metric learning method in which a series of region-specific sub-models are optimized for learning cross-view region-specific transformations. Additionally, we also present a novel feature pre-processing scheme that is designed to improve the features' discriminative power by removing weakly discriminative features. Experimental results on the publicly available VIPeR, PRID450S and QMUL GRID datasets demonstrate that the proposed method performs favorably against the state-of-the-art methods. Min Cao 0005, Chen Chen 0036, Xiyuan Hu, Silong Peng |
ICPR | 3 |
| 2017 | Complex-valued differential operator-based method for multi-component signal separation
Baokui Guo, Silong Peng, Xiyuan Hu |
Signal Process. | 3 |
| 2016 | Face spoofing detection based on 3D lighting environment analysis of image pairabstractIn this paper, we present a novel face spoofing detection method based on 3D lighting environment analysis of an image pair collected before and after the lighting environment change. Our idea is inspired from the unimpressive fact that the illumination distributions of the internal spoof face stays stable under the protection of the photo and screen plane, while that of a exposed genuine face changes accordingly to different lighting environment due to a natural response of 3D structure. After estimating two sets of lighting environment coefficients of client's face image pair with the hand of 3D Morphable Model (3DMM) and Sphere Harmonic Illumination Model (SHIM), robust liveness judgement is conducted by hypothesis tests. Experimental results show the effectiveness of proposed method on multiple kinds of face attacks including printed photo, screen photo, and video replay attack, and other advantages such as user cooperation free, loose using conditions, simple equipment demand, easy to camouflage and propitious to face recognition. Xiyuan Hu, Chen Chen 0036, Silong Peng |
ICPR | 2 |
| 2016 | Sharp image estimation from a depth-involved motion-blurred image
Yuquan Xu, Xiyuan Hu, Silong Peng |
Neurocomputing | 2 |
| 2016 | A lighting robust fitting approach of 3D morphable model for face reconstruction
Silong Peng, Xiyuan Hu |
Vis. Comput. | 3 |
| 2015 | Adaptive Integral Operators for Signal SeparationabstractThe operator-based signal separation approach uses an adaptive operator to separate a signal into a set of additive subcomponents. In this paper, we show that differential operators and their initial and boundary values can be exploited to derive corresponding integral operators. Although the differential operators and the integral operators have the same null space, the latter are more robust to noisy signals. Moreover, after expanding the kernels of Frequency Modulated (FM) signals via eigen-decomposition, the operator-based approach with the integral operator can be regarded as the matched filter approach that uses eigen-functions as the matched filters. We then incorporate the integral operator into the Null Space Pursuit (NSP) algorithm to estimate the kernel and extract the subcomponent of a signal. To demonstrate the robustness and efficacy of the proposed algorithm, we compare it with several state-of-the-art approaches in separating multiple-component synthesized signals and real-life signals. Xiyuan Hu, Silong Peng, Wen-Liang Hwang |
IEEE Signal Process. Lett. | 1 |
| 2014 | A Lighting Robust Fitting Approach of 3D Morphable Model Using Spherical Harmonic Illuminationabstract3D morph able model (3DMM) is a powerful tool to recover 3D shape and texture from a single facial image. Its foundation consists of three models (i.e. face, camera, and illumination) which can simulate the formulation process of facial images. In this paper, we adopt a new illumination model, the Sphere Harmonic Illumination Model (SHIM), to the 3DMM fitting process. The new illumination model takes more lighting factors into consideration than the Phong's model. Then, we use a new optimization algorithm to optimize the shape and texture parameters simultaneously under SHIM. Compared with the the existing methods that used SHIM to recover only texture, both the shape and texture recovered by our algorithm are improved. The experiments on he CMU-PIE database also show that, compared to other state-of-the-art methods based on the Phong's model, the proposed approach enhances the robustness of the fitting of 3DMM against lighting variations. Xiyuan Hu, Yuquan Xu, Silong Peng |
ICPR | 2 |
| 2013 | An integral operator based adaptive signal separation approachabstractThe operator-based signal separation approach uses an adaptive operator to separate a signal into additive subcomponents. And different types of operator can depict different properties of a signal. In this paper, we define a new kind of integral operator which can be derived from the second kind of Fredholm integral equation. Then, we analyze the properties of the proposed integral operator and discuss its relation to the second condition of Intrinsic Mode Function (IMF). To demonstrate the robustness and efficacy of the proposed operator, we incorporate it into the Null Space Pursuit algorithm to separate several multicomponent signals, including a real-life signal. Xiyuan Hu, Silong Peng, Wen-Liang Hwang |
ICASSP | 1 |
| 2013 | An operator-based and sparsity-based approach to adaptive signal separationabstractAn operator-based and sparsity-based approach is proposed to adaptively separate a signal into additive subcomponents. The proposed approach can be formulated as an optimization problem. Since the design of the operator can be adaptively customized to the target signal, we can propose different types of operators for different types of signals. The subcomponents are a kind of local narrow band signals in the null space of an adaptive operator and a residual signal which is a sparse signal in some sense. Our experiments, including simulated signals and a real-life signal, demonstrate the efficacy and accuracy of the proposed approach. Xiaolei Yi, Xiyuan Hu, Silong Peng |
ICASSP | 2 |
| 2012 | Single-Image Blind Deblurring for Non-uniform Camera-Shake Blur
Yuquan Xu, Xiyuan Hu, Silong Peng |
ACCV (3) | 3 |
| 2012 | Single image blind deblurring with image decompositionabstractHow to deal with themotion blurred image is a common problem in our daily life. Restoring blurred images is challenging, especially when both the blur kernel and the sharp image are unknown. In this work, we present a new algorithm for removing motion blur from a single image, which incorporates the image decomposition into the image deblurring process. Most of the existing algorithms solving the blind deblurring problem use the alternate iterative mechanism, which alternative estimates the kernel and restores the sharp image. We find that the small gradients of image are not always helpful but sometimes harmful to this kind of iterative algorithm. So we decompose the blurred image into cartoon and texture components. And we only use the cartoon part of the image, which can improve the stability and robustness of the algorithm. Our experiments show that our algorithms can achieve good results in man-made and real-life photos. Yuquan Xu, Xiyuan Hu, Silong Peng |
ICASSP | 2 |
| 2012 | Hyperspectral Imagery Denoising Using a Spatial-Spectral Domain Mixing Prior
Shaolin Chen, Xiyuan Hu, Silong Peng |
J. Comput. Sci. Technol. | 2 |
| 2011 | Multiple component predictive coding framework of still imagesabstractIn this paper, we propose a multiple component predictive coding framework. We firstly separate the reconstructed image into several subcomponents; and then predict each subcomponent independently but encode them together. To separate image into multiple subcomponents, we also propose a fast operator-based image separation algorithm. With the help of multicomponent prediction strategy, our prediction results can achieve superior performance than the H.264/AVC intra frame prediction method for images containing rich textures. By adopting the residue coding method used in H.264/AVC, we compare the compression efficacy of our proposed algorithm with the state-of-art JPEG2000 and H.264/AVC intra frame compression algorithms in the experimental part. The numerical results show that our algorithm is better than both H.264/AVC intra frame coding algorithm and JPEG2000 algorithm for images with ample textures. Xiyuan Hu, Weiping Xia, Silong Peng, Wen-Liang Hwang |
ICME | 1 |
| 2010 | Single color image dehazing using sparse priorsabstractWe present a new algorithm for removing the haze effects from a single color image. By introducing an additive noise argument in the degradation image model, we establish a unified probabilistic framework for the clear day image and the atmosphere transmission. Then we use an alternative optimization method to approximate the MAP estimators of these two variables iteratively. Experimental results demonstrate the efficiency of the proposed method on restoring the true scene colors and contrast. Xiyuan Hu, Silong Peng, Duo-Chao Wang |
ICIP | 2 |