EDBT 2026 Demo / reviewers in the wild / expert
Cunjian Chen
dblp:73/2740
· DBLP profile ↗
30ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0002-2926-9762ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 12 since 2021Security and privacy · 7 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mamba-Driven Multi-View Discriminative Clustering via Global-Local Cross-View Sequence ModelingabstractMulti-view clustering (MVC) has recently garnered increasing attention for its ability to partition unlabeled samples into distinct clusters by leveraging complementary and consistent information from different views. Existing MVC methods primarily combine deep neural networks with contrastive learning for cross-view representation learning, yet often overlook the inherent global-local structural relationships among samples. While GNN-based methods capture local structures, they struggle to model global dependencies, leading to inferior inter-cluster separability. In contrast, Transformer-based methods excel at global aggregation but suffer from quadratic complexity, and their attention smoothing effect weakens fine-grained local structures, resulting in suboptimal intra-cluster compactness. To address these limitations, we propose a novel end-to-end MVC framework called Mamba-Driven Multi-View Discriminative Clustering via Global-Local Cross-View Sequence Modeling (MGLC). By flexibly constructing multi-view sequences, MGLC fully exploits the efficient sequence modeling capabilities of Mamba to jointly model cross-view dependencies and global-local structural relationships among samples. Furthermore, MGLC introduces a Cross-Mamba Fusion module to dynamically integrate cross-view and global-local structural representations. Additionally, MGLC incorporates a Dual Calibration Contrastive Learning module, guided by high-confidence pseudo-labels, that adaptively refines both feature and semantic representations while mitigating false negatives among semantically similar samples. Extensive comparative experiments and ablation studies demonstrate the effectiveness of MGLC. Yuanyang Zhang, Xinhang Wan, Jie Xu 0044, Cunjian Chen, Tien-Tsin Wong, Li Yao 0003, Yijie Lin 0001 |
AAAI | 5 |
| 2026 | IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation
Guibao Shen, Quande Liu, Jialin Gao, Lan Du 0002, Cunjian Chen, Chi-Wing Fu, Xiaowei Hu 0001, Pheng-Ann Heng |
AAAI | 8 |
| 2026 | BlendEmo: Structured Set Prediction and Pair-Conditioned Ratio Modeling for Blended Emotion Recognition
Xiaohaoyang Lei, Cunjian Chen |
FG | 5 |
| 2026 | GlassSplat: Geometric Consistency and Pruning for Reflection-Free 3D Scene ReconstructionabstractRendering high-fidelity 3D scenes is crucial for immersive applications like virtual reality and digital twins. However, standard 3D Gaussian Splatting (3DGS) relies heavily on multi-view consistency, making it fragile in real-world scenarios plagued by glass reflections. These reflections often manifest as geometric "floaters" or severe texture artifacts, obscuring the true background. Existing solutions, which typically employ single-image priors or NeRF-based in-painting, often lack explicit 3D constraints or rely on synthetic data, failing to generalize to complex environments. To address these challenges, we first present a novel benchmark dataset of 8 real-world scenes, capturing physically paired reflective and reflection-free images. Building on this, we propose GlassSplat, a robust framework designed to eliminate view-dependent artifacts and recover clean transmission geometry. Our method initializes with a reflection prior and introduces an Affine-Based Exposure Correction module to align global photometric inconsistencies. To distinguishing valid geometry from virtual outliers, we incorporate an Epipolar Consistency Loss and an uncertainty-weighted Depth Regularization. Finally, to physically purge residual noise, we devise a Visibility-Aware Pruning strategy that dynamically filters artifacts based on multi-view statistics. Extensive experiments demonstrate that GlassSplat significantly outperforms state-of-the-art approaches, effectively recovering a clean, artifact-free 3D scene representation. Jingjiao You, Yuanyang Zhang, Yining Xu 0001, Li Yao 0003, Cunjian Chen, Tien-Tsin Wong |
ICMR | 6 |
| 2026 | Consistent and Controllable Image Animation With Linear Motion Diffusion TransformersabstractImage animation has seen significant progress, driven by the powerful generative capabilities of diffusion models. However, maintaining appearance consistency with static input images and mitigating abrupt motion transitions in generated animations remain persistent challenges. While text-to-video (T2V) generation has demonstrated impressive performance with diffusion transformer models, the image animation field still largely relies on U-Net-based diffusion models, which lag behind the latest T2V approaches. Moreover, the quadratic complexity of vanilla self-attention mechanisms in Transformers imposes heavy computational demands, making image animation particularly resource-intensive. To address these issues, we propose MiraMo, a framework designed to enhance efficiency, appearance consistency, and motion smoothness in image animation. Specifically, MiraMo introduces three key elements: (1) A foundational text-to-video architecture replacing vanilla self-attention with efficient linear attention to reduce computational overhead while preserving generation quality; (2) A novel motion residual learning paradigm that focuses on modeling motion dynamics rather than directly predicting frames, improving temporal consistency; and (3) A DCT-based noise refinement strategy during inference to suppress sudden motion artifacts, complemented by a dynamics control module to balance motion smoothness and expressiveness. Extensive experiments against state-of-the-art methods validate the superiority of MiraMo in generating consistent, smooth, and controllable animations with accelerated inference speed. Additionally, we demonstrate the versatility of MiraMo through applications in motion transfer and video editing tasks. Xin Ma 0031, Yaohui Wang 0001, Genyun Jia, Tien-Tsin Wong, Cunjian Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Consistent and Controllable Image Animation with Motion Diffusion ModelsabstractDiffusion models have achieved significant progress in the task of image animation due to their powerful generative capabilities. However, preserving appearance consistency to the static input image, and avoiding abrupt motion change in the generated animation, remains challenging. In this paper, we introduce Cinemo, a novel image animation approach that aims at achieving better appearance consistency and motion smoothness. The core of Cinemo is to focus on learning the distribution of motion residuals, rather than directly predicting frames as in existing diffusion models. During the inference, we further mitigate the sudden motion changes in the generated video by introducing a novel DCT-based noise refinement strategy. To counteract the over-smoothing of motion, we introduce a dynamics degree control design for better control of the magnitude of motion. Altogether, these strategies enable Cinemo to produce highly consistent, smooth, and motion-controllable results. Extensive experiments compared with several state-of-the-art methods demonstrate the effectiveness and superiority of our proposed approach. In the end, we also demonstrate how our model can be applied for motion transfer or video editing of any given video. The project page is available at https://maxin-cn.github.io/cinemo_project/. Xin Ma 0031, Yaohui Wang 0001, Gengyun Jia, Tien-Tsin Wong, Yuan-Fang Li, Cunjian Chen |
CVPR | 7 |
| 2025 | SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene ConsistencyabstractRecent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus on character consistency, they often overlook the crucial role of scenes in storytelling, which restricts their creativity in practice. This paper introduces scene-oriented story generation, addressing two key challenges: (i) scene planning, where current methods fail to ensure scene-level narrative coherence by relying solely on text descriptions, and (ii) scene consistency, which remains largely unexplored in terms of maintaining scene consistency across multiple stories. We propose SceneDecorator, a training-free framework that employs VLM-Guided Scene Planning to ensure narrative coherence across different scenes in a ``global-to-local'' manner, and Long-Term Scene-Sharing Attention to maintain long-term scene consistency and subject diversity across generated stories. Extensive experiments demonstrate the superior performance of SceneDecorator, highlighting its potential to unleash creativity in the fields of arts, films, and games. Quanjian Song, Fei Shen 0004, Xiaowei Hu 0001, Cunjian Chen, Pheng-Ann Heng |
NeurIPS | 7 |
| 2025 | Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention ShiftingabstractThe increase in computing power and the necessity of AI-assisted decision-making boost the growing application of large language models (LLMs). Along with this, the potential retention of sensitive data of LLMs has spurred increasing research into machine unlearning. However, existing unlearning approaches face a critical dilemma: Aggressive unlearning compromises model utility, while conservative strategies preserve utility but risk hallucinated responses. This significantly limits LLMs' reliability in knowledge-intensive applications. To address this, we introduce a novel Attention-Shifting (AS) framework for selective unlearning. AS is driven by two design objectives: (1) context-preserving suppression that attenuates attention to fact-bearing tokens without disrupting LLMs' linguistic structure; and (2) hallucination-resistant response shaping that discourages fabricated completions when queried about unlearning content. AS realizes these objectives through two attention-level interventions, which are importance-aware suppression applied to the unlearning set to reduce reliance on memorized knowledge and attention-guided retention enhancement that reinforces attention toward semantically essential tokens in the retained dataset to mitigate unintended degradation. These two components are jointly optimized via a dual-loss objective, which forms a soft boundary that localizes unlearning while preserving unrelated knowledge under representation superposition. Experimental results show that AS improves performance preservation over the state-of-the-art unlearning methods, achieving up to 15\% higher accuracy on the ToFU benchmark and 10\% on the TDEC benchmark, while maintaining competitive hallucination-free unlearning effectiveness. Compared to existing methods, AS demonstrates a superior balance between unlearning effectiveness, generalization, and response reliability. Chenchen Tan, Youyang Qu, Xinghao Li, Shujie Cui, Cunjian Chen, Longxiang Gao |
NeurIPS | 6 |
| 2025 | LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models
Yaohui Wang 0001, Xin Ma 0031, Shangchen Zhou, Yi Wang 0074, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang 0001, Yuwei Guo 0002, Tianxing Wu 0002, Chenyang Si, Yuming Jiang 0003, Cunjian Chen, Chen Change Loy, Bo Dai 0002, Dahua Lin, Yu Qiao 0001, Ziwei Liu 0002 |
Int. J. Comput. Vis. | 15 |
| 2025 | LEO: Generative Latent Image Animator for Human Video Synthesis
Yaohui Wang 0001, Xin Ma 0031, Cunjian Chen, Antitza Dantcheva, Bo Dai 0002, Yu Qiao 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | Beyond the visible: A survey on cross-spectral face recognition
David Anghelone, Cunjian Chen, Arun Ross, Antitza Dantcheva |
Neurocomputing | 2 |
| 2025 | REHRSeg: Unleashing the power of self-supervised super-resolution for resource-efficient 3D MRI segmentation
Zhiyun Song, Yinjie Zhao, Manman Fei, Xiangyu Zhao 0003, Mengjun Liu, Cunjian Chen, Chung-Hsing Yeh, Qian Wang 0001, Guoyan Zheng, Songtao Ai, Lichi Zhang |
Neurocomputing | 7 |
| 2025 | 3D surface reconstruction with enhanced high-frequency details
Shikun Zhang, Yiqun Wang 0001, Cunjian Chen, Yong Li 0023, Qiuhong Ke |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | MSFM-UNET: enhancing medical image segmentation with multi-scale and multi-view frequency fusion
Qiang Gao 0017, Yi Wang 0074, Feiyan Zhou, Yong Li 0023, Bin Fang 0001, Lan Du 0002, Cunjian Chen |
Pattern Anal. Appl. | 9 |
| 2025 | Iterative Window Mean Filter: Thwarting Diffusion-Based Adversarial PurificationabstractFace authentication systems have brought significant convenience and advanced developments, yet they have become unreliable due to their sensitivity to inconspicuous perturbations, such as adversarial attacks. Existing defenses often exhibit weaknesses when facing various attack algorithms and adaptive attacks or compromise accuracy for enhanced security. To address these challenges, we have developed a novel and highly efficient non-deep-learning-based image filter called the Iterative Window Mean Filter (IWMF) and proposed a new framework for adversarial purification, named IWMF-Diff, which integrates IWMF and denoising diffusion models. These methods can function as pre-processing modules to eliminate adversarial perturbations without necessitating further modifications or retraining of the target system. We demonstrate that our proposed methodologies fulfill four critical requirements: preserved accuracy, improved security, generalizability to various threats in different settings, and better resistance to adaptive attacks. This performance surpasses that of the state-of-the-art adversarial purification method, DiffPure. Our code is released athttps://github.com/azrealwang/iwmfdiff. Hanrui Wang 0005, Ruoxi Sun 0001, Cunjian Chen, Minhui Xue 0001, Lay-Ki Soon, Shuo Wang 0012, Zhe Jin 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | A Gray-Box Attack Against Latent Diffusion Model-Based Image Editing by Posterior CollapseabstractRecent advancements in Latent Diffusion Models (LDMs) have revolutionized image synthesis and manipulation, raising significant concerns about data misappropriation and intellectual property infringement. While adversarial attacks have been extensively explored as a protective measure against such misuse of generative AI, current approaches are severely limited by their heavy reliance on model-specific knowledge and substantial computational costs. Drawing inspiration from the posterior collapse phenomenon observed in VAE training, we propose the Posterior Collapse Attack (PCA), a novel framework for protecting images from unauthorized manipulation. Through comprehensive theoretical analysis and empirical validation, we identify two distinct collapse phenomena during VAE inference: diffusion collapse and concentration collapse. Based on this discovery, we design a unified loss function that can flexibly achieve both types of collapse through parameter adjustment, each corresponding to different protection objectives in preventing image manipulation. Our method significantly reduces dependence on model-specific knowledge by requiring access to only the VAE encoder, which constitutes less than 4% of LDM parameters. Notably, PCA achieves prompt-invariant protection by operating on the VAE encoder before text conditioning occurs, eliminating the need for empty prompt optimization required by existing methods. This minimal requirement enables PCA to maintain adequate transferability across various VAE-based LDM architectures while effectively preventing unauthorized image editing. Extensive experiments show PCA outperforms existing techniques in protection effectiveness, computational efficiency (runtime and VRAM), and generalization across VAE-based LDM variants. Our code is available at https://github.com/ZhongliangGuo/PosteriorCollapseAttack. Zhongliang Guo 0001, Chun Tong Lei, Lei Fang 0001, Shuai Zhao 0007, Yifei Qian, Zeyu Wang 0010, Cunjian Chen, Ognjen Arandjelovic, Chun Pong Lau 0001 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2024 | DiffTV: Identity-Preserved Thermal-to-Visible Face Translation via Feature Alignment and Dual-Stage ConditionsabstractThe thermal-to-visible (T2V) face translation task is essential for enabling face verification in low-light or dark conditions by converting thermal infrared faces into their visible counterparts. However, this task faces two primary challenges. First, the inherent differences between the modalities hinder the effective use of thermal information to guide RGB face reconstruction. Second, translated RGB faces often lack the identity details of the corresponding visible faces, such as skin color. To tackle these challenges, we introduce DiffTV, the first Latent Diffusion Model (LDM) specifically designed for T2V facial image translation with a focus on preserving identity. Our approach proposes a novel heterogeneous feature alignment strategy that bridges the modal gap and extracts both coarse-and fine-grained identity features consistent with visible images. Furthermore, a dual-stage condition injection strategy introduces control information to guide identity-preserved translation. Experimental results demonstrate the superior performance of DiffTV, particularly in scenarios where maintaining identity integrity is critical. Guiqin Zhao, Guoli Wang 0004, Zejin Wang, Antitza Dantcheva, Lan Du 0002, Cunjian Chen |
ACM Multimedia | 8 |
| 2024 | Uncertainty-aware image inpainting with adaptive feedback networkabstractWhile most image inpainting methods perform well on small image defects, they still struggle to deliver satisfactory results on large holes due to insufficient image guidance. To address this challenge, this paper proposes an uncertainty-aware adaptive feedback network (U2AFN), which incorporates an adaptive feedback mechanism to refine inpainting regions progressively. U2AFN predicts both an uncertainty map and an inpainting result simultaneously. During each iteration, the adaptive integration feedback block utilizes inpainting pixels with low uncertainty to guide the subsequent learning iteration. This process leads to a gradual reduction in uncertainty and produces more reliable inpainting outcomes. Our approach is extensively evaluated and compared on multiple datasets, demonstrating its superior performance over existing methods. The code is available at: https://codeocean.com/capsule/1901983/tree. Xin Ma 0031, Xiaoqiang Zhou, Huaibo Huang, Gengyun Jia, Yaohui Wang 0001, Cunjian Chen |
Expert Syst. Appl. | 7 |
| 2024 | A Multi-Task Adversarial Attack against Face AuthenticationabstractDeep learning-based identity management systems, such as face authentication systems, are vulnerable to adversarial attacks. However, existing attacks are typically designed for single-task purposes, which means they are tailored to exploit vulnerabilities unique to the individual target rather than being adaptable for multiple users or systems. This limitation makes them unsuitable for certain attack scenarios, such as morphing, universal, transferable, and counterattacks. In this article, we propose a multi-task adversarial attack algorithm called MTADV that are adaptable for multiple users or systems. By interpreting these scenarios as multi-task attacks, MTADV is applicable to both single- and multi-task attacks, and feasible in the white- and gray-box settings. Furthermore, MTADV is effective against various face datasets, including LFW, CelebA, and CelebA-HQ, and can work with different deep learning models, such as FaceNet, InsightFace, and CurricularFace. Importantly, MTADV retains its feasibility as a single-task attack targeting a single user/system. To the best of our knowledge, MTADV is the first adversarial attack method that can target all of the aforementioned scenarios in one algorithm. Hanrui Wang 0005, Shuo Wang 0012, Cunjian Chen, Massimo Tistarelli, Zhe Jin 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | TFLD: Thermal Face and Landmark Detection for Unconstrained Cross-spectral Face RecognitionabstractAutomated thermal-to-visible face recognition has received increased attention due to benefits related to low-light applications. Towards improvement of related matching accuracy, we hereby present TFLD, a detector of face and landmarks operating in the thermal spectrum. Our proposed TFLD is based on the architecture of YOLOv5, integrating sequential modules for face and landmark detection. We introduce a thermal face restoration scheme, in order to enhance thermal image quality and hence detection accuracy. We address data scarcity by transferring landmarks in paired visible and thermal images. Our experimental results showcase that our proposed detector accurately detects faces, as well as landmarks in a wide range of adversarial conditions. Further, TFLD achieves promising results on three benchmark multi-spectral face and landmark datasets, namely ARL-VTF, SF-TL54 and RWTH-Aachen; thereby improving the matching accuracy in cross-spectral face recognition by providing robust face alignment based on estimated facial landmarks. David Anghelone, Sarah Lannes, Valeriya Strizhkova, Philippe Faure, Cunjian Chen, Antitza Dantcheva |
IJCB | 5 |
| 2022 | Attention-Guided Generative Adversarial Network for Explainable Thermal to Visible Face RecognitionabstractThermal to visible face image translation aims at synthesizing high-fidelity visible face images from thermal counterparts, placing emphasis on preserving the identity of the faces. While remarkable progress has been achieved related to the quality of synthetic images, as well as related to associated face matching accuracy, interpreting the generation process from thermal to visible face images remains an open challenge. Towards tackling this challenge, we present a novel generic attention-guided generative adversarial network (AG-GAN) for thermal to visible image translation. The AG-GAN framework is based on an encoder network that directly generates attention feature maps from an input thermal image in either, supervised or unsupervised fashion. A decoder network takes the attention maps and applies adaptive layer-instance normalization, in order to reconstruct the corresponding visible image. We show that solving thermal to visible image translation tasks through AG-GAN significantly improves the cross-spectral face matching accuracy, as well as inherently supports model explanation. Cunjian Chen, David Anghelone, Philippe Faure, Antitza Dantcheva |
IJCB | 1 |
| 2021 | Explainable Thermal to Visible Face Recognition Using Latent-Guided Generative Adversarial NetworkabstractOne of the main challenges in performing thermal-to-visible face image translation is preserving the identity across different spectral bands. Existing work does not effectively disentangle the identity from other confounding factors. In this paper, we propose a Latent-Guided Generative Adversarial Network (LG-GAN) to explicitly decompose an input image into identity code that is spectral-invariant and style code that is spectral-dependent. By using such a disentanglement, we are able to analyze the identity preservation by interpreting and visualizing the identity code. We present extensive face recognition experiments on two challenging Visible-Thermal face datasets. We show that the learned identity code is effective in preserving the identity, thus offering useful insights on interpreting and explaining thermal-to-visible face image translation. David Anghelone, Cunjian Chen, Philippe Faure, Arun Ross, Antitza Dantcheva |
FG | 2 |
| 2021 | Similarity-based Gray-box Adversarial Attack Against Deep Face RecognitionabstractThe majority of adversarial attack techniques perform well against deep face recognition when the full knowledge of the system is revealed (white-box). However, such techniques act unsuccessfully in the gray-box setting where the face templates are unknown to the attackers. In this work, we propose a similarity-based gray-box adversarial attack (SGADV) technique with a newly developed objective function. SGADV utilizes the dissimilarity score to produce the optimized adversarial example, i.e., similarity-based adversarial attack. This technique applies to both white-box and gray-box attacks against authentication systems that determine genuine or imposter users using the dissimilarity score. To validate the effectiveness of SGADV, we conduct extensive experiments on face datasets of LFW, CelebA, and CelebA-HQ against deep face recognition models of FaceNet and InsightFace in both white-box and gray-box settings. The results suggest that the proposed method significantly outperforms the existing adversarial attack techniques in the gray-box setting. We hence summarize that the similarity-base approaches to develop the adversarial example could satisfactorily cater to the gray-box attack scenarios for de-authentication. Hanrui Wang 0003, Shuo Wang 0012, Zhe Jin 0001, Yandan Wang, Cunjian Chen, Massimo Tistarelli |
FG | 5 |
| 2020 | Iris Liveness Detection Competition (LivDet-Iris) - The 2020 EditionabstractLaunched in 2013, LivDet-Iris is an international competition series open to academia and industry with the aim to assess and report advances in iris Presentation Attack Detection (PAD). This paper presents results from the fourth competition of the series: LivDet-Iris 2020. This year's competition introduced several novel elements: (a) incorporated new types of attacks (samples displayed on a screen, cadaver eyes and prosthetic eyes), (b) initiated LivDet-Iris as an on-going effort, with a testing protocol available now to everyone via the Biometrics Evaluation and Testing (BEAT)* open-source platform to facilitate reproducibility and benchmarking of new algorithms continuously, and (c) performance comparison of the submitted entries with three baseline methods (offered by the University of Notre Dame and Michigan State University), and three open-source iris PAD methods available in the public domain. The best performing entry to the competition reported a weighted average APCER of 59.10% and a BPCER of 0.46% over all five attack types. This paper serves as the latest evaluation of iris PAD on a large spectrum of presentation attack instruments. Priyanka Das 0004, Joseph McGrath, Zhaoyuan Fang, Aidan Boyd, Ganghee Jang, Amir Mohammadi, Sandip Purnapatra, David Yambay, Sébastien Marcel, Mateusz Trokielewicz, Piotr Maciejewicz, Kevin W. Bowyer, Adam Czajka, Stephanie Schuckers, Juan E. Tapia, Meiling Fang, Naser Damer, Fadi Boutros, Arjan Kuijper, Renu Sharma, Cunjian Chen, Arun Ross |
IJCB | 22 |
| 2020 | Relativistic Discriminator: A One-Class Classifier for Generalized Iris Presentation Attack DetectionabstractIris based recognition systems are vulnerable to presentation attacks (PAs) where artifacts such as cosmetic contact lenses, artificial eyes and printed eyes can be used to fool the system. While many learning-based algorithms have been proposed to detect such attacks, very few are equipped to handle previously unseen or newly constructed PAs. In this research, we propose a presentation attack detection (PAD) method that utilizes a discriminator that is trained to distinguish between bonafide iris images and synthetically generated iris images. We hypothesize that such a discriminator will generate a tight boundary around the bonafide samples. This would allow the discriminator to better separate the bonafide samples from all types of PA samples. For generating synthetic irides, we train the Relativistic Average Standard Generative Adversarial Network (RaSGAN) that has been shown to generate higher resolution and better quality images than standard GANs. The relativistic discriminator (RD) component of the trained RaS-GAN is then appropriated for PA detection and is referred to as RD-PAD. Experimental results convey the efficacy of the RD-PAD as a one-class anomaly detector. Shivangi Yadav, Cunjian Chen, Arun Ross |
WACV | 2 |
| 2019 | Matching Thermal to Visible Face Images Using a Semantic-Guided Generative Adversarial NetworkabstractDesigning face recognition systems that are capable of matching face images obtained in the thermal spectrum with those obtained in the visible spectrum is a challenging problem. In this work, we propose the use of semantic-guided generative adversarial network (SG-GAN) to automatically synthesize visible face images from their thermal counterparts. Specifically, semantic labels, extracted by a face parsing network, are used to compute a semantic loss function to regularize the adversarial network during training. These semantic cues denote high-level facial component information associated with each pixel. Further, an identity extraction network is leveraged to generate multi-scale features to compute an identity loss function. To achieve photo-realistic results, a perceptual loss function is introduced during network training to ensure that the synthesized visible face is perceptually similar to the target visible face image. We extensively evaluate the benefits of individual loss functions, and combine them effectively to learn the mapping from thermal to visible face images. Experiments involving two multispectral face datasets show that the proposed method achieves promising results in both face synthesis and cross-spectral face matching. Cunjian Chen, Arun Ross |
FG | 1 |
| 2016 | Matching thermal to visible face images using hidden factor analysis in a cascaded subspace learning framework
Cunjian Chen, Arun Ross |
Pattern Recognit. Lett. | 1 |
| 2011 | Can facial metrology predict gender?abstractWe investigate the question of whether facial metrology can be exploited for reliable gender prediction. A new method based solely on metrological information from facial landmarks is developed. Here, metrological features are defined in terms of specially normalized angle and distance measures and computed based on given landmarks on facial images. The performance of the proposed metrology- based method is compared with that of a state-of-the-art appearance-based method for gender classification. Results are reported on two standard face databases, namely, MUCT and XM2VTS containing 276 and 295 images, respectively. The performance of the metrology-based approach was slightly lower than that of the appearance- based method by only about 3.8% for the MUCT database and about 5.7% for the XM2VTS database. Deng Cao, Cunjian Chen, Marco Piccirilli, Donald A. Adjeroh, Thirimachos Bourlai, Arun Ross |
IJCB | 2 |
| 2011 | Evaluation of gender classification methods on thermal and near-infrared face imagesabstractAutomatic gender classification based on face images is receiving increased attention in the biometrics community. Most gender classification systems have been evaluated only on face images captured in the visible spectrum. In this work, the possibility of deducing gender from face images obtained in the near-infrared (NIR) and thermal (THM) spectra is established. It is observed that the use of local binary pattern histogram (LBPH) features along with discriminative classifiers results in reasonable gender classification accuracy in both the NIR and THM spectra. Further, the performance of human subjects in classifying thermal face images is studied. Experiments suggest that machine-learning methods are better suited than humans for gender classification from face images in the thermal spectrum. Cunjian Chen, Arun Ross |
IJCB | 1 |
| 2008 | Information fusion of wavelet projection features for face recognitionabstractThis paper proposes a novel feature extraction method for face recognition in the wavelet domain called wavelet projection entropy (WPE). First, the projection entropy features from each wavelet subband are computed along the vertical and horizontal direction after the division. Then information fusion scheme is applied to integrate results obtained from each subband. Experiments show that WPE can extract the meaningful information from the wavelet domain. Meanwhile the decision level fusion achieves the best recognition rate among the three common information fusion methods. The proposed algorithms are validated on ORL and Yale face database for different pose and expression changes analysis. Detailed comparisons with previous published results are provided and it shows that our proposed algorithm performs very well. Cunjian Chen |
IJCNN | 1 |