EDBT 2026 Demo / reviewers in the wild / expert
Chuanwei Huang
dblp:360/6122
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Security and privacy · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fingerprint Presentation Attack Detection With Fixed Prototype Contrastive Learning and Anomaly DetectionabstractFingerprint Presentation Attack Detection (PAD) is crucial for ensuring the reliability of fingerprint recognition systems. Mainstream methods typically frame PAD as a binary classification task. Although they achieve commendable results, these approaches often fail on unseen spoofs. This is because their decision boundaries rely heavily on trainset spoofs, leading to poor generalization. To address this issue, we reformulate PAD as an anomaly detection problem and propose the Fixed-Prototype PAD method. First, we introduce Fixed-Prototype Contrastive Learning (FPCL) for extracting PAD features. FPCL clusters the features of live fingerprints near a fixed prototype while pushing the features of spoof fingerprints away. We then design a Fixed-Prototype Anomaly Detection method to replace binary classification for decision-making. Specifically, this involves comparing the distance between test fingerprint features and the fixed live prototype to detect spoofs. Our method eliminates the dependency on the trainset spoofs during decision-making, thereby enhancing generalization capability. Experimental results demonstrate that our approach achieves state-of-the-art performance on the LivDet 2019 and LivDet 2021 benchmarks. Chuanwei Huang, Hongyan Fei, Pengcheng Luo, Jufu Feng |
IEEE Signal Process. Lett. | 1 |
| 2025 | Secret Lies in Color: Enhancing AI-Generated Images Detection with Color Distribution AnalysisabstractThe advancement of Generative Adversarial Networks (GANs) and diffusion models significantly enhances the realism of synthetic images, driving progress in image processing and creative design. However, this progress also necessitates the development of effective detection methods, as synthetic images become increasingly difficult to distinguish from real ones. This difficulty leads to societal issues, such as the spread of misinformation, identity theft, and online fraud. While previous detection methods perform well on public benchmarks, they struggle with our benchmark, FakeART, particularly when dealing with the latest models and cross-domain tasks (e.g., photo-to-painting). To address this challenge, we develop a new synthetic image detection technique based on color distribution. Unlike real images, synthetic images often exhibit uneven color distribution. By employing color quantization and restoration techniques, we analyze the color differences before and after image restoration. We discover and prove that these differences closely relate to the uniformity of color distribution. Based on this finding, we extract effective color features and combine them with image features to create a detection model with only 1.4 million parameters. This model achieves state-of-the-art results across various evaluation benchmarks, including the challenging FakeART dataset. Zexi Jia, Chuanwei Huang, Yeshuang Zhu, Hongyan Fei, Xiaoyue Duan, Jiapei Zhang, Jinchao Zhang 0001, Jie Zhou 0016 |
CVPR | 2 |
| 2025 | Semantic to Structure: Learning Structural Representations for Infringement DetectionabstractStructural information in images is crucial for aesthetic assessment, and it is widely recognized in the artistic field that imitating the structure of other works significantly infringes on creators’ rights. The advancement of diffusion models has led to AI-generated content imitating artists’ structural creations, yet effective detection methods are still lacking. In this paper, we define this phenomenon as "structural infringement" and propose a corresponding detection method. Additionally, we develop quantitative metrics and create manually annotated datasets for evaluation: the SIA dataset of synthesized data, and the SIR dataset of real data. Due to the current lack of datasets for structural infringement detection, we propose a new data synthesis strategy based on diffusion models and LLM, successfully training a structural infringement detection model. Experimental results show that our method can successfully detect structural infringements and achieve notable improvements on annotated test sets. Chuanwei Huang, Zexi Jia, Hongyan Fei, Yeshuang Zhu, Jinchao Zhang 0001, Jie Zhou 0016 |
ICASSP | 1 |
| 2025 | MCID: Multi-aspect Copyright Infringement Detection for Generated Images
Chuanwei Huang, Zexi Jia, Hongyan Fei, Yeshuang Zhu, Jiapei Zhang, Xiaoyue Duan, Jinchao Zhang 0001, Jie Zhou 0016 |
ICCV | 1 |
| 2025 | A Visual Leap in Clip Compositionality Reasoning Through Generation of Counterfactual SetsabstractVision-language models (VLMs) often struggle with compositional reasoning due to insufficient high-quality image-text data. To tackle this challenge, we propose a novel block-based diffusion approach that automatically generates counterfactual datasets without manual annotation. Our method utilizes large language models to identify entities and their spatial relationships. It then independently generates image blocks as "puzzle pieces" coherently arranged according to specified compositional rules. This process creates diverse, high-fidelity counterfactual image-text pairs with precisely controlled variations. In addition, we introduce a specialized loss function that differentiates inter-set from intra-set samples, enhancing training efficiency and reducing the need for negative samples. Experiments demonstrate that fine-tuning VLMs with our counterfactual datasets significantly improves visual reasoning performance. Our approach achieves state-of-the-art results across multiple benchmarks while using substantially less training data than existing methods. Zexi Jia, Chuanwei Huang, Hongyan Fei, Yeshuang Zhu, Jiapei Zhang, Jinchao Zhang 0001, Jie Zhou 0016 |
ICCV | 2 |
| 2025 | From Imitation to Innovation: The Emergence of Ai's Unique Artistic Styles and the Challenge of Copyright Protection
Zexi Jia, Chuanwei Huang, Yeshuang Zhu, Hongyan Fei, Jiapei Zhang, Jinchao Zhang 0001, Jie Zhou 0016 |
ICCV | 2 |
| 2025 | UniVG: Towards UNIfied-modal Video GenerationabstractDiffusion based video generation has received significant attention in both the academic and industrial communities. Despite recent exploration of diverse conditional inputs for better video generation control, existing methods, primarily targeting individual tasks, often fall short in real-world scenarios where users may use any form of conditioning, either individually or combined. To address this, we propose a Unified-modal Video Generation system capable of handling multiple video generation tasks across different modalities. Our approach introduces the concept of generative freedom in the diffusion process, which allows us to reclassify video generation tasks into high-freedom and low-freedom categories based on the solution space given certain conditions. We then design different diffusion paradigms for each category. For high-freedom video generation, we present a base model that is capable of handling varied semantic combinations of text and image. For low-freedom video generation, we propose the Biased Gaussian Noise (BGN) to tackle the discrepancy of the diffusion process between the training and inference stage when using strong conditional guidance strategy. Our proposed UniVG achieves superior objective results on public datasets, surpassing the current open-source methods and is on par with the current close-source method Gen2 and Pika in human evaluations. For more samples, visit our homepage. Ludan Ruan, Chuanwei Huang, Xinyan Xiao |
ICME | 3 |
| 2025 | ArtFRD: A Fisher-Rao Mixture Metric for Generative Model Aesthetic EvaluationabstractRecent advances in generative modeling have enabled the synthesis of high-quality artistic images. Nevertheless, systematic evaluation of generative models from an aesthetic standpoint is still lacking, which hinders progress in artistic image synthesis. Existing evaluation metrics, such as Fréchet Inception Distance (FID) and CMMD, struggle with aesthetic assessment: they rely on pretrained visual features that overlook nuanced artistic attributes and employ distance functions ill-suited for modeling the diverse, multi-modal distribution of artistic styles. To address these limitations, we propose ArtFRD, a metric specifically designed for generative aesthetic evaluation. Grounded in aesthetic theory, ArtFRD extracts visual features along four key aesthetic dimensions-brushstroke, composition, lighting, and color-to capture fine-grained artistic properties. To model the multi-modal nature of artistic styles, we adopt a Gaussian Mixture Model assumption and derive an efficient approximation of the Fisher-Rao distance, which serves as the final evaluation score. Extensive experiments demonstrate that ArtFRD aligns significantly better with human aesthetic judgments than existing metrics, even across a wide range of artistic styles. These results highlight its potential as a robust and interpretable foundation for future research in generative aesthetic evaluation. Chuanwei Huang, Zexi Jia, Hongyan Fei, Yeshuang Zhu, Jinchao Zhang 0001, Jie Zhou 0016 |
ACM Multimedia | 1 |
| 2025 | Automated Framework for Extracting and Restoring Minutiae From Low-Quality FingerprintsabstractAutomated Fingerprint Identification Systems (AFIS) identify individuals swiftly and accurately by extracting distinctive features from fingerprint images. Minutiae, unique markers like ridge endings and bifurcations, are crucial for accurate identification. However, low-quality fingerprints often lack enough high-quality minutiae due to information loss. To address this, we propose a multi-stage minutiae extraction framework comprising a minutiae extractor and a minutiae repairer. The extractor adapts the Cross Stage Partial Network (CSPNet) architecture, integrating multi-scale feature modules to capture fingerprint details at different resolutions. The repairer uses a Transformer-based autoencoder with a 75% random masking strategy to restore missing or corrupted regions. We introduce an automated extraction-restoration-re-extraction process, identifying areas for repair based on the extractor's confidence map. Experiments on benchmark datasets demonstrate the superiority of our method in minutiae extraction accuracy, recall, and speed. Zexi Jia, Chuanwei Huang |
IEEE Signal Process. Lett. | 2 |
| 2024 | Fingerprint Presentation Attack Detection by Region DecompositionabstractFingerprint Presentation Attack Detection (PAD) is a crucial step in automatic fingerprint identification systems, which safeguards users from unauthorized malicious access. However, current presentation attack (i.e. spoof) techniques can forge intricate details of fingerprints (such as sweat holes), which makes the artifact evidence harder to detect. In this paper, we propose a novel PAD method from the perspective of decomposition to highlight the artifact evidence in each constituent element. Specifically, we utilize the fingerprint enhancement to decompose the fingerprint into the ridge region and the edge region. We observe that artifact evidence mainly exists in the gradient field within the ridge region, while it primarily resides in the spatial domain within the edge region. Then we propose an Orientation-Based Central Difference Convolution (OB-CDC) layer to prioritize gradient variations along the ridge direction. To further enhance robustness, we propose a Minutia Patches Random Rotation (MPRR) operation to disrupt the identity information of the fingerprint while preserving the artifact evidence. By integrating these techniques, we propose a two-stream network called Presentation-Attack-Detection-with-Region-Decomposition-Network (PADRD-Net) which integrates the processed feature of the ridge region and the edge region through a halfway fusion ResNet-18 structure. Experimental results on the LivDet 2021 dataset show that our proposed PADRD-Net can achieve 20.39% on BPCER@APCER = 1% and 87.12% on TDR@FDR = 1%, significantly outperforms the state-of-the-art. We also achieve outstanding performance in both the cross-sensor scenario and the cross-sensor and cross-material scenario. Extensive ablation studies and analysis experiments further indicate the effectiveness and robustness of our method. Hongyan Fei, Chuanwei Huang, Zheng Wang 0073, Zexi Jia, Jufu Feng |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Finger Recovery Transformer: Toward Better Incomplete Fingerprint IdentificationabstractFingerprint recognition is a crucial biometric technology extensively used in identity verification, including areas like criminal investigations, security systems, and biometric authentication. This technology encounters greater challenges when dealing with incomplete fingerprint images, especially those with significant background noise or substantial portions of the fingerprint missing. Existing incomplete fingerprint recognition technologies struggle with extensive data loss, primarily due to the significant reduction and difficulty in extracting usable features from incomplete fingerprint images. Current image processing methods or deep learning models are unable to comprehensively reconstruct fingerprint features with limited information. To address these challenges, we introduce the Finger Recovery Transformer (FingerRT), an innovative network specifically designed for recovering incomplete fingerprint information. FingerRT can simultaneously complete ambient noise cancellation and fingerprint feature information recovery, resulting in a complete and clean fingerprint image. FingerRT combines the most critical feature information in fingerprints, directional field, and minutiae, as supervision information. FingerRT inherits the denoising ability of the fingerprint enhancement networks and the powerful generative ability of the Vision Transformer architecture, enabling high-quality and robust fingerprint information recovery. By imposing constraints at multiple levels, including fingerprint features, fingerprint images, and multi-stage generation, FingerRT can complete fingerprint information accurately and effectively. Experiments demonstrate that FingerRT significantly enhances fingerprint recognition accuracy after recovery across various fingerprint datasets, including rolled, snapped, and latent fingerprints. Zexi Jia, Chuanwei Huang, Zheng Wang 0073, Hongyan Fei, Jufu Feng |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Fingerprint Presentation Attack Detection with Supervised Contrastive LearningabstractThe security of Automated Fingerprint Identification Systems (AFIS) heavily relies on the performance of the Fingerprint Presentation Attack Detection (FPAD) methods. However, the difficulty of FPAD lies in how to have strong robustness and generalization to unseen spoof fingerprints. To address this issue, we propose a novel FPAD framework with tailored Supervised Contrastive Learning (SupCon) and KNN-based OOD detection (KNN-OOD) method. We tailor the SupCon to better constrain the distribution of learned features by incorporating dynamic feature and label queues into SupCon and actively mining positive samples from the queues. In FPAD, we consider fingerprints with the same PAD label as intra-class, while those with different labels as inter-class. The tailored SupCon makes intra-class features more compact and inter-class features more dispersed. Utilizing the compact live fingerprint feature distribution, during the testing phase, we employ KNNOOD as an alternative to commonly used classification approaches. Since this approach does not rely on the distribution of trainset spoof fingerprints, it consistently achieves outstanding results even for unseen spoof fingerprints. Experiment results demonstrate that our proposed FPAD-SupCon framework achieves state-of-the-art performance on LivDet 2019 and LivDet 2021 datasets. Chuanwei Huang, Hongyan Fei, Zheng Wang 0073, Zexi Jia, Jufu Feng |
IJCB | 1 |
| 2023 | FingerSTR: Weak Supervised Transformer for Latent Fingerprint SegmentationabstractLatent fingerprint segmentation is a crucial process in contemporary biometric systems utilized in criminal investigations and security applications. Accurately segmenting the fingerprint region from the background noise and artifacts, which can be challenging due to the complexity of the surrounding environment, is the primary goal of this process. Although various methodologies, including binarization-based, texture-based, and deep learning-based segmentation approaches have been proposed, they are often limited by environmental noise and a scarcity of annotated data, resulting in a low segmentation accuracy rate. In this paper, we propose FingerSTR (Finger Segmentation Transformer), a fully Transformer-based latent fingerprint segmentation network, and introduce a new teacher-student training methodology to achieve more precise and robust segmentation results without requiring manual annotation. Based on experimental results of latent fingerprint database NIST SD27, FingerSTR surpasses both deep-learning algorithms and handcraft methods, achieving state-of-the-art performance in the latent fingerprint segmentation task. Zexi Jia, Zheng Wang 0073, Hongyan Fei, Chuanwei Huang, Jufu Feng |
IJCB | 5 |
| 2023 | Improving Latent Fingerprint Orientation Field Estimation Using Inpainting TechniquesabstractLatent fingerprints play a vital role in forensic investigations. However, accurately estimating their orientation field can be challenging due to complex noise or overlapping fingerprint regions. In this paper, we propose a method to identify and correct these regions in the orientation field estimation. Specifically, our method comprises two networks: the first is an orientation field estimation network that outputs the initial orientation field, segment, and quality map, which determines the low-quality regions, including overlapping fingerprints and unclear ridge areas. The second network refills the orientation field in low-quality regions using inpainting techniques. This effectively handles unclear ridges and overlapping fingerprints, which can disrupt orientation field estimation. We assess our method using the NIST SD 27 dataset and demonstrate superior performance compared to existing state-of-the-art latent orientation field estimation methods, achieving the average root mean square deviation of 11.20. Zheng Wang 0073, Zexi Jia, Chuanwei Huang, Hongyan Fei, Jufu Feng |
IJCB | 3 |