Zexi Jia

dblp:338/5882 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-5044-9987ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Security and privacy · 6 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model
abstract
Traditional dialogue retrieval aims to select the most appropriate utterance or image from recent dialogue history. However, they often fail to meet users’ actual needs for revisiting semantically coherent content scattered across long-form conversations. To fill this gap, we define the Fine-grained Fragment Retrieval (FFR) task, requiring models to locate query-relevant fragments, comprising both utterances and images, from multimodal long-form dialogues. As a foundation for FFR, we construct MLDR, the longest-turn multimodal dialogue retrieval dataset to date, averaging 25.45 turns per dialogue, with each naturally spanning three distinct topics. To evaluate generalization in real-world scenarios, we curate and annotate a WeChat-based test set comprising real-world multimodal dialogues with an average of 75.38 turns. Building on these resources, we explore existing generation-based Vision-Language Models (VLMs) on FFR and observe that they often retrieve incoherent utterance-image fragments. While optimized for generating responses from visual-textual inputs, these models lack explicit supervision to ensure semantic coherence within retrieved fragments. To address this, we propose F2RVLM, a generative retrieval model trained in a two-stage paradigm: (1) supervised fine-tuning to inject fragment-level retrieval knowledge, and (2) GRPO-based reinforcement learning with multi-objective rewards to encourage outputs with semantic precision, relevance, and contextual coherence. In addition, to account for difficulty variations arising from differences in intra-fragment element distribution, ranging from locally dense to sparsely scattered, we introduce a difficulty-aware curriculum sampling that ranks training instances by predicted difficulty and gradually incorporates harder examples. This strategy enhances the model’s reasoning ability in long-form, multi-turn dialogue contexts. Experiments on both in-domain and real-domain sets demonstrate that F2RVLM substantially outperforms popular VLMs, achieving superior retrieval performance.
Hanbo Bi, Zexi Jia, Jiapei Zhang, Peixiang Luo, Xiaoyue Duan, Jinchao Zhang 0001
AAAI3
2025 Secret Lies in Color: Enhancing AI-Generated Images Detection with Color Distribution Analysis
abstract
The advancement of Generative Adversarial Networks (GANs) and diffusion models significantly enhances the realism of synthetic images, driving progress in image processing and creative design. However, this progress also necessitates the development of effective detection methods, as synthetic images become increasingly difficult to distinguish from real ones. This difficulty leads to societal issues, such as the spread of misinformation, identity theft, and online fraud. While previous detection methods perform well on public benchmarks, they struggle with our benchmark, FakeART, particularly when dealing with the latest models and cross-domain tasks (e.g., photo-to-painting). To address this challenge, we develop a new synthetic image detection technique based on color distribution. Unlike real images, synthetic images often exhibit uneven color distribution. By employing color quantization and restoration techniques, we analyze the color differences before and after image restoration. We discover and prove that these differences closely relate to the uniformity of color distribution. Based on this finding, we extract effective color features and combine them with image features to create a detection model with only 1.4 million parameters. This model achieves state-of-the-art results across various evaluation benchmarks, including the challenging FakeART dataset.
Zexi Jia, Chuanwei Huang, Yeshuang Zhu, Hongyan Fei, Xiaoyue Duan, Jiapei Zhang, Jinchao Zhang 0001, Jie Zhou 0016
CVPR1
2025 Semantic to Structure: Learning Structural Representations for Infringement Detection
abstract
Structural information in images is crucial for aesthetic assessment, and it is widely recognized in the artistic field that imitating the structure of other works significantly infringes on creators’ rights. The advancement of diffusion models has led to AI-generated content imitating artists’ structural creations, yet effective detection methods are still lacking. In this paper, we define this phenomenon as "structural infringement" and propose a corresponding detection method. Additionally, we develop quantitative metrics and create manually annotated datasets for evaluation: the SIA dataset of synthesized data, and the SIR dataset of real data. Due to the current lack of datasets for structural infringement detection, we propose a new data synthesis strategy based on diffusion models and LLM, successfully training a structural infringement detection model. Experimental results show that our method can successfully detect structural infringements and achieve notable improvements on annotated test sets.
Chuanwei Huang, Zexi Jia, Hongyan Fei, Yeshuang Zhu, Jinchao Zhang 0001, Jie Zhou 0016
ICASSP2
2025 MCID: Multi-aspect Copyright Infringement Detection for Generated Images
Chuanwei Huang, Zexi Jia, Hongyan Fei, Yeshuang Zhu, Jiapei Zhang, Xiaoyue Duan, Jinchao Zhang 0001, Jie Zhou 0016
ICCV2
2025 A Visual Leap in Clip Compositionality Reasoning Through Generation of Counterfactual Sets
abstract
Vision-language models (VLMs) often struggle with compositional reasoning due to insufficient high-quality image-text data. To tackle this challenge, we propose a novel block-based diffusion approach that automatically generates counterfactual datasets without manual annotation. Our method utilizes large language models to identify entities and their spatial relationships. It then independently generates image blocks as "puzzle pieces" coherently arranged according to specified compositional rules. This process creates diverse, high-fidelity counterfactual image-text pairs with precisely controlled variations. In addition, we introduce a specialized loss function that differentiates inter-set from intra-set samples, enhancing training efficiency and reducing the need for negative samples. Experiments demonstrate that fine-tuning VLMs with our counterfactual datasets significantly improves visual reasoning performance. Our approach achieves state-of-the-art results across multiple benchmarks while using substantially less training data than existing methods.
Zexi Jia, Chuanwei Huang, Hongyan Fei, Yeshuang Zhu, Jiapei Zhang, Jinchao Zhang 0001, Jie Zhou 0016
ICCV1
2025 From Imitation to Innovation: The Emergence of Ai's Unique Artistic Styles and the Challenge of Copyright Protection
Zexi Jia, Chuanwei Huang, Yeshuang Zhu, Hongyan Fei, Jiapei Zhang, Jinchao Zhang 0001, Jie Zhou 0016
ICCV1
2025 WalkVLM: Aid Visually Impaired People Walking by Vision Language Model
abstract
Approximately 200 million individuals around the world suffer from varying degrees of visual impairment, making it crucial to leverage AI technology to offer walking assistance for these people. With the recent progress of vision-language models (VLMs), employing VLMs to improve this field has emerged as a popular research topic. However, most existing methods are studied on self-built question-answering datasets, lacking a unified training and testing benchmark for walk guidance. Moreover, in blind walking task, it is necessary to perform real-time streaming video parsing and generate concise yet informative reminders, which poses a great challenge for VLMs that suffer from redundant responses and low inference efficiency. In this paper, we firstly release a diverse, extensive, and unbiased walking awareness dataset, containing 12k video-manual annotation pairs from Europe and Asia to provide a fair training and testing benchmark for blind walking task. Furthermore, a WalkVLM model is proposed, which employs chain of thought for hierarchical planning to generate concise but informative reminders and utilizes temporal-aware adaptive prediction to reduce the temporal redundancy of reminders. Finally, we have established a solid benchmark for blind walking task and verified the advantages of WalkVLM in stream video processing for this task compared to other VLMs. Our dataset and code will be released at anonymous link https://walkvlm2024.github.io.
Yeshuang Zhu, Jiapei Zhang, Zexi Jia, Peixiang Luo, Xiaoyue Duan, Jie Zhou 0016, Jinchao Zhang 0001
ICCV6
2025 ArtFRD: A Fisher-Rao Mixture Metric for Generative Model Aesthetic Evaluation
abstract
Recent advances in generative modeling have enabled the synthesis of high-quality artistic images. Nevertheless, systematic evaluation of generative models from an aesthetic standpoint is still lacking, which hinders progress in artistic image synthesis. Existing evaluation metrics, such as Fréchet Inception Distance (FID) and CMMD, struggle with aesthetic assessment: they rely on pretrained visual features that overlook nuanced artistic attributes and employ distance functions ill-suited for modeling the diverse, multi-modal distribution of artistic styles. To address these limitations, we propose ArtFRD, a metric specifically designed for generative aesthetic evaluation. Grounded in aesthetic theory, ArtFRD extracts visual features along four key aesthetic dimensions-brushstroke, composition, lighting, and color-to capture fine-grained artistic properties. To model the multi-modal nature of artistic styles, we adopt a Gaussian Mixture Model assumption and derive an efficient approximation of the Fisher-Rao distance, which serves as the final evaluation score. Extensive experiments demonstrate that ArtFRD aligns significantly better with human aesthetic judgments than existing metrics, even across a wide range of artistic styles. These results highlight its potential as a robust and interpretable foundation for future research in generative aesthetic evaluation.
Chuanwei Huang, Zexi Jia, Hongyan Fei, Yeshuang Zhu, Jinchao Zhang 0001, Jie Zhou 0016
ACM Multimedia2
2025 Automated Framework for Extracting and Restoring Minutiae From Low-Quality Fingerprints
abstract
Automated Fingerprint Identification Systems (AFIS) identify individuals swiftly and accurately by extracting distinctive features from fingerprint images. Minutiae, unique markers like ridge endings and bifurcations, are crucial for accurate identification. However, low-quality fingerprints often lack enough high-quality minutiae due to information loss. To address this, we propose a multi-stage minutiae extraction framework comprising a minutiae extractor and a minutiae repairer. The extractor adapts the Cross Stage Partial Network (CSPNet) architecture, integrating multi-scale feature modules to capture fingerprint details at different resolutions. The repairer uses a Transformer-based autoencoder with a 75% random masking strategy to restore missing or corrupted regions. We introduce an automated extraction-restoration-re-extraction process, identifying areas for repair based on the extractor's confidence map. Experiments on benchmark datasets demonstrate the superiority of our method in minutiae extraction accuracy, recall, and speed.
Zexi Jia, Chuanwei Huang
IEEE Signal Process. Lett.1
2024 Fingerprint Presentation Attack Detection by Region Decomposition
abstract
Fingerprint Presentation Attack Detection (PAD) is a crucial step in automatic fingerprint identification systems, which safeguards users from unauthorized malicious access. However, current presentation attack (i.e. spoof) techniques can forge intricate details of fingerprints (such as sweat holes), which makes the artifact evidence harder to detect. In this paper, we propose a novel PAD method from the perspective of decomposition to highlight the artifact evidence in each constituent element. Specifically, we utilize the fingerprint enhancement to decompose the fingerprint into the ridge region and the edge region. We observe that artifact evidence mainly exists in the gradient field within the ridge region, while it primarily resides in the spatial domain within the edge region. Then we propose an Orientation-Based Central Difference Convolution (OB-CDC) layer to prioritize gradient variations along the ridge direction. To further enhance robustness, we propose a Minutia Patches Random Rotation (MPRR) operation to disrupt the identity information of the fingerprint while preserving the artifact evidence. By integrating these techniques, we propose a two-stream network called Presentation-Attack-Detection-with-Region-Decomposition-Network (PADRD-Net) which integrates the processed feature of the ridge region and the edge region through a halfway fusion ResNet-18 structure. Experimental results on the LivDet 2021 dataset show that our proposed PADRD-Net can achieve 20.39% on BPCER@APCER = 1% and 87.12% on TDR@FDR = 1%, significantly outperforms the state-of-the-art. We also achieve outstanding performance in both the cross-sensor scenario and the cross-sensor and cross-material scenario. Extensive ablation studies and analysis experiments further indicate the effectiveness and robustness of our method.
Hongyan Fei, Chuanwei Huang, Zheng Wang 0073, Zexi Jia, Jufu Feng
IEEE Trans. Inf. Forensics Secur.5
2024 Finger Recovery Transformer: Toward Better Incomplete Fingerprint Identification
abstract
Fingerprint recognition is a crucial biometric technology extensively used in identity verification, including areas like criminal investigations, security systems, and biometric authentication. This technology encounters greater challenges when dealing with incomplete fingerprint images, especially those with significant background noise or substantial portions of the fingerprint missing. Existing incomplete fingerprint recognition technologies struggle with extensive data loss, primarily due to the significant reduction and difficulty in extracting usable features from incomplete fingerprint images. Current image processing methods or deep learning models are unable to comprehensively reconstruct fingerprint features with limited information. To address these challenges, we introduce the Finger Recovery Transformer (FingerRT), an innovative network specifically designed for recovering incomplete fingerprint information. FingerRT can simultaneously complete ambient noise cancellation and fingerprint feature information recovery, resulting in a complete and clean fingerprint image. FingerRT combines the most critical feature information in fingerprints, directional field, and minutiae, as supervision information. FingerRT inherits the denoising ability of the fingerprint enhancement networks and the powerful generative ability of the Vision Transformer architecture, enabling high-quality and robust fingerprint information recovery. By imposing constraints at multiple levels, including fingerprint features, fingerprint images, and multi-stage generation, FingerRT can complete fingerprint information accurately and effectively. Experiments demonstrate that FingerRT significantly enhances fingerprint recognition accuracy after recovery across various fingerprint datasets, including rolled, snapped, and latent fingerprints.
Zexi Jia, Chuanwei Huang, Zheng Wang 0073, Hongyan Fei, Jufu Feng
IEEE Trans. Inf. Forensics Secur.1
2023 Fingerprint Presentation Attack Detection with Supervised Contrastive Learning
abstract
The security of Automated Fingerprint Identification Systems (AFIS) heavily relies on the performance of the Fingerprint Presentation Attack Detection (FPAD) methods. However, the difficulty of FPAD lies in how to have strong robustness and generalization to unseen spoof fingerprints. To address this issue, we propose a novel FPAD framework with tailored Supervised Contrastive Learning (SupCon) and KNN-based OOD detection (KNN-OOD) method. We tailor the SupCon to better constrain the distribution of learned features by incorporating dynamic feature and label queues into SupCon and actively mining positive samples from the queues. In FPAD, we consider fingerprints with the same PAD label as intra-class, while those with different labels as inter-class. The tailored SupCon makes intra-class features more compact and inter-class features more dispersed. Utilizing the compact live fingerprint feature distribution, during the testing phase, we employ KNNOOD as an alternative to commonly used classification approaches. Since this approach does not rely on the distribution of trainset spoof fingerprints, it consistently achieves outstanding results even for unseen spoof fingerprints. Experiment results demonstrate that our proposed FPAD-SupCon framework achieves state-of-the-art performance on LivDet 2019 and LivDet 2021 datasets.
Chuanwei Huang, Hongyan Fei, Zheng Wang 0073, Zexi Jia, Jufu Feng
IJCB5
2023 FingerSTR: Weak Supervised Transformer for Latent Fingerprint Segmentation
abstract
Latent fingerprint segmentation is a crucial process in contemporary biometric systems utilized in criminal investigations and security applications. Accurately segmenting the fingerprint region from the background noise and artifacts, which can be challenging due to the complexity of the surrounding environment, is the primary goal of this process. Although various methodologies, including binarization-based, texture-based, and deep learning-based segmentation approaches have been proposed, they are often limited by environmental noise and a scarcity of annotated data, resulting in a low segmentation accuracy rate. In this paper, we propose FingerSTR (Finger Segmentation Transformer), a fully Transformer-based latent fingerprint segmentation network, and introduce a new teacher-student training methodology to achieve more precise and robust segmentation results without requiring manual annotation. Based on experimental results of latent fingerprint database NIST SD27, FingerSTR surpasses both deep-learning algorithms and handcraft methods, achieving state-of-the-art performance in the latent fingerprint segmentation task.
Zexi Jia, Zheng Wang 0073, Hongyan Fei, Chuanwei Huang, Jufu Feng
IJCB1
2023 Improving Latent Fingerprint Orientation Field Estimation Using Inpainting Techniques
abstract
Latent fingerprints play a vital role in forensic investigations. However, accurately estimating their orientation field can be challenging due to complex noise or overlapping fingerprint regions. In this paper, we propose a method to identify and correct these regions in the orientation field estimation. Specifically, our method comprises two networks: the first is an orientation field estimation network that outputs the initial orientation field, segment, and quality map, which determines the low-quality regions, including overlapping fingerprints and unclear ridge areas. The second network refills the orientation field in low-quality regions using inpainting techniques. This effectively handles unclear ridges and overlapping fingerprints, which can disrupt orientation field estimation. We assess our method using the NIST SD 27 dataset and demonstrate superior performance compared to existing state-of-the-art latent orientation field estimation methods, achieving the average root mean square deviation of 11.20.
Zheng Wang 0073, Zexi Jia, Chuanwei Huang, Hongyan Fei, Jufu Feng
IJCB2
2023 Event-Based Semantic Segmentation With Posterior Attention
abstract
In the past years, attention-based Transformers have swept across the field of computer vision, starting a new stage of backbones in semantic segmentation. Nevertheless, semantic segmentation under poor light conditions remains an open problem. Moreover, most papers about semantic segmentation work on images produced by commodity frame-based cameras with a limited framerate, hindering their deployment to auto-driving systems that require instant perception and response at milliseconds. An event camera is a new sensor that generates event data at microseconds and can work in poor light conditions with a high dynamic range. It looks promising to leverage event cameras to enable perception where commodity cameras are incompetent, but algorithms for event data are far from mature. Pioneering researchers stack event data as frames so that event-based segmentation is converted to frame-based segmentation, but characteristics of event data are not explored. Noticing that event data naturally highlight moving objects, we propose a posterior attention module that adjusts the standard attention by the prior knowledge provided by event data. The posterior attention module can be readily plugged into many segmentation backbones. Plugging the posterior attention module into a recently proposed SegFormer network, we get EvSegFormer (the event-based version of SegFormer) with state-of-the-art performance in two datasets (MVSEC and DDD-17) collected for event-based segmentation. Code is available at https://github.com/zexiJia/EvSegFormer to facilitate research on event-based vision.
Zexi Jia, Kaichao You, Weihua He, Yang Tian 0002, Yongxiang Feng, Yaoyuan Wang, Xu Jia 0012, Yihang Lou, Guoqi Li 0002
IEEE Trans. Image Process.1
2022 Minutiae-awarely Learning Fingerprint Representation for Fingerprint Indexing
abstract
With compact and discriminative fingerprint representation, fingerprint indexing can effectively reduce the search space and improve the efficiency in large-scale fingerprint identification. Previous fixed-length fingerprint representations do not combine the global and minutiae local information well, leading to unsatisfactory results. In this paper, we utilize an end-to-end network to extract Minutiae-aware fingerprint RepresentationS (MaRs) that consider both the global and minutiae local information. The proposed fingerprint representation is weighted aggregated by the output feature map of the network. We hope that not only the global pattern but also the matching minutia-centered regions are similar for the paired fingerprints. We impose constraints among proposed fingerprint representations and constraints among minutia local representations aggregated from each minutia-centered region. The latter constraint strengthens the similarity of matched minutiae local regions. Experimental results show that our minutiaeaware global representation outperforms previous methods in fingerprint indexing on two benchmarks and exhibits strong indexing robustness with a 100k expanded database.
Zheng Wang 0073, Zexi Jia, Jufu Feng
IJCB4