Changtao Miao

dblp:287/9236 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0002-7634-9992ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Multilingual Text-Centric VQA
abstract
Multilingual Text-Centric Visual Question Answering (TEC-VQA) has become crucial for real-world applications, as it requires fine-grained understanding and reasoning over multilingual scene text. Recent advances in vision-language models (VLMs) have demonstrated strong potential in tackling multimodal tasks. However, most existing approaches rely primarily on textual Chain-of-Thought (CoT) and provide limited support for multilingual multimodal reasoning. To address this gap, we introduce LaV-CoT, the first Language-aware Visual CoT framework with Multi-Aspect Reward Optimization. LaV-CoT incorporates an interpretable multi-stage reasoning pipeline consisting of text summary with bounding box, language identification, spatial object-level captioning, and step-by-step logical reasoning. To improve reasoning accuracy and cross-lingual generalization, we propose a novel verifiable Multi-Aspect Reward Optimization in addition to supervised fine-tuning that incorporates rewards for linguistic consistency, structural fidelity, and response accuracy. Extensive evaluations on public datasets, including MMMB, Multilingual MMBench, and MTVQA, show that LaV-CoT outperforms open-source models of similar size by up to ~9.5% accuracy, even surpassing open-source models more than twice its size, and further exceeding several state-of-the-art proprietary models. Moreover, LaV-CoT has been integrated into our online Intelligent Document Processing platform. A further online A/B test demonstrates an \(\sim\)8.7% improvement in acceptance rate, validating its effectiveness in industrial deployment and commercial applications. Our code is available at this https://github.com/HJNVR/LaV-CoT repository.
Zhiya Tan, Shutao Gong, Fanwei Zeng, Joey Tianyi Zhou, Changtao Miao, Huazhe Tan, Weibin Yao, Jianshu Li
WWW6
2026 Prototype Memory-Based Neighboring Feature Fusion Network for Image Manipulation Localization
abstract
Image manipulation localization (IML) aims to segment manipulated regions in suspicious images. However, most existing methods rely solely on intrinsic features extracted from the input image and passively model local or global inconsistencies, making it difficult to accurately delineate manipulated regions with ambiguous boundaries. To address these challenges, we propose a prototype memory-based neighboring feature fusion network (PNF-Net), which is inspired by a biological memory mechanism. PNF-Net simulates selective preference by learning manipulation-trace prototypes as memory priors, thereby guiding representation learning toward consistent and discriminative manipulation cues. Specifically, we propose a memory-guided localization module (MLM) that models the consistencies and anomalies between manipulated regions and the background as memory priors, enabling precise localization. We then propose a neighboring feature interaction module (NFIM) that preserves fine-grained details from neighboring shallow features, enhances global semantics from neighboring deep features, and effectively fuses them. Finally, a verification fusion module (VFM) is designed to enrich contextual semantics and improve the completeness and accuracy of localization results. Extensive experiments on multiple benchmark datasets show that our PNF-Net outperforms most state-of-the-art IML models. Our code is available on https://github.com/vpsg-research/PNF-Net.
Zhiqing Guo, Changtao Miao, Wenzhong Yang, Gaobo Yang, Xin Liao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Exploiting Feature Gating and Injection For Multi-modal Manipulation Detection and Grounding
Jiazhen Wang, Bin Liu 0016, Changtao Miao, Qi Chu 0001, Nenghai Yu
ICIG (3)5
2025 Exploring Generalized Features For LLM-Generated Text Detection
Jiazhen Wang, Bin Liu 0016, Changtao Miao, Qi Chu 0001, Quanchen Zou, Deyue Zhang, Nenghai Yu
ICIG (3)3
2025 SUEDE: Shared Unified Experts for Physical- Digital Face Attack Detection Enhancement
abstract
Face recognition systems are vulnerable to physical attacks (e.g., printed photos) and digital threats (e.g., DeepFake), which are currently being studied as independent visual tasks, such as Face Anti-Spoofing and Forgery Detection. The inherent differences among various attack types present significant challenges in identifying a common feature space, making it difficult to develop a unified framework for detecting data from both attack modalities simultaneously. Inspired by the efficacy of Mixture-of-Experts (MoE) in learning across diverse domains, we explore utilizing multiple experts to learn the distinct features of various attack types. However, the feature distributions of physical and digital attacks overlap and differ. This suggests that relying solely on distinct experts to learn the unique features of each attack type may overlook shared knowledge between them. To address these issues, we propose SUEDE, the Shared Unified Experts for Physical-Digital Face Attack Detection Enhancement. SUEDE combines a shared expert (always activated) to capture common features for both attack types and multiple routed experts (selectively activated) for specific attack types. Further, we integrate CLIP as the base network to ensure the shared expert benefits from prior visual knowledge and align visual-text representations in a unified space. Extensive results demonstrate SUEDE achieves superior performance compared to state-of-the-art unified detection methods.
Zuying Xie, Changtao Miao, Ajian Liu 0001, Jiabao Guo, Feng Li 0037, Dan Guo 0001, Yunfeng Diao
ICME2
2025 GMamba: EEG Representation Learning from Spatiotemporal Perspectives via Graph Mamba
Weiwei Feng, Nanqing Xu, Changtao Miao, Tengfei Liu 0007, Weiqiang Wang 0002
ICONIP (3)4
2025 Mixture-of-Noises Enhanced Forgery-Aware Predictor for Multi-Face Manipulation Detection and Localization
abstract
With the advancement of face manipulation technology, forgery images in multi-face scenarios are gradually becoming a more complex and realistic challenge. Despite this, detection and localization methods for such multi-face manipulations remain underdeveloped. Traditional manipulation localization methods either indirectly derive detection results from localization masks, resulting in limited detection performance, or employ a naive two-branch structure to simultaneously obtain detection and localization results, which cannot effectively benefit the localization capability due to limited interaction between the two tasks. This paper proposes a new framework, namely MoNFAP, specifically tailored for multi-face manipulation detection and localization. The MoNFAP primarily introduces two novel modules: the Forgery-aware Unified Predictor (FUP) Module and the Mixture-of-Noises Module (MNM). The proposed FUP integrates detection and localization tasks using a token learning strategy and multiple forgery-aware transformers, which facilitates the use of classification information to enhance localization capability. Furthermore, to mitigate the interference from general semantic object information, we propose the MNM that leverages multiple noise extractors based on the mixture of experts concept. This allows the MNM to learn semantic-agnostic forgery features from general RGB features, further boosting the performance of our proposed framework. Finally, we establish a comprehensive benchmark for multi-face detection and localization, and the proposed MoNFAP achieves significant performance. The code is available: https://github.com/miaoct/MoNFAP.
Changtao Miao, Qi Chu 0001, Zhentao Tan, Zhenchao Jin, Wanyi Zhuang, Honggang Hu, Nenghai Yu
ACM Multimedia1
2025 MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World Scenarios
abstract
Rapid advances in Artificial Intelligence Generated Content (AIGC) have enabled increasingly sophisticated face forgeries, posing a significant threat to social security. However, current Deepfake detection methods are limited by constraints in existing datasets, which lack the diversity necessary in real-world scenarios. Specifically, these data sets fall short in four key areas: unknown of advanced forgery techniques, variability of facial scenes, richness of real data, and degradation of real-world propagation. To address these challenges, we propose the Multi-dimensional Face Forgery Image (MFFI ) dataset, tailored for real-world scenarios. MFFI enhances realism based on four strategic dimensions: 1) Wider Forgery Methods; 2) Varied Facial Scenes; 3) Diversified Authentic Data; 4) Multi-level Degradation Operations. MFFI integrates 50 different forgery methods and contains 1024K image samples. Benchmark evaluations show that MFFI outperforms existing public datasets in terms of scene complexity, cross-domain generalization capability, and detection difficulty gradients. These results validate the technical advance and practical utility of MFFI in simulating real-world conditions. The dataset and additional details are publicly available at https://github.com/inclusionConf/MFFI.
Changtao Miao, Weiwei Feng, Qi Chu 0001, Jianshu Li, Yunfeng Diao, Wei Zhou 0021, Joey Tianyi Zhou, Xiaoshuai Hao
ACM Multimedia1
2025 Towards Good Generalizations for Diffusion Generated Image Detection Using Multiple Reconstruction Contrastive Learning
abstract
A striking proficiency of diffusion models in producing and manipulating images with an unprecedented level of realism has unquestionably elicited concerns. Many methods have been proposed to detect generated images. In particular, recent studies reveal that autoencoder reconstruction error can serve as an effective indicator for distinguishing authentic and synthetic images, since most generative models adopt analogous encoder-decoder operation. However, the reliance on a single autoencoder reconstruction error provides only limited information, which is insufficient for comprehensively capturing discriminative features, resulting in restricted generalization performance. In this paper, we propose Multiple Reconstruction Contrastive Learning (MRCL), which leverages multiple reconstruction residuals to enhance the generalizability of generated image detection. Specifically, MRCL applies Dinov2-ViT with LoRA fine-tuning to extract fine-grained feature representations of origin images and their multiple VAE reconstructions. In addition, a Residual Dense Fusion module is designed to effectively combine multiple VAE reconstruction residuals. Further, a contrastive learning strategy is adopted to guide the distance of origin images and VAE reconstruction representations. Extensive experimental results demonstrate the superior generalization performance of the proposed MRCL.
Wanyi Zhuang, Qi Chu 0001, Changtao Miao, Nenghai Yu
ACM Multimedia4
2025 Parameter-free and Accessible Prompt Learning to Enhance Adversarial Robustness for Pre-trained Vision-Language Models
abstract
Xingran Zhou, Kun Yang, Changtao Miao, Bingyu Hu, Zhuoer Xu, Shiwen Cui, Changhua Meng, Dan Hong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xingran Zhou, Changtao Miao, Bingyu Hu, Zhuoer Xu, Shiwen Cui, Changhua Meng, Dan Hong
NAACL (Long Papers)3
2025 Multi-spectral Class Center Network for Face Manipulation Localization
abstract
As Deepfake content proliferates online, advancing face manipulation forensics has become crucial. To combat this emerging threat, previous methods mainly focus on studying how to distinguish authentic and manipulated face images. Although impressive, image-level classification lacks explainability and is limited to specific application scenarios, spurring recent research on pixel-level prediction for face manipulation forensics. However, existing forgery localization methods suffer from exploring frequency-based forgery traces in the localization network. In this paper, we observe that multi-frequency spectrum information is effective for identifying tampered regions. To this end, a novel Multi-spectral Class Center Network (MSCCNet) is proposed for face manipulation localization. Specifically, we design a Multi-spectral Class Center (MSCC) module to learn more generalizable and multi-frequency features. Based on the features of different frequency bands, the MSCC module collects multi-spectral class centers and computes pixel-to-class relations. Applying multi-spectral class-level representations suppresses the semantic information of the visual concepts which is insensitive to manipulated regions of forgery images. Furthermore, we propose a Multi-level Features Aggregation (MFA) module to employ more low-level forgery artifacts and structural textures. Meanwhile, we conduct a comprehensive localization benchmark based on pixel-level FF++ and Dolos datasets. Experimental results quantitatively and qualitatively demonstrate the effectiveness and superiority of the proposed MSCCNet. We expect this work to inspire more studies on pixel-level face manipulation localization. The codes are available.
Changtao Miao, Qi Chu 0001, Zhentao Tan, Zhenchao Jin, Wanyi Zhuang, Bin Liu 0016, Honggang Hu, Nenghai Yu
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Exploiting Modality-Specific Features for Multi-Modal Manipulation Detection and Grounding
abstract
AI-synthesized text and images have gained significant attention, particularly due to the widespread dissemination of multi-modal manipulations on the internet, which has resulted in numerous negative impacts on society. Existing methods for multi-modal manipulation detection and grounding primarily focus on fusing vision-language features to make predictions, while overlooking the importance of modality-specific features, leading to sub-optimal results. In this paper, we construct a simple and novel transformer-based framework for multi-modal manipulation detection and grounding tasks. Our framework simultaneously explores modality-specific features while preserving the capability for multi-modal alignment. To achieve this, we introduce visual/language pre-trained encoders and dual-branch cross-attention (DCA) to extract and fuse modality-unique features. Furthermore, we design decoupled fine-grained classifiers (DFC) to enhance modality-specific feature mining and mitigate modality competition. Moreover, we propose an implicit manipulation query (IMQ) that adaptively aggregates global contextual cues within each modality using learnable queries, thereby improving the discovery of forged details. Extensive experiments on the DGM4dataset demonstrate the superior performance of our proposed model compared to state-of-the-art approaches.
Jiazhen Wang, Bin Liu 0016, Changtao Miao, Wanyi Zhuang, Qi Chu 0001, Nenghai Yu
ICASSP3
2024 MFMS: Learning Modality-Fused and Modality-Specific Features for Deepfake Detection and Localization Tasks
abstract
This paper presents a summary of the proposed solution to the AV-Deepfake1M competition. Deepfake technology is developing fast, and realistic generation techniques of audio and videos have aroused public concerns. With this background, the AV-Deepfake1M competition aims to address the problem of audio-video Deepfake and provides a large-scale dataset named AV-Deepfake1M to boost the research in this area. In this paper, we present our solutions which have achieved top performance in this competition. We also provide more detailed experiments to prove the effectiveness of the modules used in our methods.
Changtao Miao, Jianshu Li, Wenzhong Deng, Weibin Yao, Zhe Li 0081, Bingyu Hu, Weiwei Feng, Qi Chu 0001
ACM Multimedia2
2024 Detect Text Forgery with Non-forged Image Features: A Framework for Detection and Grounding of Image-Text Manipulation
Changtao Miao, Qi Chu 0001, Dianmo Sheng, Jiazhen Wang, Bin Liu 0016, Nenghai Yu
PRCV (11)2
2023 Revisiting TENT for Test-Time Adaption Semantic Segmentation and Classification Head Adjustment
Xuanpu Zhao, Qi Chu 0001, Changtao Miao, Bin Liu 0016, Nenghai Yu
ICIG (3)3
2023 F2Trans: High-Frequency Fine-Grained Transformer for Face Forgery Detection
abstract
In recent years, face forgery detectors have aroused great interest and achieved impressive performance, but they are still struggling with generalization and robustness. In this work, we explore taking full advantage of the fine-grained forgery traces in both spatial and frequency domains to alleviate this issue. Specifically, we propose a novel High-Frequency Fine-Grained Transformer (F2Trans) network which contains two important components, namely Central Difference Attention (CDA) and High-frequency Wavelet Sampler (HWS). The premier CDA module is capable of capturing invariant fine-grained manipulation patterns by aggregating both pixel-level intensity and gradient information of the query to generate key and value pairs. Subsequently, the proposed HWS discards the low-frequency components of wavelet transformation and hierarchically explores high-frequency forgery cues of feature maps, which prevents model confusion caused by low-frequency components and pays attention to local frequency information. In addition, HWS can be employed as a special pooling layer for the F2Trans architecture to produce hierarchical feature representations in the spatial-frequency domain. Extensive experiments on multiple popular benchmarks demonstrate the generalization and robustness of the specially designed F2Trans framework is well-tailored for face forgery detection when confronting the cross-dataset, cross-manipulation, and unseen perturbations.
Changtao Miao, Zichang Tan, Qi Chu 0001, Huan Liu 0030, Honggang Hu, Nenghai Yu
IEEE Trans. Inf. Forensics Secur.1
2022 UIA-ViT: Unsupervised Inconsistency-Aware Method Based on Vision Transformer for Face Forgery Detection
Wanyi Zhuang, Qi Chu 0001, Zhentao Tan, Qiankun Liu 0001, Changtao Miao, Zixiang Luo, Nenghai Yu
ECCV (5)6
2022 Towards Intrinsic Common Discriminative Features Learning for Face Forgery Detection Using Adversarial Learning
abstract
Existing face forgery detection methods usually treat face forgery detection as a binary classification problem and adopt deep convolution neural networks to learn discriminative features. The ideal discriminative features should be only related to the real/fake labels of facial images. However, we observe that the features learned by vanilla classification networks are correlated to unnecessary properties, such as forgery methods and facial identities. Such phenomenon would limit forgery detection performance especially for the generalization ability. Motivated by this, we propose a novel method which utilizes adversarial learning to eliminate the negative effect of different forgery methods and facial identities, which helps classification network to learn intrinsic common discriminative features for face forgery detection. To leverage data lacking ground truth label of facial identities, we design a special identity discriminator based on similarity information derived from off-the-shelf face recognition model. Extensive experiments demonstrate the effectiveness of the proposed method under both intra-dataset and cross-dataset evaluation settings.
Wanyi Zhuang, Qi Chu 0001, Changtao Miao, Bin Liu 0016, Nenghai Yu
ICME4
2022 Transformer-Based Feature Compensation and Aggregation for DeepFake Detection
abstract
Deepfake detection has attracted increasing attention in recent years. In this paper, we propose a transformer-based framework with feature compensation and aggregation (Trans-FCA) to extract rich forgery cues for deepfake detection. To compensate local features for transformers, we propose a Locality Compensation Block (LCB) containing a Global-Local Cross-Attention (GLCA) to attentively fuse global transformer features and local convolutional features. To aggregate features of all layers for capturing comprehensive and various fake flaws, we propose Multi-head Clustering Projection (MCP) and Frequency-guided Fusion Module (FFM), where the MCP attentively reduces redundant features into a few concentrated clusters, and the FFM interacts all clustered features under the guidance of frequency cues. In Trans-FCA, besides global cues captured by transformer architecture, local details and rich forgery defects are also captured using the proposed fetaure compensation and aggregation. Extensive experiments show our method outperforms the state-of-the-art methods on both intra-dataset and cross-dataset testings (with AUCs of 99.85% on FaceForensics++ and 78.57% on Celeb-DF), which clearly demonstrates the superiority of our Trans-FCA for deepfake detection.
Zichang Tan, Zhichao Yang 0008, Changtao Miao, Guodong Guo
IEEE Signal Process. Lett.3
2022 Hierarchical Frequency-Assisted Interactive Networks for Face Manipulation Detection
abstract
Recently, face manipulation techniques have caused increasing trust concerns in our society. Although current face manipulation detection methods achieve impressive performance regarding intra-dataset evaluation, they are struggling to improve the generalization and robustness ability. To address this issue, we propose a novel Hierarchical Frequency-assisted Interactive Networks (HFI-Net) to explore comprehensive frequency-related forgery cues for face manipulation detection. At first, we formulate HFI-Net as a dual-branch network to take full advantage of both CNN and transformer for capturing local details and global context information, respectively. Considering the forged faces are easy to show flaws in the frequency domain, a novel Frequency-based Feature Refinement (FFR) module is proposed to learn frequency-based attention from RGB features. FFR module emphasizes forgery cues and suppresses the pristine semantics information by keeping middle-high frequency features while discarding the low-frequency ones. Based on FFR, we further develop a co-sharing Global-Local Interaction (GLI) module to conduct frequency-assisted interactions while capturing complementarity among dual branches. Lastly, we further implement the GLI module in each stage of the network to effectively explore multi-level frequency artifacts. Extensive experiments are conducted on several popular benchmarks including FaceForensics++, Celeb-DF, DeepFake-TIMIT, DFDC, UADFV, and DeeperForensics-1.0, which shows that our model outperforms the state-of-the-art, especially in unseen datasets, manipulations, and perturbations evaluation.
Changtao Miao, Zichang Tan, Qi Chu 0001, Nenghai Yu, Guodong Guo
IEEE Trans. Inf. Forensics Secur.1
2021 DFGC 2021: A DeepFake Game Competition
abstract
This paper presents a summary of the DeepFake Game Competition (DFGC) 20211. DeepFake technology is developing fast, and realistic face-swaps are increasingly deceiving and hard to detect. At the same time, DeepFake detection methods are also improving. There is a two-party game between DeepFake creators and detectors. This competition provides a common platform for benchmarking the adversarial game between current state-of-the-art DeepFake creation and detection methods. In this paper, we present the organization, results and top solutions of this competition and also share our insights obtained during this event. We also release the DFGC-21 testing dataset collected from our participants to further benefit the research community2.
Bo Peng 0002, Hongxing Fan, Wei Wang 0025, Jing Dong 0003, Yuezun Li, Siwei Lyu, Qi Li 0005, Zhenan Sun, Baoying Chen, Yanjie Hu, Shenghai Luo, Junrui Huang, Yutong Yao, Boyuan Liu, Changtao Miao, Changlei Lu, Wanyi Zhuang
IJCB19
2021 Towards Generalizable and Robust Face Manipulation Detection via Bag-of-feature
abstract
Over the past several years, to solve the problem of malicious abuse of facial manipulation technology, face manipulation detection technology has obtained considerable attention and achieved remarkable progress. However, most existing methods have very impoverished generalization ability and robustness. In this paper, we propose a novel method for face manipulation detection, which can improve the generalization ability and ro-bustness by bag-of-feature. Specifically, we extend Transformers using bag-of-feature approach to encode inter-patch relation-ships, allowing it to learn forgery features without any additional mask supervision. Extensive experiments demonstrate that our method can outperform competing for state-of-the-art methods on FaceForensics++, Celeb-DF and DeeperForensics-l.0 datasets.
Changtao Miao, Qi Chu 0001, Weihai Li, Wanyi Zhuang, Nenghai Yu
VCIP1