VLDB 2026 Research / reviewers in the wild / expert
Xiaomeng Fu
dblp:330/7473
· DBLP profile ↗
19ranked-venue papers
7as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis From Language Models to PhysicsabstractSynthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent trade-off between semantic fidelity and physical realism. Large language model (LLM)-based approaches can interpret diverse open-vocabulary instructions and compose high-level action plans, but they often generate motions that violate physical constraints. Physics-aware models improve realism through simulation or control, but they struggle with semantic complexity, fine-grained instructions, and novel concepts. To address this gap, we propose In-Context Model Predictive Generation (ICMPG), a framework that integrates language-model planning with inference-time physical feedback. ICMPG reformulates motion synthesis as a Model Predictive Control (MPC)-like process with two modules. The Context-Aware Motion Generation (CAMG) module uses an LLM as a planner to decompose textual commands and generate candidate motion sequences from motion tokens. The Model Predictive Generation (MPG) module evaluates these candidates through physical simulation and semantic alignment, estimates a composite reward, and selects the best sequence to guide subsequent generation steps. Unlike open-loop generation, this closed-loop refinement enables ICMPG to adapt motions to both the input semantics and the simulated physical environment without task-specific policy retraining. Extensive experiments across standard and zero-shot open-vocabulary settings show that ICMPG generalizes robustly to diverse commands and produces motions that are more physically plausible and semantically faithful than representative baselines on the evaluated benchmarks. The framework bridges semantic interpretation and physical simulation while remaining flexible enough to incorporate different LLM backbones, enabling more versatile and controllable text-driven motion synthesis. Xiaomeng Fu, Junfan Lin, Yang Liu 0084, Yaowei Wang 0001, Guanbin Li, Liang Lin 0004, Ziliang Chen 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | TCFG: Truncated Classifier-Free Guidance for Efficient and Scalable Text-to-Image Acceleration
Xiaomeng Fu, Jia Li 0057 |
ICCV | 1 |
| 2025 | Resolution Attack: Exploiting Image Compression to Deceive Deep Neural NetworksabstractModel robustness is essential for ensuring the stability and reliability of machine learning systems. Despite extensive research on various aspects of model robustness, such as adversarial robustness and label noise robustness, the exploration of robustness towards different resolutions, remains less explored. To address this gap, we introduce a novel form of attack: the resolution attack. This attack aims to deceive both classifiers and human observers by generating images that exhibit different semantics across different resolutions. To implement the resolution attack, we propose an automated framework capable of generating dual-semantic images in a zero-shot manner. Specifically, we leverage large-scale diffusion models for their comprehensive ability to construct images and propose a staged denoising strategy to achieve a smoother transition across resolutions. Through the proposed framework, we conduct resolution attacks against various off-the-shelf classifiers. The experimental results exhibit high attack success rate, which not only validates the effectiveness of our proposed framework but also reveals the vulnerability of current classifiers towards different resolutions. Additionally, our framework, which incorporates features from two distinct objects, serves as a competitive tool for applications such as face swapping and facial camouflage. The code is available at https://github.com/ywj1/resolution-attack. Wangjia Yu, Xiaomeng Fu, Jizhong Han, Xiaodan Zhang 0004 |
ICLR | 2 |
| 2025 | OMS: One More Step Noise Searching to Enhance Membership Inference Attacks for Diffusion ModelsabstractThe data-intensive nature of Diffusion models amplifies the risks of privacy infringements and copyright disputes, particularly when training on extensive unauthorized data scraped from the Internet. Membership Inference Attacks (MIA) aim to determine whether a data sample has been utilized by the target model during training, thereby serving as a pivotal tool for privacy preservation. Current MIA employs the prediction loss to distinguish between training member samples and non-members. These methods assume that, compared to non-members, members, having been encountered by the model during training result in a smaller prediction loss. However, this assumption proves ineffective in diffusion models due to the random noise sampled during the training process. Rather than estimating the loss, our approach examines this random noise and reformulate the MIA as a noise search problem, assuming that members are more feasible to find the noise used in the training process. We formulate this noise search process as an optimization problem and employ the fixed-point iteration to solve it. We analyze current MIA methods through the lens of the noise search framework and reveal that they rely on the first residual as the discriminative metric to differentiate members and non-members. Inspired by this observation, we introduce OMS, which augments existing MIA methods by iterating One More fixed-point Step to include a further residual, i.e., the second residual. We integrate our method into various MIA methods across different diffusion models. The experimental results validate the efficacy of our proposed approach. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han, Xingyu Gao 0001 |
IJCAI | 1 |
| 2025 | Entity Graph Alignment and Visual Reasoning for Multimodal Fake News DetectionabstractThe rise of multimodal fake news threatens reliable information dissemination by exploiting multiple modalities to create deceptive, engaging content, significantly impacting society safety. Existing methods still face challenges in cross-modal alignment (e.g., semantic inconsistencies, complex visual-semantic relations) and are vulnerable to low-quality or noisy samples. To address these, we propose Cross-Modal Alignment with Visual Reasoning Prompting (CMA-VRP) for multimodal fake news detection. Specifically, we model text and image entities with graphs to capture fine-grained semantic interactions and enhance cross-modal consistency through graph contrastive learning. Unlike methods relying on shallow image features (e.g., edges, textures), we leverage large language models (LLMs) and large vision-language models (LVLMs) to capture deep visual-semantic attributes related to reasoning (e.g., actions, scenes). Based on graph modeling and visual reasoning features, we perform graph-based cross-modal semantic fusion to unify textual and visual representations and cross-modal cycle alignment to align modality distributions by reducing semantic discrepancies, filtering modality-specific noise, and extracting invariant representations across domains. These steps enable the model to obtain semantically consistent and modality-invariant features. Extensive experiments demonstrate that our model outperforms existing methods in multimodal fake news detection and shows strong robustness against noisy samples. Guoyi Li, Die Hu 0004, Xiaomeng Fu, Qirui Tang, Yulei Wu, Xiaodan Zhang 0004, Honglei Lyu |
ACM Multimedia | 3 |
| 2025 | Zero-Shot Multimodal Fact-Checking with Conceptual ReasoningabstractIn multimodal fact-checking, advanced large multimodal models (LMMs) struggle to capture and integrate the complex relationships between text and images. A potential solution is to generate reasoning support text to optimize reasoning and integrate evidence. However, existing generation approaches rely heavily on high-quality data annotations for training, which are costly and limited in scalability, hindering responsiveness to evolving misinformation. To address these issues, we propose CoReS, a novel zero-shot multimodal fact-checking model based on Conceptual Reasoning Support-leveraging key concepts from evidence to guide the reasoning process and improve decision-making. This model includes a reasoning support text generation module that extracts key concepts (critical elements that significantly impact the judgment outcome) from raw textual evidence via retrieval and filtering. By using a Conceptual Reasoning LM, CoReS generates reasoning support texts framed around core key concepts that are semantically consistent with multimodal evidence, linking key clues, thus replacing redundant and complex evidence for fact-checking. The reasoning support texts generated by CoReS effectively distill complex evidence relationships and integrate important reasoning information, allowing the judgment model to provide clear and accurate judgments. Evaluations on benchmark datasets and the new multi-domain MultiVerify dataset demonstrate that CoReS excels in accuracy, generalization, and scalability. Guoyi Li, Die Hu 0004, Qirui Tang, Xiaomeng Fu, Yulei Wu, Xiaodan Zhang 0004, Honglei Lyu |
ACM Multimedia | 5 |
| 2025 | Unlocking Generative Priors: A New Membership Inference Framework for Diffusion ModelsabstractDiffusion models pose risks of privacy breaches and copyright disputes, primarily stemming from the potential utilization of unauthorized data during the training phase. Membership inference is aimed to determine whether a specific sample has been used in the training process of a target model, representing a critical tool for privacy violation verification. However, the increased model complexity and stochasticity inherent in diffusion renders traditional shadow-model-based or metric-based methods ineffective when applied to diffusion models. Moreover, existing methods only yield binary classification labels which lack necessary comprehensibility in practical applications. In this paper, we explore a novel perspective for membership inference by leveraging the intrinsic generative priors within the diffusion model. Compared with unseen samples, training samples exhibit stronger generative priors within the diffusion model, enabling the successful reconstruction of substantially degraded training images. Consequently, we propose the Degrade Restore Compare (DRC) framework. In this framework, an image undergoes sequential degradation and restoration, and its membership is determined by comparing it with the restored counterpart. Experimental results verify that our approach not only significantly outperforms existing methods in terms of accuracy but also provides comprehensible decision criteria, offering evidence for potential privacy violations. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han, Xingyu Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Generative Transferable Universal Adversarial Perturbation for Combating DeepfakesabstractRecently, Deepfake has posed a significant threat to our digital society. This technology allows for the modification of facial identity, expression, and attributes in facial images and videos. The misuse of Deepfake can invade personal privacy, damage individuals’ reputations, and have serious consequences. To counter this threat, researchers have proposed active defense methods using adversarial perturbation to distort Deepfake products which can hinder the dissemination of false information. However, the existing methods are primarily based on image-specific approaches, which are inefficient for large-scale data. To address these issues, we propose an end-to-end approach to generate universal perturbations for combating Deepfake. To further cope with diverse Deepfakes, we introduce an adaptive balancing strategy to combat multiple models simultaneously. Specifically, for different scenarios, we propose two types of universal perturbations. Disrupting Universal Perturbation (DUP) leads Deepfake models to generate distorted outputs. In contrast, Lapsing Universal Perturbation (LUP) tries to make the output consistent with the original image, allowing the correct information to continue propagating. Experiments demonstrate the effectiveness and better generalization of our proposed perturbation compared with state-of-the-art methods. Consequently, our proposed method offers a powerful and efficient solution for combating Deepfake, which can help preserve personal privacy and prevent reputational damage. Xi Wang 0014, Xiaomeng Fu, Jin Liu 0020, Zhaoxing Li, Yesheng Chai, Jizhong Han |
CSCWD | 3 |
| 2024 | Generative Universal Nullifying Perturbation for Countering Deepfakes Through Combined Unsupervised Feature Aggregation
Xi Wang 0014, Xiaomeng Fu, Jin Liu 0020, Zhaoxing Li, Jizhong Han |
ICANN (2) | 3 |
| 2024 | Explicit Correlation Learning for Generalizable Cross-Modal Deepfake DetectionabstractWith the rising prevalence of deepfakes, there is a growing interest in developing generalizable detection methods for various types of deepfakes. While effective in their specific modalities, traditional detection methods fall short in addressing the generalizability of detection across diverse cross-modal deepfakes. This paper aims to explicitly learn potential cross-modal correlation to enhance deepfake detection towards various generation scenarios. Our approach introduces a correlation distillation task, which models the inherent cross-modal correlation based on content information. This strategy helps to prevent the model from overfitting merely to audio-visual synchronization. Additionally, we present the Cross-Modal Deepfake Dataset (CMDFD), a comprehensive dataset with four generation methods to evaluate the detection of diverse cross-modal deepfakes. The experimental results on CMDFD and FakeAVCeleb datasets demonstrate the superior generalizability of our method over existing state-of-the-art methods. Our code and data can be found at https://github.com/ljj898/CMDFD-Dataset-and-Deepfake-Detection. Cai Yu, Shan Jia, Xiaomeng Fu, Jin Liu 0020, Jiao Dai, Xi Wang 0014, Siwei Lyu, Jizhong Han |
ICME | 3 |
| 2024 | Unveiling Structural Memorization: Structural Membership Inference Attack for Text-to-Image Diffusion ModelsabstractWith the rapid advancements of large-scale text-to-image diffusion models, various practical applications have emerged, bringing significant convenience to society. However, model developers may misuse the unauthorized data to train diffusion models. These data are at risk of being memorized by the models, thus potentially violating citizens' privacy rights. Therefore, in order to judge whether a specific image is utilized as a member of a model's training set, Membership Inference Attack (MIA) is proposed to serve as a tool for privacy protection. Current MIA methods predominantly utilize pixel-wise comparisons as distinguishing clues, considering the pixel-level memorization characteristic of diffusion models. However, it is practically impossible for text-to-image models to memorize all the pixel-level information in massive training sets. Therefore, we move to the more advanced structure-level memorization. Observations on the diffusion process show that the structures of members are better preserved compared to those of nonmembers, indicating that diffusion models possess the capability to remember the structures of member images from training sets. Drawing on these insights, we propose a simple yet effective MIA method tailored for text-to-image diffusion models. Extensive experimental results validate the efficacy of our approach. Compared to current pixel-level baselines, our approach not only achieves state-of-the-art performance but also demonstrates remarkable robustness against various distortions. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Xingyu Gao 0001, Jiao Dai, Jizhong Han |
ACM Multimedia | 2 |
| 2024 | OSM-Net: One-to-Many One-Shot Talking Head Generation With Spontaneous Head MotionsabstractOne-shot talking head generation has no explicit head movement reference, thus it is difficult to generate talking heads with head motions. Some existing works only edit the mouth area and generate still talking heads, leading to unreal talking head performance. Other works construct one-to-one mapping between audio signal and head motion sequences, introducing ambiguity correspondences into the mapping since people can behave differently in head motions when speaking the same content. This unreasonable mapping form fails to model the diversity and produces either nearly static or even exaggerated head motions, which are unnatural and strange. Therefore, the one-shot talking head generation task is actually a one-to-many ill-posed problem and people present diverse head motions when speaking. Based on the above observation, we propose OSM-Net, aone-to-manyone-shot talking head generation network with natural head motions. OSM-Net constructs a motion space that contains rich and various clip-level head motion features. Each basis of the space represents a feature of meaningful head motion in a clip rather than just a frame, thus providing more coherent and natural motion changes in talking heads. The driving audio is mapped into the motion space, around which various motion features can be sampled within a reasonable range to achieve the one-to-many mapping. Besides, the landmark constraint and time window feature input improve the accurate expression feature extraction and video generation. Extensive experiments show that OSM-Net generates more natural realistic head motions under reasonable one-to-many mapping paradigm compared with other methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Learning to Discover Forgery Cues for Face Forgery DetectionabstractLocating manipulation maps,i.e., pixel-level annotation of forgery cues, is crucial for providing interpretable detection results in face forgery detection. Related learning objects have also been widely adopted as auxiliary tasks to improve the classification performance of detectors whereas they require comparisons between paired real and forged faces to obtain manipulation maps as supervision. This requirement restricts their applicability to unpaired faces and contradicts real-world scenarios. Moreover, the used comparison methods annotate all changed pixels, including noise introduced by compression and upsampling. Using such maps as supervision hinders the learning of exploitable cues and makes models prone to overfitting. To address these issues, we introduce a weakly supervised model in this paper, named Forgery Cue Discovery (FoCus), to locate forgery cues in unpaired faces. Unlike some detectors that claim to locate forged regions in attention maps, FoCus is designed to sidestep their shortcomings of capturing partial and inaccurate forgery cues. Specifically, we propose a classification attentive regions proposal module to locate forgery cues during classification and a complementary learning module to facilitate the learning of richer cues. The produced manipulation maps can serve as better supervision to enhance face forgery detectors. Visualization of the manipulation maps of the proposed FoCus exhibits superior interpretability and robustness compared to existing methods. Experiments on five datasets and four multi-task models demonstrate the effectiveness of FoCus in both in-dataset and cross-dataset evaluations. Cai Yu, Xiaomeng Fu, Xi Wang 0014, Jiao Dai, Jizhong Han |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | OPT: One-shot Pose-Controllable Talking Head GenerationabstractOne-shot talking head generation produces lip-sync talking heads based on arbitrary audio and one source face. To guarantee the naturalness and realness, recent methods propose to achieve free pose control instead of simply editing mouth areas. However, existing methods do not preserve accurate identity of source face when generating head motions. To solve the identity mismatch problem and achieve high-quality free pose control, we present One-shot Pose-controllable Talking head generation network (OPT). Specifically, the Audio Feature Disentanglement Module separates content features from audios, eliminating the influence of speaker-specific information contained in arbitrary driving audios. Later, the mouth expression feature is extracted from the content feature and source face, during which the landmark loss is designed to enhance the accuracy of facial structure and identity preserving quality. Finally, to achieve free pose control, controllable head pose features from reference videos are fed into the Video Generator along with the expression feature and source face to generate new talking heads. Extensive quantitative and qualitative experimental results verify that OPT generates high-quality pose-controllable talking heads with no identity mismatch problem, outperforming previous SOTA methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ICASSP | 3 |
| 2023 | Large Pose Friendly Face Reenactment using subtle motionsabstractFace reenactment aims to synthesis a photo-realistic video of the source face by imitating the motion and expression of the driving video while keeping the source appearance (i.e. identity). Although good results have achieved recently, most state-of-the-art methods remain vulnerable to extreme conditions, which greatly restricts the application in the real world. Among various extreme conditions, the large pose problem is the most common one. We clarify that the large pose problem is mainly caused by the severe motion change between the source image and the current driving frame. An intuitive solution is to divide the severe motion change into a sequence of subtle motions. Therefore, we propose a new scheme that exploring the temporal coherence between previous neighbor frame and current frame. The smaller motion change between consecutive frames help to solve the large pose problem. Furthermore, a calibration net is designed to eliminate the error accumulation of the previous step. Extensive experiments demonstrate that our method performs better on large pose face reenactment than the state-of-the-art in terms of large pose cases and visual quality. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han |
ICME | 1 |
| 2023 | FONT: Flow-guided One-shot Talking Head Generation with Natural Head MotionsabstractOne-shot talking head generation has received growing attention in recent years, with various creative and practical applications. An ideal natural and vivid generated talking head video should contain natural head pose changes. However, it is challenging to map head pose sequences from driving audio since there exists a natural gap between audio-visual modalities. In this work, we propose a Flow-guided One-shot model that achieves NaTural head motions(FONT) over generated talking heads. Specifically, we design a probabilistic CVAE-based model to predict head pose sequences from driving audio and source face. Then we develop a keypoint predictor that produces unsupervised keypoints describing the facial structure information from the source face, driving audio and pose sequences. Finally, a flow- guided occlusion-aware generator is employed to produce photo-realistic talking head videos from the estimated keypoints and source face. Extensive experimental results prove that FONT generates talking heads with natural head poses and synchronized mouth shapes, outperforming other compared methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ICME | 3 |
| 2023 | MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion ModelabstractFace-to-face communication is a common scenario including roles of speakers and listeners. Most existing research methods focus on producing speaker videos, while the generation of listener heads remains largely overlooked. Responsive listening head generation is an important task that aims to model face-to-face communication scenarios by generating a listener head video given a speaker video and a listener head image. An ideal generated responsive listening video should respond to the speaker with attitude or viewpoint expressing while maintaining diversity in interaction patterns and accuracy in listener identity information. To achieve this goal, we propose the Multi-Faceted Responsive Listening Head Generation Network (MFR-Net). Specifically, MFR-Net employs the probabilistic denoising diffusion model to predict diverse head pose and expression features. In order to perform multi-faceted response to the speaker video, while maintaining accurate listener identity preservation, we design the Feature Aggregation Module to boost listener identity features and fuse them with other speaker-related features. Finally, a renderer finetuned with identity consistency loss produces the final listening head videos. Our extensive experiments demonstrate that MFR-Net not only achieves multi-faceted responses in diversity and speaker identity information but also in attitude and viewpoint expression. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ACM Multimedia | 3 |
| 2023 | CoP: Chain-of-Pose for Image Animation in Large Pose ChangesabstractImage animation involves generating a video of a source image imitating the pose of a driving video. Despite recent advancements in the image animation task, most state-of-the-art methods remain vulnerable to large pose changes. In cases of large pose changes, existing methods struggle to model the complex nonlinear motion and yield distorted results, which greatly restricts their application in the real world. To tackle this problem, we present a novel approach called Chain-of-Pose (CoP) that decomposes large pose changes into a sequence of intermediate pose changes. This enables us to handle simplified pose changes and improves the accuracy of pose estimation. Furthermore, to better preserve the appearance of the source object, we introduce the Appearance Refinement Module (ARM) that effectively integrates the appearance texture feature of the source image with the structural pose feature from the pose chain. Our experimental results demonstrate that our method qualitatively and quantitatively outperforms state-of-the-art approaches on four diverse datasets, comprising talking faces, human bodies, and pixel animals. Notably, our approach significantly improves video quality in the case of large object pose changes. Our code is attached to the supplementary material. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Shuhui Wang, Jiao Dai, Jizhong Han |
ACM Multimedia | 1 |
| 2022 | MakeItSmile: Detail-Enhanced Smiling Face ReenactmentabstractGiven a target face and a driving face, face reenactment aims to transfer attributes from the driving face to the target face. In the last decade, a great number of methods have been proposed to generate realistic reenacted faces. However, when these methods are applied to generate a smiling face, most of them can only get a mouth with blurry teeth, making the reenacted face unrealistic. This problem is mainly caused by incomplete tooth structure in the target face image under the setting of one-shot reenactment. In order to obtain smiling reenacted faces with detailed tooth structure, our method uses the tooth information from the driving face rather than the target face. Furthermore, to better represent the tooth structure and expressions of the driving face, we extract the texture with a carefully designed geometry-aware encoder. By training the encoder with tooth segmentation task and non-identity classification task, we acquire refined tooth representations and meanwhile derive the non-identity part of the driving face. We also design a specific generator to fuse the tooth texture features into the target face. Moreover, we add a mouth loss function to further ensure the high definition of the smiling reenacted face. We compare our method to existing state-of-the-art approaches. The experiments show that our method gets comparable results on non-smiling face reenactment and has superior performance on smiling face reenactment. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Wantao Liu, Jiao Dai, Jizhong Han |
IJCNN | 1 |