VLDB 2026 Research / reviewers in the wild / expert
Jiao Dai
dblp:02/9034
· DBLP profile ↗
43ranked-venue papers
1as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 22 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 11 since 2021Security and privacy · 4 · 4 since 2021Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Cognitive Distribution and Behavior-Consistent Framework for Black-Box Attacks on Recommender Systems
Hongyue Zhang, Dongqin Liu, Honglei Lv, Jiao Dai, Jizhong Han |
DASFAA (1) | 8 |
| 2026 | FreeEdit: Mask-Free Reference-Based Image Editing With Multi-Modal InstructionabstractIntroducing user-specified visual concepts in image editing is highly practical as these concepts convey the user's intent more precisely than text-based descriptions. We propose FreeEdit, a novel approach for achieving such reference-based image editing, which can accurately reproduce the visual concept from the reference image based on user-friendly language instructions. Our approach leverages the multi-modal instruction encoder to encode language instructions to guide the editing process. This implicit way of locating the editing area eliminates the need for manual editing masks. To enhance the reconstruction of reference details, we introduce the Decoupled Residual Refer-Attention (DRRA) module. This module is designed to integrate fine-grained reference features extracted by a detail extractor into the image editing process in a residual way without interfering with the original self-attention. Given that existing datasets are unsuitable for reference-based image editing tasks, particularly due to the difficulty in constructing image triplets that include a reference image, we curate a high-quality dataset, FreeBench, using a newly developed twice-repainting scheme. FreeBench comprises the images before and after editing, detailed editing instructions, as well as a reference image that maintains the identity of the edited object, encompassing tasks such as object addition, replacement, and deletion. By conducting phased training on FreeBench followed by quality tuning, FreeEdit achieves high-quality zero-shot editing through convenient language instructions. We conduct extensive experiments to evaluate the effectiveness of FreeEdit across multiple task types, demonstrating its superiority over existing methods. Runze He, Linjiang Huang, Shaofei Huang 0001, Jialin Gao, Xiaoming Wei, Jiao Dai, Jizhong Han, Si Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Diagnosing RAG vulnerabilities via sensitive knowledge recognition
Zhichen Liu, Jingkai Liu, Xiaoting Lyu, Jiao Dai, Fuqiang Hu |
Pattern Recognit. Lett. | 4 |
| 2025 | Boosting Open-Vocabulary Object Detection Performance via Class-Agnostic Pseudo-Labels and MultiModal Hybrid KnowledgeabstractOpen-vocabulary object detection (OVD) is a significant task identifying objects from categories not included in the training set. Our comprehensive analysis reveals two main issues with existing OVD models: poor generalization of localization network to novel categories and poor quality of class embedding impacting accuracy. We propose two solutions: Localization Network Enhancement based on Class-Agnostic Pseudo-Labels (LNE-CAPL) and Class Embedding Enhancement based on MultiModal Hybrid Knowledge (CEE-MMHK). These methods can be used offline to significantly improve the existing OVD models without affecting their training and inference efficiency. Extensive experiments confirm their effectiveness and universality. Code and models will soon be open-sourced. Dongqin Liu, Jiao Dai, Songlin Hu 0001 |
ICASSP | 3 |
| 2025 | Modality-Agnostic Deepfakes DetectionabstractAs AI-generated content (AIGC) thrives, deepfakes have expanded from single-modality falsification to cross-modal fake content creation, where either audio or visual components can be manipulated.While using two unimodal detectors can detect audio-visual deepfakes, cross-modal forgery clues could be overlooked.Existing multimodal deepfake detectors typically establish correspondence between the audio and visual modalities for binary real/fake classification and require the co-occurrence of both modalities.However, in real-world multi-modal applications, missing modality scenarios may occur where either modality is unavailable.In such cases, audio-visual detection methods are less practical than two independent unimodal methods.Consequently, the detector can not always obtain the number or type of manipulated modalities beforehand, necessitating a fake-modality-agnostic audio-visual detector.In this work, we introduce a comprehensive framework that is agnostic to fake modalities, which facilitates the identification of multimodal deepfakes and handles situations with missing modalities, regardless of the manipulations embedded in audio, video, or even cross-modal forms.To enhance the modeling of cross-modal forgery clues, we employ audio-visual speech recognition (AVSR) Jin Liu 0020, Jiao Dai, Xi Wang 0014, Shan Jia, Siwei Lyu, Jizhong Han |
IH&MMSec | 5 |
| 2025 | OMS: One More Step Noise Searching to Enhance Membership Inference Attacks for Diffusion ModelsabstractThe data-intensive nature of Diffusion models amplifies the risks of privacy infringements and copyright disputes, particularly when training on extensive unauthorized data scraped from the Internet. Membership Inference Attacks (MIA) aim to determine whether a data sample has been utilized by the target model during training, thereby serving as a pivotal tool for privacy preservation. Current MIA employs the prediction loss to distinguish between training member samples and non-members. These methods assume that, compared to non-members, members, having been encountered by the model during training result in a smaller prediction loss. However, this assumption proves ineffective in diffusion models due to the random noise sampled during the training process. Rather than estimating the loss, our approach examines this random noise and reformulate the MIA as a noise search problem, assuming that members are more feasible to find the noise used in the training process. We formulate this noise search process as an optimization problem and employ the fixed-point iteration to solve it. We analyze current MIA methods through the lens of the noise search framework and reveal that they rely on the first residual as the discriminative metric to differentiate members and non-members. Inspired by this observation, we introduce OMS, which augments existing MIA methods by iterating One More fixed-point Step to include a further residual, i.e., the second residual. We integrate our method into various MIA methods across different diffusion models. The experimental results validate the efficacy of our proposed approach. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han, Xingyu Gao 0001 |
IJCAI | 5 |
| 2025 | Enhancing Configuration Security in Unmanned Systems via Static Analysis and Fuzz Testing
Chongzhen Zhang, Jiao Dai, Fuqiang Hu, Dantong Yan, Hongquan Tian |
NSS | 4 |
| 2025 | PFedKD: Personalized Federated Learning via Knowledge Distillation Using Unlabeled Pseudo Data for Internet of ThingsabstractWith the rapid advancement of wearable devices and Internet of Things (IoT) technologies, sensor data generated by edge devices has surged. This data is crucial for advancing IoT applications, including health status monitoring, abnormal behavior detection, and environmental monitoring. However, traditional centralized learning requires uploading data to a central server, raising security and privacy concerns and hindering data application. Federated learning (FL) offers a solution by enabling collaborative model training on IoT devices without transferring data from the local device. In practice, edge devices generate data that is often highly heterogeneous, making it challenging for the global FL model to capture local data distributions accurately, leading to significant performance degradation. Additionally, imbalanced edge device resources and limited bandwidth can cause data transmission delays or interruptions, impacting application feasibility. To address these issues, we propose PFedKD, a novel personalized FL algorithm based on knowledge distillation, aimed at enhancing the model’s generalization ability and reducing communication overhead in heterogeneous IoT data environments. PFedKD constructs a public dataset using unlabeled pseudo data to extract knowledge from each client, training personalized models that fit local data distributions. This method controls dataset size while enhancing performance. During communication, only logits and class prototypes are transmitted, ensuring high communication efficiency. Sharpness aware minimization is introduced in local model training to optimize generalization. Additionally, we design a weight distribution mechanism based on client sample quality evaluation that optimizes knowledge aggregation and model personalization. Extensive experiments demonstrate that PFedKD significantly outperforms state-of-the-art baselines in both learning performance and communication efficiency. Bin Wang 0051, Yongsheng Zhu, Fuqiang Hu, Jiao Dai, Wei Wang 0012 |
IEEE Internet Things J. | 7 |
| 2025 | Unlocking Generative Priors: A New Membership Inference Framework for Diffusion ModelsabstractDiffusion models pose risks of privacy breaches and copyright disputes, primarily stemming from the potential utilization of unauthorized data during the training phase. Membership inference is aimed to determine whether a specific sample has been used in the training process of a target model, representing a critical tool for privacy violation verification. However, the increased model complexity and stochasticity inherent in diffusion renders traditional shadow-model-based or metric-based methods ineffective when applied to diffusion models. Moreover, existing methods only yield binary classification labels which lack necessary comprehensibility in practical applications. In this paper, we explore a novel perspective for membership inference by leveraging the intrinsic generative priors within the diffusion model. Compared with unseen samples, training samples exhibit stronger generative priors within the diffusion model, enabling the successful reconstruction of substantially degraded training images. Consequently, we propose the Degrade Restore Compare (DRC) framework. In this framework, an image undergoes sequential degradation and restoration, and its membership is determined by comparing it with the restored counterpart. Experimental results verify that our approach not only significantly outperforms existing methods in terms of accuracy but also provides comprehensible decision criteria, offering evidence for potential privacy violations. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han, Xingyu Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Customize your NeRF: Adaptive Source Driven 3D Scene Editing via Local-Global Iterative TrainingabstractIn this paper, we target the adaptive source driven 3D scene editing task by proposing a CustomNeRF model that unifies a text description or a reference image as the editing prompt. However, obtaining desired editing results conformed with the editing prompt is nontrivial since there exist two significant challenges, including accurate editing of only foreground regions and multi-view consistency given a single-view reference image. To tackle the first challenge, we propose a Local-Global Iterative Editing (LGIE) training scheme that alternates between foreground region editing and full-image editing, aimed at foreground-only manipulation while preserving the background. For the second challenge, we also design a class-guided regularization that exploits class priors within the generation model to alleviate the inconsistency problem among different views in image-driven editing. Extensive experiments show that our CustomNeRF produces precise editing results under various real scenes for both text- and image-driven settings. The code is available at: https://github.com/hrz2000/CustomNeRF. Runze He, Shaofei Huang 0001, Xuecheng Nie, Tianrui Hui, Luoqi Liu, Jiao Dai, Jizhong Han, Guanbin Li, Si Liu 0001 |
CVPR | 6 |
| 2024 | Real Appearance Modeling for More General Deepfake Detection
Cai Yu, Xi Wang 0014, Zihao Xiao 0002, Jiao Dai, Jizhong Han, Yesheng Chai |
ECCV (52) | 6 |
| 2024 | ConfR: Conflict Resolving for Generalizable Deepfake DetectionabstractDeepfake detectors often encounter performance degradation when tested on unseen forgery methods. Existing literature tries to capture common features among multiple source forgery domains. However, we show that conflict arises in the shared feature space when each domain expresses domain-specific bias. If left unresolved, this conflict might mislead the model to learn domain-specific features and lead to inferior generalization. In this paper, we propose a new learning approach, Conflict Resolving (ConfR), designed to minimize conflict and learn features that generalize across forgeries. ConfR incorporates two key elements: the Intra-Domain Consistency Preserving (ICP) loss ensures updating consistency within forgery types, and the Inter-Domain Conflict Resolving (ICR) Module resolves updating conflicts between different forgery types. Extensive experiments demonstrate that ConfR significantly improves upon the state-of-the-art method, highlighting its potential for more generalizable deepfake detection. Cai Yu, Xi Wang 0014, Zhaoxing Li, Yesheng Chai, Jiao Dai, Jizhong Han |
ICME | 7 |
| 2024 | HIDD: Human-perception-centric Incremental Deepfake DetectionabstractFacial manipulation techniques pose a significant societal threat due to the widespread dissemination of deepfake content on the internet. Existing efforts for deepfake detection exhibit inadequate generalization performance when encountering unseen or degraded samples. We attribute this limitation to the overfitting of minor forgery patterns and variations in data distribution among disparate datasets. To tackle this issue, we introduce an innovative human-perception-centric incremental deepfake detection framework to enhance the generalization capabilities of deepfake detection models through continuous learning from a limited set of new samples. Firstly, the model leverages human perceptual salience to discern and comprehend significant artifacts, thereby mitigating overfitting to minor features. Subsequently, in the incremental learning process, we utilize multi-perspective knowledge distillation and a replay strategy to maintain the performance of the old model and minimize the feature distance between old and new samples. This comprehensive approach mitigates feature-level overfitting and addresses distribution differences among various datasets in the incremental phase. We conducted thorough experiments on four benchmark datasets (FF++, DFDC-P, CDF2, and DFD), and the experimental results demonstrate the superior performance of our method. Xiaorong Ma, Yesheng Chai, Zhaoxing Li, Jiao Dai, Liangjun Zang, Jizhong Han |
ICME | 6 |
| 2024 | Explicit Correlation Learning for Generalizable Cross-Modal Deepfake DetectionabstractWith the rising prevalence of deepfakes, there is a growing interest in developing generalizable detection methods for various types of deepfakes. While effective in their specific modalities, traditional detection methods fall short in addressing the generalizability of detection across diverse cross-modal deepfakes. This paper aims to explicitly learn potential cross-modal correlation to enhance deepfake detection towards various generation scenarios. Our approach introduces a correlation distillation task, which models the inherent cross-modal correlation based on content information. This strategy helps to prevent the model from overfitting merely to audio-visual synchronization. Additionally, we present the Cross-Modal Deepfake Dataset (CMDFD), a comprehensive dataset with four generation methods to evaluate the detection of diverse cross-modal deepfakes. The experimental results on CMDFD and FakeAVCeleb datasets demonstrate the superior generalizability of our method over existing state-of-the-art methods. Our code and data can be found at https://github.com/ljj898/CMDFD-Dataset-and-Deepfake-Detection. Cai Yu, Shan Jia, Xiaomeng Fu, Jin Liu 0020, Jiao Dai, Xi Wang 0014, Siwei Lyu, Jizhong Han |
ICME | 6 |
| 2024 | HDDA: Human-perception-centric Deepfake Detection AdapterabstractFacial manipulation techniques pose a significant societal threat due to the prevalent presence of deepfake content online. Current deepfake detection methods demonstrate subpar generalization performance when applied to unseen samples. The cause of this limitation lies in the overfitting of minor forgery patterns and variations in data distribution across different datasets. To tackle this issue, we introduce an innovative Human-perception-centric Deepfake Detection Adapter, namely HDDA, to enhance the generalization ability of deepfake detection models. This adaptation primarily involves two stages. During the pre-training stage, the model utilizes human perception salience to spot significant artifacts, thus reducing overfitting to minor features. In the subsequent fine-tuning stage, we introduce an efficient parameter tuning module named Deepfake Detection Adapter. The Adapter introduces two types of lightweight yet specialized adapter modules to the pre-trained model while keeping the backbone network frozen. It fine-tunes the pre-trained model through the adapter to adapt new and unseen datasets, thereby enhancing generalization. We conducted comprehensive experiments on various standard deepfake detection benchmarks to validate the effectiveness of our approach, particularly in showcasing a compelling advantage under cross-dataset and cross-manipulation settings. Xiaorong Ma, Yesheng Chai, Jiao Dai, Zhaoxing Li, Liangjun Zang, Jizhong Han |
IJCNN | 4 |
| 2024 | Unveiling Structural Memorization: Structural Membership Inference Attack for Text-to-Image Diffusion ModelsabstractWith the rapid advancements of large-scale text-to-image diffusion models, various practical applications have emerged, bringing significant convenience to society. However, model developers may misuse the unauthorized data to train diffusion models. These data are at risk of being memorized by the models, thus potentially violating citizens' privacy rights. Therefore, in order to judge whether a specific image is utilized as a member of a model's training set, Membership Inference Attack (MIA) is proposed to serve as a tool for privacy protection. Current MIA methods predominantly utilize pixel-wise comparisons as distinguishing clues, considering the pixel-level memorization characteristic of diffusion models. However, it is practically impossible for text-to-image models to memorize all the pixel-level information in massive training sets. Therefore, we move to the more advanced structure-level memorization. Observations on the diffusion process show that the structures of members are better preserved compared to those of nonmembers, indicating that diffusion models possess the capability to remember the structures of member images from training sets. Drawing on these insights, we propose a simple yet effective MIA method tailored for text-to-image diffusion models. Extensive experimental results validate the efficacy of our approach. Compared to current pixel-level baselines, our approach not only achieves state-of-the-art performance but also demonstrates remarkable robustness against various distortions. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Xingyu Gao 0001, Jiao Dai, Jizhong Han |
ACM Multimedia | 6 |
| 2024 | OSM-Net: One-to-Many One-Shot Talking Head Generation With Spontaneous Head MotionsabstractOne-shot talking head generation has no explicit head movement reference, thus it is difficult to generate talking heads with head motions. Some existing works only edit the mouth area and generate still talking heads, leading to unreal talking head performance. Other works construct one-to-one mapping between audio signal and head motion sequences, introducing ambiguity correspondences into the mapping since people can behave differently in head motions when speaking the same content. This unreasonable mapping form fails to model the diversity and produces either nearly static or even exaggerated head motions, which are unnatural and strange. Therefore, the one-shot talking head generation task is actually a one-to-many ill-posed problem and people present diverse head motions when speaking. Based on the above observation, we propose OSM-Net, aone-to-manyone-shot talking head generation network with natural head motions. OSM-Net constructs a motion space that contains rich and various clip-level head motion features. Each basis of the space represents a feature of meaningful head motion in a clip rather than just a frame, thus providing more coherent and natural motion changes in talking heads. The driving audio is mapped into the motion space, around which various motion features can be sampled within a reasonable range to achieve the one-to-many mapping. Besides, the landmark constraint and time window feature input improve the accurate expression feature extraction and video generation. Extensive experiments show that OSM-Net generates more natural realistic head motions under reasonable one-to-many mapping paradigm compared with other methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Learning to Discover Forgery Cues for Face Forgery DetectionabstractLocating manipulation maps,i.e., pixel-level annotation of forgery cues, is crucial for providing interpretable detection results in face forgery detection. Related learning objects have also been widely adopted as auxiliary tasks to improve the classification performance of detectors whereas they require comparisons between paired real and forged faces to obtain manipulation maps as supervision. This requirement restricts their applicability to unpaired faces and contradicts real-world scenarios. Moreover, the used comparison methods annotate all changed pixels, including noise introduced by compression and upsampling. Using such maps as supervision hinders the learning of exploitable cues and makes models prone to overfitting. To address these issues, we introduce a weakly supervised model in this paper, named Forgery Cue Discovery (FoCus), to locate forgery cues in unpaired faces. Unlike some detectors that claim to locate forged regions in attention maps, FoCus is designed to sidestep their shortcomings of capturing partial and inaccurate forgery cues. Specifically, we propose a classification attentive regions proposal module to locate forgery cues during classification and a complementary learning module to facilitate the learning of richer cues. The produced manipulation maps can serve as better supervision to enhance face forgery detectors. Visualization of the manipulation maps of the proposed FoCus exhibits superior interpretability and robustness compared to existing methods. Experiments on five datasets and four multi-task models demonstrate the effectiveness of FoCus in both in-dataset and cross-dataset evaluations. Cai Yu, Xiaomeng Fu, Xi Wang 0014, Jiao Dai, Jizhong Han |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2023 | Anchor3DLane: Learning to Regress 3D Anchors for Monocular 3D Lane DetectionabstractMonocular 3D lane detection is a challenging task due to its lack of depth information. A popular solution is to first transform the front-viewed (FV) images or features into the bird-eye-view (BEV) space with inverse perspective mapping (IPM) and detect lanes from BEV features. However, the reliance of IPM on flat ground assumption and loss of context information make it inaccurate to restore 3D information from BEV representations. An attempt has been made to get rid of BEV and predict 3D lanes from FV representations directly, while it still underperforms other BEV-based methods given its lack of structured representation for 3D lanes. In this paper, we define 3D lane anchors in the 3D space and propose a BEV-free method named Anchor3DLane to predict 3D lanes directly from FV representations. 3D lane anchors are projected to the FV features to extract their features which contain both good structural and context information to make accurate predictions. In addition, we also develop a global optimization method that makes use of the equal-width property between lanes to reduce the lateral error of predictions. Extensive experiments on three popular 3D lane detection benchmarks show that our Anchor3DLane outperforms previous BEV-based methods and achieves state-of-the-art performances. The code is available at: https://github.com/tusenai/Anchor3DLane. Shaofei Huang 0001, Zhenwei Shen, Zehao Huang, Jiao Dai, Jizhong Han, Naiyan Wang, Si Liu 0001 |
CVPR | 5 |
| 2023 | Bridging Search Region Interaction with Template for RGB-T TrackingabstractRGB-T tracking aims to leverage the mutual enhancement and complement ability of RGB and TIR modalities for improving the tracking process in various scenarios, where cross-modal interaction is the key component. Some previous methods concatenate the RGB and TIR search region features directly to perform a coarse interaction process with redundant background noises introduced. Many other methods sample candidate boxes from search frames and conduct various fusion approaches on isolated pairs of RGB and TIR boxes, which limits the cross-modal interaction within local regions and brings about inadequate context modeling. To alleviate these limitations, we propose a novel Template-Bridged Search region Interaction (TBSI) module which exploits templates as the medium to bridge the cross-modal interaction between RGB and TIR search regions by gathering and distributing target-relevant object and environment contexts. Original templates are also updated with enriched multimodal contexts from the template medium. Our TBSI module is inserted into a ViT backbone for joint feature extraction, search-template matching, and cross-modal interaction. Extensive experiments on three popular RGB-T tracking benchmarks demonstrate our method achieves new state-of-the-art performances. Code is available at https://github.com/RyanHTR/TBSI. Tianrui Hui, Zizheng Xun, Fengguang Peng, Junshi Huang, Xiaoming Wei, Xiaolin Wei, Jiao Dai, Jizhong Han, Si Liu 0001 |
CVPR | 7 |
| 2023 | OPT: One-shot Pose-Controllable Talking Head GenerationabstractOne-shot talking head generation produces lip-sync talking heads based on arbitrary audio and one source face. To guarantee the naturalness and realness, recent methods propose to achieve free pose control instead of simply editing mouth areas. However, existing methods do not preserve accurate identity of source face when generating head motions. To solve the identity mismatch problem and achieve high-quality free pose control, we present One-shot Pose-controllable Talking head generation network (OPT). Specifically, the Audio Feature Disentanglement Module separates content features from audios, eliminating the influence of speaker-specific information contained in arbitrary driving audios. Later, the mouth expression feature is extracted from the content feature and source face, during which the landmark loss is designed to enhance the accuracy of facial structure and identity preserving quality. Finally, to achieve free pose control, controllable head pose features from reference videos are fed into the Video Generator along with the expression feature and source face to generate new talking heads. Extensive quantitative and qualitative experimental results verify that OPT generates high-quality pose-controllable talking heads with no identity mismatch problem, outperforming previous SOTA methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ICASSP | 6 |
| 2023 | Large Pose Friendly Face Reenactment using subtle motionsabstractFace reenactment aims to synthesis a photo-realistic video of the source face by imitating the motion and expression of the driving video while keeping the source appearance (i.e. identity). Although good results have achieved recently, most state-of-the-art methods remain vulnerable to extreme conditions, which greatly restricts the application in the real world. Among various extreme conditions, the large pose problem is the most common one. We clarify that the large pose problem is mainly caused by the severe motion change between the source image and the current driving frame. An intuitive solution is to divide the severe motion change into a sequence of subtle motions. Therefore, we propose a new scheme that exploring the temporal coherence between previous neighbor frame and current frame. The smaller motion change between consecutive frames help to solve the large pose problem. Furthermore, a calibration net is designed to eliminate the error accumulation of the previous step. Extensive experiments demonstrate that our method performs better on large pose face reenactment than the state-of-the-art in terms of large pose cases and visual quality. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han |
ICME | 4 |
| 2023 | Semantic Stage-Wise Learning for Knowledge DistillationabstractKnowledge distillation enhances the performance of the student model by transferring knowledge from the teacher model. Moreover, the attention mechanism has been introduced recently to enable each layer of the student to learn knowledge from all teacher layers, which brings about considerable optimization. However, noted that features from different layers, such as shallow and deep layers, might have a big semantic gap, and compulsively aligning one student layer to all teacher layers would mislead the learning process. To tackle this problem, an effective framework called Semantic Stage-Wise learning for Knowledge Distillation (SSWKD) is presented in this paper. We divide all layers into shallow and deep stages, and only allow feature alignment within the same stage to alleviate semantic mismatch. In addition, with the observation that the performance of deep networks relies more on some key features rather than evenly on all of them, a crucial feature enhancement method based on KL divergence is then proposed for SSWKD, forcing the student to pay more attention to critical features of the teacher. Extensive experiments and visualizations show that our SSWKD outperforms other distillation methods on CIFAR-100 and COCO2017 datasets for image classification, object detection, and instance segmentation tasks. Dongqin Liu, Wei Zhou 0019, Zhaoxing Li, Jiao Dai, Jizhong Han, Ruixuan Li 0001, Songlin Hu 0001 |
ICME | 5 |
| 2023 | FONT: Flow-guided One-shot Talking Head Generation with Natural Head MotionsabstractOne-shot talking head generation has received growing attention in recent years, with various creative and practical applications. An ideal natural and vivid generated talking head video should contain natural head pose changes. However, it is challenging to map head pose sequences from driving audio since there exists a natural gap between audio-visual modalities. In this work, we propose a Flow-guided One-shot model that achieves NaTural head motions(FONT) over generated talking heads. Specifically, we design a probabilistic CVAE-based model to predict head pose sequences from driving audio and source face. Then we develop a keypoint predictor that produces unsupervised keypoints describing the facial structure information from the source face, driving audio and pose sequences. Finally, a flow- guided occlusion-aware generator is employed to produce photo-realistic talking head videos from the estimated keypoints and source face. Extensive experimental results prove that FONT generates talking heads with natural head poses and synchronized mouth shapes, outperforming other compared methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ICME | 6 |
| 2023 | Discovering Sounding Objects by Audio Queries for Audio Visual SegmentationabstractAudio visual segmentation (AVS) aims to segment the sounding objects for each frame of a given video. To distinguish the sounding objects from silent ones, both audio-visual semantic correspondence and temporal interaction are required. The previous method applies multi-frame cross-modal attention to conduct pixel-level interactions between audio features and visual features of multiple frames simultaneously, which is both redundant and implicit. In this paper, we propose an Audio-Queried Transformer architecture, AQFormer, where we define a set of object queries conditioned on audio information and associate each of them to particular sounding objects. Explicit object-level semantic correspondence between audio and visual modalities is established by gathering object information from visual features with predefined audio queries. Besides, an Audio-Bridged Temporal Interaction module is proposed to exchange sounding object-relevant information among multiple frames with the bridge of audio features. Extensive experiments are conducted on two AVS benchmarks to show that our method achieves state-of-the-art performances, especially 7.1% M_J and 7.6% M_F gains on the MS3 setting. Shaofei Huang 0001, Hongji Zhu, Jiao Dai, Jizhong Han, Wenge Rong, Si Liu 0001 |
IJCAI | 5 |
| 2023 | Enriching Phrases with Coupled Pixel and Object Contexts for Panoptic Narrative GroundingabstractPanoptic narrative grounding (PNG) aims to segment things and stuff objects in an image described by noun phrases of a narrative caption. As a multimodal task, an essential aspect of PNG is the visual-linguistic interaction between image and caption. The previous two-stage method aggregates visual contexts from offline-generated mask proposals to phrase features, which tend to be noisy and fragmentary. The recent one-stage method aggregates only pixel contexts from image features to phrase features, which may incur semantic misalignment due to lacking object priors. To realize more comprehensive visual-linguistic interaction, we propose to enrich phrases with coupled pixel and object contexts by designing a Phrase-Pixel-Object Transformer Decoder (PPO-TD), where both fine-grained part details and coarse-grained entity clues are aggregated to phrase features. In addition, we also propose a Phrase-Object Contrastive Loss (POCL) to pull closer the matched phrase-object pairs and push away unmatched ones for aggregating more precise object contexts from more phrase-relevant object tokens. Extensive experiments on the PNG benchmark show our method achieves new state-of-the-art performance with large margins. Tianrui Hui, Junshi Huang, Xiaoming Wei, Xiaolin Wei, Jiao Dai, Jizhong Han, Si Liu 0001 |
IJCAI | 6 |
| 2023 | MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion ModelabstractFace-to-face communication is a common scenario including roles of speakers and listeners. Most existing research methods focus on producing speaker videos, while the generation of listener heads remains largely overlooked. Responsive listening head generation is an important task that aims to model face-to-face communication scenarios by generating a listener head video given a speaker video and a listener head image. An ideal generated responsive listening video should respond to the speaker with attitude or viewpoint expressing while maintaining diversity in interaction patterns and accuracy in listener identity information. To achieve this goal, we propose the Multi-Faceted Responsive Listening Head Generation Network (MFR-Net). Specifically, MFR-Net employs the probabilistic denoising diffusion model to predict diverse head pose and expression features. In order to perform multi-faceted response to the speaker video, while maintaining accurate listener identity preservation, we design the Feature Aggregation Module to boost listener identity features and fuse them with other speaker-related features. Finally, a renderer finetuned with identity consistency loss produces the final listening head videos. Our extensive experiments demonstrate that MFR-Net not only achieves multi-faceted responses in diversity and speaker identity information but also in attitude and viewpoint expression. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ACM Multimedia | 6 |
| 2023 | CoP: Chain-of-Pose for Image Animation in Large Pose ChangesabstractImage animation involves generating a video of a source image imitating the pose of a driving video. Despite recent advancements in the image animation task, most state-of-the-art methods remain vulnerable to large pose changes. In cases of large pose changes, existing methods struggle to model the complex nonlinear motion and yield distorted results, which greatly restricts their application in the real world. To tackle this problem, we present a novel approach called Chain-of-Pose (CoP) that decomposes large pose changes into a sequence of intermediate pose changes. This enables us to handle simplified pose changes and improves the accuracy of pose estimation. Furthermore, to better preserve the appearance of the source object, we introduce the Appearance Refinement Module (ARM) that effectively integrates the appearance texture feature of the source image with the structural pose feature from the pose chain. Our experimental results demonstrate that our method qualitatively and quantitatively outperforms state-of-the-art approaches on four diverse datasets, comprising talking faces, human bodies, and pixel animals. Notably, our approach significantly improves video quality in the case of large object pose changes. Our code is attached to the supplementary material. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Shuhui Wang, Jiao Dai, Jizhong Han |
ACM Multimedia | 5 |
| 2022 | Focus by Prior: Deepfake Detection Based on Prior-AttentionabstractNowadays advanced facial manipulation techniques produce deepfake videos more realistically, which makes deepfake detection more difficult. To capture subtle and intricate artifacts, recent works attempt to enhance low-level textural information by attention-based framework. However, these methods require complex simulated data or extra supervision. Highly dependent on training settings, these methods not only have high training costs but also are prone to overfitting. To address this issue, we propose a novel perspective of deepfake detection via so-called prior-attention. Specifically, we introduce prior textural information, such as edge and noise, to model the attention maps explicitly. Benefiting from these natural “attention maps”, our model significantly enhances discriminative information without additional supervision. Furthermore, we design a Feature Abstraction Block (FAB) to facilitate cross-layer features interaction and insert it into distinct layers of CNN to detect the inconsistencies at multiple spatial levels. Extensive experiments demonstrate that our method achieves performance comparable to state-of-the-art methods. Cai Yu, Jiao Dai, Xi Wang 0014, Weibo Zhang, Jin Liu 0020, Jizhong Han |
ICME | 3 |
| 2022 | MakeItSmile: Detail-Enhanced Smiling Face ReenactmentabstractGiven a target face and a driving face, face reenactment aims to transfer attributes from the driving face to the target face. In the last decade, a great number of methods have been proposed to generate realistic reenacted faces. However, when these methods are applied to generate a smiling face, most of them can only get a mouth with blurry teeth, making the reenacted face unrealistic. This problem is mainly caused by incomplete tooth structure in the target face image under the setting of one-shot reenactment. In order to obtain smiling reenacted faces with detailed tooth structure, our method uses the tooth information from the driving face rather than the target face. Furthermore, to better represent the tooth structure and expressions of the driving face, we extract the texture with a carefully designed geometry-aware encoder. By training the encoder with tooth segmentation task and non-identity classification task, we acquire refined tooth representations and meanwhile derive the non-identity part of the driving face. We also design a specific generator to fuse the tooth texture features into the target face. Moreover, we add a mouth loss function to further ensure the high definition of the smiling reenacted face. We compare our method to existing state-of-the-art approaches. The experiments show that our method gets comparable results on non-smiling face reenactment and has superior performance on smiling face reenactment. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Wantao Liu, Jiao Dai, Jizhong Han |
IJCNN | 5 |
| 2021 | Li-Net: Large-Pose Identity-Preserving Face Reenactment NetworkabstractFace reenactment is a challenging task, as it is difficult to maintain accurate expression, pose and identity simultaneously. Most existing methods directly apply driving facial landmarks to reenact source faces and ignore the intrinsic gap between two identities, resulting in the identity mismatch issue. Besides, they neglect the entanglement of expression and pose features when encoding driving faces, leading to inaccurate expressions and visual artifacts on large-pose reenacted faces. To address these problems, we propose a Large-pose Identity-preserving face reenactment network, LI-Net. Specifically, the Landmark Transformer is adopted to adjust driving landmark images, which aims to narrow the identity gap between driving and source landmark images. Then the Face Rotation Module and the Expression Enhancing Generator decouple the transformed landmark image into pose and expression features, and reenact those attributes separately to generate identity-preserving faces with accurate expressions and poses. Both qualitative and quantitative experimental results demonstrate the superiority of our method. Jin Liu 0020, Zhaoxing Li, Cai Yu, Shuqiao Zou, Jiao Dai, Jizhong Han |
ICME | 7 |
| 2021 | DLFMNet: End-to-End Detection and Localization of Face Manipulation Using Multi-Domain FeaturesabstractRecently, more and more realistic facial manipulation images and videos, known as DeepFakes, have been created and rapidly circulated in social media. Therefore, it is crucial to develop effective and efficient methods to detect the malicious DeepFakes. Previous approaches all adopt a two-step pipeline with multiple separate models, i.e., first face detection and then face forensics, and lacks robustness against compressed data. In this paper, we propose an end-to-end framework for detection and localization of face manipulation, named DLFMNet, which effectively integrates face detection and face forensics into one model, avoiding intermediate processes like image cropping and feature re-extraction. In addition, to capture richer and more robust manipulated clues, we exploit multi-domain features that takes advantages of two different but complementary domains (i.e., RGB and noise). The evaluations on FaceForensics++ dataset demonstrate the effectiveness of our proposed DLFMNet. https://github.com/LightningChan/DLFMNet. Jin Liu 0020, Cai Yu, Shuqiao Zou, Jiao Dai, Jizhong Han |
ICME | 6 |
| 2020 | FSSPOTTER: Spotting Face-Swapped Video by Spatial and Temporal CluesabstractRecent advances in face generation and manipulation have enabled the creation of sophisticated face-swapped videos, also known as DeepFakes, which brings great potential threats to our society. Hence, it is crucial to develop effective approaches to distinguish them. Currently, face-swapped videos produced by existing methods are prone to exhibit some subtle spatial and temporal manipulated traces, which can be utilized as distinctive clues for face-swapped video detection. In this paper, we propose a unified framework, named FSSpotter, to explore rich spatial and temporal information in the video simultaneously. It consists of a Spatial Feature Extractor (SFE), which aims to discover spatial evidences within a single frame, and a Temporal Feature Aggregator (TFA), which is responsible for capturing temporal inconsistencies between frames. Moreover, a novel data processing strategy is adopted to highlight the inconsistencies of forged face with its surrounding regions. The evaluations on Deepfakes of FaceForensics++, DeepfakeTIMIT, UADFV and Celeb-DF datasets demonstrate that the proposed approach achieves better or comparable performance on AUC scores. Jin Liu 0020, Guangzhi Zhou, Hongchao Gao, Jiao Dai, Jizhong Han |
ICME | 6 |
| 2020 | A Multi-head Self-relation Network for Scene Text RecognitionabstractThe text embedded in scene images can be seen everywhere in our lives. However, recognizing text from natural scene images is still a challenge because of its diverse shapes and distorted patterns. Recently, advanced recognition networks generally treat scene text recognition as a sequence prediction task. Although achieving excellent performance, these recognition networks consider the feature map cells as independent individuals and update cells state without utilizing the information of their related cells. And the local receptive field of traditional convolutional neural network (CNN) makes a single cell that cannot cover the whole text region in an image. Due to these issues, the existing recognition networks cannot extract the global context information in a visual scene. To deal with the above problems, we propose a Multi-head Self-relation Network(MSRN) for scene text recognition in this paper. The MSRN consists of several multihead self-relation layers, which are designed for extracting the global context information of a visual scene. Then the information of the related cells can be fused by multi-head self-relation layer. Furthermore, experiments over several public datasets demonstrate that our proposed recognition network achieves superior performance on several benchmark datasets including IC03, IC13, IC15, SVT-Perspective. Junwei Zhou 0005, Hongchao Gao, Jiao Dai, Dongqin Liu, Jizhong Han |
ICPR | 3 |
| 2020 | SDHF: Spotting DeepFakes with Hierarchical FeaturesabstractDeepFake videos are widely distributed on social media platforms, which has seriously affected the authenticity of digital media content, calling for robust DeepFake detection methods. Although numerous detection methods are formulated as frame-based binary classification, less attention has been paid to aggregate the features over individual frames to get a video-based judgement. We observed that for the detection of DeepFake videos, three different level forgery features from frame, clip and video can complement each other. We also found that discrete, large interval sampling strategy is more suitable for DeepFake detection, which can sample more complex video scenes, including multiple subjects, diverse facial expressions and head poses. In this work, we propose a hierarchical framework, using 2D convolutional neural networks for frame-level features extraction followed by a 1D convolutional aggregator to extract clip-level and video-level features, which can comprehensively exploit three different levels of features to make decisions. Evaluation was performed on four datasets, including DFDC, Celeb-DF, FaceForensics++ and UADFV, which provides competitive results compared to other methods. Experimental results of cross-test demonstrate that our hierarchical framework has excellent generalization performance in the face of unknown datasets. Guangzhi Zhou, Hongchao Gao, Jin Liu 0020, Zhaoxing Li, Jiao Dai |
ICTAI | 7 |
| 2020 | A Rating Bias Formulation based on Fuzzy Set for RecommendationabstractIn recommender systems, the user uncertain preference results in unexpected ratings. Previous approaches (e.g., BiasMF) only adjust the rating value based on the bias vector, ignoring the uncertainty of rating. This paper makes an initial attempt in integrating the influence of user uncertain degree and user rating bias into the matrix factorization framework, simultaneously. An approach based on fuzzy set, called fuZzy Matrix Factorization (ZMF), is proposed. Specifically, a fuzzy set of like is defined for each user, and the membership function is utilized to measure the degree of an item belonging to the fuzzy set. Then, the user uncertain preference matrix is obtained, which could explain and represent the user bias and uncertainty effectively. Furthermore, to enhance the computational impact on sparse matrix, the uncertain preference is formulated as a side-information for fusion. Besides, the proposed approach could be extended to others due to independency on additional data sources. Experimental results on three datasets show that ZMF produces an effective improvement. Fuqing Zhu, Jiao Dai, Liangjun Zang, Yipeng Su, Jizhong Han, Songlin Hu 0001 |
IJCNN | 3 |
| 2019 | A Fuzzy Set Based Approach for Rating BiasabstractIn recommender systems, the user uncertain preference results in unexpected ratings. This paper makes an initial attempt in integrating the influence of user uncertain degree into the matrix factorization framework. Specifically, a fuzzy set of like for each user is defined, and the membership function is utilized to measure the degree of an item belonging to the fuzzy set. Furthermore, to enhance the computational effect on sparse matrix, the uncertain preference is formulated as a side-information for fusion. Experimental results on three real-world datasets show that the proposed approach produces stable improvements compared with others. Jiao Dai, Fuqing Zhu, Liangjun Zang, Songlin Hu 0001, Jizhong Han |
AAAI | 2 |
| 2019 | SPL: Exploiting Unlabeled Data for Multi-label Image ClassificationabstractThe utilization of the unlabeled data provides a beneficial attempt for improving the generalization ability of the convolutional neural network (CNN) model, just as what is applied in person re-identification task. Different from that, multi-label image classification aims to predict multiple labels for each given image. The unlabeled data should be properly assigned multiple labels for regularizing the training process of CNN model. To make full use of the unlabeled data, this paper proposes a soft pseudo labeling (SPL) method for multi-label image classification. Specifically, the unlabeled samples are first generated by DCGAN and WGAN-GP. Then, the virtual multiple labels of the generated unlabeled samples are assigned based on an initial confidence value by SoftMax function. Finally, both the generated samples and original training samples are fed into the network as input, in order to learn a CNN model with stronger generalization ability. On three public multi-label image classification datasets (i.e., WIDER-Attribute, NUS-WIDE and MS-COCO), SPL provides a stable improvement over the baseline and produces a competitive performance compared with some existing multi-label image classification methods. Weibo Zhang, Fuqing Zhu, Jiao Dai, Songlin Hu 0001, Jizhong Han, Tao Guo 0006 |
ICME | 3 |
| 2019 | Dim and small target detection based on feature mapping neural networks
Zhisheng Gao, Jiao Dai, Chunzhi Xie |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | A Hybrid Model Based on the Rating Bias and Textual Bias for Recommender Systems
Jiao Dai, Songlin Hu 0001, Jizhong Han |
ICONIP (2) | 1 |
| 2017 | Dynamic Forest Model for Sentiment Classification
Jiao Dai, Jizhong Han |
ICONIP (5) | 2 |
| 2013 | HDKV: supporting efficient high-dimensional similarity search in key-value storesabstractSUMMARY Key‐value stores are widely used on large‐scale data management in the cloud environment. However, they can only naturally support key‐based queries, and do not have efficient solutions for value‐based queries. Thus, dealing with high‐dimensional data in key‐value stores is still a big challenge. State‐of‐the‐art solutions apply value‐based tree‐structure indexes to solve this issue. These methods suffer from the curse of dimensionality and cannot achieve satisfactory performance. They also bring serious load unbalancing problem among servers, and result in dramatic system scalability degradation. Meanwhile, similarity search in high‐dimensional data space becomes more and more popular in today's cloud applications. Due to the lack of efficient algorithms for value‐based queries, users have to wait for a long time before the results are returned. To address this issue, we propose a novel approach called high‐dimensional similarity query in key‐value stores (HDKV), which can generate similarity results in a short time and maintain good database scalability. In HDKV, a strict order‐preserving hash function is designed to map nearby objects in the high‐dimensional space onto adjacent keys of a continuous linear space in key‐value stores. With this strategy, many expensive random accesses are replaced with more efficient scan accesses. The experimental evaluation on real world data set shows that compared to the state‐of‐the‐art methods, HDKV can dramatically reduce the search time with little impact on the accuracy. Copyright © 2012 John Wiley & Sons, Ltd. Wei Zhou 0019, Jizhong Han, Jiao Dai, Zhiyong Xu 0003 |
Concurr. Comput. Pract. Exp. | 4 |
| 2010 | Accelerating Spatial Data Processing with MapReduceabstractMap Reduce is a key-value based programming model and an associated implementation for processing large data sets. It has been adopted in various scenarios and seems promising. However, when spatial computation is expressed straightforward by this key-value based model, difficulties arise due to unfit features and performance degradation. In this paper, we present methods as follows: 1) a splitting method for balancing workload, 2) pending file structure and redundant data partition dealing with relation between spatial objects, 3) a strip-based two-direction plane sweeping algorithm for computation accelerating. Based on these methods, ANN(All nearest neighbors) query and astronomical cross-certification are developed. Performance evaluation shows that the Map Reduce-based spatial applications outperform the traditional one on DBMS. Jizhong Han, Bibo Tu, Jiao Dai, Wei Zhou 0019 |
ICPADS | 4 |