VLDB 2026 Research / reviewers in the wild / expert
Haiwei Wu
dblp:35/8387
· DBLP profile ↗
27ranked-venue papers
10as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 7 since 2021Security and privacy · 6 · 4 first-author · 6 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Detecting diffusion-based text tampering in scene images based on multi-scale feature fusion
Qingwen Zhu, Li Dong 0006, Yuanman Li, Haiwei Wu, Yushu Zhang 0001 |
Expert Syst. Appl. | 4 |
| 2026 | FaceReclaim: Deep Traceability of Face-Swapped Images Through Feature Decoupling
Yuanman Li, Yuanchen Niu, Haiwei Wu, Yushu Zhang 0001, Jiantao Zhou 0001, Bin Li 0011 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | Single-Image Reflection Removal via Iterative Prompt Learning of Reflection LevelabstractSingle-image reflection removal (SIRR) aims to restore the latent background layer from a reflection-contaminated image. Despite the promising progress achieved by deep learning-based methods, the roles of negative training samples and descriptive prompts for the reflection severity are underexplored in most existing deep SIRR approaches, limiting their reflection removal performance and generalization capability. In this work, we introduce a novel training framework that synergistically leverages learnable prompts and image data to optimize the restoration network. To this end, we define reflection levels corresponding to varying degrees of reflection interference on the background content and learn reflection-level prompts to supervise the SIRR process. We propose an Iterative Reflection Level Reduction (IRLR) framework composed of a Restoration Network Training Module (RNTM) and a Reflection Level Learning Module (RLLM). Specifically, RNTM predicts the background layer under the guidance of prompts learned by RLLM, while RLLM in turn refines these prompts using outputs from RNTM. The two modules are trained iteratively to progressively reduce the reflection levels of estimated background layers. To initialize the prompts, we construct a dedicated reflection-level dataset for pretraining. For adaptively supervising RNTM, we design a new reflection-level-aware strategy to address the challenge of directly aligning the output background with the minimal reflection level. Comprehensive experimental results demonstrate that the proposed method significantly outperforms state-of-the-art methods on average performance across several released datasets, improving PSNR by 0.82 dB and SSIM by 0.0120, respectively. The source code and dataset are available at https://github.com/NamecantbeNULL/IRLR_SIRR. Binbin Song, Jiantao Zhou 0001, Shuning Xu, Xina Liu, Haiwei Wu, Xiaopeng Fan 0001, Bihan Wen |
IEEE Trans. Image Process. | 5 |
| 2025 | Anti-Diffusion: Preventing Abuse of Modifications of Diffusion-Based ModelsabstractAlthough diffusion-based techniques have shown remarkable success in image generation and editing tasks, their abuse can lead to severe negative social impacts. Recently, some works have been proposed to provide defense against the abuse of diffusion-based methods. However, their protection may be limited in specific scenarios by manually defined prompts or the stable diffusion (SD) version. Furthermore, these methods solely focus on tuning methods, overlooking editing methods that could also pose a significant threat. In this work, we propose Anti-Diffusion, a privacy protection system designed for general diffusion-based methods, applicable to both tuning and editing techniques. To mitigate the limitations of manually defined prompts on defense performance, we introduce the prompt tuning (PT) strategy that enables precise expression of original images. To provide defense against both tuning and editing methods, we propose the semantic disturbance loss (SDL) to disrupt the semantic information of protected images. Given the limited research on the defense against editing methods, we develop a dataset named Defense-Edit to assess the defense performance of various methods. Experiments demonstrate that our Anti-Diffusion achieves superior defense performance across a wide range of diffusion-based techniques in different scenarios. Liangbin Xie, Jiantao Zhou 0001, Xintao Wang 0002, Haiwei Wu, Jinyu Tian 0001 |
AAAI | 5 |
| 2025 | ADCD-Net: Robust Document Image Forgery Localization via Adaptive DCT Feature and Hierarchical Content DisentanglementabstractThe advancement of image editing tools has enabled malicious manipulation of sensitive document images, underscoring the need for robust document image forgery detection.Though forgery detectors for natural images have been extensively studied, they struggle with document images, as the tampered regions can be seamlessly blended into the uniform document background (BG) and structured text. On the other hand, existing document-specific methods lack sufficient robustness against various degradations, which limits their practical deployment. This paper presents ADCD-Net, a robust document forgery localization model that adaptively leverages the RGB/DCT forensic traces and integrates key characteristics of document images. Specifically, to address the DCT traces' sensitivity to block misalignment, we adaptively modulate the DCT feature contribution based on a predicted alignment score, resulting in much improved resilience to various distortions, including resizing and cropping. Also, a hierarchical content disentanglement approach is proposed to boost the localization performance via mitigating the text-BG disparities. Furthermore, noticing the predominantly pristine nature of BG regions, we construct a pristine prototype capturing traces of untampered regions, and eventually enhance both the localization accuracy and robustness. Our proposed ADCD-Net demonstrates superior forgery localization performance, consistently outperforming state-of-the-art methods by 20.79\% averaged over 5 types of distortions. The code is available at https://github.com/KAHIMWONG/ACDC-Net. Kahim Wong, Jicheng Zhou, Haiwei Wu, Yain-Whar Si, Jiantao Zhou 0001 |
ICCV | 3 |
| 2025 | IoT-Based Precision Litchi Tracking and Counting Method Using Gated MetricsabstractAccurate and efficient multi-object tracking and counting methods are designed to address the challenges of counting in complex environments.This study presents a novel tracking and counting method called LitchiCount, integrating the multi-object tracking detection model LitchiDet with a counting module to address issues such as missing counts, repeated counts, and the lack of interpretability commonly found in traditional machine learning approaches. The method is designed with the guidance of the visual interpretable method Grad-CAM++, as well as the experimental validation method based on important features. To improve the detection accuracy of small targets under dense occlusion and overlapping, we proposed LitchiDet, which combines a small target detection layer, a decoupled fully connected attention with C3Ghost module (DFC-C3Ghost) and an efficient layer aggregation network block (ELANB). Our counting module improves target tracking accuracy and robustness in dense occlusion scenes while reducing counting errors from scene changes. We propose a Distance-generalized Intersection over Union association metric using a gating mechanism(DG-GM) and an AreaC counting strategy tailored to field intricate scenes. Finally, to enhance IoT deployment, we migrated LitchiCount to the Jetson AGX Xavier platform and optimized the model with TensorRT, significantly improving computational efficiency and real-time performance, particularly in resource-limited IoT environments, meeting real-time and low-power demands. The results demonstrated that our proposed method outperforms state-of-the-art detection models, as well as DeepSort-based counting methods in detection and counting. Importantly, by applying our method to the scenario of detecting and counting litchi from multiple perspectives in a field setting, we achieved low-repetitive and reliable counting, demonstrating the robust performance of this approach in real-world applications. Jianqiang Lu, Guoqing Bao, Xiaoling Deng, Xiongzhe Han, Yubin Lan, Haiwei Wu |
IEEE Internet Things J. | 6 |
| 2025 | Coherence-Enhanced language representation learning for sequential Recommendations
Jia Xu 0005, Haiwei Wu |
Knowl. Based Syst. | 2 |
| 2025 | Rethinking Image Forgery Detection via Soft Contrastive Learning and Unsupervised ClusteringabstractImage forgery detection aims to detect and locate forged regions in an image. Most existing forgery detection algorithms formulate classification problems to classify pixels into forged or pristine. However, the definition of forged and pristine pixels is only relative within one single image, e.g., a forged region in image A is actually a pristine one in its source image B (splicing forgery). Such a relative definition has been severely overlooked by existing methods, which unnecessarily mix forged (pristine) regions across different images into the same category. To resolve this dilemma, we propose the FOrensic ContrAstive cLustering (FOCAL) method, a novel, simple yet very effective paradigm based on soft contrastive learning and unsupervised clustering for the image forgery detection. Specifically, FOCAL 1) designs a soft contrastive learning (SCL) to supervise the high-level forensic feature extraction in an image-by-image manner, explicitly reflecting the above relative definition; 2) employs an on-the-fly unsupervised clustering algorithm (instead of a trained one) to cluster the learned features into forged/pristine categories, further suppressing the cross-image influence from training data; and 3) allows to further boost the detection performance via simple feature-level concatenation without the need of retraining. Extensive experimental results over six public testing datasets demonstrate that our proposed FOCALsignificantlyoutperforms the state-of-the-art competitors by big margins: +24.8% onCoverage, +18.9% onColumbia, +17.3% onFF++, +15.3% onMISD, +15.0% onCASIAand +10.5% onNISTin terms of IoU (see also Fig. 1). The paradigm of FOCAL could bring fresh insights and serve as a novel benchmark for the image forgery detection task. The code is available athttps://github.com/HighwayWu/FOCAL. Haiwei Wu, Jiantao Zhou 0001, Yuanman Li |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | Toward Robust Learning via Core Feature-Aware Adversarial TrainingabstractDeep neural networks (DNNs) are inherently vulnerable to adversarial examples (AEs), severely deteriorating model performance on various tasks. Adversarial training (AT) is one of the most effective approaches to enhance model robustness by incorporating AEs into the training process. Notwithstanding the efficacy of AT, recent studies have unveiled that adversarial perturbations on AEs predominantly impact core features—essential for accurate predictions—more than spurious features, which are incidentally aligned with training labels but irrelevant to the model’s classification. This unequal impact induces the models trained with AT to excessively rely on spurious features, resulting in a pronouncedfeature shiftthat compromises robustness and generalization against AEs at inference. In this work, we introduce a novelCore Feature-aware Adversarial Training(COFAT) framework to cope with these challenges. COFAT employscore feature extractionto dynamically generatecore partnersby selectively retaining benign sample regions on feature maps with high-weight while masking low-weight ones, thereby ensuring the model focuses on core features. Furthermore,contrastive feature alignmentis proposed to reduce intra-class feature distances and increase inter-class separability by maintaining a center bank of class feature representations, thus mitigating reliance on spurious features. Compared to state-of-the-art AT methods, COFAT demonstrates superior performance against diverse adversarial attacks. Remarkably, COFAT improves the robustness of ResNet-18 against AutoAttack on CIFAR-10, SVHN, CIFAR-100, and Tiny ImageNet by approximately 2.14%, 3.20%, 1.69%, and 1.86%, respectively, embodying significant advancements in AT. Our code is publicized at https://github.com/Feng-peng-Li/CoFAT. Fengpeng Li, Kemou Li, Haiwei Wu, Jinyu Tian 0001, Jiantao Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Progressive Poisoned Data Isolation for Training-Time Backdoor DefenseabstractDeep Neural Networks (DNN) are susceptible to backdoor attacks where malicious attackers manipulate the model's predictions via data poisoning. It is hence imperative to develop a strategy for training a clean model using a potentially poisoned dataset. Previous training-time defense mechanisms typically employ an one-time isolation process, often leading to suboptimal isolation outcomes. In this study, we present a novel and efficacious defense method, termed Progressive Isolation of Poisoned Data (PIPD), that progressively isolates poisoned data to enhance the isolation accuracy and mitigate the risk of benign samples being misclassified as poisoned ones. Once the poisoned portion of the dataset has been identified, we introduce a selective training process to train a clean model. Through the implementation of these techniques, we ensure that the trained model manifests a significantly diminished attack success rate against the poisoned data. Extensive experiments on multiple benchmark datasets and DNN models, assessed against nine state-of-the-art backdoor attacks, demonstrate the superior performance of our PIPD method for backdoor defense. For instance, our PIPD achieves an average True Positive Rate (TPR) of 99.95% and an average False Positive Rate (FPR) of 0.06% for diverse attacks over CIFAR-10 dataset, markedly surpassing the performance of state-of-the-art methods. The code is available at https://github.com/RorschachChen/PIPD.git. Haiwei Wu, Jiantao Zhou 0001 |
AAAI | 2 |
| 2024 | DAT: Improving Adversarial Robustness via Generative Amplitude Mix-up in Frequency DomainabstractTo protect deep neural networks (DNNs) from adversarial attacks, adversarial training (AT) is developed by incorporating adversarial examples (AEs) into model training. Recent studies show that adversarial attacks disproportionately impact the patterns within the phase of the sample's frequency spectrum---typically containing crucial semantic information---more than those in the amplitude, resulting in the model's erroneous categorization of AEs. We find that, by mixing the amplitude of training samples' frequency spectrum with those of distractor images for AT, the model can be guided to focus on phase patterns unaffected by adversarial perturbations. As a result, the model's robustness can be improved. Unfortunately, it is still challenging to select appropriate distractor images, which should mix the amplitude without affecting the phase patterns. To this end, in this paper, we propose an optimized **Adversarial Amplitude Generator (AAG)** to achieve a better tradeoff between improving the model's robustness and retaining phase patterns. Based on this generator, together with an efficient AE production procedure, we design a new **Dual Adversarial Training (DAT)** strategy. Experiments on various datasets show that our proposed DAT leads to significantly improved robustness against diverse adversarial attacks. The source code is available at https://github.com/Feng-peng-Li/DAT. Fengpeng Li, Kemou Li, Haiwei Wu, Jinyu Tian 0001, Jiantao Zhou 0001 |
NeurIPS | 3 |
| 2024 | Transformer-Based Image Inpainting Detection via Label Decoupling and Constrained Adversarial TrainingabstractImage inpainting based on generative adversarial networks (GANs) has achieved great success in producing visually plausible images and plays an important role in many real tasks. However, the techniques of image inpainting might also be maliciously used, e.g., altering or removing interesting objects to report fake news. Despite the promising performance of recently developed inpainting detection algorithms, they are built on convolutional neural networks (CNNs) with limited receptive fields. Consequently, they fail to fully capture the disparity between the inpainted regions and untouched regions and thus are ineffective in obtaining fine-grained detection results. In this work, we develop a new image inpainting detection approach. First, we propose a locally enhanced transformer architecture tailored for image inpainting detection. Unlike previous CNN-based methods, our approach leverages both the short-range and long-range dependencies of pixels, enabling the learning of diverse statistical behaviors of inpainted and untouched regions. Second, to mitigate the distraction caused by near-edge pixels with a mixed nature during training, we propose decoupling the label into a body map and a soft-edge map, and then a cross-modality attention module is designed to propagate their information interactively. It demonstrates that our decoupling strategy outperforms the conventional edge supervision in enhancing detection accuracy. Finally, we devise a constrained adversarial training methodology in consideration of the confrontational generation procedure of deep image inpainting methods. It shows that our constrained adversarial training further enhances the detection performance by adaptively introducing interference noise in the inpainted regions. Extensive experiments validate the superiority of our scheme compared to existing CNN-based methods, showcasing its desirable detection generalizability for both deep inpainting and traditional inpainting algorithms. Yuanman Li, Liangpei Hu, Li Dong 0006, Haiwei Wu, Jinyu Tian 0001, Jiantao Zhou 0001, Xia Li 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Robust Camera Model Identification Over Online Social Network Shared Images via Multi-Scenario LearningabstractCamera model identification (CMI) can be widely used in image forensics such as authenticity determination, copyright protection, forgery detection, etc. Meanwhile, with the vigorous development of the Internet, online social networks (OSNs) have become the dominant channels for image sharing and transmission. However, the inevitable lossy operations on OSNs, such as compression and post-processing, impose great challenges to the existing CMI schemes, as they severely destroy the camera traces left in the images under investigation. In this work, we propose a novel CMI method that is robust against the lossy operations of various OSN platforms. Specifically, it is observed that a camera trace extractor can be easily trained on a single degradation scenario (e.g., one specific OSN platform); while much more difficult on mixed degradation scenarios (e.g., multiple OSN platforms). Inspired by this observation, we design a new multi-scenario learning (MSL) strategy, enabling us to extract robust camera traces across different OSNs. Furthermore, noticing that image smooth regions incur less distortions by OSN and less interference by image signal itself, we suggest a SmooThness-Aware Trace Extractor (STATE) that can adaptively extract camera traces according to the smoothness of the input image. The superiority of our method is verified by comparative experiments with four state-of-the-art methods, especially under various OSN transmission scenarios. Particularly, for the open-set camera model verification task, we greatly surpass the second-place by 15.30% in AUC on theFODBdataset; while for the close-set camera model classification task, we are significantly ahead of the second-place by 34.51% in F1 on theSIHDRdataset. The code of our proposed method is available athttps://github.com/HighwayWu/CameraTraceOSN. Haiwei Wu, Jiantao Zhou 0001, Jinyu Tian 0001, Weiwei Sun 0009 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Generating Robust Adversarial Examples against Online Social Networks (OSNs)abstractOnline Social Networks (OSNs) have blossomed into prevailing transmission channels for images in the modern era. Adversarial examples (AEs) deliberately designed to mislead deep neural networks (DNNs) are found to be fragile against the inevitable lossy operations conducted by OSNs. As a result, the AEs would lose their attack capabilities after being transmitted over OSNs. In this work, we aim to design a new framework for generating robust AEs that can survive the OSN transmission; namely, the AEs before and after the OSN transmission both possess strong attack capabilities. To this end, we first propose a differentiable network termed SImulated OSN (SIO) to simulate the various operations conducted by an OSN. Specifically, the SIO network consists of two modules: (1) a differentiable JPEG layer for approximating the ubiquitous JPEG compression and (2) an encoder-decoder subnetwork for mimicking the remaining operations. Based upon the SIO network, we then formulate an optimization framework to generate robust AEs by enforcing model outputs with and without passing through the SIO to be both misled. Extensive experiments conducted over Facebook, WeChat and QQ demonstrate that our attack methods produce more robust AEs than existing approaches, especially under small distortion constraints; the performance gain in terms of Attack Success Rate (ASR) could be more than 60%. Furthermore, we build a public dataset containing more than 10,000 pairs of AEs processed by Facebook, WeChat or QQ, facilitating future research in the robust AEs generation. The dataset and code are available at https://github.com/csjunjun/RobustOSNAttack.git . Jun Liu 0071, Jiantao Zhou 0001, Haiwei Wu, Weiwei Sun 0009, Jinyu Tian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Robust Image Forgery Detection over Online Social Network Shared ImagesabstractThe increasing abuse of image editing softwares, such as Photoshop and Meitu, causes the authenticity of digital images questionable. Meanwhile, the widespread availability of online social networks (OSNs) makes them the dominant channels for transmitting forged images to report fake news, propagate rumors, etc. Unfortunately, various lossy operations adopted by OSNs, e.g., compression and resizing, impose great challenges for implementing the robust image forgery detection. To fight against the OSN-shared forgeries, in this work, a novel robust training scheme is proposed. We first conduct a thorough analysis of the noise introduced by OSNs, and decouple it into two parts, i.e., predictable noise and unseen noise, which are modelled separately. The former simulates the noise introduced by the disclosed (known) operations of OSNs, while the latter is designed to not only complete the previous one, but also take into account the defects of the detector itself. We then incorporate the modelled noise into a robust training framework, significantly improving the robustness of the image forgery detector. Extensive experimental results are presented to validate the superiority of the proposed scheme compared with several state-of-the-art competitors. Finally, to promote the future development of the image forgery detection, we build a public forgeries dataset based on four existing datasets and three most popular OSNs. The designed detector recently won the top ranking in a certificate forgery detection competition11https://tianchi.aliyun.com/competition/entrance/531812/introduction. The source code and dataset are available at https://github.com/HighwayWu/lmageForensicsOSN. Haiwei Wu, Jiantao Zhou 0001, Jinyu Tian 0001, Jun Liu 0071 |
CVPR | 1 |
| 2022 | Multistage Curvature-Guided Network for Progressive Single Image Reflection RemovalabstractThanks to the powerful learning capability, deep neural networks (DNNs) have acquired broad applications in single image reflection removal. The DNN-based algorithms relax the constraints of specific priors and learn to generate visually pleasant background layers from massive training data. However, most of them employ a single network structure to recover both the semantic information and local details of the background, which may lead to obvious reflection residue or even failure. To mitigate this deficiency, in this work, we propose a Multi-stage Curvature-guided De-Reflection Network (MCDRNet), which combines multiple network architectures in a unified framework to progressively reconstruct the background layer and refine the fine-grained details. Our framework consists of three stages, where the encoder-decoders are exploited in the first two stages to recover the semantic components of background layers with lower scales and a variant ResNet is applied in the last stage to refine the background details with the original input resolution. In the first two stages, to introduce the structural guidance for the reflection removal, we cascade another decoder branch to restore the curvature map of the background. In addition, at the end of the first two stages, instead of directly passing the intermediate estimates to the next stage, we propose a Non-local Attention Module (NAM) to augment and transmit the features from decoders. Extensive experimental results on several public datasets demonstrate that the proposed MCDRNet outperforms the state-of-the-art methods quantitatively and generates visually better reflection removal results. The source code and pre-trained models are available athttps://github.com/NamecantbeNULL/MCDRNet. Binbin Song, Jiantao Zhou 0001, Haiwei Wu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | IID-Net: Image Inpainting Detection Network via Neural Architecture Search and AttentionabstractDeep learning (DL) has demonstrated its powerful capabilities in the field of image inpainting, which could produce visually plausible results. Meanwhile, the malicious use of advanced image inpainting tools (e.g. removing key objects to report fake news, erasing visible copyright watermarks, etc.) has led to increasing threats to the reliability of image data. To fight against the inpainting forgeries (not only DL-based but also traditional ones), in this work, we propose a novel end-to-end Image Inpainting Detection Network (IID-Net), to detect the inpainted regions at pixel accuracy. The proposed IID-Net consists of three sub-blocks: the enhancement block, the extraction block and the decision block. Specifically, the enhancement block aims to enhance the inpainting traces by using hierarchically combined special layers. The extraction block, automatically designed by Neural Architecture Search (NAS) algorithm, is targeted to extract features for the actual inpainting detection tasks. To further optimize the extracted latent features, we integrate global and local attention modules in the decision block, where the global attention reduces the intra-class differences by measuring the similarity of global features, while the local attention strengthens the consistency of local features. Furthermore, we thoroughly study the generalizability of our IID-Net, and find that different training data could result in vastly different generalization capability. By carefully examining 10 popular inpainting methods, we identify that the IID-Net trained on only one specific deep inpainting method exhibits desirable generalizability; namely, the obtained IID-Net can accurately detect and localize inpainting manipulations for various unseen inpainting methods as well. Extensive experimental results are presented to validate the superiority of the proposed IID-Net, compared with the state-of-the-art competitors. Our results would suggest that common artifacts are shared across diverse image inpainting methods. Finally, we build a public inpainting dataset of 10K image pairs for future research in this area. Haiwei Wu, Jiantao Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Robust Image Forgery Detection Against Transmission Over Online Social NetworksabstractThe increasing abuse of image editing software causes the authenticity of digital images questionable. Meanwhile, the widespread availability of online social networks (OSNs) makes them the dominant channels for transmitting forged images to report fake news, propagate rumors, etc. Unfortunately, various lossy operations, e.g., compression and resizing, adopted by OSNs impose great challenges for implementing the robust image forgery detection. To fight against the OSN-shared forgeries, in this work, a novel robust training scheme is proposed. Firstly, we design a baseline detector, which won the top ranking in a recent certificate forgery detection competition. Then we conduct a thorough analysis of the noise introduced by OSNs, and decouple it into two parts, i.e.,predictable noiseandunseen noise, which are modelled separately. The former simulates the noise introduced by the disclosed (known) operations of OSNs, while the latter is designed to not only complete the previous one, but also take into account the defects of the detector itself. We further incorporate the modelled noise into a robust training framework, significantly improving the robustness of the image forgery detector. Extensive experimental results are presented to validate the superiority of the proposed scheme compared with several state-of-the-art competitors, especially in the scenarios of detecting OSN-transmitted forgeries. Finally, to promote the future development of the image forgery detection, we build a public forgeries dataset based on four existing datasets through the uploading and downloading of four most popular OSNs. The data and code of this work are available athttps://github.com/HighwayWu/ImageForensicsOSN. Haiwei Wu, Jiantao Zhou 0001, Jinyu Tian 0001, Jun Liu 0071, Yu Qiao 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | Deep Generative Model for Image Inpainting With Local Binary Pattern Learning and Spatial AttentionabstractDeep learning (DL) has demonstrated its powerful capabilities in the field of image inpainting. The DL-based image inpainting approaches can produce visually plausible results, but often generate various unpleasant artifacts, especially in the boundary and highly textured regions. To tackle this challenge, in this work, we propose a new end-to-end, two-stage (coarse-to-fine) generative model through combining a local binary pattern (LBP) learning network with an actual inpainting network. Specifically, the first LBP learning network using U-Net architecture is designed to accurately predict the structural information of the missing region, which subsequently guides the second image inpainting network for better filling the missing pixels. Furthermore, an improved spatial attention mechanism is integrated into the image inpainting network, by considering the consistency not only between the known region with the generated one, but also within the generated region itself. Extensive experiments on public datasets includingCelebA-HQ,PlacesandParis StreetViewdemonstrate that our model generates better inpainting results than the state-of-the-art competing algorithms, both quantitatively and qualitatively. The source code and trained models are available athttps://github.com/HighwayWu/ImageInpainting. Haiwei Wu, Jiantao Zhou 0001, Yuanman Li |
IEEE Trans. Multim. | 1 |
| 2021 | An End-to-End Speech Accent Recognition Method Based on Hybrid CTC/Attention Transformer ASRabstractThis paper proposes a novel accent recognition system in the framework of a transformer-based end-to-end speech recognition system. To incorporate the pronunciation and linguistic knowledge into the network, we first pre-train an ASR model in a hybrid CTC/attention manner. Then, focusing on accent recognition, we extend the output token list by inserting accent labels to the transcripts and finetune the network parameters with an accented speech dataset. Our work is evaluated on the Interspeech 2020 Accented English Speech Recognition Challenge. Experiments show that our method achieves an accuracy of 72.39% on the test set and 80.98% on the development set, outperforming the baseline system by a very large margin. Our submitted system ranked second in the accent recognition task in the challenge. Haiwei Wu, Yanqing Sun, Yitao Duan |
ICASSP | 2 |
| 2021 | Transformer Based Unsupervised Pre-Training for Acoustic Representation LearningabstractRecently, a variety of acoustic tasks and related applications arised. For many acoustic tasks, the labeled data size may be limited. To handle this problem, we propose an unsupervised pre-training method using Transformer based encoder to learn a general and robust high-level representation for all acoustic tasks. Experiments have been conducted on three kinds of acoustic tasks: speech emotion recognition, sound event detection and speech translation. All the experiments have shown that pre-training using its own training data can significantly improve the performance. With a larger pre-training data combining MuST-C, Librispeech and ESC-US datasets, for speech emotion recognition, the UAR can further improve absolutely 4.3% on IEMOCAP dataset. For sound event detection, the F1 score can further improve absolutely 1.5% on DCASE2018 task5 development set and 2.1% on evaluation set. For speech translation, the BLEU score can further improve relatively 12.2% on En-De dataset and 8.4% on En-Fr dataset. Ruixiong Zhang, Haiwei Wu, Wubo Li, Dongwei Jiang, Xiangang Li |
ICASSP | 2 |
| 2021 | GIID-NET: Generalizable Image Inpainting Detection NetworkabstractDeep learning (DL) has demonstrated its powerful capabilities in the field of image inpainting, which could produce visually plausible results. Meanwhile, the malicious use of advanced image inpainting tools (e.g. removing key objects to report fake news) has led to increasing threats to the reliability of image data. To fight against the inpainting forgeries, in this work, we propose a novel end-to-end Generalizable Image Inpainting Detection Network (GIID-Net), to detect the inpainted regions at pixel accuracy. Extensive experimental results are presented to validate the superiority of the proposed GIID-Net, compared with the state-of-the-art competitors. Our results would suggest that common artifacts are shared across diverse image inpainting methods. Haiwei Wu, Jiantao Zhou 0001 |
ICIP | 1 |
| 2021 | Privacy Leakage of SIFT Features via Deep Generative Model Based Image ReconstructionabstractMany practical applications, e.g., content based image retrieval and object recognition, heavily rely on the local features extracted from the query image. As these local features are usually exposed to untrustworthy parties, the privacy leakage problem of image local features has received increasing attention in recent years. In this work, we thoroughly evaluate the privacy leakage of Scale Invariant Feature Transform (SIFT), which is one of the most widely-used image local features. We first consider the case that the adversary can fully access the SIFT features, i.e., both the SIFT descriptors and the coordinates are available. We propose a novel end-to-end, coarse-to-fine deep generative model for reconstructing the latent image from its SIFT features. The designed deep generative model consists of two networks, where the first one attempts to learn the structural information of the latent image by transforming from SIFT features to Local Binary Pattern (LBP) features, while the second one aims to reconstruct the pixel values guided by the learned LBP. Compared with the state-of-the-art algorithms, the proposed deep generative model produces much improved reconstructed results over three public datasets. Furthermore, we address more challenging cases that only partial SIFT features (either SIFT descriptors or coordinates) are accessible to the adversary. It is shown that, if the adversary can only have access to the SIFT descriptors while not their coordinates, then the modest success of reconstructing the latent image might be achieved for highly-structured images (e.g., faces) and probably would fail in general settings. In addition, the latent image usually can be reconstructed with acceptable quality solely from the SIFT coordinates. Our results would suggest that the privacy leakage problem can be avoided to a certain extent if the SIFT coordinates can be well protected. Haiwei Wu, Jiantao Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | Domain Aware Training for Far-Field Small-Footprint Keyword SpottingabstractIn this paper, we focus on the task of small-footprint keyword spotting under the far-field scenario.Far-field environments are commonly encountered in real-life speech applications, causing severe degradation of performance due to room reverberation and various kinds of noises.Our baseline system is built on the convolutional neural network trained with pooled data of both far-field and close-talking speech.To cope with the distortions, we develop three domain aware training systems, including the domain embedding system, the deep CORAL system, and the multi-task learning system.These methods incorporate domain knowledge into network training and improve the performance of the keyword classifier on far-field conditions.Experimental results show that our proposed methods manage to maintain the performance on the close-talking speech and achieve significant improvement on the far-field test set. Haiwei Wu, Yuanfei Nie, Ming Li 0026 |
INTERSPEECH | 1 |
| 2019 | Optimization of Emergency Load Shedding Based on Cultural Particle Swarm Optimization AlgorithmabstractA new optimization model is established in this paper to build power system emergency load shedding scheme adaptive to multi operation modes. Particle swarm optimization (PSO) is used to solve the model heuristically. To improve performance of PSO, the concept of belief space in cultural algorithm is introduced to form the cultural particle swarm optimization (CPSO). A new solution is generated in the belief space to replace the inferior solution in CPSO, and motion parameters of CPSO are adjusted by crowding distance to make the algorithm escape from local optimum. The proposed model and algorithm are validated with a provincial power grid model. Taoyang Xu, Changgang Li, Yutian Liu 0002, Dawei Su, Chunlei Xu, Haiwei Wu |
CEC | 7 |
| 2019 | The DKU Replay Detection System for the ASVspoof 2019 Challenge: On Data Augmentation, Feature Representation, Classification, and FusionabstractThis paper describes our DKU replay detection system for the ASVspoof 2019 challenge.The goal is to develop spoofing countermeasure for automatic speaker recognition in physical access scenario.We leverage the countermeasure system pipeline from four aspects, including the data augmentation, feature representation, classification, and fusion.First, we introduce an utterance-level deep learning framework for antispoofing.It receives the variable-length feature sequence and outputs the utterance-level scores directly.Based on the framework, we try out various kinds of input feature representations extracted from either the magnitude spectrum or phase spectrum.Besides, we also perform the data augmentation strategy by applying the speed perturbation on the raw waveform.Our best single system employs a residual neural network trained by the speed-perturbed group delay gram.It achieves EER of 1.04% on the development set, as well as EER of 1.08% on the evaluation set.Finally, using the simple average score from several single systems can further improve the performance.EER of 0.24% on the development set and 0.66% on the evaluation set is obtained for our primary system. Weicheng Cai, Haiwei Wu, Danwei Cai, Ming Li 0026 |
INTERSPEECH | 2 |
| 2019 | The DKU-LENOVO Systems for the INTERSPEECH 2019 Computational Paralinguistic Challenge
Haiwei Wu, Weiqing Wang 0004, Ming Li 0026 |
INTERSPEECH | 1 |