VLDB 2026 Research / reviewers in the wild / expert
Haiying Wu
dblp:13/7725
· DBLP profile ↗
12ranked-venue papers
0as first author
11since 2021 · last 2025
0009-0008-8816-7791ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient black-box adversarial attacks via alternate query and boundary augmentation
Jiatian Pi, Fusen Wen, Fen Xia, Haiying Wu, Qiao Liu 0001 |
Knowl. Based Syst. | 5 |
| 2025 | AEWFNet: Adaptive Enhancement and Wavelet Convolution for Hyperspectral and Multispectral Image FusionabstractHyperspectral and multispectral image fusion (HMIF) represents a highly effective approach to enhancing the spatial resolution of hyperspectral images (HSIs). However, due to the inaccurate modeling of the spectral response function (SRF) and limited capability in capturing high-frequency details, many existing methods still struggle to maintain spectral consistency and effectively represent spatial structures. To address the aforementioned issues, this study proposes an Adaptive Enhancement and Wavelet Convolution-based HMIF framework, termed AEWFNet, which aims to improve the spatial structural representation and spectral reconstruction accuracy of the fused images. First, a Spatial-Aware Response Learning (SARL) mechanism is proposed to structurally optimize and guide the spectral degradation learning module, helping the model better understand the complex interactions between spatial structures and spectral information, thereby improving the fitting accuracy to the real degradation process. Second, an Attention-based Rotation Prediction Augmentation Network (ARPAN) is designed to optimize spectral information through an adaptive enhancement strategy, effectively improving the generalization performance and feature representation of images under various rotation conditions. Finally, wavelet convolution is employed to enhance the spatial feature representation across multiple scales. Experimental results on several datasets show that AEWFNet outperforms the compared methods in both quantitative and qualitative evaluations, achieving better preservation of the spatial-spectral characteristics of the images. The code for AEWFNet will be made publicly available at https://github.com/LRuiRui517/AEWFNet. Jie Li 0089, Haiying Wu, Pan Wang 0004, Chunyu Zhu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Multi-Objective Progressive Clustering for Semi-Supervised Domain Adaptation in Speaker VerificationabstractUtilizing the pseudo-labeling algorithm with large-scale unlabeled data becomes crucial for semi-supervised domain adaptation in speaker verification tasks. In this paper, we propose a novel pseudo-labeling method named Multi-objective Progressive Clustering (MoPC), specifically designed for semi-supervised domain adaptation. Firstly, we utilize limited labeled data from the target domain to derive domain-specific descriptors based on multiple distinct objectives, namely within-graph denoising, intra-class denoising and inter-class denoising. Then, the Infomap algorithm is adopted for embedding clustering, and the descriptors are leveraged to further refine the target domain’s pseudo-labels. Moreover, to further improve the quality of pseudo labels, we introduce the subcenter-purification and progressive-merging strategy for label denoising. Our proposed MoPC method achieves 4.95% EER and ranked the 1stplace on the evaluation set of VoxSRC 2023 track 3. We also conduct additional experiments on the FFSVC dataset and yield promising results. Ze Li 0003, Yuke Lin, Xiaoyi Qin, Haiying Wu, Ming Li 0026 |
ICASSP | 6 |
| 2024 | Voxblink: A Large Scale Speaker Verification Dataset on CameraabstractIn this paper, we introduce a large-scale and high-quality audiovisual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains 1.45M utterances from 38K speakers. Due to the inherent nature of automated data collection, introducing noisy data is inevitable. Therefore, we also utilize a multi-modal purification step to generate a cleaner version of the VoxBlink, named VoxBlink-clean, comprising 18K identities and 1.02M utterances. In contrast to the VoxCeleb, the VoxBlink sources from short videos of ordinary users, and the covered scenarios can better align with real-life situations. To our best knowledge, the VoxBlink dataset is one of the largest publicly available speaker verification datasets. Leveraging the VoxCeleb and VoxBlink-clean datasets together, we employ diverse speaker verification models with multiple architectural backbones to conduct comprehensive evaluations on the VoxCeleb test sets. Experimental results indicate a substantial enhancement in performance—ranging from 12% to 30% relatively—across various backbone architectures upon incorporating the VoxBlink-clean into the training process. The details of the dataset can be found on $\color{Fuchsia} {{\text{Site}}}$. Yuke Lin, Xiaoyi Qin, Ming Cheng 0005, Haiying Wu, Ming Li 0026 |
ICASSP | 6 |
| 2023 | Improving Deep Learning Powered Auction Design
Shuyuan You, Zhiqiang Zhuang, Haiying Wu, Kewen Wang 0001, Zhe Wang 0001 |
ICONIP (8) | 3 |
| 2022 | Multi-Party Empathetic Dialogue Generation: A New Task for Dialog SystemsabstractEmpathetic dialogue assembles emotion understanding, feeling projection, and appropriate response generation.Existing work for empathetic dialogue generation concentrates on the two-party conversation scenario.Multiparty dialogues, however, are pervasive in reality.Furthermore, emotion and sensibility are typically confused; a refined empathy analysis is needed for comprehending fragile and nuanced human feelings.We address these issues by proposing a novel task called Multi-Party Empathetic Dialogue Generation in this study.Additionally, a Static-Dynamic model for Multi-Party Empathetic Dialogue Generation, SDMPED, is introduced as a baseline by exploring the static sensibility and dynamic emotion for the multi-party empathetic dialogue learning, the aspects that help SDMPED achieve the state-of-the-art performance. Lingyu Zhu 0007, Zhengkun Zhang, Jun Wang 0023, Haiying Wu, Zhenglu Yang |
ACL (1) | 5 |
| 2022 | D2S: Dynamic Distribution Supervision for Multi-Label Facial Expression RecognitionabstractDue to facial images usually evoking multiple emotions with different intensities, there exists much ambiguity in facial expression recognition (FER). Previous methods jointly optimize multi-label learning (MLL) and label distribution learning (LDL) to suppress ambiguity, which have achieved excellent performance. However, the different convergence speed of MLL and LDL make the model easily over-fitting. To address this problem, we propose a dynamic distribution supervision (D2S) method, where the label distribution information is introduced as auxiliary supervision for multilabel classification. Specifically, we develop a multi-task framework in which MLL and LDL are optimized simultaneously. The losses are dynamically weighted to overcome the inconsistency of inter-losses optimization between the two tasks. Extensive experiments on the largest benchmark dataset, i.e., RAF-ML, demonstrate the superiority of the proposed method. Haiying Wu, Jufeng Yang |
ICME | 4 |
| 2022 | EASE: Robust Facial Expression Recognition via Emotion Ambiguity-SEnsitive Cooperative NetworksabstractFacial Expression Recognition (FER) plays a crucial role in the real-world applications. However, large-scale FER datasets collected in the wild usually contain noises. More importantly, due to the ambiguity of emotion, facial images with multiple emotions are hard to be distinguished from the ones with noisy labels. Therefore, it is challenging to train a robust model for FER. To address this, we propose Emotion Ambiguity-SEnsitive cooperative networks (EASE) which contain two components. First, the ambiguity-sensitive learning module divides the training samples into three groups. The samples with small-losses in both networks are considered as clean samples, and the ones with large-losses are noisy. Note for the conflict samples that one network disagrees with the other, we distinguish the samples conveying ambiguous emotions from the ones with noises, using the polarity cues of emotions. Here, we utilize KL divergence to optimize the networks, enabling them to pay attention to the non-dominant emotions. The second part of EASE aims to enhance the diversity of the cooperative networks. With the training epochs increasing, the cooperative networks would converge to a consensus. We construct a penalty term according to the correlation between the features, which helps the networks learn diverse representations from the images. Extensive experiments on 6 popular facial expression datasets demonstrate that EASE outperforms the state-of-the-art approaches. Guoli Jia, Haiying Wu, Jufeng Yang |
ACM Multimedia | 4 |
| 2022 | TextBlock: Towards Scene Text Spotting without Fine-grained DetectionabstractScene text spotting systems which integrate text detection and recognition modules have witnessed a lot of success in recent years. Existing works mostly follow the framework of word/character-level fine-grained detection and isolated-instance recognition, which overemphasize the role of detector and ignore the rich context information in recognition. After rethinking the conventional framework, and inspired by the glimpse-focus spotting pipeline of human beings, we ask:1) "can machine spot text without accurate detection just like human beings?", and if yes, 2) "is text block another alternative for scene text spotting other than word or character?". Based on these questions, we propose a new perspective of coarse-grained detection with multi-instance recognition for text spotting. Specifically, a pioneering network termed TextBlock is developed, and a heuristic text block generation method as well as a multi-instance block-level recognition module are proposed. In this way, the burden of detection is relieved, and the contextual semantic information is well explored for recognition. To train the block-level recognizer, a synthetic dataset including about 800K images is formed. As a by-product of attention, fine-grained detection can be recovered with the recognizer. Equipped with a detector without many bells and whistles (e.g., Faster R-CNN), TextBlock achieves competitive or even better performance compared with previous sophisticated text spotters on several public benchmarks. As a primary attempt, we expect this framework will have a potential impact on scene text spotting research in the future. Yuan Zhang 0013, Yu Zhou 0015, Gangyan Zeng, Youhui Guo, Haiying Wu, Weiping Wang 0005 |
ACM Multimedia | 7 |
| 2021 | TEMP: Taxonomy Expansion with Dynamic Margin Loss through Taxonomy-PathsabstractAs an essential form of knowledge representation, taxonomies are widely used in various downstream natural language processing tasks.However, with the continuously rising of new concepts, many existing taxonomies are unable to maintain coverage by manual expansion.In this paper, we propose TEMP, a self-supervised taxonomy expansion method, which predicts the position of new concepts by ranking the generated taxonomy-paths.For the first time, TEMP employs pre-trained contextual encoders in taxonomy construction and hypernym detection problems.Experiments prove that pre-trained contextual embeddings are able to capture hypernym-hyponym relations.To learn more detailed differences between taxonomy-paths, we train the model with dynamic margin loss by a novel dynamic margin function.Extensive evaluations exhibit that TEMP outperforms prior state-of-the-art taxonomy expansion approaches by 14.3% in accuracy and 15.8% in mean reciprocal rank on three public benchmarks. Hongyuan Xu, Yanlong Wen, Haiying Wu, Xiaojie Yuan |
EMNLP (1) | 5 |
| 2021 | Dense Semantic Contrast for Self-Supervised Visual Representation LearningabstractSelf-supervised representation learning for visual pre-training has achieved remarkable success with sample (instance or pixel) discrimination and semantics discovery of instance, whereas there still exists a non-negligible gap between pre-trained model and downstream dense prediction tasks. Concretely, these downstream tasks require more accurate representation, in other words, the pixels from the same object must belong to a shared semantic category, which is lacking in the previous methods. In this work, we present Dense Semantic Contrast (DSC) for modeling semantic category decision boundaries at a dense level to meet the requirement of these tasks. Furthermore, we propose a dense cross-image semantic contrastive learning framework for multi-granularity representation learning. Specially, we explicitly explore the semantic structure of the dataset by mining relations among pixels from different perspectives. For intra-image relation modeling, we discover pixel neighbors from multiple views. And for inter-image relations, we enforce pixel representation from the same semantic class to be more similar than the representation from different classes in one mini-batch. Experimental results show that our DSC model outperforms state-of-the-art methods when transferring to downstream dense prediction tasks, including object detection, semantic segmentation, and instance segmentation. Code will be made available. Xiaoni Li, Yu Zhou 0015, Yifei Zhang 0005, Aoting Zhang, Wei Wang 0315, Haiying Wu, Weiping Wang 0005 |
ACM Multimedia | 7 |
| 2018 | Weakly supervised topic sentiment joint model with word embeddings
Xianghua Fu, Xudong Sun 0004, Haiying Wu, Laizhong Cui, Joshua Zhexue Huang |
Knowl. Based Syst. | 3 |