VLDB 2026 Research / reviewers in the wild / expert
Wangyu Wu
dblp:359/0522
· DBLP profile ↗
15ranked-venue papers
7as first author
15since 2021 · last 2026
0009-0005-8404-9489ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question AnsweringabstractWe present S-Path-RAG, a semantic-aware shortest-path Retrieval-Augmented Generation framework designed to improve multi-hop question answering over large knowledge graphs. S-Path-RAG departs from one-shot, text-heavy retrieval by enumerating bounded-length, semantically weighted candidate paths using a hybrid weighted $k$-shortest, beam, and constrained random-walk strategy, learning a differentiable path scorer together with a contrastive path encoder and lightweight verifier, and injecting a compact soft mixture of selected path latents into a language model via cross-attention. The system runs inside an iterative Neural-Socratic Graph Dialogue loop in which concise diagnostic messages produced by the language model are mapped to targeted graph edits or seed expansions, enabling adaptive retrieval when the model expresses uncertainty. This combination yields a retrieval mechanism that is both token-efficient and topology-aware while preserving interpretable path-level traces for diagnostics and intervention. We validate S-Path-RAG on standard multi-hop KGQA benchmarks and through ablations and diagnostic analyses. The results demonstrate consistent improvements in answer accuracy, evidence coverage, and end-to-end efficiency compared to strong graph- and LLM-based baselines. We further analyze trade-offs between semantic weighting, verifier filtering, and iterative updates, and report practical recommendations for deployment under constrained compute and token budgets. Yemin Wang, Tianxiang Xu 0001, Yongtai Liu, Weizhi Tang, Wangyu Wu, Simon Fong 0001 |
WWW | 6 |
| 2026 | Contrastive prompt clustering for weakly supervised semantic segmentation
Wangyu Wu, Wenqiao Zhang, Xianglin Qiu, Siqi Song, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
Expert Syst. Appl. | 1 |
| 2026 | LLM-enhanced multimodal fusion for cross-domain sequential recommendation
Wangyu Wu, Wenqiao Zhang, Siqi Song, Xianglin Qiu, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
Expert Syst. Appl. | 1 |
| 2025 | Cognitive-Inspired Hierarchical Attention Fusion With Visual and Textual for Cross-Domain Sequential Recommendation
Wangyu Wu, Siqi Song, Xianglin Qiu, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
CogSci | 1 |
| 2025 | PCFNet: Enhancing Time Series Forecasting Through Preserving Constant FrequencyabstractLong-term time series forecasting has been widely applied in finance, traffic, and other domains. The stable periodic patterns serve as the foundation for conducting long-term forecasting. However, real-world time series often consist of multi-periodic components and trend components, which poses a significant challenge to time series prediction. In this paper, we introduce PCFNet, a simple yet effective time series forecasting model, which enhances time series forecasting by preserving the constant frequency components that represent the multi-periodicity of time series during the forecasting process. Specifically, PCFNet adaptively identifies the constant frequency components through a simple gated network. Then, the residual frequency components are predicted via a single layer of complex-valued linear layer. Finally, the residual frequency components are added to the constant frequency components to obtain the final outcome. Extensive experimental results across multiple real-world time series datasets demonstrate that PCFNet achieves state-of-the-art performance as a simple architecture. Wangyu Wu, Shouguo Du, Jiyanglin Li |
ECAI | 4 |
| 2025 | Decoupling While Coupling: Towards More Accurate Stereo Image Sand Removal Beyond CertaintyabstractStereo image sand removal is crucial to improve the perceptual quality for autonomous driving perception. Existing methods often fall short in accurately estimating the uncertainty inherent in degraded images, leading to suboptimal outcomes. To address this, we introduce a novel framework named Decoupling While Coupling(DWC). DWC pioneers the integration of inter-view uncertainty estimation, cross-view uncertainty-aware interaction and block-wise uncertainty representation for superior stereo image sand removal. For cross-view information interaction, we propose an Uncertainty-aware Cross-view Attentive Interaction module(UCAI) to cope with the lack of uncertainty estimation ability in the existing cross-view information interaction mechanism. For the uncertainty perception and information interaction within the inter-view, we propose a Distribution Modeling Coupling Block(DMCB), which transmits the representation of uncertainty between each backbone module. For block-wise uncertainty estimation, we use our proposed Uncertainty-aware Distribution Feature Modulator(UDFM) as the backbone of DWC to modulate the uncertainty inside the neural network itself. Extensive experimental validations on our proposed stereo image sand removal dataset SandST confirm the efficacy of DWC. Our method not only achieves higher PSNR and SSIM, but also exhibits enhanced robustness against various sand degrees and patterns. Bingcai Wei, Hui Liu 0065, Chuang Qian 0001, Yi Jia, Wangyu Wu, Zhishan Li |
ICASSP | 5 |
| 2025 | DAF-MFNet: A Multi-Scale Driving Assistance Detection Algorithm Based on Multi-Modal FusionabstractWith the rapid development of autonomous driving technology, autonomous vehicles have become a critical direction for the future of transportation. Ensuring efficient and safe autonomous driving relies heavily on the accuracy of perception systems. However, traditional vision-based perception systems often face significant challenges in adverse visual conditions and complex road environments. To address these issues, this study proposes a multi-modal object detection method based on pixel-level fusion and a multi-scale feature focus diffusion mechanism. A multi-modal fusion module is designed at the entry point of the backbone network to integrate infrared and visible light images, significantly enhancing the model’s detection capability under poor visual conditions. Additionally, a novel multi-scale feature diffusion network is developed using a self-designed multi-scale focus module to iteratively fuse and diffuse features across three different scales, improving the model’s adaptive receptive field for features of varying scales and enhancing detection efficiency in complex road scenarios. A multi-modal dataset is also constructed to support further research in this field. Comparative experiments on three public datasets and a self-constructed dataset show that the proposed model, DAF-MFNet (Diffusion and Focus Multimodal Fusion Network), achieves state-of-the-art detection performance. On the self-constructed dataset, our model outperforms the strong baseline YOLOv8, with [email protected] and [email protected] improvements of 6.9% and 3.5%, respectively, demonstrating its strong adaptability to object detection in the autonomous driving domain. Zhuqi Li, Songtao Deng, Wangyu Wu, Jingbo Zhan, Bingcai Wei, Wuji Wei |
IJCNN | 4 |
| 2025 | Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
Xiaojian Lin, Wenxin Zhang 0005, Yuchu Jiang, Wangyu Wu, Kangxu Wang, Zongzheng Zhang, Guijin Wang, Lei Jin 0003, Hao Zhao 0002 |
ACM Multimedia | 4 |
| 2025 | Robust Modality-Incomplete Anomaly Detection: A Modality-Instructive Framework with BenchmarkabstractMultimodal Industrial Anomaly Detection (MIAD)-fusing 3D point clouds and 2D RGB for product defect detection-is critical to quality inspection. However, existing MIAD methods assume all modalities are available and paired, overlooking real-scenario modality-missing and risking overfitting to incomplete data. To address these, we conduct the first comprehensive study on Modality-Incomplete Industrial Anomaly Detection (MIIAD) and establish MIIAD Bench , a benchmark covering diverse missing settings. Meanwhile, we propose RADAR, a robust two-stage Robust modAlity-instructive fusing & Detecting frAmewoRk. RADAR integrates i) a Modality-Incomplete Instruction mechanism-guiding the multimodal Transformer to focus more on available modal info, and ii) a Double-Pseudo Hybrid Module to highlight unique modality combinations and reduce overfitting. Our results show RADAR outperforms prior methods markedly on MIIAD Bench. Bingchen Miao, Wenqiao Zhang, Juncheng Li 0006, Wangyu Wu, Siliang Tang, Zhaocheng Li, Jun Xiao 0001, Yueting Zhuang |
ACM Multimedia | 4 |
| 2025 | Robust Single Image Sand Removal by Leveraging Uncertainty-aware SAM Priors and Prompt Learning with Refined Perceptual LossabstractSand dust weather has adverse effects on image quality, making single-image sand dust removal a classic research topic in the field of image restoration. However, existing learning-based image restoration methods fail to account for uncertainties in both data and model dimensions, thus being unable to produce satisfactory results for sand dust image restoration. To address this challenge, we introduce a novel framework called the Uncertainty-aware SAM-aided Prompt-interaction Network (USPNet). USPNet comprises two key modules: the Uncertainty-aware SAM Priors Module (USPM), which addresses data-wise aleatoric uncertainties, and the Uncertainty-aware Prompt Learning Module (UPLM), which tackles model-wise epistemic uncertainties. By integrating data-wise and model-wise uncertainty learning, USPNet leverages uncertainty modeling through SAM semantic priors and distributionally representative prompts. Recognizing the unexplored uncertainties inherent in the learning process, we propose an Uncertainty-aware Perceptual Loss (UPL) to enhance the visual quality of restored images through perceptual learning. Through comprehensive perceptual studies and analysis of real sand-dust images, we propose a dataset named SanddustClearity. SanddustClearity includes daytime, nighttime synthetic, and real-world sand dust images. Our extensive experiments, conducted on both synthetic and real-world images exhibiting various levels of sand dust degradation, confirm the effectiveness and robustness of our proposed method. Our code will be available at https://github.com/WBC-ML/USPNet. Bingcai Wei, Hui Liu 0065, Chuang Qian 0001, Zijian Li 0007, Wangyu Wu, Zijie Meng |
ACM Multimedia | 5 |
| 2025 | MAC-Lookup: Multi-Axis Conditional Lookup Model for Underwater Image EnhancementabstractEnhancing underwater images is crucial for exploration. These images face Enhancing underwater images is crucial for exploration. These images face visibility and color issues due to light changes, water turbidity, and bubbles. Traditional prior-based methods and pixel-based methods often fail, while deep learning lacks sufficient high-quality datasets. We introduce the Multi-Axis Conditional Lookup (MAC-Lookup) model, which enhances visual quality by improving color accuracy, sharpness, and contrast. It includes Conditional 3D Lookup Table Color Correction (CLTCC) for preliminary color and quality correction and Multi-Axis Adaptive Enhancement (MAAE) for detail refinement. This model prevents over-enhancement and saturation while handling underwater challenges. Extensive experiments show that MAC-Lookup excels in enhancing underwater images by restoring details and colors better than existing methods. The code is https://github.com/onlycatdoraemon/MAC-Lookup. Fanghai Yi, Zehong Zheng, Zexiao Liang, Yihang Dong, Xiyang Fang, Wangyu Wu, Xuhang Chen 0002 |
SMC | 6 |
| 2025 | Adaptive Patch Contrast for Weakly Supervised Semantic Segmentation
Wangyu Wu, Tianhong Dai, Xiaowei Huang 0001, Jimin Xiao, Fei Ma 0002, Renrong Ouyang |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Generative Prompt Controlled Diffusion for weakly supervised semantic segmentationabstractWeakly supervised semantic segmentation (WSSS), aiming to train segmentation models solely using image-level labels, has received significant attention. Existing approaches mainly concentrate on creating high-quality pseudo labels by utilizing existing images and their corresponding image-level labels. However, a major challenge arises when the available dataset is limited, as the quality of pseudo labels degrades significantly. In this paper, we tackle this challenge from a different perspective by introducing a novel approach called Generative Prompt Controlled Diffusion (GPCD) for data augmentation . This approach enhances the current labeled datasets by augmenting them with a variety of images, achieved through controlled diffusion guided by Generative Pre-trained Transformer (GPT) prompts. In this process, the existing images and image-level labels provide the necessary control information , while GPT enriches the prompts to generate diverse backgrounds. Moreover, we make an original contribution by integrating data source information as tokens into the Vision Transformer (ViT) framework, which improves the ability of downstream WSSS models to recognize the origins of augmented images. Our proposed GPCD approach clearly surpasses existing state-of-the-art methods, with its advantages being more pronounced when the available data is scarce, thereby demonstrating the effectiveness of our method. Our source code will be released. Wangyu Wu, Tianhong Dai, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
Neurocomputing | 1 |
| 2024 | Image Augmentation with Controlled Diffusion for Weakly-Supervised Semantic SegmentationabstractWeakly-supervised semantic segmentation (WSSS), which aims to train segmentation models solely using image-level labels, has achieved significant attention. Existing methods primarily focus on generating high-quality pseudo labels using available images and their image-level labels. However, the quality of pseudo labels degrades significantly when the size of available dataset is limited. Thus, in this paper, we tackle this problem from a different view by introducing a novel approach called Image Augmentation with Controlled Diffusion (IACD). This framework effectively augments existing labeled datasets by generating diverse images through controlled diffusion, where the available images and image-level labels are served as the controlling information. Moreover, we also propose a high-quality image selection strategy to mitigate the potential noise introduced by the randomness of diffusion models. In the experiments, our proposed IACD approach clearly surpasses existing state-of-the-art methods. This effect is more obvious when the amount of available data is small, demonstrating the effectiveness of our method. Wangyu Wu, Tianhong Dai, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
ICASSP | 1 |
| 2024 | Top-K Pooling with Patch Contrastive Learning for Weakly-Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) using only image-level labels has gained significant attention due to cost-effectiveness. Recently, Vision Transformer (ViT) based methods without class activation map (CAM) have shown greater capability in generating reliable pseudo labels than previous methods using CAM. However, the current ViT-based methods utilize max pooling to select the patch with the highest prediction score to map the patch-level classification to the image-level one, which may affect the quality of pseudo labels due to the inaccurate classification of the patches. In this paper, we introduce a novel ViT-based WSSS method named top-K pooling with patch contrastive learning (TKP-PCL), which employs a top-K pooling layer to alleviate the limitations of previous max pooling selection. A patch contrastive error (PCE) is also proposed to enhance the patch embeddings to further improve the final results. The experimental results show that our approach is very efficient and outperforms other state-of-the-art WSSS methods on the PASCAL VOC 2012 and MS COCO 2014 dataset. Wangyu Wu, Tianhong Dai, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
SMC | 1 |