VLDB 2026 Research / reviewers in the wild / expert
Ying Fu 0003
dblp:89/1229-3
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-7358-6167ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EDVD: Cross-Modal Spatio-Temporal Fusion With Event and Diffusion for Video DeblurringabstractRestoring high-quality images from blurred videos is a highly challenging task, especially in severely blurred scenes. In recent years, event-based methods have achieved significant progress in video deblurring. However, the modal differences between the event and image increase the difficulty of feature fusion. Additionally, the sparsity of event makes it difficult to restore some local details. To address these issues, we propose a new video deblurring method. Firstly, we design a cross-modal collaborative attention mechanism to effectively fuse features from blurred frames and event frames, thereby deeply extracting motion information from event frames. Secondly, we utilize a diffusion model to generate spatial guiding prior feature, enhancing local details and textures. Furthermore, we propose an event-guided dynamic feature fusion module that adaptively integrates spatio-temporal information from neighboring frames. Experimental results on both synthetic and real datasets demonstrate that our method outperforms the current state-of-the-art approaches. The code is available at: https://github.com/Frank-Zhou-01/EDVD-main. Ying Fu 0003, Tao Wu 0010, Qing Li 0001, Xi Wu 0004, Wei Liu 0044 |
IEEE Trans. Image Process. | 1 |
| 2025 | EEG2Mesh: High-Quality 3D Mesh Generation from EEG
Ying Fu 0003, Qiaoyu Chen, Chenggang Song, Taorui Li, Dongrui Gao |
PRCV (10) | 1 |
| 2025 | An Explanation Method Based on Interpretable Linear Model With Four Key CharacteristicsabstractFor the interpretability of deep neural networks (DNNs) in visual-related tasks, existing explanation methods commonly generate a saliency map based on the linear relation between output results and input features. However, when the explanation conflicts with a human visual examination, these methods do not provide further evidence to analyze the saliency explanation. Most may fail to provide feature attribution with identifiable semantics or produce misleading explanations due to their insufficient robustness. In this paper, we first propose four key characteristics (richness, adaptivity, exclusiveness, and fairness) to evaluate the existing linear relation-based explanation method, and then construct an interpretable linear model to satisfy them. We formalize the characteristics and develop a novel explanation method based on this. We extract and reconstruct key exclusive semantic features from the feature map using the Nonnegative Matrix Factorization (NMF) algorithm, utilize the information entropy model to determine the number of features adaptively and their richness, and then linearly combine each feature with fairly assigned weights using an approximate Shapley algorithm to generate the saliency map. Compared with the state-of-the-art methods, our explanations of different datasets and DNNs are more convincing and robust in terms of Average drop (AD), Average increase (AI), Deletions (Del), and Insertions (Ins). Our supplementary experiments provide sufficient evidence that the four characteristics guarantee the feasibility of feature attribution analysis and enhance the quality of the resulting explanations. Yuecan Yuan, Zhan ao Huang, Ying Fu 0003, Xuemin Zhao, Canghong Shi, Xiaojie Li 0001, Xi Wu 0004 |
IEEE Trans. Image Process. | 4 |
| 2025 | VB-KGN: Variational Bayesian Kernel Generation Networks for Motion Image DeblurringabstractMotion blur estimation is a critical and fundamental task in scene analysis and image restoration. While most state-of-the-art deep learning-based methods for single-image motion image deblurring focus on constructing deep networks or developing training strategies, the characterization of motion blur has received less attention. In this paper, we innovatively propose a non-parametric Variational Bayesian Kernel Generation Network (VB-KGN) for characterizing motion blur in a single image. To solve this model, we employ the variational inference framework to approximate the expected statistical distribution of motion blur images in a data-driven manner. The qualitative and quantitative evaluations of our experimental results demonstrate that our proposed model can generate highly accurate motion blur kernels, significantly improving motion image deblurring performance and substantially reducing the need for extensive training sample preprocessing for deblurring tasks. Ying Fu 0003, Xiaojie Li 0001, Xin Wang 0045, Xi Wu 0004, Shu Hu 0001, Siwei Lyu, Wei Liu 0044 |
IEEE Trans. Multim. | 1 |
| 2024 | Masked Conditional Diffusion Model for Enhancing Deepfake DetectionabstractRecent studies on deepfake detection have achieved promising results when training and testing faces are from the same dataset. However, their results severely degrade when confronted with forged samples that the model has not yet seen during training. In this paper, deepfake data to help detect deepfakes. this paper present we put a new insight into diffusion model-based data augmentation, and propose a Masked Conditional Diffusion Model (MCDM) for enhancing deepfake detection. It generates a variety of forged faces from a masked pristine one, encouraging the deepfake detection model to learn generic and robust representations without overfitting to special artifacts. Extensive experiments demonstrate that forgery images generated with our method are of high quality and helpful to improve the performance of deepfake detection models. Tiewen Chen, Shanmin Yang, Shu Hu 0001, Zhenghan Fang, Ying Fu 0003, Xi Wu 0004, Xin Wang 0045 |
IJCNN | 5 |
| 2024 | LMHaze: Intensity-aware Image Dehazing with a Large-scale Multi-intensity Real Haze DatasetabstractImage dehazing has drawn a significant attention in recent years.Learning-based methods usually require paired hazy and corresponding ground truth (haze-free) images for training.However, it is difficult to collect real-world image pairs, which prevents developments of existing methods.Although several works partially alleviate this issue by using synthetic datasets or small-scale real datasets.The haze intensity distribution bias and scene homogeneity in existing datasets limit the generalization ability of these methods, particularly when encountering images with previously unseen haze intensities.In this work, we present LMHaze, a large-scale, high-quality real-world dataset.LMHaze comprises paired hazy and haze-free images captured in diverse indoor and outdoor environments, spanning multiple scenarios and haze intensities.It contains over 5K high-resolution image pairs, surpassing the size of the biggest existing real-world dehazing dataset by over 25 times.Meanwhile, to better handle images with different haze intensities, we propose a mixture-of-experts model based on Mamba (MoE-Mamba) for dehazing, which dynamically adjusts the model parameters according to the haze intensity.Moreover, with our proposed dataset, we conduct a new large multimodal model (LMM)-based benchmark study to simulate human perception for evaluating dehazed images.Experiments demonstrate that LMHaze dataset improves the dehazing performance in real scenarios and our dehazing method provides better results compared to state-of-the-art methods.The dataset and code are available at our project page. Ruikun Zhang, Hao Yang 0040, Yan Yang 0011, Ying Fu 0003, Liyuan Pan |
MMAsia | 4 |
| 2022 | Contrastive Class-Specific Encoding for Few-Shot Object DetectionabstractIn this paper, we propose a new few-shot object detection (FSOD) framework that introduces a new contrastive branch to extract the class representation of images, which improves the generalization performance of the detection model for novel classes. Additionally, we investigate the effectiveness of both self-supervised and supervised contrastive losses for class-specific encoding in our framework. Experimental results on the benchmark datasets indicate that our proposed method archives the state-of-the-art performance compared with existing FSOD methods. Dizhong Lin, Ying Fu 0003, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
ICME | 2 |