EDBT 2026 Demo / reviewers in the wild / expert
Fuyun Wang
dblp:326/5419
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2025
0009-0003-7843-5215ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scene Graph-Grounded Image GenerationabstractWith the beneft of explicit object-oriented reasoning capabilities of scene graphs, scene graph-to-image generation has made remarkable advancements in comprehending object coherence and interactive relations. Recent state-of-the-arts typically predict the scene layouts as an intermediate representation of a scene graph before synthesizing the image. Nevertheless, transforming a scene graph into an exact layout may restrict its representation capabilities, leading to discrepancies in interactive relationships (such as standing on, wearing, or covering) between the generated image and the input scene graph. In this paper, we propose a Scene Graph-Grounded Image Generation (SGG-IG) method to mitigate the above issues. Specifcally, to enhance the scene graph representation, we design a masked auto-encoder module and a relation embedding learning module to integrate structural knowledge and contextual information of the scene graph with a mask self-supervised manner. Subsequently, to bridge the scene graph with visual content, we introduce a spatial constraint and image-scene alignment constraint to capture the fne-grained visual correlation between the scene graph symbol representation and the corresponding image representation, thereby generating semantically consistent and high-quality images. Extensive experiments demonstrate the effectiveness of the method both quantitatively and qualitatively. Fuyun Wang, Tong Zhang 0021, Yuanzhi Wang, Xin Liu 0011, Zhen Cui 0001 |
AAAI | 1 |
| 2025 | Distribution Prototype Diffusion Learning for Open-set Supervised Anomaly DetectionabstractIn Open-set Supervised Anomaly Detection (OSAD), the existing methods typically generate pseudo anomalies to compensate for the scarcity of observed anomaly samples, while overlooking critical priors of normal samples, leading to less effective discriminative boundaries. To address this issue, we propose a Distribution Prototype Diffusion Learning (DPDL) method aimed at enclosing normal samples within a compact and discriminative distribution space. Specifically, we construct multiple learnable Gaussian prototypes to create a latent representation space for abundant and diverse normal samples and learn a Schrödinger bridge to facilitate a diffusive transition toward these prototypes for normal samples while steering anomaly samples away. Moreover, to enhance inter-sample separation, we design a dispersion feature learning way in hyper-spherical space, which benefits the identification of out-of-distribution anomalies. Experimental results demonstrate the effectiveness and superiority of our proposed DPDL, achieving state-of-the-art performance on 9 public datasets. Fuyun Wang, Tong Zhang 0021, Yuanzhi Wang, Yide Qiu, Xin Liu 0011, Zhen Cui 0001 |
CVPR | 1 |
| 2025 | M3Rec: Selective State Space Models with Mixture-of-Modality Experts for Multi-Modal Sequential RecommendationabstractThe rapid growth of multimedia-sharing platforms drives the development of recommender systems. While traditional ID-based methods for mining user behavior signals are well-studied, research into multimodal sequential recommendation remains nascent. Current approaches face three critical challenges: (1) inadequate modeling of user preferences across diverse modalities, (2) ineffective capture of user action sequence dependencies hinders representation learning of preferences, and (3) inefficiency in Transformer-based models due to the quadratic complexity of attention mechanisms. To address these issues, we propose M3Rec, a Mamba-based selective state space model incorporating Mixture-of-Modality experts for Multimodal sequential recommendation. M3Rec strengthens the modeling of user action sequence dependencies through shared Mamba blocks across modalities and employs modality experts to extract modality-specific user preferences. The shared Mamba blocks efficiently model long-term user preferences with fast inference and linear scalability through hardware-aware parallelism, enhancing ID-based sequence signals and filtering out non-action-dependent redundant information. This enables more accurate modeling of user preferences across heterogeneous data. Extensive experiments on three public datasets validate the model’s effectiveness. The implementation is released at https://github.com/Xu107/M3Rec-main. Tong Zhang 0021, Fuyun Wang, Zhen Cui 0001 |
ICASSP | 5 |
| 2025 | Unleashing Vecset Diffusion Model for Fast Shape Generationabstract3D shape generation has greatly flourished through the development of so-called "native" 3D diffusion, particularly through the Vecset Diffusion Model (VDM). While recent advancements have shown promising results in generating high-resolution 3D shapes, VDM still struggles with high-speed generation. Challenges exist because of difficulties not only in accelerating diffusion sampling but also VAE decoding in VDM, areas under-explored in previous works. To address these challenges, we present FlashVDM, a systematic framework for accelerating both VAE and DiT in VDM. For DiT, FlashVDM enables flexible diffusion sampling with as few as 5 inference steps and comparable quality, which is made possible by stabilizing consistency distillation with our newly introduced Progressive Flow Distillation. For VAE, we introduce a lightning vecset decoder equipped with Adaptive KV Selection, Hierarchical Volume Decoding, and Efficient Network Design. By exploiting the locality of the vecset and the sparsity of shape surface in the volume, our decoder drastically lowers FLOPs, minimizing the overall decoding overhead. We apply FlashVDM to Hunyuan3D-2 to obtain Hunyuan3D-2 Turbo. Through systematic evaluation, we show that our model significantly outperforms existing fast 3D generation methods, achieving comparable performance to the state-of-the-art while reducing inference time by over 45x for reconstruction and 32x for generation. Code and models are available at https://github.com/Tencent/FlashVDM. Zeqiang Lai, Zibo Zhao 0001, Fuyun Wang, Huiwen Shi, Xianghui Yang, Qingxiang Lin, Jie Jiang 0015, Chunchao Guo, Xiangyu Yue 0001 |
ICCV | 5 |
| 2025 | Value Diffusion Reinforcement LearningabstractModel-free reinforcement learning (RL) combined with diffusion models has achieved significant progress in addressing complex continuous control tasks. However, a persistent challenge in RL remains the accurate estimation of Q-values, which critically governs the efficacy of policy optimization. Although recent advances employ parametric distributions to model value distributions for enhanced estimation accuracy, current methodologies predominantly rely on unimodal Gaussian assumptions or quantile representations. These constraints introduce distributional bias between the learned and true value distributions, particularly in some tasks with a nonstationary policy, ultimately degrading performance. To address these limitations, we propose value diffusion reinforcement learning (VDRL), a novel model-free online RL method that utilizes the generative capacity of diffusion models to represent multimodal value distributions. The core innovation of VDRL lies in the use of the variational loss of diffusion-based value distribution, which is theoretically proven to be a tight lower bound for the optimization objective under the KL-divergence measurement. Furthermore, we introduce double value diffusion learning with sample selection to enhance training stability and further improve value estimation accuracy. Extensive experiments conducted on the MuJoCo benchmark demonstrate that VDRL significantly outperforms some SOTA model-free online RL baselines, showcasing its effectiveness and robustness. Xiaoliang Hu, Fuyun Wang, Tong Zhang 0021, Zhen Cui 0001 |
NeurIPS | 2 |
| 2025 | Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image GenerationabstractAs a new paradigm of visual content generation, autoregressive text-to-image models suffer from slow inference due to their sequential token-by-token decoding process, often requiring thousands of model forward passes to generate a single image. To address this inefficiency, we propose Speculative Jacobi-Denoising Decoding (SJD2), a framework that incorporates the denoising process into Jacobi iterations to enable parallel token generation in autoregressive models. Our method introduces a next-clean-token prediction paradigm that enables the pre-trained autoregressive models to accept noise-perturbed token embeddings and predict the next clean tokens through low-cost fine-tuning. This denoising paradigm guides the model towards more stable Jacobi trajectories. During inference, our method initializes token sequences with Gaussian noise and performs iterative next-clean-token-prediction in the embedding space. We employ a probabilistic criterion to verify and accept multiple tokens in parallel, and refine the unaccepted tokens for the next iteration with the denoising trajectory. Experiments show that our method can accelerate generation by reducing model forward passes while maintaining the visual quality of generated images. Yao Teng, Fuyun Wang, Zhekai Chen, Yu Wang 0002, Zhenguo Li, Weiyang Liu, Difan Zou, Xihui Liu |
NeurIPS | 2 |
| 2025 | MPDS: A Movie Posters Dataset for Image Generation with Diffusion Model
Tong Zhang 0021, Fuyun Wang, Xin Liu 0011, Zhen Cui 0001 |
PRCV (2) | 3 |
| 2025 | MMHCL: Multi-Modal Hypergraph Contrastive Learning for RecommendationabstractThe burgeoning presence of multimodal content-sharing platforms propels the development of personalized recommender systems. Previous works usually suffer from data sparsity and cold-start problems and may fail to adequately explore semantic user–product associations from multimodal data. To address these issues, we propose a novel Multi-Modal Hypergraph Contrastive Learning (MMHCL) framework for user recommendation. For a comprehensive information exploration from user–product relations, we construct two hypergraphs, i.e., a user-to-user (u2u) hypergraph and an item-to-item (i2i) hypergraph, to mine shared preferences among users and intricate multimodal semantic resemblance among items, respectively. This process yields denser second-order semantics that are fused with first-order user–item interaction as complementary to alleviate the data sparsity issue. Then, we design a contrastive feature enhancement paradigm by applying synergistic contrastive learning. By maximizing/minimizing the mutual information between second-order (e.g., shared preference pattern for users) and first-order (information of selected items for users) embeddings of the same/different users and items, the feature distinguishability can be effectively enhanced. Compared with using sparse primary user–item interaction only, our MMHCL obtains denser second-order hypergraphs and excavates more abundant shared attributes to explore the user–product associations, which to a certain extent alleviates the problems of data sparsity and cold-start. Extensive experiments have comprehensively demonstrated the effectiveness of our method. Our code is publicly available at https://github.com/Xu107/MMHCL . Tong Zhang 0021, Fuyun Wang, Xin Liu 0011, Zhen Cui 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image GenerationabstractRecently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote sensing (RS) images that are tremendously different from general images in terms of scale and perspective remains a formidable challenge due to the lack of a comprehensive remote sensing image generation dataset with various modalities, ground sample distances (GSD), and scenes. In this paper, we propose a Multi-modal, Multi-GSD, Multi-scene Remote Sensing (MMM-RS) dataset and benchmark for text-to-image generation in diverse remote sensing scenarios. Specifically, we first collect nine publicly available RS datasets and conduct standardization for all samples. To bridge RS images to textual semantic information, we utilize a large-scale pretrained vision-language model to automatically output text prompts and perform hand-crafted rectification, resulting in information-rich text-image pairs (including multi-modal images). In particular, we design some methods to obtain the images with different GSD and various environments (e.g., low-light, foggy) in a single sample. With extensive manual screening and refining annotations, we ultimately obtain a MMM-RS dataset that comprises approximately 2.1 million text-image pairs. Extensive experimental results verify that our proposed MMM-RS dataset allows off-the-shelf diffusion models to generate diverse RS images across various modalities, scenes, weather conditions, and GSD. The dataset is available at https://github.com/ljl5261/MMM-RS. Jialin Luo, Yuanzhi Wang, Ziqi Gu, Yide Qiu, Shuaizhen Yao, Fuyun Wang, Chunyan Xu, Zhen Cui 0001 |
NeurIPS | 6 |
| 2023 | Contrastive Multi-Level Graph Neural Networks for Session-Based RecommendationabstractSession-based recommendation (SBR) aims to predict the next item at a certain time point based on anonymous user behavior sequences. Existing methods typically model session representation based on simple item transition information. However, since session-based data consists of limited users' short-term interactions, modeling session representation by capturing fixed item transition information from a single dimension suffers from data sparsity. In this paper, we propose a novel contrastive multi-level graph neural networks (CM-GNN) to better exploit complex and high-order item transition information. Specifically, CM-GNN applies local-level graph convolutional network (L-GCN) and global-level graph convolutional network (G-GCN) on the current session and all the sessions respectively, to effectively capture pairwise relations over all the sessions by aggregation strategy. Meanwhile, CM-GNN applies hyper-level graph convolutional network (H-GCN) to capture high-order information among all the item transitions. CM-GNN further introduces an attention-based fusion module to learn pairwise relation-based session representation by fusing the item representations generated by L-GCN and G-GCN. CM-GNN averages the item representations obtained by H-GCN to obtain high-order relation-based session representation. Moreover, to convert the high-order item transition information into the pairwise relation-based session representation, CM-GNN maximizes the mutual information between the representations derived from the fusion module and the average pool layer by contrastive learning paradigm. We conduct extensive experiments on several widely used benchmark datasets to validate the efficacy of the proposed method. The encouraging results demonstrate that our proposed method outperforms the state-of-the-art SBR techniques. Fuyun Wang, Xingyu Gao 0001, Zhenyu Chen 0003, Lei Lyu 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | CGSNet: Contrastive Graph Self-Attention Network for Session-based Recommendation
Fuyun Wang, Xuequan Lu, Lei Lyu 0001 |
Knowl. Based Syst. | 1 |
| 2022 | Adaptive multi-level graph convolution with contrastive learning for skeleton-based action recognition
Pei Geng, Fuyun Wang, Lei Lyu 0001 |
Signal Process. | 3 |