VLDB 2026 Research / reviewers in the wild / expert
Fulong Ye
dblp:333/1204
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0001-9066-9717ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Generative modeling · 31% Vision and language · 23% Language models and text generation · 16% | |
| Computer graphics and multimedia
3 papers |
Visual content generation and editing · 100% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Visual content generation and editing
image generation |
1.7 | 2 | 2025 | DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning · SIGGRAPH Asia 2025 AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models · CVPR 2025 |
Machine learning › Generative modeling › diffusion model › controllable generation
controllable image generation |
0.9 | 1 | 2025 | AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models · CVPR 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models · CVPR 2025 |
Computer vision › Face, body and person analysis › face manipulation
face swapping |
0.9 | 1 | 2025 | DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning · SIGGRAPH Asia 2025 |
Visual content generation and editing › controllable generation
identity-preserving generation |
0.9 | 1 | 2025 | DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning · SIGGRAPH Asia 2025 |
Visual content generation and editing
virtual try-on |
0.9 | 1 | 2025 | AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models · CVPR 2025 |
Visual content generation and editing › image generation
text-to-image generation |
0.8 | 1 | 2024 | AltDiffusion: A Multilingual Text-to-Image Diffusion Model · AAAI 2024 |
Computer vision › Vision and language
multimodal dialogue |
0.7 | 1 | 2023 | SPRING: Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout Graph · AAAI 2023 |
Computer vision › Vision and language › visual grounding
referring expression comprehension |
0.7 | 1 | 2023 | Whether you can locate or not? Interactive Referring Expression Generation · ACM Multimedia 2023 |
Natural language and speech › Language models and text generation › text generation › sentence planning
referring expression generation |
0.7 | 1 | 2023 | Whether you can locate or not? Interactive Referring Expression Generation · ACM Multimedia 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning |
0.2 | 1 | 2023 | SPRING: Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout Graph · AAAI 2023 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 3.3triplet ID group learning · 1.7latent diffusion · 1.7feature extraction · 1.7attention mechanism · 1.7ID adapter · 1.7knowledge distillation · 1.5pre-training · 0.7layout graphs · 0.7curriculum learning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion ModelsabstractRecent advances in garment-centric image generation from text and image prompts based on diffusion models are impressive. However, existing methods lack support for various combinations of attire, and struggle to preserve the garment details while maintaining faithfulness to the text prompts, limiting their performance across diverse scenarios. In this paper, we focus on a new task, i.e., Multi-Garment Virtual Dressing, and we propose a novel Any-Dressing method for customizing characters conditioned on any combination of garments and any personalized text prompts. AnyDressing comprises two primary networks named GarmentsNet and DressingNet, which are respectively dedicated to extracting detailed clothing features and generating customized images. Specifically, we propose an efficient and scalable module called Garment-Specific Feature Extractor in GarmentsNet to individually encode garment textures in parallel. This design prevents garment confusion while ensuring network efficiency. Meanwhile, we design an adaptive Dressing-Attention mechanism and a novel Instance-Level Garment Localization Learning strategy in DressingNet to accurately inject multi-garment features into their corresponding regions. This approach efficiently integrates multi-garment texture cues into generated images and further enhances text-image consistency. Additionally, we introduce a Garment-Enhanced Texture Learning strategy to improve the fine-grained texture details of garments. Thanks to our well-craft design, Any-Dressing can serve as a plug-in module to easily integrate with any community control extensions for diffusion models, improving the diversity and controllability of synthesized images. Extensive experiments show that AnyDressing achieves state-of-the-art results. Xinghui Li, Qichao Sun, Pengze Zhang, Fulong Ye, Zhichao Liao, Wanquan Feng, Songtao Zhao |
CVPR | 4 |
| 2025 | DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group LearningabstractIn this paper, we introduce DreamID, a diffusion-based face swapping model that achieves high levels of ID similarity, attribute preservation, image fidelity, and fast inference speed. Unlike the typical face swapping training process, which often relies on implicit supervision and struggles to achieve satisfactory results. DreamID establishes explicit supervision for face swapping by constructing Triplet ID Group data, significantly enhancing identity similarity and attribute preservation. The iterative nature of diffusion models poses challenges for utilizing efficient image-space loss functions, as performing time-consuming multi-step sampling to obtain the generated image during training is impractical. To address this issue, we leverage the accelerated diffusion model SD Turbo, reducing the inference steps to a single iteration, enabling efficient pixel-level end-to-end training with explicit Triplet ID Group supervision. Additionally, we propose an improved diffusion-based model architecture comprising SwapNet, FaceNet, and ID Adapter. This robust architecture fully unlocks the power of the Triplet ID Group explicit supervision. Finally, to further extend our method, we explicitly modify the Triplet ID Group data during training to fine-tune and preserve specific attributes, such as glasses and face shape. Extensive experiments demonstrate that DreamID outperforms state-of-the-art methods in terms of identity similarity, pose and expression preservation, and image fidelity. Overall, DreamID achieves high-quality face swapping results at 512×512 resolution in just 0.6 seconds and performs exceptionally well in challenging scenarios such as complex lighting, large angles, and occlusions. Our project: https://superhero-7.github.io/DreamID/. Fulong Ye, Miao Hua, Pengze Zhang, Xinghui Li, Qichao Sun, Songtao Zhao |
SIGGRAPH Asia | 1 |
| 2024 | AltDiffusion: A Multilingual Text-to-Image Diffusion ModelabstractLarge Text-to-Image(T2I) diffusion models have shown a remarkable capability to produce photorealistic and diverse images based on text inputs. However, existing works only support limited language input, e.g., English, Chinese, and Japanese, leaving users beyond these languages underserved and blocking the global expansion of T2I models. Therefore, this paper presents AltDiffusion, a novel multilingual T2I diffusion model that supports eighteen different languages. Specifically, we first train a multilingual text encoder based on the knowledge distillation. Then we plug it into a pretrained English-only diffusion model and train the model with a two-stage schema to enhance the multilingual capability, including concept alignment and quality improvement stage on a large-scale multilingual dataset. Furthermore, we introduce a new benchmark, which includes Multilingual-General-18(MG-18) and Multilingual-Cultural-18(MC-18) datasets, to evaluate the capabilities of T2I diffusion models for generating high-quality images and capturing culture-specific concepts in different languages. Experimental results on both MG-18 and MC-18 demonstrate that AltDiffusion outperforms current state-of-the-art T2I models, e.g., Stable Diffusion in multilingual understanding, especially with respect to culture-specific concepts, while still having comparable capability for generating high-quality images. All source code and checkpoints could be found in https://github.com/superhero-7/AltDiffuson. Fulong Ye, Guang Liu 0006, Xinya Wu, Ledell Wu |
AAAI | 1 |
| 2023 | SPRING: Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout GraphabstractExisting multimodal conversation agents have shown impressive abilities to locate absolute positions or retrieve attributes in simple scenarios, but they fail to perform well when complex relative positions and information alignments are involved, which poses a bottleneck in response quality. In this paper, we propose a Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout Graph (SPRING) with abilities of reasoning multi-hops spatial relations and connecting them with visual attributes in crowded situated scenarios. Specifically, we design two types of Multimodal Question Answering (MQA) tasks to pretrain the agent. All QA pairs utilized during pretraining are generated from novel Increment Layout Graphs (ILG). QA pair difficulty labels automatically annotated by ILG are used to promote MQA-based Curriculum Learning. Experimental results verify the SPRING's effectiveness, showing that it significantly outperforms state-of-the-art approaches on both SIMMC 1.0 and SIMMC 2.0 datasets. We release our code and data at https://github.com/LYX0501/SPRING. Yuxing Long, Binyuan Hui, Fulong Ye, Yanyang Li, Zhuoxin Han, Caixia Yuan, Yongbin Li 0001, Xiaojie Wang 0006 |
AAAI | 3 |
| 2023 | Whether you can locate or not? Interactive Referring Expression GenerationabstractReferring Expression Generation (REG) aims to generate unambiguous Referring Expressions (REs) for objects in a visual scene, with a dual task of Referring Expression Comprehension (REC) to locate the referred object. Existing methods construct REG models independently by using only the REs as ground truth for model training, without considering the potential interaction between REG and REC models. In this paper, we propose an Interactive REG (IREG) model that can interact with a real REC model, utilizing signals indicating whether the object is located and the visual region located by the REC model to gradually modify REs. Our experimental results on three RE benchmark datasets, RefCOCO, RefCOCO+, and RefCOCOg show that IREG outperforms previous state-of-the-art methods on popular evaluation metrics. Furthermore, a human evaluation shows that IREG generates better REs with the capability of interaction. Fulong Ye, Yuxing Long, Fangxiang Feng, Xiaojie Wang 0006 |
ACM Multimedia | 1 |