VLDB 2026 Research / reviewers in the wild / expert
Lean Fu
dblp:225/5157
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0003-1889-9352ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ResAdapter: Domain Consistent Resolution Adapter for Diffusion ModelsabstractRecent advancement in text-to-image models and corresponding personalized technologies enables individuals to generate high-quality and imaginative images. However, they often suffer from limitations when generating images with resolutions outside of their trained domain. To overcome this limitation, we present the resolution adapter \textbf{(ResAdapter)}, a domain-consistent adapter designed for diffusion models to generate images with unrestricted resolutions and aspect ratios. Unlike other multi-resolution generation methods that process images of static resolution with complex post-process operations, ResAdapter directly generates images with the dynamical resolution. Especially, after learning a deep understanding of pure resolution priors, ResAdapter trained on the general dataset, generates resolution-free images with personalized diffusion models while preserving their original style domain. Comprehensive experiments demonstrate that ResAdapter with only 0.5M can process images with flexible resolutions for arbitrary diffusion models. More extended experiments demonstrate that ResAdapter is compatible with other modules for image generation across a broad range of resolutions, and can be integrated into other multi-resolution model for efficiently generating higher-resolution images. Jiaxiang Cheng, Pan Xie, Xin Xia 0005, Jiashi Li, Yuxi Ren, Huixia Li, Xuefeng Xiao 0001, Shilei Wen, Lean Fu |
AAAI | 10 |
| 2025 | PatchScaler: An Efficient Patch-Independent Diffusion Model for Image Super-Resolution
Yong Liu 0031, Hang Dong 0001, Jinshan Pan, Qingji Dong, Kai Chen 0023, Rongxiang Zhang, Lean Fu, Fei Wang 0008 |
ICCV | 7 |
| 2025 | Adaptive Batch-Wise Sample Scheduling for Direct Preference OptimizationabstractDirect Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its performance is highly dependent on the quality of the underlying human preference data. To address this bottleneck, prior work has explored various data selection strategies, but these methods often overlook the impact of the evolving states of the language model during the optimization process.
In this paper, we introduce a novel problem: Sample Scheduling for DPO, which aims to dynamically and adaptively schedule training samples based on the model's evolving batch-wise states throughout preference optimization. To solve this problem, we propose SamS, an efficient and effective algorithm that adaptively selects samples in each training batch based on the LLM's learning feedback to maximize the potential generalization performance.
Notably, without modifying the core DPO algorithm, simply integrating SamS significantly improves performance across tasks, with minimal additional computational overhead.
This work points to a promising new direction for improving LLM alignment through batch-wise sample selection, with potential generalization to RLHF and broader supervised learning paradigms. Zixuan Huang 0012, Yikun Ban, Lean Fu, Zhongxiang Dai, Jianxin Li 0002, Deqing Wang 0001 |
NeurIPS | 3 |
| 2024 | ByteEdit: Boost, Comply and Accelerate Generative Image Editing
Yuxi Ren, Jie Wu 0030, Yanzuo Lu, Huafeng Kuang, Xin Xia 0005, Xionghui Wang, Yixing Zhu, Pan Xie, Shiyin Wang, Xuefeng Xiao 0001, Lean Fu |
ECCV (3) | 14 |
| 2024 | UniFL: Improve Latent Diffusion Model via Unified Feedback LearningabstractLatent diffusion models (LDM) have revolutionized text-to-image generation, leading to the proliferation of various advanced models and diverse downstream applications. However, despite these significant advancements, current diffusion models still suffer from several limitations, including inferior visual quality, inadequate aesthetic appeal, and inefficient inference, without a comprehensive solution in sight. To address these challenges, we present **UniFL**, a unified framework that leverages feedback learning to enhance diffusion models comprehensively. UniFL stands out as a universal, effective, and generalizable solution applicable to various diffusion models, such as SD1.5 and SDXL.
Notably, UniFL consists of three key components: perceptual feedback learning, which enhances visual quality; decoupled feedback learning, which improves aesthetic appeal; and adversarial feedback learning, which accelerates inference.
In-depth experiments and extensive user studies validate the superior performance of our method in enhancing generation quality and inference acceleration. For instance, UniFL surpasses ImageReward by 17\% user preference in terms of generation quality and outperforms LCM and SDXL Turbo by 57\% and 20\% general preference with 4-step inference. Jie Wu 0030, Yuxi Ren, Xin Xia 0005, Huafeng Kuang, Pan Xie, Jiashi Li, Xuefeng Xiao 0001, Shilei Wen, Lean Fu, Guanbin Li |
NeurIPS | 11 |
| 2023 | Boosting Video Super Resolution with Patch-Based Temporal Redundancy Optimization
Hang Dong 0001, Jinshan Pan, Chao Zhu 0007, Boyang Liang, Yu Guo 0006, Ding Liu 0001, Lean Fu, Fei Wang 0008 |
ICANN (7) | 8 |
| 2023 | Unfolding Once is Enough: A Deployment-Friendly Transformer Unit for Super-ResolutionabstractRecent years have witnessed a few attempts of vision transformers for single image super-resolution (SISR). Since the high resolution of intermediate features in SISR models increases memory and computational requirements, efficient SISR transformers are more favored. Based on some popular transformer backbone, many methods have explored reasonable schemes to reduce the computational complexity of the self-attention module while achieving impressive performance. However, these methods only focus on the performance on the training platform (e.g., Pytorch/Tensorflow) without further optimization for the deployment platform (e.g., TensorRT). Therefore, they inevitably contain some redundant operators, posing challenges for subsequent deployment in real-world applications. In this paper, we propose a deployment-friendly transformer unit, namely UFONE (i.e., UnFolding ONce is Enough), to alleviate these problems. In each UFONE, we introduce an Inner-patch Transformer Layer (ITL) to efficiently reconstruct the local structural information from patches and a Spatial-Aware Layer (SAL) to exploit the long-range dependencies between patches. Based on UFONE, we propose a Deployment-friendly Inner-patch Transformer Network (DITN) for the SISR task, which can achieve favorable performance with low latency and memory usage on both training and deployment platforms. Furthermore, to further boost the deployment efficiency of the proposed DITN on TensorRT, we also provide an efficient substitution for layer normalization and propose a fusion optimization strategy for specific operators. Extensive experiments show that our models can achieve competitive results in terms of qualitative and quantitative performance with high deployment efficiency. Yong Liu 0031, Hang Dong 0001, Boyang Liang, Songwei Liu, Qingji Dong, Kai Chen 0023, Fangmin Chen, Lean Fu, Fei Wang 0008 |
ACM Multimedia | 8 |
| 2022 | Deep Recurrent Neural Network with Multi-Scale Bi-directional Propagation for Video DeblurringabstractThe success of the state-of-the-art video deblurring methods stems mainly from implicit or explicit estimation of alignment among the adjacent frames for latent video restoration. However, due to the influence of the blur effect, estimating the alignment information from the blurry adjacent frames is not a trivial task. Inaccurate estimations will interfere the following frame restoration. Instead of estimating alignment information, we propose a simple and effective deep Recurrent Neural Network with Multi-scale Bi-directional Propagation (RNN-MBP) to effectively propagate and gather the information from unaligned neighboring frames for better video deblurring. Specifically, we build a Multi-scale Bi-directional Propagation (MBP) module with two U-Net RNN cells which can directly exploit the inter-frame information from unaligned neighboring hidden states by integrating them in different scales. Moreover, to better evaluate the proposed algorithm and existing state-of-the-art methods on real-world blurry scenes, we also create a Real-World Blurry Video Dataset (RBVD) by a well-designed Digital Video Acquisition System (DVAS) and use it as the training and evaluation dataset. Extensive experimental results demonstrate that the proposed RBVD dataset effectively improve the performance of existing algorithms on real-world blurry videos, and the proposed algorithm performs favorably against the state-of-the-art methods on three typical benchmarks. The code is available at https://github.com/XJTU-CVLAB-LOWLEVEL/RNN-MBP. Chao Zhu 0007, Hang Dong 0001, Jinshan Pan, Boyang Liang, Lean Fu, Fei Wang 0008 |
AAAI | 6 |
| 2022 | GCFSR: a Generative and Controllable Face Super Resolution Method Without Facial and GAN PriorsabstractFace image super resolution (face hallucination) usu-ally relies on facial priors to restore realistic details and preserve identity information. Recent advances can achieve impressive results with the help of GAN prior. They ei-ther design complicated modules to modify the fixed GAN prior or adopt complex training strategies to finetune the generator. In this work, we propose a generative and controllable face SR framework, called GCFSR, which can re-construct images with faithful identity information without any additional priors. Generally, GCFSR has an encoder-generator architecture. Two modules called style modu-lation and feature modulation are designed for the multi-factor SR task. The style modulation aims to generate real-istic face details and the feature modulation dynamically fuses the multi-level encoded features and the generated ones conditioned on the upscaling factor. The simple and elegant architecture can be trained from scratch in an end-to-end manner. For small upscaling factors (≤8), GCFSR can produce surprisingly good results with only adversar-ialloss. After adding L1 and perceptual losses, GCFSR can outperform state-of-the-art methods for large upscalingfac-tors (16, 32, 64). During the test phase, we can modulate the generative strength via feature modulation by changing the conditional upscaling factor continuously to achieve various generative effects. Code is available at https://github.com/hejingwenhejingwen/GCFSR Jingwen He, Wu Shi, Kai Chen 0023, Lean Fu, Chao Dong 0005 |
CVPR | 4 |