EDBT 2026 Demo / reviewers in the wild / expert
Xiaoming Li 0002
dblp:36/3071-2
· DBLP profile ↗
26ranked-venue papers
10as first author
17since 2021 · last 2026
0000-0003-3844-9308ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 8 first-author · 12 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RefSTAR: Blind Face Image Restoration with Reference Selection, Transfer, and ReconstructionabstractIntroducing high-quality references can largely alleviate the uncertainty in blind face image restoration tasks, yet the equivocal utilization of reference priors makes it still a struggle to well preserve the human identity. We attribute the identity inconsistency to two deficiencies of existing reference-based face restoration methods, namely the inability to effectively determine which features need to be transferred, and the failure to preserve the structure and details of the selected features. This work mainly focuses on these two issues, and we present a novel blind face image restoration method that considers reference selection, transfer, and reconstruction (RefSTAR) to introduce proper features from reference images. Specifically, we construct a reference selection (RefSel) module, which can generate accurate masks to select reference features. For training the RefSel module, we construct a RefSel-HQ dataset through a mask generation pipeline, which contains annotated masks for 10,000 ground truth-reference pairs. To guarantee the exact introduction of selected reference features, a feature fusion paradigm is designed for reference feature transferring, and a Mask-Compatible Cycle-Consistency Loss is redesigned based on reference reconstruction to further ensure the presence of selected reference image features in the output image. Experiments on various backbone models demonstrate superior performance, showing better identity preservation ability and reference feature transfer quality. Zhicun Yin, Ming Liu 0018, Zhixin Wang, Renjing Pei, Xiaoming Li 0002, Rynson W. H. Lau, Wangmeng Zuo |
AAAI | 7 |
| 2026 | AITTI: Learning Adaptive Inclusive Token for Text-to-Image Generation
Xinyu Hou, Xiaoming Li 0002, Chen Change Loy |
Int. J. Comput. Vis. | 2 |
| 2025 | Omegance: A Single Parameter for Various Granularities in Diffusion-Based SynthesisabstractIn this work, we introduce a single parameter $ω$, to effectively control granularity in diffusion-based synthesis. This parameter is incorporated during the denoising steps of the diffusion model's reverse process. Our approach does not require model retraining, architectural modifications, or additional computational overhead during inference, yet enables precise control over the level of details in the generated outputs. Moreover, spatial masks or denoising schedules with varying $ω$ values can be applied to achieve region-specific or timestep-specific granularity control. Prior knowledge of image composition from control signals or reference images further facilitates the creation of precise $ω$ masks for granularity control on specific objects. To highlight the parameter's role in controlling subtle detail variations, the technique is named Omegance, combining "omega" and "nuance". Our method demonstrates impressive performance across various image and video synthesis tasks and is adaptable to advanced diffusion models. The code is available at https://github.com/itsmag11/Omegance. Xinyu Hou, Zongsheng Yue, Xiaoming Li 0002, Chen Change Loy |
ICCV | 3 |
| 2025 | Enhanced Generative Structure Prior for Chinese Text Image Super-ResolutionabstractFaithful text image super-resolution (SR) is challenging because each character has a unique structure and usually exhibits diverse font styles and layouts. While existing methods primarily focus on English text, less attention has been paid to more complex scripts like Chinese. In this paper, we introduce a high-quality text image SR framework designed to restore the precise strokes of low-resolution (LR) Chinese characters. Unlike methods that rely on character recognition priors to regularize the SR task, we propose a novel structure prior that offers structure-level guidance to enhance visual quality. Our framework incorporates this structure prior within a StyleGAN model, leveraging its generative capabilities for restoration. To maintain the integrity of character structures while accommodating various font styles and layouts, we implement a codebook-based mechanism that restricts the generative space of StyleGAN. Each code in the codebook represents the structure of a specific character, while the vector $w$w in StyleGAN controls the character's style, including typeface, orientation, and location. Through the collaborative interaction between the codebook and style, we generate a high-resolution structure prior that aligns with LR characters both spatially and structurally. Experiments demonstrate that this structure prior provides robust, character-specific guidance, enabling the accurate restoration of clear strokes in degraded characters, even for real-world LR Chinese text with irregular layouts. Xiaoming Li 0002, Wangmeng Zuo, Chen Change Loy |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | VQ-FONT: Few-Shot Font Generation with Structure-Aware Enhancement and QuantizationabstractFew-shot font generation is challenging, as it needs to capture the fine-grained stroke styles from a limited set of reference glyphs, and then transfer to other characters, which are expected to have similar styles. However, due to the diversity and complexity of Chinese font styles, the synthesized glyphs of existing methods usually exhibit visible artifacts, such as missing details and distorted strokes. In this paper, we propose a VQGAN-based framework (i.e., VQ-Font) to enhance glyph fidelity through token prior refinement and structure-aware enhancement. Specifically, we pre-train a VQGAN to encapsulate font token prior within a code-book. Subsequently, VQ-Font refines the synthesized glyphs with the codebook to eliminate the domain gap between synthesized and real-world strokes. Furthermore, our VQ-Font leverages the inherent design of Chinese characters, where structure components such as radicals and character components are combined in specific arrangements, to recalibrate fine-grained styles based on references. This process improves the matching and fusion of styles at the structure level. Both modules collaborate to enhance the fidelity of the generated fonts. Experiments on a collected font dataset show that our VQ-Font outperforms the competing methods both quantitatively and qualitatively, especially in generating challenging styles. Our code is available at https://github.com/Yaomingshuai/VQ-Font. Mingshuai Yao, Yabo Zhang, Xianhui Lin, Xiaoming Li 0002, Wangmeng Zuo |
AAAI | 4 |
| 2024 | When StyleGAN Meets Stable Diffusion: a $\mathcal{W}_{+}$ Adapter for Personalized Image GenerationabstractText-to-image diffusion models have remarkably excelled in producing diverse, high-quality, and photo-realistic images. This advancement has spurred a growing interest in incorporating specific identities into generated content. Most current methods employ an inversion approach to em-bed a target visual concept into the text embedding space using a single reference image. However, the newly synthe-sized faces either closely resemble the reference image in terms of facial attributes, such as expression, or exhibit a reduced capacity for identity preservation. Text descriptions intended to guide the facial attributes of the synthesized face may fall short, owing to the intricate entanglement of identity information with identity-irrelevant facial attributes derived from the reference image. To address these issues, we present the novel use of the extended StyleGAN embed-ding space$\mathcal{W}+$, to achieve enhanced identity preservation and disentanglement for diffusion models. By aligning this semantically meaningful human face latent space with text-to-image diffusion models, we succeed in maintaining high fidelity in identity preservation, coupled with the capacity for semantic editing. Additionally, we propose new training objectives to balance the influences of both prompt and identity conditions, ensuring that the identity-irrelevant back-ground remains negligibly affected during facial attribute modifications. Extensive experiments reveal that our method adeptly generates personalized text-to-image outputs that are not only compatible with prompt descriptions but also amenable to common StyleGAN editing directions in diverse settings. Our code and model are available at https://github.com/csxmli2016/w-plus-adapter. Xiaoming Li 0002, Xinyu Hou, Chen Change Loy |
CVPR | 1 |
| 2024 | Combining Generative and Geometry Priors for Wide-Angle Portrait Correction
Lan Yao, Chaofeng Chen, Xiaoming Li 0002, Zifei Yan, Wangmeng Zuo |
ECCV (29) | 3 |
| 2024 | GOL-SFSTS based few-shot learning mechanical anomaly detection using multi-channel audio signal
Fengqian Zou, Xiaoming Li 0002, Shengtian Sang, Haifeng Zhang 0004 |
Knowl. Based Syst. | 2 |
| 2023 | Learning Generative Structure Prior for Blind Text Image Super-resolutionabstractBlind text image super-resolution (SR) is challenging as one needs to cope with diverse font styles and unknown degradation. To address the problem, existing methods perform character recognition in parallel to regularize the SR task, either through a loss constraint or intermediate feature condition. Nonetheless, the high-level prior could still fail when encountering severe degradation. The prob-lem is further compounded given characters of complex structures, e.g., Chinese characters that combine multiple pictographic or ideographic symbols into a single charac-ter. In this work, we present a novel prior that focuses more on the character structure. In particular, we learn to encapsulate rich and diverse structures in a StyleGAN and exploit such generative structure priors for restoration. To restrict the generative space of StyleGAN so that it obeys the structure of characters yet remains flexible in handling different font styles, we store the discrete features for each character in a codebook. The code subsequently drives the StyleGAN to generate high-resolution structural details to aid text SR. Compared to priors based on character recognition, the proposed structure prior ex-erts stronger character-specific guidance to restore faithful and precise strokes of a designated character. Extensive experiments on synthetic and real datasets demonstrate the compelling performance of the proposed generative structure prior in facilitating robust text SR. Our code is available at https://github.com/csxmli2016/MARCONet. Xiaoming Li 0002, Wangmeng Zuo, Chen Change Loy |
CVPR | 1 |
| 2023 | NÜWA-LIP: Language-guided Image Inpainting with Defect-free VQGANabstractLanguage-guided image inpainting aims to fill the defective regions of an image under the guidance of text while keeping the non-defective regions unchanged. However, directly encoding the defective images is prone to have an adverse effect on the non-defective regions, giving rise to distorted structures on non-defective parts. To better adapt the text guidance to the inpainting task, this paper proposes NÜWA-LIP, which involves defect-free VQGAN (DF-VQGAN) and a multi-perspective sequence-to-sequence module (MP-S2S). To be specific, DF-VQGAN introduces relative estimation to carefully control the receptive spreading, as well as symmetrical connections to protect structure details unchanged. For harmoniously embedding text guidance into the locally defective regions, MP-S2S is employed by aggregating the complementary perspectives from low-level pixels, high-level tokens as well as the text description. Experiments show that our DF-VQGAN effectively aids the inpainting process while avoiding unexpected changes in non-defective regions. Results on three open-domain benchmarks demonstrate the superior performance of our method against state-of-the-arts. Our code, datasets, and model will be made publicly available11https://github.com/kodenii/NUWA-LIP. Minheng Ni, Xiaoming Li 0002, Wangmeng Zuo |
CVPR | 2 |
| 2023 | MetaF2N: Blind Image Super-Resolution by Learning Efficient Model Adaptation from FacesabstractDue to their highly structured characteristics, faces are easier to recover than natural scenes for blind image super-resolution. Therefore, we can extract the degradation representation of an image from the low-quality and recovered face pairs. Using the degradation representation, realistic low-quality images can then be synthesized to fine-tune the super-resolution model for the real-world low-quality image. However, such a procedure is time-consuming and laborious, and the gaps between recovered faces and the ground-truths further increase the optimization uncertainty. To facilitate efficient model adaptation towards image-specific degradations, we propose a method dubbed MetaF2N, which leverages the contained Faces to fine-tune model parameters for adapting to the whole Natural image in a Meta-learning framework. The degradation extraction and low-quality image synthesis steps are thus circumvented in our MetaF2N, and it requires only one fine-tuning step to get decent performance. Considering the gaps between the recovered faces and ground-truths, we further deploy a MaskNet for adaptively predicting loss weights at different positions to reduce the impact of low-confidence areas. To evaluate our proposed MetaF2N, we have collected a real-world low-quality dataset with one or multiple faces in each image, and our MetaF2N achieves superior performance on both synthetic and real-world datasets. Source code, pre-trained models, and collected datasets are available at https://github.com/yinzhicun/MetaF2N. Zhicun Yin, Ming Liu 0018, Xiaoming Li 0002, Longan Xiao, Wangmeng Zuo |
ICCV | 3 |
| 2023 | Learning Dual Memory Dictionaries for Blind Face RestorationabstractBlind face restoration is a challenging task due to the unknown, unsynthesizable and complex degradation, yet is valuable in many practical applications. To improve the performance of blind face restoration, recent works mainly treat the two aspects, i.e., generic and specific restoration, separately. In particular, generic restoration attempts to restore the results through general facial structure prior, while on the one hand, cannot generalize to real-world degraded observations due to the limited capability of direct CNNs' mappings in learning blind restoration, and on the other hand, fails to exploit the identity-specific details. On the contrary, specific restoration aims to incorporate the identity features from the reference of the same identity, in which the requirement of proper reference severely limits the application scenarios. Generally, it is a challenging and intractable task to improve the photo-realistic performance of blind restoration and adaptively handle the generic and specific restoration scenarios with a single unified model. Instead of implicitly learning the mapping from a low-quality image to its high-quality counterpart, this paper suggests a DMDNet by explicitly memorizing the generic and specific features through dual dictionaries. First, the generic dictionary learns the general facial priors from high-quality images of any identity, while the specific dictionary stores the identity-belonging features for each person individually. Second, to handle the degraded input with or without specific reference, dictionary transform module is suggested to read the relevant details from the dual dictionaries which are subsequently fused into the input features. Finally, multi-scale dictionaries are leveraged to benefit the coarse-to-fine restoration. The whole framework including the generic and specific dictionaries is optimized in an end-to-end manner and can be flexibly plugged into different application scenarios. Moreover, a new high-quality dataset, termed CelebRef-HQ, is constructed to promote the exploration of specific face restoration in the high-resolution space. Experimental results demonstrate that the proposed DMDNet performs favorably against the state of the arts in both quantitative and qualitative evaluation, and generates more photo-realistic results on the real-world low-quality images. The codes, models and the CelebRef-HQ dataset will be publicly available at https://github.com/csxmli2016/DMDNet. Xiaoming Li 0002, Shiguang Zhang, Shangchen Zhou, Lei Zhang 0006, Wangmeng Zuo |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Semantic-shape Adaptive Feature Modulation for Semantic Image SynthesisabstractRecent years have witnessed substantial progress in se-mantic image synthesis, it is still challenging in synthesizing photo-realistic images with rich details. Most previ-ous methods focus on exploiting the given semantic map, which just captures an object-level layout for an image. Obviously, a fine-grained part-level semantic layout will benefit object details generation, and it can be roughly in-ferred from an object's shape. In order to exploit the part-level layouts, we propose a Shape-aware Position Descrip-tor (SPD) to describe each pixel's positional feature, where object shape is explicitly encoded into the SP D feature. Fur-thermore, a Semantic-shape Adaptive Feature Modulation (SAFM) block is proposed to combine the given semantic map and our positional features to produce adaptively mod-ulated features. Extensive experiments demonstrate that the proposed SPD and SAFM significantly improve the gener-ation of objects with rich details. Moreover, our method performs favorably against the SOTA methods in terms of quantitative and qualitative evaluation. The source code and model are available at SAFM. Zhengyao Lv, Xiaoming Li 0002, Zhenxing Niu, Bing Cao 0002, Wangmeng Zuo |
CVPR | 2 |
| 2022 | From Face to Natural Image: Learning Real Degradation for Blind Image Super-Resolution
Xiaoming Li 0002, Chaofeng Chen, Xianhui Lin, Wangmeng Zuo, Lei Zhang 0006 |
ECCV (18) | 1 |
| 2021 | Progressive Semantic-Aware Style Transformation for Blind Face RestorationabstractFace restoration is important in face image processing, and has been widely studied in recent years. However, previous works often fail to generate plausible high quality (HQ) results for real-world low quality (LQ) face images. In this paper, we propose a new progressive semantic-aware style transformation framework, named PSFR-GAN, for face restoration. Specifically, instead of using an encoder-decoder framework as previous methods, we formulate the restoration of LQ face images as a multi-scale progressive restoration procedure through semantic-aware style transformation. Given a pair of LQ face image and its corresponding parsing map, we first generate a multi-scale pyramid of the inputs, and then progressively modulate different scale features from coarse-to-fine in a semantic-aware style transfer way. Compared with previous networks, the proposed PSFR-GAN makes full use of the semantic (parsing maps) and pixel (LQ images) space information from different scales of input pairs. In addition, we further introduce a semantic aware style loss which calculates the feature style loss for each semantic region individually to improve the details of face textures. Finally, we pretrain a face parsing network which can generate decent parsing maps from real-world LQ face images. Experiment results show that our model trained with synthetic data can not only produce more realistic high-resolution results for synthetic LQ inputs but also generalize better to natural LQ face images compared with state-of-the-art methods. Chaofeng Chen, Xiaoming Li 0002, Lingbo Yang, Xianhui Lin, Lei Zhang 0006, Kwan-Yee Kenneth Wong |
CVPR | 2 |
| 2021 | Learning Semantic Person Image Generation by Region-Adaptive NormalizationabstractHuman pose transfer has received great attention due to its wide applications, yet is still a challenging task that is not well solved. Recent works have achieved great success to transfer the person image from the source to the target pose. However, most of them cannot well capture the semantic appearance, resulting in inconsistent and less realistic textures on the reconstructed results. To address this issue, we propose a new two-stage framework to handle the pose and appearance translation. In the first stage, we predict the target semantic parsing maps to eliminate the difficulties of pose transfer and further benefit the latter translation of per-region appearance style. In the second one, with the predicted target semantic maps, we suggest a new person image generation method by incorporating the region-adaptive normalization, in which it takes the per-region styles to guide the target appearance generation. Extensive experiments show that our proposed SPGNet can generate more semantic, consistent, and photorealistic results and perform favorably against the state of the art methods in terms of quantitative and qualitative evaluation. The source code and model are available at https://github.com/cszy98/SPGNet.git. Zhengyao Lv, Xiaoming Li 0002, Xin Li 0106, Fu Li 0003, Dongliang He, Wangmeng Zuo |
CVPR | 2 |
| 2021 | Bearing fault diagnosis based on combined multi-scale weighted entropy morphological filtering and bi-LSTM
Fengqian Zou, Haifeng Zhang 0004, Shengtian Sang, Xiaoming Li 0002, Wanying He |
Appl. Intell. | 4 |
| 2020 | Enhanced Blind Face Restoration With Multi-Exemplar Images and Adaptive Spatial Feature FusionabstractIn many real-world face restoration applications, e.g., smartphone photo albums and old films, multiple high-quality (HQ) images of the same person usually are available for a given degraded low-quality (LQ) observation. However, most existing guided face restoration methods are based on single HQ exemplar image, and are limited in properly exploiting guidance for improving the generalization ability to unknown degradation process. To address these issues, this paper suggests to enhance blind face restoration performance by utilizing multi-exemplar images and adaptive fusion of features from guidance and degraded images. First, given a degraded observation, we select the optimal guidance based on the weighted affine distance on landmark sets, where the landmark weights are learned to make the guidance image optimized to HQ image reconstruction. Second, moving least-square and adaptive instance normalization are leveraged for {spatial} alignment and illumination translation of guidance image in the feature space. Finally, for better feature fusion, multiple adaptive spatial feature fusion (ASFF) layers are introduced to incorporate guidance features in an adaptive and progressive manner, resulting in our ASFFNet. Experiments show that our ASFFNet performs favorably in terms of quantitative and qualitative evaluation, and is effective in generating photo-realistic results on real-world LQ images. The source code and models are available at https://github.com/csxmli2016/ASFFNet. Xiaoming Li 0002, Dongwei Ren, Meng Wang 0001, Wangmeng Zuo |
CVPR | 1 |
| 2020 | Face Super-Resolution Guided by 3D Facial Priors
Xiaobin Hu, Wenqi Ren, John LaMaster, Xiaochun Cao, Xiaoming Li 0002, Zechao Li, Bjoern Menze, Wei Liu 0005 |
ECCV (4) | 5 |
| 2020 | Blind Face Restoration via Deep Multi-scale Component Dictionaries
Xiaoming Li 0002, Chaofeng Chen, Shangchen Zhou, Xianhui Lin, Wangmeng Zuo, Lei Zhang 0006 |
ECCV (9) | 1 |
| 2020 | Learning Symmetry Consistent Deep CNNs for Face CompletionabstractDeep convolutional networks (CNNs) have achieved great success in face completion to generate plausible facial structures. These methods, however, are limited in maintaining global consistency among face components and recovering fine facial details. On the other hand, reflectional symmetry is a prominent property of face images and benefits face analysis and consistency modeling, yet remaining uninvestigated in deep face completion. In this work, we leverage two kinds of symmetry-enforcing modules to form a symmetry-consistent CNN model (i.e., SymmFCNet) for effective face completion. For missing pixels on only one of the half-faces, an illumination-reweighted warping subnet is developed to guide the warping and illumination reweighting of the other half-face. As for missing pixels on both of half-faces, we present a generative reconstruction subnet together with a perceptual symmetry loss to enforce symmetry consistency of recovered structures. The SymmFCNet is constructed by stacking generative reconstruction subnet upon illumination-reweighted warping subnet, and can be learned in an end-to-end manner. Experiments show that SymmFCNet can generate globally consistent results on images with synthetic and real occlusions, and performs favorably against state-of-the-arts. Xiaoming Li 0002, Guosheng Hu, Jieru Zhu, Wangmeng Zuo, Meng Wang 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 1 |
| 2018 | Identity Preserving Face Completion for Large Ocular Region Occlusion
Weikai Chen 0001, Jun Xing, Xiaoming Li 0002, Zachary Bessinger, Fuchang Liu, Wangmeng Zuo, Ruigang Yang |
BMVC | 4 |
| 2018 | Learning Warped Guidance for Blind Face Restoration
Xiaoming Li 0002, Ming Liu 0018, Yuting Ye, Wangmeng Zuo, Liang Lin 0004, Ruigang Yang |
ECCV (13) | 1 |
| 2018 | Shift-Net: Image Inpainting via Deep Feature Rearrangement
Zhaoyi Yan, Xiaoming Li 0002, Mu Li 0005, Wangmeng Zuo, Shiguang Shan |
ECCV (14) | 2 |
| 2011 | Joint just noticeable difference model based on depth perception for stereoscopic imagesabstractJust noticeable difference (JND) model can reflect the least perceptible distortion from images, including 2D images and stereoscopic images. As we know, for the perception of human visual system (HVS), stereoscopic images have quite different characteristics from 2D images, since stereoscopic images contain not only planar information, but also depth information. This paper proposes a joint JND (JJND) model based on depth perception for stereoscopic images. Firstly, disparity estimation is performed in order to decompose the image into the occlusion region and the non-overlapped region. Then, different JND thresholds are applied on different regions, according to the depth information of the region, which can be derived from the disparity of the region. Experimental results verified our model's validity for stereoscopic images. Xiaoming Li 0002, Yue Wang 0032, Debin Zhao, Tingting Jiang 0001, Nan Zhang 0015 |
VCIP | 1 |
| 2006 | Dependency driven partitioning objects generation for hardware/software partitioningabstractHardware/software partitioning is a key issue in the design of embedded systems where performance, chip area and/or power dissipation constraints have to be met. The granularity for automatic partitioning has great impact on the run-time of the partitioning and the quality of the final implementation. In this paper we present a new approach that generates partitioning objects using the dependency of the operations in the system. By clustering the dependent operations into one object, the number of the objects and the coupling between them are lowered. Compared to the fine-grain partitioning, our approach, which is called dependency driven partitioning, prunes a great portion of invalid solution space during partitioning objects formation, so as to expedite the partitioning without loss of the quality. Experiments with simulated annealing optimization show a faster convergence than that of fine-granularity with comparable partitioning quality Shengtian Sang, Xiaoming Li 0002, Yizheng Ye |
ISCAS | 2 |