VLDB 2026 Research / reviewers in the wild / expert
Xiang Li 0177
dblp:40/1491-177
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0003-3828-9834ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Estimating Cluster Stability in Adaptive Resonance Theory for XR Image Understanding
Junjun Gu, Xiaozheng Qu, Zhaochuan Li, Zhuang Qi, Xiang Li 0177, Haibei Huang, Lei Meng 0001, Xiangxu Meng |
ICXR | 5 |
| 2025 | Towards Initialization-Agnostic Clustering with Iterative Adaptive Resonance TheoryabstractThe clustering performance of Fuzzy Adaptive Resonance Theory (Fuzzy ART) is highly dependent on the preset vigilance parameter, where deviations in its value can lead to significant fluctuations in clustering results, severely limiting its practicality for non-expert users. Existing approaches generally enhance vigilance parameter robustness through adaptive mechanisms such as particle swarm optimization and fuzzy logic rules. However, they often introduce additional hyperparameters or complex frameworks that contradict the original simplicity of the algorithm. To address this, we propose Iterative Refinement Fuzzy Adaptive Resonance Theory (IR-ART), which integrates three key phases into a unified iterative framework: (1) Cluster Stability Detection: A dynamic stability detection module that identifies unstable clusters by analyzing the change of sample size (number of samples in the cluster) in iteration. (2) Unstable Cluster Deletion: An evolutionary pruning module that eliminates low-quality clusters. (3) Vigilance Region Expansion: A vigilance region expansion mechanism that adaptively adjusts similarity thresholds. Independent of the specific execution of clustering, these three phases sequentially focus on analyzing the implicit knowledge within the iterative process, adjusting weights and vigilance parameters, thereby laying a foundation for the next iteration. Experimental evaluation demonstrates that IR-ART improves tolerance to suboptimal vigilance parameter values while preserving the parameter simplicity of Fuzzy ART. Case studies visually confirm the algorithm’s self-optimization capability through iterative refinement, making it particularly suitable for non-expert users in resource-constrained scenarios. Xiaozheng Qu, Zhaochuan Li, Zhuang Qi, Xiang Li 0177, Haibei Huang, Lei Meng 0001, Xiangxu Meng |
IJCNN | 4 |
| 2025 | LLM-Enabled Style and Content Regularization for Personalized Text-to-Image GenerationabstractThe personalized text-to-image generation has rapidly advanced with the emergence of Stable Diffusion. Existing methods, which typically fine-tune models using embedded identifiers, often struggle with insufficient stylization and inaccurate image content due to reduced textual controllability. In this paper, we propose style refinement and content preservation strategies. The style refinement strategy leverages the semantic information of visual reasoning prompts and reference images to optimize style embeddings, allowing a more precise and consistent representation of style information. The content preservation strategy addresses the content bias problem by preserving the model’s generalization capabilities, ensuring enhanced textual controllability without compromising stylization. Experimental results verify that our approach achieves superior performance in generating consistent and personalized text-to-image outputs. Anran Yu, Yaochen Zhang, Xiang Li 0177, Lei Meng 0001, Lei Wu 0002, Xiangxu Meng |
IJCNN | 4 |
| 2024 | InstantAS: Minimum Coverage Sampling for Arbitrary-Size Image GenerationabstractIn recent years, diffusion models have dominated the field of image generation with their outstanding generation quality. However, pre-trained large-scale diffusion models are generally trained using fixed-size images, and fail to maintain their performance at different aspect ratios. Existing methods for generating arbitrary-size images based on diffusion models face several issues, including the requirement for extensive finetuning or training, sluggish sampling speed, and noticeable edge artifacts. This paper presents the InstantAS method for arbitrary-size image generation. This method performs non-overlapping minimum coverage segmentation on the target image, minimizing the generation of redundant information and significantly improving sampling speed. To maintain the consistency of the generated image, we also proposed the Inter-Domain Distribution Bridging method to integrate the distribution of the entire image and suppress the separation of diffusion paths in different regions of the image. Furthermore, we propose the dynamic semantic guided cross-attention method, allowing for the control of different regions using different semantics. Experimental results show that InstantAS has better fusion capabilities compared to previous arbitrary-size image generation methods and is far ahead in sampling speed compared to them. Changshuo Wang 0003, Mingzhe Yu, Lei Wu 0002, Lei Meng 0001, Xiang Li 0177, Xiangxu Meng |
ACM Multimedia | 5 |
| 2024 | SMFS-GAN: Style-Guided Multi-class Freehand Sketch-to-Image SynthesisabstractAbstract Freehand sketch‐to‐image (S2I) is a challenging task due to the individualized lines and the random shape of freehand sketches. The multi‐class freehand sketch‐to‐image synthesis task, in turn, presents new challenges for this research area. This task requires not only the consideration of the problems posed by freehand sketches but also the analysis of multi‐class domain differences in the conditions of a single model. However, existing methods often have difficulty learning domain differences between multiple classes, and cannot generate controllable and appropriate textures while maintaining shape stability. In this paper, we propose a style‐guided multi‐class freehand sketch‐to‐image synthesis model, SMFS‐GAN, which can be trained using only unpaired data. To this end, we introduce a contrast‐based style encoder that optimizes the network's perception of domain disparities by explicitly modelling the differences between classes and thus extracting style information across domains. Further, to optimize the fine‐grained texture of the generated results and the shape consistency with freehand sketches, we propose a local texture refinement discriminator and a Shape Constraint Module, respectively. In addition, to address the imbalance of data classes in the QMUL‐Sketch dataset, we add 6K images by drawing manually and obtain QMUL‐Sketch+ dataset. Extensive experiments on SketchyCOCO Object dataset, QMUL‐Sketch+ dataset and Pseudosketches dataset demonstrate the effectiveness as well as the superiority of our proposed method. Zhenwei Cheng, Lei Wu 0002, Xiang Li 0177, Xiangxu Meng |
Comput. Graph. Forum | 3 |
| 2023 | Letter Embedding Guidance Diffusion Model for Scene Text EditingabstractScene text editing(STE) aims to modify the text in the scene image to the target text while retaining the original style. Existing models are based on GAN, where the source image and the target text are input only once during the generation process, and this approach could not fully obtain the style of the source image and content of the target text. In this paper, we propose an STE method based on the classifier-free guidance diffusion model. To our best knowledge, our model is the first work that developed diffusion models to handle the STE task. Specifically, we divide the STE task into multiple steps and extract style information and text content information in each step. In addition, we introduce the letter embedding method as guidance. We experimentally prove that our method outperforms other STE models in terms of overall realism and maintaining glyphs. Changshuo Wang 0003, Lei Wu 0002, Xu Chen 0031, Xiang Li 0177, Lei Meng 0001, Xiangxu Meng |
ICME | 4 |
| 2023 | Compositional Zero-Shot Artistic Font SynthesisabstractRecently, many researchers have made remarkable achievements in the field of artistic font synthesis, with impressive glyph style and effect style in the results. However, due to less exploration in style disentanglement, it is difficult for existing methods to envision a kind of unseen style (glyph-effect) compositions of artistic font, and thus can only learn the seen style compositions. To solve this problem, we propose a novel compositional zero-shot artistic font synthesis gan (CAFS-GAN), which allows the synthesis of unseen style compositions by exploring the visual independence and joint compatibility of encoding semantics between glyph and effect. Specifically, we propose two contrast-based style encoders to achieve style disentanglement due to glyph and effect intertwining in the image. Meanwhile, to preserve more glyph and effect detail, we propose a generator based on hierarchical dual styles AdaIN to reorganize content-styles representations from structure to texture gradually. Extensive experiments demonstrate the superiority of our model in generating high-quality artistic font images with unseen style compositions against other state-of-the-art methods. The source code and data is available at moonlight03.github.io/CAFS-GAN/. Xiang Li 0177, Lei Wu 0002, Changshuo Wang 0003, Lei Meng 0001, Xiangxu Meng |
IJCAI | 1 |
| 2023 | Anything to Glyph: Artistic Font Synthesis via Text-to-Image Diffusion ModelabstractThe automatic generation of artistic fonts is a challenging task that attracts many research interests. Previous methods specifically focus on glyph or texture style transfer. However, we often come across creative fonts composed of objects in posters or logos. These fonts have proven to be a challenge for existing methods as they struggle to generate similar designs. This paper proposes a novel method for generating creative artistic fonts using a pre-trained text-to-image diffusion model. Our model takes a shape image and a prompt describing an object as input and generates an artistic glyph image consisting of such objects. Specifically, we introduce a novel heatmap-based weak position constraint method to guide the positioning of objects in the generated image, and we also propose the Latent Space Semantic Augmentation Module that improves other information while constraining object position. Our approach is unique in that it can preserve the object’s original shape while constraining its position. And our training method requires only a small quantity of generated data, making it an efficient unsupervised learning approach. Experimental results demonstrate that our method can generate various glyphs, including Chinese, English, Japanese, and symbols, using different objects. We also conducted qualitative and quantitative comparisons with various position control methods for the diffusion model. The results indicate that our approach outperforms other methods in terms of visual quality, innovation, and user evaluation. Changshuo Wang 0003, Lei Wu 0002, Xiaole Liu, Xiang Li 0177, Lei Meng 0001, Xiangxu Meng |
SIGGRAPH Asia | 4 |
| 2022 | DSE-Net: Artistic Font Image Synthesis via Disentangled Style EncodingabstractRecently, the artistic font generation has made significant progress. However, existing methods typically treat the style of artistic font as a whole. Their performance is usually limited to the artistic fonts with complex style elements in glyph and text effect. To solve these problems, this paper presents a disentangled style encoding network, termed DSE-Net, to synthesize artistic fonts. In order to obtain the disentangled text effect features, we introduce a perspective transformation network. We propose a cross-layer fusion mechanism to improve the artistic fonts' structure and texture according to their different representations in CNN. Notably, encoding different style elements for artistic font generation is a new task, so there is no publicly-accessible dataset. Therefore, a new dataset, termed SSAF, has been constructed. Extensive experiments demonstrate that our model significantly outperforms the state-of-the-art methods, with more fine-grained text effect and accurate stroke details. Xiang Li 0177, Lei Wu 0002, Xu Chen 0031, Lei Meng 0001, Xiangxu Meng |
ICME | 1 |
| 2022 | Style-woven Attention Network for Zero-shot Ink Wash Painting Style TransferabstractTraditional Chinese painting is a unique form of artistic expression. Compared with western art painting, it pays more attention to the verve in visual effect, especially ink painting, which makes good use of lines and pays little attention to information such as texture. Some style transfer methods have recently begun to apply traditional Chinese painting style (such as ink wash style) to photorealistic. Ink stylization of different types of real-world photos in a dataset using these style transfer methods has some limitations. When the input images are animal types that have not been seen in the training set, the generated results retain some semantic features of the data in the training set, resulting in distortion. Therefore, in this paper, we attempt to separate the feature representations for styles and contents and propose a style-woven attention network to achieve zero-shot ink wash painting style transfer. Our model learns to disentangle the data representations in an unsupervised fashion and capture the semantic correlations of content and style. In addition, an ink style loss is added to improve the learning ability of the style encoder. In order to verify the ability of ink wash stylization, we augmented the publicly available dataset $ChipPhi$. Extensive experiments based on a wide validation set prove that our method achieves state-of-the-art results. Lei Wu 0002, Xiang Li 0177, Xiangxu Meng |
ICMR | 3 |