Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Liebucha Wu

dblp:391/3397 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 65% Vision and language · 35%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.622025
Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities · AAAI 2025
HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities · AAAI 2025
Computer vision › Vision and language
vision-language model
0.912025
PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models · ICCV 2025
Visual content generation and editing
image generation
0.912025
PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models · ICCV 2025
Visual content generation and editing
layout generation
0.912025
PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models · ICCV 2025
Visual content generation and editing › image generation › controllable image generation
layout-to-image generation
0.912025
PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models · ICCV 2025
Machine learning › Generative modeling › image generation › conditional image synthesis
layout-to-image generation
0.812024
HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation · NeurIPS 2024
Visual content generation and editing › image generation
controllable image generation
0.812024
HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

next-token prediction · 1.7negative layout guidance · 1.7autoregressive transformer · 1.7multi-branch diffusion · 1.5hierarchical conditioning · 1.5diffusion model · 0.9backbone-branch network · 0.9
YearPublicationVenuePosition
2025 Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities
abstract
Text-to-Image generation (TTI) technologies are advancing rapidly, especially in the English language communities. However, apart from the user input language barrier problem, English-native TTI models inherently carry biases from their English world centric training data, which creates a dilemma for development of other language-native TTI models. One common choice is to fine-tune the English-native TTI model with translated samples. It falls short of fully addressing the model bias problem. Alternatively, training non-English language native models from scratch can effectively resolve the English world bias, but model trained this way would diverge from the English TTI communities, thus not able to utilize the strides continuously gaining in the English TTI communities any more. To build Chinese TTI model meanwhile keep compatibility with the English TTI communities, we propose a novel model structure referred as "Bridge Diffusion Model" (BDM). The proposed BDM employs a backbone-branch network structure to learn the Chinese semantics while keep the latent space compatible with the English-native TTI backbone, in an end-to-end manner. The unique advantages of the proposed BDM are that it's not only adept at generating images that precisely depict Chinese semantics, but also compatible with various English-native TTI plugins, such as different checkpoints, LoRA, ControlNet, Dreambooth, and Textual Inversion, etc. Moreover, BDM can concurrently generate content seamlessly combining both Chinese-native and English-native semantics within a single image, fostering cultural interaction.
Shanyuan Liu, Bo Cheng 0016, Liebucha Wu, Ao Ma 0005, Dawei Leng, Yuhui Yin
AAAI4
2025 PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models
abstract
In this paper, we propose a unified layout planning and image generation model, PlanGen, which can pre-plan spatial layout conditions before generating images. Unlike previous diffusion-based models that treat layout planning and layout-to-image as two separate models, PlanGen jointly models the two tasks into one autoregressive transformer using only next-token prediction. PlanGen integrates layout conditions into the model as context without requiring specialized encoding of local captions and bounding box coordinates, which provides significant advantages over the previous embed-and-pool operations on layout conditions, particularly when dealing with complex layouts. Unified prompting allows PlanGen to perform multitasking training related to layout, including layout planning, layout-to-image generation, image layout understanding, etc. In addition, PlanGen can be seamlessly expanded to layout-guided image manipulation thanks to the well-designed modeling, with teacher-forcing content manipulation policy and negative layout guidance. Extensive experiments verify the effectiveness of our PlanGen in multiple layoutrelated tasks, showing its great potential. Code is available at: https://360cvgroup.github.io/PlanGen.
Runze He, Bo Cheng 0016, Qingxiang Jia, Shanyuan Liu, Ao Ma 0005, Liebucha Wu, Dawei Leng, Yuhui Yin
ICCV8
2024 HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation
abstract
The task of layout-to-image generation involves synthesizing images based on the captions of objects and their spatial positions. Existing methods still struggle in complex layout generation, where common bad cases include object missing, inconsistent lighting, conflicting view angles, etc. To effectively address these issues, we propose a \textbf{Hi}erarchical \textbf{Co}ntrollable (HiCo) diffusion model for layout-to-image generation, featuring object seperable conditioning branch structure. Our key insight is to achieve spatial disentanglement through hierarchical modeling of layouts. We use a multi branch structure to represent hierarchy and aggregate them in fusion module. To evaluate the performance of multi-objective controllable layout generation in natural scenes, we introduce the HiCo-7K benchmark, derived from the GRIT-20M dataset and manually cleaned. https://github.com/360CVGroup/HiCo_T2I.
Bo Cheng 0016, Liebucha Wu, Shanyuan Liu, Ao Ma 0005, Dawei Leng, Yuhui Yin
NeurIPS3