Ronghuan Wu

dblp:345/9761 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0001-9741-9876ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Visual content generation and editing · 84% Computer animation and physical simulation · 16%
Artificial intelligence
1 paper
Generative modeling · 87% Language models and text generation · 13%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing
vector graphics generation
1.522025
Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models · CVPR 2025
IconShop: Text-Guided Vector Icon Synthesis with Autoregressive Transformers · ACM Trans. Graph. 2023
Machine learning › Generative modeling
diffusion model
0.912025
Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models · CVPR 2025
Machine learning › Generative modeling › diffusion model
image diffusion model
0.912025
Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models · CVPR 2025
Visual content generation and editing
image vectorization
0.912025
LayerPeeler: Autoregressive Peeling for Layer-wise Image Vectorization · SIGGRAPH Asia 2025
Visual content generation and editing › vector graphics generation
Text-to-SVG generation
0.912025
Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models · CVPR 2025
Visual content generation and editing › video generation
text-to-video generation
0.912025
AniClipart: Clipart Animation with Text-to-Video Priors · Int. J. Comput. Vis. 2025
Natural language and speech › Language models and text generation
large language model
0.312025
Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models · CVPR 2025
Visual content generation and editing
image editing
0.312025
LayerPeeler: Autoregressive Peeling for Layer-wise Image Vectorization · SIGGRAPH Asia 2025
Visual content generation and editing
autoregressive generation
0.212023
IconShop: Text-Guided Vector Icon Synthesis with Autoregressive Transformers · ACM Trans. Graph. 2023

Methods — techniques the papers use, named apart from their topics

latent space optimization · 1.7large language model · 1.7image diffusion model · 1.7text-to-video diffusion model · 0.9score distillation sampling · 0.9localized attention control · 0.9diffusion model · 0.9bézier curves · 0.9autoregressive peeling · 0.9as-rigid-as-possible deformation · 0.9
YearPublicationVenuePosition
2025 Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models
abstract
Scalable Vector Graphics (SVG) has become the de facto standard for vector graphics in digital design, offering resolution independence and precise control over individual elements. Despite their advantages, creating high-quality SVG content remains challenging, as it demands technical expertise with professional editing software and a considerable time investment to craft complex shapes. Recent text-to-SVG generation methods aim to make vector graphics creation more accessible, but they still encounter limitations in shape regularity, generalization ability, and expressiveness. To address these challenges, we introduce Chat2SVG, a hybrid framework that combines the strengths of Large Language Models (LLMs) and image diffusion models for text-to-SVG generation. Our approach first uses an LLM to generate semantically meaningful SVG templates from basic geometric primitives. Guided by image diffusion models, a dual-stage optimization pipeline refines paths in latent space and adjusts point coordinates to enhance geometric complexity. Extensive experiments show that Chat2SVG outperforms existing methods in visual fidelity, path regularity, and semantic alignment. Additionally, our system enables intuitive editing through natural language instructions, making professional vector graphics creation accessible to all users. Our code is available at https://chat2svg.github.io/.
Ronghuan Wu, Wanchao Su, Jing Liao 0001
CVPR1
2025 LayerPeeler: Autoregressive Peeling for Layer-wise Image Vectorization
abstract
Image vectorization is a powerful technique that converts raster images into vector graphics, enabling enhanced flexibility and interactivity. However, popular image vectorization tools struggle with occluded regions, producing incomplete or fragmented shapes that hinder editability. While recent advancements have explored optimization-based and learning-based layer-wise image vectorization, these methods face limitations in vectorization quality and flexibility. In this paper, we introduce LayerPeeler, a novel layer-wise image vectorization approach that addresses these challenges through a progressive simplification paradigm. The key to LayerPeeler’s success lies in its autoregressive peeling strategy: by identifying and removing the topmost non-occluded layers while recovering underlying content, we generate vector graphics with complete paths and coherent layer structures. Our method leverages vision-language models to construct a layer graph that captures occlusion relationships among elements, enabling precise detection and description for non-occluded layers. These descriptive captions are used as editing instructions for a finetuned image diffusion model to remove the identified layers. To ensure accurate removal, we employ localized attention control that precisely guides the model to target regions while faithfully preserving the surrounding content. To support this, we contribute a large-scale dataset specifically designed for layer peeling tasks. Extensive quantitative and qualitative experiments demonstrate that LayerPeeler significantly outperforms existing techniques, producing vectorization results with superior path semantics, geometric regularity, and visual fidelity. Our code and dataset will be available at https://layerpeeler.github.io/.
Ronghuan Wu, Wanchao Su, Jing Liao 0001
SIGGRAPH Asia1
2025 AniClipart: Clipart Animation with Text-to-Video Priors
abstract
Abstract Clipart, a pre-made graphic art form, offers a convenient and efficient way of illustrating visual content. Traditional workflows to convert static clipart images into motion sequences are laborious and time-consuming, involving numerous intricate steps like rigging, key animation and in-betweening. Recent advancements in text-to-video generation hold great potential in resolving this problem. Nevertheless, direct application of text-to-video generation models often struggles to retain the visual identity of clipart images or generate cartoon-style motions, resulting in unsatisfactory animation outcomes. In this paper, we introduce AniClipart, a system that transforms static clipart images into high-quality motion sequences guided by text-to-video priors. To generate cartoon-style and smooth motion, we first define Bézier curves over keypoints of the clipart image as a form of motion regularization. We then align the motion trajectories of the keypoints with the provided text prompt by optimizing the Video Score Distillation Sampling (VSDS) loss, which encodes adequate knowledge of natural motion within a pretrained text-to-video diffusion model. With a differentiable As-Rigid-As-Possible shape deformation algorithm, our method can be end-to-end optimized while maintaining deformation rigidity. Experimental results show that the proposed AniClipart consistently outperforms existing image-to-video generation models, in terms of text-video alignment, visual identity preservation, and motion consistency. Furthermore, we showcase the versatility of AniClipart by adapting it to generate a broader array of animation formats, such as layered animation, which allows topological changes.
Ronghuan Wu, Wanchao Su, Kede Ma, Jing Liao 0001
Int. J. Comput. Vis.1
2023 A Fault Diagnosis Model of High-Voltage Circuit Breaker Based on Cyber-Physical Fusion
abstract
High-voltage circuit breakers are widely used in new power systems, and have the function of protecting and controlling transmission lines. With the gradual strengthening of the state perception ability of the new power system, the online monitoring ability of the mechanical fault of the high-voltage circuit breaker has been improved, which provides a relatively complete data basis for the fault diagnosis of the high-voltage circuit breaker. This research presents a method for detecting circuit breaker faults using wavelet vibration and convolutional neural network. Firstly, the continuous wavelet transform is carried out on the vibration signal of the high-voltage circuit breaker, and the wavelet energy frequency band is generated by using the discrete wavelet transform, and the characteristics of the vibration signal of the mechanical fault of the high-voltage circuit breaker are extracted. Then, the preprocessed feature maps are input into the convolutional neural network model to realize fault state diagnosis. Finally, in the simulation example, it is verified that the method proposed in this paper can effectively characterize the change of the mechanical state of the high-voltage circuit breaker, and achieve better diagnostic results for different types of faults.
Gan Tuanjie, Cao Yanzhao, Du Wenjiao, Ronghuan Wu, Chengzhi Ma
IECON4
2023 IconShop: Text-Guided Vector Icon Synthesis with Autoregressive Transformers
abstract
Scalable Vector Graphics (SVG) is a popular vector image format that offers good support for interactivity and animation. Despite its appealing characteristics, creating custom SVG content can be challenging for users due to the steep learning curve required to understand SVG grammars or get familiar with professional editing software. Recent advancements in text-to-image generation have inspired researchers to explore vector graphics synthesis using either image-based methods (i.e., text → raster image → vector graphics) combining text-to-image generation models with image vectorization, or language-based methods (i.e., text → vector graphics script) through pretrained large language models. Nevertheless, these methods suffer from limitations in terms of generation quality, diversity, and flexibility. In this paper, we introduce IconShop, a text-guided vector icon synthesis method using autoregressive transformers. The key to success of our approach is to sequentialize and tokenize SVG paths (and textual descriptions as guidance) into a uniquely decodable token sequence. With that, we are able to exploit the sequence learning power of autoregressive transformers, while enabling both unconditional and text-conditioned icon synthesis. Through standard training to predict the next token on a large-scale vector icon dataset accompanied by textural descriptions, the proposed IconShop consistently exhibits better icon synthesis capability than existing image-based and language-based methods both quantitatively (using the FID and CLIP scores) and qualitatively (through formal subjective user studies). Meanwhile, we observe a dramatic improvement in generation diversity, which is validated by the objective Uniqueness and Novelty measures. More importantly, we demonstrate the flexibility of IconShop with multiple novel icon synthesis tasks, including icon editing, icon interpolation, icon semantic combination, and icon design auto-suggestion.
Ronghuan Wu, Wanchao Su, Kede Ma, Jing Liao 0001
ACM Trans. Graph.1