Rui Yang 0011

dblp:92/1942-11 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-1996-2993ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 VideoSketcher: A Training-Free Approach for Coherent Video Sketch Transfer
abstract
Generating high-quality sketches from video requires a nuanced understanding of semantic content and visual structure, particularly for complex scenes across diverse sketch styles. Efficient and flexible video-to-sketch style transformation remains a significant challenge. We introduce VideoSketcher, a training-free framework for style-controllable sketch video generation that preserves frame structure while applying specified sketch aesthetics. Leveraging text-to-image diffusion models, VideoSketcher utilizes strong semantic priors without the need for extensive training. Our approach enforces temporal consistency by retaining latent information across frames and employs a Time-Linked Attention mechanism to capture structural elements from the source video and inject stylistic information from the reference image. To bridge the semantic gap between sketches and original video content, we introduce Sketch Directive Amplification for selective transfer of stylistic features. Additionally, a Stroke Graph Regularization strategy, comprising line and point loss, refines line consistency in the latent space. Extensive experiments validate VideoSketcher’s superior temporal stability and fidelity across diverse sketch styles and content. Video demos can be found in the supplementary materials.
Huining Li, Bangzhen Liu, Rui Yang 0011, Chenshu Xu, Xufang Pang, Shengfeng He
WACV3
2026 Large-area damage inpainting of ancient paintings with long-range contextual
Jumei Chang, Zengguo Sun, Shengfeng He, Rui Yang 0011, Mohammed Al-Madhehagi, Xiaojun Wu 0002
Eng. Appl. Artif. Intell.4
2025 One-Shot Reference-based Structure-Aware Image to Sketch Synthesis
abstract
Generating sketches that accurately reflect the content of reference images presents numerous challenges. Current methods either require paired training data or fail to accommodate a wider range and diversity of sketch styles. While pre-trained diffusion models have shown strong text-based control capabilities for reference-based content sketch generation, state-of-the-art methods still struggle with reference-based sketch generation for given content. The main difficulties lie in (1) balancing content preservation with style enhancement, and (2) representing content image textures at varying levels of abstraction to approximate the reference sketch style. In this paper, we propose a method (Ref2Sketch-SA) that transforms a given content image into a sketch based on a reference sketch. The core strategies include (1) using DDIM Inversion to enhance structural consistency in the sketch generation of content images; (2) injecting noise into the input image during the denoising process to produce a sketch that retains content attributes while aligning with, yet differing in texture from, the reference. Our model demonstrates superior performance across multiple evaluation metrics, including user style preference.
Rui Yang 0011, Honghong Yang, Qin Lei, Mianxiong Dong, Kaoru Ota, Xiaojun Wu 0002
AAAI1
2025 Stroke2Sketch: Harnessing Stroke Attributes for Training-Free Sketch Generation
abstract
Generating sketches guided by reference styles requires precise transfer of stroke attributes, such as line thickness, deformation, and texture sparsity, while preserving semantic structure and content fidelity. To this end, we propose Stroke2Sketch, a novel training-free framework that introduces cross-image stroke attention, a mechanism embedded within self-attention layers to establish fine-grained semantic correspondences and enable accurate stroke attribute transfer. This allows our method to adaptively integrate reference stroke characteristics into content images while maintaining structural integrity. Additionally, we develop adaptive contrast enhancement and semantic-focused attention to reinforce content preservation and foreground emphasis. Stroke2Sketch effectively synthesizes stylistically faithful sketches that closely resemble handcrafted results, outperforming existing methods in expressive stroke control and semantic coherence. Codes are available at https://github.com/rane7/Stroke2Sketch.
Rui Yang 0011, Huining Li, Yiyi Long, Xiaojun Wu 0002, Shengfeng He
ICCV1
2025 Semantic layout-guided diffusion model for high-fidelity image synthesis in 'The Thousand Li of Rivers and Mountains'
Rui Yang 0011, Kaoru Ota, Mianxiong Dong, Xiaojun Wu 0002
Expert Syst. Appl.1
2025 MixSA: Training-Free Reference-Based Sketch Extraction via Mixture-of-Self-Attention
abstract
Current sketch extraction methods either require extensive training or fail to capture a wide range of artistic styles, limiting their practical applicability and versatility. We introduce Mixture-of-Self-Attention (MixSA), a training-free sketch extraction method that leverages strong diffusion priors for enhanced sketch perception. At its core, MixSA employs a mixture-of-self-attention technique, which manipulates self-attention layers by substituting the keys and values with those from reference sketches. This allows for the seamless integration of brushstroke elements into initial outline images, offering precise control over texture density and enabling interpolation between styles to create novel, unseen styles. By aligning brushstroke styles with the texture and contours of colored images, particularly in late decoder layers handling local textures, MixSA addresses the common issue of color averaging by adjusting initial outlines. Evaluated with various perceptual metrics, MixSA demonstrates superior performance in sketch quality, flexibility, and applicability. This approach not only overcomes the limitations of existing methods but also empowers users to generate diverse, high-fidelity sketches that more accurately reflect a wide range of artistic expressions.
Rui Yang 0011, Xiaojun Wu 0002, Shengfeng He
IEEE Trans. Vis. Comput. Graph.1
2024 Expanding Crack Segmentation Dataset with Crack Growth Simulation and Feature Space Diversity
abstract
In this paper, we address the significant challenge of data scarcity in the field of crack segmentation, a key aspect of structural health monitoring. To tackle this, we introduce the CrackGrowDiff framework, an innovative approach for expanding crack datasets. Utilizing a two-stage controllable generation process that combines a random walk algorithm and semantic diffusion models, our framework minimizes discrepancy of misalignment between synthetic data and original data while enhancing data informativeness. We further ensure the quality and informativeness of synthetic data through feature space diversity, employing a pre-trained Variational Autoencoder (VAE) for selection based on Kullback-Leibler (KL) divergence. Comparative experiments demonstrate CrackGrowDiff’s superiority over traditional data augmentation and GANs-based methods, making it a substantial advancement in addressing the data scarcity in crack segmentation tasks. A DEMO and related code will be made public: https://huggingface.co/spaces/QinLei086/Two-stage-SDM-for-crack-dataset-expending
Qin Lei, Rui Yang 0011, Rongzhen Li, Muyang He, Mianxiong Dong, Kaoru Ota
ICME2
2024 HFA-GTNet: Hierarchical Fusion Adaptive Graph Transformer network for dance action recognition
Ru Jia, Rui Yang 0011, Honghong Yang, Xiaojun Wu 0002, Peng Li 0016, Yuping Su
J. Vis. Commun. Image Represent.3
2024 Special perceptual parsing for Chinese landscape painting scene understanding: a semantic segmentation approach
Rui Yang 0011, Honghong Yang, Ru Jia, Xiaojun Wu 0002
Neural Comput. Appl.1