Yabo Zhang

dblp:231/0624 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-1019-4334ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models
abstract
Artistic typography is a technique that enables one to visualize the meaning of an input character in an imaginable and readable manner. With powerful text-to-image diffusion models, existing methods directly design the overall geometry and texture of input character, making it challenging to ensure both creativity and legibility. In this paper, we introduce a dual-branch, training-free method called VitaGlyph, enabling flexible artistic typography with controllable geometry changes while maintaining legibility. The key insight of VitaGlyph is to treat the input character as a scene composed of a Subject and its Surrounding, which are rendered with varying degrees of geometric transformation. To enhance the visual appeal and creativity of the generated artistic typography, the Subject flexibly expresses the essential concept of the input character, while the Surrounding enriches relevant background without altering the shape. Specifically, we implement VitaGlyph through a three-phase framework: (i) Knowledge Acquisition leverages large language models to design text descriptions for the Subject and Surrounding. (ii) Regional Interpretation detects the part that matches the subject description most closely and refines the structure using Semantic Typography. (iii) Attentional Compositional Generation separately renders the textures of the Subject and Surrounding and blends them in an attention-based manner. Experiments demonstrate that VitaGlyph not only achieves better artistry and legibility, but also manages to depict multiple customized concepts, facilitating more creative and pleasing artistic typography generation. Our code is available at https://github.com/Carlofkl/VitaGlyph.
Kailai Feng, Yabo Zhang, Haodong Yu, Zhilong Ji, Jinfeng Bai, Wangmeng Zuo
WACV2
2025 VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models
abstract
Text-to-image diffusion models (T2I) have demonstrated unprecedented capabilities in creating realistic and aesthetic images. On the contrary, text-to-video diffusion models (T2V) still lag far behind in frame quality and text alignment, owing to insufficient quality and quantity of training videos. In this paper, we introduce VideoElevator, a training-free and plug-and-play method, which elevates the performance of T2V using superior capabilities of T2I. Different from conventional T2V sampling (i.e., temporal and spatial modeling), VideoElevator explicitly decomposes each sampling step into temporal motion refining and spatial quality elevating. Specifically, temporal motion refining uses encapsulated T2V to enhance temporal consistency, followed by inverting to the noise distribution required by T2I. Then, spatial quality elevating harnesses inflated T2I to directly predict less noisy latent, adding more photo-realistic details. We have conducted experiments in extensive prompts under the combination of various T2V and T2I. The results show that VideoElevator not only improves the performance of T2V baselines with foundational T2I, but also facilitates stylistic video synthesis with personalized T2I. Please watch all videos in supplementary materials for better view.
Yabo Zhang, Yuxiang Wei 0001, Xianhui Lin, Zheng Hui, Peiran Ren, Xuansong Xie, Wangmeng Zuo
AAAI1
2025 MC^2: Multi-concept Guidance for Customized Multi-concept Generation
abstract
Customized text-to-image generation, which synthesizes images based on user-specified concepts, has made significant progress in handling individual concepts. However, when extended to multiple concepts, existing methods often struggle with properly integrating different models and avoiding the unintended blending of characteristics from distinct concepts. In this paper, we propose MC2, a novel approach for multi-concept customization that enhances flexibility and fidelity through inference-time optimization. MC2enables the integration of multiple single-concept models with heterogeneous architectures. By adaptively refining attention weights between visual and textual tokens, our method ensures that image regions accurately correspond to their associated concepts while minimizing interference between concepts. Extensive experiments demonstrate that MC2outperforms training-based methods in terms of prompt-reference alignment. Furthermore, MC2can be seamlessly applied to text-to-image generation, providing robust compositional capabilities. To facilitate the evaluation of multi-concept customization, we also introduce a new benchmark, MC++. The code is available at https://github.com/jiangJiaxiu/MC-2.
Jiaxiu Jiang, Yabo Zhang, Kailai Feng, Xiaohe Wu, Wenbo Li 0002, Renjing Pei, Wangmeng Zuo
CVPR2
2025 FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
abstract
Interactive image editing allows users to modify images through visual interaction operations such as drawing, clicking, and dragging. Existing methods construct such supervision signals from videos, as they capture how objects change with various physical interactions. However, these models are usually built upon text-to-image diffusion models, so necessitate (i) massive training samples and (ii) an additional reference encoder to learn real-world dynamics and visual consistency. In this paper, we reformulate this task as an image-to-video generation problem, so that inherit powerful video diffusion priors to reduce training costs and ensure temporal consistency. Specifically, we introduce FramePainter as an efficient instantiation of this formulation. Initialized with Stable Video Diffusion, it only uses a lightweight sparse control encoder to inject editing signals. Considering the limitations of temporal attention in handling large motion between two frames, we further propose matching attention to enlarge the receptive field while encouraging dense correspondence between edited and source image tokens. We highlight the effectiveness and efficiency of FramePainter across various of editing signals: it domainantly outperforms previous state-of-the-art methods with far less training data, achieving highly seamless and coherent editing of images, \eg, automatically adjust the reflection of the cup. Moreover, FramePainter also exhibits exceptional generalization in scenarios not present in real-world videos, \eg, transform the clownfish into shark-like shape. Our code will be available at https://github.com/YBYBZhang/FramePainter.
Yabo Zhang, Xinpeng Zhou, Yihan Zeng, Hang Xu 0004, Hui Li 0035, Wangmeng Zuo
ICCV1
2025 Personalized Image Generation with Deep Generative Models: A Decade Survey
abstract
Recent advances in generative models have significantly facilitated the development of personalized content creation. Given a small set of images containing a user-specific concept, personalized image generation allows the user to create images that incorporate that concept while adhering to provided text descriptions. The technologies used for personalization have evolved alongside the development of generative models, with their distinct and interrelated components. In this survey, we present a comprehensive review of generalized personalized image generation across various generative models, including traditional GANs, contemporary text-to-image diffusion models, and emerging multi-modal autoregressive (AR) models. We first define a unified framework that standardizes the personalization process across different generative models, encompassing three key components: inversion spaces, inversion methods, and personalization schemes. This unified framework offers a structured approach to dissecting and comparing personalization techniques across different generative architectures. Building upon our framework, we provide an in-depth analysis of personalization techniques within each generative model, highlighting their unique contributions and innovations. Through comparative analysis, we elucidate the current landscape of personalized image generation, identifying commonalities and distinguishing features of existing methods. Finally, we discuss open challenges in the field and propose potential directions for future research. We keep a bibliography of related works at https://github.com/csyxwei/Awesome-Personalized-Image-Generation.
Yuxiang Wei 0001, Yiheng Zheng, Yabo Zhang, Ming Liu 0018, Zhilong Ji, Lei Zhang 0006, Wangmeng Zuo
Comput. Vis. Media3
2024 VQ-FONT: Few-Shot Font Generation with Structure-Aware Enhancement and Quantization
abstract
Few-shot font generation is challenging, as it needs to capture the fine-grained stroke styles from a limited set of reference glyphs, and then transfer to other characters, which are expected to have similar styles. However, due to the diversity and complexity of Chinese font styles, the synthesized glyphs of existing methods usually exhibit visible artifacts, such as missing details and distorted strokes. In this paper, we propose a VQGAN-based framework (i.e., VQ-Font) to enhance glyph fidelity through token prior refinement and structure-aware enhancement. Specifically, we pre-train a VQGAN to encapsulate font token prior within a code-book. Subsequently, VQ-Font refines the synthesized glyphs with the codebook to eliminate the domain gap between synthesized and real-world strokes. Furthermore, our VQ-Font leverages the inherent design of Chinese characters, where structure components such as radicals and character components are combined in specific arrangements, to recalibrate fine-grained styles based on references. This process improves the matching and fusion of styles at the structure level. Both modules collaborate to enhance the fidelity of the generated fonts. Experiments on a collected font dataset show that our VQ-Font outperforms the competing methods both quantitatively and qualitatively, especially in generating challenging styles. Our code is available at https://github.com/Yaomingshuai/VQ-Font.
Mingshuai Yao, Yabo Zhang, Xianhui Lin, Xiaoming Li 0002, Wangmeng Zuo
AAAI2
2024 ControlVideo: Training-free Controllable Text-to-video Generation
abstract
Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart lags behind due to the excessive training cost. To avert the training burden, we propose a training-free ControlVideo to produce high-quality videos based on the provided text prompts and motion sequences. Specifically, ControlVideo adapts a pre-trained text-to-image model (i.e., ControlNet) for controllable text-to-video generation. To generate continuous videos without flicker effect, we propose an interleaved-frame smoother to smooth the intermediate frames. In particular, interleaved-frame smoother splits the whole videos with successive three-frame clips, and stabilizes each clip by updating the middle frame with the interpolation among other two frames in latent space. Furthermore, a fully cross-frame interaction mechanism have been exploited to further enhance the frame consistency, while a hierarchical sampler is employed to produce long videos efficiently. Extensive experiments demonstrate that our ControlVideo outperforms the state-of-the-arts both quantitatively and qualitatively. It is worthy noting that, thanks to the efficient designs, ControlVideo could generate both short and long videos within several minutes using one NVIDIA 2080Ti. Code and videos are available at [this link](https://github.com/YBYBZhang/ControlVideo).
Yabo Zhang, Yuxiang Wei 0001, Dongsheng Jiang, Xiaopeng Zhang 0008, Wangmeng Zuo, Qi Tian 0001
ICLR1
2023 Tailoring Routing Protocols for Flying Ad Hoc Networks: Challenges and Possible Countermeasures
abstract
Implementing an resilient, efficient, and reliable network structure is crucial for highly dynamic unmanned aerial vehicle (UAV) swarms, for which flying ad hoc network (FANET) is the most suitable form. Similar to traditional ad hoc networks, the performance of FANET largely depends on the efficiency, reliability, and stability of routing schemes. However, unique characteristics of UAV make the routing design of FANET face more challenges. In order to better understand the development of FANET routing schemes, this paper attempts to clarify the current research status and grasp the future development trend of FANET routing by reviewing and analyzing relevant literatures in the past decade. Results show that geographic routing, delay tolerant network, and opportunity forward are possible countermeasures to the challenges of FANET routing.
Wei Liu 0059, Ming Xu 0016, Yabo Zhang, Yu Xia 0009, Jing Mao, Daqing Huang
APCC4
2023 Implementing Hardware-in-the-Loop Protocol Simulation for UAV Networks
abstract
Existing works on UAV network protocols generally use software simulators for performance evaluation, which makes the analysis results often differ significantly from the test results in actual environments. In order to make the analysis of UAV network protocols more realistic, it is necessary to introduce actual UAV nodes into the simulation. By drawing on the idea of hardware-in-the-loop (HIL) simulation, this paper proposes a simulation framework with real UAV nodes in the loop. To show the potentials of the proposed simulation framework, an instance for HIL simulation of routing protocols is implemented. Preliminary results indicate that using a small number of actual flying UAV nodes as an organic component of simulation could reduce the gap between simulation results and actual situations.
Ming Xu 0016, Wei Liu 0059, Yabo Zhang, Yu Xia 0009, Daqing Huang
APCC4
2023 ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation
abstract
In addition to the unprecedented ability in imaginary creation, large text-to-image models are expected to take customized concepts in image generation. Existing works generally learn such concepts in an optimization-based manner, yet bringing excessive computation or memory burden. In this paper, we instead propose a learning-based encoder, which consists of a global and a local mapping networks for fast and accurate customized text-to-image generation. In specific, the global mapping network projects the hierarchical features of a given image into multiple "new" words in the textual word embedding space, i.e., one primary word for well-editable concept and other auxiliary words to exclude irrelevant disturbances (e.g., background). In the meantime, a local mapping network injects the encoded patch features into cross attention layers to provide omitted details, without sacrificing the editability of primary concepts. We compare our method with existing optimization-based approaches on a variety of user-defined concepts, and demonstrate that our method enables high-fidelity inversion and more robust editability with a significantly faster encoding process. Our code is publicly available at https://github.com/csyxwei/ELITE.
Yuxiang Wei 0001, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo
ICCV2
2023 Misalignment Insensitive Perceptual Metric for Full Reference Image Quality Assessment
Yue Cao 0009, Yabo Zhang, Wangmeng Zuo
PRCV (11)3
2022 AI-enabled Multi-modal Network Anomaly Association: A Deep Self/Semi-Supervised Learning Approach
abstract
In nowadays large-scale networks, it is challenging for network operation and maintenance systems to analyze the reported massive network anomaly information. To handle this problem, we proposed a deep multi-modal learning approach called multi-modal anomaly root cause analysis, which enables network operation and maintenance systems to automatically and effectively associate the related network anomalies that appear from different modalities or aspects, and then locate the root causes. As a self/semi-supervised approach, our proposal is capable of realizing self-learning, self-adapting, and does not rely on a large number of manual annotations. According to the experimental results in a real large-scale network, without any annotations, our approach achieves up to 14% accuracy improvement in terms of multi-modal network anomaly association and root cause locating compared to the classical association rule mining algorithm Apriori, while its performance turns even much better when a few of labeled training samples are provided. The experiment also well proves the versatility and self-adaptability of our approach, which means our learning-based approach is able to not only achieve fast convergence but also automatically adapt itself to network changes.
Yinan Tang, Yabo Zhang, Zhifeng Yin, Jianxi Deng, Yong Cui 0001
ICC2
2022 Towards Diverse and Faithful One-shot Adaption of Generative Adversarial Networks
abstract
One-shot generative domain adaption aims to transfer a pre-trained generator on one domain to a new domain using one reference image only. However, it remains very challenging for the adapted generator (i) to generate diverse images inherited from the pre-trained generator while (ii) faithfully acquiring the domain-specific attributes and styles of the reference image. In this paper, we present a novel one-shot generative domain adaption method, i.e., DiFa, for diverse generation and faithful adaptation. For global-level adaptation, we leverage the difference between the CLIP embedding of the reference image and the mean embedding of source images to constrain the target generator. For local-level adaptation, we introduce an attentive style loss which aligns each intermediate token of an adapted image with its corresponding token of the reference image. To facilitate diverse generation, selective cross-domain consistency is introduced to select and retain domain-sharing attributes in the editing latent $\mathcal{W}+$ space to inherit the diversity of the pre-trained generator. Extensive experiments show that our method outperforms the state-of-the-arts both quantitatively and qualitatively, especially for the cases of large domain gap. Moreover, our DiFa can easily be extended to zero-shot generative domain adaption with appealing results.
Yabo Zhang, Mingshuai Yao, Yuxiang Wei 0001, Zhilong Ji, Jinfeng Bai, Wangmeng Zuo
NeurIPS1
2018 Enhanced Metric Learning via Dempster-Shafer Evidence Theory
Ying Li 0028, Yabo Zhang, Yaxin Peng
ICONIP (3)2