Yinhan Hu

dblp:326/1610 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 76% Deep learning architectures and training · 24%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.122025
StyO: Stylize Your Face in Only One-Shot · AAAI 2025
Control and Realism: Best of Both Worlds in Layout-to-Image without Training · ICML 2025
Machine learning › Generative modeling › diffusion model › guided diffusion
classifier guidance
1.012026
FreLay: Frequency-aware Energy Function for Training-free Layout-to-Image Generation · AAAI 2026
Machine learning › Generative modeling › diffusion model
diffusion sampling
1.012026
FreLay: Frequency-aware Energy Function for Training-free Layout-to-Image Generation · AAAI 2026
Machine learning › Generative modeling › image generation › conditional image synthesis
layout-to-image generation
1.012026
FreLay: Frequency-aware Energy Function for Training-free Layout-to-Image Generation · AAAI 2026
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.912025
StyO: Stylize Your Face in Only One-Shot · AAAI 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
StyO: Stylize Your Face in Only One-Shot · AAAI 2025
Visual content generation and editing › image generation › controllable image generation
layout-to-image generation
0.912025
Control and Realism: Best of Both Worlds in Layout-to-Image without Training · ICML 2025
Machine learning › Deep learning architectures and training › regularization
dropout
0.712023
DropKey for Vision Transformer · CVPR 2023
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.712023
DropKey for Vision Transformer · CVPR 2023
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.712023
DropKey for Vision Transformer · CVPR 2023
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.312025
Control and Realism: Best of Both Worlds in Layout-to-Image without Training · ICML 2025
Visual content generation and editing › face editing
face stylization
0.312025
StyO: Stylize Your Face in Only One-Shot · AAAI 2025
Visual content generation and editing
image editing
0.312025
StyO: Stylize Your Face in Only One-Shot · AAAI 2025

Methods — techniques the papers use, named apart from their topics

triple reconstruction loss · 1.7non-local attention energy function · 1.7langevin dynamics · 1.7cross-attention map constraint · 1.7contrastive text prompt · 1.7adaptive update · 1.7frequency-domain analysis · 1.0boltzmann distribution normalization · 1.0dropkey · 0.7decreasing drop ratio schedule · 0.7
YearPublicationVenuePosition
2026 FreLay: Frequency-aware Energy Function for Training-free Layout-to-Image Generation
abstract
Layout-to-Image generation has significantly advanced content creation by enabling the rendering of visual text under predefined spatial layouts. Current approaches achieve training-free layout guidance by constructing attention-based energy functions to derive correction gradients. In this paper, we demonstrate that vanilla energy functions suffer from two limitations, resulting in imprecise layout control and visually unrealistic artifacts. First, the normalizing factor of the Boltzmann distribution defined by the energy functions is non-negligible when calculating correction gradients, yet current energy functions cannot compute this factor exactly. Furthermore, while attention varies over time during the denoising process, existing approaches employ a fixed formulation. To address these challenges, we introduce FreLay, a novel training-free approach equipped with a frequency-aware energy function. Our method first reformulates the energy function to handle the normalization factor, enabling accurate computation of correction gradients. Simultaneously, leveraging the prior knowledge that low-frequency information deteriorates slower during noise addition, we design a time-specific energy function for each timestep from a frequency-domain perspective. Experimental results demonstrate that FreLay consistently outperforms existing state-of-the-art training-free methods by a large margin both qualitatively and quantitatively across multiple datasets.
Bonan Li, Yinhan Hu, Songhua Liu, Zeyu Xiao 0002, Xinchao Wang
AAAI2
2025 StyO: Stylize Your Face in Only One-Shot
abstract
This paper focuses on face stylization with a single artistic target. Existing works for this task often fail to retain the source content while achieving geometry variation. Here, we present a novel StyO model, i.e., Stylize the face in only One-shot, to solve the above problem. In particular, StyO exploits a disentanglement and recombination strategy. It first disentangles the content and style of source and target images into identifiers, which are then recombined in a cross manner to derive the stylized face image. In this way, StyO decomposes complex images into independent and specific attributes, and simplifies one-shot face stylization as the combination of different attributes from input images, thus producing results better matching face geometry of target image and content of source one. StyO is implemented with latent diffusion models (LDM) and composed of two key modules: 1) Identifier Disentanglement Learner (IDL) for disentanglement phase. It represents identifiers as contrastive text prompts, i.e. positive and negative descriptions. And it introduces a novel triple reconstruction loss to fine-tune the pre-trained LDM for encoding style and content into corresponding identifiers; 2) Fine-graind Content Controller (FCC) for recombination phase. It recombines disentangled identifiers from IDL to form an augmented text prompt for generating stylized faces. In addition, FCC also constrains the cross-attention maps of latent and text features to preserve source face details in results. The extensive evaluation shows that StyO produces high-quality images on numerous paintings of various styles and outperforms the current state-of-the-art.
Bonan Li, Xuecheng Nie, Congying Han, Yinhan Hu, Xinmin Qiu, Tiande Guo
AAAI5
2025 Control and Realism: Best of Both Worlds in Layout-to-Image without Training
abstract
Layout-to-Image generation aims to create complex scenes with precise control over the placement and arrangement of subjects. Existing works have demonstrated that pre-trained Text-to-Image diffusion models can achieve this goal without training on any specific data; however, they often face challenges with imprecise localization and unrealistic artifacts. Focusing on these drawbacks, we propose a novel training-free method, WinWinLay. At its core, WinWinLay presents two key strategies—Non-local Attention Energy Function and Adaptive Update—that collaboratively enhance control precision and realism. On one hand, we theoretically demonstrate that the commonly used attention energy function introduces inherent spatial distribution biases, hindering objects from being uniformly aligned with layout instructions. To overcome this issue, non-local attention prior is explored to redistribute attention scores, facilitating objects to better conform to the specified spatial conditions. On the other hand, we identify that the vanilla backpropagation update rule can cause deviations from the pre-trained domain, leading to out-of-distribution artifacts. We accordingly introduce a Langevin dynamics-based adaptive update scheme as a remedy that promotes in-domain updating while respecting layout constraints. Extensive experiments demonstrate that WinWinLay excels in controlling element placement and achieving photorealistic visual fidelity, outperforming the current state-of-the-art methods.
Bonan Li, Yinhan Hu, Songhua Liu, Xinchao Wang
ICML2
2023 DropKey for Vision Transformer
abstract
In this paper, we focus on analyzing and improving the dropout technique for self-attention layers of Vision Transformer, which is important while surprisingly ignored by prior works. In particular, we conduct researches on three core questions: First, what to drop in self-attention layers? Different from dropping attention weights in literature, we propose to move dropout operations forward ahead of attention matrix calculation and set the Key as the dropout unit, yielding a novel dropout-before-softmax scheme. We theoretically verify that this scheme helps keep both regularization and probability features of attention weights, alleviating the overfittings problem to specific patterns and enhancing the model to globally capture vital information; Second, how to schedule the drop ratio in consecutive layers? In contrast to exploit a constant drop ratio for all layers, we present a new decreasing schedule that gradually decreases the drop ratio along the stack of self-attention layers. We experimentally validate the proposed schedule can avoid overfittings in low-level features and missing in high-level semantics, thus improving the robustness and stableness of model training; Third, whether need to perform structured dropout operation as CNN? We attempt patch-based block-version of dropout operation and find that this useful trick for CNN is not essential for ViT. Given exploration on the above three questions, we present the novel Drop-Key method that regards Key as the drop unit and exploits decreasing schedule for drop ratio, improving ViTs in a general way. Comprehensive experiments demonstrate the effectiveness of DropKey for various ViT architectures, e.g. T2T, VOLO, CeiT and DeiT, as well as for various vision tasks, e.g., image classification, object detection, human-object interaction detection and human body shape recovery.
Bonan Li, Yinhan Hu, Xuecheng Nie, Congying Han, Xiangjian Jiang, Tiande Guo, Luoqi Liu
CVPR2
2022 DFS: A Diverse Feature Synthesis Model for Generalized Zero-Shot Learning
abstract
Generative based strategy has shown great potential in the Generalized Zero-Shot Learning task. However, it suffers severe generalization problem due to lacking of feature diversity for unseen classes to train a good classifier. In this paper, we propose to enhance the generalizability of GZSL models via improving feature diversity of unseen classes. For this purpose, we present a novel Diverse Feature Synthesis (DFS) model. Different from prior works that solely utilize semantic knowledge in the generation process, DFS leverages visual knowledge with semantic one in a unified way, thus deriving class-specific diverse feature samples and leading to robust classifier for recognizing both seen and unseen classes in the testing phase. To simplify the learning, DFS represents visual and semantic knowledge in the aligned space, making it able to produce good feature samples with a low-complexity implementation. Accordingly, DFS is composed of two consecutive generators: an aligned feature generator, transferring semantic and visual representations into aligned features; a synthesized feature generator, producing diverse feature samples of unseen classes in the aligned space. We conduct comprehensive experiments to verify the efficacy of DFS. Results demonstrate its effectiveness to generate diverse features for unseen classes, leading to superior performance on multiple benchmarks. Code will be released upon acceptance.
Bonan Li, Yinhan Hu, Congying Han, Tiande Guo
ICPR2