Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hui Zhang 0090

dblp:181/2846-90 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
0009-0004-8317-5881ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 48% Time series and sequential data · 30% Efficient and distributed learning · 14%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 91% Image and video processing · 9%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
3.542025
DiffusionAD: Norm-Guided One-Step Denoising Diffusion for Anomaly Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2025
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation · ICCV 2025
BlockDance: Reuse Structurally Similar Spatio-Temporal Features to Accelerate Diffusion Transformers · CVPR 2025
Machine learning › Efficient and distributed learning
inference acceleration
1.722025
BlockDance: Reuse Structurally Similar Spatio-Temporal Features to Accelerate Diffusion Transformers · CVPR 2025
AdaDiff: Adaptive Step Selection for Fast Diffusion Models · AAAI 2025
Machine learning › Time series and sequential data
anomaly detection
1.522025
DiffusionAD: Norm-Guided One-Step Denoising Diffusion for Anomaly Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Prototypical Residual Networks for Anomaly Detection and Localization · CVPR 2023
Machine learning › Generative modeling › diffusion model
diffusion sampling
0.912025
AdaDiff: Adaptive Step Selection for Fast Diffusion Models · AAAI 2025
Machine learning › Generative modeling › diffusion model › diffusion model acceleration
diffusion transformer acceleration
0.912025
BlockDance: Reuse Structurally Similar Spatio-Temporal Features to Accelerate Diffusion Transformers · CVPR 2025
Machine learning › Time series and sequential data › anomaly detection
generative anomaly detection
0.912025
DiffusionAD: Norm-Guided One-Step Denoising Diffusion for Anomaly Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Generative modeling › diffusion model › diffusion transformer
multimodal diffusion transformer
0.912025
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation · ICCV 2025
Visual content generation and editing › video generation
image-to-video generation
0.912025
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance · ICCV 2025
Visual content generation and editing › image generation › controllable image generation
layout-to-image generation
0.912025
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation · ICCV 2025
Visual content generation and editing
video generation
0.912025
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance · ICCV 2025
Machine learning › Time series and sequential data › anomaly detection
anomaly segmentation
0.712023
Prototypical Residual Networks for Anomaly Detection and Localization · CVPR 2023
Machine learning › Time series and sequential data › anomaly detection
industrial anomaly detection
0.712023
Prototypical Residual Networks for Anomaly Detection and Localization · CVPR 2023
Image and video processing
image restoration
0.312025
DiffusionAD: Norm-Guided One-Step Denoising Diffusion for Anomaly Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.212023
Prototypical Residual Networks for Anomaly Detection and Localization · CVPR 2023

Methods — techniques the papers use, named apart from their topics

siamese network · 1.7segmentation network · 1.7norm-guided denoising · 1.7large language model · 1.7diffusion model · 1.7spatio-temporal feature caching · 0.9reinforcement learning · 0.9policy gradient · 0.9mask conditioning · 0.9lightweight decision network · 0.9dense-to-sparse trajectory guidance · 0.9bounding box conditioning · 0.9anomaly generation · 0.7
YearPublicationVenuePosition
2025 AdaDiff: Adaptive Step Selection for Fast Diffusion Models
abstract
Diffusion models, as a type of generative model, have achieved impressive results in generating images and videos conditioned on textual conditions. However, the generation process of diffusion models involves denoising dozens of steps to produce photorealistic images/videos, which is computationally expensive. Unlike previous methods that design ``one-size-fits-all'' approaches for speed up, we argue denoising steps should be sample-specific conditioned on the richness of input texts. To this end, we introduce AdaDiff, a lightweight framework designed to learn instance-specific step usage policies, which are then used by the diffusion model for generation. AdaDiff is optimized using a policy gradient method to maximize a carefully designed reward function, balancing inference time and generation quality. We conduct experiments on three image generation and two video generation benchmarks and demonstrate that our approach achieves similar visual quality compared to the baseline using a fixed 50 denoising steps while reducing inference time by at least 33%, going as high as 40%. Furthermore, our method can be used on top of other acceleration methods to provide further speed benefits. Lastly, qualitative analysis shows that AdaDiff allocates more steps to more informative prompts and fewer steps to simpler prompts.
Hui Zhang 0090, Zuxuan Wu, Yu-Gang Jiang 0001
AAAI1
2025 BlockDance: Reuse Structurally Similar Spatio-Temporal Features to Accelerate Diffusion Transformers
abstract
Diffusion models have demonstrated impressive generation capabilities, particularly with recent advancements leveraging transformer architectures to improve both visual and artistic quality. However, Diffusion Transformers (DiTs) continue to encounter challenges related to low inference speed, primarily due to the iterative denoising process. To address this issue, we propose BlockDance, a training-free approach that explores feature similarities at adjacent time steps to accelerate DiTs. Unlike previous feature-reuse methods that lack tailored reuse strategies for features at different scales, BlockDance prioritizes the identification of the most structurally similar features, referred to as Structurally Similar Spatio-Temporal (STSS) features. These features are primarily located within the structure-focused blocks of the transformer during the later stages of denoising. Block-Dance caches and reuses these highly similar features to mitigate redundant computation, thereby accelerating DiTs while maximizing consistency with the generated results of the original model. Furthermore, considering the diversity of generated content and the varying distributions of redundant features, we introduce BlockDance-Ada, a lightweight decision-making network tailored for instance-specific acceleration. BlockDance-Ada dynamically allocates resources and provides superior content quality. Both BlockDance and BlockDance-Ada have proven effective across various generation tasks and models, achieving accelerations between 25% and 50% while maintaining generation quality.
Hui Zhang 0090, Tingwei Gao, Zuxuan Wu
CVPR1
2025 MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
abstract
Recent advances in video generation have led to remarkable improvements in visual quality and temporal coherence. Upon this, trajectory-controllable video generation has emerged to enable precise object motion control through explicitly defined spatial paths. However, existing methods struggle with complex object movements and multi-object motion control, resulting in imprecise trajectory adherence, poor object consistency, and compromised visual quality. Furthermore, these methods only support trajectory control in a single format, limiting their applicability in diverse scenarios. Additionally, there is no publicly available dataset or benchmark specifically tailored for trajectory-controllable video generation, hindering robust training and systematic evaluation. To address these challenges, we introduce MagicMotion, a novel image-to-video generation framework that enables trajectory control through three levels of conditions from dense to sparse: masks, bounding boxes, and sparse boxes. Given an input image and trajectories, MagicMotion seamlessly animates objects along defined trajectories while maintaining object consistency and visual quality. Furthermore, we present MagicData, a large-scale trajectory-controlled video dataset, along with an automated pipeline for annotation and filtering. We also introduce MagicBench, a comprehensive benchmark that assesses both video quality and trajectory control accuracy across different numbers of objects. Extensive experiments demonstrate that MagicMotion outperforms previous methods across various metrics. Our project page are publicly available at https://quanhaol.github.io/magicmotion-site.
Quanhao Li, Rui Wang 0095, Hui Zhang 0090, Qi Dai 0001, Zuxuan Wu
ICCV4
2025 CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
abstract
Diffusion models have been recognized for their ability to generate images that are not only visually appealing but also of high artistic quality. As a result, Layout-to-Image (L2I) generation has been proposed to leverage region-specific positions and descriptions to enable more precise and controllable generation. However, previous methods primarily focus on UNet-based models (\eg SD1.5 and SDXL), and limited effort has explored Multimodal Diffusion Transformers (MM-DiTs), which have demonstrated powerful image generation capabilities. Enabling MM-DiT for layout-to-image generation seems straightforward but is challenging due to the complexity of how layout is introduced, integrated, and balanced among multiple modalities. To this end, we explore various network variants to efficiently incorporate layout guidance into MM-DiT, and ultimately present SiamLayout. To inherit the advantages of MM-DiT, we use a separate set of network weights to process the layout, treating it as equally important as the image and text modalities. Meanwhile, to alleviate the competition among modalities, we decouple the image-layout interaction into a siamese branch alongside the image-text one and fuse them in the later stage. Moreover, we contribute a large-scale layout dataset, named LayoutSAM, which includes 2.7 million image-text pairs and 10.7 million entities. Each entity is annotated with a bounding box and a detailed description. We further construct the LayoutSAM-Eval benchmark as a comprehensive tool for evaluating the L2I generation quality. Finally, we introduce the Layout Designer, which taps into the potential of large language models in layout planning, transforming them into experts in layout generation and optimization. These components form CreatiLayout -- a systematic solution that integrates the layout model, dataset, and planner for creative layout-to-image generation.
Hui Zhang 0090, Dexiang Hong, Zuxuan Wu, Yu-Gang Jiang 0001
ICCV1
2025 Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic Control
abstract
Despite recent advances in diffusion models, top-tier text-to-image (T2I) models still struggle to achieve precise spatial layout control, *i.e.* accurately generating entities with specified attributes and locations. Segmentation-mask-to-image (S2I) generation has emerged as a promising solution by incorporating pixel-level spatial guidance and regional text prompts. However, existing S2I methods fail to simultaneously ensure semantic consistency and shape consistency. To address these challenges, we propose Seg2Any, a novel S2I framework built upon advanced multimodal diffusion transformers (*e.g.* FLUX). First, to achieve both semantic and shape consistency, we decouple segmentation mask conditions into regional semantic and high-frequency shape components. The regional semantic condition is introduced by a Semantic Alignment Attention Mask, ensuring that generated entities adhere to their assigned text prompts. The high-frequency shape condition, representing entity boundaries, is encoded as an Entity Contour Map and then introduced as an additional modality via multi-modal attention to guide image spatial structure. Second, to prevent attribute leakage across entities in multi-entity scenarios, we introduce an Attribute Isolation Attention Mask mechanism, which constrains each entity’s image tokens to attend exclusively to themselves during image self-attention. To support open-set S2I generation, we construct SACap-1M, a large-scale dataset containing 1 million images with 5.9 million segmented entities and detailed regional captions, along with a SACap-Eval benchmark for comprehensive S2I evaluation. Extensive experiments demonstrate that Seg2Any achieves state-of-the-art performance on both open-set and closed-set S2I benchmarks, particularly in fine-grained spatial and attribute control of entities.
Danfeng Li, Hui Zhang 0090, Zuxuan Wu
NeurIPS2
2025 DiffusionAD: Norm-Guided One-Step Denoising Diffusion for Anomaly Detection
abstract
Anomaly detection has garnered extensive applications in real industrial manufacturing due to its remarkable effectiveness and efficiency. However, previous generative-based models have been limited by suboptimal reconstruction quality, hampering their overall performance. We introduce DiffusionAD, a novel anomaly detection pipeline comprising a reconstruction sub-network and a segmentation sub-network. A fundamental enhancement lies in our reformulation of the reconstruction process using a diffusion model into a noise-to-norm paradigm. Here, the anomalous region loses its distinctive features after being disturbed by Gaussian noise and is subsequently reconstructed into an anomaly-free one. Afterward, the segmentation sub-network predicts pixel-level anomaly scores based on the similarities and discrepancies between the input image and its anomaly-free reconstruction. Additionally, given the substantial decrease in inference speed due to the iterative denoising nature of diffusion models, we revisit the denoising process and introduce a rapid one-step denoising paradigm. This paradigm achieves hundreds of times acceleration while preserving comparable reconstruction quality. Furthermore, considering the diversity in the manifestation of anomalies, we propose a norm-guided paradigm to integrate the benefits of multiple noise scales, enhancing the fidelity of reconstructions. Comprehensive evaluations on four standard and challenging benchmarks reveal that DiffusionAD outperforms current state-of-the-art approaches and achieves comparable inference speed, demonstrating the effectiveness and broad applicability of the proposed pipeline.
Hui Zhang 0090, Zheng Wang 0059, Dan Zeng 0001, Zuxuan Wu, Yu-Gang Jiang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Prototypical Residual Networks for Anomaly Detection and Localization
abstract
Anomaly detection and localization are widely used in industrial manufacturing for its efficiency and effectiveness. Anomalies are rare and hard to collect and supervised models easily over-fit to these seen anomalies with a handful of abnormal samples, producing unsatisfactory performance. On the other hand, anomalies are typically subtle, hard to discern, and of various appearance, making it difficult to detect anomalies and let alone locate anomalous regions. To address these issues, we propose a framework called Prototypical Residual Network (PRN), which learns feature residuals of varying scales and sizes between anomalous and normal patterns to accurately reconstruct the segmentation maps of anomalous regions. PRN mainly consists of two parts: multi-scale prototypes that explicitly represent the residual features of anomalies to normal patterns; a multisize self-attention mechanism that enables variable-sized anomalous feature learning. Besides, we present a variety of anomaly generation strategies that consider both seen and unseen appearance variance to enlarge and diversify anomalies. Extensive experiments on the challenging and widely used MVTec AD benchmark show that PRN outperforms current state-of-the-art unsupervised and supervised methods. We further report SOTA results on three additional datasets to demonstrate the effectiveness and generalizability of PRN.
Hui Zhang 0090, Zuxuan Wu, Zheng Wang 0059, Zhineng Chen, Yu-Gang Jiang 0001
CVPR1