Yiqiang Yan

dblp:354/8809 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2026
0009-0008-7709-0602ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ConsistentID: Portrait Generation With Multimodal Fine-Grained Identity Preserving
abstract
Diffusion-based technologies have made significant strides, particularly in personalized and customized facial generation. However, existing methods struggle to achieve high-fidelity and detailed identity (ID) consistency. This is mainly due to two challenges: insufficient fine-grained control over specific facial areas and the absence of a comprehensive strategy for ID preservation that accounts for both intricate facial details and the overall facial structure. To address these limitations, we introduce ConsistentID, an innovative method crafted for diverse identity-preserving portrait generation under fine-grained multimodal facial prompts, utilizing only a single reference image. ConsistentID comprises two core components: a multimodal facial prompt generator and an ID-preservation network. The facial prompt generator combines localized facial features, facial feature descriptions, and overall facial descriptions to enhance the precision of facial detail reconstruction. The ID-preservation network, optimized with a facial attention localization strategy, ensures consistent identity preservation across facial regions. Together, these components leverage fine-grained multimodal identity information to improve identity preservation accuracy significantly. To drive ConsistentID's training, we propose a fine-grained portrait dataset, FGID, with over 500,000 facial images, offering greater diversity and comprehensiveness than existing public facial datasets. Experimental results substantiate that our ConsistentID achieves exceptional precision and diversity in personalized facial generation, surpassing existing methods in the MyStyle dataset. In addition, although ConsistentID introduces more multimodal ID information, it still maintains rapid inference speed during the generation process.
Jiehui Huang, Wenhui Song, Zheng Chong, Zhenchao Tang, Yuhao Cheng, Long Chen 0005, Yiqiang Yan, Shengcai Liao, Xiaodan Liang
IEEE Trans. Pattern Anal. Mach. Intell.10
2026 ArtCrafter: Text-Image Aligning Artistic Attribute Transfer via Embedding Reframing
abstract
Recent years have witnessed significant advancements in text-guided style transfer, primarily attributed to innovations in diffusion models. These models excel in conditional guidance, utilizing text or images to direct the sampling process. Traditional style transfer focuses on low-level visual features, such as brushstroke textures and color distributions, and appears more like applying an artistic filter to an image. Artistic attribute transfer, however, transcends the limitations of traditional style transfer by achieving the transfer of visual concepts from color and brushstrokes to high level aesthetic attributes such as composition, pose, and key semantic elements, resulting in more natural outcomes. Therefore, we propose an innovative text-to-image artistic attribute transfer framework named ArtCrafter. Specifically, we introduce an attention-based style extraction module, meticulously engineered to capture the subtle artistic attribute elements within an image. This module features a multi-layer architecture that leverages the capabilities of perceiver attention mechanisms to integrate fine-grained information. Additionally, we present a novel text-image aligning augmentation component that adeptly balances control over both modalities, enabling the model to efficiently map image and text embeddings into a shared feature space. We achieve this through attention operations that enable smooth information flow between modalities. Lastly, we incorporate an explicit modulation that seamlessly blends multimodal enhanced embeddings with original embeddings through an embedding reframing design, empowering the model to generate diverse outputs. Extensive experiments demonstrate that ArtCrafter yields impressive results in visual stylization, exhibiting exceptional levels of artistic attribute intensity, controllability, and diversity.
Nisha Huang, Kaer Huang, Yifan Pu, Jiangshan Wang, Yiqiang Yan, Xiu Li 0001, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.6
2025 GDrag: Towards General-Purpose Interactive Editing with Anti-ambiguity Point Diffusion
abstract
Recent interactive point-based image manipulation methods have gained considerable attention for being user-friendly. However, these methods still face two types of ambiguity issues that can lead to unsatisfactory outcomes, namely, intention ambiguity which misinterprets the purposes of users, and content ambiguity where target image areas are distorted by distracting elements. To address these issues and achieve general-purpose manipulations, we propose a novel task-aware, training-free framework called GDrag. Specifically, GDrag defines a taxonomy of atomic manipulations, which can be parameterized and combined unitedly to represent complex manipulations, thereby reducing intention ambiguity. Furthermore, GDrag introduces two strategies to mitigate content ambiguity, including an anti-ambiguity dense trajectory calculation method (ADT) and a self-adaptive motion supervision method (SMS). Given an atomic manipulation, ADT converts the sparse user-defined handle points into a dense point set by selecting their semantic and geometric neighbors, and calculates the trajectory of the point set. Unlike previous motion supervision methods relying on a single global scale for low-rank adaption, SMS jointly optimizes point-wise adaption scales and latent feature biases. These two methods allow us to model fine-grained target contexts and generate precise trajectories. As a result, GDrag consistently produces precise and appealing results in different editing tasks. Extensive experiments on the challenging DragBench dataset demonstrate that GDrag outperforms state-of-the-art methods significantly. The code of GDrag will be released upon acceptance.
Xiaojian Lin, Yuhao Cheng, Yiqiang Yan, Xiaodan Liang
ICLR4
2025 ICE: Intercede Concept Erasure in Text-to-Image Diffusion Models
abstract
The success of diffusion models in text-to-image (T2I) generation has made it urgent to remove unwanted concepts, such as copyrighted, offensive, and unsafe ones, from pre-trained models in an accurate, timely, and cost-effective manner. However, limited by the inherent optimization perspective, existing methods have two major problems. Firstly, they overlook maintaining the global visual style during the erasure process, leading to significant style shifts. Secondly, excessive concept erasure causes relevant content to disappear or generates substitutes unrelated to the original object's attributes. Compared to other methods, our proposed ICE has unique advantages, as it can generate diverse visual features and achieve a balance between concept erasure and maintaining the semantic content of the target object. This is mainly achieved through our well-designed non-erasable features protector (NEFP) and augmented invariant constraints (AIC). Specifically, we enhance the protection of feature information by embedding an augmented orthogonal anchor concept matrix. Meanwhile, under controlled constraints, we introduce invariants into the embedding space to retain key semantics. This work specifically emphasizes the importance of focusing on feature expression and semantic protection in the concept erasure task for fully unleashing the performance of T2I models.
Yizhou Lin, Nisha Huang, Kaer Huang, Henglin Liu, Yiqiang Yan, Tong-Yee Lee, Xiu Li 0001
ACM Multimedia5
2025 LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
abstract
In this paper, we present LaVieID, a novel local a utoregressive vi deo diffusion framework designed to tackle the challenging id entity-preserving text-to-video task. The key idea of LaVieID is to mitigate the loss of identity information inherent in the stochastic global generation process of diffusion transformers (DiTs) from both spatial and temporal perspectives. Specifically, unlike the global and unstructured modeling of facial latent states in existing DiTs, LaVieID introduces a local router to explicitly represent latent states by weighted combinations of fine-grained local facial structures. This alleviates undesirable feature interference and encourages DiTs to capture distinctive facial characteristics. Furthermore, a temporal autoregressive module is integrated into LaVieID to refine denoised latent tokens before video decoding. This module divides latent tokens temporally into chunks, exploiting their long-range temporal dependencies to predict biases for rectifying tokens, thereby significantly enhancing inter-frame identity consistency. Consequently, LaVieID can generate high-fidelity personalized videos and achieve state-of-the-art performance. Our code and models are available at https://github.com/ssugarwh/LaVieID.
Wenhui Song, Jiehui Huang, Panwen Hu, Yuhao Cheng, Yiqiang Yan, Xiaodan Liang
ACM Multimedia7
2024 GarmentAligner: Text-to-Garment Generation via Retrieval-Augmented Multi-level Corrections
Zheng Chong, Xujie Zhang, Yuhao Cheng, Yiqiang Yan, Xiaodan Liang
ECCV (25)6
2024 Learning Interaction-aware 3D Gaussian Splatting for One-shot Hand Avatars
abstract
In this paper, we propose to create animatable avatars for interacting hands with 3D Gaussian Splatting (GS) and single-image inputs. Existing GS-based methods designed for single subjects often yield unsatisfactory results due to limited input views, various hand poses, and occlusions. To address these challenges, we introduce a novel two-stage interaction-aware GS framework that exploits cross-subject hand priors and refines 3D Gaussians in interacting areas. Particularly, to handle hand variations, we disentangle the 3D presentation of hands into optimization-based identity maps and learning-based latent geometric features and neural texture maps. Learning-based features are captured by trained networks to provide reliable priors for poses, shapes, and textures, while optimization-based identity maps enable efficient one-shot fitting of out-of-distribution hands. Furthermore, we devise an interaction-aware attention module and a self-adaptive Gaussian refinement module. These modules enhance image rendering quality in areas with intra- and inter-hand interactions, overcoming the limitations of existing GS-based methods. Our proposed method is validated via extensive experiments on the large-scale InterHand2.6M dataset, and it significantly improves the state-of-the-art performance in image quality. Code and models will be released upon acceptance.
Wanquan Liu, Xiaodan Liang, Yiqiang Yan, Yuhao Cheng, Chenqiang Gao
NeurIPS5
2024 Divide-and-conquer-based RDO-free CU Partitioning for 8K Video Compression
abstract
8K (7689×4320) ultra-high definition (UHD) videos are growing popular with the improvement of human visual experience demand. Therefore, the compression of 8K UHD videos has become a top priority in the third-generation audio video coding standard (AVS3). However, as an important part of the coding standard promotion, the real-time hardware implementation for AVS3-based 8K UHD video intra coding is severely hindered, especially in the coding unit (CU) partition stage. To break through the limitation, this article proposes a divide-and-conquer-based rate-distortion-optimization-free (RDO-free) CU partitioning algorithm for efficient hardware implementation. Aimed at the complex CU partition in AVS3, we separately design a lightweight optimization for original partitioning rules to improve division efficiency and a decision tree-based RDO-free CU decision framework to eliminate the latency caused by the waiting for rate-distortion cost calculation in RDO strategy. Afterward, a divide-and-conquer-based hardware-friendly gradient difference calculating approach is devised to accelerate the learning feature extracting speed. To ensure that the proposed algorithm is sufficient to support the real-time CU partition for 8K videos, we also develop a hardware architecture based on FPGA. Experimental results illustrate that the software coding performance of our algorithm is significantly ahead of the efficient implementation uAVS3e for AVS3 and the reference software HM-16.20 for High Efficiency Video Coding (HEVC), even though there is 9.96% loss on BD-Rate Y. Considering its importance for the hardware implementation of AVS3-based 8K real-time encoder, the coding loss is acceptable. Moreover, the hardware simulation results on VU440 FPGA with Vivado 2019 show that our algorithm can support 61.12 frames per second (fps) CU partition for 8K UHD videos with only 0.00%, 0.00%, 1.01%, and 7.78% consumption of BRAM_18K, DSP48E, FF, and LUT, respectively. Additionally, with dual-path parallelism, 122.24 fps also can be implemented with controllable resource utilization, which achieves the state-of-the-art performance.
Wei Gao 0003, Siwei Ma 0001, Yiqiang Yan
ACM Trans. Multim. Comput. Commun. Appl.4
2023 SUR-Driven Video Coding Rate Control for Jointly Optimizing Perceptual Quality and Buffer Control
abstract
Rate control plays an important role in video coding and has attracted lots of attention from researchers. However, the problems of human visual experience and buffer stability still remain. For scenes with drastic motions, parts of distortions can be masked due to the limitation of the Human Visual System (HVS), while buffers tend to suffer more overflow and underflow cases from the fluctuating bits. In this paper, we propose a novel joint rate control scheme, which is composed of the proposed SUR-based perception modeling and the proposed SUR-based Perception-Buffer Rate Control (PBRC), for HEVC to maximize human visual perception quality while preventing the underflow and overflow of buffers. First of all, to effectively model human visual quality, we introduce the perception-related Satisfied-User-Ratio (SUR) metric into the rate control process. Secondly, a time-efficient video quality prediction method called Fast Visual Multimethod Assessment Fusion (VMAF) Quality Prediction (FVQP) is designed for the generation of SUR curves within an affordable computational complexity. Thirdly, a dual-objective optimization framework is established. By jointly conducting perception modeling and PBRC, we can flexibly adjust the optimization priority between human visual quality and buffer stability, and thus the quality of achieved reconstructed videos can be effectively improved because of the decrease in frame skipping. Experimental results demonstrate that the proposed joint rate control scheme improves the human visual experience when considering frame skipping and more effectively stabilizes buffer stability than existing methods.
Zetao Yang, Wei Gao 0003, Ge Li 0002, Yiqiang Yan
IEEE Trans. Image Process.4