Chi Wang 0004

dblp:09/404-4 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-1686-2687ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization
abstract
Garment-centric fashion image generation aims to synthesize realistic and controllable human models dressing a given garment, which has attracted growing interest due to its practical applications in e-commerce. The key challenges of the task lie in two aspects: (1) faithfully preserving the garment details, and (2) gaining fine-grained controllability over the model's appearance. Existing methods typically require performing garment deformation in the generation process, which often leads to garment texture distortions. Also, they fail to control the fine-grained attributes of the generated models, due to the lack of specifically designed mechanisms. To address these issues, we propose FashionMAC, a novel diffusion-based deformation-free framework that achieves high-quality and controllable fashion showcase image generation. The core idea of our framework is to eliminate the need for performing garment deformation and directly outpaint the garment segmented from a dressed person, which enables faithful preservation of the intricate garment details. Moreover, we propose a novel region-adaptive decoupled attention (RADA) mechanism along with a chained mask injection strategy to achieve fine-grained appearance controllability over the synthesized human models. Specifically, RADA adaptively predicts the generated regions for each fine-grained text attribute and enforces the text attribute to focus on the predicted regions by a chained mask injection strategy, significantly enhancing the visual fidelity and the controllability. Extensive experiments validate the superior performance of our framework compared to existing state-of-the-art methods.
Jinxiao Li, Jingnan Wang, Zhiwen Zuo, Jianfeng Dong, Wei Li 0111, Chi Wang 0004, Weiwei Xu 0003, Xun Wang 0007
AAAI7
2026 VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
abstract
Existing for audio- and pose-driven human animation methods often struggle with stiff head movements and blurry hands, primarily due to the weak correlation between audio and head movements and the structural complexity of hands. To address these issues, we propose VividAnimator, an end-to-end framework for generating high-quality, half-body human animations driven by audio and sparse hand pose conditions. Our framework introduces three key innovations. First, to overcome the instability and high cost of online codebook training, we pre-train a Hand Clarity Codebook (HCC) that encodes rich, high-fidelity hand texture priors, significantly mitigating hand degradation. Second, we design a Dual-Stream Audio-Aware Module (DSAA) to model lip synchronization and natural head pose dynamics separately while enabling interaction. Third, we introduce a Pose Calibration Trick (PCT) that refines and aligns pose conditions by relaxing rigid constraints, ensuring smooth and natural gesture transitions. Extensive experiments demonstrate that Vivid Animator achieves state-of-the-art performance, producing videos with superior hand detail, gesture realism, and identity consistency, validated by both quantitative metrics and qualitative evaluations.
Donglin Huang, Yongyuan Li, Tianhang Liu, Junming Huang 0002, Xiaoda Yang, Chi Wang 0004, Weiwei Xu 0003
WACV6
2026 From architecture to evaluation: A comprehensive review of video generation techniques
abstract
The rapid developments of artificial intelligence have significantly impacted daily life and content production modes. In the field of video generation, researchers are now exploring this emerging technique with innovative approaches, aiming to produce videos of higher quality, longer duration, and greater diversity. Currently, numerous video generation algorithms have been developed using different architecture designs. Unlike image generation, video generation requires maintaining consistency across both spatial and temporal dimensions while ensuring aesthetic quality and dynamic coherence, making it a more challenging task. In this survey, we provide a systematic review of existing video generation methods, tracing their evolution across different architectural paradigms. We further categorize recent models by their control conditions (e.g., text-to-video, image-to-video, multi-modal guidance) and summarize their unique theoretical foundations, architectural designs, and algorithmic innovations. In the meantime, we review the commonly used video datasets and analyze their applicability to different tasks. We also present evaluations of representative models to offer a more comprehensive perspective. Our goal is to provide a clear and concise overview of these algorithms, offering insights to support future breakthroughs in video generation.
Chi Wang 0004, Guojun Lei, Weiwei Xu 0003
Virtual Real. Intell. Hardw.2
2025 AnimateAnything: Consistent and Controllable Animation for Video Generation
abstract
We present a unified controllable video generation approach AnimateAnything that facilitates precise and consistent video manipulation across various conditions, including camera trajectories, text prompts, and user motion annotations. Specifically, we carefully design a multiscale control feature fusion network to construct a common motion representation for different conditions. It explicitly converts all control information into frame-by-frame optical flows. Then we incorporate the optical flows as motion priors to guide the final video generation. In addition, to reduce the flickering issues caused by large-scale motion, we propose a frequency-based stabilization module. It can enhance temporal coherence by ensuring the video's frequency domain consistency. Experiments demonstrate that our method outperforms the state-of-the-art approaches. For more details and videos, please refer to the anonymous webpage: https://yu-shaonian.github.io/Animate_Anything/.
Guojun Lei, Chi Wang 0004, Hong Li 0016, Weiwei Xu 0003
CVPR2
2025 Detail-Preserving Latent Diffusion for Stable Shadow Removal
abstract
Achieving high-quality shadow removal with strong generalizability is challenging in scenes with complex global illumination. Due to the limited diversity in shadow removal datasets, current methods are prone to overfitting training data, often leading to reduced performance on unseen cases. To address this, we leverage the rich visual priors of a pre-trained Stable Diffusion (SD) model and propose a two-stage fine-tuning pipeline to adapt the SD model for stable and efficient shadow removal. In the first stage, we fix the VAE and fine-tune the denoiser in latent space, which yields substantial shadow removal but may lose some high-frequency details. To resolve this, we introduce a second stage, called the detail injection stage. This stage selectively extracts features from the VAE encoder to modulate the decoder, injecting fine details into the final results. Experimental results show that our method outperforms state-of-the-art shadow removal techniques. The cross-dataset evaluation further demonstrates that our method generalizes effectively to unseen data, enhancing the applicability of shadow removal methods.
Jiamin Xu, Chi Wang 0004, Renshu Gu, Weiwei Xu 0003, Gang Xu 0001
CVPR4
2025 MotionFlow: Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation
abstract
Generating videos guided by camera trajectories poses significant challenges in achieving consistency and generalizability, particularly when active objects are present. Existing approaches often attempt to learn object motions separately, which may lead to confusion regarding the relative motion between the camera and the objects. To address this challenge, we propose a novel approach that integrates both camera trajectory and object semantics and converts them into the reference motion of the corresponding pixels. Utilizing a stable diffusion network, we effectively extract reference motion maps in relation to the specified camera trajectory. These maps, along with an extracted semantic object prior, are then fed into an image-to-video network to generate the desired video that can accurately follow the designated camera trajectory while maintaining the consistency of object. Extensive experiments verify that our model outperforms SOTA methods by a large margin. Project Page: https://yu-shaonian.github.io/MotionFlow/
Guojun Lei, Chi Wang 0004, Hong Li 0016, Weiwei Xu 0003
ICME2
2025 UniTransfer: Video Concept Transfer via Progressive Spatio-Temporal Decomposition
abstract
Recent advancements in video generation models have enabled the creation of diverse and realistic videos, with promising applications in advertising and film production. However, as one of the essential tasks of video generation models, video concept transfer remains significantly challenging. Existing methods generally model video as an entirety, leading to limited flexibility and precision when solely editing specific regions or concepts. To mitigate this dilemma, we propose a novel architecture UniTransfer, which introduces both spatial and diffusion timestep decomposition in a progressive paradigm, achieving precise and controllable video concept transfer. Specifically, in terms of spatial decomposition, we decouple videos into three key components: the foreground subject, the background, and the motion flow. Building upon this decomposed formulation, we further introduce a dual-to-single-stream DiT-based architecture for supporting fine-grained control over different components in the videos. We also introduce a self-supervised pretraining strategy based on random masking to enhance the decomposed representation learning from large-scale unlabeled video data. Inspired by the Chain-of-Thought reasoning paradigm, we further revisit the denoising diffusion process and propose a Chain-of-Prompt (CoP) mechanism to achieve the timestep decomposition. We decompose the denoising process into three stages of different granularity and leverage large language models (LLMs) for stage-specific instructions to guide the generation progressively. We also curate an animal-centric video dataset called OpenAnimal to facilitate the advancement and benchmarking of research in video concept transfer. Extensive experiments demonstrate that our method achieves high-quality and controllable video concept transfer across diverse reference images and scenes, surpassing existing baselines in both visual fidelity and editability.
Guojun Lei, Tianhang Liu, Hong Li 0016, Chi Wang 0004, Weiwei Xu 0003
NeurIPS6
2024 Local Gaussian Density Mixtures for Unstructured Lumigraph Rendering
abstract
PSNR 29.47 PSNR 28.80 PSNR 27.28 PSNR 30.84 PSNR 26.69 PSNR 26.
Xiuchao Wu, Jiamin Xu, Chi Wang 0004, Yifan Peng 0001, Qixing Huang, James Tompkin 0001, Weiwei Xu 0003
SIGGRAPH Asia3
2023 CF-Font: Content Fusion for Few-Shot Font Generation
abstract
Content and style disentanglement is an effective way to achieve few-shot font generation. It allows to transfer the style of the font image in a source domain to the style defined with a few reference images in a target domain. However, the content feature extracted using a representative font might not be optimal. In light of this, we propose a content fusion module (CFM) to project the content feature into a linear space defined by the content features of basis fonts, which can take the variation of content features caused by different fonts into consideration. Our method also allows to optimize the style representation vector of reference images through a lightweight iterative style-vector refinement (ISR) strategy. Moreover, we treat the 1D projection of a character image as a probability distribution and leverage the distance between two distributions as the reconstruction loss (namely projected character loss, PCL). Compared to L2 or L1 reconstruction loss, the distribution distance pays more attention to the global shape of characters. We have evaluated our method on a dataset of 300 fonts with 6.5k characters each. Experimental results verify that our method outperforms existing state-of-the-art few-shot font generation methods by a large margin. The source code can be found at https://github.com/wangchi95/CF-Font.
Chi Wang 0004, Tiezheng Ge, Yuning Jiang 0001, Hujun Bao, Weiwei Xu 0003
CVPR1
2023 Weakly Supervised Image Matting via Patch Clustering
Yunke Zhang, Chi Wang 0004, Hujun Bao, Weiwei Xu 0003
ICIG (1)2
2022 Active Boundary Loss for Semantic Segmentation
abstract
This paper proposes a novel active boundary loss for semantic segmentation. It can progressively encourage the alignment between predicted boundaries and ground-truth boundaries during end-to-end training, which is not explicitly enforced in commonly used cross-entropy loss. Based on the predicted boundaries detected from the segmentation results using current network parameters, we formulate the boundary alignment problem as a differentiable direction vector prediction problem to guide the movement of predicted boundaries in each iteration. Our loss is model-agnostic and can be plugged in to the training of segmentation networks to improve the boundary details. Experimental results show that training with the active boundary loss can effectively improve the boundary F-score and mean Intersection-over-Union on challenging image and video object segmentation datasets.
Chi Wang 0004, Yunke Zhang, Miaomiao Cui, Peiran Ren, Yin Yang 0002, Xuansong Xie, Xian-Sheng Hua 0001, Hujun Bao, Weiwei Xu 0003
AAAI1
2021 Attention-guided Temporally Coherent Video Object Matting
abstract
This paper proposes a novel deep learning-based video object matting method that can achieve temporally coherent matting results. Its key component is an attention-based temporal aggregation module that maximizes image matting networks' strength for video matting networks. This module computes temporal correlations for pixels adjacent to each other along the time axis in feature space, which is robust against motion noises. We also design a novel loss term to train the attention weights, which drastically boosts the video matting performance. Besides, we show how to effectively solve the trimap generation problem by fine-tuning a state-of-the-art video object segmentation network with a sparse set of user-annotated keyframes. To facilitate video matting and trimap generation networks' training, we construct a large-scale video matting dataset with 80 training and 28 validation foreground video clips with ground-truth alpha mattes. Experimental results show that our method can generate high-quality alpha mattes for various videos featuring appearance change, occlusion, and fast motion. Our code and dataset can be found at: https://github.com/yunkezhang/TCVOM
Yunke Zhang, Chi Wang 0004, Miaomiao Cui, Peiran Ren, Xuansong Xie, Xian-Sheng Hua 0001, Hujun Bao, Qixing Huang, Weiwei Xu 0003
ACM Multimedia2
2019 Liver tumor segmentation in CT volumes using an adversarial densely connected network
abstract
BACKGROUND: Malignant liver tumor is one of the main causes of human death. In order to help physician better diagnose and make personalized treatment schemes, in clinical practice, it is often necessary to segment and visualize the liver tumor from abdominal computed tomography images. Due to the large number of slices in computed tomography sequence, developing an automatic and reliable segmentation method is very favored by physicians. However, because of the noise existed in the scan sequence and the similar pixel intensity of liver tumors with their surrounding tissues, besides, the size, position and shape of tumors also vary from one patient to another, automatic liver tumor segmentation is still a difficult task. RESULTS: We perform the proposed algorithm to the Liver Tumor Segmentation Challenge dataset and evaluate the segmentation results. Experimental results reveal that the proposed method achieved an average Dice score of 68.4% for tumor segmentation by using the designed network, and ASD, MSD, VOE and RVD improved from 27.8 to 21, 147 to 124, 0.52 to 0.46 and 0.69 to 0.73, respectively after performing adversarial training strategy, which proved the effectiveness of the proposed method. CONCLUSIONS: The testing results show that the proposed method achieves improved performance, which corroborated the adversarial training based strategy can achieve more accurate and robustness results on liver tumor segmentation task.
Lei Chen 0073, Hong Song 0003, Chi Wang 0004, Yutao Cui, Jian Yang 0009, Xiaohua Hu 0001, Le Zhang 0004
BMC Bioinform.3
2018 Automatic Liver Segmentation Using Multi-plane Integrated Fully Convolutional Neural Networks
Chi Wang 0004, Hong Song 0003, Lei Chen 0073, Qiang Li 0049, Jian Yang 0009, Xiaohua Hu 0001, Le Zhang 0004
BIBM1