Chaofei Wang

dblp:35/10243 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
14since 2021 · last 2024
0000-0002-3678-691XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021
YearPublicationVenuePosition
2024 Smooth Diffusion: Crafting Smooth Latent Spaces in Diffusion Models
abstract
Recently, diffusion models have made remarkable progress in text-to-image (T2I) generation, synthesizing images with highfidelity and diverse contents. Despite this advancement, latent space smoothness within diffusion models remains largely unexplored. Smooth latent spaces en-sure that a perturbation on an input latent corresponds to a steady change in the output image. This property proves beneficial in downstream tasks, including image interpolation, inversion, and editing. In this work, we expose the non-smoothness of diffusion latent spaces by observing noticeable visual fluctuations resulting from minor latent variations. To tackle this issue, we propose Smooth Diffusion, a new category of diffusion models that can be simultaneously high-performing and smooth. Specifically, we introduce Step-wise Variation Regularization to enforce the proportion between the variations of an arbitrary input latent and that of the output image is a constant at any diffusion training step. In addition, we devise an interpolation standard deviation (ISTD) metric to effectively assess the latent space smoothness of a diffusion model. Extensive quantitative and qualitative experiments demonstrate that Smooth Diffusion stands out as a more desirable solution not only in T2I generation but also across various downstream tasks. Smooth Diffusion is implemented as a plug-and-play Smooth-LoRA to work with various community models. Code is available at https://github.com/SHI-Labs/Smooth-Diffusion.
Xingqian Xu, Yifan Pu, Zanlin Ni, Chaofei Wang, Manushree Vasu, Shiji Song, Gao Huang 0001, Humphrey Shi
CVPR5
2024 ADIFT: Zero-Shot Generative Model Adaption Via Adaptive Domain-Invariant Feature Transfer
abstract
CLIP-guided zero-shot image generative model adaption methods only require textual domain labels without any target domain images, but there are some dilemmas remain unsolved, such as identity degradation and pattern overfitting. To address these issues, an adaptive domain-invariant feature transfer (ADIFT) method is proposed. It makes the target domain generator learn domain-invariant features from the source domain generator but learn domain-variant features from the CLIP space. We first introduce a local self-similarity map to represent and preserve the image identity features, and then add a parameter learnable point-wise gate module on the alignment path of the local self-similarity maps to transfer cross-domain features adaptively. Qualitative and quantitative experimental results validate that the proposed ADIFT solves the problems of identity degradation and pattern over-fitting effectively.
Chaofei Wang, Xiangan Zhao, Jiayu Xiao, Guotong Geng
ICASSP1
2024 Latency-Aware Unified Dynamic Networks for Efficient Image Recognition
abstract
Dynamic networks have become a pivotal area of study in deep learning due to their ability to selectively activate computing units (such as layers or channels) or dynamically allocate computation to information-rich regions. This capability significantly curtails unnecessary computations, adapting to varying inputs. Despite these advantages, the practical efficiency of dynamic models often falls short of theoretical computation. This discrepancy arises from three primary challenges: 1) a lack of a unified framework across different dynamic inference paradigms due to the fragmented research landscape; 2) an excessive focus on algorithm design at the expense of scheduling strategies, which are essential for optimizing resource utilization on hardware; and 3) the complexity of latency evaluation, since most current libraries cater to static operators. To tackle these issues, we introduce Latency-Aware Unified Dynamic Networks (LAUDNet), a general framework that integrates three fundamental dynamic paradigms-spatially-adaptive computation, layer skipping, and channel skipping-into a single unified formulation. LAUDNet not only refines algorithmic design but also enhances scheduling optimization with the aid of a latency predictor. This predictor efficiently and accurately predicts the inference latency of dynamic operators on specific hardware setups. Our empirical assessments across multiple vision tasks-image classification, object detection, and instance segmentation-confirm that LAUDNet significantly bridges the gap between theoretical and practical efficiency. For instance, LAUDNet cuts down the practical latency of its static counterpart, ResNet-101, by over 50% on hardware platforms like V100, RTX 3090, and TX2 GPUs. Additionally, LAUDNet excels in the accuracy-efficiency trade-off compared to other methods.
Yizeng Han, Zhihang Yuan, Yifan Pu, Chaofei Wang, Shiji Song, Gao Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 A unified framework for convolution-based graph neural networks
Xuran Pan, Xiaoyan Han, Chaofei Wang, Shiji Song, Gao Huang 0001, Cheng Wu 0002
Pattern Recognit.3
2024 FaceCLIP: Facial Image-to-Video Translation via a Brief Text Description
abstract
The existing image-to-video translation methods generally follow a frame-by-frame generative paradigm, while extracting the temporal information from a reference video or an audio stream. Inspired by the recent success in text-guided image generation, we explore a more challenging but promising task, Text-guided Image-to-Video (TI2V) translation. Given an image and a brief text description as input, TI2V aims to generate a facial expression video following the image and text. To this end, we first propose an automatic video captioning pipeline to generate dense textual descriptions for facial video datasets, using both expression labels and action units. These dense textual descriptions provide precise semantic guidance for TI2V learning. Then we design and train an efficient framework, FaceCLIP, on these datasets to deal with the TI2V translation task. FaceCLIP adopts a video autoencoder to model the temporal information of training videos, and a pretrained CLIP model to embed the video frames and the text description. We design a reconstruction loss and an embedding alignment loss to train the autoencoder to obtain the text-guided video generative ability. Recognizing that expressions are closely tied to facial landmark motions, the reconstruction loss is applied to facial landmarks rather than each video frame, significantly enhancing training efficiency. We compare FaceCLIP with several potential baseline methods, and extensively evaluate the performance using multiple metrics. Both qualitative and quantitative results validate the superiority of FaceCLIP in terms of both visual quality and expression-text consistency. Moreover, the unique ability of FaceCLIP to generate videos based on abstract texts demonstrates its stronger generalization capability.
Hayk Manukyan 0001, Chaofei Wang, Levon Khachatryan, Shant Navasardyan, Shiji Song, Humphrey Shi, Gao Huang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Zero-Shot Generative Model Adaptation via Image-Specific Prompt Learning
abstract
Recently, CLIP-guided image synthesis has shown appealing performance on adapting a pre-trained source-domain generator to an unseen target domain. It does not require any target-domain samples but only the textual domain labels. The training is highly efficient, e.g., a few minutes. However, existing methods still have some limitations in the quality of generated images and may suffer from the mode collapse issue. A key reason is that a fixed adaptation direction is applied for all cross-domain image pairs, which leads to identical supervision signals. To address this issue, we propose an Image-specific Prompt Learning (IPL) method, which learns specific prompt vectors for each source-domain image. This produces a more precise adaptation direction for every cross-domain image pair, endowing the target-domain generator with greatly enhanced flexibility. Qualitative and quantitative evaluations on various domains demonstrate that IPL effectively improves the quality and diversity of synthesized images and alleviates the mode collapse. Moreover, IPL is independent of the structure of the generative model, such as generative adversarial networks or diffusion models. Code is available at https://github.com/Picsart-AI-Research/IPL-Zero-Shot-Generative-Model-Adaptation.
Chaofei Wang, Eric J. Zhang, Kai Wang 0058, Xingqian Xu, Shiji Song, Humphrey Shi, Gao Huang 0001
CVPR2
2022 Few Shot Generative Model Adaption via Relaxed Spatial Structural Alignment
abstract
Training a generative adversarial network (GAN) with limited data has been a challenging task. A feasible solution is to start with a GAN well-trained on a large scale source domain and adapt it to the target domain with a few samples, termed as few shot generative model adaption. However, existing methods are prone to model overfitting and collapse in extremely few shot setting (less than 10). To solve this problem, we propose a relaxed spatial structural alignment (RSSA) method to calibrate the target generative models during the adaption. We design a cross-domain spatial structural consistency loss comprising the self-correlation and disturbance correlation consistency loss. It helps align the spatial structural information between the synthesis image pairs of the source and target domains. To relax the cross-domain alignment, we compress the original latent space of generative models to a subspace. Image pairs generated from the subspace are pulled closer. Qualitative and quantitative experiments show that our method consistently surpasses the state-of-the-art methods in few shot setting. Our source code: https://github.com/StevenShaw1999/RSSA.
Jiayu Xiao, Liang Li 0003, Chaofei Wang, Zhengjun Zha, Qingming Huang
CVPR3
2022 Learning to Weight Samples for Dynamic Early-Exiting Networks
Yizeng Han, Yifan Pu, Zihang Lai, Chaofei Wang, Shiji Song, Junfen Cao, Chao Deng 0002, Gao Huang 0001
ECCV (11)4
2022 Learn From the Past: Experience Ensemble Knowledge Distillation
abstract
Traditional knowledge distillation transfers "dark knowledge" of a pre-trained teacher network to a student network, and ignores the knowledge in the training process of the teacher, which we call teacher’s experience. However, in realistic educational scenarios, learning experience is often more important than learning results. In this work, we propose a novel knowledge distillation method by integrating the teacher’s experience for knowledge transfer, named experience ensemble knowledge distillation (EEKD). We save a moderate number of intermediate models from the training process of the teacher model uniformly, and then integrate the knowledge of these intermediate models by ensemble technique. A self-attention module is used to adaptively assign weights to different intermediate models in the process of knowledge transfer. Three principles of constructing EEKD on the quality, weights and number of intermediate models are explored. A surprising conclusion is found that strong ensemble teachers do not necessarily produce strong students. The experimental results on CIFAR-100 and ImageNet show that EEKD outperforms the mainstream knowledge distillation methods and achieves the state-of-the-art. In particular, EEKD even surpasses the standard ensemble distillation on the premise of saving training cost.
Chaofei Wang, Shiji Song, Gao Huang 0001
ICPR1
2022 Efficient Knowledge Distillation from Model Checkpoints
abstract
Knowledge distillation is an effective approach to learn compact models (students) with the supervision of large and strong models (teachers). As empirically there exists a strong correlation between the performance of teacher and student models, it is commonly believed that a high performing teacher is preferred. Consequently, practitioners tend to use a well trained network or an ensemble of them as the teacher. In this paper, we observe that an intermediate model, i.e., a checkpoint in the middle of the training procedure, often serves as a better teacher compared to the fully converged model, although the former has much lower accuracy. More surprisingly, a weak snapshot ensemble of several intermediate models from a same training trajectory can outperform a strong ensemble of independently trained and fully converged models, when they are used as teachers. We show that this phenomenon can be partially explained by the information bottleneck principle: the feature representations of intermediate models can have higher mutual information regarding the input, and thus contain more ``dark knowledge'' for effective distillation. We further propose an optimal intermediate teacher selection algorithm based on maximizing the total task-related mutual information. Experiments verify its effectiveness and applicability. Our code is available at https://github.com/LeapLabTHU/CheckpointKD.
Chaofei Wang, Qisen Yang, Rui Huang 0012, Shiji Song, Gao Huang 0001
NeurIPS1
2022 TC3KD: Knowledge distillation via teacher-student cooperative curriculum customization
Chaofei Wang, Gao Huang 0001, Shiji Song
Neurocomputing1
2021 Towards Learning Spatially Discriminative Feature Representations
abstract
The backbone of traditional CNN classifier is generally considered as a feature extractor, followed by a linear layer which performs the classification. We propose a novel loss function, termed as CAM-loss, to constrain the embedded feature maps with the class activation maps (CAMs) which indicate the spatially discriminative regions of an image for particular categories. CAM-loss drives the backbone to express the features of target category and suppress the features of non-target categories or background, so as to obtain more discriminative feature representations. It can be simply applied in any CNN architecture with neglectable additional parameters and calculations. Experimental results show that CAM-loss is applicable to a variety of network structures and can be combined with mainstream regularization methods to improve the performance of image classification. The strong generalization ability of CAMloss is validated in the transfer learning and few shot learning tasks. Based on CAM-loss, we also propose a novel CAAM-CAM matching knowledge distillation method. This method directly uses the CAM generated by the teacher network to supervise the CAAM generated by the student network, which effectively improves the accuracy and convergence rate of the student network.
Chaofei Wang, Jiayu Xiao, Yizeng Han, Qisen Yang, Shiji Song, Gao Huang 0001
ICCV1
2021 Point-Voting based Point Cloud Geometry Compression
abstract
The Geometry-based Point Cloud Compression (G-PCC) proposed by the Moving Picture Experts Group (MPEG) is the state-of-art point cloud compression algorithm. It provides an efficient lossy geometry compression technique called triangle soup (Trisoup) for static point clouds. Based on the pruned octree structure, Trisoup provides a local surface model consisting of multiple triangles and compresses vertices of the triangles instead of directly compressing the positions of the original points. Accordingly, we propose a point-voting based method to improve the triangle-construction within each leaf node. This new method leverages the node-based points distribution for more precise vertices determination, which better fits the local surface. Experimental results demonstrate the effectiveness of our point-voting based method for both objective and subjective quality evaluation.
Chaofei Wang, Wenjie Zhu 0004, Yingzhan Xu, Yiling Xu, Le Yang 0001
MMSP1
2021 Fine-grained few shot learning with foreground object transformation
Chaofei Wang, Shiji Song, Qisen Yang, Xiang Li 0009, Gao Huang 0001
Neurocomputing1