Xin Chen 0071

dblp:24/1518-71 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
7since 2021 · last 2024
0000-0003-1950-2468ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2024 To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy to Generate Unsafe Images ... For Now
Jinghan Jia, Xin Chen 0071, Aochuan Chen, Jiancheng Liu, Sijia Liu 0001
ECCV (57)3
2024 Fasor: A Fast Tensor Program Optimization Framework for Efficient DNN Deployment
abstract
With the growing importance of deploying deep neural networks (DNNs), there are increasing demands to improve both the efficiency and quality of tensor program optimization (TPO). TPO involves searching for possible program transformations for a given tensor program on target hardware to optimize its execution. TPO is challenging and expensive due to the exponential combinations of transformations and time-consuming on-device measurement of transformations. While prior research has primarily focused on the quality of TPO, i.e., generating high-performance tensor programs, there has been less emphasis on the efficiency of TPO, i.e., optimizing tensor programs with low optimization time overhead.
Hanxian Huang, Xin Chen 0071, Jishen Zhao
ICS2
2024 Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models
abstract
Diffusion models (DMs) have achieved remarkable success in text-to-image generation, but they also pose safety risks, such as the potential generation of harmful content and copyright violations. The techniques of machine unlearning, also known as concept erasing, have been developed to address these risks. However, these techniques remain vulnerable to adversarial prompt attacks, which can prompt DMs post-unlearning to regenerate undesired images containing concepts (such as nudity) meant to be erased. This work aims to enhance the robustness of concept erasing by integrating the principle of adversarial training (AT) into machine unlearning, resulting in the robust unlearning framework referred to as AdvUnlearn. However, achieving this effectively and efficiently is highly nontrivial. First, we find that a straightforward implementation of AT compromises DMs’ image generation quality post-unlearning. To address this, we develop a utility-retaining regularization on an additional retain set, optimizing the trade-off between concept erasure robustness and model utility in AdvUnlearn. Moreover, we identify the text encoder as a more suitable module for robustification compared to UNet, ensuring unlearning effectiveness. And the acquired text encoder can serve as a plug-and-play robust unlearner for various DM types. Empirically, we perform extensive experiments to demonstrate the robustness advantage of AdvUnlearn across various DM unlearning scenarios, including the erasure of nudity, objects, and style concepts. In addition to robustness, AdvUnlearn also achieves a balanced tradeoff with model utility. To our knowledge, this is the first work to systematically explore robust DM unlearning through AT, setting it apart from existing methods that overlook robustness in concept erasing. Codes are available at https://github.com/OPTML-Group/AdvUnlearn. Warning: This paper contains model outputs that may be offensive in nature.
Xin Chen 0071, Jinghan Jia, Chongyu Fan, Jiancheng Liu, Mingyi Hong 0001, Sijia Liu 0001
NeurIPS2
2023 Region-aware Knowledge Distillation for Efficient Image-to-Image Translation
Linfeng Zhang 0001, Xin Chen 0071, Runpei Dong, Kaisheng Ma
BMVC2
2023 Text-Visual Prompting for Efficient 2D Temporal Video Grounding
abstract
In this paper, we study the problem of temporal video grounding (TVG), which aims to predict the starting/ending time points of moments described by a text sentence within a long untrimmed video. Benefiting from fine-grained 3D visual features, the TVG techniques have achieved remarkable progress in recent years. However, the high complexity of 3D convolutional neural networks (CNNs) makes extracting dense 3D visual features time-consuming, which calls for intensive memory and computing resources. Towards efficient TVG, we propose a novel text-visual prompting (TVP) framework, which incorporates optimized perturbation patterns (that we call 'prompts') into both visual inputs and textual features of a TVG model. In sharp contrast to 3D CNNs, we show that TVP allows us to effectively co-train vision encoder and language encoder in a 2D TVG model and improves the performance of crossmodal feature fusion using only low-complexity sparse 2D visual features. Further, we propose a Temporal-Distance IoU (TDIoU) loss for efficient learning of TVG. Experiments on two benchmark datasets, Charades-STA and Activityblet Captions datasets, empirically show that the proposed TVP significantly boosts the performance of 2D TVG (e.g., 9.79% improvement on Charades-STA and 30.77% improvement on ActivityNet Captions) and achieves$5\times$inference acceleration over TVG using 3D visual features. Codes are available at Open.Intel.
Xin Chen 0071, Jinghan Jia, Sijia Liu 0001
CVPR2
2022 Wavelet Knowledge Distillation: Towards Efficient Image-to-Image Translation
abstract
Remarkable achievements have been attained with Generative Adversarial Networks (GANs) in image-to-image translation. However, due to a tremendous amount of parameters, state-of-the-art GANs usually suffer from low efficiency and bulky memory usage. To tackle this challenge, firstly, this paper investigates GANs performance from a frequency perspective. The results show that GANs, especially small GANs lack the ability to generate high-quality high frequency information. To address this problem, we propose a novel knowledge distillation method referred to as wavelet knowledge distillation. Instead of directly distilling the generated images of teachers, wavelet knowledge distillation first decomposes the images into different frequency bands with discrete wavelet transformation and then only distills the high frequency bands. As a result, the student GAN can pay more attention to its learning on high frequency bands. Experiments demonstrate that our method leads to 7.08× compression and 6.80× acceleration on CycleGAN with almost no performance drop. Additionally, we have studied the relation between discriminators and generators which shows that the compression of discriminators can promote the performance of compressed generators.
Linfeng Zhang 0001, Xin Chen 0071, Xiaobing Tu, Pengfei Wan 0001, Kaisheng Ma
CVPR2
2022 Contrastive Deep Supervision
Linfeng Zhang 0001, Xin Chen 0071, Runpei Dong, Kaisheng Ma
ECCV (26)2
2019 Beyond saliency: Understanding convolutional neural networks from saliency prediction on layer-wise relevance propagation
Yunke Tian, Klaus Mueller 0001, Xin Chen 0071
Image Vis. Comput.4
2018 Fully-Coupled Two-Stream Spatiotemporal Networks for Extremely Low Resolution Action Recognition
abstract
A major emerging challenge is how to protect people's privacy as cameras and computer vision are increasingly integrated into our daily lives, including in smart devices inside homes. A potential solution is to capture and record just the minimum amount of information needed to perform a task of interest. In this paper, we propose a fully-coupled two-stream spatiotemporal architecture for reliable human action recognition on extremely low resolution (e.g., 1216 pixel) videos. We provide an efficient method to extract spatial and temporal features and to aggregate them into a robust feature representation for an entire action video sequence. We also consider how to incorporate high resolution videos during training in order to build better low resolution action recognition models. We evaluate on two publicly-available datasets, showing significant improvements over the state-of-the-art.
Aidean Sharghi, Xin Chen 0071, David Crandall
WACV3
2016 Parallel nonparametric binarization for degraded document images
Xin Chen 0071, Yuefang Gao
Neurocomputing1
2016 Large-scale support vector machine classification with redundant data reduction
Xiangjun Shen, Lei Mu, Haoxiang Wu, Jianping Gou, Xin Chen 0071
Neurocomputing6