Yunyi Xuan

dblp:272/6809 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2024
0009-0000-6365-5475ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2024 Better Together: Data-Free Multi-Student Coevolved Distillation
Weijie Chen 0006, Yunyi Xuan, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang
Knowl. Based Syst.2
2023 Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification
abstract
Data-Free Knowledge Distillation (DFKD) has shown great potential in creating a compact student model while alleviating the dependency on real training data by synthesizing surrogate data. However, prior arts are seldom discussed under distribution shifts, which may be vulnerable in real-world applications. Recent Vision-Language Foundation Models, e.g., CLIP, have demonstrated remarkable performance in zero-shot out-of-distribution generalization, yet consuming heavy computation resources. In this paper, we discuss the extension of DFKD to Vision-Language Foundation Models without access to the billion-level image-text datasets. The objective is to customize a student model for distribution-agnostic downstream tasks with given category concepts, inheriting the out-of-distribution generalization capability from the pre-trained foundation models. In order to avoid generalization degradation, the primary challenge of this task lies in synthesizing diverse surrogate images driven by text prompts. Since not only category concepts but also style information are encoded in text prompts, we propose three novel Prompt Diversification methods to encourage image synthesis with diverse styles, namely Mix-Prompt, Random-Prompt, and Contrastive-Prompt. Experiments on out-of-distribution generalization datasets demonstrate the effectiveness of the proposed methods, with Contrastive-Prompt performing the best.
Yunyi Xuan, Weijie Chen 0006, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang
ACM Multimedia1
2022 Label Matching Semi-Supervised Object Detection
abstract
Semi-supervised object detection has made significant progress with the development of mean teacher driven self-training. Despite the promising results, the label mismatch problem is not yet fully explored in the previous works, leading to severe confirmation bias during self-training. In this paper, we delve into this problem and propose a simple yet effective LabelMatch framework from two different yet complementary perspectives, i.e., distribution-level and instance-level. For the former one, it is reasonable to approximate the class distribution of the unlabeled data from that of the labeled data according to Monte Carlo Sampling. Guided by this weakly supervision cue, we introduce a re-distribution mean teacher, which leverages adaptive label-distribution-aware confidence thresholds to generate unbiased pseudo labels to drive student learning. For the latter one, there exists an overlooked label assignment ambiguity problem across teacher-student models. To remedy this issue, we present a novel label assignment mechanism for self-training framework, namely proposal self-assignment, which injects the proposals from student into teacher and generates accurate pseudo labels to match each proposal in the student model accordingly. Experiments on both MS-COCO and PASCAL-VOC datasets demonstrate the considerable superiority of our proposed framework to other state-of-the-arts. Code will be available at https://github.com/HIK-LAB/SSOD.
Weijie Chen 0006, Shicai Yang, Yunyi Xuan, Jie Song 0011, Di Xie, Shiliang Pu, Mingli Song, Yueting Zhuang
CVPR4
2022 Simulation-and-Mining: Towards Accurate Source-Free Unsupervised Domain Adaptive Object Detection
abstract
Vanilla unsupervised domain adaptive (UDA) object detection typically requires the labeled source data for joint-training with the unlabeled target data, which is usually unavailable in real-world scenarios due to data privacy, leading to source data-free UDA object detection. Herein, we first analyze the phenomenon of cross-domain detection degradation varying from easy to hard samples (e.g. the objects with different scales or occlusion degrees), termed as domain generalization differentiation. In detail, the ability to detect easy samples is well transferred while the one to detect hard samples is dramatically degraded. To this end, we then revisit the existing self-training method, which is of great challenge to deal with the abundant false negatives (hard samples). Assumed that true positives (easy samples) labeled by the source model can be exploited as supervision cues. UDA is finally modeled into an unsupervised false negatives mining problem. Thus, we propose a Simulation-and-Mining (S&M) framework, which simulates false negatives by augmenting true positives and mines back false negatives alternatively and iteratively. Experimental results show the effectiveness.
Weijie Chen 0006, Shicai Yang, Yunyi Xuan, Di Xie, Yueting Zhuang, Shiliang Pu
ICASSP4
2021 Efficient Video Compressed Sensing Reconstruction via Exploiting Spatial-Temporal Correlation With Measurement Constraint
abstract
Recent deep learning-based video compressed sensing (VCS) methods have achieved promising results but still suffer from numerous hyper-parameters and inflexibility. This paper proposes a novel network for VCS, named STM-Net, to fast recover high-quality video frames by optionally exploiting Spatial-Temporal information with a Measurement constraint. Combining the merits of adaptive sampling and adaptive shrinkage-thresholding, we first propose an improved ISTA-Net+ for framewise independent reconstruction, called Unfolding Adaptive Shrinkage-Thresholding Network (UAST-Net). To get further non-key frames reconstruction improvement, we develop a two-phase joint deep reconstruction, including an Occlusion-Aware Temporal Alignment to avoid irrelevant information compensation and a Multiple Frames Fusion with proposed Spatial-Temporal Feature Weighting (STFW) module to guide attractive content extraction and discriminative features generation. Besides, we develop a measurement loss to reduce the solution space to facilitate network optimization. Experimental results demonstrate the superiority of the proposed STM-Net over the existing methods.
Zhichao Wei, Yunyi Xuan
ICME3
2021 Adaptive Threshold-based Sparse Representation Network for Image Compressive Sensing Reconstruction
abstract
Recently, network-based Image Compressive Sensing (ICS) algorithms show superior performance in reconstruction quality and speed, yet non-interpretable. Herein, we propose an Adaptive Threshold-based Sparse Representation Reconstruction Network (ATSR-Net), composed of the Convolutional Sparse Representation subnet (CSR-subnet) and the truly Adaptive Threshold Generation subnet (ATG-subnet). The traditional iterations are unfolded into several CSR-subnets, which can fully exploit the local and nonlocal similarities. The ATG-subnet automatically determines a threshold map based on the image intrinsic characterization for flexible feature selection. Moreover, we present a three-level consistency loss based on pixel-level, measurement-level, and feature-level, to accelerate the network convergence. Extensive experiment results demonstrate the superiority of the proposed network to the existing state-of-the-art methods by large margins, both quantitatively and qualitatively.
Yunyi Xuan
VCIP1
2020 2Ser-Vgsr-Net: A Two-Stage Enhancement Reconstruction Based On Video Group Sparse Representation Network For Compressed Video Sensing
abstract
Compressed sensing (CS) requires fewer measurements than the Nyquist theory, making it excellent potential for video compression. However, the complex computation and long-running time limit the traditional compressed video sensing (CVS) methods in real-time application. In this paper, we proposed a fast CVS reconstruction based on deep network named 2sER-VGSR-Net. We first perform ISTA-Net+ as initial reconstruction. To exploit temporal and spatial correlation intrinsically, we construct a video inter-frame group by extracting blocks from the current and reference frames while establishing a sparse representation by network, called VGSR-Net. Different from traditional CVS methods, the group proposed in this paper contains fewer blocks thanks to the accurate alignment by STMC-Net. The inter-frame reconstruction comprises two stages, of which the first stage gets primary enhancement for motion compensation, and the latter performs as residual reconstruction to recover the details. Experiments show that the proposed 2sER-VGSR-Net outperforms the existing state-of-art CVS reconstruction algorithms.
Yunyi Xuan
ICME1