VLDB 2026 Research / reviewers in the wild / expert
Yunyi Xuan
dblp:272/6809
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2024
0009-0000-6365-5475ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Better Together: Data-Free Multi-Student Coevolved Distillation
Weijie Chen 0006, Yunyi Xuan, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang |
Knowl. Based Syst. | 2 |
| 2023 | Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt DiversificationabstractData-Free Knowledge Distillation (DFKD) has shown great potential in creating a compact student model while alleviating the dependency on real training data by synthesizing surrogate data. However, prior arts are seldom discussed under distribution shifts, which may be vulnerable in real-world applications. Recent Vision-Language Foundation Models, e.g., CLIP, have demonstrated remarkable performance in zero-shot out-of-distribution generalization, yet consuming heavy computation resources. In this paper, we discuss the extension of DFKD to Vision-Language Foundation Models without access to the billion-level image-text datasets. The objective is to customize a student model for distribution-agnostic downstream tasks with given category concepts, inheriting the out-of-distribution generalization capability from the pre-trained foundation models. In order to avoid generalization degradation, the primary challenge of this task lies in synthesizing diverse surrogate images driven by text prompts. Since not only category concepts but also style information are encoded in text prompts, we propose three novel Prompt Diversification methods to encourage image synthesis with diverse styles, namely Mix-Prompt, Random-Prompt, and Contrastive-Prompt. Experiments on out-of-distribution generalization datasets demonstrate the effectiveness of the proposed methods, with Contrastive-Prompt performing the best. Yunyi Xuan, Weijie Chen 0006, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang |
ACM Multimedia | 1 |
| 2022 | Label Matching Semi-Supervised Object DetectionabstractSemi-supervised object detection has made significant progress with the development of mean teacher driven self-training. Despite the promising results, the label mismatch problem is not yet fully explored in the previous works, leading to severe confirmation bias during self-training. In this paper, we delve into this problem and propose a simple yet effective LabelMatch framework from two different yet complementary perspectives, i.e., distribution-level and instance-level. For the former one, it is reasonable to approximate the class distribution of the unlabeled data from that of the labeled data according to Monte Carlo Sampling. Guided by this weakly supervision cue, we introduce a re-distribution mean teacher, which leverages adaptive label-distribution-aware confidence thresholds to generate unbiased pseudo labels to drive student learning. For the latter one, there exists an overlooked label assignment ambiguity problem across teacher-student models. To remedy this issue, we present a novel label assignment mechanism for self-training framework, namely proposal self-assignment, which injects the proposals from student into teacher and generates accurate pseudo labels to match each proposal in the student model accordingly. Experiments on both MS-COCO and PASCAL-VOC datasets demonstrate the considerable superiority of our proposed framework to other state-of-the-arts. Code will be available at https://github.com/HIK-LAB/SSOD. Weijie Chen 0006, Shicai Yang, Yunyi Xuan, Jie Song 0011, Di Xie, Shiliang Pu, Mingli Song, Yueting Zhuang |
CVPR | 4 |
| 2022 | Simulation-and-Mining: Towards Accurate Source-Free Unsupervised Domain Adaptive Object DetectionabstractVanilla unsupervised domain adaptive (UDA) object detection typically requires the labeled source data for joint-training with the unlabeled target data, which is usually unavailable in real-world scenarios due to data privacy, leading to source data-free UDA object detection. Herein, we first analyze the phenomenon of cross-domain detection degradation varying from easy to hard samples (e.g. the objects with different scales or occlusion degrees), termed as domain generalization differentiation. In detail, the ability to detect easy samples is well transferred while the one to detect hard samples is dramatically degraded. To this end, we then revisit the existing self-training method, which is of great challenge to deal with the abundant false negatives (hard samples). Assumed that true positives (easy samples) labeled by the source model can be exploited as supervision cues. UDA is finally modeled into an unsupervised false negatives mining problem. Thus, we propose a Simulation-and-Mining (S&M) framework, which simulates false negatives by augmenting true positives and mines back false negatives alternatively and iteratively. Experimental results show the effectiveness. Weijie Chen 0006, Shicai Yang, Yunyi Xuan, Di Xie, Yueting Zhuang, Shiliang Pu |
ICASSP | 4 |
| 2021 | Efficient Video Compressed Sensing Reconstruction via Exploiting Spatial-Temporal Correlation With Measurement ConstraintabstractRecent deep learning-based video compressed sensing (VCS) methods have achieved promising results but still suffer from numerous hyper-parameters and inflexibility. This paper proposes a novel network for VCS, named STM-Net, to fast recover high-quality video frames by optionally exploiting Spatial-Temporal information with a Measurement constraint. Combining the merits of adaptive sampling and adaptive shrinkage-thresholding, we first propose an improved ISTA-Net+ for framewise independent reconstruction, called Unfolding Adaptive Shrinkage-Thresholding Network (UAST-Net). To get further non-key frames reconstruction improvement, we develop a two-phase joint deep reconstruction, including an Occlusion-Aware Temporal Alignment to avoid irrelevant information compensation and a Multiple Frames Fusion with proposed Spatial-Temporal Feature Weighting (STFW) module to guide attractive content extraction and discriminative features generation. Besides, we develop a measurement loss to reduce the solution space to facilitate network optimization. Experimental results demonstrate the superiority of the proposed STM-Net over the existing methods. Zhichao Wei, Yunyi Xuan |
ICME | 3 |
| 2021 | Adaptive Threshold-based Sparse Representation Network for Image Compressive Sensing ReconstructionabstractRecently, network-based Image Compressive Sensing (ICS) algorithms show superior performance in reconstruction quality and speed, yet non-interpretable. Herein, we propose an Adaptive Threshold-based Sparse Representation Reconstruction Network (ATSR-Net), composed of the Convolutional Sparse Representation subnet (CSR-subnet) and the truly Adaptive Threshold Generation subnet (ATG-subnet). The traditional iterations are unfolded into several CSR-subnets, which can fully exploit the local and nonlocal similarities. The ATG-subnet automatically determines a threshold map based on the image intrinsic characterization for flexible feature selection. Moreover, we present a three-level consistency loss based on pixel-level, measurement-level, and feature-level, to accelerate the network convergence. Extensive experiment results demonstrate the superiority of the proposed network to the existing state-of-the-art methods by large margins, both quantitatively and qualitatively. Yunyi Xuan |
VCIP | 1 |
| 2020 | 2Ser-Vgsr-Net: A Two-Stage Enhancement Reconstruction Based On Video Group Sparse Representation Network For Compressed Video SensingabstractCompressed sensing (CS) requires fewer measurements than the Nyquist theory, making it excellent potential for video compression. However, the complex computation and long-running time limit the traditional compressed video sensing (CVS) methods in real-time application. In this paper, we proposed a fast CVS reconstruction based on deep network named 2sER-VGSR-Net. We first perform ISTA-Net+ as initial reconstruction. To exploit temporal and spatial correlation intrinsically, we construct a video inter-frame group by extracting blocks from the current and reference frames while establishing a sparse representation by network, called VGSR-Net. Different from traditional CVS methods, the group proposed in this paper contains fewer blocks thanks to the accurate alignment by STMC-Net. The inter-frame reconstruction comprises two stages, of which the first stage gets primary enhancement for motion compensation, and the latter performs as residual reconstruction to recover the details. Experiments show that the proposed 2sER-VGSR-Net outperforms the existing state-of-art CVS reconstruction algorithms. Yunyi Xuan |
ICME | 1 |