EDBT 2026 Demo / reviewers in the wild / expert
Yixin Zhang 0007
dblp:14/5888-7
· DBLP profile ↗
11ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-4513-1106ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training-free Boosting for Few-shot Segmentation via Generalizing Semantic MiningabstractFew-shot Semantic Segmentation (FSS) aims to segment the novel target objects with the guidance of minimal annotated reference examples. The affinity-based method has great advantages in the FSS inference stage for both specialist model and foundation model. However, current affinity calculation merely relies on only support-query matching, without considering the query-specific semantic or the semantic correlation among inter-support samples, which limits the representation ability of affinity map. In this paper, we propose the Generalizing Semantic Mining (GSM) that focuses on exploiting generalizing semantic to improve the affinity calculation. Concretely, we first organize the affinity-based inference into three main steps to reveal the crucial role of affinity map. To address the low-data problem, Target Semantic Reusing module considers the query sample as a proxy reference and assigns it with proxy mask identifying its most generalizing semantic regions. Then, to generate the high-fidelity proxy mask, Query-specific Semantic Modeling module pinpoints the most generalizing regions through prior semantic analysis. Finally, Representative Re-weighting module explicitly modulates affinity calculation via generalization-aware weighting. Experiments on FSS benchmarks demonstrate that our GSM can serve as a plug-and-play free lunch for both specialist models and foundation models. Kangyu Xiao, Zilei Wang, Yixin Zhang 0007, Junjie Li 0002 |
AAAI | 3 |
| 2026 | Improving Zero-Shot Generalization for CLIP With Prompt Ensemble Self-DistillationabstractPrompt tuning has emerged as an effective alternative for adapting pre-trained Vision-Language Models (VLMs) to various downstream tasks. In our experiments utilizing prompt tuning methods, we observed that modifying the prompt initialization led to inconsistencies in the model’s predictions, particularly with pronounced variability on specific datasets. Motivated by this observation, we examine the predictive performance of two ensemble methods: prompt fusion and logits fusion. Experimental results indicate that logits fusion results in considerable performance improvements, while prompt fusion does not yield any enhancements. However, a significant downside of logits fusion is the enormous rise in inference time. To investigate a practical approach for integrating knowledge derived from multiple prompts without incurring additional inference costs, we propose a straightforward Prompt Ensemble self-Distillation (PED) framework that considerably improves the generalization capacity of prompt tuning. Specifically, we initialize multiple groups of prompts, and during the training process, we integrate the prediction outputs from each group to facilitate the learning of the fused prompts. The proposed self-distillation approach offers dual benefits: enhancing the performance of both the fused prompts and the fused logits. We utilize fused prompts for prediction during the inference process, thereby achieving performance that is comparable to that of fused logits without incurring additional inference time. We evaluate the effectiveness of our methodology across four distinct tasks. Our PED consistently demonstrates superior performance in all assessments when contrasted with numerous state-of-the-art methods. Moreover, our method can be seamlessly integrated into existing prompt learning approaches and consistently improves their performance. Our code is publicly available at https://github.com/vim-wei/PED. Zilei Wang, Yixin Zhang 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Context Sensitive Network for weakly-supervised fine-grained temporal action localization
Cerui Dong, Qinying Liu, Zilei Wang, Yixin Zhang 0007, Feng Zhao 0004 |
Neural Networks | 4 |
| 2024 | Probabilistic Contrastive Learning for Domain Adaptation
Junjie Li 0002, Yixin Zhang 0007, Zilei Wang, Saihui Hou, Keyu Tu, Man Zhang 0005 |
IJCAI | 2 |
| 2023 | Exploit Domain-Robust Optical Flow in Domain Adaptive Video Semantic SegmentationabstractDomain adaptive semantic segmentation aims to exploit the pixel-level annotated samples on source domain to assist the segmentation of unlabeled samples on target domain. For such a task, the key is to construct reliable supervision signals on target domain. However, existing methods can only provide unreliable supervision signals constructed by segmentation model (SegNet) that are generally domain-sensitive. In this work, we try to find a domain-robust clue to construct more reliable supervision signals. Particularly, we experimentally observe the domain-robustness of optical flow in video tasks as it mainly represents the motion characteristics of scenes. However, optical flow cannot be directly used as supervision signals of semantic segmentation since both of them essentially represent different information. To tackle this issue, we first propose a novel Segmentation-to-Flow Module (SFM) that converts semantic segmentation maps to optical flows, named the segmentation-based flow (SF), and then propose a Segmentation-based Flow Consistency (SFC) method to impose consistency between SF and optical flow, which can implicitly supervise the training of segmentation model. The extensive experiments on two challenging benchmarks demonstrate the effectiveness of our method, and it outperforms previous state-of-the-art methods with considerable performance improvement. Our code is available at https://github.com/EdenHazardan/SFC. Zilei Wang, Jiafan Zhuang, Yixin Zhang 0007, Junjie Li 0002 |
AAAI | 4 |
| 2023 | Revisiting Foreground and Background Separation in Weakly-supervised Temporal Action Localization: A Clustering-based ApproachabstractWeakly-supervised temporal action localization aims to localize action instances in videos with only video-level action labels. Existing methods mainly embrace a localization-by-classification pipeline that optimizes the snippet-level prediction with a video classification loss. However, this formulation suffers from the discrepancy between classification and detection, resulting in inaccurate separation of foreground and background (F&B) snippets. To alleviate this problem, we propose to explore the underlying structure among the snippets by resorting to unsupervised snippet clustering, rather than heavily relying on the video classification loss. Specifically, we propose a novel clustering-based F&B separation algorithm. It comprises two core components: a snippet clustering component that groups the snippets into multiple latent clusters and a cluster classification component that further classifies the cluster as foreground or background. As there are no ground-truth labels to train these two components, we introduce a unified self-labeling mechanism based on optimal transport to produce high-quality pseudo-labels that match several plausible prior distributions. This ensures that the cluster assignments of the snippets can be accurately associated with their F&B labels, thereby boosting the F&B separation. We evaluate our method on three benchmarks: THUMOS14, ActivityNet v1.2 and v1.3. Our method achieves promising performance on all three benchmarks while being significantly more lightweight than previous methods. Code is available at https://github.com/Qinying-Liu/CASE Qinying Liu, Zilei Wang, Shenghai Rong, Junjie Li 0002, Yixin Zhang 0007 |
ICCV | 5 |
| 2023 | Towards Effective Instance Discrimination Contrastive Loss for Unsupervised Domain AdaptationabstractDomain adaptation (DA) aims to transfer knowledge from a label-rich source domain to a related but label-scarce target domain. Recently, increasing research has focused on exploring data structure of the target domain. In light of the recent success of Instance Discrimination Contrastive (IDCo) loss in self-supervised learning, we try directly applying it to domain adaptation tasks. However, the improvement is very limited, which motivates us to rethink its underlying limitations for domain adaptation tasks. An intuitive limitation is that a pair of samples belonging to the same class could be treated as negatives. Here we argue that using low-confidence samples to construct positive and negative pairs can alleviate this issue and is more suitable for IDCo loss. Another limitation is that IDCo loss cannot capture enough semantic information. We address this by introducing domain-invariant and accurate semantic information from classifier weights and input data. Specifically, we propose a class relationship enhanced features. It uses probability weighted class prototpyes as the input features of IDCo loss, which can implicitly transfer the domain-invariant class relationship. We further propose a target-dominated cross-domain mixup that can incorporate accurate semantic information from the source domain. We evaluate the proposed method in unsupervised DA and other DA settings, and extensive experimental results reveal that our method can make IDCo loss more effective and achieve state-of-the-art performance.1 Yixin Zhang 0007, Zilei Wang, Junjie Li 0002, Jiafan Zhuang |
ICCV | 1 |
| 2023 | Exploiting Low-confidence Pseudo-labels for Source-free Object DetectionabstractSource-free object detection (SFOD) aims to adapt a source-trained detector to an unlabeled target domain without access to the labeled source data. Current SFOD methods utilize a threshold-based pseudo-label approach in the adaptation phase, which is typically limited to high-confidence pseudo-labels and results in a loss of information. To address this issue, we propose a new approach to take full advantage of pseudo-labels by introducing high and low confidence thresholds. Specifically, the pseudo-labels with confidence scores above the high threshold are used conventionally, while those between the low and high thresholds are exploited using the Low-confidence Pseudo-labels Utilization (LPU) module. The LPU module consists of Proposal Soft Training (PST) and Local Spatial Contrastive Learning (LSCL). PST generates soft labels of proposals for soft training, which can mitigate the label mismatch problem. LSCL exploits the local spatial relationship of proposals to improve the model's ability to differentiate between spatially adjacent proposals, thereby optimizing representational features further. Combining the two components overcomes the challenges faced by traditional methods in utilizing low-confidence pseudo-labels. Extensive experiments on five cross-domain object detection benchmarks demonstrate that our proposed method outperforms the previous SFOD methods, achieving state-of-the-art performance. Zilei Wang, Yixin Zhang 0007 |
ACM Multimedia | 3 |
| 2023 | Semi-supervised Domain Adaptation via Joint Contrastive Learning with SensitivityabstractSemi-supervised Domain Adaptation (SSDA) aims to learn a well-performed model using fully labeled source samples and scarcely labeled target samples, along with unlabeled target samples. Due to the dominant presence of labeled samples from the source domain in the training data, both the feature extractor and classifier can display bias towards the source domain. This can result in sub-optimal feature extraction for the challenging target samples that have notable differences from the source domain. Moreover, the source-favored classifier can hinder the classification performance of the target domain. To this end, we propose a novel Joint Contrastive Learning with Sensitivity (JCLS) in this paper, which consists of sensitivity-aware feature contrastive learning (SFCL) and class-wise probabilistic contrastive learning (CPCL). Different from the traditional contrastive learning, SFCL pays more attention to the sensitive samples during optimizing the feature extractor, and consequently the feature discrimination of unlabeled samples can be enhanced. CPCL performs class-wise contrastive learning in the probabilistic space to enforce the cross-domain classifier to match the real distribution of source and target samples. By combining these two components, our JCLS is able to extract domain-invariant and compact features and obtain a well-performed classifier. We conduct the experiments on the DomainNet and Office-Home benchmarks, and the results show that our approach achieves state-of-the-art performance. Keyu Tu, Zilei Wang, Junjie Li 0002, Yixin Zhang 0007 |
ACM Multimedia | 4 |
| 2022 | Continual Semantic Segmentation via Structure Preserving and Projected Feature Alignment
Zilei Wang, Yixin Zhang 0007 |
ECCV (29) | 3 |
| 2020 | Polynomial Regression Network for Variable-Number Lane Detection
Bingke Wang, Zilei Wang, Yixin Zhang 0007 |
ECCV (18) | 3 |