VLDB 2026 Research / reviewers in the wild / expert
Siyue Yu
dblp:280/3030
· DBLP profile ↗
26ranked-venue papers
3as first author
26since 2021 · last 2026
0009-0006-6749-4318ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 2 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | End-to-end railway obstacle detection enhanced by point cloud segmentation
Yuxing Yang, Kaizhong Xiao, Xiaolong Tuo, Liewei Wang, Siyue Yu, Jimin Xiao |
Eng. Appl. Artif. Intell. | 8 |
| 2026 | FFEvent: Fast fourier-based knowledge transfer for event cameras
Yuhui Lin, Siyue Yu, Jimin Xiao, Jiaxuan Lu |
Expert Syst. Appl. | 3 |
| 2026 | Probing 3D anomalies via multi-view registration and dual-residual analysis
Yuxing Yang, Zeyu Fu, Liewei Wang, Siyue Yu, Jimin Xiao |
Neurocomputing | 5 |
| 2026 | Unleashing the power of optimal head in CLIP and DINO for weakly supervised semantic segmentation
Xianglin Qiu, Siyue Yu, Bingfeng Zhang, Tammam Tillo, Jimin Xiao |
Pattern Recognit. | 2 |
| 2026 | DiffClick: Click-differentiated enhancement network for interactive segmentation
Siqi Song, Siyue Yu, Huiyu Zhou 0001, Xiaowei Huang 0001, Limin Yu, Jimin Xiao |
Pattern Recognit. | 2 |
| 2026 | MvP-Diff: Multivariate yet precise diffusion for anomaly images synthesis and segmentation
Siyue Yao, Eng Gee Lim, Siyue Yu, Jimin Xiao, Mingjie Sun |
Pattern Recognit. | 3 |
| 2026 | CoMasTRe+: Unleashing Disentangled Continual Segmentation With Mixture of Continual AdaptersabstractContinual Semantic Segmentation (CSS) suffers from catastrophic forgetting, particularly challenging for traditional per-pixel methods. Our prior work, CoMasTRe (CVPR 2024), introduced a query-based approach leveraging objectness by disentangling CSS into objectness learning and class recognition stages. While effective, CoMasTRe exhibited performance limitations due to feature forgetting within its pixel decoder. This paper presents CoMasTRe+, an enhanced framework specifically designed to overcome this limitation. The core contribution is a novel plugin, the Mixture of Continual Adapters (MoCA), integrated into the pixel decoder. MoCA is a dynamic architecture that mitigates feature forgetting by learning task-specific expert adapters. Crucially, MoCA employs a task-aware routing strategy and a novel adaptive routing distillation objective, tailored for continual learning, to preserve specialized feature representations across sequential tasks. CoMasTRe+ further enhances the class decoder using MoCA for improved recognition and simplicity. We extensively evaluate CoMasTRe+ on PASCAL VOC and ADE20K for continual semantic and panoptic segmentation. Experiments demonstrate that CoMasTRe+ effectively addresses the identified feature forgetting issue, significantly outperforms the original CoMasTRe, and achieves state-of-the-art results compared to both per-pixel and query-based baselines. Yizheng Gong, Siyue Yu, Liquan Shen, Jimin Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | A Training-free Synthetic Data Selection Method for Semantic SegmentationabstractTraining semantic segmenter with synthetic data has been attracting great attention due to its easy accessibility and huge quantities. Most previous methods focused on producing large-scale synthetic image-annotation samples and then training the segmenter with all of them. However, such a solution remains a main challenge in that the poor-quality samples are unavoidable, and using them to train the model will damage the training process. In this paper, we propose a training-free Synthetic Data Selection (SDS) strategy with CLIP to select high-quality samples for building a reliable synthetic dataset. Specifically, given massive synthetic image-annotation pairs, we first design a Perturbation-based CLIP Similarity (PCS) to measure the reliability of synthetic image, thus removing samples with low-quality images. Then we propose a class-balance Annotation Similarity Filter (ASF) by comparing the synthetic annotation with the response of CLIP to remove the samples related to low-quality annotations. The experimental results show that using our method significantly reduces the data size by half, while the trained segmenter achieves higher performance. Siyue Yu, Jian Pang, Bingfeng Zhang |
AAAI | 2 |
| 2025 | POT: Prototypical Optimal Transport for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) leverages Class Activation Maps (CAMs) to extract spatial information from image-level labels. However, CAMs primarily highlight the most discriminative foreground regions, leading to incomplete results. Prototype-based methods attempt to address this limitation by employing prototype CAMs instead of classifier CAMs. Nevertheless, existing prototype-based methods typically use a single prototype for each class, which is insufficient to capture all attributes of the foreground features due to the significant intra-class variations across different images. Consequently, these methods still struggle with incomplete CAM predictions. In this paper, we propose a novel framework called Prototypical Optimal Transport (POT) for WSSS. POT enhances CAM predictions by dividing features into multiple clusters and activating each cluster using its prototype. In this process, a similarity-aware optimal transport is employed to assign features to the most probable clusters. This similarity-aware strategy ensures the prioritization of significant cluster prototypes, thereby improving the accuracy of feature assignment. Additionally, we introduce an adaptive OT-based consistency loss to refine feature representations. This framework effectively overcomes the limitations of single-prototype methods, providing more complete and accurate CAM predictions. Extensive experimental results on standard WSSS benchmarks (PASCAL VOC and MS COCO) demonstrate that our method significantly improves the quality of CAMs and achieves state-of-the-art performances. The source code will be released https://github.com/jianwang91/POT. Jian Wang 0122, Tianhong Dai, Bingfeng Zhang, Siyue Yu, Eng Gee Lim, Jimin Xiao |
CVPR | 4 |
| 2025 | Feature Purification Matters: Suppressing Outlier Propagation for Training-Free Open-Vocabulary Semantic Segmentation
Siyue Yu, Bingfeng Zhang, Mingjie Sun, Yi Dong 0002, Jimin Xiao |
ICCV | 2 |
| 2025 | Class Token as Proxy: Optimal Transport-Assisted Proxy Learning for Weakly Supervised Semantic Segmentation
Jian Wang 0122, Tianhong Dai, Bingfeng Zhang, Siyue Yu, Eng Gee Lim, Jimin Xiao |
ICCV | 4 |
| 2025 | DriftRemover: Hybrid Energy Optimizations for Anomaly Images Synthesis and SegmentationabstractThis paper tackles the challenge of anomaly image synthesis and segmentation to generate various anomaly images and their segmentation labels to mitigate the issue of data scarcity. Existing approaches employ the precise mask to guide the generation, relying on additional mask generators, leading to increased computational costs and limited anomaly diversity. Although a few works use coarse masks as the guidance to expand diversity, they lack effective generation of labels for synthetic images, thereby reducing their practicality. Therefore, our proposed method simultaneously generates anomaly images and their corresponding masks by utilizing coarse masks and anomaly categories. The framework utilizes attention maps from synthesis process as mask labels and employs two optimization modules to tackle drift challenges, which are mismatches between synthetic results and real situations. Our evaluation demonstrates that our method improves pixel-level AP by 1.3% and F1-MAX by 1.8% in anomaly detection tasks on the MVTec dataset. Additionally, its successful application in practical scenarios highlights its effectiveness, improving IoU by 37.2% and F-measure by 25.1% with the Floor Dirt dataset. The code is available at https://github.com/JJessicaYao/DriftRemover. Siyue Yao, Mingjie Sun, Siyue Yu, Jimin Xiao, Eng Gee Lim |
IJCAI | 4 |
| 2025 | Segmentation guided dual-branch classification for measuring fat infiltration in paraspinal musclesabstractMuscle fat infiltration (FI) is a significant change in muscle degeneration. In particular, fat infiltration in paraspinal muscles (PSMs) indicates lumbar degenerative diseases. Thus, classifying different grades of FI in PSMs plays an important role in diagnosing the relevant lumbar diseases. Recently, many deep-learning-based methods have been introduced into medical image tasks. However, such methods for classifying the grades of FI in PSMs have not been explored. In this case, this paper aims to involve deep-learning methods in grade classification for FI in PSMs. Firstly, we construct a PSMsFIGC dataset for deep learning exploration. Our PSMsFIGC dataset contains 4 grades for classification and the corresponding PSMs segmentation masks as assistance. Additionally, we propose a segmentation guided dual-branch classification framework (SGDC) to assist radiologists in confirming the grade of FI in PSMs. The structure mainly consists of a segmentation branch and a classification branch. The segmentation branch is designed to suppress the influence of irrelevant muscles for final classification. We further design a critical area indicator based on the prediction of the segmentation branch to involve more related crucial areas for the classification branch and thus bridge the two branches. However, we find that inter-class disturbance, caused by PSMs’ similar shape and features, makes the network easily fall into local optimal. Therefore, we propose a disturbance weakening module to relieve the disturbance. Extensive experiments show that our SGDC can surpass existing classific classification networks, e.g., the proposed method achieves an impressive accuracy of 89.0% on the PSMsFIGC dataset. Our dataset and code will be released at https://github.com/myjianghao/Segmentation-Guided-Dual-branch-Classification-Framework-SGDC- . • A novel FI classification dataset constructed for exploring in FI grade. • A dual-branch framework designed to predict FI grades. • A module leveraging masks to focus on lesion-related regions. • A Gaussian-based strategy to mitigate inter-class similarity in MRI. Chengnan Jing, Hao Jiang 0054, Jimin Xiao, Siyue Yu, Minfeng Gan |
Expert Syst. Appl. | 6 |
| 2025 | Frozen CLIP-DINO: A Strong Backbone for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation has witnessed great achievements with image-level labels. Several recent approaches use the CLIP model to generate pseudo labels for training an individual segmentation model, while there is no attempt to apply the CLIP model as the backbone to directly segment objects with image-level labels. In this paper, we propose WeCLIP and its advanced version WeCLIP+, to build the single-stage pipeline for weakly supervised semantic segmentation. For WeCLIP, the frozen CLIP model is applied as the backbone for semantic feature extraction, and a new light decoder is designed to interpret extracted semantic features for final prediction. Meanwhile, we utilize the above frozen backbone to generate pseudo labels for training the decoder. Such labels are fixed during training. We then propose a refinement module (RFM) to optimize them dynamically. For WeCLIP+, we introduce the frozen DINO model to achieve more comprehensive semantic feature extraction. The frozen DINO is combined with the frozen CLIP as the backbone, followed by a shared decoder to make predictions with less training cost. Moreover, a strengthened refinement module (RFM+) is designed to revise online pseudo labels with extra guidance from DINO features. Extensive experiments show that both WeCLIP and WeCLIP+ significantly outperform other approaches with less training cost. Particularly, WeCLIP+ gets mIoU of 83.9% on VOC 2012 test set and 56.3% on COCO val set. Additionally, these two approaches also obtain promising results for fully supervised settings. Bingfeng Zhang, Siyue Yu, Jimin Xiao, Yunchao Wei, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Continual Segmentation with Disentangled Objectness Learning and Class RecognitionabstractMost continual segmentation methods tackle the prob-lem as a per-pixel classification task. However, such a paradigm is very challenging, and we find query-based seg-menters with built-in objectness have inherent advantages compared with per-pixel ones, as objectness has strong transfer ability and forgetting resistance. Based on these findings, we propose CoMasTRe by disentangling continual segmentation into two stages: forgetting-resistant continual objectness learning and well-researched continual classi-fication. CoMasTRe uses a two-stage segmenter learning class-agnostic mask proposals at the first stage and leaving recognition to the second stage. During continual learning, a simple but effective distillation is adopted to strengthen objectness. To further mitigate the forgetting of old classes, we design a multi-label class distillation strategy suited for segmentation. We assess the effectiveness of CoMas-TRe on PASCAL VOC and ADE20K. Extensive experiments show that our method outperforms per-pixel and query-based methods on both datasets. Code will be available at https://github.com/jordangong/CoMasTRe. Yizheng Gong, Siyue Yu, Xiaoyang Wang 0007, Jimin Xiao |
CVPR | 2 |
| 2024 | Frozen CLIP: A Strong Backbone for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation has witnessed great achievements with image-level labels. Several recent approaches use the CLIP model to generate pseudo labels for training an individual segmentation model, while there is no attempt to apply the CLIP model as the backbone to directly segment objects with image-level labels. In this paper, we propose WeCLIP, a CLIP-based single-stage pipeline, for weakly supervised semantic segmentation. Specifically, the frozen CLIP model is applied as the backbone for semantic feature extraction, and a new decoder is designed to interpret extracted semantic features for final prediction. Meanwhile, we utilize the above frozen backbone to generate pseudo labels for training the decoder. Such labels cannot be optimized during training. We then propose a refinement module (RFM) to rectify them dynamically. Our architecture enforces the proposed decoder and RFM to benefit from each other to boost the final performance. Extensive experiments show that our approach significantly outperforms other approaches with less training cost. Additionally, our WeCLIP also obtains promising results for fully supervised settings. The code is available at https://github.com/zbf1991/WeCLIP. Bingfeng Zhang, Siyue Yu, Yunchao Wei, Yao Zhao 0001, Jimin Xiao |
CVPR | 2 |
| 2024 | Adversarial Erasing Transformer for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation has attracted a lot of attention recently. Previous methods can be divided into two types, which are single-stage training and multi-stage training. In this paper, we focus on multi-stage training for image-level weakly supervised semantic segmentation. Many recent methods have tried to use transformer architecture as the backbone for CAM generation since it can capture global relationships to refine CAM accurately. However, we observe that such a backbone still fails to generate complete and smooth CAM. We argue that this is because the attention mechanism in the transformer can only pay attention to the most discriminative relationships. It is difficult to capture semantic-level long-range pair-wise relationships under image-level supervision. Thus, we propose an adversarial erasing transformer network called AETN, where an erasing attention mechanism is designed to establish more extensive pair-wise relationships. To cope with erasing, more target features will be forced to activate. Thus, better feature representation can be obtained for more accurate CAM generation. Besides, to further help our network learn better feature representation, we propose a self-consistent learning mechanism based on different augmentations. In this way, our AETN outperforms recent methods. Our AETN achieves 73.0 mIoU on the PASCAL VOC 2012 val set and 73.9 mIoU on the PASCAL VOC 2012 test set. Code is available a https://github.com/siyueyu/AETN. Bingfeng Zhang, Siyue Yu, Xuru Gao, Mingjie Sun, Eng Gee Lim, Jimin Xiao |
ECAI | 2 |
| 2024 | Cross-frame feature-saliency mutual reinforcing for weakly supervised video salient object detection
Jian Wang 0122, Siyue Yu, Bingfeng Zhang, Xinqiao Zhao, Ángel F. García-Fernández, Eng Gee Lim, Jimin Xiao |
Pattern Recognit. | 2 |
| 2024 | Enhanced online CAM: Single-stage weakly supervised semantic segmentation via collaborative guidance
Bingfeng Zhang, Xuru Gao, Siyue Yu, Weifeng Liu 0001 |
Pattern Recognit. | 3 |
| 2024 | Correction: ITContrast: contrastive learning with hard negative synthesis for image-text matching
Fangyu Wu 0001, Qiufeng Wang 0001, Zhao Wang 0001, Siyue Yu, Yushi Li, Eng Gee Lim |
Vis. Comput. | 4 |
| 2024 | ITContrast: contrastive learning with hard negative synthesis for image-text matching
Fangyu Wu 0001, Qiufeng Wang 0001, Zhao Wang 0001, Siyue Yu, Yushi Li, Eng Gee Lim |
Vis. Comput. | 4 |
| 2023 | Self-Compensating Learning for Few-Shot SegmentationabstractFew-shot segmentation (FSS) has witnessed rapid development. Most existing approaches extract prototypes from support images to segment query images. However, the integrity and validity of these support prototypes cannot be guaranteed. To solve the above drawbacks, we propose a self-compensating strategy, aiming to provide query-aware support information, to build more effective matching between support information and query images. Specifically, we design a prototype compensating module to mine useful information from the query prediction, to update original support prototypes as new query-aware support prototypes. Then the updated prototypes are utilized to perform the second matching with query features. In addition, we also compensate the information of original prior masks on the second matching phase, to improve the quality of prior masks. With improved prototype representations and prior knowledge, our approach can directly improve the performance of different approaches with new state-of-the-art performances. Bingfeng Zhang, Weifeng Liu 0001, Baodi Liu, Siyue Yu |
ICIP | 5 |
| 2023 | Weight-guided class complementing for long-tailed image recognition
Xinqiao Zhao, Jimin Xiao, Siyue Yu, Hui Li 0085, Bingfeng Zhang |
Pattern Recognit. | 3 |
| 2022 | Democracy Does Matter: Comprehensive Feature Mining for Co-Salient Object DetectionabstractCo-salient object detection, with the target of detecting co-existed salient objects among a group of images, is gaining popularity. Recent works use the attention mechanism or extra information to aggregate common co-salient features, leading to incomplete even incorrect responses for target objects. In this paper, we aim to mine comprehensive co-salient features with democracy and reduce background interference without introducing any extra information. To achieve this, we design a democratic prototype generation module to generate democratic response maps, covering sufficient co-salient regions and thereby involving more shared attributes of co-salient objects. Then a comprehensive prototype based on the response maps can be generated as a guide for final prediction. To suppress the noisy background information in the prototype, we propose a self-contrastive learning module, where both positive and negative pairs are formed without relying on additional classification information. Besides, we also design a democratic feature enhancement module to further strengthen the co-salient features by readjusting attention values. Extensive experiments show that our model obtains better performance than previous state-of-the-art methods, especially on challenging real-world cases (e.g., for CoCA, we obtain a gain of 2.0% for MAE, 5.4% for maximum F-measure, 2.3% for maximum E-measure, and 3.7% for S-measure) under the same settings. Source code is available at https://github.com/siyueyu/DCFM. Siyue Yu, Jimin Xiao, Bingfeng Zhang, Eng Gee Lim |
CVPR | 1 |
| 2021 | Structure-Consistent Weakly Supervised Salient Object Detection with Local Saliency CoherenceabstractSparse labels have been attracting much attention in recent years. However, the performance gap between weakly supervised and fully supervised salient object detection methods is huge, and most previous weakly supervised works adopt complex training methods with many bells and whistles. In this work, we propose a one-round end-to-end training approach for weakly supervised salient object detection via scribble annotations without pre/post-processing operations or extra supervision data. Since scribble labels fail to offer detailed salient regions, we propose a local coherence loss to propagate the labels to unlabeled regions based on image features and pixel distance, so as to predict integral salient regions with complete object structures. We design a saliency structure consistency loss as self-consistent mechanism to ensure consistent saliency maps are predicted with different scales of the same image as input, which could be viewed as a regularization technique to enhance the model generalization ability. Additionally, we design an aggregation module (AGGM) to better integrate high-level features, low-level features and global context information for the decoder to aggregate various information. Extensive experiments show that our method achieves a new state-of-the-art performance on six benchmarks (e.g. for the ECSSD dataset: Fβ = 0.8995, Eξ = 0.9079 and MAE = 0.0489), with an average gain of 4.60% for F-measure, 2.05% for E-measure and 1.88% for MAE over the previous best performing method on this task. Source code is available at http://github.com/siyueyu/SCWSSOD. Siyue Yu, Bingfeng Zhang, Jimin Xiao, Eng Gee Lim |
AAAI | 1 |
| 2021 | Fast pixel-matching for video object segmentation
Siyue Yu, Jimin Xiao, Bingfeng Zhang, Eng Gee Lim, Yao Zhao 0001 |
Signal Process. Image Commun. | 1 |