EDBT 2026 Demo / reviewers in the wild / expert
Binwei Xu
dblp:280/9626
· DBLP profile ↗
13ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0001-6763-9202ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-Anchored Progressive Framework With Noise Mitigation for Unsupervised Camouflaged Object DetectionabstractUnsupervised Camouflaged Object Detection (UCOD) presents a significant challenge due to the inherent similarity between camouflaged objects and their backgrounds, compounded by the absence of manual annotations. Although pixel-level pseudo-labeling has proven effective for unsupervised salient object detection (USOD), it is far less reliable for COD, where the concealed and ambiguous nature of camouflaged objects frequently produces noisy pseudo-labels, causing misjudgments, missed detections, and imprecise boundaries. To overcome this, we propose SAPNet, a novel self-anchored progressive framework for UCOD. Rather than depending on noisy pixel-level supervision, we leverage semantically reliable foreground and background regions as high-confidence anchors. This effectively transforms the unsupervised problem into a more robust weakly supervised paradigm, reducing learning difficulty and mitigating overfitting to noise. SAPNet learns camouflaged objects progressively by first emphasizing these confident regions and then exploiting DINO's contextual awareness to recover complete structures. Central to our framework is the semantic-driven region detector (SDRD), which employs cascaded convolutions and a residual attention projection mechanism to suppress background noise, filter erroneous information, and enhance spatial context, ensuring reliable supervision signals. Furthermore, a region-based context inference module (RCIM) is introduced to iteratively refine object boundaries by integrating multi-level semantic features under the guidance of these refined region-level anchors. Extensive experiments on four benchmark COD datasets demonstrate that SAPNet significantly outperforms state-of-the-art unsupervised methods. The source code of our SAPNet is available at https://github.com/ArloJie/SAPNet. Binwei Xu, Tuo Shen, Guanghui Yue 0001, Qiuping Jiang |
IEEE Trans. Image Process. | 2 |
| 2025 | Underwater Salient Object Detection via Dual-Stage Self-Paced Learning and Depth EmphasisabstractSalient object detection of underwater scenes (USOD) poses greater challenges than that of traditional terrestrial scenes due to the presence of diverse and complex underwater image degradation. Current deep learning-based USOD methods generally treat all samples equally while failing to account for the varying difficulty levels of different training samples, thus leading to a limited performance. To tackle this challenge, this paper introduces a novel deep USOD method which benefits from iterative Dual-stage Self-paced Learning (DSPL) and Salient Object Depth Emphasis (SODE). Specifically, a DSPL strategy, which enforces the network to only focus on simpler samples in the first stage and then shifts attention to more challenging samples in the second stage, is devised to imitate the learning process of humans. The whole network is iteratively trained with the DSPL strategy and thus gradually adapted to various underwater scenes with different difficulty levels. Additionally, the proposed method involves an SODE module, which adaptively enhances depth information to effectively locate salient objects, addressing the issue of unreliable depth data caused by underwater image quality degradation. Experimental results on two benchmark datasets demonstrate the superior performance of the proposed method against state-of-the-art methods. The source code of our method will be made available athttps://github.com/NIT-JJH/SPDE. Jianhui Jin, Qiuping Jiang, Qingyuan Wu, Binwei Xu, Runmin Cong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | A Weakly-Supervised Cross-Domain Query Framework for Video Camouflage Object DetectionabstractVCOD (Video Camouflage Object Detection) is a crucial security technology that identifies camouflaged objects in videos, bolstering security measures across diverse applications. On one hand, appearance-based VCOD methods face challenges because camouflaged appearances cause objects to blend into their surroundings, and current VCOD methods typically utilize optical flow to represent motion information. However, over-reliance on accurate estimation renders the model overly fragile. On the other hand, there is a shortage of effectively annotated camouflaged video datasets, coupled with the time-consuming and labor-intensive annotation process, severely constraining the development of this field. To address this, we propose a novel weakly-supervised framework for VCOD based on cross-domain querying of preceding and succeeding frames. Specifically, we propose a time-efficient and labor-saving manual annotation approach based on large visual models to rapidly generate pseudo-labels. Furthermore, we design a network based on Spatio-Temporal Memory (STM) that performs cross-modal feature querying with the current frame against preceding and succeeding frames to acquire useful information, thereby enhancing the focus on temporal information. Extensive experiments conducted on two common VCOD datasets have proven the effectiveness of our method, achieving state-of-the-art performance on the challenging camouflaged video data. Zelin Lu, Liang Xie 0003, Xing Zhao 0001, Binwei Xu, Haoran Liang 0001, Ronghua Liang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Multidimensional Exploration of Segment Anything Model for Weakly Supervised Video Salient Object DetectionabstractFully supervised video salient object detection (VSOD) has made considerable breakthroughs using costly and time-consuming pixel-wise annotations. Recently, to achieve a trade-off between the annotation burden and the model performance, scribble-based VSOD tasks have attracted increasing attention. However, learning the complete object structure and precise boundary details from sparse scribble annotations remains challenging. In this paper, we propose a series of strategies to effectively explore valid information from the recently proposed segmentation foundation model “Segment Anything Model (SAM)” in various perspectives to address these challenges. Specifically, due to the limited performance of SAM on videos, we propose a SAM-guided label enhancement method instead of directly using the results of SAM, which can introduce edge information while reducing the interference of erroneous information. Moreover, we propose a SAM-driven spatiotemporal network guided by general semantic features from the SAM encoder to help the model be aware of global connections. Additionally, we propose a SAM-based global-aware loss, which further considers the affinity constraint between predicted results and foreground labels or background labels from a global perspective, guiding the model to perceive the complete salient objects. Experimental results demonstrate that our method outperforms state-of-the-art weakly supervised VSOD methods and is comparable to fully supervised VSOD methods. Binwei Xu, Qiuping Jiang, Xing Zhao 0001, Chenyang Lu 0002, Haoran Liang 0001, Ronghua Liang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Learning Video Salient Object Detection Progressively From Unlabeled VideosabstractRecently, deep learning-based video salient object detection (VSOD) has achieved some breakthroughs, but these methods rely on expensive annotated videos with pixel-wise annotations or weak annotations. In this paper, based on the similarities and differences between VSOD and image salient object detection (SOD), we propose a novel VSOD method via a progressive framework that locates and segments salient objects in sequence without utilizing any video annotation. To efficiently use the knowledge learned in the SOD dataset for VSOD efficiently, we introduce dynamic saliency to compensate for the lack of motion information of SOD during the locating process while maintaining the same fine segmenting process. Specifically, we utilize the coarse locating model trained on the image dataset, to identify frames with both static and dynamic saliency. Locating results of these frames are selected as spatiotemporal location labels. Moreover, by tracking salient objects in adjacent frames, the number of spatiotemporal location labels is increased. On the basis of these location labels, a two-stream locating network with an optical flow branch is proposed to capture salient objects in videos. The results with respect to five public benchmarks demonstrate that our method outperforms the state-of-the-art weakly and unsupervised methods. Binwei Xu, Qiuping Jiang, Haoran Liang 0001, Dingwen Zhang, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Multim. | 1 |
| 2024 | Blind image quality assessment with semi-supervised learning
Xiwen Li, Binwei Xu |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | A Visual Representation-Guided Framework With Global Affinity for Weakly Supervised Salient Object DetectionabstractFully supervised salient object detection (SOD) methods have made considerable progress in performance, yet these models rely heavily on expensive pixel-wise labels. Recently, to achieve a trade-off between labeling burden and performance, scribble-based SOD methods have attracted increasing attention. Previous scribble-based models directly implement the SOD task only based on SOD training data with limited information, it is extremely difficult for them to understand the image and further achieve a superior SOD task. In this paper, we propose a simple yet effective framework guided by general visual representations with rich contextual semantic knowledge for scribble-based SOD. These general visual representations are generated by self-supervised learning based on large-scale unlabeled datasets. Our framework consists of a task-related encoder, a general visual module, and an information integration module to efficiently combine the general visual representations with task-related features to perform the SOD task based on understanding the contextual connections of images. Meanwhile, we propose a novel global semantic affinity loss to guide the model to perceive the global structure of the salient objects. Experimental results on five public benchmark datasets demonstrate that our method, which only utilizes scribble annotations without introducing any extra label, outperforms the state-of-theart weakly supervised SOD methods. Specifically, it outperforms the previous best scribble-based method on all datasets with an average gain of 5.5% for max f-measure, 5.8% for mean f-measure, 24% for MAE, and 3.1% for E-measure. Moreover, our method achieves comparable or even superior performance to the state-of-the-art fully supervised models. Binwei Xu, Haoran Liang 0001, Weihua Gong, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Synthesize Boundaries: A Boundary-Aware Self-Consistent Framework for Weakly Supervised Salient Object DetectionabstractFully supervised salient object detection (SOD) has made considerable progress based on expensive and time-consuming data with pixel-wise annotations. Recently, to relieve the labeling burden while maintaining performance, some scribble-based SOD methods have been proposed. However, learning precise boundary details from scribble annotations that lack edge information is still difficult. In this article, we propose to learn precise boundaries from our designed synthetic images and labels without introducing any extra auxiliary data. The synthetic image creates boundary information by inserting synthetic concave regions that simulate the real concave regions of salient objects. Furthermore, we propose a novel self-consistent framework that consists of a global integral branch (GIB) and a boundary-aware branch (BAB) to train a saliency detector. GIB aims to identify integral salient objects, whose input is the original image. BAB aims to help predict accurate boundaries, whose input is the synthetic image. These two branches are connected through a self-consistent loss to guide the saliency detector to predict precise boundaries while identifying salient objects. Experimental results on five benchmarks demonstrate that our method outperforms the state-of-the-art weakly supervised SOD methods and further narrows the gap with the fully supervised methods. Binwei Xu, Haoran Liang 0001, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Multim. | 1 |
| 2023 | A progressive segmentation with weight contrast label enhancement for weakly supervised video salient object detectionabstractAbstract Scribble labels have gained increasing attention in the field of weakly supervised video salient object detection (VSOD). Based on scribble labels, latest methods can spread labeled pixels to unlabeled regions using local coherence loss, but predicted objects often lose detail and boundary information. In this work, a novel method based on back‐foreground weight contrast is proposed that adds label enhancement points to facilitate the model to learn the edge, detail and location of salient object. Additionally, a new VSOD framework based on global structural localization is introduced. Enhanced scribble labels are used to assist the model for global localization, and then the located regions are finely segmented by the trained model. Extensive experiments demonstrate that the method achieves the state‐of‐the‐art performance on common VSOD datasets, with an improvement of 3.75%, 4.68%, and 0.88% in S‐measure, F‐measure, and MAE, respectively. Zelin Lu, Haoran Liang 0001, Binwei Xu, Ronghua Liang |
IET Image Process. | 3 |
| 2022 | CFN: A coarse-to-fine network for eye fixation predictionabstractAbstract Many image‐to‐image computer vision approaches have made great progress by an end‐to‐end framework with the encoder–decoder architecture. However, the same image‐to‐image eye fixation prediction task is not the same as those computer vision tasks in that it focuses more on salient regions rather than precise predictions for every pixel. Thus, it is not appropriate to directly apply the end‐to‐end encoder–decoder to the eye fixation prediction task. In addition, although high‐level feature is important, the contribution of low‐level feature should also be kept and balanced in computational model. Nevertheless, some low‐level features that attract attention are easily neglected while transiting through the deep network. Therefore, the effective way to integrate low‐level and high‐level features for improving eye fixation prediction performance is still a challenging task. In this paper, a coarse‐to‐fine network (CFN) that encompasses two pathways with different training strategies are proposed: coarse perceiving network (CFN‐Coarse) can be a simple encoder network or any of the existing pretrained network to capture the distribution of salient regions and generate high‐quality feature maps; fine integrating network (CFN‐Fine) uses fixed parameters from the CFN‐Coarse and combines features from deep to shallow in the deconvolution process by adding skip connections between down‐sampling and up‐sampling paths to efficiently integrate deep and shallow features. The saliency map obtained by the method is evaluated over 6 standard benchmark datasets, namely SALICON, MIT1003, MIT300, Toronto, OSIE, and SUN500. The results demonstrate that the method can surpass the state‐of‐the‐art accuracy of eye fixation prediction and achieves the competitive performance to date under most evaluation metrics on SALICON Saliency Prediction Challenge (LSUN2017). Binwei Xu, Haoran Liang 0001, Ronghua Liang, Peng Chen 0008 |
IET Image Process. | 1 |
| 2021 | Locate Globally, Segment Locally: A Progressive Architecture With Knowledge Review Network for Salient Object DetectionabstractSalient object location and segmentation are two different tasks in salient object detection (SOD). The former aims to globally find the most attractive objects in an image, whereas the latter can be achieved only using local regions that contain salient objects. However, previous methods mainly accomplish the two tasks simultaneously in a simple end-to-end manner, which leads to the ignorance of the differences between them. We assume that the human vision system orderly locates and segments objects, so we propose a novel progressive architecture with knowledge review network (PA-KRN) for SOD. It consists of three parts. (1) A coarse locating module (CLM) that uses body-attention label locates rough areas containing salient objects without boundary details. (2) An attention-based sampler highlights salient object regions with high resolution based on body-attention maps. (3) A fine segmenting module (FSM) finely segments salient objects. The networks applied in CLM and FSM are mainly based on our proposed knowledge review network (KRN) that utilizes the finest feature maps to reintegrate all previous layers, which can make up for the important information that is continuously diluted in the top-down path. Experiments on five benchmarks demonstrate that our single KRN can outperform state-of-the-art methods. Furthermore, our PA-KRN performs better and substantially surpasses the aforementioned methods. Binwei Xu, Haoran Liang 0001, Ronghua Liang, Peng Chen 0008 |
AAAI | 1 |
| 2021 | VSumVis: Interactive Visual Understanding and Diagnosis of Video Summarization ModelabstractWith the rapid development of mobile Internet, the popularity of video capture devices has brought a surge in multimedia video resources. Utilizing machine learning methods combined with well-designed features, we could automatically obtain video summarization to relax video resource consumption and retrieval issues. However, there always exists a gap between the summarization obtained by the model and the ones annotated by users. How to help users understand the difference, provide insights in improving the model, and enhance the trust in the model remains challenging in the current study. To address these challenges, we propose VSumVis under a user-centered design methodology, a visual analysis system with multi-feature examination and multi-level exploration, which could help users explore and analyze video content, as well as the intrinsic relationship that existed in our video summarization model. The system contains multiple coordinated views, i.e., video view, projection view, detail view, and sequential frames view. A multi-level analysis process to integrate video events and frames are presented with clusters and nodes visualization in our system. Temporal patterns concerning the difference between the manual annotation score and the saliency score produced by our model are further investigated and distinguished with sequential frames view. Moreover, we propose a set of rich user interactions that enable an in-depth, multi-faceted analysis of the features in our video summarization model. We conduct case studies and interviews with domain experts to provide anecdotal evidence about the effectiveness of our approach. Quantitative feedback from a user study confirms the usefulness of our visual system for exploring the video summarization model. Guodao Sun, Chaoqing Xu, Haoran Liang 0001, Binwei Xu, Ronghua Liang |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2020 | Video summarisation with visual and semantic cuesabstractVideo summarisation greatly improves the efficiency of people browsing videos and saves storage space. A good video summary should satisfy human visual interestingness and preserve the theme of the original video at the semantic level. Unlike many existing methods that consider only visual features to generate video summaries, this study proposes a method that combines visual and semantic cues to extract important information for dynamic video summarisation. The authors propose visual‐verbal saliency consistency to add semantic information and propose a novel attention motion, along with other visual features to fully represent visual interestingness. Based on the importance score of each frame calculated by combining these features, they select an optimal subset of segments to generate an important and interesting summary. They evaluate their method using the SumMe and TVSum datasets and experimental results show that their method generates high‐quality video summaries. Binwei Xu, Haoran Liang 0001, Ronghua Liang |
IET Image Process. | 1 |