VLDB 2026 Research / reviewers in the wild / expert
Xing Zhao 0001
dblp:87/2635-1
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0003-0870-8524ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SCLSTE: Semi-supervised Contrastive Learning-Guided Scene Text EditingabstractThe objective of scene text editing is to substitute the existing text with desirable text, while preserving the background and text styles intact. However, existing methods struggle to effectively replicate textual styles due to the complexity of real scenarios, and most are limited to training exclusively on labeled synthetic datasets. One alternative semi-supervised method incorporates unlabeled real-scene text images for training by using the original image as supervision and editing it with the same text. However, this approach risks degrading the model into an identity mapping network. To address these problems, we introduce a novel semi-supervised training strategy incorporating contrastive learning. It allows for editing real-scene text images with any text, circumventing the identity mapping issue while ensuring the accuracy of both text content and style. Moreover, we propose a robust Style-Aware Text Editing Module to address complexity of real scenarios and enhance the imitation of text styles. To the best of our knowledge, our work is the first to apply contrastive learning to the scene text editing. Extensive experiments demonstrate that our method outperforms existing models in terms of quality and quantity. Especially for the HEval metrics on both real scene (Tamper-Scene) and synthetic scene (Temper-Syn2k), we get both 12% improvement compared to state-of-the-art method. Min Yin, Liang Xie 0003, Haoran Liang 0001, Xing Zhao 0001, Ben Chen 0004, Ronghua Liang |
MMM (3) | 4 |
| 2025 | Semantic-aware representations for unsupervised Camouflaged Object Detection
Zelin Lu, Xing Zhao 0001, Liang Xie 0003, Haoran Liang 0001, Ronghua Liang |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | A Weakly-Supervised Cross-Domain Query Framework for Video Camouflage Object DetectionabstractVCOD (Video Camouflage Object Detection) is a crucial security technology that identifies camouflaged objects in videos, bolstering security measures across diverse applications. On one hand, appearance-based VCOD methods face challenges because camouflaged appearances cause objects to blend into their surroundings, and current VCOD methods typically utilize optical flow to represent motion information. However, over-reliance on accurate estimation renders the model overly fragile. On the other hand, there is a shortage of effectively annotated camouflaged video datasets, coupled with the time-consuming and labor-intensive annotation process, severely constraining the development of this field. To address this, we propose a novel weakly-supervised framework for VCOD based on cross-domain querying of preceding and succeeding frames. Specifically, we propose a time-efficient and labor-saving manual annotation approach based on large visual models to rapidly generate pseudo-labels. Furthermore, we design a network based on Spatio-Temporal Memory (STM) that performs cross-modal feature querying with the current frame against preceding and succeeding frames to acquire useful information, thereby enhancing the focus on temporal information. Extensive experiments conducted on two common VCOD datasets have proven the effectiveness of our method, achieving state-of-the-art performance on the challenging camouflaged video data. Zelin Lu, Liang Xie 0003, Xing Zhao 0001, Binwei Xu, Haoran Liang 0001, Ronghua Liang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Multidimensional Exploration of Segment Anything Model for Weakly Supervised Video Salient Object DetectionabstractFully supervised video salient object detection (VSOD) has made considerable breakthroughs using costly and time-consuming pixel-wise annotations. Recently, to achieve a trade-off between the annotation burden and the model performance, scribble-based VSOD tasks have attracted increasing attention. However, learning the complete object structure and precise boundary details from sparse scribble annotations remains challenging. In this paper, we propose a series of strategies to effectively explore valid information from the recently proposed segmentation foundation model “Segment Anything Model (SAM)” in various perspectives to address these challenges. Specifically, due to the limited performance of SAM on videos, we propose a SAM-guided label enhancement method instead of directly using the results of SAM, which can introduce edge information while reducing the interference of erroneous information. Moreover, we propose a SAM-driven spatiotemporal network guided by general semantic features from the SAM encoder to help the model be aware of global connections. Additionally, we propose a SAM-based global-aware loss, which further considers the affinity constraint between predicted results and foreground labels or background labels from a global perspective, guiding the model to perceive the complete salient objects. Experimental results demonstrate that our method outperforms state-of-the-art weakly supervised VSOD methods and is comparable to fully supervised VSOD methods. Binwei Xu, Qiuping Jiang, Xing Zhao 0001, Chenyang Lu 0002, Haoran Liang 0001, Ronghua Liang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Position Fusing and Refining for Clear Salient Object DetectionabstractMultilevel feature fusion plays a pivotal role in salient object detection (SOD). High-level features present rich semantic information but lack object position information, whereas low-level features contain object position information but are mixed with noises such as backgrounds. Appropriately addressing the gap between low- and high-level features is important in SOD. We first propose a global position embedding attention (GPEA) module to minimize the discrepancy between multilevel features in this article. We extract the position information by utilizing the semantic information at high-level features to resist noises at low-level features. Object refine attention (ORA) module is introduced to refine features used to predict saliency maps further without any additional supervision and heighten discriminative regions near the salient object, such as boundaries. Moreover, we find that the saliency maps generated by the previous methods contain some blurry regions, and we design a pixel value (PV) loss to help the model generate saliency maps with improved clarity. Experimental results on five commonly used SOD datasets demonstrated that the proposed method is effective and outperforms the state-of-the-art approaches on multiple metrics. Xing Zhao 0001, Haoran Liang 0001, Ronghua Liang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Hierarchical multi-modal video summarization with dynamic samplingabstractAbstract Previous video summarization methods often neglected inter‐frame variations during the preprocessing stage. Sampling repeated frames can lead to information redundancy, while missing key frames can result in deviations in semantic comprehension and inaccuracies in the generated summaries. This work proposes a dynamic sampling module that leverages frame‐level motion information to alleviate these issues. The module conducts high‐frequency sampling during intervals with significant changes, allowing for a finer capture of details. Combined with a hierarchical multi‐modal structure, it integrates shot‐level visual and textual information to enhance the semantic understanding of video clips and improve the accuracy of the summarized content. Extensive experiments on benchmark datasets SumMe and TVSum demonstrate the effectiveness of the proposed method. Lingjian Yu, Xing Zhao 0001, Liang Xie 0003, Haoran Liang 0001, Ronghua Liang |
IET Image Process. | 2 |
| 2024 | Salient object detection in egocentric videosabstractAbstract In the realm of video salient object detection (VSOD), the majority of research has traditionally been centered on third‐person perspective videos. However, this focus overlooks the unique requirements of certain first‐person tasks, such as autonomous driving or robot vision. To bridge this gap, a novel dataset and a camera‐based VSOD model, CaMSD , specifically designed for egocentric videos, is introduced. First, the SalEgo dataset, comprising 17,400 fully annotated frames for video salient object detection, is presented. Second, a computational model that incorporates a camera movement module is proposed, designed to emulate the patterns observed when humans view videos. Additionally, to achieve precise segmentation of a single salient object during switches between salient objects, as opposed to simultaneously segmenting two objects, a saliency enhancement module based on the Squeeze and Excitation Block is incorporated. Experimental results show that the approach outperforms other state‐of‐the‐art methods in egocentric video salient object detection tasks. Dataset and codes can be found at https://github.com/hzhang1999/SalEgo . Haoran Liang 0001, Xing Zhao 0001, Jian Liu 0053, Ronghua Liang |
IET Image Process. | 3 |
| 2024 | Motion-Aware Memory Network for Fast Video Salient Object DetectionabstractPrevious methods based on 3DCNN, convLSTM, or optical flow have achieved great success in video salient object detection (VSOD). However, these methods still suffer from high computational costs or poor quality of the generated saliency maps. To address this, we design a space-time memory (STM)-based network that employs a standard encoder-decoder architecture. During the encoding stage, we extract high-level temporal features from the current frame and its adjacent frames, which is more efficient and practical than methods reliant on optical flow. During the decoding stage, we introduce an effective fusion strategy for both spatial and temporal branches. The semantic information of the high-level features is used to improve the object details in the low-level features. Subsequently, spatiotemporal features are methodically derived step by step to reconstruct the saliency maps. Moreover, inspired by the boundary supervision prevalent in image salient object detection (ISOD), we design a motion-aware loss that predicts object boundary motion, and simultaneously perform multitask learning for VSOD and object motion prediction. This can further enhance the model's capability to accurately extract spatiotemporal features while maintaining object integrity. Extensive experiments on several datasets demonstrate the effectiveness of our method and can achieve state-of-the-art metrics on some datasets. Our proposed model does not require optical flow or additional preprocessing, and can reach an impressive inference speed of nearly 100 FPS. Xing Zhao 0001, Haoran Liang 0001, Guodao Sun, Ronghua Liang, Xiaofei He 0001 |
IEEE Trans. Image Process. | 1 |