VLDB 2026 Research / reviewers in the wild / expert
Yanbing Xue
dblp:96/7727
· DBLP profile ↗
32ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Novel Multi-View Perception and Shrinkage Aggregation Network for Inharmonious Region LocalizationabstractWith the popularity of image editing techniques, synthetic images may have inharmonious regions due to color/illumination differences between the manipulated area and the background. The inharmonious region localization task aims to find these regions, which is crucial for blind image harmonization. Existing methods rely on single-view images and do not fully explore multi-scale fusion, which limits their performance. To address these issues, in this paper, we propose a novel multi-view perception and shrinkage aggregation network (MSANet) for the inharmonious region localization task that fully utilizes multi-view images and multi-scale fusion information and can mine subtle cues between candidate objects and the background. Specifically, we first design a multi-view ensemble encoder to fully perceive the inharmonious regions by multi-view interactive learning and then aggregate the feature representations of inharmonious regions. Moreover, we propose a multi-scale shrinkage fusion decoder, where multi-scale features with multi-view prior information are utilized to aggregate adjacent features, adaptively select high-quality information, reduce background interference and gradually locate inharmonious regions. Extensive experimental results on four public datasets (HDobe5K, HCOCO, HFlickr, and Hday2Night) demonstrate that the proposed MSANet can outperform all the SOTA methods in terms of average F1 and average IoU score, while maintaining a lower computational cost1. Shenghao Chen, Chunjie Ma, Yibo Zhao 0001, Meng Liu 0006, Yanbing Xue, Zan Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | A Collaborative Hierarchical Aggregation Network for Weakly Supervised Temporal Action LocalizationabstractTemporal action localization is a fundamental task in video understanding that focuses on classifying and temporally localizing action instances in untrimmed videos. Compared to temporal action localization, the Weakly supervised Temporal Action Localization (WTAL) task presents greater challenges, as its training data lacks detailed information about action boundaries. Existing WTAL methods ignore the complementary relationship between modalities and the dependency between snippets, resulting in inaccurate localization results. To solve these issues, we propose a Collaborative Hierarchical Aggregation Network (CHA-Net). Specifically, we first use a modality complementary module to learn the synergies between modalities. Then, a collaborative enhance module is proposed to remove the information irrelevant to actions in RGB modality. Finally, a hierarchical aggregation module is proposed to capture the complete temporal information of action instances to better mine the temporal dependencies between snippets. Extensive experiments on THUMOS14, ActivityNet1.2, and ActivityNet1.3 datasets demonstrate the effectiveness of our method. Compared with F3-Net (TMM2024, Avg{0.1:0.5}) and SPCC-Net (TMM2024, Avg{0.1:0.7}) on the THUMOS14 dataset, the proposed method can achieve improvements of 3.2% and 2.4%, respectively. Zan Gao 0001, Xiaoyi Xu, Yibo Zhao 0001, Chunjie Ma, Yanbing Xue, Riwei Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | HeartOx: Efficient Multi-task Learning for Contactless Heart Rate and Blood Oxygen Estimation
Mengqi Wang, Mingyu Gu, Yanbing Xue |
ICIC (27) | 4 |
| 2025 | CACP: Covariance-Aware Cross-Domain Prototypes for Domain Adaptive Semantic SegmentationabstractDomain adaptive semantic segmentation aims to reduce domain shifts / discrepancies between source and target domains, improving the source domain model's generalization ability to the target domain. Recently, prototypical methods, which primarily use single-source or single-target domain prototypes as category centers to aggregate features from both domains, have achieved competitive performance in this task. However, due to large domain shifts, single-source domain prototypes have finite generalization ability and not all source domain knowledge is conducive to model generalization. Single-target domain prototypes are noisy because they are prematurely initialized with all features filtered by pseudo labels, which causes error accumulation in the prototypes. To address these issues, we propose a covariance-aware cross-domain prototypes method (CACP) to achieve robust domain adaptation. We propose to use both domain prototypes to dynamically rectify pseudo labels in the target domain, effectively reducing the recognition difficulty of hard target domain samples and narrowing the gap between features of the same category in both domains. In addition, to further generalize the model to the target domain, we propose two modules based on covariance correlation, FSPC (Features Selection by Prototypes Covariances) and WSPC (Weighting Source by Prototypes Coefficients), to learn discriminative characteristics. FSPC selects highly correlated features to update target domain prototypes online, denoising and enhancing discriminativeness between categories. WSPC utilizes the correlation coefficients between target domain prototypes and source domain features to weight each point in the source domain, eliminating the information interference from the source domain. In particular, CACP achieves excellent performance on the GTA5$\to$Cityscapes and SYNTHIA$\to$Cityscapes tasks with minimal computational resources and time. Yanbing Xue, Feifei Zhang 0001, Xianbin Wen, Zan Gao 0002, Shengyong Chen |
IEEE Trans. Multim. | 1 |
| 2024 | Segmentation and Quality Assessment of Continuous Fitness Movements Based on Vision
Zeying Li, Hongtao Chen, Yanbing Xue |
ICIC (11) | 4 |
| 2024 | X-CDNet: A real-time crosswalk detector based on YOLOX
Xingyuan Lu, Yanbing Xue, Xianbin Wen |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | Click-Conversion Multi-Task Model with Position Bias Mitigation for Sponsored Search in eCommerceabstractPosition bias, the phenomenon whereby users tend to focus on higher-ranked items of the search result list regardless of the actual relevance to queries, is prevailing in many ranking systems. Position bias in training data biases the ranking model, leading to increasingly unfair item rankings, click-through-rate (CTR), and conversion rate (CVR) predictions. To jointly mitigate position bias in both item CTR and CVR prediction, we propose two position-bias-free CTR and CVR prediction models: Position-Aware Click-Conversion (PACC) and PACC via Position Embedding (PACC-PE). PACC is built upon probability decomposition and models position information as a probability. PACC-PE utilizes neural networks to model product-specific position information as embedding. Experiments on the E-commerce sponsored product search dataset show that our proposed models have better ranking effectiveness and can greatly alleviate position bias in both CTR and CVR prediction. Yibo Wang 0001, Yanbing Xue, Bo Liu 0005, Musen Wen, Wenting Zhao 0006, Stephen D. Guo, Philip S. Yu |
SIGIR | 2 |
| 2023 | Hierarchical Active Learning With Qualitative Feedback on RegionsabstractLearning classification models in practice usually requires numerous labeled data for training. However, instance-based annotation can be inefficient for humans to perform. In this article, we propose and study a new type of human supervision that is fast to perform and useful for model learning. Instead of labeling individual instances, humans provide supervision to dataregions, which are subspaces of the input data space, representing subpopulations of data. Since labeling now is performed on a region level, 0/1 labeling becomes imprecise. Thus, we design the region label to be aqualitativeassessment of the class proportion, which coarsely preserves the labeling precision but is also easy for humans to do. To identify informative regions for labeling and learning, we further devise ahierarchical active learningprocess that recursively constructs a region hierarchy. This process is semisupervised in the sense that it is driven by both active learning strategies and human expertise, where humans can provide discriminative features. To evaluate our framework, we conducted extensive experiments on nine datasets as well as a real user study on a survival analysis of colorectal cancer patients. The results have clearly demonstrated the superiority of our region-based active learning framework against many instance-based active learning methods. Yazhou He, Yanbing Xue, Hongjun Wang 0002, Milos Hauskrecht, Tianrui Li 0001 |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2022 | Lightweight multi-scale convolutional neural network for real time stereo matching
Yanbing Xue, Doudou Zhang, Leida Li, Shiyin Li |
Image Vis. Comput. | 1 |
| 2021 | Parallel Cache Prefetching for LSM-Tree Based Store: From Algorithm to Evaluation
Guangping Xu, Yulei Jia, Yanbing Xue, Wenguang Zheng |
ICA3PP (1) | 4 |
| 2021 | A bus passenger re-identification dataset and a deep learning baseline using triplet embedding
Junliang Guo, Yanbing Xue, Zan Gao 0002, Guangping Xu, Hua Zhang 0003 |
Multim. Tools Appl. | 2 |
| 2020 | Attention Stereo Matching NetworkabstractDespite great progress, previous stereo matching algorithms still lack the ability to match textureless regions and slender structure areas. To tackle this problem, we propose ASM-Net, an attention stereo matching network. Attention module and disparity refinement module are constructed in the ASMNet. The attention module can improve correlation information between two images by channels and spatial attention. The feature-guided disparity refinement module learns more geometry information in different feature levels to refine the coarse prediction resolution constantly. The proposed approach was evaluated on several benchmark datasets. Experiments show that the proposed method achieves competitive results on KITTI and Scene-Flow datasets while running in real-time at 14ms. Doudou Zhang, Yanbing Xue, Hua Zhang 0003 |
ICPR | 3 |
| 2020 | Parsing human image by fusing semantic and spatial features: A deep learning approach
Ruilin Zhao, Yanbing Xue |
Inf. Process. Manag. | 2 |
| 2020 | 3D Object retrieval based on non-local graph neural networks
Yin-min Li, Ya-bin Tao, Yanbing Xue |
Multim. Tools Appl. | 5 |
| 2019 | Active Learning of Multi-Class Classification Models from Ordered Class SetsabstractIn this paper, we study the problem of learning multi-class classification models from a limited set of labeled examples obtained from human annotator. We propose a new machine learning framework that learns multi-class classification models from ordered class sets the annotator may use to express not only her top class choice but also other competing classes still under consideration. Such ordered sets of competing classes are common, for example, in various diagnostic tasks. In this paper, we first develop strategies for learning multi-class classification models from examples associated with ordered class set information. After that we develop an active learning strategy that considers such a feedback. We evaluate the benefit of the framework on multiple datasets. We show that class-order feedback and active learning can reduce the annotation cost both individually and jointly. Yanbing Xue, Milos Hauskrecht |
AAAI | 1 |
| 2019 | As-global-as-possible stereo matching with adaptive smoothness priorabstractMore global matching (MGM) overcomes the limitation of one‐dimensional scanline optimisation in semi‐global matching (SGM). Nevertheless, the possible weaknesses of the MGM algorithm are as follows: (i) only two directions are considered for each image traversal direction, which may lead to massive mismatches; (ii) disparity estimation around the object boundaries usually performs terrible since the smoothness term is designed independent of the image prior. In this research, the authors consider all of the four directions for each image traversal direction through a novel model. Besides utilising the prior of neighboured pixels' correlation, adaptive smoothness terms are modelled and augmented into the energy function. These contributions encourage ‘as‐global‐as‐possible (AGAP)’. More importantly, different from the recent works in which the aggregated algorithms have been conducted as the data term of an energy function, conversely, the authors make the energy function as a part of cost aggregation framework. Performance evaluations on Middlebury v.2 and v.3 stereo data sets demonstrate that the proposed AGAP outperforms other four most challenging stereo matching algorithms, and also performs better on Microsoft i2i stereo videos. In addition, under various strategies of parallelisation, the presented AGAP shows a near real‐time execution time. Hua Zhang 0003, Yanbing Xue, Shengyong Chen |
IET Image Process. | 3 |
| 2018 | AGO: Accelerating Global Optimization for Accurate Stereo Matching
Hua Zhang 0003, Yanbing Xue, Shengyong Chen |
MMM (1) | 3 |
| 2018 | MSCS: MeshStereo with Cross-Scale Cost Filtering for fast stereo matchingabstractMeshStereo (MS) and cross‐scale cost filtering (CSCF) are two most recently celebrated models for stereo matching. On one hand, MS model enlightens for fast solving the dense stereo correspondence problem according to a region‐based opinion. On the other hand, CSCF model could generate more robust matching cost volumes than single scale. In this study, the authors weave these two models together for attaining greater and faster disparity estimation. With CSCF, more powerful initial volumes of matching cost are computed and they are conducted as the data term of MS energy function model. More importantly, the novel‐fused stereo model also draws a closer connection between multi‐scale aggregated and global algorithms. Integrating the advantages of both stereo models, they name the presented one as MS with cross‐scale (MSCS). Performance evaluations on Middlebury v.2 and v.3 stereo data sets demonstrate that the proposed MSCS outperforms other four most challenging stereo matching algorithms; and also performs better on Microsoft i2i stereo videos. In addition, thanks to this novel‐fused model, MSCS requires fewer iteration times for optimising and makes it surprisingly possesses a much faster execution time. Hua Zhang 0003, Yanbing Xue, Shengyong Chen |
IET Comput. Vis. | 3 |
| 2018 | MMA: a multi-view and multi-modality benchmark dataset for human action recognition
Zan Gao 0002, Tao-tao Han, Hua Zhang 0003, Yanbing Xue, Guangping Xu |
Multim. Tools Appl. | 4 |
| 2018 | Semantic segmentation based on fusion of features and classifiers
Yanbing Xue, Huiqiang Geng, Hua Zhang 0003, Zhenshan Xue, Guangping Xu |
Multim. Tools Appl. | 1 |
| 2017 | Segment-tree based cost aggregation for stereo matching with enhanced segmentation advantageabstractSegment-tree (ST) based cost aggregation algorithm for stereo matching successfully integrates the information of segmentation with non-local cost aggregation framework. The tree structure which is generated by the segmentation strategy directly determines the final results for this kind of algorithms. However, the original strategy performs unreasonable due to its coarse performance and ignores to meet the disparity consistency assumption. To improve these weaknesses we propose a novel segmentation algorithm for constructing a more faithful ST with enhanced segmentation advantage according to a robust initial over-segmentation. Then we implement non-local cost aggregation framework on this new ST structure and obtain improved disparity maps. Performance evaluations on all 31 Middlebury stereo pairs show that the proposed algorithm outperforms than other five state-of-the-art aggregated based algorithms and also keeps time efficiency. Hua Zhang 0003, Yanbing Xue, Mian Zhou, Guangping Xu, Zan Gao 0002, Shengyong Chen |
ICASSP | 3 |
| 2017 | SPMVP: Spatial PatchMatch Stereo with Virtual Pixel Aggregation
Hua Zhang 0003, Yanbing Xue, Shengyong Chen |
ICONIP (3) | 3 |
| 2017 | Active Learning of Classification Models with Likert-Scale FeedbackabstractAnnotation of classification data by humans can be a time-consuming and tedious process. Finding ways of reducing the annotation effort is critical for building the classification models in practice and for applying them to a variety of classification tasks. In this paper, we develop a new active learning framework that combines two strategies to reduce the annotation effort. First, it relies on label uncertainty information obtained from the human in terms of the Likert-scale feedback. Second, it uses active learning to annotate examples with the greatest expected change. We propose a Bayesian approach to calculate the expectation and an incremental SVM solver to reduce the time complexity of the solvers. We show the combination of our active learning strategy and the Likert-scale feedback can learn classification models more rapidly and with a smaller number of labeled instances than methods that rely on either Likert-scale labels or active learning alone. Yanbing Xue, Milos Hauskrecht |
SDM | 1 |
| 2016 | REQUEST: A scalable framework for interactive construction of exploratory queriesabstractExploration over large datasets is a key first step in data analysis, as users may be unfamiliar with the underlying database schema and unable to construct precise queries that represent their interests. Such data exploration task usually involves executing numerous ad-hoc queries, which requires a considerable amount of time and human effort. In this paper, we present REQUEST, a novel framework that is designed to minimize the human effort and enable both effective and efficient data exploration. REQUEST supports the query-from-examples style of data exploration by integrating two key components: 1) Data Reduction, and 2) Query Selection. As instances of the REQUEST framework, we propose several highly scalable schemes, which employ active learning techniques and provide different levels of efficiency and effectiveness as guided by the user's preferences. Our results, on real-world datasets from Sloan Digital Sky Survey, show that our schemes on average require 1-2 orders of magnitude fewer feedback questions than the random baseline, and 3-16× fewer questions than the state-of-the-art, while maintaining interactive response time. Moreover, our schemes are able to construct, with high accuracy, queries that are often undetectable by current techniques. Xiaoyu Ge, Yanbing Xue, Mohamed A. Sharaf, Panos K. Chrysanthis |
IEEE BigData | 2 |
| 2016 | Learning of Classification Models from Noisy Soft-LabelsabstractWe develop and test a new classification model learning algorithm that relies on the soft-label information and that is able to learn classification models more rapidly and with a smaller number of labeled instances than existing approaches. Yanbing Xue, Milos Hauskrecht |
ECAI | 1 |
| 2016 | Iterative color-depth MST cost aggregation for stereo matchingabstractThe minimum spanning tree (MST) based non-local cost aggregation algorithm performs well in accuracy and time efficiency. However, it can still be improved in two aspects. First, we propose a logarithmic transformation on matching cost function to improve the matching efficiency in texture less regions. The textureless neighbors can provide effective contributions in cost aggregation by the proposed monotone increasing function. Hence the algorithm can distinguish different pixels in textureless regions. Second, MST algorithm only utilizes color information in weight function while aggregating, which leads 3D cues missing. We introduce depth weight computed from the original MST algorithm into an edge weight function. With the proposed color-depth weight, we further iteratively rebuild the tree and obtain enhanced disparity map. Performance evaluations on 19 Middlebury stereo pairs and Microsoft stereo videos show that the proposed algorithm outperforms than other five state-of-the-art cost aggregation algorithms. Hua Zhang 0003, Yanbing Xue, Mian Zhou, Guangping Xu, Zan Gao 0002 |
ICME | 3 |
| 2016 | A Fast 3D Retrieval Algorithm via Class-Statistic and Pair-Constraint ModelabstractWith the development of 3D technologies and devices, 3D model retrieval becomes a hot research topic where multi-view matching algorithms have demonstrated satisfying performance. However, exciting works overlook the common factors among objects in a single class, and they are time consuming in retrieval processing. In this paper, a class-statistics and pair-constraint model (CSPC) method is originally proposed for 3D model retrieval, which is composed of supervised class-based statistics model and pair-constraint object retrieval model. In our CSPC model, we firstly convert view-based distance measure into object-based distance measure without falling in performance, which will advance 3D model retrieval speed. Secondly, the generality of the distribution of each feature dimension in each class is computed to judge category information, and then we further adopt this distribution information to build class models. Finally, an object-based pairwise constraint is introduced on the base of the class-statistic measure, which can remove a lot of false alarm samples in retrieval. Experimental results on ETH, NTU-60, MVRED and PSB 3D datasets show that our method is fast, and its performance is also comparable with the-state-of-the-art algorithms. Zan Gao 0002, Hua Zhang 0003, Yanbing Xue, Guangping Xu |
ACM Multimedia | 4 |
| 2016 | Reverse Testing Image Set Model Based Multi-view Human Action Recognition
Yan Zhang 0154, Hua Zhang 0003, Guangping Xu, Yanbing Xue |
MMM (1) | 5 |
| 2016 | Human action recognition on depth dataset
Zan Gao 0002, Hua Zhang 0003, Anan Liu, Guangping Xu, Yanbing Xue |
Neural Comput. Appl. | 5 |
| 2015 | Single Face Image Super-Resolution via Multi-dictionary Bayesian Non-parametric Learning
Hua Zhang 0003, Yanbing Xue, Mian Zhou, Guangping Xu, Zan Gao 0002 |
ICONIP (1) | 3 |
| 2013 | Online Boosting Tracking with Fragmented Model
Dingcheng Shen, Hua Zhang 0003, Yanbing Xue, Guangping Xu, Zan Gao 0002 |
MMM (2) | 3 |
| 2012 | Human action recognition based on sparse representation induced by L1/L2 regulations
Zan Gao 0002, Anan Liu, Hua Zhang 0003, Guangping Xu, Yanbing Xue |
ICPR | 5 |