VLDB 2026 Research / reviewers in the wild / expert
Linbo Wang 0001
dblp:73/10697-1
· DBLP profile ↗
31ranked-venue papers
11as first author
21since 2021 · last 2026
0000-0001-7276-7065ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GTAvatar: Efficient garment transfer for aligned avatars simultaneously reconstructed from monocular videos
Xianyong Fang, Baofeng Zhou, Linbo Wang 0001, Zhengyi Liu |
Comput. Aided Geom. Des. | 4 |
| 2026 | Multi-view simulation for robust polyp segmentation via cross-gated decoding and soft-attention fusion
Linbo Wang 0001, Jinxian Qiu, Zhengyi Liu, Xianyong Fang, Shaohua Wan 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Text-Driven Medical Image Segmentation With LLM Semantic Bridge and LLM Prompt BridgeabstractText-driven medical image segmentation aims to accurately segment pathological regions in medical images based on textual descriptions. Existing methods face two major challenges: (a) The significant modality heterogeneity between textual and visual features leads to inefficient cross-modal feature alignment; (b) The insufficient utilization of medical shared knowledge restricts semantic understanding. To address these challenges, two large language model (LLM) bridges are constructed. LLM semantic bridge leverages the sequential modeling capability of a frozen LLM to reorganize visual features into semantically coherent units that possess linguistic logic, thereby effectively bridging vision and language. The LLM prompt bridge appends learnable prompts, which encode medical shared knowledge from the LLM, to text embeddings, thereby effectively bridging case-specificity and medical consensus knowledge. Experimental results show the predominant performance due to LLM participation.https://github.com/liuzywen/MedBridge Zhengyi Liu, Jiali Wu, Xianyong Fang, Linbo Wang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Embedding Space Decomposition Meets Invertible Networks: A New Paradigm for Unpaired Low-Light Enhancement
Linbo Wang 0001, Zhuo Yan, Zhengyi Liu, Xianyong Fang, Ping Li 0016 |
CGI (3) | 1 |
| 2025 | Discrepancy Correction Reweight Aggregation Federated Learning for Incomplete Multimodal of Brain Tumor Segmentation
Zhengyi Liu, Weikai Shi, Linbo Wang 0001, Xianyong Fang |
PRCV (14) | 3 |
| 2025 | Incomplete Multi-modal Brain Tumor Segmentation via Multi-expert Collaboration
Zhengyi Liu, Jinhai Yu, Xianyong Fang, Linbo Wang 0001 |
PRCV (14) | 5 |
| 2025 | DiffRSD: Diffusion-Based and Integrity-Aware RGB-D Rail Surface Defect Inspection
Zhengyi Liu, Junnan Zhou, Xianyong Fang, Zhengzheng Tu, Linbo Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | SSFam: Scribble Supervised Salient Object Detection FamilyabstractScribble supervised salient object detection (SSSOD) constructs segmentation ability of attractive objects from surroundings under the supervision of sparse scribble labels. For the better segmentation, depth and thermal infrared modalities serve as the supplement to RGB images in the complex scenes. Existing methods specifically design various feature extraction and multi-modal fusion strategies for RGB, RGB-Depth, RGB-Thermal, and Visual-Depth-Thermal image input respectively, leading to similar model flood. As the recently proposed Segment Anything Model (SAM) possesses extraordinary segmentation and prompt interactive capability, we propose an SSSOD family based on SAM, namedSSFam, for the combination input with different modalities. Firstly, different modal-aware modulators are designed to attain modal-specific knowledge which cooperates with modal-agnostic information extracted from the frozen SAM encoder for the better feature ensemble. Secondly, a siamese decoder is tailored to bridge the gap between the training with scribble prompt and the testing with no prompt for the stronger decoding ability. Our model demonstrates the remarkable performance among combinations of different modalities and refreshes the highest level of scribble supervised methods and comes close to the ones of fully supervised methods. Zhengyi Liu, Sheng Deng, Linbo Wang 0001, Xianyong Fang, Bin Tang 0003 |
IEEE Trans. Multim. | 4 |
| 2025 | SemanticAvatar: human surface reconstruction based on semantically consistent biplane features
Baofeng Zhou, Xianyong Fang, Linbo Wang 0001, Zhengyi Liu |
Vis. Comput. | 3 |
| 2024 | Towards Finer Human Reconstruction for Single RGB-D Images
Renlong Dai, Linbo Wang 0001, Zhengyi Liu, Xianyong Fang |
CGI (2) | 4 |
| 2024 | Label-aware Attention Network with Multi-scale Boosting for Medical Image Segmentation
Linbo Wang 0001, Peng Xu 0045, Xianfeng Cao, Michele Nappi, Shaohua Wan 0001 |
Expert Syst. Appl. | 1 |
| 2024 | LFSamba: Marry SAM With Mamba for Light Field Salient Object DetectionabstractA light field camera can reconstruct 3D scenes using captured multi-focus images that contain rich spatial geometric information, enhancing applications in stereoscopic photography, virtual reality, and robotic vision. In this work, a state-of-the-art salient object detection model for multi-focus light field images, called LFSamba, is introduced to emphasize four main insights: (a) Efficient feature extraction, where SAM is used to extract modality-aware discriminative features; (b) Inter-slice relation modeling, leveraging Mamba to capture long-range dependencies across multiple focal slices, thus extracting implicit depth cues; (c) Inter-modal relation modeling, utilizing Mamba to integrate all-focus and multi-focus images, enabling mutual enhancement; (d) Weakly supervised learning capability, developing a scribble annotation dataset from an existing pixel-level mask dataset, establishing the first scribble-supervised baseline for light field salient object detection. Zhengyi Liu, Longzhen Wang, Xianyong Fang, Zhengzheng Tu, Linbo Wang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2023 | Scribble-Supervised RGB-T Salient Object DetectionabstractSalient object detection segments attractive objects in scenes. RGB and thermal modalities provide complementary information and scribble annotations alleviate large amounts of human labor. Based on the above facts, we propose a scribble-supervised RGB-T salient object detection model. By a four-step solution (expansion, prediction, aggregation, and supervision), label-sparse challenge of scribble-supervised method is solved. To expand scribble annotations, we collect the superpixels that foreground scribbles pass through in RGB and thermal images, respectively. The expanded multi-modal labels provide the coarse object boundary. To further polish the expanded labels, we propose a prediction module to alleviate the sharpness of boundary. To play the complementary roles of two modalities, we combine the two into aggregated pseudo labels. Supervised by scribble annotations and pseudo labels, our model achieves the state-of-the-art performance on the relabeled RGBT-S dataset. Furthermore, the model is applied to RGB-D and video scribble-supervised applications, achieving consistently excellent performance.1 Zhengyi Liu, Xiaoshen Huang, Xianyong Fang, Linbo Wang 0001, Bin Tang 0003 |
ICME | 5 |
| 2023 | Sub-Band Based Attention for Robust Polyp SegmentationabstractThis article proposes a novel spectral domain based solution to the challenging polyp segmentation. The main contribution is based on an interesting finding of the significant existence of the middle frequency sub-band during the CNN process. Consequently, a Sub-Band based Attention (SBA) module is proposed, which uniformly adopts either the high or middle sub-bands of the encoder features to boost the decoder features and thus concretely improve the feature discrimination. A strong encoder supplying informative sub-bands is also very important, while we highly value the local-and-global information enriched CNN features. Therefore, a Transformer Attended Convolution (TAC) module as the main encoder block is introduced. It takes the Transformer features to boost the CNN features with stronger long-range object contexts. The combination of SBA and TAC leads to a novel polyp segmentation framework, SBA-Net. It adopts TAC to effectively obtain encoded features which also input to SBA, so that efficient sub-bands based attention maps can be generated for progressively decoding the bottleneck features. Consequently, SBA-Net can achieve the robust polyp segmentation, as the experimental results demonstrate. Xianyong Fang, Yuqing Shi, Linbo Wang 0001, Zhengyi Liu |
IJCAI | 4 |
| 2023 | Local Consensus Enhanced Siamese Network with Reciprocal Loss for Two-view Correspondence LearningabstractRecent studies of two-view correspondence learning usually establish an end-to-end network to jointly predict correspondence reliability and relative pose. We improve such a framework from two aspects. First, we propose a Local Feature Consensus (LFC) plugin block to augment the features of existing models. Given a correspondence feature, the block augments its neighboring features with mutual neighborhood consensus and aggregates them to produce an enhanced feature. As inliers obey a uniform cross-view transformation and share more consistent learned features than outliers, feature consensus strengthens inlier correlation and suppresses outlier distraction, which makes output features more discriminative for classifying inliers/outliers. Second, existing approaches supervise network training with the ground truth correspondences and essential matrix projecting one image to the other for an input image pair, without considering the information from the reverse mapping. We extend existing models to a Siamese network with a reciprocal loss that exploits the supervision of mutual projection, which considerably promotes the matching performance without introducing additional model parameters. Building upon MSA-Net [30], we implement the two proposals and experimentally achieve state-of-the-art performance on benchmark datasets. Linbo Wang 0001, Xianyong Fang, Zhengyi Liu, Chenjie Cao, Yanwei Fu 0001 |
ACM Multimedia | 1 |
| 2023 | Fine Back Surfaces Oriented Human Reconstruction for Single RGB-D ImagesabstractAbstract Current single RGB‐D image based human surface reconstruction methods generally take both the RGB images and the captured frontal depth maps together so that the 3D cues from the frontal surfaces can help infer the full surface geometries. However, we observe that the back surfaces can often be quite different from the frontal surfaces and, therefore, current methods can mess the recovery process by adopting such 3D cues, especially for the unseen back surfaces. We need to do the back surface inference without the frontal depth map. Consequently, a novel human reconstruction framework is proposed, so that human models with fine geometric details, especially for the back surfaces, can be obtained. In this approach, a progressive estimation method is introduced to effectively recover the unseen back depth maps. The coarse back depth maps are recovered by the parametric models of the subjects, with the fine ones further obtained by the normal‐maps conditioned GAN. This framework also includes a cross‐attention based denoising method for the frontal depth maps. This method adopts the cross attention between the features of the last two layers encoded from the frontal depth maps and thus suppresses the noise for fine depth maps by the attentions of features from the low‐noise and globally‐structured highest layer. Experimental results show the efficacies of the proposed ideas. Xianyong Fang, Jinshen He, Linbo Wang 0001, Zhengyi Liu |
Comput. Graph. Forum | 4 |
| 2023 | Parallel matters: Efficient polyp segmentation with parallel structured feature augmentation modulesabstractAbstract The large variations of polyp sizes and shapes and the close resemblances of polyps to their surroundings call for features with long‐range information in rich scales and strong discrimination. This article proposes two parallel structured modules for building those features. One is the Transformer Inception module (TI) which applies Transformers with different reception fields in parallel to input features and thus enriches them with more long‐range information in more scales. The other is the Local‐Detail Augmentation module (LDA) which applies the spatial and channel attentions in parallel to each block and thus locally augments the features from two complementary dimensions for more object details. Integrating TI and LDA, a new Transformer encoder based framework, Parallel‐Enhanced Network (PENet), is proposed, where LDA is specifically adopted twice in a coarse‐to‐fine way for accurate prediction. PENet is efficient in segmenting polyps with different sizes and shapes without the interference from the background tissues. Experimental comparisons with state‐of‐the‐arts methods show its merits. Xianyong Fang, Kaibing Wang, Yuqing Shi, Linbo Wang 0001, Enming Zhang, Zhengyi Liu |
IET Image Process. | 5 |
| 2023 | Adaptively feature matching via joint transformational-spatial clustering
Linbo Wang 0001, Xianyong Fang, Yanwen Guo 0001, Shaohua Wan 0001 |
Multim. Syst. | 1 |
| 2023 | LFTransNet: Light Field Salient Object Detection via a Learnable Weight DescriptorabstractLight Field Salient Object Detection (LF SOD) aims to segment the visually distinctive objects out of surroundings. Since light field images provide a multi-focus stack (many focal slices in different depth levels) and an all-focus image for the same scene, they record comprehensive but redundant information. Existing methods exploit the useful cue by long short-term memory with attention mechanism, 3D convolution, and graph learning. However, the importance of intra-slice and inter-slice in the focal stack is not well investigated. In the paper, we propose a learnable weight descriptor to simultaneously exploit different weights in slice, spatial region, and channel dimensions, and therefore propose an LF SOD method based on the learnable descriptor. The method extracts slice features and all-focus features from a weight-shared backbone and another backbone, respectively. A transformer decoder is used to learn the weight descriptor which both emphasizes the importance of each slice (inter-slice) and discriminates the spatial and channel importance of each slice (intra-slice). The learnt descriptor serves as the weight to make slice features attend to important slices, regions, and channels. Furthermore, we propose the hierarchical multi-modal fusion which aggregates high-layer features by modelling the long-range dependency to fully excavate common salient semantics and combines low-layer features by spatial constraint to eliminate the blurring effect of slice features. The experimental result exceeds the state-of-the-art methods at least 25% in terms of mean absolute error evaluation metric. It demonstrates a significant improvement in LF SOD performance via the designed learnable weight descriptor.https://github.com/liuzywen/LFTransNet Zhengyi Liu, Linbo Wang 0001, Xianyong Fang, Bin Tang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Locally Geometry-Aware Improvements of LOP for Efficient Skeleton Extraction
Xianyong Fang, Lingzhi Hu, Linbo Wang 0001 |
PRCV (3) | 4 |
| 2021 | Robust Shadow Detection by Exploring Effective Shadow ContextsabstractEffective contexts for separating shadows from non-shadow objects can appear in different scales due to different object sizes. This paper introduces a new module, Effective-Context Augmentation (ECA), to utilize these contexts for robust shadow detection with deep structures. Taking regular deep features as global references, ECA enhances the discriminative features from the parallelly computed fine-scale features and, therefore, obtains robust features embedded with effective object contexts by boosting them. We further propose a novel encoder-decoder style of shadow detection method where ECA acts as the main building block of the encoder to extract strong feature representations and the guidance to the classification process of the decoder. Moreover, the networks are optimized with only one loss, which is easy to train and does not have the instability caused by extra losses superimposed on the intermediate features among existing popular studies. Experimental results show that the proposed method can effectively eliminate fake detections. Especially, our method outperforms state-of-the-arts methods and improves over $13.97%$ and $34.67%$ on the challenging SBU and UCF datasets respectively in balance error rate. Xianyong Fang, Xiaohao He, Linbo Wang 0001, Jianbing Shen |
ACM Multimedia | 3 |
| 2020 | Geometry consistency aware confidence evaluation for feature matching
Linbo Wang 0001, Peng Xu 0045, Honglong Ren, Xianyong Fang, Shaohua Wan 0001 |
Image Vis. Comput. | 1 |
| 2019 | Single RGB-D Fitting: Total Human Modeling with an RGB-D ShotabstractExisting single shot based human modeling methods generally cannot model the complete pose details (e.g., head and hand positions) without non-trivial interactions. We explore the merits of both RGB and depth images and propose a new method called Single RGB-D Fitting (SRDF) to generate a realistic 3D human model with a single RGB-D shot from a consumer-grade depth camera. Specifically, the state-of-the-art deep learning techniques for RGB images are incorporated into SRDF, so that: 1) A compound skeleton detection method is introduced to obtain accurate 3D skeletons with refined hands based on the combination of depth and RGB images; and 2) an RGB image segmentation assisted point cloud pre-processing method is presented to obtain smooth foreground point clouds. In addition, several novel constraints are also introduced into the energy minimization model, including the shape continuity constraint, the keypoint-guided head pose prior constraint, and the penalty-enforced point cloud prior constraint. The energy model is optimized in a two-pass way so that a realistic shape can be estimated from coarse to fine. Through extensive experiments and comparisons with the state of the art methods, we demonstrate the effectiveness and efficiency of the proposed method. Xianyong Fang, Jikui Yang, Jie Rao, Linbo Wang 0001, Zhigang Deng 0001 |
VRST | 4 |
| 2019 | A unified two-parallel-branch deep neural network for joint gland contour and segmentation learning
Linbo Wang 0001, Hui Zhen, Xianyong Fang, Shaohua Wan 0001, Weiping Ding 0001, Yanwen Guo 0001 |
Future Gener. Comput. Syst. | 1 |
| 2019 | Viewpoint Assessment and Recommendation for Photographing ArchitecturesabstractThis paper studies the problem of how to assess the quality of photographing viewpoints and how to choose good viewpoints for taking photographs of architectures. We achieve this by learning from photographs of world famous landmarks that are available on the Internet and their viewpoint quality ranked by online user annotation. Unlike previous efforts devoted to photo quality assessment which mainly rely on 2D image features, we show in this paper combining 2D image features extracted from images with 3D geometric features computed on the 3D models can result in more reliable evaluation of viewpoint quality. Specifically, we collect a set of photographs for each of 15 world famous architectures as well as their 3D models from the Internet. Viewpoint recovery for images is carried out through an image-model registration process, after which a newly proposed viewpoint clustering strategy is exploited to validate users' viewpoint preferences when photographing landmarks. Finally, we extract a number of 2D and 3D features for each image based on multiple visual and geometric cues and perform viewpoint recommendation by learning from both 2D and 3D features using a specifically designed SVM-2K multi-view learner, achieving superior performance over using solely 2D or 3D features. We show the effectiveness of the proposed approach through extensive experiments. The experiments also demonstrate that our system can be used to recommend viewpoints for rendering textured 3D models of buildings for the use of architectural design, in addition to viewpoint evaluation of photographs and recommendation of viewpoints for photographing architectures in practice. Jingwu He, Linbo Wang 0001, Wenzhe Zhou, Hongjie Zhang 0002, Xiufen Cui, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | Single-Image Distance Measurement by a Smart Mobile DeviceabstractExisting distance measurement methods either require multiple images and special photographing poses or only measure the height with a special view configuration. We propose a novel image-based method that can measure various types of distance from single image captured by a smart mobile device. The embedded accelerometer is used to determine the view orientation of the device. Consequently, pixels can be back-projected to the ground, thanks to the efficient calibration method using two known distances. Then the distance in pixel is transformed to a real distance in centimeter with a linear model parameterized by the magnification ratio. Various types of distance specified in the image can be computed accordingly. Experimental results demonstrate the effectiveness of the proposed method. Shang-Wen Chen, Xianyong Fang, Jianbing Shen, Linbo Wang 0001, Ling Shao 0001 |
IEEE Trans. Cybern. | 4 |
| 2016 | Multi-view Metric Learning for Multi-view Video SummarizationabstractTraditional methods on video summarization are designed to generate summaries for single-view video records, and thus they cannot fully exploit the mutual information in multi-view video records. In this paper, we present a multiview metric learning framework for multi-view video summarization. It combines the advantages of maximum margin clustering with the disagreement minimization criterion. The learning framework thus has the ability to find a metric that best separates the input data, and meanwhile to force the learned metric to maintain underlying intrinsic structure of data points, for example geometric information. Facilitated by such a framework, a systematic solution to the multi-view video summarization problem is developed from the viewpoint of metric learning. The effectiveness of the proposed method is demonstrated by experiments. Linbo Wang 0001, Xianyong Fang, Yanwen Guo 0001, Yanwei Fu 0001 |
CW | 1 |
| 2015 | Common Visual Pattern Discovery via Nonlinear Mean Shift ClusteringabstractDiscovering common visual patterns (CVPs) from two images is a challenging task due to the geometric and photometric deformations as well as noises and clutters. The problem is generally boiled down to recovering correspondences of local invariant features, and the conventionally addressed by graph-based quadratic optimization approaches, which often suffer from high computational cost. In this paper, we propose an efficient approach by viewing the problem from a novel perspective. In particular, we consider each CVP as a common object in two images with a group of coherently deformed local regions. A geometric space with matrix Lie group structure is constructed by stacking up transformations estimated from initially appearance-matched local interest region pairs. This is followed by a mean shift clustering stage to group together those close transformations in the space. Joining regions associated with transformations of the same group together within each input image forms two large regions sharing similar geometric configuration, which naturally leads to a CVP. To account for the non-Euclidean nature of the matrix Lie group, mean shift vectors are derived in the corresponding Lie algebra vector space with a newly provided effective distance measure. Extensive experiments on single and multiple common object discovery tasks as well as near-duplicate image retrieval verify the robustness and efficiency of the proposed approach. Linbo Wang 0001, Yanwen Guo 0001, Minh N. Do |
IEEE Trans. Image Process. | 1 |
| 2014 | Confidence-driven image co-matting
Linbo Wang 0001, Tianchen Xia, Yanwen Guo 0001, Ligang Liu 0001, Jue Wang 0001 |
Comput. Graph. | 1 |
| 2014 | Video Object Co-Segmentation via Subspace Clustering and Quadratic Pseudo-Boolean Optimization in an MRF FrameworkabstractMultiple videos may share a common foreground object, for instance a family member in home videos, or a leading role in various clips of a movie or TV series. In this paper, we present a novel method for co-segmenting the common foreground object from a group of video sequences. The issue was seldom touched on in the literature. Starting from over-segmentation of each video into Temporal Superpixels (TSPs), we first propose a new subspace clustering algorithm which segments the videos into consistent spatio-temporal regions with multiple classes, such that the common foreground has consistent labels across different videos. The subspace clustering algorithm exploits the fact that across different videos the common foreground shares similar appearance features, while motions can be used to better differentiate regions within each video, making accurate extraction of object boundaries easier. We further formulate video object co-segmentation as a Markov Random Field (MRF) model which imposes the constraint of foreground model automatically computed or specified with little user effort. The Quadratic Pseudo-Boolean Optimization (QPBO) is used to generate the results. Experiments show that this video co-segmentation framework can achieve good quality foreground extraction results without user interaction for those videos with unrelated background, and with only moderate user interaction for those videos with similar background. Comparisons with previous work also show the superiority of our approach. Chuan Wang 0001, Yanwen Guo 0001, Linbo Wang 0001, Wenping Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2011 | Exploiting feature correspondence constraints for image recognitionabstractImage recognition is one of the fundamental problems in multimedia analysis. Typically in the training database, there will be more than one image for each object, however most existing bag-of-features based approaches treat them independently and completely ignore the feature correspondence relationship among them. As a result, features corresponding to the same physical point may be clustered into different clusters, which finally leads to inaccurate image representations for recognition. To tackle the problem, we present a supervised codebook construction algorithm exploiting the feature correspondence constraints in feature clustering. Features in different images of the same object are first matched, then ho-mography between images are computed to remove outliers as well as recover the feature correspondences that are not correctly matched. Features belonging to the same physical point are enforced to be in the same cluster. We show via experiments that codebook constructed using this approach can improve the recognition performance. Linbo Wang 0001, Yanwen Guo 0001, Suk Hwan Lim, Nelson L. Chang |
ICIP | 1 |