EDBT 2026 Demo / reviewers in the wild / expert
Xianyong Fang
dblp:05/4345 · also Xian-Yong Fang
· DBLP profile ↗
32ranked-venue papers
11as first author
21since 2021 · last 2026
0000-0002-6045-8430ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 8 first-author · 18 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GTAvatar: Efficient garment transfer for aligned avatars simultaneously reconstructed from monocular videos
Xianyong Fang, Baofeng Zhou, Linbo Wang 0001, Zhengyi Liu |
Comput. Aided Geom. Des. | 1 |
| 2026 | Multi-view simulation for robust polyp segmentation via cross-gated decoding and soft-attention fusion
Linbo Wang 0001, Jinxian Qiu, Zhengyi Liu, Xianyong Fang, Shaohua Wan 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Text-Driven Medical Image Segmentation With LLM Semantic Bridge and LLM Prompt BridgeabstractText-driven medical image segmentation aims to accurately segment pathological regions in medical images based on textual descriptions. Existing methods face two major challenges: (a) The significant modality heterogeneity between textual and visual features leads to inefficient cross-modal feature alignment; (b) The insufficient utilization of medical shared knowledge restricts semantic understanding. To address these challenges, two large language model (LLM) bridges are constructed. LLM semantic bridge leverages the sequential modeling capability of a frozen LLM to reorganize visual features into semantically coherent units that possess linguistic logic, thereby effectively bridging vision and language. The LLM prompt bridge appends learnable prompts, which encode medical shared knowledge from the LLM, to text embeddings, thereby effectively bridging case-specificity and medical consensus knowledge. Experimental results show the predominant performance due to LLM participation.https://github.com/liuzywen/MedBridge Zhengyi Liu, Jiali Wu, Xianyong Fang, Linbo Wang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Embedding Space Decomposition Meets Invertible Networks: A New Paradigm for Unpaired Low-Light Enhancement
Linbo Wang 0001, Zhuo Yan, Zhengyi Liu, Xianyong Fang, Ping Li 0016 |
CGI (3) | 4 |
| 2025 | Discrepancy Correction Reweight Aggregation Federated Learning for Incomplete Multimodal of Brain Tumor Segmentation
Zhengyi Liu, Weikai Shi, Linbo Wang 0001, Xianyong Fang |
PRCV (14) | 4 |
| 2025 | Incomplete Multi-modal Brain Tumor Segmentation via Multi-expert Collaboration
Zhengyi Liu, Jinhai Yu, Xianyong Fang, Linbo Wang 0001 |
PRCV (14) | 4 |
| 2025 | DiffRSD: Diffusion-Based and Integrity-Aware RGB-D Rail Surface Defect Inspection
Zhengyi Liu, Junnan Zhou, Xianyong Fang, Zhengzheng Tu, Linbo Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | SSFam: Scribble Supervised Salient Object Detection FamilyabstractScribble supervised salient object detection (SSSOD) constructs segmentation ability of attractive objects from surroundings under the supervision of sparse scribble labels. For the better segmentation, depth and thermal infrared modalities serve as the supplement to RGB images in the complex scenes. Existing methods specifically design various feature extraction and multi-modal fusion strategies for RGB, RGB-Depth, RGB-Thermal, and Visual-Depth-Thermal image input respectively, leading to similar model flood. As the recently proposed Segment Anything Model (SAM) possesses extraordinary segmentation and prompt interactive capability, we propose an SSSOD family based on SAM, namedSSFam, for the combination input with different modalities. Firstly, different modal-aware modulators are designed to attain modal-specific knowledge which cooperates with modal-agnostic information extracted from the frozen SAM encoder for the better feature ensemble. Secondly, a siamese decoder is tailored to bridge the gap between the training with scribble prompt and the testing with no prompt for the stronger decoding ability. Our model demonstrates the remarkable performance among combinations of different modalities and refreshes the highest level of scribble supervised methods and comes close to the ones of fully supervised methods. Zhengyi Liu, Sheng Deng, Linbo Wang 0001, Xianyong Fang, Bin Tang 0003 |
IEEE Trans. Multim. | 5 |
| 2025 | SemanticAvatar: human surface reconstruction based on semantically consistent biplane features
Baofeng Zhou, Xianyong Fang, Linbo Wang 0001, Zhengyi Liu |
Vis. Comput. | 2 |
| 2024 | Towards Finer Human Reconstruction for Single RGB-D Images
Renlong Dai, Linbo Wang 0001, Zhengyi Liu, Xianyong Fang |
CGI (2) | 6 |
| 2024 | FApSH: An effective and robust local feature descriptor for 3D registration and object recognition
Bao Zhao, Xiaobo Chen 0002, Xianyong Fang |
Pattern Recognit. | 4 |
| 2024 | LFSamba: Marry SAM With Mamba for Light Field Salient Object DetectionabstractA light field camera can reconstruct 3D scenes using captured multi-focus images that contain rich spatial geometric information, enhancing applications in stereoscopic photography, virtual reality, and robotic vision. In this work, a state-of-the-art salient object detection model for multi-focus light field images, called LFSamba, is introduced to emphasize four main insights: (a) Efficient feature extraction, where SAM is used to extract modality-aware discriminative features; (b) Inter-slice relation modeling, leveraging Mamba to capture long-range dependencies across multiple focal slices, thus extracting implicit depth cues; (c) Inter-modal relation modeling, utilizing Mamba to integrate all-focus and multi-focus images, enabling mutual enhancement; (d) Weakly supervised learning capability, developing a scribble annotation dataset from an existing pixel-level mask dataset, establishing the first scribble-supervised baseline for light field salient object detection. Zhengyi Liu, Longzhen Wang, Xianyong Fang, Zhengzheng Tu, Linbo Wang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | Scribble-Supervised RGB-T Salient Object DetectionabstractSalient object detection segments attractive objects in scenes. RGB and thermal modalities provide complementary information and scribble annotations alleviate large amounts of human labor. Based on the above facts, we propose a scribble-supervised RGB-T salient object detection model. By a four-step solution (expansion, prediction, aggregation, and supervision), label-sparse challenge of scribble-supervised method is solved. To expand scribble annotations, we collect the superpixels that foreground scribbles pass through in RGB and thermal images, respectively. The expanded multi-modal labels provide the coarse object boundary. To further polish the expanded labels, we propose a prediction module to alleviate the sharpness of boundary. To play the complementary roles of two modalities, we combine the two into aggregated pseudo labels. Supervised by scribble annotations and pseudo labels, our model achieves the state-of-the-art performance on the relabeled RGBT-S dataset. Furthermore, the model is applied to RGB-D and video scribble-supervised applications, achieving consistently excellent performance.1 Zhengyi Liu, Xiaoshen Huang, Xianyong Fang, Linbo Wang 0001, Bin Tang 0003 |
ICME | 4 |
| 2023 | Sub-Band Based Attention for Robust Polyp SegmentationabstractThis article proposes a novel spectral domain based solution to the challenging polyp segmentation. The main contribution is based on an interesting finding of the significant existence of the middle frequency sub-band during the CNN process. Consequently, a Sub-Band based Attention (SBA) module is proposed, which uniformly adopts either the high or middle sub-bands of the encoder features to boost the decoder features and thus concretely improve the feature discrimination. A strong encoder supplying informative sub-bands is also very important, while we highly value the local-and-global information enriched CNN features. Therefore, a Transformer Attended Convolution (TAC) module as the main encoder block is introduced. It takes the Transformer features to boost the CNN features with stronger long-range object contexts. The combination of SBA and TAC leads to a novel polyp segmentation framework, SBA-Net. It adopts TAC to effectively obtain encoded features which also input to SBA, so that efficient sub-bands based attention maps can be generated for progressively decoding the bottleneck features. Consequently, SBA-Net can achieve the robust polyp segmentation, as the experimental results demonstrate. Xianyong Fang, Yuqing Shi, Linbo Wang 0001, Zhengyi Liu |
IJCAI | 1 |
| 2023 | Local Consensus Enhanced Siamese Network with Reciprocal Loss for Two-view Correspondence LearningabstractRecent studies of two-view correspondence learning usually establish an end-to-end network to jointly predict correspondence reliability and relative pose. We improve such a framework from two aspects. First, we propose a Local Feature Consensus (LFC) plugin block to augment the features of existing models. Given a correspondence feature, the block augments its neighboring features with mutual neighborhood consensus and aggregates them to produce an enhanced feature. As inliers obey a uniform cross-view transformation and share more consistent learned features than outliers, feature consensus strengthens inlier correlation and suppresses outlier distraction, which makes output features more discriminative for classifying inliers/outliers. Second, existing approaches supervise network training with the ground truth correspondences and essential matrix projecting one image to the other for an input image pair, without considering the information from the reverse mapping. We extend existing models to a Siamese network with a reciprocal loss that exploits the supervision of mutual projection, which considerably promotes the matching performance without introducing additional model parameters. Building upon MSA-Net [30], we implement the two proposals and experimentally achieve state-of-the-art performance on benchmark datasets. Linbo Wang 0001, Xianyong Fang, Zhengyi Liu, Chenjie Cao, Yanwei Fu 0001 |
ACM Multimedia | 3 |
| 2023 | Fine Back Surfaces Oriented Human Reconstruction for Single RGB-D ImagesabstractAbstract Current single RGB‐D image based human surface reconstruction methods generally take both the RGB images and the captured frontal depth maps together so that the 3D cues from the frontal surfaces can help infer the full surface geometries. However, we observe that the back surfaces can often be quite different from the frontal surfaces and, therefore, current methods can mess the recovery process by adopting such 3D cues, especially for the unseen back surfaces. We need to do the back surface inference without the frontal depth map. Consequently, a novel human reconstruction framework is proposed, so that human models with fine geometric details, especially for the back surfaces, can be obtained. In this approach, a progressive estimation method is introduced to effectively recover the unseen back depth maps. The coarse back depth maps are recovered by the parametric models of the subjects, with the fine ones further obtained by the normal‐maps conditioned GAN. This framework also includes a cross‐attention based denoising method for the frontal depth maps. This method adopts the cross attention between the features of the last two layers encoded from the frontal depth maps and thus suppresses the noise for fine depth maps by the attentions of features from the low‐noise and globally‐structured highest layer. Experimental results show the efficacies of the proposed ideas. Xianyong Fang, Jinshen He, Linbo Wang 0001, Zhengyi Liu |
Comput. Graph. Forum | 1 |
| 2023 | Parallel matters: Efficient polyp segmentation with parallel structured feature augmentation modulesabstractAbstract The large variations of polyp sizes and shapes and the close resemblances of polyps to their surroundings call for features with long‐range information in rich scales and strong discrimination. This article proposes two parallel structured modules for building those features. One is the Transformer Inception module (TI) which applies Transformers with different reception fields in parallel to input features and thus enriches them with more long‐range information in more scales. The other is the Local‐Detail Augmentation module (LDA) which applies the spatial and channel attentions in parallel to each block and thus locally augments the features from two complementary dimensions for more object details. Integrating TI and LDA, a new Transformer encoder based framework, Parallel‐Enhanced Network (PENet), is proposed, where LDA is specifically adopted twice in a coarse‐to‐fine way for accurate prediction. PENet is efficient in segmenting polyps with different sizes and shapes without the interference from the background tissues. Experimental comparisons with state‐of‐the‐arts methods show its merits. Xianyong Fang, Kaibing Wang, Yuqing Shi, Linbo Wang 0001, Enming Zhang, Zhengyi Liu |
IET Image Process. | 2 |
| 2023 | Adaptively feature matching via joint transformational-spatial clustering
Linbo Wang 0001, Xianyong Fang, Yanwen Guo 0001, Shaohua Wan 0001 |
Multim. Syst. | 3 |
| 2023 | LFTransNet: Light Field Salient Object Detection via a Learnable Weight DescriptorabstractLight Field Salient Object Detection (LF SOD) aims to segment the visually distinctive objects out of surroundings. Since light field images provide a multi-focus stack (many focal slices in different depth levels) and an all-focus image for the same scene, they record comprehensive but redundant information. Existing methods exploit the useful cue by long short-term memory with attention mechanism, 3D convolution, and graph learning. However, the importance of intra-slice and inter-slice in the focal stack is not well investigated. In the paper, we propose a learnable weight descriptor to simultaneously exploit different weights in slice, spatial region, and channel dimensions, and therefore propose an LF SOD method based on the learnable descriptor. The method extracts slice features and all-focus features from a weight-shared backbone and another backbone, respectively. A transformer decoder is used to learn the weight descriptor which both emphasizes the importance of each slice (inter-slice) and discriminates the spatial and channel importance of each slice (intra-slice). The learnt descriptor serves as the weight to make slice features attend to important slices, regions, and channels. Furthermore, we propose the hierarchical multi-modal fusion which aggregates high-layer features by modelling the long-range dependency to fully excavate common salient semantics and combines low-layer features by spatial constraint to eliminate the blurring effect of slice features. The experimental result exceeds the state-of-the-art methods at least 25% in terms of mean absolute error evaluation metric. It demonstrates a significant improvement in LF SOD performance via the designed learnable weight descriptor.https://github.com/liuzywen/LFTransNet Zhengyi Liu, Linbo Wang 0001, Xianyong Fang, Bin Tang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Locally Geometry-Aware Improvements of LOP for Efficient Skeleton Extraction
Xianyong Fang, Lingzhi Hu, Linbo Wang 0001 |
PRCV (3) | 1 |
| 2021 | Robust Shadow Detection by Exploring Effective Shadow ContextsabstractEffective contexts for separating shadows from non-shadow objects can appear in different scales due to different object sizes. This paper introduces a new module, Effective-Context Augmentation (ECA), to utilize these contexts for robust shadow detection with deep structures. Taking regular deep features as global references, ECA enhances the discriminative features from the parallelly computed fine-scale features and, therefore, obtains robust features embedded with effective object contexts by boosting them. We further propose a novel encoder-decoder style of shadow detection method where ECA acts as the main building block of the encoder to extract strong feature representations and the guidance to the classification process of the decoder. Moreover, the networks are optimized with only one loss, which is easy to train and does not have the instability caused by extra losses superimposed on the intermediate features among existing popular studies. Experimental results show that the proposed method can effectively eliminate fake detections. Especially, our method outperforms state-of-the-arts methods and improves over $13.97%$ and $34.67%$ on the challenging SBU and UCF datasets respectively in balance error rate. Xianyong Fang, Xiaohao He, Linbo Wang 0001, Jianbing Shen |
ACM Multimedia | 1 |
| 2020 | Geometry consistency aware confidence evaluation for feature matching
Linbo Wang 0001, Peng Xu 0045, Honglong Ren, Xianyong Fang, Shaohua Wan 0001 |
Image Vis. Comput. | 5 |
| 2020 | Text Image Deblurring Using Kernel Sparsity PriorabstractPrevious methods on text image motion deblurring seldom consider the sparse characteristics of the blur kernel. This paper proposes a new text image motion deblurring method by exploiting the sparse properties of both text image itself and kernel. It incorporates the L0-norm for regularizing the blur kernel in the deblurring model, besides the L0sparse priors for the text image and its gradient. Such a L0-norm-based model is efficiently optimized by half-quadratic splitting coupled with the fast conjugate descent method. To further improve the quality of the recovered kernel, a structure-preserving kernel denoising method is also developed to filter out the noisy pixels, yielding a clean kernel curve. Experimental results show the superiority of the proposed method. The source code and results are available at: https://github.com/shenjianbing/text-image-deblur. Xianyong Fang, Jianbing Shen, Christian Jacquemin, Ling Shao 0001 |
IEEE Trans. Cybern. | 1 |
| 2019 | Single RGB-D Fitting: Total Human Modeling with an RGB-D ShotabstractExisting single shot based human modeling methods generally cannot model the complete pose details (e.g., head and hand positions) without non-trivial interactions. We explore the merits of both RGB and depth images and propose a new method called Single RGB-D Fitting (SRDF) to generate a realistic 3D human model with a single RGB-D shot from a consumer-grade depth camera. Specifically, the state-of-the-art deep learning techniques for RGB images are incorporated into SRDF, so that: 1) A compound skeleton detection method is introduced to obtain accurate 3D skeletons with refined hands based on the combination of depth and RGB images; and 2) an RGB image segmentation assisted point cloud pre-processing method is presented to obtain smooth foreground point clouds. In addition, several novel constraints are also introduced into the energy minimization model, including the shape continuity constraint, the keypoint-guided head pose prior constraint, and the penalty-enforced point cloud prior constraint. The energy model is optimized in a two-pass way so that a realistic shape can be estimated from coarse to fine. Through extensive experiments and comparisons with the state of the art methods, we demonstrate the effectiveness and efficiency of the proposed method. Xianyong Fang, Jikui Yang, Jie Rao, Linbo Wang 0001, Zhigang Deng 0001 |
VRST | 1 |
| 2019 | A Color-Pair Based Approach for Accurate Color Harmony EstimationabstractAbstract Harmonious color combinations can stimulate positive user emotional responses. However, a widely open research question is: how can we establish a robust and accurate color harmony measure for the public and professional designers to identify the harmony level of a color theme or color set. Building upon the key discovery that color pairs play an important role in harmony estimation, in this paper we present a novel color‐pair based estimation model to accurately measure the color harmony. It first takes a two‐layer maximum likelihood estimation (MLE) based method to compute an initial prediction of color harmony by statistically modeling the pair‐wise color preferences from existing datasets. Then, the initial scores are refined through a back‐propagation neural network (BPNN) with a variety of color features extracted in different color spaces, so that an accurate harmony estimation can be obtained at the end. Our extensive experiments, including performance comparisons of harmony estimation applications, show the advantages of our method in comparison with the state of the art methods. Bailin Yang, Tianxiang Wei, Xianyong Fang, Zhigang Deng 0001, Frederick W. B. Li, Xun Wang 0007 |
Comput. Graph. Forum | 3 |
| 2019 | A unified two-parallel-branch deep neural network for joint gland contour and segmentation learning
Linbo Wang 0001, Hui Zhen, Xianyong Fang, Shaohua Wan 0001, Weiping Ding 0001, Yanwen Guo 0001 |
Future Gener. Comput. Syst. | 3 |
| 2017 | Single-Image Distance Measurement by a Smart Mobile DeviceabstractExisting distance measurement methods either require multiple images and special photographing poses or only measure the height with a special view configuration. We propose a novel image-based method that can measure various types of distance from single image captured by a smart mobile device. The embedded accelerometer is used to determine the view orientation of the device. Consequently, pixels can be back-projected to the ground, thanks to the efficient calibration method using two known distances. Then the distance in pixel is transformed to a real distance in centimeter with a linear model parameterized by the magnification ratio. Various types of distance specified in the image can be computed accordingly. Experimental results demonstrate the effectiveness of the proposed method. Shang-Wen Chen, Xianyong Fang, Jianbing Shen, Linbo Wang 0001, Ling Shao 0001 |
IEEE Trans. Cybern. | 2 |
| 2016 | Multi-view Metric Learning for Multi-view Video SummarizationabstractTraditional methods on video summarization are designed to generate summaries for single-view video records, and thus they cannot fully exploit the mutual information in multi-view video records. In this paper, we present a multiview metric learning framework for multi-view video summarization. It combines the advantages of maximum margin clustering with the disagreement minimization criterion. The learning framework thus has the ability to find a metric that best separates the input data, and meanwhile to force the learned metric to maintain underlying intrinsic structure of data points, for example geometric information. Facilitated by such a framework, a systematic solution to the multi-view video summarization problem is developed from the viewpoint of metric learning. The effectiveness of the proposed method is demonstrated by experiments. Linbo Wang 0001, Xianyong Fang, Yanwen Guo 0001, Yanwei Fu 0001 |
CW | 2 |
| 2014 | A consistent pixel-wise blur measure for partially blurred imagesabstractDespite numerous efforts on blur measurement of partially blurred images, there still lacks an effective blur measure that is both pixel-wise and locally sharp consistent. The paper proposes a novel method with two contributions to overcome this limitation: 1) A new pixel-based blur metric, Multi-resolution Singular Value (MSV), which leverages the average singular value of high frequency bands to measure the blur of each pixel, and 2) a locally continuous strategy, maximum-likelihood estimation (MLE) based refinement, that ensures local continuity by imposing the local sharp consistency on pixel blur in a local correcting process. Experimental results show that our method is effective to smoothly measure the partially blurred images without local discontinuity. Xianyong Fang, Yanwen Guo 0001, Christian Jacquemin, Jian Zhou 0006, Shanchun Huang |
ICIP | 1 |
| 2013 | Fast window fusion using fuzzy equivalence relation
Xianyong Fang, Jian Zhou 0006 |
Pattern Recognit. Lett. | 1 |
| 2009 | Registration of blurred images for image mosaicabstractExisting methods for the registration of blurred images are efficient for the artificially blurred images or a planar registration. They are not suitable for image mosaic of the source images from a real camera with an almost fixed optical center. We propose a registration method so that a distortion-free registration on naturally captured images can be obtained. It adopts a multi-resolution and robust feature based inter-layer mosaic together. In each layer, Harris corner detector is chosen to effectively detect features and RANSAC is used to find reliable matches for further calibration as well as an initial homography as the initial motion of next layer. Simplex and subspace trust region methods are used consequently to estimate the stable focal length and rotation matrix through the transformation property of feature matches. Experimental results demonstrate the performance of our proposed method. Xianyong Fang, Bin Luo 0001, Jin Tang 0001, Haifeng Zhao 0001 |
CAD/Graphics | 1 |
| 2006 | A New Method of Manifold Mosaic for Large Displacement Images
Xianyong Fang |
J. Comput. Sci. Technol. | 1 |