EDBT 2026 Demo / reviewers in the wild / expert
Xiuli Shao
dblp:29/6194
· DBLP profile ↗
29ranked-venue papers
0as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prototype-based multi-view fine-grained 3D classification and ad-hoc interpretability
Shuxian Ma, Runmin Cong, Sam Kwong, Xiuli Shao |
Pattern Recognit. | 5 |
| 2026 | Incorporating Uncertainty-Guided and Top-k Codebook Matching for Real-World Blind Image Super-ResolutionabstractRecent advancements in codebook-based real image super-resolution (SR) have shown promising results in real-world applications. The core idea involves matching high-quality image features from a codebook based on low-resolution (LR) image features. However, existing methods face two major challenges: inaccurate feature matching with the codebook and poor texture detail reconstruction. To address these issues, we propose a novel Uncertainty-Guided and Top-k Codebook Matching SR (UGTSR) framework, which incorporates three key components: 1) an uncertainty learning mechanism that guides the model to focus on texture-rich regions, 2) a Top-k feature matching strategy that enhances feature matching accuracy by fusing multiple candidate features, and 3) an Align-Attention module that enhances the alignment of information between LR and HR features. Experimental results demonstrate significant improvements in texture realism and reconstruction fidelity compared to existing methods. The source code can be found at https://github.com/wwlCape/UGTSR-main. Weilei Wen, Zhaohui Zheng 0003, Chunle Guo, Xiuli Shao, Chongyi Li |
IEEE Trans. Image Process. | 6 |
| 2025 | Uncertainty-Guided Feature Learning Network for Accurate Medical Image Segmentation
Xiao-Xue Sun, Xiuli Shao, Yanding Qin, Hongpeng Wang 0001 |
ICIC (28) | 2 |
| 2025 | Efficient Scale-Uniform 3D Visual Coverage Algorithm for UAV Based on Elastic Photogrammetric ConstraintsabstractUnmanned aerial vehicles equipped with modern vision algorithms are crucial for missions such as reconstruction and target acquisition. However, when deployed in the field, undulating terrain can cause significant fluctuations in image scale and degrade the performance of vision algorithms. Instead of developing specialized image processing schemes with limited adaptability, this paper presents a novel 3D visual coverage algorithm that is compatible with existing generic vision algorithms and maintains a uniform image scale for ground targets. In detail, photogrammetric constraints are initially introduced to generate aerial waypoints, and then the negative effects of valley clustering are addressed. Elastic Photogrammetric Constraints (EPC) are further proposed to eliminate valley clustering effects induced by saddle terrain. The experimental results demonstrate that EPC reduces the traversal path length by up to 37.38 % compared to the previous work, but with a minor trade-off in scale variations. Jianping Zong, Zhongzhi Cao, Xiuli Shao, Haifeng Li 0008 |
ICRA | 5 |
| 2025 | Diffusion-Guided Domain-Adaptive Segmentation with Structure-Aware Learning via Retinal Image Noise ConditioningabstractTransferring the style from source domain to target domain for learning target models is a widely used strategy in domain adaptive segmentation. Although diffusion-based image translation has enabled flexible style transfer, it is often difficult to maintain the original structure of the image realistically during the reverse diffusion, which provides very little control over the generated image. To tackle this issue, we present the diffusion-based approach toward domain adaptive segmentation of general retinal image, which conditions diffusion models with carefully crafted input noise artifacts as explicit guidance at the inference step. Concretely, the input cross-domain image and the segmentation map of source domain are merged by summing the output of two encoders. Then, the encoder-decoder framework is adopted to iteratively refine the segmentation map by using a diffusion model. In order to enhance the general understanding of target domain distribution, we also establish the frequency-adaptive conditions for each sampling step. Moreover, this paper takes the channel-wise information and coarse semantic mask with noise of target image as guidance in the denoising process, which is different from existing approaches that input Gaussian noise and further establishes controllable conditions at the inference step. Extensive experiments on domain adaptation (DA)-based retinal image segmentation demonstrate the superiority of our approach over some state-of-the-art methods. Jinping Li, Runmin Cong, Xiuli Shao |
IJCNN | 5 |
| 2024 | Uncertainty-weighted prototype active learning in domain adaptive semantic segmentation
Sijie Niu, Xizhan Gao, Jinping Li, Xiuli Shao |
Expert Syst. Appl. | 5 |
| 2024 | GroupTransNet: Group transformer network for RGB-D salient object detection
Xian Fang, Mingfeng Jiang, Jinchao Zhu, Xiuli Shao |
Neurocomputing | 4 |
| 2024 | Coarse-to-fine online latent representations matching for one-stage domain adaptive semantic segmentation
Sijie Niu, Xizhan Gao, Xiuli Shao |
Pattern Recognit. | 4 |
| 2024 | UAFer: A Unified Model for Class-Agnostic Binary Segmentation With Uncertainty-Aware Feature ReassemblyabstractClass-agnostic binary segmentation identifies objects that are similar or very different from the complex background, including salient object detection (SOD) and camouflage object detection (COD). Most existing models only focus on a specific type of foreground and background segmentation by employing the global modeling ability of transformers, without explicitly explaining or eliminating the discrepancy between these two different distributions. They also suffer from inefficient local feature learning and inadequate feature aggregation. To make binary segmentation research more accessible and trivially generalized, we introduce a novel unified uncertainty-aware paradigm, called uncertainty-aware feature reassembly (UAFer). Specifically, the Spatial Feature Reassembly (SFR) module is presented to formulate the uncertainty of binary segmentation map as the variance of generalized Bernoulli distribution and entropy from two perspectives. Our transformer-based model is then trained to prioritize regions of higher certainty, obtaining more confident and accurate predictions during the feature upsampling. Moreover, the Channel Feature Reassembly (CFR) with adjacent feature aggregation is designed to facilitate an iterative exploration of channel integrity. This iterative learning process enhances the interaction of neighboring channel features; thus, improving universal object information decoding efficiency. Extensive quantitative and qualitative evaluations demonstrate that our proposed UAFer consistently outperforms the state-of-the-art models across three challenging domains including SOD, COD, and polyp segmentation (POLYP). The implementation codes for our approach will be publicly available at https://github.com/zihaodong/UAFR. Zizhen Liu, Runmin Cong, Tiyu Fang, Xiuli Shao, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Adaptive Blind Super-Resolution Network for Spatial-Specific and Spatial-Agnostic DegradationsabstractPrior methodologies have disregarded the diversities among distinct degradation types during image reconstruction, employing a uniform network model to handle multiple deteriorations. Nevertheless, we discover that prevalent degradation modalities, including sampling, blurring, and noise, can be roughly categorized into two classes. We classify the first class as spatial-agnostic dominant degradations, less affected by regional changes in image space, such as downsampling and noise degradation. The second class degradation type is intimately associated with the spatial position of the image, such as blurring, and we identify them as spatial-specific dominant degradations. We introduce a dynamic filter network integrating global and local branches to address these two degradation types. This network can greatly alleviate the practical degradation problem. Specifically, the global dynamic filtering layer can perceive the spatial-agnostic dominant degradation in different images by applying weights generated by the attention mechanism to multiple parallel standard convolution kernels, enhancing the network's representation ability. Meanwhile, the local dynamic filtering layer converts feature maps of the image into a spatially specific dynamic filtering operator, which performs spatially specific convolution operations on the image features to handle spatial-specific dominant degradations. By effectively integrating both global and local dynamic filtering operators, our proposed method outperforms state-of-the-art blind super-resolution algorithms in both synthetic and real image datasets. Weilei Wen, Chunle Guo, Wenqi Ren, Hongpeng Wang 0001, Xiuli Shao |
IEEE Trans. Image Process. | 5 |
| 2023 | Weakly supervised fine-grained semantic segmentation via spatial correlation-guided learning
Tiyu Fang, Jinping Li, Xiuli Shao |
Comput. Vis. Image Underst. | 4 |
| 2023 | M2RNet: Multi-modal and multi-scale refined network for RGB-D salient object detection
Xian Fang, Mingfeng Jiang, Jinchao Zhu, Xiuli Shao |
Pattern Recognit. | 4 |
| 2022 | Depth Removal Distillation for RGB-D Semantic SegmentationabstractRGB-D semantic segmentation is attracting wide attention due to its better performance than conventional RGB methods. However, most of RGB-D semantic segmentation methods need to acquire the real depth information for segmenting RGB images effectively. Therefore, it is extremely challenging to take full advantage of RGB-D semantic segmentation methods for segmenting RGB images without the depth input. To address this challenge, a general depth removal distillation method is proposed to remove depth dependence from RGB-D semantic segmentation model by knowledge distillation, which can be employed to any CNN-based segmentation network structure. Specifically, a depth-aware convolution is adopted to construct the teacher network for getting sufficient knowledge from RGB-D images. Then according to the structure consistency between depth-aware convolution and general convolution, the teacher network is used to transfer the learned knowledge to the student network with general convolutions by sharing parameters. Next, the student network makes up for the lack of depth in manner of learning by RGB images. Meantime, a Variable Temperature Cross Entropy (VTCE) loss function is proposed to further increase the accuracy of the student model by soft target distillation. Extensive experiments on NYUv2 and SUN RGB-D datasets demonstrate the superiority of our proposed approach. Tiyu Fang, Xiuli Shao, Jinping Li |
ICASSP | 3 |
| 2022 | LC3Net: Ladder context correlation complementary network for salient object detection
Xian Fang, Jinchao Zhu, Xiuli Shao, Hongpeng Wang 0001 |
Knowl. Based Syst. | 3 |
| 2022 | Wavelet-Based Texture Reformation Network for Image Super-ResolutionabstractMost reference-based image super-resolution (RefSR) methods directly leverage the raw features extracted from a pretrained VGG encoder to transfer the matched texture information from a reference image to a low-resolution image. We argue that simply operating on these raw features neglects the influence of irrelevant and redundant information and the importance of abundant high-frequency representations, leading to undesirable texture matching and transfer results. Taking the advantages of wavelet transformation, which represents the contextual and textural information of features at different scales, we propose a Wavelet-based Texture Reformation Network (WTRN) for RefSR. We first decompose the extracted texture features into low-frequency and high-frequency sub-bands and conduct feature matching on the low-frequency component. Based on the correlation map obtained from the feature matching process, we then separately swap and transfer wavelet-domain features at different stages of the network. Furthermore, a wavelet-based texture adversarial loss is proposed to make the network generate more visually plausible textures. Experiments on four benchmark datasets demonstrate that our proposed method outperforms previous RefSR methods both quantitatively and qualitatively. The source code is available at https://github.com/zskuang58/WTRN-TIP. Zhen Li 0031, Zengsheng Kuang, Zuo-Liang Zhu, Hongpeng Wang 0001, Xiuli Shao |
IEEE Trans. Image Process. | 5 |
| 2021 | Self-supervised Multi-view Clustering for Unsupervised Image Segmentation
Tiyu Fang, Xiuli Shao, Jinping Li |
ICANN (5) | 3 |
| 2021 | A Plug and Play Fast Intersection Over Union Loss for Boundary Box RegressionabstractBounding box regression is a very effective method to improve the localization accuracy of object detection. Recently, the IoU-based regression losses have been widely used in object detection algorithms. However, we observe that they degenerate seriously in the late training period, leading to slow convergence and inaccurate localization. In this paper, we design a Fast Intersection over Union (FIoU) loss, which can not only keep the advantages but also solve the weakness of IoU-based losses. Furthermore, FIoU can be directly applied to Non-Maximum Suppression (NMS) as a criterion to improve the localization performance. Numerous experiments on two popular benchmark datasets show that our method is superior to other the-state-of-art methods. Zengsheng Kuang, Xian Fang, Ruixun Zhang, Xiuli Shao |
ICASSP | 4 |
| 2021 | IBNet: Interactive Branch Network for salient object detection
Xian Fang, Jinchao Zhu, Ruixun Zhang, Xiuli Shao, Hongpeng Wang 0001 |
Neurocomputing | 4 |
| 2021 | Lightweight boundary refinement module based on point supervision for semantic segmentation
Jinping Li, Tiyu Fang, Xiuli Shao |
Image Vis. Comput. | 4 |
| 2021 | Subspace Clustering with Block Diagonal Sparse Representation
Xian Fang, Ruixun Zhang, Xiuli Shao |
Neural Process. Lett. | 4 |
| 2021 | Collaborative learning in bounding box regression for object detection
Xian Fang, Zengsheng Kuang, Ruixun Zhang, Xiuli Shao, Hongpeng Wang 0001 |
Pattern Recognit. Lett. | 4 |
| 2020 | Scale-Recursive Network with point supervision for crowd scene analysis
Ruixun Zhang, Xiuli Shao |
Neurocomputing | 3 |
| 2020 | Learning sparse features with lightweight ScatterNet for small sample training
Ruixun Zhang, Xiuli Shao, Zengsheng Kuang |
Knowl. Based Syst. | 3 |
| 2019 | Multi-scale Discriminative Location-Aware Network for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation methods aim for predicting the regions of different object categories with only a few labeled samples. It is difficult to produce segmentation results with high accuracy when a new category appears. In this paper, we propose a Multi-scale Discriminative Location-aware (MDL) network to tackle the few-shot semantic segmentation problem. In order to use information from different levels, we first keep the last three convolutional layers of FCN, and then use the VGG-16 network to extract features from the support image-label pair, which adjusts the weight of the query image segmentation branch. Discriminative location-aware architecture can improve the efficiency of few-shot segmentation, and therefore the global average pooling layer is added to produce location feature information. Finally, we evaluate our MDL model on the Pascal VOC 2012 challenge, and show that it achieves competitive mIoU score compared to methods in recent years. Ruixun Zhang, Xiuli Shao |
COMPSAC (2) | 3 |
| 2019 | Learning Deep Structured Multi-scale Features for Crisp and Object Occlusion Edge Detection
Ruixun Zhang, Xiuli Shao |
ICANN (3) | 3 |
| 2019 | A CNN-RNN Hybrid Model with 2D Wavelet Transform Layer for Image ClassificationabstractConvolutional neural networks (CNNs) have recently achieved impressive performances in image processing tasks such as image classification and object recognition. However, CNNs only process images in the spatial domain whereas spectral analysis operates in the frequency domain. In this paper, we propose the 2D wavelet transform layer. The learned features from images are viewed as two-directional sequential data, and we use two LSTM layers that sweep both horizontally and vertically across the image to compress feature matrices. Based on this, the 2D wavelet transform decomposes the above feature matrices as a learned mixing of different harmonic functions, and therefore integrating the spectral analysis into CNNs. We also select 3 × 3 convolutional mixing style with Gaussian+LSM filter to mix the output of the 2D wavelet transform to generate new output features. Finally, we combine the sequential and spectral features to build our CNN-RNN architecture with skip layers and apply it to image classification. Our proposed network is evaluated on three widely-used benchmark datasets: CIFAR-10, CIFAR-100 and Tiny ImageNet. Experiments show that our CNN-RNN hybrid model achieves better accuracy in image classification tasks. Ruixun Zhang, Xiuli Shao |
ICTAI | 3 |
| 2018 | A New Combined CNN-RNN Model for Sector Stock Price AnalysisabstractThe combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) have played important roles in deep learning in recent years to improve the prediction performance, especially in the context of temporal data analysis. Previous research has shown that certain time series could have common time-dependent characteristics. Therefore, in order to make good prediction, it is necessary to take into account the correlation between different temporal data in modeling. However, general RNN models have serious limitation to achieve this goal. In this paper, a new architecture, Deep and Wide Neural Networks (DWNN), is proposed, where CNN's convolution layer is added to the RNN's hidden state transfer process. CNN is combined with RNN to extract the correlation characteristics of different RNN models while RNNs running along the time steps. This new architecture not only has the depth of RNN in the time dimension, but also has the width of the number of temporal data. The intuition behind the DWNN model, as well as different kinds of DWNN model structures are discussed in this paper. We use stock data from the sandstorm sector of Shanghai Stock Exchange for our experiment. As shown in the result, our proposed DWNN model can reduce the prediction mean squared error by 30% compared with the general RNN model. Ruixun Zhang, Zhaozheng Yuan, Xiuli Shao |
COMPSAC (2) | 3 |
| 2018 | Multi-scale Feature Decode and Fuse Model with CRF Layer for Boundary Detection
Ruixun Zhang, Xiuli Shao, Huichao Li |
ICONIP (2) | 3 |
| 2014 | Detecting P2P botnets by discovering flow dependency in C&C traffic
Hongling Jiang, Xiuli Shao |
Peer-to-Peer Netw. Appl. | 2 |