EDBT 2026 Demo / reviewers in the wild / expert
Lianlei Shan
dblp:291/8922
· DBLP profile ↗
17ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-4648-8246ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-SoftCoT: A Self-Consistent Framework via Position-Aware Latent Space Reinforcement LearningabstractWhile Chain-of-Thought (CoT) reasoning empowers Large Language Models (LLMs) to tackle complex tasks, its reliance on discrete token decoding imposes an inherent Discreteness Bottleneck, limiting expressiveness within a restricted vocabulary space.Existing continuous reasoning approaches, such as SoftCoT (Xu et al., 2025), mitigate this but typically rely on external auxiliary models, resulting in complex deployment and fractured inference pipelines.To address these challenges, we propose Self-SoftCoT, a self-contained framework that enables a frozen LLM to internally generate and consume latent thoughts without external assistants.By establishing a singlestream "Thinking → Speaking" closed-loop, we decouple latent planning from explicit generation.Furthermore, we adopt Group Sequence Policy Optimization (GSPO) to stabilize learning and employ Position-Aware Independent Projection to mitigate representation homogenization.Experimental results on five reasoning benchmarks demonstrate that our method significantly improves the reasoning performance of frozen LLMs.Specifically, our Qwen2.5-basedmodel (Yang et al., 2024) uses only N = 2 soft tokens to outperform the Soft-CoT baseline (N = 4), improving the average accuracy from 75.06% to 78.42%.Similarly, LLaMA-3.1 (Llama Team, 2024) performance increases from 70.52% to 74.55% 1 . Liangliang Dong, Lianlei Shan, Shuaimin Li |
ACL (1) | 2 |
| 2025 | A Scene Text Detection Method Based on Supervised Contrastive Learning
Jinhong Huang, Hongrong Yin, Lianlei Shan, Subrota K. Mondal |
ICANN (2) | 3 |
| 2025 | Asymmetric Mamba-CNN Collaborative Architecture for Large-Size Remote Sensing Image Semantic SegmentationabstractLarge-size remote sensing images contain rich geographical information. Efficient and accurate semantic segmentation of these images is of significant importance in various fields. However, the massive memory requirements have hindered the development of semantic segmentation methods for large-size remote sensing images. Most existing methods struggle to balance memory usage, global modeling, and local representation accuracy. To address these issues, we propose a new semantic segmentation method for large-size remote sensing images, Mamba–CNN parallel network (MCPNet), which demonstrates impressive performance. The method is an asymmetric Mamba–convolutional neural network (CNN) hybrid architecture. Given the linear modeling complexity of Mamba, we construct the M-branch based on the visual state space (VSS) model, which processes downsampled images to reduce memory consumption while alleviating Mamba’s local forgetting problem. To further enhance the model’s capability in fine-grained detail extraction, we meticulously design a detail-preserving network (DPN) as the C-branch. This branch employs a split downsampling strategy and multiscale convolutional kernel groups to process large-size images, ensuring the preservation of spatial positional relationships while capturing fine-grained local details. Moreover, to effectively filter redundant information introduced by large-size images and bridge the semantic gap between the features extracted by CNN and Mamba, we propose a multigated feature fusion module (MG-FFM). This module progressively refines heterogeneous feature alignment through a bottom-up hierarchical refinement strategy, achieving a progressive fusion of semantics and details. Our method achieves state-of-the-art (SOTA) performance in terms of mean intersection over union (mIoU) and mF1 score on the self-constructed Yaan UAV dataset and two widely used public datasets (DeepGlobe and Inria Aerial) while consuming less GPU memory. The codes will be available athttps://github.com/fsqy-zhang/MCPNet Min Chen 0015, Lianlei Shan, Caiyi Li, Han Hu 0005, Xuming Ge, Qing Zhu 0012, Bo Xu 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | LDNET: Semantic Segmentation Of High-Resolution Images Via Learnable Patch Proposal And Dynamic RefinementabstractDue to the requirement of large GPU memory, semantic segmentation of ultra-high-resolution images faces many challenges. Previous methods are to downsample or crop into patches to fit the GPU’s memory limitations, which results in the loss of detailed information or context. To solve this problem, we propose a multi-branch network called Learnable Patch Proposal and Dynamic Refinement Network (LD-Net) specifically for semantic segmentation of high-resolution aerial images. Instead of simply putting all patches into the network for fusion, we calculate the uncertainties of each patch to select the most valuable ones to avoid invalid operations, thus improving efficiency. As for accuracy, we propose a module called Dynamic Refinement. This module uses global features (from downsampling images) as queries and local features (from patches) as keys and values when fusing features. Vice versa can also be established. In this way, different types of features can be effectively integrated. Experimental results show that our method achieves remarkable improvements over the previous state-of-the-art methods on two benchmark datasets DeepGlobe and ISIC. Yuyang Ji, Lianlei Shan |
ICME | 2 |
| 2024 | Continual Learning for Image Segmentation With Dynamic QueryabstractImage segmentation based on continual learning exhibits a critical drop of performance, mainly due to catastrophic forgetting and background shift, as they are required to incorporate new classes continually. In this paper, we propose a simple, yet effective Continual Image Segmentation method with incremental Dynamic Query (CISDQ), which decouples the representation learning of both old and new knowledge with lightweight query embedding. CISDQ mainly includes three contributions: 1) We definedynamic querieswith adaptive background class to exploit past knowledge and learn future classes naturally. 2) CISDQ proposes a class/instance-aware Query Guided Knowledge Distillation strategy to overcome catastrophic forgetting by capturing the inter-class diversity and intra-class identity. 3) Apart from semantic segmentation, CISDQ introduce the continual learning forinstance segmentationin which instance-wise labeling and supervision are considered. Extensive experiments on three datasets for two tasks (i.e. continual semantic and instance segmentation are conducted to demonstrate that CISDQ achieves the state-of-the-art performance, specifically, obtaining 4.4% and 2.9% mIoU improvements for the ADE 100-10 (6 steps) setting and ADE 100-5 (11 steps) setting. Weijia Wu 0001, Yuzhong Zhao, Zhuang Li 0002, Lianlei Shan, Zheng Shou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | End-to-End Remote Sensing Change Detection of Unregistered Bi-temporal Images for Natural Disasters
Guiqin Zhao, Lianlei Shan |
ICANN (2) | 2 |
| 2023 | A Data-Related Patch Proposal for Semantic Segmentation of Aerial ImagesabstractLarge-size images cannot be directly put into GPU for training and need to be cropped to patches due to GPU memory limitation. The commonly used cropping methods before are random cropping and sequential cropping, which are crude and fatally inefficient. Firstly, categories of datasets are often imbalanced, and just simple cropping misses an excellent opportunity to make the data distribution balanced. Secondly, the training needs to crop a large number of patches to cover all patterns, which greatly increases the training time. This problem is of great practical hazards but is often overlooked by previous works. The optimal solution is to generate valuable patches. Valuable patches refer to the value to network training, i.e., the value of this patch for the convergence of the network, and the improvement of the accuracy. To this end, we propose a data-related patch proposal strategy to sample high valuable patches. The core idea is to score each patch according to the accuracy of each category, so as to perform balanced sampling. Compared with random cropping or sequential cropping, our method can improve the segmentation accuracy and accelerate the training vastly. Moreover, our method also shows great advantages over the loss-based balanced approaches. Experiments on Deepglobe and Potsdam show the excellent effect of our method. Lianlei Shan, Guiqin Zhao, Jun Xie 0003, Peirui Cheng, Xiaobin Li 0006, Zhepeng Wang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | Boosting Semantic Segmentation of Aerial Images via Decoupled and Multilevel Compaction and DispersionabstractSemantic segmentation is a valuable task in practical applications for aerial images. Nevertheless, the segmentation performance is unsatisfactory due to aerial images’ huge intra-class variance and inter-class similarity. To solve this problem, we propose an approach to increase the distinction between classes and compact the features of the same class. Specifically, since a single aerial image contains only a small number of categories, which is fatal for previous contrastive learning, we discard InfoNCE loss in contrastive learning and use the simple Mean Square Error (MSE) loss that does not require negative samples to decouple the dispersion and compaction operations. Besides, we set up more representative prototypes for classes and extend the prototypes to the whole dataset level, which we call image- and dataset-level prototypes. Based on the calculated prototypes, we propose Multi-level intra-class Feature Compaction (MFC) and Multi-level inter-class Feature Dispersion (MFD) to compact the features of the same class and disperse the features of different classes in the latent feature space. More importantly, some measures are proposed to ensure the two do not conflict. MFC and MFD can be applied to any existing segmentation network to improve performance significantly without increasing computational complexity during inference. Moreover, we feed the calculated multi-level prototypes directly into the classifier, thus keeping the feature extraction and classifier consistent. Results on four challenging datasets, Deepglobe, iSAID, Potsdam, and Vaihingen, demonstrate the significant effect of our method, and sufficient ablation studies verify the role of each module. Lianlei Shan, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | MBNet: A Multi-Resolution Branch Network for Semantic Segmentation Of Ultra-High Resolution ImagesabstractSemantic segmentation of ultra-high resolution images is more challenging than ordinary images since high-resolution images need to be cropped into patches in training due to GPU memory limitation. To solve this problem, we design a multibranch structure to deal with multi-resolution inputs, called Multi-resolution Branch Network (MBNet). MBNet takes patches of various instead of only one resolution as inputs, so it can make the extracted features pure and different from each other so as to cover the complex scenes with tremendous variation. Moreover, to make full use of the multi-branch structure, we design a zoom module. Zoom module abandons the previous L2-norm before feature concatenation but combines different features according to the learned attention, which fully releases the advantages of multi-resolution. Results on two benchmark datasets show that our method improves significantly over the previous state-of-the-art methods. Lianlei Shan, Weiqiang Wang 0001 |
ICASSP | 1 |
| 2022 | DenseNet-Based Land Cover Classification Network With Deep FusionabstractRecently, fully convolutional network (FCN)-based (Longet al., 2015) networks have made impressive success in semantic segmentation, and these approaches achieve satisfactory results in natural images. However, in the field of high-resolution remote sensing image segmentation, the accuracy has a considerable huge gap compared with that of natural images. Through the development process of semantic segmentation, we found that the key to accurate segmentation is the context. Effective networks can always obtain large contexts, which means that context is the key to one successful segmentation network. For high-resolution remote sensing images, their elements always extend to large scope and they have no clear or regular boundaries. As a result, it needs more context to correctly classify each pixel. However, the networks designed for natural images obviously do not meet this requirement, and thus, they achieve poor segmentation results for high-resolution images. Therefore, we do some targeted improvements. Based on one powerful backbone, we add two new fusions called unit fusion and cross-level fusion, respectively. Unit fusion makes the connection from the encoder part to the decoder part not only occur in the final output of each dense block but also in the middle feature layers inside one dense block. These added fusions make feature fusion in the same level more complete, which is of great significance for complex and zigzag boundary areas. As a complement for unit fusion, cross-level fusion aims to enhance the fusion of different dense blocks. Specifically, cross-level fusion learns from the internal structure of the dense block and applies the design to the whole network level. It can incorporate nonadjacent features and rapidly increase the receptive field and context, which is very effective for the segmentation of targets with very large sizes. Experiments on Deepglobe (Demiret al., 2018) show significant improvements in our work. Lianlei Shan, Weiqiang Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Class-Incremental Learning for Semantic Segmentation in Aerial Imagery via Distillation in All AspectsabstractIncremental learning using neural networks achieves great success in semantic segmentation but still suffers from catastrophic forgetting. In this article, we propose an effective class-incremental segmentation method without storing old data. To alleviate the issue of forgetting, we present two important modules, i.e., the deep feature distillation (DFD) module and the label mixed (LM) module. The DFD module is established to learn a good feature representation of old classes by distilling a new compact feature representation from different layers of networks. The proposed LM module first identifies the examples (pixels) of old classes with high confidences utilizing the output of old models, and then, they are combined with examples of new classes to supervise the training of new models, which can achieve a good balance between learning new classing and avoiding forgetting old ones. Our ablation studies show that the DFD module and the LM module can make the learning network obtain 6.2% and 15% performance gains [mean Intersection over Union (mIOU)], respectively. Furthermore, by introducing the supervision of output distillation loss, we compare our method with several state-of-the-art methods in the extensive experiments, and the experimental results all show that our method is significantly superior to them on the dataset of aerial images. Lianlei Shan, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Class-Incremental Semantic Segmentation of Aerial Images via Pixel-Level Feature Generation and Task-Wise DistillationabstractDeep neural networks achieve significant progress in semantic segmentation but still suffer from the catastrophic forgetting problem, i.e., networks will forget old classes as they learn new ones. In this article, we propose an effective class-incremental segmentation framework without storing old data. Specifically, to alleviate the issue of catastrophic forgetting, we present two important modules, i.e., the Pixel-level Feature Generation (PFG) module, and the Task-wise Knowledge distillation (TKD) module. The PFG module is designed to constantly generate any number of features of the old classes to keep the old memory. The PFG module is the first attempt to use the generative method in class-incremental segmentation of aerial images, and it abandons the previous image generation approach but to generate pixel-level features, which is more suitable for the segmentation task. Meanwhile, the proposed TKD module is specially designed for class incremental tasks, and it only compares classes in the same learning step (task), thus avoiding the squeezing of new classes to old classes when the output is normalized (softmax), making distillation more effective. Sufficient experiments show that our method is remarkably effective and achieves more than 4.5% gains compared with state-of-the-art methods, and more than 13% compared with baselines, on all learning conditions. The ablation studies show that the PFG module and the TKD module are both indispensables. Besides, the proposed framework can be well combined with any existing class incremental learning method to achieve better performance. Lianlei Shan, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Fusing Multitask Models by Recursive Least SquaresabstractIt is easy to obtain multi-tasking models from open source platforms or various organizations. However, using these models at the same time will bring a great burden on storage and reduce computing efficiency. In this paper, we propose a transformation-based multi-task fusion method, called transformation fusion(TF), which is implemented by recursive least squares. The recursive transformation fusion not only reduces the storage burden brought by models fusion but also avoids computing the inverse matrix of high-dimensional matrix. Our multi-model fusion method can also be applied to many mainstream tasks, such as multi-task learning and offline distributed learning. Our fusion method can be applied to data-based fusion tasks as well as data-free fusion tasks. Extensive experiments demonstrate the effectiveness of our fusion method. Xiaobin Li 0006, Lianlei Shan, Weiqiang Wang 0001 |
ICASSP | 2 |
| 2021 | Decouple the High-Frequency and Low-Frequency Information of Images for Semantic SegmentationabstractAs a special kind of signal processing technology, image processing has been developed rapidly after the appearance of convolutional neural network (CNN). At present, the semantic segmentation methods are all based on CNN and ignore the advantages of traditional image processing technology. We combine the two and make them promote each other. The high frequency component of the image represents the edge part and the low frequency represents the body part. Based on this assumption, we use Fourier transform to obtain the high and low frequency component from images. Then, a multi-branch parallel network structure is designed, and the high and low frequency components are sent into two branches respectively to obtain the body and edge features. Finally, the two features are fused together through one deep feature fusion to obtain the final output. In this way, the edge information and body information are decoupled on original images and extracted separately, which not only ensures the consistency of the internal information within objects, but also strengthens the supervision of the edge part which is the most error prone area in semantic segmentation. The results on Cityscapes and KITTI fully demonstrate the effectiveness of our work. Lianlei Shan, Xiaobin Li 0006, Weiqiang Wang 0001 |
ICASSP | 1 |
| 2020 | Energy Minimum Regularization in Continual LearningabstractHow to give agents the ability of continuous learning like human and animals is still a challenge. In the regularized continual learning method OWM, the constraint of the model on the energy compression of the learned task is ignored, which results in the poor performance of the method on the dataset with a large number of learning tasks. In this paper, we propose an energy minimization regularization(EMR) method to constrain the energy of learned tasks, providing enough learning space for the following tasks that are not learned, and increasing the capacity of the model to the number of learning tasks. A large number of experiments show that our method can effectively increase the capacity of the model and reduce the sensitivity of the model to the number of tasks and the size of the network. Xiaobin Li 0006, Lianlei Shan, Minglong Li, Weiqiang Wang 0001 |
ICPR | 2 |
| 2020 | Global-Local Attention Network for Semantic Segmentation in Aerial ImagesabstractErrors in semantic segmentation could be classified into two types: the large area misclassification and inaccurate local boundaries. Previously attention-based methods typically capture rich global contextual information, which benefits the large area classification but cannot address the local errors of boundaries. In this paper, we propose a Global-Local Attention Network (GLANet) which can simultaneously consider the global context and local details. Specifically, our GLANet consists of two branches: (1) the global attention branch and (2) local attention branch. Furthermore, three different modules are embedded in GLANet for respectively modelling the semantic interdependencies in spatial, channel and boundary dimension. Lastly, we merge the outputs of different branches to enhance the feature representation further, resulting in more precise segmentation. Overall, the proposed method achieves the competitive segmentation accuracy on two public aerial image datasets, bringing significant improvements over the existing baselines. Minglong Li, Lianlei Shan, Xiaobin Li 0006, Dengji Zhou, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001, Sibao Chen 0001 |
ICPR | 2 |
| 2020 | UHRSNet: A Semantic Segmentation Network Specifically for Ultra-High-Resolution ImagesabstractSemantic segmentation is a basic task in computer vision, but only limited attention has been devoted to the ultra-high-resolution (UHR) image segmentation. Since UHR images occupy too much memory, they cannot be directly put into GPU for training. Previous methods are cropping images to small patches or downsampling the whole images. Cropping and downsampling cause the loss of contexts and details, which is essential for segmentation accuracy. To solve this problem, we improve and simplify the local and global feature fusion method in previous works. Local features are extracted from patches and global features are from downsampled images. Meanwhile, we propose one new fusion called local feature fusion for the first time, which can make patches get information from surrounding patches. We call the network with these two fusions ultra-high-resolution segmentation network (UHRSNet). These two fusions can effectively and efficiently solve the problem caused by cropping and downsampling. Experiments show a remarkable improvement on Deepglobe dataset [1]. Lianlei Shan, Minglong Li, Xiaobin Li 0006, Ke Lu 0002, Bin Luo 0001, Sibao Chen 0001, Weiqiang Wang 0001 |
ICPR | 1 |