EDBT 2026 Demo / reviewers in the wild / expert
Junli Yang
dblp:36/7185
· DBLP profile ↗
21ranked-venue papers
0as first author
13since 2021 · last 2024
0000-0001-8370-7105ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Edge-Guided Enhancement Network for Building Change Detection of Remote Sensing Images with a Hybrid CNN-Transformer ArchitectureabstractThe utilization of remote sensing images for building change detection has become a focal point of concern. Many contemporary change detection methodologies primarily focus on extracting more discriminative features while neglecting the use of prior edge information, leading to inaccurate detection results, especially in areas like building boundaries. To address this issue, we propose a method called Edge-Guided Enhancement Network(EGENet) that combines discriminative information and edge features within a unified framework. In particular, as edge represents part of the image that changes drastically and indicates high-frequency information in image, we construct an Edge Enhancement Module (EEM) to enhance high-frequency information and suppress noise. Furthermore, We design the Edge Detection Branch (EDB) to better enhance edge information and employ the detected edge information for subsequent cross attention. Our experiments on two building change detection datasets demonstrate that our approach is proper and achieves advanced performance compared to other methods. Kaiwen Jing, Chenhe Wang, Yanhan Wang, Jiarui Ban, Junli Yang |
IGARSS | 6 |
| 2024 | PCNEXT: Convolution is All You Need for Semantic Segmentation for Remote Sensing ImagesabstractSemantic segmentation of remote sensing images is crucial for various applications, including land use mapping and environmental monitoring. However, most CNNs lack the ability to capture long-range context due to their limited receptive fields. While Transformers adopt multi-head self-attention mechanism to capture long-range context for better accuracy, it often leads to high parameter volume and computational complexity. In this paper, we propose PCNeXt, a lightweight pure-convolutional neural network for semantic segmentation of remote sensing images. For the encoder part, we design a pure-convolutional lightweight module named MSWCA based on MSCA, which is the key component of the encoder of SegNeXt. For the decoder part, we design a purely Convolutional Global-Local Block (CGLB) to replace the GLTB module of UNetFormer, which utilizes the MSWCA module instead of self-attention to capture global context and a simple convolutional operation to capture local context. The experiments prove that instead of using self-attention in transformer, the proposed PCNeXt is capable of achieving competitive accuracy by using only convolutional structure. Notably, our PCNeXt achieves 91.53% aAcc and 83.9 % mIoU on the Potsdam dataset, with only 4.42MB parameters and 6.87G FLOPs calculations. Yueyao Su, Hanlu Zhen, Junli Yang, Haopeng Zhang 0001 |
IGARSS | 4 |
| 2023 | LRCDNeT: A Lightweight and Real-Time Cloud Detection Network for Remote Sensing ImagesabstractMost CNN-based cloud detection methods have high computational complexity, large parameter size, and slow inference speed, which limit their practical applications. Furthermore, cloud detection is a challenging task due to the irregular shapes and random sizes of clouds, often leading to inaccurate detection. To overcome these challenges, we propose a lightweight and real-time network (LRCDNet) tailored for cloud detection. By incorporating Short-Term Dense Concatenate (STDC) module, Multi-group Deformable Convolution (DCNv3) and multiple linear self-attention mechanisms, LRCDNet can effectively extract detail information and adaptively build long-term dependencies. In comparison to the state-of-the-art cloud detection and typical real-time semantic segmentation methods, our proposed LRCDNet strikes a better balance between accuracy and computational costs. Specifically, when tested on the GF-1 WHU dataset, LRCDNet achieves an overall accuracy (OA) of 97.37%, a F1-score of 92.42%, and a remarkable inference speed of 122.83 Frames Per Second (FPS) on a RTX A4000 GPU, with only 19.72MB parameters and 10.69G FLOPs calculations. Ruofu Li, Junli Yang, Ben Deng |
IGARSS | 2 |
| 2023 | Dlafnet: A Direct Fusion Method of 2D Aerial Image and 3D Lidar Point Cloud for Semantic SegmentationabstractSemantic segmentation of high-resolution remote sensing images (RSIs) is developing rapidly. Multispectral images can provide rich spectral information for semantic segmentation, while 3D LiDAR point cloud data can provide depth information. Thus, semantic segmentation accuracy could be improved by fusing multispectral images and 3D LiDAR point cloud. In this paper, we propose a method titled Direct LiDAR-Aerial Fusion Network (DLAFNet) which directly uses RSIs and LiDAR point cloud for semantic segmentation tasks. In particular, owing to the fact that sparse features extracted from the KPConv branch are not as essential as features from RSIs, we design LiDAR Assisted Attention Module (L-AAM). Our experiments on the modified GRSS18 dataset prove that our method is proper and can obtain the best results by comparing with its components and other methods. Wei Liu 0158, He Wang 0024, Yicheng Qiao, Junli Yang, Haopeng Zhang 0001 |
IGARSS | 5 |
| 2022 | PaRK-Detect: Towards Efficient Multi-Task Satellite Imagery Road Extraction via Patch-Wise Keypoints Detection
Shenwei Xie, Wanfeng Zheng, Zhenglin Xian, Junli Yang, Ming Wu 0001 |
BMVC | 4 |
| 2022 | Le-BEiT: A Local-Enhanced Self-Supervised Transformer for Semantic Segmentation of High Resolution Remote Sensing ImagesabstractSemantic segmentation for remote sensing images (RSI) has been a thriving research topic for a long time. Existing supervised learning methods usually require a huge amount of labeled data. Meanwhile, large size, variation in object scales, and intricate details in RSI make it essential to capture both long-range context and local information. To address these problems, we propose Le-BEiT, a self-supervised Transformer with an improved positional encoding Local-Enhanced Positional Encoding (LePE). Self-supervised learning relieves the demanding requirement of a large amount of labeled data. The self-attention mechanism in Transformer has remarkable capability in capturing long-range context. Meanwhile, we use LePE as a substitution for Relative Positional Encoding (RPE) to represent local information more effectively. Moreover, considering the domain difference between natural images and RSI, instead of ImageNet-22K, we pre-train Le-BEiT on a very small high-resolution RSI dataset—GID. To investigate the influence of pre-training dataset size on segmentation accuracy, we furtherly conduct experiments on a larger pre-training dataset called GID-DOTA, which is 1/100 of ImageNet-22K, and have observed considerable accuracy improvements. The result of our method, which relies on a much smaller pretrained dataset, achieves competitive accuracy compared to the counterpart on ImageNet-22K. Zideng Feng, Junli Yang, Zhenglin Xian |
ICIP | 3 |
| 2022 | Maskformer with Improved Encoder-Decoder Module for Semantic Segmentation of Fine-Resolution Remote Sensing ImagesabstractIn 2021, the Transformer based models have demonstrated extraordinary achievement in the field of computer vision. Among which, Maskformer, a Transformer based model adopting the mask classification method, is an outstanding model in both semantic segmentation and instance segmentation. Considering the specific characteristics of semantic segmentation of remote sensing images (RSIs), we design CADA-MaskFormer(a Mask classification-based model with Cross-shaped window self-Attention and Densely connected feature Aggregation) based on Maskformer by improving its encoder and pixel decoder. Concretely, the mask classification that generates one or even more masks for specific category to perform the elaborate segmentation is especially suitable for handling the characteristic of large within-class and small between-class variance of RSIs. Furthermore, we apply the Cross-Shaped Window self-attention mechanism to model the long-range context information contained in RSIs at maximum extent without the increasing of computational complexity. In addition, the Densely Connected Feature Aggregation Module (DCFAM) is used as the pixel decoder to incorporate multi-level feature maps from the encoder to get a finer semantic segmentation map. Extensive experiments conducted on two remotely sensed semantic segmentation datasets Potsdam and Vaihingen achieves 91.88% and 91.01% in OA index respectively, outperforming most of competitive models designed for RSIs. The code is available from https://github.com/lqwrl542293/JL-Yang_CV/tree/master/CADA_Maskformer Junli Yang |
ICIP | 2 |
| 2022 | Dual-Path Geometry-Aware Network for Semantic Segmentation of High-Resolution Aerial ImagesabstractSemantic segmentation of high-resolution aerial images is a fundamental research topic for its extensive applications. Different from natural scene datasets, the high-resolution aerial datasets provide additional elevation data such as Digital Surface Model (DSM). However, the current semantic segmentation methods of high-resolution aerial images focus on improving the feature extraction of the spectral images but fail to make full use of DSM images. Besides, the feature fusion of these two disparate data is a challenging problem. Moreover, the tremendous details and the considerable variations in scale of objects limit the representation capacity of existing segmentation networks. To address the above problems, we propose a new dual-path geometry-aware end-to-end DPGANet which consists of Multi-scale Digital Surface Model Awareness(MDSMA) path and Swin Transformer path. The MDSMA path is designed to extract multi-stage 3D geometry features from DSM images. The Res2Net modules in the MDSMA path can enhance the multi-scale representation capability of our network. The Swin Transformer path is designed to extract the multi-stage long-range dependencies from spectral images. Furthermore, for full usage of feature maps produced by corresponding stages of these two paths, we design an Attention Fusion Module(AFM) for memory-saving and computation-effective feature fusion from both spatial and channel dimensions. The segmentation results on the ISPRS Potsdam dataset achieve a competitive performance compared to other state-of-the-art methods. Zhenglin Xian, Junli Yang, Zideng Feng |
ICPR | 3 |
| 2021 | Spd-Linknet: Upgraded D-Linknet with Strip Pooling for Road ExtractionabstractIn the field of road extraction, an dominant network is D-LinkNet which won the first place in DeepGlobe 2018 challenge. Although D-LinkNet creatively proposed D-block with progressively enlarged dilated convolution and proved its efficiency, the$\mathrm{N}\times \mathrm{N}$square kernel it used still has limitations for road extraction. Road in aerial imagery usually has narrow-and-long shape, and the direction is randomly distributed. Therefore, not only large receptive field but also anisotropic long-range contextual information should be considered. Based on this intuition, we integrate a new pooling strategy named strip pooling which uses a long but narrow kernel i.e.$1\times \mathrm{N}$or$\mathrm{N}\times 1$into D-LinkNet. With strip pooling module(SPM) and mixed pooling module(MPM) designed based on strip pooling, two modifications are made to D-LinkNet: 1) We insert SPM into Res-block of the original encoder ResNet34 and name it Res-SPM-block. 2) Inspired by MPM, we connect strip pooling in parallel with D-block and name it SPD-block. We name the upgraded D-LinkNet as SPD-LinkNet. Experimental results on DeepGlobe 2018 dataset prove that SPD-LinkNet outperforms original D-LinkNet in accuracy while maintaining nearly the same inference speed. Yutao Deng, Junli Yang, Chenyi Liang, Yinuo Jing |
IGARSS | 2 |
| 2021 | Damaged Road Extraction Based on Simulated Post-Disaster Remote Sensing ImagesabstractDamaged road extraction is a challenging task in the field of remote sensing. Some existing methods include the step to extract road from pre- and post-disaster remote sensing images of the same area. In practice, it often occurs that one of these two images is missing. To solve this problem, we use CoCosNet, the model for exemplar-based image translation, to translate pre-disaster images to simulated post-disaster ones. Then we use D-LinkNet, the state-of-the-art method in road extraction, to extract road from the pre- and post-disaster images of the same area. We extract damaged road area by comparing pre-disaster road masks with post-disaster ones and output the damage level by calculating the proportion of the damaged road area. Finally, we evaluate the damaged road extraction accuracy. Experimental results on simulated post-disaster images prove the effectiveness of the simulation method and the framework for damaged road extraction and damage level evaluation. Yansong Huang, Haocai Wei, Junli Yang, Ming Wu 0001 |
IGARSS | 3 |
| 2021 | Efficient Semantic Segmentation Method with Strip Pooling for VHR Remote Sensing ImagesabstractIn this paper, we address the problem of multi-class semantic segmentation of high resolution remote sensing images with a deep convolutional neutral network based model named Strip Pooling Network(SPNet). The objects in high resolution remote sensing images vary greatly in size. Moreover, these objects have different extensibility and directions, such as long narrow roads and wide grasslands. These bring big challenge for traditional square pooling kernels. If we want to capture the long-rang dependencies and enough contextual information, we need to use large pooling kernel which definitely increases computation greatly and incorporates interference information from irrelevant regions. SPNet introduces a new pooling strategy, called strip pooling which uses a long but narrow kernel, i.e.,$1 \times \mathrm{N}$or$\mathrm{N}\times 1$. It can solve the above problem while preventing information from irrelevant regions. Meanwhile, two paths in the strip pooling module focus on the horizontal and vertical spatial dimensions respectively, which compensate for the lack of one path captured in the narrow dimension of each other. Experimental results on public available Potsdam dataset demonstrate that SPNet obtains an overall accuracy of 89.1%, which outperforms other state-of-the-art methods. Yifan Sheng, Junli Yang, Youguang Lin |
IGARSS | 2 |
| 2021 | Did-Linknet: Polishing D-Block with Dense Connection and Iterative Fusion for Road ExtractionabstractSince the D-LinkNet won the first place in CVPR2018 Deep-Globe Challenge, the deep neural network built on encoder-decoder structure has become dominant in pixel-wise road extraction [1], [2]. The primary reason for D-LinkNet's success is the creative design of D-Block, which extracts features by dilated convolutions with receptive fields of progressively growing sizes and fuses the features with element-wise addition. Although D-Block promotes road connectivity significantly’ D-LinkNet still struggles in detecting road in regions with compact road topologies and indistinguishable ground objects from the real road. Therefore, in this paper, we obligate in augmenting the extraction and fusion of feature-maps in D-Block through two architectural amendments and upgrade D-Block to DID-Block. The first amendment is introduced to maximize the information flow between any layers in D-Block, which we name dense connection. The second one, iterative fusion, is proposed to aggregate representations learned by each layer with combining the layers' outputs iteratively. To construct the DID-LinkNet, we replace the D-Block in D-LinkNet with DID-Block. We demonstrate the profit of these two structural modifications separately on DeepGlobe2018 dataset, and the experimental results show that the proposed DID-LinkNet enjoys a further road connectivity gains. Junli Yang, Ming Wu 0001 |
IGARSS | 3 |
| 2021 | Real-Time Semantic Segmentation of Aerial Videos Based on Bilateral Segmentation NetworkabstractIn recent years, deep learning algorithms have been widely used in semantic segmentation of aerial images. However, most of the current research in this field focus on images but not videos. In this paper, we address the problem of real-time aerial video semantic segmentation with BiSeNet[1]. Since BiSeNet is originally proposed for semantic segmentation of natural city scene images, we need a corresponding dataset to ensure the effect of transfer learning when applying it to aerial video segmentation. Therefore, we build a UAV streetscape sequence dataset (USSD) to fill the vacancy of dataset in this field and facilitate our research. Evaluation on USSD shows that BiSeNet outperforms other state-of-the-art methods. It achieves 79.26% mIoU and 93.37% OA with speed of 148.7 FPS on NVIDIA Tesla V100 for a 1920x1080 frame size input aerial video, which satisfies the demand of aerial video semantic segmentation with a competitive balance of accuracy and speed. The aerial video semantic segmentation results are provided at Our Repository. Yihao Zuo, Junli Yang, Yutong Zheng |
IGARSS | 2 |
| 2020 | Simple, Fast, Accurate Object Detection based on Anchor-Free Method for High Resolution Remote Sensing ImagesabstractOn object detection of remote sensing images, speed and accuracy are considered equally significant for time-sensitive tasks. Traditional object detection methods of remote sensing images use anchors to regress the bounding boxes and the position of objects which quite limit the speed of detection. In this paper, we make an attempt at employing an anchor-free method named CenterNet to address the problem of object detection for remote sensing images. This model computes the key point to regress the location, local offset and size without enumerating a nearly exhaustive list of potential object locations and size. Thus CenterNet is simpler, faster, more accurate and end-to-end differentiable than other anchor-based methods. We evaluate the performance of Cen-terNet with various backbones on remote sensing images dataset NWPU VHR-10. Among them, CenterNet based on DLA-34 backbone achieves the highest accuracy of 95.7% on the test set, which significantly outperforms most state-of-the-art methods while maintaining a real-time prediction speed of 31 FPS on NVIDIA GTX 1080. Analysis on experimental results demonstrates that this anchor-free network achieves a better balance between accuracy and speed for this task compared to other methods. Junli Yang, Wenqian Cui |
IGARSS | 2 |
| 2019 | Large Kernel Spatial Pyramid Pooling for Semantic Segmentation
Tianshi Hu, Junli Yang, Zhaoxing Zhang |
ICIG (1) | 3 |
| 2019 | Robust Real-Time Object Detection Based on Deep Learning for Very High Resolution Remote Sensing ImagesabstractRecently, the development of deep learning boosts the object detection for remote sensing images. The existing deep learning methods can be divided into two types. The region-based methods represented by Faster R-CNN have progressive performance in accuracy. However, their computational cost is massive due to the deep Convolutional Neural Network (CNN) backbones, which limits the efficiency. The regression-based methods such as YOLO and Single Shot MultiBox Detector (SSD) are advantageous in speed while the accuracy is not satisfactory. To meet the increasing demand in both speed and accuracy for object detection of remote sensing images, we employ the Reception Field Block Net (RFBNet) detector. It embeds the Receptive Field Block (RFB) module into SSD to obtain better feature representation. The experimental results on NWPU VHR-10 dataset demonstrate that the mAP of RFBNet-512 reaches 91.56%, which outperforms other state-of-the-art networks. Meanwhile, the speed is also competitive. Jinzheng Zhao, Weiyu Xiong, Qingli Li, Junli Yang |
IGARSS | 6 |
| 2019 | Efficient Multi-Class Semantic Segmentation of High Resolution Aerial Imagery with Dilated LinkNetabstractIn this paper, we address the problem of multi-class semantic segmentation of high resolution aerial imagery with a deep-convolutional-neural-network-based model named D-LinkNet, which was initially designed for the task of road extraction. D-LinkNet extends LinkNet, which is considered as an efficient method for semantic segmentation and adopts encoder-decoder architecture, by inserting dilated convolution layers between LinkNets encoder and decoder to enlarge the receptive field of kernel without reducing the resolution of the feature maps. Besides, we use the multi-class form of the original loss function combined by Dice loss and Binary Cross Entropy loss, which has proven to be effective in classification tasks with imbalanced class distribution that is prevalent for aerial imagery. Experimental results on public available Potsdam dataset demonstrate that D-LinkNet obtains an overall accuracy of 86.1%, which outperforms state-of-the-art methods. Moreover, D-LinkNet achieves a competitive speed in the experiment so it can produce relatively accurate predictions more efficiently. Qingtian Zhu, Yumin Zheng, Yulai Jiang, Junli Yang |
IGARSS | 4 |
| 2013 | Shadow Boundaries Identification in Single Natural Images via Multiple Kernels LearningabstractThe identification of shadow and shading boundaries is a key step towards reducing the imaging effects that are caused by direct illumination of the light source in the scene. Discriminating shadow boundaries from images of natural scenes has been widely applied in the field of computer vision such as object recognition, intelligent monitoring and image understanding. In this paper, we propose a method to identify shadow boundaries based on multiple kernel learning. We first extract all possible candidate boundaries and then analyze their properties. Unlike the previous proposed methods which simply combine features as a vector, we choose the optimal kernel function for every feature and learn the correct weights of different features from training database. At last, we link shadow boundaries fragments together to get longer and complete shadow boundaries. The experiment results show that the method we propose works well in shadow boundaries identification. Junli Yang, Jianwei Luo |
ICIG | 3 |
| 2011 | A Hierarchical Connection Graph Algorithm for Gable-Roof Detection in Aerial ImageabstractIn this letter, we present a hierarchical connection graph (HCG) algorithm based on a self-avoiding polygon (SAP) model for detecting and extracting gable roofs from aerial imagery. The SAP model is a deformable shape model that is capable of representing gable roofs of various shapes and appearances. The model is composed of a sequence of roof-corner templates that are connected into a SAP, which serves as a flexible shape prior. An energy function that combines features from three channels (corner, boundary, and interior area) is defined over the sequence to quantify the variability in appearances of gable roofs. To infer the most probable state of the corner sequence for an input image, we use an efficient algorithm-called HCG algorithm. The algorithm converts the solution space of a SAP model into a directed graph (which we call “HCG”) and searches for the best path using dynamic programming (DP). It is efficient for two reasons: 1) By constructing an HCG, the algorithm can quickly prune out a large amount of invalid solutions using only geometric constraints, which are inexpensive to compute, and 2) by employing DP, the algorithm decomposes the searching problem into smaller overlapping subproblems and reuses energy scores, which are expensive to compute. Experimental results on a set of challenging gable roofs show that our algorithm has good performance and is computationally effective. Qiongchen Wang, Zhiguo Jiang 0001, Junli Yang, Danpei Zhao, Zhenwei Shi 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2009 | Gable Roof Description by Self-Avoiding Polygon
Qiongchen Wang, Zhiguo Jiang 0001, Junli Yang, Danpei Zhao, Zhenwei Shi 0001 |
ACCV (3) | 3 |
| 2009 | Optimizing search engine revenue in sponsored searchabstractDisplaying sponsored ads alongside the search results is a key monetization strategy for search engine companies. Since users are more likely to click ads that are relevant to their query, it is crucial for search engine to deliver the right ads for the query and the order in which they are displayed. There are several works investigating on how to learn a ranking function to maximize the number of ad clicks. In this paper, we address a new revenue optimization Yunzhang Zhu, Gang Wang 0010, Junli Yang, Dakan Wang, Jun Yan 0001, Jian Hu 0001, Zheng Chen 0001 |
SIGIR | 3 |