Ziyi Chen 0001

dblp:37/1439-1 · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
10since 2021 · last 2026
0000-0001-5851-2779ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 SFE-CapsNet: Spatial Feature Enhanced Capsule Networks for Remote Sensing Object Detection
abstract
Remote sensing imagery often involves complex backgrounds and multi-scale targets, while variations such as rotation and scaling significantly degrade the performance of existing object detection algorithms and hinder effective modeling of spatial relationships between objects. To address these challenges, we propose SFE-CapsNet. First, we fuse scalar features extracted by convolutional neural networks (CNNs) with vector features from capsule networks through structural reorganization to form virtual capsules, approximating the dynamic routing process via a single fully connected layer. This design preserves object pose and texture information while streamlining information flow. Second, we introduce a capsule attention module that generates attention masks to dynamically enhance target-relevant features and suppress background noise, strengthening multi-level feature representations. Integrated with a feature pyramid network (FPN) architecture, our approach achieves precise detection of targets at varying scales. Experimental results demonstrate that multi-level feature fusion and the capsule attention mechanism significantly improve detection accuracy and robustness, achieving 77.65% mean Average Precision (mAP) on the DOTA dataset and 97.63% mAP on HRSC2016, highlighting its effectiveness and efficiency in complex scenes.
Ziyi Chen 0001, Wenhui Qiu, Huayou Wang, Dilong Li, Jin Gou, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.1
2026 Visible and Infrared Image Fusion Based on Adaptive Weighted Multimodal Features Extraction and Bidirectional Guidance Structure
abstract
Efficient fusion of infrared and visible images is of critical importance for real-time applications such as autonomous driving. While deep learning-based fusion methods have demonstrated significant improvements in fusion quality in recent years, current network architectures still exhibit unsatisfactory computational complexity and processing speed. To reduce computational complexity and improve fusion efficiency, mask-based methods or approaches driven by downstream tasks often prioritize key regions. However, such methods tend to overemphasize target objects, potentially overlooking contextually significant elements. To address this limitation and achieve more effective fusion, we propose BEFuse, a decoupled two-stage training strategy with an end-to-end inference framework. BEFuse extracts shallow features and gradient information from images in the first stage via a cross-modal image segmentation subnetwork. In the fusion stage, we use the Hadamard product to map features to an implicit quadratic feature space, combining feature similarity and gradient mask information, allowing automatic adjustment of loss weights and improving fusion accuracy. Experiments on four datasets (MSRS, TNO, RoadScene, and M3FD) show that BEFuse outperforms existing methods in both fusion quality and computational speed.
Ziyi Chen 0001, Gaosheng Cai, Dilong Li, Jing Wang 0049, Jin Gou, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 Leveraging Multi-View Images to Learn Domain-Invariant Discriminative Embeddings for Cross-View Geo-Localization
abstract
Cross-view geo-localization (CVGL) aims to match images of the same location captured from different viewpoints, such as those captured by Unmanned Aerial Vehicles (UAVs) and satellite platforms. The task is particularly challenging due to significant variations in scale, viewpoint, and illumination. Most existing methods employ symmetric sampling strategy to construct drone–satellite image pairs for deep metric learning, but neglect the potential of incorporating multi-view drone images to enhance the viewpoint robustness of features. To address this, we propose leveraging multi-view images to learn Domain-Invariant Discriminative Embeddings (DIDE) for CVGL. DIDE introduces an Inter-view Feature Aggregation Module (IFAM), which dynamically integrates multi-view drone information into robust embeddings. These are used in contrastive learning with satellite embeddings within batches to learn view-invariant discriminative features, while representation learning further improves scene discrimination across batches. To reduce the domain gap, DIDE constructs and aligns drone and satellite prototypes for effective cross-domain feature alignment. Furthermore, we adopt a parameter-efficient transfer learning strategy that leverages the capabilities of pre-trained foundation models while fine-tuning only dual adapters, significantly reducing the trainable parameters. DIDE achieves the state-of-the-art on University-1652 and University-160k, competitive results on SUES-200, and demonstrates strong cross-dataset transferability, with fewer training parameters and lower computational cost.
Ziyi Chen 0001, Dilong Li, Jin Gou, Cheng Wang 0003, Kyle Gao, Jonathan Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Serialization Based Point Cloud Oversegmentation
Chenghui Lu, Jianlong Kwan, Dilong Li, Ziyi Chen 0001, Haiyan Guan
ICCV4
2025 LGMamba: Large-Scale ALS Point Cloud Semantic Segmentation With Local and Global State-Space Model
abstract
The large scale and extensive coverage of point cloud data make large-scale airborne laser scanning (ALS) point cloud semantic segmentation a highly challenging task. Although transformers have shown impressive performance in large-scale point cloud semantic segmentation task, their quadratic complexity limits the processing capacity. To alleviate this issue, we propose Local and Global Mamba (LGMamba)—a novel state-space model (SSM)-based network for large-scale point cloud semantic segmentation. Specifically, we propose Local Mamba module to extract fine-grained local features by effectively capturing local dependencies. Then, we propose Global Mamba module to refine the learned local features by capturing the global long-distance dependencies of whole scenes. The validation of our method on the DALES datasets was conducted. Extensive experimental results demonstrate the effectiveness of LGMamba, with mean intersection over union (mIoU) of 82.3% and overall accuracy (OA) of 97.7% on DALES.
Dilong Li, Chongkei Chang, Ziyi Chen 0001, Jixiang Du
IEEE Geosci. Remote. Sens. Lett.4
2025 SECBNet: Semantic Segmentation-Enhanced Color Balance Network for Optical Satellite Images
abstract
Earth observation satellites can capture optical images under different temporal, climatic conditions, and platforms exhibit substantial differences in color and brightness, leading to poor visual experiences when synthesizing large-area optical satellite images. The related issue of color balancing has attracted considerable attention from researchers, yet challenges such as a lack of research data and sensitivity to model parameters persist. To address these problems, this article publishes a publicly open dataset and presents a semantic segmentation-enhanced color balance network (SECBNet). First, to mitigate the scarcity of research data, we develop a publicly available remote sensing image color balance dataset, Zhu Hai color balance image (ZHCBI), to support related research activities. Second, to improve semantic consistency between the color-balanced images and the target images, we design a dual-branch U-Net architecture guided by segmentation results and propose a novel segmentation feature loss function. Finally, to address issues of seams and unnatural transitions between blocks in segmented processing, we introduce a postprocessing module based on weighted averaging. We conducted comparative experiments and analyses with existing mainstream color balancing algorithms on the ZHCBI dataset. The results demonstrate that our proposed method achieves state-of-the-art color balancing quality, with significant improvement in visual effects and a higher peak signal-to-noise ratio (PSNR) (23.64 dB) compared with other mainstream methods.
Ziyi Chen 0001, Hanhuang Chen, Lujuan Gao, Dilong Li, Cheng Wang 0003, Linlin Xu, Somayeh Mollaee, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 DBARCT: Road Extraction Based on Double-Branch Architecture and Random Block Coding Transformer
abstract
Although transformer models are main network architectures for the delineation of roads from remote sensing imagery, they have critical limitations due to their regular patch mechanism and inefficiency in local information learning. To address these limitations for enhanced road extraction, this letter presents a novel double-branch architecture and random block coding transformer (DBARCT), with the following contributions. First, to improve local spatial details’ learning, we integrate transformer with convolutional neural network (CNN) into a novel dual-branch encoder-decoder architecture, such that the resulting model is efficient at learning both the local edge information and the global context information that are highly complementary for accurate road extraction. Second, to additionally augment the learning of global contextual information, we integrate the regular patching approach in traditional transformer models with a new irregular patching approach, such that it can better capture the global spatial information correlations that might be ignored by the regular patching approach. Third, an array of tests was carried out to meticulously scrutinize the efficacy of the fundamental elements of the suggested model. The empirical findings reveal that the intersection over union (IoU) metric attained by the proposed methodology on the LRSNY dataset stands at 88.53%, thereby corroborating the efficacy and preeminence of our approach in tasks related to road extraction.
Ziyi Chen 0001, Yucai Chen, Lujuan Gao, Dilong Li, Linlin Xu, Jonathan Li 0001, Cheng Wang 0003, Yewang Chen
IEEE Geosci. Remote. Sens. Lett.1
2024 Local Enhanced Transformer Networks for Land Cover Classification With Airborne Multispectral LiDAR Data
abstract
Transformer networks have demonstrated remarkable performance in point cloud processing tasks. However, balancing local feature aggregation with long-range dependency modeling remains a challenging issue. In this work we present a local enhanced Transformer network (LETNet) for land cover classification with multispectral LiDAR data. Specifically, we first rethink position encoding in 3D Transformers and design a novel feature encoding module that embeds comprehensive geometric and semantic information, serving a similar purpose. Then, the proposed local enhanced Transformer module is used to capture the accurate global attention weights and refine the features. Finally, to effectively extract and integrate global features across various scales, an attention-based pooling module is introduced. This module extracts global features from each encoder and decoder layer and constructs a feature pyramid to fuse these multi-scale global features. Both quantitative assessments and comparative analyses demonstrate the competitive capability and advanced performance of the LETNet in land cover classification task.
Dilong Li, Shenghong Zheng, Ziyi Chen 0001, Jonathan Li 0001, Jixiang Du
IEEE Geosci. Remote. Sens. Lett.3
2023 ASFS: A novel streaming feature selection for multi-label data based on neighborhood rough set
Yaojin Lin, Jixiang Du, Hongbo Zhang 0002, Ziyi Chen 0001, Jia Zhang 0019
Appl. Intell.5
2023 BrGAN: Blur Resist Generative Adversarial Network With Multiple Joint Dilated Residual Convolutions for Chlorophyll Color Image Restoration
abstract
This paper presents a Blur Resist Generative Adversarial Network (GAN) (BrGAN) with multiple joint dilated residual convolutions for chlorophyll image restoration of the Geostationary Ocean Color Imager (GOCI). First, a publicly available dataset was built to support this study. Second, a multiple attention perception mechanism and a multiple joint dilated residual convolution module was proposed to cope with the challenge of large missing areas in GOCI chlorophyll images. Third, a patch GAN based discrimination module was proposed to avoid the restored areas with generating mosaic and shadows. Our experimental results demonstrate that the BrGAN can reach 37.06 in the peak signal-to-noise ratio (PSNR) and 0.0485 in the Learned Perceptual Image Patch Similarity (LPIPS), respectively. The comparative study shows that the BrGAN achieves the highest effectiveness and advancement among other seven state-of-the-art methods.
Ziyi Chen 0001, Yuhua Luo, Yiping Chen 0002, Jing Wang 0049, Dilong Li, Kyle Gao, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2020 Corse-to-Fine Road Extraction Based on Local Dirichlet Mixture Models and Multiscale-High-Order Deep Learning
abstract
Road extraction from remote sensing images is an attractive but difficult task. Gray-value distribution and structure feature information are both crucial for road extraction task. However, existing methods mainly focus on structure feature information which contains morphological shape features and machine learning features, suffering from lots of false positives which are generated at positions having similar structure features but different gray-value distribution with roads. To effectively fuse the two complementary gray-value distribution and structure feature information, we propose a coarse-to-fine road extraction algorithm from remote sensing images. First, at the coarse level, we introduce a local Dirichlet mixture models (LDMM) which utilizing gray-value distribution information to pre-segment images into potential roads and backgrounds. Thus, most backgrounds having different gray-value distribution with roads can be removed firstly. Compared with original Dirichlet mixture models, the LDMM is much faster and more accurate. Next, at the fine level, we introduce a multiscale-high-order deep learning strategy based on ResNet model which can learn robust structure context features for final road extraction step. Based on the results of LDMM, the multiscale-high-order strategy can further remove false positives which have different structure features with roads. Compared with a single scanning size ResNet, our multiscale-high-order strategy can learn higher-order context information, leading to better performances. We test our algorithm on Shaoshan dataset. Experiments illustrate our better performance compared with other six state-of-the-art methods.
Ziyi Chen 0001, Wentao Fan 0001, Bineng Zhong 0001, Jonathan Li 0001, Jixiang Du, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.1
2018 Kernel correlation filters for visual tracking with adaptive fusion of heterogeneous cues
Bineng Zhong 0001, Gu Ouyang, Xin Liu 0011, Ziyi Chen 0001, Cheng Wang 0020
Neurocomputing6
2018 Semantic Labeling of Mobile LiDAR Point Clouds via Active Learning and Higher Order MRF
abstract
Using mobile Light Detection and Ranging point clouds to accomplish road scene labeling tasks shows promise for a variety of applications. Most existing methods for semantic labeling of point clouds require a huge number of fully supervised point cloud scenes, where each point needs to be manually annotated with a specific category. Manually annotating each point in point cloud scenes is labor intensive and hinders practical usage of those methods. To alleviate such a huge burden of manual annotation, in this paper, we introduce an active learning method that avoids annotating the whole point cloud scenes by iteratively annotating a small portion of unlabeled supervoxels and creating a minimal manually annotated training set. In order to avoid the biased sampling existing in traditional active learning methods, a neighbor-consistency prior is exploited to select the potentially misclassified samples into the training set to improve the accuracy of the statistical model. Furthermore, lots of methods only consider short-range contextual information to conduct semantic labeling tasks, but ignore the long-range contexts among local variables. In this paper, we use a higher order Markov random field model to take into account more contexts for refining the labeling results, despite of lacking fully supervised scenes. Evaluations on three data sets show that our proposed framework achieves a high accuracy in labeling point clouds although only a small portion of labels is provided. Moreover, comparative experiments demonstrate that our proposed framework is superior to traditional sampling methods and exhibits comparable performance to those fully supervised models.
Huan Luo 0001, Cheng Wang 0003, Chenglu Wen, Ziyi Chen 0001, Dawei Zai, Yongtao Yu, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2016 Exploiting location information to detect light pole in mobile LiDAR point clouds
abstract
With rapid development of light detection and ranging (LiDAR) technologies, three dimensional point clouds increasingly become a new approach to sense the world. In our previous work, light poles were detected from mobile LiDAR point clouds without using their locations. In this paper, we improve our previous work by considering location information between two neighboring light poles to reduce false alarm. In the proposed method, the potential light poles are first detected by the extended Hough Forest Framework. Then, a gaussian distribution is exploited to model the distance between two light poles by using locations of those detected light poles. Finally, inaccurately detected light poles are removed by considering the distance between two adjacent objects. We evaluate our proposed method on mobile LiDAR point clouds acquired by RIEGL VMX-450 system. On the basis of the experimental test instances, we demonstrate improved accuracy on light pole detection.
Huan Luo 0001, Cheng Wang 0003, Hanyun Wang, Ziyi Chen 0001, Dawei Zai, Shanxin Zhang, Jonathan Li 0001
IGARSS4
2016 Vehicle Detection in High-Resolution Aerial Images via Sparse Representation and Superpixels
abstract
This paper presents a study of vehicle detection from high-resolution aerial images. In this paper, a superpixel segmentation method designed for aerial images is proposed to control the segmentation with a low breakage rate. To make the training and detection more efficient, we extract meaningful patches based on the centers of the segmented superpixels. After the segmentation, through a training sample selection iteration strategy that is based on the sparse representation, we obtain a complete and small training subset from the original entire training set. With the selected training subset, we obtain a dictionary with high discrimination ability for vehicle detection. During training and detection, the grids of histogram of oriented gradient descriptor are used for feature extraction. To further improve the training and detection efficiency, a method is proposed for the defined main direction estimation of each patch. By rotating each patch to its main direction, we give the patches consistent directions. Comprehensive analyses and comparisons on two data sets illustrate the satisfactory performance of the proposed algorithm.
Ziyi Chen 0001, Cheng Wang 0003, Chenglu Wen, Xiuhua Teng, Yiping Chen 0002, Haiyan Guan, Huan Luo 0001, Liujuan Cao, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2016 Vehicle Detection in High-Resolution Aerial Images Based on Fast Sparse Representation Classification and Multiorder Feature
abstract
This paper presents an algorithm for vehicle detection in high-resolution aerial images through a fast sparse representation classification method and a multiorder feature descriptor that contains information of texture, color, and high-order context. To speed up computation of sparse representation, a set of small dictionaries, instead of a large dictionary containing all training items, is used for classification. To extract the context information of a patch, we proposed a high-order context information extraction method based on the proposed fast sparse representation classification method. To effectively extract the color information, the RGB color space is transformed into color name space. Then, the color name information is embedded into the grids of histogram of oriented gradient feature to represent the low-order feature of vehicles. By combining low- and high-order features together, a multiorder feature is used to describe vehicles. We also proposed a sample selection strategy based on our fast sparse representation classification method to construct a complete training subset. Finally, a set of dictionaries, which are trained by the multiorder features of the selected training subset, is used to detect vehicles based on superpixel segmentation results of aerial images. Experimental results illustrate the satisfactory performance of our algorithm.
Ziyi Chen 0001, Cheng Wang 0003, Huan Luo 0001, Hanyun Wang, Yiping Chen 0002, Chenglu Wen, Yongtao Yu, Liujuan Cao, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.1
2016 Patch-Based Semantic Labeling of Road Scene Using Colorized Mobile LiDAR Point Clouds
abstract
Semantic labeling of road scenes using colorized mobile LiDAR point clouds is of great significance in a variety of applications, particularly intelligent transportation systems. However, many challenges, such as incompleteness of objects caused by occlusion, overlapping between neighboring objects, interclass local similarities, and computational burden brought by a huge number of points, make it an ongoing open research area. In this paper, we propose a novel patch-based framework for labeling road scenes of colorized mobile LiDAR point clouds. In the proposed framework, first, three-dimensional (3-D) patches extracted from point clouds are used to construct a 3-D patch-based match graph structure (3D-PMG), which transfers category labels from labeled to unlabeled point cloud road scenes efficiently. Then, to rectify the transferring errors caused by local patch similarities in different categories, contextual information among 3-D patches is exploited by combining 3D-PMG with Markov random fields. In the experiments, the proposed framework is validated on colorized mobile LiDAR point clouds acquired by the RIEGL VMX-450 mobile LiDAR system. Comparative experiments show the superior performance of the proposed framework for accurate semantic labeling of road scenes.
Huan Luo 0001, Cheng Wang 0003, Chenglu Wen, Zhipeng Cai 0003, Ziyi Chen 0001, Hanyun Wang, Yongtao Yu, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.5
2014 Oil spill detection based on a superpixel segmentation method for SAR image
abstract
In this paper, a rapid oil spill detection approach which still maintains high detection accuracy is presented. The major contribution of the approach is using a superpixel segmentation method to subdivide the target SAR image into many approximate uniform scale pieces and preserves the boundaries well. Furthermore, a novel approach combine space distance, intensity deviation and size information together (SIS) is presented to eliminate the potential false positive, which is convenient and effective meanwhile. The proposed approach performs well and fast in both the synthetic data and RAD ARS AT-1 ScanSAR data which contain verified oil spills. The processing time is about 6s for a 512×512 image.
Ziyi Chen 0001, Cheng Wang 0003, Xiuhua Teng, Liujuan Cao, Jonathan Li 0001
IGARSS1