Zhongyuan Lu

dblp:224/4208 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0007-1695-6305ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Image recognition and object detection · 100%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
contour detection
0.912025
Enhancing Object Detection With Fourier Series · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › Image recognition and object detection
object detection
0.912025
Enhancing Object Detection With Fourier Series · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Geometric modeling and processing
shape representation
0.312025
Enhancing Object Detection With Fourier Series · IEEE Trans. Pattern Anal. Mach. Intell. 2025

Methods — techniques the papers use, named apart from their topics

rolling optimization matching · 1.7fourier series regression · 1.7
YearPublicationVenuePosition
2025 Identity Model Transformation for boosting performance and efficiency in object detection network
Zhongyuan Lu, Jin Liu 0029, Miaozhong Xu
Neural Networks1
2025 Enhancing Object Detection With Fourier Series
abstract
Traditional object detection models often lose the detailed outline information of the object. To address this problem, we propose the Fourier Series Object Detection (FSD). It encodes the object's outline closed curve into two one-dimensional periodic Fourier series. The Fourier Series Model (FSM) is constructed to regress the Fourier series for each object in the image. Thus, during inference, the detailed outline information of each object can be retrieved. We introduce Rolling Optimization Matching for Fourier loss to ensure that the model's learning process is not affected by the sequence of the starting points of the labeled contour points, speeding up the training process. The FSM demonstrates improved feature extraction and descriptive capabilities for non-rectangular or elongated object regions. The model achieves AP50 = 73.3% on the DOTA 1.5 dataset, which surpasses the state-of-the-art (SOTA) method by 6.44% at 66.86%. On the UCAS dataset, the model achieves AP50 = 97.25%, also surpassing the performance indicators of the SOTA methods. Furthermore, we introduce the object's Fourier power spectrum to describe outline features and the Fourier vector to indicate its direction. This enhances the scene semantic representation of the object detection model and paves a new pathway for the evolution of object detection methodologies.
Jin Liu 0029, Zhongyuan Lu, Yaorong Cen, Yong Hong, Miaozhong Xu
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Crisscross-Global Vision Transformers Model for Very High Resolution Aerial Image Semantic Segmentation
abstract
Semantic segmentation is a key means for understanding very-high resolution (VHR) aerial imagery. With the explosive development of deep learning, deep learning methods are being applied to the segmentation of VHR images, with convolutional neural networks (CNNs) as the basic framework. However, owing to the highly complex details present in VHR images and the high spatial dependence of geographical objects, CNN-based methods are inadequate. This is because the inherent locality of CNNs limits the size of the receptive field, thus limiting the ability to obtain long-range context information. To solve this problem, in this paper, we propose a transformer-based novel deep learning model called crisscross-global vision transformers (CGVT). CGVT exploits the transformer’s inherent ability to obtain long-range context information to solve the restricted receptive field problem. Specifically, we redesign the self-attention mechanism in the transformer and call it crisscross-global attention. It consists of two parts: crisscross transformer encoder block (CC-TEB) and global squeeze transformer encoder block (GS-TEB). CC-TEB overcomes the limitation of the traditional self-attention design (specifically, difficulty applying it to VHR aerial image segmentation) and further increases the local feature representation ability of the model. GS-TEB increases the global feature representation ability of the model. The results of experiments conducted on the popular ISPRS Vaihingen, IEEE GRSS Data Fusion Contest Zeebrugge, and LoveDA Semantic Segmentation Challenge datasets verify the effectiveness and superiority of our proposed method. Specifically, it achieved state-of-the-art performance on both Zeebrugge and LoveDA datasets, and is currently ranked second in Vaihingen dataset.
Guohui Deng, Zhaocong Wu, Miaozhong Xu, Zhiye Wang, Zhongyuan Lu
IEEE Trans. Geosci. Remote. Sens.6
2022 Hyperspectral Image Stripe Removal Network With Cross-Frequency Feature Interaction
abstract
Remote sensing images, especially hyperspectral images (HSIs), are extremely vulnerable to random noise and stripe noise. As a key aspect of HSI data quality improvement, stripe noise removal has always been a pervasive issue in remote sensing image processing. Convolutional neural networks have been applied for HSI data destriping. However, the existing methods lose the stripe-free component of the original image to a certain extent. These models also ignore the global spatial context of images and the correlation between spatial information and spectral information. Therefore, we propose a novel destriping convolutional network to overcome the problems with the existing methods. Octave convolution is used to extract cross-frequency features, and separate and compress the low-frequency information of the images, while dilation convolution (Dila-Conv) is used to reduce the amount of required calculation and also preserve the key image information. In addition, Dila-Conv can expand the receptive field to obtain multiscale features. Finally, a cross-channel enhanced spatial–spectral feature fusion module is used to acquire and integrate spatial context information and interchannel dependencies on a global scale as auxiliary information so that the network model can learn and pay attention to key feature information, specifically, “what to look for” and “where to look at,” which can facilitate the distinction between stripe and stripe-free components. Experimental results obtained using multiple datasets demonstrated that the proposed method can outperform the existing comparable methods and can produce satisfactory results in terms of visual effects and quantitative evaluation.
Miaozhong Xu, Yonghua Jiang 0001, Guohui Deng, Zhongyuan Lu, Guo Zhang 0001, Hao Cui 0002
IEEE Trans. Geosci. Remote. Sens.5