Zhengbo Luo

dblp:289/0352 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0003-1662-6166ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 RFMiD: Retinal Image Analysis for multi-Disease Detection challenge
Samiksha Pachade, Prasanna Porwal, Manesh Kokare, Girish Deshmukh, Vivek Sahasrabuddhe, Zhengbo Luo, Zitang Sun, Li Qihan, Edward Ho, Asaanth Sivajohan, Saerom Youn, Kevin Lane, Jin Chun, Yunchao Gu, Sixu Lu, Young-tack Oh, Hyunjin Park, Chia-Yen Lee, Hung Yeh, Kai-Wen Cheng, Haoyu Wang 0010, Jin Ye 0002, Junjun He, Lixu Gu, Dominik Müller, Iñaki Soto Rey, Frank Kramer 0001, Hidehisa Arai, Yuma Ochi, Takami Okada, Luca Giancardo, Gwenolé Quellec, Fabrice Mériaudeau
Medical Image Anal.6
2023 Decoupled spatiotemporal adaptive fusion network for self-supervised motion estimation
Zitang Sun, Zhengbo Luo, Shin'ya Nishida
Neurocomputing2
2022 Deep Residual Networks with Common Linear Multi-Step and Advanced Numerical Schemes
abstract
Deep Neural Networks (DNNs) are widely used state-of-the-art approaches for the massive audio signal, language and vision tasks. However, we still lack sufficient understanding of the theoretical issues inside. Based on the equivalence between deep residual networks (ResNets) and the Euler forward scheme, we present several ways to construct DNNs (ResNet as the example) in advanced numerical and linear multi-step schemes. Furthermore, we show that ResNets with various advanced schemes have better accuracy, convergence, robustness and an explainable rank against the property of these schemes. Finally, previous works are summarized and discussed; sufficient experiments are provided. Improvements in several aspects are noticeable and theoretical.
Zhengbo Luo, Weilian Zhou, Xuehui Hu
ICIP1
2022 Rethinking Unified Spectral-Spatial-Based Hyperspectral Image Classification Under 3D Configuration of Vision Transformer
abstract
Vision Transformer (ViT) has been introduced into the computer vision (CV) field with its self-attention mechanism to capture global dependency. However, simply deploying ViT on a hyperspectral image (HSI) classification task can not get satisfying results because ViT is a spatial-only self-attention model, but rich spectral information exists in HSI. Moreover, most HSI classifiers integrate spectral and spatial features in a cascaded flowchart, ignoring the internal correlation between spectral and spatial information. Furthermore, existing positional embedding (PE) methods can not fulfil the 3D configuration of ViT. Therefore, this paper proposes a unified spectral-spatial-based 3D ViT with cooperative 3D coordinate positional embedding. In the meanwhile, a novel local-global feature fusion strategy is proposed. The model does not contain convolution or recurrent units and can achieve more competitive classification performance than other state-of-the-art (SOTA) methods. Furthermore, compared with existing ViT-based HSI classifiers, our concept can get better results.
Weilian Zhou, Zhengbo Luo, Xi Xue
ICIP3
2022 Hierarchical Unified Spectral-Spatial Aggregated Transformer for Hyperspectral Image Classification
abstract
Vision Transformer (ViT) has recently been introduced into the computer vision (CV) field with its self-attention mechanism and gotten remarkable performance. However, simply applying ViT for hyperspectral image (HSI) classification is not applicable due to 1) ViT is a spatial-only self-attention model, but rich spectral information exists in HSI; 2) ViT needs sufficient training samples, but HSI suffers from limited samples; 3) ViT does not well learn local features; 4) multi-scale features for ViT are not considered. Furthermore, the methods which combine convolutional neural network (CNN) and ViT generally suffer from a large computational burden. Hence, this paper tends to design a suitable pure ViT based model for HSI classification as the following points: 1) spectral-only vision transformer with all tokens’ aggregation; 2) spatial-only local-global transformer; 3) cross-scale local-global feature fusion, and 4) a cooperative loss function to unify the spectral and spatial features. As a result, the proposed idea achieves competitive classification performance on three public datasets than other state-of-the-art methods.
Weilian Zhou, Zhengbo Luo
ICPR3
2022 Constructing infinite deep neural networks with flexible expressiveness while training
Zhengbo Luo, Zitang Sun, Weilian Zhou, Zizhang Wu
Neurocomputing1
2022 Multiscanning Strategy-Based Recurrent Neural Network for Hyperspectral Image Classification
abstract
Most methods based on the convolutional neural network show satisfying performance for hyperspectral image (HSI) classification. However, the spatial dependence among different pixels is not well learned by CNNs. A recurrent neural network (RNN) can effectively establish the dependence of nonadjacent pixels and ensure that each feature activation in its output is an activation at the specific location concerning the whole image, in contrast to the usual local context window in the CNNs. However, recent limited conversion schemes in RNN-based methods for HSI classification cannot fully capture the complete spatial dependence of an HSI patch. In this study, a novel multiscanning strategy with RNN is proposed to feature the sequential character of the HSI pixel and fully consider the spatial dependence in the HSI patch. By investigating different scanning forms, eight scanning orders are considered spatially, which flattens one local HSI patch into eight neighboring continuous pixel sequences. Moreover, considering that eight scanning orders complement one local patch with correlative dependence, the concatenated features from all scanning orders are fed into the RNN again for complementarity. As a result, the network can achieve competitive classification performance on three publicly accessible datasets using fewer parameters than other state-of-the-art methods.
Weilian Zhou, Zhengbo Luo, Haipeng Wang 0002
IEEE Trans. Geosci. Remote. Sens.3
2021 Deep Neural Networks with Flexible Complexity While Training Based on Neural Ordinary Differential Equations
abstract
Most structures of deep neural networks (DNN) are with a fixed complexity of both computational cost (parameters and FLOPs) and the expressiveness. In this work, we experimentally investigate the effectiveness of using neural ordinary differential equations (NODEs) as a component to provide further depth to relatively shallower networks rather than stacked layers (depth) which achieved improvement with fewer parameters. Moreover, we construct deep neural networks with flexible complexity based on NODEs which enables the system to adjust its complexity while training. The proposed method achieved more parameter-efficient performance than stacking standard DNNs, and it alleviates the defect of the heavy cost required by NODEs.
Zhengbo Luo, Zitang Sun, Weilian Zhou
ICASSP1
2021 HFGCNET: High-Frequency Graph Reasoning for Finer Semantic Image Segmentation
abstract
Semantic segmentation is a fundamental task in computer vision and image processing. Although existing methods based on the fully convolutional network (FCN) have greatly improved the accuracy, it still does not show satisfactory results on tiny objects and boundary regions. One of the problems is that the current FCN-based methods ignore details such as the image’s contours and edges because of over downsampling operations in the CNN encoder backbone. In signal processing, excessive down-sampling will incur spectrum aliasing, thus losing high-frequency details. This work presents a high-frequency graph convolution operation to solve the above problems. Traditional image processing generally uses the high-pass filter to extract image contours. We accordingly suppose that the high-frequency information is vital for the extractions of boundary clues and details. We implement our strategy and conduct experiments on the Cityscapes dataset, which demonstrate the effectiveness of our high-frequency graph convolution block on semantic segmentation. The proposed method achieves comparable performance and dramatically improves the performance of small objects like the rider, traffic signs, etc.
Zitang Sun, Ruojing Wang, Zhengbo Luo, Weili Chen
ICASSP3
2021 Sub-Band Grouping Spectral Feature-Attention Block for Hyperspectral Image Classification
abstract
Hyperspectral images (HSIs) consists of 2D spatial information and 1D spectral signature due to its specialty. Most models take the raw spectral signature as the input directly by regarding the spectral data as a sequence, which cannot fully explore the redundant and complementary information inside the spectral bands. In this paper, we proposed a novel sub-band grouping recurrent neural network (RNN) model with gated recurrent units (GRUs) to find the intrinsic feature in spectral information. We introduced the inter-band spectral cross-correlation measurement to see the high correlated groups of adjacent bands firstly. And then we concatenated the representative features from all groups for complementarity. The novel spectral feature-attention block was proposed to compound the mentioned steps and generated a much sparser feature representation for subsequent analysis. The experiment results illustrated the outstanding performances and got almost 1% and 5% improvement compared with the latest methods on two famous datasets.
Weilian Zhou, Zhengbo Luo
ICASSP3
2021 Transformer And Node-Compressed Dnn Based Dual-Path System For Manipulated Face Detection
abstract
Deep neural networks (DNNs) have extensively promoted data generation development; the quality of these generated content has achieved an impressive new level. Therefore, manipulated content, especially facial manipulation, is a growing concern for online information legitimacy. Most current deep learning-based methods depend on local features sampled by convolutional kernels and lack knowledge globally. To address the problem, we propose a dual-path pipeline using Neural Ordinary Differential Equations (NODE) based neural network and facial-feature biased transformer to deal with the visual content from a different view. The transformer path could link these landmarks in a long-range, moreover, we adopt an attention guided augmentation based self-ensemble for more robust performance. Extensive experiments show that our system could surpass several commonly used approaches in terms of video-level accuracy and AUC with better interpretability.
Zhengbo Luo, Zitang Sun
ICIP1
2021 Polynomial approximation based spectral dual graph convolution for scene parsing and segmentation
Zitang Sun, Ruojing Wang, Zhengbo Luo
Neurocomputing3