VLDB 2026 Research / reviewers in the wild / expert
Weilian Zhou
dblp:226/5414
· DBLP profile ↗
17ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0002-2071-1489ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FNWP: A fractal network with wavelet propagation for robust deep learningabstractTraining deep neural networks (DNNs) remains challenging due to dying activations and structural loss, especially in computer vision. While complex networks with shortcut paths often outperform plain networks empirically, a clear theoretical explanation is lacking. Moreover, these networks still suffer from structural degradation. In this paper, we provide a theoretical analysis showing that neuron survival in complex DNNs is lower bounded by the probability of their shortest path length, explaining their improved trainability. We also show that downsampling causes significant structural loss and performance decline. To address these issues, we propose a Fractal Network with Wavelet Propagation (FNWP). Guided by shortest path length theory, FNWP incorporates a contraction mapping to improve neuron survival without added complexity. Its Wavelet Propagation module performs dynamic multi-scale wavelet decomposition to reduce structural loss. Furthermore, we introduce Approximation Convolution, a drop-in upgrade to standard convolutions with no extra cost. FNWP achieves 81.9% top-1 accuracy on CIFAR-100 with 14.4M parameters and 90.5% F1 score on the Kuzushiji dataset for high-resolution optical character recognition. It also delivers strong performance on general object detection tasks and time series classification benchmarks, and shows strong robustness to depth, making it a scalable architecture for deep learning. Pengfeng Lu, Mengyunqiu Zhang, Weilian Zhou |
Neurocomputing | 4 |
| 2026 | HSIseg: Progressively enhanced extensible multi-modality framework for large patch-wise hyperspectral image segmentationabstractHyperspectral image (HSI) classification plays a critical role in remote sensing by enabling precise land-cover identification through rich spectral information. While deep learning has led to significant progress, over 90 % of existing methods rely on small patch-based networks, which suffer from two key limitations: (1) restricted receptive fields (e.g., 7 × 7 , 9 × 9 ) that hinder structural awareness and result in noisy misclassifications within homogeneous regions; and (2) undefined optimal patch sizes that lead to coarse label predictions and degraded accuracy. Inspired by large-scale image segmentation techniques such as U-Net architectures—known for their strong boundary delineation and spatial coherence—this study explores their adaptation to HSI classification. However, such adaptations remain underutilized due to challenges including performance concerns with large patches, abundant unlabeled regions, and input-shape mismatches. To address these gaps, we propose HSIseg, a segmentation-adapted framework for HSI classification. HSIseg incorporates three novel components—Dynamic Shifted Regional Transformer (DSRT), Discriminative Feature Selection (DFS), and Cross Feature Interaction (CFI)—to enhance feature representation and fusion. A progressive learning strategy with adaptive pseudo-labeling is employed to leverage unlabeled data, while multi-source data collaboration further strengthens the model’s capability. Extensive experiments on five public datasets demonstrate the effectiveness of HSIseg in overcoming the limitations of traditional patch-based approaches. Code is available at https://github.com/zhouweilian1904/HSI_Segmentation . Weilian Zhou, Weixuan Xie, Huiying (cynthia) Hou, Man Sing Wong, Haipeng Wang 0002 |
Neurocomputing | 1 |
| 2025 | Mamba-in-Mamba: Centralized Mamba-Cross-Scan in Tokenized Mamba Model for Hyperspectral image classificationabstractHyperspectral image (HSI) classification plays a crucial role in remote sensing (RS) applications, enabling the precise identification of materials and land cover based on spectral information. This supports tasks such as agricultural management and urban planning. While sequential neural models like Recurrent Neural Networks (RNNs) and Transformers have been adapted for this task, they present limitations: RNNs struggle with feature aggregation and are sensitive to noise from interfering pixels, whereas Transformers require extensive computational resources and tend to underperform when HSI datasets contain limited or unbalanced training samples. To address these challenges, Mamba architectures have emerged, offering a balance between RNNs and Transformers by leveraging lightweight, parallel scanning capabilities. Although models like Vision Mamba (ViM) and Visual Mamba (VMamba) have demonstrated improvements in visual tasks, their application to HSI classification remains underexplored, particularly in handling land-cover semantic tokens and multi-scale feature aggregation for patch-wise classifiers. In response, this study introduces the Mamba-in-Mamba (MiM) architecture for HSI classification, marking a pioneering effort in this domain. The MiM model features: (1) a novel centralized Mamba-Cross-Scan (MCS) mechanism for efficient image-to-sequence data transformation; (2) a Tokenized Mamba (T-Mamba) encoder that incorporates a Gaussian Decay Mask (GDM), Semantic Token Learner (STL), and Semantic Token Fuser (STF) for enhanced feature generation; and (3) a Weighted MCS Fusion (WMF) module with a Multi-Scale Loss Design for improved training efficiency. Experimental results on four public HSI datasets—Indian Pines, Pavia University, Houston2013, and WHU-Hi-Honghu—demonstrate that our method achieves an overall accuracy improvement of up to 3.3%, 2.7%, 1.5%, and 2.3% over state-of-the-art approaches (i.e., SSFTT, MAEST, etc.) under both fixed and disjoint training-testing settings. • A novel multi-scale pyramid Mamba model for efficient HSI classification. • Tokenized Mamba encoder enhances Mamba’s suitability for visual tasks. • Centralized Mamba-Cross-Scan improves patch-wise HSI sequential classifiers. • Satisfying classification performance with fixed and disjoint training-testing samples. Weilian Zhou, Haipeng Wang 0002, Man Sing Wong, Huiying (cynthia) Hou |
Neurocomputing | 1 |
| 2024 | Adversarial Detection Transformer For Kuzushiji RecognitionabstractKuzushiji recognition is an optical character recognition task that aims to recognize ancient Japanese characters. Kuzushiji has over 4000 categories and some characters are very similar, which poses a challenge for existing recognition methods. Moreover, Kuzushiji characters are often small and connected, making it difficult to locate character positions. To overcome these problems, we propose Adversarial Detection Transformer (Adversarial-DETR), a method that learns to fight against noisy boxes and class labels to reconstruct ground truth. In this paper, we assume that direct model predictions are noise predictions and propose Real Denoising (RDN) to leverage these prediction noises. We also introduce a target-aware focal loss (TFL) to accelerate the convergence speed. Moreover, we propose a task-driven encoder-decoder (TED) structure based on the observation that different scale features excel at different tasks. Through experiments on the Kuzushiji dataset, Adversarial-DETR achieves the best performance of 0.941 F 1 score, outperforming state-of-the-art DETRs and other detection methods. Pengfeng Lu, Mengyunqiu Zhang, Weilian Zhou |
ICIP | 4 |
| 2024 | Segmented Recurrent Transformer With Cubed 3-D-Multiscanning Strategy for Hyperspectral Image ClassificationabstractThis study introduces an innovative approach in hyperspectral imaging (HSI) classification by integrating convolution, recurrence, and self-attention mechanisms in a 3D configuration. We address several challenges such as the 1) disruption of spectral continuity by traditional dimensionality reduction methods like PCA, 2) the overlooking of band-to-band continuous features in existing spatial-only 2D multiscanning strategy, and 3) the limitations in model design by simply cascading recurrent neural networks (RNNs) with Transformers for HSI analysis. Our solution involves three core components: 1) sub-band grouping with group-wise convolution for refined dimension reduction, 2) a novel cubed 3D-multiscanning technique enabling thorough multi-directional analysis in both spectral and spatial domains, and 3) the development of a Cubic-Net framework with a specially designed Segmented Recurrent Transformer (SRT). This SRT is tailored to effectively utilize spectral continuity along with spatial contextual features, overcoming common sequential data analysis challenges seen in RNNs and Transformers. Furthermore, our feature fusion strategy successively integrates ‘short-term’ and ‘long-term’ SRT features, thereby enhancing the model’s ability to process both spectral and spatial features effectively. Experimental results from three public HSI datasets indicate our method’s improved performance over existing baselines and state-of-the-art methods. This research offers a new perspective in 3D sequential HSI classification. Weilian Zhou, Haipeng Wang 0002, Pengfeng Lu, Mengyunqiu Zhang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Learned Image Compression with Multi-Scan Based Channel FusionabstractWhile classical compression standards heavily rely on scan methods, few learned compression methods tend to use scan. The lack of scan has led to an increase in the encoding bits and the loss of relative information. Besides, recent works also struggle to choose between convolutional neural networks and transformers due to the presence of both local redundancy and global semantic information. CNN methods raise the distortion problem while the transformers increase additional encoding bits. To solve these problems, in this paper, we novelly introduce the Multi-Scan based Channel Fusion (MSCF) to the hyper-prior VAE structure to reduce the redundant bits. We propose a novel Residual Local-Global Consolidation (RLGC) module that utilizes the latest ConvNeXt and Swin transformer to enhance the image quality without additional bits. Experiments have shown that our model outperforms the state-of-the-art methods in terms of PSNR, MS-SSIM and BD-rate. Weilian Zhou, Pengfeng Lu |
ICIP | 2 |
| 2023 | Multiscanning-Based RNN-Transformer for Hyperspectral Image ClassificationabstractThe goal of hyperspectral image (HSI) classification is to assign land-cover labels to each HSI pixel in a patch-wise manner. Recently, sequential models, such as recurrent neural networks (RNN), have been developed as HSI classifiers which need to scan the HSI patch into a pixel-sequence with the scanning order first. However, RNNs have a biased ordering that cannot effectively allocate attention to each pixel in the sequence, and previous methods that use multiple scanning orders to average the features of RNNs are limited by the validity of these orders. To solve this issue, it is naturally inspired by Transformer and its self-attention to discriminatively distribute proper attention for each pixel of the pixel-sequence and each scanning order. Hence, in this study, we further develop the sequential HSI classifiers by a specially designed RNN-Transformer (RT) model to feature the multiple sequential characters of the HSI pixels in the HSI patch. Specifically, we introduce a multiscanning-controlled positional embedding strategy for the RT model to complement multiple feature fusion. Furthermore, the RT encoder is proposed for integrating ordering bias and attention re-allocation for feature generation at the sequence-level. Additionally, the spectral-spatial-based soft masked self-attention is proposed for suitable feature enhancement. Finally, an additional Fusion Transformer is deployed for scanning order-level attention allocation. As a result, the whole network can achieve competitive classification performance on four accessible datasets than other state-of-the-art methods. Our study further extends the research on sequential HSI classifiers. Weilian Zhou, Haipeng Wang 0002, Xi Xue |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Deep Residual Networks with Common Linear Multi-Step and Advanced Numerical SchemesabstractDeep Neural Networks (DNNs) are widely used state-of-the-art approaches for the massive audio signal, language and vision tasks. However, we still lack sufficient understanding of the theoretical issues inside. Based on the equivalence between deep residual networks (ResNets) and the Euler forward scheme, we present several ways to construct DNNs (ResNet as the example) in advanced numerical and linear multi-step schemes. Furthermore, we show that ResNets with various advanced schemes have better accuracy, convergence, robustness and an explainable rank against the property of these schemes. Finally, previous works are summarized and discussed; sufficient experiments are provided. Improvements in several aspects are noticeable and theoretical. Zhengbo Luo, Weilian Zhou, Xuehui Hu |
ICIP | 2 |
| 2022 | Rethinking Unified Spectral-Spatial-Based Hyperspectral Image Classification Under 3D Configuration of Vision TransformerabstractVision Transformer (ViT) has been introduced into the computer vision (CV) field with its self-attention mechanism to capture global dependency. However, simply deploying ViT on a hyperspectral image (HSI) classification task can not get satisfying results because ViT is a spatial-only self-attention model, but rich spectral information exists in HSI. Moreover, most HSI classifiers integrate spectral and spatial features in a cascaded flowchart, ignoring the internal correlation between spectral and spatial information. Furthermore, existing positional embedding (PE) methods can not fulfil the 3D configuration of ViT. Therefore, this paper proposes a unified spectral-spatial-based 3D ViT with cooperative 3D coordinate positional embedding. In the meanwhile, a novel local-global feature fusion strategy is proposed. The model does not contain convolution or recurrent units and can achieve more competitive classification performance than other state-of-the-art (SOTA) methods. Furthermore, compared with existing ViT-based HSI classifiers, our concept can get better results. Weilian Zhou, Zhengbo Luo, Xi Xue |
ICIP | 1 |
| 2022 | Hierarchical Unified Spectral-Spatial Aggregated Transformer for Hyperspectral Image ClassificationabstractVision Transformer (ViT) has recently been introduced into the computer vision (CV) field with its self-attention mechanism and gotten remarkable performance. However, simply applying ViT for hyperspectral image (HSI) classification is not applicable due to 1) ViT is a spatial-only self-attention model, but rich spectral information exists in HSI; 2) ViT needs sufficient training samples, but HSI suffers from limited samples; 3) ViT does not well learn local features; 4) multi-scale features for ViT are not considered. Furthermore, the methods which combine convolutional neural network (CNN) and ViT generally suffer from a large computational burden. Hence, this paper tends to design a suitable pure ViT based model for HSI classification as the following points: 1) spectral-only vision transformer with all tokens’ aggregation; 2) spatial-only local-global transformer; 3) cross-scale local-global feature fusion, and 4) a cooperative loss function to unify the spectral and spatial features. As a result, the proposed idea achieves competitive classification performance on three public datasets than other state-of-the-art methods. Weilian Zhou, Zhengbo Luo |
ICPR | 1 |
| 2022 | Constructing infinite deep neural networks with flexible expressiveness while training
Zhengbo Luo, Zitang Sun, Weilian Zhou, Zizhang Wu |
Neurocomputing | 3 |
| 2022 | Multiscanning Strategy-Based Recurrent Neural Network for Hyperspectral Image ClassificationabstractMost methods based on the convolutional neural network show satisfying performance for hyperspectral image (HSI) classification. However, the spatial dependence among different pixels is not well learned by CNNs. A recurrent neural network (RNN) can effectively establish the dependence of nonadjacent pixels and ensure that each feature activation in its output is an activation at the specific location concerning the whole image, in contrast to the usual local context window in the CNNs. However, recent limited conversion schemes in RNN-based methods for HSI classification cannot fully capture the complete spatial dependence of an HSI patch. In this study, a novel multiscanning strategy with RNN is proposed to feature the sequential character of the HSI pixel and fully consider the spatial dependence in the HSI patch. By investigating different scanning forms, eight scanning orders are considered spatially, which flattens one local HSI patch into eight neighboring continuous pixel sequences. Moreover, considering that eight scanning orders complement one local patch with correlative dependence, the concatenated features from all scanning orders are fed into the RNN again for complementarity. As a result, the network can achieve competitive classification performance on three publicly accessible datasets using fewer parameters than other state-of-the-art methods. Weilian Zhou, Zhengbo Luo, Haipeng Wang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Deep Neural Networks with Flexible Complexity While Training Based on Neural Ordinary Differential EquationsabstractMost structures of deep neural networks (DNN) are with a fixed complexity of both computational cost (parameters and FLOPs) and the expressiveness. In this work, we experimentally investigate the effectiveness of using neural ordinary differential equations (NODEs) as a component to provide further depth to relatively shallower networks rather than stacked layers (depth) which achieved improvement with fewer parameters. Moreover, we construct deep neural networks with flexible complexity based on NODEs which enables the system to adjust its complexity while training. The proposed method achieved more parameter-efficient performance than stacking standard DNNs, and it alleviates the defect of the heavy cost required by NODEs. Zhengbo Luo, Zitang Sun, Weilian Zhou |
ICASSP | 4 |
| 2021 | Sub-Band Grouping Spectral Feature-Attention Block for Hyperspectral Image ClassificationabstractHyperspectral images (HSIs) consists of 2D spatial information and 1D spectral signature due to its specialty. Most models take the raw spectral signature as the input directly by regarding the spectral data as a sequence, which cannot fully explore the redundant and complementary information inside the spectral bands. In this paper, we proposed a novel sub-band grouping recurrent neural network (RNN) model with gated recurrent units (GRUs) to find the intrinsic feature in spectral information. We introduced the inter-band spectral cross-correlation measurement to see the high correlated groups of adjacent bands firstly. And then we concatenated the representative features from all groups for complementarity. The novel spectral feature-attention block was proposed to compound the mentioned steps and generated a much sparser feature representation for subsequent analysis. The experiment results illustrated the outstanding performances and got almost 1% and 5% improvement compared with the latest methods on two famous datasets. Weilian Zhou, Zhengbo Luo |
ICASSP | 1 |
| 2021 | Hyperspectral Image Classification Based on Multi-stage Vision Transformer with Stacked SamplesabstractHyperspectral image classification (HSIC) is a task assigning the correct label to each pixel. It is a hot topic in the remote sensing field, which has been processed in several deep learning methods. Recently, there are some works that apply Vision Transformer (ViT) methods to the HSIC task, but the performance is not as good as some CNN-structured methods, considering that Vision Transformer uses attention to capture global information but ignores local characteristics. In this paper, a multi-stage Vision Transformer model referring to the feature extraction structure of CNN is proposed, and the result shows the realizability and reliability. Besides, experiments show that the modified ViT structure needs more samples for training. An innovative data augmentation method is used to generate extended samples with virtual yet reliable labels. The generated samples are combined with the original ones as the stacked samples, which are used for the following feature extraction process. Experiments explain the optimization of the multi-stage Vision Transformer structure with stacked samples in the accuracy term compared with other methods. Weilian Zhou |
TENCON | 3 |
| 2020 | Multi-Scanning Based Recurrent Neural Network for Hyperspectral Image ClassificationabstractAs the specialty of hyperspectral image (HSI), it consists of 2D spatial and 1D spectral information. In the field of deep learning, HSI classification is an appealing research topic. Many existing methods process the HSI in spatial or spectral domain separately, which cannot fully extract the representative features, and the most used 3D convolutional neural network (3D-CNN) will suffer from mixing up complex spectral information. In this paper, we propose a spatial-spectral unified method by using recurrent neural networks (RNN) and multi-scanning direction strategy to construct spatial-spectral information sequences for learning the spatial dependencies among the central pixel and neighboring pixels. Meanwhile, residual connections and dense connections are introduced into multi-scanning direction sequences to overcome the memory problem in the RNN. The proposed method got 99.58% and 99.81% accuracy respectively on two benchmark datasets: the Pavia University dataset and the Pavia Center dataset. It demonstrates the proposed method can achieve state-of-the-art results. Weilian Zhou |
ICPR | 1 |
| 2018 | Learning Gaussian Graphical Models Using Discriminated Hub Graphical LassoabstractWe develop a new method called Discriminated Hub Graphical Lasso (DHGL) based on Hub Graphical Lasso (HGL) by providing the prior information of hubs. We apply this new method in two situations: with known hubs and without known hubs. Then we compare DHGL with HGL using several measures of performance. When some hubs are known, we can always estimate the precision matrix better via DHGL than HGL. When no hubs are known, we use Graphical Lasso (GL) to provide information of hubs and find that the performance of DHGL will always be better than HGL if correct prior information is given, and will rarely degenerate when the prior information is incorrect. Jingtian Bai, Weilian Zhou |
ICASSP | 3 |