Zongwen Bai

dblp:52/10178 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 A unified adversarial framework for super resolution of cardiac cine magnetic resonance imaging with arbitrary scale factors
Zongwen Bai, Shipeng Cheng, Meili Zhou, Dong Wang 0022, Zhouzhou Zhu
Eng. Appl. Artif. Intell.1
2025 HyperDiff: Masked Diffusion Model with High-efficient Transformer for Hyperspectral Image Cross-Scene Classification
abstract
Hyperspectral Image (HSI) cross-scene classification is a challenging task in remote sensing, particularly when real-time processing of Target Domain (TD) HSI is required, and data cannot be reused for training. While deep learning methods have shown promising results, the generalization ability of HSI representations remains limited, mainly due to class label imbalance. This paper introduces a dual-stage learning framework based on transfer learning to enhance classification accuracy in the TD. The framework includes a self-supervised learning stage and a supervised fine-tuning stage. The self-supervised stage focuses on learning robust representations by leveraging inherent structures within HSI data, while the fine-tuning stage uses training labels to extract semantic information. A masked diffusion model predicts masked tokens from unmasked ones, capturing both high-level structures and fine details in HSI data. An efficient spatiospectral Transformer, which removes self-attention from the decoder, is proposed to enhance the self-supervised process. This design allows mask tokens to obtain information from visible tokens without interacting with each other, reducing sequence length and computational costs. By decoding each mask token conditionally independently, only a subset of masked tokens is processed. Extensive experiments on two public HSI datasets demonstrate that the proposed method outperforms state-of-the-art techniques.
Dong Wang 0022, Chanyue Wu, Jing Yang 0026, Zongwen Bai, Ying Li 0017, Qiang Shen 0001
ICASSP6
2024 Tripartite-structure transformer for hyperspectral image classification
abstract
Abstract Hyperspectral images contain rich spatial and spectral information, which provides a strong basis for distinguishing different land‐cover objects. Therefore, hyperspectral image (HSI) classification has been a hot research topic. With the advent of deep learning, convolutional neural networks (CNNs) have become a popular method for hyperspectral image classification. However, convolutional neural network (CNN) has strong local feature extraction ability but cannot deal with long‐distance dependence well. Vision Transformer (ViT) is a recent development that can address this limitation, but it is not effective in extracting local features and has low computational efficiency. To overcome these drawbacks, we propose a hybrid classification network that combines the strengths of both CNN and ViT, names Spatial‐Spectral Former(SSF). The shallow layer employs 3D convolution to extract local features and reduce data dimensions. The deep layer employs a spectral‐spatial transformer module for global feature extraction and information enhancement in spectral and spatial dimensions. Our proposed model achieves promising results on widely used public HSI datasets compared to other deep learning methods, including CNN, ViT, and hybrid models.
Liuwei Wan, Meili Zhou, Shengqin Jiang, Zongwen Bai, Haokui Zhang
Comput. Intell.4
2023 A superior image inpainting scheme using Transformer-based self-supervised attention GAN model
Meili Zhou, Xiangzhen Liu, Tingting Yi, Zongwen Bai
Expert Syst. Appl.4
2023 Denoising Aggregation of Graph Neural Networks by Using Principal Component Analysis
abstract
To avoid the overfitting phenomenon that appeared in performing graph neural networks (GNNs) on test examples, the feature encoding scheme of GNNs usually introduces the dropout procedure. However, after learning latent node representations under this scheme, Gaussian noise produced by the dropout operation is inevitably transmitted into the next neighborhood aggregation step, which necessarily hampers the unbiased aggregation ability of GNN models. To address this issue, in this article, we present a novel aggregator, denoising aggregation (DNAG), which utilizes principal component analysis (PCA) to preserve the aggregated real signals from neighboring features and simultaneously filter out the Gaussian noise. The idea is different from using PCA on traditional applications to reduce the feature dimension. We regard PCA as an aggregator to compress the neighboring node features to have better expressive denoising power. We propose new training architectures to simplify the intensive computation of PCA in DNAG. Numerical experiments show the apparent superiority of the proposed DNAG models in gaining more denoising capability and achieving the state of the art for a set of predictive tasks on several graph-structured datasets.
Wei Dong 0010, Marcin Wozniak, Junsheng Wu, Weigang Li 0005, Zongwen Bai
IEEE Trans. Ind. Informatics5
2022 Memory-Efficient Hierarchical Neural Architecture Search for Image Restoration
Haokui Zhang, Ying Li 0017, Hao Chen 0041, Chengrong Gong, Zongwen Bai, Chunhua Shen
Int. J. Comput. Vis.5
2022 Improving performance and efficiency of Graph Neural Networks by injective aggregation
Wei Dong 0010, Junsheng Wu, Xinwan Zhang, Zongwen Bai, Peng Wang 0023, Marcin Wozniak
Knowl. Based Syst.4
2022 Grafting Transformer on Automatically Designed Convolutional Neural Network for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification has been a hot topic for decides, as hyperspectral images have rich spatial and spectral information and provide strong basis for distinguishing different land-cover objects. Benefiting from the development of deep learning technologies, deep learning based HSI classification methods have achieved promising performance. Recently, several neural architecture search (NAS) algorithms have been proposed for HSI classification, which further improve the accuracy of HSI classification to a new level. In this paper, NAS and Transformer are combined for handling HSI classification task for the first time. Compared with previous work, the proposed method has two main differences. First, we revisit the search spaces designed in previous HSI classification NAS methods and propose a novel hybrid search space, consisting of the space dominated cell and the spectrum dominated cell. Compared with search spaces proposed in previous works, the proposed hybrid search space is more aligned with the characteristic of HSI data, that is, HSIs have a relatively low spatial resolution and an extremely high spectral resolution. Second, to further improve the classification accuracy, we attempt to graft the emerging transformer module on the automatically designed convolutional neural network (CNN) to add global information to local region focused features learned by CNN. Experimental results on three public HSI datasets show that the proposed method achieves much better performance than comparison approaches, including manually designed network and NAS based HSI classification methods. Especially on the most recently captured dataset Houston University, overall accuracy is improved by nearly 6 percentage points. Code is available at: https://github.com/Cecilia-xue/HyT-NAS.
Xizhe Xue, Haokui Zhang, Bei Fang, Zongwen Bai, Ying Li 0017
IEEE Trans. Geosci. Remote. Sens.4
2022 3-D-ANAS: 3-D Asymmetric Neural Architecture Search for Fast Hyperspectral Image Classification
abstract
Hyperspectral images (HSIs) provide abundant spectral and spatial information, playing an irreplaceable role in land-cover classification. Recently, based on deep learning (DL) technologies, an increasing number of HSI classification approaches have been proposed, which demonstrate promising performance. However, previous studies suffer from two major drawbacks: 1) the architecture of most DL models is manually designed, relies on specialized knowledge, and is relatively tedious. Moreover, in HSI classifications, datasets captured by different sensors have different physical properties. Correspondingly, different models need to be designed for different datasets, which further increases the workload of designing architectures and 2) the mainstream framework is a patch-to-pixel framework. The overlap regions of patches of adjacent pixels are calculated repeatedly, which increases computational cost and time cost. In addition, the classification accuracy is sensitive to the patch size, which is artificially set based on extensive investigation experiments. To overcome the issues mentioned above, we first propose a 3-D asymmetric neural network search algorithm and leverage it to automatically search for efficient architectures for HSI classifications. By analyzing the characteristics of HSIs, we specifically build a 3-D asymmetric decomposition search space, where spectral and spatial information is processed with different decomposition convolutions. Furthermore, we propose a new fast classification framework, i.e., pixel-to-pixel classification framework, which has no repetitive operations and reduces the overall cost. Experiments on three public HSI datasets captured by different sensors demonstrate the networks designed by our 3-D asymmetric neural architecture search (3-D-ANAS) achieve competitive performance compared to several state-of-the-art methods, while having a much faster inference speed. Code is available at:https://github.com/hkzhang91/3D-ANAS.
Haokui Zhang, Chengrong Gong, Yunpeng Bai, Zongwen Bai, Ying Li 0017
IEEE Trans. Geosci. Remote. Sens.4
2022 Deep Neural Network Heuristic Hierarchization for Cooperative Intelligent Transportation Fleet Management
abstract
In this article, we propose malfunction classifications for trucks, a novel idea for smart fleet management systems. In the proposed cooperative cooperative intelligent transportation (C-ITS), the developed neural network work with information from truck fleets to select the trucks that need a service. From the results returned from the deep neural network classifier, the applied heuristic algorithm uses the classification outputs to select the most important results. The proposed process is multithreaded; thus, the composed system gains additional efficiency. The implemented deep learning model achieved an accuracy above 98%, and an above 95% recall. The developed solution was tested on the Scania Truck data collection. The research results show the importance of the advances and validate our concept for potential further development.
Qiao Ke, Jakub Silka, Michal Wieczorek 0002, Zongwen Bai, Marcin Wozniak
IEEE Trans. Intell. Transp. Syst.4
2021 Orthogonal Dual Graph Regularized Nonnegative Matrix Factorization
Jinrong He, Yanxin Shi, Zongwen Bai, Zeyu Bao
ICIG (2)3
2021 Meta Transfer Learning for Few-Shot Hyperspectral Image Classification
abstract
We propose a novel meta-learning approach for few-shot hyperspectral image (HSI) classification, which learns to distil transferable prior knowledge from a base dataset with sufficient labeled samples and generalize the knowledge to an unseen dataset with extremely limited labeled samples for performance improvement. Specifically, we first construct a backbone classification model using an embedding module and a linear classifier. Then, we sample extensive synthetic few-shot tasks from the base dataset, each of which consists of a support set with limited labeled samples and a query set with some unlabeled test samples. Given these tasks, we propose to optimize the embedding module using an episode learning scheme where for each task we train the linear classier based on an initialized embedding module using the support set and ultimately optimize the embedding module based on the test error on the query set until the test error on all tasks is minimized. By doing this, the resultant embedding module is able to appropriately generalize to an unseen few-shot classification task and lead to good performance with the linear classifier. Experiments on two standard classification benchmarks under different few-shot settings demonstrate the efficacy of the proposed method.
Fei Zhou 0008, Lei Zhang 0054, Wei Wei 0008, Zongwen Bai, Yanning Zhang 0001
IGARSS4
2021 DecomVQANet: Decomposing visual question answering deep network via tensor decomposition and regression
Zongwen Bai, Ying Li 0017, Marcin Wozniak, Meili Zhou, Di Li 0006
Pattern Recognit.1
2021 MobileGCN applied to low-dimensional node feature learning
Wei Dong 0010, Junsheng Wu, Zongwen Bai, Yaoqi Hu, Weigang Li 0005, Marcin Wozniak
Pattern Recognit.3
2021 A hierarchical sampling based triplet network for fine-grained image classification
Guiqing He, Qiyao Wang, Zongwen Bai, Yuelei Xu
Pattern Recognit.4
2021 D3D: Dual 3-D Convolutional Network for Real-Time Action Recognition
abstract
Three-dimensional convolutional neural networks (3D CNNs) have been explored to learn spatio-temporal information for video-based human action recognition. Expensive computational cost and memory demand resulted from standard 3D CNNs, however, hinder their application in practical scenarios. In this article, we address the aforementioned limitations by proposing a novel dual 3-D convolutional network (D3DNet) with two complementary lightweight branches. A coarse branch maintains large temporal receptive field by a fast temporal downsampling strategy and simulates the expensive 3-D convolutions using a combination of more efficient spatial convolutions and temporal convolutions. Meanwhile, a fine branch progressively downsamples the video in the temporal domain and adopts 3-D convolutional units with reduced channel capacities to capture multiresolution spatio-temporal information. Instead of learning these two branches independently, a shallow spatiotemporal downsampling module is shared for these two branches for efficient low-level feature learning. Besides, lateral connections are learned to effectively fuse the information from the two branches at multiple stages. The proposed network makes good balance between inference speed and action recognition performance. Based on RGB information only, it achieves competing performance on five popular video-based action recognition datasets, with inference speed of 3200 FPS on a single NVIDIA GTX 2080Ti card.
Shengqin Jiang, Yuankai Qi, Haokui Zhang, Zongwen Bai, Xiaobo Lu, Peng Wang 0023
IEEE Trans. Ind. Informatics4
2021 Intelligent Internet of Things System for Smart Home Optimal Convection
abstract
The fusion of Internet of Things (IoTs) and computational intelligence makes it possible to increase energetic efficiency of our homes. Connected devices can be optimally adjusted to the needs of a family. In this article, we present our developed IoT convection installation for a small house with the developed remote platform control system. The control module is gartering readings from sensors and information from users about conditions in the house and, by the use of computational intelligence, optimizes parameters to adjust the developed IoT convection system for better comfort of a family. We have done a full convection installation, both in practical and theoretical models, together with remote control system and the proposed security model. Optimization results show increased comfort of use with lower changes in the temperature inside. The system after optimization shows significant improvement in lower changes of the temperature and lower consumption.
Adam Zielonka, Andrzej Sikora, Marcin Wozniak, Wei Wei 0006, Qiao Ke, Zongwen Bai
IEEE Trans. Ind. Informatics6
2020 Bilinear Semi-Tensor Product Attention (BSTPA) model for visual question answering
abstract
We propose a semi-tensor product attention network model as a visual question answering tool for complex interaction over image features. Proposed model performs matrix multiplication of two arbitrary dimensions, which is used to overcome possible dimensional limitations and improve recognition flexibility. In used block-wise operation we preserve spatial and temporal information but reduce the number of parameters by using low-rank pooling scheme. Applied BERT pre-train model is tuned to recognize question features. The proposed model is evaluated on the VQA2.0 dataset. Research results show that our model has good accuracy and easy reconfiguration for future research.
Zongwen Bai, Ying Li 0017, Meili Zhou, Di Li 0006, Dong Wang 0022, Dawid Polap, Marcin Wozniak
IJCNN1
2020 Design of affinity-aware encoding by embedding graph centrality for graph classification
Wei Dong 0010, Junsheng Wu, Zongwen Bai, Weigang Li 0005
Neurocomputing3