Qingzhi Zou

dblp:260/5392 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Depression and anxiety detection method based on serialized facial expression imitation
Xingyun Li, Qingzhi Zou, Qingxiang Wang
Eng. Appl. Artif. Intell.5
2024 Hybrid Deep Relation Matrix Bidirectional Approach for Relational Triple Extraction
abstract
Relation extraction is a crucial task within information extraction, and numerous models have demonstrated impressive results. However, most of the tagging-based relation triple extraction methods employ unidirectional approaches to extract subjects, objects, and relations, which may overlook crucial information. In this paper, we introduce a novel deep matrix-based bidirectional relation extraction model. Firstly, we extract forward and backward entity pairs. During the bidirectional extraction process,there may be some redundant relationships,so we use a shared encoder to connect and enhance the extraction process. Secondly, we design a low-complexity relation extraction matrix to allocate all possible relations. We assess our model using diverse benchmark datasets, and comprehensive experiments show that our approach effectively addresses subsequent triple extraction issues stemming from entity extraction failures.
Yushuai Hu, Qingzhi Zou, Ronghuan Zhang
CSCWD5
2024 FEST: Feature Enhancement Swin Transformer for Remote Sensing Image Semantic Segmentation
abstract
The global context is crucial for the precise segmentation of remote sensing images. However, the large volumes and high spatial resolutions of remote sensing images make efficient analysis of the entire scene challenging for most convolutional neural network (CNN)-based methods. To address this issue, we propose to design an innovative framework for semantic segmentation of remote sensing images called Feature Enhancement Swin Transformer (FEST). Firstly, we utilize the Swin Transformer as the encoder and incorporates a Global Information Enhancement Model (GIEM) within each Swin Transformer block to reduce information loss and enable encoding of more accurate spatial information. Secondly, we introduce an enhanced decoding structure called Enhanced Feature Fusion Module (EFFM) with added enhanced channel and spatial attention modules to retain localized information while obtaining extensive contextual information. Finally, for loss calculation, we utilize the dice and cross-entropy loss to jointly supervise the model, aiming to achieve a competitive performance. We comprehensively evaluated FEST on the ISPRS-Vaihingen and Potsdam datasets. The results indicate that our approach has achieved significant improvements in semantic segmentation tasks compared to existing methods.
Ronghuan Zhang, Qingzhi Zou
CSCWD4
2024 Medical Image Segmentation Use Convolutional Attention Augmentation TransUNet with Skip Connection Enhancement
abstract
The primary objective of medical image segmentation is the isolation of pathological tissues and distinct organs from medical images, thereby assisting in medical diagnosis. The methods employed for medical image segmentation encompass convolutional neural networks and transformer-based methods. Although the self-attention mechanism in Transformers improves its capability to capture long-range dependencies, it have limitations in learning local (contextual) relationships between pixels. Some previous research has tried to solve this problem by incorporating convolutional layers into the encoder or decoder of transformers, but sometimes feature inconsistencies arise. To better extract local features from images using convolutional neural networks and to fuse low-resolution and high-resolution features from higher-level and lower features, we propose the Convolutional Attention Augmentation TransUNet (CAA-TransUNet) model. In our model, firstly, we propose a convolutional attention augmentation module that enhances both local and global features by suppressing irrelevant background information. Secondly, we have integrated attention gates into the skip connections to aggregate feature information from various stages of the encoder during the upsampling process. Finally, we employ the technique of aggregating the loss of multi-stage features to expedite convergence speed and enhance performance. The experimental results on three public datasets demonstrate that our proposed model significantly outperforms the baseline methods.
Qingzhi Zou, Ronghuan Zhang, Yushuai Hu
CSCWD1
2024 LGDB-Net: Dual-Branch Path for Building Extraction from Remote Sensing Image
abstract
Extracting buildings from remote sensing images using deep learning techniques is a widely applied and crucial task. Convolutional Neural Networks (CNNs) adopt hierarchical feature representation, showcasing powerful capabilities in extracting local information but facing challenges in capturing global features. Transformers can address this limitation, but they perform poorly in extracting local features and significantly increase memory requirements and computational complexity. To overcome these challenges, we propose a method for building extraction from remote sensing images called LGDB-Net (Local-Global Dual-Branch Network), employing a dual-branch approach. Firstly, inspired by Swin Transformer, we designed GB-Former(Global Branch-Former) as the backbone network to model global information. We use a linear multi-head self-attention mechanism to reduce time and memory complexity while maintaining a large global receptive field. Additionally, we replace the traditional multi-layer perceptron with a convolution-enhanced multi-layer perceptron to improve channel feature representation, reduce model parameters, and enhance segmentation performance. Secondly, we use multiple Depth-wise Conv3×3 + LN (Layer Normalization) + GeLU (Gaussian Error Linear Unit) modules as the auxiliary branch for local detailed feature extraction. Finally, we adopt a multi-scale feature fusion strategy to integrate feature information from both branches. We conduct a series of experiments on three datasets: WHU, Massachusetts, and Inria. The experimental results demonstrate that the proposed method not only effectively improves segmentation accuracy with lower building omission and commission rates but also significantly reduces model parameters and computational complexity.
Ronghuan Zhang, Qingzhi Zou
ICPADS4
2024 MedX-Net: Hierarchical Transformer with Large Kernel Convolutions for 3D Medical Image Segmentation
abstract
Due to the exceptional performance of Transform-ers in 2$D$medical image segmentation, recent work has also introduced them into 3D medical segmentation tasks. For instance, Swin UNETR and other hierarchical Transformers have reintroduced prior knowledge from several convolutional networks, further enhancing the volume segmentation capabilities of models. The efficacy of these hybrid methodologies is primarily attributed to the substantial quantity of parameters and the nonlocal self-attention mechanism with a large receptive field. We argue that the behavior of these methods' large receptive fields can be simulated by employing fewer parameters through the utilization of depth-wise convolutions with large kernel. Within this manuscript, we introduce a lightweight volume segmentation model called MedX-Net, which uses convolutional network modules to simulate hierarchical Transformers for robust volume segmentation. Firstly, inspired by the hierarchical Transformer module of Swin UNETR, we investigate large-kernel depth-wise convolutions with different sizes to achieve a reduced model parameter count while maintaining a large global receptive field. Secondly, we replace the multilayer perceptron (MLP) in the hierarchical Transformer module with Inverted Bottleneck with Depthwise Convolution Enhancement(DWCE) to improve model performance with fewer activation and normalization layers, further reducing the parameter count. We validate the effectiveness and efficiency of our model for volume segmentation on three public datasets: Synapse, BTCV and ACDC. On the Synapse dataset, compared to Swin UNETR, our model achieves an improvement from 83.48% to 87.21% in Dice score. Compared to the result of 86.57% achieved by nnFormer, our model achieves superior performance while reducing the model parameter count by 64%.
Qingzhi Zou
SMC2
2023 PAT-Unet: Paired Attention Transformer for Efficient and Accurate Segmentation of 3D Medical Images
Qingzhi Zou
PRCV (13)1