VLDB 2026 Research / reviewers in the wild / expert
Qian Zhang 0003
dblp:04/2024-3
· DBLP profile ↗
13ranked-venue papers
3as first author
9since 2021 · last 2024
0000-0003-3041-643XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | LSRFormer: Efficient Transformer Supply Convolutional Neural Networks With Global Information for Aerial Image SegmentationabstractBoth local context and global context information are essential for the semantic segmentation of aerial images. Convolutional Neural Networks (CNNs) can capture local context information well but cannot model the global dependencies. Vision transformers (ViTs) are good at extracting global information but cannot retain the spatial details well. In order to leverage the advantages of these two paradigms, we integration them in one model in this study. However, global token interaction of ViT brings high computational cost, which makes it difficult to apply to large-sized aerial images. To handle this problem, we propose a novel efficient ViT block named long-short-range transformer (LSRFormer). Instead of mainstream ViTs designed as backbones, LSRFormer is a pre-training-free and plug-and-play module to be appended after CNN stages to supplement the global information. It is composed of long-range self-attention (LR-SA), short-range self-attention (SR-SA), and multi-scale-convolutional feed-forward-network (MSC-FFN). LR-SA establishes long-range dependencies at the junction of the windows and SR-SA diffuses the long-range information from window boundary to internal. MSC-FFN can capture multi-scale information inside the ViT block. We append LSRFormer block after each CNN stage of a pure convolutional network to build a model named ConvLSR-Net. Compared with existing models which combining CNN and ViTs, our model can learn both local and global representation at all stages of the model. In particular, ConvLSR-Net achieves state-of-the-art (SOTA) results on four challenging aerial image segmentation benchmarks, including iSAID, LoveDA, ISPRS Potsdam and Vaihingen. Code has been released at https://github.com/stdcoutzrh/ConvLSR-Net. Renhe Zhang, Qian Zhang 0003, Guixu Zhang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | MCL-DTI: using drug multimodal information and bi-directional cross-attention learning method for predicting drug-target interactionabstractBACKGROUND: Prediction of drug-target interaction (DTI) is an essential step for drug discovery and drug reposition. Traditional methods are mostly time-consuming and labor-intensive, and deep learning-based methods address these limitations and are applied to engineering. Most of the current deep learning methods employ representation learning of unimodal information such as SMILES sequences, molecular graphs, or molecular images of drugs. In addition, most methods focus on feature extraction from drug and target alone without fusion learning from drug-target interacting parties, which may lead to insufficient feature representation. MOTIVATION: In order to capture more comprehensive drug features, we utilize both molecular image and chemical features of drugs. The image of the drug mainly has the structural information and spatial features of the drug, while the chemical information includes its functions and properties, which can complement each other, making drug representation more effective and complete. Meanwhile, to enhance the interactive feature learning of drug and target, we introduce a bidirectional multi-head attention mechanism to improve the performance of DTI. RESULTS: To enhance feature learning between drugs and targets, we propose a novel model based on deep learning for DTI task called MCL-DTI which uses multimodal information of drug and learn the representation of drug-target interaction for drug-target prediction. In order to further explore a more comprehensive representation of drug features, this paper first exploits two multimodal information of drugs, molecular image and chemical text, to represent the drug. We also introduce to use bi-rectional multi-head corss attention (MCA) method to learn the interrelationships between drugs and targets. Thus, we build two decoders, which include an multi-head self attention (MSA) block and an MCA block, for cross-information learning. We use a decoder for the drug and target separately to obtain the interaction feature maps. Finally, we feed these feature maps generated by decoders into a fusion block for feature extraction and output the prediction results. CONCLUSIONS: MCL-DTI achieves the best results in all the three datasets: Human, C. elegans and Davis, including the balanced datasets and an unbalanced dataset. The results on the drug-drug interaction (DDI) task show that MCL-DTI has a strong generalization capability and can be easily applied to other tasks. Qian Zhang 0003 |
BMC Bioinform. | 4 |
| 2023 | DSAT-Net: Dual Spatial Attention Transformer for Building Extraction From Aerial ImagesabstractBoth local and global context dependencies are essential for building extraction from remote sensing (RS) images. Convolutional Neural Network (CNN) can extract local spatial details well but lacks the ability to model long-range dependency. In recent years, Vision Transformer (ViT) have shown great potential in modeling global context dependency. However, it usually brings huge computational cost, and spatial details can not be fully retained in the process of feature extraction. To maximize the advantages of CNNs and ViTs, we propose DSAT-Net, which combine them in one model. In DSAT-Net, we design an efficient Dual Spatial Attention Transformer (DSAFormer) to solve the defects of standard ViT. It has a dual attention structure to complement each other. Specifically, the global attention path (GAP) conducts a large scale down sampling of the feature maps before the global self-attention computing, to reduce the computational cost. The local attention path (LAP) uses efficient stripe convolution to generate local attention, which can alleviate the loss of information caused by down-sampling operation in the GAP and supplement the spatial details. In addition, we design a feature refining module called Channel Mixing Feature Refine Module (CM-FRM) to fuse low-level and high-level features. Our model achieved competitive results on three public building extraction datasets. Code will be available at: https://github.com/stdcoutzrh/BuildingExtraction. Renhe Zhang, Zhechun Wan, Qian Zhang 0003, Guixu Zhang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | SDSC-UNet: Dual Skip Connection ViT-Based U-Shaped Model for Building ExtractionabstractBenefiting from effective global information interaction, vision-transformers (ViTs) have been widely used in the building extraction task. However, buildings in remote sensing (RS) images usually differ greatly in size. Mainstream ViT-based segmentation models for RS images are based on Swin Transformer, which lacks multi-scale information inside the ViT block. In addition, they only connect the output of the entire ViT encoder block to the decoder, which ignore the similarity information of the attention maps inside the ViT encoder block, and are unable to provide better global dependencies for the decoder. To solve above problems, we introduce a novel Shunted Transformer, which enables the model to capture multi-scale information internally while fully establishing global dependencies, to build a pure ViT-based U-shaped model for building extraction. Furthermore, unlike the previous single-skip-connection structure of U-shaped methods, we build a novel dual skip connection structure inside the model. It simultaneously transmits the attention maps inside the ViT encoder block and its entire output to the decoder, thereby fully mining the information of the ViT encoder block and providing better global information guidance for the decoder. Thus, our model is named Shunted Dual Skip Connection UNet (SDSC-UNet). We also design a feature fusion module called Dual Skip Upsample Fusion Module (DSUFM) to aggregate the information. Our model has yields state-of-the-art (SOTA) performance (83.02%IoU) on the Inria Aerial Image Labeling Dataset. Code will be available. Renhe Zhang, Qian Zhang 0003, Guixu Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Semantic Segmentation Network Using Local Relationship Upsampling for Remote Sensing ImagesabstractSemantic segmentation is a fundamental task in remote sensing image processing. It provides pixel-level classification, which is important for many applications, such as building extraction and land use mapping. The development of convolutional neural network has considerably improved the performance of semantic segmentation. Most semantic segmentation networks are the encoder–decoder structure. Bilinear interpolation is an ordinary upsampling method in the decoder, but bilinear interpolation only considers its own features and inserts three times its own features. This over-simple and data-independent bilinear upsampling may lead to suboptimal results. In this work, we propose an upsampling method based on local relations to replace bilinear interpolation. Upsampling is performed by correlating the local relationship of feature maps of adjacent stages, which can better integrate local and global information. We also design a fusion module based on local similarity. Our proposed method with ResNet101 as the backbone of the segmentation network can improve the average$F_{1}$score and overall accuracy of the Vaihingen data set by 2.69% and 1.31%, respectively. Our proposed method also has fewer parameters and less inference time. Baokai Lin, Guang Yang 0068, Qian Zhang 0003, Guixu Zhang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Low-Level Feature Enhancement Network for Semantic Segmentation of BuildingsabstractIn recent years, convolutional neural networks (CNNs) have been widely used in extracting buildings from remote sensing images. Both semantic representation and spatial location details are crucial for this task. We propose methods to enhance the performance of semantic segmentation by using these low-level features considering that man-made buildings in aerial images have strong textures and edges. Texture Enhancement Attention Module (TEAM) is proposed to strengthen feature in the position with rich texture and improve the semantic representation. Edge Extraction Module (EEM) is applied for directly guiding spatial details learning, which starts with super-resolution maps created by Super-Resolution Module (SRM). Detail Supplement Module (DSM) is designed to further provide details for decoder. On this basis, we propose a low-level feature enhancement network (LFENet) for semantic segmentation of buildings. Experiment results on two aerial datasets show that our works greatly improve the accuracy over the baseline and other models. Zhechun Wan, Qian Zhang 0003, Guixu Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Homogeneous Aggregation Convolution for Building Extraction From Remote Sensing ImagesabstractAn increasing number of convolutional neural networks are being applied to various fields, and they have achieved excellent performance. However, standard convolution cannot recognize the connection between surrounding features and cannot characterize them efficiently. For example, buildings in urban remote sensing images exhibit geometric changes, such as rotation, scaling, and local changes. The concept of dynamic convolution is proposed to solve the aforementioned problem. Existing dynamic convolution methods enhance an expression by dynamically changing the sampling points or weights of convolution. However, these end-to-end training methods do not consider which sampling points are important. In this letter, we propose a homogeneous aggregation convolution (HAC) that gives more attention to the sampling points that belong to the same class as the target point. A generate probability map module is designed to generate a probability map between target and sampling points and share this probability map across convolution layers to save computational cost. Experimental results demonstrate that the proposed HAC outperforms standard convolution, and the intersection over union and F1 score are higher than the standard convolution by 2.23% and 1.28%, respectively, on the WHU and Austin datasets. Compared with other convolutions, the proposed HAC convolution is the most efficient in building extraction. Rouyu Zhang, Baokai Lin, Qian Zhang 0003, Guixu Zhang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | SPP-CPI: Predicting Compound-Protein Interactions Based On Neural NetworksabstractIdentifying interactions between compound and protein is a substantial part of the drug discovery process. Accurate prediction of interaction relationships can greatly reduce the time of drug development. The uniqueness of our method lies in three aspects:1) it represents a compound with a distance matrix. A distance matrix can capture the structural information, compared with the SMILES string. On the other hand, a distance matrix does not require complex data preprocessing for the molecular structure as the molecular graph representation, and is easier to obtain; 2) it uses SPP(Spatial pyramid pooling)-net to extract compound features, which has been successfully applied in image classification; and 3) it extracts protein features through the natural language processing method (doc2vec) to obtain sequence semantic information. We evaluated our method on three benchmark datasets-human, C.elegans, and DUDE-and the experimental results demonstrate that our proposed model presents competitive performance against state-of-the-art predictors. We also carried out drug-drug interaction (DDI) experiments to verify the strong potential of distance matrix as molecular characteristics. The source code and datasets are available at https://github.com/lxlsu/SPP_CPI. Qian Zhang 0003, Jiongmin Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Collaborative Network for Super-Resolution and Semantic Segmentation of Remote Sensing ImagesabstractIn the past few years, multitask learning (MTL) has been widely used in a single model to solve the problems of multiple businesses. MTL enables each task to achieve high performance and greatly reduces computational resource overhead. In this work, we designed a collaborative network that simultaneously solves the super-resolution semantic segmentation and super-resolution image reconstruction. This algorithm can obtain high-resolution semantic segmentation and super-resolution reconstruction results by taking relatively low-resolution images as input when high-resolution data are inconvenient or computing resources are limited. The framework consists of three parts: the semantic segmentation branch (SSB), the super-resolution branch (SRB), and the structural affinity block (SAB). Specifically, the SSB, SRB, and SAB are responsible for completing super-resolution semantic segmentation, image super-resolution reconstruction, and associated features, respectively. Our proposed method is simple and efficient, and it can replace the different branches with most of the state-of-the-art models. The International Society for Photogrammetry and Remote Sensing (ISPRS) segmentation benchmarks were used to evaluate our models. In particular, super-resolution semantic segmentation on the Potsdam dataset reduced Intersection over Union (IoU) by only 1.8% when the resolution of the input image was reduced by a factor of two. The experimental results showed that our framework can obtain more accurate semantic segmentation and super-resolution reconstruction results than the single model. Qian Zhang 0003, Guang Yang 0068, Guixu Zhang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | A Combined Deep Learning Model for the Scene Classification of High-Resolution Remote Sensing ImageabstractDeep learning now plays an important role in solving complex problems in computer vision fields. The highly challenging high-resolution remote sensing image scene classification problem can also be solved using deep learning methods. The most commonly used method of deep learning is the convolutional neural network model. In this letter, based on deep learning, a combined model named Inception-long short-term memory (LSTM) is proposed. First, we combine the deep learning feature extracted from the pretrained Inception-V3 model with a hand-crafted feature: the GIST feature. The different features are then combined and input into the batch normalization (BN) layer. Second, the BN layer plays the role of the bridge to combine the InceptionV3 model with the LSTM model, which features a softmax classifier. The LSTM model is used to analyze the features and classify the different high-resolution remote sensing scene images. The proposed model, as a whole, can be uniformly trained. Three different datasets-the NWPU-RESISC45 dataset, the UC Merced dataset, and the SIRI-WHU dataset-were used to verify the effectiveness of the proposed model. The results show that the proposed Inception-LSTM model shows an outstanding performance in the scene classification task. Yunya Dong, Qian Zhang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | A Morphological Building Detection Framework for High-Resolution Optical Imagery Over Urban AreasabstractThis letter proposes an efficient framework for building detection from coarse to fine using morphological technique for high-resolution optical satellite imagery over urban areas. First, the preliminary result of building regions is obtained by the recently developed morphological building index (MBI) method, which is able to detect potential building structures. However, the raw results derived from the MBI can be subject to a number of false alarms, which are caused by bright soil, roads, and open areas. In this letter, we propose to use morphological spatial pattern analysis as a postprocessing to further optimize the MBI result and remove the commission errors. The original MBI result is then separated into seven mutually exclusive categories-core, islet, loop, bridge, perforation, edge, and branch-by applying a series of morphological transformations such as erosions, geodesic dilation, reconstruction by dilation, anchored skeletonization, etc. The objects corresponding to the generic categories are then analyzed, and the categories corresponding to building parts are maintained, while the others are abandoned. After this postprocessing, the small noisy patches and narrow roads, which were wrongly extracted by the MBI, can be removed. In addition, the shape of the buildings can also be regularized by removing the branches, and the holes contained in the building objects can be identified and filled. Extensive experiments performed on GeoEye-1 and WorldView-2 images confirm the effectiveness and robustness of the proposed morphological building detection framework. Qian Zhang 0003, Xin Huang 0002, Guixu Zhang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2014 | Adaptive image segmentation by using mean-shift and evolutionary optimisationabstractUndersegmentation or oversegmentation is a challenge faced in image segmentation methods, and it is extreme important to determine the optimal number of regions (clusters) of an image in real‐world applications. In this study, we introduce an adaptive strategy to do so. The basic idea is to firstly oversegment an image by using the Mean‐shift (MS) method, and then segment the obtained oversegmented results by using an evolutionary algorithm. In the second stage, a feature is extracted for each region obtained by the MS method, and a new fitness function is designed to determine the optimal number of clusters. The adaptive approach is applied to a variety of images, and the experimental results show that our method is both efficient and effective for image segmentation. Cong Liu 0011, Aimin Zhou, Qian Zhang 0003, Guixu Zhang |
IET Image Process. | 3 |
| 2013 | An Energy-Driven Total Variation Model for Segmentation and Classification of High Spatial Resolution Remote-Sensing ImageryabstractAn energy-driven total variation (TV) formulation is proposed for the segmentation of high spatial resolution remote-sensing imagery. The TV model is an effective tool for image processing operations such as restoration, enhancement, reconstruction, and diffusion. Due to the relationship between the TV model and the segmentation problem, in this letter, a TV-based approach is investigated for segmentation of high-spatial-resolution remote-sensing imagery. Subsequently, an object-based classification method, i.e., majority voting, is used to classify the segmented results. In experiments, the proposed TV-based method is compared with the widely used fractal net evolution approach and the clustering segmentation methods such as the expectation–maximization and$k$-means. The performances of the segmentation and the classification are evaluated based on both thematic and geometric indices. Qian Zhang 0003, Xin Huang 0002, Liangpei Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |