VLDB 2026 Research / reviewers in the wild / expert
Wenkai Zhang 0002
dblp:153/2367-2
· DBLP profile ↗
29ranked-venue papers
1as first author
20since 2021 · last 2024
0000-0002-8903-2708ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 12 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Vigen500k: A Sustainable-Expansion Image-Text Aligned Dataset For Remote SensingabstractRecently, large-scale Vision-Language Models (VLMs) have gained widely attention in the field of remote sensing. However, the researching on VLM requires a substantial amount of data, which is relatively scarce in the remote sensing domain. To overcome this limitation, in this paper, we present ViGen500K, a larger and more challenging image-text dataset. Nearly 500,000 images have been collected, accompanied by over 1 million annotations to adapt to the diverse requirements of various image-text tasks in remote sensing. Besides, a promising, efficient, low-cost, and highly automated data annotation method is proposed to make our dataset could be easily extensive by keeping adding extra unlabeled remote sensing images. Theoretically, ViGen500K is an infinitely large dataset. From a quantitative point of view, compared with traditional image caption datasets, ViGen500K not only has more images but also covers more object categories, which enables the model trained on our dataset could have a wider range of target-text alignment capabilities. Several experiments have been conducted to provide benchmarks for our dataset. Boyuan Tong, Runyan Du, Wenkai Zhang 0002, Shuoke Li, Zhi Guo, Xian Sun 0001, Guangluan Xu |
IGARSS | 3 |
| 2024 | Spatial guided image captioning: Guiding attention with object's spatial interactionabstractAbstract Nowadays relational position embedding is widely used in many large multi‐modal models. It begins with relational captioning (a branch of image captioning) and contains two procedures: geometric modelling and prior attention. However, there are some problems that remain unsolved in the conventional procedures. This paper reviews the shortcomings of geometric modelling and prior attention. Then, a new framework called relational guided transformer (RGT) is proposed to verify the authors' conclusion from the origin of relational position embedding—relational captioning. Specifically, RGT has two simple but effective improvements in geometric modelling and prior attention: (1) A machine‐learned geometric modelling strategy called multi‐task geometric modelling (MTG) is used under multi‐task learning, replacing the original hand‐made geometric feature. (2) The effectiveness of multiple kinds of prior attention is discussed and preserved in a better form, which is called spatial guided attention (SGA) to integrate the geometric prior knowledge. Extensive experiments on MSCOCO and Flickr30k have been performed to investigate the effectiveness of each module and prove our argument. The superiority of the model comparing to the authors' baseline has also been proven on the offline evaluation with the “Karpathy” test split of both datasets. Runyan Du, Wenkai Zhang 0002, Shuoke Li, Zhi Guo |
IET Image Process. | 2 |
| 2024 | Injecting Linguistic Into Visual Backbone: Query-Aware Multimodal Fusion Network for Remote Sensing Visual GroundingabstractThe remote sensing visual grounding (RSVG) task focuses on accurately identifying and localizing specific targets in remote sensing (RS) images using descriptive query expressions. Existing methods independently extract visual and textual features, ignoring early complementary information between image and text. This leads to information loss and misalignment, limiting the model’s ability to distinguish similar targets. To address this challenge, we propose the query-aware multimodal fusion network (QAMFN), which introduces an innovative query-guided visual attention (QGVA) mechanism in the early stages of the visual encoder. This mechanism integrates textual information during the early visual feature extraction process, thereby resolving the issue of missing image-text complementary information. QGVA ensures that the visual backbone accurately focuses on local features highly relevant to the query by injecting textual information into the visual encoding process. Additionally, to enhance the model’s ability to integrate multimodal information and adapt to more complex RS images, we introduce the text-semantic attention-guided masking (TAM) module. TAM aggregates multimodal features processed by the backbones and filters out redundant information, producing high-quality fused features. Experiments demonstrate that our approach sets a new record on the DIOR-RSVG dataset, improving accuracy to 81.67% (an absolute increase of 4.98%). Wenkai Zhang 0002, Hanbo Bi, Shuoke Li, Haichen Yu, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | RingMo-SAM: A Foundation Model for Segment Anything in Multimodal Remote-Sensing ImagesabstractThe proposal of Segment Anything Model (SAM) has created a new paradigm for deep learning-based semantic segmentation field, and has shown amazing generalization performance. However, we find it may fail or perform poorly on multimodal remote sensing scenarios, especially the Synthetic Aperture Radar (SAR) images. Besides, SAM does not provide category information of objects. In this paper, we propose a foundation model for multimodal remote sensing image segmentation called RingMo-SAM, which can not only segment anything in optical and SAR remote sensing data, but also identify object categories. First, a large-scale dataset containing millions of segmentation instances is constructed by collecting multiple open-source datasets in this field to train the model. Then, by constructing an instance-type and terrain-type category-decoupling mask decoder, the category-wise segmentation of various objects is achieved. In addition, a prompt encoder embedded with the characteristics of multimodal remote sensing data is designed. It not only supports multi-box prompts to improve the segmentation accuracy of multi-objects in complicated remote sensing scenes, but also supports SAR characteristics prompts to improve the segmentation performance on SAR images. Extensive experimental results on several datasets including iSAID, ISPRS Vaihingen, ISPRS Potsdam, AIR-PolSAR-Seg, etc. have demonstrated the effectiveness of our method. Junxi Li, Xuexue Li, Ruixue Zhou, Wenkai Zhang 0002, Yingchao Feng, Wenhui Diao, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Efficient and Controllable Remote Sensing Fake Sample Generation Based on Diffusion Model
Chongyang Hao, Ruixue Zhou, Wenkai Zhang 0002, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Hypersphere-Based Remote Sensing Cross-Modal Text-Image Retrieval via Curriculum LearningabstractRemote sensing cross-modal text-image retrieval (RSCTIR) is a flexible and human-centered approach to retrieving rich information from different modalities, which has attracted plenty of attention in recent years. It remains challenging because the current methods usually ignore the varying difficulty levels of different sample pairs, stemming from the large image distribution difference and the high text similarity in the remote sensing (RS) field. Therefore, in this paper, we propose an innovative hypersphere-based visual semantic alignment (HVSA) network via curriculum learning. Specifically, we first design an adaptive alignment strategy based on curriculum learning, that aligns RS image-text pairs from easy to hard. Sample pairs with different levels of difficulty are treated unequally, and we obtain a better embedding representation when projecting the features onto the unit hypersphere. Then, to measure the robustness of cross-modal feature alignment on the unit hypersphere, we introduce the feature uniformity strategy. It reduces the occurrence of mismatching cases and improves generalization performance. Finally, we design the key-entity attention (KEA) mechanism to alleviate the problem of information imbalance among different modalities. KEA has the ability to extract information about the key entity which is aligned with textual information. Despite its conciseness, our framework achieves state-of-the-art performance on classical datasets of RSCTIR tasks while enjoying faster inference. The summed recall of HVSA on the RISCD and RSITMD is 120.97 and 198.94, 2.50 and 10.49 points ahead of the current best methods, respectively. Extensive experiments demonstrate the competitiveness of our method. The code has been released at https://github.com/ZhangWeihang99/HVSA. Shuoke Li, Wenkai Zhang 0002, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Mimicking the Brain's Cognition of Sarcasm From Multidisciplines for Twitter Sarcasm DetectionabstractSarcasm is a sophisticated construct to express contempt or ridicule. It is well-studied in multiple disciplines (e.g., neuroanatomy and neuropsychology) but is still in its infancy in computational science (e.g., Twitter sarcasm detection). In contrast to previous methods that are usually geared toward a single discipline, we focus on the multidisciplinary cross-innovation, i.e., improving embryonic sarcasm detection in computational science by leveraging the advanced knowledge of sarcasm cognition in neuroanatomy and neuropsychology. In this work, we are oriented toward sarcasm detection in social media and correspondingly propose a multimodal, multi-interactive, and multihierarchical neural network ($M_{3}N_{2} $). We select Twitter, image, text in image, and image caption as the input of$M_{3}N_{2} $since the brain’s perception of sarcasm requires multiple modalities. To reasonably address the multimodalities, we introduce singlewise, pairwise, triplewise, and tetradwise modality interactions incorporating gate mechanism and guide attention (GA) to simulate the interactions and collaborations of involved regions in the brain while perceiving multiple modes. Specifically, we exploit a multihop process for each modality interaction to extract modal information multiple times using GA for obtaining multiperspective information. Also, we adopt a two-hierarchical structure leveraging self-attention accompanied by attention pooling to integrate multimodal semantic information from different levels mimicking the brain’s first- and second-order comprehensions of sarcasm. Experimental results show that$M_{3}N_{2} $achieves competitive performance in sarcasm detection and displays powerful generalization ability in multimodal sentiment analysis and emotion recognition. Fanglong Yao, Xian Sun 0001, Wenkai Zhang 0002, Kun Fu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Without detection: Two-step clustering features with local-global attention for image captioningabstractAbstract The current image captioning methods usually integrate an object detection network to obtain image features at the level of objects and other salient regions. However, the detection network needs to be independently pre‐trained on additional data. Thus, mainly due to the demand for extra training data and computing resources, the detection network's utilization will impose higher training costs on the overall captioning model. In this work, the authors propose a local–global attention model based on two‐step clustering features for image captioning. The two‐step clustering features can be obtained at a relatively low cost and have the presentation ability in objects or other salient image regions. To make the model perceive the image better, the authors introduce a novel local–global attention mechanism. The model will analyse the clustering features from local perspectives to global ones at each time step, making the model better understand the image contents. The authors evaluate the proposed method on the MSCOCO test server, achieving BLEU‐4/METEOR/ROUGE‐L scores of 36.8, 27.4, and 57.2, respectively. With the benefit of reducing training costs, the authors' model also achieves closing results compared with the models using detection features. Wenkai Zhang 0002, Xian Sun 0001 |
IET Comput. Vis. | 2 |
| 2022 | Semantic-meshed and content-guided transformer for image captioningabstractAbstract The transformer architecture has been the dominant framework for today's image captioning tasks because of its superior performance. However, existing methods based on transformer often lack the integrated use of multi‐level semantic information and are weak in maintaining the relevance of captions to the image. In this paper, a semantic‐meshed and content‐guided transformer network is introduced for image captioning to solve these problems. The semantic‐meshed mechanism allows the model to generate words by selecting semantic information of multiple interaction levels adaptively through attention‐based reconstruction. And the content‐guided module guides the words generation by using attribute features that represent the image content, which aims to keep the generated caption consistent with the main content of the image. Experiments on dataset on the MSCOCO captioning dataset are conducted to validate the authors’ model and achieve superior results compared to other state‐of‐the‐art method approaches. Wenkai Zhang 0002, Xian Sun 0001 |
IET Comput. Vis. | 2 |
| 2022 | Associatively Segmenting Semantics and Estimating Height From Monocular Remote-Sensing ImageryabstractNumerous deep-learning methods have been successfully applied to semantic segmentation and height estimation of remote-sensing imagery. It has also been proved that such framework can be reusable for multiple tasks to reduce computational resource overhead. However, there are still some technical limitations due to the semantic inconsistency between 3-D and 2-D features and strong interference of different objects with similar spectral-spatial properties. Previous works have sought to address these issues through hard parameter sharing or soft parameter sharing schemes. But due to unintentional integration, the specific information transmitted between multiple tasks is not clear or in a lot of redundancy. Furthermore, tuning the weights by hand between classification and regression loss function is challenging. In this paper, a novel multi-task learning method, termed ASSEH, is proposed to associatively segment semantics and estimate height from monocular remote-sensing imagery. First, considering semantic inconsistency across tasks, we design a task-specific distillation (TSD) module containing a set of task-specific gating units for each task at the cost of fewer parameters. The module allows for task-specific features to be tailored from backbone, whilst allowing for task-shared features to be transmitted. Second, we leverage the proposed cross-task propagation (CTP) module to construct and diffuse the local pattern graphlets at the common positions across tasks. Such a high-order recursive method can bridge two tasks explicitly to effectively settle semantic ambiguities caused by similar spectral characteristics with less computational burden and memory requirements. Third, a dynamic weighted geometric mean (DWGeoMean) strategy is introduced to dynamically learn the weights of each task and be more robust to the magnitude of the loss function. Finally, the results on ISPRS Vaihingen and Urban Semantic 3D data set well demonstrate that our ASSEH achieves the state-of-the-art performance. Wenjie Liu 0016, Xian Sun 0001, Wenkai Zhang 0002, Zhi Guo, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image RetrievalabstractRemote sensing (RS) cross-modal text–image retrieval has attracted extensive attention for its advantages of flexible input and efficient query. However, traditional methods ignore the characteristics of multiscale and redundant targets in RS image, leading to the degradation of retrieval accuracy. To cope with the problem of multiscale scarcity and target redundancy in RS multimodal retrieval task, we come up with a novel asymmetric multimodal feature matching network (AMFMN). Our model adapts to multiscale feature inputs, favors multisource retrieval methods, and can dynamically filter redundant features. AMFMN employs the multiscale visual self-attention (MVSA) module to extract the salient features of RS image and utilizes visual features to guide the text representation. Furthermore, to alleviate the positive samples ambiguity caused by the strong intraclass similarity in RS image, we propose a triplet loss function with dynamic variable margin based on prior similarity of sample pairs. Finally, unlike the traditional RS image-text dataset with coarse text and higher intraclass similarity, we construct a fine-grained and more challenging Remote sensing Image-Text Match dataset (RSITMD), which supports RS image retrieval through keywords and sentence separately and jointly. Experiments on four RS text–image datasets demonstrate that the proposed model can achieve state-of-the-art performance in cross-modal RS text–image retrieval task. Wenkai Zhang 0002, Kun Fu 0001, Chubo Deng, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Learning to Evaluate Performance of Multimodal Semantic LocalizationabstractSemantic localization (SeLo) refers to the task of obtaining the most relevant locations in large-scale remote sensing (RS) images using semantic information such as text. As an emerging task based on cross-modal retrieval, SeLo achieves semantic-level retrieval with only caption-level annotation, which demonstrates its great potential in unifying downstream tasks. Although SeLo has been carried out successively, but there is currently no work has systematically explores and analyzes this urgent direction. In this paper, we thoroughly study this field and provide a complete benchmark in terms of metrics and testdata to advance the SeLo task. Firstly, based on the characteristics of this task, we propose multiple discriminative evaluation metrics to quantify the performance of the SeLo task. The devised significant area proportion, attention shift distance, and discrete attention distance are utilized to evaluate the generated SeLo map from pixel-level and region-level. Next, to provide standard evaluation data for the SeLo task, we contribute a diverse, multi-semantic, multi-objective Semantic Localization Testset (AIR-SLT). AIR-SLT consists of 22 large-scale RS images and 59 test cases with different semantics, which aims to provide a comprehensive evaluations for retrieval models. Finally, we analyze the SeLo performance of RS cross-modal retrieval models in detail, explore the impact of different variables on this task, and provide a complete benchmark for the SeLo task. We have also established a new paradigm for RS referring expression comprehension, and demonstrated the great advantage of SeLo in semantics through combining it with tasks such as detection and road extraction. The proposed evaluation metrics, semantic localization testsets, and corresponding scripts have been open to access at https://github.com/xiaoyuan1996/SemanticLocalizationMetrics. Wenkai Zhang 0002, Zhaoying Pan, Yongqiang Mao, Shuoke Li, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Lightweight Multi-Scale Crossmodal Text-Image Retrieval Method in Remote SensingabstractRemote sensing (RS) crossmodal text-image retrieval has become a research hotspot in recent years for its application in semantic localization. However, since multiple inferences on slices are demanded in semantic localization, designing a crossmodal retrieval model with less computation but well performance becomes an emergent and challenging task. In this article, considering the characteristics of multi-scale and target redundancy in RS, a concise but effective crossmodal retrieval model (LW-MCR) is designed. The proposed model incorporates multi-scale information and dynamically filters out redundant features when encoding RS image, while text features are obtained via lightweight group convolution. To improve the retrieval performance of LW-MCR, we come up with a novel hidden supervised optimization method based on knowledge distillation. This method enables the proposed model to acquire dark knowledge of the multi-level layers and representation layers in the teacher network, which significantly improves the accuracy of our lightweight model. Finally, on the basis of contrast learning, we present a method employing unlabeled data to boost the performance of RS retrieval model further. The experiment results on four RS image-text datasets demonstrate the efficiency of LW-MCR in RS crossmodal retrieval (RSCR) tasks. We have released some codes of the semantic localization and made it open to access athttps://github.com/xiaoyuan1996/retrievalSystem. Wenkai Zhang 0002, Xuee Rong, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local InformationabstractCross-modal remote sensing text-image retrieval (RSCTIR) has recently become an urgent research hotspot due to its ability of enabling fast and flexible information extraction on remote sensing (RS) images. However, current RSCTIR methods mainly focus on global features of RS images, which leads to the neglect of local features that reflect target relationships and saliency. In this article, we first propose a novel RSCTIR framework based on global and local information (GaLR), and design a multi-level information dynamic fusion (MIDF) module to efficaciously integrate features of different levels. MIDF leverages local information to correct global information, utilizes global information to supplement local information, and uses the dynamic addition of the two to generate prominent visual representation. To alleviate the pressure of the redundant targets on the graph convolution network (GCN) and to improve the model’s attention on salient instances during modeling local features, the denoised representation matrix and the enhanced adjacency matrix (DREA) are devised to assist GCN in producing superior local representations. DREA not only filters out redundant features with high similarity, but also obtains more powerful local features by enhancing the features of prominent objects. Finally, to make full use of the information in the similarity matrix during inference, we come up with a plug-and-play multivariate rerank (MR) algorithm. The algorithm utilizes the$k$nearest neighbors of the retrieval results to perform a reverse search, and improves the performance by combining multiple components of bidirectional retrieval. Extensive experiments on public datasets strongly demonstrate the state-of-the-art performance of GaLR methods on the RSCTIR task. The code of GaLR method, MR algorithm, and corresponding files have been made available at:https://github.com/xiaoyuan1996/GaLR. Wenkai Zhang 0002, Changyuan Tian 0001, Xuee Rong, Zhengyuan Zhang 0003, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Global Visual Feature and Linguistic State Guided Attention for Remote Sensing Image CaptioningabstractThe encoder–decoder framework is prevalent in existing remote-sensing image captioning (RSIC) models. The appearance of attention mechanisms brings significant results. However, current attention-based caption models only build up the relationships between the local features without introducing the global visual feature and removing redundant feature components. It will cause caption models to generate descriptive sentences that are weakly related to the scene of images. To solve the problems, this article proposed a global visual feature-guided attention (GVFGA) mechanism. First, GVFGA introduces the global visual feature and fuses them with local visual features to build up their relationships between them. Second, an attention gate utilizing the global visual feature is proposed in GVFGA to filter out redundant feature components in the fused image features and provide more salient image features. In addition, to relieve the hidden state’s burden, a linguistic state (LS) is proposed to specifically provide textual features, making the hidden state only guiding visual–textual attention process. What’s more, to further refine the fusion of visual features and textual features, a LS-Guided Attention (LSGA) mechanism is proposed. It can also filter out the irrelevant information in the fused visual–textual feature with the help of an attention gate. The experimental results show that this proposed image captioning model can achieve better results on three RSIC datasets, UCM-Captions, Sydney-Captions, and RSICD datasets. Zhengyuan Zhang 0003, Wenkai Zhang 0002, Menglong Yan, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Weakly Supervised Semantic Segmentation in Aerial Imagery via Explicit Pixel-Level ConstraintsabstractIn recent years, image-level weakly supervised semantic segmentation (WSSS) has developed rapidly in natural scenes due to the easy availability of classification tags. However, limited to complex backgrounds, multi-category scenes, and dense small targets in remote sensing (RS) images, relatively little research has been conducted in this field. To alleviate the impact of the above problems in RS scenes, a self-supervised Siamese network based on an explicit pixel-level constraints framework is proposed, which greatly improves the quality of class activation maps and the positioning accuracy in multi-category RS scenes. Specifically, there are three novel devices in this paper to promote performance to a new level: (a) A pixel-soft classification loss is proposed, which realizes explicit constraints on pixels during the image-level training; (b) A pixel global awareness module, which captures high-level semantic context and low-level pixel spatial information, is constructed to improve the consistency and accuracy of RS object segmentation; (c) A dynamic multi-scale fusion module with a gating mechanism is devised, which enhances feature representation and improves the positioning accuracy of RS objects, particularly on small and dense objects. Experiments on two RS challenge datasets demonstrate that these proposed modules achieve new state-of-the-art results by only using image-level labels, which improve mIoU to 36.79% on iSAID and 45.43% on ISPRS in the WSSS task. To the best of our knowledge, this is the first work to perform image-level WSSS on multi-class RS scenes. Ruixue Zhou, Wenkai Zhang 0002, Xuee Rong, Wenjie Liu 0016, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | An enhanced dynamic interaction network for claim verification
Peiguang Li, Xian Sun 0001, Wenkai Zhang 0002, Guangluan Xu |
Neurocomputing | 4 |
| 2021 | D-MmT: A concise decoder-only multi-modal transformer for abstractive summarization in videos
Nayu Liu, Xian Sun 0001, Wenkai Zhang 0002, Guangluan Xu |
Neurocomputing | 4 |
| 2021 | Commonalities-, specificities-, and dependencies-enhanced multi-task learning network for judicial decision prediction
Fanglong Yao, Xian Sun 0001, Wenkai Zhang 0002, Kun Fu 0001 |
Neurocomputing | 4 |
| 2021 | Reasoning like Humans: On Dynamic Attention Prior in Image Captioning
Yong Wang 0051, Xian Sun 0001, Wenkai Zhang 0002 |
Knowl. Based Syst. | 4 |
| 2020 | Multistage Fusion with Forget Gate for Multimodal Summarization in Open-Domain VideosabstractMultimodal summarization for open-domain videos is an emerging task, aiming to generate a summary from multisource information (video, audio, transcript).Despite the success of recent multiencoder-decoder frameworks on this task, existing methods lack finegrained multimodality interactions of multisource inputs.Besides, unlike other multimodal tasks, this task has longer multimodal sequences with more redundancy and noise.To address these two issues, we propose a multistage fusion network with the fusion forget gate module, which builds upon this approach by modeling fine-grained interactions between the multisource modalities through a multistep fusion schema and controlling the flow of redundant information between multimodal long sequences via a forgetting module.Experimental results on the How2 dataset show that our proposed model achieves a new state-of-the-art performance.Comprehensive analysis empirically verifies the effectiveness of our fusion schema and forgetting module on multiple encoder-decoder architectures.Specially, when using high noise ASR transcripts (W ER>30%), our model still achieves performance close to the ground-truth transcript model, which reduces manual annotation cost. Nayu Liu, Xian Sun 0001, Wenkai Zhang 0002, Guangluan Xu |
EMNLP (1) | 4 |
| 2020 | Improving Intra- and Inter-Modality Visual Relation for Image CaptioningabstractIt is widely shared that capturing relationships among multi-modality features would be helpful for representing and ultimately describing an image. In this paper, we present a novel Intra- and Inter-modality visual Relation Transformer to improve connections among visual features, termed I2RT. Firstly, we propose Relation Enhanced Transformer Block (RETB) for image feature learning, which strengthens intra-modality visual relations among objects. Moreover, to bridge the gap between inter-modality feature representations, we align them explicitly via Visual Guided Alignment (VGA) module. Finally, an end-to-end formulation is adopted to train the whole model jointly. Experiments on the MS-COCO dataset show the effectiveness of our model, leading to improvements on all commonly used metrics on the "Karpathy" test split. Extensive ablation experiments are conducted for the comprehensive analysis of the proposed method. Yong Wang 0051, Wenkai Zhang 0002, Qing Liu 0021, Zhengyuan Zhang 0003, Xian Sun 0001 |
ACM Multimedia | 2 |
| 2020 | SA-NLI: A Supervised Attention based framework for Natural Language Inference
Peiguang Li, Wenkai Zhang 0002, Guangluan Xu, Xian Sun 0001 |
Neurocomputing | 3 |
| 2020 | Gated hierarchical multi-task learning network for judicial decision prediction
Fanglong Yao, Xian Sun 0001, Wenkai Zhang 0002, Kun Fu 0001 |
Neurocomputing | 5 |
| 2019 | Geometrical Model for the Layover of Gable-Roofed Buildings and its Application in Building ReconstructionabstractBuilding reconstruction from SAR images is a hot topic in recent years. Currently, related methods mainly deal with on flat-roofed buildings. In this paper, we extend the research scope to gable-roofed buildings, and try to present a parameterized geometrical model for the layover of gable-roofed buildings. Based on this model, a top-down building reconstruction technique based on MCMC method is proposed. Through representing the layover with parameterized geometrical models, building reconstruction is converted into an optimization problem under the Bayesian scheme. In order to obtain global optima, simulated annealing algorithm with MCMC is used in the optimization stage. Two groups of transmission kernels which are responsible for model updates are designed according to the model. Experiments show that the layover model is accurate and the reconstruction method is effective. Yue Zhang 0016, Zhirui Wang 0003, Liangjin Zhao, Wenkai Zhang 0002, Menglong Yan, Xian Sun 0001 |
IGARSS | 4 |
| 2019 | Effective Fusion of Multi-Modal Data with Group Convolutions for Semantic Segmentation of Aerial ImageryabstractIn this paper, we achieve a semantic segmentation of aerial imagery based on the fusion of multi-modal data in an effective way. The multi-modal data contains a true orthophoto and the corresponding normalized Digital Surface Model (nDSM), which are stacked together before they are fed into a Convolutional Neural Network (CNN). Though the two modalities are fused at the early stage, their features are learned independently with group convolutions firstly and then the learned features of different modalities are fused at multiple scales with standard convolutions. Therefore, the multi-scale fusion of multi-modal features is completed in a single-branch convolutional network. In this way, the computational cost is reduced while the experimental results reveal that we can still get promising results. Kaiqiang Chen, Kun Fu 0001, Menglong Yan, Wenkai Zhang 0002, Yue Zhang 0016, Xian Sun 0001 |
IGARSS | 5 |
| 2019 | Aerial Image and Map Synthesis Using Generative Adversarial NetworksabstractAccurate automatic conversion between aerial images and maps is a valuable and challenging task in computer vision and computer graphics. Deep convolutional neural networks (CNN) have achieved promising results on this task but the results accuracy is not ideal. In this paper, we propose a solution to improve the precision and quality of the transforming results. The core learning method is based on generative adversarial networks (GANs). A novel generator and a multi-scale discriminator are introduced in our network. The generator operates at the progressive method to gurantee the spatial consistency between the inputs and outputs, and our multi-scale discriminator focuses on increasing the capacity of the network and guides the generator to generate better results. In particular, our architecture can also be used as a general neural network for style translation. Analytic experiments on the aerial-to-map dataset show that our network outperforms the existing method, advancing both accuracy and visual appearance. Yue Zhang 0016, Wenkai Zhang 0002, Siyue Wang, Yaoling Wang, Lei Wang 0077 |
IGARSS | 3 |
| 2019 | Effective Classification of Local Climate Zones Based on Multi-Source Remote Sensing DataabstractThe local climate zone (LCZ) classification divides the urban areas into 17 categories, which are composed of 10 manmade structures and 7 natural landscapes. Though originally designed for temperature study, LCZ classification can be used for studies on economy and population. In this paper, we achieve a LCZ classification with convolutional neural networks based on the multi-source remote sensing data, including the polarimetric synthetic aperture radar (PolSAR) data and the corresponding multi-spectral imagery (MSI). Through experiments we attempt to reveal the contributions of the SAR data and the MSI to the classification performance. Furthermore, we emphasize the crucial importance of the preprocessing on the training data to derive a balanced dataset. We are ranked second in the Tianchi competition rankings when we submit our results. Yingchao Feng, Wenkai Zhang 0002, Yue Zhang 0016, Siyue Wang, Kun Fu 0001, Kaiqiang Chen |
IGARSS | 3 |
| 2019 | Joint optimisation convex-negative matrix factorisation for multi-modal image collection summarisation based on images and tagsabstractImage collection summarisation aims to represent a large‐scale multi‐modal collection with a small subset of images and tags, helping navigate a large image dataset. Most extant methods leverage the contributions of text‐to‐visual summaries, ignoring the visual contribution to the textual topic. When the tags are weakly labelled, the textual topic cannot accurately reflect the visual summary. To solve this, the authors propose a novel model, joint optimisation of convex non‐negative matrix factorisation, which incorporates images and tags in a beneficial way. The objective function contains visual and textual error functions, sharing the same indicator matrix, connecting different modal relations. Then, they propose an iterative algorithm to optimise the proposed model. Finally, they explore the effects of different visual feature representations (e.g. bag‐of‐words and deep learning) on multi‐modal collection summary. Our proposed method is then compared with state‐of‐the‐art algorithms using two multi‐modal datasets (i.e. MIRFlickr and NUS‐WIDE‐SCENE). Experimental results demonstrate the effectiveness of their proposed approach. Wenkai Zhang 0002, Kun Fu 0001, Xian Sun 0001, Yuhang Zhang 0006, Hao Sun 0009 |
IET Comput. Vis. | 1 |