EDBT 2026 Demo / reviewers in the wild / expert
Yuxi Sun 0002
dblp:254/4385-2
· DBLP profile ↗
11ranked-venue papers
7as first author
11since 2021 · last 2024
0000-0002-3040-5880ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Cross-Modal Hashing With Feature Semi-Interaction and Semantic Ranking for Remote Sensing Ship Image RetrievalabstractCross-modal hashing plays a pivotal role in large-scale remote sensing (RS) ship image retrieval. RS ship images often exhibit similar overall appearance with subtle differences. Existing hashing methods typically employ feature non-interaction strategies to generate common hash codes, which may not effectively capture the correlations between cross-modal ship images to reduce intermodality discrepancies. To address this issue, we propose a novel cross-modal hashing approach based on feature semi-interaction and semantic ranking (FSISR) for RS ship image retrieval. Our FSISR approach not only captures intricate correlations between different ship image modalities, but also enables the construction of hash tables for large-scale retrieval. FSISR comprises a feature semi-interaction module and a semantic ranking objective function. The semi-interaction module utilizes clustering centers from one modality to learn the correlations between two modalities and generate robust shared representations. The objective function optimizes these representations in a common Hamming space, consisting of a shared semantic alignment loss and a margin-free ranking loss. The alignment loss employs a shared semantic layer to preserve label-level similarity, while the ranking loss incorporates hard examples to establish a margin-free loss that captures similarity ranking relationships. We evaluate the performance of our method on benchmark datasets and demonstrate its effectiveness for cross-modal RS ship image retrieval.https://github.com/sunyuxi/FSISR. Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Yifang Ban, Sebastian Hafner, Xutao Li 0003, Chuyao Luo, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Toward a Variation-Aware and Interpretable Model for Radar Image Sequence PredictionabstractRadar image sequence prediction (RISP) aims to predict future radar images based on historical observations. In the past few years, neural network-based methods have shown impressive performance for RISP. However, two limitations stills exist. 1) They fail to exploit variation information when capturing spatial dependencies. 2) They neglect to analyze and interpret the model. In this article, we propose a variation-aware prediction model for the first limitation, and develop a relevance propagation technique for the second one. Specifically, 1) we recustomize the vanilla convolution by introducing a variation-aware item. The new convolution unit yields two advantages when capturing spatial dependencies, i.e., exploiting variation information and offering spatially-varying kernels. As a result, it can learn the diverse and complex radar echo patterns. By equipping the unit into a typical network (PredRNN), we propose a novel prediction model, dubbed as VA-PredRNN. 2) As for analyzing our model, we propagate the output backward layer by layer till the input. Hence, we can reveal the relevance between the output and the intermediate states. To the best of the authors' knowledge, this is the first work to study the interpretability of a multilayer RISP model. We conduct extensive experiments on two datasets, and the results demonstrate the effectiveness of our VA-PredRNN. We also carry out a series of analyses using the proposed relevance propagation technique. According to the results, we discover the importance of different states. Yunming Ye, Bowen Zhang 0005, Huiwei Lin, Yuxi Sun 0002, Xutao Li 0003, Chuyao Luo |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | Knowledge-enhanced Prompt-tuning for Stance DetectionabstractInvestigating public attitudes on social media is important in opinion mining systems. Stance detection aims to analyze the attitude of an opinionated text (e.g., favor, neutral, or against) toward a given target. Existing methods mainly address this problem from the perspective of fine-tuning. Recently, prompt-tuning has achieved success in natural language processing tasks. However, conducting prompt-tuning methods for stance detection in real-world remains a challenge for several reasons: (1) The text form of stance detection is usually short and informal, which makes it difficult to design label words for the verbalizer. (2) The tweet text may not explicitly give the attitude. Instead, users may use various hashtags or background knowledge to express stance-aware perspectives. In this article, we first propose a prompt-tuning-based framework that performs stance detection in a cloze question manner. Specifically, a knowledge-enhanced prompt-tuning framework (KEprompt) method is designed, which consists of an automatic verbalizer (AutoV) and background knowledge injection (BKI). Specifically, in AutoV, we introduce a semantic graph to build a better mapping from the predicted word of the pretrained language model and detection labels. In BKI, we first propose a topic model for learning hashtag representation and introduce ConceptGraph as the supplement of the target. At last, we present a challenging dataset for stance detection, where all stance categories are expressed in an implicit manner. Extensive experiments on a large real-world dataset demonstrate the superiority of KEprompt over state-of-the-art methods. Hu Huang 0009, Bowen Zhang 0005, Xiang-Yang Li 0001, Baoquan Zhang, Yuxi Sun 0002, Chuyao Luo, Cheng Peng 0003 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2023 | Cross-View Object Geo-Localization in a Local Region With Satellite ImageryabstractCross-view geo-localization is a critical task in various applications, such as smart city management and disaster monitoring. Current methods typically divide a satellite image into patches and use these patches to identify the geographic location of a query image. However, these methods can only provide the location of an image rather than the location of a specific object of interest. This makes it difficult to link these methods to GeoDatabases to obtain detailed information about a target object, such as its name and construction time. To overcome this limitation, we propose a novel problem of cross-view object geo-localization in a local region with high-resolution satellite images. This problem includes two main challenges: accurately identifying the location of an object and distinguishing the target object from others in satellite images. To address these challenges, we present a new Detection-based Geo-localization method called DetGeo, which consists of an object detection-based framework with a two-branch encoder and a query-aware cross-view fusion module. DetGeo uses cross-view images as input to the detector to provide object-level geo-localization. The fusion module employs cross-view spatial attention to focus on relevant areas of target objects during cross-view feature fusion. To evaluate our method, we constructed a new Cross-View Object Geo-Localization dataset called CVOGL, which comprises ground-view or drone-view images as query images and satellite-view images as geo-tagged reference images. Comprehensive experiments are conducted to demonstrate the effectiveness of our method on CVOGL. https://github.com/sunyuxi/DetGeo. Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Shanshan Feng 0001, Xutao Li 0003, Chuyao Luo, Puzhao Zhang, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Consistency Center-Based Deep Cross-Modal Hashing for Multisource Remote Sensing Image RetrievalabstractCross-modal hashing aims to retrieve similar images from large-scale Earth Observation (EO) data archives, which typically contain multiple satellite sources of remote sensing (RS) images. However, existing cross-modal hashing methods primarily focus on dual-source RS images and often face two main limitations when retrieving multi-source RS images. Firstly, these methods exhibit significant redundancy as they require handling all possible dual-source combinations in multi-source RS images. Secondly, they often rely on pairwise or triplet image sources to construct objective functions, which are not significantly effective in reducing the discrepancies among multiple RS image sources. To address these limitations, we propose a novel Consistency Center-based deep cross-modal Hashing method called C2Hash for multi-source RS image retrieval. Our C2Hash employs a multi-branch hashing network to directly encode multi-source RS images into unified hash codes, thereby offering higher processing efficiency. Furthermore, C2Hash introduces consistency centers to construct a novel objective function. The consistency center represents the shared semantic features among similar multi-source RS images and is generated by a label hashing network. The objective function encourages similar multi-source RS images to approach the same consistency center to align all image sources in a unified Hamming space. Our method can effectively reduce the discrepancies across multiple image sources and generate unified hash codes. To evaluate its effectiveness, we construct a new Multi-Source RS Image dataset called MSRSI, comprising five different types of image sources. We conduct comprehensive experiments to demonstrate the superior performance of our method on the MSRSI dataset. https://github.com/sunyuxi/C2Hash. Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Xutao Li 0003, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Visual Grounding in Remote Sensing ImagesabstractGround object retrieval from a large-scale remote sensing image is very important for lots of applications. We present a novel problem of visual grounding in remote sensing images. Visual grounding aims to locate the particular objects (in the form of the bounding box or segmentation mask) in an image by a natural language expression. The task already exists in the computer vision community. However, existing benchmark datasets and methods mainly focus on natural images rather than remote sensing images. Compared with natural images, remote sensing images contain large-scale scenes and the geographical spatial information of ground objects (e.g., longitude, latitude). The existing method cannot deal with these challenges. In this paper, we collect a new visual grounding dataset, called RSVG, and design a new method, namely GeoVG. In particular, the proposed method consists of a language encoder, image encoder, and fusion module. The language encoder is used to learn numerical geospatial relations and represent a complex expression as a geospatial relation graph. The image encoder is applied to learn large-scale remote sensing scenes with adaptive region attention. The fusion module is used to fuse the text and image feature for visual grounding. We evaluate the proposed method by comparing it to the state-of-the-art methods on RSVG. Experiments show that our method outperforms the previous methods on the proposed datasets. https://sunyuxi.github.io/publication/GeoVG Yuxi Sun 0002, Shanshan Feng 0001, Xutao Li 0003, Yunming Ye, Jian Kang 0005 |
ACM Multimedia | 1 |
| 2022 | Unsupervised deep hashing through learning soft pseudo label for remote sensing image retrieval
Yuxi Sun 0002, Yunming Ye, Xutao Li 0003, Shanshan Feng 0001, Bowen Zhang 0005, Jian Kang 0005, Kuai Dai |
Knowl. Based Syst. | 1 |
| 2022 | PredRANN: The spatiotemporal attention Convolution Recurrent Neural Network for precipitation nowcasting
Chuyao Luo, Xinyue Zhao, Yuxi Sun 0002, Xutao Li 0003, Yunming Ye |
Knowl. Based Syst. | 3 |
| 2022 | Better Visual Interpretation for Remote Sensing Scene ClassificationabstractDeep learning-based methods have been widely applied in remote sensing scene classification tasks. Recently, researchers focus more on clarifying the basis of a decision. For example, class activation mapping (CAM) can provide us the evidence by highlighting the related area in an image. However, the interpretability of remote sensing scene classification is more challenging than natural images, since remote sensing images usually contain more complicated objects. As a result, the CAM visual interpretation with traditional convolutional neural networks cannot accurately locate all target objects, which leads to some important objects are ignored. In this letter, we propose a novel model, named encoder-classifier-reconstruction CAM (ECR-CAM) neural network, to provide a more precise visual explanation. Specifically, ECR-CAM consists of four modules: an encoder module, a classifier module, a reconstruction module, and a CAM module. Encoder module is utilized to extract image features, and classifier module accounts for generating predictions. The reconstruction module is the key to locate more target objects. It employs the extracted features to reconstruct the input images, which is a pixel-level process. The reconstruction process allows the features to retain important information about all objects, which cannot be achieved by the classification task alone. Finally, the CAM module can show more target objects with more informative features. Experimental results show that our model not only improves the classification performance but also can locate the target objects more accurately. Yuxi Sun 0002, Shanshan Feng 0001, Yunming Ye, Xutao Li 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Multisensor Fusion and Explicit Semantic Preserving-Based Deep Hashing for Cross-Modal Remote Sensing Image RetrievalabstractCross-modal hashing is an important tool for retrieving useful information from very-high-resolution (VHR) optical images and synthetic aperture radar (SAR) images. Dealing with the intermodal discrepancies, including both spatial–spectral and visual semantic aspects, between VHR and SAR images is extremely vital to generate high-quality common hash codes in the Hamming space. However, existing cross-modal hashing methods ignore the spatial–spectral discrepancy when representing VHR and SAR images. Moreover, existing methods employ derived supervised signals, such as pairwise training images, to implicitly guide hashing learning, which fails to effectively deal with the visual semantic discrepancy, i.e., cannot adequately preserve the intraclass similarity and interclass discrimination between VHR and SAR images. To address these drawbacks, this article proposes a multisensor fusion and explicit semantic preserving-based deep Hashing method, termed as MsEspH, which can effectively deal with the discrepancies. Specifically, we design a novel cross-modal hashing network to eliminate the spatial–spectral discrepancies by fusing extra multispectral images (MSIs), which are generated in real time by a generative adversarial network. Then, we propose an explicit semantic preserving-based objective function by analyzing the connection between classification and hash learning. The objective function can preserve the intraclass similarity and interclass discrimination with class labels directly. Moreover, we theoretically verify that hash learning and classification can be unified into a learning framework under certain conditions. To evaluate our method, we construct and release a large-scale VHR-SAR image dataset. Extensive experiments on the dataset demonstrate that our method outperforms various state-of-the-art cross-modal hashing methods. Yuxi Sun 0002, Shanshan Feng 0001, Yunming Ye, Xutao Li 0003, Jian Kang 0005, Zhichao Huang 0001, Chuyao Luo |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Multisource Data Reconstruction-Based Deep Unsupervised Hashing for Unisource Remote Sensing Image RetrievalabstractUnsupervised hashing for remote sensing (RS) image retrieval first extracts image features and then use these features to construct supervised information (e.g., pseudo-labels) to train hashing networks. Existing methods usually regard RS images as natural images to extract unisource features. However, these features only contain partial information about ground objects and cannot produce reliable pseudo-labels. In addition, existing methods only generate a pseudo single-label to annotate each RS image, which cannot accurately represent multiple scenes in a RS image. To address these drawbacks, this paper proposes a new Multisource data reconstruction-based deep unsupervised Hashing method, called MrHash, which explores the characteristics of RS images to construct reliable pseudo-labels. In particular, we first use geographic coordinates to obtain different satellite images and develop a novel autoencoder network to extract multisource features from these images. Then pseudo multi-labels are designed to deal with the coexistence of multiple scenes in a single image. These labels are generated by a custom probability function with extracted multisource features. Finally, we propose a novel multi-semantic hash loss by using the Kull-back–Leibler (KL) divergence to preserve the semantic similarity of these pseudo multi-labels in Hamming space. Our newly developed MrHash only uses multisource images to construct supervised information, and hash code generation still relies on a unisource input image. Experiments on benchmark datasets clearly show the superiority of the proposed method over state-of-the-art baselines. https://github.com/sunyuxi/MrHash. Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Yifang Ban, Xutao Li 0003, Bowen Zhang 0005, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |