VLDB 2026 Research / reviewers in the wild / expert
Yu Gu 0015
dblp:15/4208-15
· DBLP profile ↗
21ranked-venue papers
2as first author
17since 2021 · last 2025
0000-0003-3634-2275ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 10 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TDMER: A Task-Driven Method for Multimodal Emotion RecognitionabstractIn multimodal emotion recognition, disentangled representation learning method effectively address the inherent heterogeneity among modalities. To facilitate the flexible integration of enhanced disentangled features into multimodal emotional features, we propose a task-driven multimodal emotion recognition method TDMER. Its Cross-Modal Learning module promotes adaptive cross-modal learning of features disentangled into modality-invariant and modality-specific subspaces, based on their contributions to emotional classification probabilities. The Task-Contribution Fusion mechanism then assigns controllable weights to the enhanced features according to their task objectives, generating multimodal fusion features that improve the emotion classifier’s discriminative ability. The proposed TDMER approach has been evaluated on two widely-used multimodal emotion recognition benchmarks and demonstrated significant performance improvements compared with other state-of the-art methods. Yu Gu 0015, Hai-Xiang Lin, Linsong Liu |
ICASSP | 2 |
| 2024 | Fourier Domain Adaptive Multi-Modal Remote Sensing Image Template Matching Based on Siamese NetworkabstractMulti-modal remote sensing image template matching is a meaningful and crucial topic in remote sensing image processing. However, due to different imaging mechanisms, there are significant nonlinear radiometric variations among multi-modal remote sensing images, increasing the matching challenge and leading to poor matching performances. To tackle this issue, this paper proposes a Fourier Domain Adaptive Network (FDANet) for multi-modal remote sensing image matching. Firstly, FDANet randomly swaps the low-frequency spectrum information between multi-modal images through the Fourier transform to reduce differences among multi-modal images, enhancing network adaptability to different image modalities and improving the multi-modal image matching performance. Secondly, FDANet extracts domain-invariant features from the transformed images through a deep Siamese network. After that, FDANet performs template matching and achieves high-precision multi-modal remote sensing image matching. In addition, we adopt the contrastive learning loss to optimize the FDANet. Extensive experiments on multi-modal remote sensing image matching demonstrate the effectiveness and advantages of the proposed FDANet. Chonghua Lv, Dou Quan, Shuang Wang 0001, Xiangming Jiang, Yu Gu 0015, Licheng Jiao |
IGARSS | 7 |
| 2024 | Continual learning for cross-modal image-text retrieval based on domain-selective attention
Rui Yang 0038, Shuang Wang 0001, Yu Gu 0015, Jihui Wang, Yingzhi Sun, Yu Liao, Licheng Jiao |
Pattern Recognit. | 3 |
| 2023 | Multi-Source Fusion Network for Remote Sensing Image Segmentation with Hierarchical TransformerabstractRecently, due to the limitations of single sensor, it is hard to improve the performance of land cover classification. The traditional image segmentation methods can not process the optical remote sensing images effectively, especially when optical sensor is affected by complex weather conditions. However, as an active radar, synthetic aperture radar(SAR) has the advantage of not being restricted by weather conditions with the the penetrability of electromagnetic radiation. So multi-sensor data fusion provides a great potential for land cover classification. In this paper, a new fusion network called SegFusion is proposed to improve the performance of land cover classification. There are two main components in SegFusion which are hierarchical Transformer encoder and Swin-Fusion(SW-Fusion) module. First, a hierarchical Transformer encoder is used to extract multilevel feature of optical and SAR images. By integrating features from different layers, we can obtain powerful representation that combines both low-resolution fine-grained features and high-resolution coarse-grained features. Second, SW-Fusion module is used to fuse the features of optical and SAR data. In SW-Fusion, we use modified Swin Transformer [1] block with multi-head cross-attention mechanism to exchange information between features from different sources. Bo Liu 0009, Bo Ren 0001, Biao Hou, Yu Gu 0015 |
IGARSS | 4 |
| 2023 | Domain Distribution Alignment for Boosting Multi-Modal Remote Sensing Image MatchingabstractMulti-modal images can obtain complementary and rich information images, which are more widely used in various applications. However, due to the different imaging mechanisms of different sensors, there are significant domain distribution differences between multi-modal images. In multi-modal image matching, existing deep learning methods should deal with the image content difference caused by rotation transformation and the domain distribution difference caused by different sensors, which are very difficult for the deep network. To address this issue, we propose to combine an instance comparison and a batch comparison to deal with image content differences and domain distribution differences, respectively. We design a new domain distribution alignment method to explicitly constrain the sample domain distribution of the multi-modal images are consistent through the domain distribution alignment loss. Extensive multi-modal remote sensing image patch matching experiments have shown the effectiveness of the proposed method. Furthermore, the proposed multi-modal domain distribution alignment method has more obvious advantages when there are significant content differences and distribution differences. Dou Quan, Chonghua Lv, Yanhe Guo, Shuang Wang 0001, Yu Gu 0015, Licheng Jiao |
IGARSS | 6 |
| 2023 | Incremental Land Cover Classification via Strategies for Edge Removal and Feature Point AggregationabstractConvolutional neural networks will face the problem of catastrophic forgetting in the process of incremental learning. To solve this problem, we propose an incremental learning method called the strategy of edge removal and feature point aggregation, or ERFPA for short. In the cross-entropy loss, we perform edge detection and removal on the labels generated by the old model predictions, and then fuse them with the new labels. We calculate the mean point of different classes, and make the model learn features better by narrowing the distance with similar pixels. As demonstrated by the results of our experiment, on two remote sensing image datasets: CCF and Vaihingen, our method achieves state-of-the-art results. Zhao Wang 0011, Bo Ren 0001, Biao Hou, Yu Gu 0015 |
IGARSS | 4 |
| 2023 | Deep Continuous Matching Network for more Robust Multi-Modal Remote Sensing Image Patch MatchingabstractDue to the powerful feature extraction capabilities of deep neural networks, traditional approaches are gradually replaced by deep learning approaches for image matching tasks. For multi-modal image patch matching, the deep model should mainly learn the modality-invariant features. For multi-modal images with rotation transformation (RT), the deep model should learn the modality-invariant features and rotation-invariant features simultaneously. However, the performance of the latter trained model is degraded for the former task. The main reason is that the modality invariance of the features degenerates. This paper proposes a deep multi-modal remote sensing image matching network (DCMNet) that combines descriptor learning and continuous learning to solve this problem. Firstly, DCMNet is trained for learning modality-invariant features in multi-modal image patch matching. Then, DCMNet is optimized for multi-modal image patch matching with RT. In the later learning process, we reduce the change of important parameters for the modality-invariant features learning. Experiments demonstrate the effectiveness and robustness of DCMNet in alleviating the modal invariance degradation problem of features. Rufan Zhou, Dou Quan, Chonghua Lv, Yanhe Guo, Shuang Wang 0001, Yu Gu 0015, Licheng Jiao |
IGARSS | 6 |
| 2023 | A Novel Coarse-to-Fine Deep Learning Registration Framework for Multimodal Remote Sensing ImagesabstractMulti-modal remote sensing images with large rotation transformation (RT) are challenging to be registered. It needs to deal with the global geometric deformation caused by great RT and significant local appearance differences caused by different imaging mechanisms. Existing deep learning methods mainly use a single deep descriptor learning (DDL) network to extract invariant features for identifying matching samples and discriminative feature descriptors for separating non-matching samples. However, it is difficult to extract local invariant feature descriptors to RT and modality change through a single DDL network. This paper proposes a novel coarse-to-fine deep learning image registration framework for multi-modal remote sensing images based on two task-specific deep models. Specifically, in the coarse registration stage, this paper designs an effective deep ordinal regression (DOR) network for rotation correction, which can reduce the difficulty of multi-modal image registration and boost image registration. The proposed DOR network transforms the rotation correction task into a rotation ordinal regression problem, which can exploit the potential relationship between the rotation ordinals to improve the accuracy of rotation estimation. In the fine registration stage, we adopt the DDL network to deal with the image modality change based on the rotation-corrected images. Extensive experimental results on multi-modal image datasets demonstrate the significant advantages of the proposed coarse-to-fine deep learning registration framework. The DOR network achieves higher rotation correction accuracy, which can significantly improve the multi-modal image registration performances. Dou Quan, Huiyuan Wei, Shuang Wang 0001, Yu Gu 0015, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | A joint framework for mining discriminative and frequent visual representation
Ying Zhou 0021, Xuefeng Liang, Zhihui Liang, Yu Gu 0015, Yifei Yin |
Neurocomputing | 6 |
| 2022 | Multi-Classifier Interactive Learning for Ambiguous Speech Emotion RecognitionabstractIn recent years, speech emotion recognition technology is of great significance in widespread applications such as call centers, social robots and health care. Thus, the speech emotion recognition has been attracted much attention in both industry and academic. Since emotions existing in an entire utterance may have varied probabilities, speech emotion is likely to be ambiguous, which poses great challenges to recognition tasks. However, previous studies commonly assigned a single-label or multi-label to each utterance in certain. Therefore, their algorithms result in low accuracies because of the inappropriate representation. Inspired by the optimally interacting theory, we address the ambiguous speech emotions by proposing a novel multi-classifier interactive learning (MCIL) method. In MCIL, multiple different classifiers first mimic several individuals, who have inconsistent cognitions of ambiguous emotions, and construct new ambiguous labels (the emotion probability distribution). Then, they are retrained with the new labels to interact with their cognitions. This procedure enables each classifier to learn better representations of ambiguous data from others, and further improves the recognition ability. The experiments on three benchmark corpora (MAS, IEMOCAP, and FAU-AIBO) demonstrate that MCIL does not only improve each classifier’s performance, but also raises their recognition consistency from moderate to substantial. Ying Zhou 0021, Xuefeng Liang, Yu Gu 0015, Yifei Yin, Longshan Yao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Deep Feature Correlation Learning for Multi-Modal Remote Sensing Image RegistrationabstractDeep descriptors have advantages over handcrafted descriptors on local image patch matching. However, due to the complex imaging mechanism of remote sensing images and the significant differences in appearance between multi-modal images, existing deep learning descriptors are unsuitable for multi-modal remote sensing image registration directly. To solve this problem, this paper proposes a deep feature correlation learning network (Cnet) for multi-modal remote sensing image registration. Firstly, Cnet builds a feature learning network based on the deep convolutional network with the attention learning module, to enhance the feature representation by focusing on meaningful features. Secondly, this paper designs a novel feature correlation loss function for Cnet optimization. It focuses on the relative feature correlation between matching and non-matching samples, which can improve the stability of network training and decrease the risk of overfitting. Additionally, the proposed feature correlation loss with a scale factor can further enhance the network training and accelerate the network convergence. Extensive experimental results on image patch matching (Brown, HPatches), cross-spectral image registration (VIS-NIR), multi-modal remote sensing image registration, and single-modal remote sensing image registration have demonstrated the effectiveness and robustness of the proposed method. Dou Quan, Shuang Wang 0001, Yu Gu 0015, Ruiqi Lei, Bowu Yang, Shaowei Wei, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Joint-Training Two-Stage Method For Remote Sensing Image CaptioningabstractCompared with remote sensing image (RSI) captioning methods based on the traditional encoder-decoder model, two-stage RSI captioning methods include an auxiliary remote sensing task to provide prior information, which enables them to generate more accurate descriptions. In previous two-stage RSI captioning methods, however, the image captioning and the auxiliary remote sensing tasks are handled separately, which is time-consuming and ignores mutual interference between tasks. To solve this problem, we propose a novel joint-training two-stage (JTTS) RSI captioning method. We use multi-label classification to provide prior information, and we design a differentiable sampling operator to replace the traditional non-differentiable sampling operation to index the multi-label classification result. In contrast to previous two-stage RSI captioning methods, our method can implement joint-training, and the joint loss allows the error of the generated description to flow into the optimization of the multi-label classification via back-propagation. Specifically, we approximate the Heaviside step function with the steep logistic function to implement a differentiable sampling operator for the multi-label classification. We propose a dynamic contrast loss function for multi-label classification task to ensure that a certain margin is maintained between the probabilities of the positive label and the negative label during sampling. We design an attribute-guided decoder to filter the multi-label prior information obtained by the sampling operator to generate more accurate image captions. The results of extensive experiments show that the JTTS method achieves state-of-the-art performance on the RSICD, the UCM-Captions, and the Sydney-Captions datasets. Xiutiao Ye, Shuang Wang 0001, Yu Gu 0015, Jihui Wang, Biao Hou, Fausto Giunchiglia, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Cluster Alignment With Target Knowledge Mining for Unsupervised Domain Adaptation Semantic SegmentationabstractUnsupervised domain adaptation (UDA) carries out knowledge transfer from the labeled source domain to the unlabeled target domain. Existing feature alignment methods in UDA semantic segmentation achieve this goal by aligning the feature distribution between domains. However, these feature alignment methods ignore the domain-specific knowledge of the target domain. In consequence, 1) the correlation among pixels of the target domain is not explored; and 2) the classifier is not explicitly designed for the target domain distribution. To conquer these obstacles, we propose a novel cluster alignment framework, which mines the domain-specific knowledge when performing the alignment. Specifically, we design a multi-prototype clustering strategy to make the pixel features within the same class tightly distributed for the target domain. Subsequently, a contrastive strategy is developed to align the distributions between domains, with the clustered structure maintained. After that, a novel affinity-based normalized cut loss is devised to learn task-specific decision boundaries. Our method enhances the model's adaptability in the target domain, and can be used as a pre-adaptation for self-training to boost its performance. Sufficient experiments prove the effectiveness of our method against existing state-of-the-art methods on representative UDA benchmarks. Shuang Wang 0001, Dong Zhao 0007, Yuwei Guo 0001, Qi Zang, Yu Gu 0015, Yi Li 0054, Licheng Jiao |
IEEE Trans. Image Process. | 6 |
| 2021 | Cascade Attention Fusion for Fine-Grained Image Captioning Based on Multi-Layer LSTMabstractThe conventional visual attention-based image captioning approaches typically use image information to guide caption generation. Results from these models tend to be coarse and ignore the details in the image, such as objects, attributes and the distinguishing aspects of each image. In this paper, we propose a visual and semantic fusion network with a margin-based training guidance mechanism to generate fine image descriptions that depict more objects, attributes and other distinguishing aspects of images. In our model, the visual attention layer introduces more low-level visual information, the semantic attention layer provides more high-level semantic attributes. Furthermore, the proposed margin-based loss encourages our model to produce more discriminative descriptions. Extensive experiments are conducted on COCO and Flickr30K image captioning datasets to validate our method, and the results show its superior performance at captioning. Our method achieves a state-of-the-art 70.6 CIDEr-D on the Flickr30K dataset, and a competitive 123.5 CIDEr-D on the MS-COCO dataset. Shuang Wang 0001, Yun Meng, Yu Gu 0015, Xiutiao Ye, Jingxian Tian, Licheng Jiao |
ICASSP | 3 |
| 2021 | Progressive Co-Teaching for Ambiguous Speech Emotion RecognitionabstractSpeech emotion recognition is a challenging task due to the ambiguity of emotion, which makes it difficult to learn the features of emotion data using machine learning algorithms. However, previous studies conventionally ignore the ambiguity of emotion and treat the emotion data as the same difficulty level, which results in low recognition accuracy. Motivated by human and animal learning studies, we propose a novel method named Progressive Co-teaching (PCT) to learn speech emotion features from simple to difficult. PCT method automatically identifies the difficulty level of data by itself using loss values, and then each network exchanges easy instances with small loss to peer network for early training. The rest instances with large loss are added gradually for later training. The experiment results demonstrate that our method achieves an improvement of 3.8% and 1.27% on MAS and IEMOCAP database than the state-of-the-arts, respectively. Yifei Yin, Yu Gu 0015, Longshan Yao, Ying Zhou 0021, Xuefeng Liang |
ICASSP | 2 |
| 2021 | Multi-View Attention Network for Remote Sensing Image CaptioningabstractIn traditional remote sensing image captioning models, the attention mechanism plays a dominant role and has been used to integrate image features to infer the latent visual-semantic alignment. However, the scenes of remote sensing image are complex and diverse, using only one attention module to capture features often leads to insufficient semantic representation. In our work, we present a novel Multi-view Attention Network (MAN) model to realize feature integration from different views. With MAN, more semantically rich ensemble attended features can be obtained by different attention modules. Specifically, we enforce the weights of attention modules to be diverse through a cosine distance loss. This will provide the model with distinct views to make semantic predictions for each feature. Extensive experiments on benchmark datasets demonstrate the effectiveness of the proposed model for the task of remote sensing image captioning. Yun Meng, Yu Gu 0015, Xiutiao Ye, Jingxian Tian, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 2 |
| 2021 | Cross-Modal Feature Fusion Retrieval for Remote Sensing Image-Voice RetrievalabstractWith the increasing popularity of remote sensing technology applications, some emergency scenarios require rapid retrieval of remote sensing images, such as earthquake rescue, etc. Due to the high efficiency of voice input, researchers have focused on cross-modal remote sensing image-voice retrieval methods. However, these methods have two major drawbacks: speech input lacks discrimination and the intra-modal semantic information is under used. To address these drawbacks, we propose a novel cross-modal feature fusion retrieval model. Our model provides a more optimized cross-modal common feature space than previous models and thus optimizes the retrieval performance. First, our model adds the extra textual keyword information to the audio feature for remote sensing image retrieval. Second, it introduces inter-modality adversarial learning and intra-modality semantic discrimination into the remote sensing image-voice retrieval task. We conducted experiments on two datasets modified from the UCM-Captions dataset and the Remote Sensing Image Caption Dataset. The experimental results show that our model outperforms state-of-the-art models in this task. Rui Yang 0038, Yu Gu 0015, Yu Liao, Yingzhi Sun, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 2 |
| 2020 | Jointly Discriminating and Frequent Visual Representation Mining
Qiannan Wang, Ying Zhou 0021, Zhaoyan Zhu, Xuefeng Liang, Yu Gu 0015 |
ACCV (3) | 5 |
| 2017 | Full polarization SAR image classification using deep learning with shallow featureabstractThe classification of the POL-SAR image become more and more important with the development of the polarization of synthetic aperture radar system. Generally, the classification of POL-SAR images are based on polarization feature, such as support vector machine (SVM), Wishart clustering and other methods. Specifically, some ground objects usually have some weak scattering characteristics which cannot obtain good results by only using the traditional classification based on polarization features. So, the deep learning based on T matrix is used to mine the powerful feature of SAR data. In order to speed up computation and improve classification accuracy, a classification of full-polarization SAR images based on Deep Learning with Shallow features is proposed in this paper. The proposed method can get better classification for those weak scatter objects than those methods only using polarization features. Debo Li, Yu Gu 0015, Shuiping Gou, Licheng Jiao |
IGARSS | 2 |
| 2016 | Speech Emotion Recognition Using Voiced Segment Selection AlgorithmabstractSpeech emotion recognition (SER) poses one of the major challenges in human-machine interaction. We propose a new algorithm, the Voiced Segment Selection (VSS) algorithm, which can produce an accurate segmentation of speech signals. The VSS algorithm deals with the voiced signal segment as the texture image processing feature which is different from the traditional method. It uses the Log-Gabor filters to extract the voiced and unvoiced features from spectrogram to make the classification. The finding shows that the VSS method is a more accurate algorithm for voiced segment detection. Therefore, it has potential to improve performance of emotion recognition from speech. Yu Gu 0015, Eric O. Postma, Hai-Xiang Lin, H. Jaap van den Herik |
ECAI | 1 |
| 2016 | Speech Emotion Recognition with Log-Gabor Filters
Yu Gu 0015, Eric O. Postma, Hai-Xiang Lin, H. Jaap van den Herik |
ICAART (2) | 1 |