EDBT 2026 Demo / reviewers in the wild / expert
Xiaodong Gu 0001
dblp:71/4467-1
· DBLP profile ↗
113ranked-venue papers
10as first author
54since 2021 · last 2026
0000-0002-7096-1830ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 93 · 10 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Audio-visual segmentation via hierarchical side tuning with state space model
Wenyi Xia, Qingwei Geng, Xiaodong Gu 0001 |
Appl. Intell. | 3 |
| 2026 | Knowledge distillation from single-label to multi-label with class activation maps
Boyu Guan, Hengwei Liu, Youcai Zhang, Yuzhuo Qin, Xiaodong Gu 0001 |
Expert Syst. Appl. | 5 |
| 2026 | MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models
Yuncheng Guo, Xiaodong Gu 0001 |
Int. J. Comput. Vis. | 2 |
| 2026 | Vision-language alignment with sigmoid loss and dual-token contrastive change localizer for precise change captioning
Xiaodong Gu 0001 |
Neurocomputing | 2 |
| 2025 | MMRL: Multi-Modal Representation Learning for Vision-Language ModelsabstractLarge-scale pre-trained Vision-Language Models (VLMs) have become essential for transfer learning across diverse tasks. However, adapting these models with limited few-shot data often leads to overfitting, diminishing their performance on new tasks. To tackle this issue, we propose a novel Multi-Modal Representation Learning (MMRL) framework that introduces a shared, learnable, and modality-agnostic representation space. MMRL projects the space tokens to text and image representation tokens, facilitating more effective multi-modal interactions. Unlike previous approaches that solely optimize class token features, MMRL integrates representation tokens at higher layers of the encoders—where dataset-specific features are more prominent—while preserving generalized knowledge in the lower layers. During training, both representation and class features are optimized, with trainable projection layer applied to the representation tokens, whereas the class token projection layer remains frozen to retain pre-trained knowledge. Furthermore, a regularization term is introduced to align the class features and text features with the zero-shot features from the frozen VLM, thereby safeguarding the model’s generalization capacity. For inference, a decoupling strategy is employed, wherein both representation and class features are utilized for base classes, while only the class features, which retain more generalized knowledge, are used for new tasks. Extensive experiments across 15 datasets demonstrate that MMRL outperforms state-of-the-art methods, achieving a balanced trade-off between task-specific adaptation and generalization. Code is available at https://github.com/yunncheng/MMRL. Yuncheng Guo, Xiaodong Gu 0001 |
CVPR | 2 |
| 2025 | Hybrid Visual Adapter and Drop-View Training for Change CaptioningabstractIn this paper, we address the task of change captioning, which aims to generate sentences describing fine-grained differences between similar images while resisting distractors like viewpoint variations. To overcome the challenges of learning robust difference representations and avoiding over-reliance on partial features, we propose two key innovations. First, the Hybrid Visual Adapter (HVA) enhances the Vision Transformer (ViT) backbone by explicitly modeling cross-image interactions through attention mechanisms. By attending to counterpart regions during feature extraction, HVA suppresses irrelevant distractors and learns viewpoint-invariant representations, significantly improving robustness against spatial misalignment. Second, we introduce drop-view training, a strategy that randomly masks one of the three input features (before, after, or difference features) during caption generation. This forces the decoder to fully explore complementary visual cues. Extensive experimental results show that our method achieves superior performance on three public datasets with different change types. Xiaodong Gu 0001 |
IJCNN | 2 |
| 2025 | Intra-frame scan-free video state spaces model for video moment retrieval
Fengzhen Yu, Xiaodong Gu 0001 |
Appl. Intell. | 2 |
| 2025 | Exploiting explicit item-item correlations from knowledge graphs for enhanced sequential recommendationabstractIn recent years, the research of employing knowledge graphs (KGs) in sequential recommendation (SR) has received a lot of attention, since the side information extracted from KGs, especially the information of the correlations between items, indeed helps the SR models achieve better performance. However, many previous KG-based SR models tend to introduce some noise information when learning item embeddings, or insufficiently fuse item–item correlations into their sequential modeling, thus limiting their performance improvements . In this paper, we propose a D istance- A ware K nowledge-based S equential R ecommendation model ( DAKSR ), which exploits the explicit item–item correlations from KGs to achieve enhanced SR. Specifically, as one critical component in our DAKSR, the distance score matrix (DSM) is first obtained to indicate the correlations between items, and then leveraged in the following three major modules of DAKSR. First, in the Item-Set Embedding layer (ISE) all item embeddings are learned based on DSM, in which the noise information is eliminated effectively. Meanwhile, the Knowledge-Infused Transformer (KIT) incorporates DSM into its attention mechanism to improve the feature extraction. Furthermore, the Knowledge Contrastive Learning module (KCL) also leverages the item–item correlations presented in DSM to generate two credible sequence views, which are used to refine sample representations through a contrastive learning strategy, and thus improve the model’s robustness. Our extensive experiments on three SR benchmarks obviously demonstrate our DAKSR’s superior performance over the state-of-the-art (SOTA) KG-based recommendation models. The implementation of our DAKSR is available at https://github.com/Easonsi/DAKSR for reproducing our experiment results conveniently. Yanlin Zhang, Deqing Yang, Xiaodong Gu 0001 |
Inf. Syst. | 4 |
| 2025 | Towards photorealistic face generation using text-guided Semantic-Spatial FaceGAN
Xiaodong Gu 0001 |
Multim. Tools Appl. | 2 |
| 2025 | Attention-Based Gating Network for Robust Segmentation TrackingabstractVisual object tracking is a challenging task that aims to accurately estimate the scale and position of a designated target. Recently, segmentation networks have proven effective in visual tracking, producing outstanding results for target scale estimation. However, segmentation-based trackers still lack robustness due to the presence of similar distractors. To mitigate this issue, we propose an Attention-based Gating Network (AGNet) that produces gating weights to diminish the impact of feature maps linked to similar distractors. Subsequently, we incorporate the AGNet into the segmentation-based tracking paradigm to achieve accurate and robust tracking. Specifically, the AGNet utilizes three cascading Multi-Head Cross-Attention (MHCA) modules to generate gating weights that govern the generation of feature maps in the baseline tracker. The proficiency of the MHCA in modeling global semantic information effectively suppresses feature maps associated with similar distractors. Additionally, we introduce a distractor-aware training strategy that leverages distractor masks to train our model. To alleviate the issue of partial occlusion, we introduce a box refinement module to enhance the accuracy of the predicted target box. Comprehensive experiments conducted on 11 challenging tracking benchmarks show that our approach significantly surpasses the baseline tracker across all metrics and achieves excellent results on multiple tracking benchmarks. Yijin Yang, Xiaodong Gu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Integrating Audio-Visual Contexts with Refinement for Segmentation
Qingwei Geng, Xiaodong Gu 0001 |
ICANN (3) | 2 |
| 2024 | Boundary-Aware Noise-Resistant Video Moment Retrieval
Fengzhen Yu, Xiaodong Gu 0001 |
ICANN (3) | 2 |
| 2024 | Masked co-attention model for audio-visual event localization
Hengwei Liu, Xiaodong Gu 0001 |
Appl. Intell. | 2 |
| 2024 | Enhancing accuracy, diversity, and random input compatibility in face attribute manipulation
Xiaodong Gu 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Co-saliency detection with two-stage co-attention mining and individual calibration
Zhenshan Tan, Xiaodong Gu 0001, Qingrong Cheng |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | An improved StyleGAN-based TextToFace model with Local-Global information Fusion
Xiaodong Gu 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Enhancing user and item representation with collaborative signals for KG-based recommendation
Yanlin Zhang, Xiaodong Gu 0001 |
Neural Comput. Appl. | 2 |
| 2024 | Semantic Pre-Alignment and Ranking Learning With Unified Framework for Cross-Modal RetrievalabstractCross-modal retrieval aims at retrieving highly semantic relevant information among multi-modalities. Existing cross-modal retrieval methods mainly explore the semantic consistency between image and text while rarely consider the rankings of positive instances in the retrieval results. Moreover, these methods seldom take into account the cross-interaction between image and text, which leads to the deficiency of learning their semantic relations. In this paper, we propose a Unified framework with Ranking Learning (URL) for cross-modal retrieval. The unified framework consists of three sub-networks, visual network, textual network, and interaction network. Visual network and textual network project the image feature and text feature into their corresponding hidden spaces respectively. Then, the interaction network forces the target image-text representation to align in the common space. For unifying both semantics and rankings, we propose a new optimization paradigm including pre-alignment for semantic knowledge transfer and ranking learning for final retrieval, which can decouple semantic alignment and ranking learning. The former focuses on the semantic pre-alignment optimized by semantic classification and the latter revolves around the retrieval rankings. For the ranking learning, we introduce a cross-AP loss which can directly optimize the retrieval metric average precision for cross-modal retrieval. We conduct experiments on four widely-used benchmarks, including Wikipedia dataset, Pascal Sentence dataset, NUS-WIDE-10k dataset, and PKU XMediaNet dataset respectively. Extensive experimental results show that the proposed method can obtain higher retrieval precision. Qingrong Cheng, Zhenshan Tan, Keyu Wen, Xiaodong Gu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Learning Dynamical Position Embedding for Discriminative Segmentation TrackingabstractVisual tracking plays a pivotal role in intelligent transportation systems and has a wide range of practical applications such as autonomous driving and traffic counting. Recently, the attention mechanism in Transformers has been successfully applied to the field of visual tracking, leading to a significant improvement in tracking performance. However, Transformer-based trackers directly flatten two-dimensional image features into one-dimensional vectors to compute attention scores. This process unavoidably results in the omission of crucial position distribution information necessary for precise target localization. To address this issue, we propose a novel cross-attention based tracking-by-segmentation framework, called Dynamical Position Embedding based Tracking framework (DPET). DPET incorporates an additional network for modeling position information to complement the cross-attention module. To be specific, a dynamical position embedding network is introduced to adaptively encode position information. This network is then integrated into the cross-attention based feature fusion network to compensate for the loss of position distribution information. As a result, the fused feature incorporates abundant contextual semantic cues for target classification and precise position information for target localization simultaneously. To overcome the constraints imposed by bounding-boxes, a segmentation network that takes the fused feature as input is designed to achieve accurate pixel-wise tracking. Extensive experiments on eight challenging tracking benchmarks show that our DPET tracker enables real-time operations and achieves promising tracking performance on the GOT-10K benchmark. Especially, DPET tracker achieves the top accuracy scores on VOT2016, VOT2018 and VOT2019 benchmarks. Yijin Yang, Xiaodong Gu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Generating Distinctive Facial Images from Natural Language Descriptions via Spatial Map Fusion
Xiaodong Gu 0001 |
ICANN (5) | 2 |
| 2023 | A Cross-Modal View to Utilize Label Semantics for Enhancing Student Network in Multi-label Classification
Yuzhuo Qin, Hengwei Liu, Xiaodong Gu 0001 |
ICANN (1) | 3 |
| 2023 | A Unified Video Semantics Extraction and Noise Object Suppression Network for Video Saliency Detection
Zhenshan Tan, Xiaodong Gu 0001 |
ICANN (7) | 2 |
| 2023 | Triplet Spatiotemporal Aggregation Network for Video Saliency DetectionabstractThe effective aggregation of spatiotemporal information to accommodate real-world complex scenes is a fundamental issue in video saliency detection. In this paper, we propose a Triplet Spatiotemporal Aggregation Network (TSAN) to address it from the aggregation of spatiotemporal interaction, spatiotemporal information distribution, and multi-level spatiotemporal features. Firstly, we propose an interactive aggregation gate (IAG) module to model spatial and temporal global context information and perform inter-modal information transfer. Secondly, we employ an information distribution consistency (IDC) module to enhance the consistency of spatiotemporal representation by maximizing the correlation of spatiotemporal high-level features. Finally, we design a multi-level spatiotemporal feature aggregation (MSF) framework to merge cross-level and cross-modal features. These three modules are combined into a unified framework to jointly optimize spatiotemporal information for more precise results. Experimental results on five prevailing datasets show that TSAN outperforms previous competitors. Zhenshan Tan, Xiaodong Gu 0001 |
ICME | 3 |
| 2023 | Temporal Modeling Approach for Video Action Recognition Based on Vision-language Models
Xiaodong Gu 0001 |
ICONIP (3) | 2 |
| 2023 | Multi Feature Representation and Aggregation Network for Accurate and Robust Visual TrackingabstractSegmentation-based tracking paradigm has been successfully applied to tracking field and significantly improves the tracking performance. Although segmentation-based trackers are effective for target scale estimation, it makes the trackers have high requirements for the extracted target features due to the need for pixel-level segmentation. Therefore, in this article, we propose a novel multi feature representation and aggregation network and introduce it into tracking-by-segmentation framework to extract and integrate rich features for segmentation-based tracking. To be specific, the proposed approach firstly models three complementary feature representations through cross-attention, cross-correlation and dilated involution mechanisms respectively and employ a simple feature aggregation network to fuse these features. And then feeding those fusion features into a segmentation network obtains the accurate target state estimation. In addition, we introduce a bounding box refinement module to further refine the target box to alleviate the issues of partial occlusion and surrounding distractors. The extensive experimental results show that the proposed tracker achieves very promising tracking performance on seven challenging visual tracking benchmarks. Code and models are available at https://github.com/Yang428/FEAST. Yijin Yang, Xiaodong Gu 0001 |
IJCNN | 2 |
| 2023 | Learning rich feature representation and aggregation for accurate visual tracking
Yijin Yang, Xiaodong Gu 0001 |
Appl. Intell. | 2 |
| 2023 | Adversarial pre-optimized graph representation learning with double-order sampling for cross-modal retrieval
Qingrong Cheng, Xiaodong Gu 0001 |
Expert Syst. Appl. | 3 |
| 2023 | Learning joint relationship attention network for image captioning
Changzhi Wang, Xiaodong Gu 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Learning Double-Level Relationship Networks for image captioning
Changzhi Wang, Xiaodong Gu 0001 |
Inf. Process. Manag. | 2 |
| 2023 | Task-aware prototype refinement for improved few-shot learning
Xiaodong Gu 0001 |
Neural Comput. Appl. | 2 |
| 2023 | High-speed Scene Text Detection with Attention and Multi-scale Label Generation
Xiaodong Gu 0001 |
Neural Process. Lett. | 2 |
| 2023 | Accurate and robust visual tracking using bounding box refinement and online sample filtering
Yijin Yang, Xiaodong Gu 0001 |
Signal Process. Image Commun. | 2 |
| 2023 | Joint Correlation and Attention Based Feature Fusion Network for Accurate Visual TrackingabstractCorrelation operation and attention mechanism are two popular feature fusion approaches which play an important role in visual object tracking. However, the correlation-based tracking networks are sensitive to location information but loss some context semantics, while the attention-based tracking networks can make full use of rich semantic information but ignore the position distribution of the tracked object. Therefore, in this paper, we propose a novel tracking framework based on joint correlation and attention networks, termed as JCAT, which can effectively combine the advantages of these two complementary feature fusion approaches. Concretely, the proposed JCAT approach adopts parallel correlation and attention branches to generate position and semantic features. Then the fusion features are obtained by directly adding the location feature and semantic feature. Finally, the fused features are fed into the segmentation network to generate the pixel-wise state estimation of the object. Furthermore, we develop a segmentation memory bank and an online sample filtering mechanism for robust segmentation and tracking. The extensive experimental results on eight challenging visual tracking benchmarks show that the proposed JCAT tracker achieves very promising tracking performance and sets a new state-of-the-art on the VOT2018 benchmark. Yijin Yang, Xiaodong Gu 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Vision-Language Matching for Text-to-Image Synthesis via Generative Adversarial NetworksabstractText-to-image synthesis is an attractive but challenging task that aims to generate a photo-realistic and semantic consistent image from a specific text description. The images synthesized by off-the-shelf models usually contain limited components compared with the corresponding image and text description, which decreases the image quality and the textual-visual consistency. To address this issue, we propose a novel Vision-Language Matching strategy for text-to-image synthesis, named VLMGAN*, which introduces a dual vision-language matching mechanism to strengthen the image quality and semantic consistency. The dual vision-language matching mechanism considers textual-visual matching between the generated image and the corresponding text description, and visual-visual consistent constraints between the synthesized image and the real image. Given a specific text description, VLMGAN* firstly encodes it into textual features and then feeds them to a dual vision-language matching-based generative model to synthesize a photo-realistic and textual semantic consistent image. Besides, the popular evaluation metrics for text-to-image synthesis are borrowed from simple image generation, which mainly evaluate the reality and diversity of the synthesized images. Therefore, we introduce a metric named Vision-Language Matching Score (VLMS) to evaluate the performance of text-to-image synthesis which can consider both the image quality and the semantic consistency between the synthesized image and the description. The proposed dual multi-level vision-language matching strategy can be applied to other text-to-image synthesis methods. We implement this strategy on two popular baselines, which are marked with${\text{VLMGAN}_{+\text{AttnGAN}}}$and${\text{VLMGAN}_{+\text{DFGAN}}}$. The experimental results on two widely-used datasets show that the model achieves significant improvements over other state-of-the-art methods. Qingrong Cheng, Keyu Wen, Xiaodong Gu 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | A StyleCLIP-Based Facial Emotion Manipulation Method for Discrepant Emotion Transitions
Xiaodong Gu 0001 |
ACCV (4) | 2 |
| 2022 | Building Joint Relationship Attention Network for Image-Text GenerationabstractAttention based methods for image-text generation often focus on visual features individually, while ignoring relationship information among image features that provides important guidance for generating sentences. To alleviate this issue, in this work we propose the Joint Relationship Attention Network (JRAN) that novelly explores the relationships among the features. Specifically, different from the previous relationship based approaches that only explore the single relationship in the image, our JRAN can effectively learn two relationships, the visual relationships among region features and the visual-semantic relationships between region features and semantic features, and further make a dynamic trade-off between them during outputting the relationship representation. Moreover, we devise a new relationship based attention, which can adaptively focus on the output relationship representation when predicting different words. Extensive experiments on large-scale MSCOCO and small-scale Flickr30k datasets show that JRAN achieves state-of-the-art performance. More remarkably, JRAN achieves new 28.3% and 58.2% performance in terms of BLEU4 and CIDEr metric on Flickr30k dataset. Changzhi Wang, Xiaodong Gu 0001 |
COLING | 2 |
| 2022 | UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual DialogabstractVisual Dialog aims to answer multi-round, interactive questions based on the dialog history and image content. Existing methods either consider answer ranking and generating individually or only weakly capture the relation across the two tasks implicitly by two separate models. The research on a universal framework that jointly learns to rank and generate answers in a single model is seldom explored. In this paper, we propose a contrastive learning-based framework UTC to unify and facilitate both discriminative and generative tasks in visual dialog with a single model. Specifically, considering the inherent limitation of the previous learning paradigm, we devise two inter-task contrastive losses i.e., context contrastive loss and answer contrastive loss to make the discriminative and generative tasks mutually reinforce each other. These two com-plementary contrastive losses exploit dialog context and target answer as anchor points to provide representation learning signals from different perspectives. We evaluate our proposed UTC on the VisDial v1.0 dataset, where our method outperforms the state-of-the-art on both discriminative and generative tasks and surpasses previous state-of-the-art generative methods by more than 2 absolute points on Recall@1. Zhenshan Tan, Qingrong Cheng, Xin Jiang 0002, Qun Liu 0001, Yudong Zhu, Xiaodong Gu 0001 |
CVPR | 7 |
| 2022 | A Unified Multiple Inducible Co-attentions and Edge Guidance Network for Co-saliency Detection
Zhenshan Tan, Xiaodong Gu 0001 |
ICANN (1) | 2 |
| 2022 | Feature Recalibration Network for Salient Object Detection
Zhenshan Tan, Xiaodong Gu 0001 |
ICANN (4) | 2 |
| 2022 | A Unified Two-Stage Group Semantics Propagation and Contrastive Learning Network for Co-Saliency DetectionabstractCo-saliency detection (CoSOD) aims at discovering the repetitive salient objects from multiple images. Two primary challenges are group semantics extraction and noise object suppression. In this paper, we present a unified Two-stage grOup semantics PropagatIon and Contrastive learning NETwork (TopicNet) for CoSOD. TopicNet can be decomposed into two substructures, including a two-stage group semantics propagation module (TGSP) to address the first challenge and a contrastive learning module (CLM) to address the second challenge. Concretely, for TGSP, we design an image-to-group propagation module (IGP) to capture the consensus representation of intra-group similar features and a group-to-pixel propagation module (GPP) to build the relevancy of consensus representation. For CLM, with the design of positive samples, the semantic consistency is enhanced. With the design of negative samples, the noise objects are suppressed. Experimental results on three prevailing benchmarks reveal that TopicNet outperforms other competitors in terms of various evaluation metrics. Zhenshan Tan, Keyu Wen, Yuzhuo Qin, Xiaodong Gu 0001 |
ICME | 5 |
| 2022 | Image Captioning with Local-Global Visual Interaction Network
Changzhi Wang, Xiaodong Gu 0001 |
ICONIP (6) | 2 |
| 2022 | Image captioning with adaptive incremental global context attention
Changzhi Wang, Xiaodong Gu 0001 |
Appl. Intell. | 2 |
| 2022 | Dynamic-balanced double-attention fusion for image captioning
Changzhi Wang, Xiaodong Gu 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2022 | Co-saliency detection with intra-group two-stage group semantics propagation and inter-group contrastive learning
Zhenshan Tan, Xiaodong Gu 0001 |
Knowl. Based Syst. | 2 |
| 2022 | Visual context learning based on textual knowledge for image-text retrieval
Yuzhuo Qin, Xiaodong Gu 0001, Zhenshan Tan |
Neural Networks | 2 |
| 2021 | Towards Image Retrieval with Noisy Labels via Non-deterministic Features
Hengwei Liu, Jinyu Ma, Xiaodong Gu 0001 |
ICANN (3) | 3 |
| 2021 | Enhancing Separate Encoding with Multi-layer Feature Alignment for Image-Text Matching
Keyu Wen, Linyang Li, Xiaodong Gu 0001 |
ICANN (1) | 3 |
| 2021 | Subspace Constraint for Single Image Super-Resolution
Yanlin Zhang, Ding Qin, Xiaodong Gu 0001 |
ICANN (3) | 3 |
| 2021 | UED: A Unified Encoder Decoder Network for Visual Dialog
Xiaodong Gu 0001 |
ICONIP (6) | 2 |
| 2021 | An Image Captioning Approach Using Dynamical AttentionabstractIn recent years, as an active topic in the field of vision and language, image captioning has made great progress. Previous approaches have demonstrated the superiority of spatial and channel attentions in image captioning task. However, such attention-based approaches ignore the difference between function words (e.g., “to”, “for” and “out”) and notional words (e.g., “girl”, “teddy” and “bear”). To address above issue, in this paper we propose a dynamical balancing attention model (BAM) based on attention variation for image captioning, which uses attention variation to fuse channel attention and region attention. Generating function and notional words, it effectively balances the contribution of image channel feature and that of image region feature. Further, the proposed approach dynamically focuses on the most relevant attention features in word prediction. Extensive experimental results on typical datasets show our approach outperforms the attention based approaches and achieves competitive performance over existing end-to-end leading approaches. Changzhi Wang, Xiaodong Gu 0001 |
IJCNN | 2 |
| 2021 | Depth scale balance saliency detection with connective feature pyramid and edge guidance
Zhenshan Tan, Xiaodong Gu 0001 |
Appl. Intell. | 2 |
| 2021 | Context-aware network with foreground recalibration for grounding natural language in video
Xiaodong Gu 0001 |
Neural Comput. Appl. | 2 |
| 2021 | Bridging multimedia heterogeneity gap via Graph Representation Learning for cross-modal retrieval
Qingrong Cheng, Xiaodong Gu 0001 |
Neural Networks | 2 |
| 2021 | Learning Dual Semantic Relations With Graph Attention for Image-Text MatchingabstractImage-Text Matching is one major task in cross-modal information processing. The main challenge is to learn the unified visual and textual representations. Previous methods that perform well on this task primarily focus on not only the alignment between region features in images and the corresponding words in sentences, but also the alignment between relations of regions and relational words. However, the lack of joint learning of regional features and global features will cause the regional features to lose contact with the global context, leading to the mismatch with those non-object words which have global meanings in some sentences. In this work, in order to alleviate this issue, it is necessary to enhance the relations between regions and the relations between regional and global concepts to obtain a more accurate visual representation so as to be better correlated to the corresponding text. Thus, a novel multi-level semantic relations enhancement approach namedDual Semantic Relations Attention Network(DSRAN)is proposed which mainly consists of two modules, separate semantic relations module and the joint semantic relations module. DSRAN performs graph attention in both modules respectively for region-level relations enhancement and regional-global relations enhancement at the same time. With these two modules, different hierarchies of semantic relations are learned simultaneously, thus promoting the image-text matching process by providing more information for the final visual representation. Quantitative experimental results have been performed on MS-COCO and Flickr30K and our method outperforms previous approaches by a large margin due to the effectiveness of the dual semantic relations learning scheme. Keyu Wen, Xiaodong Gu 0001, Qingrong Cheng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Salient Object Detection with Edge Recalibration
Zhenshan Tan, Yikai Hua, Xiaodong Gu 0001 |
ICANN (1) | 3 |
| 2020 | Enriched Feature Representation and Combination for Deep Saliency Detection
Lecheng Zhou, Xiaodong Gu 0001 |
ICANN (1) | 2 |
| 2020 | End-to-end Saliency-Guided Deep Image Retrieval
Jinyu Ma, Xiaodong Gu 0001 |
ICONIP (4) | 2 |
| 2020 | SBN: Scale Balance Network for Accurate Salient Object DetectionabstractRecent great progress has been made on Salient Object Detection (SOD) by deep Convolutional Neural Networks (CNNs). However, most SOD methods still suffer from scale imbalance issue, which pays more attention on large salient areas but ignores small salient areas though they belong to the same object. To address this issue, this paper proposes a Scale Balance Network (SBN) to accurately locate large salient areas and recognize small salient areas. Firstly, a backbone network specifically designed for object detection is adopted in this paper, which captures larger resolution with more spatial features in deeper layers. Secondly, to focus on the balance between the large salient areas and the small salient areas, this paper proposes a novel Connective Feature Pyramid Module (CFPM) for sufficiently leveraging the multi-scale features and the multi-level features, which includes Feature Coherence Enhancement (FCE) and Feature Progressive Extraction (FPE). FCE is designed to enhance the coherence between high-level and low-level features, and FPE is designed to extract the progressive features in different convolutional layers. Finally, an Edge Enhancement Architecture with Various Kernels (EEAVK) is proposed to refine the edge features. Experimental results on five benchmark datasets show that the proposed method outperforms or achieves consistently superior performance in comparison with other methods under different evaluation metrics. Zhenshan Tan, Xiaodong Gu 0001 |
IJCNN | 2 |
| 2020 | Dual Semantic Relationship Attention Network for Image-Text MatchingabstractImage-Text Matching is one major task in cross-modal information processing. The main challenge is to learn the unified vision and language representations. Previous methods that perform well on this task primarily focus on the region features in images corresponding to the words in sentences. However, this will cause the regional features to lose contact with the global context, leading to the mismatch with those non-object words in some sentences. In this work, in order to alleviate this problem, a novel Dual Semantic Relationship Attention Network is proposed which mainly consists of two modules, separate semantic relationship module and the joint semantic relationship module. With these two modules, different hierarchies of semantic relationships are learned simultaneously, thus promoting the image-text matching process. Quantitative experiments have been performed on MS-COCO and Flickr-30K and our method outperforms previous approaches by a large margin due to the effectiveness of the dual semantic relationship attention scheme. Keyu Wen, Xiaodong Gu 0001 |
IJCNN | 2 |
| 2020 | Semantic Modulation Based Residual Network for Temporal Language Queries Grounding in Video
Xiaodong Gu 0001 |
ISNN | 2 |
| 2020 | Algorithm using supervised subspace learning and non-local representation for pose variation recognitionabstractPose variation has been one of the challenges of face recognition. To solve this challenge, the authors propose a classification algorithm using supervised subspace learning and non‐local representation (SSLNR). In SSLNR, they first propose a supervised subspace learning algorithm (SSLA). SSLA includes three different terms. The first term is the difference term, which can reduce the intra‐class differences. The second term is the block‐diagonal regularisation term, which promotes the samples to be represented by intra‐class samples. The last one is the noise robust term. Then, the original samples are mapped to the learned subspace by using SSLA. Thus, the intra‐class differences of the samples mapped to the learned subspace are reduced. Finally, those mapped samples are classified by proposed non‐local constraint‐based extended sparse representation classifier. SSLNR is extensively evaluated using four databases, namely Georgia Tech, Label faces in the wild, FEI and CVL. Experimental results show that SSLNR achieves better performance than some state‐of‐the‐art algorithms, such as DARG and RRNN. Mengmeng Liao, Changzhi Wang, Xiaodong Gu 0001 |
IET Comput. Vis. | 3 |
| 2020 | Face recognition approach by subspace extended sparse representation and discriminative feature learning
Mengmeng Liao, Xiaodong Gu 0001 |
Neurocomputing | 2 |
| 2020 | Scene image retrieval with siamese spatial attention pooling
Jinyu Ma, Xiaodong Gu 0001 |
Neurocomputing | 2 |
| 2020 | Subspace clustering based on alignment and graph embedding
Mengmeng Liao, Xiaodong Gu 0001 |
Knowl. Based Syst. | 2 |
| 2020 | Deep attentional fine-grained similarity network with adversarial learning for cross-modal retrieval
Qingrong Cheng, Xiaodong Gu 0001 |
Multim. Tools Appl. | 2 |
| 2020 | Single-image super-resolution with multilevel residual attention network
Ding Qin, Xiaodong Gu 0001 |
Neural Comput. Appl. | 2 |
| 2020 | Embedding topological features into convolutional neural network salient object detection
Lecheng Zhou, Xiaodong Gu 0001 |
Neural Networks | 2 |
| 2019 | FCN Salient Object Detection Using Region Cropping
Yikai Hua, Xiaodong Gu 0001 |
ICANN (3) | 2 |
| 2019 | Adversarial Learning for Cross-Modal Retrieval with Wasserstein Distance
Qingrong Cheng, Youcai Zhang, Xiaodong Gu 0001 |
ICONIP (1) | 3 |
| 2019 | Group Loss: An Efficient Strategy for Salient Object Detection
Yikai Hua, Xiaodong Gu 0001 |
ICONIP (4) | 2 |
| 2019 | Deep Residual-Dense Attention Network for Image Super-Resolution
Ding Qin, Xiaodong Gu 0001 |
ICONIP (5) | 2 |
| 2019 | Face recognition algorithm based on feature descriptor and weighted linear sparse representationabstractGenerally, the commonly used sparse‐based methods, such as sparse representation classifier, have achieved a good recognition result in face recognition. However, there exist several problems in those methods. First, those methods think that the importance of each atom is the same in representing other query samples. This is not reasonable because different atoms contain different amounts of information, their importance should be different when they together represent the query samples. Second, those methods cannot meet the real‐time requirement when dealing with large data set. In this study, on the one hand, the authors propose a fast extended sparse‐weighted representation classifier (FESWRC) by considering the different importance of atoms and using primal augmented Lagrangian method as well as principal component analysis. On the other hand, the authors propose a distinctive feature descriptor, named logarithmic‐weighted sum (LWS) feature descriptor. The authors combine FESWRC and LWS and used for face recognition, this method is called face recognition algorithm based on feature descriptor and weighted linear sparse representation (FDWLSR). Experimental results show that FDWLSR can realise real‐time recognition and the recognition rate can achieve 100.0, 100.0, 91.6, 93.4 and 87.4%, respectively, on the Yale, Olivetti Research Laboratory (ORL), faculdade de engenharia industrial (FEI), face recognition technology program (FERET) and labelled face in the wild datasets. Mengmeng Liao, Xiaodong Gu 0001 |
IET Image Process. | 2 |
| 2018 | Two-Stream Convolutional Neural Network for Multimodal Matching
Youcai Zhang, Yiwei Gu, Xiaodong Gu 0001 |
ICANN (1) | 3 |
| 2018 | Face Recognition Using Improved Extended Sparse Representation Classifier and Feature Descriptor
Mengmeng Liao, Xiaodong Gu 0001 |
ICIC (3) | 2 |
| 2018 | Data-Driven and Collision-Free Hybrid Crowd Simulation Model for Real Scenario
Qingrong Cheng, Zhiping Duan, Xiaodong Gu 0001 |
ICONIP (7) | 3 |
| 2018 | Deep Neural Network Based Salient Object Detection with Image Enhancement
Lecheng Zhou, Xiaodong Gu 0001 |
ICONIP (4) | 2 |
| 2018 | Hybrid classification approach using extreme learning machine and sparse representation classifier with adaptive thresholdabstractHere, the authors propose a hybrid classification approach using extreme learning machine (ELM) and sparse representation classifier (SRC) with adaptive threshold, which they called ATELMSRC. ATELMSRC can adaptively adjust the threshold, and make more test images correctly classified by ELM compared with ELMSRC, which not only reduces the classification time greatly but also improves the classification accuracy. In addition, primal augmented Lagrangian method is used in ATELMSRC to speed up the solution of ‐norm, which also speeds up the classification process. Experimental results on USPS handwritten digits data set and UMIST face data set show that the total classification time of the authors ATELMSRC is very short for large data sets, only 1/310 of SRC, 1/805 of extended SRC (ESRC), and 1/41 of ELMSRC. Meanwhile, the classification accuracy of the authors’ ATELMSRC is 97.80% on USPS handwritten digits data set, and 99.27% on UMIST face data set, which are higher than those of ELM, SRC, ESRC, ELMSRC etc. Mengmeng Liao, Xiaodong Gu 0001 |
IET Signal Process. | 2 |
| 2017 | Word Embedding Dropout and Variable-Length Convolution Window in Convolutional Neural Network for Sentiment Classification
Shangdi Sun, Xiaodong Gu 0001 |
ICANN (2) | 2 |
| 2017 | A Supervised Term Weighting Scheme for Multi-class Text Categorization
Yiwei Gu, Xiaodong Gu 0001 |
ICIC (3) | 2 |
| 2017 | Facial Expression Recognition Using Double-Stage Sample-Selected SVM
Xiaodong Gu 0001 |
ICIC (1) | 2 |
| 2017 | An Approach to Pulse Coupled Neural Network Based Vein Recognition
Xiaodong Gu 0001 |
ICONIP (6) | 2 |
| 2017 | Graph Embedding Learning for Cross-Modal Information Retrieval
Youcai Zhang, Xiaodong Gu 0001 |
ICONIP (3) | 2 |
| 2017 | FCN and Unit-Linking PCNN Based Image Saliency Detection
Lecheng Zhou, Xiaodong Gu 0001 |
ICONIP (3) | 2 |
| 2017 | Balancing between over-weighting and under-weighting in supervised term weighting
Haibing Wu, Xiaodong Gu 0001, Yiwei Gu |
Inf. Process. Manag. | 2 |
| 2017 | Cascaded Convolutional Neural Networks for Aspect-Based Opinion Summary
Xiaodong Gu 0001, Yiwei Gu, Haibing Wu |
Neural Process. Lett. | 1 |
| 2016 | Aspect-based Opinion Summarization with Convolutional Neural NetworksabstractThis paper studies Aspect-based Opinion Summarization (AOS) of reviews on particular products. In practice, an AOS system needs to address two core subtasks, aspect extraction and sentiment classification. Most existing approaches to aspect extraction, using linguistic analysis or topic modeling, are general across different products but not precise enough or suitable for particular products. Instead we take a less general but more precise scheme, which directly maps each review sentence into pre-defined aspects. To tackle aspect mapping and sentiment classification, we propose two Convolutional Neural Network (CNN) based methods, cascaded CNN and multitask CNN. Cascaded CNN contains two levels of convolutional networks. Multiple CNNs at level 1 deal with aspect mapping task, and a single CNN at level 2 deals with sentiment classification. Multitask CNN also contains multiple aspect CNNs and a sentiment CNN, but different networks share the same word embeddings. Experimental results show that both cascaded and multitask CNNs with pre-trained word embedding outperform linear classifiers, and multitask CNN generally performs better than cascaded CNN. Haibing Wu, Yiwei Gu, Shangdi Sun, Xiaodong Gu 0001 |
IJCNN | 4 |
| 2016 | Tracking Based on Unit-Linking Pulse Coupled Neural Network Image Icon and Particle Filter
Xiaodong Gu 0001 |
ISNN | 2 |
| 2015 | Moving Target Tracking Based on Pulse Coupled Neural Network and Optical Flow
Qiling Ni, Jianchen Wang, Xiaodong Gu 0001 |
ICONIP (3) | 3 |
| 2015 | Max-Pooling Dropout for Regularization of Convolutional Neural Networks
Haibing Wu, Xiaodong Gu 0001 |
ICONIP (1) | 2 |
| 2015 | A Graph Community and Bag of Categorized Visual Words Based Image Retrieval
Haoyuan Lu, Shangdi Sun, Xiaodong Gu 0001 |
ICONIP (4) | 4 |
| 2015 | Towards dropout training for convolutional neural networks
Haibing Wu, Xiaodong Gu 0001 |
Neural Networks | 2 |
| 2014 | Reducing Over-Weighting in Supervised Term Weighting for Sentiment Analysis
Haibing Wu, Xiaodong Gu 0001 |
COLING | 2 |
| 2014 | Image Retrieval Using a Novel Color Similarity Measurement and Neural Networks
Xiaodong Gu 0001 |
ICONIP (3) | 2 |
| 2013 | Attention selection using global topological properties based on pulse coupled neural network
Xiaodong Gu 0001, Yuanyuan Wang 0001 |
Comput. Vis. Image Underst. | 1 |
| 2012 | Vehicle License Plate Localization and License Number Recognition Using Unit-Linking Pulse Coupled Neural Network
Xiaodong Gu 0001 |
ICONIP (5) | 2 |
| 2011 | Attention selection model using Weight Adjusted Topological properties and quantification evaluating criterionabstractTopological properties have important function in human beings visual attention. TPQFT (Topological properties based Phase spectrum of Quaternion Fourier Transform) model is an attention selection model using topological properties expression we have introduced. A new quantification criterion to evaluate every channel's contribution and model's performance is proposed in this paper and used in TPQFT. This paper improves TPQFT model in several aspects and WTPQFT (Weight Adjusted Topological properties based Phase spectrum of Quaternion Fourier Transform) model is introduced. The experimental results show that WTPQFT model reflects the real attention selection more accurately than TPQFT and PQFT (Phase spectrum of Quaternion Fourier Transform) method. Xiaodong Gu 0001, Yuanyuan Wang 0001 |
IJCNN | 2 |
| 2010 | Histogram similarity measure using variable bin size distance
Xiaodong Gu 0001, Yuanyuan Wang 0001 |
Comput. Vis. Image Underst. | 2 |
| 2009 | Fault diagnosis of power electronic system based on fault gradation and neural network group
Chengcai Ma, Xiaodong Gu 0001, Yuanyuan Wang 0001 |
Neurocomputing | 2 |
| 2009 | Color discrimination enhancement for dichromats using self-organizing color transformation
Xiaodong Gu 0001, Yuanyuan Wang 0001 |
Inf. Sci. | 2 |
| 2008 | A Neural Oscillation Model for Contour Separation in Color Images
Xiaodong Gu 0001, Yuanyuan Wang 0001 |
ICONIP (2) | 2 |
| 2008 | Happy-Sad Expression Recognition Using Emotion Geometry Feature and Support Vector Machine
Linlu Wang, Xiaodong Gu 0001, Yuanyuan Wang 0001, Liming Zhang 0001 |
ICONIP (2) | 2 |
| 2008 | Image quality assessment using edge and contrast similarityabstractMeasurement of visual quality is of fundamental importance to some image processing applications. And the perceived image distortion of any image strongly depends on the local features, such as edges, flats and textures. Since edges often convey much information of an image, we propose a novel algorithm for image quality assessment based on the edge and contrast similarity between the distorted image and the reference(perfect) image. We demonstrate its promise through a set of intuitive examples, as well as validate its performance with subjective ratings. We also compare our method with two other state-of-the-art objective ones, which uses 550 images with different distortion types and BP neural network. Xiaodong Gu 0001, Yuanyuan Wang 0001 |
IJCNN | 2 |
| 2007 | A Fixed Transformation of Color Images for Dichromats Based on Similarity Matrices
Yinhui Deng, Yuanyuan Wang 0001, Jibin Bao, Xiaodong Gu 0001 |
ICIC (1) | 5 |
| 2007 | Classification Using Multi-valued Pulse Coupled Neural Network
Xiaodong Gu 0001 |
ICONIP (2) | 1 |
| 2007 | Anti-aircraft Missile Deployment Optimization Using Hopfield Neural NetworkabstractIn this paper, we construct a novel HNN energy function, and use the HNN with this energy function to solve anti-aircraft missile deployment, which is a constrained layout problem. A near-optimum solution is obtained when HNN reaches a stable state, i.e, the minimum of the energy function is reached. We have studied the convergence of the network and the relationship between parameters of the network and stability. A group of suitable parameters are obtained by lots of experiences. The simulation results in different scales show that our approach can obtain the near-optimum solution. In addition, our approach can get better solutions in larger scales than other methods such as the divide-and-conquer algorithm which is often used to solve constrained layout problems, and it can also be extended to other constrained layout problems, such as mobile base station planning and integrated circuit layout design. Xiaodong Gu 0001, Yuanyuan Wang 0001 |
IJCNN | 2 |
| 2006 | Object Detection Using Unit-Linking PCNN Image Icons
Xiaodong Gu 0001, Yuanyuan Wang 0001, Liming Zhang 0001 |
ISNN (2) | 1 |
| 2006 | A New Color Blindness Cure Model Based on BP Neural Network
Xiaodong Gu 0001, Yuanyuan Wang 0001 |
ISNN (2) | 2 |
| 2005 | General design approach to unit-linking PCNN for image processingabstractPCNN (pulse coupled neural network), a phenomenological model exhibiting synchronous pulse bursts, can be used in image processing efficiently by the specified algorithm corresponding to the specified application, but so far there has been no general design approach. This paper describes that the parallel pulse-spreading behavior of unit-linking PCNN for image processing, based on PCNN, is equal to the operation of mathematic morphology. Hereby we propose the general unit-linking PCNN design approach for binary image processing. In the meantime, in order to explain how to apply this new general design approach in the specified application, binary image denoising, binary image edge detection, binary image hole-filter based on unit-linking PCNN are analyzed respectively as examples. Xiaodong Gu 0001, Liming Zhang 0001, Daoheng Yu |
IJCNN | 1 |
| 2005 | Global Icons and Local Icons of Images Based Unit-Linking PCNN and Their Application to Robot Navigation
Xiaodong Gu 0001, Liming Zhang 0001 |
ISNN (2) | 1 |
| 2005 | Image shadow removal using pulse coupled neural networkabstractThis paper introduces an approach for image shadow removal by using pulse coupled neural network (PCNN), based on the phenomena of synchronous pulse bursts in the animal visual cortexes. Two shadow-removing criteria are proposed. These two criteria decide how to choose the optimal parameter (the linking strength beta). The computer simulation results of shadow removal based on PCNN show that if these two criteria are satisfied, shadows are removed completely and the shadow-removed images are almost as the same as the original nonshadowed images. The shadow removal results are independent of changes of intensities of shadows in some range and variations of the places of shadows. When the first criterion is satisfied, even if the second criterion is not satisfied, as to natural grey images that have abundant grey levels, shadows also can be removed and PCNN shadow-removed images retain the shapes of the objects in original images. These two criteria also can be used for color images by dividing a color image into three channels (R, G, B). For shadows varying drastically, such as the noisy points in images, these two criteria are still right, but difficult to satisfy. Therefore, this approach can efficiently remove shadows that do not include the random noise. Xiaodong Gu 0001, Daoheng Yu, Liming Zhang 0001 |
IEEE Trans. Neural Networks | 1 |
| 2004 | Simplified PCNN and Its Periodic Solutions
Xiaodong Gu 0001, Liming Zhang 0001, Daoheng Yu |
ISNN (1) | 1 |
| 2004 | Delay PCNN and Its Application for Optimization
Xiaodong Gu 0001, Liming Zhang 0001, Daoheng Yu |
ISNN (1) | 1 |
| 2004 | Image thinning using pulse coupled neural network
Xiaodong Gu 0001, Daoheng Yu, Liming Zhang 0001 |
Pattern Recognit. Lett. | 1 |