VLDB 2026 Research / reviewers in the wild / expert
Xinjian Gao
dblp:170/5806
· DBLP profile ↗
15ranked-venue papers
10as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust multimodal face anti-spoofing via frequency-domain feature refinement and aggregation
Rui Sun 0004, Xiaolu Yu, Xinjian Gao |
Pattern Recognit. Lett. | 4 |
| 2024 | Multimodal Visual-Semantic Representations Learning for Scene Text RecognitionabstractScene Text Recognition (STR), the critical step in OCR systems, has attracted much attention in computer vision. Recent research on modeling textual semantics with Language Model (LM) has witnessed remarkable progress. However, LM only optimizes the joint probability of the estimated characters generated from the Vision Model (VM) in a single language modality, ignoring the visual-semantic relations in different modalities. Thus, LM-based methods can hardly generalize well to some challenging conditions, in which the text has weak or multiple semantics, arbitrary shape, and so on. To migrate the above issue, in this paper, we propose Multimodal Visual-Semantic Representations Learning for Text Recognition Network (MVSTRN) to reason and combine the multimodal visual-semantic information for accurate Scene Text Recognition. Specifically, our MVSTRN builds a bridge between vision and language through its unified architecture and has the ability to reason visual semantics by guiding the network to reconstruct the original image from the latent text representation, breaking the structural gap between vision and language. Finally, the tailored multimodal Fusion (MMF) module is motivated to combine the multimodal visual and textual semantics from VM and LM to make the final predictions. Extensive experiments demonstrate our MVSTRN achieves state-of-the-art performance on several benchmarks. Xinjian Gao, Ye Pang, Yuyu Liu, Maokun Han, Jun Yu 0001, Wei Wang 0496, Yuanxu Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Macroscopic-and-Microscopic Rain Streaks Disentanglement Network for Single-Image DerainingabstractSingle-image deraining aims to restore the image that is degraded by the rain streaks, where the long-standing bottleneck lies in how to disentangle the rain streaks from the given rainy image. Despite the progress made by substantial existing works, several crucial questions - e.g., How to distinguish rain streaks and clean image, while how to disentangle rain streaks from low-frequency pixels, and further prevent the blurry edges - have not been well investigated. In this paper, we attempt to solve all of them under one roof. Our observation is that rain streaks are bright stripes with higher pixel values that are evenly distributed in each color channel of the rainy image, while the disentanglement of the high-frequency rain streaks is equivalent to decreasing the standard deviation of the pixel distribution for the rainy image. To this end, we propose a self-supervised rain streaks learning network to characterize the similar pixel distribution of the rain streaks from a macroscopic viewpoint over various low-frequency pixels of gray-scale rainy images, coupling with a supervised rain streaks learning network to explore the specific pixel distribution of the rain streaks from a microscopic viewpoint between each paired rainy and clean images. Building on this, a self-attentive adversarial restoration network comes up to prevent the further blurry edges. These networks compose an end-to-end Macroscopic-and-Microscopic Rain Streaks Disentanglement Network, named [Formula: see text]RSD-Net, to learn rain streaks, which is further removed for single image deraining. The experimental results validate its advantages on deraining benchmarks against the state-of-the-arts. The code is available at: https://github.com/xinjiangaohfut/MMRSD-Net. Xinjian Gao, Yang Wang 0023, Meng Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | DBCAN: Dual-Branch Cross-Attention Network for Scene Text RecognitionabstractScene text recognition, especially irregular text recognition, is a challenging task due to the large variance in text appearance. Although some existing methods have achieved state-of-the-art performance with the attention-based encoder-decoder framework, they always perform poorly on some challenging text such as severely curved, blurred, and incomplete-semantic text. To address these issues, we propose a Dual-Branch Cross-Attention Network (DBCAN). Different from the previous methods heavily relying on semantic information, DBCAN can enhance the position clues and learn semantic relations with two separate branches and fuse them by a tailored Cross-Attention Module (CAM). Furthermore, a Convolution-Based 2D Positional Embedding (CBPE) is introduced to describe the 2D spatial dependencies of characters. Extensive experiments demonstrate our DBCAN is more accurate and robust than the previous methods and achieves state-of-the-art performance on several benchmarks, particularly CUTE (93.4%). Our code is made publicly available at https://github.com/GaoXinJian-USTC/DBCAN. Xinjian Gao, Ye Pang, Yuyu Liu, Jun Yu 0001, Maokun Han, Wei Wang 0496 |
ICME | 1 |
| 2022 | Attention in Attention: Modeling Context Correlation for Efficient Video ClassificationabstractAttention mechanisms have significantly boosted the performance of video classification neural networks thanks to the utilization of perspective contexts. However, the current research on video attention generally focuses on adopting a specific aspect of contexts (e.g., channel, spatial/temporal, or global context) to refine the features and neglects their underlying correlation when computing attentions. This leads to incomplete context utilization and hence bears the weakness of limited performance improvement. To tackle the problem, this paper proposes an efficient attention-in-attention (AIA) method for element-wise feature refinement, which investigates the feasibility of inserting the channel context into the spatio-temporal attention learning module, referred to as CinST, and also its reverse variant, referred to as STinC. Specifically, we instantiate the video feature contexts as dynamics aggregated along a specific axis with global average and max pooling operations. The workflow of an AIA module is that the first attention block uses one kind of context information to guide the gating weights calculation of the second attention that targets at the other context. Moreover, all the computational operations in attention units act on the pooled dimension, which results in quite few computational cost increase (https://github.com/haoyanbin918/Attention-in-Attention. Yanbin Hao, Shuo Wang 0008, Pei Cao 0001, Xinjian Gao, Tong Xu 0001, Jinmeng Wu, Xiangnan He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Improving Image Similarity Learning by Adding External MemoryabstractThe type of neural networks widely used in artificial intelligence applications mixes its computation and memory modules in neuron weights and activities. The previously learned information are stored in network weights. When dealing with complex data, e.g., those possessing diverse content or containing long-sequences, some information stored in the weights can be altered drastically or wiped as the training goes, but they are not necessarily unimportant. External memory is a recent technique proposed to prevent from forgetting significant previously learned information. In this work, we aim at taking advantage of this recent technique to advance the similarity learning task that is critical in many real-world artificial intelligence applications. We propose suitable external memory design supported by extended attention mechanism. Two different kinds of memory modules are proposed so that the similarity learning process can dynamically shift focus over a wide range of diverse content contained by the training data. Effectiveness of the proposed method is demonstrated through evaluations based on different image retrieval tasks and compared against various state-of-the-art algorithms in the field. Xinjian Gao, Tingting Mu, John Yannis Goulermas, Jingkuan Song, Meng Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | Multi-Scale Spatial Transform Network for Atmospheric Polarization Prediction
Tianyi Dang, Xinjian Gao, Linfang Xie, Yuan Yan, Jun Gao 0006 |
ICIG (1) | 3 |
| 2021 | Radar Object Detection Using Data Merging, Enhancement and FusionabstractCompared to visible images, radar images are generally considered to be an active and robust solution, even in adverse driving situations, for object detection. However, the accuracy of radar object detection (ROD) is always poor. Owing to taking full advantage of data merging, enhancement and fusion, this paper proposes an effective ROD system with only radar images as the input. First, an aggregation module is designed to merge the data from all chirps in the same frame. Then, various gaussian noises with different parameters are employed to increase data diversity and reduce over-fitting based on the analysis of training data. Moreover, due to the process of inference with default parameters is not accurate enough, some hyperparameters are changed to increase the accuracy performance. Finally, a combination strategy is adopted to benefit from multi-model fusion. ROD2021 Challenge is supported by ACM ICMR 2021, and our team (ustc-nelslip) ranked 2nd in the test stage of this challenge. Diverse evaluations also verify the superiority of the proposed system. Jun Yu 0001, Xinlong Hao, Xinjian Gao, Yuyu Liu, Peng Chang 0002, Fang Gao 0001, Feng Shuang 0002 |
ICMR | 3 |
| 2021 | Meta-learning based relation and representation learning networks for single-image deraining
Xinjian Gao, Yang Wang 0023, Jun Cheng 0002, Mingliang Xu 0001, Meng Wang 0001 |
Pattern Recognit. | 1 |
| 2020 | Weakly-Supervised Video Object Grounding by Exploring Spatio-Temporal ContextsabstractGrounding objects in visual context from natural language queries is a crucial yet challenging vision-and-language task, which has gained increasing attention in recent years. Existing work has primarily investigated this task in the context of still images. Despite their effectiveness, these methods cannot be directly migrated into the video context, mainly due to 1) the complex spatio-temporal structure of videos and 2) the scarcity of fine-grained annotations of videos. To effectively ground objects in videos is profoundly more challenging and less explored. Xun Yang 0001, Xueliang Liu, Meng Jian, Xinjian Gao, Meng Wang 0001 |
ACM Multimedia | 4 |
| 2020 | Self-attention driven adversarial similarity learning network
Xinjian Gao, Zhao Zhang 0001, Tingting Mu, Chaoran Cui, Meng Wang 0001 |
Pattern Recognit. | 1 |
| 2020 | An Interpretable Deep Architecture for Similarity Learning Built Upon Hierarchical ConceptsabstractIn general, development of adequately complex mathematical models, such as deep neural networks, can be an effective way to improve the accuracy of learning models. However, this is achieved at the cost of reduced post-hoc model interpretability, because what is learned by the model can become less intelligible and tractable to humans as the model complexity increases. In this paper, we target a similarity learning task in the context of image retrieval, with a focus on the model interpretability issue. An effective similarity neural network (SNN) is proposed to offer not only to seek robust retrieval performance but also to achieve satisfactory post-hoc interpretability. The network is designed by linking the neuron architecture with the organization of a concept tree and by formulating neuron operations to pass similarity information between concepts. Various ways of understanding and visualizing what is learned by the SNN neurons are proposed. We also exhaustively evaluate the proposed approach using a number of relevant datasets against a number of state-of-the-art approaches to demonstrate the effectiveness of the proposed network. Our results show that the proposed approach can offer superior performance when compared against state-of-the-art approaches. Neuron visualization results are demonstrated to support the understanding of the trained neurons. Xinjian Gao, Tingting Mu, John Yannis Goulermas, Jeyan Thiyagalingam, Meng Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Attention driven multi-modal similarity learning
Xinjian Gao, Tingting Mu, John Yannis Goulermas, Meng Wang 0001 |
Inf. Sci. | 1 |
| 2018 | Topic driven multimodal similarity learning with multi-view voted convolutional features
Xinjian Gao, Tingting Mu, John Yannis Goulermas, Meng Wang 0001 |
Pattern Recognit. | 1 |
| 2016 | Local voting based multi-view embedding
Xinjian Gao, Tingting Mu, Meng Wang 0001 |
Neurocomputing | 1 |