VLDB 2026 Research / reviewers in the wild / expert
Gaoyun An
dblp:48/2349
· DBLP profile ↗
59ranked-venue papers
5as first author
28since 2021 · last 2026
0000-0002-2843-843XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 13 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CKCR: Context-aware knowledge construction and retrieval for knowledge-based visual question answering
Zhengxue Li, Gaoyun An |
J. Vis. Commun. Image Represent. | 6 |
| 2026 | DePoint: Improving rotation robustness of 3D point cloud analysis via decreasing entropy
Lu Shi 0004, Gaoyun An, Yi-Gang Cen, Yansen Huang, Fei Gan |
Neural Networks | 2 |
| 2026 | SHC: Deeply Activating Human-Like Cognitive Ability for Visual Question AnsweringabstractHuman cognitive mechanism depends on a sophisticated information processing framework, including perception, attention, memory, language, reasoning, problem solving and decision-making. However, current research only focuses on isolated process rather than systematically simulating human cognitive mechanism. Meanwhile, with the rapid development of large language models, related works have predominantly centered on language-level exploration, while in-depth mining of visual information remains insufficient. Here, to deeply activate the multi-modal understanding ability, a Systematic Human-like Cognitive (SHC) method is proposed for visual question answering, where the above mentioned sophisticated seven processes are systematically modeled as three core modules: hierarchical perception, semantic refinement and dynamic reasoning. The Hierarchical Perception Module (HPM) extracts hierarchical features from different levels to simulate the incremental integration mode of biological neural system. Based on the selective attention theory, one Semantic Refinement Module (SRM) is designed as a key-value accumulation optimization mechanism that enhances high-level semantics from low-level features via a multi-level cascaded attention structure. Finally, the Dynamic Reasoning Module (DRM), following the utility maximization decision theory, employs a dual weighting mechanism to dynamically fuse high-level semantic features and low-level fine-grained features, forming a unified high-quality visual representation that is then fed into the large language model for reasoning together with the text input. Experimental results demonstrate that SHC achieves competitive performance on multiple visual question answering benchmarks, including VQA-v2, Text-VQA, GQA, and ScienceQA, as well as multimodal evaluation benchmarks such as POPE, MMB, MME, and MM-Vet. Comparative experiments with multiple models of the same-scale validate the latent capacity of SHC to prompt the performance of multi-modal understanding tasks and its superiority in fine-grained visual information perception, and even surpasses multimodal models with larger-scale on certain tasks. Zhenxue Wang, Gaoyun An, Congyan Lang, Dapeng Oliver Wu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Human-Inspired Scene Understanding: A Grounded Cognition Method for Unbiased Scene Graph GenerationabstractScene Graph Generation (SGG) is a critical cross-modal task for scene understanding, which aims to detect visual relations in an image. Most SGG methods are significantly affected by highly skewed long-tailed bias, and prefer predicates with sufficient samples regardless of the semantic accuracy. Current unbiased SGG methods focus on compensating for the imbalanced long-tailed distribution, but they are fragile to dataset changes. The fundamental cause for this problem is the limited generalization ability, thus the diversity of classes needs to be modeled explicitly. By imitating the human cognition, a Grounded Cognition Method (GCM) for unbiased scene graph generation is proposed here, where the simulation, bodily states, and situated action are modeled. For simulations, an Out Domain Knowledge Injection module is proposed to expand the model's visual perception by reducing the reliance on an isolated class. Meanwhile, a Semantic Group Aware Synthesizer is proposed for linguistic perception modeling by categorizing specific predicate classes into a high-level semantic group. For bodily states, the modalities are erased separately to imitate the limited state of physical senses, which forces the model to rely on the remaining modality to compensate for the understanding of the whole scene. For situated actions, a Shapley Enhanced Multimodal Counterfactual module is proposed to model the dynamic interaction with the environment and cope with diverse contexts. Experiments on Visual Genome, GQA, and Open Images V6 demonstrate the effectiveness of our GCM, which outperforms state-of-the-art methods and achieves a better trade-off. Yiqing Hao, Gaoyun An, Binyang Song, Dapeng Oliver Wu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Hippocampal Memory-Like Separation-Completion Collaborative Network for Unbiased Scene Graph GenerationabstractScene Graph Generation (SGG) is a challenging cross-modal task, which aims to identify entities and relationships in a scene simultaneously. Due to the highly skewed long-tailed distribution, the generated scene graphs are dominated by relation categories of head samples. Current works address this problem by designing re-balancing strategies at the data level or refining relation representations at the feature level. Different from them, we attribute this impact to catastrophic interference, that is, the subsequent learning of dominant relations tends to overwrite the earlier learning of rare relations. To address it at the modeling level, a Hippocampal Memory-Like Separation-Completion Collaborative Network (HMSC2) is proposed here, which imitates the hippocampal encoding and retrieval process. Inspired by the pattern separation of dentate gyrus during memory encoding, a Gradient Separation Classifier and a Prototype Separation Learning module are proposed to relieve the catastrophic interference of tail categories by modeling the separated classifier and prototypes. In addition, inspired by the pattern completion of area CA3 of the hippocampus during memory retrieval, a Prototype Completion Module is designed to supplement the incomplete information of prototypes by introducing relation representations as cues. Finally, the completed prototype and relation representations are connected within a hypersphere space by a Contrastive Connected Module. Experimental results on the Visual Genome and GQA datasets show our HMSC2 achieves state-of-the-art performance on the unbiased SGG task, effectively relieving the long-tailed problem. The source codes are released on GitHub: https://github.com/Nora-Zhang98/HMSC2. Gaoyun An, Yiqing Hao, Dapeng Oliver Wu |
IEEE Trans. Image Process. | 2 |
| 2025 | Injecting Cross-modal Fine-Grained Perception into LLMs for 3D Object-of-Interest UnderstandingabstractRecent advancements in 3D Large Language Models (LLMs) have revealed significant potential in enhancing the understanding of 3D scenes. However, previous methods have struggled with extracting and utilizing fine-grained information of 3D objects for the coarsness of point clouds, resulting in limitations in understanding object-of-interested (OoI) within the scene. To address this issue, we introduce the object-centric 2D-3D interaction module for enhancing the ability of LLMs for 3D understanding tasks, which consists of the fine-grained 2D representation perception and the object-centric 3D scene representation perception. Specifically, the 2D representation associated with 3D objects is captured based on cross-modal semantic consistency without any spatial projector. Experimental results show that our model significantly outperforms existing methods on benchmarks including ScanRefer and ScanQA. Qianqian Sun, Lu Shi 0004, Linna Zhang, Gaoyun An, Yi Jin 0001, Yidong Li, Yi-Gang Cen |
ICME | 4 |
| 2025 | VCF: An effective Vision-Centric Framework for Visual Question Answering
Longkun Peng, Shan Cao 0002, Zhaoqilin Yang, Gaoyun An |
Neurocomputing | 6 |
| 2025 | EIRA: an explicit-implicit representation alignment for multimodal relation extraction
Gaoyun An, Zhaoqilin Yang, Xingyu Ren, Qiuqi Ruan |
Multim. Syst. | 2 |
| 2025 | Attention redirection transformer with semantic oriented learning for unbiased scene graph generation
Gaoyun An, Yi-Gang Cen, Qiuqi Ruan |
Pattern Recognit. | 2 |
| 2025 | Cross-scene visual context parsing with large vision-language modelabstractRelation analysis is crucial for image-based applications such as visual reasoning and visual question answering . Current relation analysis such as scene graph generation (SGG) only focuses on building relationships among objects within a single image. However, in real-world applications, relationships among objects across multiple images, as seen in video understanding , may hold greater significance as they can capture global information. This is still a challenging and unexplored task. In this paper, we aim to explore the technique of Cross-Scene Visual Context Parsing (CS-VCP) using a large vision-language model. To achieve this, we first introduce a cross-scene dataset comprising 10,000 pairs of cross-scene visual instruction data, with each instruction describing the common knowledge of a pair of cross-scene images. We then propose a Cross-Scene Visual Symbiotic Linkage (CS-VSL) model to understand both cross-scene relationships and objects by analyzing the rationales in each scene. The model is pre-trained on 100,000 cross-scene image pairs and validated on 10,000 image pairs. Both quantitative and qualitative experiments demonstrate the effectiveness of the proposed method. Our method has been released on GitHub: https://github.com/gavin-gqzhang/CS-VSL . Shichao Kan, Lu Shi 0004, Wanru Xu, Gaoyun An, Yi-Gang Cen |
Pattern Recognit. | 5 |
| 2025 | Human motion prediction model with decoupled behavior and pose
Langtian Lang, Yinshuo Sun, Bowen Ni, Gaoyun An |
Pattern Recognit. Lett. | 5 |
| 2024 | Synergetic Prototype Learning Network for Unbiased Scene Graph GenerationabstractScene Graph Generation (SGG) is an important cross-modal task in scene understanding, aiming to detect visual relations in an image. However, due to the various appearance features, the feature distributions of different categories have suffered from a severe overlap, which makes the decision boundaries ambiguous. The current SGG methods mainly attempt to re-balance the data distribution, which is dataset-dependent and limits the generalization. To solve this problem, a Synergetic Prototype Learning Network (SPLN) is proposed here, where the generalized semantic space is modeled and the synergetic effect among different semantic subspaces is delved into. In SPLN, a Collaboration-induced Prototype Learning method is proposed to model the interaction of visual semantics and structural semantics. The conventional visual semantics is focused on with a residual-driven representation enhancement module to capture details. And the intersection of structural semantics and visual semantics is explicitly modeled as conceptual semantics, which has been ignored by existing methods. Meanwhile, to alleviate the noise of unrelated and meaningless words, an Intersection-induced Prototype Learning method is also proposed specially for conceptual semantics with an essence-driven prototype enhancement module. Moreover, a Selective Fusion Module is proposed to synergetically integrate the results of visual, structural, conceptual branches and the generalized semantics projection. Experiments on VG and GQA datasets show that our method achieves state-of-the-art performance on the unbiased metrics. Ziwei Shang, Zhaoqilin Yang, Shan Cao 0002, Yi-Gang Cen, Gaoyun An |
ACM Multimedia | 7 |
| 2024 | EPK-CLIP: External and Priori Knowledge CLIP for action recognition
Zhaoqilin Yang, Gaoyun An, Zhenxing Zheng, Shan Cao 0002 |
Expert Syst. Appl. | 2 |
| 2024 | Facial action unit detection with emotion consistency: a cross-modal learning approach
Wenyu Song, Dongxin Liu, Gaoyun An, Yun Duan, Laifu Wang |
Multim. Syst. | 3 |
| 2024 | Bridging Visual and Textual Semantics: Towards Consistency for Unbiased Scene Graph GenerationabstractScene Graph Generation (SGG) aims to detect visual relationships in an image. However, due to long-tailed bias, SGG is far from practical. Most methods depend heavily on the assistance of statistics co-occurrence to generate a balanced dataset, so they are dataset-specific and easily affected by noises. The fundamental cause is that SGG is simplified as a classification task instead of a reasoning task, thus the ability capturing the fine-grained details is limited and the difficulty in handling ambiguity is increased. By imitating the way of dual process in cognitive psychology, a Visual-Textual Semantics Consistency Network (VTSCN) is proposed to model the SGG task as a reasoning process, and relieve the long-tailed bias significantly. In VTSCN, as the rapid autonomous process (Type1 process), we design a Hybrid Union Representation (HUR) module, which is divided into two steps for spatial awareness and working memories modeling. In addition, as the higher order reasoning process (Type2 process), a Global Textual Semantics Modeling (GTS) module is designed to individually model the textual contexts with the word embeddings of pairwise objects. As the final associative process of cognition, a Heterogeneous Semantics Consistency (HSC) module is designed to balance the type1 process and the type2 process. Lastly, our VTSCN raises a new way for SGG model design by fully considering human cognitive process. Experiments on Visual Genome, GQA and PSG datasets show our method is superior to state-of-the-art methods, and ablation studies validate the effectiveness of our VTSCN. Gaoyun An, Yiqing Hao, Dapeng Oliver Wu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | CAST: Cross-Modal Retrieval and Visual Conditioning for image captioning
Shan Cao 0002, Gaoyun An, Yi-Gang Cen, Zhaoqilin Yang, Weisi Lin |
Pattern Recognit. | 2 |
| 2024 | GBC: Guided Alignment and Adaptive Boosting CLIP Bridging Vision and Language for Robust Action RecognitionabstractThe Contrastive Language-Image Pre-training (CLIP) model achieves strong generalization by using a large number of text-image pairs for contrastive learning. However, when it is transferred to action recognition, the following two questions remain to be solved: 1) How to guide the model to focus more on human-body-related regions to better align actions and text, and 2) How to make the model strengthen itself in a targeted manner to deal with difficult-to-classify categories. To solve these problems, a Guided alignment and adaptive Boosting CLIP (GBC) is proposed, which employs visual prior knowledge and benefits from both feature and decision aggregation in a boosting manner. During early training, visual prior knowledge related to human body is adopted, which enables the model to better align human actions with category text to be robust to distribution shift. At the later stage of training, the CLIP encoder is frozen, and multiple downstream feature & decision aggregation modules are sequentially generated and trained. In such way, the model is able to boost the performance from different perspectives in the Boosting manner and at a linearly increasing cost. Moreover, a class-adaptive re-weighting strategy is proposed to make the model focus more on optimizing categories that are difficult to classify. The effectiveness of our model is validated on six action recognition datasets (Kinetics-600, Kinetics-400, Jester, HMDB-51, UCF-101, and Mini-Kinetics-200), including both fully supervised and zero-shot experiments. Our model achieves superior results compared to state-of-the-art methods on all datasets. Zhaoqilin Yang, Gaoyun An, Zhenxing Zheng, Shan Cao 0002, Qiuqi Ruan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Multi-step Prediction of LTE-R Communication Quality based on CA-TCN and Differential EvolutionabstractWith the continuous development of heavy-haul railway technology, the demand for the faster transmission and higher bandwidth capacity of wireless communication network is also increasing. Therefore, Long Term Evolution for Railway (LTE-R) begin to replace Global System for Mobile Communications for Railway (GSM-R) to burden the core wireless communication services. In order to improve the operation and maintenance efficiency of LTE-R network, this paper proposes a multi-step LTE-R communication quality prediction method based on Differential Evolution algorithm and Temporal Convolutional Network with Coordinate Attention (TCNCA). Firstly, this method uses differential evolution and permutation importance index to filter the features of multivariate LTE-R communication quality data. Then, the coordinate attention mechanism and TCN network were fused to build the prediction model. Finally, this method still used Differential Evolution to adjust the model parameters based on Mann—Kendall test, so that the model could take into account the prediction accuracy and the trend of data change, thus realize the multi-step prediction for LTE-R communication quality. Experimental results on real data set show that the proposed method can provide decision support for the active maintenance of LTE-R network, and has high application value. Jiantao Qu, Chunyu Qi, Gaoyun An, He |
TrustCom | 3 |
| 2023 | A dual-modal graph attention interaction network for person Re-identificationabstractAbstract Person Re‐identification (Re‐ID) is a task of matching target pedestrians under cross‐camera surveillance. Learning discriminative feature representations is the main issue for person Re‐ID. A few recent methods introduce text descriptions as auxiliary information to enhance feature representations, as it offers richer semantic information and perspective consistency. However, these works usually process text and images separately, which leads to the absence of cross‐modal interactions. In this article, a Dual‐modal Graph Attention Interaction Network (Dual‐GAIN) is proposed to integrate visual features and textual features into a heterogeneous graph to model the relationship between them, simultaneously. The proposed Dual‐GAIN mainly consists of two components: a dual‐stream feature extractor and a Graph Attention Interaction Network (GAIN). Specifically, the two‐stream feature extractor is utilised to extract visual features and textual features respectively. Then, visual local features and textual features are treated as nodes to construct a multi‐modal graph. Cosine similarity constrained attention weights are introduced in GAIN, which is designed for cross‐modal interaction and feature fusion on this heterogeneous multi‐modal graph. Experiments on public large‐scale datasets, that is, Market‐1501, CUHK03 labelled, and CUHK03 detected, demonstrate our method achieves the state‐of‐the‐art performance. Gaoyun An, Qiuqi Ruan |
IET Comput. Vis. | 2 |
| 2023 | SRI3D: Two-stream inflated 3D ConvNet based on sparse regularization for action recognitionabstractAbstract Although most state‐of‐the‐art action recognition models have adopted a two‐stream 3D convolutional structure as a backbone network, few works have studied the impact of loss functions on action recognition models. In addition, sparsity is used as a key prior knowledge in many fields. However, as far as is known, no one has studied the influence of the sparsity of network output on the output of deep learning‐based action recognition models. Therefore, this paper proposes a novel two‐stream inflated 3D ConvNet based on the sparse regularization (SRI3D) model for action recognition. In order to allow the network to learn the sparsity of output, the ℓ 1 norm is embedded in the loss function in regularization form in a plug‐and‐play manner. It can make the classification result after the fusion of the two‐stream network only be the category with the highest confidence in one of the streams and not the other cases. The proposed loss function based on sparse regularization makes the output vector of the neural network as sparse as possible so that the classification results will not be ambiguous. Experimental results show that compared with other state‐of‐the‐art models, this SRI3D has a competitive advantage on Kinetics‐400, Something‐Something V2, UCF‐101 and HMDB‐51. Zhaoqilin Yang, Gaoyun An, Zhenxing Zheng, Qiuqi Ruan |
IET Image Process. | 2 |
| 2023 | Collaborative and Multilevel Feature Selection Network for Action RecognitionabstractThe feature pyramid has been widely used in many visual tasks, such as fine-grained image classification, instance segmentation, and object detection, and had been achieving promising performance. Although many algorithms exploit different-level features to construct the feature pyramid, they usually treat them equally and do not make an in-depth investigation on the inherent complementary advantages of different-level features. In this article, to learn a pyramid feature with the robust representational ability for action recognition, we propose a novel collaborative and multilevel feature selection network (FSNet) that applies feature selection and aggregation on multilevel features according to action context. Unlike previous works that learn the pattern of frame appearance by enhancing spatial encoding, the proposed network consists of the position selection module and channel selection module that can adaptively aggregate multilevel features into a new informative feature from both position and channel dimensions. The position selection module integrates the vectors at the same spatial location across multilevel features with positionwise attention. Similarly, the channel selection module selectively aggregates the channel maps at the same channel location across multilevel features with channelwise attention. Positionwise features with different receptive fields and channelwise features with different pattern-specific responses are emphasized respectively depending on their correlations to actions, which are fused as a new informative feature for action recognition. The proposed FSNet can be inserted into different backbone networks flexibly, and extensive experiments are conducted on three benchmark action datasets, Kinetics, UCF101, and HMDB51. Experimental results show that FSNet is practical and can be collaboratively trained to boost the representational ability of existing networks. FSNet achieves superior performance against most top-tier models on Kinetics and all models on UCF101 and HMDB51. Zhenxing Zheng, Gaoyun An, Shan Cao 0002, Dapeng Oliver Wu, Qiuqi Ruan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Causal Property Based Anti-conflict Modeling with Hybrid Data Augmentation for Unbiased Scene Graph Generation
Gaoyun An |
ACCV (4) | 2 |
| 2022 | PromptLearner-CLIP: Contrastive Multi-Modal Action Representation Learning with Context Optimization
Zhenxing Zheng, Gaoyun An, Shan Cao 0002, Zhaoqilin Yang, Qiuqi Ruan |
ACCV (4) | 2 |
| 2022 | Dual-attention guided network for facial action unit detectionabstractAbstract Attention mechanism has recently aroused increasing concerns in the field of computer vision like Action Unit (AU) detection. Because facial AU exists in a fixed local area of a human face, it is advantageous to apply the attention mechanism to AU detection. A Dual‐Attention Guided Network (DAGNet) is proposed for automatically AU detection, which introduces dual attention to selectively extract deep features. Dual attention refers to predefined explicitly models feature dependencies from spatial and channel attention, respectively, based on the semantics of AU label. In addition, since the global and local features show different facial attributes and supplement mutually, the proposed DAGNet learns feature representations from global and local perspectives, respectively. Learning global and local features simultaneously during training can lead to better generalization performance; a fusion module is designed for aggregating all the learned information to construct a unified architecture for end‐to‐end AU detection. Extensive experiments on two challenging datasets, BP4D and DISFA, result in F1‐scores of 64.0% and 62.6%, respectively, which shows that the proposed DAGNet achieves the performance of the state‐of‐the‐art in the field of image‐based AU detection. Wenyu Song, Shuze Shi, Gaoyun An |
IET Image Process. | 4 |
| 2022 | Heterogeneous spatio-temporal relation learning network for facial action unit detection
Wenyu Song, Shuze Shi, Gaoyun An |
Pattern Recognit. Lett. | 4 |
| 2022 | Vision-Enhanced and Consensus-Aware Transformer for Image CaptioningabstractImage captioning generates descriptions in a natural language for a given image. Due to its great potential for a wide range of applications, many deep learning based-methods have been proposed. The co-occurrence of words such as mouse and keyboard, constitutes commonsense knowledge, which is referred to as consensus. However, it is challenging to consider commonsense knowledge in producing captions that have rich, natural, and meaningful semantics. In this paper, a Vision-enhanced and Consensus-aware Transformer (VCT) is proposed to exploit both visual information and consensus knowledge for image captioning with three key components: a vision-enhanced encoder, consensus-aware knowledge representation generator, and consensus-aware decoder. The vision-enhanced encoder extends the vanilla self-attention module with a memory-based attention module and a visual perception module for learning better visual representation of an image. Specifically, the relationships between regions in an image and the image’s global context are leveraged with scene memory in the memory-based attention module. The visual perception module further enhances the correlation among neighboring tokens in both the spatial and channel-wise dimensions. To learn consensus-aware representations, a word correlation graph is constructed by computing the statistical co-occurrence between semantic concepts. Then consensus knowledge can be acquired using a graph convolutional network in the consensus-aware knowledge representation generator. Finally, such consensus knowledge is integrated into the consensus-aware decoder through consensus memory and a knowledge-based control module to produce a caption. Experimental results on two popular benchmark datasets (MSCOCO and Flickr30k) demonstrate that our proposed model achieves state-of-the-art performance. Extensive ablation studies also validate the effectiveness of each component. Shan Cao 0002, Gaoyun An, Zhenxing Zheng, Zhiyong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Facial Action Unit Detection Based on Transformer and Attention Mechanism
Wenyu Song, Shuze Shi, Gaoyun An |
ICIG (2) | 3 |
| 2021 | Global and Local Knowledge-Aware Attention Network for Action RecognitionabstractConvolutional neural networks (CNNs) have shown an effective way to learn spatiotemporal representation for action recognition in videos. However, most traditional action recognition algorithms do not employ the attention mechanism to focus on essential parts of video frames that are relevant to the action. In this article, we propose a novel global and local knowledge-aware attention network to address this challenge for action recognition. The proposed network incorporates two types of attention mechanism called statistic-based attention (SA) and learning-based attention (LA) to attach higher importance to the crucial elements in each video frame. As global pooling (GP) models capture global information, while attention models focus on the significant details to make full use of their implicit complementary advantages, our network adopts a three-stream architecture, including two attention streams and a GP stream. Each attention stream employs a fusion layer to combine global and local information and produces composite features. Furthermore, global-attention (GA) regularization is proposed to guide two attention streams to better model dynamics of composite features with the reference to the global information. Fusion at the softmax layer is adopted to make better use of the implicit complementary advantages between SA, LA, and GP streams and get the final comprehensive predictions. The proposed network is trained in an end-to-end fashion and learns efficient video-level features both spatially and temporally. Extensive experiments are conducted on three challenging benchmarks, Kinetics, HMDB51, and UCF101, and experimental results demonstrate that the proposed network outperforms most state-of-the-art methods. Zhenxing Zheng, Gaoyun An, Dapeng Oliver Wu, Qiuqi Ruan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Subgraph and object context-masked network for scene graph generationabstractScene graph generation is to recognise objects and their semantic relationships in an image and can help computers understand visual scene. To improve relationship prediction, geometry information is essential and usually incorporated into relationship features. Existing methods use coordinates of objects to encode their spatial layout. However, in this way, they neglect the context of objects. In this study, to take full use of spatial knowledge efficiently, the authors propose a novel subgraph and object context‐masked network (SOCNet) consisting of spatial mask relation inference (SMRI) and hierarchical message passing (HMP) modules to address the scene graph generation task. In particular, to take advantage of spatial knowledge, SMRI masks partial context of object features depending on their spatial layout of objects and corresponding subgraph to facilitate their relationship recognition. To refine the features of objects and subgraphs, they also propose HMP that passes highly correlated messages from both microcosmic and macroscopic aspects through a triple‐path structure including subgraph–subgraph, object–object, and subgraph–object paths. Finally, statistical co‐occurrence probability is used to regularise relationship prediction. SOCNet integrates HMP and SMRI into a unified network, and comprehensive experiments on visual relationship detection and visual genome datasets indicate that SOCNet outperforms several state‐of‐the‐art methods on two common tasks. Zhenxing Zheng, Gaoyun An, Songhe Feng |
IET Comput. Vis. | 3 |
| 2020 | E2-capsule neural networks for facial expression recognition using AU-aware attentionabstractCapsule neural network is a new and popular technique in deep learning. However, the traditional capsule neural network does not extract features sufficiently before the dynamic routing between capsules. In this study, one double enhanced capsule neural network (E2‐Capsnet) that uses AU‐aware attention for facial expression recognition (FER) is proposed. The E2‐Capsnet takes advantage of dynamic routing between capsules and has two enhancement modules which are beneficial to FER. The first enhancement module is the convolutional neural network with AU‐aware attention, which can focus on the active areas of the expression. The second enhancement module is the capsule neural network with multiple convolutional layers, which enhances the ability of the feature representation. Finally, the squashing function is used to classify the facial expression. The authors demonstrate the effectiveness of E2‐Capsnet on the two public benchmark datasets, RAF‐DB and EmotioNet. The experimental results show that their E2‐Capsnet is superior to the state‐of‐the‐art methods. The code is available at https://github.com/ShanCao18/E2‐Capsnet . Shan Cao 0002, Yuqian Yao, Gaoyun An |
IET Image Process. | 3 |
| 2020 | Interactions Guided Generative Adversarial Network for unsupervised image captioning
Shan Cao 0002, Gaoyun An, Zhenxing Zheng, Qiuqi Ruan |
Neurocomputing | 2 |
| 2020 | View-specific subspace learning and re-ranking for semi-supervised person re-identification
Jieru Jia, Qiuqi Ruan, Yi Jin 0001, Gaoyun An, Shiming Ge |
Pattern Recognit. | 4 |
| 2020 | Aligned Dynamic-Preserving Embedding for Zero-Shot Action RecognitionabstractZero-shot learning (ZSL) typically explores a shared semantic space in order to recognize novel categories in the absence of any labeled training data. However, the traditional ZSL methods always suffer from serious domain shift problem in human action recognition. This is because: 1) existing ZSL methods are specifically designed for object recognition from static images, which do not capture the temporal dynamics of video sequences, and poor performances are always generated if those methods are directly applied to zero-shot action recognition; 2) these methods always blindly project the target data into a shared space using a semantic mapping obtained by the source data without any adaptation, in which the underlying structures of target data are ignored; and 3) severe inter-class variations exist in various action categories. The traditional ZSL methods do not take relationships across different categories into consideration. In this paper, we propose a novel aligned dynamic-preserving embedding (ADPE) model for zero-shot action recognition in a transductive setting. In our model, an adaptive embedding of target videos is learned, exploring the distributions of both the source and target data. An aligned regularization is further proposed to couple the centers of target semantic representations with their corresponding label prototypes in order to preserve the relationships across different categories. Most significantly, during our embedding, the temporal dynamics of video sequences are simultaneously preserved via exploiting the temporal consistency of video sequences and capturing the temporal evolution of successive segments of actions. Our model can effectively overcome the domain shift problem in zero-shot action recognition. The experiments on Olympic sports, HMDB51, and UCF101 datasets demonstrate the effectiveness of our model. Yu Kong 0001, Qiuqi Ruan, Gaoyun An, Yun Fu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Residual Joint Attention Network with Graph Structure Inference for Object Detection
Chuansheng Xu, Gaoyun An, Qiuqi Ruan |
ICIG (1) | 2 |
| 2019 | Semi-supervised feature selection analysis with structured multi-view sparse regularization
Caijuan Shi, Changyu Duan, Zhibin Gu, Qi Tian 0001, Gaoyun An, Ruizhen Zhao |
Neurocomputing | 5 |
| 2019 | Spatial-temporal pyramid based Convolutional Neural Network for action recognition
Zhenxing Zheng, Gaoyun An, Dapeng Oliver Wu, Qiuqi Ruan |
Neurocomputing | 2 |
| 2019 | Deep spectral feature pyramid in the frequency domain for long-term action recognition
Gaoyun An, Zhenxing Zheng, Dapeng Oliver Wu |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | FERLrTc: 2D+3D facial expression recognition via low-rank tensor completion
Yunfang Fu, Qiuqi Ruan, Ziyan Luo, Yi Jin 0001, Gaoyun An, Jun Wan 0001 |
Signal Process. | 5 |
| 2019 | Robust discriminant low-rank representation for subspace clustering
Gaoyun An, Yi-Gang Cen, Hengyou Wang, Ruizhen Zhao |
Soft Comput. | 2 |
| 2018 | Hierarchical and Spatio-Temporal Sparse Representation for Human Action RecognitionabstractIn this paper, we present a novel two-layer video representation for human action recognition employing hierarchical group sparse encoding technique and spatio-temporal structure. In the first layer, a new sparse encoding method named locally consistent group sparse coding (LCGSC) is proposed to make full use of motion and appearance information of local features. LCGSC method not only encodes global layouts of features within the same video-level groups, but also captures local correlations between them, which obtains expressive sparse representations of video sequences. Meanwhile, two kinds of efficient location estimation models, namely an absolute location model and a relative location model, are developed to incorporate spatio-temporal structure into LCGSC representations. In the second layer, action-level group is established, where a hierarchical LCGSC encoding scheme is applied to describe videos at different levels of abstractions. On the one hand, the new layer captures higher order dependency between video sequences; on the other hand, it takes label information into consideration to improve discrimination of videos' representations. The superiorities of our hierarchical framework are demonstrated on several challenging datasets. Yu Kong 0001, Qiuqi Ruan, Gaoyun An, Yun Fu 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Multiple metric learning with query adaptive weights and multi-task re-weighting for person re-identification
Jieru Jia, Qiuqi Ruan, Gaoyun An, Yi Jin 0001 |
Comput. Vis. Image Underst. | 3 |
| 2017 | A sparse neighborhood preserving non-negative tensor factorization algorithm for facial expression recognition
Gaoyun An, Shuai Liu 0003, Qiuqi Ruan |
Pattern Anal. Appl. | 1 |
| 2017 | Multiview Hessian Semisupervised Sparse Feature Selection for Multimedia AnalysisabstractFacing a large number of unlabeled data and a small number of labeled data, semisupervised sparse feature selection has received increasing attention in recent years. However, most semisupervised feature selection algorithms are developed for single-view data and cannot naturally handle multiview data. Moreover, most existing semisupervised sparse feature selection methods are based on Laplacian regularization, which is a lack of extrapolating power. To overcome the above-mentioned drawbacks, we present a multiview Hessian semi-supervised sparse feature selection (MHSFS) framework in this paper. MHSFS can directly accomplish multiview sparse feature selection by exploiting multiview learning to reveal and leverage the correlated and complemental information among different views. In addition, MHSFS can achieve better performance based on Hessian regularization, which favors functions whose values linearly vary with respect to geodesic distance and preserves the local manifold structure well. A simple yet efficient iterative method is proposed to solve the objective function, followed by convergence analysis. We apply the proposed method into different multimedia analysis tasks, such as image annotation, video concept detection, and 3D motion analysis. The results show that MHSFS outperforms the state-of-the-art sparse feature selection methods and achieves good performance. Caijuan Shi, Gaoyun An, Ruizhen Zhao, Qiuqi Ruan, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Action Recognition Using Local Consistent Group Sparse Coding with Spatio-Temporal StructureabstractThis paper presents a novel and efficient framework for human action recognition through integrating the local consistent group sparse representation with spatio-temporal structure of each video sequence. We firstly propose a sparse encoding scheme named local consistent group sparse coding (LCGSC) to generate the sparse representation of each video sequence. The novel encoding scheme takes global structural information of features belonging to one group into consideration as well as the local correlations between similar features. In order to incorporate the spatio-temporal structures, an average location (AL) model is proposed to describe the distribution of each visual word along the spatio-temporal coordinates on the basis of the obtained sparse codes. Eventually, each video sequence is jointly represented by the sparse representation and the spatio-temporal layouts which fully model its motion, appearance and spatio-temporal information. Our framework is computationally efficient and achieves comparable performance on the challenging datasets with state-of-the-art methods. Qiuqi Ruan, Gaoyun An, Yun Fu 0001 |
ACM Multimedia | 3 |
| 2016 | Facial expression recognition using sparse local Fisher discriminant analysis
Qiuqi Ruan, Gaoyun An |
Neurocomputing | 3 |
| 2015 | Multiple strategies to enhance automatic 3D facial expression recognition
Xiaoli Li 0009, Qiuqi Ruan, Gaoyun An, Yi Jin 0001, Ruizhen Zhao |
Neurocomputing | 3 |
| 2015 | Context and locality constrained linear coding for human action recognition
Qiuqi Ruan, Gaoyun An, Wanru Xu |
Neurocomputing | 3 |
| 2015 | Semi-supervised sparse feature selection based on multi-view Laplacian regularization
Caijuan Shi, Qiuqi Ruan, Gaoyun An |
Image Vis. Comput. | 3 |
| 2015 | Projection-optimal local Fisher discriminant analysis for feature extraction
Qiuqi Ruan, Gaoyun An |
Neural Comput. Appl. | 3 |
| 2015 | Fully automatic 3D facial expression recognition using polytypic multi-block local binary patterns
Xiaoli Li 0009, Qiuqi Ruan, Yi Jin 0001, Gaoyun An, Ruizhen Zhao |
Signal Process. | 4 |
| 2015 | Hessian Semi-Supervised Sparse Feature Selection Based on ${L_{2, 1/2}}$ -Matrix NormabstractSemi-supervised sparse feature selection, which can exploit the small number labeled data and large number unlabeled data simultaneously, has become an important technique in many applications on large-scale web image owing to its high efficiency and effectiveness. Recently, graph Laplacian-based semi-supervised sparse feature selection has obtained considerable attention, but it suffers with only few labeled data because Laplacian regularization is short of extrapolating power. In this paper we propose a novel semi-supervised sparse feature selection framework based on Hessian regularization and l2,1/2- matrix norm, namely Hessian sparse feature selection based on L2,1/2- matrix norm (HFSL). Hessian regularization favors functions whose values vary linearly with respect to geodesic distance and preserves the local manifold structure well, leading to good extrapolating power to boost semi-supervised learning, and then to enhance HFSL performance. The l2,1/2-matrix norm model makes HFSL select the most discriminative sparse features with good robustness. An efficient iterative algorithm is designed to optimize the objective function. We apply our algorithm into the image annotation task and conduct extensive experiments on two web image datasets. The results demonstrate that our algorithm outperforms state-of-the-art sparse feature selection methods and is promising for large-scale web image applications. Caijuan Shi, Qiuqi Ruan, Gaoyun An, Ruizhen Zhao |
IEEE Trans. Multim. | 3 |
| 2014 | Sparse feature selection based on graph Laplacian for web image annotation
Caijuan Shi, Qiuqi Ruan, Gaoyun An |
Image Vis. Comput. | 3 |
| 2012 | Tensor rank one differential graph preserving analysis for facial expression recognition
Shuai Liu 0003, Qiuqi Ruan, Chuantao Wang, Gaoyun An |
Image Vis. Comput. | 4 |
| 2010 | An illumination normalization model for face recognition under varied lighting conditions
Gaoyun An, Jiying Wu, Qiuqi Ruan |
Pattern Recognit. Lett. | 1 |
| 2009 | Independent Gabor Analysis of Discriminant Features Fusion for Face RecognitionabstractA discriminant feature fusion model is proposed for face recognition with large variations of pose, expression, lighting, etc. Discriminant features are extracted by the wavelet transform-based method from two source images. One source image is a holistic gray value image and the other is an illumination invariant geometric image. Face sample is reconstructed by the adaptive fused discriminant feature. Then a bank of Gabor filters is built to extract Gabor representations of the reconstructed samples. Finally higher-order statistical relationships among variables of samples are extracted for classifier. According to experiments, the model outperforms conventional algorithms under complex conditions (large variations of lighting, expression, accessory, etc.). Jiying Wu, Gaoyun An, Qiuqi Ruan |
IEEE Signal Process. Lett. | 2 |
| 2008 | Gabor-based multi-scale Illumination Normalization model for face recognitionabstractA novel Gabor-based multi-scale illumination normalization (GMSIN) model is proposed and applied to face recognition. GMSIN uses Total variation under different norm constraints. It removes the lighting effect in two scale parts of image and fuses the multi-scaled illumination invariant features. Then a bank of Gabor filters is built to extract lighting invariant Gabor face representations. Finally the higher-order statistical relationships among variables of samples are extracted for classifier. According to the experiments on the large scale CAS-PEAL face database, GMSIN could outperform conventional algorithms when they face most outliers (lighting, expression, masking etc.). Jiying Wu, Gaoyun An, Qiuqi Ruan |
ICIP | 2 |
| 2008 | Independent Gabor Analysis of Multiscale Total Variation-Based Quotient ImageabstractA new algorithm for independent Gabor analysis of multiscale total variation-based quotient image is proposed and applied to face recognition with only one sample per subject here. With our preproposed multiscale TV-based quotient image (TVQI) model, the large-scale and small-scale features are firstly fused to produce the most expressive lighting invariant face. Then a bank of Gabor filters is built to extract lighting invariant Gabor face representations with specified scales and orientations. Last, an information maximization algorithm is adopted to extract higher-order statistical relationships among variables of samples for classifier. According to the experiments on the large-scale CAS-PEAL face database, our approach could outperform Gabor-based ICA, Gabor-based KPCA, and TVQI when they face most outliers (lighting, expression, masking, etc.). Gaoyun An, Jiying Wu, Qiuqi Ruan |
IEEE Signal Process. Lett. | 1 |
| 2007 | A Novel Image Interpolation Method Based on Both Local and Global Information
Jiying Wu, Qiuqi Ruan, Gaoyun An |
ICIC (1) | 3 |
| 2006 | A Novel Model for Gabor-Based Independent Radial Basis Function Neural Networks and Its Application to Face Recognition
Gaoyun An, Qiuqi Ruan |
ICONIP (2) | 1 |