VLDB 2026 Research / reviewers in the wild / expert
Guixian Zhang
dblp:311/8752
· DBLP profile ↗
33ranked-venue papers
11as first author
33since 2021 · last 2027
0000-0002-7632-8411ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 4 first-author · 19 since 2021Databases, data management, data science and information retrieval · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Beyond conceptual path bias: A unified hierarchical framework for learning path recommendation in online educational systems
Guan Yuan, Guixian Zhang, Shang Liu 0001 |
Expert Syst. Appl. | 3 |
| 2026 | Noise-Aware Graph-Based Cognitive Diagnostic Framework Through Low-Rank AlignmentabstractGraph Neural Networks (GNNs) have effectively improved the performance of Cognitive Diagnosis Models (CDMs). Existing works have proposed a series of Graph-based Cognitive Diagnosis Frameworks (GCDFs) to enhance robustness to noise. However, these robust designs are often general methods for GNNs and are not designed for cognitive diagnosis, which undermines real cognitive information during the denoising process. Interestingly, a noteworthy phenomenon has been overlooked: even without robustness designs, GCDFs can still learn correct information in noisy environments. In this paper, we conduct a comprehensive empirical analysis of this issue. We found that noise primarily accumulates in lower singular components. Even in noisy environments, the principal subspaces of representations still remain stable. Based on these findings, we propose a Noise-aware Cognitive Diagnostic framework based on Low-rank Alignment, named NCDLA. The framework first performs low-rank reconstruction of the interaction matrix between students and exercises, retaining only larger singular values to achieve noise reduction. Then, the reconstructed interaction matrix and the original interaction matrix are combined with the Q matrix to form a noise-reduced heterogeneous graph and an original heterogeneous graph. In order to distinguish between the interaction patterns of correct and incorrect responses, we decompose the heterogeneous graph according to the type of response. NCDLA achieves denoising of student representations and exercises representations through a self-supervised strategy based on low-rank reconstruction and a spectral anchor regularisation method. Extensive experiments on three datasets demonstrate that NCDLA achieves optimal prediction performance and robustness. Guixian Zhang, Guan Yuan, Shang Liu 0001, Xiaojing Du, Debo Cheng |
AAAI | 1 |
| 2026 | Harnessing LLM for Noise-Robust Cognitive Diagnosis in Web-Based Intelligent Education SystemsabstractCognitive diagnostics in the Web-based Intelligent Education System (WIES) aims to assess students' mastery of knowledge concepts from heterogeneous, noisy interactions. Recent work has tried to utilize Large Language Models (LLMs) for cognitive diagnosis, yet LLMs struggle with structured data and are prone to noise-induced misjudgments. Specially, WIES's open environment continuously attracts new students and produces vast amounts of response logs, exacerbating the data imbalance and noise issues inherent in traditional educational systems. To address these challenges, we propose DLLM, a Diffusion-based LLM framework for noise-robust cognitive diagnosis. DLLM first constructs independent subgraphs based on response correctness, then applies relation augmentation alignment module to mitigate data imbalance. The two subgraph representations are then fused and aligned with LLM-derived, semantically augmented representations. Importantly, before each alignment step, DLLM employs a two-stage denoising diffusion module to eliminate intrinsic noise while assisting structural representation alignment. Specifically, unconditional denoising diffusion first removes erroneous information, followed by conditional denoising diffusion based on graph signal to eliminate misleading information. Finally, the noise-robust representation that integrates semantic knowledge and structural information is fed into existing cognitive diagnosis models for prediction. Experimental results on three publicly available web-based educational platform datasets demonstrate that our DLLM achieves optimal predictive performance across varying noise levels, which demonstrates that DLLM achieves noise robustness while effectively leveraging semantic knowledge from LLM. Guixian Zhang, Guan Yuan, Ziqi Xu 0001, Jing Ren 0001, Zhenyun Deng, Debo Cheng |
WWW | 1 |
| 2026 | A multi-view graph neural network with subgraph variational autoencoder for class-Imbalanced node classification
Longqing Du, Zhirong Huang, Jiecheng Li, Guixian Zhang, Debo Cheng, Guangquan Lu, Shichao Zhang 0001 |
Knowl. Based Syst. | 4 |
| 2026 | Learning instrumental variable representation for debiasing in recommender systemsabstractRecommender systems are essential for filtering content to match user preferences. However, traditional recommender systems often suffer from biases inherent in the data, such as popularity bias. These biases, particularly those stemming from latent confounders, can result in inaccurate recommendations and reduce both the diversity and effectiveness of the system. Existing debiasing methods for recommender systems, however, either fail to account for latent confounders or rely on predefined instrumental variables (IVs). To address this research gap, we propose a novel causality-based recommendation algorithm, Data-driven IV representation learning for debiasing in Recommender System (DIVRS), which enables the learning of IV representation directly from user-item interaction data. By leveraging the learned IV representation, DIVRS decomposes user behaviour into causal and confounding relationships to address potential bias in recommender systems. Additionally, we introduce Orthogonal Promotion Regularisation (OPR) for DIVRS to address the problem that Graph Convolutional Networks (GCNs) amplify bias. We also propose a variant of GCNs for DIVRS, called DIVRS-GCN. Experimental results on the Douban-Movie and Movielens-10M datasets demonstrate that both DIVRS and DIVRS-GCN effectively mitigate confounding bias while outperform the state-of-the-art methods in recommendation performance. For example, on both datasets, our DIVRS and DIVRS-GCN improve Recall@20 by up to 10.98 %. This validates their effectiveness and robustness. Our approaches improve recommendation accuracy while delivering more balanced and diverse suggestions, effectively addressing the limitations of existing IV-based recommender systems. Zhirong Huang, Shichao Zhang 0001, Debo Cheng, Jiuyong Li, Lin Liu 0003, Guangquan Lu, Guixian Zhang |
Neural Networks | 7 |
| 2026 | Toward fair graph neural networks via dual-teacher knowledge distillation
Debo Cheng, Guixian Zhang, Yi Li 0083, Shichao Zhang 0001 |
Neural Networks | 3 |
| 2026 | Fairness-aware graph learning via adaptive manifold structure and representation alignment
Boyan Chen, Guixian Zhang, Jinyi Jie, Debo Cheng, Shichao Zhang 0001 |
Pattern Recognit. | 2 |
| 2026 | Graph Unlearning System with Subgraph De-Isolation MeasuresabstractGraph unlearning system offers a promising solution for securely erasing specific data points and their associated influences from Graph Neural Networks (GNNs). However, existing approaches often treat the problem as multiple isolated and disjoint sub-problems by partitioning graph data into isolated subgraphs, which overlooks the native graph structure information between subgraphs. This results in biased representations that hinder the accurate modeling of key connections and relationships within the data, leading to a notable reduction in model utility due to this loss of information. To address these issues, we propose an innovative framework called N on- I solated G raph Eraser (NIGEraser) that decomposes the unlearning task into multiple non-isolated, intersecting sub-problems. Specifically, a novel non-isolated graph partitioning strategy is proposed for NIGEraser that mitigates isolation by replicating key nodes across multiple neighboring subgraphs, along with an attention-based sub-model aggregation technique in that global graph structure information is employed. By this design, a broader natural neighborhood is explored, capturing and effectively utilizing the critical graph structure features lost between subgraphs during partitioning, thereby reducing information loss during task decomposition and aggregation. Additionally, it is demonstrated that graph unlearning methods can overcome the limitations of traditional isolated partitioning strategies, providing an effective theoretical constraint on time consumption. Extensive experiments on four real-world graph-structured datasets show that NIGEraser consistently outperforms existing unlearning methods, offering superior model utility while ensuring efficient and deterministic data removal. Yi Li 0083, Debo Cheng, Guixian Zhang, Shichao Zhang 0001 |
ACM Trans. Auton. Adapt. Syst. | 3 |
| 2026 | Towards Fair Graph Representation Learning by Overcoming Social HomophilyabstractWith the widespread use of Graph Neural Networks (GNNs) for representation learning from network data, the fairness of GNN models has raised great attention lately. Fair GNNs aim to ensure that node representations can be accurately classified, but not easily associated with a specific group. Existing advanced approaches essentially enhance the generalisation of node representation in combination with data augmentation strategy and do not directly impose constraints on the fairness of GNNs. In this work, we identify that a fundamental reason for the unfairness of GNNs is the phenomenon of social homophily , i.e., users in the same group are more inclined to congregate. The message-passing mechanism of GNNs can cause users in the same group to have similar representations due to social homophily, leading model predictions to establish spurious correlations with sensitive attributes. Inspired by this reason, we propose a method called Equity-Aware GNN (EAGNN) towards fair graph representation learning. Specifically, to ensure that model predictions are independent of sensitive attributes while maintaining prediction performance, we introduce constraints for fair representation learning based on three principles: sufficiency, independence and separation. We theoretically demonstrate that our EAGNN method can effectively achieve group fairness. Extensive experiments on three datasets with varying levels of social homophily illustrate that our EAGNN method achieves the state-of-the-art performance across two fairness metrics and offers competitive effectiveness. Guixian Zhang, Guan Yuan, Debo Cheng, Lin Liu 0003, Jiuyong Li, Shichao Zhang 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2025 | Community-Centric Graph UnlearningabstractGraph unlearning technology has become increasingly important since the advent of the `right to be forgotten' and the growing concerns about the privacy and security of artificial intelligence. Graph unlearning aims to quickly eliminate the effects of specific data on graph neural networks (GNNs). However, most existing deterministic graph unlearning frameworks follow a balanced partition-submodel training-aggregation paradigm, resulting in a lack of structural information between subgraph neighborhoods and redundant unlearning parameter calculations. To address this issue, we propose a novel Graph Structure Mapping Unlearning paradigm (GSMU) and a novel method based on it named Community-centric Graph Eraser (CGE). CGE maps community subgraphs to nodes, thereby enabling the reconstruction of a node-level unlearning operation within a reduced mapped graph. CGE makes the exponential reduction of both the amount of training data and the number of unlearning parameters. Extensive experiments conducted on five real-world datasets and three widely used GNN backbones have verified the high performance and efficiency of our CGE method, highlighting its potential in the field of graph unlearning. Yi Li 0083, Shichao Zhang 0001, Guixian Zhang, Debo Cheng |
AAAI | 3 |
| 2025 | Causality-Inspired Disentanglement for Fair Graph Neural NetworksabstractFair graph neural networks aim to eliminate discriminatory biases in predictions. Existing approaches often rely on adversarial learning to mitigate dependencies between sensitive attributes and labels but face challenges due to optimisation difficulties. A key limitation lies in neglecting intrinsic causality, which may lead to the entanglement of sensitive and causal factors, discarding causal factors or retaining sensitive factors in the final prediction, especially on unbalanced datasets. To address this issue, we propose a Causality-inspired Disentangled framework for Fair Graph neural networks (CDFG). In CDFG, node representations are conceptualised as a combination of causal and sensitive factors, enabling fair representation learning by only utilising the causal factors. We first use a counterfactual data generation mechanism to generate counterfactual data with similar causal factors but completely different sensitive factors. Then, we input real-world data and counterfactual data into the factor disentanglement module to achieve independence and disentanglement between the causal factors and sensitive factors. Finally, an adaptive mask module extracts the causal representation for fair and accurate graph-based predictions. Extensive experiments on three widely used datasets demonstrate that CDFG consistently outperforms existing methods, achieving competitive utility and significantly improved fairness. Guixian Zhang, Debo Cheng, Guan Yuan, Shang Liu 0001 |
IJCAI | 1 |
| 2025 | Injective edge chromatic index of sparse graphs
Guixian Zhang |
Discret. Appl. Math. | 2 |
| 2025 | Certainty-aware support vector machines for fair classification
Boyan Chen, Guixian Zhang, Debo Cheng, Shichao Zhang 0001 |
Neurocomputing | 3 |
| 2025 | Identifying local useful information for attribute graph anomaly detection
Penghui Xi, Debo Cheng, Guangquan Lu, Zhenyun Deng, Guixian Zhang, Shichao Zhang 0001 |
Neurocomputing | 5 |
| 2025 | Deconfounding representation learning for mitigating latent confounding effects in recommendation
Guixian Zhang, Guan Yuan, Debo Cheng, Lin Liu 0003, Jiuyong Li, Ziqi Xu 0001, Shichao Zhang 0001 |
Knowl. Inf. Syst. | 1 |
| 2025 | Multi-Cause Deconfounding for Recommender Systems with Latent Confounders
Zhirong Huang, Debo Cheng, Jiuyong Li, Lin Liu 0003, Guixian Zhang, Shichao Zhang 0001 |
Knowl. Based Syst. | 6 |
| 2025 | Disentangled contrastive learning for fair graph representations
Guixian Zhang, Guan Yuan, Debo Cheng, Lin Liu 0003, Jiuyong Li, Shichao Zhang 0001 |
Neural Networks | 1 |
| 2025 | Multibehavior Intent Disentangled Learning for Fine-Grained Interest Discovery in RecommendationabstractThe multibehavior recommendation aims at alleviating the data sparsity problem and improving recommendation accuracy by exploiting the rich knowledge in auxiliary behaviors. However, existing methods focus on modeling the relationships between behaviors while ignoring users’ interaction intents, making it difficult to capture users’ fine-grained interest requirements, which leads to a decline in user experience. To address this issue, we propose a multibehavior intent disentangled recommendation (MBIDR) model. First, we design an intent-aware interaction classifier that automatically identifies various intents based on user and item characteristics, classifying interactions into different categories to better explore users’ fine-grained interests. Second, we develop an adaptive relation learning approach that enables the model to better capture the varying importance of different interaction patterns in relation to user preferences. Third, we introduce multitask learning and nonsampling loss, which effectively leverage richer supervisory signals to enhance model training performance. Finally, extensive experiments on two real datasets demonstrate the effectiveness of MBIDR over baselines, with the best improvement reaching 16.20%. Guan Yuan, Guixian Zhang, Rui Bing |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | Latent Representation Learning for Attributed Graph Anomaly DetectionabstractAnomaly detection in attributed graph data has been widely applied in real applications. However, the intricate topology of graph data, high-dimensional attributes, and class imbalance inherent in anomaly detection tasks render attributed graph anomaly detection a challenging task. To detect anomalies using the intricate topology information of graph data, a dual-masked autoencoders is proposed for attributed graph anomaly detection, denoted as MAGAD. Specifically, in the MAGAD, the class imbalance in attributed graph data is dealt with by randomly masking the original graph data to obtain masked graph data for the anomaly detection task. And then, a latent representation of the graph data is obtained by training dual autoencoders, where one autoencoder is developed for reconstructing the original graph data, and another for reconstructing randomly masked graph data. This assists in identifying abnormal nodes in the attributed graph data. Subsequently, to capture anomalous information from relevant features, MAGAD uses a random re-masking strategy for latent representations learned from the masked graph. Finally, the anomaly scores of the nodes are calculated using the learned latent representations from the decoders of the dual autoencoders. Experimental results on five real-world datasets demonstrate that the MAGAD algorithm outperforms state-of-the-art anomaly detection algorithms. Shichao Zhang 0001, Penghui Xi, Mengqi Jiang, Guixian Zhang, Debo Cheng |
ACM Trans. Knowl. Discov. Data | 4 |
| 2025 | Mitigating Propensity Bias of Large Language Models for Recommender SystemsabstractThe rapid development of Large Language Models (LLMs) creates new opportunities for recommender systems, especially by exploiting the side information (e.g., descriptions and analyses of items) generated by these models. However, aligning this side information with collaborative information from historical interactions poses significant challenges. The inherent biases within LLMs can skew recommendations, resulting in distorted and potentially unfair user experiences. On the other hand, propensity bias causes side information to be aligned in such a way that it often tends to represent all inputs in a low-dimensional subspace, leading to a phenomenon known as dimensional collapse, which severely restricts the recommender system’s ability to capture user preferences and behaviors. To address these issues, we introduce a novel framework named Counterfactual LLM Recommendation (CLLMR). Specifically, we propose a spectrum-based side information encoder that implicitly embeds structural information from historical interactions into the side information representation, thereby circumventing the risk of dimension collapse. Furthermore, our CLLMR approach explores the causal relationships inherent in LLM-based recommender systems. By leveraging counterfactual inference, we counteract the biases introduced by LLMs. Extensive experiments demonstrate that our CLLMR approach consistently enhances the performance of various recommender models. Guixian Zhang, Guan Yuan, Debo Cheng, Lin Liu 0003, Jiuyong Li, Shichao Zhang 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2024 | Quantum Guess and Determine Attack on Stream CiphersabstractAbstract To guarantee the security of symmetric key schemes against quantum adversary, developing quantum cryptanalytic techniques becomes a major worldwide challenge in the post-quantum world. In this paper, we present a general framework of classical guess and determine attack on stream ciphers, and then convert it into quantum guess and determine attack. It shows that, for a given stream cipher with a key size of $k$ bits and an internal state size of $n$ bits, if a basic guess and determine attack with a time complexity below $O ( {{{2}^{{3k}/{2}}}}/{n} )$ is available, there is a quantum guess and determine attack with multiple data that can recover all $n$ internal state bits of the cipher with complexity below $O ( {{2^{k / 2}}} )$. As applications, we present quantum guess and determine attacks on the SNOW-like stream ciphers. The results show that all of SNOW 1.0 with 128-bit key, SNOW 2.0 with 128-bit key and SOSEMANUK are insecure against quantum guess and determine attack. The resource requirements for implementing a quantum guess and determine attack on SNOW 3G are evaluated as a case study. To the best of our knowledge, this is the first time that the general quantum guess and determine attack is formally proposed and applied to the SNOW-like stream ciphers. Lin Ding 0001, Guixian Zhang, Tairong Shi |
Comput. J. | 3 |
| 2024 | Learning fair representations via rebalancing graph structure
Guixian Zhang, Debo Cheng, Guan Yuan, Shichao Zhang 0001 |
Inf. Process. Manag. | 1 |
| 2024 | Contrastive learning for fair graph representations via counterfactual graph augmentation
Debo Cheng, Guixian Zhang, Shichao Zhang 0001 |
Knowl. Based Syst. | 3 |
| 2024 | Bayesian Graph Local Extrema Convolution with Long-tail Strategy for Misinformation DetectionabstractIt has become a cardinal task to identify fake information (misinformation) on social media, because it has significantly harmed the government and the public. There are many spam bots maliciously retweeting misinformation. This study proposes an efficient model for detecting misinformation with self-supervised contrastive learning. A B ayesian graph L ocal extrema C onvolution (BLC) is first proposed to aggregate node features in the graph structure. The BLC approach considers unreliable relationships and uncertainties in the propagation structure, and the differences between nodes and neighboring nodes are emphasized in the attributes. Then, a new long-tail strategy for matching long-tail users with the global social network is advocated to avoid over-concentration on high-degree nodes in graph neural networks. Finally, the proposed model is experimentally evaluated with two public Twitter datasets and demonstrates that the proposed long-tail strategy significantly improves the effectiveness of existing graph-based methods in terms of detecting misinformation. The robustness of BLC has also been examined on three graph datasets and demonstrates that it consistently outperforms traditional algorithms when perturbed by 15% of a dataset. Guixian Zhang, Shichao Zhang 0001, Guan Yuan |
ACM Trans. Knowl. Discov. Data | 1 |
| 2024 | Nonlocal Hybrid Network for Long-tailed Image ClassificationabstractIt is a significant issue to deal with long-tailed data when classifying images. A nonlocal hybrid network (NHN) that takes into account both feature learning and classifier learning is proposed. The NHN can capture the existence of dependencies between two locations that are far away from each other as well as alleviate the impact of long-tailed data on the model to some extent. The dependency relationship between distant pixels is obtained first through a nonlocal module to extract richer feature representations. Then, a learnable soft class center is proposed to balance the supervised contrastive loss and reduce the impact of long-tailed data on feature learning. For efficiency, a logit adjustment strategy is adopted to correct the bias caused by the different label distributions between the training and test sets and obtain a classifier that is more suitable for long-tailed data. Finally, extensive experiments are conducted on two benchmark datasets, the long-tailed CIFAR and the large-scale real-world iNaturalist 2018, both of which have imbalanced label distributions. The experimental results show that the proposed NHN model is efficient and promising. Rongjiao Liang, Shichao Zhang 0001, Guixian Zhang, Jinyun Tang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Multi-head Similarity Feature Representation and Filtration for Image-Text Matching
Mengqi Jiang, Shichao Zhang 0001, Debo Cheng, Leyuan Zhang, Guixian Zhang |
ADMA (2) | 5 |
| 2023 | LRAGAD: Local Information Recognition for Attribute Graph Anomaly DetectionabstractAnomaly detection is a crucial technique for comprehending the intricate structures of data, and it has garnered significant attention in various real-life applications, including finance, transportation, and network security. Presently, existing anomaly detection methods primarily concentrate on shallow techniques like residual analysis and community discovery, as well as deep learning methods that employ self-encoders as the underlying framework. However, these current methods overlook the local information during training, resulting in subpar and unexplained results. To recognize the local information, in this paper, we propose a novel Local Information Recognition method for Attribute Graph Anomaly Detection (LRAGAD). Specifically, to use the contextual structural information, LRAGAD first constructs a contrastive learning representation by generating different substructures from the target nodes, while reconstructing the whole graph using self-encoders that can utilize the neighborhood information of target nodes. Moreover, to better understand the complex graph structure, LRAGAD uses anomaly score estimation for outlier prediction. Experimental results on five real-world datasets demonstrate that the proposed LRAGAD method obtains better performance on AUC scores. Penghui Xi, Debo Cheng, Zhenyun Deng, Guixian Zhang, Shichao Zhang 0001 |
ICTAI | 4 |
| 2023 | FPGNN: Fair path graph neural network for mitigating discrimination
Guixian Zhang, Debo Cheng, Shichao Zhang 0001 |
World Wide Web (WWW) | 1 |
| 2022 | Multi-View Gated Graph Convolutional Network for Aspect-Level Sentiment Classification
Guixian Zhang, Zhi Lei, Zhirong Huang, Guangquan Lu |
ADMA (1) | 2 |
| 2022 | Personalized Headline Generation with Enhanced User Interest Perception
Guangquan Lu, Guixian Zhang, Zhi Lei |
ICANN (2) | 3 |
| 2022 | Bilateral-Branch Network for Imbalanced Visual RegressionabstractImbalanced visual regression is a practical and pressing issue, but current research is in its early stages. We propose an end-to-end Bilateral-Branch Network (BILBN) for dealing with imbalanced visual regression tasks. The BILBN consists of feature learning and regressor learning branches. The cumulative learning strategy is employed to gradually transition from feature learning to regressor learning in the BILBN model. Furthermore, we propose the Balanced MSESPL loss function, which allows the feature learning to learn simple features first and then progress to learn difficult ones. We also use feature distribution smoothing in the feature learning branch to learn a better feature representation. Compared with feature learning, regressor learning is quite simple, and we only use absolute error in the regressor branch. Finally, extensive experiments are conducted on the IMDB-WIKI-DIR and AgeDB-DIR to show the efficiency and superiority of our proposed methods and the BILBN model. Rongjiao Liang, Guixian Zhang, Zhi Lei, Shichao Zhang 0001 |
ICTAI | 2 |
| 2022 | Rumour Detection on Social Media with Long-Tail StrategyabstractWith the popularity of the mobile Internet, the proliferation of false rumors on social media has caused significant losses to the government and the public. Rumor detection on social media has become the critical research content. Recently, scholars have taken advantage of Graph Neural Networks (GNNs) to learn the textual features and propagation structure of rumors. On social media, only a few users have lots of retweets, and most real users belong to the long-tail users, who are low degree nodes in the social network. However, the existing graph convolution network would be more biased towards high degree nodes and ignore low degree nodes. We propose a new long-tail strategy based on the improved transformer that enhances the model's learning about long-tail users' features. Our long-tail strategy is a general approach that works well with existing GNN approaches. We also perform contrastive learning by combining sparse and dense attention to capture interaction features. Through extensive comparative and ablation experiments, we achieved state-of-the-art results and demonstrated the effectiveness of each module on two real-world datasets. Guixian Zhang, Rongjiao Liang, Zhongyi Yu, Shichao Zhang 0001 |
IJCNN | 1 |
| 2022 | A Multi-level Mesh Mutual Attention Model for Visual Question AnsweringabstractAbstract Visual question answering is a complex multimodal task involving images and text, with broad application prospects in human–computer interaction and medical assistance. Therefore, how to deal with the feature interaction and multimodal feature fusion between the critical regions in the image and the keywords in the question is an important issue. To this end, we propose a neural network based on the encoder–decoder structure of the transformer architecture. Specifically, in the encoder, we use multi-head self-attention to mine word–word connections within question features and stack multiple layers of attention to obtain multi-level question features. We propose a mutual attention module to perform information exchange between modalities for better question features and image features representation on the decoder side. Besides, we connect the encoder and decoder in a meshed manner, perform mutual attention operations with multi-level question features, and aggregate information in an adaptive way. We propose a multi-scale fusion module in the fusion stage, which utilizes feature information at different scales to complete modal fusion. We test and validate the model effectiveness on VQA v1 and VQA v2 datasets. Our model achieves better results than state-of-the-art methods. Zhi Lei, Guixian Zhang, Rongjiao Liang |
Data Sci. Eng. | 2 |