EDBT 2026 Demo / reviewers in the wild / expert
Fangzhi Xu
dblp:308/2941
· DBLP profile ↗
25ranked-venue papers
9as first author
25since 2021 · last 2026
0000-0003-4712-0547ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 7 first-author · 19 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAPS: Multi-Agent Personality Shaping for Collaborative ReasoningabstractCollaborative reasoning with multiple agents offers the potential for more robust and diverse problem-solving. However, existing approaches often suffer from homogeneous agent behaviors and lack of reflective and rethinking capabilities. We propose Multi-Agent Personality Shaping ((MAPS), a novel framework that enhances reasoning through agent diversity and internal critique. Inspired by the Big Five personality theory, MAPS assigns distinct personality traits to individual agents, shaping their reasoning styles and promoting heterogeneous collaboration. To enable deeper and more adaptive reasoning, MAPS introduces a Critic agent that reflects on intermediate outputs, revisits flawed steps, and guides iterative refinement. This integration of personality-driven agent design and structured collaboration improves both reasoning depth and flexibility. Empirical evaluations across three benchmarks demonstrate the strong performance of MAPS, with further analysis confirming its generalizability across different large language models and validating the benefits of multi-agent collaboration. Jian Zhang 0087, Zhangqi Wang, Fangzhi Xu, Qika Lin, Lingling Zhang 0005, Rui Mao 0010, Erik Cambria, Jun Liu 0002 |
AAAI | 4 |
| 2026 | OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic WorkflowsabstractQiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, Lingpeng Kong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie 0002, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zichen Ding 0002, Qi Liu 0049, Zhiyong Wu 0003, Zhuosheng Zhang 0001, Ben Kao, Lingpeng Kong |
ACL (1) | 5 |
| 2026 | MUR: Momentum Uncertainty guided Reasoning for Large Language ModelsabstractHang Yan, Fangzhi Xu, Rongman Xu, Yifei Li, Jian Zhang, Haoran Luo, Xiaobao Wu, Anh Tuan Luu, Haiteng Zhao, Qika Lin, Jun Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hang Yan 0010, Fangzhi Xu, Rongman Xu, Yifei Li 0006, Jian Zhang 0087, Haoran Luo 0001, Xiaobao Wu, Anh Tuan Luu, Haiteng Zhao, Qika Lin, Jun Liu 0002 |
ACL (1) | 2 |
| 2026 | OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using AgentsabstractBowen Yang, Kaiming Jin, Zhenyu Wu, Zhaoyang Liu, Qiushi Sun, Zehao Li, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Yian Wang, Qingyun Li, Yu Qiao, Zun Wang, Zichen Ding. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kaiming Jin, Zhaoyang Liu 0001, Qiushi Sun, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Yian Wang 0003, Qingyun Li, Yu Qiao 0001, Zun Wang 0001, Zichen Ding 0002 |
ACL (1) | 9 |
| 2025 | Self-supervised Quantized Representation for Seamlessly Integrating Knowledge Graphs with Large Language ModelsabstractQika Lin, Tianzhe Zhao, Kai He, Zhen Peng, Fangzhi Xu, Ling Huang, Jingying Ma, Mengling Feng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Qika Lin, Tianzhe Zhao, Kai He 0001, Zhen Peng 0005, Fangzhi Xu, Ling Huang 0003, Jingying Ma, Mengling Feng |
ACL (1) | 5 |
| 2025 | OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task SynthesisabstractGraphical User Interface (GUI) agents powered by Vision-Language Models (VLMs) have demonstrated human-like computer control capability. Despite their utility in advancing digital automation, a critical bottleneck persists: collecting high-quality trajectory data for training. Common practices for collecting such data rely on human supervision or synthetic data generation through executing pre-defined tasks, which are either resource-intensive or unable to guarantee data quality. Moreover, these methods suffer from limited data diversity and significant gaps between synthetic data and real-world environments. To address these challenges, we propose OS-Genesis, a novel GUI data synthesis pipeline that reverses the conventional trajectory collection process. Instead of relying on pre-defined tasks, OS-Genesis enables agents first to perceive environments and perform step-wise interactions, then retrospectively derive high-quality tasks to enable trajectory-level exploration. A trajectory reward model is then employed to ensure the quality of the generated trajectories. We demonstrate that training GUI agents with OS-Genesis significantly improves their performance on highly challenging online benchmarks. In-depth analysis further validates OS-Genesis's efficiency and its superior data quality and diversity compared to existing synthesis methods. Our codes, data, and checkpoints are available at OS-Genesis Homepage. Qiushi Sun, Kanzhi Cheng, Zichen Ding 0002, Chuanyang Jin, Yian Wang 0003, Fangzhi Xu, Chengyou Jia, Zhoumianze Liu, Ben Kao, Guohao Li 0001, Junxian He, Yu Qiao 0001, Zhiyong Wu 0003 |
ACL (1) | 6 |
| 2025 | φ-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
Fangzhi Xu, Hang Yan 0010, Haiteng Zhao, Jun Liu 0002, Qika Lin, Zhiyong Wu 0003 |
ACL (1) | 1 |
| 2025 | Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced ReasoningabstractFangzhi Xu, Hang Yan, Chang Ma, Haiteng Zhao, Qiushi Sun, Kanzhi Cheng, Junxian He, Jun Liu, Zhiyong Wu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Fangzhi Xu, Hang Yan 0010, Haiteng Zhao, Qiushi Sun, Kanzhi Cheng, Junxian He, Jun Liu 0002, Zhiyong Wu 0003 |
ACL (1) | 1 |
| 2025 | Interactive Evolution: A Neural-Symbolic Self-Training Framework For Large Language ModelsabstractOne of the primary driving forces contributing to the superior performance of Large Language Models (LLMs) is the extensive availability of human-annotated natural language data, which is used for alignment fine-tuning. This inspired researchers to investigate self-training methods to mitigate the extensive reliance on human annotations. However, the current success of self-training has been primarily observed in natural language scenarios, rather than in the increasingly important neural-symbolic scenarios. To this end, we propose an environment-guided neural-symbolic self-training framework named ENVISIONS. It aims to overcome two main challenges: (1) the scarcity of symbolic data, and (2) the limited proficiency of LLMs in processing symbolic language. Extensive evaluations conducted on three distinct domains demonstrate the effectiveness of our approach. Additionally, we have conducted a comprehensive analysis to uncover the factors contributing to ENVISIONS’s success, thereby offering valuable insights for future research in this area. Fangzhi Xu, Qiushi Sun, Kanzhi Cheng, Jun Liu 0002, Yu Qiao 0001, Zhiyong Wu 0003 |
ACL (1) | 1 |
| 2025 | OS-ATLAS: Foundation Action Model for Generalist GUI AgentsabstractExisting efforts in building GUI agents heavily rely on the availability of robust commercial Vision-Language Models (VLMs) such as GPT-4o and GeminiProVision. Practitioners are often reluctant to use open-source VLMs due to their significant performance lag compared to their closed-source counterparts, particularly in GUI grounding and Out-Of-Distribution (OOD) scenarios. To facilitate future research in this area, we developed OS-Atlas—a foundational GUI action model that excels at GUI grounding and OOD agentic tasks through innovations in both data and modeling.
We have invested significant engineering effort in developing an open-source toolkit for synthesizing GUI grounding data across multiple platforms, including Windows, Linux, MacOS, Android, and the web. Leveraging this toolkit, we are releasing the largest open-source cross-platform GUI grounding corpus to date, which contains over 13 million GUI elements. This dataset, combined with innovations in model training, provides a solid foundation for OS-Atlas to understand GUI screenshots and generalize to unseen interfaces.
Through extensive evaluation across six benchmarks spanning three different platforms (mobile, desktop, and web), OS-Atlas demonstrates significant performance improvements over previous state-of-the-art models. Our evaluation also uncovers valuable insights into continuously improving and scaling the agentic capabilities of open-source VLMs. Zhiyong Wu 0003, Fangzhi Xu, Yian Wang 0003, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding 0002, Paul Pu Liang, Yu Qiao 0001 |
ICLR | 3 |
| 2025 | Vision-Language Models Can Self-Improve Reasoning via ReflectionabstractKanzhi Cheng, Li YanTao, Fangzhi Xu, Jianbing Zhang, Hao Zhou, Yang Liu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kanzhi Cheng, Yantao Li 0003, Fangzhi Xu, Hao Zhou 0012, Yang Liu 0005 |
NAACL (Long Papers) | 3 |
| 2025 | ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart UnderstandingabstractCharts are high-density visualization carriers for complex data, serving as a crucial medium for information extraction and analysis. Automated chart understanding poses significant challenges to existing multimodal large language models (MLLMs) due to the need for precise and complex visual reasoning. Current step-by-step reasoning models primarily focus on text-based logical reasoning for chart understanding. However, they struggle to refine or correct their reasoning when errors stem from flawed visual understanding, as they lack the ability to leverage multimodal interaction for deeper comprehension. Inspired by human cognitive behavior, we propose ChartSketcher, a multimodal feedback-driven step-by-step reasoning method designed to address these limitations. ChartSketcher is a chart understanding model that employs Sketch-CoT, enabling MLLMs to annotate intermediate reasoning steps directly onto charts using a programmatic sketching library, iteratively feeding these visual annotations back into the reasoning process. This mechanism enables the model to visually ground its reasoning and refine its understanding over multiple steps. We employ a two-stage training strategy: a cold start phase to learn sketch-based reasoning patterns, followed by off-policy reinforcement learning to enhance reflection and generalization. Experiments demonstrate that ChartSketcher achieves promising performance on chart understanding benchmarks and general vision tasks, providing an interactive and interpretable approach to chart comprehension. Muye Huang, Lingling Zhang 0005, Jie Ma 0001, Han Lai, Fangzhi Xu, Yifei Li 0006, Yaqiang Wu, Jun Liu 0002 |
NeurIPS | 5 |
| 2025 | Cross-Modal Knowledge Diffusion-Based Generation for Difference-Aware Medical VQAabstractMultimodal medical applications have garnered considerable attention due to their potential to offer comprehensive and robust support for medical assistance. Specifically, within this domain, difference-aware medical Visual Question Answering (VQA) has emerged as a topic of increasing interest that enables the recognition of changes in physical conditions over time when compared to previous states and provides customized suggestions accordingly. However, it is challenging because samples usually exhibit characteristics of complexity, diversity, and inherent noise. Besides, there is a need for multimodal knowledge understanding of the medical domain. The difference-aware setting requiring image comparison further intensifies these situations. To this end, we propose a cross-Modal knowlEdge diffusioN-baseD gEneration netwoRk (MENDER), where the diffusion mechanism with multi-step denoising and knowledge injection from global to local level are employed to tackle the aforementioned challenges, respectively. The diffusion process is to gradually generate answers with the sequence input of questions, random noises for the answer masks and virtual vision prompts of images. The strategy of answer nosing and knowledge cascading is specifically tailored for this task and is implemented during forward and reverse diffusion processes. Moreover, the visual and structure knowledge injection are proposed to learn virtual vision prompts to guide the diffusion process, where the former is realized using a pre-trained medical image-text network and the latter is modeled with spatial and semantic graph structures processed by the heterogeneous graph Transformer models. Experiment results demonstrate the effectiveness of MENDER for difference-aware medical VQA. Furthermore, it also exhibits notable performance in the low-resource setting and conventional medical VQA tasks. Qika Lin, Kai He 0001, Yifan Zhu 0001, Fangzhi Xu, Erik Cambria, Mengling Feng |
IEEE Trans. Image Process. | 4 |
| 2025 | Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and BeyondabstractLogical reasoning consistently plays a fundamental and significant role in the domains of knowledge engineering and artificial intelligence. Recently, Large Language Models (LLMs) have emerged as a noteworthy innovation in natural language processing (NLP). However, the question of whether LLMs can effectively address the task of logical reasoning, which requires gradual cognitive inference similar to human intelligence, remains unanswered. To this end, we aim to bridge this gap and provide comprehensive evaluations in this paper. First, to offer systematic evaluations, we select fifteen typical logical reasoning datasets and organize them into deductive, inductive, abductive and mixed-form reasoning settings. Considering the comprehensiveness of evaluations, we include 3 early-era representative LLMs and 4 trending LLMs. Second, different from previous evaluations relying only on simple metrics (e.g.,accuracy), we propose fine-level evaluations in objective and subjective manners, covering both answers and explanations, includinganswer correctness,explain correctness,explain completenessandexplain redundancy. Additionally, to uncover the logical flaws of LLMs, problematic cases will be attributed to five error types from two dimensions, i.e.,evidence selection processandreasoning process. Third, to avoid the influences of knowledge bias and concentrate purely on benchmarking the logical reasoning capability of LLMs, we propose a new dataset with neutral content. Based on the in-depth evaluations, this paper finally forms a general evaluation scheme of logical reasoning capability from six dimensions (i.e.,Correct,Rigorous,Self-aware,Active,OrientedandNo hallucination). It reflects the pros and cons of LLMs and gives guiding directions for future works. Fangzhi Xu, Qika Lin, Jiawei Han 0010, Tianzhe Zhao, Jun Liu 0002, Erik Cambria |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsabstractKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Li YanTao, Jianbing Zhang, Zhiyong Wu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li 0003, Zhiyong Wu 0003 |
ACL (1) | 4 |
| 2024 | PathReasoner: Modeling Reasoning Path with Equivalent Extension for Logical Question AnsweringabstractLogical reasoning task has attracted great interest since it was proposed.Faced with such a task, current competitive models, even large language models (e.g., ChatGPT and PaLM 2), still perform badly.Previous promising LMs struggle in logical consistency modeling and logical structure perception.To this end, we model the logical reasoning task by transforming each logical sample into reasoning paths and propose an architecture PathReasoner.It addresses the task from the views of both data and model.To expand the diversity of the logical samples, we propose an atom extension strategy supported by equivalent logical formulas, to form new reasoning paths.From the model perspective, we design a stack of transformer-style blocks.In particular, we propose a path-attention module to joint model in-atom and cross-atom relations with the highorder diffusion strategy.Experiments show that PathReasoner achieves competitive performances on two logical reasoning benchmarks and great generalization abilities. Fangzhi Xu, Qika Lin, Tianzhe Zhao, Jiawei Han 0010, Jun Liu 0002 |
ACL (1) | 1 |
| 2024 | Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language ModelsabstractFangzhi Xu, Zhiyong Wu, Qiushi Sun, Siyu Ren, Fei Yuan, Shuai Yuan, Qika Lin, Yu Qiao, Jun Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Fangzhi Xu, Zhiyong Wu 0003, Qiushi Sun, Fei Yuan 0006, Shuai Yuan 0018, Qika Lin, Yu Qiao 0001, Jun Liu 0002 |
ACL (1) | 1 |
| 2024 | A Semantic Mention Graph Augmented Model for Document-Level Event Argument ExtractionabstractDocument-level Event Argument Extraction (DEAE) aims to identify arguments and their specific roles from an unstructured document. The advanced approaches on DEAE utilize prompt-based methods to guide pre-trained language models (PLMs) in extracting arguments from input documents. They mainly concentrate on establishing relations between triggers and entity mentions within documents, leaving two unresolved problems: a) independent modeling of entity mentions; b) document-prompt isolation. To this end, we propose a semantic mention Graph Augmented Model (GAM) to address these two problems in this paper. Firstly, GAM constructs a semantic mention graph that captures relations within and between documents and prompts, encompassing co-existence, co-reference and co-type relations. Furthermore, we introduce an ensemble graph transformer module to address mentions and their three semantic relations effectively. Later, the graph-augmented encoder-decoder module incorporates the relation-specific graph into the input embedding of PLMs and optimizes the encoder section with topology information, enhancing the relations comprehensively. Extensive experiments on the RAMS and WikiEvents datasets demonstrate the effectiveness of our approach, surpassing baseline methods and achieving a new state-of-the-art performance. Jian Zhang 0087, Changlin Yang, Qika Lin, Fangzhi Xu, Jun Liu 0002 |
LREC/COLING | 5 |
| 2024 | Contrastive Graph Representations for Logical Formulas Embedding (Extended Abstract)abstractEmbedding symbolic logical formulas into a low-dimensional continuous space provides an effective way for the Neural-Symbolic system. However, current studies are all constrained by the syntactic structure modeling and fail to preserve intrinsic semantics. To this end, we propose a novel model of Contrastive Graph Representations (ConGR) for logical formulas embedding. Firstly, it introduces a densely connected graph convolutional network (GCN) with an attention mechanism to process syntax parsing graphs of formulas. Secondly, the contrastive instances for each anchor formula are generated by the transformation under the guidance of logical properties. Two types of contrast, global-local and global-global, are carried out to refine formula embeddings with semantic information. Extensive experiments demonstrate that ConGR obtains superior performance against state-of-the-art baselines. Qika Lin, Jun Liu 0002, Lingling Zhang 0005, Yudai Pan, Fangzhi Xu, Hongwei Zeng 0001 |
ICDE | 6 |
| 2024 | Mind Reasoning Manners: Enhancing Type Perception for Generalized Zero-Shot Logical Reasoning Over TextabstractLogical reasoning task involves diverse types of complex reasoning over text, based on the form of multiple-choice question answering (MCQA). Given the context, question and a set of options as the input, previous methods achieve superior performances on the full-data setting. However, the current benchmark dataset has the ideal assumption that the reasoning type distribution on the train split is close to the test split, which is inconsistent with many real application scenarios. To address it, there remain two problems to be studied: 1) how is the zero-shot capability of the models (train on seen types and test on unseen types)? and 2) how to enhance the perception of reasoning types for the models? For problem 1, we propose a new benchmark for generalized zero-shot logical reasoning, named ZsLR. It includes six splits based on the three type sampling strategies. For problem 2, a type-aware model TaCo is proposed. It utilizes the heuristic input reconstruction and builds a text graph with a global node. Incorporating graph reasoning and contrastive learning, TaCo can improve the type perception in the global representation. Extensive experiments on both the zero-shot and full-data settings prove the superiority of TaCo over the state-of-the-art (SOTA) methods. Also, we experiment and verify the generalization capability of TaCo on other logical reasoning dataset. Fangzhi Xu, Jun Liu 0002, Qika Lin, Tianzhe Zhao, Jian Zhang 0087, Lingling Zhang 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | TECHS: Temporal Logical Graph Networks for Explainable Extrapolation ReasoningabstractExtrapolation reasoning on temporal knowledge graphs (TKGs) aims to forecast future facts based on past counterparts.There are two main challenges: (1) incorporating the complex information, including structural dependencies, temporal dynamics, and hidden logical rules;(2) implementing differentiable logical rule learning and reasoning for explainability.To this end, we propose an explainable extrapolation reasoning framework TEemporal logiCal grapH networkS (TECHS), which mainly contains a temporal graph encoder and a logical decoder.The former employs a graph convolutional network with temporal encoding and heterogeneous attention to embed topological structures and temporal dynamics.The latter integrates propositional reasoning and first-order reasoning by introducing a reasoning graph that iteratively expands to find the answer.A forward message-passing mechanism is also proposed to update node representations, and their propositional and first-order attention scores.Experimental results demonstrate that it outperforms state-of-the-art baselines. Qika Lin, Jun Liu 0002, Rui Mao 0010, Fangzhi Xu, Erik Cambria |
ACL (1) | 4 |
| 2023 | MoCA: Incorporating domain pretraining and cross attention for textbook question answering
Fangzhi Xu, Qika Lin, Jun Liu 0002, Lingling Zhang 0005, Tianzhe Zhao, Qi Chai, Yudai Pan, Yi Huang 0017, Qianying Wang 0002 |
Pattern Recognit. | 1 |
| 2023 | Contrastive Graph Representations for Logical Formulas EmbeddingabstractCurrently, the non-transparent computing process of deep learning has become a significant reason hindering its further development. The Neural-Symbolic (NS) system formed by integrating logic rules into neural networks has attracted increasing attention owing to its direct interpretability. Embedding symbolic logical formulas into a low-dimensional continuous space provides an effective way for the NS system. However, current studies are all constrained by the modeling ability for its syntactic structure and fail to preserve the intrinsic semantics in embeddings, which causes poor performance on downstream reasoning tasks. To this end, this paper proposes a novel method ofContrastiveGraphRepresentations (ConGR) for logical formulas embedding. First, to improve the modeling ability for the syntactic structure, ConGR introduces a densely connected graph convolutional network (GCN) with an attention mechanism to process syntax parsing graphs of formulas. In this way, discriminative local and global embeddings of formulas are obtained at the syntax level. Second, the contrastive instances (positive or negative) for each anchor formula are generated by the transformation under the guidance of logical properties. To preserve semantic information, two types of contrast, global-local and global-global, are carried out to refine formula embeddings. Extensive experiments demonstrate that ConGR obtains superior performance against state-of-the-art baselines on entailment checking and premise selection datasets. Qika Lin, Jun Liu 0002, Lingling Zhang 0005, Yudai Pan, Fangzhi Xu, Hongwei Zeng 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | Incorporating Context Graph with Logical Reasoning for Inductive Relation PredictionabstractRelation prediction on knowledge graphs (KGs) aims to infer missing valid triples from observed ones. Although this task has been deeply studied, most previous studies are limited to the transductive setting and cannot handle emerging entities. Actually, the inductive setting is closer to real-life scenarios because it allows entities in the testing phase to be unseen during training. However, it is challenging to precisely conduct inductive relation prediction as there exists requirements of entity-independent relation modeling and discrete logical reasoning for interoperability. To this end, we propose a novel model ConGLR to incorporate context graph with logical reasoning. Firstly, the enclosing subgraph w.r.t. target head and tail entities are extracted and initialized by the double radius labeling. And then the context graph involving relational paths, relations and entities is introduced. Secondly, two graph convolutional networks (GCNs) with the information interaction of entities and relations are carried out to process the subgraph and context graph respectively. Considering the influence of different edges and target relations, we introduce edge-aware and relation-aware attention mechanisms for the subgraph GCN. Finally, by treating the relational path as rule body and target relation as rule head, we integrate neural calculating and logical reasoning to obtain inductive scores. And to focus on the specific modeling goals of each module, the stop-gradient is utilized in the information interaction between context graph and subgraph GCNs in the training process. In this way, ConGLR satisfies two inductive requirements at the same time. Extensive experiments demonstrate that ConGLR obtains outstanding performance against state-of-the-art baselines on twelve inductive dataset versions of three common KGs. Qika Lin, Jun Liu 0002, Fangzhi Xu, Yudai Pan, Yifan Zhu 0001, Lingling Zhang 0005, Tianzhe Zhao |
SIGIR | 3 |
| 2022 | Logiformer: A Two-Branch Graph Transformer Network for Interpretable Logical ReasoningabstractMachine reading comprehension has aroused wide concerns, since it explores the potential of model for text understanding. To further equip the machine with the reasoning capability, the challenging task of logical reasoning is proposed. Previous works on logical reasoning have proposed some strategies to extract the logical units from different aspects. However, there still remains a challenge to model the long distance dependency among the logical units. Also, it is demanding to uncover the logical structures of the text and further fuse the discrete logic to the continuous text embedding. To tackle the above issues, we propose an end-to-end model Logiformer which utilizes a two-branch graph transformer network for logical reasoning of text. Firstly, we introduce different extraction strategies to split the text into two sets of logical units, and construct the logical graph and the syntax graph respectively. The logical graph models the causal relations for the logical branch while the syntax graph captures the co-occurrence relations for the syntax branch. Secondly, to model the long distance dependency, the node sequence from each graph is fed into the fully connected graph transformer structures. The two adjacent matrices are viewed as the attention biases for the graph transformer layers, which map the discrete logical structures to the continuous text embedding space. Thirdly, a dynamic gate mechanism and a question-aware self-attention module are introduced before the answer prediction to update the features. The reasoning process provides the interpretability by employing the logical units, which are consistent with human cognition. The experimental results show the superiority of our model, which outperforms the state-of-the-art single model on two logical reasoning benchmarks. Fangzhi Xu, Jun Liu 0002, Qika Lin, Yudai Pan, Lingling Zhang 0005 |
SIGIR | 1 |