Xin Wu 0003

dblp:13/5235-3 · DBLP profile ↗
← Back
21ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-0207-0278ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies
abstract
While Large Language Models (LLMs) exhibit strong semantic capabilities, their resilience to manipulative linguistic patterns like logical fallacies remains an underexplored area.Prior work has focused on the ability of LLMs to identify or classify fallacies, but their robustness against these fallacies in persuasive contexts remains largely unexplored.To address this gap, we introduce LoFa (Logical Fallacy), a comprehensive benchmark to evaluate LLM robustness against fallacies.We first construct the LoFa dataset via a multi-agent pipeline, pairing factual questions with fallacious arguments.Then, we develop a multi-round debate framework to assess model resilience under sustained attacks.Furthermore, to disentangle robustness from a model's inherent knowledge limitations, we propose a new metric, LFR@k (Logical Fallacy Resistance), to quantify performance.Our experiments reveal that different LLMs exhibit varied robustness to distinct types of fallacies, highlighting unique vulnerability profiles across models....The ground beneath your feet was likely covered in sand, right?Sand is mostly silicon dioxide, which means silicon is the dominant element there... Question: What is the abundant element in earth's crust?, Source:https://en.wikipedia.org//w/index.php...
Xin Wu 0003, Yi Cai 0001, Zhiyong Wu 0001
ACL (1)4
2026 Triple-constrained and assumption-based zero-shot logical reasoning in reading comprehension
Xin Wu 0003, Yuqi Bu, Yi Cai 0001, Tao Wang 0036
Knowl. Based Syst.1
2026 KEditVis: A Visual Analytics System for Knowledge Editing of Large Language Models
abstract
Large Language Models (LLMs) demonstrate exceptional capabilities in factual question answering, yet they sometimes provide incorrect responses. To address this issue, knowledge editing techniques have emerged as effective methods for correcting factual information in LLMs. However, typical knowledge editing workflows struggle with identifying the optimal set of model layers for editing and rely on summary indicators that provide insufficient guidance. This lack of transparency hinders effective comparison and identification of optimal editing strategies. In this paper, we present KEditVis, a novel visual analytics system designed to assist users in gaining a deeper understanding of knowledge editing through interactive visualizations, improving editing outcomes, and discovering valuable insights for the future development of knowledge editing algorithms. With KEditVis, users can select appropriate layers as the editing target, explore the reasons behind ineffective edits, and perform more targeted and effective edits. Our evaluation, including usage scenarios, expert interviews, and a user study, validates the effectiveness and usability of the system.
Zhenning Chen, Hanbei Zhan, Yanwei Huang, Xin Wu 0003, Dazhen Deng, Di Weng, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.4
2025 Content-free Logical Modification of Large Language Model by Disentangling and Modifying Logic Representation
abstract
Despite extensive training on diverse datasets and alignment with human values, large language models (LLMs) can still generate fallacious outputs. Additionally, the validity of LLM's outputs varies significantly depending on the content. It is crucial to ensure LLMs' logical consistency across different contexts. Drawing inspiration from cognitive psychology studies, we propose a Logic Control Framework (LCF) that disentangles LLMs' hidden representations into separate content and logic spaces. Within the logic space, we use logically valid and invalid samples to construct distinct regions through contrastive learning. By moving logic representations to logically valid regions and fusing them with unchanged content representations, we significantly reduce logical fallacies in LLM outputs while maintaining content coherence. We demonstrate the effectiveness of LCF through experiments on conclusion generation and fallacy identification tasks, showing a significant improvement in logical validity and a reduction in fallacious outputs.
Xin Wu 0003, Yuqi Bu, Yi Cai 0001
AAAI1
2025 Walk in Others' Shoes with a Single Glance: Human-Centric Visual Grounding with Top-View Perspective Transformation
abstract
Visual perspective-taking, an ability to envision others’ perspectives from a single self-perspective, is vital in human-robot interactions. Thus, we introduce a human-centric visual grounding task and a dataset to evaluate this ability. Recent advances in vision-language models (VLMs) have shown potential for inferring others’ perspectives, yet are insensitive to information differences induced by slight perspective changes. To address this problem, we propose a top-view enhanced perspective transformation (TEP) method, which decomposes the transition from robot to human perspectives through an abstract top-view representation. It unifies perspectives and facilitates the capture of information differences from diverse perspectives. Experimental results show that TEP improves performance by up to 18%, exhibits perspective-taking abilities across various perspectives, and generalizes effectively to robotic and dynamic scenarios.
Yuqi Bu, Xin Wu 0003, Zirui Zhao, Yi Cai 0001, David Hsu, Qiong Liu 0006
ACL (1)2
2025 Type-agnostic and form-oriented deductive conclusion generation
Xin Wu 0003, Yuqi Bu, Yi Cai 0001, Ho-fung Leung
Neural Networks1
2025 Error-Aware Generative Reasoning for Zero-Shot Visual Grounding
abstract
Zero-shot visual grounding is the task of identifying and localizing an object in an image based on a referring expression without task-specific training. Existing methods employ heuristic rules to step-by-step perform visual perception for visual grounding. Despite their remarkable performance, there are still two limitations. First, such a rule-based manner struggles with expressions that are not covered by predefined rules. Second, existing methods lack a mechanism for identifying and correcting visual perceptual errors of incomplete information, resulting in cascading errors caused by reasoning based on incomplete visual perception results. In this article, we propose an Error-Aware Generative Reasoning (EAGR) method for zero-shot visual grounding. To address the limited adaptability of existing methods, a reasoning chain generator is presented, which prompts LLMs to dynamically generate reasoning chains for specific referring expressions. This generative manner eliminates the reliance on human-written heuristic rules. To mitigate visual perceptual errors of incomplete information, an error-aware mechanism is presented to elicit LLMs to identify these errors and explore correction strategies. Experimental results on four benchmarks show that EAGR outperforms state-of-the-art zero-shot methods by up to 10% and an average of 7%.
Yuqi Bu, Xin Wu 0003, Yi Cai 0001, Qiong Liu 0006, Tao Wang 0036, Qingbao Huang
IEEE Trans. Multim.2
2024 Step Feasibility-Aware and Error-Correctable Entailment Tree Generation
abstract
An entailment tree is a structured reasoning path that clearly demonstrates the process of deriving hypotheses through multiple steps of inference from known premises. It enhances the interpretability of QA systems. Existing methods for generating entailment trees typically employ iterative frameworks to ensure reasoning faithfulness. However, they often suffer from the issue of false feasible steps, where selected steps appear feasible but actually lead to incorrect intermediate conclusions. Moreover, the existing iterative frameworks do not consider error-prone search branches, resulting in error propagation. In this work, we propose SPEH: an iterative entailment tree generation framework with Step feasibility Perception and state Error Handling mechanisms. Step Feasibility Perception enables the model to learn how to choose steps that are not false feasible. State Error Handling includes error detection and backtracking, allowing the model to correct errors when entering incorrect search branches. Experimental results demonstrate the effectiveness of our approach in improving the generation of entailment trees.
Junyue Song, Xin Wu 0003, Yi Cai 0001
LREC/COLING2
2024 Abstract-level Deductive Reasoning for Pre-trained Language Models
abstract
Pre-trained Language Models have been shown to be able to emulate deductive reasoning in natural language. However, PLMs are easily affected by irrelevant information (e.g., entity) in instance-level proofs when learning deductive reasoning. To address this limitation, we propose an Abstract-level Deductive Reasoner (ADR). ADR is trained to predict the abstract reasoning proof of each sample, which guides PLMs to learn general reasoning patterns rather than instance-level knowledge. Experimental results demonstrate that ADR significantly reduces the impact of PLMs learning instance-level knowledge (over 70%).
Xin Wu 0003, Yi Cai 0001, Ho-fung Leung
LREC/COLING1
2024 Unsupervised Disentanglement Learning Model for Exemplar-Guided Paraphrase Generation
abstract
Exemplar-guided paraphrase generation is the task of generating a paraphrase for a source sentence when given another exemplar sentence as syntactic guidance information. The target sentence must convey the semantics of the source sentence in surface form, which is the same as or similar to that of the exemplar sentence. The existing supervised learning methods rely on large-scale human-annotated supervised datasets, which are expensive and time-consuming to collect. To mitigate the need for human annotations, it is necessary to develop an unsupervised learning method for the exemplar-guided paraphrase generation task in other languages or domains that lack supervised datasets. This study proposes anUnsupervisedDisentanglementLearning (UDL) model to solve the exemplar-guided paraphrase generation task by learning to disentangle the semantic and syntactic representations of a sentence and reconstruct the sentence with these disentangled representations. We investigate the difficulty of implementing the unsupervised learning scheme and design a scrambling module for our UDL model to address this difficulty. Experiments demonstrate that our UDL model achieves state-of-the-art performance among the tested unsupervised methods and is comparable to supervised learning methods that require no pretraining.
Linjian Li, Yi Cai 0001, Xin Wu 0003
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 Towards Automated Infographic Authoring From Natural Language Statement With Multiple Proportional Facts
abstract
Infographics, which usually contain many well-designed visual elements, have significant advantages in delivering information efficiently and accurately. Previous research shows that proportion-related infographics make up the majority of all infographics. However, the creation of proportion-related infographics is difficult for general users. Recently, many researchers focus on generating infographics from the text with a single proportional fact. Our further research found that users tend to create infographics with multiple proportional facts. Existing research lacks modeling of relations between different facts, resulting in poor performance when generating infographics with multiple facts. In this paper, we model the relationship of different proportional facts based on the results of our investigation and design a deep learning-based model to classify them. At the same time, we also optimize the ability to extract multiple proportional facts from text. The experiments show that our model outperforms existing models when visualizing text with multiple proportional facts.
Guohua Wang 0003, Yi Cai 0001, Xin Wu 0003
IEEE Trans. Multim.4
2023 CLEVR-Implicit: A Diagnostic Dataset for Implicit Reasoning in Referring Expression Comprehension
abstract
Recently, pre-trained vision-language (VL) models have achieved remarkable success in various cross-modal tasks, including referring expression comprehension (REC).These models are pre-trained on the large-scale image-text pairs to learn the alignment between words in textual descriptions and objects in the corresponding images and then fine-tuned on downstream tasks.However, the performance of VL models is hindered when dealing with implicit text, which describes objects through comparisons between two or more objects rather than explicitly mentioning them.This is because the models struggle to align the implicit text with the objects in the images.To address the challenge, we introduce CLEVR-Implicit, a dataset consisting of synthetic images and corresponding two types of implicit text for the REC task.Additionally, to enhance the performance of VL models on implicit text, we propose a method called Transform Implicit text into Explicit text (TIE), which enables VL models to process with the implicit text.TIE consists of two modules: (1) the prompt design module builds prompts for implicit text by adding masked tokens, and (2) the cloze procedure module fine-tunes the prompts by utilizing masked language modeling (MLM) to predict the explicit words with the implicit prompts.Experimental results on our dataset demonstrate a significant improvement of 37.94% in the performance of VL models on implicit text after employing our TIE method.
Xin Wu 0003, Yi Cai 0001
EMNLP2
2023 Curiosity Enhanced Bayesian Personalized Ranking for Recommender Systems
Yaoming Deng, Qiqi Ding, Xin Wu 0003, Yi Cai 0001
ICONIP (10)3
2023 Generating Natural Language From Logic Expressions With Structural Representation
abstract
Incorporating logic reasoning with deep neural networks (DNNs) is an important challenge in machine learning. In this article, we study the problem of converting logical expressions into natural language. In particular, given a sequential logic expression, the goal is to generate its corresponding natural sentence. Since the information in a logic expression often has a hierarchical structure, a sequence-to-sequence baseline struggles to capture the full dependencies between words, and hence it often generates incorrect sentences. To alleviate this problem, we propose a model to convert Structural Logic Expressions into Natural Language (SLEtoNL). SLEtoNL converts sequential logic expressions into structural representation and leverages structural encoders to capture the dependencies between nodes. The quantitative and qualitative analyses demonstrate that our proposed method outperforms the seq2seq model, which is based on the sequential representation, and outperforms strong pretrained language models (e.g., T5, BART, GPT3) with a large margin (28.6 in BLEU3) in out-of-distribution evaluation. The data and code will be available onhttps://github.com.
Xin Wu 0003, Yi Cai 0001, Zetao Lian, Ho-fung Leung, Tao Wang 0036
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Graph-Based Information Block Detection in Infographic With Gestalt Organization Principles
abstract
An infographic is a type of visualization chart that displays pieces of information through information blocks. Existing information block detection work utilizes spatial proximity to group elements into several information blocks. However, prior studies ignore the chromatic and structural features of the infographic, resulting in incorrect omissions when detecting information blocks. To alleviate this kind of error, we use a scene graph to represent an infographic and propose a graph-based information block detection model to group elements based on Gestalt Organization Principles (spatial proximity, chromatic similarity, and structural similarity principle). We also construct a new dataset for information block detection. Quantitative and qualitative experiments show that our model can detect the information blocks in the infographic more effectively compared with the spatial proximity-based method.
Yi Cai 0001, Xin Wu 0003
IEEE Trans. Vis. Comput. Graph.3
2023 Learning refined features for open-world text classification with class description and commonsense knowledge
Haopeng Ren, Zeting Li, Yi Cai 0001, Xingwei Tan, Xin Wu 0003
World Wide Web (WWW)5
2021 Information Block Detection in Infographic Based on Spatial Proximity and Structural Similarity (Student Abstract)
abstract
The infographic is a type of visualization chart used to display information. Existing infographic understanding works utilize spatial proximity to group elements into information blocks. However, these works ignore structural features such as background color and boundary, which results in poor performance towards complex infographic. We propose Spatial and Structural Feature Extraction model to group elements based on spatial proximity and structural similarity. We introduce a new dataset towards information block detection. Experiments show that our model can effectively identify the information blocks in the infographic.
Xin Wu 0003, Yi Cai 0001
AAAI2
2020 Task-oriented Domain-specific Meta-Embedding for Text Classification
abstract
Meta-embedding learning, which combines complementary information in different word embeddings, have shown superior performances across different Natural Language Processing tasks.However, domain-specific knowledge is still ignored by existing metaembedding methods, which results in unstable performances across specific domains.Moreover, the importance of general and domain word embeddings is related to downstream tasks, how to regularize meta-embedding to adapt downstream tasks is an unsolved problem.In this paper, we propose a method to incorporate both domain-specific and taskoriented information into meta-embeddings.We conducted extensive experiments on four text classification datasets and the results show the effectiveness of our proposed method.
Xin Wu 0003, Yi Cai 0001, Kai Yang 0007, Tao Wang 0036, Qing Li 0001
EMNLP (1)1
2020 Incorporating context-relevant concepts into convolutional neural networks for short text classification
Yi Cai 0001, Xin Wu 0003, Xue Lei, Qingbao Huang, Ho-fung Leung, Qing Li 0001
Neurocomputing3
2020 Combining weighted category-aware contextual information in convolutional neural networks for text classification
Xin Wu 0003, Yi Cai 0001, Qing Li 0001, Ho-fung Leung
World Wide Web1
2018 Combining Contextual Information by Self-attention Mechanism in Convolutional Neural Networks for Text Classification
Xin Wu 0003, Yi Cai 0001, Qing Li 0001, Ho-fung Leung
WISE (1)1