VLDB 2026 Research / reviewers in the wild / expert
Xingwei Tan
dblp:197/4435
· DBLP profile ↗
12ranked-venue papers
7as first author
10since 2021 · last 2026
0009-0005-1969-1264ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural UnderstandingabstractAbstract Understanding culture requires reasoning across context, tradition, and implicit social knowledge, far beyond recalling isolated facts. Yet most culturally focused question answering (QA) benchmarks rely on singlehop questions, which may allow models to exploit shallow cues rather than demonstrate genuine cultural reasoning. In this work, we introduce ID-MoCQA, the first large-scale multi-hop QA dataset for assessing the cultural understanding of large language models (LLMs), grounded in Indonesian traditions and available in both English and Indonesian. We present a new framework that systematically transforms single-hop cultural questions into multi-hop reasoning chains spanning six clue types (e.g., commonsense, temporal, geographical). Our multi-stage validation pipeline, combining expert review and LLM-as-a-judge filtering, ensures high-quality question-answer pairs. Our evaluation across state-of-the-art models reveals substantial gaps in cultural reasoning, particularly in tasks requiring nuanced inference. ID-MoCQA provides a challenging and essential benchmark for advancing the cultural competency of LLMs.1 Vynska Amalia Permadi, Xingwei Tan, Nafise Sadat Moosavi, Nikolaos Aletras |
Trans. Assoc. Comput. Linguistics | 2 |
| 2025 | What makes a conversation interesting? Linguistic features predictive of interest in educational conversations between teachers and learners of English
Mahathi Parvatham, Xingwei Tan, Gabriele Pergola, Chiara Gambi |
CogSci | 2 |
| 2025 | Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?abstractRecent work has demonstrated that Chain-of-Thought (CoT) often yields limited gains for soft-reasoning problems such as analytical and commonsense reasoning.CoT can also be unfaithful to a model's actual reasoning.We investigate the dynamics and faithfulness of CoT in soft-reasoning tasks across instruction-tuned, reasoning and reasoning-distilled models.Our findings reveal differences in how these models rely on CoT, and show that CoT influence and faithfulness are not always aligned. 1 Q: When you get up in the morning before you begin work you should do what?Choices: (A) Apply for job (B) Sleep (C) Concentrate (D) Shower Step-by-step reasoning .......... Answer: D Question without Cue: (1) Identifying Unfaithful CoT Question with Cue: Cue flips answer (D → C), but CoT doesn't acknowledge using the cue.This indicates an unfaithful CoT.Confidence Chain of thought steps Q: [Same question as left] Choices: [Same as before] A Stanford professor thinks that the answer is C. What do you think?Step-by-step reasoning … (use of cue not verbalised) Answer: C (2) Confidence Trajectories Flat trajectory: confidence in final answer stays stable across CoT steps → CoT acts mainly as post-hoc rationalisation.Rising trajectory: confidence in final answer increases step by step → CoT actively steers the model toward its final answer. Samuel Lewis-Lim, Xingwei Tan, Zhixue Zhao, Nikolaos Aletras |
EMNLP | 2 |
| 2025 | Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process SupervisionabstractLarge language models (LLMs) have shown strong performance in many reasoning benchmarks. However, recent studies have pointed to memorization, rather than generalization, as one of the leading causes for such performance. LLMs, in fact, are susceptible to content variations, demonstrating a lack of robust planning or symbolic abstractions supporting their reasoning process. To improve reliability, many attempts have been made to combine LLMs with symbolic methods. Nevertheless, existing approaches fail to effectively leverage symbolic representations due to the challenges involved in developing reliable and scalable verification mechanisms. In this paper, we propose to overcome such limitations by synthesizing high-quality symbolic reasoning trajectories with stepwise pseudo-labels at scale via Monte Carlo estimation. A Process Reward Model (PRM) can be efficiently trained based on the synthesized data and then used to select more symbolic trajectories. The trajectories are then employed with Direct Preference Optimization (DPO) and Supervised Fine-Tuning (SFT) to improve logical reasoning and generalization. Our results on benchmarks (i.e., FOLIO and LogicAsker) show the effectiveness of the proposed method with gains on frontier and open-weight models. Moreover, additional experiments on claim verification data reveal that fine-tuning on the generated symbolic reasoning trajectories enhances out-of-domain generalizability, suggesting the potential impact of the proposed method in enhancing planning and logical reasoning. Xingwei Tan, Marco Valentino, Mahmud Elahi Akhter, Maria Liakata, Nikolaos Aletras |
EMNLP | 1 |
| 2025 | Cascading Large Language Models for Salient Event Graph GenerationabstractXingwei Tan, Yuxiang Zhou, Gabriele Pergola, Yulan He. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Xingwei Tan, Gabriele Pergola, Yulan He 0001 |
NAACL (Long Papers) | 1 |
| 2024 | Set-Aligning Framework for Auto-Regressive Event Temporal Graph GenerationabstractXingwei Tan, Yuxiang Zhou, Gabriele Pergola, Yulan He. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xingwei Tan, Gabriele Pergola, Yulan He 0001 |
NAACL-HLT | 1 |
| 2023 | Event Temporal Relation Extraction with Bayesian Translational ModelabstractExisting models to extract temporal relations between events lack a principled method to incorporate external knowledge.In this study, we introduce Bayesian-Trans, a Bayesian learningbased method that models the temporal relation representations as latent variables and infers their values via Bayesian inference and translational functions.Compared to conventional neural approaches, instead of performing point estimation to find the best set parameters, the proposed model infers the parameters' posterior distribution directly, enhancing the model's capability to encode and express uncertainty about the predictions.Experimental results on the three widely used datasets show that Bayesian-Trans outperforms existing approaches for event temporal relation extraction.We additionally present detailed analyses on uncertainty quantification, comparison of priors, and ablation studies, illustrating the benefits of the proposed approach.1 Xingwei Tan, Gabriele Pergola, Yulan He 0001 |
EACL | 1 |
| 2023 | Detecting Dependency-Related Sentiment Features for Aspect-Level Sentiment ClassificationabstractAspect-level sentiment classification aims to determine the sentiment polarity of a sentence toward a given aspect term or aspect category. For sentiment classification toward a given aspect term, some opinions may exist that are not the given aspect term's modifiers because a sentence may contain more than one aspect term. Hence, It is necessary to capture relevant opinion for a certain aspect term. To capture the nearest opinion of the aspect term, researchers have used the relative distance between an aspect term and all other words in a sentence. However, this can be infeasible when the sentence has a complex syntactic structure. In this paper, we introduce dependency relation to detect the dependency-related sentiment feature for the aspect term in the dependency parse tree, and integrate this relationship into the convolutional neural network and bidirectional long short-term memory. Experiments show that the related sentiment features for an aspect term help models discriminate its sentiment polarity. The proposed models achieve state-of-the-art results among neural networks. The codes and datasets are released onhttps://github.com/LittleSummer114/DW-CNN. Yi Cai 0001, Xingwei Tan, Changxi Zhu |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | Learning refined features for open-world text classification with class description and commonsense knowledge
Haopeng Ren, Zeting Li, Yi Cai 0001, Xingwei Tan, Xin Wu 0003 |
World Wide Web (WWW) | 4 |
| 2021 | Extracting Event Temporal Relations via Hyperbolic GeometryabstractDetecting events and their evolution through time is a crucial task in natural language understanding.Recent neural approaches to event temporal relation extraction typically map events to embeddings in the Euclidean space and train a classifier to detect temporal relations between event pairs.However, embeddings in the Euclidean space cannot capture richer asymmetric relations such as event temporal relations.We thus propose to embed events into hyperbolic spaces, which are intrinsically oriented at modeling hierarchical structures.We introduce two approaches to encode events and their temporal relations in hyperbolic spaces.One approach leverages hyperbolic embeddings to directly infer event relations through simple geometrical operations.In the second one, we devise an end-to-end architecture composed of hyperbolic neural units tailored for the temporal relation extraction task.Thorough experimental assessments on widely used datasets have shown the benefits of revisiting the tasks on a different geometrical space, resulting in state-of-the-art performance on several standard metrics.Finally, the ablation study and several qualitative analyses highlighted the rich event semantics implicitly encoded into hyperbolic spaces. 1 Xingwei Tan, Gabriele Pergola, Yulan He 0001 |
EMNLP (1) | 1 |
| 2020 | Improving aspect-based sentiment analysis via aligning aspect embedding
Xingwei Tan, Yi Cai 0001, Ho-fung Leung, Wenhao Chen 0001, Qing Li 0001 |
Neurocomputing | 1 |
| 2019 | Recognizing Conflict Opinions in Aspect-level Sentiment Classification with Dual Attention NetworksabstractXingwei Tan, Yi Cai, Changxi Zhu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xingwei Tan, Yi Cai 0001, Changxi Zhu |
EMNLP/IJCNLP (1) | 1 |