Weiwen Xu

dblp:57/4640 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0002-6090-2984ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Revealing Procedural Reasoning Structures in Chain-of-Thought Training via Span-Level Gradient Organization
abstract
Jia Liu, Jiaxin Luo, Weiwen Xu, Jonathan M. Garibaldi, Xiao-Kun Wu, Yixue Hao, Min Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jia Liu 0009, Jiaxin Luo, Weiwen Xu, Jonathan M. Garibaldi, Xiaokun Wu 0004, Yixue Hao, Min Chen 0003
ACL (1)3
2026 MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation
abstract
Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengyuan Liu, Tanmoy Chakraborty 0002, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu 0071, Xing Xie 0001, Xiaoyuan Yi, Jing Yao 0003, Chaojun Wang, Rui Liu 0019, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Lingyu Ye, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen
ACL (1)4
2025 FineReason: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving
abstract
Guizhen Chen, Weiwen Xu, Hao Zhang, Hou Pong Chan, Chaoqun Liu, Lidong Bing, Deli Zhao, Anh Tuan Luu, Yu Rong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Guizhen Chen, Weiwen Xu, Hao Zhang 0048, Hou Pong Chan, Chaoqun Liu, Lidong Bing, Deli Zhao, Anh Tuan Luu, Yu Rong 0001
ACL (1)2
2025 Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks
abstract
Large language models excel at problemsolving but often struggle with complex reasoning and factual accuracy.While chainof-thought and retrieval-augmented generation help break down problems and retrieve knowledge, they still falter on challenging tasks like competitive programming due to frequent reasoning errors and irrelevant retrieval.To address this, we introduce Critic-guided planning with Retrieval-augmentation, CR-Planner, a novel framework that leverages fine-tuned critic models to guide both reasoning and retrieval processes through planning.CR-Planner iteratively selects and executes sub-goals, guided by critic models.A sub-goal critic identifies promising sub-goals from reasoning, query generation, and retrieval, while an execution critic evaluates outputs of sub-goal executions.We employ Monte Carlo Tree Search to collect data for critic training, allowing systematic exploration of action sequences and effective navigation toward the final answer.We evaluate CR-Planner on challenging domain-knowledgeintensive and reasoning-heavy tasks, including competitive programming, theorem-driven math reasoning, and complex domain retrieval problems.It significantly outperforms baselines, demonstrating effectiveness in both reasoning and retrieval.Our code is available at https://github.com/xingxuanli/CR-Planner.
Xingxuan Li, Weiwen Xu, Fangkai Jiao, Shafiq R. Joty, Lidong Bing
ACL (1)2
2025 Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
abstract
As LLMs continuously evolve, there is an urgent need for a reliable evaluation method that delivers trustworthy results promptly.Currently, static benchmarks suffer from inflexibility and unreliability, leading users to prefer human voting platforms like Chatbot Arena.However, human evaluations require significant manual effort.Therefore, we propose Auto-Arena, an innovative framework that automates the entire evaluation process using LLM-powered agents.Firstly, an LLM examiner generates questions.Then, two LLM candidates engage in a multi-round peer battle based on the questions, aiming at revealing their true performance differences.Finally, a committee of LLM judges collaboratively discusses and decides the winner, reducing bias and enhancing fairness.During the peer battles, we observe intriguing scenarios where the LLM candidates display competitive behaviors and learn from the opponents.In our extensive experiments involving 15 recent LLMs, Auto-Arena shows a 92.14% correlation with human preferences, surpassing all previous expert-annotated benchmarks without any manual efforts.Auto-Arena offers a promising alternative to current human evaluation platforms for evaluating LLMs automatically. 1
Wenxuan Zhang 0001, Yew Ken Chia, Weiwen Xu, Deli Zhao, Lidong Bing
ACL (1)4
2025 ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
abstract
Yu Sun, Xingyu Qian, Weiwen Xu, Hao Zhang, Chenghao Xiao, Long Li, Deli Zhao, Wenbing Huang, Tingyang Xu, Qifeng Bai, Yu Rong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Xingyu Qian, Weiwen Xu, Hao Zhang 0048, Chenghao Xiao, Deli Zhao, Wenbing Huang 0001, Tingyang Xu, Qifeng Bai, Yu Rong 0001
EMNLP3
2025 Scaling Language-centric Omnimodal Representation Learning
abstract
Recent multimodal embedding approaches leveraging multimodal large language models (MLLMs) fine-tuned with contrastive learning (CL) have shown promising results, yet the underlying reasons behind their superiority remain underexplored. This work argues that a crucial advantage of MLLM-based approaches stems from implicit cross-modal alignment achieved during generative pretraining, where the language decoder learns to exploit multimodal signals within a shared representation space for generating unimodal outputs. Through analysis of anisotropy and kernel similarity structure, we empirically confirm that latent alignment emerges within MLLM representations, allowing CL to serve as a lightweight refinement stage. Leveraging this insight, we propose a Language-Centric Omnimodal Embedding framework, termed LCO-Embed. Extensive experiments across diverse backbones and benchmarks demonstrate its effectiveness, achieving state-of-the-art performance across modalities. Furthermore, we identify a Generation-Representation Scaling Law (GRSL), showing that the representational capabilities gained through contrastive refinement scale positively with the MLLM's generative capabilities. This suggests that improving generative abilities evolves as an effective paradigm for enhancing representation quality. We provide a theoretical explanation of GRSL, which formally links the MLLM's generative quality to the upper bound on its representation performance, and validate it on a challenging, low-resource visual-document retrieval task, showing that continual generative pretraining before CL can further enhance the potential of a model's embedding capabilities. Codes, models, and resources are available at https://github.com/LCO-Embedding/LCO-Embedding.
Chenghao Xiao, Hou Pong Chan, Hao Zhang 0048, Weiwen Xu, Mahani Aljunied, Yu Rong 0001
NeurIPS4
2024 Nonfactoid Question Answering as Query-Focused Summarization With Graph-Enhanced Multihop Inference
abstract
Nonfactoid question answering (QA) is one of the most extensive yet challenging applications and research areas in natural language processing (NLP). Existing methods fall short of handling the long-distance and complex semantic relations between the question and the document sentences. In this work, we propose a novel query-focused summarization method, namely a graph-enhanced multihop query-focused summarizer (GMQS), to tackle the nonfactoid QA problem. Specifically, we leverage graph-enhanced reasoning techniques to elaborate the multihop inference process in nonfactoid QA. Three types of graphs with different semantic relations, namely semantic relevance, topical coherence, and coreference linking, are constructed for explicitly capturing the question-document and sentence-sentence interrelationships. Relational graph attention network (RGAT) is then developed to aggregate the multirelational information accordingly. In addition, the proposed method can be adapted to both extractive and abstractive applications as well as be mutually enhanced by joint learning. Experimental results show that the proposed method consistently outperforms both existing extractive and abstractive methods on two nonfactoid QA datasets, WikiHow and PubMedQA, and possesses the capability of performing explainable multihop reasoning.
Yang Deng 0002, Wenxuan Zhang 0001, Weiwen Xu, Ying Shen 0001, Wai Lam
IEEE Trans. Neural Networks Learn. Syst.3
2024 VisCI: A visualization framework for anomaly detection and interactive optimization of composite index
abstract
Composite index is always derived with the weighted aggregation of hierarchical components, which is widely utilized to distill intricate and multidimensional matters in economic and business statistics. However, the composite indices always present inevitable anomalies at different levels oriented from the calculation and expression processes of hierarchical components, thereby impairing the precise depiction of specific economic issues. In this paper, we propose VisCI, a visualization framework for anomaly detection and interactive optimization of composite index. First, LSTM-AE model is performed to detect anomalies from the lower level to the higher level of the composite index. Then, a comprehensive array of visual cues are designed to visualize anomalies, such as hierarchy and anomaly visualization. In addition, an interactive operation is provided to ensure accurate and efficient index optimization, mitigating the adverse impact of anomalies on index calculation and representation. Finally, we implement a visualization framework with interactive interfaces, facilitating both anomaly detection and intuitive composite index optimization. Case studies based on real-world datasets and expert interviews are conducted to demonstrate the effectiveness of our VisCI in commodity index anomaly exploration and anomaly optimization.
Zhiguang Zhou, Yuna Ni, Weiwen Xu, Guoting Hu, Ying Lai, Peixiong Chen, Weihua Su
Vis. Informatics4
2023 PeerDA: Data Augmentation via Modeling Peer Relation for Span Identification Tasks
abstract
Span identification aims at identifying specific text spans from text input and classifying them into pre-defined categories.Different from previous works that merely leverage the Subordinate (SUB) relation (i.e. if a span is an instance of a certain category) to train models, this paper for the first time explores the Peer (PR) relation, which indicates that two spans are instances of the same category and share similar features.Specifically, a novel Peer Data Augmentation (PeerDA) approach is proposed which employs span pairs with the PR relation as the augmentation data for training.PeerDA has two unique advantages: (1) There are a large number of PR span pairs for augmenting the training data.(2) The augmented data can prevent the trained model from over-fitting the superficial span-category mapping by pushing the model to leverage the span semantics.Experimental results on ten datasets over four diverse tasks across seven domains demonstrate the effectiveness of PeerDA.Notably, PeerDA achieves state-of-the-art results on six of them. 1
Weiwen Xu, Xin Li 0056, Yang Deng 0002, Wai Lam, Lidong Bing
ACL (1)1
2023 From Cloze to Comprehension: Retrofitting Pre-trained Masked Language Models to Pre-trained Machine Reader
abstract
We present Pre-trained Machine Reader (PMR), a novel method for retrofitting pre-trained masked language models (MLMs) to pre-trained machine reading comprehension (MRC) models without acquiring labeled data. PMR can resolve the discrepancy between model pre-training and downstream fine-tuning of existing MLMs. To build the proposed PMR, we constructed a large volume of general-purpose and high-quality MRC-style training data by using Wikipedia hyperlinks and designed a Wiki Anchor Extraction task to guide the MRC-style pre-training. Apart from its simplicity, PMR effectively solves extraction tasks, such as Extractive Question Answering and Named Entity Recognition. PMR shows tremendous improvements over existing approaches, especially in low-resource scenarios. When applied to the sequence classification task in the MRC formulation, PMR enables the extraction of high-quality rationales to explain the classification process, thereby providing greater prediction explainability. PMR also has the potential to serve as a unified model for tackling various extraction and classification tasks in the MRC formulation.
Weiwen Xu, Xin Li 0056, Wenxuan Zhang 0001, Wai Lam, Luo Si, Lidong Bing
NeurIPS1
2023 A Unified Multi-task Learning Framework for Multi-goal Conversational Recommender Systems
abstract
Recent years witnessed several advances in developing multi-goal conversational recommender systems (MG-CRS) that can proactively attract users’ interests and naturally lead user-engaged dialogues with multiple conversational goals and diverse topics. Four tasks are often involved in MG-CRS, including Goal Planning, Topic Prediction, Item Recommendation, and Response Generation. Most existing studies address only some of these tasks. To handle the whole problem of MG-CRS, modularized frameworks are adopted where each task is tackled independently without considering their interdependencies. In this work, we propose a novel Unified MultI-goal conversational recommeNDer system (UniMIND). Specifically, we unify these four tasks with different formulations into the same sequence-to-sequence paradigm. Prompt-based learning strategies are investigated to endow the unified model with the capability of multi-task learning. Finally, the overall learning and inference procedure consists of three stages, including multi-task learning, prompt-based tuning, and inference. Experimental results on two MG-CRS benchmarks (DuRecDial and TG-ReDial) show that UniMIND achieves state-of-the-art performance on all tasks with a unified model. Extensive analyses and discussions are provided for shedding some new perspectives for MG-CRS.
Yang Deng 0002, Wenxuan Zhang 0001, Weiwen Xu, Wenqiang Lei, Tat-Seng Chua, Wai Lam
ACM Trans. Inf. Syst.3
2022 ConReader: Exploring Implicit Relations in Contracts for Contract Clause Extraction
abstract
We study automatic Contract Clause Extraction (CCE) by modeling implicit relations in legal contracts.Existing CCE methods mostly treat contracts as plain text, creating a substantial barrier to understanding contracts of high complexity.In this work, we first comprehensively analyze the complexity issues of contracts and distill out three implicit relations commonly found in contracts, namely, 1) Long-range Context Relation that captures the correlations of distant clauses; 2) Term-Definition Relation that captures the relation between important terms with their corresponding definitions; and 3) Similar Clause Relation that captures the similarities between clauses of the same type.Then we propose a novel framework ConReader to exploit the above three relations for better contract understanding and improving CCE.Experimental results show that ConReader makes the prediction more interpretable and achieves new state-of-the-art on two CCE tasks in both conventional and zero-shot settings.1
Weiwen Xu, Yang Deng 0002, Wenqiang Lei, Wenlong Zhao 0012, Tat-Seng Chua, Wai Lam
EMNLP1
2019 Revisit Automatic Error Detection for Wrong and Missing Translation - A Supervised Approach
abstract
Wenqiang Lei, Weiwen Xu, Ai Ti Aw, Yuanxin Xiang, Tat Seng Chua. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Wenqiang Lei, Weiwen Xu, AiTi Aw, Yuanxin Xiang, Tat-Seng Chua
EMNLP/IJCNLP (1)2
2009 Timed verification of the generic architecture of a memory circuit using parametric timed automata
Remy Chevallier, Emmanuelle Encrenaz-Tiphène, Laurent Fribourg, Weiwen Xu
Formal Methods Syst. Des.4