Xintao Wang 0001

dblp:205/3991-1 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-1428-8677ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
abstract
Xintao Wang, Jian Yang, Weiyuan Li, Rui Xie, Jen-tse Huang, Jun Gao, Shuai Huang, Yueping Kang, Yuanli Guo, Hongwei Feng, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xintao Wang 0001, Jian Yang 0003, Weiyuan Li, Rui Xie 0005, Jen-tse Huang 0001, Yueping Kang, Yuanli Guo, Hongwei Feng, Yanghua Xiao
ACL (1)1
2026 Can LLMs Learn to Map the World from Local Descriptions?
abstract
Sirui Xia, Aili Chen, Xintao Wang, Tinghui Zhu, Yikai Zhang, Jiangjie Chen, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Sirui Xia, Aili Chen, Xintao Wang 0001, Tinghui Zhu, Yikai Zhang 0004, Jiangjie Chen, Yanghua Xiao
ACL (1)3
2025 BOOKWORLD: From Novels to Interactive Agent Societies for Story Creation
abstract
Recent advances in large language models (LLMs) have enabled social simulation through multi-agent systems.Prior efforts focus on agent societies created from scratch, assigning agents with newly defined personas.However, simulating established fictional worlds and characters remain largely underexplored, despite its significant practical value.In this paper, we introduce BOOKWORLD, a comprehensive system for constructing and simulating book-based multi-agent societies.BOOK-WORLD's design covers comprehensive realworld intricacies, including diverse and dynamic characters, fictional worldviews, geographical constraints and changes, e.t.c.BOOK-WORLD enables diverse applications including story generation, interactive games and social simulation, offering novel ways to extend and explore beloved fictional works.Through extensive experiments, we demonstrate that BOOKWORLD generates creative, high-quality stories while maintaining fidelity to the source books, surpassing previous methods with a win rate of 75.36%.The code and demo of this paper can be found at the project page: https://bookworld2025.github.io/.
Yiting Ran, Xintao Wang 0001, Jiaqing Liang, Yanghua Xiao, Deqing Yang
ACL (1)2
2025 CoSER: Coordinating LLM-Based Persona Simulation of Established Roles
abstract
Role-playing language agents (RPLAs) have emerged as promising applications of large language models (LLMs). However, simulating established characters presents a challenging task for RPLAs, due to the lack of authentic character datasets and nuanced evaluation methods using such data. In this paper, we present CoSER, a collection of a high-quality dataset, open models, and an evaluation protocol towards effective RPLAs of established characters. The CoSER dataset covers 17,966 characters from 771 renowned books. It provides authentic dialogues with real-world intricacies, as well as diverse data types such as character experiences and internal thoughts. Drawing from acting methodology, we introduce given-circumstance acting for training and evaluating role-playing LLMs, where LLMs sequentially portray multiple characters in book scenes. Using our dataset, we develop CoSER 8B and CoSER 70B, i.e., advanced open role-playing LLMs built on LLaMA-3.1 models. Extensive experiments demonstrate the value of the CoSER dataset for RPLA training, evaluation and retrieval. Moreover, CoSER 70B exhibits state-of-the-art performance surpassing or matching GPT-4o on our evaluation and three existing benchmarks, i.e., achieving 75.80% and 93.47% accuracy on the InCharacter and LifeChoice benchmarks respectively. Our code, dataset and models are available at: https://github.com/Neph0s/CoSER.
Xintao Wang 0001, Xinfeng Yuan, Rui Xu 0026, Jen-tse Huang 0001, Haoran Guo, Jiangjie Chen, Shuchang Zhou 0003, Wei Wang 0009, Yanghua Xiao
ICML1
2025 ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
abstract
Spatial reasoning is a key capability in the field of artificial intelligence, especially crucial in areas such as robotics, computer vision, and natural language understanding. However, evaluating the ability of multimodal large language models (MLLMs) in complex spatial reasoning still faces challenges, particularly in scenarios requiring multi-step reasoning and precise mathematical constraints. This paper introduces ORIGAMISPACE, a new dataset and benchmark designed to evaluate the multi-step spatial reasoning ability and the capacity to handle mathematical constraints of MLLMs through origami tasks. The dataset contains 350 data instances, each comprising a strictly formatted crease pattern (CP diagram), the Compiled Flat Pattern, the complete Folding Process, and the final Folded Shape Image. We propose four evaluation tasks: Pattern Prediction, Multi-step Spatial Reasoning, Spatial Relationship Prediction, and End-to-End CP Code Generation. For the CP code generation task, we design an interactive environment and explore the possibility of using reinforcement learning methods to train MLLMs. Through experiments on existing MLLMs, we initially reveal the strengths and weaknesses of these models in handling complex spatial reasoning tasks.
Rui Xu 0026, Dakuan Lu, Zicheng Zhao, Xiaoyu Tan, Xintao Wang 0001, Jiangjie Chen, Yinghui Xu 0001
NeurIPS5
2025 ARIA: Training Language Agents with Intention-driven Reward Aggregation
abstract
Large language models (LLMs) have enabled agents to perform complex reasoning and decision-making through free-form language interactions. However, in open-ended language action environments (e.g., negotiation or question-asking games), the action space can be formulated as a joint distribution over tokens, resulting in an extremely large and combinatorial action space. Sampling actions in such a space can lead to extreme reward sparsity, which brings large reward variance, hindering effective reinforcement learning (RL). To address this, we propose **ARIA**, a method that **A**ggregates **R**ewards in **I**ntention space to enable efficient and effective language **A**gents training. ARIA aims to project natural language actions from the high-dimensional joint token distribution space into a low-dimensional intention space, where semantically similar actions are clustered and assigned shared rewards. This intention-aware reward aggregation reduces reward variance by densifying reward signals, fostering efficient and effective policy optimization. Extensive experiments demonstrate that ARIA not only significantly reduces gradient variance, but also delivers substantial performance gains of average 9.95% across four downstream tasks (e.g., negotiation and text-based games), consistently outperforming strong offline and online RL baselines.
Ruihan Yang, Yikai Zhang 0004, Aili Chen, Xintao Wang 0001, Jiangjie Chen, Deqing Yang, Yanghua Xiao
NeurIPS4
2024 InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews
abstract
Xintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, Jiangjie Chen, Cheng Li, Yanghua Xiao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xintao Wang 0001, Yunze Xiao, Jen-tse Huang 0001, Rui Xu 0026, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang 0009, Jiangjie Chen, Yanghua Xiao
ACL (1)1
2024 Source Prompt: Coordinated Pre-training of Language Models on Diverse Corpora from Multiple Sources
abstract
Pre-trained language models (PLMs) have established the new paradigm in the field of NLP. For more powerful PLMs, one of the most popular and successful ways is to continuously scale up sizes of the models and the pre-training corpora. These large corpora, typically obtained by converging smaller ones from multiple sources, are thus growing increasingly diverse. However, colossal converged corpora don't always enhance PLMs' performance. In this paper, we identify the disadvantage of heterogeneous corpora from multiple sources for pre-training PLMs. Towards coordinated pre-training on diverse corpora, we further propose Source Prompt (SP), which explicitly prompt the model with the source of data at the pre-training and fine-tuning stages. Extensive experimental results show that pre-training PLMs with SP on diverse corpora significantly improves performance in various downstream tasks.
Yipei Xu, Dakuan Lu, Jiaqing Liang, Jin Zhao 0004, Xintao Wang 0001, Hengkui Wu, Liujiang Liu, Yingsi Xin, Xuepeng Liu, Yanghua Xiao, Zhixu Li
CIKM5
2024 Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional Works
abstract
Large language models (LLMs) have demonstrated impressive performance and spurred numerous AI applications, in which role-playing agents (RPAs) are particularly popular, especially for fictional characters.The prerequisite for these RPAs lies in the capability of LLMs to understand characters from fictional works.Previous efforts have evaluated this capability via basic classification tasks or characteristic imitation, failing to capture the nuanced character understanding with LLMs.In this paper, we propose evaluating LLMs' character understanding capability via the character profiling task, i.e., summarizing character profiles from corresponding materials, a widely adopted yet understudied practice for RPA development.Specifically, we construct the CROSS dataset from literature experts and assess the generated profiles by comparing them with ground truth references and evaluating their applicability in downstream tasks.Our experiments, which cover various summarization methods and LLMs, have yielded promising results.These results strongly validate the character understanding capability of LLMs.Resources are available at https://github. com/Joanna0123/character_profiling.
Xinfeng Yuan, Yuhan Cui, Tianhe Lin, Xintao Wang 0001, Rui Xu 0026, Jiangjie Chen, Deqing Yang
EMNLP5
2023 MAPS-KB: A Million-Scale Probabilistic Simile Knowledge Base
abstract
The ability to understand and generate similes is an imperative step to realize human-level AI. However, there is still a considerable gap between machine intelligence and human cognition in similes, since deep models based on statistical distribution tend to favour high-frequency similes. Hence, a large-scale symbolic knowledge base of similes is required, as it contributes to the modeling of diverse yet unpopular similes while facilitating additional evaluation and reasoning. To bridge the gap, we propose a novel framework for large-scale simile knowledge base construction, as well as two probabilistic metrics which enable an improved understanding of simile phenomena in natural language. Overall, we construct MAPS-KB, a million-scale probabilistic simile knowledge base, covering 4.3 million triplets over 0.4 million terms from 70 GB corpora. We conduct sufficient experiments to justify the effectiveness and necessity of the methods of our framework. We also apply MAPS-KB on three downstream tasks to achieve state-of-the-art performance, further demonstrating the value of MAPS-KB. Resources of MAPS-KB are publicly available at https://github.com/Abbey4799/MAPS-KB.
Qianyu He, Xintao Wang 0001, Jiaqing Liang, Yanghua Xiao
AAAI2
2023 Prototypical Concept Representation
abstract
Concepts are building blocks of human thinking. For machines, concept understanding has also been increasingly important, which makes concept representation a fundamental problem in artificial intelligence. While many concepts have their instances, the massive amount of information carried by instances has long been ignored in current concept representation, which limits the usage of these concepts in applications. In this paper, inspired by prototype theory in cognitive science, we propose prototypical concept representation for machines, which represents each concept with a distributed prototype derived from representations of its instances. For prototypical representation learning, we further introduce a novel model named Prototypical Siamese Network (PSN). PSN is trained under the supervision ofisAdetermination, one of the most important concept-related applications. Results of extensive experiments demonstrate that, our method achieves state-of-the-art performance, thus validating the effectiveness of prototypical concept representation.
Xintao Wang 0001, Jiaqing Liang, Yanghua Xiao, Wei Wang 0009
IEEE Trans. Knowl. Data Eng.1
2022 Language Models as Knowledge Embeddings
abstract
Knowledge embeddings (KE) represent a knowledge graph (KG) by embedding entities and relations into continuous vector spaces. Existing methods are mainly structure-based or description-based. Structure-based methods learn representations that preserve the inherent structure of KGs. They cannot well represent abundant long-tail entities in real-world KGs with limited structural information. Description-based methods leverage textual information and language models. Prior approaches in this direction barely outperform structure-based ones, and suffer from problems like expensive negative sampling and restrictive description demand. In this paper, we propose LMKE, which adopts Language Models to derive Knowledge Embeddings, aiming at both enriching representations of long-tail entities and solving problems of prior description-based methods. We formulate description-based KE learning with a contrastive learning framework to improve efficiency in training and evaluation. Experimental results show that LMKE achieves state-of-the-art performance on KE benchmarks of link prediction and triple classification, especially for long-tail entities.
Xintao Wang 0001, Qianyu He, Jiaqing Liang, Yanghua Xiao
IJCAI1