VLDB 2026 Research / reviewers in the wild / expert
Yunze Xiao
dblp:310/1568
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AniTales: End-to-End Multimodal Story Generation Through Natural Language Prompting (Student Abstract)abstractWe present AniTales, a system designed to generate multimodal visual novels from natural language prompts. Our system integrates large language models for story generation, diffusion models for character art, and text-to-speech for voice acting. This paper describes the system's architecture and presents findings from a pilot user study. We evaluated the system with general users (n=10) and domain experts (n=5), focusing on usability, coherence, and visual consistency. General users reported high usability (SUS: 84/100) and strong character-dialogue consistency (4.2/5), along with an average score of 82/100 for their intention to continue using the platform. These initial results suggest AniTales is a promising approach for bridging the gap between text-based AI storytelling and end-to-end multimedia content creation. Mrigendra Agrawal, Yunze Xiao |
AAAI | 2 |
| 2026 | The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use AgentsabstractAutonomous agents based on large language models (LLMs) are rapidly evolving to handle multi-turn tasks, but ensuring their trustworthiness remains a critical challenge.A fundamental pillar of this trustworthiness is calibration, which refers to an agent's ability to express confidence that reliably reflects its actual performance.While calibration is well-established for static models, its dynamics in tool-integrated agentic workflows remain under-explored.In this work, we systematically investigate verbalized calibration in tooluse agents, revealing a fundamental confidence dichotomy driven by tool type.Specifically, our pilot study identifies that evidence tools (e.g., web search) systematically induce severe overconfidence due to inherent noise in retrieved information, while verification tools (e.g., code interpreters) can ground reasoning through deterministic feedback and mitigate miscalibration.To robustly improve calibration across tool types, we propose a reinforcement learning (RL) fine-tuning framework that jointly optimizes task accuracy and calibration, supported by a holistic benchmark of reward designs.We demonstrate that our trained agents not only achieve superior calibration but also exhibit robust generalization from local training environments to noisy web settings and to distinct domains such as mathematical reasoning.Our results highlight the necessity of domain-specific calibration strategies for tooluse agents.More broadly, this work establishes a foundation for building self-aware agents that can reliably communicate uncertainty in highstakes, real-world deployments. Weihao Xuan, Qingcheng Zeng, Heli Qi, Yunze Xiao, Naoto Yokoya |
ACL (1) | 4 |
| 2025 | Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion DynamicsabstractJiarui Liu, Yueqi Song, Yunze Xiao, Mingqian Zheng, Lindia Tjuatja, Jana Schaich Borg, Mona T. Diab, Maarten Sap. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jiarui Liu 0004, Yueqi Song, Yunze Xiao, Mingqian Zheng, Lindia Tjuatja, Jana Schaich Borg, Mona T. Diab, Maarten Sap |
EMNLP | 3 |
| 2025 | Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of DesignabstractLarge Language Models (LLMs) increasingly exhibit anthropomorphism characteristicshuman-like qualities portrayed across their outlook, language, behavior, and reasoning functions.Such characteristics enable more intuitive and engaging human-AI interactions.However, current research on anthropomorphism remains predominantly risk-focused, emphasizing over-trust and user deception while offering limited design guidance.We argue that anthropomorphism should instead be treated as a concept of design that can be intentionally tuned to support user goals.Drawing from multiple disciplines, we propose that the anthropomorphism of an LLM-based artifact should reflect the interaction between artifact designers and interpreters.This interaction is facilitated by cues embedded in the artifact by the designers and the (cognitive) responses of the interpreters to the cues.Cues are categorized into four dimensions: perceptive, linguistic, behavioral, and cognitive.By analyzing the manifestation and effectiveness of each cue, we provide a unified taxonomy with actionable levers for practitioners.Consequently, we advocate for function-oriented evaluations of anthropomorphic design. Yunze Xiao, Lynnette Hui Xian Ng, Jiarui Liu 0004, Mona T. Diab |
EMNLP | 1 |
| 2025 | MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model EvaluationabstractWeihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Weihao Xuan, Rui Yang 0016, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing 0001, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li 0079, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen 0001, Douglas Teodoro, Nan Liu 0003, Randy Goebel, Lei Ma 0003, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li |
EMNLP | 5 |
| 2025 | Embracing Contradiction: Theoretical Inconsistency Will Not Impede the Road of Building Responsible AI SystemsabstractThis position paper argues that the theoretical inconsistency often observed among Responsible AI (RAI) metrics, such as differing fairness definitions or trade-offs between accuracy and privacy, should be embraced as a valuable feature rather than a flaw to be eliminated. We contend that navigating these inconsistencies, by treating metrics as divergent objectives, yields three key benefits: (1) Normative Pluralism: maintaining a full suite of potentially contradictory metrics ensures that the diverse moral stances and stakeholder values inherent in RAI are adequately represented; (2) Epistemological Completeness: using multiple, sometimes conflicting, metrics captures multifaceted ethical concepts more fully and preserves greater informational fidelity than any single, simplified definition; (3) Implicit Regularization: jointly optimizing for theoretically conflicting objectives discourages overfitting to any one metric, steering models toward solutions with better generalization and robustness under real-world complexities. In contrast, enforcing theoretical consistency by simplifying or pruning metrics risks narrowing value diversity, losing conceptual depth, and degrading model performance. We therefore advocate a shift in RAI theory and practice: from getting trapped by metric inconsistencies to establishing practice-focused theories, documenting the normative provenance and inconsistency levels of inconsistent metrics, and elucidating the mechanisms that permit robust, approximated consistency in practice. Gordon Dai, Yunze Xiao |
NeurIPS | 2 |
| 2024 | InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological InterviewsabstractXintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, Jiangjie Chen, Cheng Li, Yanghua Xiao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xintao Wang 0001, Yunze Xiao, Jen-tse Huang 0001, Rui Xu 0026, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang 0009, Jiangjie Chen, Yanghua Xiao |
ACL (1) | 2 |
| 2024 | Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMsabstractLexical-syntactic flexibility, in the form of conversion (or zero-derivation) is a hallmark of English morphology. In conversion, a word with one part of speech is placed in a non-prototypical context, where it is coerced to behave as if it had a different part of speech. However, while this process affects a large part of the English lexicon, little work has been done to establish the degree to which language models capture this type of generalization. This paper reports the first study on the behavior of large language models with reference to conversion. We design a task for testing lexical-syntactic flexibility—the degree to which models can generalize over words in a construction with a non-prototypical part of speech. This task is situated within a natural language inference paradigm. We test the abilities of five language models—two proprietary models (GPT-3.5 and GPT-4), three open source model (Mistral 7B, Falcon 40B, and Llama 2 70B). We find that GPT-4 performs best on the task, followed by GPT-3.5, but that the open source language models are also able to perform it and that the 7-billion parameter Mistral displays as little difference between its baseline performance on the natural language inference task and the non-prototypical syntactic category task, as the massive GPT-4. David R. Mortensen, Valentina Izrailevitch, Yunze Xiao, Hinrich Schütze, Leonie Weissweiler |
LREC/COLING | 3 |
| 2024 | ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking PerturbationsabstractDetecting hate speech and offensive language is essential for maintaining a safe and respectful digital environment.This study examines the limitations of state-of-the-art large language models (LLMs) in identifying offensive content within systematically perturbed data, with a focus on Chinese, a language particularly susceptible to such perturbations.We introduce ToxiCloakCN 1 , an enhanced dataset derived from ToxiCN, augmented with homophonic substitutions and emoji transformations, to test the robustness of LLMs against these cloaking perturbations.Our findings reveal that existing models significantly underperform in detecting offensive content when these perturbations are applied.We provide an in-depth analysis of how different types of offensive content are affected by these perturbations and explore the alignment between human and model explanations of offensiveness.Our work highlights the urgent need for more advanced techniques in offensive language detection to combat the evolving tactics used to evade detection mechanisms. Yunze Xiao, Kenny T. W. Choo, Roy Ka-Wei Lee |
EMNLP | 1 |
| 2022 | Detailed Facial Geometry Recovery from Multi-View Images by Learning an Implicit FunctionabstractRecovering detailed facial geometry from a set of calibrated multi-view images is valuable for its wide range of applications. Traditional multi-view stereo (MVS) methods adopt an optimization-based scheme to regularize the matching cost. Recently, learning-based methods integrate all these into an end-to-end neural network and show superiority of efficiency. In this paper, we propose a novel architecture to recover extremely detailed 3D faces within dozens of seconds. Unlike previous learning-based methods that regularize the cost volume via 3D CNN, we propose to learn an implicit function for regressing the matching cost. By fitting a 3D morphable model from multi-view images, the features of multiple images are extracted and aggregated in the mesh-attached UV space, which makes the implicit function more effective in recovering detailed facial shape. Our method outperforms SOTA learning-based MVS in accuracy by a large margin on the FaceScape dataset. The code and data are released in https://github.com/zhuhao-nju/mvfr. Yunze Xiao, Hao Zhu 0004, Zhengyu Diao, Xiangju Lu, Xun Cao |
AAAI | 1 |