VLDB 2026 Research / reviewers in the wild / expert
Yingxiu Zhao
dblp:194/4993
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0001-8236-3486ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How Foundational Skills Influence VLM-based Embodied Agents: A Native PerspectiveabstractRecent advances in vision–language models (VLMs) have shed light on human-level embodied intelligence. However, existing benchmarks for VLM-driven embodied agents still rely on high-level commands or discretised action spaces—``non-native'' settings that diverge markedly from the real world. Moreover, current benchmarks focus exclusively on high-level tasks, while lacking joint evaluation and analysis on both low- and high-level. To bridge these gaps, we present \textbf{NativeEmbodied}, a challenging benchmark for VLM-driven embodied agents that adopts a unified, native low-level action space. Built upon diverse simulated scenes, NativeEmbodied first designs three representative high-level tasks in complex scenarios to evaluate overall performance. For more detailed and comprehensive performance analysis, we further decouple the entangled skills behind complex tasks and construct four types of low-level tasks, each corresponding to a key fundamental embodied skill. This joint evaluation across task and skill granularities enables a fine-grained assessment of embodied agent. Comprehensive experiments on the best VLMs reveal pronounced deficiencies in certain fundamental embodied skills. Further analysis shows that these bottlenecks severely constrain performance on high-level tasks. Our NativeEmbodied not only pinpoints the key challenges faced by current VLM-driven embodied agents, but also provides valuable insight for future development of this field. Pi Bu, Keyu Pan, Xinrun Xu, Yingxiu Zhao, Tong Xu 0001 |
AAAI | 5 |
| 2026 | Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic TrainingabstractJihao Gu, Qihang Ai, Yingyao Wang, Pi Bu, Jingxuan Xing, Yue Cao, Zekun Zhu, Wei Jiang, Ziming Wang, Yingxiu Zhao, Ming-Liang Zhang, Jun Song, Yuning Jiang, Bo Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jihao Gu, Qihang Ai, Yingyao Wang, Pi Bu, Jingxuan Xing, Zekun Zhu, Yingxiu Zhao, Yuning Jiang 0001, Bo Zheng 0007 |
ACL (1) | 10 |
| 2025 | CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing GamesabstractRecent advances in Vision-Language-Action models (VLAs) have expanded the capabilities of embodied intelligence. However, significant challenges remain in real-time decision-making in complex 3D environments, which demand second-level responses, high-resolution perception, and tactical reasoning under dynamic conditions. To advance the field, we introduce CombatVLA, an efficient VLA model optimized for combat tasks in 3D action role-playing games(ARPGs). Specifically, our CombatVLA is a 3B model trained on video-action pairs collected by an action tracker, where the data is formatted as action-of-thought (AoT) sequences. Thereafter, CombatVLA seamlessly integrates into an action execution framework, allowing efficient inference through our truncated AoT strategy. Experimental results demonstrate that CombatVLA not only outperforms all existing models on the combat understanding benchmark but also achieves a 50-fold acceleration in game combat. Moreover, it has a higher task success rate than human players. We will open-source all resources, including the action tracker, dataset, benchmark, model weights, training code, and the implementation of the framework at https://combatvla.github.io/. Pi Bu, Yingyao Wang, Yingxiu Zhao, Siran Yang, Jiamang Wang |
ICCV | 7 |
| 2025 | Are Large Language Models Ready for Multi-Turn Tabular Data Analysis?abstractConversational Tabular Data Analysis, a collaboration between humans and machines, enables real-time data exploration for informed decision-making. The challenges and costs of collecting realistic conversational logs for tabular data analysis hinder comprehensive quantitative evaluation of Large Language Models (LLMs) in this task. To mitigate this issue, we introduce CoTA, a new benchmark to evaluate LLMs on conversational tabular data analysis. CoTA contains 1013 conversations, covering 4 practical scenarios: Normal, Action, Private, and Private Action. Notably, CoTA is constructed by an economical multi-agent environment, Decision Company, with few human efforts. This environment ensures efficiency and scalability of generating new conversational data. Our comprehensive study, conducted by data analysis experts, demonstrates that Decision Company is capable of producing diverse and high-quality data, laying the groundwork for efficient data annotation. We evaluate popular and advanced LLMs in CoTA, which highlights the challenges of conversational tabular data analysis. Furthermore, we propose Adaptive Conversation Reflection (ACR), a self-generated reflection strategy that guides LLMs to learn from successful histories. Experiments demonstrate that ACR can evolve LLMs into effective conversational data analysis agents, achieving a relative performance improvement of up to 35.14%. Jinyang Li 0003, Nan Huo, Yan Gao 0002, Yingxiu Zhao, Ge Qu, Bowen Qin, Yurong Wu, Xiaodong Li 0009, Chenhao Ma 0001, Jian-Guang Lou, Reynold Cheng |
ICML | 5 |
| 2024 | Tree-Instruct: A Preliminary Study of the Intrinsic Relationship between Complexity and AlignmentabstractTraining large language models (LLMs) with open-domain instruction data has yielded remarkable success in aligning to end tasks and human preferences. Extensive research has highlighted the importance of the quality and diversity of instruction data. However, the impact of data complexity, as a crucial metric, remains relatively unexplored from three aspects: (1)where the sustainability of performance improvements with increasing complexity is uncertain; (2)whether the improvement brought by complexity merely comes from introducing more training tokens; and (3)where the potential benefits of incorporating instructions from easy to difficult are not yet fully understood. In this paper, we propose Tree-Instruct to systematically enhance the instruction complexity in a controllable manner. By adding a specified number of nodes to instructions’ semantic trees, this approach not only yields new instruction data from the modified tree but also allows us to control the difficulty level of modified instructions. Our preliminary experiments reveal the following insights: (1)Increasing complexity consistently leads to sustained performance improvements of LLMs. (2)Under the same token budget, a few complex instructions outperform diverse yet simple instructions. (3)Curriculum instruction tuning might not yield the anticipated results; focusing on increasing complexity appears to be the key. Yingxiu Zhao, Bowen Yu 0002, Binyuan Hui, Haiyang Yu 0003, Fei Huang 0002, Nevin Lianwen Zhang, Yongbin Li 0001 |
LREC/COLING | 1 |
| 2024 | Automatic Instruction Evolving for Large Language ModelsabstractFine-tuning large pre-trained language models with Evol-Instruct has achieved encouraging results across a wide range of tasks.However, designing effective evolving methods for instruction evolution requires substantial human expertise.This paper proposes Auto Evol-Instruct, an end-to-end framework that evolves instruction datasets using large language models without any human effort.The framework automatically analyzes and summarizes suitable evolutionary strategies for the given instruction data and iteratively improves the evolving method based on issues exposed during the instruction evolution process.Our extensive experiments demonstrate that the best method optimized by Auto Evol-Instruct outperforms human-designed methods on various benchmarks, including MT-Bench, AlpacaEval, GSM8K, and HumanEval. Can Xu 0002, Yingxiu Zhao, Jian-Guang Lou, Weizhu Chen |
EMNLP | 3 |
| 2023 | API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMsabstractRecent research has demonstrated that Large Language Models (LLMs) can enhance their capabilities by utilizing external tools.However, three pivotal questions remain unanswered: (1) How effective are current LLMs in utilizing tools?(2) How can we enhance LLMs' ability to utilize tools?(3) What obstacles need to be overcome to leverage tools?To address these questions, we introduce API-Bank, a groundbreaking benchmark, specifically designed for tool-augmented LLMs.For the first question, we develop a runnable evaluation system consisting of 73 API tools.We annotate 314 tool-use dialogues with 753 API calls to assess the existing LLMs' capabilities in planning, retrieving, and calling APIs.For the second question, we construct a comprehensive training set containing 1,888 tool-use dialogues from 2,138 APIs spanning 1,000 distinct domains.Using this dataset, we train Lynx, a tool-augmented LLM initialized from Alpaca.Experimental results demonstrate that GPT-3.5 exhibits improved tool utilization compared to GPT-3, while GPT-4 excels in planning.However, there is still significant potential for further improvement.Moreover, Lynx surpasses Alpaca's tool utilization performance by more than 26 pts and approaches the effectiveness of GPT-3.5.Through error analysis, we highlight the key challenges for future research in this field to answer the third question 1 . Yingxiu Zhao, Bowen Yu 0002, Feifan Song 0001, Hangyu Li 0003, Haiyang Yu 0003, Fei Huang 0002, Yongbin Li 0001 |
EMNLP | 2 |
| 2023 | Causal Document-Grounded Dialogue Pre-trainingabstractThe goal of document-grounded dialogue (DocGD) is to generate a response by anchoring the evidence in a supporting document in accordance with the dialogue context.This entails four causally interconnected variables.While task-specific pre-training has significantly enhanced performances on numerous downstream tasks, existing DocGD methods still rely on general pre-trained language models without a specifically tailored pre-training approach that explicitly captures the causal relationships.To address this, we present the first causallycomplete dataset construction strategy for developing million-scale DocGD pre-training corpora.Additionally, we propose a causallyperturbed pre-training strategy to better capture causality by introducing perturbations on the variables and optimizing the overall causal effect.Experiments conducted on three benchmark datasets demonstrate that our causal pretraining yields substantial and consistent improvements in fully-supervised, low-resource, few-shot, and zero-shot settings 1 . Yingxiu Zhao, Bowen Yu 0002, Bowen Li 0002, Haiyang Yu 0003, Jinyang Li 0003, Fei Huang 0002, Yongbin Li 0001, Nevin Lianwen Zhang |
EMNLP | 1 |
| 2022 | Improving Meta-learning for Low-resource Text Classification and Generation via Memory ImitationabstractYingxiu Zhao, Zhiliang Tian, Huaxiu Yao, Yinhe Zheng, Dongkyu Lee, Yiping Song, Jian Sun, Nevin Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yingxiu Zhao, Zhiliang Tian, Huaxiu Yao, Yinhe Zheng, Yiping Song, Jian Sun 0021, Nevin Lianwen Zhang |
ACL (1) | 1 |
| 2022 | Hard Gate Knowledge Distillation - Leverage Calibration for Robust and Reliable Language ModelabstractIn knowledge distillation, a student model is trained with supervisions from both knowledge from a teacher and observations drawn from a training data distribution.Knowledge of a teacher is considered a subject that holds interclass relations which send a meaningful supervision to a student; hence, much effort has been put to find such knowledge to be distilled.In this paper, we explore a question that has been given little attention: "when to distill such knowledge."The question is answered in our work with the concept of model calibration; we view a teacher model not only as a source of knowledge but also as a gauge to detect miscalibration of a student.This simple and yet novel view leads to a hard gate knowledge distillation scheme that switches between learning from a teacher model and training data.We verify the gating mechanism in the context of natural language generation at both the token-level and the sentence-level.Empirical comparisons with strong baselines show that hard gate knowledge distillation not only improves model generalization, but also significantly lowers model calibration error. Zhiliang Tian, Yingxiu Zhao, Ka Chun Cheung, Nevin Lianwen Zhang |
EMNLP | 3 |
| 2022 | Prompt Conditioned VAE: Enhancing Generative Replay for Lifelong Learning in Task-Oriented DialogueabstractLifelong learning (LL) is vital for advanced task-oriented dialogue (ToD) systems.To address the catastrophic forgetting issue of LL, generative replay methods are widely employed to consolidate past knowledge with generated pseudo samples.However, most existing generative replay methods use only a single taskspecific token to control their models.This scheme is usually not strong enough to constrain the generative model due to insufficient information involved.In this paper, we propose a novel method, prompt conditioned VAE for lifelong learning (PCLL), to enhance generative replay by incorporating tasks' statistics.PCLL captures task-specific distributions with a conditional variational autoencoder, conditioned on natural language prompts to guide the pseudo-sample generation.Moreover, it leverages a distillation process to further consolidate past knowledge by alleviating the noise in pseudo samples.Experiments on natural language understanding tasks of ToD systems demonstrate that PCLL significantly outperforms competitive baselines in building lifelong learning models.We release the code and data at GitHub. Yingxiu Zhao, Yinhe Zheng, Zhiliang Tian, Jian Sun 0021, Nevin Lianwen Zhang |
EMNLP | 1 |
| 2022 | SeqPATE: Differentially Private Text Generation via Knowledge DistillationabstractProtecting the privacy of user data is crucial for text generation models, which can leak sensitive information during generation. Differentially private (DP) learning methods provide guarantees against identifying the existence of a training sample from model outputs. PATE is a recent DP learning algorithm that achieves high utility with strong privacy protection on training samples. However, text generation models output tokens sequentially in a large output space; the classic PATE algorithm is not customized for this setting. Furthermore, PATE works well to protect sample-level privacy, but is not designed to protect phrases in samples. In this paper, we propose SeqPATE, an extension of PATE to text generation that protects the privacy of individual training samples and sensitive phrases in training data. To adapt PATE to text generation, we generate pseudo-contexts and reduce the sequence generation problem to a next-word prediction problem. To handle the large output space, we propose a candidate filtering strategy to dynamically reduce the output space, and refine the teacher aggregation of PATE to avoid low agreement due to voting for a large number of candidates. To further reduce privacy losses, we use knowledge distillation to reduce the number of teacher queries. The experiments verify the effectiveness of SeqPATE in protecting both training samples and sensitive phrases. Zhiliang Tian, Yingxiu Zhao, Yu-Xiang Wang 0003, Nevin Lianwen Zhang |
NeurIPS | 2 |