VLDB 2026 Research / reviewers in the wild / expert
Xinbei Ma
dblp:301/8959
· DBLP profile ↗
9ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0003-1505-8603ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 7 first-author · 9 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Planning, search and constraint satisfaction · 24% Language models and text generation · 21% Question answering and dialogue systems · 20% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 50% Query processing and optimization · 50% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
robustness |
1.6 | 2 | 2025 | Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions · ACL (1) 2025 On the Robustness of Editing Large Language Models · EMNLP 2024 |
Natural language and speech › Question answering and dialogue systems
dialogue understanding |
1.2 | 2 | 2023 | Enhanced Speaker-Aware Multi-Party Multi-Turn Dialogue Comprehension · IEEE ACM Trans. Audio Speech Lang. Process. 2023 Structural Characterization for Dialogue Disentanglement · ACL (1) 2022 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
constraint-based planning |
0.9 | 1 | 2025 | Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints · NeurIPS 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › language-based planning
LLM-based planning |
0.9 | 1 | 2025 | Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions · ACL (1) 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
planning evaluation |
0.9 | 1 | 2025 | Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints · NeurIPS 2025 |
Robotics › Autonomous driving › autonomous vehicle testing
simulation-based testing |
0.9 | 1 | 2025 | Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints · NeurIPS 2025 |
Natural language and speech › Language models and text generation
knowledge editing |
0.8 | 1 | 2024 | On the Robustness of Editing Large Language Models · EMNLP 2024 |
Natural language and speech › Language models and text generation
retrieval-augmented language models |
0.7 | 1 | 2023 | Query Rewriting in Retrieval-Augmented Large Language Models · EMNLP 2023 |
Query processing and optimization
query rewriting |
0.7 | 1 | 2023 | Query Rewriting in Retrieval-Augmented Large Language Models · EMNLP 2023 |
Information retrieval
retrieval-augmented generation |
0.7 | 1 | 2023 | Query Rewriting in Retrieval-Augmented Large Language Models · EMNLP 2023 |
Natural language and speech › Question answering and dialogue systems › multi-party dialogue
dialogue disentanglement |
0.6 | 1 | 2022 | Structural Characterization for Dialogue Disentanglement · ACL (1) 2022 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2025 | LESA: Learnable LLM Layer Scaling-Up · ACL (1) 2025 |
Natural language and speech › Question answering and dialogue systems › dialogue modeling
dialogue structure modeling |
0.2 | 1 | 2022 | Structural Characterization for Dialogue Disentanglement · ACL (1) 2022 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.3language model prompting · 1.3singular value decomposition · 0.9neural network parameter prediction · 0.9multimodal large language model · 0.9large language model · 0.9continual pre-training · 0.9model editing · 0.8empirical study · 0.8heterogeneous graph network · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental DistractionsabstractXinbei Ma, Yiting Wang, Yao Yao, Tongxin Yuan, Aston Zhang, Zhuosheng Zhang, Hai Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xinbei Ma, Yao Yao 0008, Tongxin Yuan, Aston Zhang, Zhuosheng Zhang 0001, Hai Zhao 0001 |
ACL (1) | 1 |
| 2025 | LESA: Learnable LLM Layer Scaling-UpabstractTraining Large Language Models (LLMs) from scratch requires immense computational resources, making it prohibitively expensive.Model scaling-up offers a promising solution by leveraging the parameters of smaller models to create larger ones.However, existing depth scaling-up methods rely on empirical heuristic rules for layer duplication, which result in poorer initialization and slower convergence during continual pre-training.We propose LESA, a novel learnable method for depth scaling-up.By concatenating parameters from each layer and applying Singular Value Decomposition, we uncover latent patterns between layers, suggesting that inter-layer parameters can be learned.LESA uses a neural network to predict the parameters inserted between adjacent layers, enabling better initialization and faster training.Experiments show that LESA outperforms existing baselines, achieving superior performance with less than half the computational cost during continual pre-training.Extensive analyses demonstrate its effectiveness across different model sizes and tasks. 1 Zouying Cao, Xinbei Ma, Yao Yao 0008, Zhi Chen 0006, Libo Qin 0001, Hai Zhao 0001 |
ACL (1) | 3 |
| 2025 | Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted ConstraintsabstractUnlike reasoning, which often entails a deep sequence of deductive steps, complex real-world planning is characterized by the need to synthesize a broad spectrum of parallel and potentially conflicting information and constraints. For example, in travel planning scenarios, it requires the integration of diverse real-world information and user preferences. While LLMs show promise, existing methods with long-horizon thinking struggle with handling multifaceted constraints, leading to suboptimal solutions. Motivated by the challenges of real-world travel planning, this paper introduces the Multiple Aspects of Planning (MAoP), empowering LLMs with "wide-horizon thinking" to solve planning problems with multifaceted constraints. Instead of direct planning, MAoP leverages the strategist to conduct pre-planning from various aspects and provide the planning blueprint for planners, enabling strong inference-time scalability by scaling aspects to consider various constraints. In addition, existing benchmarks for multi-constraint planning are flawed because they assess constraints in isolation, ignoring causal dependencies within the constraints, e.g, travel planning, where past activities dictate future itinerary. To address this, we propose Travel-Sim, an agent-based benchmark assessing plans via real-world simulation, thereby inherently resolving these causal dependencies. This paper advances LLM capabilities in complex planning and offers novel insights for evaluating sophisticated scenarios through simulation. Dongjie Yang, Chengqiang Lu, Qimeng Wang, Xinbei Ma, Yan Gao 0017, Yao Hu 0002, Hai Zhao 0001 |
NeurIPS | 4 |
| 2024 | PROM: A Phrase-level Copying Mechanism with Pre-training for Abstractive SummarizationabstractBased on the remarkable achievements of pre-trained language models in abstractive summarization, the copying mechanism has proved helpful by improving the factuality, stability, and overall performance. This work proposes PROM, a new PhRase-level cOpying Mechanism that enhances attention on n-grams, which can be applied to zero-shot summarization with pre-training. PROM adds an indicator layer to explicitly pick up tokens in n-gram that can be copied from the source, and calculates an auxiliary loss for the copying prediction. Empirical studies show that PROM makes significant improvements in fine-tuning on benchmarks. In the zero-shot setting, PROM is utilized in the self-supervised pre-training on raw corpora and provides new general baselines on a wide range of summarization datasets. Further analysis shows that PROM performs more reasonable copying and contributes to faithfulness. Our code is publicly available at https://github.com/xbmxb/PROM. Xinbei Ma, Yeyun Gong, Hai Zhao 0001, Nan Duan 0001 |
LREC/COLING | 1 |
| 2024 | On the Robustness of Editing Large Language ModelsabstractLarge language models (LLMs) have played a pivotal role in building communicative AI, yet they encounter the challenge of efficient updates.Model editing enables the manipulation of specific knowledge memories and the behavior of language generation without retraining.However, the robustness of model editing remains an open question.This work seeks to understand the strengths and limitations of editing methods, facilitating practical applications of communicative AI.We focus on three key research questions.RQ1: Can edited LLMs behave consistently resembling communicative AI in realistic situations?RQ2: To what extent does the rephrasing of prompts lead LLMs to deviate from the edited knowledge memory?RQ3: Which knowledge features are correlated with the performance and robustness of editing?Our empirical studies uncover a substantial disparity between existing editing methods and the practical application of LLMs.On rephrased prompts that are flexible but common in realistic applications, the performance of editing experiences a significant decline.Further analysis shows that more popular knowledge is memorized better, easier to recall, and more challenging to edit effectively. Xinbei Ma, Tianjie Ju, Jiyang Qiu, Zhuosheng Zhang 0001, Hai Zhao 0001, Lifeng Liu, Yulong Wang 0004 |
EMNLP | 1 |
| 2024 | Multi-turn dialogue comprehension from a topic-aware perspective
Xinbei Ma, Hai Zhao 0001, Zhuosheng Zhang 0001 |
Neurocomputing | 1 |
| 2023 | Query Rewriting in Retrieval-Augmented Large Language ModelsabstractLarge Language Models (LLMs) play powerful, black-box readers in the retrieve-thenread pipeline, making remarkable progress in knowledge-intensive tasks.This work introduces a new framework, Rewrite-Retrieve-Read instead of the previous retrieve-then-read for the retrieval-augmented LLMs from the perspective of the query rewriting.Unlike prior studies focusing on adapting either the retriever or the reader, our approach pays attention to the adaptation of the search query itself, for there is inevitably a gap between the input text and the needed knowledge in retrieval.We first prompt an LLM to generate the query, then use a web search engine to retrieve contexts.Furthermore, to better align the query to the frozen modules, we propose a trainable scheme for our pipeline.A small language model is adopted as a trainable rewriter to cater to the black-box LLM reader.The rewriter is trained using the feedback of the LLM reader by reinforcement learning.Evaluation is conducted on downstream tasks, open-domain QA and multiple-choice QA.Experiments results show consistent performance improvement, indicating that our framework is proven effective and scalable, and brings a new framework for retrieval-augmented LLM 1 . Xinbei Ma, Yeyun Gong, Hai Zhao 0001, Nan Duan 0001 |
EMNLP | 1 |
| 2023 | Enhanced Speaker-Aware Multi-Party Multi-Turn Dialogue ComprehensionabstractMulti-party multi-turn dialogue comprehension brings unprecedented challenges in handling complicated scenarios, as the co-occurrence of multiple speakers causes complexity and inconsistency. As a result of the multiple participation, the shift of speaker roles and crisscrossed discourse relations among utterances hinder reading comprehension. Motivated by this, we further integrate the enhancements of speaker-related features for dialogue comprehension performance. This work proposes a novel model with enhancement from both sides of speaker roles and speaker-aware relations. At the token level, we apply a speaker mask for attention, while at the discourse level, we utilize heterogeneous graph networks for comprehensive speaker-aware discourse clues. Experimental results show that ourEnhancedSpeaker-Aware method (ESA) helps achieve state-of-the-art performance on the Molweni dataset, as well as significant improvements on the FriendsQA dataset. We find that our method makes steady improvements on stronger backbones. Analysis shows that our model enhances the connections between utterances and their own speakers and captures the speaker-aware discourse relations. Discussions on data features and error cases are presented, and a visualized case is displayed. The findings reveal the importance of speaker-aware signals in dialogue comprehension. Xinbei Ma, Zhuosheng Zhang 0001, Hai Zhao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Structural Characterization for Dialogue DisentanglementabstractTangled multi-party dialogue contexts lead to challenges for dialogue reading comprehension, where multiple dialogue threads flow simultaneously within a common dialogue record, increasing difficulties in understanding the dialogue history for both human and machine.Previous studies mainly focus on utterance encoding methods with carefully designed features but pay inadequate attention to characteristic features of the structure of dialogues.We specially take structure factors into account and design a novel model for dialogue disentangling.Based on the fact that dialogues are constructed on successive participation and interactions between speakers, we model structural information of dialogues in two aspects: 1)speaker property that indicates whom a message is from, and 2) reference dependency that shows whom a message may refer to.The proposed method achieves new state-of-the-art on the Ubuntu IRC benchmark dataset and contributes to dialogue-related comprehension. Xinbei Ma, Zhuosheng Zhang 0001, Hai Zhao 0001 |
ACL (1) | 1 |