Xinbei Ma

dblp:301/8959 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0003-1505-8603ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 7 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Planning, search and constraint satisfaction · 24% Language models and text generation · 21% Question answering and dialogue systems · 20%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 50% Query processing and optimization · 50%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
robustness
1.622025
Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions · ACL (1) 2025
On the Robustness of Editing Large Language Models · EMNLP 2024
Natural language and speech › Question answering and dialogue systems
dialogue understanding
1.222023
Enhanced Speaker-Aware Multi-Party Multi-Turn Dialogue Comprehension · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Structural Characterization for Dialogue Disentanglement · ACL (1) 2022
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
constraint-based planning
0.912025
Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints · NeurIPS 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › language-based planning
LLM-based planning
0.912025
Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions · ACL (1) 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
planning evaluation
0.912025
Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints · NeurIPS 2025
Robotics › Autonomous driving › autonomous vehicle testing
simulation-based testing
0.912025
Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints · NeurIPS 2025
Natural language and speech › Language models and text generation
knowledge editing
0.812024
On the Robustness of Editing Large Language Models · EMNLP 2024
Natural language and speech › Language models and text generation
retrieval-augmented language models
0.712023
Query Rewriting in Retrieval-Augmented Large Language Models · EMNLP 2023
Query processing and optimization
query rewriting
0.712023
Query Rewriting in Retrieval-Augmented Large Language Models · EMNLP 2023
Information retrieval
retrieval-augmented generation
0.712023
Query Rewriting in Retrieval-Augmented Large Language Models · EMNLP 2023
Natural language and speech › Question answering and dialogue systems › multi-party dialogue
dialogue disentanglement
0.612022
Structural Characterization for Dialogue Disentanglement · ACL (1) 2022
Machine learning › Efficient and distributed learning
model compression
0.312025
LESA: Learnable LLM Layer Scaling-Up · ACL (1) 2025
Natural language and speech › Question answering and dialogue systems › dialogue modeling
dialogue structure modeling
0.212022
Structural Characterization for Dialogue Disentanglement · ACL (1) 2022

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.3language model prompting · 1.3singular value decomposition · 0.9neural network parameter prediction · 0.9multimodal large language model · 0.9large language model · 0.9continual pre-training · 0.9model editing · 0.8empirical study · 0.8heterogeneous graph network · 0.7
YearPublicationVenuePosition
2025 Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions
abstract
Xinbei Ma, Yiting Wang, Yao Yao, Tongxin Yuan, Aston Zhang, Zhuosheng Zhang, Hai Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xinbei Ma, Yao Yao 0008, Tongxin Yuan, Aston Zhang, Zhuosheng Zhang 0001, Hai Zhao 0001
ACL (1)1
2025 LESA: Learnable LLM Layer Scaling-Up
abstract
Training Large Language Models (LLMs) from scratch requires immense computational resources, making it prohibitively expensive.Model scaling-up offers a promising solution by leveraging the parameters of smaller models to create larger ones.However, existing depth scaling-up methods rely on empirical heuristic rules for layer duplication, which result in poorer initialization and slower convergence during continual pre-training.We propose LESA, a novel learnable method for depth scaling-up.By concatenating parameters from each layer and applying Singular Value Decomposition, we uncover latent patterns between layers, suggesting that inter-layer parameters can be learned.LESA uses a neural network to predict the parameters inserted between adjacent layers, enabling better initialization and faster training.Experiments show that LESA outperforms existing baselines, achieving superior performance with less than half the computational cost during continual pre-training.Extensive analyses demonstrate its effectiveness across different model sizes and tasks. 1
Zouying Cao, Xinbei Ma, Yao Yao 0008, Zhi Chen 0006, Libo Qin 0001, Hai Zhao 0001
ACL (1)3
2025 Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints
abstract
Unlike reasoning, which often entails a deep sequence of deductive steps, complex real-world planning is characterized by the need to synthesize a broad spectrum of parallel and potentially conflicting information and constraints. For example, in travel planning scenarios, it requires the integration of diverse real-world information and user preferences. While LLMs show promise, existing methods with long-horizon thinking struggle with handling multifaceted constraints, leading to suboptimal solutions. Motivated by the challenges of real-world travel planning, this paper introduces the Multiple Aspects of Planning (MAoP), empowering LLMs with "wide-horizon thinking" to solve planning problems with multifaceted constraints. Instead of direct planning, MAoP leverages the strategist to conduct pre-planning from various aspects and provide the planning blueprint for planners, enabling strong inference-time scalability by scaling aspects to consider various constraints. In addition, existing benchmarks for multi-constraint planning are flawed because they assess constraints in isolation, ignoring causal dependencies within the constraints, e.g, travel planning, where past activities dictate future itinerary. To address this, we propose Travel-Sim, an agent-based benchmark assessing plans via real-world simulation, thereby inherently resolving these causal dependencies. This paper advances LLM capabilities in complex planning and offers novel insights for evaluating sophisticated scenarios through simulation.
Dongjie Yang, Chengqiang Lu, Qimeng Wang, Xinbei Ma, Yan Gao 0017, Yao Hu 0002, Hai Zhao 0001
NeurIPS4
2024 PROM: A Phrase-level Copying Mechanism with Pre-training for Abstractive Summarization
abstract
Based on the remarkable achievements of pre-trained language models in abstractive summarization, the copying mechanism has proved helpful by improving the factuality, stability, and overall performance. This work proposes PROM, a new PhRase-level cOpying Mechanism that enhances attention on n-grams, which can be applied to zero-shot summarization with pre-training. PROM adds an indicator layer to explicitly pick up tokens in n-gram that can be copied from the source, and calculates an auxiliary loss for the copying prediction. Empirical studies show that PROM makes significant improvements in fine-tuning on benchmarks. In the zero-shot setting, PROM is utilized in the self-supervised pre-training on raw corpora and provides new general baselines on a wide range of summarization datasets. Further analysis shows that PROM performs more reasonable copying and contributes to faithfulness. Our code is publicly available at https://github.com/xbmxb/PROM.
Xinbei Ma, Yeyun Gong, Hai Zhao 0001, Nan Duan 0001
LREC/COLING1
2024 On the Robustness of Editing Large Language Models
abstract
Large language models (LLMs) have played a pivotal role in building communicative AI, yet they encounter the challenge of efficient updates.Model editing enables the manipulation of specific knowledge memories and the behavior of language generation without retraining.However, the robustness of model editing remains an open question.This work seeks to understand the strengths and limitations of editing methods, facilitating practical applications of communicative AI.We focus on three key research questions.RQ1: Can edited LLMs behave consistently resembling communicative AI in realistic situations?RQ2: To what extent does the rephrasing of prompts lead LLMs to deviate from the edited knowledge memory?RQ3: Which knowledge features are correlated with the performance and robustness of editing?Our empirical studies uncover a substantial disparity between existing editing methods and the practical application of LLMs.On rephrased prompts that are flexible but common in realistic applications, the performance of editing experiences a significant decline.Further analysis shows that more popular knowledge is memorized better, easier to recall, and more challenging to edit effectively.
Xinbei Ma, Tianjie Ju, Jiyang Qiu, Zhuosheng Zhang 0001, Hai Zhao 0001, Lifeng Liu, Yulong Wang 0004
EMNLP1
2024 Multi-turn dialogue comprehension from a topic-aware perspective
Xinbei Ma, Hai Zhao 0001, Zhuosheng Zhang 0001
Neurocomputing1
2023 Query Rewriting in Retrieval-Augmented Large Language Models
abstract
Large Language Models (LLMs) play powerful, black-box readers in the retrieve-thenread pipeline, making remarkable progress in knowledge-intensive tasks.This work introduces a new framework, Rewrite-Retrieve-Read instead of the previous retrieve-then-read for the retrieval-augmented LLMs from the perspective of the query rewriting.Unlike prior studies focusing on adapting either the retriever or the reader, our approach pays attention to the adaptation of the search query itself, for there is inevitably a gap between the input text and the needed knowledge in retrieval.We first prompt an LLM to generate the query, then use a web search engine to retrieve contexts.Furthermore, to better align the query to the frozen modules, we propose a trainable scheme for our pipeline.A small language model is adopted as a trainable rewriter to cater to the black-box LLM reader.The rewriter is trained using the feedback of the LLM reader by reinforcement learning.Evaluation is conducted on downstream tasks, open-domain QA and multiple-choice QA.Experiments results show consistent performance improvement, indicating that our framework is proven effective and scalable, and brings a new framework for retrieval-augmented LLM 1 .
Xinbei Ma, Yeyun Gong, Hai Zhao 0001, Nan Duan 0001
EMNLP1
2023 Enhanced Speaker-Aware Multi-Party Multi-Turn Dialogue Comprehension
abstract
Multi-party multi-turn dialogue comprehension brings unprecedented challenges in handling complicated scenarios, as the co-occurrence of multiple speakers causes complexity and inconsistency. As a result of the multiple participation, the shift of speaker roles and crisscrossed discourse relations among utterances hinder reading comprehension. Motivated by this, we further integrate the enhancements of speaker-related features for dialogue comprehension performance. This work proposes a novel model with enhancement from both sides of speaker roles and speaker-aware relations. At the token level, we apply a speaker mask for attention, while at the discourse level, we utilize heterogeneous graph networks for comprehensive speaker-aware discourse clues. Experimental results show that ourEnhancedSpeaker-Aware method (ESA) helps achieve state-of-the-art performance on the Molweni dataset, as well as significant improvements on the FriendsQA dataset. We find that our method makes steady improvements on stronger backbones. Analysis shows that our model enhances the connections between utterances and their own speakers and captures the speaker-aware discourse relations. Discussions on data features and error cases are presented, and a visualized case is displayed. The findings reveal the importance of speaker-aware signals in dialogue comprehension.
Xinbei Ma, Zhuosheng Zhang 0001, Hai Zhao 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Structural Characterization for Dialogue Disentanglement
abstract
Tangled multi-party dialogue contexts lead to challenges for dialogue reading comprehension, where multiple dialogue threads flow simultaneously within a common dialogue record, increasing difficulties in understanding the dialogue history for both human and machine.Previous studies mainly focus on utterance encoding methods with carefully designed features but pay inadequate attention to characteristic features of the structure of dialogues.We specially take structure factors into account and design a novel model for dialogue disentangling.Based on the fact that dialogues are constructed on successive participation and interactions between speakers, we model structural information of dialogues in two aspects: 1)speaker property that indicates whom a message is from, and 2) reference dependency that shows whom a message may refer to.The proposed method achieves new state-of-the-art on the Ubuntu IRC benchmark dataset and contributes to dialogue-related comprehension.
Xinbei Ma, Zhuosheng Zhang 0001, Hai Zhao 0001
ACL (1)1