Mohsen Mesgar

dblp:140/3476 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Question answering and dialogue systems · 42% Language models and text generation · 34% Information extraction and text analysis · 16%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
table question answering
1.012026
Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems › dialogue
argumentative dialogue
0.712023
A Dataset of Argumentative Dialogues on Scientific Papers · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
argument mining
0.712023
A Dataset of Argumentative Dialogues on Scientific Papers · ACL (1) 2023
Natural language and speech › Question answering and dialogue systems › question generation
clarifying question
0.712023
Python Code Generation by Asking Clarification Questions · ACL (1) 2023
Natural language and speech › Question answering and dialogue systems › knowledge-grounded dialogue
document-grounded dialogue
0.712023
A Dataset of Argumentative Dialogues on Scientific Papers · ACL (1) 2023
Program synthesis and code generation
code generation from natural language
0.712023
Python Code Generation by Asking Clarification Questions · ACL (1) 2023
Natural language and speech › Question answering and dialogue systems › dialogue evaluation
dialogue coherence evaluation
0.412020
Dialogue Coherence Assessment Without Explicit Dialogue Act Labels · ACL 2020
Natural language and speech › Language models and text generation › text summarization
document summarization
0.412019
Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation · IJCAI 2019
Natural language and speech › Language models and text generation › text summarization
extractive summarization
0.412019
Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation · IJCAI 2019
Machine learning › Reinforcement learning
reward learning
0.412019
Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation · IJCAI 2019
Natural language and speech › Information extraction and text analysis › text classification › automated scoring
automated essay scoring
0.312018
A Neural Local Coherence Model for Text Quality Assessment · EMNLP 2018
Natural language and speech › Language models and text generation › text evaluation
coherence modeling
0.312018
A Neural Local Coherence Model for Text Quality Assessment · EMNLP 2018
Natural language and speech › Information extraction and text analysis › text classification
readability assessment
0.312018
A Neural Local Coherence Model for Text Quality Assessment · EMNLP 2018
Information retrieval › search engines › structured data search
table retrieval
0.312026
Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation · ACL (1) 2026
Natural language and speech › Language models and text generation › text summarization › long document summarization
scientific paper summarization
0.212016
Generating Coherent Summaries of Scientific Articles Using Coherence Patterns · EMNLP 2016
Natural language and speech › Language models and text generation
text summarization
0.212016
Generating Coherent Summaries of Scientific Articles Using Coherence Patterns · EMNLP 2016
Natural language and speech › Language models and text generation
pre-trained language model
0.212023
Python Code Generation by Asking Clarification Questions · ACL (1) 2023
Machine learning › Learning paradigms
multi-task learning
0.112020
Dialogue Coherence Assessment Without Explicit Dialogue Act Labels · ACL 2020
Machine learning › Reinforcement learning
policy learning
0.112019
Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation · IJCAI 2019

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 2.4large language model · 2.0pre-trained language model · 0.7fine-tuning · 0.7multi-task learning · 0.4dialogue act prediction · 0.4learning to rank · 0.4sentence embeddings · 0.3semantic state modeling · 0.3coherence patterns · 0.2
YearPublicationVenuePosition
2026 Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation
abstract
Table Question Answering (TQA) aims to answer natural language questions about tabular data, often accompanied by additional contexts such as text passages.The task spans diverse settings, varying in table representation, question/answer complexity, modality involved, and domain.While recent advances in large language models (LLMs) have led to substantial progress in TQA, the field still lacks a systematic organization and understanding of task formulations, core challenges, and methodological trends, particularly in light of emerging research directions such as reinforcement learning.This survey addresses this gap by providing a comprehensive and structured overview of TQA research with a focus on LLM-based methods.We provide a comprehensive categorization of existing benchmarks and task setups.We group current modeling strategies according to the challenges they target, and analyze their strengths and limitations.Furthermore, we highlight underexplored but timely topics that have not been systematically covered in prior research.By unifying disparate research threads and identifying open problems, our survey offers a consolidated foundation for the TQA community, enabling a deeper understanding of the state of the art and guiding future developments in this rapidly evolving area.
Wei Zhou 0067, Bolei Ma, Annemarie Friedrich, Mohsen Mesgar
ACL (1)4
2024 FREB-TQA: A Fine-Grained Robustness Evaluation Benchmark for Table Question Answering
abstract
Wei Zhou, Mohsen Mesgar, Heike Adel, Annemarie Friedrich. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Wei Zhou 0067, Mohsen Mesgar, Heike Adel, Annemarie Friedrich
NAACL-HLT2
2023 Python Code Generation by Asking Clarification Questions
abstract
Code generation from text requires understanding the user's intent from a natural language description and generating an executable code snippet that satisfies this intent.While recent pretrained language models demonstrate remarkable performance for this task, these models fail when the given natural language description is under-specified.In this work, we introduce a novel and more realistic setup for this task.We hypothesize that the underspecification of a natural language description can be resolved by asking clarification questions.Therefore, we collect and introduce a new dataset named CodeClarQA containing pairs of natural language descriptions and code with created synthetic clarification questions and answers.The empirical results of our evaluation of pretrained language model performance on code generation show that clarifications result in more precisely generated code, as shown by the substantial improvement of model performance in all evaluation metrics.Alongside this, our task and dataset introduce new challenges to the community, including when and what clarification questions should be asked.Our code and dataset are available on GitHub.1
Haau-Sing Li, Mohsen Mesgar, André F. T. Martins, Iryna Gurevych
ACL (1)2
2023 A Dataset of Argumentative Dialogues on Scientific Papers
abstract
With recent advances in question-answering models, various datasets have been collected to improve and study the effectiveness of these models on scientific texts.Questions and answers in these datasets explore a scientific paper by seeking factual information from the paper's content.However, these datasets do not tackle the argumentative content of scientific papers, which is of huge importance in persuasiveness of a scientific discussion.We introduce ArgSciChat, a dataset of 41 argumentative dialogues between scientists on 20 NLP papers.The unique property of our dataset is that it includes both exploratory and argumentative questions and answers in a dialogue discourse on a scientific paper.Moreover, the size of ArgSciChat demonstrates the difficulties in collecting dialogues for specialized domains.Thus, our dataset is a challenging resource to evaluate dialogue agents in low-resource domains, in which collecting training data is costly.We annotate all sentences of dialogues in ArgSciChat and analyze them extensively.The results confirm that dialogues in ArgSci-Chat include exploratory and argumentative interactions.Furthermore, we use our dataset to fine-tune and evaluate a pre-trained documentgrounded dialogue agent.The agent achieves a low performance on our dataset, motivating a need for dialogue agents with a capability to reason and argue about their answers.We publicly release ArgSciChat 1 .
Federico Ruggeri, Mohsen Mesgar, Iryna Gurevych
ACL (1)2
2023 The Devil is in the Details: On Models and Training Regimes for Few-Shot Intent Classification
abstract
In task-oriented dialog (ToD) new intents emerge on regular basis, with a handful of available utterances at best.This renders effective Few-Shot Intent Classification (FSIC) a central challenge for modular ToD systems.Recent FSIC methods appear to be similar: they use pretrained language models (PLMs) to encode utterances and predominantly resort to nearestneighbor-based inference.However, they also differ in major components: they start from different PLMs, use different encoding architectures and utterance similarity functions, and adopt different training regimes.Coupling of these vital components together with the lack of informative ablations prevents the identification of factors that drive the (reported) FSIC performance.We propose a unified framework to evaluate these components along the following key dimensions: (1) Encoding architectures: Cross-Encoder vs Bi-Encoders; (2) Similarity function: Parameterized (i.e., trainable) vs non-parameterized; (3) Training regimes: Episodic meta-learning vs conventional (i.e., non-episodic) training.Our experimental results on seven FSIC benchmarks reveal three new important findings.First, the unexplored combination of cross-encoder architecture and episodic meta-learning consistently yields the best FSIC performance.Second, episodic training substantially outperforms its non-episodic counterpart.Finally, we show that splitting episodes into support and query sets has a limited and inconsistent effect on performance.Our findings show the importance of ablations and fair comparisons in FSIC.We publicly release our code and data 1 .
Mohsen Mesgar, Thy Thy Tran, Goran Glavas, Iryna Gurevych
EACL1
2021 Improving Factual Consistency Between a Response and Persona Facts
abstract
Neural models for response generation produce responses that are semantically plausible but not necessarily factually consistent with facts describing the speaker's persona.These models are trained with fully supervised learning where the objective function barely captures factual consistency.We propose to finetune these models by reinforcement learning and an efficient reward function that explicitly captures the consistency between a response and persona facts as well as semantic plausibility 1 .Our automatic and human evaluations on the PersonaChat corpus confirm that our approach increases the rate of responses that are factually consistent with persona facts over its supervised counterpart while retaining the language quality of responses.
Mohsen Mesgar, Edwin Simpson, Iryna Gurevych
EACL1
2020 Dialogue Coherence Assessment Without Explicit Dialogue Act Labels
abstract
Recent dialogue coherence models use the coherence features designed for monologue texts, e.g.nominal entities, to represent utterances and then explicitly augment them with dialogue-relevant features, e.g., dialogue act labels.It indicates two drawbacks, (a) semantics of utterances is limited to entity mentions, and (b) the performance of coherence models strongly relies on the quality of the input dialogue act labels.We address these issues by introducing a novel approach to dialogue coherence assessment.We use dialogue act prediction as an auxiliary task in a multi-task learning scenario to obtain informative utterance representations for coherence assessment.Our approach alleviates the need for explicit dialogue act labels during evaluation.The results of our experiments show that our model substantially (more than 20 accuracy points) outperforms its strong competitors on the Dai-lyDialogue corpus, and performs on par with them on the SwitchBoard corpus for ranking dialogues concerning their coherence.We release our source code 1 .
Mohsen Mesgar, Sebastian Bücker, Iryna Gurevych
ACL1
2019 Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation
abstract
Document summarisation can be formulated as a sequential decision-making problem, which can be solved by Reinforcement Learning (RL) algorithms. The predominant RL paradigm for summarisation learns a cross-input policy, which requires considerable time, data and parameter tuning due to the huge search spaces and the delayed rewards. Learning input-specific RL policies is a more efficient alternative, but so far depends on handcrafted rewards, which are difficult to design and yield poor performance. We propose RELIS, a novel RL paradigm that learns a reward function with Learning-to-Rank (L2R) algorithms at training time and uses this reward function to train an input-specific RL policy at test time. We prove that RELIS guarantees to generate near-optimal summaries with appropriate L2R and RL algorithms. Empirically, we evaluate our approach on extractive multi-document summarisation. We show that RELIS reduces the training time by two orders of magnitude compared to the state-of-the-art models while performing on par with them.
Yang Gao 0021, Christian M. Meyer, Mohsen Mesgar, Iryna Gurevych
IJCAI3
2018 A Neural Local Coherence Model for Text Quality Assessment
abstract
We propose a local coherence model that captures the flow of what semantically connects adjacent sentences in a text.We represent the semantics of a sentence by a vector and capture its state at each word of the sentence.We model what relates two adjacent sentences based on the two most similar semantic states, each of which is in one of the sentences.We encode the perceived coherence of a text by a vector, which represents patterns of changes in salient information that relates adjacent sentences.Our experiments demonstrate that our approach is beneficial for two downstream tasks: Readability assessment, in which our model achieves new state-of-the-art results; and essay scoring, in which the combination of our coherence vectors and other taskdependent features significantly improves the performance of a strong essay scorer.
Mohsen Mesgar, Michael Strube 0001
EMNLP1
2016 Generating Coherent Summaries of Scientific Articles Using Coherence Patterns
abstract
Previous work on automatic summarization does not thoroughly consider coherence while generating the summary.We introduce a graph-based approach to summarize scientific articles.We employ coherence patterns to ensure that the generated summaries are coherent.The novelty of our model is twofold: we mine coherence patterns in a corpus of abstracts, and we propose a method to combine coherence, importance and non-redundancy to generate the summary.We optimize these factors simultaneously using Mixed Integer Programming.Our approach significantly outperforms baseline and state-of-the-art systems in terms of coherence (summary coherence assessment) and relevance (ROUGE scores).
Daraksha Parveen, Mohsen Mesgar, Michael Strube 0001
EMNLP2
2016 Lexical Coherence Graph Modeling Using Word Embeddings
abstract
Coherence is established by semantic connections between sentences of a text which can be modeled by lexical relations.In this paper, we introduce the lexical coherence graph (LCG), a new graph-based model to represent lexical relations among sentences.The frequency of subgraphs (coherence patterns) of this graph captures the connectivity style of sentence nodes in this graph.The coherence of a text is encoded by a vector of these frequencies.We evaluate the LCG model on the readability ranking task.The results of the experiments show that the LCG model obtains higher accuracy than state-of-the-art coherence models.Using larger subgraphs yields higher accuracy, because they capture more structural information.However, larger subgraphs can be sparse.We adapt Kneser-Ney smoothing to smooth subgraphs' frequencies.Smoothing improves performance.
Mohsen Mesgar, Michael Strube 0001
HLT-NAACL1