VLDB 2026 Research / reviewers in the wild / expert
Saab Mansour
dblp:03/8053
· DBLP profile ↗
27ranked-venue papers
3as first author
19since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented GenerationabstractMaría Andrea Cruz Blandón, Jayasimha Talur, Bruno Charron, Dong Liu, Saab Mansour, Marcello Federico. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. María Andrea Cruz Blandón, Jayasimha Talur, Bruno Charron, Saab Mansour, Marcello Federico |
ACL (1) | 5 |
| 2025 | DeAL: Decoding-time Alignment for Large Language ModelsabstractJames Y. Huang, Sailik Sengupta, Daniele Bonadiman, Yi-An Lai, Arshit Gupta, Nikolaos Pappas, Saab Mansour, Katrin Kirchhoff, Dan Roth. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. James Y. Huang, Sailik Sengupta, Daniele Bonadiman, Yi-An Lai, Arshit Gupta, Nikolaos Pappas 0004, Saab Mansour, Katrin Kirchhoff, Dan Roth 0001 |
ACL (1) | 7 |
| 2025 | Structured List-Grounded Question AnsweringabstractDocument-grounded dialogue systems aim to answer user queries by leveraging external information. Previous studies have mainly focused on handling free-form documents, often overlooking structured data such as lists, which can represent a range of nuanced semantic relations. Motivated by the observation that even advanced language models like GPT-3.5 often miss semantic cues from lists, this paper aims to enhance question answering (QA) systems for better interpretation and use of structured lists. To this end, we introduce the LIST2QA dataset, a novel benchmark to evaluate the ability of QA systems to respond effectively using list information. This dataset is created from unlabeled customer service documents using language models and model-based filtering processes to enhance data quality, and can be used to fine-tune and evaluate QA models. Apart from directly generating responses through fine-tuned models, we further explore the explicit use of Intermediate Steps for Lists (ISL), aligning list items with user backgrounds to better reflect how humans interpret list items before generating responses. Our experimental results demonstrate that models trained on LIST2QA with our ISL approach outperform baselines across various metrics. Specifically, our fine-tuned Flan-T5-XL model shows increases of 3.1% in ROUGE-L, 4.6% in correctness, 4.5% in faithfulness, and 20.6% in completeness compared to models without applying filtering and the proposed ISL method. Mujeen Sung, Song Feng 0001, James Gung, Raphael Shu, Yi Zhang 0053, Saab Mansour |
COLING | 6 |
| 2025 | Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary EvaluationabstractMahnaz Koupaee, Jake W. Vincent, Saab Mansour, Igor Shalyminov, Han He, Hwanjun Song, Raphael Shu, Jianfeng He, Yi Nian, Amy Wing-mei Wong, Kyu J. Han, Hang Su. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Mahnaz Koupaee, Jake W. Vincent, Saab Mansour, Igor Shalyminov, Han He, Hwanjun Song, Raphael Shu, Yi Nian, Amy Wing-mei Wong, Kyu J. Han |
NAACL (Long Papers) | 3 |
| 2024 | Eliciting Better Multilingual Structured Reasoning from LLMs through CodeabstractThe development of large language models (LLM) has shown progress on reasoning, though studies have largely considered either English or simple reasoning tasks.To address this, we introduce a multilingual structured reasoning and explanation dataset, termed xSTREET, that covers four tasks across six languages.xSTREET exposes a gap in base LLM performance between English and non-English reasoning tasks. 1 We then propose two methods to remedy this gap, building on the insight that LLMs trained on code are better reasoners.First, at training time, we augment a code dataset with multilingual comments using machine translation while keeping program code as-is.Second, at inference time, we bridge the gap between training and inference by employing a prompt structure that incorporates step-by-step code primitives to derive new facts and find a solution.Our methods show improved multilingual performance on xSTREET, most notably on the scientific commonsense reasoning subtask.Furthermore, the models show no regression on non-reasoning tasks, thus demonstrating our techniques maintain general-purpose abilities. Bryan Li, Tamer Alkhouli, Daniele Bonadiman, Nikolaos Pappas 0004, Saab Mansour |
ACL (1) | 5 |
| 2024 | FineSurE: Fine-grained Summarization Evaluation using LLMsabstractAutomated evaluation is crucial for streamlining text summarization benchmarking and model development, given the costly and timeconsuming nature of human evaluation.Traditional methods like ROUGE do not correlate well with human judgment, while recently proposed LLM-based metrics provide only summary-level assessment using Likertscale scores.This limits deeper model analysis, e.g., we can only assign one hallucination score at the summary level, while at the sentence level, we can count sentences containing hallucinations.To remedy those limitations, we propose FineSurE, a fine-grained evaluator specifically tailored for the summarization task using large language models (LLMs).It also employs completeness and conciseness criteria, in addition to faithfulness, enabling multi-dimensional assessment.We compare various open-source and proprietary LLMs as backbones for FineSurE.In addition, we conduct extensive benchmarking of FineSurE against SOTA methods including NLI-, QA-, and LLM-based methods, showing improved performance especially on the completeness and conciseness dimensions.The code is available at https://github.com/ DISL-Lab/FineSurE-ACL24. Hwanjun Song, Igor Shalyminov, Jason Cai, Saab Mansour |
ACL (1) | 5 |
| 2024 | Can Your Model Tell a Negation from an Implicature? Unravelling Challenges With Intent EncodersabstractYuwei Zhang, Siffi Singh, Sailik Sengupta, Igor Shalyminov, Hang Su, Hwanjun Song, Saab Mansour. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Siffi Singh, Sailik Sengupta, Igor Shalyminov, Hwanjun Song, Saab Mansour |
ACL (1) | 7 |
| 2024 | MAGID: An Automated Pipeline for Generating Synthetic Multi-modal DatasetsabstractHossein Aboutalebi, Hwanjun Song, Yusheng Xie, Arshit Gupta, Lijia Sun, Hang Su, Igor Shalyminov, Nikolaos Pappas, Siffi Singh, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Hossein Aboutalebi, Hwanjun Song, Yusheng Xie, Arshit Gupta, Lijia Sun, Igor Shalyminov, Nikolaos Pappas 0004, Siffi Singh, Saab Mansour |
NAACL-HLT | 10 |
| 2024 | CERET: Cost-Effective Extrinsic Refinement for Text GenerationabstractJason Cai, Hang Su, Monica Sunkara, Igor Shalyminov, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jason Cai, Monica Sunkara, Igor Shalyminov, Saab Mansour |
NAACL-HLT | 5 |
| 2024 | Semi-Supervised Dialogue Abstractive Summarization via High-Quality Pseudolabel SelectionabstractJianfeng He, Hang Su, Jason Cai, Igor Shalyminov, Hwanjun Song, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jason Cai, Igor Shalyminov, Hwanjun Song, Saab Mansour |
NAACL-HLT | 6 |
| 2024 | FLAP: Flow-Adhering Planning with Constrained Decoding in LLMsabstractShamik Roy, Sailik Sengupta, Daniele Bonadiman, Saab Mansour, Arshit Gupta. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Shamik Roy, Sailik Sengupta, Daniele Bonadiman, Saab Mansour, Arshit Gupta |
NAACL-HLT | 4 |
| 2024 | TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue SummarizationabstractLiyan Tang, Igor Shalyminov, Amy Wong, Jon Burnsky, Jake Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, Lijia Sun, Yi Zhang, Saab Mansour, Kathleen McKeown. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Liyan Tang, Igor Shalyminov, Amy Wing-mei Wong, Jon Burnsky, Jake W. Vincent, Siffi Singh, Song Feng 0001, Hwanjun Song, Lijia Sun, Yi Zhang 0053, Saab Mansour, Kathy McKeown |
NAACL-HLT | 13 |
| 2023 | DFEE: Interactive DataFlow Execution and Evaluation KitabstractDataFlow has been emerging as a new paradigm for building task-oriented chatbots due to its expressive semantic representations of the dialogue tasks. Despite the availability of a large dataset SMCalFlow and a simplified syntax, the development and evaluation of DataFlow-based chatbots remain challenging due to the system complexity and the lack of downstream toolchains. In this demonstration, we present DFEE, an interactive DataFlow Execution and Evaluation toolkit that supports execution, visualization and benchmarking of semantic parsers given dialogue input and backend database. We demonstrate the system via a complex dialog task: event scheduling that involves temporal reasoning. It also supports diagnosing the parsing results via a friendly interface that allows developers to examine dynamic DataFlow and the corresponding execution results. To illustrate how to benchmark SoTA models, we propose a novel benchmark that covers more sophisticated event scheduling scenarios and a new metric on task success evaluation. The codes of DFEE have been released on https://github.com/amazonscience/dataflow-evaluation-toolkit. Han He, Song Feng 0001, Daniele Bonadiman, Yi Zhang 0053, Saab Mansour |
AAAI | 5 |
| 2023 | Robustification of Multilingual Language Models to Real-world Noise in Crosslingual Zero-shot Settings with Robust Contrastive PretrainingabstractAdvances in neural modeling have achieved state-of-the-art (SOTA) results on public natural language processing (NLP) benchmarks, at times surpassing human performance.However, there is a gap between public benchmarks and real-world applications where noise, such as typographical or grammatical mistakes, is abundant and can result in degraded performance.Unfortunately, works which evaluate the robustness of neural models on noisy data and propose improvements, are limited to the English language.Upon analyzing noise in different languages, we observe that noise types vary greatly across languages.Thus, existing investigations do not generalize trivially to multilingual settings.To benchmark the performance of pretrained multilingual language models, we construct noisy datasets covering five languages and four NLP tasks and observe a clear gap in the performance between clean and noisy data in the zero-shot cross-lingual setting.After investigating several ways to boost the robustness of multilingual models in this setting, we propose Robust Contrastive Pretraining (RCP).RCP combines data augmentation with a contrastive loss term at the pretraining stage and achieves large improvements on noisy (& original test data) across two sentencelevel (+3.2%) and two sequence-labeling (+10 F1-score) multilingual classification tasks.Language Noise Injection Ratio Realistic Utt. % Realistic Examples (test-set) Unrealistic Examples (test-set) French (fr) 0.1 95.4% Me montré les vols directs de Charlotte à Minneapolis mardi matin .Quelle compagnie aérienne fut YX Me montré des vols entre Détroit er St. Louis sur Delta Northwest US Air est United Airlines .Lister des vols de Las Vegas à Son Diego German (de) 0.2 94.5% Zeige mir der Flüge zwischen Housten und Orlando Welche Flüge gibt es vom Tacoma nach San Jose Zeige mit alle Flüge vor Charlotte nach Minneapolis zum Dienstag morgen Zeige mit Flüge an Milwaukee nach Washington DC v. 12 Uhr Asa Cooper Stickland, Sailik Sengupta, Jason Krone, Saab Mansour |
EACL | 4 |
| 2023 | Conversation Style Transfer using Few-Shot LearningabstractShamik Roy, Raphael Shu, Nikolaos Pappas, Elman Mansimov, Yi Zhang, Saab Mansour, Dan Roth. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shamik Roy, Raphael Shu, Nikolaos Pappas 0004, Elman Mansimov, Yi Zhang 0001, Saab Mansour, Dan Roth 0001 |
IJCNLP (1) | 6 |
| 2022 | Label Semantic Aware Pre-training for Few-shot Text ClassificationabstractAaron Mueller, Jason Krone, Salvatore Romeo, Saab Mansour, Elman Mansimov, Yi Zhang, Dan Roth. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Aaron Mueller, Jason Krone, Salvatore Romeo, Saab Mansour, Elman Mansimov, Yi Zhang 0001, Dan Roth 0001 |
ACL (1) | 4 |
| 2022 | Injecting Domain Knowledge in Language Models for Task-oriented Dialogue SystemsabstractPre-trained language models (PLM) have advanced the state-of-the-art across NLP applications, but lack domain-specific knowledge that does not naturally occur in pre-training data.Previous studies augmented PLMs with symbolic knowledge for different downstream NLP tasks.However, knowledge bases (KBs) utilized in these studies are usually large-scale and static, in contrast to small, domain-specific, and modifiable knowledge bases that are prominent in real-world task-oriented dialogue (TOD) systems.In this paper, we showcase the advantages of injecting domain-specific knowledge prior to fine-tuning on TOD tasks.To this end, we utilize light-weight adapters that can be easily integrated with PLMs and serve as a repository for facts learned from different KBs.To measure the efficacy of proposed knowledge injection methods, we introduce Knowledge Probing using Response Selection (KPRS) -a probe designed specifically for TOD models.Experiments 1 on KPRS and the response generation task show improvements of knowledge injection with adapters over strong baselines. * Work performed while at AWS AI Labs 1 https://github.com/amazon-research/ domain-knowledge-injection Denis Emelin, Daniele Bonadiman, Sawsan Alqahtani, Saab Mansour |
EMNLP | 5 |
| 2021 | Nearest Neighbour Few-Shot Learning for Cross-lingual ClassificationabstractEven though large pre-trained multilingual models (e.g.mBERT, XLM-R) have led to significant performance gains on a wide range of cross-lingual NLP tasks, success on many downstream tasks still relies on the availability of sufficient annotated data.Traditional fine-tuning of pre-trained models using only a few target samples can cause over-fitting.This can be quite limiting as most languages in the world are under-resourced.In this work, we investigate cross-lingual adaptation using a simple nearest neighbor few-shot (< 15 samples) inference technique for classification tasks.We experiment using a total of 16 distinct languages across two NLP tasks-XNLI and PAWS-X.Our approach consistently improves traditional fine-tuning using only a handful of labeled samples in target locales.We also demonstrate its generalization capability across tasks.* Work done while Saiful was interning at Amazon AI 1 We loosely use the term LM to describe unsupervised pretrained models including Masked-LMs and Causal-LMs Saiful Bari, Batool Haider, Saab Mansour |
EMNLP (1) | 3 |
| 2021 | Knowledge-Driven Slot Constraints for Goal-Oriented Dialogue SystemsabstractPiyawat Lertvittayakumjorn, Daniele Bonadiman, Saab Mansour. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Piyawat Lertvittayakumjorn, Daniele Bonadiman, Saab Mansour |
NAACL-HLT | 3 |
| 2020 | End-to-End Slot Alignment and Recognition for Cross-Lingual NLUabstractNatural language understanding (NLU) in the context of goal-oriented dialog systems typically includes intent classification and slot labeling tasks.Existing methods to expand an NLU system to new languages use machine translation with slot label projection from source to the translated utterances, and thus are sensitive to projection errors.In this work, we propose a novel end-to-end model that learns to align and predict target slot labels jointly for cross-lingual transfer.We introduce MultiATIS++, a new multilingual NLU corpus that extends the Multilingual ATIS corpus to nine languages across four language families, and evaluate our method using the corpus.Results show that our method outperforms a simple label projection method using fast-align on most languages, and achieves competitive performance to the more complex, state-of-the-art projection method with only half of the training time.We release our MultiATIS++ corpus to the community to continue future research on cross-lingual NLU. Weijia Xu, Batool Haider, Saab Mansour |
EMNLP (1) | 3 |
| 2015 | Spelling Correction of User Search Queries through Statistical Machine TranslationabstractWe use character-based statistical machine translation in order to correct user search queries in the e-commerce domain.The training data is automatically extracted from event logs where users re-issue their search queries with potentially corrected spelling within the same session.We show results on a test set which was annotated by humans and compare against online autocorrection capabilities of three additional web sites.Overall, the methods presented in this paper outperform fully productized spellchecking and autocorrection services in terms of accuracy and F1 score.We also propose novel evaluation steps based on retrieved search results of the corrected queries in terms of quantity and relevance. Sasa Hasan, Carmen Heger, Saab Mansour |
EMNLP | 3 |
| 2014 | Translation model based weighting for phrase extraction
Saab Mansour, Hermann Ney |
EAMT | 1 |
| 2013 | Phrase Training Based Adaptation for Statistical Machine Translation
Saab Mansour, Hermann Ney |
HLT-NAACL | 1 |
| 2012 | A Holistic Approach to Bilingual Sentence Fragment Extraction from Comparable Corpora
Mahdi Khademian, Kaveh Taghipour, Saab Mansour, Shahram Khadivi |
LREC | 3 |
| 2012 | Arabic-Segmentation Combination Strategies for Statistical Machine Translation
Saab Mansour, Hermann Ney |
LREC | 1 |
| 2012 | A comparison of segmentation methods and extended lexicon models for Arabic statistical machine translation
Sasa Hasan, Saab Mansour, Hermann Ney |
Mach. Transl. | 2 |
| 2009 | Recent advances in SRI'S IraqCommTM Iraqi Arabic-English speech-to-speech translation systemabstractWe summarize recent progress on SRI's IraqCommtrade Iraqi Arabic-English two-way speech-to-speech translation system. In the past year we made substantial developments in our speech recognition and machine translation technology, leading to significant improvements in both accuracy and speed of the IraqComm system. On the 2008 NIST-evaluation dataset our twoway speech-to-text (S2T) system achieved 6% to 8% absolute improvement in BLEU in both directions, compared to our previous year system. Murat Akbacak, Horacio Franco, Michael W. Frandsen, Sasa Hasan, Huda Jameel, Andreas Kathol, Shahram Khadivi, Arindam Mandal, Saab Mansour, Kristin Precoda, Colleen Richey, Dimitra Vergyri, Wen Wang 0001, Jing Zheng 0001 |
ICASSP | 10 |