Saab Mansour

dblp:03/8053 · DBLP profile ↗
← Back
27ranked-venue papers
3as first author
19since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2025 MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation
abstract
María Andrea Cruz Blandón, Jayasimha Talur, Bruno Charron, Dong Liu, Saab Mansour, Marcello Federico. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
María Andrea Cruz Blandón, Jayasimha Talur, Bruno Charron, Saab Mansour, Marcello Federico
ACL (1)5
2025 DeAL: Decoding-time Alignment for Large Language Models
abstract
James Y. Huang, Sailik Sengupta, Daniele Bonadiman, Yi-An Lai, Arshit Gupta, Nikolaos Pappas, Saab Mansour, Katrin Kirchhoff, Dan Roth. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
James Y. Huang, Sailik Sengupta, Daniele Bonadiman, Yi-An Lai, Arshit Gupta, Nikolaos Pappas 0004, Saab Mansour, Katrin Kirchhoff, Dan Roth 0001
ACL (1)7
2025 Structured List-Grounded Question Answering
abstract
Document-grounded dialogue systems aim to answer user queries by leveraging external information. Previous studies have mainly focused on handling free-form documents, often overlooking structured data such as lists, which can represent a range of nuanced semantic relations. Motivated by the observation that even advanced language models like GPT-3.5 often miss semantic cues from lists, this paper aims to enhance question answering (QA) systems for better interpretation and use of structured lists. To this end, we introduce the LIST2QA dataset, a novel benchmark to evaluate the ability of QA systems to respond effectively using list information. This dataset is created from unlabeled customer service documents using language models and model-based filtering processes to enhance data quality, and can be used to fine-tune and evaluate QA models. Apart from directly generating responses through fine-tuned models, we further explore the explicit use of Intermediate Steps for Lists (ISL), aligning list items with user backgrounds to better reflect how humans interpret list items before generating responses. Our experimental results demonstrate that models trained on LIST2QA with our ISL approach outperform baselines across various metrics. Specifically, our fine-tuned Flan-T5-XL model shows increases of 3.1% in ROUGE-L, 4.6% in correctness, 4.5% in faithfulness, and 20.6% in completeness compared to models without applying filtering and the proposed ISL method.
Mujeen Sung, Song Feng 0001, James Gung, Raphael Shu, Yi Zhang 0053, Saab Mansour
COLING6
2025 Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation
abstract
Mahnaz Koupaee, Jake W. Vincent, Saab Mansour, Igor Shalyminov, Han He, Hwanjun Song, Raphael Shu, Jianfeng He, Yi Nian, Amy Wing-mei Wong, Kyu J. Han, Hang Su. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Mahnaz Koupaee, Jake W. Vincent, Saab Mansour, Igor Shalyminov, Han He, Hwanjun Song, Raphael Shu, Yi Nian, Amy Wing-mei Wong, Kyu J. Han
NAACL (Long Papers)3
2024 Eliciting Better Multilingual Structured Reasoning from LLMs through Code
abstract
The development of large language models (LLM) has shown progress on reasoning, though studies have largely considered either English or simple reasoning tasks.To address this, we introduce a multilingual structured reasoning and explanation dataset, termed xSTREET, that covers four tasks across six languages.xSTREET exposes a gap in base LLM performance between English and non-English reasoning tasks. 1 We then propose two methods to remedy this gap, building on the insight that LLMs trained on code are better reasoners.First, at training time, we augment a code dataset with multilingual comments using machine translation while keeping program code as-is.Second, at inference time, we bridge the gap between training and inference by employing a prompt structure that incorporates step-by-step code primitives to derive new facts and find a solution.Our methods show improved multilingual performance on xSTREET, most notably on the scientific commonsense reasoning subtask.Furthermore, the models show no regression on non-reasoning tasks, thus demonstrating our techniques maintain general-purpose abilities.
Bryan Li, Tamer Alkhouli, Daniele Bonadiman, Nikolaos Pappas 0004, Saab Mansour
ACL (1)5
2024 FineSurE: Fine-grained Summarization Evaluation using LLMs
abstract
Automated evaluation is crucial for streamlining text summarization benchmarking and model development, given the costly and timeconsuming nature of human evaluation.Traditional methods like ROUGE do not correlate well with human judgment, while recently proposed LLM-based metrics provide only summary-level assessment using Likertscale scores.This limits deeper model analysis, e.g., we can only assign one hallucination score at the summary level, while at the sentence level, we can count sentences containing hallucinations.To remedy those limitations, we propose FineSurE, a fine-grained evaluator specifically tailored for the summarization task using large language models (LLMs).It also employs completeness and conciseness criteria, in addition to faithfulness, enabling multi-dimensional assessment.We compare various open-source and proprietary LLMs as backbones for FineSurE.In addition, we conduct extensive benchmarking of FineSurE against SOTA methods including NLI-, QA-, and LLM-based methods, showing improved performance especially on the completeness and conciseness dimensions.The code is available at https://github.com/ DISL-Lab/FineSurE-ACL24.
Hwanjun Song, Igor Shalyminov, Jason Cai, Saab Mansour
ACL (1)5
2024 Can Your Model Tell a Negation from an Implicature? Unravelling Challenges With Intent Encoders
abstract
Yuwei Zhang, Siffi Singh, Sailik Sengupta, Igor Shalyminov, Hang Su, Hwanjun Song, Saab Mansour. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Siffi Singh, Sailik Sengupta, Igor Shalyminov, Hwanjun Song, Saab Mansour
ACL (1)7
2024 MAGID: An Automated Pipeline for Generating Synthetic Multi-modal Datasets
abstract
Hossein Aboutalebi, Hwanjun Song, Yusheng Xie, Arshit Gupta, Lijia Sun, Hang Su, Igor Shalyminov, Nikolaos Pappas, Siffi Singh, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Hossein Aboutalebi, Hwanjun Song, Yusheng Xie, Arshit Gupta, Lijia Sun, Igor Shalyminov, Nikolaos Pappas 0004, Siffi Singh, Saab Mansour
NAACL-HLT10
2024 CERET: Cost-Effective Extrinsic Refinement for Text Generation
abstract
Jason Cai, Hang Su, Monica Sunkara, Igor Shalyminov, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jason Cai, Monica Sunkara, Igor Shalyminov, Saab Mansour
NAACL-HLT5
2024 Semi-Supervised Dialogue Abstractive Summarization via High-Quality Pseudolabel Selection
abstract
Jianfeng He, Hang Su, Jason Cai, Igor Shalyminov, Hwanjun Song, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jason Cai, Igor Shalyminov, Hwanjun Song, Saab Mansour
NAACL-HLT6
2024 FLAP: Flow-Adhering Planning with Constrained Decoding in LLMs
abstract
Shamik Roy, Sailik Sengupta, Daniele Bonadiman, Saab Mansour, Arshit Gupta. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Shamik Roy, Sailik Sengupta, Daniele Bonadiman, Saab Mansour, Arshit Gupta
NAACL-HLT4
2024 TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
abstract
Liyan Tang, Igor Shalyminov, Amy Wong, Jon Burnsky, Jake Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, Lijia Sun, Yi Zhang, Saab Mansour, Kathleen McKeown. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Liyan Tang, Igor Shalyminov, Amy Wing-mei Wong, Jon Burnsky, Jake W. Vincent, Siffi Singh, Song Feng 0001, Hwanjun Song, Lijia Sun, Yi Zhang 0053, Saab Mansour, Kathy McKeown
NAACL-HLT13
2023 DFEE: Interactive DataFlow Execution and Evaluation Kit
abstract
DataFlow has been emerging as a new paradigm for building task-oriented chatbots due to its expressive semantic representations of the dialogue tasks. Despite the availability of a large dataset SMCalFlow and a simplified syntax, the development and evaluation of DataFlow-based chatbots remain challenging due to the system complexity and the lack of downstream toolchains. In this demonstration, we present DFEE, an interactive DataFlow Execution and Evaluation toolkit that supports execution, visualization and benchmarking of semantic parsers given dialogue input and backend database. We demonstrate the system via a complex dialog task: event scheduling that involves temporal reasoning. It also supports diagnosing the parsing results via a friendly interface that allows developers to examine dynamic DataFlow and the corresponding execution results. To illustrate how to benchmark SoTA models, we propose a novel benchmark that covers more sophisticated event scheduling scenarios and a new metric on task success evaluation. The codes of DFEE have been released on https://github.com/amazonscience/dataflow-evaluation-toolkit.
Han He, Song Feng 0001, Daniele Bonadiman, Yi Zhang 0053, Saab Mansour
AAAI5
2023 Robustification of Multilingual Language Models to Real-world Noise in Crosslingual Zero-shot Settings with Robust Contrastive Pretraining
abstract
Advances in neural modeling have achieved state-of-the-art (SOTA) results on public natural language processing (NLP) benchmarks, at times surpassing human performance.However, there is a gap between public benchmarks and real-world applications where noise, such as typographical or grammatical mistakes, is abundant and can result in degraded performance.Unfortunately, works which evaluate the robustness of neural models on noisy data and propose improvements, are limited to the English language.Upon analyzing noise in different languages, we observe that noise types vary greatly across languages.Thus, existing investigations do not generalize trivially to multilingual settings.To benchmark the performance of pretrained multilingual language models, we construct noisy datasets covering five languages and four NLP tasks and observe a clear gap in the performance between clean and noisy data in the zero-shot cross-lingual setting.After investigating several ways to boost the robustness of multilingual models in this setting, we propose Robust Contrastive Pretraining (RCP).RCP combines data augmentation with a contrastive loss term at the pretraining stage and achieves large improvements on noisy (& original test data) across two sentencelevel (+3.2%) and two sequence-labeling (+10 F1-score) multilingual classification tasks.Language Noise Injection Ratio Realistic Utt. % Realistic Examples (test-set) Unrealistic Examples (test-set) French (fr) 0.1 95.4% Me montré les vols directs de Charlotte à Minneapolis mardi matin .Quelle compagnie aérienne fut YX Me montré des vols entre Détroit er St. Louis sur Delta Northwest US Air est United Airlines .Lister des vols de Las Vegas à Son Diego German (de) 0.2 94.5% Zeige mir der Flüge zwischen Housten und Orlando Welche Flüge gibt es vom Tacoma nach San Jose Zeige mit alle Flüge vor Charlotte nach Minneapolis zum Dienstag morgen Zeige mit Flüge an Milwaukee nach Washington DC v. 12 Uhr
Asa Cooper Stickland, Sailik Sengupta, Jason Krone, Saab Mansour
EACL4
2023 Conversation Style Transfer using Few-Shot Learning
abstract
Shamik Roy, Raphael Shu, Nikolaos Pappas, Elman Mansimov, Yi Zhang, Saab Mansour, Dan Roth. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shamik Roy, Raphael Shu, Nikolaos Pappas 0004, Elman Mansimov, Yi Zhang 0001, Saab Mansour, Dan Roth 0001
IJCNLP (1)6
2022 Label Semantic Aware Pre-training for Few-shot Text Classification
abstract
Aaron Mueller, Jason Krone, Salvatore Romeo, Saab Mansour, Elman Mansimov, Yi Zhang, Dan Roth. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Aaron Mueller, Jason Krone, Salvatore Romeo, Saab Mansour, Elman Mansimov, Yi Zhang 0001, Dan Roth 0001
ACL (1)4
2022 Injecting Domain Knowledge in Language Models for Task-oriented Dialogue Systems
abstract
Pre-trained language models (PLM) have advanced the state-of-the-art across NLP applications, but lack domain-specific knowledge that does not naturally occur in pre-training data.Previous studies augmented PLMs with symbolic knowledge for different downstream NLP tasks.However, knowledge bases (KBs) utilized in these studies are usually large-scale and static, in contrast to small, domain-specific, and modifiable knowledge bases that are prominent in real-world task-oriented dialogue (TOD) systems.In this paper, we showcase the advantages of injecting domain-specific knowledge prior to fine-tuning on TOD tasks.To this end, we utilize light-weight adapters that can be easily integrated with PLMs and serve as a repository for facts learned from different KBs.To measure the efficacy of proposed knowledge injection methods, we introduce Knowledge Probing using Response Selection (KPRS) -a probe designed specifically for TOD models.Experiments 1 on KPRS and the response generation task show improvements of knowledge injection with adapters over strong baselines. * Work performed while at AWS AI Labs 1 https://github.com/amazon-research/ domain-knowledge-injection
Denis Emelin, Daniele Bonadiman, Sawsan Alqahtani, Saab Mansour
EMNLP5
2021 Nearest Neighbour Few-Shot Learning for Cross-lingual Classification
abstract
Even though large pre-trained multilingual models (e.g.mBERT, XLM-R) have led to significant performance gains on a wide range of cross-lingual NLP tasks, success on many downstream tasks still relies on the availability of sufficient annotated data.Traditional fine-tuning of pre-trained models using only a few target samples can cause over-fitting.This can be quite limiting as most languages in the world are under-resourced.In this work, we investigate cross-lingual adaptation using a simple nearest neighbor few-shot (< 15 samples) inference technique for classification tasks.We experiment using a total of 16 distinct languages across two NLP tasks-XNLI and PAWS-X.Our approach consistently improves traditional fine-tuning using only a handful of labeled samples in target locales.We also demonstrate its generalization capability across tasks.* Work done while Saiful was interning at Amazon AI 1 We loosely use the term LM to describe unsupervised pretrained models including Masked-LMs and Causal-LMs
Saiful Bari, Batool Haider, Saab Mansour
EMNLP (1)3
2021 Knowledge-Driven Slot Constraints for Goal-Oriented Dialogue Systems
abstract
Piyawat Lertvittayakumjorn, Daniele Bonadiman, Saab Mansour. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Piyawat Lertvittayakumjorn, Daniele Bonadiman, Saab Mansour
NAACL-HLT3
2020 End-to-End Slot Alignment and Recognition for Cross-Lingual NLU
abstract
Natural language understanding (NLU) in the context of goal-oriented dialog systems typically includes intent classification and slot labeling tasks.Existing methods to expand an NLU system to new languages use machine translation with slot label projection from source to the translated utterances, and thus are sensitive to projection errors.In this work, we propose a novel end-to-end model that learns to align and predict target slot labels jointly for cross-lingual transfer.We introduce MultiATIS++, a new multilingual NLU corpus that extends the Multilingual ATIS corpus to nine languages across four language families, and evaluate our method using the corpus.Results show that our method outperforms a simple label projection method using fast-align on most languages, and achieves competitive performance to the more complex, state-of-the-art projection method with only half of the training time.We release our MultiATIS++ corpus to the community to continue future research on cross-lingual NLU.
Weijia Xu, Batool Haider, Saab Mansour
EMNLP (1)3
2015 Spelling Correction of User Search Queries through Statistical Machine Translation
abstract
We use character-based statistical machine translation in order to correct user search queries in the e-commerce domain.The training data is automatically extracted from event logs where users re-issue their search queries with potentially corrected spelling within the same session.We show results on a test set which was annotated by humans and compare against online autocorrection capabilities of three additional web sites.Overall, the methods presented in this paper outperform fully productized spellchecking and autocorrection services in terms of accuracy and F1 score.We also propose novel evaluation steps based on retrieved search results of the corrected queries in terms of quantity and relevance.
Sasa Hasan, Carmen Heger, Saab Mansour
EMNLP3
2014 Translation model based weighting for phrase extraction
Saab Mansour, Hermann Ney
EAMT1
2013 Phrase Training Based Adaptation for Statistical Machine Translation
Saab Mansour, Hermann Ney
HLT-NAACL1
2012 A Holistic Approach to Bilingual Sentence Fragment Extraction from Comparable Corpora
Mahdi Khademian, Kaveh Taghipour, Saab Mansour, Shahram Khadivi
LREC3
2012 Arabic-Segmentation Combination Strategies for Statistical Machine Translation
Saab Mansour, Hermann Ney
LREC1
2012 A comparison of segmentation methods and extended lexicon models for Arabic statistical machine translation
Sasa Hasan, Saab Mansour, Hermann Ney
Mach. Transl.2
2009 Recent advances in SRI'S IraqCommTM Iraqi Arabic-English speech-to-speech translation system
abstract
We summarize recent progress on SRI's IraqCommtrade Iraqi Arabic-English two-way speech-to-speech translation system. In the past year we made substantial developments in our speech recognition and machine translation technology, leading to significant improvements in both accuracy and speed of the IraqComm system. On the 2008 NIST-evaluation dataset our twoway speech-to-text (S2T) system achieved 6% to 8% absolute improvement in BLEU in both directions, compared to our previous year system.
Murat Akbacak, Horacio Franco, Michael W. Frandsen, Sasa Hasan, Huda Jameel, Andreas Kathol, Shahram Khadivi, Arindam Mandal, Saab Mansour, Kristin Precoda, Colleen Richey, Dimitra Vergyri, Wen Wang 0001, Jing Zheng 0001
ICASSP10