Mohammad Saleh

dblp:67/1649 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0003-3165-3015ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 70% Reinforcement learning · 16% Representation and self-supervised learning · 6%

Topics — the 19 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
2.532025
Building Math Agents with Multi-Turn Iterative Preference Learning · ICLR 2025
RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025
Statistical Rejection Sampling Improves Preference Optimization · ICLR 2024
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
1.122025
Building Math Agents with Multi-Turn Iterative Preference Learning · ICLR 2025
RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025
Natural language and speech › Language models and text generation
text summarization
0.932020
PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization · ICML 2020
Generating Wikipedia by Summarizing Long Sequences · ICLR (Poster) 2018
Assessing The Factual Accuracy of Generated Text · KDD 2019
Natural language and speech › Language models and text generation
mathematical reasoning
0.912025
Building Math Agents with Multi-Turn Iterative Preference Learning · ICLR 2025
Machine learning › Reinforcement learning
preference learning
0.912025
Building Math Agents with Multi-Turn Iterative Preference Learning · ICLR 2025
Natural language and speech › Language models and text generation › alignment
reward hacking
0.912025
RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025
Machine learning › Reinforcement learning › reward learning
reward model training
0.912025
RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025
Natural language and speech › Language models and text generation › agentic language model › tool-augmented language models
tool-augmented reasoning
0.912025
Building Math Agents with Multi-Turn Iterative Preference Learning · ICLR 2025
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
0.822020
PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization · ICML 2020
Generating Wikipedia by Summarizing Long Sequences · ICLR (Poster) 2018
Natural language and speech › Language models and text generation
preference optimization
0.812024
Statistical Rejection Sampling Improves Preference Optimization · ICLR 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.812024
Statistical Rejection Sampling Improves Preference Optimization · ICLR 2024
Natural language and speech › Language models and text generation › language modeling
conditional language model
0.712023
Out-of-Distribution Detection and Selective Generation for Conditional Language Models · ICLR 2023
Natural language and speech › Language models and text generation › text generation
conditional text generation
0.712023
Calibrating Sequence likelihood Improves Conditional Language Generation · ICLR 2023
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.712023
Out-of-Distribution Detection and Selective Generation for Conditional Language Models · ICLR 2023
Machine learning › Representation and self-supervised learning
pre-training
0.412020
PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization · ICML 2020
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.412020
PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization · ICML 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
fact extraction
0.412019
Assessing The Factual Accuracy of Generated Text · KDD 2019
Natural language and speech › Language models and text generation
text generation
0.212023
Out-of-Distribution Detection and Selective Generation for Conditional Language Models · ICLR 2023
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty calibration
0.212023
Calibrating Sequence likelihood Improves Conditional Language Generation · ICLR 2023

Methods — techniques the papers use, named apart from their topics

direct preference optimization · 1.6sequence likelihood calibration · 1.4supervised fine-tuning · 0.9data augmentation · 0.9chain-of-thought · 0.9causal framework · 0.9KTO · 0.9rejection sampling · 0.8selective generation · 0.7out-of-distribution detection · 0.7
YearPublicationVenuePosition
2025 RRM: Robust Reward Model Training Mitigates Reward Hacking
abstract
Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. However, traditional RM training, which relies on response pairs tied to specific prompts, struggles to disentangle prompt-driven preferences from prompt-independent artifacts, such as response length and format. In this work, we expose a fundamental limitation of current RM training methods, where RMs fail to effectively distinguish between contextual signals and irrelevant artifacts when determining preferences. To address this, we introduce a causal framework that learns preferences independent of these artifacts and propose a novel data augmentation technique designed to eliminate them. Extensive experiments show that our approach successfully filters out undesirable artifacts, yielding a more robust reward model (RRM). Our RRM improves the performance of a pairwise reward model trained on Gemma-2-9b-it, on Reward-Bench, increasing accuracy from 80.61% to 84.15%. Additionally, we train two DPO policies using both the RM and RRM, demonstrating that the RRM significantly enhances DPO-aligned policies, improving MT-Bench scores from 7.27 to 8.31 and length-controlled win-rates in AlpacaEval-2 from 33.46% to 52.49%.
Tianqi Liu 0002, Wei Xiong 0015, Jie Ren 0006, Lichang Chen, Rishabh Joshi, Zhen Qin 0001, Tianhe Yu, Daniel Sohn, Anastasia Makarova, Jeremiah Z. Liu, Bilal Piot, Abraham Ittycheriah, Aviral Kumar, Mohammad Saleh
ICLR18
2025 Building Math Agents with Multi-Turn Iterative Preference Learning
abstract
Recent studies have shown that large language models' (LLMs) mathematical problem-solving capabilities can be enhanced by integrating external tools, such as code interpreters, and employing multi-turn Chain-of-Thought (CoT) reasoning. While current methods focus on synthetic data generation and Supervised Fine-Tuning (SFT), this paper studies the complementary direct preference learning approach to further improve model performance. However, existing direct preference learning algorithms are originally designed for the single-turn chat task, and do not fully address the complexities of multi-turn reasoning and external tool integration required for tool-integrated mathematical reasoning tasks. To fill in this gap, we introduce a multi-turn direct preference learning framework, tailored for this context, that leverages feedback from code interpreters and optimizes trajectory-level preferences. This framework includes multi-turn DPO and multi-turn KTO as specific implementations. The effectiveness of our framework is validated through training of various language models using an augmented prompt set from the GSM8K and MATH datasets. Our results demonstrate substantial improvements: a supervised fine-tuned Gemma-1.1-it-7B model's performance increased from 77.5% to 83.9% on GSM8K and from 46.1% to 51.2% on MATH. Similarly, a Gemma-2-it-9B model improved from 84.1% to 86.3% on GSM8K and from 51.0% to 54.5% on MATH.
Wei Xiong 0015, Chengshuai Shi, Aviv Rosenberg 0002, Zhen Qin 0001, Daniele Calandriello, Misha Khalman, Rishabh Joshi, Bilal Piot, Mohammad Saleh, Tong Zhang 0001, Tianqi Liu 0002
ICLR10
2025 LiPO: Listwise Preference Optimization through Learning-to-Rank
abstract
Tianqi Liu, Zhen Qin, Junru Wu, Jiaming Shen, Misha Khalman, Rishabh Joshi, Yao Zhao, Mohammad Saleh, Simon Baumgartner, Jialu Liu, Peter J Liu, Xuanhui Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Tianqi Liu 0002, Zhen Qin 0001, Misha Khalman, Rishabh Joshi, Mohammad Saleh, Simon Baumgartner, Peter J. Liu, Xuanhui Wang
NAACL (Long Papers)8
2024 Statistical Rejection Sampling Improves Preference Optimization
abstract
Improving the alignment of language models with human preferences remains an active research challenge. Previous approaches have primarily utilized online Reinforcement Learning from Human Feedback (RLHF). Recently, offline methods such as Sequence Likelihood Calibration (SLiC) and Direct Preference Optimization (DPO) have emerged as attractive alternatives, offering improvements in stability and scalability while maintaining competitive performance. SLiC refines its loss function using sequence pairs sampled from a supervised fine-tuned (SFT) policy, while DPO directly optimizes language models based on preference data, foregoing the need for a separate reward model. However, the maximum likelihood estimator (MLE) of the target optimal policy requires labeled preference pairs sampled from that policy. The absence of a reward model in DPO constrains its ability to sample preference pairs from the optimal policy. Meanwhile, SLiC can only sample preference pairs from the SFT policy. To address these limitations, we introduce a novel approach called Statistical Rejection Sampling Optimization (RSO) designed to source preference data from the target optimal policy using rejection sampling, enabling a more accurate estimation of the optimal policy. We also propose a unified framework that enhances the loss functions used in both SLiC and DPO from a preference modeling standpoint. Through extensive experiments across diverse tasks, we demonstrate that RSO consistently outperforms both SLiC and DPO as evaluated by both Large Language Models (LLMs) and human raters.
Tianqi Liu 0002, Rishabh Joshi, Misha Khalman, Mohammad Saleh, Peter J. Liu
ICLR5
2023 Out-of-Distribution Detection and Selective Generation for Conditional Language Models
Jie Ren 0006, Jiaming Luo, Kundan Krishna, Mohammad Saleh, Balaji Lakshminarayanan, Peter J. Liu
ICLR5
2023 Calibrating Sequence likelihood Improves Conditional Language Generation
Misha Khalman, Rishabh Joshi, Shashi Narayan, Mohammad Saleh, Peter J. Liu
ICLR5
2022 Virtual Interactive Imagery: Virtual Reality in the Treatment of Obesity
abstract
This paper proposes a multidisciplinary framework that applies psychological treatment. The framework uses virtual reality to apply guided imagery scenarios, where a person imagines a certain situation positively, to better deal with similar ones in reality. The idea is to use virtual reality to achieve self-efficacy; a self-belief that empowers success. The proposed concept of "Virtual Interactive Imagery" represents immersing users in virtual images to live and interact in a predeveloped scenario positively, instead of having people depend on a facilitator to guide them through imagination. The paper focuses on obesity treatment and the role of imagination in human weight loss. The methodology of creating a virtual system is presented, and a proof of concept for the system is developed and tested through a usability test.
Mariam Salim, Mohammad Saleh, Osama Halabi
CW2
2021 Augmented reality flavor: cross-modal mapping across gustation, olfaction, and vision
abstract
Gustatory display research is still in its infancy despite being one of the essential everyday senses that human practice while eating and drinking. Indeed, the most important and frequent tasks that our brain deals with every day are foraging and feeding. The recent studies by psychologists and cognitive neuroscientist revealed how complex multisensory rely on the integration of cues from all the human senses in any flavor experiences. The perception of flavor is multisensory and involves combinations of gustatory and olfactory stimuli. The cross-modal mapping between these modalities needs to be more explored in the virtual environment and simulation, especially in liquid food. In this paper, we present a customized wearable Augmented Reality (AR) system and olfaction display to study the effect of vision and olfaction on the gustatory sense. A user experiment and extensive analysis conducted to study the influence of each stimulus on the overall flavor, including other factors like age, previous experience in Virtual Reality (VR)/AR, and beverage consumption. The result showed that smell contributes strongly to the flavor with less contribution to the vision. However, the combination of these stimuli can deliver richer experience and a higher belief rate. Beverage consumption had a significant effect on the flavor belief rate. Experience is correlated with stimulus and age is correlated with belief rate, and both indirectly affected the belief rate.
Osama Halabi, Mohammad Saleh
Multim. Tools Appl.2
2020 PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization
abstract
Recent work pre-training Transformers with self-supervised objectives on large text corpora has shown great success when fine-tuned on downstream NLP tasks including text summarization. However, pre-training objectives tailored for abstractive text summarization have not been explored. Furthermore there is a lack of systematic evaluation across diverse domains. In this work, we propose pre-training large Transformer-based encoder-decoder models on massive text corpora with a new self-supervised objective. In PEGASUS, important sentences are removed/masked from an input document and are generated together as one output sequence from the remaining sentences, similar to an extractive summary. We evaluated our best PEGASUS model on 12 downstream summarization tasks spanning news, science, stories, instructions, emails, patents, and legislative bills. Experiments demonstrate it achieves state-of-the-art performance on all 12 downstream datasets measured by ROUGE scores. Our model also shows surprising performance on low-resource summarization, surpassing previous state-of-the-art results on 6 datasets with only 1000 examples. Finally we validated our results using human evaluation and show that our model summaries achieve human performance on multiple datasets.
Jingqing Zhang, Mohammad Saleh, Peter J. Liu
ICML3
2019 Assessing The Factual Accuracy of Generated Text
abstract
We propose a model-based metric to estimate the factual accuracy of generated text that is complementary to typical scoring schemes like ROUGE (Recall-Oriented Understudy for Gisting Evaluation) and BLEU (Bilingual Evaluation Understudy). We introduce and release a new large-scale dataset based on Wikipedia and Wikidata to train relation classifiers and end-to-end fact extraction models. The end-to-end models are shown to be able to extract complete sets of facts from datasets with full pages of text. We then analyse multiple models that estimate factual accuracy on a Wikipedia text summarization task, and show their efficacy compared to ROUGE and other model-free variants by conducting a human evaluation study.
Ben Goodrich, Vinay Rao, Peter J. Liu, Mohammad Saleh
KDD4
2018 Generating Wikipedia by Summarizing Long Sequences
Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, Noam Shazeer
ICLR (Poster)2
2017 Biomedical Epilepsy Multiprocessor Network on Chip EEG Utilizing IEEE802.11n Systems
abstract
This paper presents a novel wearable biomedical Network on Chip (NoC) concept development to monitor and predict irregular brain waves as advanced sensitive portable for an electroencephalogram (EEG) analysis device. The proposed device will monitor brain's spontaneous electrical activity in normal and abnormal situations for specific patients suffering from different types of epilepsy. This NOC would be able to predict the severity of the forthcoming epileptic attack. Meanwhile, this device will alert epileptic patients by giving an alarm on detection of any kind of abnormal brain electrical activity. The EEG NoC reads brain signals from a wireless sensor on the located on the patient's scalp, and runs parallel processing and filtering for the brainwaves. This process makes the detection of brain abnormalities possible, and many patients can be saved by predicting the time of epileptic seizure. When such a prediction occurs, alarm signals are sent to the patient to take protective measures. In this way, it will help patients to prevent themselves from different kinds of injuries and risky behaviors occurring during epilepsy attack or after that.
Mohammad Saleh, Jaafar Gaber, Maxime Wack
MASS1
2014 Performance Comparison of classification algorithms for EEG-based remote epileptic seizure detection in Wireless Sensor Networks
abstract
Identification of epileptic seizure remotely by analyzing the electroencephalography (EEG) signal is very important for scalable sensor-based health systems. Classification is the most important technique for wide-ranging applications to categorize the items according to its features with respect to predefined set of classes. In this paper, we conduct a performance evaluation based on the noiseless and noisy EEG-based epileptic seizure data using various classification algorithms including BayesNet, DecisionTable, IBK, J48/C4.5, and VFI. The reconstructed and noisy EEG data are decomposed with discrete cosine transform into several sub-bands. In addition, some of statistical features are extracted from the wavelet coefficients to represent the whole EEG data inputs into the classifiers. Benchmark on widely used dataset is utilized for automatic epileptic seizure detection including both normal and epileptic EEG datasets. The classification accuracy results confirm that the selected classifiers have greater potentiality to identify the noisy epileptic disorders.
Khalid Abualsaud, Massudi Mahmuddin, Mohammad Saleh, Amr Mohamed 0001
AICCSA3
2014 ConProve: A conceptual prover system
abstract
ConProve is an automated prover for propositional logic. It takes, as an input, a set of propositional formulas and proves whether a goal holds or not. ConProve converts each formula to its corresponding Truth Table Binary Relation (TTBR) considered also as a formal context (FC). The objects in FC correspond to all possible formulas interpretations (in terms of their truth value assignments), and the properties in FC correspond to the terms. When the function the 'BuildContext' function, ConProve starts the new goal proving. Firstly, it adds the goal negation to the set of formulas and constructs the formal contexts (FCs) relating formulas to terms. Secondly, it makes the FCs grouping and deduces, based on the conceptual reasoning, if the goal holds. The tool offers a user-friendly interface allowing the editing of the set of formulas as well as the visualization of the reasoning steps. Besides the tool, the paper illustrates the importance of the conceptual reasoning in deriving new conclusions as well as in discovering new, possibly implications by applying the extended Galois Connection.
Samir Elloumi, Ali Jaoua, Bilel Boulifa, Mohammad Saleh, Jameela Al Otaibi
AICCSA4