VLDB 2026 Research / reviewers in the wild / expert
Mohammad Saleh
dblp:67/1649
· DBLP profile ↗
14ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0003-3165-3015ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Language models and text generation · 70% Reinforcement learning · 16% Representation and self-supervised learning · 6% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
2.5 | 3 | 2025 | Building Math Agents with Multi-Turn Iterative Preference Learning · ICLR 2025 RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 Statistical Rejection Sampling Improves Preference Optimization · ICLR 2024 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
1.1 | 2 | 2025 | Building Math Agents with Multi-Turn Iterative Preference Learning · ICLR 2025 RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Natural language and speech › Language models and text generation
text summarization |
0.9 | 3 | 2020 | PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization · ICML 2020 Generating Wikipedia by Summarizing Long Sequences · ICLR (Poster) 2018 Assessing The Factual Accuracy of Generated Text · KDD 2019 |
Natural language and speech › Language models and text generation
mathematical reasoning |
0.9 | 1 | 2025 | Building Math Agents with Multi-Turn Iterative Preference Learning · ICLR 2025 |
Machine learning › Reinforcement learning
preference learning |
0.9 | 1 | 2025 | Building Math Agents with Multi-Turn Iterative Preference Learning · ICLR 2025 |
Natural language and speech › Language models and text generation › alignment
reward hacking |
0.9 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Machine learning › Reinforcement learning › reward learning
reward model training |
0.9 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Natural language and speech › Language models and text generation › agentic language model › tool-augmented language models
tool-augmented reasoning |
0.9 | 1 | 2025 | Building Math Agents with Multi-Turn Iterative Preference Learning · ICLR 2025 |
Natural language and speech › Language models and text generation › text summarization
abstractive summarization |
0.8 | 2 | 2020 | PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization · ICML 2020 Generating Wikipedia by Summarizing Long Sequences · ICLR (Poster) 2018 |
Natural language and speech › Language models and text generation
preference optimization |
0.8 | 1 | 2024 | Statistical Rejection Sampling Improves Preference Optimization · ICLR 2024 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.8 | 1 | 2024 | Statistical Rejection Sampling Improves Preference Optimization · ICLR 2024 |
Natural language and speech › Language models and text generation › language modeling
conditional language model |
0.7 | 1 | 2023 | Out-of-Distribution Detection and Selective Generation for Conditional Language Models · ICLR 2023 |
Natural language and speech › Language models and text generation › text generation
conditional text generation |
0.7 | 1 | 2023 | Calibrating Sequence likelihood Improves Conditional Language Generation · ICLR 2023 |
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection |
0.7 | 1 | 2023 | Out-of-Distribution Detection and Selective Generation for Conditional Language Models · ICLR 2023 |
Machine learning › Representation and self-supervised learning
pre-training |
0.4 | 1 | 2020 | PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization · ICML 2020 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.4 | 1 | 2020 | PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization · ICML 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
fact extraction |
0.4 | 1 | 2019 | Assessing The Factual Accuracy of Generated Text · KDD 2019 |
Natural language and speech › Language models and text generation
text generation |
0.2 | 1 | 2023 | Out-of-Distribution Detection and Selective Generation for Conditional Language Models · ICLR 2023 |
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty calibration |
0.2 | 1 | 2023 | Calibrating Sequence likelihood Improves Conditional Language Generation · ICLR 2023 |
Methods — techniques the papers use, named apart from their topics
direct preference optimization · 1.6sequence likelihood calibration · 1.4supervised fine-tuning · 0.9data augmentation · 0.9chain-of-thought · 0.9causal framework · 0.9KTO · 0.9rejection sampling · 0.8selective generation · 0.7out-of-distribution detection · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RRM: Robust Reward Model Training Mitigates Reward HackingabstractReward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. However, traditional RM training, which relies on response pairs tied to specific prompts, struggles to disentangle prompt-driven preferences from prompt-independent artifacts, such as response length and format. In this work, we expose a fundamental limitation of current RM training methods, where RMs fail to effectively distinguish between contextual signals and irrelevant artifacts when determining preferences. To address this, we introduce a causal framework that learns preferences independent of these artifacts and propose a novel data augmentation technique designed to eliminate them. Extensive experiments show that our approach successfully filters out undesirable artifacts, yielding a more robust reward model (RRM). Our RRM improves the performance of a pairwise reward model trained on Gemma-2-9b-it, on Reward-Bench, increasing accuracy from 80.61% to 84.15%. Additionally, we train two DPO policies using both the RM and RRM, demonstrating that the RRM significantly enhances DPO-aligned policies, improving MT-Bench scores from 7.27 to 8.31 and length-controlled win-rates in AlpacaEval-2 from 33.46% to 52.49%. Tianqi Liu 0002, Wei Xiong 0015, Jie Ren 0006, Lichang Chen, Rishabh Joshi, Zhen Qin 0001, Tianhe Yu, Daniel Sohn, Anastasia Makarova, Jeremiah Z. Liu, Bilal Piot, Abraham Ittycheriah, Aviral Kumar, Mohammad Saleh |
ICLR | 18 |
| 2025 | Building Math Agents with Multi-Turn Iterative Preference LearningabstractRecent studies have shown that large language models' (LLMs) mathematical problem-solving capabilities can be enhanced by integrating external tools, such as code interpreters, and employing multi-turn Chain-of-Thought (CoT) reasoning. While current methods focus on synthetic data generation and Supervised Fine-Tuning (SFT), this paper studies the complementary direct preference learning approach to further improve model performance. However, existing direct preference learning algorithms are originally designed for the single-turn chat task, and do not fully address the complexities of multi-turn reasoning and external tool integration required for tool-integrated mathematical reasoning tasks. To fill in this gap, we introduce a multi-turn direct preference learning framework, tailored for this context, that leverages feedback from code interpreters and optimizes trajectory-level preferences. This framework includes multi-turn DPO and multi-turn KTO as specific implementations. The effectiveness of our framework is validated through training of various language models using an augmented prompt set from the GSM8K and MATH datasets. Our results demonstrate substantial improvements: a supervised fine-tuned Gemma-1.1-it-7B model's performance increased from 77.5% to 83.9% on GSM8K and from 46.1% to 51.2% on MATH. Similarly, a Gemma-2-it-9B model improved from 84.1% to 86.3% on GSM8K and from 51.0% to 54.5% on MATH. Wei Xiong 0015, Chengshuai Shi, Aviv Rosenberg 0002, Zhen Qin 0001, Daniele Calandriello, Misha Khalman, Rishabh Joshi, Bilal Piot, Mohammad Saleh, Tong Zhang 0001, Tianqi Liu 0002 |
ICLR | 10 |
| 2025 | LiPO: Listwise Preference Optimization through Learning-to-RankabstractTianqi Liu, Zhen Qin, Junru Wu, Jiaming Shen, Misha Khalman, Rishabh Joshi, Yao Zhao, Mohammad Saleh, Simon Baumgartner, Jialu Liu, Peter J Liu, Xuanhui Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Tianqi Liu 0002, Zhen Qin 0001, Misha Khalman, Rishabh Joshi, Mohammad Saleh, Simon Baumgartner, Peter J. Liu, Xuanhui Wang |
NAACL (Long Papers) | 8 |
| 2024 | Statistical Rejection Sampling Improves Preference OptimizationabstractImproving the alignment of language models with human preferences remains an active research challenge. Previous approaches have primarily utilized online Reinforcement Learning from Human Feedback (RLHF). Recently, offline methods such as Sequence Likelihood Calibration (SLiC) and Direct Preference Optimization (DPO) have emerged as attractive alternatives, offering improvements in stability and scalability while maintaining competitive performance. SLiC refines its loss function using sequence pairs sampled from a supervised fine-tuned (SFT) policy, while DPO directly optimizes language models based on preference data, foregoing the need for a separate reward model. However, the maximum likelihood estimator (MLE) of the target optimal policy requires labeled preference pairs sampled from that policy. The absence of a reward model in DPO constrains its ability to sample preference pairs from the optimal policy. Meanwhile, SLiC can only sample preference pairs from the SFT policy. To address these limitations, we introduce a novel approach called Statistical Rejection Sampling Optimization (RSO) designed to source preference data from the target optimal policy using rejection sampling, enabling a more accurate estimation of the optimal policy. We also propose a unified framework that enhances the loss functions used in both SLiC and DPO from a preference modeling standpoint. Through extensive experiments across diverse tasks, we demonstrate that RSO consistently outperforms both SLiC and DPO as evaluated by both Large Language Models (LLMs) and human raters. Tianqi Liu 0002, Rishabh Joshi, Misha Khalman, Mohammad Saleh, Peter J. Liu |
ICLR | 5 |
| 2023 | Out-of-Distribution Detection and Selective Generation for Conditional Language Models
Jie Ren 0006, Jiaming Luo, Kundan Krishna, Mohammad Saleh, Balaji Lakshminarayanan, Peter J. Liu |
ICLR | 5 |
| 2023 | Calibrating Sequence likelihood Improves Conditional Language Generation
Misha Khalman, Rishabh Joshi, Shashi Narayan, Mohammad Saleh, Peter J. Liu |
ICLR | 5 |
| 2022 | Virtual Interactive Imagery: Virtual Reality in the Treatment of ObesityabstractThis paper proposes a multidisciplinary framework that applies psychological treatment. The framework uses virtual reality to apply guided imagery scenarios, where a person imagines a certain situation positively, to better deal with similar ones in reality. The idea is to use virtual reality to achieve self-efficacy; a self-belief that empowers success. The proposed concept of "Virtual Interactive Imagery" represents immersing users in virtual images to live and interact in a predeveloped scenario positively, instead of having people depend on a facilitator to guide them through imagination. The paper focuses on obesity treatment and the role of imagination in human weight loss. The methodology of creating a virtual system is presented, and a proof of concept for the system is developed and tested through a usability test. Mariam Salim, Mohammad Saleh, Osama Halabi |
CW | 2 |
| 2021 | Augmented reality flavor: cross-modal mapping across gustation, olfaction, and visionabstractGustatory display research is still in its infancy despite being one of the essential everyday senses that human practice while eating and drinking. Indeed, the most important and frequent tasks that our brain deals with every day are foraging and feeding. The recent studies by psychologists and cognitive neuroscientist revealed how complex multisensory rely on the integration of cues from all the human senses in any flavor experiences. The perception of flavor is multisensory and involves combinations of gustatory and olfactory stimuli. The cross-modal mapping between these modalities needs to be more explored in the virtual environment and simulation, especially in liquid food. In this paper, we present a customized wearable Augmented Reality (AR) system and olfaction display to study the effect of vision and olfaction on the gustatory sense. A user experiment and extensive analysis conducted to study the influence of each stimulus on the overall flavor, including other factors like age, previous experience in Virtual Reality (VR)/AR, and beverage consumption. The result showed that smell contributes strongly to the flavor with less contribution to the vision. However, the combination of these stimuli can deliver richer experience and a higher belief rate. Beverage consumption had a significant effect on the flavor belief rate. Experience is correlated with stimulus and age is correlated with belief rate, and both indirectly affected the belief rate. Osama Halabi, Mohammad Saleh |
Multim. Tools Appl. | 2 |
| 2020 | PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationabstractRecent work pre-training Transformers with self-supervised objectives on large text corpora has shown great success when fine-tuned on downstream NLP tasks including text summarization. However, pre-training objectives tailored for abstractive text summarization have not been explored. Furthermore there is a lack of systematic evaluation across diverse domains. In this work, we propose pre-training large Transformer-based encoder-decoder models on massive text corpora with a new self-supervised objective. In PEGASUS, important sentences are removed/masked from an input document and are generated together as one output sequence from the remaining sentences, similar to an extractive summary. We evaluated our best PEGASUS model on 12 downstream summarization tasks spanning news, science, stories, instructions, emails, patents, and legislative bills. Experiments demonstrate it achieves state-of-the-art performance on all 12 downstream datasets measured by ROUGE scores. Our model also shows surprising performance on low-resource summarization, surpassing previous state-of-the-art results on 6 datasets with only 1000 examples. Finally we validated our results using human evaluation and show that our model summaries achieve human performance on multiple datasets. Jingqing Zhang, Mohammad Saleh, Peter J. Liu |
ICML | 3 |
| 2019 | Assessing The Factual Accuracy of Generated TextabstractWe propose a model-based metric to estimate the factual accuracy of generated text that is complementary to typical scoring schemes like ROUGE (Recall-Oriented Understudy for Gisting Evaluation) and BLEU (Bilingual Evaluation Understudy). We introduce and release a new large-scale dataset based on Wikipedia and Wikidata to train relation classifiers and end-to-end fact extraction models. The end-to-end models are shown to be able to extract complete sets of facts from datasets with full pages of text. We then analyse multiple models that estimate factual accuracy on a Wikipedia text summarization task, and show their efficacy compared to ROUGE and other model-free variants by conducting a human evaluation study. Ben Goodrich, Vinay Rao, Peter J. Liu, Mohammad Saleh |
KDD | 4 |
| 2018 | Generating Wikipedia by Summarizing Long Sequences
Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, Noam Shazeer |
ICLR (Poster) | 2 |
| 2017 | Biomedical Epilepsy Multiprocessor Network on Chip EEG Utilizing IEEE802.11n SystemsabstractThis paper presents a novel wearable biomedical Network on Chip (NoC) concept development to monitor and predict irregular brain waves as advanced sensitive portable for an electroencephalogram (EEG) analysis device. The proposed device will monitor brain's spontaneous electrical activity in normal and abnormal situations for specific patients suffering from different types of epilepsy. This NOC would be able to predict the severity of the forthcoming epileptic attack. Meanwhile, this device will alert epileptic patients by giving an alarm on detection of any kind of abnormal brain electrical activity. The EEG NoC reads brain signals from a wireless sensor on the located on the patient's scalp, and runs parallel processing and filtering for the brainwaves. This process makes the detection of brain abnormalities possible, and many patients can be saved by predicting the time of epileptic seizure. When such a prediction occurs, alarm signals are sent to the patient to take protective measures. In this way, it will help patients to prevent themselves from different kinds of injuries and risky behaviors occurring during epilepsy attack or after that. Mohammad Saleh, Jaafar Gaber, Maxime Wack |
MASS | 1 |
| 2014 | Performance Comparison of classification algorithms for EEG-based remote epileptic seizure detection in Wireless Sensor NetworksabstractIdentification of epileptic seizure remotely by analyzing the electroencephalography (EEG) signal is very important for scalable sensor-based health systems. Classification is the most important technique for wide-ranging applications to categorize the items according to its features with respect to predefined set of classes. In this paper, we conduct a performance evaluation based on the noiseless and noisy EEG-based epileptic seizure data using various classification algorithms including BayesNet, DecisionTable, IBK, J48/C4.5, and VFI. The reconstructed and noisy EEG data are decomposed with discrete cosine transform into several sub-bands. In addition, some of statistical features are extracted from the wavelet coefficients to represent the whole EEG data inputs into the classifiers. Benchmark on widely used dataset is utilized for automatic epileptic seizure detection including both normal and epileptic EEG datasets. The classification accuracy results confirm that the selected classifiers have greater potentiality to identify the noisy epileptic disorders. Khalid Abualsaud, Massudi Mahmuddin, Mohammad Saleh, Amr Mohamed 0001 |
AICCSA | 3 |
| 2014 | ConProve: A conceptual prover systemabstractConProve is an automated prover for propositional logic. It takes, as an input, a set of propositional formulas and proves whether a goal holds or not. ConProve converts each formula to its corresponding Truth Table Binary Relation (TTBR) considered also as a formal context (FC). The objects in FC correspond to all possible formulas interpretations (in terms of their truth value assignments), and the properties in FC correspond to the terms. When the function the 'BuildContext' function, ConProve starts the new goal proving. Firstly, it adds the goal negation to the set of formulas and constructs the formal contexts (FCs) relating formulas to terms. Secondly, it makes the FCs grouping and deduces, based on the conceptual reasoning, if the goal holds. The tool offers a user-friendly interface allowing the editing of the set of formulas as well as the visualization of the reasoning steps. Besides the tool, the paper illustrates the importance of the conceptual reasoning in deriving new conclusions as well as in discovering new, possibly implications by applying the extended Galois Connection. Samir Elloumi, Ali Jaoua, Bilel Boulifa, Mohammad Saleh, Jameela Al Otaibi |
AICCSA | 4 |