VLDB 2026 Research / reviewers in the wild / expert
Ehsan Shareghi
dblp:09/7859
· DBLP profile ↗
28ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0001-9119-8638ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 5 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical ReasoningabstractLarge language models have demonstrated remarkable capabilities in complex mathematical reasoning tasks, but they inevitably generate errors throughout multi-step solutions. Process-level Reward Models (PRMs) have shown great promise by providing supervision and evaluation at each intermediate step, thereby effectively improving the models’ reasoning abilities. However, training effective PRMs requires high-quality process reward data, yet existing methods for constructing such data are often labour-intensive or inefficient. In this paper, we propose an uncertainty-driven framework for automated process reward data construction, encompassing both data generation and annotation processes for PRMs. Additionally, we identify the limitations of both majority vote and PRMs, and introduce two generic uncertainty-aware output aggregation methods: Hybrid Majority Reward Vote and Weighted Reward Frequency Vote, which combine the strengths of majority vote with PRMs. Extensive experiments on ProcessBench, MATH, and GSMPlus show the effectiveness and efficiency of the proposed PRM data construction framework, and demonstrate that the two output aggregation methods further improve the mathematical reasoning abilities across diverse PRMs. Jiuzhou Han, Wray L. Buntine, Ehsan Shareghi |
AAAI | 3 |
| 2026 | Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement LearningabstractZheng Hui, Yijiang River Dong, Sanhanat Sivapiromrat, Ehsan Shareghi, Nigel Collier. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zheng Hui, Yijiang River Dong, Sanhanat Sivapiromrat, Ehsan Shareghi, Nigel Collier |
ACL (1) | 4 |
| 2026 | Could language models win the International Linguistics Olympiad?abstractLinguistic puzzles, wherein the solver must deduce rules of an unfamiliar language purely in-context, represent a uniquely perplexing problem format even for state-of-the-art large language models.Yet by exploring various inference-time scaling methods, we demonstrate that language models' performance on these problems can be improved without the need for fine-tuning or providing supplementary linguistic context.To this end, this paper introduces the first domain-specific inference-time scaling framework for linguistic puzzles, which we use to improve the performance of three model families -R1 (Deepseek), Gemini 2.5 Flash (Google), and Llama 3.3 70B Instruct (Meta) -on a challenging Linguistics Olympiad-based benchmark by 4.9, 13.1, and 4.9 percentage points, respectively.Nonetheless, even when multiple optimisations are applied, we find that LLMs' linguistic puzzle performance remains well below comparable mathematical and commonsense benchmarks, and we speculate as to why linguistic reasoning continues to pose a distinctive challenge for even the most capable large language models. 1 Jamie Garnham, Ehsan Shareghi |
CoNLL | 2 |
| 2025 | Logical Reasoning with Outcome Reward Models for Test-Time ScalingabstractLogical reasoning is a critical benchmark for evaluating the capabilities of large language models (LLMs), as it reflects their ability to derive valid conclusions from given premises.While the combination of test-time scaling with dedicated outcome or process reward models has opened up new avenues to enhance LLMs performance in complex reasoning tasks, this space is under-explored in deductive logical reasoning.We present a set of Outcome Reward Models (ORMs) for deductive reasoning.To train the ORMs we mainly generate data using Chain-of-Thought (CoT) with single and multiple samples.Additionally, we propose a novel tactic to further expand the type of errors covered in the training dataset of the ORM.In particular, we propose an echo generation technique that leverages LLMs' tendency to reflect incorrect assumptions made in prompts to extract additional training data, covering previously unexplored error types.While a standard CoT chain may contain errors likely to be made by the reasoner, the echo strategy deliberately steers the model toward incorrect reasoning.We show that ORMs trained on CoT and echoaugmented data demonstrate improved performance on the FOLIO, JustLogic, and ProverQA datasets across four different LLMs. Ramya Keerthy Thatikonda, Wray L. Buntine, Ehsan Shareghi |
EMNLP | 3 |
| 2025 | Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language ModelsabstractLarge Audio Language Models (LALMs) have extended the capabilities of Large Language Models (LLMs) by enabling audio-based human interactions.However, recent research has revealed that LALMs remain vulnerable to harmful queries due to insufficient safetyalignment.Despite advances in defence measures for text and vision LLMs, effective safetyalignment strategies and audio-safety dataset specifically targeting LALMs are notably absent.Meanwhile defence measures based on Supervised Fine-tuning (SFT) struggle to address safety improvement while avoiding overrejection issues, significantly compromising helpfulness.In this work, we propose an unsupervised safety-fine-tuning strategy as remedy that reshapes model's representation space to enhance existing LALMs safety-alignment while balancing the risk of over-rejection.Our experiments, conducted across three generations of Qwen LALMs, demonstrate that our approach significantly improves LALMs safety under three modality input conditions (audiotext, text-only, and audio-only) while increasing over-rejection rate by only 0.88% on average. 1 Warning: this paper contains harmful examples. Lizhen Qu, Ehsan Shareghi, Gholamreza Haffari |
EMNLP | 3 |
| 2025 | All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model ReasoningabstractConfidence estimation is essential for the reliable deployment of large language models (LLMs).Existing methods are primarily designed for factual QA tasks and often fail to generalize to reasoning tasks.To address this gap, we propose a set of training-free, graph-based confidence estimation methods tailored to reasoning tasks.Our approach models reasoning paths as directed graphs and estimates confidence by exploiting graph properties such as centrality, path convergence, and path weighting.Experiments with two LLMs on three reasoning datasets demonstrate improved confidence estimation and enhanced performance on two downstream tasks. Caiqi Zhang, Ehsan Shareghi, Nigel Collier |
EMNLP | 3 |
| 2025 | Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language ModelsabstractLarge Language Models (LLMs) are expected to be predictable and trustworthy to support reliable decision-making systems. Yet current LLMs often show inconsistencies in their judgments. In this work, we examine \textit{logical preference consistency} as a foundational requirement for building more dependable LLM systems, ensuring stable and coherent decision-making while minimizing erratic or contradictory outputs.
To quantify the logical preference consistency, we propose a universal evaluation framework based on three fundamental properties: *transitivity*, *commutativity* and *negation invariance*.
Through extensive experimentation across diverse LLMs, we demonstrate that these properties serve as strong indicators of judgment robustness.
Furthermore, we introduce a data refinement and augmentation technique, REPAIR, that enhances logical consistency while maintaining alignment with human preferences. Finally, we show that improving consistency leads to better performance in LLM-driven logic-based algorithms, reinforcing stability and coherence in decision-making systems. Yinhong Liu, Zhijiang Guo, Tianya Liang, Ehsan Shareghi, Ivan Vulic, Nigel Collier |
ICML | 4 |
| 2025 | SpeechDialogueFactory: A Framework for Natural Speech Dialogue Generation
Minghan Wang, Ye Bai 0002, Thuy-Trang Vu, Ehsan Shareghi, Gholamreza Haffari |
INTERSPEECH | 5 |
| 2025 | Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal ModelsabstractHao Yang, Lizhen Qu, Ehsan Shareghi, Gholamreza Haffari. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Lizhen Qu, Ehsan Shareghi, Gholamreza Haffari |
NAACL (Long Papers) | 3 |
| 2024 | Harnessing the Power of Large Language Models for Natural Language to First-Order Logic TranslationabstractAdvancements in logical reasoning, utilizing LLMs to convert natural language into logical symbolism, combined with the use of external theorem provers, have repositioned the symbolic approach as a central point of interest. The main challenge within this paradigm lies in the LLMs’ capability to accurately translate natural language (NL) statements into first-order-logic (FOL) expressions. Although LLMs have shown notable success, there remains a gap in understanding the limitations and challenges they encounter in NL-FOL translation. This is primarily due to the absence of datasets and evaluation test beds at the required fine-grained level. We present MALLS, a dataset of 28K diverse and verified sentence-level NL-FOL pairs collected from GPT4. We utilize a combined strategy of FOL rule parsing, human annotation, and automatic filtering to ensure quality. We also present LogicLLaMA, a LLaMA2-7B/13B fine-tuned on MALLS for NL-FOL translation, which can be used standalone or to correct previously generated rules by GPT3.5 after being further fine-tuned via a novel reinforcement learning with human feedback (RLHF) framework. We benchmark a wide range of LLMs on MALLS and previous datasets, highlighting weaknesses in them in NL-FOL translation and demonstrating the advantages of MALLS. We also show that LogicLLaMA achieves GPT4-level performance and can generalize to other datasets. Project repo is available at https://github.com/gblackout/LogicLLaMA Yuan Yang 0007, Siheng Xiong, Ali Payani, Ehsan Shareghi, Faramarz Fekri |
ACL (1) | 4 |
| 2024 | Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and InsightsabstractLarge Multimodal Models (LMMs) have achieved great success recently, demonstrating a strong capability to understand multimodal information and to interact with human users.Despite the progress made, the challenge of detecting high-risk interactions in multimodal settings, and in particular in speech modality, remains largely unexplored.Conventional research on risk for speech modality primarily emphasises the content (e.g., what is captured as transcription).However, in speechbased interactions, paralinguistic cues in audio can significantly alter the intended meaning behind utterances.In this work, we propose a speech-specific risk taxonomy, covering 8 risk categories under hostility (malicious sarcasm and threats), malicious imitation (age, gender, ethnicity), and stereotypical biases (age, gender, ethnicity).Based on the taxonomy, we create a small-scale dataset for evaluating current LMMs capability in detecting these categories of risk.We observe even the latest models remain ineffective to detect various paralinguistic-specific risks in speech (e.g., Gemini 1.5 Pro is performing only slightly above random baseline).1 Warning: this paper contains biased and offensive examples. A Experimental ResultsWe provide complete experimental results including accuracy and macro-averaged F1 score as metrics in Table 6. B Examples for Sub-categoriesWe provide examples from our text sets for each sub-category in Table 7. C Description of Speech Generation from AudioboxWe provide the examples for speech generation from Audiobox in Table 8. D Prompting StrategiesWe provide a complete list covering prompting strategies used in our evaluation experiments and analysis in Table 9 and Table 10, respectively. E Computational Hardware and APIWe conduct all our evaluation experiments and analysis on 4×A100 GPUs.No fine-tuning was done and the experiments only involved inference.For Gemini 1.5 Pro we used gemini-1.5-proAPI, and for GPT-4 we used gpt-4-turbo API.Temperature was set to 0 and sampling at decoding was switched off. Lizhen Qu, Ehsan Shareghi, Reza Haf |
EMNLP | 3 |
| 2023 | On Reality and the Limits of Language Data: Aligning LLMs with Human Norms
Nigel Collier, Fangyu Liu 0001, Ehsan Shareghi |
CogSci | 3 |
| 2023 | A Minimal Approach for Natural Language Action Space in Text-based GamesabstractText-based games (TGs) are language-based interactive environments for reinforcement learning.While language models (LMs) and knowledge graphs (KGs) are commonly used for handling large action space in TGs, it is unclear whether these techniques are necessary or overused.In this paper, we revisit the challenge of exploring the action space in TGs and propose ϵ-admissible exploration, a minimal approach of utilizing admissible actions, for training phase.Additionally, we present a textbased actor-critic (TAC) agent that produces textual commands for game, solely from game observations, without requiring any KG or LM.Our method, on average across 10 games from Jericho, outperforms strong baselines and stateof-the-art agents that use LM and KG.Our approach highlights that a much lighter model design, with a fresh perspective on utilizing the information within the environments, suffices for an effective exploration of exponentially large action spaces. 1 Dongwon Ryu, Gholamreza Haffari, Shirui Pan, Ehsan Shareghi |
CoNLL | 5 |
| 2023 | Investigating Pre-trained Audio Encoders in the Low-Resource ConditionabstractPre-trained speech encoders have been central to pushing state-of-the-art results across various speech understanding and generation tasks. Nonetheless, the capabilities of these encoders in low-resource settings are yet to be thoroughly explored. To address this, we conduct a comprehensive set of experiments using a representative set of 3 state-of-the-art encoders (Wav2vec2, WavLM, Whisper) in the low-resource setting across 7 speech understanding and generation tasks. We provide various quantitative and qualitative analyses on task performance, convergence speed, and representational properties of the encoders. We observe a connection between the pre-training protocols of these encoders and the way in which they capture information in their internal layers. In particular, we observe the Whisper encoder exhibits the greatest low-resource capabilities on content-driven tasks in terms of performance and convergence speed. Jinming Zhao, Gholamreza Haffari, Ehsan Shareghi |
INTERSPEECH | 4 |
| 2022 | Rewire-then-Probe: A Contrastive Recipe for Probing Biomedical Knowledge of Pre-trained Language ModelsabstractZaiqiao Meng, Fangyu Liu, Ehsan Shareghi, Yixuan Su, Charlotte Collins, Nigel Collier. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Zaiqiao Meng, Fangyu Liu 0001, Ehsan Shareghi, Yixuan Su, Charlotte Collins, Nigel Collier |
ACL (1) | 3 |
| 2022 | Self-supervised Graph Masking Pre-training for Graph-to-Text GenerationabstractLarge-scale pre-trained language models (PLMs) have advanced Graph-to-Text (G2T) generation by processing the linearised version of a graph.However, the linearisation is known to ignore the structural information.Additionally, PLMs are typically pre-trained on free text which introduces domain mismatch between pre-training and downstream G2T generation tasks.To address these shortcomings, we propose graph masking pre-training strategies that neither require supervision signals nor adjust the architecture of the underlying pre-trained encoder-decoder model.When used with a pre-trained T5, our approach achieves new state-of-the-art results on WebNLG+2020 and EventNarrative G2T generation datasets.Our method also shows to be very effective in the low-resource setting. 1 Jiuzhou Han, Ehsan Shareghi |
EMNLP | 2 |
| 2022 | M-Adapter: Modality Adaptation for End-to-End Speech-to-Text TranslationabstractEnd-to-end speech-to-text translation models are often initialized with pre-trained speech encoder and pre-trained text decoder. This leads to a significant training gap between pretraining and fine-tuning, largely due to the modality differences between speech outputs from the encoder and text inputs to the decoder. In this work, we aim to bridge the modality gap between speech and text to improve translation quality. We propose M-Adapter, a novel Transformer-based module, to adapt speech representations to text. While shrinking the speech sequence, M-Adapter produces features desired for speech-to-text translation via modelling global and local dependencies of a speech sequence. Our experimental results show that our model outperforms a strong baseline by up to 1 BLEU score on the Must-C En→DE dataset. Jinming Zhao, Gholamreza Haffari, Ehsan Shareghi |
INTERSPEECH | 4 |
| 2021 | A Closer Look at Few-Shot Crosslingual Transfer: The Choice of Shots MattersabstractMengjie Zhao, Yi Zhu, Ehsan Shareghi, Ivan Vulić, Roi Reichart, Anna Korhonen, Hinrich Schütze. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ehsan Shareghi, Ivan Vulic, Roi Reichart, Anna Korhonen, Hinrich Schütze |
ACL/IJCNLP (1) | 3 |
| 2021 | Combining Deep Generative Models and Multi-lingual Pretraining for Semi-supervised Document ClassificationabstractSemi-supervised learning through deep generative models and multi-lingual pretraining techniques have orchestrated tremendous success across different areas of NLP.Nonetheless, their development has happened in isolation, while the combination of both could potentially be effective for tackling task-specific labelled data shortage.To bridge this gap, we combine semi-supervised deep generative models and multi-lingual pretraining to form a pipeline for document classification task.Compared to strong supervised learning baselines, our semi-supervised classification framework is highly competitive and outperforms the state-of-the-art counterparts in lowresource settings across several languages.1 Ehsan Shareghi, Yingzhen Li, Roi Reichart, Anna Korhonen |
EACL | 2 |
| 2021 | Mixture-of-Partitions: Infusing Large Biomedical Knowledge Graphs into BERTabstractInfusing factual knowledge into pretrained models is fundamental for many knowledgeintensive tasks.In this paper, we propose Mixture-of-Partitions (MoP), an infusion approach that can handle a very large knowledge graph (KG) by partitioning it into smaller subgraphs and infusing their specific knowledge into various BERT models using lightweight adapters.To leverage the overall factual knowledge for a target task, these sub-graph adapters are further fine-tuned along with the underlying BERT through a mixture layer.We evaluate our MoP with three biomedical BERTs (SciBERT, BioBERT, PubmedBERT) on six downstream tasks (inc.NLI, QA, Classification), and the results show that our MoP consistently enhances the underlying BERTs in task performance, and achieves new SOTA performances on five evaluated datasets.1 Zaiqiao Meng, Fangyu Liu 0001, Thomas Hikaru Clark, Ehsan Shareghi, Nigel Collier |
EMNLP (1) | 4 |
| 2021 | It Is Not As Good As You Think! Evaluating Simultaneous Machine Translation on Interpretation DataabstractMost existing simultaneous machine translation (SiMT) systems are trained and evaluated on offline translation corpora.We argue that SiMT systems should be trained and tested on real interpretation data.To illustrate this argument, we propose an interpretation test set and conduct a realistic evaluation of SiMT trained on offline translations.Our results, on our test set along with 3 existing smaller scale language pairs, highlight the difference of up-to 13.83 BLEU score when SiMT models are evaluated on translation vs interpretation data.In the absence of interpretation training data, we propose a translationto-interpretation (T2I) style transfer method which allows converting existing offline translations into interpretation-style data, leading to up-to 2.8 BLEU improvement.However, the evaluation gap remains notable, calling for constructing large-scale interpretation corpora better suited for evaluating and developing SiMT systems. 1 Jinming Zhao, Philip Arthur, Gholamreza Haffari, Trevor Cohn, Ehsan Shareghi |
EMNLP (1) | 5 |
| 2021 | Self-Alignment Pretraining for Biomedical Entity RepresentationsabstractFangyu Liu, Ehsan Shareghi, Zaiqiao Meng, Marco Basaldella, Nigel Collier. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Fangyu Liu 0001, Ehsan Shareghi, Zaiqiao Meng, Marco Basaldella, Nigel Collier |
NAACL-HLT | 2 |
| 2020 | COMETA: A Corpus for Medical Entity Linking in the Social MediaabstractWhilst there has been growing progress in Entity Linking (EL) for general language, existing datasets fail to address the complex nature of health terminology in layman's language.Meanwhile, there is a growing need for applications that can understand the public's voice in the health domain.To address this we introduce a new corpus called COMETA, consisting of 20k English biomedical entity mentions from Reddit expert-annotated with links to SNOMED CT, a widely-used medical knowledge graph.Our corpus satisfies a combination of desirable properties, from scale and coverage to diversity and quality, that to the best of our knowledge has not been met by any of the existing resources in the field.Through benchmark experiments on 20 EL baselines from string-to neural-based models we shed light on the ability of these systems to perform complex inference on entities and concepts under 2 challenging evaluation scenarios.Our experimental results on COMETA illustrate that no golden bullet exists and even the best mainstream techniques still have a significant performance gap to fill, while the best solution relies on combining different views of data. Marco Basaldella, Fangyu Liu 0001, Ehsan Shareghi, Nigel Collier |
EMNLP (1) | 3 |
| 2017 | Compressed Nonparametric Language ModellingabstractHierarchical Pitman-Yor Process priors are compelling for learning language models, outperforming point-estimate based methods. However, these models remain unpopular due to computational and statistical inference issues, such as memory and time usage, as well as poor mixing of sampler. In this work we propose a novel framework which represents the HPYP model compactly using compressed suffix trees. Then, we develop an efficient approximate inference scheme in this framework that has a much lower memory footprint compared to full HPYP and is fast in the inference time. The experimental results illustrate that our model can be built on significantly larger datasets compared to previous HPYP models, while being several orders of magnitudes smaller, fast for training and inference, and outperforming the perplexity of the state-of-the-art Modified Kneser-Ney count-based LM smoothing by up to 15%. Ehsan Shareghi, Gholamreza Haffari, Trevor Cohn |
IJCAI | 1 |
| 2016 | Richer Interpolative Smoothing Based on Modified Kneser-Ney Language ModelingabstractIn this work we present a generalisation of the Modified Kneser-Ney interpolative smoothing for richer smoothing via additional discount parameters.We provide mathematical underpinning for the estimator of the new discount parameters, and showcase the utility of our rich MKN language models on several European languages.We further explore the interdependency among the training data size, language model order, and number of discount parameters.Our empirical results illustrate that larger number of discount parameters, i) allows for better allocation of mass in the smoothing process, particularly on small data regime where statistical sparsity is severe, and ii) leads to significant reduction in perplexity, particularly for out-of-domain test sets which introduce higher ratio of out-ofvocabulary words. 1 Ehsan Shareghi, Trevor Cohn, Gholamreza Haffari |
EMNLP | 1 |
| 2016 | Fast, Small and Exact: Infinite-order Language Modelling with Compressed Suffix TreesabstractEfficient methods for storing and querying are critical for scaling high-order m-gram language models to large corpora. We propose a language model based on compressed suffix trees, a representation that is highly compact and can be easily held in memory, while supporting queries needed in computing language model probabilities on-the-fly. We present several optimisations which improve query runtimes up to 2500×, despite only incurring a modest increase in construction time and memory usage. For large corpora and high Markov orders, our method is highly competitive with the state-of-the-art KenLM package. It imposes much lower memory requirements, often by orders of magnitude, and has runtimes that are either similar (for training) or comparable (for querying). Ehsan Shareghi, Matthias Petri, Gholamreza Haffari, Trevor Cohn |
Trans. Assoc. Comput. Linguistics | 1 |
| 2015 | Compact, Efficient and Unlimited Capacity: Language Modeling with Compressed Suffix TreesabstractEfficient methods for storing and querying language models are critical for scaling to large corpora and high Markov orders.In this paper we propose methods for modeling extremely large corpora without imposing a Markov condition.At its core, our approach uses a succinct index -a compressed suffix tree -which provides near optimal compression while supporting efficient search.We present algorithms for on-the-fly computation of probabilities under a Kneser-Ney language model.Our technique is exact and although slower than leading LM toolkits, it shows promising scaling properties, which we demonstrate through ∞-order modeling over the full Wikipedia collection. Ehsan Shareghi, Matthias Petri, Gholamreza Haffari, Trevor Cohn |
EMNLP | 1 |
| 2015 | Structured Prediction of Sequences and Trees Using Infinite Contexts
Ehsan Shareghi, Gholamreza Haffari, Trevor Cohn, Ann E. Nicholson |
ECML/PKDD (2) | 1 |