Mohammad Hossein Saliminabi

dblp:421/1615 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0005-0065-0972ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Failing Forward: Understanding Query Failure in Retrieval, Judgment, and Generation
abstract
Modern information retrieval pipelines combine retrieval, LLM-based generation, and LLM-based judgment, and a poor outcome may originate in any of the three stages. Existing work studies these failures in isolation. This paper instead asks whether query difficulty itself transfers across the three stages: are the same queries hard to retrieve, hard to generate for, and hard to judge? Using four years of TREC Deep Learning benchmarks (2019–2022), we define hard-to-retrieve, hard-to-generate, and hard-to-judge query sets under a unified quartile-based operationalization and analyze their overlap, their stability across system configurations, and the linguistic and semantic causes of failure in each task. We find that the three sets overlap only weakly; three-way overlap is at or below the level expected under independence, indicating that difficulty is largely task-conditioned and does not transfer reliably across stages. The overlap structure is nonetheless stable across retrievers, generators, and judging setups, suggesting that task-specific difficulty is driven by query characteristics interacting with each task's inductive biases rather than by model choice. We further induce a data-driven typology of failure causes and show that conditioning generation on task-relevant difficulty cues yields consistent gains in answer quality.
Negar Arabzadeh, Mohammad Hossein Saliminabi, Dimitrios Androutsos, Morteza Zihayat, Ebrahim Bagheri
SIGIR3
2026 LearnDCG: End-to-End Joint Optimization of Ranker and Loss in Neural Ranking
Mohammad Hossein Saliminabi, Dimitrios Androutsos, Ebrahim Bagheri
SIGIR1
2026 Diffusion-based generative modeling for expert team formation
abstract
Forming effective expert teams is central to domains where solving complex problems requires diverse, complementary skills. However, automating this task is highly challenging due to sparse co-occurrence data, long-tailed expert participation, and the combinatorial complexity of unseen skill configurations. Existing graph-based, probabilistic, and neural approaches often struggle with generalization, fairness, and robustness, leading to biased selections that favor historically popular experts over more suitable candidates. To address these challenges, we propose a generative framework for expert team formation based on denoising diffusion probabilistic models. We cast team formation as skill-conditioned imputation (i.e., inpainting), where skills are treated as observed context and the expert component is generated via conditional diffusion sampling. This design enables our method to preserve semantic skill–expert alignment, mitigate data sparsity, and generate diverse yet contextually coherent teams. Extensive experiments on DBLP and DOTA2 datasets show that our model consistently outperforms state-of-the-art baselines, achieving over 3 × higher recall (16.4% vs. 5.0%) and MAP (9.7% vs. 2.2%) on DBLP, while delivering more than 5 × improvement in MRR (13.3% vs. 2.5%) on DOTA2. Fairness analysis further demonstrates that our method reduces average overlap with the top-100 most popular experts to 2.6, compared to 86.7 for the strongest baseline, and achieves near-optimal diversity with NDKL ≈ 0.1 under high non-popular expert ratios. For reproducibility purposes , we made our code and model publicly available at https://github.com/17shiraz/DiffTF .
Mohammad Hossein Saliminabi, Sajad Ebrahimi 0001, Radin Hamidi Rad, Dimitrios Androutsos, Fattane Zarrinkalam, Ebrahim Bagheri
Inf. Process. Manag.1
2025 LLM-as-a-Judge in Entity Retrieval: Assessing Explicit and Implicit Relevance
abstract
Entity retrieval plays a critical role in information access systems, yet the development and evaluation of retrieval models remain constrained by the limited availability of high-quality supervision. While recent work has demonstrated the utility of large language models (LLMs) as relevance assessors in passage and document retrieval, their reliability in the context of entity retrieval-where targets are abstract, underspecified, and often semantically sparse-remains unexplored. In this work, we evaluate LLM-based judgments against two complementary supervision signals: human-annotated relevance labels from the DBpedia-Entity benchmark and implicit feedback from user clicks in the LaQuE dataset. We show that LLMs exhibit strong agreement with expert annotations and replicate user click patterns with over 91% agreement, suggesting alignment with behavioral judgments despite noisy input queries. We further identify and analyze systematic mismatches for user clicks on irrelevant entities. Our findings establish LLMs not only as effective annotators for entity relevance judgment-even when given only the entity title-but also as powerful tools for predicting click-through behavior and simulating explainable user intent. Our code, prompts, and data are publicly available at: https://github.com/17shiraz/ClickLLM
Mohammad Hossein Saliminabi, Negar Arabzadeh, Dimitrios Androutsos, Morteza Zihayat, Ebrahim Bagheri
CIKM1