Hosein Azarbonyad

dblp:118/2623 · DBLP profile ↗
← Back
18ranked-venue papers in the field
8as first author
7since 2021 · last 2026
0000-0003-0841-7516ORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 15 (6 first)Data Mining & Knowledge Discovery · 2 (1 first)Database Systems & Data Management · 1 (1 first)
YearPublicationVenuePosition
2026 CLEF 2026 SimpleText Track - Simplify Scientific Text (and Nothing More)
Liana Ermakova, Hosein Azarbonyad, Jan Bakker, Gautam Kishore Shahi, Benjamin Vendeville, Jaap Kamps
ECIR (4)2
2025 CLEF 2025 SimpleText Track - Simplify Scientific Text (and Nothing More)
Liana Ermakova, Hosein Azarbonyad, Jan Bakker, Benjamin Vendeville, Jaap Kamps
ECIR (5)2
2025 Question-Answer Extraction from Scientific Articles Using Knowledge Graphs and Large Language Models
abstract
When deciding to read an article or incorporate it into their research, scholars often seek to quickly identify and understand its main ideas.In this paper, we aim to extract these key concepts and contributions from scientific articles in the form of Question and Answer (QA) pairs.We propose two distinct approaches for generating QAs.The first approach involves selecting salient paragraphs, using a Large Language Model (LLM) to generate questions, ranking these questions by the likelihood of obtaining meaningful answers, and subsequently generating answers.This method relies exclusively on the content of the articles.However, assessing an article's novelty typically requires comparison with the existing literature.Therefore, our second approach leverages a Knowledge Graph (KG) for QA generation.We construct a KG by fine-tuning an Entity Relationship (ER) extraction model on scientific articles and using it to build the graph.We then employ a salient triplet extraction method to select the most pertinent ERs per article, utilizing metrics such as the centrality of entities based on a triplet TF-IDF-like measure.This measure assesses the saliency of a triplet based on its importance within the article compared to its prevalence in the literature.For evaluation, we generate QAs using both approaches and have them assessed by Subject Matter Experts (SMEs) through a set of predefined metrics to evaluate the quality of both questions and answers.Our evaluations demonstrate that the KG-based approach effectively captures the main ideas discussed in the articles.Furthermore, our findings indicate that fine-tuning the ER extraction model on our scientific corpus is crucial for extracting high-quality triplets from such documents.
Hosein Azarbonyad, Zi Long Zhu, Georgios Cheirmpos, Zubair Afzal, Vikrant Yadav, George Tsatsaronis 0001
SIGIR1
2024 CLEF 2024 SimpleText Track - Improving Access to Scientific Texts for Everyone
Liana Ermakova, Eric SanJuan, Stéphane Huet, Hosein Azarbonyad, Giorgio Maria Di Nunzio, Federica Vezzani, Jennifer D'Souza 0001, Salomon Kabongo, Hamed Babaei Giglou, Yue Zhang 0069, Sören Auer, Jaap Kamps
ECIR (6)4
2024 ScienceDirect Topic Pages: A Knowledge Base of Scientific Concepts Across Various Science Domains
Artemis Çapari, Hosein Azarbonyad, George Tsatsaronis 0001, Zubair Afzal, Judson Dunham
SIGIR2
2023 Generating Topic Pages for Scientific Concepts Using Scientific Publications
Hosein Azarbonyad, Zubair Afzal, George Tsatsaronis 0001
ECIR (2)1
2023 CLEF 2023 SimpleText Track - What Happens if General Users Search Scientific Texts?
Liana Ermakova, Eric SanJuan, Stéphane Huet, Olivier Augereau, Hosein Azarbonyad, Jaap Kamps
ECIR (3)5
2019 Domain Adaptation for Commitment Detection in Email
abstract
People often make commitments to perform future actions. Detecting commitments made in email (e.g., "I'll send the report by end of day'') enables digital assistants to help their users recall promises they have made and assist them in meeting those promises in a timely manner. In this paper, we show that commitments can be reliably extracted from emails when models are trained and evaluated on the same domain (corpus). However, their performance degrades when the evaluation domain differs. This illustrates the domain bias associated with email datasets and a need for more robust and generalizable models for commitment detection. To learn a domain-independent commitment model, we first characterize the differences between domains (email corpora) and then use this characterization to transfer knowledge between them. We investigate the performance of domain adaptation, namely transfer learning, at different granularities: feature-level adaptation and sample-level adaptation. We extend this further using a neural autoencoder trained to learn a domain-independent representation for training samples. We show that transfer learning can help remove domain bias to obtain models with less domain dependence. Overall, our results show that domain differences can have a significant negative impact on the quality of commitment detection models and that transfer learning has enormous potential to address this issue.
Hosein Azarbonyad, Robert Sim, Ryen W. White
WSDM1
2019 Learning to Transform, Combine, and Reason in Open-Domain Question Answering
abstract
Users seek direct answers to complex questions from large open-domain knowledge sources like the Web. Open-domain question answering has become a critical task to be solved for building systems that help address users' complex information needs. Most open-domain question answering systems use a search engine to retrieve a set of candidate documents, select one or a few of them as context, and then apply reading comprehension models to extract answers. Some questions, however, require taking a broader context into account, e.g., by considering low-ranked documents that are not immediately relevant, combining information from multiple documents, and reasoning over multiple facts from these documents to infer the answer. In this paper, we propose a model based on the Transformer architecture that is able to efficiently operate over a larger set of candidate documents by effectively combining the evidence from these documents during multiple steps of reasoning, while it is robust against noise from low-ranked non-relevant documents included in the set. We use our proposed model, called TraCRNet, on two public open-domain question answering datasets, SearchQA and Quasar-T, and achieve results that meet or exceed the state-of-the-art.
Mostafa Dehghani 0001, Hosein Azarbonyad, Jaap Kamps, Maarten de Rijke
WSDM2
2019 HiTR: Hierarchical Topic Model Re-Estimation for Measuring Topical Diversity of Documents
Hosein Azarbonyad, Mostafa Dehghani 0001, Tom Kenter, Maarten Marx, Jaap Kamps, Maarten de Rijke
IEEE Trans. Knowl. Data Eng.1
2017 Telling How to Narrow it Down: Browsing Path Recommendation for Exploratory Search
abstract
Supporting exploratory search tasks with the help of structured data is an effective way to go beyond keyword search, as it provides an overview of the data, enables users to zoom in on their intent, and provides assistance during their navigation trails. However, finding a good starting point for a search episode in the given structure can still pose a considerable challenge, as users tend to be unfamiliar with exact, complex hierarchical structure. Thus, providing lookahead clues can be of great help and allow users to make better decisions on their search trajectory.
Mostafa Dehghani 0001, Glorianna Jagfeld, Hosein Azarbonyad, Alex Olieman, Jaap Kamps, Maarten Marx
CHIIR3
2017 Words are Malleable: Computing Semantic Shifts in Political and Media Discourse
abstract
Recently, researchers started to pay attention to the detection of temporal shifts in the meaning of words. However, most (if not all) of these approaches restricted their efforts to uncovering change over time, thus neglecting other valuable dimensions such as social or political variability. We propose an approach for detecting semantic shifts between different viewpoints---broadly defined as a set of texts that share a specific metadata feature, which can be a time-period, but also a social entity such as a political party. For each viewpoint, we learn a semantic space in which each word is represented as a low dimensional neural embedded vector. The challenge is to compare the meaning of a word in one space to its meaning in another space and measure the size of the semantic shifts. We compare the effectiveness of a measure based on optimal transformations between the two spaces with a measure based on the similarity of the neighbors of the word in the respective spaces. Our experiments demonstrate that the combination of these two performs best. We show that the semantic shifts not only occur over time but also along different viewpoints in a short period of time. For evaluation, we demonstrate how this approach captures meaningful semantic shifts and can help improve other tasks such as the contrastive viewpoint summarization and ideology detection (measured as classification accuracy) in political texts. We also show that the two laws of semantic change which were empirically shown to hold for temporal shifts also hold for shifts across viewpoints. These laws state that frequent words are less likely to shift meaning while words with many senses are more likely to do so.
Hosein Azarbonyad, Mostafa Dehghani 0001, Kaspar Beelen, Alexandra Arkut, Maarten Marx, Jaap Kamps
CIKM1
2017 Hierarchical Re-estimation of Topic Models for Measuring Topical Diversity
Hosein Azarbonyad, Mostafa Dehghani 0001, Tom Kenter, Maarten Marx, Jaap Kamps, Maarten de Rijke
ECIR1
2016 Generalized Group Profiling for Content Customization
abstract
There is an ongoing debate on personalization, adapting results to the unique user exploiting a user's personal history, versus customization, adapting results to a group profile sharing one or more characteristics with the user at hand. Personal profiles are often sparse, due to cold start problems and the fact that users typically search for new items or information, necessitating to back-off to customization, but group profiles often suffer from accidental features brought in by the unique individual contributing to the group. In this paper we propose a generalized group profiling approach that teases apart the exact contribution of the individual user level and the `abstract' group level by extracting a latent model that captures all, and only, the essential features of the whole group. Our main findings are the followings.
Mostafa Dehghani 0001, Hosein Azarbonyad, Jaap Kamps, Maarten Marx
CHIIR2
2016 Luhn Revisited: Significant Words Language Models
abstract
Users tend to articulate their complex information needs in only a few keywords, making underspecified statements of request the main bottleneck for retrieval effectiveness. Taking advantage of feedback information is one of the best ways to enrich the query representation, but can also lead to loss of query focus and harm performance in particular when the initial query retrieves only little relevant information when overfitting to accidental features of the particular observed feedback documents. Inspired by the early work of Luhn [23], we propose significant words language models of feedback documents that capture all, and only, the significant shared terms from feedback documents. We adjust the weights of common terms that are already well explained by the document collection as well as the weight of rare terms that are only explained by specific feedback documents, which eventually results in having only the significant terms left in the feedback model.
Mostafa Dehghani 0001, Hosein Azarbonyad, Jaap Kamps, Djoerd Hiemstra, Maarten Marx
CIKM2
2016 Measuring Interestingness of Political Documents
abstract
Political texts are pervasive on the Web covering laws and policies in national and supranational jurisdictions. Access to this data is crucial for government transparency and accountability to the population. The main aim of our research is developing a ranking method for political documents which captures the interesting content within political documents. Text interestingness is a measure of assessing the quality of documents from users' perspective which shows their willingness to read a document. Different approaches are proposed for measuring the interestingness of texts. In this research we focus on measuring political texts' interestingness. As political data sources, we use publicly available parliamentary proceedings.
Hosein Azarbonyad
SIGIR1
2015 Sources of Evidence for Automatic Indexing of Political Texts
Mostafa Dehghani 0001, Hosein Azarbonyad, Maarten Marx, Jaap Kamps
ECIR2
2015 Time-Aware Authorship Attribution for Short Text Streams
abstract
Identifying authors of short texts on Internet or social media based communication systems is an important tool against fraud and cybercrimes. Besides the challenges raised by the limited length of these short messages, evolving language and writing styles of authors of these texts makes authorship attribution difficult. Most current short text authorship attribution approaches only address the challenge of limited text length. However, neglecting the second challenge may lead to poor performance of authorship attribution for authors who change their writing styles.
Hosein Azarbonyad, Mostafa Dehghani 0001, Maarten Marx, Jaap Kamps
SIGIR1