Zubair Afzal

dblp:06/8909 · DBLP profile ↗
← Back
6ranked-venue papers in the field
1as first author
5since 2021 · last 2025
0000-0001-7812-2399ORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 5Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
YearPublicationVenuePosition
2025 Question-Answer Extraction from Scientific Articles Using Knowledge Graphs and Large Language Models
abstract
When deciding to read an article or incorporate it into their research, scholars often seek to quickly identify and understand its main ideas.In this paper, we aim to extract these key concepts and contributions from scientific articles in the form of Question and Answer (QA) pairs.We propose two distinct approaches for generating QAs.The first approach involves selecting salient paragraphs, using a Large Language Model (LLM) to generate questions, ranking these questions by the likelihood of obtaining meaningful answers, and subsequently generating answers.This method relies exclusively on the content of the articles.However, assessing an article's novelty typically requires comparison with the existing literature.Therefore, our second approach leverages a Knowledge Graph (KG) for QA generation.We construct a KG by fine-tuning an Entity Relationship (ER) extraction model on scientific articles and using it to build the graph.We then employ a salient triplet extraction method to select the most pertinent ERs per article, utilizing metrics such as the centrality of entities based on a triplet TF-IDF-like measure.This measure assesses the saliency of a triplet based on its importance within the article compared to its prevalence in the literature.For evaluation, we generate QAs using both approaches and have them assessed by Subject Matter Experts (SMEs) through a set of predefined metrics to evaluate the quality of both questions and answers.Our evaluations demonstrate that the KG-based approach effectively captures the main ideas discussed in the articles.Furthermore, our findings indicate that fine-tuning the ER extraction model on our scientific corpus is crucial for extracting high-quality triplets from such documents.
Hosein Azarbonyad, Zi Long Zhu, Georgios Cheirmpos, Zubair Afzal, Vikrant Yadav, George Tsatsaronis 0001
SIGIR4
2024 ScienceDirect Topic Pages: A Knowledge Base of Scientific Concepts Across Various Science Domains
Artemis Çapari, Hosein Azarbonyad, George Tsatsaronis 0001, Zubair Afzal, Judson Dunham
SIGIR4
2023 Generating Topic Pages for Scientific Concepts Using Scientific Publications
Hosein Azarbonyad, Zubair Afzal, George Tsatsaronis 0001
ECIR (2)2
2022 The ChEMU 2022 Evaluation Campaign: Information Extraction in Chemical Patents
Yuan Li 0012, Biaoyan Fang, Estrid He, Hiyori Yoshikawa, Saber A. Akhondi, Christian Druckenbrodt, Camilo Thorne, Zenan Zhai, Zubair Afzal, Trevor Cohn, Timothy Baldwin, Karin Verspoor
ECIR (2)9
2021 ChEMU 2021: Reaction Reference Resolution and Anaphora Resolution in Chemical Patents
Estrid He, Biaoyan Fang, Hiyori Yoshikawa, Yuan Li 0012, Saber A. Akhondi, Christian Druckenbrodt, Camilo Thorne, Zubair Afzal, Zenan Zhai, Lawrence Cavedon, Trevor Cohn, Timothy Baldwin, Karin Verspoor
ECIR (2)8
2016 Learning Domain Labels Using Conceptual Fingerprints: An In-Use Case Study in the Neurology Domain
Zubair Afzal, George Tsatsaronis 0001, Marius A. Doornenbal, Pascal Coupet, Michelle Gregory
EKAW1