Mohamad Yaser Jaradeh

dblp:210/9086 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-8777-2780ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Towards Context-Aware Search: Dynamic Facet Generation in Digital Libraries
Mutahira Khalid, Mohamad Yaser Jaradeh, Sören Auer, Markus Stocker
ESWC (2)2
2026 SCI-IDEA: Context-Aware Scientific Ideation Using Token and Sentence Embeddings
abstract
Abstract Generating context-aware, high-quality, and innovative scientific ideas remains a central challenge in AI-supported research. We introduce SCI-IDEA , a two-stage framework combining large language models (LLMs) with a specialised Aha-Moment Detection module for iterative idea refinement. The first stage extracts structured facets: objectives, methodology, evaluation, and future work from previous publications, compressing each paper to $${\sim }$$ 200 tokens to enable scalable processing of researcher profiles that would otherwise exceed frontier LLM context windows. The second stage integrates these facets with a token-level embedding approach to identify novelty and context gaps, indicating candidate ideas as Aha moments when they surpass predefined thresholds for both novelty and surprise. We evaluate SCI-IDEA across 100 computer-science researcher profiles, 4 LLMs (GPT-4o, GPT−4.5, DeepSeek-32B, DeepSeek-70B), three embedding strategies, and 5 prompting configurations via a hybrid protocol of automated LLM-as-judge scoring with GPT−4.1 and a human evaluation with 15 PhD-level domain experts. These two evaluation signals diverge substantially: LLM-based scores exceed expert ratings by 3–4 points on a 10-point scale, with near-zero inter-rater correlation ( $$r = 0.02$$ –0.17, $$p > 0.05$$ ), indicating that automated scores reflect relative architectural comparisons rather than human-verified measures of absolute idea quality. Within the LLM-evaluated setting, SCI-IDEA with token-level embeddings achieves a mean quality score of 6.97, a modest but constant improvement over the LLM-only baseline ( $$\Delta = +0.20$$ points; $$p < 0.05$$ ) and driven primarily by improved feasibility scores ( $$+0.21$$ points, $$p < 0.001$$ ). Expert evaluators independently confirmed this trend while assigning lower absolute scores. These findings suggest an architectural advantage of facet-based context modelling and token-level novelty detection, though their practical importance deserves validation through longitudinal studies with larger domain experts. This work is validated in the computer science domain, but its application to other fields requires further investigation.
Farhana Keya, Gollam Rabby, Sören Auer, Sahar Vahdati, Prasenjit Mitra 0001, Mohamad Yaser Jaradeh
Mach. Learn.6
2025 Neuro-Symbolic Federated Research Artifact Search
Farhana Keya, Sören Auer, Mohamad Yaser Jaradeh
TPDL3
2025 Introducing ORKG ASK: An AI-Driven Scholarly Literature Search and Exploration System Taking a Neuro-Symbolic Approach
Allard Oelen, Mohamad Yaser Jaradeh, Sören Auer
ICWE2
2023 Information extraction pipelines for knowledge graphs
abstract
In the last decade, a large number of knowledge graph (KG) completion approaches were proposed. Albeit effective, these efforts are disjoint, and their collective strengths and weaknesses in effective KG completion have not been studied in the literature. We extend Plumber, a framework that brings together the research community's disjoint efforts on KG completion. We include more components into the architecture of Plumber to comprise 40 reusable components for various KG completion subtasks, such as coreference resolution, entity linking, and relation extraction. Using these components, Plumber dynamically generates suitable knowledge extraction pipelines and offers overall 432 distinct pipelines. We study the optimization problem of choosing optimal pipelines based on input sentences. To do so, we train a transformer-based classification model that extracts contextual embeddings from the input and finds an appropriate pipeline. We study the efficacy of Plumber for extracting the KG triples using standard datasets over three KGs: DBpedia, Wikidata, and Open Research Knowledge Graph. Our results demonstrate the effectiveness of Plumber in dynamically generating KG completion pipelines, outperforming all baselines agnostic of the underlying KG. Furthermore, we provide an analysis of collective failure cases, study the similarities and synergies among integrated components and discuss their limitations.
Mohamad Yaser Jaradeh, Kuldeep Singh 0001, Markus Stocker, Andreas Both 0001, Sören Auer
Knowl. Inf. Syst.1
2021 Better Call the Plumber: Orchestrating Dynamic Information Extraction Pipelines
Mohamad Yaser Jaradeh, Kuldeep Singh 0001, Markus Stocker, Andreas Both 0001, Sören Auer
ICWE1
2021 Triple Classification for Scholarly Knowledge Graph Completion
abstract
structured information representing knowledge encoded in scientific publications. With the sheer volume of published scientific literature comprising a plethora of inhomogeneous entities and relations to describe scientific concepts, these KGs are inherently incomplete. We present exBERT, a method for leveraging pre-trained transformer language models to perform scholarly knowledge graph completion. We model triples of a knowledge graph as text and perform triple classification (i.e., belongs to KG or not). The evaluation shows that exBERT outperforms other baselines on three scholarly KG completion datasets in the tasks of triple classification, link prediction, and relation prediction. Furthermore, we present two scholarly datasets as resources for the research community, collected from public KGs and online resources.
Mohamad Yaser Jaradeh, Kuldeep Singh 0001, Markus Stocker, Sören Auer
K-CAP1
2020 Challenges of Linking Organizational Information in Open Government Data to Knowledge Graphs
Jan Portisch, Omaima Fallatah, Sebastian Neumaier, Mohamad Yaser Jaradeh, Axel Polleres
EKAW4
2020 Question Answering on Scholarly Knowledge Graphs
Mohamad Yaser Jaradeh, Markus Stocker, Sören Auer
TPDL1
2020 The STEM-ECR Dataset: Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources
abstract
We introduce the STEM (Science, Technology, Engineering, and Medicine) Dataset for Scientific Entity Extraction, Classification, and Resolution, version 1.0 (STEM-ECR v1.0). The STEM-ECR v1.0 dataset has been developed to provide a benchmark for the evaluation of scientific entity extraction, classification, and resolution tasks in a domain-independent fashion. It comprises abstracts in 10 STEM disciplines that were found to be the most prolific ones on a major publishing platform. We describe the creation of such a multidisciplinary corpus and highlight the obtained findings in terms of the following features: 1) a generic conceptual formalism for scientific entities in a multidisciplinary scientific context; 2) the feasibility of the domain-independent human annotation of scientific entities under such a generic formalism; 3) a performance benchmark obtainable for automatic extraction of multidisciplinary scientific entities using BERT-based neural models; 4) a delineated 3-step entity resolution procedure for human annotation of the scientific entities via encyclopedic entity linking and lexicographic word sense disambiguation; and 5) human evaluations of Babelfy returned encyclopedic links and lexicographic senses for our entities. Our findings cumulatively indicate that human annotation and automatic learning of multidisciplinary scientific concepts as well as their semantic disambiguation in a wide-ranging setting as STEM is reasonable.
Jennifer D'Souza 0001, Anett Hoppe, Arthur Brack, Mohamad Yaser Jaradeh, Sören Auer, Ralph Ewerth
LREC4
2019 Open Research Knowledge Graph: A System Walkthrough
Mohamad Yaser Jaradeh, Allard Oelen, Manuel Prinz, Markus Stocker, Sören Auer
TPDL1
2019 Open Research Knowledge Graph: Next Generation Infrastructure for Semantic Scholarly Knowledge
abstract
Despite improved digital access to scholarly knowledge in recent decades, scholarly communication remains exclusively document-based. In this form, scholarly knowledge is hard to process automatically. We present the first steps towards a knowledge graph based infrastructure that acquires scholarly knowledge in machine actionable form thus enabling new possibilities for scholarly knowledge curation, publication and processing. The primary contribution is to present, evaluate and discuss multi-modal scholarly knowledge acquisition, combining crowdsourced and automated techniques. We present the results of the first user evaluation of the infrastructure with the participants of a recent international conference. Results suggest that users were intrigued by the novelty of the proposed infrastructure and by the possibilities for innovative scholarly knowledge processing it could enable.
Mohamad Yaser Jaradeh, Allard Oelen, Kheir Eddine Farfar, Manuel Prinz, Jennifer D'Souza 0001, Gábor Kismihók, Markus Stocker, Sören Auer
K-CAP1
2017 Capturing Knowledge in Semantically-typed Relational Patterns to Enhance Relation Linking
abstract
Transforming natural language questions into formal queries is an integral task in Question Answering (QA) systems. QA systems built on knowledge graphs like DBpedia, require a step after natural language processing for linking words, specifically including named entities and relations, to their corresponding entities in a knowledge graph. To achieve this task, several approaches rely on background knowledge bases containing semantically-typed relations, e.g., PATTY, for an extra disambiguation step. Two major factors may affect the performance of relation linking approaches whenever background knowledge bases are accessed: a) limited availability of such semantic knowledge sources, and b) lack of a systematic approach on how to maximize the benefits of the collected knowledge. We tackle this problem and devise SIBKB, a semantic-based index able to capture knowledge encoded on background knowledge bases like PATTY. SIBKB represents a background knowledge base as a bi-partite and a dynamic index over the relation patterns included in the knowledge base. Moreover, we develop a relation linking component able to exploit SIBKB features. The benefits of SIBKB are empirically studied on existing QA benchmarks and observed results suggest that SIBKB is able to enhance the accuracy of relation linking by up to three times.
Kuldeep Singh 0001, Isaiah Onando Mulang', Ioanna Lytra, Mohamad Yaser Jaradeh, Ahmad Sakor, Maria-Esther Vidal, Christoph Lange 0002, Sören Auer
K-CAP4