Sergio José Rodríguez Méndez

dblp:156/5803 · also Sergio Jose Rodriguez Mendez · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0001-7203-8399ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
YearPublicationVenuePosition
2024 Doc-KG: Unstructured documents to knowledge graph construction, identification and validation with Wikidata
abstract
Abstract The exponential growth of textual data in the digital era underlines the pivotal role of Knowledge Graphs (KGs) in effectively storing, managing, and utilizing this vast reservoir of information. Despite the copious amounts of text available on the web, a significant portion remains unstructured, presenting a substantial barrier to the automatic construction and enrichment of KGs. To address this issue, we introduce an enhanced Doc‐KG model, a sophisticated approach designed to transform unstructured documents into structured knowledge by generating local KGs and mapping these to a target KG, such as Wikidata. Our model innovatively leverages syntactic information to extract entities and predicates efficiently, integrating them into triples with improved accuracy. Furthermore, the Doc‐KG model's performance surpasses existing methodologies by utilizing advanced algorithms for both the extraction of triples and their subsequent identification within Wikidata, employing Wikidata's Unified Resource Identifiers for precise mapping. This dual capability not only facilitates the construction of KGs directly from unstructured texts but also enhances the process of identifying triple mentions within Wikidata, marking a significant advancement in the domain. Our comprehensive evaluation, conducted using the renowned WebNLG benchmark dataset, reveals the Doc‐KG model's superior performance in triple extraction tasks, achieving an unprecedented accuracy rate of 86.64%. In the domain of triple identification, the model demonstrated exceptional efficacy by mapping 61.35% of the local KG to Wikidata, thereby contributing 38.65% of novel information for KG enrichment. A qualitative analysis based on a manually annotated dataset further confirms the model's excellence, outshining baseline methods in extracting high‐fidelity triples. This research embodies a novel contribution to the field of knowledge extraction and management, offering a robust framework for the semantic structuring of unstructured data and paving the way for the next generation of KGs.
Armin Haller, Sergio José Rodríguez Méndez, Usman Naseem
Expert Syst. J. Knowl. Eng.3
2022 An Analysis of Links in Wikidata
Armin Haller, Axel Polleres, Daniil Dobriy, Nicolas Ferranti, Sergio José Rodríguez Méndez
ESWC5
2022 Active knowledge graph completion
abstract
Enterprise and public Knowledge Graphs (KGs) are known to be incomplete. Methods for automatic completion, sometimes by rule learning, scale well. While previous rule-based methods learn closed (non-existential) rules, we introduce Open Path (OP) rules that are constrained existential rules. We present a novel algorithm, OPRL, for learning OP rules. Closed rules complete a KG by answering queries of unclear origin, usually derived from a holdback test set in experimental settings. However, OP rules can generate relevant queries for KG completion. OPRL generates queries even when there is no closed rule to answer the query, or when the correct answer is a missing entity that is not present in the KG. For OPRL to scale well, we propose a novel embedding-based fitness function to efficiently estimate rule quality. Additionally, we introduce a novel, efficient vector computation to formally assess rule quality. We evaluate OPRL using adaptations of Freebase, YAGO2, Wikidata, and a synthetic Poker KG. We find that OPRL mines hundreds of accurate rules from massive KGs with up to 8 M facts. The OP rules generate queries with precision as high as 98% and recall of 62% on a complete KG, demonstrating the first solution for active knowledge graph completion.
Pouya Ghiasnezhad Omran, Kerry L. Taylor, Sergio José Rodríguez Méndez, Armin Haller
Inf. Sci.3
2021 TNNT: The Named Entity Recognition Toolkit
abstract
Extraction of categorised named entities from text is a complex task given the availability of a variety of Named Entity Recognition (NER) models and the unstructured information encoded in different source document formats. Processing the documents to extract text, identifying suitable NER models for a task, and obtaining statistical information is important in data analysis to make informed decisions. This paper presents\footnoteThe manuscript follows guidelines to showcase a demonstration that introduces an overview of how the toolkit works: input document set, initial settings, processing, and output set. The input document set is artificial in order to show various toolkit capabilities. TNNT, a toolkit that automates the extraction of categorised named entities from unstructured information encoded in source documents, using diverse state-of-the-art (SOTA) Natural Language Processing (NLP) tools and NER models.TNNT integrates 21 different NER models as part of a Knowledge Graph Construction Pipeline (KGCP) that takes a document set as input and processes it based on the defined settings, applying the selected blocks of NER models to output the results. The toolkit generates all results with an integrated summary of the extracted entities, enabling enhanced data analysis to support the KGCP, and also, to aid further NLP tasks.
Sandaru Seneviratne, Sergio José Rodríguez Méndez, Xuecheng Zhang, Pouya Ghiasnezhad Omran, Kerry L. Taylor, Armin Haller
K-CAP2
2020 HDGI: A Human Device Gesture Interaction Ontology for the Internet of Things
Madhawa Perera, Armin Haller, Sergio José Rodríguez Méndez, Matt Adcock
ISWC (2)3
2020 Schímatos: A SHACL-Based Web-Form Generator for Knowledge Graph Editing
Jesse Wright, Sergio José Rodríguez Méndez, Armin Haller, Kerry L. Taylor, Pouya Ghiasnezhad Omran
ISWC (2)2
2014 Augmented Brain Computer Interaction Based on Fog Computing and Linked Data
abstract
An augmented brain computer interface that can detect users' brain states in real-life situations has been developed using wireless EEG headsets, smart phones and ubiquitous computing services. This kind of wearable natural user interfaces will have a wide-range of potential applications in future smart environments. This paper describes its ubiquitous system architecture and introduces its enabling technologies, which include machine-to-machine publish/subscribe protocols, multi-tier fog/cloud computing infrastructure and a linked data web. Its real-time responsiveness and easiness-of-use will be demonstrated by playing a multi-player on-line BCI game EEG Tractor Beam at the Intelligent Environment Conference.
John K. Zao, Tchin Tze Gan, Chun Kai You, Sergio José Rodríguez Méndez, Cheng En Chung, Yu-Te Wang, Tim R. Mullen, Tzyy-Ping Jung
Intelligent Environments4