Daniel Garijo

dblp:36/10486 · DBLP profile ↗
← Back
23ranked-venue papers in the field
4as first author
8since 2021 · last 2025
0000-0003-0454-7145ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 14 (3 first)Other / Interdisciplinary · 6 (1 first)Information Retrieval & Web Search · 2Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2025 Good practice versus reality: A landscape analysis of Research Software metadata adoption in European Open Science Clusters
abstract
Research Software has become a key asset to support the results described in academic publications, enabling effective data analysis and reproducibility. In order to ensure adherence of Research Software to the Findable, Accessible, Interoperable, and Reusable (FAIR) principles, the scientific community has proposed metadata guidelines and best practices. However, it is unclear how these practices have been adopted so far. This paper examines how different scientific communities describe Research Software with metadata to support FAIR, how do they adopt existing good practices regarding citation, documentation or versioning, and what is the current adoption of archival services for long-term preservation. We carry out our analysis in the software registries of five science clusters (in domains ranging from Physics to Environmental Sciences), together with a multi-domain collaborative software registry. Our results highlight the main gaps in metadata adoption in the different communities, opening an opportunity for future contributions to aid researchers in adopting good FAIR and Open Science practices.
Anas El Hounsri, Daniel Garijo
MSR2
2024 Bidirectional Paper-Repository Tracing in Software Engineering
abstract
While computer science papers frequently include their associated code repositories, establishing a clear link between papers and their corresponding implementations may be challenging due to the number of code repositories used in research publications. In this paper we describe a lightweight method for effectively identifying bidirectional links between papers and repositories from both LaTeX and PDF sources. We have used our approach to analyze more than 14000 PDF and Latex files in the Software Engineering category of Arxiv, generating a dataset of more than 1400 paper-code implementations and assessing current citation practices on it.
Daniel Garijo, Miguel Arroyo, Esteban González, Christoph Treude, Nicola Tarocco
MSR1
2023 TEC: Transparent Emissions Calculation Toolkit
abstract
Greenhouse gas emissions have become a common means for determining the carbon footprint of any commercial activity, ranging from booking a trip or manufacturing a product to training a machine learning model. However, calculating the amount of emissions associated with these activities can be a difficult task, involving estimations of energy used and considerations of location and time period. In this paper, we introduce the Transparent Emissions Calculation (TEC) toolkit, an open source effort aimed at addressing this challenge. Our contributions include two ontologies (ECFO and PECO) that represent emissions conversion factors and the provenance traces of carbon emissions calculations (respectively), a public knowledge graph with thousands of conversion factors (with their corresponding YARRRML and RML mappings) and a prototype carbon emissions calculator which uses our knowledge graph to produce a transparent emissions report. Resource permanent URL: https://w3id.org/tec-toolkit .
Milan Markovic, Daniel Garijo, Stefano Germano, Iman Naja
ISWC2
2022 Extending Ontology Engineering Practices to Facilitate Application Development
Paola Espinoza-Arias, Daniel Garijo, Óscar Corcho
EKAW2
2022 FAIROs: Towards FAIR Assessment in Research Objects
Esteban González, Alejandro Benítez, Daniel Garijo
TPDL3
2022 Inspect4py: A Knowledge Extraction Framework for Python Code Repositories
abstract
This work presents inspect4py, a static code analysis framework designed to automatically extract the main features, metadata and documentation of Python code repositories. Given an input folder with code, inspect4py uses abstract syntax trees and state of the art tools to find all functions, classes, tests, documentation, call graphs, module dependencies and control flows within all code files in that repository. Using these findings, inspect4py infers different ways of invoking a software component. We have evaluated our framework on 95 annotated repositories, obtaining promising results for software type classification (over 95% F1-score). With inspect4py, we aim to ease the understandability and adoption of software repositories by other researchers and developers.
Rosa Filgueira, Daniel Garijo
MSR2
2022 A study of the quality of Wikidata
Kartik Shenoy, Filip Ilievski, Daniel Garijo, Daniel Schwabe 0001, Pedro A. Szekely
J. Web Semant.3
2021 Crossing the chasm between ontology engineering and application development: A survey
abstract
The adoption of Knowledge Graphs (KGs) by public and private organizations to integrate and publish data has increased in recent years. Ontologies play a crucial role in providing the structure for KGs, but are usually disregarded when designing Application Programming Interfaces (APIs) to enable browsing KGs in a developer-friendly manner. In this paper we provide a systematic review of the state of the art on existing approaches to ease access to ontology-based KG data by application developers. We propose two comparison frameworks to understand specifications, technologies and tools responsible for providing APIs for KGs. Our results reveal several limitations on existing API-based specifications, technologies and tools for KG consumption, which outline exciting research challenges including automatic API generation, API resource path prediction, ontology-based API versioning, and API validation and testing.
Paola Espinoza-Arias, Daniel Garijo, Óscar Corcho
J. Web Semant.2
2020 Coming to Terms with FAIR Ontologies
María Poveda-Villalón, Paola Espinoza-Arias, Daniel Garijo, Óscar Corcho
EKAW3
2020 OBA: An Ontology-Based Framework for Creating REST APIs for Knowledge Graphs
Daniel Garijo, Maximiliano Osorio
ISWC (2)1
2020 KGTK: A Toolkit for Large Knowledge Graph Manipulation and Analysis
Filip Ilievski, Daniel Garijo, Hans Chalupsky, Naren Teja Divvala, Yixiang Yao, Craig Milo Rogers, Ronpeng Li, Daniel Schwabe 0001, Pedro A. Szekely
ISWC (2)2
2019 SoMEF: A Framework for Capturing Scientific Software Metadata from its Documentation
abstract
Scientific software has become a key asset to reproduce and understand the products of scientific research in many disciplines. However, scientific software is becoming increasingly complex and, as a result, researchers need to spend a significant amount of time finding, reading and understanding software documentation to set it up. In this paper we describe SoMEF, a Software Metadata Extraction Framework designed to help highlighting the most important parts of scientific software documentation. SoMEF processes the README files in GitHub repositories to automatically extract which parts of their text refer to the description, installation, invocation, or citation of a software component. Despite its simple features, SoMEF successfully categorizes README excerpts with a minimum 0.92 precision and 0.90 ROC AUC. These results, tested on a corpus of over 70 scientific software repositories, are a promising start towards automatically generating knowledge graphs of scientific software metadata.
Allen Mao, Daniel Garijo, Shobeir Fakhraei
IEEE BigData2
2019 T2WML: Table To Wikidata Mapping Language
abstract
The web contains millions of useful spreadsheets and CSV files, but these files are difficult to use in applications because they use a wide variety of data layouts and terminology. We present Table To Wikidata Mapping Language (T2WML), a language that makes it easy to map and link arbitrary spreadsheets and CSV files to the Wikidata data model. The output of T2WML consists of Wikidata statements that can be loaded in the public Wikidata knowledge base or in a Wikidata clone repository, creating an augmented Wikidata knowledge graph that application developers can query using SPARQL.
Pedro A. Szekely, Daniel Garijo, Divij Bhatia, Yixiang Yao, Jay Pujara
K-CAP2
2019 Automating ontology engineering support activities with OnToology
Ahmad Alobaid, Daniel Garijo, María Poveda-Villalón, Idafen Santana-Pérez, Alba Fernández-Izquierdo, Óscar Corcho
J. Web Semant.2
2018 PSM-Flow: Probabilistic Subgraph Mining for Discovering Reusable Fragments in Workflows
abstract
Scientific workflows define computational processes needed for carrying out scientific experiments. Existing workflow repositories contain hundreds of scientific workflows, where scientists can find materials and knowledge to facilitate workflow design for running related experiments. Identifying reusable fragments in growing workflow repositories has become increasingly important. In this paper, we present PSM-Flow, a probabilistic subgraph mining algorithm designed to discover commonly occurring fragments in a workflow corpus using a modified version of the Latent Dirichlet Allocation algorithm. The proposed model encodes the geodesic distance between workflow steps into the model for implicitly modeling fragments. PSM-Flow captures variations of frequent fragments while maintaining its space complexity bounded polynomially, as it requires no candidate generation. We applied PSM-Flow to three real-world scientific workflow datasets containing more than 750 workflows for neuroimaging analysis. Our results show that PSM-Flow outperforms three state of the art frequent subgraph mining techniques. We also discuss other potential future improvements of the proposed method.
Chin Wang Cheong, Daniel Garijo, William Kwok-Wai Cheung, Yolanda Gil
WI2
2017 WIDOCO: A Wizard for Documenting Ontologies
Daniel Garijo
ISWC (2)1
2017 A Controlled Crowdsourcing Approach for Practical Ontology Extensions and Metadata Annotations
Yolanda Gil, Daniel Garijo, Varun Ratnakar, Deborah Khider, Julien Emile-Geay, Nicholas McKay
ISWC (2)2
2015 OntoSoft: Capturing Scientific Software Metadata
abstract
This paper presents OntoSoft, an ontology to describe metadata for scientific software. The ontology is designed considering how scientists would approach the reuse and sharing of software. This includes supporting a scientist to: 1) identify software, 2) understand and assess software, 3) execute software, 4) get support for the software, 5) do research with the software, and 6) update the software. The ontology is available in OWL and contains more than fifty terms. We are using OntoSoft to structure a software registry for geosciences, and to develop user interfaces to capture its metadata.
Yolanda Gil, Varun Ratnakar, Daniel Garijo
K-CAP3
2015 Using a suite of ontologies for preserving workflow-centric research objects
abstract
Scientific workflows are a popular mechanism for specifying and automating data-driven in silico experiments. A significant aspect of their value lies in their potential to be reused. Once shared, workflows become useful building blocks that can be combined or modified for developing new experiments. However, previous studies have shown that storing workflow specifications alone is not sufficient to ensure that they can be successfully reused, without being able to understand what the workflows aim to achieve or to re-enact them. To gain an understanding of the workflow, and how it may be used and repurposed for their needs, scientists require access to additional resources such as annotations describing the workflow, datasets used and produced by the workflow, and provenance traces recording workflow executions. In this article, we present a novel approach to the preservation of scientific workflows through the application of research objects—aggregations of data and metadata that enrich the workflow specifications. Our approach is realised as a suite of ontologies that support the creation of workflow-centric research objects. Their design was guided by requirements elicited from previous empirical analyses of workflow decay and repair. The ontologies developed make use of and extend existing well known ontologies, namely the Object Reuse and Exchange (ORE) vocabulary, the Annotation Ontology (AO) and the W3C PROV ontology (PROVO). We illustrate the application of the ontologies for building Workflow Research Objects with a case-study that investigates Huntington’s disease, performed in collaboration with a team from the Leiden University Medial Centre (HG-LUMC). Finally we present a number of tools developed for creating and managing workflow-centric research objects.
Khalid Belhajjame, Jun Zhao 0003, Daniel Garijo, Matthew Gamble, Kristina M. Hettne, Raúl Palma, Eleni Mina, Óscar Corcho, José Manuél Gómez-Pérez, Sean Bechhofer, Graham Klyne, Carole A. Goble
J. Web Semant.3
2013 From Preserving Data to Preserving Research: Curation of Process and Context
Rudolf Mayer, Stefan Pröll, Andreas Rauber, Raúl Palma, Daniel Garijo
TPDL5
2013 Detecting common scientific workflow fragments using templates and execution provenance
abstract
Provenance plays a major role when understanding and reusing the methods applied in a scientific experiment, as it provides a record of inputs, the processes carried out and the use and generation of intermediate and final results. In the specific case of in-silico scientific experiments, a large variety of scientific workflow systems (e.g., Wings, Taverna, Galaxy, Vistrails) have been created to support scientists. All of these systems produce some sort of provenance about the executions of the workflows that encode scientific experiments. However, provenance is normally recorded at a very low level of detail, which complicates the understanding of what happened during execution. In this paper we propose an approach to automatically obtain abstractions from low-level provenance data by finding common workflow fragments on workflow execution provenance and relating them to templates. We have tested our approach with a dataset of workflows published by the Wings workflow system. Our results show that by using these kinds of abstractions we can highlight the most common abstract methods used in the executions of a repository, relating different runs and workflow templates with each other.
Daniel Garijo, Óscar Corcho, Yolanda Gil
K-CAP1
2011 Extending DCAM for Metadata Provenance
Kai Eckert 0001, Daniel Garijo, Michael Panzer
Dublin Core Conference2
2011 Metadata Provenance: Dublin Core on the Next Level
Kai Eckert 0001, Daniel Garijo, Michael Panzer, Omer Percin
Dublin Core Conference2