José Manuél Gómez-Pérez

dblp:74/9922 · DBLP profile ↗
← Back
34ranked-venue papers
8as first author
7since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 16 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 8 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 2 first-authorSystems, architecture and hardware · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 SPACE-IDEAS: A Dataset for Salient Information Detection in Space Innovation
abstract
Detecting salient parts in text using natural language processing has been widely used to mitigate the effects of information overflow. Nevertheless, most of the datasets available for this task are derived mainly from academic publications. We introduce SPACE-IDEAS, a dataset for salient information detection from innovation ideas related to the Space domain. The text in SPACE-IDEAS varies greatly and includes informal, technical, academic and business-oriented writing styles. In addition to a manually annotated dataset we release an extended version that is annotated using a large generative language model. We train different sentence and sequential sentence classifiers, and show that the automatically annotated dataset can be leveraged using multitask learning to train better classifiers.
Andrés García-Silva, Cristian Berrio, José Manuél Gómez-Pérez
LREC/COLING3
2023 Capturing Pertinent Symbolic Features for Enhanced Content-Based Misinformation Detection
abstract
Preventing the spread of misinformation is challenging. The detection of misleading content presents a significant hurdle due to its extreme linguistic and domain variability. Content-based models have managed to identify deceptive language by learning representations from textual data such as social media posts and web articles. However, aggregating representative samples of this heterogeneous phenomenon and implementing effective real-world applications is still elusive. Based on analytical work on the language of misinformation, this paper analyzes the linguistic attributes that characterize this phenomenon and how representative of such features some of the most popular misinformation datasets are. We demonstrate that the appropriate use of pertinent symbolic knowledge in combination with neural language models is helpful in detecting misleading content. Our results achieve state-of-the-art performance in misinformation datasets across the board, showing that our approach offers a valid and robust alternative to multi-task transfer learning without requiring any additional training data. Furthermore, our results show evidence that structured knowledge can provide the extra boost required to address a complex and unpredictable real-world problem like misinformation detection, not only in terms of accuracy but also time efficiency and resource utilization.
Flavio Merenda, José Manuél Gómez-Pérez
K-CAP2
2023 Textual Entailment for Effective Triple Validation in Object Prediction
Andrés García-Silva, Cristian Berrio, José Manuél Gómez-Pérez
ISWC3
2022 SpaceQA: Answering Questions about the Design of Space Missions and Space Craft Concepts
abstract
We present SpaceQA, to the best of our knowledge the first open-domain QA system in Space mission design. SpaceQA is part of an initiative by the European Space Agency (ESA) to facilitate the access, sharing and reuse of information about Space mission design within the agency and with the public. We adopt a state-of-the-art architecture consisting of a dense retriever and a neural reader and opt for an approach based on transfer learning rather than fine-tuning due to the lack of domain-specific annotated data. Our evaluation on a test set produced by ESA is largely consistent with the results originally reported by the evaluated retrievers and confirms the need of fine tuning for reading comprehension. As of writing this paper, ESA is piloting SpaceQA internally.
Andrés García-Silva, Cristian Berrio, José Manuél Gómez-Pérez, José Antonio Martínez Heras, Alessandro Donati, Ilaria Roma
SIGIR3
2021 Classifying Scientific Publications with BERT - Is Self-attention a Feature Selection Method?
Andrés García-Silva, José Manuél Gómez-Pérez
ECIR (1)2
2021 Weaving a Semantic Web of Credibility Reviews for Explainable Misinformation Detection (Extended Abstract)
abstract
This paper summarises work where we combined semantic web technologies with deep learning systems to obtain state-of-the art explainable misinformation detection. We proposed a conceptual and computational model to describe a wide range of misinformation detection systems based around the concepts of credibility and reviews. We described how Credibility Reviews (CRs) can be used to build networks of distributed bots that collaborate for misinformation detection which we evaluated by building a prototype based on publicly available datasets and deep learning models.
Ronald Denaux, Martino Mensio, José Manuél Gómez-Pérez, Harith Alani
IJCAI3
2021 On the impact of knowledge-based linguistic annotations in the quality of scientific embeddings
Andrés García-Silva, Ronald Denaux, José Manuél Gómez-Pérez
Future Gener. Comput. Syst.3
2020 ISAAQ - Mastering Textbook Questions with Pre-trained Transformers and Bottom-Up and Top-Down Attention
abstract
Textbook Question Answering is a complex task in the intersection of Machine Comprehension and Visual Question Answering that requires reasoning with multimodal information from text and diagrams.For the first time, this paper taps on the potential of transformer language models and bottom-up and top-down attention to tackle the language and visual understanding challenges this task entails.Rather than training a language-visual transformer from scratch we rely on pretrained transformers, fine-tuning and ensembling.We add bottom-up and top-down attention to identify regions of interest corresponding to diagram constituents and their relationships, improving the selection of relevant visual information for each question and answer options.Our system ISAAQ reports unprecedented success in all TQA question types, with accuracies of 81.36%, 71.11% and 55.12% on true/false, text-only and diagram multiple choice questions.ISAAQ also demonstrates its broad applicability, obtaining state-of-theart results in other demanding datasets.
José Manuél Gómez-Pérez, Raúl Ortega 0001
EMNLP (1)1
2020 Making Metadata Fit for Next Generation Language Technology Platforms: The Metadata Schema of the European Language Grid
abstract
The current scientific and technological landscape is characterised by the increasing availability of data resources and processing tools and services. In this setting, metadata have emerged as a key factor facilitating management, sharing and usage of such digital assets. In this paper we present ELG-SHARE, a rich metadata schema catering for the description of Language Resources and Technologies (processing and generation services and tools, models, corpora, term lists, etc.), as well as related entities (e.g., organizations, projects, supporting documents, etc.). The schema powers the European Language Grid platform that aims to be the primary hub and marketplace for industry-relevant Language Technology in Europe. ELG-SHARE has been based on various metadata schemas, vocabularies, and ontologies, as well as related recommendations and guidelines.
Penny Labropoulou, Katerina Gkirtzou, Maria Gavrilidou, Miltos Deligiannis, Dimitrios Galanis, Stelios Piperidis, Georg Rehm, Maria Berger, Valérie Mapelli, Mickaël Rigault, Victoria Arranz, Khalid Choukri, Gerhard Backfried, José Manuél Gómez-Pérez, Andrés García-Silva
LREC14
2020 European Language Grid: An Overview
abstract
With 24 official EU and many additional languages, multilingualism in Europe and an inclusive Digital Single Market can only be enabled through Language Technologies (LTs). European LT business is dominated by hundreds of SMEs and a few large players. Many are world-class, with technologies that outperform the global players. However, European LT business is also fragmented – by nation states, languages, verticals and sectors, significantly holding back its impact. The European Language Grid (ELG) project addresses this fragmentation by establishing the ELG as the primary platform for LT in Europe. The ELG is a scalable cloud platform, providing, in an easy-to-integrate way, access to hundreds of commercial and non-commercial LTs for all European languages, including running tools and services as well as data sets and resources. Once fully operational, it will enable the commercial and non-commercial European LT community to deposit and upload their technologies and data sets into the ELG, to deploy them through the grid, and to connect with other resources. The ELG will boost the Multilingual Digital Single Market towards a thriving European LT community, creating new jobs and opportunities. Furthermore, the ELG project organises two open calls for up to 20 pilot projects. It also sets up 32 national competence centres and the European LT Council for outreach and coordination purposes.
Georg Rehm, Maria Berger, Ela Elsholz, Stefanie Hegele, Florian Kintzel, Katrin Marheinecke, Stelios Piperidis, Miltos Deligiannis, Dimitrios Galanis, Katerina Gkirtzou, Penny Labropoulou, Kalina Bontcheva, Jan Hajic 0001, Jana Hamrlová, Lukás Kacena, Khalid Choukri, Victoria Arranz, Andrejs Vasiljevs, Orians Anvari, Andis Lagzdins, Julija Melnika, Gerhard Backfried, Erinç Dikici, Miroslav Jánosík, Katja Prinz, Christoph Prinz, Severin Stampler, Dorothea Thomas-Aniola, José Manuél Gómez-Pérez, Andrés García-Silva, Cristian Berrio, Ulrich Germann, Steve Renals, Ondrej Klejch
LREC31
2020 The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual Europe
abstract
Multilingualism is a cultural cornerstone of Europe and firmly anchored in the European treaties including full language equality. However, language barriers impacting business, cross-lingual and cross-cultural communication are still omnipresent. Language Technologies (LTs) are a powerful means to break down these barriers. While the last decade has seen various initiatives that created a multitude of approaches and technologies tailored to Europe’s specific needs, there is still an immense level of fragmentation. At the same time, AI has become an increasingly important concept in the European Information and Communication Technology area. For a few years now, AI – including many opportunities, synergies but also misconceptions – has been overshadowing every other topic. We present an overview of the European LT landscape, describing funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. We present a brief overview of the main LT-related activities on the EU level in the last ten years and develop strategic guidance with regard to four key dimensions.
Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajic 0001, Khalid Choukri, Andrejs Vasiljevs, Gerhard Backfried, Christoph Prinz, José Manuél Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriute, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavrilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Marko Tadic, Dan Tufis, Tamás Váradi, Kadri Vider, Andy Way, François Yvon
LREC11
2020 Linked Credibility Reviews for Explainable Misinformation Detection
Ronald Denaux, José Manuél Gómez-Pérez
ISWC (1)2
2019 Assessing the Lexico-Semantic Relational Knowledge Captured by Word and Concept Embeddings
abstract
Deep learning currently dominates the benchmarks for various NLP tasks and, at the basis of such systems, words are frequently represented as embeddings --vectors in a low dimensional space-- learned from large text corpora and various algorithms have been proposed to learn both word and concept embeddings. One of the claimed benefits of such embeddings is that they capture knowledge about semantic relations. Such embeddings are most often evaluated through tasks such as predicting human-rated similarity and analogy which only test a few, often ill-defined, relations. In this paper, we propose a method for (i) reliably generating word and concept pair datasets for a wide number of relations by using a knowledge graph and (ii) evaluating to what extent pre-trained embeddings capture those relations. We evaluate the approach against a proprietary and a public knowledge graph and analyze the results, showing which lexico-semantic relational knowledge is captured by current embedding learning approaches.
Ronald Denaux, José Manuél Gómez-Pérez
K-CAP2
2019 Look, Read and Enrich - Learning from Scientific Figures and their Captions
abstract
Compared to natural images, understanding scientific figures is particularly hard for machines. However, there is a valuable source of information in scientific literature that until now has remained untapped: the correspondence between a figure and its caption. In this paper we investigate what can be learnt by looking at a large number of figures and reading their captions, and introduce a figure-caption correspondence learning task that makes use of our observations. Training visual and language networks without supervision other than pairs of unconstrained figures and captions is shown to successfully solve this task. We also show that transferring lexical and semantic knowledge from a knowledge graph significantly enriches the resulting features. Finally, we demonstrate the positive impact of such features in other tasks involving scientific text and figures, like multi-modal classification and machine comprehension for question answering, outperforming supervised baselines and ad-hoc approaches.
José Manuél Gómez-Pérez, Raúl Ortega 0001
K-CAP1
2019 Enabling FAIR research in Earth Science through research objects
Andrés García-Silva, José Manuél Gómez-Pérez, Raúl Palma, Marcin Krystek, Simone Mantovani, Federica Foglini, Valentina Grande, Francesco De Leo, Stefano Salvi, Elisa Trasatti, Vito Romaniello, Mirko Albani, Cristiano Silvagni, Rosemarie Leone, Fulvio Marelli, Sergio Albani, Michele Lazzarini, Hazel J. Napier, Ilkay Altintas
Future Gener. Comput. Syst.2
2018 A Research Object-Based Toolkit to Support the Earth Science Research Lifecycle
abstract
Data-intensive science disciplines, like Earth Science, are increasingly producing and consuming a variety of digital resources during the course of a scientific investigation. Instead of having these resources in isolated repositories, scientists are seeking ways for managing and making these resources available from a single place, and at the same time they are also increasingly interested in the adoption of FAIR principles to enhance the visibility and reusability of scientific results. This has called for new methods to improve the access and communication of results. Research Objects are a key building block towards realising this vision. They provide a structured way (a model) to describe scientific resources related to an investigation, along with the context in which they were used and the people involved. But research objects are as useful in practice as the availability of tools supporting their adoption. In this paper, we present a toolkit, tailored for Earth Sciences, comprising a set of services and applications around research objects that support scientists throughout the research lifecycle to manage, share, find and reuse scientific results, and we discuss initial insights into the community adoption.
Raúl Palma, Andrés García-Silva, José Manuél Gómez-Pérez, Marcin Krystek
eScience3
2017 Ensuring the Quality of Research Objects in the Earth Science Domain
abstract
Research objects were designed in data-intensive science under the premises of interoperability and machine-readability to describe scientific processes and findings including all the resources that were used in the research endeavour. In this poster we present our work with Earth Science communities, that have embraced the research object model for long-term preservation and reuse of knowledge, to design checklists, the main tool available to assess the quality of research objects and upon which quality metrics such completeness, stability and reliability are calculated.
Andrés García-Silva, Raúl Palma, José Manuél Gómez-Pérez
eScience3
2017 Semantic Technologies and Text Analysis in Support of Scientific Knowledge Reuse
abstract
Research objects act as a semantically rich container of all the information that lead to a scientific result, and have the potential to change how data intensive science shares and reuses methods, datasets and results. Despite a comprehensive set of vocabularies to describe semantically the aggregated resources, users often limit themselves to provide metadata at the container level, ignoring the valuable information that the aggregated resources contain. In this poster we explore how the combination of semantic technologies and natural language processing can be used to enrich research objects with structured metadata aiming at enhancing their findability as a crucial aspect towards their reuse.
Andrés García-Silva, Raúl Palma, José Manuél Gómez-Pérez
eScience3
2017 Research Objects for Interworkability among Global Environmental and Geophysical Data
abstract
The provisioning and exploitation at a global scale of environmental and geophysical data requires advanced automation and governance mechanisms that enable (meta)data interoperability but also the exchange of formalized scientific concepts and methods. In this paper we introduce recent efforts in such direction, based on scientific workflows and research objects as enablers of such vision. The former enable the integration of different web services in a higher-level data processing artifact while the latter enhances governance around data product validation and consistency, result reproducibility and credit to the principal investigators and data providers. This paper provides a concise overview of our project, current status and next steps.
José Manuél Gómez-Pérez, Chuck Meertens, Fran Boler, Henry Loescher, Christine Laney, Daniel Crawl, Ilkay Altintas
eScience1
2017 Towards a Human-Machine Scientific Partnership Based on Semantically Rich Research Objects
abstract
A research object is a single information unit encapsulating all the knowledge relevant to a particular scientific investigation, their associated metadata and the context where such resources were produced and came into play. Aimed at enhancing the preservation, reuse and scholarly communication of data-intensive science, research objects are both technical and social artifacts that represent a partnership between scientific communities and the computational support required in nowadays science. In this paper, we explore such partnership, identifying the lack of appropriate machine-readable metadata as one of its main inhibitors, and address the semantic enrichment of research objects as one key aspect towards its establishment. Focused on the specific needs of Earth Science communities, we propose extensions to research object representation models and present novel methods and tools to enrich research object metadata through automatic means. Finally, we validate the approach through the implementation of a recommender system that exploits the resulting metadata to facilitate research object discovery and reuse, enabling humans and machines to work together and accelerate the research life cycle.
José Manuél Gómez-Pérez, Raúl Palma, Andrés García-Silva
eScience1
2017 Supporting Research and Operational Earth Science Portals through ROHub
abstract
Earth science disciplines are increasingly producing large amounts of heterogeneous data that needs to be accessed, integrated and processed during the course of a research or operational process. In this setting, there is a growing need from earth scientists to collect and manage effectively these resources, including the data used, the methods applied, and the results produced. This is even more challenging if we consider i) the dynamic and collaborative nature of the earth science processes; ii) the interfaces typically used by scientists where they should be able to accomplish core management tasks with a minimal overhead. In this poster, we present our approach to address these challenges, which is based on research objects and on the associated technological support provided by ROHub platform. We show how ROHub is used as the underlying technology for supporting earth scientists work through different interfaces tailored to their communities needs and usage scenario.
Raúl Palma, Marcin Krystek, José Manuél Gómez-Pérez, Andrés García-Silva, Sergio Ferraresi, Simone Mantovani, Serena Avolio, Sergio De Gioia
eScience3
2015 Troubleshooting and Optimizing Named Entity Resolution Systems in the Industry
Panos Alexopoulos, Ronald Denaux, José Manuél Gómez-Pérez
ESWC3
2015 Using a suite of ontologies for preserving workflow-centric research objects
abstract
Scientific workflows are a popular mechanism for specifying and automating data-driven in silico experiments. A significant aspect of their value lies in their potential to be reused. Once shared, workflows become useful building blocks that can be combined or modified for developing new experiments. However, previous studies have shown that storing workflow specifications alone is not sufficient to ensure that they can be successfully reused, without being able to understand what the workflows aim to achieve or to re-enact them. To gain an understanding of the workflow, and how it may be used and repurposed for their needs, scientists require access to additional resources such as annotations describing the workflow, datasets used and produced by the workflow, and provenance traces recording workflow executions. In this article, we present a novel approach to the preservation of scientific workflows through the application of research objects—aggregations of data and metadata that enrich the workflow specifications. Our approach is realised as a suite of ontologies that support the creation of workflow-centric research objects. Their design was guided by requirements elicited from previous empirical analyses of workflow decay and repair. The ontologies developed make use of and extend existing well known ontologies, namely the Object Reuse and Exchange (ORE) vocabulary, the Annotation Ontology (AO) and the W3C PROV ontology (PROVO). We illustrate the application of the ontologies for building Workflow Research Objects with a case-study that investigates Huntington’s disease, performed in collaboration with a team from the Leiden University Medial Centre (HG-LUMC). Finally we present a number of tools developed for creating and managing workflow-centric research objects.
Khalid Belhajjame, Jun Zhao 0003, Daniel Garijo, Matthew Gamble, Kristina M. Hettne, Raúl Palma, Eleni Mina, Óscar Corcho, José Manuél Gómez-Pérez, Sean Bechhofer, Graham Klyne, Carole A. Goble
J. Web Semant.9
2013 Interactive acquisition of fuzzy ontological knowledge indialogue systems
abstract
In this paper we present a novel semi-automatic framework for acquiring and maintaining fuzzy ontological knowledge in dialogue systems. The main goal is to narrow the vagueness interpretation gap between the systems and their users and thus offer better information services to the latter.
Panos Alexopoulos, José Manuél Gómez-Pérez
K-CAP2
2013 When History Matters - Assessing Reliability for the Reuse of Scientific Workflows
José Manuél Gómez-Pérez, Esteban García-Cuesta, Aleix Garrido, José Enrique Ruiz, Jun Zhao 0003, Graham Klyne
ISWC (2)1
2013 A Formalism and Method for Representing and Reasoning with Process Models Authored by Subject Matter Experts
abstract
Enabling Subject Matter Experts (SMEs) to formulate knowledge without the intervention of Knowledge Engineers (KEs) requires providing SMEs with methods and tools that abstract the underlying knowledge representation and allow them to focus on modeling activities. Bridging the gap between SME-authored models and their representation is challenging, especially in the case of complex knowledge types like processes, where aspects like frame management, data, and control flow need to be addressed. In this paper, we describe how SME-authored process models can be provided with an operational semantics and grounded in a knowledge representation language like F-logic to support process-related reasoning. The main results of this work include a formalism for process representation and a mechanism for automatically translating process diagrams into executable code following such formalism. From all the process models authored by SMEs during evaluation 82 percent were well formed, all of which executed correctly. Additionally, the two optimizations applied to the code generation mechanism produced a performance improvement at reasoning time of 25 and 30 percent with respect to the base case, respectively.
José Manuél Gómez-Pérez, Michael Erdmann, Mark Greaves, Óscar Corcho
IEEE Trans. Knowl. Data Eng.1
2012 Why workflows break - Understanding and combating decay in Taverna workflows
abstract
Workflows provide a popular means for preserving scientific methods by explicitly encoding their process. However, some of them are subject to a decay in their ability to be re-executed or reproduce the same results over time, largely due to the volatility of the resources required for workflow executions. This paper provides an analysis of the root causes of workflow decay based on an empirical study of a collection of Taverna workflows from the myExperiment repository. Although our analysis was based on a specific type of workflow, the outcomes and methodology should be applicable to workflows from other systems, at least those whose executions also rely largely on accessing third-party resources. Based on our understanding about decay we recommend a minimal set of auxiliary resources to be preserved together with the workflows as an aggregation object and provide a software tool for end-users to create such aggregations and to assess their completeness.
Jun Zhao 0003, José Manuél Gómez-Pérez, Khalid Belhajjame, Graham Klyne, Esteban García-Cuesta, Aleix Garrido, Kristina M. Hettne, Marco Roos, David De Roure, Carole A. Goble
eScience2
2012 Open Innovation in an Enterprise 3.0 framework: Three case studies
Francesco Carbone, Jesús Contreras, Josefa Z. Hernández, José Manuél Gómez-Pérez
Expert Syst. Appl.4
2011 miKrow: Semantic Intra-enterprise Micro-Knowledge Management System
Víctor Penela, Guillermo Álvaro, Carlos Ruiz Moreno, Carmen Córdoba, Francesco Carbone, Michelangelo Castagnone, José Manuél Gómez-Pérez, Jesús Contreras
ESWC (2)7
2011 A Novel Approach to Visualizing and Navigating Ontologies
Enrico Motta, Paul Mulholland, Silvio Peroni, Mathieu d'Aquin, José Manuél Gómez-Pérez, Victor Mendez, Fouad Zablith
ISWC (1)5
2010 miKrow: Enabling Knowledge Management One Update at a Time
Guillermo Álvaro, Víctor Penela, Francesco Carbone, Carmen Córdoba, Michelangelo Castagnone, José Manuél Gómez-Pérez, Jesús Contreras
IC3K6
2010 A framework and computer system for knowledge-level acquisition, representation, and reasoning with process knowledge
José Manuél Gómez-Pérez, Michael Erdmann, Mark Greaves, Óscar Corcho, V. Richard Benjamins
Int. J. Hum. Comput. Stud.1
2007 Applying problem solving methods for process knowledge acquisition, representation, and reasoning
abstract
In this paper we present an approach towards knowledgeacquisition of process knowledge for the natural sciences.The work has been conducted within Project Halo, whichis creating advanced knowledge authoring and questionanswering systems for the natural sciences. An analysis of AP®-level questions for Biology, Chemistry and Physicsuncovered that process knowledge is the single most frequenttype of knowledge required. Thus, we developedmeans to acquire process knowledge, to formally representit, and to reason about it in order to answer novel questionsabout the domains.All these tasks are supported by an abstract process metamodel.It provides the terminology for user-tailored processdiagrams, which are automatically translated into executableFLogic code. The meta-model and the code generationare based on the notion of Problem Solving Methods(PSM) which represent an abstract formalization of thereasoning strategies needed for processes.
José Manuél Gómez-Pérez, Michael Erdmann, Mark Greaves
K-CAP1
2005 Solving Collaborative Fuzzy Agents Problems with CLP(FD)
Susana Muñoz-Hernández, José Manuél Gómez-Pérez
PADL2