Andrés García-Silva

dblp:59/4344 · DBLP profile ↗
← Back
17ranked-venue papers
12as first author
5since 2021 · last 2024
0000-0002-5664-488XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 7 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 first-authorArtificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2024 SPACE-IDEAS: A Dataset for Salient Information Detection in Space Innovation
abstract
Detecting salient parts in text using natural language processing has been widely used to mitigate the effects of information overflow. Nevertheless, most of the datasets available for this task are derived mainly from academic publications. We introduce SPACE-IDEAS, a dataset for salient information detection from innovation ideas related to the Space domain. The text in SPACE-IDEAS varies greatly and includes informal, technical, academic and business-oriented writing styles. In addition to a manually annotated dataset we release an extended version that is annotated using a large generative language model. We train different sentence and sequential sentence classifiers, and show that the automatically annotated dataset can be leveraged using multitask learning to train better classifiers.
Andrés García-Silva, Cristian Berrio, José Manuél Gómez-Pérez
LREC/COLING1
2023 Textual Entailment for Effective Triple Validation in Object Prediction
Andrés García-Silva, Cristian Berrio, José Manuél Gómez-Pérez
ISWC1
2022 SpaceQA: Answering Questions about the Design of Space Missions and Space Craft Concepts
abstract
We present SpaceQA, to the best of our knowledge the first open-domain QA system in Space mission design. SpaceQA is part of an initiative by the European Space Agency (ESA) to facilitate the access, sharing and reuse of information about Space mission design within the agency and with the public. We adopt a state-of-the-art architecture consisting of a dense retriever and a neural reader and opt for an approach based on transfer learning rather than fine-tuning due to the lack of domain-specific annotated data. Our evaluation on a test set produced by ESA is largely consistent with the results originally reported by the evaluated retrievers and confirms the need of fine tuning for reading comprehension. As of writing this paper, ESA is piloting SpaceQA internally.
Andrés García-Silva, Cristian Berrio, José Manuél Gómez-Pérez, José Antonio Martínez Heras, Alessandro Donati, Ilaria Roma
SIGIR1
2021 Classifying Scientific Publications with BERT - Is Self-attention a Feature Selection Method?
Andrés García-Silva, José Manuél Gómez-Pérez
ECIR (1)1
2021 On the impact of knowledge-based linguistic annotations in the quality of scientific embeddings
Andrés García-Silva, Ronald Denaux, José Manuél Gómez-Pérez
Future Gener. Comput. Syst.1
2020 Making Metadata Fit for Next Generation Language Technology Platforms: The Metadata Schema of the European Language Grid
abstract
The current scientific and technological landscape is characterised by the increasing availability of data resources and processing tools and services. In this setting, metadata have emerged as a key factor facilitating management, sharing and usage of such digital assets. In this paper we present ELG-SHARE, a rich metadata schema catering for the description of Language Resources and Technologies (processing and generation services and tools, models, corpora, term lists, etc.), as well as related entities (e.g., organizations, projects, supporting documents, etc.). The schema powers the European Language Grid platform that aims to be the primary hub and marketplace for industry-relevant Language Technology in Europe. ELG-SHARE has been based on various metadata schemas, vocabularies, and ontologies, as well as related recommendations and guidelines.
Penny Labropoulou, Katerina Gkirtzou, Maria Gavrilidou, Miltos Deligiannis, Dimitrios Galanis, Stelios Piperidis, Georg Rehm, Maria Berger, Valérie Mapelli, Mickaël Rigault, Victoria Arranz, Khalid Choukri, Gerhard Backfried, José Manuél Gómez-Pérez, Andrés García-Silva
LREC15
2020 European Language Grid: An Overview
abstract
With 24 official EU and many additional languages, multilingualism in Europe and an inclusive Digital Single Market can only be enabled through Language Technologies (LTs). European LT business is dominated by hundreds of SMEs and a few large players. Many are world-class, with technologies that outperform the global players. However, European LT business is also fragmented – by nation states, languages, verticals and sectors, significantly holding back its impact. The European Language Grid (ELG) project addresses this fragmentation by establishing the ELG as the primary platform for LT in Europe. The ELG is a scalable cloud platform, providing, in an easy-to-integrate way, access to hundreds of commercial and non-commercial LTs for all European languages, including running tools and services as well as data sets and resources. Once fully operational, it will enable the commercial and non-commercial European LT community to deposit and upload their technologies and data sets into the ELG, to deploy them through the grid, and to connect with other resources. The ELG will boost the Multilingual Digital Single Market towards a thriving European LT community, creating new jobs and opportunities. Furthermore, the ELG project organises two open calls for up to 20 pilot projects. It also sets up 32 national competence centres and the European LT Council for outreach and coordination purposes.
Georg Rehm, Maria Berger, Ela Elsholz, Stefanie Hegele, Florian Kintzel, Katrin Marheinecke, Stelios Piperidis, Miltos Deligiannis, Dimitrios Galanis, Katerina Gkirtzou, Penny Labropoulou, Kalina Bontcheva, Jan Hajic 0001, Jana Hamrlová, Lukás Kacena, Khalid Choukri, Victoria Arranz, Andrejs Vasiljevs, Orians Anvari, Andis Lagzdins, Julija Melnika, Gerhard Backfried, Erinç Dikici, Miroslav Jánosík, Katja Prinz, Christoph Prinz, Severin Stampler, Dorothea Thomas-Aniola, José Manuél Gómez-Pérez, Andrés García-Silva, Cristian Berrio, Ulrich Germann, Steve Renals, Ondrej Klejch
LREC32
2019 Enabling FAIR research in Earth Science through research objects
Andrés García-Silva, José Manuél Gómez-Pérez, Raúl Palma, Marcin Krystek, Simone Mantovani, Federica Foglini, Valentina Grande, Francesco De Leo, Stefano Salvi, Elisa Trasatti, Vito Romaniello, Mirko Albani, Cristiano Silvagni, Rosemarie Leone, Fulvio Marelli, Sergio Albani, Michele Lazzarini, Hazel J. Napier, Ilkay Altintas
Future Gener. Comput. Syst.1
2018 A Research Object-Based Toolkit to Support the Earth Science Research Lifecycle
abstract
Data-intensive science disciplines, like Earth Science, are increasingly producing and consuming a variety of digital resources during the course of a scientific investigation. Instead of having these resources in isolated repositories, scientists are seeking ways for managing and making these resources available from a single place, and at the same time they are also increasingly interested in the adoption of FAIR principles to enhance the visibility and reusability of scientific results. This has called for new methods to improve the access and communication of results. Research Objects are a key building block towards realising this vision. They provide a structured way (a model) to describe scientific resources related to an investigation, along with the context in which they were used and the people involved. But research objects are as useful in practice as the availability of tools supporting their adoption. In this paper, we present a toolkit, tailored for Earth Sciences, comprising a set of services and applications around research objects that support scientists throughout the research lifecycle to manage, share, find and reuse scientific results, and we discuss initial insights into the community adoption.
Raúl Palma, Andrés García-Silva, José Manuél Gómez-Pérez, Marcin Krystek
eScience2
2017 Ensuring the Quality of Research Objects in the Earth Science Domain
abstract
Research objects were designed in data-intensive science under the premises of interoperability and machine-readability to describe scientific processes and findings including all the resources that were used in the research endeavour. In this poster we present our work with Earth Science communities, that have embraced the research object model for long-term preservation and reuse of knowledge, to design checklists, the main tool available to assess the quality of research objects and upon which quality metrics such completeness, stability and reliability are calculated.
Andrés García-Silva, Raúl Palma, José Manuél Gómez-Pérez
eScience1
2017 Semantic Technologies and Text Analysis in Support of Scientific Knowledge Reuse
abstract
Research objects act as a semantically rich container of all the information that lead to a scientific result, and have the potential to change how data intensive science shares and reuses methods, datasets and results. Despite a comprehensive set of vocabularies to describe semantically the aggregated resources, users often limit themselves to provide metadata at the container level, ignoring the valuable information that the aggregated resources contain. In this poster we explore how the combination of semantic technologies and natural language processing can be used to enrich research objects with structured metadata aiming at enhancing their findability as a crucial aspect towards their reuse.
Andrés García-Silva, Raúl Palma, José Manuél Gómez-Pérez
eScience1
2017 Towards a Human-Machine Scientific Partnership Based on Semantically Rich Research Objects
abstract
A research object is a single information unit encapsulating all the knowledge relevant to a particular scientific investigation, their associated metadata and the context where such resources were produced and came into play. Aimed at enhancing the preservation, reuse and scholarly communication of data-intensive science, research objects are both technical and social artifacts that represent a partnership between scientific communities and the computational support required in nowadays science. In this paper, we explore such partnership, identifying the lack of appropriate machine-readable metadata as one of its main inhibitors, and address the semantic enrichment of research objects as one key aspect towards its establishment. Focused on the specific needs of Earth Science communities, we propose extensions to research object representation models and present novel methods and tools to enrich research object metadata through automatic means. Finally, we validate the approach through the implementation of a recommender system that exploits the resulting metadata to facilitate research object discovery and reuse, enabling humans and machines to work together and accelerate the research life cycle.
José Manuél Gómez-Pérez, Raúl Palma, Andrés García-Silva
eScience3
2017 Supporting Research and Operational Earth Science Portals through ROHub
abstract
Earth science disciplines are increasingly producing large amounts of heterogeneous data that needs to be accessed, integrated and processed during the course of a research or operational process. In this setting, there is a growing need from earth scientists to collect and manage effectively these resources, including the data used, the methods applied, and the results produced. This is even more challenging if we consider i) the dynamic and collaborative nature of the earth science processes; ii) the interfaces typically used by scientists where they should be able to accomplish core management tasks with a minimal overhead. In this poster, we present our approach to address these challenges, which is based on research objects and on the associated technological support provided by ROHub platform. We show how ROHub is used as the underlying technology for supporting earth scientists work through different interfaces tailored to their communities needs and usage scenario.
Raúl Palma, Marcin Krystek, José Manuél Gómez-Pérez, Andrés García-Silva, Sergio Ferraresi, Simone Mantovani, Serena Avolio, Sergio De Gioia
eScience4
2013 Semantic Characterization of Tweets Using Topic Models: A Use Case in the Entertainment Domain
abstract
In the entertainment domain users tweet about their expectations and opinions regarding upcoming, current and past experiences, while companies advertise and promote the shows. This characterization, important for customers and companies, goes beyond traditional sentiment analysis where the polarity of the sentiments expressed in opinions is usually identified as positive, negative or neutral. The authors investigate different tweet representation models, including bags of words and probabilistic topic models, to shed light on the semantics of the messages. Their experiments show that topic-based models generated with Latent Dirichlet Allocation (LDA) yield, most of the times, better categorizations when compared to TF-IDF based features, particularly when these models are enriched with natural language features and specific Twitter slang.
Andrés García-Silva, Víctor Rodríguez-Doncel, Óscar Corcho
Int. J. Semantic Web Inf. Syst.1
2012 Characterising Emergent Semantics in Twitter Lists
Andrés García-Silva, Jeon-Hyung Kang, Kristina Lerman, Óscar Corcho
ESWC1
2012 Enabling Folksonomies for Knowledge Extraction: A Semantic Grounding Approach
abstract
Folksonomies emerge as the result of the free tagging activity of a large number of users over a variety of resources. They can be considered as valuable sources from which it is possible to obtain emerging vocabularies that can be leveraged in knowledge extraction tasks. However, when it comes to understanding the meaning of tags in folksonomies, several problems mainly related to the appearance of synonymous and ambiguous tags arise, specifically in the context of multilinguality. The authors aim to turn folksonomies into knowledge structures where tag meanings are identified, and relations between them are asserted. For such purpose, they use DBpedia as a general knowledge base from which they leverage its multilingual capabilities.
Andrés García-Silva, Iván Cantador, Óscar Corcho
Int. J. Semantic Web Inf. Syst.1
2011 Multipedia: enriching DBpedia with multimedia information
abstract
Enriching knowledge bases with multimedia information makes it possible to complement textual descriptions with visual and audio information. Such complementary information can help users to understand the meaning of assertions, and in general improve the user experience with the knowledge base. In this paper we address the problem of how to enrich ontology instances with candidate images retrieved from existing Web search engines. DBpedia has evolved into a major hub in the Linked Data cloud, interconnecting millions of entities organized under a consistent ontology. Our approach taps into the Wikipedia corpus to gather context information for DBpedia instances and takes advantage of image tagging information when this is available to calculate semantic relatedness between instances and candidate images. We performed experiments with focus on the particularly challenging problem of highly ambiguous names. Both methods presented in this work outperformed the baseline. Our best method leveraged context words from Wikipedia, tags from Flickr and type information from DBpedia to achieve an average precision of 80%.
Andrés García-Silva, Max Jakob, Pablo N. Mendes, Christian Bizer
K-CAP1