VLDB 2026 Research / reviewers in the wild / expert
Rosa Filgueira
dblp:37/4672 · also Rosa Filgueira Vicente
· DBLP profile ↗
31ranked-venue papers
13as first author
15since 2021 · last 2026
0000-0002-5715-3046ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 19 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 6 first-author · 11 since 2021Systems, architecture and hardware · 11 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mapping Change: A Temporal and Semantic Knowledge Base of Scottish Gazetteers
Lilin Yu, Angelo A. Salatino, Rosa Filgueira |
ESWC (2) | 3 |
| 2025 | ImpactLens: Using AI and Data to Communicate Research Impact EffectivelyabstractAlthough the impact science has on society is a dimension that has been receiving increasing attention in recent years, surprisingly little work has been done on developing software tools that track it systematically by integrating information from different data sources and trying to capture all of its dimensions explicitly (e.g. impact on the economy, health and wellbeing, public understanding, the environment etc).In this paper, we present the prototypical ImpactLens system, which aims to leverage modern-day data-driven and AI technologies to extract ‘impact stories’ from rich, yet diverse data sources. We argue that the development of such tools can significantly enhance the ability of research communities to communicate the value of their work to society, while also providing opportunities to benefit researcher development, strategic research planning, and better exploitation of research results. Rosa Filgueira, Lilin Yu, Xingran Ruan, Sypro Nita, Laura Moran, Michael Rovatsos |
eScience | 1 |
| 2025 | Frances++: LLM-Based Semantic Enrichment and Spatial Graphs for Digitized Historical CollectionsabstractThis paper presents major enhancements to frances, a platform for exploring digitized historical collections using LLM-based semantic and spatial enrichments. We focus on three corpora from the National Library of Scotland (NLS): the Gazetteers of Scotland, the Broadsides, and the Encyclopaedia Britannica. Using a flexible, prompt-based LLM pipeline, we extract over 50,000 structured articles from the Gazetteers, adapting to diverse typographic layouts. For Broadsides, we align OCR-derived text (extracted using defoe) with manually corrected transcriptions published by NLS, linking both sources for improved access and reuse. Historical text is further improved using a fine-tuned Llama2 model selected after evaluating several LLMs for post-OCR correction. To enrich geographic information, we combine Stanza NER, in-text coordinate parsing, and the Edinburgh Geoparser, enabling both modern georesolution and preservation of historical geography. We extend the Heritage Textual Ontology (HTO) to support article-level records, spatiotemporal entities, and in-text annotations via CRMgeo and Web Annotation standards. Resulting RDF knowledge graphs are deployed on a GeoSPARQL-enabled Fuseki server and indexed in Elasticsearch to support full-text, semantic, and spatial search. The upgraded frances interface offers entity highlighting, historical maps, and provenance tracking, demonstrating a scalable pipeline for structured, enriched access to OCRed heritage texts. Lilin Yu, Rosa Filgueira |
eScience | 2 |
| 2024 | Advancing frances: New Heritage Textual Ontology, Enhanced Knowledge Graphs, and Refined Search CapabilitiesabstractThis paper presents significant enhancements to the frances platform, incorporating the Heritage Textual Ontology (HTO), advanced knowledge graphs, and sophisticated search capabilities, along with innovative data visualization methods. The HTO integrates diverse historical collections and unifies various sources. Leveraging this ontology, the new knowledge graphs connect data across different sources and editions, linking to external resources like Wikipedia and Dbpedia to enrich semantic relationships. We employed deep-learning-based spell correction for OCR error correction. Enhanced search functionalities, powered by Elasticsearch and semantic technologies, enable precise retrieval and analysis. Additionally, new data visualization approaches offer multifaceted interpretations of search results. A case study tracking slavery references in historical editions of the Encyclopaedia Britannica 1768-1860 demonstrates the platform’s effectiveness in analyzing historical text and validating frances’s capabilities. Lilin Yu, Ash Charlton, Melissa Terras, Rosa Filgueira |
e-Science | 4 |
| 2024 | Multi-Level AI-Driven Analysis of Software Repository SimilaritiesabstractThis paper introduces significant enhancements to RepoSim4Py and RepoSnipy, advanced semantic tools for deep analysis of software repositories. RepoSim4Py commandline toolbox now supports multi-level embedding, encompassing code, documentation, requirements, README, and comprehensive repository analysis, which enable the understanding of repository dynamics. Concurrently, RepoSnipy webbased search engine facilitates sophisticated repository similarity searches and introduces clustering based on both repository tags (topic_cluster) and code embeddings (code_cluster). We also introduce SimilarityCal, a novel binary classification model trained on these clusters, to predict and quantify repository similarities with high accuracy. These developments provide researchers and developers with powerful tools to navigate the complex landscape of software repositories, improving efficiency in software development and fostering innovation through better reuse of existing resources. Honglin Zhang, Leyu Zhang, Lei Fang 0001, Rosa Filgueira |
e-Science | 4 |
| 2023 | Mapping the Repository Landscape: Harnessing Similarity with RepoSim and RepoSnipyabstractThe rapid growth of scientific software development has led to the emergence of large and complex codebases, making it challenging to search, find, and compare software repositories within the scientific research community. In this paper, we propose a solution by leveraging deep learning techniques to learn embeddings that capture semantic similarities among repositories. Our approach focuses on identifying repositories with similar semantics, even when their code fragments and documentation exhibit different syntax. To address this challenge, we introduce two complementary open-source tools: RepoSim and RepoSnipy. RepoSim is a command-line toolbox designed to represent repositories at both the source code and documentation levels. It utilizes the UniXcoder pre-trained language model, which has demonstrated remarkable performance in code-related understanding tasks. RepoSnipy is a web-based neural semantic search engine that utilizes the powerful capabilities of RepoSim and offers a user-friendly search interface, allowing researchers and practitioners to query public repositories hosted on GitHub and discover semantically similar repositories. RepoSim and RepoSnipy empower researchers, developers, and practitioners by facilitating the comparison and analysis of software repositories. They not only enable efficient collaboration and code reuse but also accelerate the development of scientific software. Rosa Filgueira |
e-Science | 2 |
| 2023 | Message from the IEEE eScience 2023 Conference Leadership eScience 2023abstractThe 19th IEEE Conference on eScience (eScience 2023), which took place from October 9th to 13th in Limassol (Cyprus), provided a platform where researchers, developers, and users of eScience applications and enabling IT technologies delved into the realms of interdisciplinary collaboration. The mission of the eScience conference was to drive innovation in data- and compute-intensive research, spanning a wide array of disciplines, including the physical and biological sciences, as well as the social sciences, arts, and humanities. IEEE eScience 2023 successfully dismantled traditional barriers, nurturing collaboration among interdisciplinary research communities, developers, and eScience application users. Throughout that five-day event, the conference focus was on advancing all aspects of eScience and its associated technologies, applications, algorithms, and tools, with a strong emphasis on practical solutions to real-world challenges. New Directions and Communities: This year, we have introduced four key topics: “Computational Science for sustainable development,” “FAIR,” “Research Infrastructures for eScience,” and “Continuum Computing: Convergence between Cloud Computing and the Internet of Things (IoT).” These tracks enriched our discussions and investigations. Keynote Speakers: In addition to these exciting tracks, we were delighted to present our distinguished keynote speakers: Dr. İlkay Altıntaş illuminated the convergence of machine learning, AI, and scientific research, showcasing innovative approaches and real-world applications. Professor Ian T. Foster took us on a journey into the realm of global science services, demonstrating how they can reshape collaborative research. Finally, Professor Paul Watson shared insights from the National Innovation Centre for Data, highlighting successful projects that leverage data science and AI to drive impact. Program Overview: As we embarked on this eScience journey, we invited participants to engage, collaborate, and explore the transformative potential of eScience and its associated technologies. Together, we addressed practical solutions, open challenges, and pushed the boundaries of interdisciplinary research. This year's conference featured three keynotes from different domains, presented 40 peer-reviewed papers, showcased 23 posters, featured 15 invited talks from renowned researchers, hosted 6 workshops, and provided 5 tutorials. The conference was held in person and was co-located with the 4th Global Research Platform Workshop (4GRP). Our sincere appreciation goes out to the authors who submitted exceptional papers, and we are immensely thankful for the dedicated efforts of our numerous volunteers. The Technical Paper Committee diligently reviewed 96 papers and 25 posters, with over 70 volunteers conducting 308 reviews. This collaborative endeavour has resulted in the exceptional collection of papers featured in this volume. We would also like to express our deep gratitude to the IEEE Computer Society and IEEE's Technical Committee on High-Performance Computing (TCHPC) for their unwavering sponsorship and support, which have been instrumental in making this conference possible. Finally, we extend our gratitude to you, our valued readers, for your interest in this volume. We are confident that the contents within will greatly contribute to the advancement of your research, whether it pertains to eScience or other domains. George Angelos Papadopoulos, Rafael Ferreira da Silva, Rosa Filgueira |
e-Science | 3 |
| 2023 | Building Lightweight Semantic Search EnginesabstractDespite significant advances in methods for processing large volumes of structured and unstructured data, surprisingly little attention has been devoted to developing general practical methodologies that leverage state-of-the-art technologies to build domain-specific semantic search engines tailored to use cases where they could provide substantial benefits. This paper presents a methodology for developing these kinds of systems in a lightweight, modular, and flexible way with a particular focus on providing powerful search tools in domains where non-expert users encounter challenges in exploring the data repository at hand. Using an academic expertise finder tool as a case study, we demonstrate how this methodology allows us to leverage powerful off-the-shelf technology to enable the rapid, low-cost development of semantic search engines, while also affording developers with the necessary flexibility to embed user-centric design in their development in order to maximise uptake and application value. Michael Rovatsos, Rosa Filgueira |
e-Science | 2 |
| 2023 | RepoGraph: A Novel Semantic Code Exploration Tool for Python Repositories Based on Knowledge Graphs and Deep LearningabstractThis work presents RepoGraph, an integrated semantic code exploration web tool that combines information extraction, knowledge graphs, and deep learning models. It offers new capabilities for software developers (from academia and industry) to represent and query Python repositories. Unlike existing tools, RepoGraph not only provides a novel search interface powered by deep learning techniques but also exposes the underlying features and representations of repositories to users. Additionally, it offers several interactive visualizations. We also introduce RepoPyOnto, a new ontology that captures the features of Python code repositories and is used by RepoGraph for representing the captured knowledge. Finally, we successfully evaluate RepoGraph against several criteria, including function summarization performance, the correctness and relevance of search results, as well as the processing time for constructing graphs of various sizes. Rosa Filgueira |
e-Science | 2 |
| 2023 | frances: Cloud-Based Historical Text Mining with Deep Learning and Parallel ProcessingabstractFrances is an advanced cloud-based text mining digital platform that leverages information extraction, knowledge graphs, natural language processing (NLP), deep learning, and parallel processing techniques. It has been specifically designed to unlock the full potential of historical digital textual collections, such as those from the National Library of Scotland, offering cloud-based capabilities and extended support for complex NLP analyses and data visualizations. frances enables realtime recurrent operational text mining and provides robust capabilities for temporal analysis, accompanied by automatic visualizations for easy result inspection. In this paper, we present the motivation behind the development of frances, emphasizing its innovative design and novel implementation aspects. We also outline future development directions, and we evaluate the platform through two comprehensive case studies in history and publishing history. Lilin Yu, Ash Charlton, Wilfrid Askins, Melissa Terras, Rosa Filgueira |
e-Science | 5 |
| 2022 | SparkFlow: Towards High-Performance Data Analytics for Spark-based Genome AnalysisabstractThe recent advances in DNA sequencing technology triggered next-generation sequencing (NGS) research in full scale. Big Data (BD) is becoming the main driver in analyzing these large-scale bioinformatics data. However, this complicated process has become the system bottleneck, requiring an amal-gamation of scalable approaches to deliver the needed performance and hide the deployment complexity. Utilizing cutting-edge scientific workflows can robustly address these challenges. This paper presents a Spark-based alignment workflow called SparkFlow for massive NGS analysis over singularity containers. SparkFlow is highly scalable, reproducible, and capable of parallelizing computation by utilizing data-level parallelism and load balancing techniques in HPC and Cloud environments. The proposed workflow capitalizes on benchmarking two state-of-art NGS workflows, i.e., Base Recalibrator and ApplyBQSR. SparkFlow realizes the ability to accelerate large-scale cancer genomic analysis by scaling vertically (HyperThreading) and horizontally (provisions on-demand). Our result demonstrates a trade-off inevitably between the targeted applications and proces-sor architecture. SparkFlow achieves a decisive improvement in NGS computation performance, throughput, and scalability while maintaining deployment complexity. The paper's findings aim to pave the way for a wide range of revolutionary enhancements and future trends within the High-performance Data Analytics (HPDA) genome analysis realm. Rosa Filgueira, Feras M. Awaysheh, Adam C. Carter, Darren J. White, Omer F. Rana |
CCGRID | 1 |
| 2022 | frances: A Deep Learning NLP and Text Mining Web Tool to Unlock Historical Digital Collections: A Case Study on the Encyclopaedia BritannicaabstractThis work presents frances, an integrated text mining tool that combines information extraction, knowledge graphs, NLP, deep learning, parallel processing and Semantic Web techniques to unlock the full value of historical digital textual collections, offering new capabilities for researchers to use powerful analysis methods without being distracted by the technology and middleware details. To demonstrate these capabilities, we use the first eight editions of the Encyclopaedia Britannica offered by the National Library of Scotland (NLS) as an example digital collection to mine and analyse. We have developed novel parallel heuristics to extract terms from the original collection (alongside metadata), which provides a mix of unstructured and semi-structured input data, and populated a new knowledge graph with this information. Our Natural Language Processing models enable frances to perform advanced analyses that go significantly beyond simple search using the information stored in the knowledge graph. Furthermore, frances also allows for creating and running complex text mining analyses at scale. Our results show that the novel computational techniques developed within frances provide a vehicle for researchers to formalize and connect findings and insights derived from the analysis of large-scale digital corpora such as the Encyclopaedia Britannica. Rosa Filgueira |
e-Science | 1 |
| 2022 | Inspect4py: A Knowledge Extraction Framework for Python Code RepositoriesabstractThis work presents inspect4py, a static code analysis framework designed to automatically extract the main features, metadata and documentation of Python code repositories. Given an input folder with code, inspect4py uses abstract syntax trees and state of the art tools to find all functions, classes, tests, documentation, call graphs, module dependencies and control flows within all code files in that repository. Using these findings, inspect4py infers different ways of invoking a software component. We have evaluated our framework on 95 annotated repositories, obtaining promising results for software type classification (over 95% F1-score). With inspect4py, we aim to ease the understandability and adoption of software repositories by other researchers and developers. Rosa Filgueira, Daniel Garijo |
MSR | 1 |
| 2022 | Scalable adaptive optimizations for stream-based workflows in multi-HPC-clusters and cloud infrastructures
Rosa Filgueira, Thomas Heinis |
Future Gener. Comput. Syst. | 2 |
| 2021 | Extending defoe for the Efficient Analysis of Historical Texts at ScaleabstractThis paper presents the new facilities provided in defoe, a parallel toolbox for querying a wealth of digitised newspapers and books at scale. defoe has been extended to work with further Natural Language Processing () tools such as the Edinburgh Geoparser, to store the preprocessed text in several storage facilities and to support different types of queries and analyses. We have also extended the collection of XML schemas supported by defoe, increasing the versatility of the tool for the analysis of digital historical textual data at scale. Finally, we have conducted several studies in which we worked with humanities and social science researchers who posed complex and interested questions to large-scale digital collections. Results shows that defoe allows researchers to conduct their studies and obtain results faster, while all the large-scale text mining complexity is automatically handled by defoe. Rosa Filgueira, Claire Grover, Vasilios Karaiskos, Beatrice Alex, Sarah Van Eyndhoven, Lisa Gotthard, Melissa Terras |
e-Science | 1 |
| 2019 | Comprehensible Control for Researchers and Developers Facing Data ChallengesabstractThe DARE platform enables researchers and their developers to exploit more capabilities to handle complexity and scale in data, computation and collaboration. Today's challenges pose increasing and urgent demands for this combination of capabilities. To meet technical, economic and governance constraints, application communities must use use shared digital infrastructure principally via virtualisation and mapping. This requires precise abstractions that retain their meaning while their implementations and infrastructures change. Giving specialists direct control over these capabilities with detail relevant to each discipline is necessary for adoption. Research agility, improved power and retained return on intellectual investment incentivise that adoption. We report on an architecture for establishing and sustaining the necessary optimised mappings and early evaluations of its feasibility with two application communities. Malcolm P. Atkinson 0001, Rosa Filgueira, Iraklis A. Klampanos, Antonis Koukourikos, Amrey Krause, Federica Magnoni, Christian Pagé, Andreas Rietbrock, Alessandro Spinuso |
eScience | 2 |
| 2019 | defoe: A Spark-Based Toolbox for Analysing Digital Historical Textual DataabstractThis work presents defoe, a new scalable and portable digital eScience toolbox that enables historical research. It allows for running text mining queries across large datasets, such as historical newspapers and books in parallel via Apache Spark. It handles queries against collections that comprise several XML schemas and physical representations. The proposed tool has been successfully evaluated using five different large-scale historical text datasets and two HPC environments, as well as on desktops. Results shows that defoe allows researchers to query multiple datasets in parallel from a single command-line interface and in a consistent way, without any HPC environment-specific requirements. Rosa Filgueira, Mariona Coll Ardanuy, Giovanni Colavizza, James Hetherington, Melissa Terras, Anna Roubícková, Amrey Krause, Ruth Ahnert, Tessa Hauswedell, Julianne Nyhan, David Beavan, Timothy Hobson |
eScience | 1 |
| 2019 | DARE: A Reflective Platform Designed to Enable Agile Data-Driven Research on the CloudabstractThe DARE platform has been designed to help research developers deliver user-facing applications and solutions over diverse underlying e-infrastructures, data and computational contexts. The platform is Cloud-ready, and relies on the exposure of APIs, which are suitable for raising the abstraction level and hiding complexity. At its core, the platform implements the cataloguing and execution of fine-grained and Python-based dispel4py workflows as services. Reflection is achieved via a logical knowledge base, comprising multiple internal catalogues, registries and semantics, while it supports persistent and pervasive data provenance. This paper presents design and implementation aspects of the DARE platform, as well as it provides directions for future development. Iraklis A. Klampanos, Federica Magnoni, Emanuele Casarotti, Christian Pagé, Mike Lindner, Andreas Ikonomopoulos, Vangelis Karkaletsis, Athanasios Davvetas, André Gemünd, Malcolm P. Atkinson 0001, Antonis Koukourikos, Rosa Filgueira, Amrey Krause, Alessandro Spinuso, Angelos Charalambidis |
eScience | 12 |
| 2019 | Using simple PID-inspired controllers for online resilient resource management of distributed scientific workflows
Rafael Ferreira da Silva, Rosa Filgueira, Ewa Deelman, Erola Pairo-Castineira, Ian Michael Overton, Malcolm P. Atkinson 0001 |
Future Gener. Comput. Syst. | 2 |
| 2018 | Establishing Core Concepts for Information-Powered Collaborations
Luca Trani, Malcolm P. Atkinson 0001, Daniele Bailo, Rossana Paciello, Rosa Filgueira |
Future Gener. Comput. Syst. | 5 |
| 2017 | A characterization of workflow management systems for extreme-scale applications
Rafael Ferreira da Silva, Rosa Filgueira, Ilia Pietri, Ming Jiang 0005, Rizos Sakellariou, Ewa Deelman |
Future Gener. Comput. Syst. | 2 |
| 2016 | Automating environmental computing applications with scientific workflowsabstractComputational environmental science applications have evolved and become more complex over the last decade. In order to cope with the needs of such applications, computational methods and technologies have emerged to support the execution of these applications on heterogeneous, distributed systems. Among them are workflow management systems such as Pegasus. Pegasus is being used by researchers to model seismic wave propagation, to discover new celestial objects, to study RNA critical to human brain development, and to investigate other important research questions. This paper provides an introduction to scientific workflows and describes Pegasus and its main features. The paper highlights how the environmental science community has used Pegasus to automate their scientific workflow executions on high performance and high throughput computing systems by presenting three use cases: two Earth science workflows, and a climate science workflow. Rafael Ferreira da Silva, Ewa Deelman, Rosa Filgueira, Karan Vahi, Mats Rynge, Rajiv Mayani, Benjamin Mayer |
eScience | 3 |
| 2015 | VERCE Delivers a Productive E-science Environment for Seismology ResearchabstractThe VERCE project has pioneered an e-Infrastructure to support researchers using established simulation codes on high-performance computers in conjunction with multiple sources of observational data. This is accessed and organised via the VERCE science gateway that makes it convenient for seismologists to use these resources from any location via the Internet. Their data handling is made flexible and scalable by two Python libraries, ObsPy and dispel4py and by data services delivered by ORFEUS and EUDAT. Provenance driven tools enable rapid exploration of results and of the relationships between data, which accelerates understanding and method improvement. These powerful facilities are integrated and draw on many other e-Infrastructures. This paper presents the motivation for building such systems, it reviews how solid-Earth scientists can make significant research progress using them and explains the architecture and mechanisms that make their construction and operation achievable. We conclude with a summary of the achievements to date and identify the crucial steps needed to extend the capabilities for seismologists, for solid-Earth scientists and for similar disciplines. Malcolm P. Atkinson 0001, Michele Carpenè, Emanuele Casarotti, Steffen Claus, Rosa Filgueira, Anton Frank, Michelle Galea, Tom Garth, André Gemünd, Heiner Igel, Iraklis A. Klampanos, Amrey Krause, Lion Krischer, Siew Hoon Leong, Federica Magnoni, Jonas Matser, Alberto Michelini, Andreas Rietbrock, Horst Schwichtenberg, Alessandro Spinuso, Jean-Pierre Vilotte |
e-Science | 5 |
| 2015 | dispel4py: An Agile Framework for Data-Intensive eScienceabstractWe present dispel4py a versatile data-intensive kit presented as a standard Python library. It empowers scientists to experiment and test ideas using their familiar rapid-prototyping environment. It delivers mappings to diverse computing infrastructures, including cloud technologies, HPC architectures and specialised data-intensive machines, to move seamlessly into production with large-scale data loads. The mappings are fully automated, so that the encoded data analyses and data handling are completely unchanged. The underpinning model is lightweight composition of fine-grained operations on data, coupled together by data streams that use the lowest cost technology available. These fine-grained workflows are locally interpreted during development and mapped to multiple nodes and systems such as MPI and Storm for production. We explain why such an approach is becoming more essential in order that data-driven research can innovate rapidly and exploit the growing wealth of data while adapting to current technical trends. We show how provenance management is provided to improve understanding and reproducibility, and how a registry supports consistency and sharing. Three application domains are reported and measurements on multiple infrastructures show the optimisations achieved. Finally we present the next steps to achieve scalability and performance. Rosa Filgueira, Amrey Krause, Malcolm P. Atkinson 0001, Iraklis A. Klampanos, Alessandro Spinuso, Susana Sánchez-Expósito |
e-Science | 1 |
| 2014 | eScience Gateway Stimulating Collaboration in Rock Physics and VolcanologyabstractEarth scientist observe many facets of the planet's crust and integrate their resulting data to better understand the processes at work. We report on a new data-intensive science gateway designed to bring rock physicists and volcanologists into a collaborative framework that enables them to accelerate their research and integrate well with other Earth scientists. The science gateway supports three major functions: 1) sharing data from laboratories and observatories, experimental facilities and computational model runs, 2) sharing computational models and methods for analysing experimental and observational data, and 3) supporting recurrent tasks, such as data collection and running application in real time. Our prototype gateway has worked with two exemplar projects giving experience of data gathering, model sharing and data analysis. The geoscientists found that the gateway accelerated their work, triggered new practices and provided a good platform for long-term collaboration. Rosa Filgueira, Malcolm P. Atkinson 0001, Ian G. Main, Steve Boon, Christopher Kilburn, Philip Meredith |
eScience | 1 |
| 2014 | Applying Selectively Parallel I/O Compression to Parallel Storage Systems
Rosa Filgueira, Malcolm P. Atkinson 0001, Yusuke Tanimura, Isao Kojima |
Euro-Par | 1 |
| 2013 | MPI collective I/O based on advanced reservations to obtain performance guarantees from shared storage systemsabstractAs more data-intensive computing applications are executed on high performance computing clusters, resource contention on the shared storage system attached to the clusters becomes significant. The contention might cause I/O performance degradation and spoil performance improvement of coordinated parallel I/O by the MPI-IO implementation. In order to solve this problem, an advanced reservation approach where storage resources are managed based on the reservations to satisfy the I/O performance requirements, has been proposed. In this paper, we apply the concept of reserved data access to MPI-IO, in particular to Two-Phase collective I/O which is primarily used for I/O aggregation in non-contiguous access by MPI applications. We developed a prototype by using Dynamic-CoMPI which supports further improvement of Two-Phase I/O by using a locality aware strategy, and Papio which is a parallel storage system providing performance reservation functionality. After describing our prototype design and implementation, we show leverage of the concept by comparing our implementation with other existing MPI-IO implementations backed by OrangeFS and Lustre. The evaluation experiment confirms that the optimization benefit of Two-Phase I/O can be preserved by our approach, under the resource contention situation. Yusuke Tanimura, Rosa Filgueira, Isao Kojima, Malcolm P. Atkinson 0001 |
CLUSTER | 2 |
| 2012 | An Adaptive, Scalable, and Portable Technique for Speeding Up MPI-Based Applications
Rosa Filgueira, Malcolm P. Atkinson 0001, Alberto Nuñez, Javier Fernández 0001 |
Euro-Par | 1 |
| 2012 | Dynamic-CoMPI: dynamic optimization techniques for MPI parallel applications
Rosa Filgueira, Jesús Carretero 0001, David E. Singh, Alejandro Calderón 0001, Alberto Nuñez |
J. Supercomput. | 1 |
| 2008 | Exploiting data compression in collective I/O techniquesabstractThis paper presents Two-Phase Compressed I/O (TPC I/O,) an optimization of the Two-Phase collective I/O technique from ROMIO, the most popular MPI-IO implementation. In order to reduce network traffic, TPC I/O employs LZO algorithm to compress and decompress exchanged data in the inter-node communication operations. The compression algorithm has been fully implemented in the MPI collective technique, allowing to dynamically use (or not) compression. Compared with Two-Phase I/O, Two-Phase Compressed I/O obtains important improvements in the overall execution time for many of the considered scenarios. Rosa Filgueira, David E. Singh, Juan Carlos Pichel, Jesús Carretero 0001 |
CLUSTER | 1 |
| 2007 | Optimization and evaluation of parallel I/O in BIPS3D parallel irregular applicationabstractThis paper presents the optimization and evaluation of parallel I/O for the BIPS3D parallel irregular application, a 3-dimensional simulation of BJT and HBT bipolar devices. The parallel version of BIPS3D employs Metis, a library for partitioning graphs, finite element meshes, or sparse matrices. First, we show how the partitioning information provided by Metis can be used in order to improve the performance of parallel I/O. Second, we propose a novel technique, called Interval Data Grouping (IDG), which exploits the data replication of mesh nodes for optimizing the scheduling of the parallel file operations. Finally, we evaluate the parallel I/O version of BIPS3D for various existing parallel I/O techniques and present an in-depth analysis of the IDG performance. Rosa Filgueira, David E. Singh, Florin Isaila, Jesús Carretero 0001, Antonio J. García-Loureiro |
IPDPS | 1 |