Helena Galhardas

dblp:g/HelenaGalhardas · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-9330-3910ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 20 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Efficient and Scalable Search for Statistics
abstract
International audience
Antoine Gauquier, Simon Ebel, Helena Galhardas, Théo Galizzi, Ioana Manolescu, Aurélien Peden, Pierre Senellart
ICDE3
2024 STaR: Space and Time-aware Statistic Query Answering
abstract
High-quality data is essential for informed public debate. High-quality statistical data sources provide valuable reference information for verifying claims. To assist journalists and fact-checkers, user queries about specific claims should be automatically answered using statistical tables. However, the large number and variety of these sources make this task challenging.
Oana Balalau, Simon Ebel, Helena Galhardas, Théo Galizzi, Ioana Manolescu
CIKM3
2022 DeepData: Machine learning in the marine ecosystems
abstract
Based on environmental and species monitoring data, Species Distribution Modelling (SDM) tries to build a model to predict the distribution of a species across a geographic area. These models can then be used to manage the activities in the area in order to prevent negative economic and environmental impacts. In marine ecosystems, SDM can be used to regulate fishing practices or manage protected areas. This paper presents DeepData, a new no-code web-based machine learning platform to facilitate the work of marine biologists with SDM. The DeepData tool enables to automate SDM, by automating the creation and validation of the model by marine biologists. Biologists mostly use probabilistic algorithms, such as maximum entropy, generalized linear models and generalized additive models. The DeepData tool also allows the use of machine learning algorithms, such as classification and regression trees, random forests and support vector machines. Moreover, besides the usage of machine learning algorithms, other steps in SDM, such as data preparation and model evaluation, are also discussed in the paper. Furthermore, a concrete explanation of the use of the DeepData tool is presented, as well as the details of implementation and evaluation.
Leonor Silva, Magda Resende, Helena Galhardas, Vasco Manquinho, Inês Lynce
Expert Syst. Appl.3
2022 Graph integration of structured, semistructured and unstructured data for data journalism
Angelos-Christos G. Anadiotis, Oana Balalau, Catarina Conceição, Helena Galhardas, Mhd Yamen Haddad, Ioana Manolescu, Tayeb Merabti, Jingmao You
Inf. Syst.4
2022 AcX: System, Techniques, and Experiments for Acronym Expansion
abstract
In this information-accumulating world, each of us must learn continuously. To participate in a new field, or even a sub-field, one must be aware of the terminology including the acronyms that specialists know so well, but newcomers do not. Building on state-of-the art acronym tools, our end-to-end acronym expander system called AcX takes a document, identifies its acronyms, and suggests expansions that are either found in the document or appropriate given the subject matter of the document. As far as we know, AcX is the first open source and extensible system for acronym expansion that allows mixing and matching of different inference modules. As of now, AcX works for English, French, and Portuguese with other languages in progress. This paper describes the design and implementation of AcX, proposes three new acronym expansion benchmarks , compares state-of-the-art techniques on them, and proposes ensemble techniques that improve on any single technique. Finally, the paper evaluates the performance of AcX and related work MadDog system in end-to-end experiments on a new human-annotated dataset of Wikipedia documents. Our experiments show that AcX outperforms MadDog but that human performance is still substantially better than the best automated approaches. Thus, achieving Acronym Expansion at a human level is still a rich and open challenge.
João L. M. Pereira, João Casanova, Helena Galhardas, Dennis E. Shasha
Proc. VLDB Endow.3
2021 Discovering Conflicts of Interest across Heterogeneous Data Sources with ConnectionLens
abstract
Investigative Journalism (IJ, in short) requires combining highly heterogeneous digital datasets coming from a wide variety of sources. We have developed ConnectionLens, a system that integrates such sources into a single heterogeneous graph and enables users to query the graph using keywords. The first iteration of the system [7] followed a mediator architecture which severely constrained its query scalability. Thus, we fully re-engineered the system, moving it to a warehouse architecture, and replacing its core components (information extraction, data querying, and interactive interfaces), which allowed us to handle uses cases orders of magnitude larger than the previous platform. In a consortium of computer scientists and investigative journalists, we propose to demonstrate ConnectionLens' capability to integrate arbitrary heterogeneous datasets and query them flexibly by means of keywords. Among several scenarios, our main focus will be on a real-world journalistic use case about situations which may lead to Conflicts of Interest between biomedical experts and various organizations, such as corporations, lobbies, etc. The demonstration will showcase the end-to-end data analysis pipeline, illustrate each system component, and the different parameters governing graph creation and querying.
Angelos-Christos G. Anadiotis, Oana Balalau, Théo Bouganim, Francesco Chimienti, Helena Galhardas, Mhd Yamen Haddad, Stephane Horel, Ioana Manolescu, Youssr Youssef
CIKM5
2021 UNIANO: robust and efficient anomaly consensus in time series sensitive to cross-correlated anomaly profiles
abstract
Time series anomaly detection is an active research area, combining dozens of state-of-the-art methods that place heterogeneous views on what is an anomaly.This diversity of views -local and global, point and segment, univariate and multivariate, context-free and context-aware anomalies -is associated with moderate-to-high output divergences between methods.As a result, the user is faced with the difficult and laborious task of selecting the most appropriate methods and identifying cross-method consensus in an attempt to optimize recall and precision.Despite the relevance of establishing agreement criteria, existing principles are scarce and suffer from major problems: 1) show biases towards methods with correlated/redundant anomaly profiles; 2) depend on anomaly score thresholding; 3) prevent online detection; and 4) offer consensus not subjected to sound statistical testing.This work proposes UNIANO (UNIfied ANOmaly), an approach that combines simple yet effective empirical multivariate distribution statistics to address these drawbacks, guaranteeing a parameter-free and statistically robust integration of heterogeneous anomaly views.In this context, anomalies detected by less prevalent and concordant anomaly profiles, such as context-aware profiles in the presence of complementary variables, are not undervalued.Given a n-length time series and m views, UNIANO is aided by adequate data structures to achieve O(n log m 2 n) training time and linear O(m) testing-and-updating time.The gathered results confirm the relevance of the proposed approach.
Leonor Silva, Helena Galhardas, Vasco Manquinho, Rui Henriques
SDM2
2019 On-demand big data integration - A hybrid ETL approach for reproducible scientific research
Pradeeban Kathiravelu, Ashish Sharma 0001, Helena Galhardas, Peter Van Roy, Luís Veiga
Distributed Parallel Databases3
2018 Adaptive Execution of Continuous and Data-intensive Workflows with Machine Learning
abstract
To extract value from evergrowing volumes of data and to drive decision making, organizations frequently resort to the composition of data processing workflows. The typical workflow model enforces strict temporal synchronization across processing steps without accounting the actual effect of intermediate computations on the final workflow output. However, this is not the most desirable in a multitude of scenarios. We identify a class of applications for continuous data processing where the workflow output changes slowly and without great significance in a short time window, thus squandering compute resources with current approaches.
Sérgio Esteves, Helena Galhardas, Luís Veiga
Middleware2
2018 ConnectionLens: Finding Connections Across Heterogeneous Data Sources
abstract
Nowadays, journalism is facilitated by the existence of large amounts of publicly available digital data sources. In particular, journalists can do investigative work, which typically consists on keyword-based searches over many heterogeneous, independently produced and dynamic data sources, to obtain useful, interconnecting and traceable information. We propose to demonstrate C onnection L ens , a system based on a novel algorithm for keyword search across heterogeneous data sources. Our demonstration scenarios are based on use cases suggested by journalists from the french journal Le Monde, with whom we collaborate.
Camille Chanial, Rédouane Dziri, Helena Galhardas, Julien Leblay, Minh-Huong Le Nguyen, Ioana Manolescu
Proc. VLDB Endow.3
2015 A Benchmark for Relation Extraction Kernels
João L. M. Pereira, Helena Galhardas, Bruno Martins 0001
ADBIS2
2015 Learning to Rank Adaptively for Scalable Information Extraction
abstract
Information extraction systems extract structured data from natural language text, to support richer querying and anal-ysis of the data than would be possible over the unstruc-tured text. Unfortunately, information extraction is a com-putationally expensive task, so exhaustively processing all documents of a large collection might be prohibitive. Such exhaustive processing is generally unnecessary, though, be-cause many times only a small set of documents in a collec-tion is useful for a given information extraction task. There-fore, by identifying these useful documents, and not process-ing the rest, we could substantially improve the efficiency and scalability of an extraction task. Existing approaches for identifying such documents often miss useful documents and also lead to the processing of useless documents unnec-essarily, which in turn negatively impacts the quality and efficiency of the extraction process. To address these limita-tions of the state-of-the-art techniques, we propose a prin-cipled, learning-based approach for ranking documents ac-cording to their potential usefulness for an extraction task. Our low-overhead, online learning-to-rank methods exploit the information collected during extraction, as we process new documents and the fine-grained characteristics of the useful documents are revealed. Then, these methods decide when the ranking model should be updated, hence signifi-cantly improving the document ranking quality over time. Our experiments show that our approach achieves higher ac-curacy than the state-of-the-art alternatives. Importantly, our approach is lightweight and efficient, and hence is a sub-stantial step towards scalable information extraction. 1.
Pablo Barrio 0002, Gonçalo Simões, Helena Galhardas, Luis Gravano
EDBT3
2014 Specifying complex correspondences between relational schemas and RDF models for generating customized R2RML mappings
abstract
The W3C RDB2RDF Working Group proposed a standard language to map relational data into RDF triples, called R2RML. However, creating R2RML mappings may sometimes be a difficult task because it involves the creation of views (within the mappings or not) and referring to them in the R2RML mapping. To overcome such difficulty, this paper first proposes algebraic correspondence assertions, which simplify the definition of relational-to-RDF mappings and yet are expressive enough to cover a wide range of mappings. Algebraic correspondence assertions include data-metadata mappings (where data elements in one schema serve as metadata components in the other), mappings containing custom value functions (e.g., data format transformation functions) and union, intersection and difference between tables. Then, the paper shows how to automatically compile algebraic correspondence assertions into R2RML mappings.
Valéria Magalhães Pequeno, Vânia M. P. Vidal, Marco A. Casanova, Luís Eufrasio T. Neto, Helena Galhardas
IDEAS5
2013 When Speed Has a Price: Fast Information Extraction Using Approximate Algorithms
abstract
A wealth of information produced by individuals and organizations is expressed in natural language text. This is a problem since text lacks the explicit structure that is necessary to support rich querying and analysis. Information extraction systems are sophisticated software tools to discover structured information in natural language text. Unfortunately, information extraction is a challenging and time-consuming task. In this paper, we address the limitations of state-of-the-art systems for the optimization of information extraction programs, with the objective of producing efficient extraction executions. Our solution relies on exploiting a wide range of optimization opportunities. For efficiency, we consider a wide spectrum of execution plans, including approximate plans whose results differ in their precision and recall. Our optimizer accounts for these characteristics of the competing execution plans, and uses accurate predictors of their extraction time, recall, and precision. We demonstrate the efficiency and effectiveness of our optimizer through a large-scale experimental evaluation over real-world datasets and multiple extraction tasks and approaches.
Gonçalo Simões, Helena Galhardas, Luis Gravano
Proc. VLDB Endow.2
2011 Support for User Involvement in Data Cleaning
Helena Galhardas, Antónia Lopes, Emanuel Santos
DaWaK1
2010 An Argumentation-based Approach to Database Repair
Emanuel Santos, João Pavão Martins, Helena Galhardas
ECAI3
2010 Subspace tree: high dimensional multimedia indexing with logarithmic temporal complexity
Andreas Wichert, Helena Galhardas
J. Intell. Inf. Syst.4
2007 One-to-many data transformations through data mappers
Paulo Carreira 0001, Helena Galhardas, Antónia Lopes, João Pereira 0002
Data Knowl. Eng.2
2005 Data Mapper: An Operator for Expressing One-to-Many Data Transformations
Paulo Carreira 0001, Helena Galhardas, João Pereira 0002, Antónia Lopes
DaWaK2
2004 Efficient Development of Data Migration Transformations
abstract
In this paper, we present a data migration tool named DATA FUSION. Its main features are: A domain specific language designed to conveniently model complex data transformations; an integrated development environment that assists users on managing complex data transformation projects and an auditing facility that provides relevant information to project managers and external auditors.
Paulo Carreira 0001, Helena Galhardas
SIGMOD Conference2
2001 Declarative Data Cleaning: Language, Model, and Algorithms
Helena Galhardas, Daniela Florescu, Dennis E. Shasha, Eric Simon, Cristian-Augustin Saita
VLDB1
2000 An Extensible Framework for Data Cleaning
abstract
Projet CARAVEL
Helena Galhardas, Daniela Florescu, Dennis E. Shasha, Eric Simon
ICDE1
2000 AJAX: An Extensible Data Cleaning Tool
abstract
@@@@ groups together matching pairs with a high similarity value by applying a given grouping criteria (e.g. by transitive closure). Finally, ging collapses each individual cluster into a tuple of the resulting data source. AJAX provides @@@@ for specifying data cleaning programs, which consists of SQL statements enriched with a set of specific primitives to express these transformations.
Helena Galhardas, Daniela Florescu, Dennis E. Shasha, Eric Simon
SIGMOD Conference1