EDBT 2026 Demo / reviewers in the wild / expert
Gaël Dias
dblp:59/1816 · also Gaël Harry Dias
· DBLP profile ↗
26ranked-venue papers in the field
5as first author
4since 2021 · last 2025
0000-0002-5840-1603ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 20 (2 first)Other / Interdisciplinary · 4 (2 first)Data Mining & Knowledge Discovery · 1Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multilingual Evaluation of Main Content Extractors for Web PagesabstractTools designed to extract main content from web pages require thorough evaluation, yet existing benchmarks disproportionately focus on English-language datasets. Consequently, previous studies have shown that while these extractors are well-optimized for English, their effectiveness partially or entirely diminishes in other languages. This study reproduces and extends recent benchmarks by incorporating multilingual datasets as a key factor. We analyze extractor performance across five languages-Greek, English, Polish, Russian, and Chinese-highlighting the need to adapt extraction models to linguistic variations. Our results show that while some extractors maintain stable performance, others suffer significant drops in precision and recall on non-English or structurally irregular pages. Aurélien Bournonville, Gaël Dias, Thomas Largillier, Emmanuel Marchand, Fabrice Maurel, Guillaume Pitel, François Rioult |
SIGIR | 2 |
| 2022 | Multimodal Web Page Segmentation Using Self-organized Multi-objective ClusteringabstractWeb page segmentation (WPS) aims to break a web page into different segments with coherent intra- and inter-semantics. By evidencing the morpho-dispositional semantics of a web page, WPS has traditionally been used to demarcate informative from non-informative content, but it has also evidenced its key role within the context of non-linear access to web information for visually impaired people. For that purpose, a great deal of ad hoc solutions have been proposed that rely on visual, logical, and/or text cues. However, such methodologies highly depend on manually tuned heuristics and are parameter-dependent. To overcome these drawbacks, principled frameworks have been proposed that provide the theoretical bases to achieve optimal solutions. However, existing methodologies only combine few discriminant features and do not define strategies to automatically select the optimal number of segments. In this article, we present a multi-objective clustering technique called MCS that relies on \( K \) -means, in which (1) visual, logical, and text cues are all combined in a early fusion manner and (2) an evolutionary process automatically discovers the optimal number of clusters (segments) as well as the correct positioning of seeds. As such, our proposal is parameter-free, combines many different modalities, does not depend on manually tuned heuristics, and can be run on any web page without any constraint. An exhaustive evaluation over two different tasks, where (1) the number of segments must be discovered or (2) the number of clusters is fixed with respect to the task at hand, shows that MCS drastically improves over most competitive and up-to-date algorithms for a wide variety of external and internal validation indices. In particular, results clearly evidence the impact of the visual and logical modalities towards segmentation performance. Srivatsa Ramesh Jayashree, Gaël Dias, Judith Jeyafreeda Andrew, Sriparna Saha 0001, Fabrice Maurel, Stéphane Ferrari |
ACM Trans. Inf. Syst. | 2 |
| 2021 | Time-Matters: Temporal Unfolding of Texts
Ricardo Campos 0001, Jorge Duque, Tiago Cândido, Gaël Dias, Alípio Mário Jorge, Celia Nunes |
ECIR (2) | 5 |
| 2021 | Improving Neural Text Style Transfer by Introducing Loss Function SequentialityabstractText style transfer is an important issue for conversational agents as it may adapt utterance production to specific dialogue situations. It consists in introducing a given style within a sentence while preserving its semantics. Within this scope, different strategies have been proposed that either rely on parallel data or take advantage of non-supervised techniques. In this paper, we follow the latter approach and show that the sequential introduction of different loss functions into the learning process can boost the performance of a standard model. We also evidence that combining different style classifiers that either focus on global or local textual information improves sentence generation. Experiments on the Yelp dataset show that our methodology strongly competes with the current state-of-the-art models across style accuracy, grammatical correctness, and content preservation. Chinmay Rane, Gaël Dias, Alexis Lechervy, Asif Ekbal |
SIGIR | 2 |
| 2020 | Patch-Based Identification of Lexical Semantic Relations
Nesrine Bannour, Gaël Dias, Youssef Chahir, Houssam Akhmouch |
ECIR (1) | 2 |
| 2019 | Learning Lexical-Semantic Relations Using Intuitive Cognitive Links
Georgios Balikas, Gaël Dias, Rumen Moraliyski, Houssam Akhmouch, Massih-Reza Amini |
ECIR (1) | 2 |
| 2017 | Overview of the 4th HistoInformatics WorkshopabstractIn line with global trends, historical records are increasingly available in forms that computer can process. These ever expanding records (such as scanned books, large-scale corpora, academic papers, maps, photos, audios, videos)---either digitally born or reconstructed through digitization pipelines---are too big to be read or viewed manually. Historians, like other humanities researchers, have a keen interest in computational approaches to process and study digitized historical information for research, writing, and dissemination of historical knowledge. In Computer Science, experimental tools and methods are challenged to be validated regarding their relevance for real-world questions and applications. The HistoInformatics workshop series is focused on the challenges and opportunities of data-driven humanities and brings together scientists and scholars at the forefront of this emerging field, at the interface between History, Anthropology, Archaeology, Computer Science and associated disciplines as well as the cultural heritage sector. The 4th HistoInformatics Workshop was a half day workshop co-located with the 26th ACM International Conference on Information and Knowledge Management (CIKM 2017) in Singapore. Mohammed Hasanuzzaman, Gaël Dias, Adam Jatowt, Marten Düring, Antal van den Bosch |
CIKM | 2 |
| 2017 | Identifying top relevant dates for implicit time sensitive queries
Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes |
Inf. Retr. J. | 2 |
| 2016 | Multi-objective Word Sense Induction Using Content and Interlink Connections
Sudipta Acharya, Asif Ekbal, Sriparna Saha 0001, Prabhakaran Santhanam, José G. Moreno 0001, Gaël Dias |
NLDB | 6 |
| 2016 | GTE-Rank: A time-aware search engine to answer time-sensitive queries
Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes |
Inf. Process. Manag. | 2 |
| 2015 | Learning Pretopological Spaces for Lexical Taxonomy Acquisition
Guillaume Cleuziou, Gaël Dias |
ECML/PKDD (2) | 2 |
| 2015 | Understanding Temporal Query IntentabstractUnderstanding the temporal orientation of web search queries is an important issue for the success of information access systems. In this paper, we propose a multi-objective ensemble learning solution that (1) allows to accurately classify queries along their temporal intent and (2) identifies a set of performing solutions thus offering a wide range of possible applications. Experiments show that correct representation of the problem can lead to great classification improvements when compared to recent state-of-the-art solutions and baseline ensemble techniques. Mohammed Hasanuzzaman, Sriparna Saha 0001, Gaël Dias, Stéphane Ferrari |
SIGIR | 3 |
| 2015 | Adapted B-CUBED Metrics to Unbalanced DatasetsabstractB-CUBED metrics have recently been adopted in the evaluation of clustering results as well as in many other related tasks. However, this family of metrics is not well adapted when datasets are unbalanced. This issue is extremely frequent in Web results, where classes are distributed following a strong unbalanced pattern. In this paper, we present a modified version of B-CUBED metrics to overcome this situation. Results in toy and real datasets indicate that the proposed adaptation correctly considers the particularities of unbalanced cases. José G. Moreno 0001, Gaël Dias |
SIGIR | 2 |
| 2014 | GTE-Rank: Searching for Implicit Temporal Query ResultsabstractTemporal information retrieval has been a topic of great interest in recent years. Despite the efforts that have been conducted so far, most popular search engines remain underdeveloped when it comes to explicitly considering the use of temporal information in their search process. In this paper we present GTE-Rank, an online searching tool that takes time into account when ranking time-sensitive query web search results. GTE-Rank is defined as a linear combination of topical and temporal scores to reflect the relevance of any web page both in topical and temporal dimensions. The resulting system can be explored graphically through a search interface made available for research purposes. Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes |
CIKM | 2 |
| 2014 | GTE-Cluster: A Temporal Search Interface for Implicit Temporal Queries
Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes |
ECIR | 2 |
| 2014 | Query log driven web search results clusteringabstractDifferent important studies in Web search results clustering have recently shown increasing performances motivated by the use of external resources. Following this trend, we present a new algorithm called Dual C-Means, which provides a theoretical background for clustering in different representation spaces. Its originality relies on the fact that external resources can drive the clustering process as well as the labeling task in a single step. To validate our hypotheses, a series of experiments are conducted over different standard datasets and in particular over a new dataset built from the TREC Web Track 2012 to take into account query logs information. The comprehensive empirical evaluation of the proposed approach demonstrates its significant advantages over traditional clustering and labeling techniques. José G. Moreno 0001, Gaël Dias, Guillaume Cleuziou |
SIGIR | 2 |
| 2013 | Using Text-Based Web Image Search Results Clustering to Minimize Mobile Devices Wasted Space-Interface
José G. Moreno 0001, Gaël Dias |
ECIR | 2 |
| 2012 | GTE: a distributional second-order co-occurrence approach to improve the identification of top relevant dates in web snippetsabstractIn this paper, we present an approach to identify top relevant dates in Web snippets with respect to a given implicit temporal query. Our approach is two-fold. First, we propose a generic temporal similarity measure called GTE, which evaluates the temporal similarity between a query and a date. Second, we propose a classification model to accurately relate relevant dates to their corresponding query terms and withdraw irrelevant ones. We suggest two different solutions: a threshold-based classification strategy and a supervised classifier based on a combination of multiple similarity measures. We evaluate both strategies over a set of real-world text queries and compare the performance of our Web snippet approach with a query log approach over the same set of queries. Experiments show that determining the most relevant dates of any given implicit temporal query can be improved with GTE combined with the second order similarity measure InfoSimba, the Dice coefficient and the threshold-based strategy compared to (1) first-order similarity measures and (2) the query log based approach. Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes |
CIKM | 2 |
| 2012 | Temporal Web Image Retrieval
Gaël Dias, José G. Moreno 0001, Adam Jatowt, Ricardo Campos 0001 |
SPIRE | 1 |
| 2012 | Disambiguating Implicit Temporal Queries by Clustering Top Relevant Dates in Web SnippetsabstractWith the growing popularity of research in Temporal Information Retrieval (T-IR), a large amount of temporal data is ready to be exploited. The ability to exploit this information can be potentially useful for several tasks. For example, when querying "Football World Cup Germany", it would be interesting to have two separate clusters {1974,2006} corresponding to each of the two temporal instances. However, clustering of search results by time is a non-trivial task that involves determining the most relevant dates associated to a query. In this paper, we propose a first approach to flat temporal clustering of search results. We rely on a second order co-occurrence similarity measure approach which first identifies top relevant dates. Documents are grouped at the year level, forming the temporal instances of the query. Experimental tests were performed using real-world text queries. We used several measures for evaluating the performance of the system and compared our approach with Carrot Web-snippet clustering engine. Both experiments were complemented with a user survey. Ricardo Campos 0001, Alípio Mário Jorge, Gaël Dias, Celia Nunes |
Web Intelligence | 3 |
| 2011 | A pretopological framework for the automatic construction of lexical-semantic structures from textsabstractWe present in this paper a new approach for the automatic generation of lexical structures from texts. This tedious task is based on the strong hypothesis that simple statistical observations on textual usages can provide pieces of semantics about the lexicon. Using such "naive" observations only, we propose a (pre)-topological framework to formalize and combine various hypothesis on textual data usages and then to derive a structure similar to usual lexical knowledge basis such as WordNet. In addition we also consider the evaluation problem for obtained lexical structures ; a multi-level evaluation strategy is proposed that measures the fitting between a given reference structure and automatically generated structures on different point of views : intrinsic/structural and application-based points of view. The evaluation strategy is then used to quantify the contribution of the new structuring approach with respect to the corresponding solution proposed by (Sanderson et al. 2000) on two case studies that differs on the domain and the size of the lexicon. Guillaume Cleuziou, Davide Buscaldi, Vincent Levorato, Gaël Dias |
CIKM | 4 |
| 2011 | Informative Polythetic Hierarchical Ephemeral ClusteringabstractEphemeral clustering has been studied for more than a decade, although with low user acceptance. According to us, this situation is mainly due to (1) an excessive number of generated clusters, which makes browsing difficult and (2) low quality labeling, which introduces imprecision within the search process. In this paper, our motivation is twofold. First, we propose to reduce the number of clusters of Web page results, but keeping all different query meanings. For that purpose, we propose a new polythetic methodology based on an informative similarity measure, the InfoSimba, and a new hierarchical clustering algorithm, the HISGK-means. Second, a theoretical background is proposed to define meaningful cluster labels embedded in the definition of the HISGK-means algorithm, which may elect as best label, words outside the given cluster. To confirm our intuitions, we propose a new evaluation framework, which shows that we are able to extract most of the important query meanings but generating much less clusters than state-of-the-art systems. Gaël Dias, Guillaume Cleuziou, David Machado |
Web Intelligence | 1 |
| 2011 | Recognizing Textual Entailment by Generality Using Informative Asymmetric Measures and Multiword Unit Identification to Summarize Ephemeral ClustersabstractIn the context of Ephemeral Clustering of web Pages, it can be interesting to label each cluster with a small summary instead of just a label. Within this scope, we introduce the paradigm of Textual Entailment by Generality, which can be defined as the entailment from a specific web snippet towards a more general web snippet. The subjacent idea is to find the best web snippet, which summarizes and subsumes all the other web snippets within an ephemeral cluster. To reach this objective, we first propose a new informative asymmetric similarity measure called the Simplified Asymmetric InfoSimba (AISs), which can be combined with different asymmetric association measures. In particular, the AISs proposes an unsupervised language-independent solution to infer Textual Entailment by Generality and as such can help to encounter the web snippet with maximum semantic coverage. This new methodology is tested against the first Recognizing Textual Entailment data set (RTE-1)1 for an exhaustive number of asymmetric association measures with and without the identification of Multiword Units. The comparative experiments with existing state-of-the-art methodologies show promising results. Gaël Dias, Sebastião Pais, Katarzyna Wegrzyn-Wolska, Robert Mahl |
Web Intelligence | 1 |
| 2009 | High-level Features for Learning Subjective Language across Domains
Gaël Dias, Dinko Dimchev Lambov, Veska Noncheva |
ICWSM | 1 |
| 2008 | Mapping General-Specific Noun Relationships to WordNet Hypernym/Hyponym Relations
Gaël Dias, Raycho Mukelov, Guillaume Cleuziou |
EKAW | 1 |
| 2006 | WISE: Hierarchical Soft Clustering of Web Page Search Results Based on Web Content Mining TechniquesabstractTypically, search engines are low precision in response to a query, retrieving lots of useless Web pages, and missing some other important ones. In this paper, we study the problem of the hierarchical clustering of Web pages search results. In particular, we propose an architecture called WISE, a meta-search engine that automatically builds clusters of related Web pages embodying one meaning of the query. These clusters are then hierarchically organized and labeled with a phrase representing the key concept of the cluster and the corresponding Web documents. The system which is a Web-based interface (soon available at wise.di.ubi.pt), introduces some interesting new ideas, such as the preselection of the retrieved Web pages, the capacity to statistically detect phrases within documents and the representation of documents based on their most relevant key concepts by using Web content mining techniques. The final step of the system is supported by a graph-based overlapping clustering algorithm which groups the selected documents into a hierarchy of clusters Ricardo Campos 0001, Gaël Dias, Celia Nunes |
Web Intelligence | 2 |