VLDB 2026 Research / reviewers in the wild / expert
Elöd Egyed-Zsigmond
dblp:e/ElodEgyedZsigmond
· DBLP profile ↗
14ranked-venue papers in the field
0as first author
7since 2021 · last 2026
0000-0002-1218-8026ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 9Data Mining & Knowledge Discovery · 3Database Systems & Data Management · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Unified Approach for Sexism Detection in Social Media Memes Under Hard and Soft Evaluation Settings
Lorenzo Calogiuri, Elöd Egyed-Zsigmond, Nathan Nowakowski, Luca Cagliero |
ICWE | 2 |
| 2025 | Instruct-to-SPARQL: A text-to-SPARQL dataset for training SPARQL AgentsabstractThe rapid adoption of Large Language Models (LLMs) for search engines and fact-checking platforms necessitates enhancing their output accuracy.Retrieval Augmented Generation (RAG) mitigates hallucinations but requires semantically rich repositories like Wikidata.However, there is a lack of high-quality data to fine-tune LLMs for querying such knowledge bases.To address this gap, we propose a curated dataset with 2,771 unique queries for fine-tuning LLMs to generate accurate and syntactically valid SPARQL queries from natural language instructions.This dataset, customized for interaction with Wikidata, also serves as a robust benchmark for text-to-SPARQL task evaluation.Key findings show that models generally perform better on queries with lower complexity. Mehdi Ben Amor, Alexis Strappazzon, Michael Granitzer, Elöd Egyed-Zsigmond, Jelena Mitrovic |
CHIIR | 4 |
| 2025 | Measuring temporal gains in assisted document transcriptionabstractTranscribing ancient manuscripts is a very time consuming task mainly done by specialists like historians and paleographers. In order to minimize time spent on transcribing, computing tools can be used to assist researchers in this process, but, to our knowledge, there are no studies that evaluate and quantify precisely their actual benefits in terms of human time spent on the transcription. This paper presents a study on quantifying the temporal efficiency of different transcription and segmentation workflows. The main contribution of our work is the extension of a popular existing open-source text transcription web platform: eScriptorium with tools and methods to measure time spent on the transcription task. We explore and compare the efficiency of three workflows: fully manual, using a default annotation model, and using a finetuned model for both ancient manuscript segmentation and transcription. In our experiments, we aim to observe and compare the temporal gains achieved by each workflow, highlighting the trade-offs between manual and automated processing of transcriptions. The paper describes the design of the tracing layer, presents the methodology and the results. This work is conducted as part of the ChEDiL1 French ANR project. Shad Mohammad, Elöd Egyed-Zsigmond, Frank Lebourgeois, Michiel Streijger, Michela Bussotti, Luis Tovar Pimentel, Vincent Paillusson |
DocEng | 2 |
| 2023 | Linked-DocRED - Enhancing DocRED with Entity-Linking to Evaluate End-To-End Document-Level Information Extraction PipelinesabstractInformation Extraction (IE) pipelines aim to extract meaningful entities and relations from documents and structure them into a knowledge graph that can then be used in downstream applications. Training and evaluating such pipelines requires a dataset annotated with entities, coreferences, relations, and entity-linking. However, existing datasets either lack entity-linking labels, are too small, not diverse enough, or automatically annotated (that is, without a strong guarantee of the correction of annotations). Therefore, we propose Linked-DocRED, to the best of our knowledge, the first manually-annotated, large-scale, document-level IE dataset. We enhance the existing and widely-used DocRED dataset with entity-linking labels that are generated thanks to a semi-automatic process that guarantees high-quality annotations. In particular, we use hyperlinks in Wikipedia articles to provide disambiguation candidates. We also propose a complete framework of metrics to benchmark end-to-end IE pipelines, and we define an entity-centric metric to evaluate entity-linking. The evaluation of a baseline shows promising results while highlighting the challenges of an end-to-end IE pipeline. Linked-DocRED, the source code for the entity-linking, the baseline, and the metrics are distributed under an open-source license and can be downloaded from a public repository. Pierre-Yves Genest, Pierre-Edouard Portier, Elöd Egyed-Zsigmond, Martino Lovisetto |
SIGIR | 3 |
| 2022 | PromptORE - A Novel Approach Towards Fully Unsupervised Relation ExtractionabstractUnsupervised Relation Extraction (RE) aims to identify relations between entities in text, without having access to labeled data during training. This setting is particularly relevant for domain specific RE where no annotated dataset is available and for open-domain RE where the types of relations are a priori unknown. Although recent approaches achieve promising results, they heavily depend on hyperparameters whose tuning would most often require labeled data. To mitigate the reliance on hyperparameters, we propose PromptORE, a "Prompt-based Open Relation Extraction" model. We adapt the novel prompt-tuning paradigm to work in an unsupervised setting, and use it to embed sentences expressing a relation. We then cluster these embeddings to discover candidate relations, and we experiment different strategies to automatically estimate an adequate number of clusters. To the best of our knowledge, PromptORE is the first unsupervised RE model that does not need hyperparameter tuning. Results on three general and specific domain datasets show that PromptORE consistently outperforms state-of-the-art models with a relative gain of more than 40% in B3, V-measure and ARI. Qualitative analysis also indicates PromptORE's ability to identify semantically coherent clusters that are very close to true relations. Pierre-Yves Genest, Pierre-Edouard Portier, Elöd Egyed-Zsigmond, Laurent-Walter Goix |
CIKM | 3 |
| 2022 | French translation of a dialogue dataset and text-based emotion detection
Pierre-Yves Genest, Laurent-Walter Goix, Yasser Khalafaoui, Elöd Egyed-Zsigmond, Nistor Grozavu |
Data Knowl. Eng. | 4 |
| 2021 | A Stacking Approach for Cross-Domain Argument Identification
Alaa Alhamzeh, Mohamed Bouhaouel, Elöd Egyed-Zsigmond, Jelena Mitrovic, Lionel Brunie, Harald Kosch |
DEXA (1) | 3 |
| 2019 | An End-User Pipeline for Scraping and Visualizing Semi-Structured Data over the Web
Gabriela Alejandra Bosetti, Sergio Firmenich, Marco Winckler, Gustavo Rossi, Ulises Cornejo Fandos, Elöd Egyed-Zsigmond |
ICWE | 6 |
| 2015 | I-Louvain: An Attributed Graph Clustering Method
David Combe, Christine Largeron, Mathias Géry, Elöd Egyed-Zsigmond |
IDA | 4 |
| 2012 | Getting Clusters from Structure Data and Attribute DataabstractIf the clustering task is widely studied both in graph clustering and in non supervised learning, combined clustering which exploits simultaneously the relationships between the vertices and attributes describing them, is quite new. In this paper, we present different scenarios for this task and, we evaluate their performances and their results on a dataset, with ground truth, built from several sources and containing a scientific social network in which textual data is associated to each vertex and the classes are known. We argue that, depending on the kind of data we have and the type of results we want, the choice of the clustering method is important and we present some concrete examples for underlining this. David Combe, Christine Largeron, Elöd Egyed-Zsigmond, Mathias Géry |
ASONAM | 3 |
| 2012 | Combining Relations and Text in Scientific Network ClusteringabstractIn this paper, we present different combined clustering methods and we evaluate their performances and their results on a dataset with ground truth. This dataset, built from several sources, contains a scientific social network in which textual data is associated to each vertex and the classes are known. Indeed, while the clustering task is widely studied both in graph clustering and in non supervised learning, combined clustering which exploits simultaneously the relationships between the vertices and attributes describing them, is quite new. We argue that, depending on the kind of data we have and the type of results we want, the choice of the clustering method is important and we present some concrete examples for underlining this. David Combe, Christine Largeron, Elöd Egyed-Zsigmond, Mathias Géry |
ASONAM | 3 |
| 2012 | Geo-based automatic image annotationabstractA huge number of user-tagged images are daily uploaded to the web. Recently, a growing number of those images are also geotagged. These provide new opportunities for solutions to automatically tag images so that efficient image management and retrieval can be achieved. In this paper an automatic image annotation approach is proposed. It is based on a statistical model that combines two different kinds of information: high level information represented by user tags of images captured in the same location as a new unlabeled image (input image); and low level information represented by the visual similarity between the input image and the collection of geographically similar images. To maximize the number of images that are visually similar to the input image, an iterative visual matching approach is proposed and evaluated. The results show that a significant recall improvement can be achieved with an increasing number of iterations. The quality of the recommended tags has also been evaluated and an overall good performance has been observed. Hatem Mousselly Sergieh, Gabriele Gianini, Mario Döller, Harald Kosch, Elöd Egyed-Zsigmond, Jean-Marie Pinon |
ICMR | 5 |
| 2012 | Modeling, encoding and querying multi-structured documents
Pierre-Edouard Portier, Noureddine Chatti, Sylvie Calabretto, Elöd Egyed-Zsigmond, Jean-Marie Pinon |
Inf. Process. Manag. | 4 |
| 2008 | Online ancient documents: ArmariusabstractMany museums and libraries digitize their collections of historical manuscripts to preserve the historic documents and to facilitate their browsing. The collections are available as digital images and they need annotation to be accessible and exploitable. The annotations can be created manually, automatically or semi-automatically. Manual annotation is expensive and tedious; hence the reuse of users' experiences, by tracing their actions during the annotation process, helps other users to accomplish repetitive tasks in a semi-automatic manner. In this article we present a digital archive model and prototype of a collaborative system for the management of online ancient manuscript. The application offers an online annotation service, an assistant for semi-automatic annotation, and a tracing system that saves traces of important actions in order to reuse them in a recommender system afterward. Reim Doumat, Elöd Egyed-Zsigmond, Jean-Marie Pinon, Emese Csiszar |
ACM Symposium on Document Engineering | 2 |