VLDB 2026 Research / reviewers in the wild / expert
Sérgio Nunes 0001
dblp:21/5534-1
· DBLP profile ↗
15ranked-venue papers in the field
3as first author
6since 2021 · last 2026
0000-0002-2693-988XORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 15 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CitiLink-Minutes: A Multilayer Annotated Dataset of Municipal Meeting Minutes
Ricardo Campos 0001, Ana Filipa Pacheco, Ana Luísa Fernandes, Inês Cantante, Rute Rebouças, Luís Filipe Cunha, José Isidro, José Pedro Evans, Miguel Marques, Rodrigo Batista, Evelin Amorim, Alípio Mário Jorge, Nuno Guimarães, Sérgio Nunes 0001, Antonio Leal-Millán, Purificação Silvano |
ECIR (4) | 14 |
| 2026 | ClaimPT: A Portuguese Dataset of Annotated Claims in News Articles
Ricardo Campos 0001, Raquel Sequeira, Sara Nerea, Inês Cantante, Diogo Folques, Luís Filipe Cunha, João Canavilhas, António Branco, Alípio Mário Jorge, Sérgio Nunes 0001, Nuno Guimarães, Purificação Silvano |
ECIR (4) | 10 |
| 2026 | CitiLink: Enhancing Municipal Transparency and Citizen Engagement Through Searchable Meeting Minutes
José Pedro Evans, José Isidro, Miguel Marques, Afonso Fonseca, Ricardo Morais, João Canavilhas, Arian Pasquali, Purificação Silvano, Alípio Mário Jorge, Nuno Guimarães, Sérgio Nunes 0001, Ricardo Campos 0001 |
ECIR (4) | 12 |
| 2026 | Labadain Chat: A Conversational Agent for the Tetun LanguageabstractLarge language model (LLM)-based conversational assistants are designed for general-purpose conversation tasks and are primarily optimized for high-resource languages. Although these systems support some low-resource languages (LRLs), their responses often fall short of user expectations. Consequently, speakers of LRLs remain marginalized and unable to fully benefit from advances in LLMs. These challenges underscore the need for targeted, language-specific solutions that can effectively serve underrepresented language communities. This study presents Labadain Chat, a conversational agent for Tetun, a low-resource language spoken by over 932,000 people in Timor-Leste. We adapt existing LLMs to Tetun using language-specific prompting strategies and report on the system's architecture, features, applications, and utility for the Tetun-speaking community. Results from the user study show a high task success rate for Labadain Chat (91%, with substantial inter-annotator agreement, Cohen's κ=0.67) and high user satisfaction (4.30 out of 5, with Cohen's weighted κ=0.75), demonstrating the effectiveness of language-specific LLM customization for Tetun. Overall, this study provides a practical pathway toward promoting equitable access to AI-powered information services for the Tetun-speaking community and suggests an adaptable methodology that can be applied to other under-resourced languages in similar contexts. The system is publicly available at https://www.labadain.com, with mobile applications for both iOS and Android. Gabriel de Jesus, Sérgio Nunes 0001 |
SIGIR | 2 |
| 2026 | CitiLink-Summ: A Dataset of Discussion Subjects Summaries in European Portuguese Municipal Meeting MinutesabstractMunicipal meeting minutes are formal records documenting the discussions and decisions of local government, yet their content is often lengthy, dense, and difficult for citizens to navigate. Automatic summarization can help address this challenge by producing concise summaries for each discussion subject. Despite its potential, research on summarizing discussion subjects in municipal meeting minutes remains largely unexplored, especially in low-resource languages, where the inherent complexity of these documents adds further challenges. A major bottleneck is the scarcity of datasets containing high-quality, manually crafted summaries, which limits the development and evaluation of effective summarization models for this domain. In this paper, we present CitiLink-Summ, a new corpus of European Portuguese municipal meeting minutes, comprising 120 documents and 2,880 manually hand-written summaries, each corresponding to a distinct discussion subject. Leveraging this dataset, we establish baseline results for automatic summarization in this domain, employing state-of-the-art generative models (e.g., BART, PRIMERA) as well as large language models (LLMs), evaluated with both lexical and semantic metrics such as ROUGE, BLEU, METEOR, and BERTScore. CitiLink-Summ provides the first benchmark for municipal-domain summarization in European Portuguese, offering a valuable resource for advancing NLP research on complex administrative texts. Miguel Marques, Ana Luísa Fernandes, Ana Filipa Pacheco, Rute Rebouças, Inês Cantante, José Isidro, Luís Filipe Cunha, Alípio Mário Jorge, Nuno Guimarães, Sérgio Nunes 0001, António Leal, Purificação Silvano, Ricardo Campos 0001 |
WWW | 10 |
| 2023 | From ISAD(G) to Linked Data Archival Descriptions
Inês Koch, Catarina Pires, Carla Teixeira Lopes, Cristina Ribeiro 0001, Sérgio Nunes 0001 |
TPDL | 5 |
| 2020 | Army ANT: A Workbench for Innovation in Entity-Oriented Search
José Luís Devezas, Sérgio Nunes 0001 |
ECIR (2) | 2 |
| 2019 | Information Processing & Management Journal Special Issue on Narrative Extraction from Texts (Text2Story): Preface
Alípio Mário Jorge, Ricardo Campos 0001, Adam Jatowt, Sérgio Nunes 0001 |
Inf. Process. Manag. | 4 |
| 2015 | Summarization of changes in dynamic text collections using Latent Dirichlet Allocation model
Manika Kar, Sérgio Nunes 0001, Cristina Ribeiro 0001 |
Inf. Process. Manag. | 2 |
| 2012 | Studying a Personality Coreference Network in a News Stories Photo Collection
José Luís Devezas, Filipe Coelho, Sérgio Nunes 0001, Cristina Ribeiro 0001 |
ECIR | 3 |
| 2011 | Using the H-Index to Estimate Blog Authority
José Luís Devezas, Sérgio Nunes 0001, Cristina Ribeiro 0001 |
ICWSM | 2 |
| 2011 | Term weighting based on document revision historyabstractIn real-world information retrieval systems, the underlying document collection is rarely stable or definitive. This work is focused on the study of signals extracted from the content of documents at different points in time for the purpose of weighting individual terms in a document. The basic idea behind our proposals is that terms that have existed for a longer time in a document should have a greater weight. We propose 4 term weighting functions that use each document's history to estimate a current term score. To evaluate this thesis, we conduct 3 independent experiments using a collection of documents sampled from Wikipedia. In the first experiment, we use data from Wikipedia to judge each set of terms. In a second experiment, we use an external collection of tags from a popular social bookmarking service as a gold standard. In the third experiment, we crowdsource user judgments to collect feedback on term preference. Across all experiments results consistently support our thesis. We show that temporally aware measures, specifically the proposed revision term frequency and revision term frequency span, outperform a term-weighting measure based on raw term frequency alone. Sérgio Nunes 0001, Cristina Ribeiro 0001, Gabriel David |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | Term frequency dynamics in collaborative articlesabstractDocuments on the World Wide Web are dynamic entities. Mainstream information retrieval systems and techniques are primarily focused on the latest version a document, generally ignoring its evolution over time. In this work, we study the term frequency dynamics in web documents over their lifespan. We use the Wikipedia as a document collection because it is a broad and public resource and, more important, because it provides access to the complete revision history of each document. We investigate the progression of similarity values over two projection variables, namely revision order and revision date. Based on this investigation we find that term frequency in encyclopedic documents - i.e. comprehensive and focused on a single topic - exhibits a rapid and steady progression towards the document's current version. The content in early versions quickly becomes very similar to the present version of the document. Sérgio Nunes 0001, Cristina Ribeiro 0001, Gabriel David |
ACM Symposium on Document Engineering | 1 |
| 2009 | Characterizing the Portuguese Blogosphere
Telmo Couto, Cristina Ribeiro 0001, Sérgio Nunes 0001 |
ICWSM | 3 |
| 2008 | Use of Temporal Expressions in Web Search
Sérgio Nunes 0001, Cristina Ribeiro 0001, Gabriel David |
ECIR | 1 |