Sérgio Nunes 0001

dblp:21/5534-1 · DBLP profile ↗
← Back
15ranked-venue papers in the field
3as first author
6since 2021 · last 2026
0000-0002-2693-988XORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 15 (3 first)
YearPublicationVenuePosition
2026 CitiLink-Minutes: A Multilayer Annotated Dataset of Municipal Meeting Minutes
Ricardo Campos 0001, Ana Filipa Pacheco, Ana Luísa Fernandes, Inês Cantante, Rute Rebouças, Luís Filipe Cunha, José Isidro, José Pedro Evans, Miguel Marques, Rodrigo Batista, Evelin Amorim, Alípio Mário Jorge, Nuno Guimarães, Sérgio Nunes 0001, Antonio Leal-Millán, Purificação Silvano
ECIR (4)14
2026 ClaimPT: A Portuguese Dataset of Annotated Claims in News Articles
Ricardo Campos 0001, Raquel Sequeira, Sara Nerea, Inês Cantante, Diogo Folques, Luís Filipe Cunha, João Canavilhas, António Branco, Alípio Mário Jorge, Sérgio Nunes 0001, Nuno Guimarães, Purificação Silvano
ECIR (4)10
2026 CitiLink: Enhancing Municipal Transparency and Citizen Engagement Through Searchable Meeting Minutes
José Pedro Evans, José Isidro, Miguel Marques, Afonso Fonseca, Ricardo Morais, João Canavilhas, Arian Pasquali, Purificação Silvano, Alípio Mário Jorge, Nuno Guimarães, Sérgio Nunes 0001, Ricardo Campos 0001
ECIR (4)12
2026 Labadain Chat: A Conversational Agent for the Tetun Language
abstract
Large language model (LLM)-based conversational assistants are designed for general-purpose conversation tasks and are primarily optimized for high-resource languages. Although these systems support some low-resource languages (LRLs), their responses often fall short of user expectations. Consequently, speakers of LRLs remain marginalized and unable to fully benefit from advances in LLMs. These challenges underscore the need for targeted, language-specific solutions that can effectively serve underrepresented language communities. This study presents Labadain Chat, a conversational agent for Tetun, a low-resource language spoken by over 932,000 people in Timor-Leste. We adapt existing LLMs to Tetun using language-specific prompting strategies and report on the system's architecture, features, applications, and utility for the Tetun-speaking community. Results from the user study show a high task success rate for Labadain Chat (91%, with substantial inter-annotator agreement, Cohen's κ=0.67) and high user satisfaction (4.30 out of 5, with Cohen's weighted κ=0.75), demonstrating the effectiveness of language-specific LLM customization for Tetun. Overall, this study provides a practical pathway toward promoting equitable access to AI-powered information services for the Tetun-speaking community and suggests an adaptable methodology that can be applied to other under-resourced languages in similar contexts. The system is publicly available at https://www.labadain.com, with mobile applications for both iOS and Android.
Gabriel de Jesus, Sérgio Nunes 0001
SIGIR2
2026 CitiLink-Summ: A Dataset of Discussion Subjects Summaries in European Portuguese Municipal Meeting Minutes
abstract
Municipal meeting minutes are formal records documenting the discussions and decisions of local government, yet their content is often lengthy, dense, and difficult for citizens to navigate. Automatic summarization can help address this challenge by producing concise summaries for each discussion subject. Despite its potential, research on summarizing discussion subjects in municipal meeting minutes remains largely unexplored, especially in low-resource languages, where the inherent complexity of these documents adds further challenges. A major bottleneck is the scarcity of datasets containing high-quality, manually crafted summaries, which limits the development and evaluation of effective summarization models for this domain. In this paper, we present CitiLink-Summ, a new corpus of European Portuguese municipal meeting minutes, comprising 120 documents and 2,880 manually hand-written summaries, each corresponding to a distinct discussion subject. Leveraging this dataset, we establish baseline results for automatic summarization in this domain, employing state-of-the-art generative models (e.g., BART, PRIMERA) as well as large language models (LLMs), evaluated with both lexical and semantic metrics such as ROUGE, BLEU, METEOR, and BERTScore. CitiLink-Summ provides the first benchmark for municipal-domain summarization in European Portuguese, offering a valuable resource for advancing NLP research on complex administrative texts.
Miguel Marques, Ana Luísa Fernandes, Ana Filipa Pacheco, Rute Rebouças, Inês Cantante, José Isidro, Luís Filipe Cunha, Alípio Mário Jorge, Nuno Guimarães, Sérgio Nunes 0001, António Leal, Purificação Silvano, Ricardo Campos 0001
WWW10
2023 From ISAD(G) to Linked Data Archival Descriptions
Inês Koch, Catarina Pires, Carla Teixeira Lopes, Cristina Ribeiro 0001, Sérgio Nunes 0001
TPDL5
2020 Army ANT: A Workbench for Innovation in Entity-Oriented Search
José Luís Devezas, Sérgio Nunes 0001
ECIR (2)2
2019 Information Processing & Management Journal Special Issue on Narrative Extraction from Texts (Text2Story): Preface
Alípio Mário Jorge, Ricardo Campos 0001, Adam Jatowt, Sérgio Nunes 0001
Inf. Process. Manag.4
2015 Summarization of changes in dynamic text collections using Latent Dirichlet Allocation model
Manika Kar, Sérgio Nunes 0001, Cristina Ribeiro 0001
Inf. Process. Manag.2
2012 Studying a Personality Coreference Network in a News Stories Photo Collection
José Luís Devezas, Filipe Coelho, Sérgio Nunes 0001, Cristina Ribeiro 0001
ECIR3
2011 Using the H-Index to Estimate Blog Authority
José Luís Devezas, Sérgio Nunes 0001, Cristina Ribeiro 0001
ICWSM2
2011 Term weighting based on document revision history
abstract
In real-world information retrieval systems, the underlying document collection is rarely stable or definitive. This work is focused on the study of signals extracted from the content of documents at different points in time for the purpose of weighting individual terms in a document. The basic idea behind our proposals is that terms that have existed for a longer time in a document should have a greater weight. We propose 4 term weighting functions that use each document's history to estimate a current term score. To evaluate this thesis, we conduct 3 independent experiments using a collection of documents sampled from Wikipedia. In the first experiment, we use data from Wikipedia to judge each set of terms. In a second experiment, we use an external collection of tags from a popular social bookmarking service as a gold standard. In the third experiment, we crowdsource user judgments to collect feedback on term preference. Across all experiments results consistently support our thesis. We show that temporally aware measures, specifically the proposed revision term frequency and revision term frequency span, outperform a term-weighting measure based on raw term frequency alone.
Sérgio Nunes 0001, Cristina Ribeiro 0001, Gabriel David
J. Assoc. Inf. Sci. Technol.1
2010 Term frequency dynamics in collaborative articles
abstract
Documents on the World Wide Web are dynamic entities. Mainstream information retrieval systems and techniques are primarily focused on the latest version a document, generally ignoring its evolution over time. In this work, we study the term frequency dynamics in web documents over their lifespan. We use the Wikipedia as a document collection because it is a broad and public resource and, more important, because it provides access to the complete revision history of each document. We investigate the progression of similarity values over two projection variables, namely revision order and revision date. Based on this investigation we find that term frequency in encyclopedic documents - i.e. comprehensive and focused on a single topic - exhibits a rapid and steady progression towards the document's current version. The content in early versions quickly becomes very similar to the present version of the document.
Sérgio Nunes 0001, Cristina Ribeiro 0001, Gabriel David
ACM Symposium on Document Engineering1
2009 Characterizing the Portuguese Blogosphere
Telmo Couto, Cristina Ribeiro 0001, Sérgio Nunes 0001
ICWSM3
2008 Use of Temporal Expressions in Web Search
Sérgio Nunes 0001, Cristina Ribeiro 0001, Gabriel David
ECIR1