VLDB 2026 Research / reviewers in the wild / expert
Cinthia Sánchez
dblp:297/0492
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0001-5977-3559ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large Language Models in Crisis Informatics for Zero and Few-Shot ClassificationabstractThis article presents an exploration of the use of pre-trained Large Language Models (LLMs) for crisis classification to address labeled data dependency issues. We present a methodology that enhances open LLMs through fine-tuning, creating zero-shot and few-shot classifiers that approach traditional supervised models in classifying crisis-related messages. A comparative study evaluates crisis classification tasks using general domain pre-trained LLMs, crisis-specific LLMs, and traditional supervised learning methods, establishing a benchmark in the field. Our task-specific fine-tuned Llama model achieved a 69% macro F1 score in classifying humanitarian information–a remarkable 26% improvement compared to the Llama baseline, even with limited training data. Moreover, it outperformed ChatGPT4 by 3% in macro F1. This improvement increased to 71% macro F1 when fine-tuning Llama with multitask data. For the binary classification of messages as related vs. not related to crises, we observed that pre-trained LLMs, such as Llama 2 and ChatGPT4, performed well without fine-tuning, achieving an 87% macro F1 score with ChatGPT4. This research expands our knowledge of how to exploit the potential of LLMs for crisis classification, representing a great opportunity for crisis scenarios that lack labeled data. The findings emphasize the potential of LLMs in crisis informatics to address cold start challenges, especially critical in the initial phases of a disaster, while also showcasing their capacity to attain high accuracy even with limited training data. Cinthia Sánchez, Andrés Abeliuk, Barbara Poblete |
ACM Trans. Web | 1 |
| 2023 | Cross-Lingual and Cross-Domain Crisis Classification for Low-Resource ScenariosabstractSocial media data has emerged as a useful source of timely information about real-world crisis events. One of the main tasks related to the use of social media for disaster management is the automatic identification of crisis-related messages. Most of the studies on this topic have focused on the analysis of data for a particular type of event in a specific language. This limits the possibility of generalizing existing approaches because models cannot be directly applied to new types of events or other languages. In this work, we study the task of automatically classifying messages that are related to crisis events by leveraging cross-language and cross-domain labeled data. Our goal is to make use of labeled data from high-resource languages to classify messages from other (low-resource) languages and/or of new (previously unseen) types of crisis situations. For our study we consolidated from the literature a large unified dataset containing multiple crisis events and languages. Our empirical findings show that it is indeed possible to leverage data from crisis events in English to classify the same type of event in other languages, such as Spanish and Italian (80.0% F1-score). Furthermore, we achieve good performance for the cross-domain task (80.0% F1-score) in a cross-lingual setting. Overall, our work contributes to improving the data scarcity problem that is so important for multilingual crisis classification. In particular, mitigating cold-start situations in emergency events, when time is of essence. Cinthia Sánchez, Hernan Sarmiento, Andrés Abeliuk, Jorge Pérez 0001, Barbara Poblete |
ICWSM | 1 |
| 2021 | Transfer Learning for the Multilingual and Multi-Domain Classification of Messages Relating to CrisesabstractSocial media plays an important role as a source of information during crisis events. It allows for more rapid dissemination of critical information than traditional news media, as its users can provide immediate information from the locations where events are unfolding. Several studies have addressed the automatic detection of crisis-related messages to contribute to disaster management and humanitarian assistance. However, most of them have focused on a particular language (usually English) or type of event, which limits their applicability to other contexts. The lack of labeled data in different languages and types of disasters poses a major obstacle to the application of supervised learning-based approaches to more diverse scenarios. To address this problem, this research aims to characterize messages related to diverse crisis domains in a language-agnostic manner in order to construct multilingual crisis detectors. To achieve this, we propose a comprehensive evaluation of transfer learning performance in terms of crisis domain and language, comparing different data representations and classification techniques. Cinthia Sánchez |
SIGIR | 1 |