VLDB 2026 Research / reviewers in the wild / expert
Javier Osorio
dblp:40/5866
· DBLP profile ↗
6ranked-venue papers in the field
0as first author
5since 2021 · last 2025
0000-0002-4548-7794ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 4Data Mining & Knowledge Discovery · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Conflict Event Actor Prediction Using Spatial Graph Neural Networks
Niamat Zawad, Patrick T. Brandt, Latifur Khan, Vito D'Orazio, Javier Osorio |
IEEE Big Data | 5 |
| 2024 | ConfliLPC: Logits and Parameter Calibration for Political Conflict Analysis in Continual LearningabstractThe ConfliLPC framework introduces an innovative integration of Logits and Parameter Calibration (LPC) with the ConfliBERT model, tailored specifically for the nuanced analysis of political conflict and violence. This paper details the development and application of ConfliLPC, highlighting its robust capability to adapt to evolving data landscapes without succumbing to catastrophic forgetting (CF), a common challenge in machine learning models applied to dynamic domains such as political science. ConfliLPC enhances accuracy and adaptability by continually adjusting its parameters to accommodate new information while retaining valuable historical insights. The framework has been rigorously tested across various conflict scenarios, demonstrating superior performance in real-time analysis and predictive tasks. This work serves as a significant contribution to the fields of political science, conflict research, and applied machine learning, providing a powerful tool for analysts and policymakers engaged in the understanding and resolution of political conflict. The experimental results highlight the efficiency of the ConfliLPC method and its capability to minimize CF. Our code is publicly available1 Xiaodi Li 0002, Niamat Zawad, Patrick T. Brandt, Javier Osorio, Vito D'Orazio, Latifur Khan |
IEEE Big Data | 4 |
| 2022 | Confli-T5: An AutoPrompt Pipeline for Conflict Related Text AugmentationabstractRecent advances in natural language processing (NLP) and Big Data technologies have been crucial for scientists to analyze political unrest and violence, prevent harm, and promote global conflict management. Government agencies and public security organizations have invested heavily in deep learning-based applications to study global conflicts and political violence. However, such applications involving text classification, information extraction, and other NLP-related tasks require extensive human efforts in annotating/labeling texts. While limited labeled data may drastically hurt the models’ performance (over-fitting), large demands on annotation tasks may turn real-world applications impracticable. To address this problem, we propose Confli-T5, a prompt-based method that leverages the domain knowledge from existing political science ontology to generate synthetic but realistic labeled text samples in the conflict and mediation domain. Our model allows generating textual data from the ground up and employs our novel Double Random Sampling mechanism to improve the quality (coherency and consistency) of the generated samples. We conduct experiments over six standard datasets relevant to political science studies to show the superiority of Confli-T5. Our codes are publicly available1. Erick Skorupa Parolin, Yibo Hu 0002, Latifur Khan, Patrick T. Brandt, Javier Osorio, Vito D'Orazio |
IEEE Big Data | 5 |
| 2021 | CoMe-KE: A New Transformers Based Approach for Knowledge Extraction in Conflict and Mediation DomainabstractKnowledge discovery and extraction approaches attract special attention across industries and areas moving toward the 5V Era. In the political and social sciences, scholars and governments dedicate considerable resources to develop intelligent systems for monitoring, analyzing and predicting conflicts and affairs involving political entities across the globe. Such systems rely on background knowledge from external knowledge bases, that conflict experts commonly maintain manually. The high costs and extensive human efforts associated with updating and extending these repositories often compromise their correctness of. Here we introduce CoMe-KE (Conflict and Mediation Knowledge Extractor) to extend automatically knowledge bases about conflict and mediation events. We explore state-of-the-art natural language models to discover new political entities, their roles and status from news. We propose a distant supervised method and propose an innovative zero-shot approach based on a dynamic hypothesis procedure. Our methods leverage pre-trained models through transfer learning techniques to obtain excellent results with no need for a labeled data. Finally, we demonstrate the superiority of our method through a comprehensive set of experiments involving two study cases in the social sciences domain. CoMe-KE significantly outperforms the existing baseline, with (on average) double of the performance retrieving new political entities. Erick Skorupa Parolin, Yibo Hu 0002, Latifur Khan, Javier Osorio, Patrick T. Brandt, Vito D'Orazio |
IEEE BigData | 4 |
| 2021 | 3M-Transformers for Event Coding on Organized Crime DomainabstractPolitical scientists and security agencies increasingly rely on computerized event data generation to track conflict processes and violence around the world. However, most of these approaches rely on pattern-matching techniques constrained by large dictionaries that are too costly to develop, update, or expand to emerging domains or additional languages. In this paper, we provide an effective solution to those challenges. Here we develop the 3M-Transformers (Multilingual, Multi-label, Multitask) approach for Event Coding from domain specific multilingual corpora, dispensing external large repositories for such task, and expanding the substantive focus of analysis to organized crime, an emerging concern for security research. Our results indicate that our 3M-Transformers configurations outperform state-of-the-art usual Transformers models (BERT and XLM-RoBERTa) for coding events on actors, actions and locations in English, Spanish, and Portuguese languages. Erick Skorupa Parolin, Latifur Khan, Javier Osorio, Patrick T. Brandt, Vito D'Orazio, Jennifer S. Holmes |
DSAA | 3 |
| 2020 | HANKE: Hierarchical Attention Networks for Knowledge Extraction in Political Science DomainabstractExtracting structured metadata from unstructured text in different domains is gaining strong attention from multiple research communities. In Political Science, these metadata play a significant role on studying intra and inter-state interactions between political entities. The process of extracting such metadata usually relies on domain specific ontologies and knowledge-based repositories. In particular, Political Scientists regularly use the well-defined ontology CAMEO, which is designed for capturing conflict and mediation relations. Since CAMEO repositories are currently human maintained, the high cost and extensive human effort associated with updating them makes it difficult to include new entries on a regular basis. This paper introduces HANKE: an innovative framework for automatically extracting knowledge representations from unstructured sources, in order to extend CAMEO ontology both in the same domain and towards other related domains in political science. HANKE combines Hierarchical Attention Networks as engine for identifying relevant structures in raw-text and the novel Frequency-Based Ranker approach to obtain a collection of candidate entries for CAMEO's repositories. To show the efficiency of the proposed framework, we evaluate its performance on capturing existing CAMEO representations in a soft-labelled dataset. We also empirically demonstrate the versatility and superiority of HANKE method by applying it to two case studies related to CAMEO extension on its actual domain and towards organized crime domain. Erick Skorupa Parolin, Latifur Khan, Javier Osorio, Vito D'Orazio, Patrick T. Brandt, Jennifer S. Holmes |
DSAA | 3 |