Vito D'Orazio

dblp:149/9947 · also Vito J. D'Orazio · DBLP profile ↗
← Back
9ranked-venue papers in the field
1as first author
6since 2021 · last 2025
—ORCID · none

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 6 (1 first)Data Mining & Knowledge Discovery · 2Database Systems & Data Management · 1
YearPublicationVenuePosition
2025 Conflict Event Actor Prediction Using Spatial Graph Neural Networks
Niamat Zawad, Patrick T. Brandt, Latifur Khan, Vito D'Orazio, Javier Osorio
IEEE Big Data4
2024 ConfliLPC: Logits and Parameter Calibration for Political Conflict Analysis in Continual Learning
abstract
The ConfliLPC framework introduces an innovative integration of Logits and Parameter Calibration (LPC) with the ConfliBERT model, tailored specifically for the nuanced analysis of political conflict and violence. This paper details the development and application of ConfliLPC, highlighting its robust capability to adapt to evolving data landscapes without succumbing to catastrophic forgetting (CF), a common challenge in machine learning models applied to dynamic domains such as political science. ConfliLPC enhances accuracy and adaptability by continually adjusting its parameters to accommodate new information while retaining valuable historical insights. The framework has been rigorously tested across various conflict scenarios, demonstrating superior performance in real-time analysis and predictive tasks. This work serves as a significant contribution to the fields of political science, conflict research, and applied machine learning, providing a powerful tool for analysts and policymakers engaged in the understanding and resolution of political conflict. The experimental results highlight the efficiency of the ConfliLPC method and its capability to minimize CF. Our code is publicly available1
Xiaodi Li 0002, Niamat Zawad, Patrick T. Brandt, Javier Osorio, Vito D'Orazio, Latifur Khan
IEEE Big Data5
2022 Confli-T5: An AutoPrompt Pipeline for Conflict Related Text Augmentation
abstract
Recent advances in natural language processing (NLP) and Big Data technologies have been crucial for scientists to analyze political unrest and violence, prevent harm, and promote global conflict management. Government agencies and public security organizations have invested heavily in deep learning-based applications to study global conflicts and political violence. However, such applications involving text classification, information extraction, and other NLP-related tasks require extensive human efforts in annotating/labeling texts. While limited labeled data may drastically hurt the models’ performance (over-fitting), large demands on annotation tasks may turn real-world applications impracticable. To address this problem, we propose Confli-T5, a prompt-based method that leverages the domain knowledge from existing political science ontology to generate synthetic but realistic labeled text samples in the conflict and mediation domain. Our model allows generating textual data from the ground up and employs our novel Double Random Sampling mechanism to improve the quality (coherency and consistency) of the generated samples. We conduct experiments over six standard datasets relevant to political science studies to show the superiority of Confli-T5. Our codes are publicly available1.
Erick Skorupa Parolin, Yibo Hu 0002, Latifur Khan, Patrick T. Brandt, Javier Osorio, Vito D'Orazio
IEEE Big Data6
2021 CoMe-KE: A New Transformers Based Approach for Knowledge Extraction in Conflict and Mediation Domain
abstract
Knowledge discovery and extraction approaches attract special attention across industries and areas moving toward the 5V Era. In the political and social sciences, scholars and governments dedicate considerable resources to develop intelligent systems for monitoring, analyzing and predicting conflicts and affairs involving political entities across the globe. Such systems rely on background knowledge from external knowledge bases, that conflict experts commonly maintain manually. The high costs and extensive human efforts associated with updating and extending these repositories often compromise their correctness of. Here we introduce CoMe-KE (Conflict and Mediation Knowledge Extractor) to extend automatically knowledge bases about conflict and mediation events. We explore state-of-the-art natural language models to discover new political entities, their roles and status from news. We propose a distant supervised method and propose an innovative zero-shot approach based on a dynamic hypothesis procedure. Our methods leverage pre-trained models through transfer learning techniques to obtain excellent results with no need for a labeled data. Finally, we demonstrate the superiority of our method through a comprehensive set of experiments involving two study cases in the social sciences domain. CoMe-KE significantly outperforms the existing baseline, with (on average) double of the performance retrieving new political entities.
Erick Skorupa Parolin, Yibo Hu 0002, Latifur Khan, Javier Osorio, Patrick T. Brandt, Vito D'Orazio
IEEE BigData6
2021 3M-Transformers for Event Coding on Organized Crime Domain
abstract
Political scientists and security agencies increasingly rely on computerized event data generation to track conflict processes and violence around the world. However, most of these approaches rely on pattern-matching techniques constrained by large dictionaries that are too costly to develop, update, or expand to emerging domains or additional languages. In this paper, we provide an effective solution to those challenges. Here we develop the 3M-Transformers (Multilingual, Multi-label, Multitask) approach for Event Coding from domain specific multilingual corpora, dispensing external large repositories for such task, and expanding the substantive focus of analysis to organized crime, an emerging concern for security research. Our results indicate that our 3M-Transformers configurations outperform state-of-the-art usual Transformers models (BERT and XLM-RoBERTa) for coding events on actors, actions and locations in English, Spanish, and Portuguese languages.
Erick Skorupa Parolin, Latifur Khan, Javier Osorio, Patrick T. Brandt, Vito D'Orazio, Jennifer S. Holmes
DSAA5
2021 An Ecosystem of Applications for Modeling Political Violence
abstract
Conflict researchers face many challenges, including (1) how to model conflicts, (2) how to measure them, (3) how to manage their spatio-temporal character, and (4) how to handle a potential abundance of information and explanation. In this paper, we describe an ecosystem of tools designed for use by subject matter experts that addresses these challenges. Three case studies show workflows that are facilitated by this ecosystem.
Aline Bessa, Sonia Castelo Quispe, Rémi Rampin, Aécio S. R. Santos, Michael Shoemate, Vito D'Orazio, Juliana Freire
SIGMOD Conference6
2020 HANKE: Hierarchical Attention Networks for Knowledge Extraction in Political Science Domain
abstract
Extracting structured metadata from unstructured text in different domains is gaining strong attention from multiple research communities. In Political Science, these metadata play a significant role on studying intra and inter-state interactions between political entities. The process of extracting such metadata usually relies on domain specific ontologies and knowledge-based repositories. In particular, Political Scientists regularly use the well-defined ontology CAMEO, which is designed for capturing conflict and mediation relations. Since CAMEO repositories are currently human maintained, the high cost and extensive human effort associated with updating them makes it difficult to include new entries on a regular basis. This paper introduces HANKE: an innovative framework for automatically extracting knowledge representations from unstructured sources, in order to extend CAMEO ontology both in the same domain and towards other related domains in political science. HANKE combines Hierarchical Attention Networks as engine for identifying relevant structures in raw-text and the novel Frequency-Based Ranker approach to obtain a collection of candidate entries for CAMEO's repositories. To show the efficiency of the proposed framework, we evaluate its performance on capturing existing CAMEO representations in a soft-labelled dataset. We also empirically demonstrate the versatility and superiority of HANKE method by applying it to two case studies related to CAMEO extension on its actual domain and towards organized crime domain.
Erick Skorupa Parolin, Latifur Khan, Javier Osorio, Vito D'Orazio, Patrick T. Brandt, Jennifer S. Holmes
DSAA4
2019 Modeling and Forecasting Armed Conflict: AutoML with Human-Guided Machine Learning
abstract
Machine learning has made slow inroads into quantitative social science due to both a mismatch of machine learning's strengths to the causal and inferential tasks domain researchers pursue [1] and also a lack of algorithmic training among many domain experts [2]. However, conflict research- the empirical examination of political unrest, violence and civil war-has seen a growing emphasis on prediction and forecasting models. We describe automated machine learning (AutoML) to identify models, and human-guided machine learning (HGML), and show how these can incorporate domain knowledge and research requirements into model selection and assessment, and provide high quality machine learning pipelines to domain experts comparable to state-of-the-literature solutions. We examine three peer-reviewed papers with predictive models of conflict [3, 4, 5] and run their data through our HGML system using multiple AutoML engines and find this system produces slightly elevated performance on each paper's model, without any ML expertise required of the user. Our research has three takeaways for computational social science. First, predictive models of conflict would benefit from even minimal applications of AutoML; Secondly, human-guided machine learning offers the attractive option of constraining AutoML systems to address the kinds of questions conflict researchers assess with predictive models; Finally, current existing AutoML implementations produce divergent solutions and so can be productively harnessed in parallel.
Vito D'Orazio, James Honaker, Raman Prasady, Michael Shoemate
IEEE BigData1
2017 RePAIR: Recommend political actors in real-time from news websites
abstract
Extracting a structured representation of political events from news reports is at the intersection of the computational and social sciences. A traditional approach is to use dictionary-based pattern lookups to identify actors and actions involved in potential events. A key complication of this approach is updating the dictionaries with new actors (e.g., when a new president takes office). Currently, the dictionaries are curated by humans, updated infrequently, and at a high cost. This means that tools dependent on the actor dictionaries (e.g., PETRARCH) overlook events when actors are missing in the dictionary. Since these tools use only the syntactic structure of the sentence (e.g., parse tree, etc.) for their event coding, missing actors will generate events which fail to capture actual political interaction. To overcome these issues, we propose a framework RePAIR to recommend new political actors in real-time from the political news articles with RSS feeds related to national/international politics across the globe. The framework identifies semantic structure of a sentence using an Automatic Content Extraction (ACE) method and uses a frequency based actor ranking algorithm to recommend the most frequent new political actors over multiple time windows. We also suggest the associated role of recommended new actors from the role of co-occurred political actors in the existing CAMEO actor dictionary. Further we integrate an external knowledge base (e.g., Wikipedia) into our framework to capture the evolving roles of existing actors over time and recommend new roles for them. Furthermore, we consider PETRARCH and BBN ACCENT event coders for actor recommendation, and a graph-based actor role recommendation using weighted label propagation as baselines and compare them with our framework. Experimental results show our approaches outperform them significantly.
Mohiuddin Solaimani, Sayeed Salam, Latifur Khan, Patrick T. Brandt, Vito D'Orazio
IEEE BigData5