Jennifer S. Holmes

dblp:232/5562 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
1since 2021 · last 2021
0000-0002-7672-5935ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Theory of computation · 2 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2021 3M-Transformers for Event Coding on Organized Crime Domain
abstract
Political scientists and security agencies increasingly rely on computerized event data generation to track conflict processes and violence around the world. However, most of these approaches rely on pattern-matching techniques constrained by large dictionaries that are too costly to develop, update, or expand to emerging domains or additional languages. In this paper, we provide an effective solution to those challenges. Here we develop the 3M-Transformers (Multilingual, Multi-label, Multitask) approach for Event Coding from domain specific multilingual corpora, dispensing external large repositories for such task, and expanding the substantive focus of analysis to organized crime, an emerging concern for security research. Our results indicate that our 3M-Transformers configurations outperform state-of-the-art usual Transformers models (BERT and XLM-RoBERTa) for coding events on actors, actions and locations in English, Spanish, and Portuguese languages.
Erick Skorupa Parolin, Latifur Khan, Javier Osorio, Patrick T. Brandt, Vito D'Orazio, Jennifer S. Holmes
DSAA6
2020 HANKE: Hierarchical Attention Networks for Knowledge Extraction in Political Science Domain
abstract
Extracting structured metadata from unstructured text in different domains is gaining strong attention from multiple research communities. In Political Science, these metadata play a significant role on studying intra and inter-state interactions between political entities. The process of extracting such metadata usually relies on domain specific ontologies and knowledge-based repositories. In particular, Political Scientists regularly use the well-defined ontology CAMEO, which is designed for capturing conflict and mediation relations. Since CAMEO repositories are currently human maintained, the high cost and extensive human effort associated with updating them makes it difficult to include new entries on a regular basis. This paper introduces HANKE: an innovative framework for automatically extracting knowledge representations from unstructured sources, in order to extend CAMEO ontology both in the same domain and towards other related domains in political science. HANKE combines Hierarchical Attention Networks as engine for identifying relevant structures in raw-text and the novel Frequency-Based Ranker approach to obtain a collection of candidate entries for CAMEO's repositories. To show the efficiency of the proposed framework, we evaluate its performance on capturing existing CAMEO representations in a soft-labelled dataset. We also empirically demonstrate the versatility and superiority of HANKE method by applying it to two case studies related to CAMEO extension on its actual domain and towards organized crime domain.
Erick Skorupa Parolin, Latifur Khan, Javier Osorio, Vito D'Orazio, Patrick T. Brandt, Jennifer S. Holmes
DSAA6
2018 SPERG: Scalable Political Event Report Geoparsing in Big Data
abstract
Digital newspaper archives accumulated over the last few decades serve as an easily accessible, rich source of information for researchers to conduct analytic studies. Extracting unambiguous geographic identifiers, such as the geographic coordinates solely from text descriptions, a process also known as geoparsing, has proven to be a major challenge when applied to massive corpora like newspaper archives. We focus primarily on archived newspaper reports on political events and aim to parse the exact event location with high accuracy. We identify all the focus locations, which includes all locations mentioned in a report along with the coordinates and task ourselves with recognizing the location where the event actually occurred which we define as our primary focus location. Our objective is to extract the latitude-longitude information of these primary focus locations. Existing geoparsers only partially serve the purpose and are not robust enough to process large data archives in reasonable time. In this paper we propose a framework to extract geolocation information of primary focus location from 76 million documents in a distributed environment.
Aswin Krishna Gunasekaran, Maryam Bahojb Imani, Latifur Khan, Christan Grant, Patrick T. Brandt, Jennifer S. Holmes
ISI6