Jana Diesner

dblp:40/3892 · DBLP profile ↗
← Back
12ranked-venue papers in the field
2as first author
6since 2021 · last 2024
0000-0001-8183-7109ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 8 (1 first)Data Mining & Knowledge Discovery · 3 (1 first)Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2024 Topics, Temporal Patterns, and Network Characteristics of AI-Related Discourse on Reddit
Pingjing Yang, Kanyao Han, Jana Diesner
ASONAM (4)3
2024 Detection and Categorization of Needs during Crises Based on Twitter Data
abstract
The Ukraine-Russia conflict has brought sizable detrimental impact to the global energy, food, finance, and manufacturing industries, as well to many affected people. In this paper, we use Twitter (now X) to automatically identify who needs what from text data and how the types of needs that we categorized and standardize evolved throughout this conflict. Our findings suggest that the Ukraine expresses a need for weapons, Russia for land, Europe for gas, and America for leadership. The majority of needs expressed on Twitter during this conflict are related to the categories transportation, military, health & medical, financial and money, energy, and essential items (food, water, shelter, non-food items). Stated needs changed as the conflict escalated or fell into stalemate. Needs also varied depending on the tweet's location, with tweets from Ukraine's neighboring countries being related to food and medicine, while tweets from non-neighboring countries stated needs for clothing and tents. Tweets written in Ukrainian and Russian shared similar need terms, such as medicines and kits, compared to English tweets, which expressed needs such as ammunition and humanitarian aid. Our comparison of needs across four different disaster events, namely this conflict, an earthquake, a major hurricane, and the COVID-19 pandemic, showed how needs differ depending on the nature of the crisis and how domain-adjustment of needs categories is necessary. We contribute to the crisis informatics literature by (1) validating a methodology for using tweets to study the demand and supply of things that different stakeholders need during crisis events and (2) testing, comparing, and improving the fit of widely used need classification schemas for studying crisis from different domains.
Pingjing Yang, Ly Dinh, Alex Stratton, Jana Diesner
ICWSM4
2023 An expert-in-the-loop method for domain-specific document categorization based on small training data
abstract
Abstract Automated text categorization methods are of broad relevance for domain experts since they free researchers and practitioners from manual labeling, save their resources (e.g., time, labor), and enrich the data with information helpful to study substantive questions. Despite a variety of newly developed categorization methods that require substantial amounts of annotated data, little is known about how to build models when (a) labeling texts with categories requires substantial domain expertise and/or in‐depth reading, (b) only a few annotated documents are available for model training, and (c) no relevant computational resources, such as pretrained models, are available. In a collaboration with environmental scientists who study the socio‐ecological impact of funded biodiversity conservation projects, we develop a method that integrates deep domain expertise with computational models to automatically categorize project reports based on a small sample of 93 annotated documents. Our results suggest that domain expertise can improve automated categorization and that the magnitude of these improvements is influenced by the experts' understanding of categories and their confidence in their annotation, as well as data sparsity and additional category characteristics such as the portion of exclusive keywords that can identify a category.
Kanyao Han, Rezvaneh Rezapour, Katia Nakamura, Dikshya Devkota, Daniel C. Miller, Jana Diesner
J. Assoc. Inf. Sci. Technol.6
2022 Information Extraction from Social Media: A Hands-on Tutorial on Tasks, Data, and Open Source Tools
abstract
Information extraction (IE) is a common sub-area of natural language processing that focuses on identifying structured data from unstructured data. One application domain of IE is Information Retrieval (IR), which relies on accurate and high-performance IE to retrieve high quality results from massive datasets. Another example of IE is to identify named entities in a text. For example, in the the sentence "Katy Perry lives in the USA", Katy Perry and USA are named entities of types of PERSON and LOCATION, respectively. Also, identify the sentiment expressed in a text is another instance of IE: in the sentence, "This movie was awesome", the expressed sentiment is positive. Finally, IE is concerned with identifying various linguistic aspects of text data, e.g., part of speech of words, noun phrases, dependency parses, etc., which can serve as features for additional IE tasks. This tutorial introduces participants to a) the usage of Python based, open-source tools that support IE from social media data (mainly Twitter), and b) best practices for ensuring the responsible use of IE and research data. Participants will learn and practice various lexical, semantic, and syntactic IE techniques that are commonly used for analyzing tweets. Participants will also be familiarized with the landscape of publicly available social media data (including popular NLP and IE benchmarks) and methods for collecting and preparing them for analysis. Furthermore, participants will be trained to use a suite of open source tools (SAIL for active learning, TwitterNER for named entity recognition, TweetNLP for transformer based NLP, and SocialMediaIE for multi task learning), which utilize advanced machine learning techniques (e.g., deep learning, active learning with human-in-the-loop, multi-lingual, and multi-task learning) to perform IE on their own or existing datasets. Participants will also learn how social contexts of text production and usage of results can be integrated into IE systems to improve these systems and to consider the role of time in improving social media IE quality. Finally, participants will learn about the governance of social media data for research purposes. The tools introduced in the tutorial will focus on the three main stages of IE, namely, collection of data (including annotation), data processing and analytics, and visualization of the extracted information. More details can be found at: https://socialmediaie.github.io/tutorials/
Shubhanshu Mishra, Rezvaneh Rezapour, Jana Diesner
CIKM3
2022 Information Extraction from Social Media: A Hands-On Tutorial on Tasks, Data, and Open Source Tools
Shubhanshu Mishra, Rezvaneh Rezapour, Jana Diesner
ECIR (2)3
2021 Variation in Situational Awareness Information due to Selection of Data Source, Summarization Method, and Method Implementation
Maria Janina Sarol, Ly Dinh, Jana Diesner
ICWSM3
2019 Adversarial perturbations to manipulate the perception of power and influence in networks
abstract
Observed social networks are often considered as proxies for underlying social networks. The analysis of observed networks oftentimes involves the identification of influential nodes via various centrality metrics. Our work is motivated by recent research on the investigation and design of adversarial attacks on machine learning systems. We apply the concept of adversarial attacks to social networks by studying strategies by which an adversary can minimally perturb the observed network structure to achieve their target function of modifying the ranking of nodes according to centrality measures. This can represent the attempts of an adversary to boost or demote the degree to which others perceive them as influential or powerful. It also allows us to study the impact of adversarial attacks on targets and victims, and to design metrics and security measures that help to identify and mitigate adversarial network attacks. We conduct a series of experiments on synthetic network data to identify attacks that allow the adversarial node to achieve their objective with a single move. We test this approach on different common network topologies and for common centrality metrics. We find that there is a small set of moves that result in the adversary achieving their objective, and this set is smaller for decreasing centrality metrics than for increasing them. These results can help with assessing the robustness of centrality measures. The notion of changing social network data to yield adversarial outcomes has practical implications, e.g., for information diffusion on social media, influence and power dynamics in social systems, and improving network security.
Mihai Valentin Avram, Shubhanshu Mishra, Nikolaus Nova Parulian, Jana Diesner
ASONAM4
2016 Distortive effects of initial-based name disambiguation on measurements of large-scale coauthorship networks
abstract
Scholars have often relied on name initials to resolve name ambiguities in large‐scale coauthorship network research. This approach bears the risk of incorrectly merging or splitting author identities. The use of initial‐based disambiguation has been justified by the assumption that such errors would not affect research findings too much. This paper tests that assumption by analyzing coauthorship networks from five academic fields—biology, computer science, nanoscience, neuroscience, and physics—and an interdisciplinary journal, PNAS. Name instances in data sets of this study were disambiguated based on heuristics gained from previous algorithmic disambiguation solutions. We use disambiguated data as a proxy of ground‐truth to test the performance of three types of initial‐based disambiguation. Our results show that initial‐based disambiguation can misrepresent statistical properties of coauthorship networks: It deflates the number of unique authors, number of components, average shortest paths, clustering coefficient, and assortativity, while it inflates average productivity, density, average coauthor number per author, and largest component size. Also, on average, more than half of top 10 productive or collaborative authors drop off the lists. Asian names were found to account for the majority of misidentification by initial‐based disambiguation due to their common surname and given name initials.
Jinseok Kim 0001, Jana Diesner
J. Assoc. Inf. Sci. Technol.2
2015 Little Bad Concerns: Using Sentiment Analysis to Assess Structural Balance in Communication Networks
abstract
We present and test a scalable approach for assigning valence to links in unsigned graphs with the ultimate goal of enabling triadic balanced assessment in communication networks. We do this by applying domain-adjusted sentiment analysis to the content of communication data and translating aggregated sentiment scores for information exchanged between network members into link signs. This approach facilitates fast, informed and systematic balance testing (we generate link signs for 166,670 triads in our data); allowing for empirical hypothesis testing and theory building based on current or archival communication data. The proposed technique eliminates the need for manually labeling text data, and overcomes limitations with inferring valence from self-reported or user-generated (meta-) data in situations where historical context and ground truth valence data might be unavailable or limited. We test this approach on corporate email data to complement the large amount of prior work based on social media data and the limited knowledge on sentiment in professional settings. Our results suggest that sentiment is overall slightly positive and emotionality is low, which reflects conventions of language use in a corporate environment. We observe that people draw from (the top of) a smaller pool of positive terms more frequently than from a larger set of negative terms. The ratio of balanced triads (on average about 88%) to unbalanced triads (12%) remains relatively stable despite changes in corporate performance. The labor-intense adjustment of a given lexical resource to some dataset and domain pays off as it generates more empirical evidence with lower variance.
Jana Diesner, Craig S. Evans
ASONAM1
2015 Impact of Entity Disambiguation Errors on Social Network Properties
Jana Diesner, Craig S. Evans, Jinseok Kim 0001
ICWSM1
2015 Coauthorship networks: A directed network approach considering the order and number of coauthors
abstract
In many scientific fields, the order of coauthors on a paper conveys information about each individual's contribution to a piece of joint work. We argue that in prior network analyses of coauthorship networks, the information on ordering has been insufficiently considered because ties between authors are typically symmetrized. This is basically the same as assuming that each coauthor has contributed equally to a paper. We introduce a solution to this problem by adopting a coauthorship credit allocation model proposed by Kim and Diesner (2014), which in its core conceptualizes coauthoring as a directed, weighted, and self‐looped network. We test and validate our application of the adopted framework based on a sample data of 861 authors who have published in the journal Psychometrika. The results suggest that this novel sociometric approach can complement traditional measures based on undirected networks and expand insights into coauthoring patterns such as the hierarchy of collaboration among scholars. As another form of validation, we also show how our approach accurately detects prominent scholars in the Psychometric Society affiliated with the journal.
Jinseok Kim 0001, Jana Diesner
J. Assoc. Inf. Sci. Technol.2
2014 Why name ambiguity resolution matters for scholarly big data research
abstract
This paper illustrates how data pre-processing choices about author name disambiguation can affect research findings about scholarly networks and hypotheses about underlying social mechanisms. We have analyzed three big scholarly datasets that were disambiguated algorithmically and via two common initial-based disambiguation methods; namely first-initial and all-initials disambiguation. The comparison of resulting bibliometric and network properties revealed that initial-disambiguation bears the prevalent risks of incorrectly merging author identities, underestimating the number of unique authors and inflating the average productivity and number of collaborators per author. The gaps between outcomes of name ambiguity resolution methods range from −4.23% to −87.36% per dataset for the number of unique authors, from 3.75% to 691.20% for average productivity, and from 5.06% to 285.28% for degree centrality for initial based methods compared to algorithmic disambiguation. This calls for special attention to data pre-processing choices in scholarly big data research.
Jinseok Kim 0001, Jana Diesner, Amirhossein Aleyasen, Hwan-Min Kim
IEEE BigData2