EDBT 2026 Demo / reviewers in the wild / expert
Lisa Kaati
dblp:34/5632
· DBLP profile ↗
20ranked-venue papers in the field
3as first author
8since 2021 · last 2024
0000-0002-3724-7504ORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 11 (2 first)Big Data, Cloud & Distributed Data Systems · 7 (1 first)Other / Interdisciplinary · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Visions of Violence : Threatful Communication in Incel CommunitiesabstractThe incel subculture has gained increasing attention due to its toxic nature and its association with real-world violence. This paper investigates the prevalence and characteristics of violent threatful communication within incel forums, focusing on a platform known as Blackpill. We have trained a machine learning model to detect violent threatful language and analyzed the posts. The analysis concentrated on three key aspects: the identity of perpetrators (categorized into first-person, third-person, or generalized), the targets (individuals, groups, or general targets), and the types of violence described (general violence, sexual violence, self-harm, and military violence). The analysis showed that the most common type violent threatful communication involved generalized perpetrators targeting groups. Additionally, 13.5% of the violent threatful communication contained coded language, including references to video games to obscure violent intentions. A smaller proportion of the posts (4.1%) glorified past mass shooters and violent criminals.This research highlights the complexities of identifying violent rhetoric in online forums and the use of coded language to evade detection, emphasizing the need for refined models in threat detection. Lukas Lundmark, Lisa Kaati, Amendra Shrestha |
IEEE Big Data | 2 |
| 2023 | Linguistic Alignments: Detecting Similarities in Language Use in Written CommunicationabstractHuman language has many functions. Our communication on social media carries information about how we relate to ourselves and others, that is our identity, and we adjust our language to become more similar to our community - in the same way as we dress and style and act to show our commitment to the groups we belong to. Within a community, members adopt the community's language, and the common language becomes a unifying factor. Amendra Shrestha, Lisa Kaati, Nazar Akrami |
ASONAM | 2 |
| 2023 | Harmful Communication: Detection of Toxic Language and Threats on SwedishabstractHarmful communication, such as toxic language and threats directed toward individuals or groups, is a common problem on most social media platforms and online spaces. While several approaches exist for detecting toxic language and threats in English, few attempts have detected such communication in Swedish. Thus, we used transfer learning and BERT to train two machine learning models: one that detects toxic language and one that detects threats in Swedish. We also examined the intersection between toxicity and threat. The models are trained on data from several different sources, with authentic social media posts and data translated from English. Our models perform well on test data with an F1-score above 0.94 for detecting toxic language and 0.86 for detecting threats. However, the models' performance decreases significantly when they are applied to new unseen social media data. Examining the intersection between toxic language and threats, we found that 20% of the threats on social media are not toxic, which means that they would not be detected using only methods for detecting toxic language. Our finding highlights the difficulties with harmful language and the need to use different methods to detect different kinds of harmful language. Amendra Shrestha, Lisa Kaati, Nazar Akrami, Kevin Lindén, Arvin Moshfegh |
ASONAM | 2 |
| 2023 | General Risk Index : A Measure for Predicting Violent Behavior Through Written CommunicationabstractOne of the most challenging threats to the security of society is attacks from violent lone offenders. Identifying potential offenders is difficult since they act alone and do not necessarily communicate with others. However, several targeted violent attacks have been preceded by communication published on social media and the internet. Such communication is a valuable component when conducting risk and threat assessments.In this paper, we introduce a diagnostic measure of the risk of violent behavior based on text analysis. Using automated text analysis, we extract psychological variables and warning indicators from a given text and summarize these in an index that we denote as the general risk index. When developing the general risk index, we analyzed data (text) from 208 288 users on 32 online environments with diverse ideologies/orientations, including 76 previous violent lone offenders. A receiver operating characteristics (ROC) analysis showed that, when using the general risk index, it was possible to correctly classify between 90% and 96% of the cases depending on the comparison sample. These results support the predictive validity of the general risk index, suggesting that the risk index can be used to identify individuals with an increased risk of committing violent attacks that need further investigation. Lisa Kaati, Amendra Shrestha, Nazar Akrami |
IEEE Big Data | 1 |
| 2022 | Predicting Targeted Violence from Social Media CommunicationabstractFor decades, threat assessment professionals have used structured professional judgment instruments to make decisions about, for example, the likelihood of violent behavior of an individual. However, with the increased use of social media, most people use online digital platforms to communicate, which is also the case for potential violent offenders. For example, many mass shootings in recent years have been preceded by communication in online forums. In this paper, we introduce methods to identify markers of the warning behaviors Leakage, Fixation, Identification, and Affiliation and examine their discriminant validity. Our results show that violent offenders score higher on these markers and that these markers were present among a significantly higher proportion of violent offenders as compared to the normal population. We argue that our method can be used to predict potential planned, purposeful, or instrumental targeted violence in written communication. Automated methods for detecting warning behavior from written communication can serve as a complement to traditional threat assessment and provides unique opportunities for threat assessment beyond traditional methods. Lisa Kaati, Amendra Shrestha, Nazar Akrami |
ASONAM | 1 |
| 2022 | A Machine Learning Approach to Identify Toxic Language in the Online SpaceabstractIn this study, we trained three machine learning models to detect toxic language on social media. These models were trained using data from diverse sources to ensure that the models have a broad understanding of toxic language. Next, we evaluate the performance of our models on a dataset with samples of data from a large number of diverse online forums. The test dataset was annotated by three independent annotators. We also compared the performance of our models with Perspective API - a toxic language detection model created by Jigsaw and Google's Counter Abuse Technology team. The results showed that our classification models performed well on data from the domains they were trained on (Fl = 0.91, 0.91, & 0.84, for the RoBERTa, BERT, & SVM respectively), but the performance decreased when they were tested on annotated data from new domains (Fl = 0.80, 0.61, 0.49, & 0.77, for the RoBERTa, BERT, SVM, & Google perspective, respectively). Finally, we used the best-performing model on the test data (RoBERTa, ROC = 0.86) to examine the frequency (/proportion) of toxic language in 21 diverse forums. The results of these analyses showed that forums for general discussions with moderation (e.g., Alternate history) had much lower proportions of toxic language compared to those with minimal moderation (e.g., 8Kun). Although highlighting the complexity of detecting toxic language, our results show that model performance can be improved by using a diverse dataset when building new models. We conclude by discussing the implication of our findings and some directions for future research. Lisa Kaati, Amendra Shrestha, Nazar Akrami |
ASONAM | 1 |
| 2021 | Similarity ranking using handcrafted stylometric traits in a swedish contextabstractIn this paper we introduce a new type of handcrafted textual features called stylometric traits, used to create a stylistic writeprint of an author's writing style. These can be divided into four categories: (i) word variations, (ii) abbreviations, (iii) internet jargon, and (iv) numbers. Johan Fernquist, Björn Pelzer, Lukas Lundmark, Lisa Kaati |
ASONAM | 4 |
| 2021 | Words of Suicide: Identifying Suicidal Risk in Written CommunicationsabstractSuicide is a global health problem with more than 700,000 individuals dying by self-destruction each year, yet it is classified as a low base rate behavior that is difficult to prognosticate. Aiming to advance suicide prediction and prevention, we examined the potential use of machine learning and text analyses models to predict suicide risk based on written communications. Specifically, we used a dataset consisting of more than 27,000 general writings unrelated to suicide, 193 genuine suicide notes from individuals who committed suicide, and an additional 89 suicide posts shared on sub-Reddits for an in-the-wild test to examine the prediction accuracy of two machine learning models (SVM & RoBERTa) and a linguistic marker model. Our tests showed that the machine learning models performed better than the linguistic marker model when examined on the test data. However, the linguistic marker model achieved higher results in the wild, correctly classifying 88% of written communications as a "high risk of suicide" versus 56% and 70% of the machine learning models. The best in-the-wild performing model was adopted in an online suicide risk assessment tool called Edwin to honor Edwin Shneidman for his numerous contributions to the field of suicidology. Finally, discrepancies between training and real-world data, vocabulary variation across domains, and the limited number of benchmarks constitute limitations that need to be addressed in future research. Amendra Shrestha, Nazar Akrami, Lisa Kaati, Julia Kupper, Matthew R. Schumacher |
IEEE BigData | 3 |
| 2020 | Introducing Digital-7 Threat Assessment of Individuals in Digital EnvironmentsabstractOne of the most challenging threats towards the security of the society is attacks from violent lone offenders, individuals that act alone or with minimal help from others without any economic gains or direct orders from organizations. Over the past few years, several terror attacks have been accompanied by manifestos published on social media platforms that outline ideology, motivation, and in some cases tactical choices. The trend in publishing manifestos and other communication on social media sites before committing an attack has increased the need for threat assessment in digital environments. Most existing methods for threat assessment are developed to be used in offline settings where information about an individual is accessible and cases where the individual is present and can answer questions. In this paper, we present seven indicators that can be used to assess the potential threat of violence based on digital communication only. The seven indicators are designed to be used when analyzing texts and can be seen as a complement to other risk assessment protocols. Amendra Shrestha, Nazar Akrami, Lisa Kaati |
ASONAM | 3 |
| 2020 | The Applicability of Authorship Verification to Swedish Discussion ForumsabstractThe authorship verification problem of determining whether two collections of textual content have been written by the same author or not is relevant in several contexts, e.g., when law enforcement officers try to find out whether a suspect with a known user account has other user accounts in the same or other web forums. In this paper, we evaluate how well the recently suggested attention-based hierarchical neural network approach AdHominem works and if it can be used to link user accounts on Swedish discussion forums. The results are encouraging and show that using AdHominem is a promising way forward when linking user accounts both on the same discussion forum and in cross-domain settings in which users write on a large variety of topics. Lukas Lundmark, Björn Pelzer, Lisa Kaati, Johan Fernquist |
IEEE BigData | 4 |
| 2019 | Levels of hate in online environmentsabstractHate speech in online environments is a severe problem for many reasons. The space for reasoning and argumentation shrinks, individuals refrain from expressing their opinions, and polarization of views increases. Hate speech contributes to a climate where threats and even violence are increasingly regarded as acceptable. Tor Berglind, Björn Pelzer, Lisa Kaati |
ASONAM | 3 |
| 2019 | Automatic Extraction of Personality from Text: Challenges and OpportunitiesabstractIn this study we examined the possibility to extract personality traits from a text. We created an extensive dataset by having experts annotate personality traits in a large number of texts from multiple online sources. From these annotated texts we selected a sample and made further annotations ending up with a large low-reliability dataset and a small high-reliability dataset. We then used the two datasets to train and test several machine learning models to extract personality from text, including a language model. Finally, we evaluated our best models in the wild, on datasets from different domains. Our results show that the models based on the small high-reliability dataset performed better (in terms of R2) than models based on large low-reliability dataset. Also, the language model based on the small high-reliability dataset performed better than the random baseline. Finally, and more importantly, the results showed our best model did not perform better than the random baseline when tested in the wild. Taken together, our results show that determining personality traits from a text remains a challenge and that no firm conclusions can be made on model performance before testing in the wild. Nazar Akrami, Johan Fernquist, Tim Isbister, Lisa Kaati, Björn Pelzer |
IEEE BigData | 4 |
| 2019 | A Study on the Feasibility to Detect Hate Speech in SwedishabstractHate speech in digital environments is becoming a societal challenge. To deal with the problem, techniques that automatically detect hate speech have been developed by social media companies as well as researchers. Hate can be expressed in many different ways, which makes it difficult to detect automatically using algorithms. Also, how hate is expressed depends heavily on the language. The effectiveness of automatic detection techniques is still to be improved in many languages. In this paper, we attempt to detect hate speech in Swedish using machine learning. We compare different pre-trained language models that are fine-tuned on a corpus of hateful comments. To examine how well our models would work in a real scenario, we used a set of randomly selected comments from a Swedish discussion forum. The results showed that using pre-trained language models provides a better result than using a baseline SVM model, but it also reveals that detecting hate speech in the wild is challenge that need more research. Johan Fernquist, Oskar Lindholm, Lisa Kaati, Nazar Akrami |
IEEE BigData | 3 |
| 2019 | PRAT - a Tool for Assessing Risk in Written CommunicationabstractIn this paper, we present a tool for assessing the risk of targeted violence in written communication: the profile risk assessment tool (PRAT). The tool reads a text and extracts a profile based on a set of assessment factors including, for example, personality, emotionality, identification, fixation, and leakage. These factors have all been shown to have relevance in risk assessment by previous research. In PRAT, each assessment factor is broken down into a set of indicators to enable creating a profile for subsequent threat. To our knowledge, PRAT is the first threat assessment tool for digital communication. Amendra Shrestha, Lisa Kaati, Nazar Akrami |
IEEE BigData | 2 |
| 2014 | Activity profiles in online social mediaabstractAnalysis and mining of social media has become an important research area. A challenging problem in this area consists in the identification of a group of users with similar patterns. In this paper, we propose the classification of users based on their activity profiles (e.g., periods of the day when the user is most and least active in online communications). Activity profiles can be useful for many purposes, such as marketing and user behavior analysis. They can also serve as a basis for other techniques such as stylometric and time analysis in order to increase the precision and scalability of multiple aliases identification techniques. We have implemented a prototype tool and applied it on a dataset from the ICWSM data set Boards.ie, showing the usefulness of our classification. Mohamed Faouzi Atig, Sofia Cassel, Lisa Kaati, Amendra Shrestha |
ASONAM | 3 |
| 2013 | Detecting multiple aliases in social mediaabstractMonitoring and analysis of web forums is becoming important for intelligence analysts around the globe since terrorists and extremists are using forums for spreading propaganda and communicating with each other. Various tools for analyzing the content of forum postings and identifying aliases that need further inspection by analysts have been proposed throughout literature, but a problem related to this is that individuals can make use of several aliases. In this paper we propose a number of matching techniques for detecting forum users who make use of multiple aliases. By combining different techniques such as time profiling and stylometric analysis of messages the accuracy of recognizing users with multiple aliases increases, as shown in experiments conducted on the ICWSM dataset boards.ie. Lisa Kaati, Amendra Shrestha |
ASONAM | 2 |
| 2012 | Combining Entity Matching Techniques for Detecting Extremist Behavior on Discussion BoardsabstractMany extremist groups and terrorists use the Web for various purposes such as exchanging and reinforcing their beliefs, making monitoring and analysis of discussion boards an important task for intelligence analysts in order to detect individuals that might pose a threat towards society. In this work we focus on how to automatically analyze discussion boards in an effective manner. More specifically, we propose a method for fusing several alias (entity) matching techniques that can be used to identify authors with multiple aliases. This is one part of a larger system, where the aim is to provide the analyst with a list of potential extremist worth investigating further. Johan Dahlin, Lisa Kaati, Christian Mårtenson, Pontus Svenson |
ASONAM | 3 |
| 2012 | Aspects of plan operators in a tree automata framework
Johanna Björklund, Eric Jonsson, Lisa Kaati |
FUSION | 3 |
| 2010 | Detecting Social Positions Using SimulationabstractDescribing social positions and roles is an important topic within social network analysis. One approach is to compute a suitable equivalence relation on the nodes of the target network. One relation that is often used for this purpose is regular equivalence, or bisimulation, as it is known within the field of computer science. In this paper we consider a relation from computer science called simulation relation. Simulation creates a partial order on the set of actors in a network and we can use this order to identify actors that have characteristic properties. The simulation relation can also be used to compute simulation equivalence which is a less restrictive equivalence relation than regular equivalence but is still computable in polynomial time. This paper primarily considers weighted directed networks and we present definitions of both weighted simulation equivalence and weighted regular equivalence. Weighted networks can be used to model a number of network domains, including information flow, trust propagation, and communication channels. Many of these domains have applications within homeland security and in the military, where one wants to survey and elicit key roles within an organization. Identifying social positions can be difficult when the target organization lacks a formal structure or is partially hidden. Joel Brynielsson, Johanna Björklund, Lisa Kaati, Christian Mårtenson, Pontus Svenson |
ASONAM | 3 |
| 2010 | Weighted unranked tree automata as a framework for plan recognition
Johanna Björklund, Lisa Kaati |
FUSION | 2 |