EDBT 2026 Demo / reviewers in the wild / expert
Nazar Akrami
dblp:217/1503
· DBLP profile ↗
10ranked-venue papers in the field
1as first author
6since 2021 · last 2023
0000-0002-9641-6275ORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 5Big Data, Cloud & Distributed Data Systems · 5 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Linguistic Alignments: Detecting Similarities in Language Use in Written CommunicationabstractHuman language has many functions. Our communication on social media carries information about how we relate to ourselves and others, that is our identity, and we adjust our language to become more similar to our community - in the same way as we dress and style and act to show our commitment to the groups we belong to. Within a community, members adopt the community's language, and the common language becomes a unifying factor. Amendra Shrestha, Lisa Kaati, Nazar Akrami |
ASONAM | 3 |
| 2023 | Harmful Communication: Detection of Toxic Language and Threats on SwedishabstractHarmful communication, such as toxic language and threats directed toward individuals or groups, is a common problem on most social media platforms and online spaces. While several approaches exist for detecting toxic language and threats in English, few attempts have detected such communication in Swedish. Thus, we used transfer learning and BERT to train two machine learning models: one that detects toxic language and one that detects threats in Swedish. We also examined the intersection between toxicity and threat. The models are trained on data from several different sources, with authentic social media posts and data translated from English. Our models perform well on test data with an F1-score above 0.94 for detecting toxic language and 0.86 for detecting threats. However, the models' performance decreases significantly when they are applied to new unseen social media data. Examining the intersection between toxic language and threats, we found that 20% of the threats on social media are not toxic, which means that they would not be detected using only methods for detecting toxic language. Our finding highlights the difficulties with harmful language and the need to use different methods to detect different kinds of harmful language. Amendra Shrestha, Lisa Kaati, Nazar Akrami, Kevin Lindén, Arvin Moshfegh |
ASONAM | 3 |
| 2023 | General Risk Index : A Measure for Predicting Violent Behavior Through Written CommunicationabstractOne of the most challenging threats to the security of society is attacks from violent lone offenders. Identifying potential offenders is difficult since they act alone and do not necessarily communicate with others. However, several targeted violent attacks have been preceded by communication published on social media and the internet. Such communication is a valuable component when conducting risk and threat assessments.In this paper, we introduce a diagnostic measure of the risk of violent behavior based on text analysis. Using automated text analysis, we extract psychological variables and warning indicators from a given text and summarize these in an index that we denote as the general risk index. When developing the general risk index, we analyzed data (text) from 208 288 users on 32 online environments with diverse ideologies/orientations, including 76 previous violent lone offenders. A receiver operating characteristics (ROC) analysis showed that, when using the general risk index, it was possible to correctly classify between 90% and 96% of the cases depending on the comparison sample. These results support the predictive validity of the general risk index, suggesting that the risk index can be used to identify individuals with an increased risk of committing violent attacks that need further investigation. Lisa Kaati, Amendra Shrestha, Nazar Akrami |
IEEE Big Data | 3 |
| 2022 | Predicting Targeted Violence from Social Media CommunicationabstractFor decades, threat assessment professionals have used structured professional judgment instruments to make decisions about, for example, the likelihood of violent behavior of an individual. However, with the increased use of social media, most people use online digital platforms to communicate, which is also the case for potential violent offenders. For example, many mass shootings in recent years have been preceded by communication in online forums. In this paper, we introduce methods to identify markers of the warning behaviors Leakage, Fixation, Identification, and Affiliation and examine their discriminant validity. Our results show that violent offenders score higher on these markers and that these markers were present among a significantly higher proportion of violent offenders as compared to the normal population. We argue that our method can be used to predict potential planned, purposeful, or instrumental targeted violence in written communication. Automated methods for detecting warning behavior from written communication can serve as a complement to traditional threat assessment and provides unique opportunities for threat assessment beyond traditional methods. Lisa Kaati, Amendra Shrestha, Nazar Akrami |
ASONAM | 3 |
| 2022 | A Machine Learning Approach to Identify Toxic Language in the Online SpaceabstractIn this study, we trained three machine learning models to detect toxic language on social media. These models were trained using data from diverse sources to ensure that the models have a broad understanding of toxic language. Next, we evaluate the performance of our models on a dataset with samples of data from a large number of diverse online forums. The test dataset was annotated by three independent annotators. We also compared the performance of our models with Perspective API - a toxic language detection model created by Jigsaw and Google's Counter Abuse Technology team. The results showed that our classification models performed well on data from the domains they were trained on (Fl = 0.91, 0.91, & 0.84, for the RoBERTa, BERT, & SVM respectively), but the performance decreased when they were tested on annotated data from new domains (Fl = 0.80, 0.61, 0.49, & 0.77, for the RoBERTa, BERT, SVM, & Google perspective, respectively). Finally, we used the best-performing model on the test data (RoBERTa, ROC = 0.86) to examine the frequency (/proportion) of toxic language in 21 diverse forums. The results of these analyses showed that forums for general discussions with moderation (e.g., Alternate history) had much lower proportions of toxic language compared to those with minimal moderation (e.g., 8Kun). Although highlighting the complexity of detecting toxic language, our results show that model performance can be improved by using a diverse dataset when building new models. We conclude by discussing the implication of our findings and some directions for future research. Lisa Kaati, Amendra Shrestha, Nazar Akrami |
ASONAM | 3 |
| 2021 | Words of Suicide: Identifying Suicidal Risk in Written CommunicationsabstractSuicide is a global health problem with more than 700,000 individuals dying by self-destruction each year, yet it is classified as a low base rate behavior that is difficult to prognosticate. Aiming to advance suicide prediction and prevention, we examined the potential use of machine learning and text analyses models to predict suicide risk based on written communications. Specifically, we used a dataset consisting of more than 27,000 general writings unrelated to suicide, 193 genuine suicide notes from individuals who committed suicide, and an additional 89 suicide posts shared on sub-Reddits for an in-the-wild test to examine the prediction accuracy of two machine learning models (SVM & RoBERTa) and a linguistic marker model. Our tests showed that the machine learning models performed better than the linguistic marker model when examined on the test data. However, the linguistic marker model achieved higher results in the wild, correctly classifying 88% of written communications as a "high risk of suicide" versus 56% and 70% of the machine learning models. The best in-the-wild performing model was adopted in an online suicide risk assessment tool called Edwin to honor Edwin Shneidman for his numerous contributions to the field of suicidology. Finally, discrepancies between training and real-world data, vocabulary variation across domains, and the limited number of benchmarks constitute limitations that need to be addressed in future research. Amendra Shrestha, Nazar Akrami, Lisa Kaati, Julia Kupper, Matthew R. Schumacher |
IEEE BigData | 2 |
| 2020 | Introducing Digital-7 Threat Assessment of Individuals in Digital EnvironmentsabstractOne of the most challenging threats towards the security of the society is attacks from violent lone offenders, individuals that act alone or with minimal help from others without any economic gains or direct orders from organizations. Over the past few years, several terror attacks have been accompanied by manifestos published on social media platforms that outline ideology, motivation, and in some cases tactical choices. The trend in publishing manifestos and other communication on social media sites before committing an attack has increased the need for threat assessment in digital environments. Most existing methods for threat assessment are developed to be used in offline settings where information about an individual is accessible and cases where the individual is present and can answer questions. In this paper, we present seven indicators that can be used to assess the potential threat of violence based on digital communication only. The seven indicators are designed to be used when analyzing texts and can be seen as a complement to other risk assessment protocols. Amendra Shrestha, Nazar Akrami, Lisa Kaati |
ASONAM | 2 |
| 2019 | Automatic Extraction of Personality from Text: Challenges and OpportunitiesabstractIn this study we examined the possibility to extract personality traits from a text. We created an extensive dataset by having experts annotate personality traits in a large number of texts from multiple online sources. From these annotated texts we selected a sample and made further annotations ending up with a large low-reliability dataset and a small high-reliability dataset. We then used the two datasets to train and test several machine learning models to extract personality from text, including a language model. Finally, we evaluated our best models in the wild, on datasets from different domains. Our results show that the models based on the small high-reliability dataset performed better (in terms of R2) than models based on large low-reliability dataset. Also, the language model based on the small high-reliability dataset performed better than the random baseline. Finally, and more importantly, the results showed our best model did not perform better than the random baseline when tested in the wild. Taken together, our results show that determining personality traits from a text remains a challenge and that no firm conclusions can be made on model performance before testing in the wild. Nazar Akrami, Johan Fernquist, Tim Isbister, Lisa Kaati, Björn Pelzer |
IEEE BigData | 1 |
| 2019 | A Study on the Feasibility to Detect Hate Speech in SwedishabstractHate speech in digital environments is becoming a societal challenge. To deal with the problem, techniques that automatically detect hate speech have been developed by social media companies as well as researchers. Hate can be expressed in many different ways, which makes it difficult to detect automatically using algorithms. Also, how hate is expressed depends heavily on the language. The effectiveness of automatic detection techniques is still to be improved in many languages. In this paper, we attempt to detect hate speech in Swedish using machine learning. We compare different pre-trained language models that are fine-tuned on a corpus of hateful comments. To examine how well our models would work in a real scenario, we used a set of randomly selected comments from a Swedish discussion forum. The results showed that using pre-trained language models provides a better result than using a baseline SVM model, but it also reveals that detecting hate speech in the wild is challenge that need more research. Johan Fernquist, Oskar Lindholm, Lisa Kaati, Nazar Akrami |
IEEE BigData | 4 |
| 2019 | PRAT - a Tool for Assessing Risk in Written CommunicationabstractIn this paper, we present a tool for assessing the risk of targeted violence in written communication: the profile risk assessment tool (PRAT). The tool reads a text and extracts a profile based on a set of assessment factors including, for example, personality, emotionality, identification, fixation, and leakage. These factors have all been shown to have relevance in risk assessment by previous research. In PRAT, each assessment factor is broken down into a set of indicators to enable creating a profile for subsequent threat. To our knowledge, PRAT is the first threat assessment tool for digital communication. Amendra Shrestha, Lisa Kaati, Nazar Akrami |
IEEE BigData | 3 |