EDBT 2026 Demo / reviewers in the wild / expert
Andrea L. Bertozzi
dblp:80/2099
· DBLP profile ↗
16ranked-venue papers in the field
0as first author
8since 2021 · last 2024
0000-0003-0396-7391ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 14Information Retrieval & Web Search · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | The 8th Workshop on Graph Techniques for Adversarial Activity Analytics (GTA3 2024)abstractGraphs are powerful analytic tools for modeling adversarial activities across a wide range of domains and applications. Examples include identifying and responding to cybersecurity systems' threats and vulnerabilities, strengthening critical infrastructure's resilience and robustness, and combating covert illicit activities that span various domains like finance, communication, and transportation. With the rapid development of generative AI, the lifecycle and throughput of adversarial activities, such as generating attacks or synthesizing deceptive signals, have accelerated significantly. For instance, a malicious actor can generate a large number of malware variants to flood defense systems or create agents to disseminate misleading signals, obscuring their activities. Consequently, there is a pressing need for novel and effective technology to autonomously handle these adversarial activities and keep pace with the evolving threats. The purpose of this workshop is to provide a forum to discuss emerging research problems and novel approaches in graph analysis for modeling adversarial activities in the age of generative AI. Jiejun Xu, Hanghang Tong, Andrea L. Bertozzi |
CIKM | 3 |
| 2023 | Hate speech and hate crimes: a data-driven study of evolving discourse around marginalized groupsabstractThis study explores the dynamic relationship between online discourse, as observed in tweets, and physical hate crimes, focusing on marginalized groups. Leveraging natural language processing techniques, including keyword extraction and topic modeling, we analyze the evolution of online discourse after events affecting these groups. Examining sentiment and polarizing tweets, we establish correlations with hate crimes in Black and LGBTQ+ communities. Using a knowledge graph, we connect tweets, users, topics, and hate crimes, enabling network analyses. Our findings reveal divergent patterns in the evolution of user communities for Black and LGBTQ+ groups, with notable differences in sentiment among influential users. This analysis sheds light on distinctive online discourse patterns and emphasizes the need to monitor hate speech to prevent hate crimes, especially following significant events impacting marginalized communities. Malvina Bozhidarova, Jonathn Chang, Aaishah Ale-rasool, Chongyao Ma, Andrea L. Bertozzi, P. Jeffrey Brantingham, Junyuan Lin, Sanjukta Krishnagopal |
IEEE Big Data | 6 |
| 2023 | AutoKG: Efficient Automated Knowledge Graph Generation for Language ModelsabstractTraditional methods of linking large language models (LLMs) to knowledge bases via the semantic similarity search often fall short of capturing complex relational dynamics. To address these limitations, we introduce AutoKG, a lightweight and efficient approach for automated knowledge graph (KG) construction. For a given knowledge base consisting of text blocks, AutoKG first extracts keywords using a LLM and then evaluates the relationship weight between each pair of keywords using graph Laplace learning. We employ a hybrid search scheme combining vector similarity and graph-based associations to enrich LLM responses. Preliminary experiments demonstrate that AutoKG offers a more comprehensive and interconnected knowledge retrieval mechanism compared to the semantic similarity search, thereby enhancing the capabilities of LLMs in generating more insightful and relevant outputs. Andrea L. Bertozzi |
IEEE Big Data | 2 |
| 2023 | Intentional Youth Development Activities and Peer Effects in a Gang Prevention ProgramabstractWe analyze the impact of group activities targeting Social-Emotional Learning (SEL) and peer effects on the risk and protective factors associated with gang involvement among youth participating in the Los Angeles Mayor’s Office of Gang Reduction and Youth Development (GRYD) Prevention program. We compare the impact of targeted and non-targeted activities in decreasing Internal Risk, External Risk, and Family Norms Risk, as measured by a standardized questionnaire. We show that targeted activities are effective in decreasing Internal Risk and that activities focusing on Emotional Management are the most beneficial. Since targeted activities involve group interactions, we investigate the impact of peer network effects on outcomes using both a linear-in-means model and dynamic mode decomposition with control (DMDc). Our analysis suggests that peer network effects contribute to beneficial changes in risk and protective factors above and beyond the skill-building content of the activities. Xiaoxian Shen, Zichun Liao, Andrea L. Bertozzi, P. Jeffrey Brantingham, Jona Lelmi |
IEEE Big Data | 7 |
| 2022 | Knowledge Graphs of the QAnon Twitter NetworkabstractUsing Knowledge Graphs to understand noisy naturalistic data has gained significant prominence in recent years. In this paper, we apply Knowledge Graphs to a new dataset of tweets of an ideologically far-right Twitter network by sourcing tweet histories of users who discussed QAnon in the summer of 2018 [1]. We further develop a new method that arms topic models with relational information from Knowledge Graphs and apply the new technique to study this dataset. Our analysis shows that users do not form a monolithic belief or social network, but rather comprise many smaller interlinking communities which discuss unique key political events (e.g., the January 6thCapitol riots). Clay Adams, Malvina Bozhidarova, Andrew Gao, Zhengtong Liu, John Priniski, Junyuan Lin, Rishi Sonthalia, Andrea L. Bertozzi, P. Jeffrey Brantingham |
IEEE Big Data | 9 |
| 2022 | Combining Dynamic Mode Decomposition and Difference-in-Differences in an Analysis of At-Risk YouthabstractWe analyze the impact of the Los Angeles Mayor’s Office of Gang Reduction Youth Development (GRYD) prevention programming using quasi-experimental data. We model the evolution of questionnaire scores and apply Dynamic Mode Decomposition (DMD) to describe the asymptotic behavior of the dynamical system. The analysis indicates that risk decreased for youth who enrolled in GRYD prevention services, while it increased or remained the same for those who were in the control group. We augment these observations using a difference-in-differences (DID) model, showing that the decrease in risk can be attributed to enrolment in prevention services. We draw a connection between DMD and DID using both mathematical analysis and empirical evidence from the questionnaire data. Combining DMD and DID with factor analysis, we investigate the effectiveness of prevention services with respect to different attitudinal domains. We conclude that gang prevention is most effective in impacting attitudes towards negative peer obedience and least effective in impacting attitudes towards violence for self defense. Our analytical approach can be extended to other types of repeated questionnaires. Marc Andrew Choi, Siyu Huang, Hengyuan Qi, Marco Scialanga, Emerson McMullen, Axel Sanchez Moreno, Yifei Lou, Andrea L. Bertozzi, P. Jeffrey Brantingham |
IEEE Big Data | 8 |
| 2021 | An Analysis of COVID-19 Knowledge Graph Construction and ApplicationsabstractThe construction and application of knowledge graphs have seen a rapid increase across many disciplines in re-cent years. Additionally, the problem of uncovering relationships between developments in the COVID-19 pandemic and social me-dia behavior is of great interest to researchers hoping to curb the spread of the disease. In this paper we present a knowledge graph constructed from COVID-19 related tweets in the Los Angeles area, supplemented with federal and state policy announcements and disease spread statistics. By incorporating dates, topics, and events as entities, we construct a knowledge graph that describes the connections between these useful information. We use natural language processing and change point analysis to extract tweet-topic, tweet-date, and event-date relations. Further analysis on the constructed knowledge graph provides insight into how tweets reflect public sentiments towards COVID-19 related topics and how changes in these sentiments correlate with real-world events. Dominic Flocco, Bryce Palmer-Toy, Ruixiao Wang 0001, Rishi Sonthalia, Junyuan Lin, Andrea L. Bertozzi, P. Jeffrey Brantingham |
IEEE BigData | 7 |
| 2021 | Active Learning for the Subgraph Matching ProblemabstractThe subgraph matching problem arises in a number of modern machine learning applications including segmented images and meshes of 3D objects for pattern recognition, bio-chemical reactions and security applications. This graph-based problem can have a very large and complex solution space especially when the world graph has many more nodes and edges than the template. In a real use-case scenario, analysts may need to query additional information about template nodes or world nodes to reduce the problem size and the solution space. Currently, this query process is done by hand, based on the personal experience of analysts. By analogy to the well-known active learning problem in machine learning classification problems, we present a machine-based active learning problem for the subgraph match problem in which the machine suggests optimal template target nodes that would be most likely to reduce the solution space when it is otherwise overly large and complex. The humans in the loop can then include additional information about those target nodes. We present some case studies for both synthetic and real world datasets for multichannel subgraph matching. Yurun Ge, Andrea L. Bertozzi |
IEEE BigData | 2 |
| 2020 | Who killed Lilly Kane? A case study in applying knowledge graphs to crime fictionabstractWe present a preliminary study of a knowledge graph created from season one of the television show Veronica Mars, which follows the eponymous young private investigator as she attempts to solve the murder of her best friend Lilly Kane. We discuss various techniques for mining the knowledge graph for clues and potential suspects. We also discuss best practice for collaboratively constructing knowledge graphs from television shows. Mariam Alaverdian, William Gilroy, Veronica Kirgios, Carolina Matuk, Daniel McKenzie, Tachin Ruangkriengsin, Andrea L. Bertozzi, P. Jeffrey Brantingham |
IEEE BigData | 8 |
| 2020 | Inexact Attributed Subgraph MatchingabstractWe present an approach for inexact subgraph matching on attributed graphs optimizing the graph edit distance. By combining lower bounds on the cost of individual assignments, we obtain a heuristic for a backtracking tree search to identify optimal solutions. We evaluate our algorithm on a knowledge graph dataset derived from real-world data, and analyze the space of optimal solutions. Thomas K. Tu, Jacob D. Moorman, Dominic Yang, Qinyi Chen, Andrea L. Bertozzi |
IEEE BigData | 5 |
| 2020 | Analyzing Effectiveness of Gang Interventions using Koopman Operator TheoryabstractKoopman operator theory, applied via numerical techniques such as dynamic mode decomposition (DMD) and autoencoders, has recently emerged as an interesting mathematical framework for understanding how complex, high-dimensional dynamical systems evolve. In this paper, we apply several DMD and autoencoder algorithms to a dataset of gang involvement and activity to assess the effectiveness City of Los Angeles Mayor's Office of Gang Reduction and Youth Development's (GRYD) Intervention Family Case Management Program. We compare various subsets of the data to explore differences in sub-populations. We then control for different covariates in our analysis of dynamical changes in population characteristics over time. Statistically significant results suggest the efficacy of the GRYD FCM Program. Sian Wen, Tanishq Bhatia, Nicholas Liskij, David Hyde 0001, Andrea L. Bertozzi, P. Jeffrey Brantingham |
IEEE BigData | 6 |
| 2020 | Emotion Classification and Textual Clustering Techniques for Gang Intervention DataabstractWe study a recent dataset documenting the nature of gang involvement among 14-25 year-olds participating in the Los Angeles Mayor's Office of Gang Reduction and Youth Development (GRYD) Intervention Family Case Management Program. We use natural language processing techniques, including emotion classification and textual clustering, to perform quantitative analyses of free-form responses in the data. These analyses yield insights into the effectiveness of the program and provide a better understanding of its participants. We also compare several computational techniques and remark on their relative effectiveness in application to this dataset. Ruofei Wu, Chenxin Yang, David Hyde 0001, Andrea L. Bertozzi, P. Jeffrey Brantingham |
IEEE BigData | 4 |
| 2019 | Noisy Subgraph Isomorphisms on Multiplex NetworksabstractWe address the problem of finding noisy subgraph isomorphisms on large multiplex networks. Our goal is to find as many subgraph matches as possible within a noise tolerance. We propose a novel approach based on the well-known A* search algorithm. Our approach employs new heuristics to estimate the number of missing edges of subgraph matches. This method is verified on one of the synthetic multiplex networks from the Modeling Adversarial Activity program of the Defense Advanced Research Projects Agency. Yanghui Wang, Andrea L. Bertozzi |
IEEE BigData | 5 |
| 2019 | Applications of Structural Equivalence to Subgraph Isomorphism on Multichannel MultigraphsabstractStructural Equivalence refers to the ability to exchange two vertices in a graph without changing the structure of the graph. We provide basic definitions and properties applicable to the subgraph isomorphism problem. We show examples of structural equivalence that reduce the size of the search tree for subgraph isomorphism counting and enumeration, applied to multichannel networks. Thien Nguyen 0004, Dominic Yang, Yurun Ge, Andrea L. Bertozzi |
IEEE BigData | 5 |
| 2019 | Learning to Predict Human Stress Level with Incomplete Sensor Data from Wearable DevicesabstractStress is a common problem in modern life that can bring both psychological and physical disorder. Wearable sensors are commonly used to study the relationship between physical records and mental status. Although sensor data generated by wearable devices provides an opportunity to identify stress in people for predictive medicine, in practice, the data are typically complicated and vague and also often fragmented. In this paper, we propose DataCompletion with Diurnal Regularizers (DCDR) and TemporallyHierarchical Attention Network (THAN) to address the fragmented data issue and predict human stress level with recovered sensor data. We model fragmentation as a sparsity issue. The nuclear norm minimization method based on the low-rank assumption is first applied to derive unobserved sensor data with diurnal patterns of human behaviors. A hierarchical recurrent neural network with the attention mechanism then models temporally structural information in the reconstructed sensor data, thereby inferring the predicted stress level. Data for this study were from 75 undergraduate students (taken from a sample of a larger study) who provided sensor data from smart wristbands. They also completed weekly stress surveys as ground-truth labels about their stress levels. This survey lasted 12 weeks and the sensor records are also in this period. The experimental results demonstrate that our approach significantly outperforms conventional methods in both data completion and stress level prediction. Moreover, an in-depth analysis further shows the effectiveness and robustness of our approach. Jyun-Yu Jiang, Zehan Chao, Andrea L. Bertozzi, Wei Wang 0010, Sean D. Young, Deanna Needell |
CIKM | 3 |
| 2018 | Filtering Methods for Subgraph Matching on Multiplex NetworksabstractWe present filtering methods for finding all sub-graphs of a large multiplex network that are isomorphic to a smaller template network. These methods are shown to be effective on a set of synthetic transaction networks from the DARPA Modeling Adversarial Activity (MAA) program. In some cases, filtering allows us to identify and enumerate all possible isomorphisms. We observe that in some of the MAA networks, the number of subgraphs isomorphic to the template is orders of magnitude larger than the size of the network. Jacob D. Moorman, Qinyi Chen, Thomas K. Tu, Zachary M. Boyd, Andrea L. Bertozzi |
IEEE BigData | 5 |