VLDB 2026 Research / reviewers in the wild / expert
Reihaneh Rabbany
dblp:94/9024 · also Reihaneh Rabbany Khorasgani
· DBLP profile ↗
26ranked-venue papers in the field
4as first author
16since 2021 · last 2026
0000-0003-2348-0353ORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 22 (4 first)Information Retrieval & Web Search · 3Database Systems & Data Management · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deepfakes in the 2025 Canadian Election: Prevalence, Partisanship, and Platform DynamicsabstractConcerns about AI-generated political content are growing, yet there is limited empirical evidence on how deepfakes appear and circulate across social platforms during major events in democratic countries. We analyze the 2025 Canadian federal election across X, Bluesky, and Reddit using a high-accuracy detector trained on diverse modern generative models, covering 187,778 posts. We find that 5.9% of election-related images were deepfakes. Right-leaning accounts shared them more often (9.2% of images flagged) than left-leaning users (3.9%), with flagged content more frequently defamatory or conspiratorial. Yet, most detected deepfakes were benign or non-political, and harmful ones drew little attention, accounting for only 0.1% of all views on X. Overall, deepfakes were present in the election conversation, but their reach was modest, and realistic fabricated images, although less common, drew higher engagement, highlighting growing concerns about their misuses. Victor Livernoche, Andreea Musulan, Zachary Yang, Jean-François Godbout, Reihaneh Rabbany |
WWW | 5 |
| 2025 | Responsible AI DayabstractThis special day event on Responsible Artificial Intelligence (AI) brings together researchers, practitioners, and policymakers to explore how data mining and machine learning systems can be designed to align with ethical principles, societal values, and human well-being. As AI technologies increasingly influence decisions in healthcare, finance, governance, and social systems, there is a critical need to develop frameworks that embed fairness, accountability, and privacy directly into the foundations of knowledge discovery. This full-day event will feature a mix of invited talks, interactive debates, expert panels, and peer-reviewed research presentations, all focused on the practical integration of ethical design into data-driven systems. The Responsible AI Day builds on the success of Canada's NSERC CREATE Program on Responsible AI, an interdisciplinary initiative training the next generation of AI researchers across computer science, law, bioethics, public health, and media studies. Topics will span scalable AI governance, privacy-preserving computation, algorithmic bias mitigation, and the socio-legal tensions emerging in generative AI. By positioning responsible AI as a sociotechnical challenge, this special day aligns with KDD's mission of advancing data science that is not only technically robust but also socially conscious. Ebrahim Bagheri, Faezeh Ensan, Calvin Hillis, Reihaneh Rabbany, Robin Cohen, Benjamin C. M. Fung, Sébastien Gambs |
KDD (2) | 4 |
| 2025 | Temporal Graph Learning WorkshopabstractThe Temporal Graph Learning (TGL) workshop, now in its third edition at KDD 2025, offers an interdisciplinary platform for researchers to explore the evolving applications of temporal networks in various domains, including recommender systems, social network analysis, traffic analytics, and epidemiological data analysis.The workshop aims to facilitate the exchange of ideas across disciplines, highlight successes and challenges in TGL, and outline future research directions.The workshop welcomes diverse contributions, offers keynote talks from academic and industry experts, and is complemented by a panel discussion on emerging aspects of TGL. Shenyang Huang, Daniele Zambon, Andrea Cini, Farimah Poursafaei, Jacob Chmura, Julia Gastinger, Reihaneh Rabbany, Michael M. Bronstein |
KDD (2) | 7 |
| 2025 | A Guide to Misinformation Detection Data and EvaluationabstractMisinformation is a complex societal issue, and mitigating solutions are difficult to create due to data deficiencies. To address this, we have curated the largest collection of (mis)information datasets in the literature, totaling 75. From these, we evaluated the quality of 36 datasets that consist of statements or claims, as well as the 9 datasets that consist of data in purely paragraph form. We assess these datasets to identify those with solid foundations for empirical work and those with flaws that could result in misleading and non-generalizable results, such as spurious correlations, or examples that are ambiguous or otherwise impossible to assess for veracity. We find the latter issue is particularly severe and affects most datasets in the literature. We further provide state-of-the-art baselines on all these datasets, but show that regardless of label quality, categorical labels may no longer give an accurate evaluation of detection model performance. Finally, we propose and highlight Evaluation Quality Assurance (EQA) as a tool to guide the field toward systemic solutions rather than inadvertently propagating issues in evaluation. Overall, this guide aims to provide a roadmap for higher quality data and better grounded evaluations, ultimately improving research in misinformation detection. All datasets and other artifacts are available at misinfo-datasets.complexdatalab.com. The extended paper, including the appendices, can be accessed via arXiv at arxiv.org/abs/2411.05060. Camille Thibault, Jacob-Junqi Tian, Gabrielle Péloquin-Skulski, Taylor Lynn Curtis, James Zhou, Florence Laflamme, Luke Yuxiang Guan, Reihaneh Rabbany, Jean-François Godbout, Kellin Pelrine |
KDD (2) | 8 |
| 2025 | Unified Game Moderation: Soft-Prompting and LLM-Assisted Label Transfer for Resource-Efficient Toxicity DetectionabstractToxicity detection in gaming communities faces significant scaling challenges when expanding across multiple games and languages, particularly in real-time environments where computational efficiency is crucial. We present two key findings to address these challenges while building upon our previous work on ToxBuster, a BERT-based real-time toxicity detection system. First, we introduce a soft-prompting approach that enables a single model to effectively handle multiple games by incorporating game-context tokens, matching the performance of more complex methods like curriculum learning while offering superior scalability. Second, we develop an LLM-assisted label transfer framework using GPT-4o-mini to extend support to seven additional languages. Evaluations on real game chat data across French, German, Portuguese, and Russian achieve macro F1-scores ranging from 32.96% to 58.88%, with particularly strong performance in German, surpassing the English benchmark of 45.39%. In production, this unified approach significantly reduces computational resources and maintenance overhead compared to maintaining separate models for each game and language combination. At Ubisoft, this model successfully identifies an average of 50 players, per game, per day engaging in sanctionable behavior. Zachary Yang, Domenico Tullo, Reihaneh Rabbany |
KDD (2) | 3 |
| 2024 | Party Prediction for TwitterabstractA large number of studies on social media compare the behaviour of users from different political parties. As a basic step, they employ a predictive model for inferring their political affiliation. The accuracy of this model can change the conclusions of a downstream analysis significantly, yet the choice between different models seems to be made arbitrarily. In this paper, we provide a comprehensive survey and an empirical comparison of the current party prediction practices and propose several new approaches which are competitive with or outperform state-of-the-art methods, yet require less computational resources. Party prediction models rely on the content generated by the users (e.g., tweet texts), the relations they have (e.g., who they follow), or their activities and interactions (e.g., which tweets they like). We examine all of these and compare their signal strength for the party prediction task. This paper lets the practitioner select from a wide range of data types that all give strong performance. Finally, we conduct extensive experiments on different aspects of these methods, such as data collection speed and transfer capabilities, which can provide further insights for both applied and methodological research. Kellin Pelrine, Anne Imouza, Zachary Yang, Jacob-Junqi Tian, Sacha Levy, Gabrielle Desrosiers-Brisebois, Aarash Feizi, Cécile Amadoro, André Blais, Jean-François Godbout, Reihaneh Rabbany |
ICWSM | 11 |
| 2024 | Temporal Graph Analysis with TGXabstractReal-world networks, with their evolving relations, are best captured as temporal graphs. However, existing software libraries are largely designed for static graphs where the dynamic nature of temporal graphs is ignored. Bridging this gap, we introduce TGX, a Python package specially designed for analysis of temporal networks that encompasses an automated pipeline for data loading, data processing, and analysis of evolving graphs. TGX provides access to eleven built-in datasets and eight external Temporal Graph Benchmark (TGB) datasets as well as any novel datasets in the .csv format. Beyond data loading, TGX facilitates data processing functionalities such as discretization of temporal graphs and node sub-sampling to accelerate working with larger datasets. For comprehensive investigation, TGX offers network analysis by providing a diverse set of measures, including average node degree and the evolving number of nodes and edges per timestamp. Additionally, the package consolidates meaningful visualization plots indicating the evolution of temporal patterns, such as Temporal Edge Appearance (TEA) and Temporal Edge Traffic (TET) plots. The TGX package is a robust tool for examining the features of temporal graphs and can be used in various areas like studying social networks, citation networks, and tracking user interactions. We plan to continuously support and update TGX based on community feedback. TGX is publicly available on: https://github.com/ComplexData-MILA/TGX. Razieh Shirzadkhani, Shenyang Huang, Elahe Kooshafar, Reihaneh Rabbany, Farimah Poursafaei |
WSDM | 4 |
| 2024 | Laplacian Change Point Detection for Single and Multi-view Dynamic GraphsabstractDynamic graphs are rich data structures that are used to model complex relationships between entities over time. In particular, anomaly detection in temporal graphs is crucial for many real-world applications such as intrusion identification in network systems, detection of ecosystem disturbances, and detection of epidemic outbreaks. In this article, we focus on change point detection in dynamic graphs and address three main challenges associated with this problem: (i) how to compare graph snapshots across time, (ii) how to capture temporal dependencies, and (iii) how to combine different views of a temporal graph. To solve the above challenges, we first propose Laplacian Anomaly Detection (LAD) which uses the spectrum of graph Laplacian as the low dimensional embedding of the graph structure at each snapshot. LAD explicitly models short-term and long-term dependencies by applying two sliding windows. Next, we propose MultiLAD, a simple and effective generalization of LAD to multi-view graphs. MultiLAD provides the first change point detection method for multi-view dynamic graphs. It aggregates the singular values of the normalized graph Laplacian from different views through the scalar power mean operation. Through extensive synthetic experiments, we show that (i) LAD and MultiLAD are accurate and outperforms state-of-the-art baselines and their multi-view extensions by a large margin, (ii) MultiLAD’s advantage over contenders significantly increases when additional views are available, and (iii) MultiLAD is highly robust to noise from individual views. In five real-world dynamic graphs, we demonstrate that LAD and MultiLAD identify significant events as top anomalies such as the implementation of government COVID-19 interventions which impacted the population mobility in multi-view traffic networks. Shenyang Huang, Samy Coulombe, Yasmeen Hitti, Reihaneh Rabbany, Guillaume Rabusseau |
ACM Trans. Knowl. Discov. Data | 4 |
| 2023 | Fast and Attributed Change Detection on Dynamic Graphs with Density of States
Shenyang Huang, Jacob Danovitch, Guillaume Rabusseau, Reihaneh Rabbany |
PAKDD (1) | 4 |
| 2023 | DisKeyword: Tweet Corpora Exploration for Keyword SelectionabstractHow to accelerate the search for relevant topical keywords within a tweet corpus? Computational social scientists conducting topical studies employ large, self-collected or crowdsourced social media datasets such as tweet corpora. Comprehensive sets of relevant keywords are often necessary to sample or analyze these data sources. However, naively skimming through thousands of keywords can quickly become a daunting task. In this study, we present a web-based application to simplify the search for relevant topical hashtags in a tweet corpus. DisKeyword allows users to grasp high-level trends in their dataset, while iteratively labeling keywords recommended based on their links to prior labeled hashtags. We open-source our code under the MIT license. Sacha Levy, Reihaneh Rabbany |
WSDM | 2 |
| 2023 | DeltaShield: Information Theory for Human- Trafficking DetectionabstractGiven a million escort advertisements, how can we spot near-duplicates? Such micro-clusters of ads are usually signals of human trafficking (HT). How can we summarize them to convince law enforcement to act? Spotting micro-clusters of near-duplicate documents is useful in multiple, additional settings, including spam-bot detection in Twitter ads, plagiarism, and more. We present InfoShield , which makes the following contributions: practical , being scalable and effective on real data; parameter-free and principled , requiring no user-defined parameters; interpretable , finding a document to be the cluster representative, highlighting all the common phrases, and automatically detecting “slots” (i.e., phrases that differ in every document); and generalizable , beating or matching domain-specific methods in Twitter bot detection and HT detection, respectively, as well as being language independent. Interpretability is particularly important for the anti-HT domain, where law enforcement must visually inspect ads. Our experiments on real data show that InfoShield correctly identifies Twitter bots with an F1 score over 90% and detects HT ads with 84% precision. Moreover, it is scalable, requiring about 8 hours for 4 million documents on a stock laptop. Our incremental version, DeltaShield , allows for fast, incremental updates, with minor loss of accuracy. Catalina Vajiac, Meng-Chieh Lee, Aayushi Kulshrestha, Sacha Levy, Namyong Park 0001, Andreas M. Olligschlaeger, Cara Jones, Reihaneh Rabbany, Christos Faloutsos |
ACM Trans. Knowl. Discov. Data | 8 |
| 2022 | A Strong Node Classification Baseline for Temporal GraphsabstractMany real-world complex systems can be modelled by temporal networks. Representation learning on these networks often captures their dynamic evolution and is a first step for performing further analysis, e.g. node classification. Node classification is a fundamental task for graph analysis in general and in the context of temporal graph, is often employed to categories nodes based on their activity patterns. Analysis of existing real world networks from different high-stake domains reveals that the rate of the malicious activities is on uptick, resulting in catastrophic social or economic consequences. This strongly motivates designing accurate node classification methods for temporal graphs. In this paper, we propose TGbase, for node classification on weighted temporal networks. TGbase efficiently extracts key features to consider the structural characteristics of each node and its neighborhood as well as the intensity and timestamp of the interactions among node pairs. These features accurately differentiate different classes of nodes, as shown on eight real-world benchmark datasets, outperforming multiple state-of-the-art (SOTA) deep/complex models. Our strong yet simple model is also generic, whereas the SOTA contenders are designed often for their specific (class of) datasets. Farimah Poursafaei, Zeljko Zilic, Reihaneh Rabbany |
SDM | 3 |
| 2021 | INFOSHIELD: Generalizable Information-Theoretic Human-Trafficking DetectionabstractGiven a million escort advertisements, how can we spot near-duplicates? Such micro-clusters of ads are usually signals of human trafficking. How can we summarize them, visually, to convince law enforcement to act? Can we build a general tool that works for different languages? Spotting micro-clusters of near-duplicate documents is useful in multiple, additional settings, including spam-bot detection in Twitter ads, plagiarism, and more.We present INFOSHIELD, which makes the following contributions: (a) Practical, being scalable and effective on real data, (b) Parameter-free and Principled, requiring no user-defined parameters, (c) Interpretable, finding a document to be the cluster representative, highlighting all the common phrases, and automatically detecting "slots", i.e. phrases that differ in every document; and (d) Generalizable, beating or matching domain-specific methods in Twitter bot detection and human trafficking detection respectively, as well as being language-independent finding clusters in Spanish, Italian, and Japanese. Interpretability is particularly important for the anti human-trafficking domain, where law enforcement must visually inspect ads.Our experiments on real data show that INFOSHIELD correctly identifies Twitter bots with an F1 score over 90% and detects human-trafficking ads with 84% precision. Moreover, it is scalable, requiring about 8 hours for 4 million documents on a stock laptop. Meng-Chieh Lee, Catalina Vajiac, Aayushi Kulshrestha, Sacha Levy, Namyong Park 0001, Cara Jones, Reihaneh Rabbany, Christos Faloutsos |
ICDE | 7 |
| 2021 | Graph Attention Networks with Positional Embeddings
Liheng Ma, Reihaneh Rabbany, Adriana Romero-Soriano |
PAKDD (1) | 2 |
| 2021 | SigTran: Signature Vectors for Detecting Illicit Activities in Blockchain Transaction Networks
Farimah Poursafaei, Reihaneh Rabbany, Zeljko Zilic |
PAKDD (1) | 2 |
| 2021 | The Surprising Performance of Simple Baselines for Misinformation DetectionabstractAs social media becomes increasingly prominent in our day to day lives, it is increasingly important to detect informative content and prevent the spread of disinformation and unverified rumours. While many sophisticated and successful models have been proposed in the literature, they are often compared with older NLP baselines such as SVMs, CNNs, and LSTMs. In this paper, we examine the performance of a broad set of modern transformer-based language models and show that with basic fine-tuning, these models are competitive with and can even significantly outperform recently proposed state-of-the-art methods. We present our framework as a baseline for creating and evaluating new methods for misinformation detection. We further study a comprehensive set of benchmark datasets, and discuss potential data leakage and the need for careful design of the experiments and understanding of datasets to account for confounding variables. As an extreme case example, we show that classifying only based on the first three digits of tweet ids, which contain information on the date, gives state-of-the-art performance on a commonly used benchmark dataset for fake news detection –Twitter16. We provide a simple tool to detect this problem and suggest steps to mitigate it in future datasets. Kellin Pelrine, Jacob Danovitch, Reihaneh Rabbany |
WWW | 3 |
| 2020 | Laplacian Change Point Detection for Dynamic GraphsabstractDynamic and temporal graphs are rich data structures that are used to model complex relationships between entities over time. In particular, anomaly detection in temporal graphs is crucial for many real world applications such as intrusion identification in network systems, detection of ecosystem disturbances and detection of epidemic outbreaks. In this paper, we focus on change point detection in dynamic graphs and address two main challenges associated with this problem: I) how to compare graph snapshots across time, II) how to capture temporal dependencies. To solve the above challenges, we propose Laplacian Anomaly Detection (LAD) which uses the spectrum of the Laplacian matrix of the graph structure at each snapshot to obtain low dimensional embeddings. LAD explicitly models short term and long term dependencies by applying two sliding windows. In synthetic experiments, LAD outperforms the state-of-the-art method. We also evaluate our method on three real dynamic networks: UCI message network, US senate co-sponsorship network and Canadian bill voting network. In all three datasets, we demonstrate that our method can more effectively identify anomalous time points according to significant real world events. Shenyang Huang, Yasmeen Hitti, Guillaume Rabusseau, Reihaneh Rabbany |
KDD | 4 |
| 2018 | Active Search of Connections for Case Building and Combating Human TraffickingabstractHow can we help an investigator to efficiently connect the dots and uncover the network of individuals involved in a criminal activity based on the evidence of their connections, such as visiting the same address, or transacting with the same bank account? We formulate this problem as Active Search of Connections, which finds target entities that share evidence of different types with a given lead, where their relevance to the case is queried interactively from the investigator. We present RedThread, an efficient solution for inferring related and relevant nodes while incorporating the user's feedback to guide the inference. Our experiments focus on case building for combating human trafficking, where the investigator follows leads to expose organized activities, i.e. different escort advertisements that are connected and possibly orchestrated. RedThread is a local algorithm and enables online case building when mining millions of ads posted in one of the largest classified advertising websites. The results of RedThread are interpretable, as they explain how the results are connected to the initial lead. We experimentally show that RedThread learns the importance of the different types and different pieces of evidence, while the former could be transferred between cases. Reihaneh Rabbany, David Bayani, Artur Dubrawski |
KDD | 1 |
| 2018 | Social-Affiliation Networks: Patterns and the SOAR Model
Dhivya Eswaran, Reihaneh Rabbany, Artur Dubrawski, Christos Faloutsos |
ECML/PKDD (2) | 2 |
| 2017 | Beyond Assortativity: Proclivity Index for Attributed Networks (ProNe)
Reihaneh Rabbany, Dhivya Eswaran, Artur Dubrawski, Christos Faloutsos |
PAKDD (1) | 1 |
| 2015 | Generalization of clustering agreements and distances for overlapping clusters and network communities
Reihaneh Rabbany, Osmar R. Zaïane |
Data Min. Knowl. Discov. | 1 |
| 2014 | Community Dynamics: Event and Role Analysis in Social Network Analysis
Justin Fagnan, Reihaneh Rabbany, Mansoureh Takaffoli, Eric Verbeek 0002, Osmar R. Zaïane |
ADMA | 2 |
| 2014 | SSRM: Structural social role mining for dynamic social networksabstractA social role is a special position an individual possesses within a network, which indicates his or her behaviours, expectations, and responsibilities. Identifying the roles that individuals play in a social network has various direct applications, such as detecting influential members, trustworthy people, idea innovators, etc. Roles can also be used for further analyses of the network, e.g. community detection, temporal event prediction, and summarization. In this paper, we propose a structural social role mining framework (SSRM), which is built to identify roles, study their changes, and analyze their impacts on the underlying social network. We define fundamental roles in a social network (namely leader, outermost, mediator, and outsider), and then propose methodologies to identify them, and track their changes. To identify these roles, we leverage the traditional social network analyses and metrics, as well as proposing new measures, including community-based variants for the Betweenness centrality. Our results indicate how the changes in the structural roles, in combination with the changes in the community structure of a network, can provide additional clues into the dynamics of networks. Afra Abnar, Mansoureh Takaffoli, Reihaneh Rabbany, Osmar R. Zaïane |
ASONAM | 3 |
| 2014 | Community evolution prediction in dynamic social networksabstractFinding patterns of interaction and predicting the future structure of networks has many important applications, such as recommendation systems and customer targeting. Community structure of social networks may undergo different temporal events and transitions. In this paper, we propose a framework to predict the occurrence of different events and transition for communities in dynamic social networks. Our framework incorporates key features related to a community - its structure, history, and influential members, and automatically detects the most predictive features for each event and transition. Our experiments on real world datasets confirms that the evolution of communities can be predicted with a very high accuracy, while we further observe that the most significant features vary for the predictability of each event and transition. Mansoureh Takaffoli, Reihaneh Rabbany, Osmar R. Zaïane |
ASONAM | 2 |
| 2013 | Incremental local community identification in dynamic social networksabstractSocial networks are usually drawn from the interactions between individuals, and therefore are temporal and dynamic in essence. Examining how the structure of these networks changes over time provides insights into their evolution patterns, factors that trigger the changes, and ultimately predict the future structure of these networks. One of the key structural characteristics of networks is their community structure --groups of densely interconnected nodes. Communities in a dynamic social network span over periods of time and are affected by changes in the underlying population, i.e. they have fluctuating members and can grow and shrink over time. In this paper, we introduce a new incremental community mining approach, in which communities in the current time are obtained based on the communities from the past time frame. Compared to previous independent approaches, this incremental approach is more effective at detecting stable communities over time. Extensive experimental studies on real datasets, demonstrate the applicability, effectiveness, and soundness of our proposed framework. Mansoureh Takaffoli, Reihaneh Rabbany, Osmar R. Zaïane |
ASONAM | 2 |
| 2012 | Relative Validity Criteria for Community Mining AlgorithmsabstractGrouping data points is one of the fundamental tasks in data mining, which is commonly known as clustering if data points are described by attributes. When dealing with interrelated data that does not have any attributes and is represented in the form of nodes and their relationships, this task is also referred to as community mining. There has been a considerable number of approaches proposed in recent years for mining communities in a given network. But little work has been done on how to evaluate community mining results. The common practice is to use an agreement measure to compare the mining result against a ground truth, however, the ground truth is not known in most of the real world applications. In this paper, we investigate relative clustering quality measures defined for evaluation of clustering data points with attributes and propose proper adaptations to make them applicable in the context of social networks. Not only these relative criteria could be used as metrics for evaluating quality of the groupings but also they could be used as objectives for designing new community mining algorithms. Reihaneh Rabbany, Mansoureh Takaffoli, Justin Fagnan, Osmar R. Zaïane, Ricardo J. G. B. Campello |
ASONAM | 1 |