VLDB 2026 Research / reviewers in the wild / expert
Paras Sheth
dblp:285/5124
· DBLP profile ↗
18ranked-venue papers in the field
9as first author
18since 2021 · last 2026
0000-0002-6186-6946ORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 7 (3 first)Information Retrieval & Web Search · 6 (2 first)Big Data, Cloud & Distributed Data Systems · 4 (3 first)Database Systems & Data Management · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Workshop on Benchmarking Causal Models (CausalBench)
K. Selçuk Candan, Huan Liu 0001, Ruocheng Guo, Paras Sheth |
WSDM | 4 |
| 2026 | Causality Guided Representation Learning for Cross-Style Hate Speech Detection
Chengshuai Zhao, Shu Wan 0002, Paras Sheth, Karan Patwa, K. Selçuk Candan, Huan Liu 0001 |
WWW | 3 |
| 2025 | CausalBench-ER: Causally-Informed Explanations and Recommendations for Reproducible BenchmarkingabstractDue to the critical role causality plays in decision-making, the state of-the-art in machine learning for causality is rapidly evolving. With rapid development and deployment of new models, datasets, and metrics, it is increasingly difficult for researchers and practitioners to identify the most suitable approach for their problem. Models exhibit different performances when they train on different data or even when they are used under different hardware/software platforms, making it challenging for users to select the appropriate setup pertinent to their problem. To address these difficulties, we present a computing framework, CausalBench-ER that serves, not only as a benchmarking platform for causal machine learning models, but also as a resource that can explain benchmarking results across different metrics, software, and hardware setups. Furthermore, CausalBench-ER recommends additional scenarios to consider to help pave the way towards more robust benchmarking. Ahmet Kapkiç, Pratanu Mandal, Abhinav Gorantla, Shu Wan 0002, Ertugrul Çoban, Paras Sheth, Huan Liu 0001, K. Selçuk Candan |
CIKM | 6 |
| 2025 | CausalBench: Causal Learning Research StreamlinedabstractRecent advances in causal machine learning introduced a plethora of new causal discovery and causal inference models to tackle decision support problems. Yet, these models exhibit different performance when they train on different data, and even different hardware/software platforms, making it challenging for users to select the appropriate setup pertinent to their specific problem instance. The situation is complicated by the fact that, until recently, the field lacked a unified, publicly available, and configurable platform that supports all major causal inference tasks, including causal discovery, causal effect estimation, and causal inference. CausalBench is a comprehensive benchmarking tool for causal machine learning that facilitates accurate and reproducible benchmarking of causal models across metrics and deployment contexts and helps users to select the most appropriate set up (such as hyper-parameter configuration) for the specific problem setting. This tutorial is intended to familiarize attendees from diverse backgrounds, who are interested in causal learning models and with the capabilities of CausalBench. The tutorial begins with an introduction to ''causality'' and causal machine learning, and then provides hands-on experience with CausalBench to equip attendees with the knowledge necessary to utilize CausalBench for their causal learning problems. Ahmet Kapkiç, Pratanu Mandal, Abhinav Gorantla, Shu Wan 0002, Ertugrul Çoban, Paras Sheth, Huan Liu 0001, K. Selçuk Candan |
KDD (2) | 6 |
| 2024 | Cross-Platform Hate Speech Detection with Weakly Supervised Causal DisentanglementabstractContent moderation on social media faces increasing challenges due to the rapid evolution of hate speech. Identifying hate speech is challenging, especially as it constantly evolves to evade detection. To address this, current methods often rely on auxiliary data like target labels, which specify the particular group targeted by hate speech, to improve detection accuracy. While these target labels can enhance model performance, they are often scarce, inconsistent across platforms, and unable to capture the full spectrum of hate speech variations. To overcome these limitations, we introduce HATE-WATCH, a novel weakly supervised framework that adapts to the fluid nature of hate speech without relying heavily on explicit target labels. By employing confidence-based reweighting and contrastive regularization, HATE-WATCH effectively disentangles input features into universal and platform-specific representations, enabling robust detection even in the absence of detailed target labels. This approach significantly advances cross-platform hate speech detection, offering a more adaptable and scalable solution that contributes to safer online communities by addressing the real-world complexities of content moderation. Paras Sheth, Tharindu Kumarage, Raha Moraffah, Aman Chadha, Huan Liu 0001 |
IEEE Big Data | 1 |
| 2024 | Introducing CausalBench: A Flexible Benchmark Framework for Causal Analysis and Machine Learning
Ahmet Kapkiç, Pratanu Mandal, Shu Wan 0002, Paras Sheth, Abhinav Gorantla, Yoonhyuk Choi, Huan Liu 0001, K. Selçuk Candan |
CIKM | 4 |
| 2024 | Exploring Platform Migration Patterns between Twitter and Mastodon: A User Behavior StudyabstractA recent surge of users migrating from Twitter to alternative platforms, such as Mastodon, raised questions regarding what migration patterns are, how different platforms impact user behaviors, and how migrated users settle in the migration process. In this study, we elaborate on how we investigate these questions by collecting data over 10,000 users who migrated from Twitter to Mastodon within the first ten weeks following the ownership change of Twitter. Our research is structured in three primary steps. First, we develop algorithms to extract and analyze migration patterns. Second, by leveraging behavioral analysis, we examine the distinct architectures of Twitter and Mastodon to learn how user behaviors correspond with the characteristics of each platform. Last, we determine how particular behavioral factors influence users to stay on Mastodon. We share our findings of user migration, insights, and lessons learned from the user behavior study. Ujun Jeong, Paras Sheth, Anique Tahir, Faisal Alatawi, H. Russell Bernard, Huan Liu 0001 |
ICWSM | 2 |
| 2024 | Causality Guided Disentanglement for Cross-Platform Hate Speech Detectionabstractespite their value in promoting open discourse, social media plat- forms are often exploited to spread harmful content. Current deep learning and natural language processing models used for detect- ing this harmful content rely on domain-specific terms affecting their ability to adapt to generalizable hate speech detection. This is because they tend to focus too narrowly on particular linguistic signals or the use of certain categories of words. Another signifi- cant challenge arises when platforms lack high-quality annotated data for training, leading to a need for cross-platform models that can adapt to different distribution shifts. Our research introduces a cross-platform hate speech detection model capable of being trained on one platform's data and generalizing to multiple unseen platforms. One way to achieve good generalizability across plat- forms is to disentangle the input representations into invariant and platform-dependent features. We also argue that learning causal relationships, which remain constant across diverse environments, can significantly aid in understanding invariant representations in hate speech. By disentangling input into platform-dependent fea- tures (useful for predicting hate targets) and platform-independent features (used to predict the presence of hate), we learn invariant representations resistant to distribution shifts. These features are then used to predict hate speech across unseen platforms. Our ex- tensive experiments across four platforms highlight our model's enhanced efficacy compared to existing state-of-the-art methods in detecting generalized hate speech Paras Sheth, Raha Moraffah, Tharindu Kumarage, Aman Chadha, Huan Liu 0001 |
WSDM | 1 |
| 2023 | Quantifying the Echo Chamber Effect: An Embedding Distance-based ApproachabstractThe rise of social media platforms has facilitated the formation of echo chambers, which are online spaces where users predominantly encounter viewpoints that reinforce their existing beliefs while excluding dissenting perspectives. This phenomenon significantly hinders information dissemination across communities and fuels societal polarization. Therefore, it is crucial to develop methods for quantifying echo chambers. In this paper, we present the Echo Chamber Score (ECS), a novel metric that assesses the cohesion and separation of user communities by measuring distances between users in the embedding space. In contrast to existing approaches, ECS is able to function without labels for user ideologies and makes no assumptions about the structure of the interaction graph. To facilitate measuring distances between users, we propose EchoGAE, a self-supervised graph autoencoder-based user embedding model that leverages users' posts and the interaction graph to embed them in a manner that reflects their ideological similarity. To assess the effectiveness of ECS, we use a Twitter dataset consisting of four topics - two polarizing and two non-polarizing. Our results showcase ECS's effectiveness as a tool for quantifying echo chambers and shedding light on the dynamics of online discourse. Faisal Alatawi, Paras Sheth, Huan Liu 0001 |
ASONAM | 2 |
| 2023 | STREAMS: Towards Spatio-Temporal Causal Discovery with Reinforcement Learning for Streamflow Rate PredictionabstractThe capacity to anticipate streamflow is critical to the efficient functioning of reservoir systems as it gives vital information to reservoir operators about water release quantities as well as help quantify the impact of environmental factors on downstream water quality. Yet, streamflow modelling is difficult owing to the intricate interactions between different watershed outlets. In this paper, we argue that one possible solution to this problem is to identify the causal structure of these outlets, which would allow for the identification of crucial watershed outlets while capturing the spatiotemporally informed complex relationships leading to improved hydrological resource management. However, due to the inherent complexity of spatiotemporal causal learning problems, extending existing causal discovery methods to a whole basin is a major hurdle. To address these issues, we offer STREAMS, a new framework that uses Reinforcement Learning (RL) to optimize the search space for causal discovery and an LSTM-GCN based autoencoder to infer spatiotemporal causal features for streamflow rate prediction. We conduct extensive experiments on the Brazos river basin carried out within the scope of a US Army Corps of Engineers, Engineering With Nature Initiative project, including empirical studies of generalization performance to verify the nature of the inferred relationships. Paras Sheth, Ahmadreza Mosallanezhad, Kaize Ding, Reepal Shah, John Sabo, Huan Liu 0001, K. Selçuk Candan |
CIKM | 1 |
| 2023 | PEACE: Cross-Platform Hate Speech Detection - A Causality-Guided Framework
Paras Sheth, Tharindu Kumarage, Raha Moraffah, Aman Chadha, Huan Liu 0001 |
ECML/PKDD (1) | 1 |
| 2023 | Causal Disentanglement for Implicit Recommendations with Network InformationabstractOnline user engagement is highly influenced by various machine learning models, such as recommender systems. These systems recommend new items to the user based on the user’s historical interactions. Implicit recommender systems reflect a binary setting showing whether a user interacted (e.g., clicked on) with an item or not. However, the observed clicks may be due to various causes such as user’s interest, item’s popularity, and social influence factors. Traditional recommender systems consider these causes under a unified representation, which may lead to the emergence and amplification of various biases in recommendations. However, recent work indicates that by disentangling the unified representations, one can mitigate bias (e.g., popularity bias) in recommender systems and help improve recommendation performance. Yet, prior work in causal disentanglement in recommendations does not consider a crucial factor, that is, social influence. Social theories such as homophily and social influence provide evidence that a user’s decision can be highly influenced by the user’s social relations. Thus, accounting for the social relations while disentangling leads to less biased recommendations. To this end, we identify three separate causes behind an effect (e.g., clicks): (a) user’s interest, (b) item’s popularity, and (c) user’s social influence. Our approach seeks to causally disentangle the user and item latent features to mitigate popularity bias in implicit feedback–based social recommender systems. To achieve this goal, we draw from causal inference theories and social network theories and propose a causality-aware disentanglement method that leverages both the user–item interaction network and auxiliary social network information. Experiments on real-world datasets against various state-of-the-art baselines validate the effectiveness of the proposed model for mitigating popularity bias and generating de-biased recommendations. Paras Sheth, Ruocheng Guo, Lu Cheng 0001, Huan Liu 0001, K. Selçuk Candan |
ACM Trans. Knowl. Discov. Data | 1 |
| 2022 | Exploring the Target Distribution for Surrogate-Based Black-Box AttacksabstractDeep Neural Networks are shown to be prone to adversarial attacks. In the black-box setting, where no information about the target is available, surrogate-based black-box attacks train a surrogate on samples queried from the target to imitate the black-box’s behavior. The trained surrogate is then attacked to generate adversarial examples. Existing surrogate-based attacks suffer from low success rates because they fail to accurately capture the target’s behavior, i.e., their surrogates only mimic the target’s outputs for a given set of inputs. Moreover, their attack strategy relies on noisy estimations of high dimensional gradients w.r.t. the inputs (i.e., surrogate’s gradients) to generate adversarial examples. Ideally, a successful surrogate-based attack should possess two properties: (1) Train and employ a surrogate that accurately imitates the target behavior for every pair of input and output, i.e., the joint distribution of the target over its input and outputs; and (2) Generate adversarial examples by directly manipulating the class-dependent factors of the input, i.e., factors that affect the target’s output, rather than relying on noisy estimations of gradients. We propose a novel surrogate-based attack framework with a surrogate architecture that learns the target distribution over its inputs and outputs while disentangling the class-dependent factors from class-irrelevant ones. The framework is equipped with a novel attack strategy that fully utilizes the target distribution captured by the surrogate while generating adversarial examples by directly manipulating the class-dependent factors. Extensive experiments demonstrate the efficacy of our attack in generating highly successful adversarial examples compared to state-of-the-art methods. Raha Moraffah, Paras Sheth, Huan Liu 0001 |
IEEE Big Data | 2 |
| 2022 | Causal Discovery for Feature Selection in Physical Process-Based Hydrological SystemsabstractPhysical process-based hydrological models are widely adopted to simulate the water quantity or quality. One of the most commonly used hydrological models is Soil and Water Assessment Tool (SWAT). SWAT models for a large watershed can have over tens of thousands of Hydrological Resource Units (HRUs) which necessitates considerable computational resources. One way to speed up applications of the SWAT model could be to leverage machine learning techniques to identify the crucial features for the prediction task – feature selection. However, majority of the feature selection techniques rely on correlations or some form of a score metric (e.g. mutual information). Furthermore, since correlation does not imply causation, it is important to identify the causal features to improve the prediction accuracy while enhancing the interpretability of machine learning models. However, the SWAT model uses multiple data inputs and features that typically vary by space/HRUs, but may or may not vary over time. This makes it difficult to directly utilize causal discovery models to infer the causal relations. Furthermore, due to the lack of the ground truth causal graph for the SWAT model it is difficult to comment on the validity of the learned causal relations. To overcome these problems, we propose a novel framework that first infers the causal relations for the daily scale of the SWAT data using causal discovery algorithms. Then, it utilizes a community detection module to group similar features together for better interpretability. Finally, it identifies the stable causal relations that appear most often across all the timesteps and leverage them for the prediction of the water quantity. By utilizing only the causal features for the prediction of the target variable can lead to high accuracy as it removes the reliance on spurious correlations. Furthermore, we conduct extensive experiments to validate the effectiveness of the proposed framework along with a real-world case study to evaluate whether the selected features are interpretable or not. Paras Sheth, Durmus Doner, Yuhang Wei, Rebecca Muenich, John Sabo, K. Selçuk Candan, Huan Liu 0001 |
IEEE Big Data | 1 |
| 2022 | STCD: A Spatio-Temporal Causal Discovery Framework for Hydrological SystemsabstractCausal learning has become an essential attribute in majority of the machine learning models. One of the widely studied fields in causal learning is causal discovery which aims to identify potential cause-effect relationships from observational data. Temporal causal discovery models are specifically curated to enforece the temporal constraints while discovering the causal relationships. However, in physical systems such as hydrological systems, there are additional constraints such as spatial constraints that play a crucial role in deciding whether a node is a causal parent for another node or not. Failing to enforce these additional constraints may mislead the model to classify an irrelevant relationship as a causal relationship. Furthermore, causal discovery models are evaluated against a ground truth causal graph. However, the hydrological systems contain a huge number of features making it challenging to obtain a ground-truth causal graph. To deal with the aforementioned problems, in this study we propose a new Spatio-Temporal Causal Discovery Framework named, STCD. By enforcing temporal and spatial constraints STCD aims at identifying meaningful causal relationships. Furthermore, to evaluate the causal relations inferred by STCD in the absence of the ground-truth causal graph, we utilize only the causal parents of a target variable for prediction across different years. We demonstrate that utilizing only the causal features identified by STCD to predict the flow-rate for a target location attains superior performance. Paras Sheth, Reepal Shah, John Sabo, K. Selçuk Candan, Huan Liu 0001 |
IEEE Big Data | 1 |
| 2022 | Causal Disentanglement with Network Information for Debiased Recommendations
Paras Sheth, Ruocheng Guo, Kaize Ding, Lu Cheng 0001, K. Selçuk Candan, Huan Liu 0001 |
SISAP | 1 |
| 2021 | CauseBox: A Causal Inference Toolbox for BenchmarkingTreatment Effect Estimators with Machine Learning MethodsabstractCausal inference is a critical task in various fields such as healthcare, economics, marketing and education. Recently, there have been significant advances through the application of machine learning techniques, especially deep neural networks. Unfortunately, to-date many of the proposed methods are evaluated on different (data, software/hardware, hyperparameter) setups and consequently it is nearly impossible to compare the efficacy of the available methods or reproduce results presented in original research manuscripts. In this paper, we propose a causal inference toolbox (CauseBox) that addresses the aforementioned problems. At the time of publication, the toolbox includes seven state of the art causal inference methods and two benchmark datasets. By providing convenient command-line and GUI-based interfaces, the CauseBox toolbox helps researchers fairly compare the state of the art methods in their chosen application context against benchmark datasets. The code is made public at github.com/paras2612/CauseBox. Paras Sheth, Ujun Jeong, Ruocheng Guo, Huan Liu 0001, K. Selçuk Candan |
CIKM | 1 |
| 2021 | Causal inference for time series analysis: problems, methods and evaluation
Raha Moraffah, Paras Sheth, Mansooreh Karami, Anchit Bhattacharya, Qianru Wang, Anique Tahir, Adrienne Raglin, Huan Liu 0001 |
Knowl. Inf. Syst. | 2 |