EDBT 2026 Demo / reviewers in the wild / expert
Hridoy Sankar Dutta
dblp:148/4303
· DBLP profile ↗
14ranked-venue papers
8as first author
8since 2021 · last 2026
0000-0002-4254-8299ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 6 first-author · 5 since 2021Security and privacy · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-author · 2 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GONets: A First-Look Into GitHub Organisation NetworksabstractThe rapid growth of open-source software has made GitHub a central platform for studying large-scale collaboration for software development. Existing datasets and event logs are typically limited to individual projects or user–repository collaboration at the GitHub level. We introduce GONets, a publicly available dataset that provides a direct bipartite representation of user–repository interactions at the organisational level. GONets comprises over 230K GitHub organisations, 2.52M unique users, 4.37M repositories, and nearly 30M contribution edges, enriched with detailed organisational metadata including organisation size, creation date and follower count. We present an automated data collection pipeline that retrieves organisational and membership information via GitHub's API, aggregates user-repository contribution data, and constructs large-scale bipartite networks suitable for network analysis and modelling. Using this dataset, we conduct the first large-scale characterisation of the GitHub organisational ecosystem, analysing organisational attributes and user-repository contribution patterns across a diverse range of organisations. We release the complete dataset and collection pipeline to support reproducible research and facilitate future studies on topics including code reuse diffusion, collaboration dynamics and the security ecosystem of GitHub organisations. The entire dataset can be accessed at: https://zenodo.org/records/18472549. Hridoy Sankar Dutta, Biswadeep Khan, Parth Mitesh Shah, Amit A. Nanavati |
SIGIR | 1 |
| 2026 | iOSModZoo: A Large-Scale Study of Third-Party iOS App Markets
Luis Adan Saavedra, Hridoy Sankar Dutta, Alastair R. Beresford, Alice Hutchings |
WISEC | 2 |
| 2026 | Rethinking Targeted Data Poisoning in Voice Authentication: A Critique and Defense MechanismabstractRecent deep learning techniques have significantly improved voice authentication systems. However, they remain vulnerable to threats, including data poisoning and Man-in-the-Middle (MitM) attacks. This paper reevaluates the attack and defense mechanisms of “The Guardian”, examining their assumptions and modifying the proposal. We conduct experiments with real-world datasets to validate their effectiveness under realistic conditions and assess the feasibility of executing such attacks. Additionally, we introduce a defense mechanism that improves resilience after redefining a threat model grounded in operational feasibility, specifically by isolating the enrollment phase from training phase assumptions. By analyzing the current literature, we identify open challenges and suggest directions for further improving the security of voice authentication systems. Kamel Kamel, Keshav Sood, Hridoy Sankar Dutta, Sunil Aryal |
IEEE Internet Things J. | 3 |
| 2025 | App-solutely Modded: Surveying Modded App Market Operators and Original App DevelopersabstractApp-solutely Modded: Surveying Modded App Market Operators and Original App Developers Luis Adan Saavedra, Hridoy Sankar Dutta, Alastair R. Beresford, Alice Hutchings |
AsiaCCS | 2 |
| 2025 | YTCommentVerse: A Multi-Category Multi-Lingual YouTube Comment CorpusabstractIn this paper, we introduce YTCommentVerse, a large-scale multilingual and multi-category dataset of YouTube comments. It contains over 32 million comments from 178,000 videos contributed by more than 20 million unique users spanning 15 distinct YouTube content categories such as Music, News, Education and Entertainment. Each comment in the dataset includes video and comment IDs, user channel details, upvotes and category labels. With comments in over 50 languages, YTCommentVerse provides a rich resource for exploring sentiment, toxicity and engagement patterns across diverse cultural and topical contexts. This dataset helps fill a major gap in publicly available social media datasets particularly for analyzing video sharing platforms by combining multiple languages, detailed categories and other metadata. Hridoy Sankar Dutta, Biswadeep Khan |
CIKM | 1 |
| 2022 | Weakening the Inner Strength: Spotting Core Collusive Users in YouTube Blackmarket Network
Hridoy Sankar Dutta, Nirav Diwan, Tanmoy Chakraborty 0002 |
ICWSM | 1 |
| 2021 | ABOME: A Multi-platform Data Repository of Artificially Boosted Online Media Entities
Hridoy Sankar Dutta, Udit Arora, Tanmoy Chakraborty 0002 |
ICWSM | 1 |
| 2021 | Detecting and Analyzing Collusive Entities on YouTubeabstractYouTube sells advertisements on the posted videos, which in turn enables the content creators to monetize their videos. As an unintended consequence, this has proliferated various illegal activities such as artificial boosting of views, likes, comments, and subscriptions. We refer to such videos (gaining likes and comments artificially) and channels (gaining subscriptions artificially) as “collusive entities.” Detecting such collusive entities is an important yet challenging task. Existing solutions mostly deal with the problem of spotting fake views, spam comments, fake content, and so on, and oftentimes ignore how such fake activities emerge via collusion. Here, we collect a large dataset consisting of two types of collusive entities on YouTube— videos submitted to gain collusive likes and comment requests and channels submitted to gain collusive subscriptions. We begin by providing an in-depth analysis of collusive entities on YouTube fostered by various blackmarket services . Following this, we propose models to detect three types of collusive YouTube entities: videos seeking collusive likes, channels seeking collusive subscriptions, and videos seeking collusive comments. The third type of entity is associated with temporal information. To detect videos and channels for collusive likes and subscriptions, respectively, we utilize one-class classifiers trained on our curated collusive entities and a set of novel features. The SVM-based model shows significant performance with a true positive rate of 0.911 and 0.910 for detecting collusive videos and collusive channels, respectively. To detect videos seeking collusive comments, we propose CollATe , a novel end-to-end neural architecture that leverages time-series information of posted comments along with static metadata of videos. CollATe is composed of three components: metadata feature extractor (which derives metadata-based features from videos), anomaly feature extractor (which utilizes the time-series data to detect sudden changes in the commenting activity), and comment feature extractor (which utilizes the text of the comments posted during collusion and computes a similarity score between the comments). Extensive experiments show the effectiveness of CollATe (with a true positive rate of 0.905) over the baselines. Hridoy Sankar Dutta, Mayank Jobanputra, Himani Negi, Tanmoy Chakraborty 0002 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2020 | Blackmarket-Driven Collusion Among Retweeters-Analysis, Detection, and CharacterizationabstractThe growth of online social media has led to a huge increase in the number of users who want to share and publicize various kinds of information. Twitter, the most popular micro-blogging platform, has become a hotbed for users who are involved in different activities such as news publishing, job hunting, recruiting, advertising and publicity. Retweeting a tweet is a major action to broadcast a user's message out to millions of users. Retweet action has two major advantages: (i) gaining quick exposure to the content, and (ii) increasing likelihood of gaining new Twitter followers in return. The organic way of gaining a larger number of retweets is a time consuming process, which leads to the creation of unfair methods to gain retweets. Thus, Twitter users often approach various blackmarket services to gain retweets inorganically in a short duration. Blackmarkets spread their collusive ecosystem in such a way that Twitter is unable to detect them even after devoting significant effort to purge the platform off bots, trolls, and fake accounts. One major reason behind the evasion is that the collusive users involved in blackmarket services exhibit a mix of organic and inorganic behavior - they organically reweet some genuine tweets; at the same time, they inorganically retweet tweets submitted to blackmarket services. This paper is the first attempt to provide a thorough study of the collusive users involved in two types of blackmarket services - Premium and Freemium. We collect a novel dataset of collusive users comprising of users from both types of blackmarket services. We provide network-centric, profile-centric, timeline-centric and retweet-centric characteristics of these users and show how users involved in premium blackmarket services exhibit diverse behavior as compared to those involved in freemium services. We further employ human annotators to label collusive users into three types: bots, promotional customers, and normal customers. We then curate 63 novel features to run state-of-the-art classifiers in two settings - binary classification (collusive vs. genuine) and multi-class classification (bot, promotional customers, normal customers, and genuine users). Bagging achieves the best accuracy (macro F1-score of 0.892) in the former setting, whereas Random Forest outperforms others (macro F1-score of 0.791) in the latter setting. We also develop a chrome extension, SCoRe++ which can detect collusive retweeters in real time. Hridoy Sankar Dutta, Tanmoy Chakraborty 0002 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | HawkesEye: Detecting Fake Retweeters Using Hawkes Process and Topic ModelingabstractRetweets are essential to boost the popularity of a tweet, and a large number of fake retweeters can contribute heavily to this aspect. We define a fake retweeter as a Twitter account that retweets spammy tweets, retweets an abnormally large amount of tweets in a short period, or misuses a trending hashtag to promote events irrelevant to the topic of discussion. We introduce an up-to-date, temporally diverse, trend-oriented labeled dataset to address the problem of fake retweeter detection. We develop a novel classifier, called HawkesEye which makes predictions based on a temporal window, in contrast to existing approaches which require agraph-likerelationship between tweet entities, or the presence of theentire retweeting timelineof a retweeter. HawkesEye utilizes both temporal and textual information using a class-specific topic model and Hawkes processes. Experiments on our curated dataset show significant improvement over four state-of-the-art methods, with precision and recall scores of 0.964 and 0.960 on a balanced dataset, respectively – HawkesEye beats the best baseline by 6.16% and 25.98% relative improvement in terms of precision and recall, respectively. We also diagnose our model to understand the advantages and pitfalls of the underlying mechanism. We believe that the extent of this study is not restricted to Twitter, but generalizable to other social media systems such as Facebook and Instagram with similar reposting capabilities. Hridoy Sankar Dutta, Vishal Raj Dutta, Aditya Adhikary, Tanmoy Chakraborty 0002 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | Analyzing and Detecting Collusive Users Involved in Blackmarket Retweeting ActivitiesabstractWith the rise in popularity of social media platforms like Twitter, having higher influence on these platforms has a greater value attached to it, since it has the power to influence many decisions in the form of brand promotions and shaping opinions. However, blackmarket services that allow users to inorganically gain influence are a threat to the credibility of these social networking platforms. Twitter users can gain inorganic appraisals in the form of likes, retweets, and follows through these blackmarket services either by paying for them or by joining syndicates wherein they gain such appraisals by providing similar appraisals to other users. These customers tend to exhibit a mix of organic and inorganic retweeting behavior, making it tougher to detect them. In this article, we investigate these blackmarket customers engaged in collusive retweeting activities. We collect and annotate a novel dataset containing various types of information about blackmarket customers and use these sources of information to construct multiple user representations. We adopt Weighted Generalized Canonical Correlation Analysis (WGCCA) to combine these individual representations to derive user embeddings that allow us to effectively classify users as: genuine users, bots, promotional customers, and normal customers. Our method significantly outperforms state-of-the-art approaches (32.95% better macro F1-score than the best baseline). Udit Arora, Hridoy Sankar Dutta, Brihi Joshi, Aditya Chetan, Tanmoy Chakraborty 0002 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2019 | CoReRank: Ranking to Detect Users Involved in Blackmarket-Based Collusive Retweeting ActivitiesabstractTwitter's popularity has fostered the emergence of various illegal user activities - one such activity is to artificially bolster visibility of tweets by gaining large number of retweets within a short time span. The natural way to gain visibility is time-consuming. Therefore, users who want their tweets to get quick visibility try to explore shortcuts - one such shortcut is to approach the blackmarket services, and gain retweets for their own tweets by retweeting other customers' tweets. Thus the users intrinsically become a part of a collusive ecosystem controlled by these services. In this paper, we propose CoReRank, an unsupervised framework to detect collusive users (who are involved in producing artificial retweets), and suspicious tweets (which are submitted to the blackmarket services) simultaneously. CoReRank leverages the retweeting (or quoting) patterns of users, and measures two scores - the 'credibility' of a user and the 'merit' of a tweet. We propose a set of axioms to derive the interdependency between these two scores, and update them in a recursive manner. The formulation is further extended to handle the cold start problem. CoReRank is guaranteed to converge in a finite number of iterations and has linear time complexity. We also propose a semi-supervised version of CoReRank (called CoReRank+) which leverages a partial ground-truth labeling of users and tweets. Extensive experiments are conducted to show the superiority of CoReRank compared to six baselines on a novel dataset we collected and annotated. CoReRank beats the best unsupervised baseline method by 269% (20%) (relative) average precision and 300% (22.22%) (relative) average recall in detecting collusive (genuine) users. CoReRank+ beats the best supervised baseline method by 33.18% AUC. CoReRank also detects suspicious tweets with 0.85 (0.60) average precision (recall). To our knowledge, CoReRank is the first unsupervised method to detect collusive users and suspicious tweets simultaneously with theoretical guarantees. Aditya Chetan, Brihi Joshi, Hridoy Sankar Dutta, Tanmoy Chakraborty 0002 |
WSDM | 3 |
| 2019 | CRIMP: Here crisis mapping goes offline
Partha Sarathi Paul 0001, Bishakh Chandra Ghosh, Hridoy Sankar Dutta, Kingshuk De, Arka Prava Basu, Prithviraj Pramanik, Sujoy Saha, Sandip Chakraborty 0001, Niloy Ganguly, Subrata Nandi |
J. Netw. Comput. Appl. | 3 |
| 2018 | Retweet Us, We will Retweet You: Spotting Collusive Retweeters Involved in Blackmarket ServicesabstractTwitter has increasingly become a popular platform to share news and user opinion. A tweet is considered to be important if it receives high number of affirmative reactions from other Twitter users via Retweets. Retweet count is thus considered as a surrogate measure for positive crowd-sourced reactions - high number of retweets of a tweet not only help the tweet being broadcasted, but also aid in making its topic trending. This in turn bolsters the social reputation of the author of the tweet. Since social reputation/impact of users/t weets influences many decisions (such as promoting brands, advertisement etc.), several blackmarket syndicates have actively been engaged in producing fake retweets in a collusive manner. Users who want to boost the impact of their tweets approach the blackmarket services, and gain retweets for their own tweets by retweeting other customers' tweets. Thus they become customers of blackmarket syndicates and engage in fake activities. Interestingly, these customers are neither bots, nor even fake users - they are usually normal human beings; they express a mix of organic and inorganic retweeting activities, and there is no synchronicity across their behaviors. In this paper, we make a first attempt to investigate such blackmarket customers engaged in producing fake retweets. We collected and annotated a novel dataset comprising of customers of many blackmarket services and characterize them using a set of 64 novel features. We show how their social behavior differs from genuine users. We then use state-of-the-art supervised models to detect three types of customers (bots, promotional, normal) and genuine users. We achieve a Macro Fl-score of 0.87 with SVM, outperforming four other baselines significantly. We further design a browser extension, SCoRe which, given the link of a tweet, spots its fake retweeters in real-time. We also collected users' feedback on the performance of SCoRe and obtained 85% accuracy. Hridoy Sankar Dutta, Aditya Chetan, Brihi Joshi, Tanmoy Chakraborty 0002 |
ASONAM | 1 |