VLDB 2026 Research / reviewers in the wild / expert
Biswadeep Khan
dblp:139/9219
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2026
0000-0003-2568-7198ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Empirical software engineering › mining software repositories
dataset construction |
1.0 | 1 | 2026 | GONets: A First-Look Into GitHub Organisation Networks · SIGIR 2026 |
Empirical software engineering
mining software repositories |
1.0 | 1 | 2026 | GONets: A First-Look Into GitHub Organisation Networks · SIGIR 2026 |
Methods — techniques the papers use, named apart from their topics
data collection pipeline · 1.0bipartite network analysis · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GONets: A First-Look Into GitHub Organisation NetworksabstractThe rapid growth of open-source software has made GitHub a central platform for studying large-scale collaboration for software development. Existing datasets and event logs are typically limited to individual projects or user–repository collaboration at the GitHub level. We introduce GONets, a publicly available dataset that provides a direct bipartite representation of user–repository interactions at the organisational level. GONets comprises over 230K GitHub organisations, 2.52M unique users, 4.37M repositories, and nearly 30M contribution edges, enriched with detailed organisational metadata including organisation size, creation date and follower count. We present an automated data collection pipeline that retrieves organisational and membership information via GitHub's API, aggregates user-repository contribution data, and constructs large-scale bipartite networks suitable for network analysis and modelling. Using this dataset, we conduct the first large-scale characterisation of the GitHub organisational ecosystem, analysing organisational attributes and user-repository contribution patterns across a diverse range of organisations. We release the complete dataset and collection pipeline to support reproducible research and facilitate future studies on topics including code reuse diffusion, collaboration dynamics and the security ecosystem of GitHub organisations. The entire dataset can be accessed at: https://zenodo.org/records/18472549. Hridoy Sankar Dutta, Biswadeep Khan, Parth Mitesh Shah, Amit A. Nanavati |
SIGIR | 2 |
| 2025 | YTCommentVerse: A Multi-Category Multi-Lingual YouTube Comment CorpusabstractIn this paper, we introduce YTCommentVerse, a large-scale multilingual and multi-category dataset of YouTube comments. It contains over 32 million comments from 178,000 videos contributed by more than 20 million unique users spanning 15 distinct YouTube content categories such as Music, News, Education and Entertainment. Each comment in the dataset includes video and comment IDs, user channel details, upvotes and category labels. With comments in over 50 languages, YTCommentVerse provides a rich resource for exploring sentiment, toxicity and engagement patterns across diverse cultural and topical contexts. This dataset helps fill a major gap in publicly available social media datasets particularly for analyzing video sharing platforms by combining multiple languages, detailed categories and other metadata. Hridoy Sankar Dutta, Biswadeep Khan |
CIKM | 2 |