VLDB 2026 Research / reviewers in the wild / expert
Sourav Chatterji
dblp:50/65
· DBLP profile ↗
5ranked-venue papers
3as first author
1since 2021 · last 2023
0009-0004-2525-4506ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Recommender systems · 58% Web and social media mining · 38% Query processing and optimization · 4% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% | |
| Theoretical computer science
1 paper |
Computational complexity · 77% Graph algorithms and graph theory · 23% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Recommender systems
cold-start recommendation |
0.2 | 1 | 2016 | Discovery of Topical Authorities in Instagram · WWW 2016 |
Web and social media mining
social network analysis |
0.2 | 1 | 2016 | Discovery of Topical Authorities in Instagram · WWW 2016 |
Recommender systems
user recommendation |
0.2 | 1 | 2016 | Discovery of Topical Authorities in Instagram · WWW 2016 |
Bioinformatics and computational biology › metagenomics
binning |
0.1 | 1 | 2008 | CompostBin: A DNA Composition-Based Algorithm for Binning Environmental Shotgun Reads · RECOMB 2008 |
Bioinformatics and computational biology
metagenomics |
0.1 | 1 | 2008 | CompostBin: A DNA Composition-Based Algorithm for Binning Environmental Shotgun Reads · RECOMB 2008 |
Bioinformatics and computational biology › genomics › DNA sequencing
shotgun sequencing |
0.1 | 1 | 2008 | CompostBin: A DNA Composition-Based Algorithm for Binning Environmental Shotgun Reads · RECOMB 2008 |
Web and social media mining › online social networks
instagram |
0.1 | 1 | 2016 | Discovery of Topical Authorities in Instagram · WWW 2016 |
Bioinformatics and computational biology › genome annotation
gene prediction |
0.0 | 1 | 2004 | Multiple organism gene finding by collapsed gibbs sampling · RECOMB 2004 |
Bioinformatics and computational biology
gibbs sampling |
0.0 | 1 | 2004 | Multiple organism gene finding by collapsed gibbs sampling · RECOMB 2004 |
Bioinformatics and computational biology
sequence analysis |
0.0 | 1 | 2004 | Multiple organism gene finding by collapsed gibbs sampling · RECOMB 2004 |
Query processing and optimization › query optimization
join ordering |
0.0 | 1 | 2002 | On the Complexity of Approximate Query Optimization · PODS 2002 |
Computational complexity
hardness of approximation |
0.0 | 1 | 2002 | On the Complexity of Approximate Query Optimization · PODS 2002 |
Methods — techniques the papers use, named apart from their topics
wikipedia grounding · 0.2label propagation · 0.2DNA composition analysis · 0.1gibbs sampling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Making Data Engineering Declarative
Michael Armbrust, Ali Ghodsi 0002, Reynold Xin, Vuk Ercegovac, Sourav Chatterji, Eun-Gyu Kim, Paul Lappas, Yannis Papakonstantinou, Yingyi Bu, Yijia Cui, Rahul Govind, Aakash Japi, Kiavash Kianfar, Jon Mio, Mukul Murthy, Supun Nakandala, Yannis Sismanis, Justin Tang, Joseph Torres |
CIDR | 5 |
| 2016 | Discovery of Topical Authorities in InstagramabstractInstagram has more than 400 million monthly active accounts who share more than 80 million pictures and videos daily. This large volume of user-generated content is the application's notable strength, but also makes the problem of finding the authoritative users for a given topic challenging. Discovering topical authorities can be useful for providing relevant recommendations to the users. In addition, it can aid in building a catalog of topics and top topical authorities in order to engage new users, and hence provide a solution to the cold-start problem. In this paper, we present a novel approach that we call the Authority Learning Framework (ALF) to find topical authorities in Instagram. ALF is based on the self-described interests of the follower base of popular accounts. We infer regular users' interests from their self-reported biographies that are publicly available and use Wikipedia pages to ground these interests as fine-grained, disambiguated concepts. We propose a generalized label propagation algorithm to propagate the interests over the follower graph to the popular accounts. We show that even if biography-based interests are sparse at an individual user level they provide strong signals to infer the topical authorities and let us obtain a high precision authority list per topic. Our experiments demonstrate that ALF performs significantly better at user recommendation task compared to fine-tuned and competitive methods, via controlled experiments, in-the-wild tests, and over an expert-curated list of topical authorities. Aditya Pal, Amac Herdagdelen, Sourav Chatterji, Sumit Taank, Deepayan Chakrabarti |
WWW | 3 |
| 2008 | CompostBin: A DNA Composition-Based Algorithm for Binning Environmental Shotgun Reads
Sourav Chatterji, Ichitaro Yamazaki, Zhaojun Bai, Jonathan A. Eisen |
RECOMB | 1 |
| 2004 | Multiple organism gene finding by collapsed gibbs samplingabstractThe Gibbs sampling method has been widely used for sequence analysis after it was successfully applied to the problem of identifying regulatory motif sequences upstream of genes. Since then numerous variants of the original idea have emerged, however in all cases the application has been to finding short motifs in collections of short sequences (typically less than 100 nucleotides long). In this paper we introduce a Gibbs sampling approach for identifying genes in multiple large genomic sequences up to hundreds of kilobases long. This approach leverages the evolutionary relationships between the sequences to improve the gene predictions, without explicitly aligning the sequences. We have applied our method to the analysis of genomic sequence from 14 genomic regions, totaling roughly 1.8Mb of sequence in each organism. We show that our approach compares favorably with existing ab-initio approaches to gene finding, including pairwise comparison based gene prediction methods which make explicit use of alignments. Furthermore, excellent performance can be obtained with as little as 4 organisms, and the method overcomes a number of difficulties of previous comparison based gene finding approaches: it is robust with respect to genomic rearrangements, can work with draft sequence, and is fast (linear in the number and length of the sequences). It can also be seamlessly integrated with Gibbs sampling motif detection methods. Sourav Chatterji, Lior Pachter |
RECOMB | 1 |
| 2002 | On the Complexity of Approximate Query OptimizationabstractIn this work, we study the complexity of the problem of approximate query optimization. We show that, for any δ > 0, the problem of finding a join order sequence whose cost is within a factor 2Θ(log1-δ(K)) of K, where K is the cost of the optimal join order sequence is NP-Hard. The complexity gap remains if the number of edges in the query graph is constrained to be a given function e(n) of the number of vertices n of the query graph, where n(n - 1)/2 - Θ(nτ) ≥ e(n) ≥ n + Θ(nτ) and τ is any constant between 0 and 1. These results show that, unless P=NP, the query optimization problem cannot be approximately solved by an algorithm that runs in polynomial time and has a competitive ratio that is within some polylogarithmic factor of the optimal cost. Sourav Chatterji, Sai Surya Kiran Evani, Sumit Ganguly, Mahesh Datt Yemmanuru |
PODS | 1 |