EDBT 2026 Demo / reviewers in the wild / expert
Paramveer S. Dhillon
dblp:35/3993 · also Paramveer Dhillon
· DBLP profile ↗
9ranked-venue papers in the field
3as first author
5since 2021 · last 2025
0000-0002-0994-9488ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 5 (1 first)Data Mining & Knowledge Discovery · 3 (2 first)Database Systems & Data Management · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging Group-Level Signals for Robust Many-Domain Generalization
Yachuan Liu, Qiaozhu Mei, Paramveer S. Dhillon |
PAKDD (3) | 4 |
| 2025 | Recommendation and TemptationabstractPeer Reviewed Md Sanzeed Anwar, Paramveer S. Dhillon, Grant Schoenebeck |
RecSys | 2 |
| 2024 | Filter Bubble or Homogenization? Disentangling the Long-Term Effects of Recommendations on User Consumption PatternsabstractRecommendation algorithms play a pivotal role in shaping our media choices, which makes it crucial to comprehend their long-term impact on user behavior. These algorithms are often linked to two critical outcomes: homogenization, wherein users consume similar content despite disparate underlying preferences, and the filter bubble effect, wherein individuals with differing preferences only consume content aligned with their preferences (without much overlap with other users). Prior research assumes a trade-off between homogenization and filter bubble effects and then shows that personalized recommendations mitigate filter bubbles by fostering homogenization. However, because of this assumption of a tradeoff between these two effects, prior work cannot develop a more nuanced view of how recommendation systems may independently impact homogenization and filter bubble effects. We develop a more refined definition of homogenization and the filter bubble effect by decomposing them into two key metrics: how different the average consumption is between users (inter-user diversity) and how varied an individual's consumption is (intra-user diversity). We then use a novel agent-based simulation framework that enables a holistic view of the impact of recommendation systems on homogenization and filter bubble effects. Our simulations show that traditional recommendation algorithms (based on past behavior) mainly reduce filter bubbles by affecting inter-user diversity without significantly impacting intra-user diversity. Building on these findings, we introduce two new recommendation algorithms that take a more nuanced approach by accounting for both types of diversity. Md Sanzeed Anwar, Grant Schoenebeck, Paramveer S. Dhillon |
WWW | 3 |
| 2023 | Unique in What Sense? Heterogeneous Relationships between Multiple Types of Uniqueness and Popularity in MusicabstractHow does our society appreciate the uniqueness of cultural products? This fundamental puzzle has intrigued scholars in many fields, including psychology, sociology, anthropology, and marketing. It has been theorized that cultural products that balance familiarity and novelty are more likely to become popular. However, a cultural product's novelty is typically multifaceted. This paper uses songs as a case study to study the multiple facets of uniqueness and their relationship with success. We first unpack the multiple facets of a song's novelty or uniqueness and, next, measure its impact on a song's popularity. We employ a series of statistical models to study the relationship between a song's popularity and novelty associated with its lyrics, chord progressions, or audio properties. Our analyses performed on a dataset of over fifty thousand songs find a consistently negative association between all types of song novelty and popularity. Overall we found a song's lyrics uniqueness to have the most significant association with its popularity. However, audio uniqueness was the strongest predictor of a song's popularity, conditional on the song's genre. We further found the theme and repetitiveness of a song's lyrics to mediate the relationship between the song's popularity and novelty. Broadly, our results contradict the "optimal distinctiveness theory'' (balance between novelty and familiarity) and call for an investigation into the multiple dimensions along which a cultural product's uniqueness could manifest. Yulin Yu, Pui Yin Cheung, Yong-Yeol Ahn, Paramveer S. Dhillon |
ICWSM | 4 |
| 2022 | Judging a Book by Its Cover: Predicting the Marginal Impact of Title on Reddit Post Popularity
Evan Weissburg, Arya Kumar, Paramveer S. Dhillon |
ICWSM | 3 |
| 2013 | Learning to explore scientific workflow repositoriesabstractScientific workflows are gaining popularity, and repositories of workflows are starting to emerge. In this paper we describe TopicsExplorer, a data exploration approach for myExperiment.org, a collaborative platform for the exchange of scientific workflows and experimental plans. Our approach uses a variant of topic modeling with tags as features, and generates a browsable view of the repository. TopicsExplorer has been fully integrated into the open-source platform of myExperiment.org, and is available to users at www.myexperiment.org/topics. We also present our recently developed personalization component that customizes topics based on user feedback. Finally, we discuss our ongoing performance optimization efforts that make computing and managing personalized topic views of the myExperiment.org repository feasible. Julia Stoyanovich, Paramveer S. Dhillon, Susan B. Davidson, Brian Lyons |
SSDBM | 2 |
| 2011 | Semi-supervised multi-task learning of structured prediction models for web information extractionabstractExtracting information from web pages is an important problem; it has several applications such as providing improved search results and construction of databases to serve user queries. In this paper we propose a novel structured prediction method to address two important aspects of the extraction problem: (1) labeled data is available only for a small number of sites and (2) a machine learned global model does not generalize adequately well across many websites. For this purpose, we propose a weight space based graph regularization method. This method has several advantages. First, it can use unlabeled data to address the limited labeled data problem and falls in the class of graph regularization based semi-supervised learning approaches. Second, to address the generalization inadequacy of a global model, this method builds a local model for each website. Viewing the problem of building a local model for each website as a task, we learn the models for a collection of sites jointly; thus our method can also be seen as a graph regularization based multi-task learning approach. Learning the models jointly with the proposed method is very useful in two ways: (1) learning a local model for a website can be effectively influenced by labeled and unlabeled data from other websites; and (2) even for a website with only unlabeled examples it is possible to learn a decent local model. We demonstrate the efficacy of our method on several real-life data; experimental results show that significant performance improvement can be obtained by combining semi-supervised and multi-task learning in a single framework. Paramveer S. Dhillon, Sundararajan Sellamanickam, S. Sathiya Keerthi |
CIKM | 1 |
| 2009 | Multi-task Feature Selection Using the Multiple Inclusion Criterion (MIC)
Paramveer S. Dhillon, Brian Tomasik, Dean P. Foster, Lyle H. Ungar |
ECML/PKDD (1) | 1 |
| 2008 | Efficient Feature Selection in the Presence of Multiple Feature ClassesabstractWe present an information theoretic approach to feature selection when the data possesses feature classes. Feature classes are pervasive in real data. For example, in gene expression data, the genes which serve as features may be divided into classes based on their membership in gene families or pathways. When doing word sense disambiguation or named entity extraction, features fall into classes including adjacent words, their parts of speech, and the topic and venue of the document the word is in. When predictive features occur predominantly in a small number of feature classes, our information theoretic approach significantly improves feature selection. Experiments on real and synthetic data demonstrate substantial improvement in predictive accuracy over the standard L0penalty-based stepwise and stream wise feature selection methods as well as over Lasso and Elastic Nets, all of which are oblivious to the existence of feature classes. Paramveer S. Dhillon, Dean P. Foster, Lyle H. Ungar |
ICDM | 1 |