EDBT 2026 Demo / reviewers in the wild / expert
Josh Attenberg
dblp:87/1377 · also Joshua Attenberg
· DBLP profile ↗
10ranked-venue papers
4as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-authorDatabases, data management, data science and information retrieval · 8 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
7 papers |
Information retrieval · 49% Data mining · 16% Machine learning and data management · 12% | |
| Artificial intelligence
2 papers |
Efficient and distributed learning · 49% Learning paradigms · 26% Probabilistic and Bayesian machine learning · 25% |
Topics — the 23 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › ranking
learning to rank |
0.2 | 1 | 2016 | Images Don't Lie: Transferring Deep Visual Semantic Features to Large-Scale Multimodal Learning to Rank · KDD 2016 |
Recommender systems › collaborative filtering
matrix factorization |
0.2 | 1 | 2014 | Style in the long tail: discovering unique interests with latent variable models in large scale social E-commerce · KDD 2014 |
Machine learning › Efficient and distributed learning
active learning |
0.1 | 1 | 2011 | Online active inference and learning · KDD 2011 |
Machine learning › Efficient and distributed learning › active learning › active data collection
stream-based active learning |
0.1 | 1 | 2011 | Online active inference and learning · KDD 2011 |
Query processing and optimization › query execution
batch query processing |
0.1 | 1 | 2011 | Batch query processing for web search engines · WSDM 2011 |
Information retrieval
query processing |
0.1 | 1 | 2011 | Batch query processing for web search engines · WSDM 2011 |
Machine learning and data management
active learning |
0.1 | 1 | 2010 | Why label when you can search?: alternatives to active learning for applying human resources to build classification models under extreme class imbalance · KDD 2010 |
Data mining › predictive modeling
classification |
0.1 | 1 | 2010 | Why label when you can search?: alternatives to active learning for applying human resources to build classification models under extreme class imbalance · KDD 2010 |
Data mining › predictive modeling › classification
class imbalance |
0.1 | 1 | 2010 | Why label when you can search?: alternatives to active learning for applying human resources to build classification models under extreme class imbalance · KDD 2010 |
Information retrieval › indexing › inverted index
document identifier assignment |
0.1 | 1 | 2010 | Scalable techniques for document identifier assignment ininverted indexes · WWW 2010 |
Information retrieval › indexing
index compression |
0.1 | 1 | 2010 | Scalable techniques for document identifier assignment ininverted indexes · WWW 2010 |
Information retrieval › indexing
inverted index |
0.1 | 1 | 2010 | Scalable techniques for document identifier assignment ininverted indexes · WWW 2010 |
Machine learning › Learning paradigms
multi-task learning |
0.1 | 1 | 2009 | Feature hashing for large scale multitask learning · ICML 2009 |
Information retrieval › user behavior › search behavior
click model |
0.1 | 1 | 2009 | Modeling and predicting user behavior in sponsored search · KDD 2009 |
Data mining
dimensionality reduction |
0.1 | 1 | 2009 | Feature hashing for large scale multitask learning · ICML 2009 |
Machine learning and data management › feature transformation
feature hashing |
0.1 | 1 | 2009 | Feature hashing for large scale multitask learning · ICML 2009 |
Information retrieval › online advertising
sponsored search |
0.1 | 1 | 2009 | Modeling and predicting user behavior in sponsored search · KDD 2009 |
Web and social media mining › user behavior analysis
user behavior modeling |
0.1 | 1 | 2009 | Modeling and predicting user behavior in sponsored search · KDD 2009 |
Web and social media mining
e-commerce |
0.1 | 1 | 2014 | Style in the long tail: discovering unique interests with latent variable models in large scale social E-commerce · KDD 2014 |
Machine learning › Learning paradigms
cost-sensitive learning |
0.0 | 1 | 2011 | Online active inference and learning · KDD 2011 |
Information retrieval
search engines |
0.0 | 1 | 2010 | Scalable techniques for document identifier assignment ininverted indexes · WWW 2010 |
Machine learning and data management › training data management
training data collection |
0.0 | 1 | 2010 | Why label when you can search?: alternatives to active learning for applying human resources to build classification models under extreme class imbalance · KDD 2010 |
Information retrieval
ranking |
0.0 | 1 | 2009 | Modeling and predicting user behavior in sponsored search · KDD 2009 |
Methods — techniques the papers use, named apart from their topics
transfer learning · 0.2convolutional neural network · 0.2feature hashing · 0.2exponential tail bounds · 0.2latent variable model · 0.2online density estimation · 0.1decision theory · 0.1cost-sensitive learning · 0.1search-based sampling · 0.1active learning · 0.1mixture of pareto distributions · 0.1generative model · 0.1classifier · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Images Don't Lie: Transferring Deep Visual Semantic Features to Large-Scale Multimodal Learning to RankabstractSearch is at the heart of modern e-commerce. As a result, the task of ranking search results automatically (learning to rank) is a multibillion dollar machine learning problem. Traditional models optimize over a few hand-constructed features based on the item's text. In this paper, we introduce a multimodal learning to rank model that combines these traditional features with visual semantic features transferred from a deep convolutional neural network. In a large scale experiment using data from the online marketplace Etsy, we verify that moving to a multimodal representation significantly improves ranking quality. We show how image features can capture fine-grained style information not available in a text-only representation. In addition, we show concrete examples of how image information can successfully disentangle pairs of highly different items that are ranked similarly by a text-only model. Corey Lynch, Kamelia Aryafar, Josh Attenberg |
KDD | 3 |
| 2014 | Exploring User Behaviour on Etsy through Dominant ColorsabstractThe online realm has become a driving force in the retail marketplace. E-Commerce websites can provide a level of diversity and uniqueness that is impossible in the world of brick-and-mortar retail. Etsy is an online marketplace1 for artisans selling unique handcrafted goods, and vintage wares that couldn't be found elsewhere. Etsy caters to the long tail of online retail [1]. Intuitively, online retail is a visual experience-shoppers have particular styles that they find appealing, often images are used as first order information when making shopping decisions. There are a variety of signals extracted from the images representing those items for sale by shoppers. Amongst these, color composition is an important cue for visual search and image ranking-often shoppers have a palette of favorite colors, or a mental image of what they're looking for, partially determined by color. In this paper, we introduce a novel dataset for user behaviour prediction. We address the problem of inferring dominant color composition from the pixel-level color distribution of listed images on Etsy. We explore the dominant colors of favorited listings and investigate the entropy of colors distribution among Etsy users. Kamelia Aryafar, Corey Lynch, Josh Attenberg |
ICPR | 3 |
| 2014 | Style in the long tail: discovering unique interests with latent variable models in large scale social E-commerceabstractPurchasing decisions in many product categories are heavily influenced by the shopper's aesthetic preferences. It's insufficient to simply match a shopper with popular items from the category in question; a successful shopping experience also identifies products that match those aesthetics. The challenge of capturing shoppers' styles becomes more difficult as the size and diversity of the marketplace increases. At Etsy, an online marketplace for handmade and vintage goods with over 30 million diverse listings, the problem of capturing taste is particularly important -- users come to the site specifically to find items that match their eclectic styles. Diane Hu, Rob Hall 0001, Josh Attenberg |
KDD | 3 |
| 2011 | Online active inference and learningabstractWe present a generalized framework for active inference, the selective acquisition of labels for cases at prediction time in lieu of using the estimated labels of a predictive model. We develop techniques within this framework for classifying in an online setting, for example, for classifying the stream of web pages where online advertisements are being served. Stream applications present novel complications because (i) at the time of label acquisition, we don't know the set of instances that we will eventually see, (ii) instances repeat based on some unknown (and possibly skewed) distribution. We combine ideas from decision theory, cost-sensitive learning, and online density estimation. We also introduce a method for on-line estimation of the utility distribution, which allows us to manage the budget over the stream. The resulting model tells which instances to label so that by the end of each budget period, the budget is best spent (in expectation). The main results show that: (1) our proposed approach to active inference on streams can indeed reduce error costs substantially over alternative approaches, (2) more sophisticated online estimations achieve larger reductions in error. We next discuss simultaneously conducting active inference and active learning. We show that our expected-utility active inference strategy also selects good examples for learning. We close by pointing out that our utility-distribution estimation strategy can also be applied to convert pool-based active learning techniques into budget-sensitive online active learning techniques. Josh Attenberg, Foster J. Provost |
KDD | 1 |
| 2011 | Batch query processing for web search enginesabstractLarge web search engines are now processing billions of queries per day. Most of these queries are interactive in nature, requiring a response in fractions of a second. However, there are also a number of important scenarios where large batches of queries are submitted for various web mining and system optimization tasks that do not require an immediate response. Given the significant cost of executing search queries over billions of web pages, it is a natural question to ask if such batches of queries can be more efficiently executed than interactive queries. Shuai Ding 0006, Josh Attenberg, Ricardo Baeza-Yates, Torsten Suel |
WSDM | 2 |
| 2010 | Why label when you can search?: alternatives to active learning for applying human resources to build classification models under extreme class imbalanceabstractThis paper analyses alternative techniques for deploying low-cost human resources for data acquisition for classifier induction in domains exhibiting extreme class imbalance - where traditional labeling strategies, such as active learning, can be ineffective. Consider the problem of building classifiers to help brands control the content adjacent to their on-line advertisements. Although frequent enough to worry advertisers, objectionable categories are rare in the distribution of impressions encountered by most on-line advertisers - so rare that traditional sampling techniques do not find enough positive examples to train effective models. An alternative way to deploy human resources for training-data acquisition is to have them "guide" the learning by searching explicitly for training examples of each class. We show that under extreme skew, even basic techniques for guided learning completely dominate smart (active) strategies for applying human resources to select cases for labeling. Therefore, it is critical to consider the relative cost of search versus labeling, and we demonstrate the tradeoffs for different relative costs. We show that in cost/skew settings where the choice between search and active labeling is equivocal, a hybrid strategy can combine the benefits. Josh Attenberg, Foster J. Provost |
KDD | 1 |
| 2010 | A Unified Approach to Active Dual Supervision for Labeling Features and Examples
Josh Attenberg, Prem Melville, Foster J. Provost |
ECML/PKDD (1) | 1 |
| 2010 | Scalable techniques for document identifier assignment ininverted indexesabstractWeb search engines depend on the full-text inverted index data structure. Because the query processing performance is so dependent on the size of the inverted index, a plethora of research has focused on fast end effective techniques for compressing this structure. Recently, several authors have proposed techniques for improving index compression by optimizing the assignment of document identifiers to the documents in the collection, leading to significant reduction in overall index size. Shuai Ding 0006, Josh Attenberg, Torsten Suel |
WWW | 2 |
| 2009 | Feature hashing for large scale multitask learningabstractEmpirical evidence suggests that hashing is an effective strategy for dimensionality reduction and practical nonparametric estimation. In this paper we provide exponential tail bounds for feature hashing and show that the interaction between random subspaces is negligible with high probability. We demonstrate the feasibility of this approach with experimental results for a new use case --- multitask learning with hundreds of thousands of tasks. Kilian Q. Weinberger, Anirban Dasgupta 0001, John Langford 0001, Alexander J. Smola, Josh Attenberg |
ICML | 5 |
| 2009 | Modeling and predicting user behavior in sponsored searchabstractImplicit user feedback, including click-through and subsequent browsing behavior, is crucial for evaluating and improving the quality of results returned by search engines. Several recent studies [1, 2, 3, 13, 25] have used post-result browsing behavior including the sites visited, the number of clicks, and the dwell time on site in order to improve the ranking of search results. In this paper, we first study user behavior on sponsored search results (i.e., the advertisements displayed by search engines next to the organic results), and compare this behavior to that of organic results. Second, to exploit post-result user behavior for better ranking of sponsored results, we focus on identifying patterns in user behavior and predict expected on-site actions in future instances. In particular, we show how post-result behavior depends on various properties of the queries, advertisement, sites, and users, and build a classifier using properties such as these to predict certain aspects of the user behavior. Additionally, we develop a generative model to mimic trends in observed user activity using a mixture of pareto distributions. We conduct experiments based on billions of real navigation trails collected by a major search engine's browser toolbar. Josh Attenberg, Sandeep Pandey, Torsten Suel |
KDD | 1 |