EDBT 2026 Demo / reviewers in the wild / expert
Diane Hu
dblp:26/9953
· DBLP profile ↗
11ranked-venue papers
4as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
6 papers |
Recommender systems · 60% Information retrieval · 32% Data mining · 7% | |
| Artificial intelligence
4 papers |
Trustworthy machine learning · 44% Reinforcement learning · 26% Kernel, tree and ensemble methods · 19% | |
| Theoretical computer science
2 papers |
Information theory · 48% Mathematical optimization · 28% Algorithmic game theory and mechanism design · 24% | |
| Software engineering, system software, and programming languages
2 papers |
Empirical software engineering · 68% Software maintenance and evolution · 32% |
Topics — the 24 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Recommender systems › sequential recommendation
next-item recommendation |
0.9 | 2 | 2020 | Attentive Sequential Models of Latent Intent for Next Item Recommendation · WWW 2020 Time to Shop for Valentine's Day: Shopping Occasions and Sequential Recommendation in E-commerce · WSDM 2020 |
Recommender systems
sequential recommendation |
0.9 | 2 | 2020 | Attentive Sequential Models of Latent Intent for Next Item Recommendation · WWW 2020 Time to Shop for Valentine's Day: Shopping Occasions and Sequential Recommendation in E-commerce · WSDM 2020 |
Machine learning › Reinforcement learning
multi-objective reinforcement learning |
0.6 | 1 | 2022 | Toward Pareto Efficient Fairness-Utility Trade-off in Recommendation through Reinforcement Learning · WSDM 2022 |
Recommender systems
fairness-aware recommendation |
0.6 | 1 | 2022 | Toward Pareto Efficient Fairness-Utility Trade-off in Recommendation through Reinforcement Learning · WSDM 2022 |
Machine learning › Kernel, tree and ensemble methods › decision tree learning
decision tree optimization |
0.4 | 1 | 2020 | Generalized and Scalable Optimal Sparse Decision Trees · ICML 2020 |
Machine learning › Trustworthy machine learning
interpretability |
0.4 | 1 | 2020 | Generalized and Scalable Optimal Sparse Decision Trees · ICML 2020 |
Machine learning › Trustworthy machine learning › interpretability › sparse decision trees
optimal sparse decision trees |
0.4 | 1 | 2020 | Generalized and Scalable Optimal Sparse Decision Trees · ICML 2020 |
Mathematical optimization
combinatorial optimization |
0.4 | 1 | 2020 | Generalized and Scalable Optimal Sparse Decision Trees · ICML 2020 |
Algorithmic game theory and mechanism design
multi-armed bandit |
0.4 | 1 | 2019 | A Sequential Test for Selecting the Better Variant: Online A/B testing, Adaptive Allocation, and Continuous Monitoring · WSDM 2019 |
Information theory › statistical inference
sequential analysis |
0.4 | 1 | 2019 | A Sequential Test for Selecting the Better Variant: Online A/B testing, Adaptive Allocation, and Continuous Monitoring · WSDM 2019 |
Information theory › hypothesis testing
sequential hypothesis testing |
0.4 | 1 | 2019 | A Sequential Test for Selecting the Better Variant: Online A/B testing, Adaptive Allocation, and Continuous Monitoring · WSDM 2019 |
Information retrieval
e-commerce search |
0.3 | 1 | 2018 | Turning Clicks into Purchases: Revenue Optimization for Product Search in E-Commerce · SIGIR 2018 |
Information retrieval › ranking
learning to rank |
0.3 | 1 | 2018 | Turning Clicks into Purchases: Revenue Optimization for Product Search in E-Commerce · SIGIR 2018 |
Information retrieval
ranking |
0.3 | 1 | 2018 | Turning Clicks into Purchases: Revenue Optimization for Product Search in E-Commerce · SIGIR 2018 |
Information retrieval › online advertising
revenue optimization |
0.3 | 1 | 2018 | Turning Clicks into Purchases: Revenue Optimization for Product Search in E-Commerce · SIGIR 2018 |
Information retrieval
web search |
0.3 | 1 | 2018 | Turning Clicks into Purchases: Revenue Optimization for Product Search in E-Commerce · SIGIR 2018 |
Recommender systems › collaborative filtering
matrix factorization |
0.2 | 1 | 2014 | Style in the long tail: discovering unique interests with latent variable models in large scale social E-commerce · KDD 2014 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.1 | 1 | 2020 | Attentive Sequential Models of Latent Intent for Next Item Recommendation · WWW 2020 |
Machine learning › Trustworthy machine learning › interpretability
explainable AI |
0.1 | 1 | 2020 | Generalized and Scalable Optimal Sparse Decision Trees · ICML 2020 |
Empirical software engineering › controlled experiment › online controlled experiments
a/b testing |
0.1 | 1 | 2019 | A Sequential Test for Selecting the Better Variant: Online A/B testing, Adaptive Allocation, and Continuous Monitoring · WSDM 2019 |
Empirical software engineering › controlled experiment
online controlled experiments |
0.1 | 1 | 2019 | A Sequential Test for Selecting the Better Variant: Online A/B testing, Adaptive Allocation, and Continuous Monitoring · WSDM 2019 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.1 | 1 | 2010 | Latent Variable Models for Predicting File Dependencies in Large-Scale Software Development · NIPS 2010 |
Recommender systems › collaborative filtering
implicit feedback |
0.1 | 1 | 2018 | Turning Clicks into Purchases: Revenue Optimization for Product Search in E-Commerce · SIGIR 2018 |
Web and social media mining
e-commerce |
0.1 | 1 | 2014 | Style in the long tail: discovering unique interests with latent variable models in large scale social E-commerce · KDD 2014 |
Methods — techniques the papers use, named apart from their topics
deep deterministic policy gradient · 1.1conditioned network · 1.1temporal convolutional network · 0.9self-attention · 0.9dynamic programming · 0.9branch-and-bound · 0.9thompson sampling · 0.8sequential girshick test · 0.8imputation · 0.8gating layer · 0.4attention mechanism · 0.4style-aware embeddings · 0.4deep neural network · 0.4learning to rank · 0.3implicit feedback modeling · 0.3exponential family PCA · 0.2binary matrix completion · 0.2bernoulli mixture model · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Toward Pareto Efficient Fairness-Utility Trade-off in Recommendation through Reinforcement LearningabstractThe issue of fairness in recommendation is becoming increasingly essential as Recommender Systems (RS) touch and influence more and more people in their daily lives. In fairness-aware recommendation, most of the existing algorithmic approaches mainly aim at solving a constrained optimization problem by imposing a constraint on the level of fairness while optimizing the main recommendation objective, e.g., click through rate (CTR). While this alleviates the impact of unfair recommendations, the expected return of an approach may significantly compromise the recommendation accuracy due to the inherent trade-off between fairness and utility. This motivates us to deal with these conflicting objectives and explore the optimal trade-off between them in recommendation. One conspicuous approach is to seek aPareto efficient/optimal solution to guarantee optimal compromises between utility and fairness. Moreover, considering the needs of real-world e-commerce platforms, it would be more desirable if we can generalize the wholePareto Frontier, so that the decision-makers can specify any preference of one objective over another based on their current business needs. Therefore, in this work, we propose a fairness-aware recommendation framework usingmulti-objective reinforcement learning (MORL), called MoFIR (pronounced "more fair ''), which is able to learn a single parametric representation for optimal recommendation policies over the space of all possible preferences. Specially, we modify traditional Deep Deterministic Policy Gradient (DDPG) by introducingconditioned network (CN) into it, which conditions the networks directly on these preferences and outputs Q-value-vectors. Experiments on several real-world recommendation datasets verify the superiority of our framework on both fairness metrics and recommendation measures when compared with all other baselines. We also extract the approximate Pareto Frontier on real-world datasets generated by MoFIR and compare to state-of-the-art fairness methods. Yingqiang Ge, Xiaoting Zhao, Lucia Yu, Saurabh Paul, Diane Hu, Chu-Cheng Hsieh, Yongfeng Zhang 0003 |
WSDM | 5 |
| 2020 | Generalized and Scalable Optimal Sparse Decision TreesabstractDecision tree optimization is notoriously difficult from a computational perspective but essential for the field of interpretable machine learning. Despite efforts over the past 40 years, only recently have optimization breakthroughs been made that have allowed practical algorithms to find optimal decision trees. These new techniques have the potential to trigger a paradigm shift, where, it is possible to construct sparse decision trees to efficiently optimize a variety of objective functions, without relying on greedy splitting and pruning heuristics that often lead to suboptimal solutions. The contribution in this work is to provide a general framework for decision tree optimization that addresses the two significant open problems in the area: treatment of imbalanced data and fully optimizing over continuous variables. We present techniques that produce optimal decision trees over variety of objectives including F-score, AUC, and partial area under the ROC convex hull. We also introduce a scalable algorithm that produces provably optimal results in the presence of continuous variables and speeds up decision tree construction by several order of magnitude relative to the state-of-the art. Jimmy Lin, Chudi Zhong, Diane Hu, Cynthia Rudin, Margo I. Seltzer |
ICML | 3 |
| 2020 | Time to Shop for Valentine's Day: Shopping Occasions and Sequential Recommendation in E-commerceabstractCurrently, most sequence-based recommendation models aim to predict a user's next actions (e.g. next purchase) based on their past actions. These models either capture users' intrinsic preference (e.g. a comedy lover, or a fan of fantasy) from their long-term behavior patterns or infer their current needs by emphasizing recent actions. However, in e-commerce, intrinsic user behavior may be shifted by occasions such as birthdays, anniversaries, or gifting celebrations (Valentine's Day or Mother's Day), leading to purchases that deviate from long-term preferences and are not related to recent actions. In this work, we propose a novel next-item recommendation system which models a user's default, intrinsic preference, as well as two different kinds of occasion-based signals that may cause users to deviate from their normal behavior. More specifically, this model is novel in that it: (1) captures a personal occasion signal using an attention layer that models reoccurring occasions specific to that user (e.g. a birthday); (2) captures a global occasion signal using an attention layer that models seasonal or reoccurring occasions for many users (e.g. Christmas); (3) balances the user's intrinsic preferences with the personal and global occasion signals for different users at different timestamps with a gating layer. We explore two real-world e-commerce datasets (Amazon and Etsy) and show that the proposed model outperforms state-of-the-art models by 7.62% and 6.06% in predicting users' next purchase. Jianling Wang, Raphael Louca, Diane Hu, Caitlin Cellier, James Caverlee, Liangjie Hong |
WSDM | 3 |
| 2020 | Attentive Sequential Models of Latent Intent for Next Item RecommendationabstractUsers exhibit different intents across e-commerce services (e.g. discovering items, purchasing gifts, etc.) which drives them to interact with a wide variety of items in multiple ways (e.g. click, add-to-cart, add-to-favorites, purchase). To give better recommendations, it is important to capture user intent, in addition to considering their historic interactions. However these intents are by definition latent, as we observe only a user’s interactions, and not their underlying intent. To discover such latent intents, and use them effectively for recommendation, in this paper we propose an Attentive Sequential model of Latent Intent (ASLI in short). Our model first learns item similarities from users’ interaction histories via a self-attention layer, then uses a Temporal Convolutional Network layer to obtain a latent representation of the user’s intent from her actions on a particular category. We use this representation to guide an attentive model to predict the next item. Results from our experiments show that our model can capture the dynamics of user behavior and preferences, leading to state-of-the-art performance across datasets from two major e-commerce platforms, namely Etsy and Alibaba. Md. Mehrab Tanjim, Congzhe Su, Ethan Benjamin, Diane Hu, Liangjie Hong, Julian J. McAuley |
WWW | 4 |
| 2019 | Understanding the Role of Style in E-commerce ShoppingabstractAesthetic style is the crux of many purchasing decisions. When considering an item for purchase, buyers need to be aligned not only with the functional aspects (e.g. description, category, ratings) of an item's specification, but also its stylistic and aesthetic aspects (e.g. modern, classical, retro) as well. Style becomes increasingly important on e-commerce sites like Etsy, an online marketplace for handmade and vintage goods, where hundreds of thousands of items can differ by style and aesthetic alone. As such, it is important for industry recommender systems to properly model style when understanding shoppers' buying preference. In past work, because of its abstract nature, style is often approached in an unsupervised manner, represented by nameless latent factors or embeddings. As a result, there has been no previous work on predictive models nor analysis devoted to understanding how style, or even the presence of style, impacts a buyer's purchase decision. In this paper, we discuss a novel process by which we leverage 43 named styles given by merchandising experts in order to bootstrap large-scale style prediction and analysis of how style impacts purchase decision. We train a supervised, style-aware deep neural network that is shown to predict item style with high accuracy, while generating style-aware embeddings that can be used in downstream recommendation tasks. We share in our analysis, based on over a year's worth of transaction data and show that these findings are crucial to understanding how to more explicitly leverage style signal in industry-scale recommender systems. Aakash Sabharwal, Adam Henderson, Diane Hu, Liangjie Hong |
KDD | 4 |
| 2019 | A Sequential Test for Selecting the Better Variant: Online A/B testing, Adaptive Allocation, and Continuous MonitoringabstractOnline A/B tests play an instrumental role for Internet companies to improve products and technologies in a data-driven manner. An online A/B test, in its most straightforward form, can be treated as a static hypothesis test where traditional statistical tools such as p-values and power analysis might be applied to help decision makers determine which variant performs better. However, a static A/B test presents both time cost and the opportunity cost for rapid product iterations. For time cost, a fast-paced product evolution pushes its shareholders to consistently monitor results from online A/B experiments, which usually invites peeking and altering experimental designs as data collected. It is recognized that this flexibility might harm statistical guarantees if not introduced in the right way, especially when online tests are considered as static hypothesis tests. For opportunity cost, a static test usually entails a static allocation of users into different variants, which prevents an immediate roll-out of the better version to larger audience or risks of alienating users who may suffer from a bad experience. While some works try to tackle these challenges, no prior method focuses on a holistic solution to both issues. In this paper, we propose a unified framework utilizing sequential analysis and multi-armed bandit to address time cost and the opportunity cost of static online tests simultaneously. In particular, we present an imputed sequential Girshick test that accommodates online data and dynamic allocation of data. The unobserved potential outcomes are treated as missing data and are imputed using empirical averages. Focusing on the binomial model, we demonstrate that the proposed imputed Girshick test achieves Type-I error and power control with both a fixed allocation ratio and an adaptive allocation such as Thompson Sampling through extensive experiments. In addition, we also run experiments on historical Etsy.com A/B tests to show the reduction in opportunity cost when using the proposed method. Nianqiao Ju, Diane Hu, Adam Henderson, Liangjie Hong |
WSDM | 2 |
| 2018 | Learning within-session budgets from browsing trajectoriesabstractBuilding price- and budget-aware recommender systems is critical in settings where one wishes to produce recommendations that balance users' preferences (what they like) with a model of purchase likelihood (what they will buy). A trivial solution consists of learning global budget terms for each user based on their past expenditure. To more accurately model user budgets, we also consider a user's within-session budget, which may deviate from their global budget depending on their shopping context. In this paper, we find that users implicitly reveal their session-specific budgets through the sequence of items they browse within that session. Specifically, we find that some users "browse down," by purchasing the cheapest item among alternatives under consideration, others "browse up" (selecting the most expensive), and others ultimately purchase items around the middle. Surprisingly, this mixture of behaviors is difficult to observe globally, as individual users tend to belong firmly to one of the three segments. To model this behavior, we develop an interpretable budget model that combines a clustering component to detect different user segments, with a model of segment-specific purchase profiles. We apply our model on a dataset of browsing and purchasing sessions from Etsy, a large e-commerce website focused on handmade and vintage goods, where it outperforms strong baselines and existing production systems. Diane Hu, Raphael Louca, Liangjie Hong, Julian J. McAuley |
RecSys | 1 |
| 2018 | Turning Clicks into Purchases: Revenue Optimization for Product Search in E-CommerceabstractIn recent years, product search engines have emerged as a key factor for online businesses. According to a recent survey, over 55% of online customers begin their online shopping journey by searching on an E-Commerce (EC) website like Amazon as opposed to a generic web search engine like Google. Information retrieval research to date has been focused on optimizing search ranking algorithms for web documents while little attention has been paid to product search. There are several intrinsic differences between web search and product search that make the direct application of traditional search ranking algorithms to EC search platforms difficult. First, the success of web and product search is measured differently; one seeks to optimize for relevance while the other must optimize for both relevance and revenue. Second, when using real-world EC transaction data, there is no access to manually annotated labels. In this paper, we address these differences with a novel learning framework for EC product search called LETORIF (LEarning TO Rank with Implicit Feedback). In this framework, we utilize implicit user feedback signals (such as user clicks and purchases) and jointly model the different stages of the shopping journey to optimize for EC sales revenue. We conduct experiments on real-world EC transaction data and introduce a a new evaluation metric to estimate expected revenue after re-ranking. Experimental results show that LETORIF outperforms top competitors in improving purchase rates and total revenue earned. Liang Wu 0006, Diane Hu, Liangjie Hong, Huan Liu 0001 |
SIGIR | 2 |
| 2014 | Style in the long tail: discovering unique interests with latent variable models in large scale social E-commerceabstractPurchasing decisions in many product categories are heavily influenced by the shopper's aesthetic preferences. It's insufficient to simply match a shopper with popular items from the category in question; a successful shopping experience also identifies products that match those aesthetics. The challenge of capturing shoppers' styles becomes more difficult as the size and diversity of the marketplace increases. At Etsy, an online marketplace for handmade and vintage goods with over 30 million diverse listings, the problem of capturing taste is particularly important -- users come to the site specifically to find items that match their eclectic styles. Diane Hu, Rob Hall 0001, Josh Attenberg |
KDD | 1 |
| 2011 | Toward Robust Material Recognition for Everyday ObjectsabstractMaterial recognition is a fundamental problem in perception that is receiving increasing attention. Following the recent work using Flickr [16, 23], we empirically study material recognition of real-world objects using a rich set of local features. We use the Kernel Descriptor framework [5] and extend the set of descriptors to include material-motivated attributes using variances of gradient orientation and magnitude. Large-Margin Nearest Neighbor learning is used for a 30-fold dimension reduction. We improve the state-of-the-art accuracy on the Flickr dataset [16] from 45 % to 54%. We also introduce two new datasets using ImageNet and macro photos, extensively evaluating our set of features and showing promising connections between material and object recognition. Diane Hu, Liefeng Bo, Xiaofeng Ren |
BMVC | 1 |
| 2010 | Latent Variable Models for Predicting File Dependencies in Large-Scale Software DevelopmentabstractWhen software developers modify one or more files in a large code base, they must also identify and update other related files. Many file dependencies can be detected by mining the development history of the code base: in essence, groups of related files are revealed by the logs of previous workflows. From data of this form, we show how to detect dependent files by solving a problem in binary matrix completion. We explore different latent variable models (LVMs) for this problem, including Bernoulli mixture models, exponential family PCA, restricted Boltzmann machines, and fully Bayesian approaches. We evaluate these models on the development histories of three large, open-source software systems: Mozilla Firefox, Eclipse Subversive, and Gimp. In all of these applications, we find that LVMs improve the performance of related file prediction over current leading methods. Diane Hu, Laurens van der Maaten, Youngmin Cho, Lawrence K. Saul, Sorin Lerner |
NIPS | 1 |