EDBT 2026 Demo / reviewers in the wild / expert
Frank Yang
dblp:16/11180
· DBLP profile ↗
15ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Theory of computation · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Delayed Feedback Modeling with Influence FunctionsabstractIn online advertising under the cost-per-conversion (CPA) model, accurate conversion rate (CVR) prediction is crucial. A major challenge is delayed feedback, where conversions may occur long after user interactions, leading to incomplete recent data and biased model training. Existing solutions partially mitigate this issue but often rely on auxiliary models, making them computationally inefficient and less adaptive to user interest shifts. We propose IF-DFM, an Influence Function-empowered for Delayed Feedback Modeling which estimates the impact of newly arrived and delayed conversions on model parameters, enabling efficient updates without full retraining. By reformulating the inverse Hessian-vector-product as an optimization problem, IF-DFM achieves a favorable trade-off between scalability and effectiveness. Experiments on benchmark datasets show that IF-DFM outperforms prior methods in both accuracy and adaptability. Chenlu Ding, Jiancan Wu, Yancheng Yuan, Cunchun Li, Xiang Wang 0010, Dingxian Wang, Frank Yang, Andrew Rabinovich |
AAAI | 7 |
| 2026 | Enhancing Job Search Effectiveness with LLM-Powered Context-Aware Query Reformulation
Quang Hieu Vu, Behnaz Nojavanasghari, Frank Yang, Andrew Rabinovich |
ECIR (4) | 3 |
| 2026 | Bi-Level Optimization for Generative Recommendation: Bridging Tokenization and GenerationabstractGenerative recommendation is emerging as a transformative paradigm by directly generating recommended items, rather than relying on matching. Building such a system typically involves two key components: (1) optimizing the tokenizer to derive suitable item identifiers, and (2) training the recommender based on those identifiers. Existing approaches often treat these components separately—either sequentially or in alternation—overlooking their interdependence. This separation can lead to misalignment: the tokenizer is trained without direct guidance from the recommendation objective, potentially yielding suboptimal identifiers that degrade recommendation performance. To address this, we propose BLOGER, a Bi-Level Optimization for GEnerative Recommendation framework, which explicitly models the interdependence between the tokenizer and the recommender in a unified optimization process. The lower level trains the recommender using tokenized sequences, while the upper level optimizes the tokenizer based on both the tokenization loss and recommendation loss. We adopt a meta-learning approach to solve this bi-level optimization efficiently, and introduce gradient surgery to mitigate gradient conflicts in the upper-level updates, thereby ensuring that item identifiers are both informative and recommendation-aligned. Extensive experiments on multiple real-world datasets demonstrate that BLOGER consistently outperforms state-of-the-art generative recommendation methods while maintaining practical efficiency with no significant additional computational overhead, effectively bridging the gap between item tokenization and autoregressive generation. We release our code at https://github.com/Ten-Mao/BLOGER. Yimeng Bai, Yang Zhang 0072, Dingxian Wang, Frank Yang, Andrew Rabinovich, Wenge Rong, Fuli Feng |
SIGIR | 5 |
| 2026 | FilterRec: An Intent-Aware Framework for Dynamic Filter Recommendation
Dingxian Wang, Jiacheng Dong, Jiaqi Deng 0001, Jing Long, Ted Liu, George Barelas, Arya Taylor, Spyros Kapnissis, Frank Yang, Andrew Rabinovich, Guandong Xu |
WWW | 9 |
| 2025 | Dynamic Threats to Credible AuctionsabstractWe study the design of credible auctions when a seller has private information about her costs and cannot commit to public announcements. Akbarpour and Li [2020] show that when the reserve price is common knowledge, the first-price auction is the unique mechanism that is simultaneously optimal, static, and credible. Our paper demonstrates that this conclusion is overturned when the seller is privately informed about her cost, a common feature in many real-world markets. Martino Banchio, Andrzej Skrzypacz, Frank Yang |
EC | 3 |
| 2025 | Multidimensional Screening with ReturnsabstractA monopolist wants to sell multiple goods, accounting for the possibility of returns. For any bundle purchased, a buyer can return any part of the bundle for a partial refund specified by a return policy. The seller is constrained to offer additive return policies: the total price of a bundle must be divided into return values for each of the goods in the bundle. We show that if the buyer's values are additive and independent across goods, then selling each good separately is optimal. The result applies even if the seller can use stochastic mechanisms, and holds for general multidimensional screening problems. We show that selling separately remains approximately optimal if the buyer experiences small return costs, as well as if the return policy only has to be close to additive or the buyer's values are weakly correlated. Applying our analysis to multiproduct pricing without returns, we introduce a new concept of bundle discounts and show that obtaining revenue gains from using any complex mechanism requires discounting bundles in our sense. Alexander Haberman, Ravi Jagadeesan, Frank Yang |
EC | 3 |
| 2025 | Multidimensional Monotonicity and Economic ApplicationsabstractWe study the set of multidimensional monotone functions from [0, 1]n to [0, 1] as well as their one-dimensional marginals. We characterize the extreme points of these convex sets subject to finitely many linear constraints. These characterizations lead to new results in various mechanism design and information design problems, including public good provision with interdependent values; interim efficient bilateral trade mechanisms; asymmetric reduced form auctions; and optimal private private information structure. As another application, we also present a mechanism anti-equivalence theorem for two-agent, two-alternative social choice problems: A mechanism is payoff-equivalent to a deterministic DIC mechanism if and only if they are ex-post equivalent. Frank Yang, Kai-Hao Yang |
EC | 1 |
| 2024 | Case Study: Runtime Safety Verification of Neural Network Controlled System
Frank Yang, Simon Sinong Zhan, Yixuan Wang 0001, Chao Huang 0015, Qi Zhu 0002 |
RV | 1 |
| 2023 | Comparison of Screening DevicesabstractPublic agencies are often tasked with allocating scarce resources (such as public housing or financial aid) to a target population. In many such cases, the goal is to maximize social welfare, which requires identifying agents who have the highest social value for the resource. The challenge is that while public agencies may be able to access some data about potential beneficiaries (for example, through means testing), they generally lack information necessary to achieve perfect targeting. When monetary transfers are unavailable or ineffective in targeting, public agencies often rely on "ordeals" instead. Natural examples include standing in line, filing out complicated forms, dealing with "red tape," waiting, visiting an office at an inconvenient time, or traveling to a registration site. But what makes one costly screening device better than another? Mohammad Akbarpour, Piotr Dworczak, Frank Yang |
EC | 3 |
| 2022 | Costly Multidimensional ScreeningabstractA screening instrument is costly if it is socially wasteful andproductive otherwise. A principal screens an agent with multidimensional private information and quasilinear preferences that are additively separable across two components: a one-dimensional productive component and a multidimensional costly component. Can the principal improve upon simple one-dimensional mechanisms by also using the costly instruments? We show that if the agent has preferences between the two components that are positively correlated in a suitably defined sense, then simply screening the productive component is optimal. The result holds for general type and allocation spaces, and allows for nonlinear and interdependent valuations. We discuss applications to optimal regulation, labor market screening, and pricing and bundling by a multiple-good monopolist. Frank Yang |
EC | 1 |
| 2020 | 2-limited broadcast domination in subcubic graphs
Michael A. Henning, Gary MacGillivray, Frank Yang |
Discret. Appl. Math. | 3 |
| 2018 | Customized Regression Model for Airbnb Dynamic PricingabstractThis paper describes the pricing strategy model deployed at Airbnb, an online marketplace for sharing home and experience. The goal of price optimization is to help hosts who share their homes on Airbnb set the optimal price for their listings. In contrast to conventional pricing problems, where pricing strategies are applied to a large quantity of identical products, there are no "identical" products on Airbnb, because each listing on our platform offers unique values and experiences to our guests. The unique nature of Airbnb listings makes it very difficult to estimate an accurate demand curve that's required to apply conventional revenue maximization pricing strategies. Julian Qian, Chen-Hung Wu, Spencer De Mars, Frank Yang |
KDD | 7 |
| 2018 | k-broadcast domination and k-multipacking
Michael A. Henning, Gary MacGillivray, Frank Yang |
Discret. Appl. Math. | 3 |
| 2015 | Why Big Data Industrial Systems Need Rules and What We Can Do About ItabstractBig Data industrial systems that address problems such as classification, information extraction, and entity matching very commonly use hand-crafted rules. Today, however, little is understood about the usage of such rules. In this paper we explore this issue. We discuss how these systems differ from those considered in academia. We describe default solutions, their limitations, and reasons for using rules. We show examples of extensive rule usage in industry. Contrary to popular perceptions, we show that there is a rich set of research challenges in rule generation, evaluation, execution, optimization, and maintenance. We discuss ongoing work at WalmartLabs and UW-Madison that illustrate these challenges. Our main conclusions are (1) using rules (together with techniques such as learning and crowdsourcing) is fundamental to building semantics-intensive Big Data systems, and (2) it is increasingly critical to address rule management, given the tens of thousands of rules industrial systems often manage today in an ad-hoc fashion. Paul Suganthan G. C., Krishna Gayatri K., Haojun Zhang, Frank Yang, Narasimhan Rampalli, Shishir Prasad, Esteban Arcaute, Ganesh Krishnan, Rohit Deep, Vijay Raghavendra, AnHai Doan |
SIGMOD Conference | 5 |
| 2014 | Chimera: Large-Scale Classification using Machine Learning, Rules, and CrowdsourcingabstractLarge-scale classification is an increasingly critical Big Data problem. So far, however, very little has been published on how this is done in practice. In this paper we describe Chimera, our solution to classify tens of millions of products into 5000+ product types at WalmartLabs. We show that at this scale, many conventional assumptions regarding learning and crowdsourcing break down, and that existing solutions cease to work. We describe how Chimera employs a combination of learning, rules (created by in-house analysts), and crowdsourcing to achieve accurate, continuously improving, and cost-effective classification. We discuss a set of lessons learned for other similar Big Data systems. In particular, we argue that at large scales crowdsourcing is critical, but must be used in combination with learning, rules, and in-house analysts. We also argue that using rules (in conjunction with learning) is a must, and that more research attention should be paid to helping analysts create and manage (tens of thousands of) rules more effectively. Narasimhan Rampalli, Frank Yang, AnHai Doan |
Proc. VLDB Endow. | 3 |