EDBT 2026 Demo / reviewers in the wild / expert
Tianyi Peng
dblp:243/6511
· DBLP profile ↗
18ranked-venue papers
2as first author
15since 2021 · last 2025
0000-0002-9046-3206ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Computer networks · 2Theory of computation · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dark Experience for Incremental Keyword SpottingabstractSpoken keyword spotting (KWS) is crucial for identifying keywords within audio inputs and is widely used in applications like Apple Siri and Google Home, particularly on edge devices. Current deep learning-based KWS systems, which are typically trained on a limited set of keywords, can suffer from performance degradation when encountering new domains, a challenge often addressed through few-shot fine-tuning. However, this adaptation frequently leads to catastrophic forgetting, where the model’s performance on original data deteriorates. Progressive continual learning (CL) strategies have been proposed to overcome this, but they face limitations such as the need for task-ID information and increased storage, making them less practical for lightweight devices. To address these challenges, we introduce Dark Experience for Keyword Spotting (DE-KWS), a novel CL approach that leverages dark knowledge to distill past experiences throughout the training process. DE-KWS combines rehearsal and distillation, using both ground truth labels and logits stored in a memory buffer to maintain model performance across tasks. Evaluations on the Google Speech Command dataset show that DE-KWS outperforms existing CL baselines in average accuracy without increasing model size, offering an effective solution for resource-constrained edge devices. The scripts are available on GitHub1for future research. Tianyi Peng, Yang Xiao 0019 |
ICASSP | 1 |
| 2025 | Speeding up Policy Simulation in Supply Chain RLabstractSimulating a single trajectory of a dynamical system under some state-dependent policy is a core bottleneck in policy optimization (PO) algorithms. The many inherently serial policy evaluations that must be performed in a single simulation constitute the bulk of this bottleneck. In applying PO to supply chain optimization (SCO) problems, simulating a single sample path corresponding to one month of a supply chain can take several hours. We present an iterative algorithm to accelerate policy simulation, dubbed Picard Iteration. This scheme carefully assigns policy evaluation tasks to independent processes. Within an iteration, any given process evaluates the policy only on its assigned tasks while assuming a certain cached’ evaluation for other tasks; the cache is updated at the end of the iteration. Implemented on GPUs, this scheme admits batched evaluation of the policy across a single trajectory. We prove that the structure afforded by many SCO problems allows convergence in a small number of iterations independent of the horizon. We demonstrate practical speedups of 400x on large-scale SCO problems even with a single GPU, and also demonstrate practical efficacy in other RL environments. Vivek F. Farias, Joren Gijsbrechts, Aryan I. Khojandi, Tianyi Peng, Andrew Zheng |
ICML | 4 |
| 2025 | AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation
Yang Xiao 0019, Tianyi Peng, Yanghao Zhou, Rohan Kumar Das |
INTERSPEECH | 2 |
| 2025 | Multi-agent Markov EntanglementabstractValue decomposition has long been a fundamental technique in multi-agent reinforcement learning and dynamic programming. Specifically, the value function of a global state $(s_1,s_2,\ldots,s_N)$ is often approximated as the sum of local functions: $V(s_1,s_2,\ldots,s_N)\approx\sum_{i=1}^N V_i(s_i)$. This approach has found various applications in modern RL systems. However, the theoretical justification for why this decomposition works so effectively remains underexplored. In this paper, we uncover the underlying mathematical structure that enables value decomposition. We demonstrate that a Markov decision process (MDP) permits value decomposition *if and only if* its transition matrix is not "entangled"—a concept analogous to quantum entanglement in quantum physics. Drawing inspiration from how physicists measure quantum entanglement, we introduce how to measure the "Markov entanglement" and show that this measure can be used to bound the decomposition error in general multi-agent MDPs. Using the concept of Markov entanglement, we proved that a widely-used class of policies, the index policy, is weakly-entangled and enjoys a sublinear $\mathcal O(\sqrt{N})$ scale of decomposition error for $N$-agent systems. Finally, we show Markov entanglement can be efficiently estimated, guiding practitioners on the feasibility of value decomposition. Shuze Chen, Tianyi Peng |
NeurIPS | 2 |
| 2025 | LLM Generated Persona is a Promise with a CatchabstractThe use of large language models (LLMs) to simulate human behavior has gained significant attention, particularly through personas that approximate individual characteristics. Persona-based simulations hold promise for transforming disciplines that rely on population-level feedback, including social science, economic analysis, marketing research, and business operations. Traditional methods to collect realistic persona data face significant challenges. They are prohibitively expensive and logistically challenging due to privacy constraints, and often fail to capture multi-dimensional attributes, particularly subjective qualities. Consequently, synthetic persona generation with LLMs offers a scalable, cost-effective alternative. However, current approaches rely on ad hoc and heuristic generation techniques that do not guarantee methodological rigor or simulation precision, resulting in systematic biases in downstream tasks. Through extensive large-scale experiments including presidential election forecasts and general opinion surveys of the U.S. population, we reveal that these biases can lead to significant deviations from real-world outcomes. Based on the experimental results, this position paper argues that a rigorous and systematic science of persona generation is needed to ensure the reliability of LLM-driven simulations of human behavior. We call for not only methodological innovations and empirical foundations but also interdisciplinary organizational and institutional support for the development of this field. To support further research and development in this area, we have open-sourced approximately one million generated personas, available for public access and analysis. Leon Li, Hongseok Namkoong, Tianyi Peng |
NeurIPS | 4 |
| 2025 | Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian FrameworkabstractCareful curation of data sources can significantly improve the performance of LLM pre-training, but predominant approaches rely heavily on intuition or costly trial-and-error, making them difficult to generalize across different data domains and downstream tasks. Although scaling laws can provide a principled and general approach for data curation, standard deterministic extrapolation from small-scale experiments to larger scales requires strong assumptions on the reliability of such extrapolation, whose brittleness has been highlighted in prior works. In this paper, we introduce a probabilistic extrapolation framework for data mixture optimization that avoids rigid assumptions and explicitly models the uncertainty in performance across decision variables. We formulate data curation as a sequential decision-making problem–multi-fidelity, multi-scale Bayesian optimization–where {data mixtures, model scale, training steps} are adaptively selected to balance training cost and potential information gain. Our framework naturally gives rise to algorithm prototypes that leverage noisy information from inexpensive experiments to systematically inform costly training decisions. To accelerate methodological progress, we build a simulator based on 472 language model pre-training runs with varying data compositions from the SlimPajama dataset. We observe that even simple kernels and acquisition functions can enable principled decisions across training models from 20M to 1B parameters and achieve 2.6x and 3.3x speedups compared to multi-fidelity BO and random search baselines. Taken together, our framework underscores potential efficiency gains achievable by developing principled and transferable data mixture optimization methods. Our code is publicly available at https://github.com/namkoong-lab/data-recipes. Thomson Yen, Andrew Wei Tung Siah, C. Guetta, Tianyi Peng, Hongseok Namkoong |
NeurIPS | 5 |
| 2025 | Tail-Optimized Caching for LLM InferenceabstractPrompt caching is critical for reducing latency and cost in LLM inference---OpenAI and Anthropic report up to 50–90\% cost savings through prompt reuse. Despite its widespread success, little is known about what constitutes an optimal prompt caching policy, particularly when optimizing tail latency—a metric of central importance to practitioners. The widely used Least Recently Used (LRU) policy can perform arbitrarily poor on this metric, as it is oblivious to the heterogeneity of conversation lengths. To address this gap, we propose Tail-Optimized LRU, a simple two-line modification that reallocates KV cache capacity to prioritize high-latency conversations by evicting cache entries that are unlikely to affect future turns. Though the implementation is simple, we prove its optimality under a natural stochastic model of conversation dynamics, providing the first theoretical justification for LRU in this setting---a result that may be of independent interest to the caching community.
Experimentally, on real conversation data WildChat~\citep{zhao2024wildchat}, Tail-Optimized LRU achieves up to 27.5\% reduction in P90 tail Time to First Token latency and 23.9\% in P95 tail latency compared to LRU, along with up to 38.9\% decrease in SLO violations of 200ms.
We believe this provides a practical and theoretically grounded option for practitioners seeking to optimize tail latency in real-world LLM deployments. Ciamac C. Moallemi, Tianyi Peng |
NeurIPS | 4 |
| 2025 | Differences-in-Neighbors for Network Interference in ExperimentsabstractExperiments in online platforms frequently suffer from network interference, where treatments applied to one unit affect outcomes of connected ones, violating the Stable Unit Treatment Value Assumption (SUTVA) and substantially biasing treatment effect estimations. A common solution is to cluster connected units and randomize treatments at the cluster level, typically followed by estimation using either a simple difference-in-means (DM) estimator, which ignores remaining interference and suffers from O(δ) bias where δ measures interference strength; or the unbiased Horvitz-Thompson (HT) estimator, which eliminates bias through importance sampling but incurs exponentially high variance scaling with d, the maximum network degree. This fundamental limitation persists even with sophisticated clustering designs, creating narrow bias-variance tradeoffs often inadequate for practical applications. Tianyi Peng, Naimeng Ye, Andrew Zheng |
EC | 1 |
| 2024 | QGym: Scalable Simulation and Benchmarking of Queuing Network ControllersabstractQueuing network control allows allocation of scarce resources to manage congestion, a fundamental problem in manufacturing, communications, and healthcare. Compared to standard RL problems, queueing problems are distinguished by unique challenges: i) a system operating in continuous time, ii) high stochasticity, and iii) long horizons over which the system can become unstable (exploding delays). To provide the empirical foundations for methodological development tackling these challenges, we present an open-sourced queueing simulation framework, QGym, that benchmark queueing policies across realistic problem instances. Our modular framework allows the researchers to build on our initial instances, which provide a wide range of environments including parallel servers, criss-cross, tandem, and re-entrant networks, as well as a realistically calibrated hospital queuing system. From these, various policies can be easily tested, including both model-free RL methods and classical queuing policies. Our testbed significantly expands the scope of empirical benchmarking in prior work, and complements thetraditional focus on evaluating algorithms based on mathematical guarantees in idealized settings. QGym code is open-sourced at https://github.com/namkoong-lab/QGym. Ethan Che, Tianyi Peng, Hongseok Namkoong |
NeurIPS | 5 |
| 2023 | Correcting for Interference in Experiments: A Case Study at DouyinabstractInterference is a ubiquitous problem in experiments conducted on two-sided content marketplaces, such as Douyin (China’s analog of TikTok). In many cases, creators are the natural unit of experimentation, but creators interfere with each other through competition for viewers’ limited time and attention. “Naive” estimators currently used in practice simply ignore the interference, but in doing so incur bias on the order of the treatment effect. We formalize the problem of inference in such experiments as one of policy evaluation. Off-policy estimators, while unbiased, are impractically high variance. We introduce a novel Monte-Carlo estimator, based on “Differences-in-Qs” (DQ) techniques, which achieves bias that is second-order in the treatment effect, while remaining sample-efficient to estimate. On the theoretical side, our contribution is to develop a generalized theory of Taylor expansions for policy evaluation, which extends DQ theory to all major MDP formulations. On the practical side, we implement our estimator on Douyin’s experimentation platform, and in the process develop DQ into a truly “plug-and-play” estimator for interference in real-world settings: one which provides robust, low-bias, low-variance treatment effect estimates; admits computationally cheap, asymptotically exact uncertainty quantification; and reduces MSE by 99% compared to the best existing alternatives in our applications. Vivek F. Farias, Hao Li 0191, Tianyi Peng, Xinyuyang Ren, Andrew Zheng |
RecSys | 3 |
| 2022 | Uncertainty Quantification for Low-Rank Matrix Completion with Heterogeneous and Sub-Exponential NoiseabstractThe problem of low-rank matrix completion with heterogeneous and sub-exponential (as opposed to homogeneous Gaussian) noise is particularly relevant to a number of applications in modern commerce. Examples include panel sales data and data collected from web-commerce systems such as recommendation engines. An important unresolved question for this problem is characterizing the distribution of estimated matrix entries under common low-rank estimators. Such a characterization is essential to any application that requires quantification of uncertainty in these estimates and has heretofore only been available under the assumption of homogenous Gaussian noise. Here we characterize the distribution of estimated matrix entries when the observation noise is heterogeneous sub-Exponential and provide, as an application, explicit formulas for this distribution when observed entries are Poisson or Binary distributed. Vivek F. Farias, Andrew A. Li, Tianyi Peng |
AISTATS | 3 |
| 2022 | Markovian Interference in ExperimentsabstractWe consider experiments in dynamical systems where interventions on some experimental units impact other units through a limiting constraint (such as a limited supply of products). Despite outsize practical importance, the best estimators for this `Markovian' interference problem are largely heuristic in nature, and their bias is not well understood. We formalize the problem of inference in such experiments as one of policy evaluation. Off-policy estimators, while unbiased, apparently incur a large penalty in variance relative to state-of-the-art heuristics. We introduce an on-policy estimator: the Differences-In-Q's (DQ) estimator. We show that the DQ estimator can in general have exponentially smaller variance than off-policy evaluation. At the same time, its bias is second order in the impact of the intervention. This yields a striking bias-variance tradeoff so that the DQ estimator effectively dominates state-of-the-art alternatives. From a theoretical perspective, we introduce three separate novel techniques that are of independent interest in the theory of Reinforcement Learning (RL). Our empirical evaluation includes a set of experiments on a city-scale ride-hailing simulator. Vivek F. Farias, Andrew A. Li, Tianyi Peng, Andrew Zheng |
NeurIPS | 3 |
| 2021 | Near-Optimal Entrywise Anomaly Detection for Low-Rank Matrices with Sub-Exponential NoiseabstractWe study the problem of identifying anomalies in a low-rank matrix observed with sub-exponential noise, motivated by applications in retail and inventory management. State of the art approaches to anomaly detection in low-rank matrices apparently fall short, since they require that non-anomalous entries be observed with vanishingly small noise (which is not the case in our problem, and indeed in many applications). So motivated, we propose a conceptually simple entrywise approach to anomaly detection in low-rank matrices. Our approach accommodates a general class of probabilistic anomaly models. We extend recent work on entrywise error guarantees for matrix completion, establishing such guarantees for sub-exponential matrices, where in addition to missing entries, a fraction of entries are corrupted by (an also unknown) anomaly model. Viewing the anomaly detection as a classification task, to the best of our knowledge, we are the first to achieve the min-max optimal detection rate (up to log factors). Using data from a massive consumer goods retailer, we show that our approach provides significant improvements over incumbent approaches to anomaly detection. Vivek F. Farias, Andrew A. Li, Tianyi Peng |
ICML | 3 |
| 2021 | Learning Treatment Effects in Panels with General Intervention PatternsabstractThe problem of causal inference with panel data is a central econometric question. The following is a fundamental version of this problem: Let $M^*$ be a low rank matrix and $E$ be a zero-mean noise matrix. For a `treatment' matrix $Z$ with entries in $\{0,1\}$ we observe the matrix $O$ with entries $O_{ij} := M^*_{ij} + E_{ij} + \mathcal{T}_{ij} Z_{ij}$ where $\mathcal{T}_{ij} $ are unknown, heterogenous treatment effects. The problem requires we estimate the average treatment effect $\tau^* := \sum_{ij} \mathcal{T}_{ij} Z_{ij} / \sum_{ij} Z_{ij}$. The synthetic control paradigm provides an approach to estimating $\tau^*$ when $Z$ places support on a single row. This paper extends that framework to allow rate-optimal recovery of $\tau^*$ for general $Z$, thus broadly expanding its applicability. Our guarantees are the first of their type in this general setting. Computational experiments on synthetic and real-world data show a substantial advantage over competing estimators. Vivek F. Farias, Andrew A. Li, Tianyi Peng |
NeurIPS | 3 |
| 2021 | The Limits to Learning a Diffusion ModelabstractThis paper provides the first sample complexity lower bounds for the estimation of simple diffusion models which seek to explain the diffusion of an epidemic in a network. The Susceptible-Infected-Recovered (SIR) model is a classic example, proposed nearly a century ago [2]. The SIR model remains a cornerstone for the forecasting of epidemics. The so-called Bass model [1] remains a basic building block in forecasting consumer adoption of new products and services. The durability of these models arises from the fact that they have shown an excellent fit to data, in numerous studies spanning both the epidemiology and marketing literatures. Somewhat paradoxically, using these same models as reliable forecasting tools presents a challenge. Jackie Baek, Vivek F. Farias, Andreea Georgescu, Retsef Levi, Tianyi Peng, Deeksha Sinha, Joshua Wilde, Andrew Zheng |
EC | 5 |
| 2020 | Optimal Remote Entanglement DistributionabstractDistributing entanglement between distant nodes is an essential task in quantum networks. To achieve this task, quantum repeaters have been introduced to perform entanglement swapping. This paper offers a design of remote entanglement distribution (RED) protocols that maximize the entanglement distribution rate (EDR). We introduce the concept of enodes, representing the entangled quantum bit (qubit) pairs in the network. This concept enables us to design the optimal RED protocols based on the solutions of some linear programming problems. Moreover, we investigate RED in a homogeneous repeater chain, which is a building block for many quantum networks. In particular, we determine the maximum EDR for homogeneous repeater chains in a closed form. Our results enable the distribution of long-distance entanglement with noisy intermediate-scale quantum (NISQ) technologies and provide insights into the design and implementation of general quantum networks. Wenhan Dai, Tianyi Peng, Moe Z. Win |
IEEE J. Sel. Areas Commun. | 2 |
| 2020 | Quantum Queuing DelayabstractQueuing delay is an essential topic in the design of quantum networks. This paper introduces a tractable model for analyzing the queuing delay of quantum data, referred to as quantum queuing delay (QQD). The model employs a dynamic programming formalism and accounts for practical aspects such as the finite memory size. Using this model, we develop a cognitive-memory-based policy for memory management and show that this policy can decrease the average queuing delay exponentially with respect to memory size. Such a significant reduction can be traced back to the use of entanglement, a peculiar quantum phenomenon that has no classical counterpart. Numerical results validate the theoretical analysis and demonstrate the near-optimal performance of the developed policy. Wenhan Dai, Tianyi Peng, Moe Z. Win |
IEEE J. Sel. Areas Commun. | 2 |
| 2019 | Remote State Preparation for Multiple PartiesabstractRemote state preparation (RSP) is a technique to transmit quantum states with classical communication and previously shared entanglement. In this paper, we consider RSP in a multiparty setting. A simple yet nontrivial case is studied, where there is one sender and two receivers. We put forth a broadcasting method for developing multiparty RSP protocols. We show that this method can remotely prepare arbitrary states by consuming classical and quantum resources. For preparing highly entangled states, the proposed method achieves the lower bound for the amount of consumed resources asymptotically. Wenhan Dai, Tianyi Peng, Moe Z. Win |
ICASSP | 2 |