Assaf Zeevi

dblp:85/2086 · also Assaf J. Zeevi · DBLP profile ↗
← Back
28ranked-venue papers
4as first author
15since 2021 · last 2025
0000-0003-1075-6664ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 2 first-author · 13 since 2021Theory of computation · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Learning the Pareto Front Using Bootstrapped Observation Samples
abstract
We consider Pareto front identification (PFI) for linear bandits (PFILin), i.e., the goal is to identify a set of arms with undominated mean reward vectors when the mean reward vector is a linear function of the context. PFILin includes the best arm identification problem and multi-objective active learning as special cases. The sample complexity of our proposed algorithm is optimal up to a logarithmic factor. In addition, the regret incurred by our algorithm during the estimation is within a logarithmic factor of the optimal regret among all algorithms that identify the Pareto front. Our key contribution is a new estimator that in every round updates the estimate for the unknown parameter along \emph{multiple} context directions – in contrast to the conventional estimator that only updates the parameter estimate along the chosen context. This allows us to use low-regret arms to collect information about Pareto optimal arms. Our key innovation is to reuse the exploration samples multiple times; in contrast to conventional estimators that use each sample only once. Numerical experiments demonstrate that the proposed algorithm successfully identifies the Pareto front while controlling the regret.
Wonyoung Kim, Garud Iyengar, Assaf Zeevi
AISTATS3
2025 Linear Bandits with Partially Observable Features
abstract
We study the linear bandit problem that accounts for partially observable features. Without proper handling, unobserved features can lead to linear regret in the decision horizon $T$, as their influence on rewards is unknown. To tackle this challenge, we propose a novel theoretical framework and an algorithm with sublinear regret guarantees. The core of our algorithm consists of (i) feature augmentation, by appending basis vectors that are orthogonal to the row space of the observed features; and (ii) the introduction of a doubly robust estimator. Our approach achieves a regret bound of $\tilde{O}(\sqrt{(d + d\_h)T})$, where $d$ is the dimension of the observed features and $d_h$ depends on the extent to which the unobserved feature space is contained in the observed one, thereby capturing the intrinsic difficulty of the problem. Notably, our algorithm requires no prior knowledge of the unobserved feature space, which may expand as more features become hidden. Numerical experiments confirm that our algorithm outperforms both non-contextual multi-armed bandits and linear bandit algorithms depending solely on observed features.
Wonyoung Kim, Garud Iyengar, Assaf Zeevi, Min-hwan Oh
ICML4
2025 Bayesian Design Principles for Frequentist Sequential Learning
abstract
We develop a general theory to optimize the frequentist regret for sequential learning problems, from which efficient bandit and reinforcement learning algorithms can be derived via unified Bayesian principles. Building on the recent Decision-Estimation Coefficient (DEC) framework, we propose a novel optimization approach to generate “algorithmic beliefs” at each round and use Bayesian posteriors for decision-making. The optimization objective, termed “Algorithmic Information Ratio” (AIR), represents an intrinsic complexity measure that effectively characterizes the frequentist regret of any algorithm. Although AIR’s minimax regret aligns with that provided by DEC, it additionally offers an algorithm-dependent perspective–distinct from a minimax complexity–facilitating algorithm design and analysis. Specifically, AIR enables deriving explicit algorithms via belief parameterization and provides clear approximation guidelines with provable guarantees. Moreover, the resulting algorithms have a simple structure and are computationally efficient for several representative problems. We illustrate our framework with a novel algorithm for multi-armed bandits that performs strongly across stochastic, adversarial, and non-stationary environments, and demonstrate applicability to linear bandits, convex bandits, and reinforcement learning.
Yunbei Xu, Assaf Zeevi
J. ACM2
2024 A Doubly Robust Approach to Sparse Reinforcement Learning
abstract
We propose a new regret minimization algorithm for episodic sparse linear Markov decision process (SMDP) where the state-transition distribution is a linear function of observed features. The only previously known algorithm for SMDP requires the knowledge of the sparsity parameter and oracle access to an unknown policy. We overcome these limitations by combining the doubly robust method that allows one to use feature vectors of \emph{all} actions with a novel analysis technique that enables the algorithm to use data from all periods in all episodes. The regret of the proposed algorithm is $\tilde{O}(\sigma^{-1}_{\min}s_{\star} H \sqrt{N})$, where $\sigma_{\min}$ denotes the restrictive the minimum eigenvalue of the average Gram matrix of feature vectors, $s_\star$ is the sparsity parameter, $H$ is the length of an episode, and $N$ is the number of rounds. We provide a lower regret bound that matches the upper bound to logarithmic factors on a newly identified subclass of SMDPs. Our numerical experiments support our theoretical results and demonstrate the superior performance of our algorithm.
Wonyoung Kim, Garud Iyengar, Assaf Zeevi
AISTATS3
2023 Complexity Analysis of a Countable-armed Bandit Problem
abstract
We consider a stochastic multi-armed bandit (MAB) problem motivated by “large” action spaces, and endowed with a population of arms containing exactly $K$ arm-types, each characterized by a distinct mean reward. The decision maker is oblivious to the statistical properties of reward distributions as well as the population-level distribution of different arm-types, and is precluded also from observing the type of an arm after play. We study the classical problem of minimizing the expected cumulative regret over a horizon of play $n$, and propose algorithms that achieve a rate-optimal finite-time instance-dependent regret of $\mathcal{O}\left( \log n \right)$. We also show that the instance-independent (minimax) regret is $\tilde{\mathcal{O}}\left( \sqrt{n} \right)$ when $K=2$. While the order of regret and complexity of the problem suggests a great degree of similarity to the classical MAB problem, properties of the performance bounds and salient aspects of algorithm design are quite distinct from the latter, as are the key primitives that determine complexity along with the analysis tools needed to study them.
Anand Kalvit, Assaf Zeevi
ALT2
2023 Last Switch Dependent Bandits with Monotone Payoff Functions
abstract
In a recent work, Laforgue et al. introduce the model of last switch dependent (LSD) bandits, in an attempt to capture nonstationary phenomena induced by the interaction between the player and the environment. Examples include satiation, where consecutive plays of the same action lead to decreased performance, or deprivation, where the payoff of an action increases after an interval of inactivity. In this work, we take a step towards understanding the approximability of planning LSD bandits, namely, the (NP-hard) problem of computing an optimal arm-pulling strategy under complete knowledge of the model. In particular, we design the first efficient constant approximation algorithm for the problem and show that, under a natural monotonicity assumption on the payoffs, its approximation guarantee (almost) matches the state-of-the-art for the special and well-studied class of recharging bandits (also known as delay-dependent). In this attempt, we develop new tools and insights for this class of problems, including a novel higher-dimensional relaxation and the technique of mirroring the evolution of virtual states. We believe that these novel elements could potentially be used for approaching richer classes of action-induced nonstationary bandits (e.g., special instances of restless bandits). In the case where the model parameters are initially unknown, we develop an online learning adaptation of our algorithm for which we provide sublinear regret guarantees against its full-information counterpart.
Ayoub Foussoul, Vineet Goyal, Orestis Papadigenopoulos, Assaf Zeevi
ICML4
2023 Improved Algorithms for Multi-period Multi-class Packing Problems with Bandit Feedback
abstract
We consider the linear contextual multi-class multi-period packing problem (LMMP) where the goal is to pack items such that the total vector of consumption is below a given budget vector and the total value is as large as possible. We consider the setting where the reward and the consumption vector associated with each action is a class-dependent linear function of the context, and the decision-maker receives bandit feedback. LMMP includes linear contextual bandits with knapsacks and online revenue management as special cases. We establish a new estimator which guarantees a faster convergence rate, and consequently, a lower regret in LMMP. We propose a bandit policy that is a closed-form function of said estimated parameters. When the contexts are non-degenerate, the regret of the proposed policy is sublinear in the context dimension, the number of classes, and the time horizon $T$ when the budget grows at least as $\sqrt{T}$. We also resolve an open problem posed in Agrawal & Devanur (2016) and extend the result to a multi-class setting. Our numerical experiments clearly demonstrate that the performance of our policy is superior to other benchmarks in the literature.
Wonyoung Kim, Garud Iyengar, Assaf Zeevi
ICML3
2023 Bayesian Design Principles for Frequentist Sequential Learning
abstract
We develop a general theory to optimize the frequentist regret for sequential learning problems, where efficient bandit and reinforcement learning algorithms can be derived from unified Bayesian principles. We propose a novel optimization approach to create "algorithmic beliefs" at each round, and use Bayesian posteriors to make decisions. This is the first approach to make Bayesian-type algorithms prior-free and applicable to adversarial settings, in a generic and optimal manner. Moreover, the algorithms are simple and often efficient to implement. As a major application, we present a novel algorithm for multi-armed bandits that achieves the "best-of-all-worlds" empirical performance in the stochastic, adversarial, and non-stationary environments. And we illustrate how these principles can be used in linear bandits, convex bandits, and reinforcement learning.
Yunbei Xu, Assaf Zeevi
ICML2
2022 Dynamic Learning in Large Matching Markets
abstract
We study a sequential matching problem faced by "large" centralized platforms where "jobs" must be matched to "workers" subject to uncertainty about worker skill proficiencies. Jobs arrive at discrete times with "job-types" observable upon arrival. To capture the "choice overload" phenomenon, we posit an unlimited supply of workers where each worker is characterized by a vector of attributes (aka "worker-types") drawn from an underlying population-level distribution. The distribution as well as mean payoffs for possible worker-job type-pairs are unobservables and the platform's goal is to sequentially match incoming jobs to workers in a way that maximizes its cumulative payoffs over the planning horizon. We establish lower bounds on the "regret" of any matching algorithm in this setting and propose a novel rate-optimal learning algorithm that adapts to aforementioned primitives "online." Our learning guarantees highlight a distinctive characteristic of the problem: achievable performance only has a "second-order" dependence on worker-type distributions; we believe this finding may be of interest more broadly.
Anand Kalvit, Assaf Zeevi
NeurIPS2
2022 Online Allocation and Learning in the Presence of Strategic Agents
abstract
We study the problem of allocating $T$ sequentially arriving items among $n$ homogenous agents under the constraint that each agent must receive a prespecified fraction of all items, with the objective of maximizing the agents' total valuation of items allocated to them. The agents' valuations for the item in each round are assumed to be i.i.d. but their distribution is apriori unknown to the central planner.vTherefore, the central planner needs to implicitly learn these distributions from the observed values in order to pick a good allocation policy. However, an added challenge here is that the agents are strategic with incentives to misreport their valuations in order to receive better allocations. This sets our work apart both from the online auction mechanism design settings which typically assume known valuation distributions and/or involve payments, and from the online learning settings that do not consider strategic agents. To that end, our main contribution is an online learning based allocation mechanism that is approximately Bayesian incentive compatible, and when all agents are truthful, guarantees a sublinear regret for individual agents' utility compared to that under the optimal offline allocation policy.
Steven Yin, Shipra Agrawal 0001, Assaf Zeevi
NeurIPS3
2022 Practical Nonparametric Sampling Strategies for Quantile-Based Ordinal Optimization
abstract
Given a finite number of stochastic systems, the goal of our problem is to dynamically allocate a finite sampling budget to maximize the probability of selecting the “best” system. Systems are encoded with the probability distributions that govern sample observations, which are unknown and only assumed to belong to a broad family of distributions that need not admit any parametric representation. The best system is defined as the one with the highest quantile value. The objective of maximizing the probability of selecting this best system is not analytically tractable. In lieu of that, we use the rate function for the probability of error relying on large deviations theory. Our point of departure is an algorithm that naively combines sequential estimation and myopic optimization. This algorithm is shown to be asymptotically optimal; however, it exhibits poor finite-time performance and does not lead itself to implementation in settings with a large number of systems. To address this, we propose practically implementable variants that retain the asymptotic performance of the former while dramatically improving its finite-time performance.
Mark Broadie, Assaf Zeevi
INFORMS J. Comput.3
2021 Learning to Stop with Surprisingly Few Samples
abstract
We consider a discounted infinite horizon optimal stopping problem. If the underlying distribution is known a priori, the solution of this problem is obtained via dynamic programming (DP) and is given by a well known threshold rule. When information on this distribution is lacking, a natural (though naive) approach is “explore-then-exploit," whereby the unknown distribution or its parameters are estimated over an initial exploration phase, and this estimate is then used in the DP to determine actions over the residual exploitation phase. We show: (i) with proper tuning, this approach leads to performance comparable to the full information DP solution; and (ii) despite common wisdom on the sensitivity of such “plug in" approaches in DP due to propagation of estimation errors, a surprisingly “short" (logarithmic in the horizon) exploration horizon suffices to obtain said performance. In cases where the underlying distribution is heavy-tailed, these observations are even more pronounced: a single sample exploration phase suffices.
Daniel Russo 0001, Assaf Zeevi
COLT2
2021 Sparsity-Agnostic Lasso Bandit
abstract
We consider a stochastic contextual bandit problem where the dimension $d$ of the feature vectors is potentially large, however, only a sparse subset of features of cardinality $s_0 \ll d$ affect the reward function. Essentially all existing algorithms for sparse bandits require a priori knowledge of the value of the sparsity index $s_0$. This knowledge is almost never available in practice, and misspecification of this parameter can lead to severe deterioration in the performance of existing methods. The main contribution of this paper is to propose an algorithm that does not require prior knowledge of the sparsity index $s_0$ and establish tight regret bounds on its performance under mild conditions. We also comprehensively evaluate our proposed algorithm numerically and show that it consistently outperforms existing methods, even when the correct sparsity index is revealed to them but is kept hidden from our algorithm.
Min-hwan Oh, Garud Iyengar, Assaf Zeevi
ICML3
2021 A Closer Look at the Worst-case Behavior of Multi-armed Bandit Algorithms
abstract
One of the key drivers of complexity in the classical (stochastic) multi-armed bandit (MAB) problem is the difference between mean rewards in the top two arms, also known as the instance gap. The celebrated Upper Confidence Bound (UCB) policy is among the simplest optimism-based MAB algorithms that naturally adapts to this gap: for a horizon of play n, it achieves optimal O(log n) regret in instances with "large" gaps, and a near-optimal O(\sqrt{n log n}) minimax regret when the gap can be arbitrarily "small." This paper provides new results on the arm-sampling behavior of UCB, leading to several important insights. Among these, it is shown that arm-sampling rates under UCB are asymptotically deterministic, regardless of the problem complexity. This discovery facilitates new sharp asymptotics and a novel alternative proof for the O(\sqrt{n log n}) minimax regret of UCB. Furthermore, the paper also provides the first complete process-level characterization of the MAB problem in the conventional diffusion scaling. Among other things, the "small" gap worst-case lens adopted in this paper also reveals profound distinctions between the behavior of UCB and Thompson Sampling, such as an "incomplete learning" phenomenon characteristic of the latter.
Anand Kalvit, Assaf Zeevi
NeurIPS2
2021 Dynamic Pricing and Learning under the Bass Model
abstract
Most of the dynamic pricing and learning literature has focused on a relatively simple setting where given current pricing decision, demand is independent of past actions and demand values. With the evolution of online platforms and marketplaces, the focus on such homogeneous modeling environments is becoming increasingly less realistic. For example, platforms now rely more and more on online reviews and ratings to inform and guide consumers. Product quality information is also increasingly available on online blogs, discussion forums, and social networks, that create further word-of-mouth effects. One clear implication on the dynamic pricing and learning problem is that the demand environment can no longer be assumed to be static; for example, in the context of online reviews, sales of the product trigger reviews/ratings, and these in turn influence subsequent demand behavior etc. To that end, product diffusion models, such as the popular Bass model [1, 2], are known to be extremely robust and parsimonious, capturing aforementioned word-of-mouth and imitation effects on the growth in sales of a new product. The Bass model describes the process by which new products get adopted as an interaction between existing users and potential new users. It creates a state-dependent evolution of market response which is well aligned with the impact of recent technological developments, such as online review platforms, on the customer purchase behavior.
Shipra Agrawal 0001, Steven Yin, Assaf Zeevi
EC3
2020 From Finite to Countable-Armed Bandits
abstract
We consider a stochastic bandit problem with countably many arms that belong to a finite set of types, each characterized by a unique mean reward. In addition, there is a fixed distribution over types which sets the proportion of each type in the population of arms. The decision maker is oblivious to the type of any arm and to the aforementioned distribution over types, but perfectly knows the total number of types occurring in the population of arms. We propose a fully adaptive online learning algorithm that achieves O(log n) distribution-dependent expected cumulative regret after any number of plays n, and show that this order of regret is best possible. The analysis of our algorithm relies on newly discovered concentration and convergence properties of optimism-based policies like UCB in finite-armed bandit problems with zero gap, which may be of independent interest.
Anand Kalvit, Assaf Zeevi
NeurIPS2
2020 Towards Problem-dependent Optimal Learning Rates
abstract
We study problem-dependent rates, i.e., generalization errors that scale tightly with the variance or the effective loss at the "best hypothesis." Existing uniform convergence and localization frameworks, the most widely used tools to study this problem, often fail to simultaneously provide parameter localization and optimal dependence on the sample size. As a result, existing problem-dependent rates are often rather weak when the hypothesis class is "rich" and the worst-case bound of the loss is large. In this paper we propose a new framework based on a "uniform localized convergence" principle. We provide the first (moment-penalized) estimator that achieves the optimal variance-dependent rate for general "rich" classes; we also establish improved loss-dependent rate for standard empirical risk minimization.
Yunbei Xu, Assaf Zeevi
NeurIPS2
2018 A General Approach to Multi-Armed Bandits Under Risk Criteria
abstract
Different risk-related criteria have received recent interest in learning problems, where typically each case is treated in a customized manner. In this paper we provide a more systematic approach to analyzing such risk criteria within a stochastic multi-armed bandit (MAB) formulation. We identify a set of general conditions that yield a simple characterization of the oracle rule (which serves as the regret benchmark), and facilitate the design of upper confidence bound (UCB) learning policies. The conditions are derived from problem primitives, primarily focusing on the relation between the arm reward distributions and the (risk criteria) performance metric. Among other things, the work highlights some (possibly non-intuitive) subtleties that differentiate various criteria in conjunction with statistical properties of the arms. Our main findings are illustrated on several widely used objectives such as conditional value-at-risk, mean-variance, Sharpe-ratio, and more.
Asaf B. Cassel, Shie Mannor, Assaf Zeevi
COLT3
2017 Thompson Sampling for the MNL-Bandit
abstract
We consider a sequential subset selection problem under parameter uncertainty, where at each time step, the decision maker selects a subset of cardinality $K$ from $N$ possible items (arms), and observes a (bandit) feedback in the form of the index of one of the items in said subset, or none. Each item in the index set is ascribed a certain value (reward), and the feedback is governed by a Multinomial Logit (MNL) choice model whose parameters are a priori unknown. The objective of the decision maker is to maximize the expected cumulative rewards over a finite horizon $T$, or alternatively, minimize the regret relative to an oracle that knows the MNL parameters. We refer to this as the MNL-Bandit problem. This problem is representative of a larger family of exploration-exploitation problems that involve a combinatorial objective, and arise in several important application domains. We present an approach to adapt Thompson Sampling to this problem and show that it achieves near-optimal regret as well as attractive numerical performance.
Shipra Agrawal 0001, Vashist Avadhanula, Vineet Goyal, Assaf Zeevi
COLT4
2016 A Near-Optimal Exploration-Exploitation Approach for Assortment Selection
abstract
We consider an online assortment optimization problem, where in every round, the retailer offers a K-cardinality subset (assortment) of N substitutable products to a consumer, and observes the response. We model consumer choice behavior using the widely used multinomial logit (MNL) model, and consider the retailer's problem of dynamically learning the model parameters, while optimizing cumulative revenues over the selling horizon T. Formulating this as a variant of a multi-armed bandit problem, we present an algorithm based on the principle of "optimism in the face of uncertainty." A naive MAB formulation would treat each of the N choose K possible assortments as a distinct "arm", leading to regret bounds that are exponential in K. We show that by exploiting the specific characteristics of the MNL model it is possible to design an algorithm with Õ(√NT) regret, under a mild assumption. We demonstrate that this performance is nearly optimal, by providing a (randomized) instance of this problem on which any online algorithm would incur at least ΩOmega(√NT/K) regret.
Shipra Agrawal 0001, Vashist Avadhanula, Vineet Goyal, Assaf Zeevi
EC4
2015 Online Time Series Prediction with Missing Data
abstract
We consider the problem of time series prediction in the presence of missing data. We cast the problem as an online learning problem in which the goal of the learner is to minimize prediction error. We then devise an efficient algorithm for the problem, which is based on autoregressive model, and does not assume any structure on the missing data nor on the mechanism that generates the time series. We show that our algorithm’s performance asymptotically approaches the performance of the best AR predictor in hindsight, and corroborate the theoretic results with an empirical study on synthetic and real-world data.
Oren Anava, Elad Hazan, Assaf Zeevi
ICML3
2014 Stochastic Multi-Armed-Bandit Problem with Non-stationary Rewards
Yonatan Gur, Assaf Zeevi, Omar Besbes
NIPS2
2011 A Note on Performance Limitations in Bandit Problems With Side Information
abstract
We consider a sequential adaptive allocation problem which is formulated as a traditional two armed bandit problem but with one important modification: at each time step t, before selecting which arm to pull, the decision maker has access to a random variable Xtwhich provides information on the reward in each arm. Performance is measured as the fraction of time an inferior arm (generating lower mean reward) is pulled. We derive a minimax lower bound that proves that in the absence of sufficient statistical "diversity" in the distribution of the covariate X, a property that we shall refer to as lack of persistent excitation, no policy can improve on the best achievable performance in the traditional bandit problem without side information.
Alexander Goldenshluger, Assaf Zeevi
IEEE Trans. Inf. Theory2
2010 Nonparametric Bandits with Covariates
Philippe Rigollet, Assaf Zeevi
COLT2
1998 On the Performance of Vector Quantizers Empirically Designed from Dependent Sources
abstract
Suppose we are given n real valued samples Z/sub 1/, Z/sub 2/, ..., Z/sub n/ from a stationary source P. We consider the following question. For a compression scheme that uses blocks of length k, what is the minimal distortion (for encoding the true source P) induced by a vector quantizer of fixed rate R, designed from the training sequence. For a certain class of dependent sources, we derive conditions ensuring that the empirically designed quantizer performs as well (on the average) as the optimal quantizer, for almost every training sequence emitted by the source. In particular, we observe that for a code rate R, the optimal way to choose the dimension of the quantizer is k/sub n/=[(1-/spl delta/)R/sup -1/ log n]. The problem of empirical design of a vector quantizer of fixed dimension k based on a vector valued training sequence X/sub 1/, X/sub 2/, ..., X/sub n/ is also considered. For a class of dependent sources, it is shown that the mean squared error (MSE) of the empirically designed quantizer w.r.t the true source distribution converges to the minimum possible MSE at a rate of O(/spl radic/(log n/n)), for almost every training sequence emitted by the source. In addition, the expected value of the distortion redundancy-the difference between the MSEs of the quantizers-converges to zero for a sequence of increasing block lengths k, if we have at our disposal corresponding training sequences whose length grows as n=2/sup (R+/spl delta/)k/. Some of the derivations extend results in empirical quantizer design using an i.i.d. Training sequence, obtained by Linder et al. (see IEEE Trans. on Info. Theory, vol.40, p.1728-40, 1994) and Merhav and Ziv (see IEEE Trans. on Info. Theory, vol.43, p.1112-23, 1997). Proof of the techniques rely on the results in the theory of empirical processes, indexed by VC function classes.
Assaf Zeevi
Data Compression Conference1
1998 Error Bounds for Functional Approximation and Estimation Using Mixtures of Experts
abstract
We examine some mathematical aspects of learning unknown mappings with the mixture of experts model (MEM). Specifically, we observe that the MEM is at least as powerful as a class of neural networks, in a sense that will be made precise. Upper bounds on the approximation error are established for a wide class of target functions. The general theorem states that /spl par/f-f/sub n//spl par//sub p//spl les/c/n/sup r/d/ for f/spl isin/W/sub p//sup r/(L) (a Sobolev class over [-1,1]/sup d/), and f/sub n/ belongs to an n-dimensional manifold of normalized ridge functions. The same bound holds for the MEM as a special case of the above. The stochastic error, in the context of learning from independent and identically distributed (i.i.d.) examples, is also examined. An asymptotic analysis establishes the limiting behavior of this error, in terms of certain pseudo-information matrices. These results substantiate the intuition behind the MEM, and motivate applications.
Assaf Zeevi, Ron Meir, Vitaly Maiorov
IEEE Trans. Inf. Theory1
1997 Density Estimation Through Convex Combinations of Densities: Approximation and Estimation Bounds
Assaf Zeevi, Ron Meir
Neural Networks1
1996 Time Series Prediction using Mixtures of Experts
Assaf Zeevi, Ron Meir, Robert J. Adler
NIPS1