Dan Kushnir

dblp:87/231 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-authorComputer networks · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 70% Machine learning and data management · 30%
Computer graphics and multimedia
1 paper
Image and video processing · 50% Geometric modeling and processing · 50%
Artificial intelligence
1 paper
Representation and self-supervised learning · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management
active learning
0.212014
Active-transductive learning with label-adapted kernels · KDD 2014
Data mining › predictive modeling
classification
0.212014
Active-transductive learning with label-adapted kernels · KDD 2014
Data mining
clustering
0.112010
Efficient Multilevel Eigensolvers with Applications to Data Analysis Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Data mining › clustering
spectral clustering
0.112010
Efficient Multilevel Eigensolvers with Applications to Data Analysis Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Image and video processing
image segmentation
0.112010
Efficient Multilevel Eigensolvers with Applications to Data Analysis Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Geometric modeling and processing
spectral methods
0.112010
Efficient Multilevel Eigensolvers with Applications to Data Analysis Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.112008
Manifold Learning: The Price of Normalization · J. Mach. Learn. Res. 2008
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.112008
Manifold Learning: The Price of Normalization · J. Mach. Learn. Res. 2008
Data mining
dimensionality reduction
0.012010
Efficient Multilevel Eigensolvers with Applications to Data Analysis Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2010

Methods — techniques the papers use, named apart from their topics

multigrid interpolation · 0.2lanczos algorithm · 0.2algebraic multigrid · 0.2spectral graph theory · 0.2random walk · 0.2minimal-cut · 0.2spectral methods · 0.1normalization · 0.1
YearPublicationVenuePosition
2020 Active Community Detection with Maximal Expected Model Change
abstract
We present a novel active learning algorithm for community detection on networks. Our proposed algorithm uses a Maximal Expected Model Change (MEMC) criterion for querying network nodes label assignments. MEMC detects nodes that maximally change the community assignment likelihood model following a query. Our method is inspired by detection in the benchmark Stochastic Block Model (SBM), where we provide sample complexity analysis and empirical study with SBM and real network data for binary as well as for the multi-class settings. The analysis also covers the most challenging case of sparse degree and below-detection-threshold SBMs, where we observe a super-linear error reduction. MEMC is shown to be superior to the random selection baseline and other state-of-the-art active learners.
Dan Kushnir, Benjamin Mirabelli
AISTATS1
2019 Towards Clustering High-dimensional Gaussian Mixture Clouds in Linear Running Time
abstract
Clustering mixtures of Gaussian distributions is a fundamental and challenging problem. State-of-the-art theoretical work on learning Gaussian mixture models has mostly focused on estimating the mixture parameters, where clustering is given as a byproduct. These methods have focused mostly on improving separation bounds for different mixture classes, and doing so in polynomial time and sample complexity. Less emphasis has been given to aligning these algorithms to the challenges of big data. In this paper, we focus on clustering $n$ samples from an arbitrary mixture of $c$-separated Gaussians in $\mathbb{R}^p$ in time that is linear in $p$ and $n$, and sample complexity that is independent of $p$. Our analysis suggests that for sufficiently separated Gaussians after $o(\log{p})$ random projections a good direction is found that yields a small clustering error. Specifically, for a user-specified error $e$, the expected number of such projections is small and bounded by $o(\ln p)$ when $\gamma\leq c\sqrt{\ln{\ln{p}}}$ and $\gamma=Q^{-1}(e)$ is the separation of the Gaussians with $Q$ as the tail distribution function of the normal distribution. Consequently, the expected overall running time of the algorithm is linear in $n$ and quasi-linear in $p$ at $o(\ln{p})O(np)$, and the sample complexity is independent of $p$. Unlike the methods that are based on $k$-means, our analysis is applicable to any mixture class (spherical or non-spherical). Finally, an extension to $k>2$ components is also provided.
Dan Kushnir, Shirin Jalali, Iraj Saniee
AISTATS1
2018 Predicting Outages in Radio Networks with Alarm Data
abstract
Modern cellular networks are complex systems offering a wide range of services and present challenges in detecting anomalous events when they do occur. The networks are engineered for high reliability and, hence, the data from these networks is predominantly normal with a small proportion being anomalous. From an operations perspective, it is important to detect these anomalies in a timely manner in order to mitigate them and preclude the occurrence of major failure events. In telecommunications diverse set of data such as KPIs, logs and alarms are generated to monitor the health and stability of the network element and the services carried over it [1]–[4].
Dan Kushnir, Gautam Gohil, Zulfiquar Sayeed, Hüseyin Uzunalioglu
IWQoS1
2017 Detecting and predicting outages in mobile networks with log data
abstract
Modern cellular networks are complex systems offering a wide range of services and present challenges in detecting anomalous events when they do occur. The networks are engineered for high reliability and, hence, the data from these networks is predominantly normal with a small proportion being anomalous. From an operations perspective, it is important to detect these anomalies in a timely manner, to correct vulnerabilities in the network and preclude the occurrence of major failure events. The objective of our work is anomaly detection in cellular networks in near real-time to improve network performance and reliability. We use performance data from a 4G LTE network to develop a methodology for anomaly detection in such networks. Two rigorous prediction models are proposed: a non-parametric approach (Chi-Square test), and a parametric one (Gaussian Mixture Models). These models are trained to detect differences between distributions to classify a target distribution as belonging to a normal period or abnormal period with high accuracy. We discuss the merits between the approaches and show that both provide a more nuanced view of the network than simple thresh-olds of success/failure used by operators in production networks today.
Vijay K. Gurbani, Dan Kushnir, Veena B. Mendiratta, Chitra Phadke, Eric Falk, Radu State
ICC2
2014 Active-transductive learning with label-adapted kernels
abstract
This paper presents an efficient active-transductive approach for classification. A common approach of active learning algorithms is to focus on querying points near the class boundary in order to refine it. However, for certain data distributions, this approach has been shown to lead to uninformative samples. More recent approaches consider combining data exploration with traditional refinement techniques. These techniques typically require tuning sampling of unexplored regions with refinement of detected class boundaries. They also involve significant computational costs for the exploration of informative query candidates. We present a novel iterative active learning algorithm designed to overcome these shortcomings by using a linear running-time active-transductive learning approach that naturally switches from exploration to refinement. The passive classifier employed in our algorithm builds a random-walk on the data graph based on a modified graph geometry that combines the data distribution with current label hypothesis; while the query component uses the uncertainty of the evolving hypothesis. Our supporting theory draws the link between the spectral properties of our iteration matrix and a solution to the minimal-cut problem for a fused hypothesis-data graph. Experiments demonstrate computational complexity that is orders of magnitude lower than state-of-the-art, and competitive results on benchmark data and real churn prediction data.
Dan Kushnir
KDD1
2010 Efficient Multilevel Eigensolvers with Applications to Data Analysis Tasks
abstract
Multigrid solvers proved very efficient for solving massive systems of equations in various fields. These solvers are based on iterative relaxation schemes together with the approximation of the "smooth" error function on a coarser level (grid). We present two efficient multilevel eigensolvers for solving massive eigenvalue problems that emerge in data analysis tasks. The first solver, a version of classical algebraic multigrid (AMG), is applied to eigenproblems arising in clustering, image segmentation, and dimensionality reduction, demonstrating an order of magnitude speedup compared to the popular Lanczos algorithm. The second solver is based on a new, much more accurate interpolation scheme. It enables calculating a large number of eigenvectors very inexpensively.
Dan Kushnir, Meirav Galun, Achi Brandt
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Manifold Learning: The Price of Normalization
Yair Goldberg, Alon Zakai, Dan Kushnir, Yaacov Ritov
J. Mach. Learn. Res.3
2006 Fast multiscale clustering and manifold identification
Dan Kushnir, Meirav Galun, Achi Brandt
Pattern Recognit.1