Michael Nett

dblp:56/7967 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8Artificial intelligence and machine learning · 3Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Spatial and temporal data management · 31% Query processing and optimization · 31% Data mining · 23%
Theoretical computer science
3 papers
Algorithms and data structures · 76% Information theory · 24%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › incremental computation
incremental nearest-neighbor search
0.312017
Dimensional Testing for Reverse k-Nearest Neighbor Search · Proc. VLDB Endow. 2017
Spatial and temporal data management › reverse nearest neighbor
reverse k-nearest neighbor search
0.312017
Dimensional Testing for Reverse k-Nearest Neighbor Search · Proc. VLDB Endow. 2017
Algorithms and data structures
similarity search
0.322015
Rank-Based Similarity Search: Reducing the Dimensional Dependence · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Estimating Local Intrinsic Dimensionality · KDD 2015
Data mining › dimensionality reduction
intrinsic dimensionality estimation
0.212015
Estimating Local Intrinsic Dimensionality · KDD 2015
Algorithms and data structures › similarity search › nearest neighbor search
k-nearest neighbors
0.212015
Rank-Based Similarity Search: Reducing the Dimensional Dependence · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Information theory
intrinsic dimensionality
0.222015
Dimensional Testing for Multi-step Similarity Search · ICDM 2012
Rank-Based Similarity Search: Reducing the Dimensional Dependence · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Information retrieval
similarity search
0.112012
Dimensional Testing for Multi-step Similarity Search · ICDM 2012
Algorithms and data structures › numerical linear algebra
dimensionality reduction
0.112012
Dimensional Testing for Multi-step Similarity Search · ICDM 2012

Methods — techniques the papers use, named apart from their topics

probability weighted moments · 0.4method of moments · 0.4maximum likelihood estimation · 0.4extreme value theory · 0.4heuristic search · 0.3generalized expansion dimension · 0.3termination tests · 0.3pruning · 0.3intrinsic dimensionality characterization · 0.3rank cover tree · 0.2non-metric pruning · 0.2
YearPublicationVenuePosition
2018 Extreme-value-theoretic estimation of local intrinsic dimensionality
Laurent Amsaleg, Oussama Chelly, Teddy Furon, Stéphane Girard, Michael E. Houle, Ken-ichi Kawarabayashi, Michael Nett
Data Min. Knowl. Discov.7
2017 Dimensional Testing for Reverse k-Nearest Neighbor Search
abstract
Given a query object q, reverse k -nearest neighbor (R k NN) search aims to locate those objects of the database that have q among their k -nearest neighbors. In this paper, we propose an approximation method for solving R k NN queries, where the pruning operations and termination tests are guided by a characterization of the intrinsic dimensionality of the data. The method can accommodate any index structure supporting incremental (forward) nearest-neighbor search for the generation and verification of candidates, while avoiding impractically-high preprocessing costs. We also provide experimental evidence that our method significantly outperforms its competitors in terms of the tradeoff between execution time and the quality of the approximation. Our approach thus addresses many of the scalability issues surrounding the use of previous methods in data mining.
Guillaume Casanova, Elias Englmeier, Michael E. Houle, Peer Kröger, Michael Nett, Erich Schubert, Arthur Zimek
Proc. VLDB Endow.5
2015 Estimating Local Intrinsic Dimensionality
abstract
This paper is concerned with the estimation of a local measure of intrinsic dimensionality (ID) recently proposed by Houle. The local model can be regarded as an extension of Karger and Ruhl's expansion dimension to a statistical setting in which the distribution of distances to a query point is modeled in terms of a continuous random variable. This form of intrinsic dimensionality can be particularly useful in search, classification, outlier detection, and other contexts in machine learning, databases, and data mining, as it has been shown to be equivalent to a measure of the discriminative power of similarity functions. Several estimators of local ID are proposed and analyzed based on extreme value theory, using maximum likelihood estimation (MLE), the method of moments (MoM), probability weighted moments (PWM), and regularly varying functions (RV). An experimental evaluation is also provided, using both real and artificial data.
Laurent Amsaleg, Oussama Chelly, Teddy Furon, Stéphane Girard, Michael E. Houle, Ken-ichi Kawarabayashi, Michael Nett
KDD7
2015 Rank-Based Similarity Search: Reducing the Dimensional Dependence
abstract
This paper introduces a data structure for k-NN search, the Rank Cover Tree (RCT), whose pruning tests rely solely on the comparison of similarity values; other properties of the underlying space, such as the triangle inequality, are not employed. Objects are selected according to their ranks with respect to the query object, allowing much tighter control on the overall execution costs. A formal theoretical analysis shows that with very high probability, the RCT returns a correct query result in time that depends very competitively on a measure of the intrinsic dimensionality of the data set. The experimental results for the RCT show that non-metric pruning strategies for similarity search can be practical even when the representational dimension of the data is extremely high. They also show that the RCT is capable of meeting or exceeding the level of performance of state-of-the-art methods that make use of metric pruning or other selection tests involving numerical constraints on distance values.
Michael E. Houle, Michael Nett
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Rank Cover Trees for Nearest Neighbor Search
Michael E. Houle, Michael Nett
SISAP2
2012 Dimensional Testing for Multi-step Similarity Search
abstract
In data mining applications such as subspace clustering or feature selection, changes to the underlying feature set can require the reconstruction of search indices to support fundamental data mining tasks. For such situations, multi-step search approaches have been proposed that can accommodate changes in the underlying similarity measure without the need to rebuild the index. In this paper, we present a heuristic multi-step search algorithm that utilizes a measure of intrinsic dimension, the generalized expansion dimension (GED), as the basis of its search termination condition. Compared to the current state-of-the-art method, experimental results show that our heuristic approach is able to obtain significant improvements in both the number of candidates and the running time, while losing very little in the accuracy of the query results.
Michael E. Houle, Xiguo Ma, Michael Nett, Vincent Oria
ICDM3
2012 Fast Similarity Computation in Factorized Tensors
Michael E. Houle, Hisashi Kashima, Michael Nett
SISAP3
2010 Air-Indexing on Error Prone Communication Channels
Emmanuel Müller, Philipp Kranen, Michael Nett, Felix Reidl, Thomas Seidl 0001
DASFAA (1)3
2009 SEICOS: semantically enriched interactive collaborative online shopping
abstract
In this paper we investigate the potential benefits from integrating Semantic Web technologies into the context of collaboration in virtual 3D environments. We provide a prototype e-commerce solution, the Seicos system, based on the SecondLife client and the OpenSimulator server. In the system, semantic information is retrieved and integrated from multiple heterogeneous sources and subsequently cached and accessed using Semantic Web methods. The semantic information gathered by our system can support the customer and is integrated into a knowledge base which is used for processing queries.
Andreas Budde, Michael Nett, Florian Wagner 0001
iiWAS2