VLDB 2026 Research / reviewers in the wild / expert
Milos Radovanovic 0001
dblp:79/5222
· DBLP profile ↗
38ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0003-2225-7803ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 23 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Security and privacy · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unveiling distinguishing properties of adversarial examples through neural tangent kernelsabstractDeep neural networks (DNNs), in addition to great power, exhibit significant weaknesses, with one of the most notable being the susceptibility to adversarial attacks by carefully crafted, minimally perturbed, adversarial examples. Despite numerous efforts towards creating reliable detectors and defense mechanisms, the underlying factors that make such attacks possible have not yet been fully understood. In this paper, we further explore the use of local intrinsic dimensionality (LID), and, for the first time, hubness, as indicators for detection of adversarial examples. Furthermore, we involve neural tangent kernels (NTKs) in this task, showing that NTK-induced distance measures offer more information on how to detect adversarial examples, through both LID and hubness, than distances computed on raw input data or one network layer. Furthermore, we combine LID and hubness as features into a well-performing classifier that exhibits better detection accuracy than any of the two features used on their own. Finally, we combine the LID and hubness features with state-of-the-art adversarial attack detectors, demonstrating their general usefulness for the task. Our findings offer insight into the properties and nature of adversarial examples of various types, paving the way to the construction of better detectors and, possibly, the construction of models that are more robust to such attacks. Alexandros Nanopoulos, Milos Radovanovic 0001 |
Neurocomputing | 2 |
| 2025 | Dynamic Graph Embedding Through Hub-Aware Random Walks
Aleksandar Tomcic, Milos Savic 0001, Dusan Simic, Milos Radovanovic 0001 |
SISAP | 4 |
| 2025 | Domain adaptation for improving automatic airborne pollen classification with expert-verified measurementsabstractAbstract This study presents a novel approach to enhance the accuracy of automatic classification systems for airborne pollen particles by integrating domain adaptation techniques. Our method incorporates expert-verified measurements into the convolutional neural network (CNN) training process to address the discrepancy between laboratory test data and real-world environmental measurements. We systematically fine-tuned CNN models, initially developed on standard reference datasets, with these expert-verified measurements. A comprehensive exploration of hyperparameters was conducted to optimize the CNN models, ensuring their robustness and adaptability across various environmental conditions and pollen types. Empirical results indicate a significant improvement, evidenced by a 22.52% increase in correlation and a 38.05% reduction in standard deviation across 29 cases of different pollen classes over multiple study years. This research highlights the potential of domain adaptation techniques in environmental monitoring, particularly in contexts where the integrity and representativeness of reference datasets are difficult to verify. Predrag Matavulj, Slobodan Jelic, Domagoj Severdija, Sanja Brdar, Milos Radovanovic 0001, Danijela Tesendic, Branko Sikoparija |
Appl. Intell. | 5 |
| 2024 | Dimensionality-Aware Outlier DetectionabstractWe present a nonparametric method for outlier detection that takes full account of local variations in intrinsic dimensionality within the dataset. Using the theory of Local Intrinsic Dimensionality (LID), our ‘dimensionality-aware’ outlier detection method, DAO, is derived as an estimator of an asymptotic local expected density ratio involving the query point and a close neighbor drawn at random. The dimensionality-aware behavior of DAO is due to its use of local estimation of LID values in a theoretically-justified way. Through comprehensive experimentation on more than 800 synthetic and real datasets, we show that DAO significantly outperforms three popular and important benchmark outlier detection methods: Local Outlier Factor (LOF), Simplified LOF, and kNN. Alastair Anderberg, James Bailey 0001, Ricardo J. G. B. Campello, Michael E. Houle, Henrique O. Marques, Milos Radovanovic 0001, Arthur Zimek |
SDM | 6 |
| 2024 | Multi-class boosting for the analysis of multiple incomplete views on microbiome dataabstractBACKGROUND: Microbiome dysbiosis has recently been associated with different diseases and disorders. In this context, machine learning (ML) approaches can be useful either to identify new patterns or learn predictive models. However, data to be fed to ML methods can be subject to different sampling, sequencing and preprocessing techniques. Each different choice in the pipeline can lead to a different view (i.e., feature set) of the same individuals, that classical (single-view) ML approaches may fail to simultaneously consider. Moreover, some views may be incomplete, i.e., some individuals may be missing in some views, possibly due to the absence of some measurements or to the fact that some features are not available/applicable for all the individuals. Multi-view learning methods can represent a possible solution to consider multiple feature sets for the same individuals, but most existing multi-view learning methods are limited to binary classification tasks or cannot work with incomplete views. RESULTS: We propose irBoost.SH, an extension of the multi-view boosting algorithm rBoost.SH, based on multi-armed bandits. irBoost.SH solves multi-class classification tasks and can analyze incomplete views. At each iteration, it identifies one winning view using adversarial multi-armed bandits and uses its predictions to update a shared instance weight distribution in a learning process based on boosting. In our experiments, performed on 5 multi-view microbiome datasets, the model learned by irBoost.SH always outperforms the best model learned from a single view, its closest competitor rBoost.SH, and the model learned by a multi-view approach based on feature concatenation, reaching an improvement of 11.8% of the F1-score in the prediction of the Autism Spectrum disorder and of 114% in the prediction of the Colorectal Cancer disease. CONCLUSIONS: The proposed method irBoost.SH exhibited outstanding performances in our experiments, also compared to competitor approaches. The obtained results confirm that irBoost.SH can fruitfully be adopted for the analysis of microbiome data, due to its capability to simultaneously exploit multiple feature sets obtained through different sequencing and preprocessing pipelines. Andrea Simeon, Milos Radovanovic 0001, Tatjana Loncar-Turukalo, Michelangelo Ceci, Sanja Brdar, Gianvito Pio |
BMC Bioinform. | 2 |
| 2023 | Local intrinsic dimensionality measures for graphs, with applications to graph embeddings
Milos Savic 0001, Vladimir Kurbalija, Milos Radovanovic 0001 |
Inf. Syst. | 3 |
| 2022 | Evaluation of LID-Aware Graph Embedding Methods for Node Clustering
Dusica Knezevic, Jela Babic, Milos Savic 0001, Milos Radovanovic 0001 |
SISAP | 4 |
| 2022 | Elastic distances for time-series classification: Itakura versus Sakoe-Chiba constraints
Zoltan Geler, Vladimir Kurbalija, Mirjana Ivanovic, Milos Radovanovic 0001 |
Knowl. Inf. Syst. | 4 |
| 2021 | Local Intrinsic Dimensionality and Graphs: Towards LID-aware Graph Embedding Algorithms
Milos Savic 0001, Vladimir Kurbalija, Milos Radovanovic 0001 |
SISAP | 3 |
| 2021 | High Intrinsic Dimensionality Facilitates Adversarial Attack: Theoretical EvidenceabstractMachine learning systems are vulnerable to adversarial attack. By applying to the input object a small, carefully-designed perturbation, a classifier can be tricked into making an incorrect prediction. This phenomenon has drawn wide interest, with many attempts made to explain it. However, a complete understanding is yet to emerge. In this paper we adopt a slightly different perspective, still relevant to classification. We consider retrieval, where the output is a set of objects most similar to a user-supplied query object, corresponding to the set of k-nearest neighbors. We investigate the effect of adversarial perturbation on the ranking of objects with respect to a query. Through theoretical analysis, supported by experiments, we demonstrate that as the intrinsic dimensionality of the data domain rises, the amount of perturbation required to subvert neighborhood rankings diminishes, and the vulnerability to adversarial attack rises. We examine two modes of perturbation of the query: either `closer' to the target point, or `farther' from it. We also consider two perspectives: `query-centric', examining the effect of perturbation on the query's own neighborhood ranking, and `target-centric', considering the ranking of the query point in the target's neighborhood set. All four cases correspond to practical scenarios involving classification and retrieval. Laurent Amsaleg, James Bailey 0001, Amélie Barbe, Sarah M. Erfani, Teddy Furon, Michael E. Houle, Milos Radovanovic 0001, Xuan Vinh Nguyen |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2020 | Time-Series Classification with Constrained DTW Distance and Inverse-Square Weighted k-NNabstractThe problem of time-series classification witnessed the application of many techniques for data mining and machine learning, including neural networks, support vector machines, and Bayesian approaches. Somewhat surprisingly, the simple 1-nearest neighbor (1NN) classifier, in combination with the Dynamic Time Warping (DTW) distance measure, is still competitive and not rarely superior to more advanced classification methods, which includes the majority-voting k-nearest neighbor (kNN) classifier. In this paper we focus on the kNN classifier combined with the inverse-squared weighting scheme, and its interaction with constrained DTW distance. By performing experiments on the entire UCR Time Series Classification Archive we show that with proper selection of the constraint parameter r and neighborhood size k, inverse-square weighted kNN consistently outperforms 1NN. Zoltan Geler, Vladimir Kurbalija, Mirjana Ivanovic, Milos Radovanovic 0001 |
INISTA | 4 |
| 2020 | An MQTT-based Resource Management Framework for Edge Computing SystemsabstractThe complexity of IoT systems and tasks that are put before them require shifts in the way resources and service provisioning are managed. The concept of edge computing is introduced to enhance IoT systems' scalability, reactivity, efficiency, and privacy. In this paper, we present an edge computing solution for resource management of context-aware decision-making processes distributed between IoT gateways. The solution performs decision-making process management for smart actuation, based on analysis of sensory data streams, and context-informed edge computing resource and service provisioning management based on topology and operational changes. Our architectural solution showcases the first version of a Resource Management Framework - a generic framework for software resource orchestration best-suited to IoT platforms with event-driven, publish-subscribe communication mechanisms. Proof of concept experiments that are executed in a simulated edge computing testbed validate our solution's performance in improving the resilience and responsiveness of the edge computing system when there are operational and topology changes. Furthermore, the framework addresses the recovery of failed decision-making processes, impacting the overall health of the underlying IoT system. Sasa Pesic, Milos Radovanovic 0001, Mirjana Ivanovic |
INISTA | 2 |
| 2020 | Weighted kNN and constrained elastic distances for time-series classification
Zoltan Geler, Vladimir Kurbalija, Mirjana Ivanovic, Milos Radovanovic 0001 |
Expert Syst. Appl. | 4 |
| 2019 | Dynamic Time Warping: Itakura vs Sakoe-ChibaabstractIn the domain of time-series classification, one simple but persistently successful method is the 1-nearest neighbour (1NN) classifier coupled with an elastic distance measure such as Dynamic Time Warping (DTW). In this paper we evaluate the performance of DTW when constrained using the Itakura parallelogram, and compare it with the more commonly used Sakoe-Chiba band, as well as with the unconstrained DTW. Results show that although the Itakura parallelogram is generally inferior to the Sakoe-Chiba band, it is still superior to unconstrained DTW. Furthermore, on individual data sets the Itakura parallelogram can produce superior results, warranting further investigation into the merits of its use with DTW and other elastic distance measures for time-series classification. Zoltan Geler, Vladimir Kurbalija, Mirjana Ivanovic, Milos Radovanovic 0001, Weihui Dai |
INISTA | 4 |
| 2019 | Hyperledger Fabric Blockchain as a Service for the IoT: Proof of Concept
Sasa Pesic, Milos Radovanovic 0001, Mirjana Ivanovic, Milenko Tosic, Ognjen Ikovic, Dragan Boskovic |
MEDI | 2 |
| 2019 | CAAVI-RICS Model for Analyzing the Security of Fog Computing Systems: AuthenticationabstractThe overarching connectivity of "things" in the Internet of Things presents an appealing environment for innovation and business ventures, but also brings a certain set of security challenges. Engineering secure Internet of Things systems requires addressing the peculiar circumstances under which they operate: constraints due to limited resources, high node churn, decentralized decision making, direct interfacing with end users etc. Thus, techniques and methodologies for building secure and robust Internet of Things systems should support these conditions. In this paper, we are presenting a description of the CAAVI-RICS framework, a novel security review methodology tightly coupled with distributed, Internet of Things and fog computing systems. With CAAVI-RICS we are exploring credibility, authentication, authorization, verification, and integrity (CAAVI) through explaining the rationale, influence, concerns and security solutions (RICS) that accompany them. Our contribution is a thorough systematic categorization and rationalization of security issues, covering the security landscape of Internet of Things/fog computing systems, as well as contributing to the discussion on the aspects of fog computing security and state-of-the-art solutions. Specifically, in this paper we explore the Authentication in Internet of Things systems through the RICS review methodology. Sasa Pesic, Milos Radovanovic 0001, Mirjana Ivanovic, Costin Badica, Milenko Tosic, Ognjen Ikovic, Dragan Boskovic |
PDCAT | 2 |
| 2019 | Intrinsic Dimensionality Estimation within Tight LocalitiesabstractAccurate estimation of Intrinsic Dimensionality (ID) is of crucial importance in many data mining and machine learning tasks, including dimensionality reduction, outlier detection, similarity search and subspace clustering. However, since their convergence generally requires sample sizes (that is, neighborhood sizes) on the order of hundreds of points, existing ID estimation methods may have only limited usefulness for applications in which the data consists of many natural groups of small size. In this paper, we propose a local ID estimation strategy stable even for ‘tight’ localities consisting of as few as 20 sample points. The estimator applies MLE techniques over all available pairwise distances among the members of the sample, based on a recent extreme-value-theoretic model of intrinsic dimensionality, the Local Intrinsic Dimension (LID). Our experimental results show that our proposed estimation technique can achieve notably smaller variance, while maintaining comparable levels of bias, at much smaller sample sizes than state-of-the-art estimators. Laurent Amsaleg, Oussama Chelly, Michael E. Houle, Ken-ichi Kawarabayashi, Milos Radovanovic 0001, Weeris Treeratanajaru |
SDM | 5 |
| 2018 | Gender-Based Analysis of Intra-Institutional Research Productivity and CollaborationabstractCurrent Research Information Systems (CRISs) offer great opportunities for assessments of institutional research outputs and extraction of useful and actionable knowledge based on various data-analysis techniques. However, many of these opportunities have not been explored in depth, especially in c ulture-sensitive areas such as gender-based analysis of research productivity and collaboration. In this paper we present GERBER, a network-based methodology and accompanying tool for gender-based analysis of publication data stored in institutional CRISs. GERBER relies on statistically robust techniques applied on weighted co-authorship networks whose nodes are enriched with different types of researcher evaluation metrics. The functionality of GERBER is demonstrated on publication data stored in the institutional CRIS of the Faculty of Sciences, University of Novi Sad, Serbia. The obtained results show that GERBER enables institutional research managers and policy makers to detect gender inequalities and homophily in research productivity and collaboration. Finally, we discuss different possibilities to integrate GERBER with CRISs in order to facilitate continuous gender-based evaluation of researchers. Milos Savic 0001, Mirjana Ivanovic, Milos Radovanovic 0001, Bojana Dimic Surla |
Fundam. Informaticae | 3 |
| 2016 | Flattening the Density Gradient for Eliminating Spatial Centrality to Reduce HubnessabstractSpatial centrality, whereby samples closer to the center of a dataset tend to be closer to all other samples, is regarded as one source of hubness. Hubness is well known to degrade k-nearest-neighbor (k-NN) classification. Spatial centrality can be removed by centering, i.e., shifting the origin to the global center of the dataset, in cases where inner product similarity is used. However, when Euclidean distance is used, centering has no effect on spatial centrality because the distance between the samples is the same before and after centering. As described in this paper, we propose a solution for the hubness problem when Euclidean distance is considered. We provide a theoretical explanation to demonstrate how the solution eliminates spatial centrality and reduces hubness. We then present some discussion of the reason the proposed solution works, from a viewpoint of density gradient, which is regarded as the origin of spatial centrality and hubness. We demonstrate that the solution corresponds to flattening the density gradient. Using real-world datasets, we demonstrate that the proposed method improves k-NN classification performance and outperforms an existing hub-reduction method. Kazuo Hara, Ikumi Suzuki, Kei Kobayashi, Kenji Fukumizu, Milos Radovanovic 0001 |
AAAI | 5 |
| 2016 | Towards Culture-Sensitive Extensions of CRISs: Gender-Based Researcher Evaluation
Milos Savic 0001, Mirjana Ivanovic, Milos Radovanovic 0001, Bojana Dimic Surla |
MEDI | 3 |
| 2016 | Comparison of different weighting schemes for the kNN classifier on time-series data
Zoltan Geler, Vladimir Kurbalija, Milos Radovanovic 0001, Mirjana Ivanovic |
Knowl. Inf. Syst. | 3 |
| 2015 | Localized Centering: Reducing Hubness in Large-Sample DataabstractHubness has been recently identified as a problematic phenomenon occurring in high-dimensional space. In this paper, we address a different type of hubness that occurs when the number of samples is large. We investigate the difference between the hubness in high-dimensional data and the one in large-sample data. One finding is that centering, which is known to reduce the former, does not work for the latter. We then propose a new hub-reduction method, called localized centering. It is an extension of centering, yet works effectively for both types of hubness. Using real-world datasets consisting of a large number of documents, we demonstrate that the proposed method improves the accuracy of k-nearest neighbor classification. Kazuo Hara, Ikumi Suzuki, Masashi Shimbo, Kei Kobayashi, Kenji Fukumizu, Milos Radovanovic 0001 |
AAAI | 6 |
| 2015 | Introducing cultural issues and cultural awareness in conceptual modelling educationabstractConceptual modelling education is often a part of any Informatics and Computer Science curricula but is also included in other study programs where conceptual modelling knowledge is needed. Conceptual modelling is well defined and numerous sources are available in existing literature. The steps to develop conceptual models are easy to learn and support is possible also via different tools, based on a theoretical background, mostly through the Peter Chen approach. In general, such an educational approach can be described as a standard conceptual modelling education. But, unfortunately, standard conceptual modeling education does not inform culturally-aware developers of conceptual models. The confidence and assurance that cross cultural issues and cultural awareness will be considered in the phase of conceptual modelling could be reached on one side by upgrading the conceptual modeling approach with a cultural point of view, as well as by introducing cultural issues into conceptual modelling education. Students need to be aware of basic cultural issues and concepts and their influence on conceptual modeling. In the long term, the mentioned content should become part of standard conceptual modeling education. This means changes in the content of curricula as well as a change in thinking by students and teachers. Tatjana Welzer, Marjan Druzovec, Lili Nemec Zlatolas, Marko Hölbl, Hannu Jaakkola, Mirjana Ivanovic, Milos Radovanovic 0001 |
EJC | 7 |
| 2015 | Reducing Hubness for Kernel Regression
Kazuo Hara, Ikumi Suzuki, Kei Kobayashi, Kenji Fukumizu, Milos Radovanovic 0001 |
SISAP | 5 |
| 2015 | Reverse Nearest Neighbors in Unsupervised Distance-Based Outlier DetectionabstractOutlier detection in high-dimensional data presents various challenges resulting from the “curse of dimensionality.” A prevailing view is that distance concentration, i.e., the tendency of distances in high-dimensional data to become indiscernible, hinders the detection of outliers by making distance-based methods label all points as almost equally good outliers. In this paper, we provide evidence supporting the opinion that such a view is too simple, by demonstrating that distance-based methods can produce more contrasting outlier scores in high-dimensional settings. Furthermore, we show that high dimensionality can have a different impact, by reexamining the notion of reverse nearest neighbors in the unsupervised outlier-detection context. Namely, it was recently observed that the distribution of points' reverse-neighbor counts becomes skewed in high dimensions, resulting in the phenomenon known as hubness. We provide insight into how some points (antihubs) appear very infrequently in k-NN lists of other points, and explain the connection between antihubs, outliers, and existing unsupervised outlier-detection methods. By evaluating the classic k-NN method, the angle-based technique designed for high-dimensional data, the density-based local outlier factor and influenced outlierness methods, and antihub-based methods on various synthetic and real-world data sets, we offer novel insight into the usefulness of reverse neighbor counts in unsupervised outlier detection. Milos Radovanovic 0001, Alexandros Nanopoulos, Mirjana Ivanovic |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | Impact of the Sakoe-Chiba Band on the DTW Time Series Distance Measure for kNN Classification
Zoltan Geler, Vladimir Kurbalija, Milos Radovanovic 0001, Mirjana Ivanovic |
KSEM | 3 |
| 2014 | The influence of global constraints on similarity measures for time-series databases
Vladimir Kurbalija, Milos Radovanovic 0001, Zoltan Geler, Mirjana Ivanovic |
Knowl. Based Syst. | 2 |
| 2014 | The Role of Hubness in Clustering High-Dimensional DataabstractHigh-dimensional data arise naturally in many domains, and have regularly presented a great challenge for traditional data mining techniques, both in terms of effectiveness and efficiency. Clustering becomes difficult due to the increasing sparsity of such data, as well as the increasing difficulty in distinguishing distances between data points. In this paper, we take a novel perspective on the problem of clustering high-dimensional data. Instead of attempting to avoid the curse of dimensionality by observing a lower dimensional feature subspace, we embrace dimensionality by taking advantage of inherently high-dimensional phenomena. More specifically, we show that hubness, i.e., the tendency of high-dimensional data to contain points (hubs) that frequently occur in k-nearest-neighbor lists of other points, can be successfully exploited in clustering. We validate our hypothesis by demonstrating that hubness is a good measure of point centrality within a high-dimensional data cluster, and by proposing several hubness-based clustering algorithms, showing that major hubs can be used effectively as cluster prototypes or as guides during the search for centroid-based cluster configurations. Experimental results demonstrate good performance of our algorithms in multiple settings, particularly in the presence of large quantities of noise. The proposed methods are tailored mostly for detecting approximately hyperspherical clusters and need to be extended to properly handle clusters of arbitrary shapes. Nenad Tomasev, Milos Radovanovic 0001, Dunja Mladenic, Mirjana Ivanovic |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2011 | A probabilistic approach to nearest-neighbor classification: naive hubness bayesian kNNabstractMost machine-learning tasks, including classification, involve dealing with high-dimensional data. It was recently shown that the phenomenon of hubness, inherent to high-dimensional data, can be exploited to improve methods based on nearest neighbors (NNs). Hubness refers to the emergence of points (hubs) that appear among the k NNs of many other points in the data, and constitute influential points for kNN classification. In this paper, we present a new probabilistic approach to kNN classification, naive hubness Bayesian k-nearest neighbor (NHBNN), which employs hubness for computing class likelihood estimates. Experiments show that NHBNN compares favorably to different variants of the kNN classifier, including probabilistic kNN (PNN) which is often used as an underlying probabilistic framework for NN classification, signifying that NHBNN is a promising alternative framework for developing probabilistic NN algorithms. Nenad Tomasev, Milos Radovanovic 0001, Dunja Mladenic, Mirjana Ivanovic |
CIKM | 2 |
| 2011 | The Role of Hubness in Clustering High-Dimensional Data
Nenad Tomasev, Milos Radovanovic 0001, Dunja Mladenic, Mirjana Ivanovic |
PAKDD (1) | 2 |
| 2010 | Time-Series Classification in Many Intrinsic DimensionsabstractIn the context of many data mining tasks, high dimensionality was shown to be able to pose significant problems, commonly referred to as different aspects of the curse of dimensionality. In this paper, we investigate in the time-series domain one aspect of the dimensionality curse called hubness, which refers to the tendency of some instances in a data set to become hubs by being included in unexpectedly many k-nearest neighbor lists of other instances. Through empirical measurements on a large collection of time-series data sets we demonstrate that the hubness phenomenon is caused by high intrinsic dimensionality of time-series data, and shed light on the mechanism through which hubs emerge, focusing on the popular and successful dynamic time warping (DTW) distance. Also, the interaction between hubness and the information provided by class labels is investigated, by considering label matches and mismatches between neighboring time series. Following our findings we formulate a framework for categorizing time-series data sets based on measurements that reflect hubness and the diversity of class labels among nearest neighbors. The framework allows one to assess whether hubness can be successfully used to improve the performance of k-NN classification. Finally, the merits of the framework are demonstrated through experimental evaluation of 1-NN and k-NN classifiers, including a proposed weighting scheme that is designed to make use of hubness information. Our experimental results show that the examined framework, in the majority of cases, is able to correctly reflect the circumstances in which hubness information can effectively be employed in k-NN time-series classification. Milos Radovanovic 0001, Alexandros Nanopoulos, Mirjana Ivanovic |
SDM | 1 |
| 2010 | On the existence of obstinate results in vector space modelsabstractThe vector space model (VSM) is a popular and widely applied model in information retrieval (IR). VSM creates vector spaces whose dimensionality is usually high (e.g., tens of thousands of terms). This may cause various problems, such as susceptibility to noise and difficulty in capturing the underlying semantic structure, which are commonly recognized as different aspects of the "curse of dimensionality." In this paper, we investigate a novel aspect of the dimensionality curse, which is referred to as hubness and manifested by the tendency of some documents (called hubs) to be included in unexpectedly many search result lists. Hubness may impact VSM considerably since hubs can become obstinate results, irrelevant to a large number of queries, thus harming the performance of an IR system and the experience of its users. We analyze the origins of hubness, showing it is primarily a consequence of high (intrinsic) dimensionality of data, and not a result of other factors such as sparsity and skewness of the distribution of term frequencies. We describe the mechanisms through which hubness emerges by exploring the behavior of similarity measures in high-dimensional vector spaces. Our consideration begins with the classical VSM (tf-idf term weighting and cosine similarity), but the conclusions generalize to more advanced variations, such as Okapi BM25. Moreover, we explain why hubness may not be easily mitigated by dimensionality reduction, and propose a similarity adjustment scheme that takes into account the existence of hubs. Experimental results over real data indicate that significant improvement can be obtained through consideration of hubness. Milos Radovanovic 0001, Alexandros Nanopoulos, Mirjana Ivanovic |
SIGIR | 1 |
| 2010 | Hubs in Space: Popular Nearest Neighbors in High-Dimensional Data
Milos Radovanovic 0001, Alexandros Nanopoulos, Mirjana Ivanovic |
J. Mach. Learn. Res. | 1 |
| 2009 | Nearest neighbors in high-dimensional data: the emergence and influence of hubsabstractHigh dimensionality can pose severe difficulties, widely recognized as different aspects of the curse of dimensionality. In this paper we study a new aspect of the curse pertaining to the distribution of k-occurrences, i.e., the number of times a point appears among the k nearest neighbors of other points in a data set. We show that, as dimensionality increases, this distribution becomes considerably skewed and hub points emerge (points with very high k-occurrences). We examine the origin of this phenomenon, showing that it is an inherent property of high-dimensional vector space, and explore its influence on applications based on measuring distances in vector spaces, notably classification, clustering, and information retrieval. Milos Radovanovic 0001, Alexandros Nanopoulos, Mirjana Ivanovic |
ICML | 1 |
| 2009 | How does high dimensionality affect collaborative filtering?abstractA crucial operation in memory-based collaborative filtering (CF) is determining nearest neighbors (NNs) of users/items. This paper addresses two phenomena that emerge when CF algorithms perform NN search in high-dimensional spaces that are typical in CF applications. The first is similarity concentration and the second is the appearance of hubs (i.e. points which appear in $k$-NN lists of many other points). Through theoretical analysis and experimental evaluation we show that these phenomena are inherent properties of high-dimensional space, unrelated to other data properties like sparsity, and that they can impact CF algorithms by questioning the meaning and representativeness of discovered NNs. Moreover, we show that it is not easy to mitigate the phenomena using dimensionality reduction. Studying these phenomena aims to provide a better understanding of the limitations of memory-based CF and motivate the development of new algorithms that would overcome them. Alexandros Nanopoulos, Milos Radovanovic 0001, Mirjana Ivanovic |
RecSys | 2 |
| 2007 | Automatic Categorization of Human-Coded and Evolved CoreWar Warriors
Nenad Tomasev, Doni Pracner, Milos Radovanovic 0001, Mirjana Ivanovic |
PKDD | 3 |
| 2006 | Document Representations for Classification of Short Web-Page Descriptions
Milos Radovanovic 0001, Mirjana Ivanovic |
DaWaK | 1 |
| 2006 | Interactions Between Document Representation and Feature Selection in Text Categorization
Milos Radovanovic 0001, Mirjana Ivanovic |
DEXA | 1 |