EDBT 2026 Demo / reviewers in the wild / expert
Tomás Skopal
dblp:s/TomasSkopal
· DBLP profile ↗
64ranked-venue papers in the field
22as first author
8since 2021 · last 2025
0000-0002-6591-0879ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 46 (15 first)Information Retrieval & Web Search · 15 (6 first)Data Mining & Knowledge Discovery · 3 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VISAnt: Unsupervised Data Exploration with Chernoff Faces
Ivaná Sixtova, Ladislav Peska, Jakub Lokoc, David Bernhauer, Tomás Skopal |
SISAP | 5 |
| 2025 | Class Representatives Selection in non-metric spaces for nearest prototype classification
Jaroslav Hlavác, Martin Kopp, Tomás Skopal |
Inf. Syst. | 3 |
| 2024 | Visualizations for universal deep-feature representations: survey and taxonomyabstractAbstract In data science and content-based retrieval, we find many domain-specific techniques that employ a data processing pipeline with two fundamental steps. First, data entities are represented by some visualizations, while in the second step, the visualizations are used with a machine learning model to extract deep features. Deep convolutional neural networks (DCNN) became the standard and reliable choice. The purpose of using DCNN is either a specific classification task or just a deep feature representation of visual data for additional processing (e.g., similarity search). Whereas the deep feature extraction is a domain-agnostic step in the pipeline (inference of an arbitrary visual input), the visualization design itself is domain-dependent and ad hoc for every use case. In this paper, we survey and analyze many instances of data visualizations used with deep learning models (mostly DCNN) for domain-specific tasks. Based on the analysis, we synthesize a taxonomy that provides a systematic overview of visualization techniques suitable for usage with the models. The aim of the taxonomy is to enable the future generalization of the visualization design process to become completely domain-agnostic, leading to the automation of the entire feature extraction pipeline. As the ultimate goal, such an automated pipeline could lead to universal deep feature data representations for content-based retrieval. Tomás Skopal, Ladislav Peska, David Hoksza, Ivaná Sixtova, David Bernhauer |
Knowl. Inf. Syst. | 1 |
| 2023 | Class Representatives Selection in Non-metric Spaces for Nearest Prototype Classification
Jaroslav Hlavác, Martin Kopp, Jan Kohout, Tomás Skopal |
SISAP | 4 |
| 2022 | Open dataset discovery using context-enhanced similarity search
David Bernhauer, Martin Necaský, Petr Skoda 0001, Jakub Klímek, Tomás Skopal |
Knowl. Inf. Syst. | 5 |
| 2021 | Videolytics: System for Data Analytics of Video StreamsabstractWe present Videolytics, a web-based system for advanced analytics over recorded video streams. Video cameras have become widely used for indoor and outdoor surveillance. Covering even more public space in cities, the cameras serve various purposes ranging from security to traffic monitoring, urban life, and marketing. The goal is to obtain effective and efficient models to process the video data automatically and produce the desired features for data analytics. Videolytics combines the best of deep learning and hand-designed analytical models to create a solution applicable in real-life situations. The architecture of the Videolytics framework is centered around a database of video features and detected objects, where new higher-level objects result from fusion of (lower-level) objects and features already stored in the database. The system provides a number of visualization options, an SQL-based analytics module as well as a real-time surveillance mode. Tomás Skopal, Dominika Durisková, Petr Pechman, Marek Dobranský, Vladislav Khachaturian |
CIKM | 1 |
| 2021 | Similarity vs. Relevance: From Simple Searches to Complex Discovery
Tomás Skopal, David Bernhauer, Petr Skoda 0001, Jakub Klímek, Martin Necaský |
SISAP | 1 |
| 2021 | On augmenting database schemas by latent visual attributesabstractAbstract Decision-making in our everyday lives is surrounded by visually important information. Fashion, housing, dating, food or travel are just a few examples. At the same time, most commonly used tools for information retrieval operate on relational and text-based search models which are well understood by end users, but unable to directly cover visual information contained in images or videos. Researcher communities have been trying to reveal the semantics of multimedia in the last decades with ever-improving results, dominated by the success of deep learning. However, this does not close the gap to relational retrieval model on its own and often rather solves a very specialized task like assigning one of pre-defined classes to each object within a closed application ecosystem. Retrieval models based on these novel techniques are difficult to integrate in existing application-agnostic environments built around relational databases, and therefore, they are not so widely used in the industry. In this paper, we address the problem of closing the gap between visual information retrieval and relational database model. We propose and formalize a model for discovering candidates for new relational attributes by analysis of available visual content. We design and implement a system architecture supporting the attribute extraction, suggestion and acceptance processes. We apply the solution in the context of e-commerce and show how it can be seamlessly integrated with SQL environments widely used in the industry. At last, we evaluate the system in a user study and discuss the obtained results. Tomás Grosup, Ladislav Peska, Tomás Skopal |
Knowl. Inf. Syst. | 3 |
| 2020 | Evaluation Framework for Search Methods Focused on Dataset Findability in Open Data CatalogsabstractMany institutions publish datasets as Open Data in catalogs, however, their retrieval remains problematic issue due to the absence of dataset search benchmarking. We propose a framework for evaluating findability of datasets, regardless of retrieval models used. As task-agnostic labeling of datasets by ground truth turns out to be infeasible in the general domain of open data datasets, the proposed framework is based on evaluation of entire retrieval scenarios that mimic complex retrieval tasks. In addition to the framework we present a proof of concept specification and evaluation on several similarity-based retrieval models and several dataset discovery scenarios within a catalog, using our experimental evaluation tool. Instead of traditional matching of query with metadata of all the datasets, in similarity-based retrieval the query is formulated using a set of datasets (query by example) and the most similar datasets to the query set are retrieved from the catalog as a result. Petr Skoda 0001, David Bernhauer, Martin Necaský, Jakub Klímek, Tomás Skopal |
iiWAS | 5 |
| 2020 | On Visualizations in the Role of Universal Data RepresentationabstractThe deep learning revolution changed the world of machine learning and boosted the AI industry as such. In particular, the most effective models for image retrieval are based on deep convolutional neural networks (DCNN), outperforming the traditional "hand-engineered" models by far. However, this tremendous success was redeemed by a high cost in the form of an exhaustive gathering of labeled data, followed by designing and training the DCNN models. In this paper, we outline a vision of a framework for instant transfer learning, where a generic pre-trained DCNN model is used as a universal feature extraction method for visualized unstructured data in many (non-visual) domains. The deep feature descriptors are then usable in similarity search tasks (database queries, joins) and in other parts of the data processing pipeline. The envisioned framework should enable practitioners to instantly use DCNN-based data representations in their new domains without the need for the costly training step. Moreover, by use of the framework the information visualization community could acquire a versatile metric for measuring the quality of data visualizations, which is generally a difficult task. Tomás Skopal |
ICMR | 1 |
| 2020 | Analysing Indexability of Intrinsically High-Dimensional Data Using TriGen
David Bernhauer, Tomás Skopal |
SISAP | 2 |
| 2020 | Visualizer of Dataset Similarity Using Knowledge Graph
Petr Skoda 0001, Jakub Matejík, Tomás Skopal |
SISAP | 3 |
| 2019 | SIMILANT: An Analytic Tool for Similarity Modeling
David Bernhauer, Tomás Skopal, Irena Holubová, Ladislav Peska, Martin Svoboda |
CIKM | 2 |
| 2019 | Towards Augmented Database Schemes by Discovery of Latent Visual Attributes
Tomás Grosup, Ladislav Peska, Tomás Skopal |
EDBT | 3 |
| 2019 | Improving Findability of Open Data Beyond Data CatalogsabstractThere is a vast amount of datasets available as Open Data on the Web. However, it is challenging for consumers to find datasets relevant to their goals. This is because the available metadata in catalogs is not descriptive enough. Nevertheless, datasets exist in various types of contexts not expressed in the metadata. These may include information about the data publisher, the legislation related to dataset publication, etc. In this paper we describe an idea of a data model that enables consumers to better understand the data. We propose to define a formal model for representation of the datasets and their contexts, and we propose to apply existing similarity techniques, adjust them to fit each identified dataset context type and combine them together to measure similarity of datasets in new ways, improving their findability. Tomás Skopal, Jakub Klímek, Martin Necaský |
iiWAS | 1 |
| 2019 | Explainable Similarity of Datasets Using Knowledge Graph
Petr Skoda 0001, Jakub Klímek, Martin Necaský, Tomás Skopal |
SISAP | 4 |
| 2019 | Non-metric Similarity Search Using Genetic TriGen
David Bernhauer, Tomás Skopal |
SISAP | 2 |
| 2018 | Interactive Product Search Based on Global and Local Visual-Semantic Features
Tomás Skopal, Ladislav Peska, Tomás Grosup |
SISAP | 1 |
| 2018 | Advanced Analytics of Large Connected Data Based on Similarity Modeling
Tomás Skopal, Ladislav Peska, Irena Holubová, Petr Pascenko, Jan Hucín |
SISAP | 1 |
| 2017 | Analyzing Mathematical Content to Detect Academic PlagiarismabstractThis paper presents, to our knowledge, the first study on analyzing mathematical expressions to detect academic plagiarism. We make the following contributions. First, we investigate confirmed cases of plagiarism to categorize the similarities of mathematical content commonly found in plagiarized publications. From this investigation, we derive possible feature selection and feature comparison strategies for developing math-based detection approaches and a ground truth for our experiments. Second, we create a test collection by embedding confirmed cases of plagiarism into the NTCIR-11 MathIR Task dataset, which contains approx. 60 million mathematical expressions in 105,120 documents from arXiv.org. Third, we develop a first math-based detection approach by implementing and evaluating different feature comparison approaches using an open source parallel data processing pipeline built using the Apache Flink framework. The best performing approach identifies all but two of our real-world test cases at the top rank and achieves a mean reciprocal rank of 0.86. The results show that mathematical expressions are promising text-independent features to identify academic plagiarism in large collections. To facilitate future research on math-based plagiarism detection, we make our source code and data available. Norman Meuschke, Moritz Schubotz, Felix Hamborg, Tomás Skopal, Bela Gipp |
CIKM | 4 |
| 2017 | Product Exploration based on Latent Visual AttributesabstractIn this demo paper, we present a prototype web application of a product search engine of a fashion e-shop. Although e-shop products consist of full-text description, relational attributes (e.g., price, type, size, color, etc.) as well as visual information (product photo), traditional search engines in e-shops only provide full-text and relational attributes for product filtering. In our retrieval model, we incorporate also the visual information into the search by extracting visual-semantic features using deep convolutional neural networks. Furthermore, visual exploration of the product space using the visual-semantic features (multi-example queries) is used to dynamically discover latent visual attributes that could enhance the original relational schema by fuzzy attributes (e.g., a floral pattern in product). In the demo, we show how these latent attributes could be used to recommend the user preferred products and even outfits (e.g., shoes, bag, jacket) that fit a certain visual style. Tomás Skopal, Ladislav Peska, Gregor Kovalcík, Tomás Grosup, Jakub Lokoc |
CIKM | 1 |
| 2017 | Malware Discovery Using Behaviour-Based Exploration of Network Traffic
Jakub Lokoc, Tomás Grosup, Premysl Cech, Tomás Pevný, Tomás Skopal |
SISAP | 5 |
| 2015 | Evaluating Multilayer Multimedia Exploration
Juraj Mosko, Jakub Lokoc, Tomás Grosup, Premysl Cech, Tomás Skopal, Jan Lansky |
SISAP | 5 |
| 2014 | On Effective Known Item Video Search Using Feature SignaturesabstractIn this demo paper, we present a video retrieval and browsing tool inspired by the natural human ability to memorize visual stimuli of color regions in video frames. Our tool utilizes feature signatures that can be used to represent both significant color regions in the key-frames and simple query sketches. As recently shown at the video browser showdown, such simple representation enables both effective end efficient interactive retrieval and browsing in video. Jakub Lokoc, Adam Blazek, Tomás Skopal |
ICMR | 3 |
| 2014 | Video Retrieval with Feature Signature Sketches
Adam Blazek, Jakub Lokoc, Tomás Skopal |
SISAP | 3 |
| 2014 | Analyzing and dynamically indexing the query set
Juan Manuel Barrios, Benjamin Bustos, Tomás Skopal |
Inf. Syst. | 3 |
| 2014 | On indexing metric spaces using cut-regions
Jakub Lokoc, Juraj Mosko, Premysl Cech, Tomás Skopal |
Inf. Syst. | 4 |
| 2013 | Dynamic multimedia exploration using SIFT matchingabstractIn this demo paper, we focus on the dynamic multimedia exploration techniques which are an intuitive, effective and entertaining way to present a pre-selected subset of a multimedia database to the users. More specifically, we present an exploration schema employing a similarity model based on SIFT descriptors that can be used to explore image database according to regions in the images. We also provide a simple mechanism to reduce the number of nonrelevant SIFT descriptors in the query image. The reduction of SIFTs in the query image improves the speed and fluency of the exploration process as demonstrated in our demo application. Jakub Lokoc, Lukás Navrátil, Jachym Tousek, Tomás Skopal |
ICMR | 4 |
| 2013 | Designing Similarity Indexes with Parallel Genetic Programming
Tomás Bartos, Tomás Skopal |
SISAP | 2 |
| 2013 | On Scalable Approximate Search with the Signature Quadratic Form Distance
Jakub Lokoc, Tomás Grosup, Tomás Skopal |
SISAP | 3 |
| 2013 | Ptolemaic access methods: Challenging the reign of the metric space model
Magnus Lie Hetland, Tomás Skopal, Jakub Lokoc, Christian Beecks |
Inf. Syst. | 2 |
| 2012 | Similarity search in 3D object-based video dataabstractIn this paper, we present the vision of the usage of an object-based video data storage format for similarity search. The efficient (fast) and effective (accurate) search in video streams is an ongoing and still unsolved problem. Using an object-based format of multimedia data, all the information that is needed to answer queries is already available in a machine accessible format. This way, the process of creating (video) descriptors as well as the similarity search becomes easier, because the data is already organized in a manner that allows fast access to specific information. To demonstrate the concept of similarity search process using the object-based 3D video format, we present experiments conducted on generated clouds of points (an abstraction of 3D video data). Jakub Lokoc, Jürgen Wünschmann, Tomás Skopal, Albrecht Rothermel |
CIKM | 3 |
| 2012 | Image exploration using online feature extraction and rerankingabstractWe present an image meta-search engine that allows content-based exploration of the results obtained from various sources (mostly based on keyword query). The online feature extraction and the particle physics model are the two key features of our demo application that shows very promising results. Jakub Lokoc, Tomás Grosup, Tomás Skopal |
ICMR | 3 |
| 2012 | Snake Table: A Dynamic Pivot Table for Streams of k-NN Searches
Juan Manuel Barrios, Benjamin Bustos, Tomás Skopal |
SISAP | 3 |
| 2012 | Revisiting Techniques for Lowerbounding the Dynamic Time Warping Distance
Tomás Bartos, Tomás Skopal |
SISAP | 2 |
| 2012 | Cut-Region: A Compact Building Block for Hierarchical Metric Indexing
Jakub Lokoc, Premysl Cech, Tomás Skopal |
SISAP | 4 |
| 2012 | SIR: The Smart Image Retrieval Engine
Jakub Lokoc, Tomás Grosup, Tomás Skopal |
SISAP | 3 |
| 2012 | Visual Image Search: Feature Signatures or/and Global Descriptors
Jakub Lokoc, David Novak, Michal Batko, Tomás Skopal |
SISAP | 4 |
| 2012 | SimTandem: Similarity Search in Tandem Mass Spectra
Jakub Galgonek, David Hoksza, Tomás Skopal |
SISAP | 4 |
| 2012 | Algorithmic Exploration of Axiom Spaces for Efficient Similarity Search at Large Scale
Tomás Skopal, Tomás Bartos |
SISAP | 1 |
| 2012 | Combining CPU and GPU architectures for fast similarity search
Martin Krulis, Tomás Skopal, Jakub Lokoc, Christian Beecks |
Distributed Parallel Databases | 2 |
| 2012 | D-Cache: Universal Distance Cache for Metric Access MethodsabstractThe caching of accessed disk pages has been successfully used for decades in database technology, resulting in effective amortization of I/O operations needed within a stream of query or update requests. However, in modern complex databases, like multimedia databases, the I/O cost becomes a minor performance factor. In particular, metric access methods (MAMs), used for similarity search in complex unstructured data, have been designed to minimize rather the number of distance computations than I/O cost (when indexing or querying). Inspired by I/O caching in traditional databases, in this paper we introduce the idea of distance caching for usage with MAMs—a novel approach to streamline similarity search. As a result, we present the D-cache, a main-memory data structure which can be easily implemented into any MAM, in order to spare the distance computations spent by queries/updates. In particular, we have modified two state-of-the-art MAMs to make use of D-cache—the M-tree and Pivot tables. Moreover, we present the D-file, an index-free MAM based on simple sequential search augmented by D-cache. The experimental evaluation shows that performance gain achieved due to D-cache is significant for all the MAMs, especially for the D-file. Tomás Skopal, Jakub Lokoc, Benjamin Bustos |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2011 | Processing the signature quadratic form distance on many-core GPU architecturesabstractThe Signature Quadratic Form Distance on feature signatures represents a flexible distance-based similarity model for effective content-based multimedia retrieval. Although metric indexing approaches are able to speed up query processing by two orders of magnitude, their applicability to large-scale multimedia databases containing billions of images is still a challenging issue. In this paper, we propose the utilization of GPUs for efficient query processing with the Signature Quadratic Form Distance. We show how to process multiple distance computations in parallel and demonstrate efficient query processing by comparing many-core GPU with multi-core CPU implementations. Martin Krulis, Jakub Lokoc, Christian Beecks, Tomás Skopal, Thomas Seidl 0001 |
CIKM | 4 |
| 2011 | On (not) indexing quadratic form distance by metric access methodsabstractThe quadratic form distance (QFD) has been utilized as an effective similarity function in multimedia retrieval, in particular, when a histogram representation of objects is used. Unlike the widely used Euclidean distance, the QFD allows to arbitrarily correlate the histogram bins (dimensions), allowing thus to better model the similarity between histograms. However, unlike Euclidean distance, which is of linear time complexity, the QFD requires quadratic time to evaluate the similarity of two objects. In consequence, indexing and querying a database under QFD are expensive operations. In this paper we show that, given static correlations between dimensions, the QFD space can be transformed into an equivalent Euclidean space. Thus, the overall complexity of indexing and searching in the QFD similarity model can be reduced qualitatively. Besides the theoretical time complexity analysis of our approach applied to several metric access methods, in experimental evaluation we show the real-time speedup on a real-world image database. Tomás Skopal, Tomás Bartos, Jakub Lokoc |
EDBT | 1 |
| 2011 | Non-metric similarity search problems in very large collectionsabstractThis tutorial surveys domains employing non-metric functions for effective similarity search, and methods for efficient non-metric similarity search in very large collections. Benjamin Bustos, Tomás Skopal |
ICDE | 2 |
| 2011 | Indexing the signature quadratic form distance for efficient content-based multimedia retrievalabstractThe Signature Quadratic Form Distance has been introduced as an adaptive similarity measure coping with flexible content representations of various multimedia data. Although the Signature Quadratic Form Distance has shown good retrieval performance with respect to their qualities of effectiveness and efficiency, its applicability to index structures remains a challenging issue due to its dynamic nature. In this paper, we investigate the indexability of the Signature Quadratic Form Distance regarding metric access methods. We show how the distance's inherent parameters determine the indexability and analyze the relationship between effectiveness and efficiency on numerous image databases. Christian Beecks, Jakub Lokoc, Thomas Seidl 0001, Tomás Skopal |
ICMR | 4 |
| 2011 | Fuzzy approach to non-metric similarity indexingabstractThe task of similarity search becomes more complex when the distance measure is not a metric. In this paper, we investigated the recently proposed fuzzy approach to similarity search in non-metric databases where the triangle inequality might not hold. In summary, we took nine fuzzy T-norms, proposed a tuning algorithm for the fuzzy T-norm operators (Lambda Tuning Algorithm), and applied this approach to the pivot-based search. We present the results focusing on the efficiency and effectiveness of the suggested method. Tomás Bartos, Alan Eckhardt, Tomás Skopal |
SISAP | 3 |
| 2011 | Parameterized earth mover's distance for efficient metric space indexingabstractThe Earth Mover's Distance is a well-known distance measure employed in various domains, especially for content-based retrieval in multimedia databases. However, the distance evaluation is a considerably expensive task and thus for large multimedia databases, efficient query processing becomes a challenging problem. In this paper, we introduce a parameterized version of the Earth Mover's Distance that can be used by database experts to change the distance distribution in the derived distance space in order to improve the indexability. We empirically show, that we can significantly improve the indexability of the distance space and that we can tune the retrieval quality by adapting the parameterized Earth Mover's Distance. Jakub Lokoc, Christian Beecks, Thomas Seidl 0001, Tomás Skopal |
SISAP | 4 |
| 2011 | Ptolemaic indexing of the signature quadratic form distanceabstractThe signature quadratic form distance has been introduced as an adaptive similarity measure coping with flexible content representations of multimedia data. While this distance has shown high retrieval quality, its high computational complexity underscores the need for efficient search methods. Recent research has shown that a huge improvement in search efficiency is achieved when using metric indexing. In this paper, we analyze the applicability of Ptolemaic indexing to the signature quadratic form distance. We show that it is a Ptolemaic metric and present an application of Ptolemaic pivot tables to image databases, resolving queries nearly four times as fast as the state-of-the-art metric solution, and up to 300 times as fast as sequential scan. Jakub Lokoc, Magnus Lie Hetland, Tomás Skopal, Christian Beecks |
SISAP | 3 |
| 2011 | Clustered pivot tables for I/O-optimized similarity searchabstractThe pivot tables are a popular metric access method, primarily designed as a main-memory index structure. It has been many times proven that pivot tables are very efficient in terms of distance computations, hence, when assuming a computationally expensive distance function. However, for cheaper distance functions and/or huge datasets exceeding the capacity of the main memory, the classic pivot tables become inefficient. The situation is dramatically changing with the rise of solid state disks that decrease the seek times, so we can now efficiently access also small fragments of data stored in the secondary memory. In this paper, we propose a persistent variant of pivot tables, the clustered pivot tables, focusing on minimizing I/O cost when accessing small data blocks (a few kilobytes). The clustered pivot tables employs a preprocessing method utilizing the M-tree in the role of clustering technique and an original heuristic for I/O-optimized kNN query processing. In the experiments we empirically show that our proposed method significantly reduces the number of necessary I/O operations during query processing. Juraj Mosko, Jakub Lokoc, Tomás Skopal |
SISAP | 3 |
| 2011 | Protein sequences identification using NM-treeabstractWe have generalized a method for tandem mass spectra interpretation, based on the parameterized Hausdorff distance dHP. Instead of just peptides (short pieces of proteins), in this paper we describe the interpretation of whole protein sequences. For this purpose, we employ the recently introduced NM-tree to index the database of hypothetical mass spectra for exact or fast approximate search. The NM-tree combines the M-tree with the TriGen algorithm in a way that allows to dynamically control the retrieval precision at query time. A scheme for protein sequences identification using the NM-tree is proposed. Tomás Skopal, David Hoksza, Jakub Lokoc, Jakub Galgonek |
SISAP | 2 |
| 2011 | Preface
Vlastislav Dohnal, Tomás Skopal |
Inf. Syst. | 2 |
| 2009 | On Index-Free Similarity Search in Metric Spaces
Tomás Skopal, Benjamin Bustos |
DEXA | 1 |
| 2009 | On Fuzzy vs. Metric Similarity Search in Complex Databases
Alan Eckhardt, Tomás Skopal, Peter Vojtás |
FQAS | 2 |
| 2008 | NM-Tree: Flexible Approximate Similarity Search in Metric and Non-metric Spaces
Tomás Skopal, Jakub Lokoc |
DEXA | 1 |
| 2007 | Improving the Performance of M-Tree Family by Nearest-Neighbor Graphs
Tomás Skopal, David Hoksza |
ADBIS | 1 |
| 2007 | Construction of Tree-Based Indexes for Level-Contiguous Buffering Support
Tomás Skopal, David Hoksza, Jaroslav Pokorný |
DASFAA | 1 |
| 2007 | Unified framework for fast exact and approximate search in dissimilarity spacesabstractIn multimedia systems we usually need to retrieve database (DB) objects based on their similarity to a query object, while the similarity assessment is provided by a measure which defines a (dis)similarity score for every pair of DB objects. In most existing applications, the similarity measure is required to be a metric, where the triangle inequality is utilized to speed up the search for relevant objects by use of metric access methods (MAMs), for example, the M-tree. A recent research has shown, however, that nonmetric measures are more appropriate for similarity modeling due to their robustness and ease to model a made-to-measure similarity. Unfortunately, due to the lack of triangle inequality, the nonmetric measures cannot be directly utilized by MAMs. From another point of view, some sophisticated similarity measures could be available in a black-box nonanalytic form (e.g., as an algorithm or even a hardware device), where no information about their topological properties is provided, so we have to consider them as nonmetric measures as well. From yet another point of view, the concept of similarity measuring itself is inherently imprecise and we often prefer fast but approximate retrieval over an exact but slower one. To date, the mentioned aspects of similarity retrieval have been solved separately, that is, exact versus approximate search or metric versus nonmetric search. In this article we introduce a similarity retrieval framework which incorporates both of the aspects into a single unified model. Based on the framework, we show that for any dissimilarity measure (either a metric or nonmetric) we are able to change the “amount” of triangle inequality, and so obtain an approximate or full metric which can be used for MAM-based retrieval. Due to the varying “amount” of triangle inequality, the measure is modified in a way suitable for either an exact but slower or an approximate but faster retrieval. Additionally, we introduce the TriGen algorithm aimed at constructing the desired modification of any black-box distance automatically, using just a small fraction of the database. Tomás Skopal |
ACM Trans. Database Syst. | 1 |
| 2006 | On Fast Non-metric Similarity Search by Metric Access Methods
Tomás Skopal |
EDBT | 1 |
| 2006 | A new range query algorithm for Universal B-trees
Tomás Skopal, Michal Krátký, Jaroslav Pokorný, Václav Snásel |
Inf. Syst. | 1 |
| 2005 | Nearest Neighbours Search Using the PM-Tree
Tomás Skopal, Jaroslav Pokorný, Václav Snásel |
DASFAA | 1 |
| 2005 | Modified LSI Model for Efficient Search by Metric Access Methods
Tomás Skopal, Pavel Moravec 0001 |
ECIR | 1 |
| 2004 | Metric Indexing for the Vector Model in Text Retrieval
Tomás Skopal, Pavel Moravec 0001, Jaroslav Pokorný, Václav Snásel |
SPIRE | 1 |
| 2003 | Revisiting M-Tree Building Principles
Tomás Skopal, Jaroslav Pokorný, Michal Krátký, Václav Snásel |
ADBIS | 1 |