EDBT 2026 Demo / reviewers in the wild / expert
Daniel S. Kaster
dblp:k/DanielSKaster · also Daniel dos Santos Kaster
· DBLP profile ↗
24ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0001-9628-3313ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 18 · 3 since 2021Artificial intelligence and machine learning · 8Human-computer interaction and ubiquitous computing · 5Applied, interdisciplinary, general and emerging computing · 4Graphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | The Dataset-Similarity-Based Approach to Select Datasets for Evaluation in Similarity Retrieval
Matheus A. L. Matiazzo, Vitor de Castro-Silva, Rafael Seidi Oyamada, Daniel S. Kaster |
SISAP | 4 |
| 2023 | A meta-learning configuration framework for graph-based similarity search indexes
Rafael Seidi Oyamada, Larissa Capobianco Shimomura, Sylvio Barbon Junior, Daniel S. Kaster |
Inf. Syst. | 4 |
| 2021 | A survey on graph-based methods for similarity searches in metric spaces
Larissa Capobianco Shimomura, Rafael Seidi Oyamada, Marcos R. Vieira, Daniel S. Kaster |
Inf. Syst. | 4 |
| 2020 | Towards Proximity Graph Auto-configuration: An Approach Based on Meta-learning
Rafael Seidi Oyamada, Larissa Capobianco Shimomura, Sylvio Barbon Junior, Daniel S. Kaster |
ADBIS | 4 |
| 2019 | HGraph: A Connected-Partition Approach to Proximity Graphs for Similarity Search
Larissa Capobianco Shimomura, Daniel S. Kaster |
DEXA (1) | 2 |
| 2018 | On the Support of the Similarity-Aware Division Operator in a Commercial RDBMS
Guilherme Q. Vasconcelos, Daniel S. Kaster, Robson L. F. Cordeiro |
ADBIS | 2 |
| 2018 | Standard SQL Approaches for Similarity SearchingabstractThis paper addresses complex data storage and retrieval in RDBMS, which depends on metric distance functions for the assessment of data dissimilarity. However, both the empirical analysis of strategies for complex data storage and the definition of a suitable representation for similarity query operators are still open issues in the literature. Here, we fulfill those gaps through the classification, implementation, and evaluation of existing approaches for complex data storage according to four structures found in standard SQL, namely relational, object-relational, binary and semi-structured. Moreover, we also discuss a comprehensive model for complex data retrieval, whose conception of similarity operators is consistent with standard SQL representations. Accordingly, a distance function representation is presented, which enables the RDBMS query processor to interpret and execute physical similarity operators. Experimental results indicate: (i) relational and object-relational structures outperform the other two competitors in the majority of scenarios, whereas (ii) object-relational strategy enables the use of a broader representation. Pedro Henrique Braga Siqueira, Paulo H. Oliveira, Marcos V. N. Bedo, Daniel S. Kaster |
CLEI | 4 |
| 2018 | What Lies Beyond Structured Data? A Comparison Study for Metric Data Storage
Pedro Henrique Braga Siqueira, Paulo H. Oliveira, Marcos V. N. Bedo, Daniel S. Kaster |
DEXA (2) | 4 |
| 2018 | Performance Analysis of Graph-Based Methods for Exact and Approximate Similarity Search in Metric Spaces
Larissa Capobianco Shimomura, Marcos R. Vieira, Daniel S. Kaster |
SISAP | 3 |
| 2018 | The Merkurion approach for similarity searching optimization in Database Management SystemsabstractModern Database Management Systems (DBMSs) retrieve songs that resemble those in a music dataset, identify plagiarism in a set of documents, or provide past cases to physicians by taking into account the characteristics of a query exam. All such tasks require the comparison of data by similarity, which can be expressed in terms of distance-based queries in metric spaces. Traditional query processing relies mostly on histograms for describing the data distribution space and choosing a data retrieval path that quickly leads to the answer, discarding comparisons of most unwanted data. However, DBMSs still lack adequate support for selectivity estimation of query operators for data types embedded in metric spaces. This article addresses a novel strategy that extends the query optimizer of a DBMS, so that it can also perform both logical and physical query plan optimizations in searches that include similarity predicates. The proposal, named Merkurion, updates the concept of Data Distribution Space and captures data distributions according to the distances between the elements within a dataset. Moreover, it employs concise representations of such distributions, called synopses, for the definition of rules that enable similarity searching optimization. An extensive evaluation of Merkurion in real-world datasets has proven its effectiveness and broad applicability to many data domains. Marcos V. N. Bedo, Daniel S. Kaster, Agma J. M. Traina, Caetano Traina Jr. |
Data Knowl. Eng. | 2 |
| 2017 | BREATH: Heat Maps Assisting the Detection of Abnormal Lung Regions in CT ScansabstractComputed Tomography (CT) scans are often employed to diagnose lung diseases, as abnormal tissue regions may indicate whether proper treatment is required. However, detecting specific regions containing abnormalities in a CT scan demands time and effort of specialists. Moreover, different parts of a single lung image may present both normal and abnormal characteristics, what makes inaccurate the classification of a single lung as healthy (normal) or not. In this paper we propose the BREATH method, capable of detecting abnormalities in lung tissue regions, highlighting them by means of a heat map visualization. The method starts by segmenting lung tissues using a superpixel-based approach, followed by the training of a statistical model to represent normal tissues and, finally, the generation of a heat map showing abnormal regions that require attention from the physicians. We validated our statistical model using a dataset with 246 lung CT scans, where 40 are healthy and the remaining present varying diseases. Experimental results show that BREATH is accurate for lung segmentation with F-Measure of up to 0.99. The statistical modeling of healthy and abnormal lung regions has shown almost no overlap, and the detection of superpixels containing abnormalities presented precision values higher than 86%, for all values of recall. These values support our claim that the heat map representation of BREATH for the abnormal detection can be used as an intuitive method to assist physicians during the diagnosis. Mirela Teixeira Cazzolato, Lucas C. Scabora, Alceu Ferraz Costa, Marcos Roberto Nesso Junior, Luis Fernando Milano Oliveira, Daniel S. Kaster, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 6 |
| 2017 | Defining Similarity Spaces for Large-Scale Image Retrieval Through Scientific WorkflowsabstractContent-Based Image Retrieval (CBIR) employs visual features from images for searching and retrieving of data. Systems based on this concept depend on a similarity space instance definition, but achieving an ideal instance is a very complex process and is dependent on domain knowledge. At the same time, domain experts are often unable to interact fully with systems because of technicalities. In this paper, we propose an architecture, based on scientific workflows, which allows users with no prior programming experience to build processes on images, creating Similarity Spaces and evaluating them when running similarity queries. Through this architecture, they can use domain expertise to improve image retrieval in a coordinated, auditable and reproducible manner, while being able to process very large image collections. We describe a prototype system and carry out experiments evaluating its performance in various scenarios. The current implementation supports both similarity space definition and querying workflows, achieving suitable speedups with the increase in the number of machines. Luis Fernando Milano Oliveira, Daniel S. Kaster |
IDEAS | 2 |
| 2017 | CLAP, ACIR and SCOOP: Novel techniques for improving the performance of dynamic Metric Access MethodsabstractConstant technological advances in electronic devices have led to the growth of elaborated data such as large texts, time series, georeferenced imagery, genetic sequences, photos, videos and several other types of complex data. Differently from scalar, traditional data types such as numbers and strings, complex data do not present the order relation property, which allows identifying whether an element precedes another according to some criterion. Therefore, these data are usually compared by the similarity degree among them. The Metric Access Methods (MAMs) are recognized as well-suited to perform similarity queries over such kind of data more efficiently than other access methods. MAMs can be considered dynamic or static depending on the pivot type used to construct them. Pivots are often employed to narrow the search for data. Global pivots can be employed to look into elements in the whole dataset, thus they have a high impact in the process of pruning irrelevant elements, since a single global pivot can be used to discard a large amount of irrelevant elements. Nevertheless, MAMs based on global pivots may have their dynamicity compromised by the fact that eventual pivot-related updates must be propagated through the entire structure. Local pivots, on the other hand, allow the maintenance to occur locally at the price of a lower pruning ability. In this paper, we propose novel techniques for improving the performance of dynamic MAMs without harming their dynamicity, once that several applications handle online complex data and, consequently, demand efficient dynamic indexes to be successful. Specifically, our main contributions are three techniques: (i) CLAP, which consists of employing local additional pivots to reduce distance calculations; (ii) ACIR, which is combined with CLAP and anticipates information from child nodes to reduce unnecessary disk accesses; and (iii) SCOOP, which is combined with CLAP as an extended version of ACIR, anticipating a larger amount of information from child nodes. The techniques have been applied to a dynamic MAM and evaluated over real datasets ranging from moderate to high dimensionality and cardinality. The experimental results show that our techniques were able to reduce query execution time in up to 63% for point queries and up to 53% for queries retrieving multiple elements. Paulo H. Oliveira, Caetano Traina Jr., Daniel S. Kaster |
Inf. Syst. | 3 |
| 2015 | Improving the Pruning Ability of Dynamic Metric Access Methods with Local Additional Pivots and Anticipation of Information
Paulo H. Oliveira, Caetano Traina Jr., Daniel S. Kaster |
ADBIS | 3 |
| 2015 | Improving Metric Access Methods with Bucket Files
Ives Rene Venturini Pola, Agma J. M. Traina, Caetano Traina Jr., Daniel S. Kaster |
SISAP | 4 |
| 2015 | Compact distance histogram: a novel structure to boost k-nearest neighbor queriesabstractThe k-Nearest Neighbor query (k-NNq) is one of the most useful similarity queries. Elaborated k-NNq algorithms depend on an initial radius to prune regions of the search space that cannot contribute to the answer. Therefore, estimating a suitable starting radius is of major importance to accelerate k-NNq execution. This paper presents a new technique to estimate a tight initial radius. Our approach, named CDH-kNN, relies on Compact Distance Histograms (CDHs), which are pivot-based histograms defined as piecewise linear functions. Such structures approximate the distance distribution and are compressed according to a given constraint, which can be a desired number of buckets and/or a maximum allowed error. The covering radius of a k-NNq is estimated based on the relationship between the query element and the CDHs' joint frequencies. The paper presents a complete specification of CDH-kNN, including CDH's construction and radii estimation. Extensive experiments on both real and synthetic datasets highlighted the efficiency of our approach, showing that it was up to 72% faster than existing algorithms, outperforming every competitor in all the setups evaluated. In fact, the experiments showed that our proposal was just 20% slower than the theoretical lower bound. Marcos V. N. Bedo, Daniel S. Kaster, Agma J. M. Traina, Caetano Traina Jr. |
SSDBM | 2 |
| 2013 | Does a CBIR system really impact decisions of physicians in a clinical environment?abstractContent-based image retrieval systems are employed in several areas. One of the most prominent area is the medical field, due to the huge volume of digital images daily generated in healthcare institutions employed for decision making. There are several works applying CBIR techniques over medical images. However, the great majority of them do not verify whether the systems are actually considered by the specialists as a pontential aid in a real environment. In order to fill this research void in the literature, this work explores user experiments in a CBIR system involving resident physicians and radiologists. To do so, we developed a CBIR system according to requirements provided by the specialists and employed a methodology to analyze the effectiveness of the system for supporting them in clinical routine. The methodology aims at evaluating the system's impact in the user's decision, inquiring the specialists about the image classification and their degree of certainty in different situations using the system. By analyzing the obtained results we can argue that the proposed methodology joined with our medical CBIR system presented a high acceptance and viability rate regarding the radiologists interests in the clinical practice domain, providing a novel approach to analyze CBIR systems under realistic conditions. Marcelo Ponciano-Silva, Juliana P. Souza, Pedro Henrique Bugatti, Marcos V. N. Bedo, Daniel S. Kaster, Rosana T. V. Braga, Angela D. Bellucci, Paulo Mazzoncini de Azevedo Marques, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 5 |
| 2013 | A Similarity-Based Approach for Financial Time Series Analysis and Forecasting
Marcos V. N. Bedo, Davi Pereira dos Santos, Daniel S. Kaster, Caetano Traina Jr. |
DEXA (2) | 3 |
| 2013 | Efficient Execution of Conjunctive Complex Queries on Big Multimedia DatabasesabstractThis paper proposes an approach to efficientlyexecute conjunctive queries on big complex data together withtheir related conventional data. The basic idea is to horizontallyfragment the database according to criteria frequently usedin query predicates. The collection of fragments is indexed toefficiently find the fragment(s) whose contents satisfy some querypredicate(s). The contents of each fragment are then indexed aswell, to support efficient filtering of the fragment data according to other query predicate s) conjunctively connected to the former. This strategy has been applied to a collection of more than 106 million images together with their related conventional data. Experimental results show considerable performance gain of the proposed approach for queries with conventional and similaritybasedpredicates, compared to the use of a unique metric index for the entire database contents. Karina Fasolin, Renato Fileto, Marcelo Krüger, Daniel S. Kaster, Mônica Ribeiro Porto Ferreira, Robson L. F. Cordeiro, Agma J. M. Traina, Caetano Traina Jr. |
ISM | 4 |
| 2011 | Using Visual Analysis to Weight Multiple Signatures to Discriminate Complex DataabstractComplex data is usually represented through signatures, which are sets of features describing the data content. Several kinds of complex data allow extracting different signatures from an object, representing complementary data characteristics. However, there is no ground truth of how balancing these signatures to reach an ideal similarity distribution. It depends on the analyst intent, that is, according to the job he/she is performing, a few signatures should have more impact in the data distribution than others. This work presents a new technique, called Visual Signature Weighting (ViSW), which allows interactively analyzing the impact of each signature in the similarity of complex data represented through multiple signatures. Our method provides means to explore the tradeoff of prioritizing signatures over the others, by dynamically changing their weight relation. We also present case studies showing that the technique is useful for global dataset analysis as well as for inspecting subspaces of interest. Renato Bueno, Daniel S. Kaster, Humberto Luiz Razente, Maria Camila Nardini Barioni, Agma J. M. Traina, Caetano Traina Jr. |
IV | 2 |
| 2010 | Metric Data Analysis Enhanced through Temporal VisualizationabstractThe human vision can naturally interpret data in spaces of 2 or 3 dimensions. When data is in higher dimensional spaces, in most cases the visualization is not intuitive. Regarding metric spaces, the interpretation is even harder, since they often do not have a direct spatial representation. However, the need to analyze how metric-represented data evolve over time is pretty common when one needs to understand several phenomena and in decision making processes, as it occurs in medical and agrometeorological applications. This paper presents three interactive techniques to visualize metric data that vary over time. Each one focus on a different way to interpret the temporal information. The first technique shows data evolving in a timeline axis. The second overlaps evolving snapshots of the space showing how the space varies regarding time. The last one does not treat temporal data as a dimension, it is used instead to define the similarity among complex data, employing the new concept of metric-temporal spaces, which seamlessly integrate time and metric data into a single similarity space. Visualization examples with real datasets are presented to show the usefulness of the proposed techniques. Renato Bueno, Humberto Luiz Razente, Daniel S. Kaster, Maria Camila Nardini Barioni, Agma J. M. Traina, Caetano Traina Jr. |
IV | 3 |
| 2009 | Unsupervised scaling of multi-descriptor similarity functions for medical image datasetsabstractContent-based search has proven to be a proper complement to textual queries over medical image databases. In many applications, employing multiple image descriptors and combining the respective distance functions using adequate scale factors improves the retrieval accuracy. However, the existing weighting methods are either exhaustive or supervised. In this paper, we present the Fractal-scaled Product Metric, an unsupervised method to determine a scale factor among features in multi-descriptor image similarity assessment based on the fractal theory. The composite distance function obtained is not limited to dimensional image descriptors and enables using scalable indexing structures. Experiments have shown that the proposed method determines near-optimal scale factors for the descriptors involved, and always improves the precision of the results, outperforming the individual descriptors up to 31% on the average precision. Renato Bueno, Daniel S. Kaster, Adriano Arantes Paterlini, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 2 |
| 2009 | Time-Aware Similarity Search: A Metric-Temporal Representation for Complex Data
Renato Bueno, Daniel S. Kaster, Agma J. M. Traina, Caetano Traina Jr. |
SSTD | 2 |
| 2008 | A New Approach for Optimization of Dynamic Metric Access Methods Using an Algorithm of Effective Deletion
Renato Bueno, Daniel S. Kaster, Agma J. M. Traina, Caetano Traina Jr. |
SSDBM | 2 |